EDBT 2026 Demo / reviewers in the wild / expert
Raquel Fernández
dblp:02/5384
· DBLP profile ↗
61ranked-venue papers
7as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Playpen: An Environment for Exploring Learning From Dialogue Game FeedbackabstractNicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia |
EMNLP | 11 |
| 2025 | Reading Between the Prompts: How Stereotypes Shape LLM's Implicit PersonalizationabstractGenerative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -a phenomenon called implicit personalization.Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups, even when no demographic information is explicitly provided.In this work, we systematically explore how LLMs respond to stereotypical cues using controlled synthetic conversations, by analyzing the models' latent user representations through both model internals and generated answers to targeted user questions.Our findings reveal that LLMs do infer demographic attributes based on these stereotypical signals, which for a number of groups even persists when the user explicitly identifies with a different demographic group.Finally, we show that this form of stereotypedriven implicit personalization can be effectively mitigated by intervening on the model's internal representations using a trained linear probe to steer them toward the explicitly stated identity.Our results highlight the need for greater transparency and control in how LLMs represent user identity. Vera Neplenbroek, Arianna Bisazza, Raquel Fernández |
EMNLP | 3 |
| 2025 | RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMsabstractAmbiguity resolution is key to effective communication.While humans effortlessly address ambiguity through conversational grounding strategies, the extent to which current language models can emulate these strategies remains unclear.In this work, we examine referential ambiguity in image-based question answering by introducing RACQUET, a carefully curated dataset targeting distinct aspects of ambiguity.Through a series of evaluations, we reveal significant limitations and problems of overconfidence of state-of-the-art large multimodal language models in addressing ambiguity in their responses.The overconfidence issue becomes particularly relevant for RACQUET-BIAS, a subset designed to analyze a critical yet underexplored problem: failing to address ambiguity leads to stereotypical, socially biased responses.Our results underscore the urgency of equipping models with robust strategies to deal with uncertainty without resorting to undesirable stereotypes. Alberto Testoni, Barbara Plank, Raquel Fernández |
EMNLP | 3 |
| 2024 | Speakers align both their gestures and words not only to establish but also to maintain reference to create shared labels for novel objects in interaction
Sho Akamine, Esam Ghaleb, Marlou Rasenberg, Raquel Fernández, Antje Meyer, Asli Özyürek |
CogSci | 4 |
| 2024 | Analysing Cross-Speaker Convergence in Face-to-Face Dialogue through the Lens of Automatically Detected Shared Linguistic Constructions
Esam Ghaleb, Marlou Rasenberg, Wim T. J. L. Pouw, Ivan Toni, Judith Holler, Asli Özyürek, Raquel Fernández |
CogSci | 7 |
| 2024 | Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic ProcessesabstractThere is an intricate relation between the properties of an image and how humans behave while describing the image.This behavior shows ample variation, as manifested in human signals such as eye movements and when humans start to describe the image.Despite the value of such signals of visuo-linguistic variation, they are virtually disregarded in the training of current pretrained models, which motivates further investigation.Using a corpus of Dutch image descriptions with concurrently collected eye-tracking data, we explore the nature of the variation in visuo-linguistic signals, and find that they correlate with each other.Given this result, we hypothesize that variation stems partly from the properties of the images, and explore whether image representations encoded by pretrained vision encoders can capture such variation.Our results indicate that pretrained models do so to a weak-to-moderate degree, suggesting that the models lack biases about what makes a stimulus complex for humans and what leads to variations in human outputs. Ece Takmaz, Sandro Pezzelle, Raquel Fernández |
EACL (1) | 3 |
| 2024 | Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification QuestionsabstractClarification questions are an essential dialogue tool to signal misunderstanding, ambiguities, and under-specification in language use.While humans are able to resolve uncertainty by asking questions since childhood, modern dialogue systems struggle to generate effective questions.To make progress in this direction, in this work we take a collaborative dialogue task as a testbed and study how model uncertainty relates to human uncertainty-an as yet underexplored problem.We show that model uncertainty does not mirror human clarificationseeking behavior, which suggests that using human clarification questions as supervision for deciding when to ask may not be the most effective way to resolve model uncertainty.To address this issue, we propose an approach to generating clarification questions based on model uncertainty estimation, compare it to several alternatives, and show that it leads to significant improvements in terms of task success.Our findings highlight the importance of equipping dialogue systems with the ability to assess their own uncertainty and exploit in interaction. Alberto Testoni, Raquel Fernández |
EACL (1) | 2 |
| 2024 | Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationabstractEnsuring the verifiability of model answers is a fundamental challenge for retrieval-augmented generation (RAG) in the question answering (QA) domain.Recently, self-citation prompting was proposed to make large language models (LLMs) generate citations to supporting documents along with their answers.However, self-citing LLMs often struggle to match the required format, refer to non-existent sources, and fail to faithfully reflect LLMs' context usage throughout the generation.In this work, we present MIRAGE -Model Internals-based RAG Explanations -a plug-and-play approach using model internals for faithful answer attribution in RAG applications.MIRAGE detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction via saliency methods.We evaluate our proposed approach on a multilingual extractive QA dataset, finding high agreement with human answer attribution.On open-ended QA, MIRAGE achieves citation quality and efficiency comparable to self-citation while also allowing for a finer-grained control of attribution parameters.Our qualitative evaluation highlights the faithfulness of MIRAGE's attributions and underscores the promising application of model internals for RAG answer attribution. 1 Jirui Qi, Gabriele Sarti, Raquel Fernández, Arianna Bisazza |
EMNLP | 3 |
| 2024 | Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic EvaluationabstractIn face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture representation learning challenging. How can we learn meaningful gestures representations considering gestures’ variability and relationship with speech? This paper tackles this challenge by employing self-supervised contrastive learning techniques to learn gesture representations from skeletal and speech information. We propose an approach that includes both unimodal and multimodal pre-training to ground gesture representations in co-occurring speech. For training, we utilize a face-to-face dialogue dataset rich with representational iconic gestures. We conduct thorough intrinsic evaluations of the learned representations through comparison with human-annotated pairwise gesture similarity. Moreover, we perform a diagnostic probing analysis to assess the possibility of recovering interpretable gesture features from the learned representations. Our results show a significant positive correlation with human-annotated gesture similarity and reveal that the similarity between the learned representations is consistent with well-motivated patterns related to the dynamics of dialogue interaction. Moreover, our findings demonstrate that several features concerning the form of gestures can be recovered from the latent representations. Overall, this study shows that multimodal contrastive learning is a promising approach for learning gesture representations, which opens the door to using such representations in larger-scale gesture analysis studies. Esam Ghaleb, Bulat Khaertdinov, Wim T. J. L. Pouw, Marlou Rasenberg, Judith Holler, Asli Özyürek, Raquel Fernández |
ICMI | 7 |
| 2024 | Co-Speech Gesture Detection through Multi-Phase Sequence LabelingabstractGestures are integral components of face-to-face communication. They unfold over time, often following predictable movement phases of preparation, stroke, and retraction. Yet, the prevalent approach to automatic gesture detection treats the problem as binary classification, classifying a segment as either containing a gesture or not, thus failing to capture its inherently sequential and contextual nature. To address this, we introduce a novel framework that reframes the task as a multi-phase sequence labeling problem rather than binary classification. Our model processes sequences of skeletal movements over time windows, uses Transformer encoders to learn contextual embeddings, and leverages Conditional Random Fields to perform sequence labeling. We evaluate our proposal on a large dataset of diverse co-speech gestures in task-oriented face-to-face dialogues. The results consistently demonstrate that our method significantly outperforms strong baseline models in detecting gesture strokes. Furthermore, applying Transformer encoders to learn contextual embeddings from movement sequences substantially improves gesture unit detection. These results highlight our framework’s capacity to capture the fine-grained dynamics of co-speech gesture phases, paving the way for more nuanced and accurate gesture detection and analysis. Esam Ghaleb, Ilya Burenko, Marlou Rasenberg, Wim T. J. L. Pouw, Peter Uhrig, Judith Holler, Ivan Toni, Asli Özyürek, Raquel Fernández |
WACV | 9 |
| 2023 | Interpretable Word Sense Representations via Definition Generation: The Case of Semantic Change AnalysisabstractWe propose using automatically generated natural language definitions of contextualised word usages as interpretable word and word sense representations.Given a collection of usage examples for a target word, and the corresponding data-driven usage clusters (i.e., word senses), a definition is generated for each usage with a specialised Flan-T5 language model, and the most prototypical definition in a usage cluster is chosen as the sense label.We demonstrate how the resulting sense labels can make existing approaches to semantic change analysis more interpretable, and how they can allow users-historical linguists, lexicographers, or social scientists-to explore and intuitively explain diachronic trajectories of word meaning.Semantic change analysis is only one of many possible applications of the 'definitions as representations' paradigm.Beyond being human-readable, contextualised definitions also outperform token or usage sentence embeddings in word-in-context semantic similarity judgements, making them a new promising type of lexical representation for NLP. Mario Giulianelli, Iris Luden, Raquel Fernández, Andrey Kutuzov |
ACL (1) | 3 |
| 2023 | The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal ModelsabstractDespite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction.In this work, we explore to what extent they handle basic linguistic constructions-active-passive voice, coordination, and relative clauses-that even preschool children can typically master.We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities.We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings.Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples.Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting.This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities. Xinyi Chen 0005, Raquel Fernández, Sandro Pezzelle |
EMNLP | 2 |
| 2023 | What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityabstractIn Natural Language Generation (NLG) tasks, for any input, multiple communicative goals are plausible, and any goal can be put into words, or produced, in multiple ways.We characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, connecting human production variability to aleatoric or data uncertainty.We then inspect the space of output strings shaped by a generation system's predicted probability distribution and decoding algorithm to probe its uncertainty.For each test input, we measure the generator's calibration to human production variability.Following this instance-level approach, we analyse NLG models and decoding strategies, demonstrating that probing a generator with multiple samples and, when possible, multiple references, provides the level of detail necessary to gain understanding of a model's representation of uncertainty. 1 * Equal contribution. 1 https://github.com/dmg-illc/nlg-uncertainty-probes Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, Barbara Plank |
EMNLP | 4 |
| 2023 | Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesabstractWe present information value, a measure which quantifies the predictability of an utterance relative to a set of plausible alternatives.We introduce a method to obtain interpretable estimates of information value using neural text generators, and exploit their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.Information value is a stronger predictor of utterance acceptability in written and spoken dialogue than aggregates of token-level surprisal and it is complementary to surprisal for predicting eye-tracked reading times. 1 Mario Giulianelli, Sarenne Wallbridge, Raquel Fernández |
EMNLP | 3 |
| 2023 | Cross-Lingual Consistency of Factual Knowledge in Multilingual Language ModelsabstractMultilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages.With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs.To this end, we propose a Rankingbased Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy.Using this metric, we conduct an in-depth analysis of the determining factors for CLC, both at model level and at language-pair level.Among other results, we find that increasing model size leads to higher factual probing accuracy in most languages, but does not improve cross-lingual consistency.Finally, we conduct a case study on CLC when new factual associations are inserted in the PLMs via model editing.Results on a small sample of facts inserted in English reveal a clear pattern whereby the new piece of knowledge transfers only to languages with which English has a high RankC score. 1 Jirui Qi, Raquel Fernández, Arianna Bisazza |
EMNLP | 2 |
| 2023 | GROOViST: A Metric for Grounding Objects in Visual StorytellingabstractA proper evaluation of stories generated for a sequence of images-the task commonly referred to as visual storytelling-must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding.In this work, we focus on evaluating the degree of grounding, that is, the extent to which a story is about the entities shown in the images.We analyze current metrics, both designed for this purpose and for general vision-text alignment.Given their observed shortcomings, we propose a novel evaluation tool, GROOViST, that accounts for cross-modal dependencies, temporal misalignments (the fact that the order in which entities appear in the story and the image sequence may not match), and human intuitions on visual grounding.An additional advantage of GROOViST is its modular design, where the contribution of each component can be assessed and interpreted individually. Aditya K. Surikuchi, Sandro Pezzelle, Raquel Fernández |
EMNLP | 3 |
| 2022 | Time Alignment between Gaze and Speech in Image Descriptions: Exploring Theories of Linearization
Ece Takmaz, Sandro Pezzelle, Raquel Fernández |
CogSci | 3 |
| 2022 | Stop Measuring Calibration When Humans DisagreeabstractCalibration is a popular framework to evaluate whether a classifier knows when it does not know-i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct.Correctness is commonly estimated against the human majority class.Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies.We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instancelevel measures of calibration that capture key statistical properties of human judgementsclass frequency, ranking and entropy. 1 Joris Baan, Wilker Aziz, Barbara Plank, Raquel Fernández |
EMNLP | 4 |
| 2022 | Structural Persistence in Language Models: Priming as a Window into Abstract Language RepresentationsabstractAbstract We investigate the extent to which modern neural language models are susceptible to structural priming, the phenomenon whereby the structure of a sentence makes the same structure more probable in a follow-up sentence. We explore how priming can be used to study the potential of these models to learn abstract structural information, which is a prerequisite for good performance on tasks that require natural language understanding skills. We introduce a novel metric and release Prime-LM, a large corpus where we control for various linguistic factors that interact with priming strength. We find that Transformer models indeed show evidence of structural priming, but also that the generalizations they learned are to some extent modulated by semantic information. Our experiments also show that the representations acquired by the models may not only encode abstract sequential structure but involve certain level of hierarchical syntactic information. More generally, our study shows that the priming paradigm is a useful, additional tool for gaining insights into the capacities of language models and opens the door to future priming-based investigations that probe the model’s internal states.1 Arabella Sinclair, Jaap Jumelet, Willem H. Zuidema, Raquel Fernández |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Analysing Human Strategies of Information Transmission as a Function of Discourse ContextabstractSpeakers are thought to use rational information transmission strategies for efficient communication (Genzel and Charniak, 2002;Aylett and Turk, 2004;Jaeger and Levy, 2007).Previous work analysing these strategies in sentence production has failed to take into account how the information content of sentences varies as a function of the available discourse context.In this study, we estimate sentence information content within discourse context.We find that speakers transmit information at a stable rate-i.e., rationally-in English newspaper articles but that this rate decreases in spoken open domain and written task-oriented dialogues.We also observe that speakers' choices are not oriented towards local uniformity of information, which is another hypothesised rational strategy.We suggest that a more faithful model of communication should explicitly include production costs and goal-oriented rewards. Mario Giulianelli, Raquel Fernández |
CoNLL | 2 |
| 2021 | Is Information Density Uniform in Task-Oriented Dialogues?abstractThe Uniform Information Density principle states that speakers plan their utterances to reduce fluctuations in the density of the information transmitted.In this paper, we test whether, and within which contextual units this principle holds in task-oriented dialogues.We show that there is evidence supporting the principle in written dialogues where participants play a cooperative reference game as well as in spoken dialogues involving instruction giving and following.Our study underlines the importance of identifying the relevant contextual components, showing that information content increases particularly within topically and referentially related contextual units. Mario Giulianelli, Arabella Sinclair, Raquel Fernández |
EMNLP (1) | 3 |
| 2021 | Word Representation Learning in Multimodal Pre-Trained Transformers: An Intrinsic EvaluationabstractAbstract This study carries out a systematic intrinsic evaluation of the semantic representations learned by state-of-the-art pre-trained multimodal Transformers. These representations are claimed to be task-agnostic and shown to help on many downstream language-and-vision tasks. However, the extent to which they align with human semantic intuitions remains unclear. We experiment with various models and obtain static word representations from the contextualized ones they learn. We then evaluate them against the semantic judgments provided by human speakers. In line with previous evidence, we observe a generalized advantage of multimodal representations over language- only ones on concrete word pairs, but not on abstract ones. On the one hand, this confirms the effectiveness of these models to align language and vision, which results in better semantic representations for concepts that are grounded in images. On the other hand, models are shown to follow different representation learning patterns, which sheds some light on how and when they perform multimodal integration. Sandro Pezzelle, Ece Takmaz, Raquel Fernández |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Analysing Lexical Semantic Change with Contextualised Word RepresentationsabstractThis paper presents the first unsupervised approach to lexical semantic change that makes use of contextualised word representations.We propose a novel method that exploits the BERT neural language model to obtain representations of word usages, clusters these representations into usage types, and measures change along time with three proposed metrics.We create a new evaluation dataset and show that the model representations and the detected semantic shifts are positively correlated with human judgements.Our extensive qualitative analysis demonstrates that our method captures a variety of synchronic and diachronic linguistic phenomena.We expect our work to inspire further research in this direction. Mario Giulianelli, Marco Del Tredici, Raquel Fernández |
ACL | 3 |
| 2020 | Asking questions with a big impact: Adapting to other interpretations of gradable adjectives
Sandro Pezzelle, Raquel Fernández |
CogSci | 2 |
| 2020 | Words are the Window to the Soul: Language-based User Representations for Fake News DetectionabstractCognitive and social traits of individuals are reflected in language use.Moreover, individuals who are prone to spread fake news online often share common traits.Building on these ideas, we introduce a model that creates representations of individuals on social media based only on the language they produce, and use them to detect fake news.We show that language-based user representations are beneficial for this task.We also present an extended analysis of the language of fake news spreaders, showing that its main features are mostly domain independent and consistent across two English datasets.Finally, we exploit the relation between language use and connections in the social graph to assess the presence of the Echo Chamber effect in our data. Marco Del Tredici, Raquel Fernández |
COLING | 2 |
| 2020 | Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational ContextsabstractDialogue participants often refer to entities or situations repeatedly within a conversation, which contributes to its cohesiveness.Subsequent references exploit the common ground accumulated by the interlocutors and hence have several interesting properties, namely, they tend to be shorter and reuse expressions that were effective in previous mentions.In this paper, we tackle the generation of first and subsequent references in visually grounded dialogue.We propose a generation model that produces referring utterances grounded in both the visual and the conversational context.To assess the referring effectiveness of its output, we also implement a reference resolution system.Our experiments and analyses show that the model produces better, more effective referring utterances than a model not grounded in the dialogue context, and generates subsequent references that exhibit linguistic patterns akin to humans. Ece Takmaz, Mario Giulianelli, Sandro Pezzelle, Arabella Sinclair, Raquel Fernández |
EMNLP (1) | 5 |
| 2020 | Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human GazeabstractWhen speakers describe an image, they tend to look at objects before mentioning them.In this paper, we investigate such sequential crossmodal alignment by modelling the image description generation process computationally.We take as our starting point a state-of-theart image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production.In particular, we propose the first approach to image description generation where visual processing is modelled sequentially.Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production.We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more naturalparticularly when gaze is encoded with a dedicated recurrent component. Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández |
EMNLP (1) | 4 |
| 2019 | Psycholinguistics Meets Continual Learning: Measuring Catastrophic Forgetting in Visual Question AnsweringabstractWe study the issue of catastrophic forgetting in the context of neural multimodal approaches to Visual Question Answering (VQA).Motivated by evidence from psycholinguistics, we devise a set of linguistically-informed VQA tasks, which differ by the types of questions involved (Wh-questions and polar questions).We test what impact task difficulty has on continual learning, and whether the order in which a child acquires question types facilitates computational models.Our results show that dramatic forgetting is at play and that task difficulty and order matter.Two well-known current continual learning methods mitigate the problem only to a limiting degree. Claudio Greco 0002, Barbara Plank, Raquel Fernández, Raffaella Bernardi |
ACL (1) | 3 |
| 2019 | The PhotoBook Dataset: Building Common Ground through Visually-Grounded DialogueabstractThis paper introduces the PhotoBook dataset, a large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.Taking inspiration from seminal work on dialogue analysis, we propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context as well as previously established referring expressions.We provide a detailed description of the task setup and a thorough analysis of the 2,500 dialogues collected.To further illustrate the novel features of the dataset, we propose a baseline model for reference resolution which uses a simple method to take into account shared information accumulated in a reference chain.Our results show that this information is particularly important to resolve later descriptions and underline the need to develop more sophisticated models of common ground in dialogue interaction. 1 Janosch Haber, Tim Baumgärtner, Ece Takmaz, Lieke Gelderloos, Elia Bruni, Raquel Fernández |
ACL (1) | 6 |
| 2019 | Is the Red Square Big? MALeViC: Modeling Adjectives Leveraging Visual ContextsabstractSandro Pezzelle, Raquel Fernández. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sandro Pezzelle, Raquel Fernández |
EMNLP/IJCNLP (1) | 2 |
| 2019 | You Shall Know a User by the Company It Keeps: Dynamic Representations for Social Media Users in NLPabstractMarco Del Tredici, Diego Marcheggiani, Sabine Schulte im Walde, Raquel Fernández. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Marco Del Tredici, Diego Marcheggiani, Sabine Schulte im Walde, Raquel Fernández |
EMNLP/IJCNLP (1) | 4 |
| 2018 | Ask No More: Deciding when to guess in referential visual dialogueabstractOur goal is to explore how the abilities brought in by a dialogue manager can be included in end-to-end visually grounded conversational agents. We make initial steps towards this general goal by augmenting a task-oriented visual dialogue model with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess. Our analyses show that adding a decision making component produces dialogues that are less repetitive and that include fewer unnecessary questions, thus potentially leading to more efficient and less unnatural interactions. Ravi Shekhar, Tim Baumgärtner, Aashish Venkatesh, Elia Bruni, Raffaella Bernardi, Raquel Fernández |
COLING | 6 |
| 2018 | The Road to Success: Assessing the Fate of Linguistic Innovations in Online CommunitiesabstractWe investigate the birth and diffusion of lexical innovations in a large dataset of online social communities. We build on sociolinguistic theories and focus on the relation between the spread of a novel term and the social role of the individuals who use it, uncovering characteristics of innovators and adopters. Finally, we perform a prediction task that allows us to anticipate whether an innovation will successfully spread within a community. Marco Del Tredici, Raquel Fernández |
COLING | 2 |
| 2018 | Automatic Evaluation of Neural Personality-based ChatbotsabstractStylistic variation is critical to render the utterances generated by conversational agents natural and engaging.In this paper, we focus on sequence-to-sequence models for open-domain dialogue response generation and propose a new method to evaluate the extent to which such models are able to generate responses that reflect different personality traits. Yujie Xing, Raquel Fernández |
INLG | 2 |
| 2017 | Adversarial evaluation for open-domain dialogue generationabstractWe investigate the potential of adversarial evaluation methods for open-domain dialogue generation systems, comparing the performance of a discriminative agent to that of humans on the same task.Our results show that the task is hard, both for automated models and humans, but that a discriminative agent can learn patterns that lead to above-chance performance. Elia Bruni, Raquel Fernández |
SIGDIAL Conference | 2 |
| 2016 | The LAMBADA dataset: Word prediction requiring a broad discourse contextabstractDenis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández |
ACL (1) | 9 |
| 2016 | A Data-driven Investigation of Corrective Feedback on Subject Omission Errors in First Language AcquisitionabstractWe investigate implicit corrections in the form of contrastive discourse in childadult interaction, which have been argued to contribute to language learning.In contrast to previous work in psycholinguistics, we adopt a data-driven methodology, using comparably large amounts of data and leveraging computational methods.We conduct a corpus study on the use of parental corrective feedback and show that its presence in child directed speech is associated with a reduction of child subject omission errors in English. Sarah Hiller, Raquel Fernández |
CoNLL | 2 |
| 2016 | On the Influence of Gender on Interruptions in Multiparty DialogueabstractDuring conversations, participants do not always alternate turns smoothly. One cause of disturbance particularly prominent in multiparty dialogue is the presence of interruptions: interventions that prevent current speakers from finishing their turns. Previous work, mostly within the field of sociolinguistics, has suggested that the gender of the dialogue participants plays an important role in their interruptive behaviour. We investigate existing hypotheses in this respect by systematically analysing interruptions in a corpus of spoken multiparty meetings that include a minimum of two male and two female participants. We find a number of significant differences, including the fact that women are more often interrupted overall and that men interrupt more often women than other men, in particular using speech overlap to grab the floor. We do not find evidence for the hypothesis that women interrupt other women more frequently than they interrupt men. Paul Van Eecke, Raquel Fernández |
INTERSPEECH | 2 |
| 2016 | PentoRef: A Corpus of Spoken References in Task-oriented Dialogues
Sina Zarrieß, Julian Hough, Casey Kennington, Ramesh R. Manuvinakurike, David DeVault, Raquel Fernández, David Schlangen |
LREC | 6 |
| 2016 | Questioning Arbitrariness in Language: a Data-Driven Study of Conventional IconicityabstractThis paper presents a data-driven investigation of phonesthemes, phonetic units said to carry meaning associations, thus challenging the traditionally assumed arbitrariness of language.Phonesthemes have received a substantial amount of attention within the cognitive science literature on sound iconicity, but nevertheless remain a controversial and understudied phenomenon.Here we employ NLP techniques to address two main questions: How can the existence of phonesthemes be tested at a large scale with quantitative methods?And how can the meaning arguably carried by a phonestheme be induced automatically from word embeddings?We develop novel methods to make progress on these fronts and compare our results to previous work, obtaining substantial improvements. Ekaterina Abramova, Raquel Fernández |
HLT-NAACL | 2 |
| 2016 | Multimodal Semantic Learning from Child-Directed InputabstractAngeliki Lazaridou, Grzegorz Chrupała, Raquel Fernández, Marco Baroni. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Angeliki Lazaridou, Grzegorz Chrupala, Raquel Fernández, Marco Baroni |
HLT-NAACL | 3 |
| 2014 | Quantifying Categorical and Conceptual Convergence in Child-Adult Dialogue
Raquel Fernández, Robert Grimm 0003 |
CogSci | 1 |
| 2014 | Empirical Analysis of Aggregation Methods for Collective Annotation
Ciyang Qing, Ulle Endriss, Raquel Fernández, Justin Kruger |
COLING | 3 |
| 2014 | The Role of Polarity in Inferring Acceptance and Rejection in DialogueabstractWe study the role that logical polarity plays in determining the rejection or ac-ceptance function of an utterance in dia-logue. We develop a model inspired by re-cent work on the semantics of negation and polarity particles and test it on annota-ted data from two spoken dialogue corpo-ra: the Switchboard Corpus and the AMI Meeting Corpus. Our experiments show that taking into account the relative pola-rity of a proposal under discussion and of its response greatly helps to distinguish re-jections from acceptances in both corpora. 1 Julian J. Schlöder, Raquel Fernández |
SIGDIAL Conference | 2 |
| 2013 | Collective Annotation of Linguistic Resources: Basic Principles and a Formal Model
Ulle Endriss, Raquel Fernández |
ACL (1) | 2 |
| 2013 | Automatic Labeling of Phonesthemic Senses
Ekaterina Abramova, Raquel Fernández, Federico Sangati |
CogSci | 2 |
| 2012 | Building a Corpus of Indefinite Uses Annotated with Fine-grained Semantic Functions
Maria Aloni, Andreas van Cranenburgh, Raquel Fernández, Marta Sznajder |
LREC | 3 |
| 2010 | The CALO Meeting Assistant SystemabstractThe CALO Meeting Assistant (MA) provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper presents the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, topic identification and segmentation, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, Stanley Peters, Dilek Hakkani-Tür, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
IEEE Trans. Speech Audio Process. | 8 |
| 2009 | Who is "You"? Combining Linguistic and Gaze Features to Resolve Second-Person References in Dialogue
Matthew Frampton, Raquel Fernández, Patrick Ehlen, C. Mario Christoudias, Trevor Darrell, Stanley Peters |
EACL | 2 |
| 2009 | Cascaded Lexicalised Classifiers for Second-Person Reference Resolution
Matthew Purver, Raquel Fernández, Matthew Frampton, Stanley Peters |
SIGDIAL Conference | 2 |
| 2008 | Identifying relevant phrases to summarize decisions in spoken meetingsabstractWe address the problem of identifying words and phrases that accurately capture, or contribute to, the semantic gist of decisions made in multi-party human-human meetings. We first describe our approach to modelling decision discussions in spoken meetings and then compare two approaches to extracting information from these discussions. The first one uses an opendomain semantic parser that identifies candidate phrases for decision summaries and then employs machine learning techniques to select from those candidate phrases. The second one uses categorical and sequential classifiers that exploit simple syntactic and semantic features to identify words and phrases relevant for decision summarization. Raquel Fernández, Matthew Frampton, John Dowding, Anish Adukuzhiyil, Patrick Ehlen, Stanley Peters |
INTERSPEECH | 1 |
| 2008 | The CALO meeting speech recognition and understanding systemabstractThe CALO Meeting Assistant provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper summarizes the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Dilek Hakkani-Tür, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Stanley Peters, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
SLT | 6 |
| 2007 | Speaking through a noisy channel - experiments on inducing clarification behaviour in human-human dialogueabstractWe report results of an experiment on inducing communication problems in human-human dialogue.We set up a voice-only cooperative task where we manipulated one channel by replacing (in real-time, at random points) all signal with noise.Altogether around 10% of the speaker's signal was thus removed.We found an increase in clarification requests of a form that has previously been hypothesised to be used mainly for clarifying acoustic problems.We also found a correlation between the percentage of an utterance being manipulated and the use of devices for pointing out error locations.From our findings, we derive a gold-standard policy for clarification behaviour. David Schlangen, Raquel Fernández |
INTERSPEECH | 2 |
| 2007 | Classifying Non-Sentential Utterances in Dialogue: A Machine Learning ApproachabstractIn this article we use well-known machine learning methods to tackle a novel task, namely the classification of non-sentential utterances (NSUs) in dialogue. We introduce a fine-grained taxonomy of NSU classes based on corpus work, and then report on the results of several machine learning experiments. First, we present a pilot study focused on one of the NSU classes in the taxonomy—bare wh-phrases or “sluices”—and explore the task of disambiguating between the different readings that sluices can convey. We then extend the approach to classify the full range of NSU classes, obtaining results of around an 87% weighted F-score. Thus our experiments show that, for the taxonomy adopted, the task of identifying the right NSU class can be successfully learned, and hence provide a very encouraging basis for the more general enterprise of fully processing NSUs. Raquel Fernández, Jonathan Ginzburg, Shalom Lappin |
Comput. Linguistics | 1 |
| 2006 | Interaction in Task-Oriented Human-Human Dialogue: the Effects of Different turn-Taking PoliciesabstractIn human-human dialogue, the allocation of turns between the participants is normally managed smoothly, without the participants paying much attention to it. In contrast, for spoken dialogue systems turn allocation is a difficult task, and often technical restrictions are introduced to simplify it. In this paper we investigate, by comparing two experimentally collected corpora of human-human task oriented dialogue, what the consequences are of imposing one particular kind of restriction, namely that of using a simplex channel managed by push-to-talk (PTT). We found, as expected, a loss of interactivity in the PTT condition (fewer, longer turns; more silences), but surprisingly, no loss of efficiency; in fact, the subjects in the PTT condition were able to finish their task in roughly the same time, while using fewer words along the way. We analyse here the differences in the interaction patterns and the interplay of 'naturalness' and efficiency as relevant factors for practical system development. Raquel Fernández, Tatjana Lucht, Kepa Joseba Rodríguez, David Schlangen |
SLT | 1 |
| 2005 | Scaling up from Dialogue to Multilogue: Some Principles and BenchmarksabstractThe paper considers how to scale up dialogue protocols to multilogue, settings with multiple conversationalists. We extract two benchmarks to evaluate scaled up protocols based on the long distance resolution possibilities of non-sentential utterances in dialogue and multilogue in the British National Corpus. In light of these benchmarks, we then consider three possible transformations to dialogue protocols, formulated within an issue-based approach to dialogue management. We show that one such transformation yields protocols for querying and assertion that fulfill these benchmarks. Jonathan Ginzburg, Raquel Fernández |
ACL | 2 |
| 2004 | Classifying Ellipsis in Dialogue: A Machine Learning Approach
Raquel Fernández, Jonathan Ginzburg, Shalom Lappin |
COLING | 1 |
| 2003 | A Dynamic Logic Formalisation of the Dialogue Gameboard
Raquel Fernández |
EACL | 1 |
| 2002 | Non-Sentential Utterances: Grammar and Dialogue Dynamics in Corpus Annotation
Raquel Fernández, Jonathan Ginzburg |
COLING | 1 |
| 1998 | Object following and obstacle avoidance using a laser scanner in the outdoor mobile robot Auriga-αabstractA low computational cost method for terrestrial mobile robots that uses a laser scanner for following mobile objects and avoiding obstacles is presented. In particular, the technique has been successfully implemented in the outdoor mobile robot Auriga-/spl alpha/. The measurement data from the laser scanner and the vehicle position are used as input variables. The outputs are the new curvature and velocity of the robot in order to follow a mobile object, or track a previously recorded path and avoid any possible obstacle in its way. Jorge L. Martínez, Ana Pozo-Ruz, Salvador Pedraza, Raquel Fernández |
IROS | 4 |
| 1997 | Dynamic speed planning for safe navigationabstractThe paper deals with speed control for autonomous mobile robots. First a trajectory planner attaches a speed component to the postures of a path by considering speed limitations introduced by the vehicle and by task specifications. In order to provide robustness during task execution in the real world, environment feedback is used to dynamically adjust the speed of the vehicle to the presence of unexpected obstacles, both static and mobile. This is accomplished by combining the planned speed profile with the outputs of two reactive controllers. The system has been successfully implemented within the control architecture of the RAM-2 mobile robot. Anthony Mandow, Victor F. Muñoz, Raquel Fernández, Alfonso García-Cerezo |
IROS | 3 |