VLDB 2026 Research / reviewers in the wild / expert
Emiel Krahmer
dblp:61/3236 · also Emiel J. Krahmer
· DBLP profile ↗
132ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0002-6304-7549ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 120 · 13 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 42 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 10 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Watt Impact? Making the Environmental Impact of Generative AI Tangible Through Participatory Data PhysicalizationsabstractThe rapid adoption and use of generative AI tools, like ChatGPT, has led to an exponential increase in its environmental impact. However, most users are unaware of this impact. To investigate and raise public awareness of the impact, we designed ecoAIware, a set of five tangible participatory data physicalizations (PDPs) highlighting different environmental impacts of genAI. ecoAIware was deployed at a library and an academic conference, where visitors could interact with the PDPs by contributing tokens representing their awareness of the different impacts. In this pictorial, we present our findings on current awareness levels of visitors, their engagement with ecoAIware and how it contributed to raising their awareness. This pictorial contributes to (i) the discourse within HCI communities around the environmental impact of genAI and (ii) by showcasing how creative methods such as PDPs can be used to engage users with critical topics and collect data in the process. Anniek Jansen, Supraja Sankaran, Emiel Krahmer |
Creativity & Cognition | 3 |
| 2025 | How Well Can Large Language Models Reflect? A Human Evaluation of LLM-generated Reflections for Motivational Interviewing DialoguesabstractMotivational Interviewing (MI) is a counseling technique that promotes behavioral change through reflective responses to mirror or refine client statements. While advanced Large Language Models (LLMs) can generate engaging dialogues, challenges remain for applying them in a sensitive context such as MI. This work assesses the potential of LLMs to generate MI reflections via three LLMs: GPT-4, Llama-2, and BLOOM, and explores the effect of dialogue context size and integration of MI strategies for reflection generation by LLMs. We conduct evaluations using both automatic metrics and human judges on four criteria: appropriateness, relevance, engagement, and naturalness, to assess whether these LLMs can accurately generate the nuanced therapeutic communication required in MI. While we demonstrate LLMs’ potential in generating MI reflections comparable to human therapists, content analysis shows that significant challenges remain. By identifying the strengths and limitations of LLMs in generating empathetic and contextually appropriate reflections in MI, this work contributes to the ongoing dialogue in enhancing LLM’s role in therapeutic counseling. Mustafa Erkan Basar, Xin Sun 0016, Iris Hendrickx, Jan de Wit, Tibor Bosse, Gert-Jan de Bruijn, Jos A. Bosch, Emiel Krahmer |
COLING | 8 |
| 2024 | Use of spatial reference frames for motion events in Balinese co-speech gesture
Danielle Naegeli, David Peeters, Emiel Krahmer, Marieke Schouwstra, Made Sri Satyawati, Connie de Vos |
CogSci | 3 |
| 2024 | Eliciting Motivational Interviewing Skill Codes in Psychotherapy with LLMs: A Bilingual Dataset and Analytical StudyabstractBehavioral coding (BC) in motivational interviewing (MI) holds great potential for enhancing the efficacy of MI counseling. However, manual coding is labor-intensive, and automation efforts are hindered by the lack of data due to the privacy of psychotherapy. To address these challenges, we introduce BiMISC, a bilingual dataset of MI conversations in English and Dutch, sourced from real counseling sessions. Expert annotations in BiMISC adhere strictly to the motivational interviewing skills code (MISC) scheme, offering a pivotal resource for MI research. Additionally, we present a novel approach to elicit the MISC expertise from Large language models (LLMs) for MI coding. Through the in-depth analysis of BiMISC and the evaluation of our proposed approach, we demonstrate that the LLM-based approach yields results closely aligned with expert annotations and maintains consistent performance across different languages. Our contributions not only furnish the MI community with a valuable bilingual dataset but also spotlight the potential of LLMs in MI coding, laying the foundation for future MI research. Xin Sun 0016, Jiahuan Pei, Jan de Wit, Mohammad Aliannejadi, Emiel Krahmer, Jos T. P. Dobber, Jos A. Bosch |
LREC/COLING | 5 |
| 2023 | Neural Data-to-Text Generation Based on Small Datasets: Comparing the Added Value of Two Semi-Supervised Learning Approaches on Top of a Large Language ModelabstractAbstract This study discusses the effect of semi-supervised learning in combination with pretrained language models for data-to-text generation. It is not known whether semi-supervised learning is still helpful when a large-scale language model is also supplemented. This study aims to answer this question by comparing a data-to-text system only supplemented with a language model, to two data-to-text systems that are additionally enriched by a data augmentation or a pseudo-labeling semi-supervised learning approach. Results show that semi-supervised learning results in higher scores on diversity metrics. In terms of output quality, extending the training set of a data-to-text system with a language model using the pseudo-labeling approach did increase text quality scores, but the data augmentation approach yielded similar scores to the system without training set extension. These results indicate that semi-supervised learning approaches can bolster output quality and diversity, even when a language model is also present. Chris van der Lee, Thiago Castro Ferreira, Chris Emmery, Travis J. Wiltshire, Emiel Krahmer |
Comput. Linguistics | 5 |
| 2023 | The Design and Observed Effects of Robot-performed Manual Gestures: A Systematic ReviewabstractCommunication using manual (hand) gestures is considered a defining property of social robots, and their physical embodiment and presence, therefore, we see a need for a comprehensive overview of the state-of-the-art in social robots that use gestures. This systematic literature review aims to address this need by (1) describing the gesture production process of a social robot, including the design and planning steps, and (2) providing a survey of the effects of robot-performed gestures on human-robot interactions in a multitude of domains. We identify patterns and themes from the existing body of literature, resulting in nine outstanding questions for research on robot-performed gestures regarding: developments in sensor technology and AI, structuring the gesture design and evaluation process, the relationship between physical appearance and gestures, the effects of planning on the overall interaction, standardizing measurements of gesture “quality,” individual differences, gesture mirroring, whether human-likeness is desirable, and universal accessibility of robots. We also reflect on current methodological practices in studies of robot-performed gestures and suggest improvements regarding replicability, external validity, measurement instruments used, and connections with other disciplines. These outstanding questions and methodological suggestions can guide future work in this field of research. Jan de Wit, Paul Vogt, Emiel Krahmer |
ACM Trans. Hum. Robot Interact. | 3 |
| 2022 | Cross-cultural differences in the emergence of referential strategies in artificial sign languages
Danielle Naegeli, David Peeters, Emiel Krahmer, Marieke Schouwstra, Yasamin Motamedi, Connie de Vos |
CogSci | 3 |
| 2022 | Hints of Independence in a Pre-scripted World: On Controlled Usage of Open-domain Language Models for Chatbots in Highly Sensitive DomainsabstractContains fulltext : 250750.pdf (Publisher’s version ) (Open Access) Mustafa Erkan Basar, Iris Hendrickx, Emiel Krahmer, Gert-Jan de Bruijn, Tibor Bosse |
ICAART (1) | 3 |
| 2022 | Affective Words and the Company They Keep: Studying the Accuracy of Affective Word Lists in Determining Sentence and Word Valence in a Domain-Specific CorpusabstractIn this article, we explore whether and how linguistic and pragmatic context can change individual word valence and emotionality in two parts. In the first part, we investigate whether sentence contexts retrieved from a domain-specific corpus (soccer) bias individual word affect. We then examine whether word valence with and without context accurately indicates sentence valence. In the second part, we compare word ratings from the first part to four different existing affective lexicons, with different levels of sensitivity to semantic and pragmatic context, and examine their accuracy in determining sentence valence. Results show a significant difference between words with and without context, the former more accurate in determining sentence valence than the latter. The preexisting lexicons were found to be similar to the individual word ratings collected in the first part of the study, with human-evaluated, context-sensitive lexicons being the most accurate in determining sentence valence. We discuss implications for emotion theory and bag-of-words approaches to sentiment analysis. Nadine Braun, Martijn Goudbeek, Emiel Krahmer |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Preregistering NLP researchabstractEmiel van Miltenburg, Chris van der Lee, Emiel Krahmer. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Emiel van Miltenburg, Chris van der Lee, Emiel Krahmer |
NAACL-HLT | 3 |
| 2021 | Human evaluation of automatically generated text: Current trends and best practice guidelinesabstractCurrently, there is little agreement as to how Natural Language Generation (NLG) systems should be evaluated, with a particularly high degree of variation in the way that human evaluation is carried out. This paper provides an overview of how (mostly intrinsic) human evaluation is currently conducted and presents a set of best practices, grounded in the literature. These best practices are also linked to the stages that researchers go through when conducting an evaluation research (planning stage; execution and release stage), and the specific steps in these stages. With this paper, we hope to contribute to the quality and consistency of human evaluations in NLG. Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Emiel Krahmer |
Comput. Speech Lang. | 4 |
| 2021 | Designing and Evaluating Iconic Gestures for Child-Robot Second Language LearningabstractAbstract In this paper, we examine the process of designing robot-performed iconic hand gestures in the context of a long-term study into second language tutoring with children of approximately 5 years old. We explore four factors that may relate to their efficacy in supporting second language tutoring: the age of participating children; differences between gestures for various semantic categories, e.g. measurement words, such as small, versus counting words, such as five; the quality (comprehensibility) of the robot’s gestures; and spontaneous reenactment or imitation of the gestures. Age was found to relate to children’s learning outcomes, with older children benefiting more from the robot’s iconic gestures than younger children, particularly for measurement words. We found no conclusive evidence that the quality of the gestures or spontaneous reenactment of said gestures related to learning outcomes. We further propose several improvements to the process of designing and implementing a robot’s iconic gesture repertoire. Jan de Wit, Bram Willemsen, Mirjam de Haas, Rianne van den Berghe, Paul M. Leseman, Ora Oudgenoeg-Paz, Josje Verhagen, Paul Vogt, Emiel Krahmer |
Interact. Comput. | 9 |
| 2021 | The interplay of prosodic cues in the L2: How intonation, rhythm, and speech rate in speech by Spanish learners of Dutch contribute to L1 Dutch perceptions of accentedness and comprehensibilityabstractThis study investigates the relative contribution of L2 intonation, rhythm, and speech rate to L1 perceptions of accentedness and comprehensibility. The intonation, rhythm, and speech rate of an L1 speaker of Dutch was transferred onto the segmental string of four Spanish learners of Dutch, resulting in eight conditions that reflect all possible combinations of these prosodic cues. Our results show that improving the prosody of L2 speakers positively affects L1 perceptions of L2 speech, but not to the same extent for both measures: Concerning accentedness, Dutch listeners were influenced by intonation and speech rate transfer, but not by rhythm transfer. No interactions were found between prosodic cues. Comprehensibility ratings were also affected by intonation and speech rate, but for this measure, the effects were mediated by interactions between the prosodic features. Our study reaffirms the importance of differentiating between different aspects of perception and provides insight into those features that are most likely to affect each L1 perception type. Lieke van Maastricht, Tim Zee, Emiel Krahmer, Marc Swerts |
Speech Commun. | 3 |
| 2020 | Emotional Words - The Relationship of Self- and Other-Annotation of Affect in Written Text
Nadine Braun, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2020 | Effects of domain size during reference production in photo-realistic scenes
Ruud Koolen, Emiel Krahmer |
CogSci | 2 |
| 2020 | Varied Human-Like Gestures for Social Robots: Investigating the Effects on Children's Engagement and Language LearningabstractTo investigate whether a humanoid robot's use of gestures improves children's learning of second language vocabulary, and if variation in gestures strengthens this effect, we conducted a field study where a total of 94 children (aged 4-6 years old) played a language learning game with a NAO robot. The robot either used no gestures at all, repeated the same gesture every time a target word was presented, or produced a different gesture for each occurrence of a target word. We found that, contrary to what the majority of existing research suggests, the robot's use of gestures did not result in increased learning outcomes, compared to a robot that did not use gestures. However, engagement between child and robot was higher in both the repeated and varied gesture conditions, compared to the condition without gestures. An exploratory analysis showed that age played a role: the older children in the study learned more than the younger children when the robot used gestures. It is therefore important to carefully consider the design and application of robot gestures to support the learning process. The contribution of this work is twofold: it is a conceptual reproduction of a previous study, and we have taken first steps towards exploring the role of variation in gestures. The study was preregistered, and all materials are made publicly available. Jan de Wit, Arold Brandse, Emiel Krahmer, Paul Vogt |
HRI | 3 |
| 2020 | The CACAPO Dataset: A Multilingual, Multi-Domain Dataset for Neural Pipeline and End-to-End Data-to-Text GenerationabstractThis paper describes the CACAPO dataset, built for training both neural pipeline and endto-end data-to-text language generation systems.The dataset is multilingual (Dutch and English), and contains almost 10,000 sentences from human-written news texts in the sports, weather, stocks, and incidents domain, together with aligned attribute-value paired data.The dataset is unique in that the linguistic variation and indirect ways of expressing data in these texts reflect the challenges of real world NLG tasks. Chris van der Lee, Chris Emmery, Sander Wubben, Emiel Krahmer |
INLG | 4 |
| 2020 | Gradations of Error Severity in Automatic Image DescriptionsabstractEarlier research has shown that evaluation metrics based on textual similarity (e.g., BLEU, CIDEr, Meteor) do not correlate well with human evaluation scores for automatically generated text.We carried out an experiment with Chinese speakers, where we systematically manipulated image descriptions to contain different kinds of errors.Because our manipulated descriptions form minimal pairs with the reference descriptions, we are able to assess the impact of different kinds of errors on the perceived quality of the descriptions.Our results show that different kinds of errors elicit significantly different evaluation scores, even though all erroneous descriptions differ in only one character from the reference descriptions.Evaluation metrics based solely on textual similarity are unable to capture these differences, which (at least partially) explains their poor correlation with human judgments.Our work provides the foundations for future work, where we aim to understand why different errors are seen as more or less severe. Emiel van Miltenburg, Wei-Ting Lu, Emiel Krahmer, Albert Gatt, Guanyi Chen, Kees van Deemter |
INLG | 3 |
| 2020 | Query-based summarization of discussion threadsabstractAbstract In this paper, we address query-based summarization of discussion threads. New users can profit from the information shared in the forum, Please check if the inserted city and country names in the affiliations are correct. if they can find back the previously posted information. However, discussion threads on a single topic can easily comprise dozens or hundreds of individual posts. Our aim is to summarize forum threads given real web search queries. We created a data set with search queries from a discussion forum’s search engine log and the discussion threads that were clicked by the user who entered the query. For 120 thread–query combinations, a reference summary was made by five different human raters. We compared two methods for automatic summarization of the threads: a query-independent method based on post features, and Maximum Marginal Relevance (MMR), a method that takes the query into account. We also compared four different word embeddings representations as alternative for standard word vectors in extractive summarization. We find (1) that the agreement between human summarizers does not improve when a query is provided that: (2) the query-independent post features as well as a centroid-based baseline outperform MMR by a large margin; (3) combining the post features with query similarity gives a small improvement over the use of post features alone; and (4) for the word embeddings, a match in domain appears to be more important than corpus size and dimensionality. However, the differences between the models were not reflected by differences in quality of the summaries created with help of these models. We conclude that query-based summarization with web queries is challenging because the queries are short, and a click on a result is not a direct indicator for the relevance of the result. Suzan Verberne, Emiel Krahmer, Sander Wubben, Antal van den Bosch |
Nat. Lang. Eng. | 2 |
| 2019 | Lifting the Curse of Knowing: How Feedback Improves Readers' Perspective-Taking
Debby Damen, Marije van Amelsvoort, Per van der Wijst, Emiel Krahmer |
CogSci | 4 |
| 2019 | Neural data-to-text generation: A comparison between pipeline and end-to-end architecturesabstractThiago Castro Ferreira, Chris van der Lee, Emiel van Miltenburg, Emiel Krahmer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Thiago Castro Ferreira, Chris van der Lee, Emiel van Miltenburg, Emiel Krahmer |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Second Language Tutoring Using Social Robots: L2TOR - The MovieabstractThis video illustrates the large-scale experiment of the L2TOR project that will be presented at the HRI 2019 conference. The experiment aimed to investigate how 192 Dutch 5-year-old children could learn 34 English words from a NAO robot in 7 lessons. The experiment compared 4 conditions: 1) robot using iconic gestures, 2) robot without iconic gestures, 3) tablet only, and 4) a control group. The results revealed that children could learn more English words in all experimental conditions compared to the control group. The three experimental conditions did not show any significant differences regarding the learning outcomes. Paul Vogt, Rianne van den Berghe, Mirjam de Haas, Laura Kunold, Junko Kanero, Ezgi Mamus, Jean-Marc Montanier, Cansu Oranç, Ora Oudgenoeg-Paz, Daniel Hernández García, Fotios Papadopoulos, Thorsten Schodde, Josje Verhagen, Christopher D. Wallbridge, Bram Willemsen, Jan de Wit, Tony Belpaeme, Tilbe Göksun, Stefan Kopp, Emiel Krahmer, Aylin C. Küntay, Paul M. Leseman, Amit Kumar Pandey |
HRI | 20 |
| 2019 | Second Language Tutoring Using Social Robots: A Large-Scale StudyabstractWe present a large-scale study of a series of seven lessons designed to help young children learn English vocabulary as a foreign language using a social robot. The experiment was designed to investigate 1) the effectiveness of a social robot teaching children new words over the course of multiple interactions (supported by a tablet), 2) the added benefit of a robot's iconic gestures on word learning and retention, and 3) the effect of learning from a robot tutor accompanied by a tablet versus learning from a tablet application alone. For reasons of transparency, the research questions, hypotheses and methods were preregistered. With a sample size of 194 children, our study was statistically well-powered. Our findings demonstrate that children are able to acquire and retain English vocabulary words taught by a robot tutor to a similar extent as when they are taught by a tablet application. In addition, we found no beneficial effect of a robot's iconic gestures on learning gains. Paul Vogt, Rianne van den Berghe, Mirjam de Haas, Laura Kunold, Junko Kanero, Ezgi Mamus, Jean-Marc Montanier, Cansu Oranç, Ora Oudgenoeg-Paz, Daniel Hernández García, Fotios Papadopoulos, Thorsten Schodde, Josje Verhagen, Christopher D. Wallbridge, Bram Willemsen, Jan de Wit, Tony Belpaeme, Tilbe Göksun, Stefan Kopp, Emiel Krahmer, Aylin C. Küntay, Paul M. Leseman, Amit Kumar Pandey |
HRI | 20 |
| 2019 | Playing Charades with a Robot: Collecting a Large Dataset of Human Gestures Through HRIabstractThis work documents a playful human-robot interaction, in the form of a game of charades, through which a humanoid robot is able to learn how to produce and recognize gestures by interacting with human participants. We describe an extensive dataset of gesture recordings, which can be used for future research into gestures, specifically for human-robot interaction applications. Jan de Wit, Bram Willemsen, Mirjam de Haas, Emiel Krahmer, Paul Vogt, Marije Merckens, Reinjet Oostdijk, Chani Savelberg, Sabine Verdult, Pieter Wolfert |
HRI | 4 |
| 2019 | A Personalized Data-to-Text Support Tool for Cancer PatientsabstractIn this paper, we present a novel data-to-text system for cancer patients, providing information on quality of life implications after treatment, which can be embedded in the context of shared decision making.Currently, information on quality of life implications is often not discussed, partly because (until recently) data has been lacking.In our work, we rely on a newly developed prediction model, which assigns patients to scenarios.Furthermore, we use data-to-text techniques to explain these scenario-based predictions in personalized and understandable language.We highlight the possibilities of NLG for personalization, discuss ethical implications and also present the outcomes of a first evaluation with clinicians. Saar Hommes, Chris van der Lee, Felix J. Clouth, Jeroen K. Vermunt, Xander Verbeek, Emiel Krahmer |
INLG | 6 |
| 2019 | Best practices for the human evaluation of automatically generated textabstractCurrently, there is little agreement as to how Natural Language Generation (NLG) systems should be evaluated, with a particularly high degree of variation in the way that human evaluation is carried out.This paper provides an overview of how human evaluation is currently conducted, and presents a set of best practices, grounded in the literature.With this paper, we hope to contribute to the quality and consistency of human evaluations in NLG. Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, Emiel Krahmer |
INLG | 5 |
| 2019 | On task effects in NLG corpus elicitation: a replication study using mixed effects modelingabstractTask effects in NLG corpus elicitation recently started to receive more attention, but are usually not modeled statistically.We present a controlled replication of the study by Van Miltenburg et al. (2018b), contrasting spoken with written descriptions.We collected additional written Dutch descriptions to supplement the spoken data from the DIDEC corpus, and analyzed the descriptions using mixed effects modeling to account for variation between participants and items.Our results show that the effects of modality largely disappear in a controlled setting. Emiel van Miltenburg, Merel van de Kerkhof, Ruud Koolen, Martijn Goudbeek, Emiel Krahmer |
INLG | 5 |
| 2018 | NeuralREG: An end-to-end approach to referring expression generationabstractThiago Castro Ferreira, Diego Moussallem, Ákos Kádár, Sander Wubben, Emiel Krahmer. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Thiago Castro Ferreira, Diego Moussallem, Ákos Kádár, Sander Wubben, Emiel Krahmer |
ACL (1) | 5 |
| 2018 | Changing Minds: The Effect of Stimulated Attention to Another's Different Point of View on Visual Perspective-Taking
Debby Damen, Marije van Amelsvoort, Per van der Wijst, Emiel Krahmer |
CogSci | 4 |
| 2018 | The Curse of Knowing: The Influence of Explicit Perspective-Awareness Instructions on Perceivers' Perspective-Taking
Debby Damen, Per van der Wijst, Marije van Amelsvoort, Emiel Krahmer |
CogSci | 4 |
| 2018 | Aspect-based summarization of pros and cons in unstructured product reviewsabstractWe developed three systems for generating pros and cons summaries of product reviews. Automating this task eases the writing of product reviews, and offers readers quick access to the most important information. We compared SynPat, a system based on syntactic phrases selected on the basis of valence scores, against a neural-network-based system trained to map bag-of-words representations of reviews directly to pros and cons, and the same neural system trained on clusters of word-embedding encodings of similar pros and cons. We evaluated the systems in two ways: first on held-out reviews with gold-standard pros and cons, and second by asking human annotators to rate the systems’ output on relevance and completeness. In the second evaluation, the gold-standard pros and cons were assessed along with the system output. We find that the human-generated summaries are not deemed as significantly more relevant or complete than the SynPat systems; the latter are scored higher than the human-generated summaries on a precision metric. The neural approaches yield a lower performance in the human assessment, and are outperformed by the baseline. Florian Kunneman, Sander Wubben, Antal van den Bosch, Emiel Krahmer |
COLING | 4 |
| 2018 | Evaluating the text quality, human likeness and tailoring component of PASS: A Dutch data-to-text system for soccerabstractWe present an evaluation of PASS, a data-to-text system that generates Dutch soccer reports from match statistics which are automatically tailored towards fans of one club or the other. The evaluation in this paper consists of two studies. An intrinsic human-based evaluation of the system’s output is described in the first study. In this study it was found that compared to human-written texts, computer-generated texts were rated slightly lower on style-related text components (fluency and clarity) and slightly higher in terms of the correctness of given information. Furthermore, results from the first study showed that tailoring was accurately recognized in most cases, and that participants struggled with correctly identifying whether a text was written by a human or computer. The second study investigated if tailoring affects perceived text quality, for which no results were garnered. This lack of results might be due to negative preconceptions about computer-generated texts which were found in the first study. Chris van der Lee, Bart Verduijn, Emiel Krahmer, Sander Wubben |
COLING | 3 |
| 2018 | DIDEC: The Dutch Image Description and Eye-tracking CorpusabstractWe present a corpus of spoken Dutch image descriptions, paired with two sets of eye-tracking data: Free viewing, where participants look at images without any particular purpose, and Description viewing, where we track eye movements while participants produce spoken descriptions of the images they are viewing. This paper describes the data collection procedure and the corpus itself, and provides an initial analysis of self-corrections in image descriptions. We also present two studies showing the potential of this data. Though these studies mainly serve as an example, we do find two interesting results: (1) the eye-tracking data for the description viewing task is more coherent than for the free-viewing task; (2) variation in image descriptions (also called ‘image specificity’; Jas and Parikh, 2015) is only moderately correlated across different languages. Our corpus can be used to gain a deeper understanding of the image description task, particularly how visual attention is correlated with the image description process. Emiel van Miltenburg, Ákos Kádár, Ruud Koolen, Emiel Krahmer |
COLING | 4 |
| 2018 | The Effect of a Robot's Gestures and Adaptive Tutoring on Children's Acquisition of Second Language VocabulariesabstractThis paper presents a study in which children, four to six years old, were taught words in a second language by a robot tutor. The goal is to evaluate two ways for a robot to provide scaffolding for students: the use of iconic gestures, combined with adaptively choosing the next learning task based on the child»s past performance. The results show a positive effect on long-term memorization of novel words, and an overall higher level of engagement during the learning activities when gestures are used. The adaptive tutoring strategy reduces the extent to which the level of engagement is diminishing during the later part of the interaction. Jan de Wit, Thorsten Schodde, Bram Willemsen, Kirsten Bergmann, Mirjam de Haas, Stefan Kopp, Emiel Krahmer, Paul Vogt |
HRI | 7 |
| 2018 | Enriching the WebNLG corpusabstractThis paper describes the enrichment of WebNLG corpus (Gardent et al., 2017a,b), with the aim to further extend its usefulness as a resource for evaluating common NLG tasks, including Discourse Ordering, Lexicalization and Referring Expression Generation.We also produce a silverstandard German translation of the corpus to enable the exploitation of NLG approaches to other languages than English.The enriched corpus is publicly available 1 . Thiago Castro Ferreira, Diego Moussallem, Emiel Krahmer, Sander Wubben |
INLG | 3 |
| 2018 | Automated learning of templates for data-to-text generation: comparing rule-based, statistical and neural methodsabstractThe current study investigated novel techniques and methods for trainable approaches to data-to-text generation. Neural Machine Translation was explored for the conversion from data to text as well as the addition of extra templatization steps of the data input and text output in the conversion process. Evaluation using BLEU did not find the Neural Machine Translation technique to perform any better compared to rule-based or Statistical Machine Translation, and the templatization method seemed to perform similarly or sometimes worse compared to direct data-to-text conversion. However, the human evaluation metrics indicated that Neural Machine Translation yielded the highest quality output and that the templatization method was able to increase text quality in multiple situations. Chris van der Lee, Emiel Krahmer, Sander Wubben |
INLG | 2 |
| 2018 | Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluationabstractThis paper surveys the current state of the art in Natural Language Generation (NLG), defined as the task of generating text or speech from non-linguistic input. A survey of NLG is timely in view of the changes that the field has undergone over the past two decades, especially in relation to new (usually data-driven) methods, as well as new applications of NLG technology. This survey therefore aims to (a) give an up-to-date synthesis of research on the core tasks in NLG and the architectures adopted in which such tasks are organised; (b) highlight a number of recent research topics that have arisen partly as a result of growing synergies between NLG and other areas of artificial intelligence; (c) draw attention to the challenges in NLG evaluation, relating them to similar challenges faced in other areas of NLP, with an emphasis on different evaluation methods and the relationships between them. Albert Gatt, Emiel Krahmer |
J. Artif. Intell. Res. | 2 |
| 2017 | Automatic Summarization of Domain-specific Forum Threads: Collecting Reference DataabstractWe create and analyze two sets of reference summaries for discussion threads on a patient support forum: expert summaries and crowdsourced, non-expert summaries. Ideally, reference summaries for discussion forum threads are created by expert members of the forum community. When there are few or no expert members available, crowdsourcing the reference summaries is an alternative. In this paper we investigate whether domain-specific forum data requires the hiring of domain experts for creating reference summaries. We analyze the inter-rater agreement for both data-sets and we train summarization models using the two types of reference summaries. The inter-rater agreement in crowdsourced reference summaries is low, close to random, while domain experts achieve a considerably higher, fair, agreement. The trained models however are similar to each other. We conclude that it is possible to train an extractive summarization model on crowdsourced data that is similar to an expert model, even if the inter-rater agreement for the crowdsourced data is low. Suzan Verberne, Antal van den Bosch, Sander Wubben, Emiel Krahmer |
CHIIR | 4 |
| 2017 | Perspective-Taking in Referential Communication: Does Stimulated Attention to Addressee's Perspective Influence Speakers' Reference Production?
Debby Damen, Per van der Wijst, Marije van Amelsvoort, Emiel Krahmer |
CogSci | 4 |
| 2017 | Do Speaker's Emotions influence their Language Production? Studying the Influence of Disgust and Amusement on Alignment in Interactive Reference
Charlotte Out, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2017 | Generating flexible proper name references in text: Data, models and evaluationabstractThis study introduces a statistical model able to generate variations of a proper name by taking into account the person to be mentioned, the discourse context and variation.The model relies on the REGnames corpus, a dataset with 53,102 proper name references to 1,000 people in different discourse contexts.We evaluate the versions of our model from the perspective of how human writers produce proper names, and also how human readers process them.The corpus 1 and the model 2 are publicly available. Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben |
EACL (1) | 2 |
| 2017 | Linguistic realisation as machine translation: Comparing different MT models for AMR-to-text generationabstractIn this paper, we study AMR-to-text generation, framing it as a translation task and comparing two different MT approaches (Phrasebased and Neural MT).We systematically study the effects of 3 AMR preprocessing steps (Delexicalisation, Compression, and Linearisation) applied before the MT phase.Our results show that preprocessing indeed helps, although the benefits differ for the two MT models.The implementation of the models are publicly available 1 . Thiago Castro Ferreira, Iacer Calixto, Sander Wubben, Emiel Krahmer |
INLG | 4 |
| 2017 | PASS: A Dutch data-to-text system for soccer, targeted towards specific audiencesabstractWe present PASS, a data-to-text system that generates Dutch soccer reports from match statistics.One of the novel elements of PASS is the fact that the system produces corpusbased texts tailored towards fans of one club or the other, which can most prominently be observed in the tone of voice used in the reports.Furthermore, the system is open source and uses a modular design, which makes it relatively easy for people to add extensions.Human-based evaluation shows that people are generally positive towards PASS in regards to its clarity and fluency, and that the tailoring is accurately recognized in most cases. Chris van der Lee, Emiel Krahmer, Sander Wubben |
INLG | 2 |
| 2017 | L1 Perceptions of L2 Prosody: The Interplay Between Intonation, Rhythm, and Speech Rate and Their Contribution to Accentedness and ComprehensibilityabstractThis study investigates the cumulative effect of (non-)native intonation, rhythm, and speech rate in utterances produced by Spanish learners of Dutch on Dutch native listeners’ perceptions. In order to assess the relative contribution of these language-specific properties to perceived accentedness and comprehensibility, speech produced by Spanish learners of Dutch was manipulated using transplantation and resynthesis techniques. Thus, eight manipulation conditions reflecting all possible combinations of L1 and L2 intonation, rhythm, and speech rate were created, resulting in 320 utterances that were rated by 50 Dutch natives on their degree of foreign accent and ease of comprehensibility. Our analyses show that all manipulations result in lower accentedness and higher comprehensibility ratings. Moreover, both measures are not affected in the same way by different combinations of prosodic features: For accentedness, Dutch listeners appear most influenced by intonation, and intonation combined with speech rate. This holds for comprehensibility ratings as well, but here the combination of all three properties, including rhythm, also significantly affects ratings by native speakers. Thus, our study reaffirms the importance of differentiating between different aspects of perception and provides insight into those features that are most likely to affect how native speakers perceive second language learners. Lieke van Maastricht, Tim Zee, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 3 |
| 2016 | Towards more variation in text generation: Developing and evaluating variation models for choice of referential formabstractIn this study, we introduce a nondeterministic method for referring expression generation. We describe two models that account for individual variation in the choice of referential form in automatically generated text: a Naive Bayes model and a Recurrent Neural Network. Both are evaluated using the VaREG corpus. Then we select the best performing model to generate referential forms in texts from the GREC-2.0 corpus and conduct an evaluation experiment in which humans judge the coherence and comprehensibility of the generated texts, comparing them both with the original references and those produced by a random baseline model. Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben |
ACL (1) | 2 |
| 2016 | Referential choice in identification and route directions
Adriana Alexandra Baltaretu, Emiel Krahmer, Alfons Maes |
CogSci | 2 |
| 2016 | A connectionist model for automatic generation of child-adult interaction patterns
Moinuddin M. Haque, Paul Vogt, Afra Alishahi, Emiel Krahmer |
CogSci | 4 |
| 2016 | Viewing time affects overspecification: Evidence for two strategies of attribute selection during reference production
Ruud Koolen, Albert Gatt, Roger P. G. van Gompel, Emiel Krahmer, Kees van Deemter |
CogSci | 4 |
| 2016 | The Influence of Language-specific Auditory Cues on the Learnability of Center-embedded Recursion
Jun Lai, Chiara de Jong, Dingguo Gao, Ren Huang, Emiel Krahmer, Jan Sprenger |
CogSci | 5 |
| 2016 | Learning How To Throw Darts: The Effect Of Modeling Type And Reflection On Dart-Throwing Skills
Janneke van der Loo, Eefje Frissen, Emiel Krahmer |
CogSci | 3 |
| 2016 | The Multilingual Affective Soccer Corpus (MASC): Compiling a biased parallel corpus on soccer reportage in English, German and DutchabstractThe emergence of the internet has led to a whole range of possibilities to not only collect large, but also highly specified text corpora for linguistic research.This paper introduces the Multilingual Affective Soccer Corpus.MASC is a collection of soccer match reports in English, German and Dutch.Parallel texts are collected manually from the involved soccer clubs' homepages with the aim of investigating the role of affect in sports reportage in different languages and cultures, taking into account the different perspectives of the teams and possible outcomes of a match.The analyzed aspects of emotional language will open up new approaches for biased automatic generation of texts. Nadine Braun, Martijn Goudbeek, Emiel Krahmer |
INLG | 3 |
| 2016 | Towards proper name generation: a corpus analysisabstractWe introduce a corpus for the study of proper name generation.The corpus consists of proper name references to people in webpages, extracted from the Wikilinks corpus.In our analyses, we aim to identify the different ways, in terms of length and form, in which a proper names are produced throughout a text. Thiago Castro Ferreira, Sander Wubben, Emiel Krahmer |
INLG | 3 |
| 2016 | Abstractive Compression of Captions with Attentive Recurrent Neural NetworksabstractIn this paper we introduce the task of abstractive caption or scene description compression.We describe a parallel dataset derived from the FLICKR30K and MSCOCO datasets.With this data we train an attention-based bidirectional LSTM recurrent neural network and compare the quality of its output to a Phrasebased Machine Translation (PBMT) model and a human generated short description.An extensive evaluation is done using automatic measures and human judgements.We show that the neural model outperforms the PBMT model.Additionally, we show that automatic measures are not very well suited for evaluating this text-to-text generation task. Sander Wubben, Emiel Krahmer, Antal van den Bosch, Suzan Verberne |
INLG | 2 |
| 2016 | Individual Variation in the Choice of Referential FormabstractThiago Castro Ferreira, Emiel Krahmer, Sander Wubben. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben |
HLT-NAACL | 2 |
| 2016 | Native speaker perceptions of (non-)native prominence patterns: Effects of deviance in pitch accent distributions on accentedness, comprehensibility, intelligibility, and nativeness
Lieke van Maastricht, Emiel Krahmer, Marc Swerts |
Speech Commun. | 2 |
| 2015 | Landmarks in motion: Unstable entities in route directions
Adriana Alexandra Baltaretu, Emiel Krahmer, Alfons Maes |
CogSci | 2 |
| 2015 | The learnability of Auditory Center-embedded Recursion
Jun Lai, Emiel Krahmer, Jan Sprenger |
CogSci | 2 |
| 2015 | On what happens in gesture when communication is unsuccessful
Marieke Hoetjes, Emiel Krahmer, Marc Swerts |
Speech Commun. | 2 |
| 2014 | On what happens in speech and gesture when communication is unsuccessful
Marieke Hoetjes, Emiel Krahmer, Marc Swerts |
CogSci | 2 |
| 2014 | How perceived distractor distance influences reference production: Effects of perceptual grouping in 2D and 3D scenes
Ruud Koolen, Eugene Houben, Jan Huntjens, Emiel Krahmer |
CogSci | 4 |
| 2014 | Studying Frequency Effects in Learning Center-embedded Recursion
Jun Lai, Emiel Krahmer, Jan Sprenger |
CogSci | 2 |
| 2014 | On the automaticity of reduction in dialogue: Cognitive load and repeated multimodal references
Ingrid Masson-Carro, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2014 | Pantomime Strategies: On Regularities in How People Translate Mental Representations into the Gesture Modality
Karin van Nispen, Mieke van de Sandt-Koenderman, Lisette Mol, Emiel Krahmer |
CogSci | 4 |
| 2014 | The Use of Colour in Reference Production: A Comparison between Dutch and Greek
Mirjana Sekicki, Jette Viethen, Martijn Goudbeek, Emiel Krahmer |
CogSci | 4 |
| 2014 | Nonverbal Cues of Meta-Memory Awareness in Older Adults
Mandy Visser, Marie Postma, Emiel Krahmer, Marc Swerts |
CogSci | 3 |
| 2014 | Creating and using large monolingual parallel corpora for sentential paraphrase generation
Sander Wubben, Antal van den Bosch, Emiel Krahmer |
LREC | 3 |
| 2014 | Does our speech change when we cannot gesture?
Marieke Hoetjes, Emiel Krahmer, Marc Swerts |
Speech Commun. | 2 |
| 2013 | Workshop Proposal: PRE-CogSci 2013: Bridging the gap between cognitive and computational approaches to reference
Albert Gatt, Roger P. G. van Gompel, Ellen Gurman Bard, Emiel Krahmer, Kees van Deemter |
CogSci | 4 |
| 2013 | Production of referring expressions: Preference trumps discrimination
Albert Gatt, Emiel Krahmer, Roger P. G. van Gompel, Kees van Deemter |
CogSci | 2 |
| 2013 | The object without qualities: referring with negative properties
Martijn Goudbeek, Inge Haagmans, Emiel Krahmer |
CogSci | 3 |
| 2013 | The impact of bottom-up and top-down saliency cues on reference production
Ruud Koolen, Emiel Krahmer, Marc Swerts |
CogSci | 2 |
| 2013 | How big is the BFG? The impact of redundant size adjectives on size perception
Emiel Krahmer, Marret Noordewier, Martijn Goudbeek, Ruud Koolen |
CogSci | 1 |
| 2013 | Cognitive load does not decrease pronoun use when speaker's and addressee's perspectives are dissociated
Jorrig Vogels, Emiel Krahmer, Alfons Maes |
CogSci | 2 |
| 2013 | Crosslinguistic priming in interactive reference: evidence for conceptual alignment in speech production
Anne Vullinghs, Martijn Goudbeek, Emiel Krahmer |
INTERSPEECH | 3 |
| 2013 | Positive Affective Interactions: The Role of Repeated Exposure and CopresenceabstractWe describe and evaluate a new interface to induce positive emotions in users: a digital, interactive adaptive mirror. We study whether the induced affect is repeatable after a fixed interval (Study 1) and how copresence influences the emotion induction (Study 2). Results show that participants systematically feel more positive after an affective mirror session, that this effect is repeatable, and stronger when a friend is copresent. Suleman Shahid, Emiel Krahmer, Mark A. Neerincx, Marc Swerts |
IEEE Trans. Affect. Comput. | 2 |
| 2012 | Sentence Simplification by Monolingual Machine Translation
Sander Wubben, Antal van den Bosch, Emiel Krahmer |
ACL (1) | 3 |
| 2012 | Does domain size impact speech onset time during reference production?
Albert Gatt, Roger P. G. van Gompel, Emiel Krahmer, Kees van Deemter |
CogSci | 3 |
| 2012 | Do repeated references result in sign reduction?
Marieke Hoetjes, Emiel Krahmer, Marc Swerts |
CogSci | 2 |
| 2012 | Conceptual alignment in reference with artificial and human dialogue partners
Koen van Lierop, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2012 | The Impact of Colour Difference and Colour Codability on Reference Production
Jette Viethen, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2012 | Factors influencing children's display of surprise
Mandy Visser, Emiel Krahmer, Marc Swerts |
CogSci | 2 |
| 2012 | Learning Preferences for Referring Expression Generation: Effects of Domain, Language and Algorithm
Ruud Koolen, Emiel Krahmer, Mariët Theune |
INLG | 2 |
| 2012 | Contrastive intonation in autism: The effect of speaker- and listener-perspective
Constantijn Kaland, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 2 |
| 2012 | Computational Generation of Referring Expressions: A SurveyabstractThis article offers a survey of computational research on referring expression generation (REG). It introduces the REG problem and describes early work in this area, discussing what basic assumptions lie behind it, and showing how its remit has widened in recent years. We discuss computational frameworks underlying REG, and demonstrate a recent trend that seeks to link REG algorithms with well-established Knowledge Representation techniques. Considerable attention is given to recent efforts at evaluating REG algorithms and the lessons that they allow us to learn. The article concludes with a discussion of the way forward in REG, focusing on references in larger and more realistic settings. Emiel Krahmer, Kees van Deemter |
Comput. Linguistics | 1 |
| 2012 | Video-mediated and co-present gameplay: Effects of mutual gaze on game experience, expressiveness and perceived social presenceabstractWe study how pairs of children interact socially and express their emotions while playing games in different communicative settings. In particular, we study how such interactions can vary for environments that differ regarding the level of mediation and the associated feelings of social presence. Overall, the study compared three conditions (one face-to-face gameplay condition, and two video-mediated gameplay conditions; one allowing for mutual gaze, the other not) and focused on the social presence and non-verbal behavior of children in three conditions. The results show that the presence of mutual eye-gaze enriches the feelings of social presence, fun and game experience; conversely, the absence of mutual eye-gaze dramatically effects the quality of interaction in the video-mediated environment. The results of this study stress the importance of mutual gaze, and we therefore argue that it should become an integral component of future VMC systems, particularly in those designed for playful settings and children. Suleman Shahid, Emiel Krahmer, Marc Swerts |
Interact. Comput. | 2 |
| 2011 | PRE-CogSci 2011 - Bridging the gap between computational, empirical and theoretical approaches to reference
Kees van Deemter, Albert Gatt, Roger P. G. van Gompel, Emiel Krahmer |
CogSci | 4 |
| 2011 | Attribute preference and priming in reference production: Experimental evidence and computational modeling
Albert Gatt, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2011 | GREEBLES Greeble greeb. On reduction in speech and gesture in repeated references
Marieke Hoetjes, Ruud Koolen, Martijn Goudbeek, Emiel Krahmer, Marc Swerts |
CogSci | 4 |
| 2011 | Salient in the mind, salient in prosody
Constantijn Kaland, Emiel Krahmer, Marc Swerts |
CogSci | 2 |
| 2011 | Effects of scene variation on referential overspecification
Ruud Koolen, Martijn Goudbeek, Emiel Krahmer |
CogSci | 3 |
| 2011 | Gesturing by aphasic speakers, how does it compare?
Lisette Mol, Emiel Krahmer, Mieke van de Sandt-Koenderman |
CogSci | 2 |
| 2011 | How visual saliency affects referent accessibility
Jorrig Vogels, Emiel Krahmer, Alfons Maes |
CogSci | 2 |
| 2011 | Who is more expressive during child-robot interaction: Pakistani or Dutch children?abstractIn this study we have tried to determine if the cultural background of children has an influence on how they interact with robots. Children of different age groups and cultures played a card guessing game with a robot (iCat). By using perception tests to evaluate the children's emotional response it was revealed that children from South Asia (Pakistani) were much more expressive than European children (Dutch) and younger children were more expressive than the older ones in the context of child robot interaction. Suleman Shahid, Emiel Krahmer, Marc Swerts, Omar Mubin |
HRI | 2 |
| 2010 | Automatic analysis of semantic similarity in comparable text through syntactic tree matching
Erwin Marsi, Emiel Krahmer |
COLING | 2 |
| 2010 | Cross-linguistic Attribute Selection for REG: Comparing Dutch and English
Mariët Theune, Ruud Koolen, Emiel Krahmer |
INLG | 3 |
| 2010 | Paraphrase Generation as Monolingual Translation: Data and Evaluation
Sander Wubben, Antal van den Bosch, Emiel Krahmer |
INLG | 3 |
| 2010 | The D-TUNA Corpus: A Dutch Dataset for the Evaluation of Referring Expression Generation Algorithms
Ruud Koolen, Emiel Krahmer |
LREC | 2 |
| 2010 | Human Language Technology and Communicative Disabilities: Requirements and Possibilities for the Future
Marina B. Ruiter, Toni C. M. Rietveld, Catia Cucchiarini, Emiel Krahmer, Helmer Strik |
LREC | 4 |
| 2010 | Using child-robot interaction to investigate the user acceptance of constrained and artificial languagesabstractThe possibility of improving speech recognition accuracy within human computer and robot interaction by having users interact in a constrained natural language or an artificial language has been explored. However, what has not been evaluated yet is the user acceptance of such forms of interaction. In this paper we discuss two separate but similar studies which were aimed at assessing the usability of constrained and artificial languages in contrast to natural languages. The interaction context was implemented in a game played between children and the iCat robot. We subjectively measured various variables related to gaming experience and interacting in the new languages. Our results reveal that there were no significant differences in the user experience across the two interaction mediums in comparison to natural languages. Omar Mubin, Suleman Shahid, Eva van de Sande, Emiel Krahmer, Marc Swerts, Christoph Bartneck, Loe M. G. Feijs |
RO-MAN | 4 |
| 2010 | What Computational Linguists Can Learn from Psychologists (and Vice Versa)abstractSometimes I am amazed by how much the field of computational linguistics haschanged in the past 15 to 20 years. In the mid-nineties, I was working in a researchinstitute where language and speech technologists worked in relatively close quarters.Speech technology seemed on the verge of a major breakthrough; this was around thetime that Bill Gates was quoted in Business Week as saying that speech was not justthe future of Windows, but the future of computing itself. At the same time, languagetechnology was, well, nowhere. Bill Gates certainly wasn’t championing language tech-nology in those days. And while the possible applications of speech technology seemedendless (who would use a keyboard in 2010, when speech-driven user interfaces wouldhave replaced traditional computers?) the language people were thinking hard aboutpossible applications for their admittedly somewhat immature technologies.Predicting the future is a tricky thing. No major breakthrough came for speechtechnology — I am still typing this. However, language technology did change almostbeyond recognition. Perhaps one of the main reasons for this has been the explosivegrowth of the internet, which helped language technology in two different ways. Onthe one hand it instigated the development and refinement of techniques needed forsearching in document collections of unprecedented size, on the other it resulted in alarge increase of freely available text data. Recently, language technology has been par-ticularly successful for tasks where huge amounts of textual data is available to whichstatistical machine learning techniques can be applied (Halevy, Norvig, and Pereira2009). As a result of these developments, mainstream computational linguistics is nowa successful, application-oriented discipline which is particularly good at extractinginformation from sequences of words.But there is more to language than that. For speakers, words are the result of acomplex speech production process; for listeners they are what starts off the similarlycomplex comprehension process. However, in many current applications no attentionis given to the processes by which words are produced nor to the processes by whichthey can be understood. Language is treated as a product not as a process, in theterminology of Clark (1996). In addition, we use language not only as a vehicle forfactual information exchange; speakers may have all sorts of other intentions with theirwords; they may want to convince others to do or buy something, they may want toinduce a particular emotion in the addressee etc. These days, most of computationallinguistics (with a few notable exceptions, more about which below) has little to say Emiel Krahmer |
Comput. Linguistics | 1 |
| 2008 | GRAPH: The Costs of Redundancy in Referring Expressions
Emiel Krahmer, Mariët Theune, Jette Viethen, Iris Hendrickx |
INLG | 1 |
| 2008 | On the role of acting skills for the collection of simulated emotional speechabstractWe experimentally compared non-simulated with simulated expressions of emotion produced both by inexperienced and by experienced actors.Contrary to our expectations, in a perception experiment participants rated the expressions of experienced actors as more extreme and less like non-simulated ("real") expressions than those produced by non-professional actors. Emiel Krahmer, Marc Swerts |
INTERSPEECH | 1 |
| 2008 | Nonverbal responses to social inclusion and exclusionabstractApplying an elaborate and novel experimental paradigm we collected non-verbal responses to social inclusion and exclusion. A judgement experiment revealed that it was possible to determine whether a person was included or excluded, based solely on non-verbal, behavioral cues, especially if the person was male. A detailed coding using the ECSI scale suggests that included speakers displayed more cues associated with Affiliation and Relaxation, whereas excluded speakers gave more evidence of Flight and Displacement. Emiel Krahmer, Juliette Schaafsma, Marc Swerts, Ad Vingerhoets |
INTERSPEECH | 1 |
| 2008 | Gender-related differences in the production and perception of emotion
Marc Swerts, Emiel Krahmer |
INTERSPEECH | 2 |
| 2008 | Controlling Redundancy in Referring Expressions
Jette Viethen, Robert Dale, Emiel Krahmer, Mariët Theune, Pascal Touset |
LREC | 3 |
| 2007 | Incremental perception of acted and real emotional speechabstractThis paper reports on an experiment using the gating paradigm to test the recognition speed for various emotional expressions from a speaker’s face. In a perception experiment, subjects were presented with video clips of speakers who displayed negative or positive emotions, which were either acted or real. The clips were shown in successive segments (gates) of increasing duration. Results show that subjects are surprisingly accurate in their recognition of the various emotions, as they already reach high recognition scores in the first gate (after only 160 milliseconds). Interestingly, the recognition speed is faster for positive than negative emotions, in line with comparable valency effects reported by Leppanen and Hietanen (2003). Finally, the gating results confirm earlier findings that acted emotions are perceived as more intense than true emotions (Wilting et al., 2006), as the former get more extreme recognition scores than the latter, already after a short period of exposure. Pashiera Barkhuysen, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 2 |
| 2007 | Using eye movements for online evaluation of speech synthesisabstractThis paper * describes an eye tracking experiment to study the processing of diphone synthesis, unit selection synthesis, and human speech taking segmental and suprasegmental speech quality into account. The results showed that both factors influenced the processing of human and synthetic speech, and confirmed that eye tracking is a promising albeit time consuming research method to evaluate synthetic speech. Index Terms: eye movements, human speech, diphone synthesis, unit selection synthesis. Charlotte van Hooijdonk, Edwin Commandeur, Reinier Cozijn, Emiel Krahmer, Erwin Marsi |
INTERSPEECH | 4 |
| 2007 | Audiovisual emotional speech of game playing children: effects of age and cultureabstractIn this paper we study how children of different age groups (8 and 12 years old) and with different cultural backgrounds (Dutch and Pakistani) signal positive and negative emotions in audiovisual speech. Data was collected in an ethical way using a simple but surprisingly effective game in which pairs of participants have to guess whether an upcoming card will contain a higher or lower number than a reference card. The data thus collected was used in a series of cross-cultural perception studies, in which Dutch and Pakistani observers classified emotional expressions of Dutch and Pakistani speakers. Results show that classification accuracy is uniformly high for Pakistani children, but drops for older and for winning Dutch children 1. Suleman Shahid, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 2 |
| 2006 | How auditory and visual prosody is used in end-of-utterance detectionabstractIn this paper, we describe a series of perception studies using visual and auditory cues to end-of-utterance. Fragments were taken from a recorded interview session, consisting of the parts in which speakers provided answers. Final and non-final parts of these fragments were used, varying in length. The subjects had to assess whether the speaker had finished his or her turn, based upon these fragments. The fragments were presented in 3 modalities: either a bimodal presentation mode (both auditory and visually), or in only the auditory or the visual mode. Results show that the audio-visual condition evoked the highest proportion of correct classifications and the auditory condition the lowest. Thus, the combination of modalities clearly works best. Also, non-final fragments are classified better than final ones, and longer fragments are classified better than short ones. It furthermore appears that these factors are different for different modalities: longer fragments are better classified in the auditory modality, while for short fragments the visual modality works better. This suggests that people may make more use of global cues in the auditory modality, while for the visual modality local cues are sufficient. Index Terms: Audiovisual speech, prosody, end-of-utterance detection, speech production, speech perception. Pashiera Barkhuysen, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 2 |
| 2006 | Testing the effect of audiovisual cues to prominence via a reaction-time experimentabstractThis article discusses a perception experiment to investigate the relation between auditory and visual cues for marking prosodic prominence.The methodology makes use of a reaction-time experiment.For this experiment, recordings of a sentence with 3 accents were systematically manipulated in such a way that auditory and visual markers of prominence were either congruent (occurring on the same word) or incongruent (in that the auditory and the visual cues were positioned on different words).Subjects were instructed to indicate as fast as possible which word they perceived as the most prominent one.Classification results show first of all that subjects' responses were much more dependent on auditory than on visual cues.In addition, however, we found that incongruent stimuli lead to slower reaction times than congruent stimuli, showing that visual cues do have an impact on the cognitive processing of prosodic prominence. Emiel Krahmer, Marc Swerts |
INTERSPEECH | 1 |
| 2006 | The importance of different facial areas for signalling visual prominenceabstractThis article discusses the processing of facial markers of promi-nence in spoken utterances. In particular, it investigates which area of a speaker’s face contains the strongest cues to promi-nence, using stimuli with the entire face visible or versions in which participants could only see the upper or lower half, or the right or left part of the face. To compensate for potential ceiling effects, subjects were positioned at a distance of either 50cm, 250cm or 380cm from the screen which displayed the film frag-ments. The task of the subjects was to indicate for each stimulus which word they perceived as the most prominent one. Results show that, while prominence detection becomes more difficult at longer distances, the upper facial area has stronger cue value for prominence detection than the bottom part, and that the left part of the face is more important than the right part. Results of mirror-images of the original fragments show that this latter result is due both to a speaker and an observer effect. Index Terms: prominence, facial areas, audiovisual speech 1. Marc Swerts, Emiel Krahmer |
INTERSPEECH | 2 |
| 2006 | Real vs. acted emotional speechabstractEven though the use of actors is a popular method for researching the expression of emotion, little is known about the relation between acted and real emotions. To shed some light on this, we set up a novel experiment, based on the Velten mood induction procedure, during which participants have to utter pre-defined sentences with a strong emotional content. In one group of participants, real positive or negative emotions were induced, while another group was instructed to act positive or negative while uttering Velten sentences. Results of a mood questionnaire revealed that participants in the real emotion condition, indeed felt positive or negative, depending on whether they read positive or negative sentences, while participants in the acted emotion condition felt neutral afterwards. In a second, perception experiment, it was found that acted emotions (especially negative ones) were perceived more strongly than the real emotions. This suggests that actors do not feel the acted emotion, and may engage in overacting, which casts doubt on the usefulness of actors as a way to study real emotions. Janneke Wilting, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 2 |
| 2005 | Predicting end of utterance in multimodal and unimodal conditionsabstractIn this paper, we describe a series of perception studies on uni-and multimodal cues to end of utterance. Stimuli were fragments taken from a recorded interview session, consisting of the parts in which speakers provided answers. The answers varied in length and were presented without the preceding question of the interviewer. The subjects had to predict when the speaker would finish his turn, based on video material and/or auditory material. The experiment consisted of 3 conditions: in one condition, the stimuli were presented as they were recorded (both audio and vision), in the two remaining conditions stimuli were presented in only the auditory or the visual channel.Results show that the audiovisual condition evoked the fastest reaction times and the visual condition the slowest. Arguably, the combination of cues from different modalities function as complementary sources and might thus improve prediction.determine the relative weight of the different modalities used for end of utterance marking. We compare a multimodal condition in which subjects have both auditory and visual cues at their disposal (stimuli presented as they were produced) with two unimodal conditions where subjects could only use auditory or visual cues. Pashiera Barkhuysen, Emiel Krahmer, Marc Swerts |
INTERSPEECH | 2 |
| 2005 | Real versus Template-Based Natural Language Generation: A False Opposition?abstractThis article challenges the received wisdom that template-based approaches to the generation of language are necessarily inferior to other approaches as regards their maintainability, linguistic well-foundedness, and quality of output. Some recent NLG systems that call themselves “template-based” will illustrate our claims. Kees van Deemter, Mariët Theune, Emiel Krahmer |
Comput. Linguistics | 3 |
| 2005 | Problem detection in human-machine interactions based on facial expressions of users
Pashiera Barkhuysen, Emiel Krahmer, Marc Swerts |
Speech Commun. | 2 |
| 2004 | Signaling and detecting uncertainty in audiovisual speech by children and adultsabstractWe describe two experiments on signaling and detecting uncertainty in audiovisual speech by adults and children. In the first study, ut-terances from adult speakers and child speakers (aged 7-8) were elicitated and annotated with a set of six audiovisual features. It was found that when adult speakers are uncertain about their answer they are more likely to produce filled pauses, delays, high intonation, eyebrow movements, smiles and funny faces. The basic picture for the child speakers is similar, in that the presence of an audiovisual cue in an answer correlates with uncertainty, but the differences are relatively small and only significant for the features delay, eyebrow and funny face. In the second study both adult and child judges watched answers from adult and child speakers selected from the first study to find out whether they were able to correctly estimate a speakers ’ level of uncertainty. It was found that both child and adult judges give more accurate scores for answers from adult speakers than from child speakers and that child judges overall provide less accurate scores than adult judges. 1. Emiel Krahmer, Marc Swerts |
INTERSPEECH | 1 |
| 2004 | The influence of target size and distance on the production of speech and gesture in multimodal referring expressionsabstractIn this paper we report on a production experiment for multimodal referring expressions. Subjects performed an object identification task in an interactive setting. 20 subjects participated and were asked if they could identify 30 countries on a world map on the wall. Subjects performed their tasks on two distances: close (10 subjects) and at a distance of 2.5 meters (10 subjects). The assumption is that these conditions yield precise and imprecise pointing gestures respectively. In addition we varied the ’size’ of target objects (large or isolated objects versus small objects). This study resulted in a corpus of 600 multimodal referring expressions. A statistical analysis (ANOVA) revealed a main effect of distance (subjects adapt their language to the kind of pointing gesture) and also a main effect of target (smaller objects are more difficult to describe than large or isolated objects). Ielka van der Sluis, Emiel Krahmer |
INTERSPEECH | 2 |
| 2004 | Evaluating Multimodal NLG Using Production Experiments
Ielka van der Sluis, Emiel Krahmer |
LREC | 2 |
| 2003 | Graph-Based Generation of Referring ExpressionsabstractThis article describes a new approach to the generation of referring expressions. We propose to formalize a scene (consisting of a set of objects with various properties and relations) as a labeled directed graph and describe content selection (which properties to include in a referring expression) as a subgraph construction problem. Cost functions are used to guide the search process and to give preference to some solutions over others. The current approach has four main advantages: (1) Graph structures have been studied extensively, and by moving to a graph perspective we get direct access to the many theories and algorithms for dealing with graphs; (2) many existing generation algorithms can be reformulated in terms of graphs, and this enhances comparison and integration of the various approaches; (3) the graph perspective allows us to solve a number of problems that have plagued earlier algorithms for the generation of referring expressions; and (4) the combined use of graphs and cost functions paves the way for an integration of rule-based generation techniques with more recent stochastic approaches. Emiel Krahmer, Sebastiaan van Erk, Andre Verleg |
Comput. Linguistics | 1 |
| 2002 | Perceptual evaluation of audiovisual cues for prominenceabstractThis paper reports on two experiments with a Talking Head that explore the ability of eyebrow movements to cue focus. The first experiment tests how listeners react to synthetic stimuli in which the eyebrow movements coincide with pitch accents versus those in which these two occur on different words. Results show that subjects prefer those utterances in which pitch and eyebrow movements are aligned on the same word. The second experiment investigates whether listeners are sensitive to eyebrow movements when they have to rate the prominence of particular words in audiovisual stimuli. This experiment shows that eyebrow movements both boost the perceived prominence of words that also receive a pitch accent, and downscale the prominence of unaccented words in the immediate context of the accented word. 1. Emiel Krahmer, Zsófia Ruttkay, Marc Swerts, Wieger Wesselink |
INTERSPEECH | 1 |
| 2002 | The dual of denial: Two uses of disconfirmations in dialogue and their prosodic correlates
Emiel Krahmer, Marc Swerts, Mariët Theune, Mieke F. Weegels |
Speech Commun. | 1 |
| 2001 | Detecting Problematic Turns in Human-Machine Interactions: Rule-induction Versus Memory-based Learning ApproachesabstractWe address the issue of on-line detection of communication problems in spoken dialogue systems. The usefulness is investigated of the sequence of system question types and the word graphs corresponding to the respective user utterances. By applying both rule-induction and memory-based learning techniques to data obtained with a Dutch train time-table information system, the current paper demonstrates that the aforementioned features indeed lead to a method for problem detection that performs significantly above baseline. The results are interesting from a dialogue perspective since they employ features that are present in the majority of spoken dialogue systems and can be obtained with little or no computational overhead. The results are interesting from a machine learning perspective, since they show that the rule-based method performs significantly better than the memory-based method, because the former is better capable of representing interactions between features. Antal van den Bosch, Emiel Krahmer, Marc Swerts |
ACL | 2 |
| 2001 | Reconstructing dialogue historyabstractThis paper deals with a perceptual analysis of accent structure in Dutch to see to what extent listeners are able to reconstruct information from the prior discourse context on the basis of prosodic properties of the current utterance. Using data collected in an earlier dialogue game experiment, subjects were asked to perform a perceptual task in which they had to try and reconstruct what the previous utterance was on the basis of input utterances with different accent patterns. Our results reveal that listeners are able to correctly guess the prior context for a significant number of cases, but that performance depends on the type of intonation contour of the input utterance. 1. Marc Swerts, Emiel Krahmer |
INTERSPEECH | 2 |
| 2001 | From data to speech: a general approach
Mariët Theune, Esther Klabbers, Jan-Roelof de Pijper, Emiel Krahmer, Jan Odijk |
Nat. Lang. Eng. | 4 |
| 2001 | On the alleged existence of contrastive accents
Emiel Krahmer, Marc Swerts |
Speech Commun. | 1 |
| 2000 | Preferred modalities in dialogue systemsabstractThis research describes which modalities are preferred in particular contexts when interacting with a multi-modal dialogue system. The trade-off between three factors is investigated: (i) speech recognition performance, (ii) efficiency of input modality and (iii) the system 's output modality. Four versions were developed of a multimodal examinator to be used in elementary school. The versions differed in recognition performance (`perfect' vs. realistic) and output modality (speech or text). In all systems, subjects could provide input via speaking or typing. Answer length in characters was used as a measure of efficiency. Results show that both speech recognition performance and efficiency have a strong impact on preferred modalities. No effect was found of the system's output modality. 1. INTRODUCTION "Speech is the bicycle of user-interface design," according to Shneiderman (1998:328), "it is great fun to use (. . . ), but it can carry only a light load. Sober advocates know that ... Vildan Bilici, Emiel Krahmer, Saskia te Riele, Raymond N. J. Veldhuis |
INTERSPEECH | 2 |
| 2000 | On the Use of Prosody for On-line Evaluation of Spoken Dialogue Systems
Marc Swerts, Emiel Krahmer |
LREC | 2 |
| 1999 | Problem spotting in human-machine interactionabstractIn human-human communication, dialogue participants are con-tinuously sending and receiving signals on the status of the inform-ation being exchanged. We claim that if spoken dialogue systems were able to detect such cues and change their strategy accordingly, the interaction between user and systemwould improve. Therefore, the goals of the present study are as follows: (i) to find out which positive and negative cues people actually use in human-machine interaction in response to explicit and implicit verification questions and (ii) to see which (combinations of) cues have the best predictive potential for spotting the presence or absence of problems. It was found that subjects systematically use negative/marked cues (more words, marked word order, more repetitions and corrections, less new information etc.) when there are communication problems. Using precision and recall matrices it was found that various combinations of cues are accurate problem spotters. This kind of information may turn out to be highly relevant for spoken dia-logue systems, e.g., by providing quantitative criteria for changing the dialogue strategy or speech recognition engine. Emiel Krahmer, Marc Swerts, Mariët Theune, Mieke F. Weegels |
EUROSPEECH | 1 |
| 1998 | A generic algorithm for generating spoken monologuesabstractThe defining property of a Concept-to-Speech system is that it combines language and speech generation. Language generation converts the input concepts into natural language, which speech generation subsequently transforms into speech. Potentially, this leads to a more `natural sounding' output than can be achieved in a plain Text-to-Speech system, since the correct placement of pitch accents and intonational boundaries ---an important factor contributing to the `naturalness' of the generated speech--- is co-determined by syntactic and discourse information, which is typically available in the language generation module. In this paper, a generic algorithm for the generation of coherent spoken monologues is discussed, called D2S. Language generation is done by a module called LGM which is based on TAG-like syntactic structures with open slots, combined with conditions which determine when the syntactic structure can be used properly. A speech generation module (SGM) converts the output of the LGM into speech using either phrase-concatenation or diphone-synthesis. Esther Klabbers, Emiel Krahmer, Mariët Theune |
ICSLP | 2 |
| 1998 | Reconciling two competing views on contrastivenessabstractSpeakers may use pitch accents as pointers to new information, or as signals of a contrast relation between the accented item and a limited set of alternatives. Some people claim that contrastive accents are more emphatic than newness accents and have a different melodic shape. Others, however, maintain that contrastiveness can only be determined by looking at how accents are distributed in an utterance. In this paper it is argued that these two competing views can be reconciled by showing that they apply on different levels. To this end, accent patterns were obtained in a (semi-)spontaneous way via a dialogue game (Dutch) in which two participants had to describe coloured figures in consecutive turns. By varying the sequential order, target descriptions ("blue square") were collected in four contexts: no contrast (all new), contrast in the adjective, contrast in the noun, all contrast. A distributional analysis revealed that both all new and all contrast situations correspond with dou... Emiel Krahmer, Marc Swerts |
ICSLP | 1 |
| 1998 | Context sensitive generation of descriptionsabstractProbably the best current algorithm for generating definite descriptions is the Incremental Algorithm due to Dale and Reiter. If we want to use this algorithm in a Concept-to-Speech system, however, we encounter two limitations: (i) the algorithm is insensitive to the linguistic context and thus always produces the same description for an object, (ii) the output is a list of properties which uniquely determine one object from a set of objects: how this list is to be expressed in spoken natural language is not addressed. We propose a modification of the Incremental Algorithm based on the idea that a definite description refers to the most salient element in the current context satisfying the descriptive content. We show that the modified algorithm allows for the context-sensitive generation of both distinguishing and anaphoric descriptions, while retaining the attractive properties of Dale and Reiter's original algorithm. 1. INTRODUCTION In their interesting 1995 paper, Dale & Reiter ... Emiel Krahmer, Mariët Theune |
ICSLP | 1 |
| 1997 | Robust spoken dialogue management for driver information systemsabstractConsidering the limitations of Speech Recognition for the development of user-system dialogues in real applications, robustness is a primary objective. In this paper, we describe the most essential characteristics of the Dialogue Manager of a driver information system that is controlled by voice, mainly showing how its design has been driven by the characteristics of voice in such a dialogue. We present the main methods used by the Dialogue Manager to come to an effective balance between robustness and efficiency. We illustrate them with examples from the first implementation of the system. 1 Introduction In recent years, there has been a growing tendency to include speech as an in- and output mode for user interfaces. This is due to the expectation that including speech broadens the `bandwidth' of interaction between user and system, and allows for a more `natural' communication. To translate these expected advantages into actual characteristics of the intended system, a number of p... Xavier Pouteau, Emiel Krahmer, Jan Landsbergen |
EUROSPEECH | 2 |