Casey Kennington

dblp:73/8163 · also Casey Redd Kennington · DBLP profile ↗
← Back
51ranked-venue papers
18as first author
22since 2021 · last 2026
0000-0001-6654-8966ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 14 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 17 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization
abstract
Emotion expression is essential for human-robot interaction, yet current systems rely on static models that cannot adapt to individual users. We present an online reinforcement learning framework that adapts robot emotional behavior policy during live dialogue using binary human feedback. The system integrates a DeBERTa-v3-base emotion classifier and applies Group Relative Policy Optimization (GRPO) in a human-robot dialogue system. At each dialogue turn, the classifier samples a group of emotion candidates and the selected emotion is passed to a generative model that synthesizes a novel robot emotional behavior. We evaluate the system in three experiments: (1) offline supervised fine-tuning followed by GRPO on synthetic dialogue data, (2) a live GRPO training with a human teacher and (3) a final experiment with human participants. Results indicate that the robot was perceived as responsive and emotionally consistent, with high ratings for personality coherence and contextual appropriateness of emotional behaviors. Results further show that online GRPO with human feedback enables effective real-time emotion adaptation in embodied interaction.
Anna Manaseryan, Casey Kennington
SIGDIAL2
2026 Reporting Guidelines for Large Language Models in Human-Robot Interaction
abstract
The comparatively recent advent of Large Language Models (LLMs) has resulted in a wide array of new capabilities and components relevant to Human–Robot Interaction (HRI) researchers. LLMs are being applied to vision, manipulation, planning, reasoning, learning, and HRI problems, frequently as “Scarecrows,” in which LLMs serve as black box modules integrated into robot architectures for the purpose of quickly enabling full-pipeline solutions. However, despite this explosion of applications, general questions remain about the best ways to incorporate LLMs into robot architectures, appropriate safety and guardrail considerations, and, critically, how to report properly on HRI research that involves LLMs. In this article, we explore the question of reporting guidelines for HRI researchers who utilize Scarecrows in robot architectures. We identify five key stakeholder groups in the HRI research process, discuss what information each group needs from HRI researchers, and identify appropriate mechanisms for conveying that information from HRI researchers to stakeholders either directly or indirectly. We contribute a set of suggested guidelines regarding what information should be included when researchers disseminate information about HRI research that uses LLMs.
Cynthia Matuszek, Tom Williams 0001, Nick DePalma, Ross Mead, Ruchen Wen, Eike Schneiders, Casey Kennington, Alemitu Mequanint Bezabih
ACM Trans. Hum. Robot Interact.7
2025 Retico: A Framework for Robot/IVA-ready Spoken Dialogue
Casey Kennington, Catherine Henry
HAI1
2025 Recognizing and Generating Novel Emotional Behaviors on Two Robotic Platforms
abstract
Recent advancements in language modeling have enabled robots to more easily generate complex behaviors. However, ensuring that the generated behaviors align with the intended emotional states of the robot is necessary in many domains where robots are used. In this paper, we present an adversarial-like training regime in which a generative model of emotional behavior is enhanced through feedback from both an emotion discriminator and a novelty loss, to ensure that the generated behaviors are non-redundant. Our generative model, fine-tuned on a dataset of robot behaviors labeled with emotions, generates behavior sequences perceived as reflecting the emotional qualities of the input emotion labels. Through our training regime, the generative model is refined by minimizing the discrepancies in both emotion classification and behavioral novelty. We evaluated our approach through multiple experiments and human evaluations, where participants were asked to appraise the emotions conveyed by robot behaviors and rate the novelty of the behaviors. Experimental results demonstrate that our two models, one for classifying and one for generating emotional behaviors, are effective, with the generative model producing emotionally rich behaviors that differ from previously generated outputs.
Rista Baral, Bethany Grenz, Casey Kennington
IROS3
2025 Learning to Speak Like a Child: Reinforcing and Evaluating a Child-level Generative Language Model
abstract
A language model that can generate utterances that are appraised as being within a specific age of a young child who is beginning their language learning journey can be useful in scenarios where child-level language is needed, for example in virtual avatars, interactions with individuals who have disabilities, or developmental robotics. In this paper, we focus on an age range that is not represented in prior work: emergent speakers. We use the CHILDES database to train and tune language models of different parameter sizes using a group relative policy optimization reinforcement learning regime. Our goal is to find the most coherent, yet child-like language model while keeping the number of parameters to as few as possible. We evaluate using metrics of coherency, “toddlerality,” and an evaluation using human subjects who interact with two robot platforms. Our experiments show that even small language models (under 1 billion parameters) can be used effectively to generate child-like utterances.
Enoch Levandovsky, Anna Manaseryan, Casey Kennington
SIGDIAL3
2025 rrSDS 2.0: Incremental, Modular, Distributed, Multimodal Spoken Dialogue with Robotic Platforms
abstract
This demo will showcase updates made to the ‘robot-ready spoken dialogue system’ built on the Retico framework. Updates include new modules, logging and real-time monitoring tools, integrations with the Coppelia Sim virtual robot platfrom, integrations with a benchmark, improved documentation, and pypi environment usage.
Anna Manaseryan, Porter Rigby, Brooke Matthews, Catherine Henry, Josue Torres-Fonseca, Ryan Whetten, Enoch Levandovsky, Casey Kennington
SIGDIAL8
2024 Kid Query: Co-designing an Application to Scaffold Query Formulation
abstract
In this work, we discuss the findings emerging from co-design sessions between children ages 6 to 11 and adults, which were conducted to advance knowledge on how to best support children using well-known search tools for online information discovery. Specifically, we argue that by leveraging scaffolding, gamification techniques, and design choices via an application, it is possible to enhance children’s habits related to query formulation. Outcomes from this preliminary exploration reveal that gameplay incentives (e.g. levels, points, and other incentives like customization) are needed and effective in motivating further interaction with the application, which in turn leads to further utilization of the scaffolding needed to positively impact query formulation.
Benjamin Bettencourt, Maria Soledad Pera, Casey Kennington, Katherine Landau Wright, Jerry Alan Fails
IDC3
2024 How Readability Cues Affect Children's Navigation of Search Engine Result Pages
abstract
Children often interact with search engines within a classroom context to complete assignments or discover new information. To successfully identify relevant resources among those presented on a search engine results page (SERP), users must first be able to comprehend the text included in SERP snippets. While this task may be straightforward for an adult user, children may encounter obstacles in terms of readability and comprehension when attempting to navigate a SERP. Previous research has demonstrated the positive impact of including visual cues on a SERP as relevance signals to guide children toward appropriate resources. In this work, we explore the effect of supplying visual cues related to readability and text difficulty on children’s (ages 6-12) navigation of a SERP. Using quantitative data collected from user-interface interactions and qualitative data gathered from participant interviews, we analyze the impact of these visual cues on children’s selection of results on a SERP when carrying out information discovery tasks.
Christine Pinney, Benjamin Bettencourt, Jerry Alan Fails, Casey Kennington, Katherine Landau Wright, Maria Soledad Pera
IDC4
2024 Conceptual Pacts for Reference Resolution Using Small, Dynamically Constructed Language Models: A Study in Puzzle Building Dialogues
abstract
Using Brennan and Clark’s theory of a Conceptual Pact, that when interlocutors agree on a name for an object, they are forming a temporary agreement on how to conceptualize that object, we present an extension to a simple reference resolver which simulates this process over time with different conversation pairs. In a puzzle construction domain, we model pacts with small language models for each referent which update during the interaction. When features from these pact models are incorporated into a simple bag-of-words reference resolver, the accuracy increases compared to using a standard pre-trained model. The model performs equally to a competitor using the same data but with exhaustive re-training after each prediction, while also being more transparent, faster and less resource-intensive. We also experiment with reducing the number of training interactions, and can still achieve reference resolution accuracies of over 80% in testing from observing a single previous interaction, over 20% higher than a pre-trained baseline. While this is a limited domain, we argue the model could be applicable to larger real-world applications in human and human-robot interaction and is an interpretable and transparent model.
Julian Hough, Sina Zarrieß, Casey Kennington, David Schlangen, Massimo Poesio
LREC/COLING3
2024 Incorporating Word-level Phonemic Decoding into Readability Assessment
abstract
Current approaches in automatic readability assessment have found success with the use of large language models and transformer architectures. These techniques lead to accuracy improvement, but they do not offer the interpretability that is uniquely required by the audience most often employing readability assessment tools: teachers and educators. Recent work that employs more traditional machine learning methods has highlighted the linguistic importance of considering semantic and syntactic characteristics of text in readability assessment by utilizing handcrafted feature sets. Research in Education suggests that, in addition to semantics and syntax, phonetic and orthographic instruction are necessary for children to progress through the stages of reading and spelling development; children must first learn to decode the letters and symbols on a page to recognize words and phonemes and their connection to speech sounds. Here, we incorporate this word-level phonemic decoding process into readability assessment by crafting a phonetically-based feature set for grade-level classification for English. Our resulting feature set shows comparable performance to much larger, semantically- and syntactically-based feature sets, supporting the linguistic value of orthographic and phonetic considerations in readability assessment.
Christine Pinney, Casey Kennington, Maria Soledad Pera, Katherine Landau Wright, Jerry Alan Fails
LREC/COLING2
2024 A Systematic Evaluation of Code-generating Chatbots for Use in Undergraduate Computer Science Education
abstract
This research paper focuses on evaluating code-generating chatbots. Chatbots like ChatGPT released in the past three years have proven capable of a wide variety of tasks within a conversational interaction, including writing code and answering code-related questions. With these recent advances, chatbots have many potential uses in education, including computer science education. However, before these chatbots are used in CS curricula, their capabilities and limitations must be systematically tested and understood. In this work, we evaluate the capabilities and limitations of four known, open-source, code-based chatbots in programming tasks by performing a standardized study in which different chatbots are tasked with providing answers for a variety of assignments from Boise State University's computer science program. We found that while all of the chatbots can write code and provide explanations, some do better than others, and each of them work differently in conversations. Moreover, all of them suffered similar and important limitations, which has implications for adoption in curriculum. As a second experiment, we used the Llama chatbot to perform a human evaluation by enabling student novice and experienced programmers to use it as a coding assistant to complete specific tasks in a common software development environment. We found that the coding assistant can help novice programmers accomplish simple tasks in comparable time and code efficacy as more experienced programmers. Given these experiments, and given feedback from participants in our studies, we see a clear picture emerge: new programmers should learn important concepts about programming without the help of code assistants so students can (1) demonstrate their understanding of important concepts and (2) have enough experience to assess code assistant output as useful or erroneous. Then, once intermediate skills are mastered (e.g., object oriented programming and data structures), it seems appropriate to introduce students systematically to coding assistants to help with specific assignments throughout the undergraduate computer science curriculum. We conclude by addressing ethical considerations for the use of code-based chatbots in computer science education and future directions of research.
Adam J. Torek, Elijah Sorensen, Natalie Hahle, Casey Kennington
FIE4
2023 Exploring Transformers as Compact, Data-efficient Language Models
abstract
Large scale transformer models, trained with massive datasets have become the standard in natural language processing.The huge size of most transformers make research with these models impossible for those with limited computational resources.Additionally, the enormous pretraining data requirements of transformers exclude pretraining them with many smaller datasets that might provide enlightening results.In this study, we show that transformers can be significantly reduced in size, with as few as 5.7 million parameters, and still retain most of their downstream capability.Further we show that transformer models can retain comparable results when trained on human-scale datasets, as few as 5 million words of pretraining data.Overall, the results of our study suggest transformers function well as compact, data efficient language models and that complex model compression methods, such as model distillation are not necessarily superior to pretraining reduced size transformer models from scratch.
Clayton Fields, Casey Kennington
CoNLL2
2022 Searching for Engagement: Child Engagement and Search Engine Result Pages
abstract
In this paper, we explore how children engage with search engine result pages (SERP) generated by a popular search API in response to their online inquiries. We do so to further understand children navigation behaviour. To accomplish this goal, we examine search logs produced as a result of children (ages 6 to 12), using metrics commonly used to operationalize engagement, including: position of clicks, time spent hovering, and the sequence of navigation on a SERP. We also investigate the potential connection between the text complexity of SERP snippets and engagement. Our findings verify that children engage more frequently with SERP results in higher ranking positions, but that engagement does not decrease linearly as children navigate to lower ranking positions. They also reveal that children generally spend more time hovering on snippets with more complex readability levels (i.e., harder to read) than snippets on the lower end of the readability spectrum.
Benjamin Bettencourt, Arif Ahmed 0004, Nic Way, Casey Kennington, Katherine Landau Wright, Jerry Alan Fails
IDC4
2022 Supercalifragilisticexpialidocious: Why Using the "Right" Readability Formula in Children's Web Search Matters
Garrett Allen, Ashlee Milton, Katherine Landau Wright, Jerry Alan Fails, Casey Kennington, Maria Soledad Pera
ECIR (1)5
2022 HADREB: Human Appraisals and (English) Descriptions of Robot Emotional Behaviors
abstract
Humans sometimes anthropomorphize everyday objects, but especially robots that have human-like qualities and that are often able to interact with and respond to humans in ways that other objects cannot. Humans especially attribute emotion to robot behaviors, partly because humans often use and interpret emotions when interacting with other humans, and they apply that capability when interacting with robots. Moreover, emotions are a fundamental part of the human language system and emotions are used as scaffolding for language learning, making them an integral part of language learning and meaning. However, there are very few datasets that explore how humans perceive the emotional states of robots and how emotional behaviors relate to human language. To address this gap we have collected HADREB, a dataset of human appraisals and English descriptions of robot emotional behaviors collected from over 30 participants. These descriptions and human emotion appraisals are collected using the Mistyrobotics Misty II and the Digital Dream Labs Cozmo (formerly Anki) robots. The dataset contains English descriptions and emotion appraisals of more than 500 descriptions and graded valence labels of 8 emotion pairs for each behavior and each robot. In this paper we describe the process of collecting and cleaning the data, give a general analysis of the data, and evaluate the usefulness of the dataset in two experiments, one using a language model to map descriptions to emotions, the other maps robot behaviors to emotions.
Josue Torres-Fonseca, Casey Kennington
LREC2
2022 Understanding Intention for Machine Theory of Mind: a Position Paper
abstract
Theory of Mind is often characterized as the ability to recognize desires, beliefs, and intentions of others. In this position paper, I look at the literature on modeling Theory of Mind in machines and find that, to date, intention is not usually a focus. I define what I mean by intention— choice with commitment—following prior work. Intention has a long history of research in some communities, and I offer one theoretical framework for modeling intention as a starting point. I take inspiration from how children learn intention through joint attention with others and how that leads to Theory of Mind. I argue that though models of machine Theory of Mind need not follow the same learning progression as children, intention is an aspect of Theory of Mind that should be more explicit.
Casey Kennington
RO-MAN1
2022 Symbol and Communicative Grounding through Object Permanence with a Mobile Robot
abstract
Object permanence is the ability to form and recall mental representations of objects even when they are not in view.Despite being a crucial developmental step for children, object permanence has had only some exploration as it relates to symbol and communicative grounding in spoken dialogue systems.In this paper, we leverage SLAM as a module for tracking object permanence and use a robot platform to move around a scene where it discovers objects and learns how they are denoted.We evaluated by comparing our system's effectiveness at learning words from human dialogue partners both with and without object permanence.We found that with object permanence, human dialogue partners spoke with the robot and the robot correctly identified objects it had learned about significantly more than without object permanence, which suggests that object permanence helped facilitate communicative and symbol grounding.
Josue Torres-Fonseca, Catherine Henry, Casey Kennington
SIGDIAL3
2022 Spoken language interaction with robots: Recommendations for future research
abstract
With robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with.
Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005
Comput. Speech Lang.13
2021 Engage!: Co-designing Search Engine Result Pages to Foster Interactions
abstract
In this paper, we take a step towards understanding how to design search engine results pages (SERP) that encourage children’s engagement as they seek for online resources. For this, we conducted a participatory design session to enable us to elicit children’s preferences and determine what children (ages 6–12) find lacking in more traditional SERP. We learned that children want more dynamic means of navigating results and additional ways to interact with results via icons. We use these findings to inform the design of a new SERP interface, which we denoted CHIRP. To gauge the type of engagement that a SERP incorporating interactive elements–CHIRP–can foster among children, we conducted a user study at a public school. Analysis of children’s interactions with CHIRP, in addition to responses to a post-task survey, reveals that adding additional interaction points results in a SERP interface that children prefer, but one that does not necessarily change engagement levels through clicks or time spent on SERP.
Garrett Allen, Benjamin L. Peterson, Dhanush kumar Ratakonda, Mostofa Najmus Sakib, Jerry Alan Fails, Casey Kennington, Katherine Landau Wright, Maria Soledad Pera
IDC6
2021 Enriching Language Models with Visually-grounded Word Vectors and the Lancaster Sensorimotor Norms
abstract
Language models are trained only on text despite the fact that humans learn their first language in a highly interactive and multimodal environment where the first set of learned words are largely concrete, denoting physical entities and embodied states.To enrich language models with some of this missing experience, we leverage two sources of information: (1) the Lancaster Sensorimotor norms, which provide ratings (means and standard deviations) for over 40,000 English words along several dimensions of embodiment, and which capture the extent to which something is experienced across 11 different sensory modalities, and (2) vectors from coefficients of binary classifiers trained on images for the BERT vocabulary.We pre-trained the ELECTRA model and fine-tuned the RoBERTa model with these two sources of information then evaluate using the established GLUE benchmark and the Visual Dialog benchmark.We find that enriching language models with the Lancaster norms and image vectors improves results in both tasks, with some implications for robust language models that capture holistic linguistic meaning in a language learning context.
Casey Kennington
CoNLL1
2021 BiGBERT: Classifying Educational Web Resources for Kindergarten-12th Grades
Garrett Allen, Brody Downs, Aprajita Shukla, Casey Kennington, Jerry Alan Fails, Katherine Landau Wright, Maria Soledad Pera
ECIR (2)4
2021 In-Game Social Interactions to Facilitate ESL Students' Morphological Awareness, Language and Literacy Skills
abstract
Video games that require players to utilize a target or second language to complete tasks have emerged as alternative pedagogical tools for Second Language Acquisition (SLA). With the exception of vocabulary acquisition, much of the prior research in game-based SLA fails to gauge students' literacy skills, specifically their morphological awareness or understanding of the smallest meaningful linguistic units (e.g., prefixes, suffixes, and roots). Given this shortcoming, we utilize a two-player online game to facilitate social interactions between Native English Speakers (NES) and English as a Second Language (ESL) students as a mechanism to generate ESL students' written output in the targeted language and draw attention to their morphological awareness. Analysis of chat logs demonstrates the game's potential to enhance ESL students' morphological awareness and other important L2 literacy skills such as word reading accuracy. Both NES and ESL students' reflections of their gameplay experiences suggest game design modifications that promote ESL students' willingness to communicate with NES while developing their morphological awareness and practicing their L2 communication and literacy skills.
Yolanda A. Rankin, Sana Tibi, Casey Kennington, Na-eun Han
Proc. ACM Hum. Comput. Interact.3
2020 Guiding the selection of child spellchecker suggestions using audio and visual cues
abstract
Spellchecking functionality embedded in existing search tools can assist children by offering a list of spelling alternatives when a spelling error is detected. Unfortunately, children tend to generally select the first alternative when presented with a list of options, as opposed to the one that matches their intent. In this paper, we describe a study we conducted with 191 children ages 6-12 in order to offer empirical evidence of: (1) their selection habits when identifying spelling suggestions that match the word they meant to type, and (2) the degree of influence multimodal cues, i.e., synthesized speech and images, have in prompting children to select the correct spelling suggestion. The results from our study reveal that multimodal cues, primarily synthesized speech, have a positive impact on the children's ability to identify their intended word from a list of spelling suggestions.
Brody Downs, Aprajita Shukla, Mikey Krentz, Maria Soledad Pera, Katherine Landau Wright, Casey Kennington, Jerry Alan Fails
IDC6
2020 Evaluating and Improving Child-Directed Automatic Speech Recognition
abstract
Speech recognition has seen dramatic improvements in the last decade, though those improvements have focused primarily on adult speech. In this paper, we assess child-directed speech recognition and leverage a transfer learning approach to improve child-directed speech recognition by training the recent DeepSpeech2 model on adult data, then apply additional tuning to varied amounts of child speech data. We evaluate our model using the CMU Kids dataset as well as our own recordings of child-directed prompts. The results from our experiment show that even a small amount of child audio data improves significantly over a baseline of adult-only or child-only trained models. We report a final general Word-Error-Rate of 29% over a baseline of 62% that uses the adult-trained model. Our analyses show that our model adapts quickly using a small amount of data and that the general child model works better than school grade-specific models. We make available our trained model and our data collection tool.
Eric G. Booth, Jake Carns, Casey Kennington, Nader Rafla 0001
LREC3
2020 KidSpell: A Child-Oriented, Rule-Based, Phonetic Spellchecker
abstract
For help with their spelling errors, children often turn to spellcheckers integrated in software applications like word processors and search engines. However, existing spellcheckers are usually tuned to the needs of traditional users (i.e., adults) and generally prove unsatisfactory for children. Motivated by this issue, we introduce KidSpell, an English spellchecker oriented to the spelling needs of children. KidSpell applies (i) an encoding strategy for mapping both misspelled words and spelling suggestions to their phonetic keys and (ii) a selection process that prioritizes candidate spelling suggestions that closely align with the misspelled word based on their respective keys. To assess the effectiveness of, we compare the model’s performance against several popular, mainstream spellcheckers in a number of offline experiments using existing and novel datasets. The results of these experiments show that KidSpell outperforms existing spellcheckers, as it accurately prioritizes relevant spelling corrections when handling misspellings generated by children in both essay writing and online search tasks. As a byproduct of our study, we create two new datasets comprised of spelling errors generated by children from hand-written essays and web search inquiries, which we make available to the research community.
Brody Downs, Oghenemaro Anuyah, Aprajita Shukla, Jerry Alan Fails, Maria Soledad Pera, Katherine Landau Wright, Casey Kennington
LREC7
2020 rrSDS: Towards a Robot-ready Spoken Dialogue System
abstract
Spoken interaction with a physical robot requires a dialogue system that is modular, multimodal, distributive, incremental and temporally aligned.In this demo paper, we make significant contributions towards fulfilling these requirements by expanding upon the ReTiCo incremental framework.We outline the incremental and multimodal modules and how their computation can be distributed.We demonstrate the power and flexibility of our robotready spoken dialogue system to be integrated with almost any robot.
Casey Kennington, Daniele Moro, Lucas Marchand, Jake Carns, David McNeill
SIGdial1
2020 Learning Word Groundings from Humans Facilitated by Robot Emotional Displays
abstract
In working towards accomplishing a humanlevel acquisition and understanding of language, a robot must meet two requirements: the ability to learn words from interactions with its physical environment, and the ability to learn language from people in settings for language use, such as spoken dialogue.In a live interactive study, we test the hypothesis that emotional displays are a viable solution to the cold-start problem of how to communicate without relying on language the robot does not-indeed, cannot-yet know.We explain our modular system that can autonomously learn word groundings through interaction and show through a user study with 21 participants that emotional displays improve the quantity and quality of the inputs provided to the robot.
David McNeill, Casey Kennington
SIGdial2
2019 Searching for spellcheckers: What kids want, what kids need
abstract
Misspellings in queries used to initiate online searches is an everyday occurrence. When this happens, users either rely on the search engine's ability to understand their query or they turn to spellcheckers. Spellcheckers are usually based on popular dictionaries or past query logs, leading to spelling suggestions that often better resonate with adult users because that data is more readily available. Based on an educational perspective, previous research reports, and initial analyses of sample search logs, we hypothesize that existing spellcheckers are not suitable for young users who frequently encounter spelling challenges when searching for information online. We present early results of our ongoing research focused on identifying the needs and expectations children have regarding spellcheckers.
Brody Downs, Tyler French, Maria Soledad Pera, Katherine Landau Wright, Casey Kennington, Jerry Alan Fails
IDC5
2019 Query Formulation Assistance for Kids: What is Available, When to Help & What Kids Want
abstract
Children use popular web search tools, which are generally designed for adult users. Because children have different developmental needs than adults, these tools may not always adequately support their search for information. Moreover, even though search tools offer support to help in query formulation, these too are aimed at adults and may hinder children rather than help them. This calls for the examination of existing technologies in this area, to better understand what remains to be done when it comes to facilitating query-formulation tasks for young users. In this paper, we investigate interaction elements of query formulation--including query suggestion algorithms--for children. The primary goals of our research efforts are to: (i) examine existing plug-ins and interfaces that explicitly aid children's query formulation; (ii) investigate children's interactions with suggestions offered by a general-purpose query suggestion strategy vs. a counterpart designed with children in mind; and (iii) identify, via participatory design sessions, their preferences when it comes to tools / strategies that can help children find information and guide them through the query formulation process. Our analysis shows that existing tools do not meet children's needs and expectations; the outcomes of our work can guide researchers and developers as they implement query formulation strategies for children.
Jerry Alan Fails, Maria Soledad Pera, Oghenemaro Anuyah, Casey Kennington, Katherine Landau Wright, William Bigirimana
IDC4
2018 Placing Objects in Gesture Space: Toward Incremental Interpretation of Multimodal Spatial Descriptions
abstract
When describing routes not in the current environment, a common strategy is to anchor the description in configurations of salient landmarks, complementing the verbal descriptions by "placing" the non-visible landmarks in the gesture space. Understanding such multimodal descriptions and later locating the landmarks from real world is a challenging task for the hearer, who must interpret speech and gestures in parallel, fuse information from both modalities, build a mental representation of the description, and ground the knowledge to real world landmarks. In this paper, we model the hearer's task, using a multimodal spatial description corpus we collected. To reduce the variability of verbal descriptions, we simplified the setup to use simple objects as landmarks. We describe a real-time system to evaluate the separate and joint contribution of the modalities. We show that gestures not only help to improve the overall system performance, even if to a large extent they encode redundant information, but also result in earlier final correct interpretations. Being able to build and apply representations incrementally will be of use in more dialogical settings, we argue, where it can enable immediate clarification in cases of mismatch.
Ting Han 0003, Casey Kennington, David Schlangen
AAAI2
2018 Predicting Perceived Age: Both Language Ability and Appearance are Important
abstract
When interacting with robots in a situated spoken dialogue setting, human dialogue partners tend to assign anthropomorphic and social characteristics to those robots.In this paper, we explore the age and educational level that human dialogue partners assign to three different robotic systems, including an un-embodied spoken dialogue system.We found that how a robot speaks is as important to human perceptions as the way the robot looks.Using the data from our experiment, we derived prosodic, emotional, and linguistic features from the participants to train and evaluate a classifier that predicts perceived intelligence, age, and education level.
Sarah Plane, Ariel Marvasti, Tyler Egan, Casey Kennington
SIGDIAL Conference4
2017 A Graphical Digital Personal Assistant that Grounds and Learns Autonomously
abstract
We present a speech-driven digital personal assistant that is robust despite little or no training data and autonomously improves as it interacts with users. The system is able to establish and build common ground between itself and users by signaling understanding and by learning a mapping via interaction between the words that users actually speak and the system actions. We evaluated our system with real users and found an overall positive response. We further show through objective measures that autonomous learning improves performance in a simple itinerary filling task.
Casey Kennington, Aprajita Shukla
HAI1
2017 Temporal alignment using the incremental unit framework
abstract
We propose a method for temporal alignment--a precondition of meaningful fusion--of multimodal systems, using the incremental unit dialogue system framework, which gives the system flexibility in how it handles alignment: either by delaying a modality for a specified amount of time, or by revoking (i.e., backtracking) processed information so multiple information sources can be processed jointly. We evaluate our approach in an offline experiment with multimodal data and find that using the incremental framework is flexible and shows promise as a solution to the problem of temporal alignment in multimodal systems.
Casey Kennington, Ting Han 0003, David Schlangen
ICMI1
2017 A simple generative model of incremental reference resolution for situated dialogue
Casey Kennington, David Schlangen
Comput. Speech Lang.1
2016 Resolving References to Objects in Photographs using the Words-As-Classifiers Model
abstract
A common use of language is to refer to visually present objects.Modelling it in computers requires modelling the link between language and perception.The "words as classifiers" model of grounded semantics views words as classifiers of perceptual contexts, and composes the meaning of a phrase through composition of the denotations of its component words.It was recently shown to perform well in a game-playing scenario with a small number of object types.We apply it to two large sets of real-world photographs that contain a much larger variety of object types and for which referring expressions are available.Using a pre-trained convolutional neural network to extract image region features, and augmenting these with positional information, we show that the model achieves performance competitive with the state of the art in a reference resolution task (given expression, find bounding box of its referent), while, as we argue, being conceptually simpler and more flexible.
David Schlangen, Sina Zarrieß, Casey Kennington
ACL (1)3
2016 PentoRef: A Corpus of Spoken References in Task-oriented Dialogues
Sina Zarrieß, Julian Hough, Casey Kennington, Ramesh R. Manuvinakurike, David DeVault, Raquel Fernández, David Schlangen
LREC3
2016 Supporting Spoken Assistant Systems with a Graphical User Interface that Signals Incremental Understanding and Prediction State
abstract
Arguably, spoken dialogue systems are most often used not in hands/eyes-busy situations, but rather in settings where a graphical display is also available, such as a mobile phone.We explore the use of a graphical output modality for signalling incremental understanding and prediction state of the dialogue system.By visualising the current dialogue state and possible continuations of it as a simple tree, and allowing interaction with that visualisation (e.g., for confirmations or corrections), the system provides both feedback on past user actions and guidance on possible future ones, and it can span the continuum from slot filling to full prediction of user intent (such as GoogleNow).We evaluate our system with real users and report that they found the system intuitive and easy to use, and that incremental and adaptive settings enable users to accomplish more tasks.
Casey Kennington, David Schlangen
SIGDIAL Conference1
2016 Real-Time Understanding of Complex Discriminative Scene Descriptions
abstract
Real-world scenes typically have complex structure, and utterances about them consequently do as well.We devise and evaluate a model that processes descriptions of complex configurations of geometric shapes and can identify the described scenes among a set of candidates, including similar distractors.The model works with raw images of scenes, and by design can work word-by-word incrementally.Hence, it can be used in highly-responsive interactive and situated settings.Using a corpus of descriptions from game-play between human subjects (who found this to be a challenging task), we show that reconstruction of description structure in our system contributes to task success and supports the performance of the word-based model of grounded semantics that we use.
Ramesh R. Manuvinakurike, Casey Kennington, David DeVault, David Schlangen
SIGDIAL Conference2
2015 Simple Learning and Compositional Application of Perceptually Grounded Word Meanings for Incremental Reference Resolution
abstract
Casey Kennington, David Schlangen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Casey Kennington, David Schlangen
ACL (1)1
2015 Incrementally Tracking Reference in Human/Human Dialogue Using Linguistic and Extra-Linguistic Information
abstract
Casey Kennington, Ryu Iida, Takenobu Tokunaga, David Schlangen. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Casey Kennington, Ryu Iida, Takenobu Tokunaga, David Schlangen
HLT-NAACL1
2014 Better Driving and Recall When In-car Information Presentation Uses Situationally-Aware Incremental Speech Output Generation
abstract
It is established that driver distraction is the result of sharing cognitive resources between the primary task (driving) and any other secondary task. In the case of holding conversations, a human passenger who is aware of the driving conditions can choose to interrupt his speech in situations potentially requiring more attention from the driver, but in-car information systems typically do not exhibit such sensitivity. We have designed and tested such a system in a driving simulation environment. Unlike other systems, our system delivers information via speech (calendar entries with scheduled meetings) but is able to react to signals from the environment to interrupt when the driver needs to be fully attentive to the driving task and subsequently resume its delivery. Distraction is measured by a secondary short-term memory task. In both tasks, drivers perform significantly worse when the system does not adapt its speech, while they perform equally well to control conditions (no concurrent task) when the system intelligently interrupts and resumes.
Casey Kennington, Spyros Kousidis, Timo Baumann, Hendrik Buschmeier, Stefan Kopp, David Schlangen
AutomotiveUI1
2014 Situated Incremental Natural Language Understanding using a Multimodal, Linguistically-driven Update Model
Casey Kennington, Spyros Kousidis, David Schlangen
COLING1
2014 Probabilistic multiparty dialogue management for a game master robot
abstract
We present our ongoing research on multiparty dialogue management for a game master robot which engages multiple human participants to play a quiz game. The robot invites passing people to join the game, instructs participants on the rules of the game, and leads them in the game. The robot has to manage people leaving and coming at arbitrary times. Our approach maintains a dialogue manager for each participant, and a module takes a final action with each decision cycle; responsible to decide "what/whom/when to say". We have implemented the dialogue manager with a probabilistic rules approach [4] and made preliminary evaluations with our multiparty human-robot game dialogue data that was collected in a WoZ fashion.
Casey Kennington, Kotaro Funakoshi, Yuki Takahashi, Mikio Nakano
HRI1
2014 A Multimodal In-Car Dialogue System That Tracks The Driver's Attention
abstract
When a passenger speaks to a driver, he or she is co-located with the driver, is generally aware of the situation, and can stop speaking to allow the driver to focus on the driving task. In-car dialogue systems ignore these important aspects, making them more distracting than even cell-phone conversations. We developed and tested a "situationally-aware" dialogue system that can interrupt its speech when a situation which requires more attention from the driver is detected, and can resume when driving conditions return to normal. Furthermore, our system allows driver-controlled resumption of interrupted speech via verbal or visual cues (head nods). Over two experiments, we found that the situationally-aware spoken dialogue system improves driving performance and attention to the speech content, while driver-controlled speech resumption does not hinder performance in either of these two tasks
Spyros Kousidis, Casey Kennington, Timo Baumann, Hendrik Buschmeier, Stefan Kopp, David Schlangen
ICMI2
2014 InproTKs: A Toolkit for Incremental Situated Processing
abstract
In order to process incremental situated dialogue, it is necessary to accept information from various sensors, each tracking, in real-time, different aspects of the physical situation.We present extensions of the incremental processing toolkit IN-PROTK which make it possible to plug in such multimodal sensors and to achieve situated, real-time dialogue.We also describe a new module which enables the use in INPROTK of the Google Web Speech API, which offers speech recognition with a very large vocabulary and a wide choice of languages.We illustrate the use of these extensions with a description of two systems handling different situated settings.
Casey Kennington, Spyros Kousidis, David Schlangen
SIGDIAL Conference1
2014 Situated incremental natural language understanding using Markov Logic Networks
Casey Kennington, David Schlangen
Comput. Speech Lang.1
2013 Interpreting Situated Dialogue Utterances: an Update Model that Uses Speech, Gaze, and Gesture Information
Casey Kennington, Spyros Kousidis, David Schlangen
SIGDIAL Conference1
2013 Investigating speaker gaze and pointing behaviour in human-computer interaction with the mint.tools collection
Spyros Kousidis, Casey Kennington, David Schlangen
SIGDIAL Conference2
2012 Suffix Trees as Language Models
Casey Kennington, Martin Kay, Annemarie Friedrich
LREC1
2012 Markov Logic Networks for Situated Incremental Natural Language Understanding
Casey Kennington, David Schlangen
SIGDIAL Conference1
2008 Elicited Imitation as an Oral Proficiency Measure with ASR Scoring
C. Ray Graham, Deryle W. Lonsdale, Casey Kennington, Jeremiah McGhee
LREC3