Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Justine Cassell

dblp:61/1634 · DBLP profile ↗
← Back
60ranked-venue papers
17as first author
6since 2021 · last 2024
0000-0003-2770-7359ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 35 · 10 first-author · 1 since 2021Artificial intelligence and machine learning · 30 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 10 · 1 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Language models and text generation · 40% Question answering and dialogue systems · 28% Information extraction and text analysis · 17%
Human-computer interaction and pervasive computing
8 papers
Human-AI interaction · 43% Learning and educational technologies · 30% Human-robot interaction · 23%
Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 91% Computer animation and physical simulation · 7% Image and video coding · 2%

Topics — the 22 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › dialogue modeling
conversational grounding
0.822024
Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding · EMNLP 2024
Towards a Model of Face-to-Face Grounding · ACL 2003
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding · EMNLP 2024
Multimedia analysis and retrieval
affective computing
0.312017
Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia Content · ACM Multimedia 2017
Multimedia analysis and retrieval › affective computing
sentiment analysis
0.312017
Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia Content · ACM Multimedia 2017
Human-robot interaction › affective interaction
rapport
0.312017
Cognitive-Inspired Conversational-Strategy Reasoner for Socially-Aware Agents · IJCAI 2017
Natural language and speech › Question answering and dialogue systems
conversational agents
0.222023
How About Kind of Generating Hedges using End-to-End Neural Models? · ACL (1) 2023
Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents · SIGGRAPH 1994
Learning and educational technologies › pedagogical agents
teachable agent
0.112012
"Oh dear stacy!": social interaction, elaboration, and learning with teachable agents · CHI 2012
Learning and educational technologies
language learning
0.112011
Brick by brick: iterating interventions to bridge the achievement gap with virtual peers · CHI 2011
Learning and educational technologies › pedagogical agents
virtual peer
0.112011
Brick by brick: iterating interventions to bridge the achievement gap with virtual peers · CHI 2011
Machine learning › Deep learning architectures and training
attention mechanism
0.112017
Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia Content · ACM Multimedia 2017
Machine learning › Deep learning architectures and training › attention mechanism
temporal attention
0.112017
Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia Content · ACM Multimedia 2017
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
embodied conversational agents
0.022001
Non-Verbal Cues for Discourse Structure · ACL 2001
Relational agents: a model and implementation of building user trust · CHI 2001
Learning and educational technologies › STEM education
science education
0.012011
Brick by brick: iterating interventions to bridge the achievement gap with virtual peers · CHI 2011
Human-AI interaction › conversational agents
embodied conversational agents
0.022003
Embodiment in Conversational Interfaces: Rea · CHI 1999
Towards a Model of Face-to-Face Grounding · ACL 2003
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse structure
0.012001
Non-Verbal Cues for Discourse Structure · ACL 2001
Human-robot interaction › social robot
social dialogue
0.012001
Relational agents: a model and implementation of building user trust · CHI 2001
Usability and user experience research › user perception
user trust
0.012001
Relational agents: a model and implementation of building user trust · CHI 2001
Human-robot interaction › robot communication
multimodal dialogue
0.011999
Embodiment in Conversational Interfaces: Rea · CHI 1999
Computer vision › Video understanding and tracking
gesture recognition
0.011997
Temporal Classification of Natural Gesture and Application to Video Coding · CVPR 1997
Interaction techniques and input › input modality
multimodal input
0.011999
Embodiment in Conversational Interfaces: Rea · CHI 1999
Image and video coding › coding for machines
semantic video compression
0.011997
Temporal Classification of Natural Gesture and Application to Video Coding · CVPR 1997
Image and video coding
video compression
0.011997
Temporal Classification of Natural Gesture and Application to Video Coding · CVPR 1997

Methods — techniques the papers use, named apart from their topics

pre-trained language model · 1.1model explainability · 1.1human evaluation · 0.8reranking · 0.7hedge classifier · 0.7fine-tuning · 0.7cognitive modeling · 0.6attention model · 0.6LSTM · 0.6speaker-distribution loss · 0.3decision-making · 0.3decision making · 0.3think-aloud study · 0.1corpus analysis · 0.1rule-based behavior assignment · 0.0model implementation · 0.0linguistic analysis · 0.0experiment · 0.0
YearPublicationVenuePosition
2024 Conversational Grounding: Annotation and Analysis of Grounding Acts and Grounding Units
abstract
Successful conversations often rest on common understanding, where all parties are on the same page about the information being shared. This process, known as conversational grounding, is crucial for building trustworthy dialog systems that can accurately keep track of and recall the shared information. The proficiencies of an agent in grounding the conveyed information significantly contribute to building a reliable dialog system. Despite recent advancements in dialog systems, there exists a noticeable deficit in their grounding capabilities. Traum (Traum, 1995) provided a framework for conversational grounding introducing Grounding Acts and Grounding Units, but substantial progress, especially in the realm of Large Language Models, remains lacking. To bridge this gap, we present the annotation of two dialog corpora employing Grounding Acts, Grounding Units, and a measure of their degree of grounding. We discuss our key findings during the annotation and also provide a baseline model to test the performance of current Language Models in categorizing the grounding acts of the dialogs. Our work aims to provide a useful resource for further research in making conversations with machines better understood and more reliable in natural day-to-day collaborative dialogs.
Biswesh Mohapatra, Seemab Hassan, Laurent Romary, Justine Cassell
LREC/COLING4
2024 Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding
abstract
Conversational grounding, vital for building effective dialogue between people and between people and dialogue systems, involves ensuring a mutual understanding of shared information.Despite its importance, there has been limited research on this aspect of conversation in recent years, especially after the advent of Large Language Models (LLMs).Previous studies have highlighted the shortcomings of some pre-trained language models in conversational grounding.However, most testing for conversational grounding capabilities involves human evaluations that are costly and time-consuming.This has led to a lack of testing across multiple models of varying sizes, a critical need given the rapid rate of new model releases.This gap in research becomes more significant considering recent advances in language models, which have led to new emergent capabilities.In this paper, we evaluate the performance of LLMs in various aspects of conversational grounding and analyze why some models perform better than others.We demonstrate a direct correlation between the size of the pre-training dataset, size of the model and conversational grounding abilities, suggesting that they have independently acquired some pragmatic capabilities from larger pre-training datasets.Finally, we propose ways to enhance the capabilities of the models that lag in our tests.
Biswesh Mohapatra, Manav Nitin Kapadnis, Laurent Romary, Justine Cassell
EMNLP4
2023 How About Kind of Generating Hedges using End-to-End Neural Models?
abstract
Hedging is a strategy for softening the impact of a statement in conversation.In reducing the strength of an expression, it may help to avoid embarrassment (more technically, "face threat") to one's listener.For this reason, it is often found in contexts of instruction, such as tutoring.In this work, we develop a model of hedge generation based on i) fine-tuning stateof-the-art language models trained on humanhuman tutoring data, followed by ii) reranking to select the candidate that best matches the expected hedging strategy within a candidate pool using a hedge classifier.We apply this method to a natural peer-tutoring corpus containing a significant number of disfluencies, repetitions, and repairs.The results show that generation in this noisy environment is feasible with reranking.By conducting an error analysis for both approaches, we reveal the challenges faced by systems attempting to accomplish both social and task-oriented goals in conversation.
Alafate Abulimiti, Chloé Clavel, Justine Cassell
ACL (1)3
2023 When to generate hedges in peer-tutoring interactions
abstract
This paper explores the application of machine learning techniques to predict where hedging occurs in peer-tutoring interactions.The study uses a naturalistic face-to-face dataset annotated for natural language turns, conversational strategies, tutoring strategies, and nonverbal behaviors.These elements are processed into a vector representation of the previous turns, which serves as input to several machine learning models, including MLP and LSTM.The results show that embedding layers, capturing the semantic information of the previous turns, significantly improves the model's performance.Additionally, the study provides insights into the importance of various features, such as interpersonal rapport and nonverbal behaviors, in predicting hedges by using Shapley values (Hart, 1989) for feature explanation.We discover that the eye gaze of both the tutor and the tutee has a significant impact on hedge prediction.We further validate this observation through a follow-up ablation study.
Alafate Abulimiti, Chloé Clavel, Justine Cassell
SIGDIAL3
2022 "You might think about slightly revising the title": Identifying Hedges in Peer-tutoring Interactions
abstract
Hedges play an important role in the management of conversational interaction.In peertutoring, they are notably used by tutors in dyads (pairs of interlocutors) experiencing low rapport to tone down the impact of instructions and negative feedback.Pursuing the objective of building a tutoring agent that manages rapport with students in order to improve learning, we used a multimodal peer-tutoring dataset to construct a computational framework for identifying hedges.We compared approaches relying on pre-trained resources with others that integrate insights from the social science literature.Our best performance involved a hybrid approach that outperforms the existing baseline while being easier to interpret.We employ a model explainability tool to explore the features that characterize hedges in peer-tutoring conversations, and we identify some novel features, and the benefits of such a hybrid model approach.Models RB MLP (KDF) MLP (PTE) MLP (K+P) CNN (PTE) LSTM (KDF) LSTM(PTE) LSTM (K+P) BERT (PTE) LGB (KDF) LGB (PTE) LGB (K+P) Rule-based No Yes No No No Yes No No Yes Yes
Yann Raphalen, Chloé Clavel, Justine Cassell
ACL (1)3
2022 The Future of the Body in Tomorrow's Workplace
abstract
Even in the most hectic or time-conscious workplace, employees gather in person to chat. And even in the most networked workplace, employees still make time for face-to-face collaboration. This isn’t surprising if we consider that the ability and desire to engage in face-to-face communication (using eye gaze to manage turn-taking, head nods to indicate listening, and smiles to indicate attention) starts soon after birth – well before infants even learn to talk. We might say that we are built to communicate face-to-face! But what role will embodied interaction play in the future workplace, when we will be interacting with autonomous robots, engaging with other people through presence robots, and working in a virtual world where we and our colleagues are represented by avatars? In this talk I will describe some of the ways that embodied interaction is likely to change in the future, and some of the ways that we need to take those future scenarios into account as we design and implement multimodal interfaces.
Justine Cassell
ICMI1
2019 "You give a little of yourself": family support for children's use of an IVR literacy system
abstract
Low levels of childhood literacy in global contexts may be mitigated by educational technologies, however, these technologies often rely on parents of sufficient literacy to effectively support their children. Given low levels of adult literacy in many low-resource contexts, we investigate the nature of low-literate adult support for children's use of a literacy technology designed to foster early literacy precursors. We deployed an interactive voice response (IVR) system with 38 families in a rural village in Côte d'Ivoire using the IVR for 5 weeks in their homes. Using call log data and grounded theory analyses of IVR observations and interviews, we find evidence that families leverage complex support networks where family members support children's use of the IVR in different ways, via a collective network of intermediaries. These results suggest opportunities to scaffold low-literate family supporters for educational technologies.
Michael A. Madaio, Vikram Kamath Cannanure, Evelyn Yarzebinski, Shelby Zasacky, Fabrice Tanoh, Joelle Hannon-Cropp, Justine Cassell, Kaja Jasinska, Amy Ogan
COMPASS7
2019 A Model of Social Explanations for a Conversational Movie Recommendation System
abstract
A critical aspect of any recommendation process is explaining the reasoning behind each recommendation. These explanations can not only improve users' experiences, but also change their perception of the recommendation quality. This work describes our human-centered design for our conversational movie recommendation agent, which explains its decisions as humans would. After exploring and analyzing a corpus of dyadic interactions, we developed a computational model of explanations. We then incorporated this model in the architecture of a conversational agent and evaluated the resulting system via a user experiment. Our results show that social explanations can improve the perceived quality of both the system and the interaction, regardless of the intrinsic quality of the recommendations.
Florian Pecune, Shruti Murali, Vivian Tsai, Yoichi Matsuyama, Justine Cassell
HAI5
2018 Predicting the Temporal and Social Dynamics of Curiosity in Small Group Learning
Bhargavi Paranjape, Justine Cassell
AIED (1)3
2018 A User Simulator Architecture for Socially-Aware Conversational Agents
abstract
Over the last two decades, Reinforcement Learning (RL) has emerged the method of choice for data-driven dialog management. However, one of the limitations of RL methods for the optimization of dialog managers in the context of virtual conversational agents, is that they require a large amount of data, which is often unavailable, particularly when the dialog deals with complex discourse phenomena. User simulators help address this problem by generating synthetic data to train RL agents in an online fashion. In this work, we extend user simulators to the case of socially-aware conversational agents, that combine task and social functions. We propose a novel architecture that takes into consideration the user's conversational goals and generates both task and social behaviour. Our proposed architecture is general enough to be useful for training socially-aware conversational agents in any domain. As a proof of concept, we construct a user simulator for training a conversational recommendation agent and provide evidence towards the effectiveness of the approach.
Alankar Jain, Florian Pecune, Yoichi Matsuyama, Justine Cassell
IVA4
2018 Towards Automatic Generation of Peer-Targeted Science Talk in Curiosity-Evoking Virtual Agent
abstract
Curiosity is a critical skill that spurs learning, but is often found to decline with age and schooling. Recent research has shown that peer interaction may serve a special role in inducing curiosity through increased uncertainty and conceptual conflicts, since peers have similar authority in knowledge. For a virtual agent to stimulate curiosity, it should be able to generate curiosity-eliciting verbal behaviors such as hypothesis verbalization and argumentation, in the manner that simulates peer-like cognitive and behavioral abilities. In this paper, we design and implement a virtual peer that can carry out key curiosity-eliciting science talk during a dialog-based multi-party board game. We propose a child-centered and data-driven approach to simulate the latent reasoning process of young children and age-appropriate language during open-ended game play. In particular, we use a combination of child knowledge-graph construction and child-child interaction driven modeling to generate game appropriate behaviors that are compatible with 9-14 year old children. Encouraging human evaluation of the generated behaviors and generalizability of the generation framework to other tasks opens up new directions in incorporating open-endedness and science talk in virtual agents that will make them truly play a peer role in learning.
Bhargavi Paranjape, Yubin Ge, Jessica Hammer, Justine Cassell
IVA5
2017 A New Theoretical Framework for Curiosity for Learning in Social Contexts
Tanmay Sinha, Justine Cassell
EC-TEL3
2017 Curious Minds Wonder Alike: Studying Multimodal Behavioral Dynamics to Design Social Scaffolding of Curiosity
Tanmay Sinha, Justine Cassell
EC-TEL3
2017 Using Temporal Association Rule Mining to Predict Dyadic Rapport in Peer Tutoring
Michael A. Madaio, Rae Lasko, Justine Cassell, Amy Ogan
EDM3
2017 Cognitive-Inspired Conversational-Strategy Reasoner for Socially-Aware Agents
abstract
In this work we propose a novel module for a dialogue system that allows a conversational agent to utter phrases that do not just meet the system's task intentions, but also work towards achieving the system's social intentions. The module - a Social Reasoner - takes the task goals the system must achieve and decides the appropriate conversational style and strategy with which the dialogue system describes the information the user desires so as to boost the strength of the relationship between the user and system (rapport), and therefore the user's engagement and willingness to divulge the information the agent needs to efficiently and effectively achieve the user's goals. Our Social Reasoner is inspired both by analysis of empirical data of friends and stranger dyads engaged in a task, and by prior literature in fields as diverse as reasoning processes in cognitive and social psychology, decision-making, sociolinguistics and conversational analysis. Our experiments demonstrated that, when using the Social Reasoner in a Dialogue System, the rapport level between the user and system increases in more than 35% in comparison with those cases where no Social Reasoner is used.
Oscar J. Romero, Justine Cassell
IJCAI3
2017 Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia Content
abstract
The sheer amount of human-centric multimedia content has led to increased research on human behavior understanding. Most existing methods model behavioral sequences without considering the temporal saliency. This work is motivated by the psychological observation that temporally selective attention enables the human perceptual system to process the most relevant information. In this paper, we introduce a new approach, named Temporally Selective Attention Model (TSAM), designed to selectively attend to salient parts of human-centric video sequences. Our TSAM models learn to recognize affective and social states using a new loss function called speaker-distribution loss. Extensive experiments show that our model achieves the state-of-the-art performance on rapport detection and multimodal sentiment analysis. We also show that our speaker-distribution loss function can generalize to other computational models, improving the prediction performance of deep averaging network and Long Short Term Memory (LSTM).
Liangke Gui, Michael A. Madaio, Amy Ogan, Justine Cassell, Louis-Philippe Morency
ACM Multimedia5
2016 The Effect of Friendship and Tutoring Roles on Reciprocal Peer Tutoring Strategies
Michael A. Madaio, Amy Ogan, Justine Cassell
ITS3
2016 Socially-Aware Virtual Agents: Automatically Assessing Dyadic Rapport from Temporal Patterns of Behavior
Tanmay Sinha, Alan W. Black, Justine Cassell
IVA4
2016 Socially-Aware Animated Intelligent Personal Assistant Agent
abstract
SARA (Socially-Aware Robot Assistant) is an embodied intelligent personal assistant that analyses the user's visual (head and face movement), vocal (acoustic features) and verbal (conversational strategies) behaviours to estimate its rapport level with the user, and uses its own appropriate visual, vocal and verbal behaviors to achieve task and social goals.The presented agent aids conference attendees by eliciting their preferences through building rapport, and then making informed personalized recommendations about sessions to attend and people to meet.
Yoichi Matsuyama, Arjun Bhardwaj, Oscar Romeo, Sushma Akoju, Justine Cassell
SIGDIAL Conference6
2016 Automatic Recognition of Conversational Strategies in the Service of a Socially-Aware Dialog System
abstract
In this work, we focus on automatically recognizing social conversational strategies that in human conversation contribute to building, maintaining or sometimes destroying a budding relationship.These conversational strategies include self-disclosure, reference to shared experience, praise and violation of social norms.By including rich contextual features drawn from verbal, visual and vocal modalities of the speaker and interlocutor in the current and previous turn, we can successfully recognize these dialog phenomena with an accuracy of over 80% and kappa ranging from 60-80%.Our findings have been successfully integrated into an end-to-end socially aware dialog system, with implications for virtual agents that can use rapport between user and system to improve task-oriented assistance.
Tanmay Sinha, Alan W. Black, Justine Cassell
SIGDIAL Conference4
2015 Fine-Grained Analyses of Interpersonal Processes and Their Effect on Learning
Tanmay Sinha, Justine Cassell
AIED2
2015 Connecting the Dots: Predicting Student Grade Sequences from Bursty MOOC Interactions over Time
abstract
In this work, we track the interaction of students across multiple Massive Open Online Courses (MOOCs) on edX. Leveraging the ``burstiness" factor of three of the most commonly exhibited interaction forms made possible by online learning (i.e, video lecture viewing, coursework access and discussion forum posting), we take on the task of predicting student performance (operationalized as grade) across these courses. Specifically, we utilize the probabilistic framework of Conditional Random Fields (CRF) to formalize the problem of predicting the sequence of grades achieved by a student in different MOOCs, taking into account the contextual dependency of this outcome measure on students' general interaction trend across courses. Based on a comparative analysis of the combination of interaction features, our best CRF model can achieve a precision of 0.581, recall of 0.660 and a weighted F-score of 0.560, outweighing several baseline discriminative classifiers applied at each sequence position. These findings have implications for initiating early instructor intervention, so as to engage students along less active interaction dimensions that could be associated with low grades.
Tanmay Sinha, Justine Cassell
L@S2
2014 Towards a Computational Architecture of Dyadic Rapport Management for Virtual Agents
Alexandros Papangelis, Justine Cassell
IVA3
2014 Towards a Dyadic Computational Model of Rapport Management for Human-Virtual Agent Interaction
Alexandros Papangelis, Justine Cassell
IVA3
2013 The Effects of Culturally Congruent Educational Technologies on Student Achievement
Samantha L. Finkelstein, Evelyn Yarzebinski, Callie Vaughn, Amy Ogan, Justine Cassell
AIED5
2013 Automatic Prediction of Friendship via Multi-model Dyadic Features
Zhou Yu 0005, David Gerritsen, Amy Ogan, Alan W. Black, Justine Cassell
SIGDIAL Conference5
2012 "Oh dear stacy!": social interaction, elaboration, and learning with teachable agents
abstract
Understanding how children perceive and interact with teachable agents (systems where children learn through teaching a synthetic character embedded in an intelligent tutoring system) can provide insight into the effects of so-cial interaction on learning with intelligent tutoring systems. We describe results from a think-aloud study where children were instructed to narrate their experience teaching Stacy, an agent who can learn to solve linear equations with the student's help. We found treating her as a partner, primarily through aligning oneself with Stacy using pronouns like you or we rather than she or it significantly correlates with student learning, as do playful face-threatening comments such as teasing, while elaborate explanations of Stacy's behavior in the third-person and formal tutoring statements reduce learning gains. Additionally, we found that the agent's mistakes were a significant predictor for students shifting away from alignment with the agent.
Amy Ogan, Samantha L. Finkelstein, Elijah Mayfield, Claudia D'Adamo, Noboru Matsuda, Justine Cassell
CHI6
2012 Rudeness and Rapport: Insults and Learning Gains in Peer Tutoring
Amy Ogan, Samantha L. Finkelstein, Erin Walker, Ryan Carlson, Justine Cassell
ITS5
2012 "Love ya, jerkface": Using Sparse Log-Linear Models to Build Positive and Impolite Relationships with Teens
William Yang Wang, Samantha L. Finkelstein, Amy Ogan, Alan W. Black, Justine Cassell
SIGDIAL Conference5
2011 Brick by brick: iterating interventions to bridge the achievement gap with virtual peers
abstract
We lay out one strand of a continuing investigation into the development of a virtual peer to help children learn to use "school English" and "school-ratified science talk". In this paper we describe a detailed analysis of a corpus of child-child language use, and report our findings on the ways children shift dialects and ways of discussing science depending on the social context and task. We discuss the implications of these results for the redesign of a virtual peer that can evoke language behaviors associated with student achievement. Furthermore, our results allow us to describe the ways in which this virtual agent can tailor its level of interaction based on a child's current aptitude in this area.
Emilee Rader, Margaret Echelbarger, Justine Cassell
CHI3
2010 Report on the Second NLG Challenge on Generating Instructions in Virtual Environments (GIVE-2)
Alexander Koller, Kristina Striegnitz, Andrew Gargett, Donna Byron, Justine Cassell, Robert Dale, Johanna D. Moore, Jon Oberlander
INLG5
2009 Modeling culturally authentic style shifting with virtual peers
abstract
We report on a new kind of culturally-authentic embodied conversational agent more in line with the ways that culture and ethnicity function in the real world. On the basis of the careful analysis of a corpus of verbal and nonverbal behavior, we found that children shift dialects and ways of using their body depending on social context and task. Based on these results, we implemented a culturally authentic African American virtual peer capable of "code-switching" between African American English and Mainstream American English, and of using nonverbal behavior differently, depending on context. An evaluation of the agent revealed that the virtual peer elicited the same style changes in real children as real children did in one another.
Justine Cassell, Kathleen Geraghty, Berto Gonzalez, John Borland
ICMI1
2008 Designing virtual peers for assessment and intervention for children with autism
abstract
Our research focuses on the use and design of virtual peers (life-sized, computer-animated children) as intervention and assessment tools for the social and communication skills of children with social skills deficits, such as autism. To best design a virtual peer that simulates human interaction, we observe and analyze the behaviors of both typically-developing children and children with autism as they play with peers. Later, we apply these behavioral characteristics to the behavioral repertoire of the virtual peers. This analysis identifies the key design attributes of a virtual peer that best elicits the social and communication skills we are interested in evaluating and addressing during the assessment and treatment procedures.
Julia Merryman, Andrea Tartaro, Miri Arie, Justine Cassell
IDC4
2008 Modelling rapport in embodied conversational agents
Justine Cassell
INTERSPEECH1
2007 Ethnic Identity and Engagement in Embodied Conversational Agents
Francisco Iacobelli, Justine Cassell
IVA2
2007 The Behavior Markup Language: Recent Developments and Challenges
Hannes Högni Vilhjálmsson, Nathan Cantelmo, Justine Cassell, Nicolas Ech Chafai, Michael Kipp, Stefan Kopp, Maurizio Mancini, Stacy Marsella, Andrew N. Marshall, Catherine Pelachaud, Zsófia Ruttkay, Kristinn R. Thórisson, Herwin van Welbergen, Rick J. van der Werf
IVA3
2005 Learning with Virtual Peers
Justine Cassell
AIED1
2005 Oral tradition, aboral coordination: building rapport with embodied conversational agents
abstract
Oral tradition, aboral coordination: building rapport with embodied conversational agents Harmony or rapport between people is essential for relationships as diverse as seller-buyer and teacher-learner. In this talk I describe the kinds of verbal behaviors -- such as common interactional structures and narrative resonance -- and non-verbal behaviors -- such as attention, positivity, and coordination -- that function together to establish a sense of rapport between two people in conversation. These studies are used as the basis for the implementation of virtual peers -- adults, but also more recently embodied conversational virtual children who are capable of acting as friends and learning partners with real children from different ethnic traditions, collaborating to tell stories from the child's own cultural context, and aiding children in making the transition between home and school language.
Justine Cassell
IUI1
2004 Towards integrated microplanning of language and iconic gesture for multimodal output
abstract
When talking about spatial domains, humans frequently accompany their explanations with iconic gestures to depict what they are referring to. For example, when giving directions, it is common to see people making gestures that indicate the shape of buildings, or outline a route to be taken by the listener, and these gestures are essential to the understanding of the directions. Based on results from an ongoing study on language and gesture in direction-giving, we propose a framework to analyze such gestural images into semantic units (image description features), and to link these units to morphological features (hand shape, trajectory, etc.). This feature-based framework allows us to generate novel iconic gestures for embodied conversational agents, without drawing on a lexicon of canned gestures. We present an integrated microplanner that derives the form of both coordinated natural language and iconic gesture directly from given communicative goals, and serves as input to the speech and gesture realization engine in our NUMACK project.
Stefan Kopp, Paul Tepper, Justine Cassell
ICMI3
2003 Towards a Model of Face-to-Face Grounding
abstract
We investigate the verbal and nonverbal means for grounding, and propose a design for embodied conversational agents that relies on both kinds of signals to establish common ground in human-computer interaction. We analyzed eye gaze, head nods and attentional focus in the context of a direction-giving task. The distribution of nonverbal behaviors differed depending on the type of dialogue move being grounded, and the overall pattern reflected a monitoring of lack of negative feedback. Based on these results, we present an ECA that uses verbal and nonverbal grounding acts to update dialogue state.
Yukiko I. Nakano, Gabe Reinstein, Tom Stocky, Justine Cassell
ACL4
2003 Negotiated Collusion: Modeling Social Languageand its Relationship Effects in Intelligent Agents
Justine Cassell, Timothy W. Bickmore
User Model. User Adapt. Interact.1
2002 Shared reality: spatial intelligence in intuitive user interfaces
abstract
In this paper, we describe an interface that demonstrates spatial intelligence. This interface, an embodied conversational kiosk, builds on research in embodied conversational agents (ECAs) and on information displays in mixed reality and kiosk format. ECAs leverage people's abilities to coordinate information displayed in multiple modalities, particularly information conveyed in speech and gesture. Mixed reality depends on users' interactions with everyday objects that are enhanced with computational overlays. We describe an implementation, MACK (Media lab Autonomous Conversational Kiosk), an ECA who can answer questions about and give directions to the MIT Media Lab's various research groups, projects and people. MACK uses a combination of speech, gesture, and indications on a normal paper map that users place on a table between themselves and MACK. Research issues involve users' differential attention to hand gestures, speech and the map, and how reference using these modalities can be fused in input and generation.
Tom Stocky, Justine Cassell
IUI2
2002 Levels of Detail for Crowds and Groups
abstract
Abstract Work on levels of detail for human simulation has occurred mainly on a geometrical level, either by reducing the numbers of polygons representing a virtual human, or replacing them with a two‐dimensional imposter. Approaches that reduce the complexity of motions generated have also been proposed. In this paper, we describe ongoing development of a framework for Adaptive Level Of Detail for Human Animation (ALOHA), which incorporates levels of detail for not only geometry and motion, but also includes a complexity gradient for natural behaviour, both conversational and social. ACM CSS: I.3.7 Three‐Dimensional Graphics and Realism—Animation
Carol O'Sullivan, Justine Cassell, Hannes Högni Vilhjálmsson, John Dingliana, Simon Dobbyn, B. McNamee, Christopher Peters 0001, Thanh Giang
Comput. Graph. Forum2
2001 Non-Verbal Cues for Discourse Structure
abstract
This paper addresses the issue of designing embodied conversational agents that exhibit appropriate posture shifts during dialogues with human users. Previous research has noted the importance of hand gestures, eye gaze and head nods in conversations between embodied agents and humans. We present an analysis of human monologues and dialogues that suggests that postural shifts can be predicted as a function of discourse state in monologues, and discourse and conversation state in dialogues. On the basis of these findings, we have implemented an embodied conversational agent that uses Collagen in such a way as to generate postural shifts.
Justine Cassell, Yukiko I. Nakano, Timothy W. Bickmore, Candace L. Sidner, Charles Rich
ACL1
2001 Relational agents: a model and implementation of building user trust
abstract
Building trust with users is crucial in a wide range of applications, such as financial transactions, and some minimal degree of trust is required in all applications to even initiate and maintain an interaction with a user. Humans use a variety of relational conversational strategies, including small talk, to establish trusting relationships with each other. We argue that such strategies can also be used by interface agents, and that embodied conversational agents are ideally suited for this task given the myriad cues available to them for signaling trustworthiness. We describe a model of social dialogue, an implementation in an embodied conversation agent, and an experiment in which social dialogue was demonstrated to have an effect on trust, for users with a disposition to be extroverts.
Timothy W. Bickmore, Justine Cassell
CHI2
2001 BEAT: the Behavior Expression Animation Toolkit
abstract
The Behavior Expression Animation Toolkit (BEAT) allows animators to input typed text that they wish to be spoken by an animated human figure, and to obtain as output appropriate and synchronized nonverbal behaviors and synthesized speech in a form that can be sent to a number of different animation systems. The nonverbal behaviors are assigned on the basis of actual linguistic and contextual analysis of the typed text, relying on rules derived from extensive research into human conversational behavior. The toolkit is extensible, so that new rules can be quickly added. It is designed to plug into larger systems that may also assign personality profiles, motion characteristics, scene constraints, or the animation styles of particular animators.
Justine Cassell, Hannes Högni Vilhjálmsson, Timothy W. Bickmore
SIGGRAPH1
2001 More than just a pretty face: conversational protocols and the affordances of embodiment
Justine Cassell, Timothy W. Bickmore, Lee Campbell, Hannes Högni Vilhjálmsson, Hao Yan 0003
Knowl. Based Syst.1
2001 Making Space for Voice: Technologies to Support Children's Fantasy and Storytelling
Justine Cassell, Kimiko Ryokai
Pers. Ubiquitous Comput.1
2000 Gesture in Conversation: Problems in Description and Interpretation
Justine Cassell
FG1
2000 Coordination and context-dependence in the generation of embodied conversation
abstract
We describe the generation of communicative actions in an implemented embodied conversational agent. Our agent plans each utterance so that multiple communicative goals may be realized opportunistically by a composite action including not only speech but also coverbal gesture that fits the context and the ongoing speech in ways representative of natural human conversation. We accomplish this by reasoning from a grammar which describes gesture declaratively in terms of its discourse function, semantics and synchrony with speech.
Justine Cassell, Matthew Stone, Hao Yan 0003
INLG1
2000 More than just a pretty face: affordances of embodiment
abstract
Prior research into embodied interface agents has found that users like them and find them engaging. In this paper, we argue that embodiment can serve an even stronger function if system designers use actual human conversational protocols in the design of the interface. Communicative behaviors such as salutations and farewells, conversational turn-taking with interruptions, and referring to objects using pointing gestures are examples of protocols that all native speakers of a language already know how to perform and that can thus be leveraged in an intelligent interface. We discuss how these protocols are integrated into Rea, an embodied, multi-modal conversational interface agent who acts as a real-estate salesperson, and we show why embodiment is required for their successful implementation.
Justine Cassell, Timothy W. Bickmore, Hannes Högni Vilhjálmsson, Hao Yan 0003
IUI1
1999 Embodiment in Conversational Interfaces: Rea
abstract
In this paper, we argue for embodied corrversational characters as the logical extension of the metaphor of human - computer interaction as a conversation. We argue that the only way to fully model the richness of human I&+ to-face communication is to rely on conversational analysis that describes sets of conversational behaviors as fi~lfilling conversational functions, both interactional and propositional. We demonstrate how to implement this approach in Rea, an embodied conversational agent that is capable of both multimodal input understanding and output generation in a limited application domain. Rea supports both social and task-oriented dialogue. We discuss issues that need to be addressed in creating embodied conversational agents, and describe the architecture of the Rea interface.
Justine Cassell, Timothy W. Bickmore, Mark Billinghurst, Lee Campbell, K. Chang, Hannes Högni Vilhjálmsson, Hao Yan 0003
CHI1
1999 StoryMat: A Play Space with Narrative Memories
abstract
In this paper, we present the design and the prototype of a work-in-progress, StoryMat: a soft intelligent play mat that records and recalls children's storytelling activities.
Kimiko Ryokai, Justine Cassell
IUI2
1999 Fully Embodied Conversational Avatars: Making Communicative Behaviors Autonomous
Justine Cassell, Hannes Högni Vilhjálmsson
Auton. Agents Multi Agent Syst.1
1998 Interactive Storytelling Environments: Coping with Cardiac Illness at Boston's Children's Hospital
abstract
This paper describes exploration of uses of a computational storytelling environment on the Cardiology Unit of the Children's Hospital in Boston during the summer of 1997.Young cardiac patients ranging from age 7 to 16 used the SAGE environment to tell personal stories and create interactive characters, as a way of coping with cardiac ihness, hospitalizations, and invasive medical procedures.This pilot study is part of a larger collaborative effort between Children's Hospital and hZERL -A Mitsubishi Electric Research L&oratory to develop a web-based application, de Experience Journal, to assist patients and their families in dealing with serious medical illness.The focus of the paper is on young patients' uses of SAGE, on SAGE's aftordances in the context of the hospital, and on design recommendations for the development of future computational play kits.Preliminary analysis of children's stories indicates that children used different modes of interaction-direct, media&, and differ&-depending upon what personae the narrator chooses to take on.These modes seem to vary with the mindset and health condition of the Child.
Marina Umaschi Bers, Edith Ackermann, Justine Cassell, Beth Donegan, Joseph Gonzalez-Heydrich, David Ray DeMaso, Carol Strohecker, Sarah Lualdi, Dennis Bromley, Judith Karlin
CHI3
1997 Temporal Classification of Natural Gesture and Application to Video Coding
abstract
A method for the temporal classification of natural gesture from video imagery is presented. The work is motivated by recent developments in the theory of natural gesture which have identified several key temporal aspects of gesture important to communication. In particular gesticulation during conversation can be coarsely characterized as periods of bi-phasic or tri-phasic gesture separated by a rest state. We first present an automatic procedure for hypothesizing plausible rest state configurations of a speaker. Second, we develop a state-based parsing algorithm used to both select among candidate rest states and to parse an incoming video stream into bi-phasic and tri-phasic gestures. Finally, we demonstrate the use of the bi-phasic/tri-phasic labeling to select semantically significant static images for low bandwidth coding of video of story-telling speakers.
Andrew D. Wilson, Aaron F. Bobick, Justine Cassell
CVPR3
1997 The implications of a theory of play for the design of computer toys (panel)
abstract
Article Free Access Share on The implications of a theory of play for the design of computer toys (panel) Authors: Bill Kolomyjec View Profile , Justine Cassell MIT Media Lab MIT Media LabView Profile , Yasmine B. Kafai University of California, Los Angeles University of California, Los AngelesView Profile , Mary Williamson University of California, Berkeley University of California, BerkeleyView Profile Authors Info & Claims SIGGRAPH '97: Proceedings of the 24th annual conference on Computer graphics and interactive techniquesAugust 1997 Pages 431–433https://doi.org/10.1145/258734.258898Published:03 August 1997Publication History 3citation611DownloadsMetricsTotal Citations3Total Downloads611Last 12 Months52Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Bill Kolomyjec, Justine Cassell, Yasmin B. Kafai, Mary Williamson
SIGGRAPH2
1996 What You Need to Know about Spontaneous Gesture, and You Need to Know It
Justine Cassell, David McNeil
FG1
1996 Recovering the Temporal Structure of Natural Gesture
abstract
A method for the recovery of the temporal structure and phases in natural gesture is presented. The work is motivated by recent developments in the theory of natural gesture which have identified several key aspects of gesture important to communication. In particular, gesticulation during conversation can be coarsely characterized as periods of bi-phasic or tri-phasic gesture separated by a rest state. We first present an automatic procedure for hypothesizing plausible rest state configurations of a speaker; the method uses the repetition of subsequences to indicate potential rest states. Second, we develop a state-based parsing algorithm used to both select among candidate rest stares and to parse an incoming video stream into bi-phasic and multi-phasic gestures. We present results from examples of story-telling speakers.
Andrew D. Wilson, Aaron F. Bobick, Justine Cassell
FG3
1994 Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
abstract
We describe an implemented system which automatically generates and animates conversations between multiple human-like agents with appropriate and synchronized speech, intonation, facial expressions, and hand gestures. Conversation is created by a dialogue planner that produces the text as well as the intonation of the utterances. The speaker/listener relationship, the text, and the intonation in turn drive facial expressions, lip motions, eye gaze, head motion, and arm gestures generators. Coordinated arm, wrist, and hand motions are invoked to create semantically meaningful gestures. Throughout we will use examples from an actual synthesized, fully animated conversation.
Justine Cassell, Catherine Pelachaud, Norman I. Badler, Mark Steedman, Brett Achorn, Tripp Becket, Brett Douville, Scott Prevost, Matthew Stone
SIGGRAPH1