Sean Andrist

dblp:23/8725 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
5since 2021 · last 2025
0000-0003-4972-7027ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 20 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Converting Spatial to Social: Using Persistent Homology to Understand Social Groups
abstract
ICMI ’25, Canberra, ACT, Australia
Valerie K. Chen, Claire Liang, Julie A. Shah, Sean Andrist
ICMI4
2024 "Uh, This One?": Leveraging Behavioral Signals for Detecting Confusion during Physical Tasks
abstract
A longstanding goal in the AI and HCI research communities is building intelligent assistants to help people with physical tasks. To be effective in this, AI assistants must be aware of not only the physical environment, but also the human user and their cognitive states. In this paper, we specifically consider the detection of confusion, which we operationalize as the moments when a user is “stuck” and needs assistance. We explore how behavioral features such as gaze, head pose, and hand movements differ between periods of confusion vs no-confusion. We present various modeling approaches for detecting confusion that combine behavioral features, length of time, instructional text embeddings, and egocentric video. Although deep networks (e.g., V-Jepa) trained on full video streams perform well in distinguishing confusion from non-confusion, simpler models leveraging lighter weight behavioral features exhibit similarly high performance, even when generalizing to unseen tasks.
Maia Stiber, Dan Bohus, Sean Andrist
ICMI3
2024 Interaction-Shaping Robotics: Robots That Influence Interactions between Other Agents
abstract
Work in Human–Robot Interaction (HRI) has investigated interactions between one human and one robot as well as human–robot group interactions. Yet the field lacks a clear definition and understanding of the influence a robot can exert on interactions between other group members (e.g., human-to-human). In this article, we define Interaction-Shaping Robotics (ISR), a subfield of HRI that investigates robots that influence the behaviors and attitudes exchanged between two (or more) other agents. We highlight key factors of interaction-shaping robots that include the role of the robot, the robot-shaping outcome, the form of robot influence, the type of robot communication, and the timeline of the robot’s influence. We also describe three distinct structures of human–robot groups to highlight the potential of ISR in different group compositions and discuss targets for a robot’s interaction-shaping behavior. Finally, we propose areas of opportunity and challenges for future research in ISR.
Sarah Gillet, Marynel Vázquez, Sean Andrist, Iolanda Leite, Sarah Sebo
ACM Trans. Hum. Robot Interact.3
2023 HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World
abstract
Building an interactive AI assistant that can perceive, reason, and collaborate with humans in the real world has been a long-standing pursuit in the AI community. This work is part of a broader research effort to develop intelligent agents that can interactively guide humans through performing tasks in the physical world. As a first step in this direction, we introduce HoloAssist, a large-scale egocentric human interaction dataset, where two people collaboratively complete physical manipulation tasks. The task performer executes the task while wearing a mixed-reality headset that captures seven synchronized data streams. The task instructor watches the performer’s egocentric video in real time and guides them verbally. By augmenting the data with action and conversational annotations and observing the rich behaviors of various participants, we present key insights into how human assistants correct mistakes, intervene in the task completion procedure, and ground their instructions to the environment. HoloAssist spans 166 hours of data captured by 350 unique instructor-performer pairs. Furthermore, we construct and present benchmarks on mistake detection, intervention type prediction, and hand forecasting, along with detailed analysis. We expect HoloAssist will provide an important resource for building AI assistants that can fluidly collaborate with humans in the real world. Data can be downloaded at https://holoassist.github.io/.
Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, Marc Pollefeys
ICCV6
2022 Continual Learning about Objects in the Wild: An Interactive Approach
abstract
We introduce a mixed-reality, interactive approach for continually learning to recognize an open-ended set of objects in a user’s surrounding environment. The proposed approach leverages the multimodal sensing, interaction, and rendering affordances of a mixed-reality headset, and enables users to label nearby objects via speech, gaze, and gestures. Image views of each labeled object are automatically captured from varying viewpoints over time, as the user goes about their everyday tasks. The labels provided by the user can be propagated forward and backwards in time and paired with the collected views to update an object recognition model, in order to continually adapt it to the user’s specific objects and environment. We review key challenges for the proposed interactive continual learning approach, present details of an end-to-end system implementation, and report on results and lessons learned from an initial, exploratory case study using the system.
Dan Bohus, Sean Andrist, Ashley Feniello, Nick Saw, Eric Horvitz
ICMI2
2020 Metareasoning in Modular Software Systems: On-the-Fly Configuration Using Reinforcement Learning with Rich Contextual Representations
Aditya Modi 0002, Debadeepta Dey, Alekh Agarwal, Adith Swaminathan, Besmira Nushi, Sean Andrist, Eric Horvitz
AAAI6
2020 Lessons Learned in Designing AI for Autistic Adults
abstract
Through an iterative design process using Wizard of Oz (WOz) prototypes, we designed a video calling application for people with Autism Spectrum Disorder. Our Video Calling for Autism prototype provided an Expressiveness Mirror that gave feedback to autistic people on how their facial expressions might be interpreted by their neurotypical conversation partners. This feedback was in the form of emojis representing six emotions and a bar indicating the amount of overall expressiveness demonstrated by the user. However, when we built a working prototype and conducted a user study with autistic participants, their negative feedback caused us to reconsider how our design process led to a prototype that they did not find useful. We reflect on the design challenges around developing AI technology for an autistic user population, how Wizard of Oz prototypes can be overly optimistic in representing AI-driven prototypes, how autistic research participants can respond differently to user experience prototypes of varying fidelity, and how designing for people with diverse abilities needs to include that population in the development process.
Andrew Begel, John C. Tang, Sean Andrist, Michael Barnett 0001, Tony Carbary, Piali Choudhury, Edward Cutrell, Alberto Fung, Sasa Junuzovic, Daniel McDuff, Kael Rowan, Shibashankar Sahoo, Jennifer Frances Waldern, Jessica Wolk, Annuska Z. Perkins
ASSETS3
2020 REFORM: Recognizing F-formations for Social Robots
abstract
Recognizing and understanding conversational groups, or F-formations, is a critical task for situated agents designed to interact with humans. F-formations contain complex structures and dynamics, yet are used intuitively by people in everyday face-to-face conversations. Prior research exploring ways of identifying F-formations has largely relied on heuristic algorithms that may not capture the rich dynamic behaviors employed by humans. We introduce REFORM (REcognize F-FORmations with Machine learning), a data-driven approach for detecting F-formations given human and agent positions and orientations. REFORM decomposes the scene into all possible pairs and then reconstructs F-formations with a voting-based scheme. We evaluated our approach across three datasets: the SALSA dataset, a newly collected human-only dataset, and a new set of acted human-robot scenarios, and found that REFORM yielded improved accuracy over a state-of-the-art F-formation detection algorithm. We also introduce symmetry and tightness as quantitative measures to characterize F-formations.
Hooman Hedayati, Annika Muehlbradt, Daniel Szafir, Sean Andrist
IROS4
2019 Demonstrating a Framework for Rapid Development of Physically Situated Interactive Systems
abstract
We demonstrate an open, extensible framework for enabling faster development and study of physically situated interactive systems. The framework provides a programming model for parallel coordinated computation centered on temporal streams of data, a set of tools for data visualization and processing, and an open ecosystem of components. The demonstration showcases an interaction toolkit of components for systems that interact with people via natural language in the open world.
Sean Andrist, Dan Bohus, Ashley Feniello
HRI1
2019 Recognizing F-Formations in the Open World
abstract
A key skill for social robots in the wild will be to understand the structure and dynamics of conversational groups in order to fluidly participate in them. Social scientists have long studied the rich complexity underlying such focused encounters, or F-formations. However, current state-of-the-art algorithms that robots might use to recognize F-formations are highly heuristic and quite brittle. In this report, we explore a data-driven approach to detect F-formations from sets of tracked human positions and orientations, trained and evaluated on two openly available human-only datasets and a small human-robot dataset that we collected. We also discuss the potential for further computational characterization of F-formations beyond simply detecting their occurrence.
Hooman Hedayati, Daniel Szafir, Sean Andrist
HRI3
2019 Managing Stress: The Needs of Autistic Adults in Video Calling
abstract
Video calling (VC) aims to create multi-modal, collaborative environments that are "just like being there." However, we found that autistic individuals, who exhibit atypical social and cognitive processing, may not share this goal. We interviewed autistic adults about their perceptions of VC compared to other computer- mediated communications (CMC) and face-to-face interactions. We developed a neurodiversity-sensitive model of CMC that describes how stressors such as sensory sensitivities, cognitive load, and anxiety, contribute to their preferences for CMC channels. We learned that they apply significant effort to construct coping strategies to support their sensory, cognitive, and social needs. These strategies include moderating their sensory inputs, creating mental models of conversation partners, and attempting to mask their autism by adopting neurotypical behaviors. Without effective strategies, interviewees experience more stress, have less capacity to interpret verbal and non-verbal cues, and feel less empowered to participate. Our findings reveal critical needs for autistic users. We suggest design opportunities to support their ability to comfortably use VC, and in doing so, point the way towards making VC more comfortable for all.
Annuska Z. Perkins, Andrew Begel, Jennifer Frances Waldern, John C. Tang, Michael Barnett 0001, Edward Cutrell, Daniel McDuff, Sean Andrist, Meredith Ringel Morris
Proc. ACM Hum. Comput. Interact.8
2017 Looking Coordinated: Bidirectional Gaze Mechanisms for Collaborative Interaction with Virtual Characters
abstract
Successful collaboration relies on the coordination and alignment of communicative cues. In this paper, we present mechanisms of bidirectional gaze - the coordinated production and detection of gaze cues - by which a virtual character can coordinate its gaze cues with those of its human user. We implement these mechanisms in a hybrid stochastic/heuristic model synthesized from data collected in human-human interactions. In three lab studies wherein a virtual character instructs participants in a sandwich-making task, we demonstrate how bidirectional gaze can lead to positive outcomes in error rate, completion time, and the agent's ability to produce quick, effective nonverbal references. The first study involved an on-screen agent and the participant wearing eye-tracking glasses. The second study demonstrates that these positive outcomes can be achieved using head-pose estimation in place of full eye tracking. The third study demonstrates that these effects also transfer into virtual-reality interactions.
Sean Andrist, Michael Gleicher, Bilge Mutlu
CHI1
2017 Rapid development of multimodal interactive systems: a demonstration of platform for situated intelligence
abstract
We demonstrate an open, extensible platform for developing and studying multimodal, integrative-AI systems. The platform provides a time-aware, stream-based programming model for parallel coordinated computation, a set of tools for data visualization, processing, and learning, and an ecosystem of pluggable AI components. The demonstration will showcase three applications built on this platform and highlight how the platform can significantly accelerate development and research in multimodal interactive systems.
Dan Bohus, Sean Andrist, Mihai Jalobeanu
ICMI2
2016 Are You Messing with Me?: Querying about the Sincerity of Interactions in the Open World
abstract
When interacting with robots deployed in the open world, people may often attempt to engage with them in a playful manner or test their competencies. Such engagements are often associated with language and behaviors that fall outside of designed task capabilities and can lead to interaction failures. Detecting when users are driven by play and curiosity can help a robot to understand why some interactions are breaking down, respond more appropriately by conveying its capabilities to its users, and enhance perceptions of its situational awareness and social intelligence. We have been studying the intentions of everyday users in their engagement with a long-lived robot system that provides directions within an office building. We report on a pilot field-study exploring the use of direct queries to elicit the sincerity of user requests, in terms of their actual need for directions. We discuss early results from this initial study and frame research directions and design implications for robots deployed in the wild.
Sean Andrist, Dan Bohus, Eric Horvitz
HRI1
2015 Look Like Me: Matching Robot Personality via Gaze to Increase Motivation
abstract
Socially assistive robots are envisioned to provide social and cognitive assistance where they will seek to motivate and engage people in therapeutic activities. Due to their physicality, robots serve as a powerful technology for motivating people. Prior work has shown that effective motivation requires adaption to user needs and characteristics, but how robots might successfully achieve such adaptation is still unknown. In this paper, we present work on matching a robot's personality-expressed via its gaze behavior-to that of its users. We confirmed in an online study with 22 participants that the robot's gaze behavior can successfully express either an extroverted or introverted personality. In a laboratory study with 40 participants, we demonstrate the positive effect of personality matching on a user's motivation to engage in a repetitive task. These results have important implications for the design of adaptive robot behaviors in assistive human-robot interaction.
Sean Andrist, Bilge Mutlu, Adriana Tapus
CHI1
2015 Effects of Culture on the Credibility of Robot Speech: A Comparison between English and Arabic
abstract
As social robots begin to enter our lives as providers of information, assistance, companionship, and motivation, it becomes increasingly important that these robots are capable of interacting effectively with human users across different cultural settings worldwide. A key capability in establishing acceptance and usability is the way in which robots structure their speech to build credibility and express information in a meaningful and persuasive way. Previous work has established that robots can use speech to improve credibility in two ways: expressing practical knowledge and using rhetorical linguistic cues. In this paper, we present two studies that build on prior work to explore the effects of language and cultural context on the credibility of robot speech. In the first study (n=96), we compared the relative effectiveness of knowledge and rhetoric on the credibility of robot speech between Arabic-speaking robots in Lebanon and English-speaking robots in the USA, finding the rhetorical linguistic cues to be more important in Arabic than in English. In the second study (n=32), we compared the effectiveness of credible robot speech between robots speaking either Modern Standard Arabic or the local Arabic dialect, finding the expression of both practical knowledge and rhetorical ability to be most important when using the local dialect. These results reveal nuanced cultural differences in perceptions of robots as credible agents and have important implications for the design of human-robot interactions across Arabic and Western cultures.
Sean Andrist, Micheline Ziadee, Halim Boukaram, Bilge Mutlu, Majd F. Sakr
HRI1
2015 A Review of Eye Gaze in Virtual Agents, Social Robotics and HCI: Behaviour Generation, User Interaction and Perception
abstract
Abstract A person's emotions and state of mind are apparent in their face and eyes. As a Latin proverb states: ‘The face is the portrait of the mind; the eyes, its informers’. This presents a significant challenge for Computer Graphics researchers who generate artificial entities that aim to replicate the movement and appearance of the human eye, which is so important in human–human interactions. This review article provides an overview of the efforts made on tackling this demanding task. As with many topics in computer graphics, a cross‐disciplinary approach is required to fully understand the workings of the eye in the transmission of information to the user. We begin with a discussion of the movement of the eyeballs, eyelids and the head from a physiological perspective and how these movements can be modelled, rendered and animated in computer graphics applications. Furthermore, we present recent research from psychology and sociology that seeks to understand higher level behaviours, such as attention and eye gaze, during the expression of emotion or during conversation. We discuss how these findings are synthesized in computer graphics and can be utilized in the domains of Human–Robot Interaction and Human–Computer Interaction for allowing humans to interact with virtual agents and other artificial entities. We conclude with a summary of guidelines for animating the eye and head from the perspective of a character animator.
Kerstin Ruhland, Christopher Peters 0001, Sean Andrist, Jeremy B. Badler, Norman I. Badler, Michael Gleicher, Bilge Mutlu, Rachel McDonnell
Comput. Graph. Forum3
2015 Gaze and Attention Management for Embodied Conversational Agents
abstract
To facilitate natural interactions between humans and embodied conversational agents (ECAs), we need to endow the latter with the same nonverbal cues that humans use to communicate. Gaze cues in particular are integral in mechanisms for communication and management of attention in social interactions, which can trigger important social and cognitive processes, such as establishment of affiliation between people or learning new information. The fundamental building blocks of gaze behaviors are gaze shifts : coordinated movements of the eyes, head, and body toward objects and information in the environment. In this article, we present a novel computational model for gaze shift synthesis for ECAs that supports parametric control over coordinated eye, head, and upper body movements. We employed the model in three studies with human participants. In the first study, we validated the model by showing that participants are able to interpret the agent’s gaze direction accurately. In the second and third studies, we showed that by adjusting the participation of the head and upper body in gaze shifts, we can control the strength of the attention signals conveyed, thereby strengthening or weakening their social and cognitive effects. The second study shows that manipulation of eye--head coordination in gaze enables an agent to convey more information or establish stronger affiliation with participants in a teaching task, while the third study demonstrates how manipulation of upper body coordination enables the agent to communicate increased interest in objects in the environment.
Tomislav Pejsa, Sean Andrist, Michael Gleicher, Bilge Mutlu
ACM Trans. Interact. Intell. Syst.2
2014 Conversational gaze aversion for humanlike robots
abstract
Gaze aversion-the intentional redirection away from the face of an interlocutor-is an important nonverbal cue that serves a number of conversational functions, including signaling cognitive effort, regulating a conversation's intimacy level, and managing the conversational floor. In prior work, we developed a model of how gaze aversions are employed in conversation to perform these functions. In this paper, we extend the model to apply to conversational robots, enabling them to achieve some of these functions in conversations with people. We present a system that addresses the challenges of adapting human gaze aversion movements to a robot with very different affordances, such as a lack of articulated eyes. This system, implemented on the NAO platform, autonomously generates and combines three distinct types of robot head movements with different purposes: face-tracking movements to engage in mutual gaze, idle head motion to increase lifelikeness, and purposeful gaze aversions to achieve conversational functions. The results of a human-robot interaction study with 30 participants show that gaze aversions implemented with our approach are perceived as intentional, and robots can use gaze aversions to appear more thoughtful and effectively manage the conversational floor.
Sean Andrist, Xiang Zhi Tan, Michael Gleicher, Bilge Mutlu
HRI1
2013 Fun and fair: influencing turn-taking in a multi-party game with a virtual agent
abstract
Language-based interfaces for children hold great promise in education, therapy, and entertainment. An important subset of these interfaces includes those with a virtual agent that mediates the interaction. When participants are groups of children, the agent will need to exert a certain amount of turn-taking control to ensure that all group members participate and benefit from the experience, but must do so without being so overtly directive as to undermine the children's enjoyment of and engagement in the task. We present a hierarchy of nonverbal and verbal behaviors that a virtual agent can employ flexibly when passing the conversational turn. When used effectively, these behaviors can equalize participation, and potentially decrease the amount of overlapping speech among participants, improving automatic speech recognition in turn. We evaluated the behaviors by having children play a language-based game twice, once with a flexible host and once with an inflexible host that did not have access to the behaviors. Post-game opinion cards revealed no difference between the conditions with respect to fun or likability of the host, despite the flexible agent eliciting more evenly distributed play.
Sean Andrist, Iolanda Leite, Jill Fain Lehman
IDC1
2013 Rhetorical robots: making robots more effective speakers using linguistic cues of expertise
Sean Andrist, Erin Spannan, Bilge Mutlu
HRI1
2013 Controllable models of gaze behavior for virtual agents and humanlike robots
abstract
Embodied social agents, through their ability to afford embodied interaction using nonverbal human communicative cues, hold great promise in application areas such as education, training, rehabilitation, and collaborative work. Gaze cues are particularly important for achieving significant social and communicative goals. In this research, I explore how agents - both virtual agents and humanlike robots - might achieve such goals through the use of various gaze mechanisms. To this end, I am developing computational control models of gaze behavior that treat gaze as the output of a system with a number of multimodal inputs. These inputs can be characterized at different levels of interaction, from non-interactive (e.g., physical characteristics of the agent itself) to fully interactive (e.g., speech and gaze behavior of a human interlocutor). This research will result in a number of control models that each focus on a different gaze mechanism, combined into an open-source library of gaze behaviors that will be usable by both human-robot and human-virtual agent interaction designers. System-level evaluations in naturalistic settings will validate this gaze library for its ability to evoke positive social and cognitive responses in human users.
Sean Andrist
ICMI1
2013 Managing chaos: models of turn-taking in character-multichild interactions
abstract
Turn-taking decisions in multiparty settings are complex, especially when the participants are children. Our goal is to endow an interactive character with appropriate turn-taking behavior using visual, audio and contextual features. To that end, we investigate three distinct turn-taking models: a baseline model grounded in established turn-taking rules for adults and two machine learning models, one trained with data collected in situ and the other trained with data collected in more controlled conditions. The three models are shown to have different profiles of behavior during silences, overlapping speech, and at the end of participants' turns. An exploratory user evaluation focusing on the decision points where the models differ showed clear preference for the machine learning models over the baseline model. The results indicate that the rules for language interactions with small groups of children are not simply an extension of the rules for interacting with small groups of adults.
Iolanda Leite, Hannaneh Hajishirzi, Sean Andrist, Jill Fain Lehman
ICMI3
2013 Conversational Gaze Aversion for Virtual Agents
Sean Andrist, Bilge Mutlu, Michael Gleicher
IVA1
2012 Designing effective gaze mechanisms for virtual agents
abstract
Virtual agents hold great promise in human-computer interaction with their ability to afford embodied interaction using nonverbal human communicative cues. Gaze cues are particularly important to achieve significant high-level outcomes such as improved learning and feelings of rapport. Our goal is to explore how agents might achieve such outcomes through seemingly subtle changes in gaze behavior and what design variables for gaze might lead to such positive outcomes. Drawing on research in human physiology, we developed a model of gaze behavior to capture these key design variables. In a user study, we investigated how manipulations in these variables might improve affiliation with the agent and learning. The results showed that an agent using affiliative gaze elicited more positive feelings of connection, while an agent using referential gaze improved participants' learning. Our model and findings offer guidelines for the design of effective gaze behaviors for virtual agents.
Sean Andrist, Tomislav Pejsa, Bilge Mutlu, Michael Gleicher
CHI1