Sharon Mozgai

dblp:205/1845 · also Sharon A. Mozgai · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-3308-2474ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Multi-Platform Intelligent Agents with the Virtual Human Toolkit
abstract
The research and development (R&D) of intelligent virtual agents (IVAs) is inherently complex.We aim to manage this complexity by releasing a major update to the Virtual Human Toolkit, combining the best aspects of academic and commercial approaches into a principled R&D platform that emphasizes interoperability, extendability, re-use, and support for multiple hardware targets.We demonstrate the current status of the VHToolkit on Windows desktop and Quest 3 AR/VR headset.
Arno Hartholt, Edward Fast, Kevin Kim, Andrew Leeds, Edwin Sookiassian, Sharon Mozgai
IVA6
2024 Estuary: A Framework For Building Multimodal Low-Latency Real-Time Socially Interactive Agents
abstract
The rise in capability and ubiquity of generative artificial intelligence (AI) technologies has enabled its application to the field of Socially Interactive Agents (SIAs). Despite significant interest in leveraging modern AI-powered components in real-time SIA research, substantial friction remains due to the absence of a standardized and universal SIA framework. As such, we developed Estuary: a multimodal (text, audio, and soon video) framework which facilitates the development of low-latency, real-time SIAs. Estuary seeks to reduce repeat work between studies and to provide a flexible platform that can be run entirely off-cloud to maximize configurability, controllability, reproducibility of studies, and speed of agent response times. We achieve this by constructing a robust multimodal framework which incorporates current and future components seamlessly into a modular and interoperable architecture.
Spencer Lin, Basem Rizk, Miru Jun, Andy Artze, Caitlin Sullivan, Sharon Mozgai, Scott S. Fisher
IVA6
2023 Toward a Scoping Review of Social Intelligence in Virtual Humans
abstract
As the demand for socially intelligent Virtual Humans (VHs) increases, so follows the demand for effective and efficient cross-discipline collaboration that is required to bring these VHs “to life”. One avenue for increasing cross-discipline fluency is the aggregation and organization of seemingly disparate areas of research and development (e.g., graphics and emotion models) that are essential to the field of VH research. Our initial investigation (1) identifies and catalogues research streams concentrated in three multidisciplinary VH topic clusters within the domain of social intelligence, Emotion, Social Behavior, and The Face, (2) brings to the forefront key themes and prolific authors within each topic cluster, and (3) provides evidence that a full scoping review is warranted to further map the field, aggregate research findings, and identify gaps in the research. To enable collaboration, we provide full access to the refined VH cluster datasets, key word and author word clouds, as well as interactive evidence maps.
Sharon Mozgai, Sarah Beland, Andrew Leeds, Jade G. Winn, Cari Kaurloto, Dirk Heylen, Arno Hartholt
FG1
2022 Re-architecting the virtual human toolkit: towards an interoperable platform for embodied conversational agent research and development
abstract
The research and development (R&D) of intelligent virtual agents (IVAs) is inherently complex. We aim to manage this complexity by combining the best aspects of academic and commercial approaches into a principled R&D platform that emphasizes interoperability, ex-tendability, re-use, and support for multiple hardware targets. This IVA platform, the Virtual Human Toolkit 2.0, is a re-architecture of our earlier work and combines a modular message passing architecture with that of a microservices architecture. This paper discusses our approach, design decisions, lessons learned, and current status of this ongoing effort. We illustrate the strengths of the architecture, how best to use commodity AI cloud services in one's own work, and how to port legacy stand-alone software to a web service.
Arno Hartholt, Edward Fast, Zongjian Li, Kevin Kim, Andrew Leeds, Sharon Mozgai
IVA6
2021 Introducing VHMason: A Visual, Integrated, Multimodal Virtual Human Authoring Tool
abstract
A major impediment to the success of virtual agents is the inability of non-technical experts to easily author content. To address this barrier we present VHMason, a multimodal authoring tool designed to help creative authors build embodied conversational agents. We introduce the novel aspects of this authoring tool and explore a use case of the creation of an agent-led educational experience implemented at Children's Hospital Los Angeles (CHLA).
Arno Hartholt, Edward Fast, Andrew Leeds, Sharon Mozgai
IVA4
2020 Introducing Canvas: Combining Nonverbal Behavior Generation with User-Generated Content to Rapidly Create Educational Videos
abstract
Rapidly creating educational content that is effective, engaging, and low-cost is a challenge. We present Canvas, a tool for educators that addresses this challenge by enabling the generation of educational video, led by an intelligent virtual agent, that combines rapid nonverbal behavior generation techniques with end-user facing authoring tools. With Canvas, educators can easily produce compelling educational videos with a minimum of investment by leveraging existing content provided by the tool (e.g., characters and environments) while incorporating their own custom content (e.g., images and video clips). Canvas has been delivered to the Smithsonian Science Education Center and is currently being evaluated internally before wider release. We discuss the system, feature set, design process, and lessons learned.
Arno Hartholt, Adam Reilly, Edward Fast, Sharon Mozgai
IVA4
2020 The Effects of Experience on Deception in Human-Agent Negotiation
Johnathan Mell, Gale M. Lucas, Sharon Mozgai, Jonathan Gratch
J. Artif. Intell. Res.3
2019 Virtual Humans in Augmented Reality: A First Step towards Real-World Embedded Virtual Roleplayers
abstract
We present one of the first applications of virtual humans in Augmented Reality (AR), which allows young adults with Autism Spectrum Disorder (ASD) the opportunity to practice job interviews. It uses the Magic Leap's AR hardware sensors to provide users with immediate feedback on six different metrics, including eye gaze, blink rate and head orientation. The system provides two characters, with three conversational modes each. Ported from an existing desktop application, the main development lessons learned were: 1) provide users with navigation instructions in the user interface, 2) avoid dark colors as they are rendered transparently, 3) use dynamic gaze so characters maintain eye contact with the user, 4) use hardware sensors like eye gaze to provide user feedback, and 5) use surface detection to place characters dynamically in the world.
Arno Hartholt, Sharon Mozgai, Edward Fast, Matt Liewer, Adam Reilly, Wendy R. Whitcup, Albert A. Rizzo
HAI2
2019 Virtual Job Interviewing Practice for High-Anxiety Populations
abstract
We present a versatile system for training job interviewing skills that focuses specifically on segments of the population facing increased challenges during the job application process. In particular, we target those with Autism Spectrum Disorder (ADS), veterans transitioning to civilian life, and former convicts integrating back into society. The system itself follows the SAIBA framework and contains several interviewer characters, who each represent a different type of vocational field, (e.g. service industry, retail, office, etc.) Each interviewer can be set to one of three conversational modes, which not only affects what they say and how they say it, but also their supporting body language. This approach offers varying difficulties, allowing users to start practicing with interviewers who are more encouraging and accommodating before moving on to personalities that are more direct and indifferent. Finally, the user can place the interviewers in different environmental settings (e.g. conference room, restaurant, executive office, etc.), allowing for many different combinations in which to practice.
Arno Hartholt, Sharon Mozgai, Albert A. Rizzo
IVA2
2018 Towards a Repeated Negotiating Agent that Treats People Individually: Cooperation, Social Value Orientation, & Machiavellianism
abstract
We present the results of a study in which humans negotiate with computerized agents employing varied tactics over a repeated number of economic ultimatum games. We report that certain agents are highly effective against particular classes of humans: several individual difference measures for the human participant are shown to be critical in determining which agents will be successful. Asking for favors works when playing with pro-social people but backfires with more selfish individuals. Further, making poor offers invites punishment from Machiavellian individuals. These factors may be learned once and applied over repeated negotiations, which means user modeling techniques that can detect these differences accurately will be more successful than those that don't. Our work additionally shows that a significant benefit of cooperation is also present in repeated games---after sufficient interaction. These results have deep significance to agent designers who wish to design agents that are effective in negotiating with a broad swath of real human opponents. Furthermore, it demonstrates the effectiveness of techniques which can reason about negotiation over time.
Johnathan Mell, Gale M. Lucas, Sharon Mozgai, Jill Boberg, Ron Artstein, Jonathan Gratch
IVA3
2018 NADiA: Neural Network Driven Virtual Human Conversation Agents
abstract
Advances in artificial intelligence and in particular machine learning and neural networks have given rise to a new generation of virtual assistants and chatbots. Within this work, we present NADiA - Neurally Animated Dialog Agent - that leverages both the user's verbal input as well as their facial expressions to respond in a meaningful way. NADiA combines a neural language model that generates appropriate responses to user prompts, a convolutional neural network for facial expression analysis, and virtual human technology that is deployed on a mobile phone. Here, we evaluate NADiA's anthropomorphic characteristics and its ability to understand the human interlocutor using both subjective as well as objective measures. We find that NADiA significantly outperforms state of the art chatbot technology and produces comparable behavior to human generated reference outputs.
Jason Wu 0001, Sayan Ghosh 0004, Mathieu Chollet, Steven Ly, Sharon Mozgai, Stefan Scherer
IVA5
2017 Manual and automatic measures confirm - Intranasal oxytocin increases facial expressivity
abstract
The effects of oxytocin on facial emotional expressivity were investigated in individuals with schizophrenia and age-matched healthy controls during the completion of a Social Judgment Task (SJT) with a double-blind, placebo-controlled, cross-over design. Although pharmacological interventions exist to help alleviate some symptoms of schizophrenia, currently available agents are not effective at improving the severity of blunted facial affect. Participant facial expressivity was previously quantified from video recordings of the SJT using a well-validated manual approach (Facial Expression Coding System; FACES). We confirm these findings using an automated computer-based approach. Using both methods we found that the administration of oxytocin significantly increased total facial expressivity in individuals with schizophrenia and increased facial expressivity at trend level in healthy controls. Secondary analysis showed that oxytocin also significantly increased the frequency of negative valence facial expressions in individuals with schizophrenia but not in healthy controls and that oxytocin did not significantly increase positive valence facial expressions in either group. Both manual coding and automatic facial analysis revealed the same pattern of findings. Considering manual annotation can be expensive and time-consuming, these results suggest that automatic facial analysis may be an efficient and cost-effective alternative to currently utilized manual approaches and may be ready for use in clinical settings.
Catherine Neubauer, Sharon Mozgai, Brandon Chuang, Joshua Woolley, Stefan Scherer
ACII2
2017 The relationship between task-induced stress, vocal changes, and physiological state during a dyadic team task
abstract
It is commonly known that a relationship exists between the human voice and various emotional states. Past studies have demonstrated changes in a number of vocal features, such as fundamental frequency f0 and peakSlope, as a result of varying emotional state. These voice characteristics have been shown to relate to emotional load, vocal tension, and, in particular, stress. Although much research exists in the domain of voice analysis, few studies have assessed the relationship between stress and changes in the voice during a dyadic team interaction. The aim of the present study was to investigate the multimodal interplay between speech and physiology during a high-workload, high-stress team task. Specifically, we studied task-induced effects on participants' vocal signals, specifically, the f0 and peakSlope features, as well as participants' physiology, through cardiovascular measures. Further, we assessed the relationship between physiological states related to stress and changes in the speaker's voice. We recruited participants with the specific goal of working together to diffuse a simulated bomb. Half of our sample participated in an "Ice Breaker" scenario, during which they were allowed to converse and familiarize themselves with their teammate prior to the task, while the other half of the sample served as our "Control". Fundamental frequency (f0), peakSlope, physiological state, and subjective stress were measured during the task. Results indicated that f0 and peakSlope significantly increased from the beginning to the end of each task trial, and were highest in the last trial, which indicates an increase in emotional load and vocal tension. Finally, cardiovascular measures of stress indicated that the vocal and emotional load of speakers towards the end of the task mirrored a physiological state of psychological "threat".
Catherine Neubauer, Mathieu Chollet, Sharon Mozgai, Mark Dennison, Peter Khooshabeh, Stefan Scherer
ICMI3
2017 To Tell the Truth: Virtual Agents and Morning Morality
Sharon Mozgai, Gale M. Lucas, Jonathan Gratch
IVA1