Zubin Datta Choudhary

dblp:265/2251 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0003-4303-5759ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Teleportation Destination Previews Support Memory Retention During Virtual Navigation
Zubin Datta Choudhary, Ferran Argelaguet, Gerd Bruder, Greg Welch
VR1
2024 Investigating the relationships between user behaviors and tracking factors on task performance and trust in augmented reality
Matthew Gottsacker, Hiroshi Furuya, Zubin Datta Choudhary, Austin Erickson, Ryan Schubert, Gerd Bruder, Michael P. Browne, Greg Welch
Comput. Graph.3
2023 Exploring the Social Influence of Virtual Humans Unintentionally Conveying Conflicting Emotions
abstract
The expression of human emotion is integral to social interaction, and in virtual reality it is increasingly common to develop virtual avatars that attempt to convey emotions by mimicking these visual and aural cues, i.e. the facial and vocal expressions. However, errors in (or the absence of) facial tracking can result in the rendering of incorrect facial expressions on these virtual avatars. For example, a virtual avatar may speak with a happy or unhappy vocal inflection while their facial expression remains otherwise neutral. In circumstances where there is conflict between the avatar's facial and vocal expressions, it is possible that users will incorrectly interpret the avatar's emotion, which may have unintended consequences in terms of social influence or in terms of the outcome of the interaction. In this paper, we present a human-subjects study (N = 22) aimed at understanding the impact of conflicting facial and vocal emotional expressions. Specifically we explored three levels of emotional valence (unhappy, neutral, and happy) expressed in both visual (facial) and aural (vocal) forms. We also investigate three levels of head scales (down-scaled, accurate, and up-scaled) to evaluate whether head scale affects user interpretation of the conveyed emotion. We find significant effects of different multimodal expressions on happiness and trust perception, while no significant effect was observed for head scales. Evidence from our results suggest that facial expressions have a stronger impact than vocal expressions. Additionally, as the difference between the two expressions increase, the less predictable the multimodal expression becomes. For example, for the happy-looking and happy-sounding multimodal expression, we expect and see high happiness rating and high trust, however if one of the two expressions change, this mismatch makes the expression less predictable. We discuss the relationships, implications, and guidelines for social applications that aim to leverage multimodal social cues.
Zubin Datta Choudhary, Nahal Norouzi, Austin Erickson, Ryan Schubert, Gerd Bruder, Greg Welch
VR1
2023 Visual Hearing Aids: Artificial Visual Speech Stimuli for Audiovisual Speech Perception in Noise
abstract
Speech perception is optimal in quiet environments, but noise can impair comprehension and increase errors. In these situations, lip reading can help, but it is not always possible, such as during an audio call or when wearing a face mask. One approach to improve speech perception in these situations is to use an artificial visual lip reading aid. In this paper, we present a user study (N = 17) in which we compared three levels of audio stimuli visualizations and two levels of modulating the appearance of the visualization based on the speech signal, and we compared them against two control conditions: an audio-only condition, and a real human speaking. We measured participants’ speech reception thresholds (SRTs) to understand the effects of these visualizations on speech perception in noise. These thresholds indicate the decibel levels of the speech signal that are necessary for a listener to receive the speech correctly 50% of the time. Additionally, we measured the usability of the approaches and the user experience. We found that the different artificial visualizations improved participants’ speech reception compared to the audio-only baseline condition, but they were significantly poorer than the real human condition. This suggests that different visualizations can improve speech perception when the speaker’s face is not available. However, we also discuss limitations of current plug-and-play lip sync software and abstract representations of the speaker in the context of speech perception.
Zubin Datta Choudhary, Gerd Bruder, Greg Welch
VRST1
2023 Virtual Big Heads in Extended Reality: Estimation of Ideal Head Scales and Perceptual Thresholds for Comfort and Facial Cues
abstract
Extended reality (XR) technologies, such as virtual reality (VR) and augmented reality (AR), provide users, their avatars, and embodied agents a shared platform to collaborate in a spatial context. Although traditional face-to-face communication is limited by users’ proximity, meaning that another human’s non-verbal embodied cues become more difficult to perceive the farther one is away from that person, researchers and practitioners have started to look into ways to accentuate or amplify such embodied cues and signals to counteract the effects of distance with XR technologies. In this article, we describe and evaluate the Big Head technique, in which a human’s head in VR/AR is scaled up relative to their distance from the observer as a mechanism for enhancing the visibility of non-verbal facial cues, such as facial expressions or eye gaze. To better understand and explore this technique, we present two complimentary human-subject experiments in this article. In our first experiment, we conducted a VR study with a head-mounted display to understand the impact of increased or decreased head scales on participants’ ability to perceive facial expressions as well as their sense of comfort and feeling of “uncannniness” over distances of up to 10 m. We explored two different scaling methods and compared perceptual thresholds and user preferences. Our second experiment was performed in an outdoor AR environment with an optical see-through head-mounted display. Participants were asked to estimate facial expressions and eye gaze, and identify a virtual human over large distances of 30, 60, and 90 m. In both experiments, our results show significant differences in minimum, maximum, and ideal head scales for different distances and tasks related to perceiving faces, facial expressions, and eye gaze, and we also found that participants were more comfortable with slightly bigger heads at larger distances. We discuss our findings with respect to the technologies used, and we discuss implications and guidelines for practical applications that aim to leverage XR-enhanced facial cues.
Zubin Datta Choudhary, Austin Erickson, Nahal Norouzi, Kangsoo Kim, Gerd Bruder, Greg Welch
ACM Trans. Appl. Percept.1
2023 Visual Facial Enhancements Can Significantly Improve Speech Perception in the Presence of Noise
abstract
Human speech perception is generally optimal in quiet environments, however it becomes more difficult and error prone in the presence of noise, such as other humans speaking nearby or ambient noise. In such situations, human speech perception is improved by speech reading, i.e., watching the movements of a speaker's mouth and face, either consciously as done by people with hearing loss or subconsciously by other humans. While previous work focused largely on speech perception of two-dimensional videos of faces, there is a gap in the research field focusing on facial features as seen in head-mounted displays, including the impacts of display resolution, and the effectiveness of visually enhancing a virtual human face on speech perception in the presence of noise. In this paper, we present a comparative user study ( N=21) in which we investigated an audio-only condition compared to two levels of head-mounted display resolution ( 1832×1920 or 916×960 pixels per eye) and two levels of the native or visually enhanced appearance of a virtual human, the latter consisting of an up-scaled facial representation and simulated lipstick (lip coloring) added to increase contrast. To understand effects on speech perception in noise, we measured participants' speech reception thresholds (SRTs) for each audio-visual stimulus condition. These thresholds indicate the decibel levels of the speech signal that are necessary for a listener to receive the speech correctly 50% of the time. First, we show that the display resolution significantly affected participants' ability to perceive the speech signal in noise, which has practical implications for the field, especially in social virtual environments. Second, we show that our visual enhancement method was able to compensate for limited display resolution and was generally preferred by participants. Specifically, our participants indicated that they benefited from the head scaling more than the added facial contrast from the simulated lipstick. We discuss relationships, implications, and guidelines for applications that aim to leverage such enhancements.
Zubin Datta Choudhary, Gerd Bruder, Greg Welch
IEEE Trans. Vis. Comput. Graph.1
2022 The Effects of an Embodied Pedagogical Agent's Synthetic Speech Accent on Learning Outcomes
abstract
Modern text-to-speech engines can be an effective speech choice for embodied virtual pedagogical agents. However, it is not known how synthesized accents influence learning outcomes and perceptions of the agent. In this paper, we conducted a between-subjects experiment (n=60) to determine the effect of a pedagogical agent’s machine synthesized text-to-speech accent (United States English or Indian English) on learning outcomes and perceptions of the agent for students in the United States. Our results indicate that learner gender interacts with synthesized speech accent to significantly affect learning outcomes and perceptions of the agent. Our results reveal that a foreign synthetic speech accent may affect the learning outcomes of female university students (n=30), but not male university students (n=30). Finally, our results indicate that learner gender interacts with synthesized speech accent to affect perceptions of the pedagogical agent’s human-likeness. We provide novel insights on the differences between male and female learners for interactions with pedagogical agents with synthetic TTS accents.
Tiffany D. Do, Mamtaj Akter, Zubin Datta Choudhary, Roger Azevedo, Ryan P. McMahan
ICMI3
2021 Revisiting Distance Perception with Scaled Embodied Cues in Social Virtual Reality
abstract
Previous research on distance estimation in virtual reality (VR) has well established that even for geometrically accurate virtual objects and environments users tend to systematically mis-estimate distances. This has implications for Social VR, where it introduces variables in personal space and proxemics behavior that change social behaviors compared to the real world. One yet unexplored factor is related to the trend that avatars' embodied cues in Social VR are often scaled, e.g., by making one's head bigger or one's voice louder, to make social cues more pronounced over longer distances. In this paper we investigate how the perception of avatar distance is changed based on two means for scaling embodied social cues: visual head scale and verbal volume scale. We conducted a human-subject study employing a mixed factorial design with two Social VR avatar representations (full-body, head-only) as a between factor as well as three visual head scales and three verbal volume scales (up-scaled, accurate, down-scaled) as within factors. For three distances from social to far-public space, we found that visual head scale had a significant effect on distance judgments and should be tuned for Social VR, while conflicting verbal volume scales did not, indicating that voices can be scaled in Social VR without immediate repercussions on spatial estimates. We discuss the interactions between the factors and implications for Social VR.
Zubin Datta Choudhary, Matthew Gottsacker, Kangsoo Kim, Ryan Schubert, Jeanine K. Stefanucci, Gerd Bruder, Greg Welch
VR1
2020 Virtual Big Heads: Analysis of Human Perception and Comfort of Head Scales in Social Virtual Reality
abstract
Virtual reality (VR) technologies provide a shared platform for collaboration among users in a spatial context. To enhance the quality of social signals during interaction between users, researchers and practitioners started augmenting users’ interpersonal space with different types of virtual embodied social cues. A prominent example is commonly referred to as the "Big Head" technique, in which the head scales of virtual interlocutors are slightly increased to leverage more of the display’s visual space to convey facial social cues. While beneficial in improving interpersonal social communication, the benefits and thresholds of human perception of facial cues and comfort in such Big Head environments are not well understood, limiting their usefulness and subjective experience.In this paper, we present a human-subject study that we conducted to understand the impact of an increased or decreased head scale in social VR on participants’ ability to perceive facial expressions as well as their sense of comfort and feeling of "uncanniness." We explored two head scaling methods and compared them with respect to perceptual thresholds and user preferences. We further show that the distance to interlocutors has an important effect on the results. We discuss implications and guidelines for practical applications that aim to leverage VR-enhanced social cues.
Zubin Datta Choudhary, Kangsoo Kim, Ryan Schubert, Gerd Bruder, Greg Welch
VR1