Dhruv Jain

dblp:126/2498 · DBLP profile ↗
← Back
40ranked-venue papers
15as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 37 · 15 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People
abstract
As virtual 3D environments become more prevalent, equitable access is essential for blind and low-vision (BLV) users, who face challenges with spatial awareness, navigation, and interaction. Prior work has explored supplementing visual information with auditory or haptic modalities, but these methods are static and offer limited support for dynamic, in-context adaptation. Recent advances in generative AI allow users to query and modify 3D scenes via natural language, introducing a paradigm that offers greater flexibility and control for accessibility. We present RAVEN, a system that enables BLV users to issue queries and modification prompts to improve the runtime accessibility of 3D virtual scenes. We evaluated RAVEN with eight BLV people and six Unity developers, generating empirical insights into how conversational programming can support personalized accessibility in 3D environments. Our work highlights both the promise of natural language interaction—intuitive, flexible, and empowering—and the challenges of ensuring reliability, transparency, and trust in generative AI–driven accessibility systems.
Xinyun Cao, Kexin Ju 0001, Venkatesh Potluri, Dhruv Jain
CHI5
2025 Demo of RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People
abstract
Figure 1: RAVEN is an interactive system that empowers BLV users to query and modify 3D scenes via natural language.The above image illustrates an example of an accessibility modification: A) A low-vision user types in a modification text.B) The system integrates runtime code generation LLM agent with dynamic scene information and instructions to apply accessibilityenhancing changes at runtime.C) The system compiles LLM-produced code to achieve modification while providing spoken response to the user.
Xinyun Cao, Kexin Ju 0001, Venkatesh Potluri, Dhruv Jain
ASSETS5
2025 CapTune: Adapting Non-Speech Captions With Anchored Generative Models
abstract
Non-speech captions are essential to the video experience of deaf and hard of hearing (DHH) viewers, yet conventional approaches often overlook the diversity of their preferences. We present CapTune, a system that enables customization of non-speech captions based on DHH viewers' needs while preserving creator intent. CapTune allows caption authors to define safe transformation spaces using concrete examples and empowers viewers to personalize captions across four dimensions: level of detail, expressiveness, sound representation method, and genre alignment. Evaluations with seven caption creators and twelve DHH participants showed that CapTune supported creators' creative control while enhancing viewers' emotional engagement with content. Our findings also reveal trade-offs between information richness and cognitive load, tensions between interpretive and descriptive representations of sound, and the context-dependent nature of caption preferences.
Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu, Dhruv Jain
ASSETS4
2025 Demo of CapTune: Adapting Non-Speech Captions with Anchored Generative Models
Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu, Dhruv Jain
ASSETS4
2025 SoundNarratives: Rich Auditory Scene Descriptions to Support Deaf and Hard of Hearing People
abstract
state-of-the-art audio language model.A user study with 10 DHH participants demonstrated a significant preference for SoundNarratives over a baseline model, along with a potential for improved confidence and situational awareness.
Liang-Yuan Wu, Dhruv Jain
ASSETS2
2025 EvolveCaptions: Real-Time Collaborative ASR Adaptation for DHH Speakers
abstract
Figure 1: Overview of EvolveCaptions.(1) Hearing users correct live captions of the DHH speaker's voice.(2) The DHH speaker records targeted phrases generated from the corrected terms.(3) The Whisper ASR model is fine-tuned with the recordings and adapts to the speaker over time.
Liang-Yuan Wu, Dhruv Jain
ASSETS2
2025 CARTGPT: Real-Time Correction of CART Captions Using Large Language Models
abstract
Communication Access Realtime Translation (CART) is a widely used captioning technology among deaf and hard of hearing (DHH) individuals, valued for its high accuracy and ability to convey speaker cues and contextual sounds in real time. However, CART performance can degrade in challenging conditions such as background noise, technical jargon, or rapid speech—reducing caption quality and impacting comprehension. We introduce CARTGPT, a real-time captioning system that enhances CART transcripts by leveraging large language models (LLMs) and automatic speech recognition (ASR) input to detect and correct transcription errors. To inform the design of CARTGPT, we conducted a formative study with 10 professional CART captioners to identify common sources of error and their perspectives on using AI for caption correction. We evaluated CARTGPT on a 39.7-hour speech dataset spanning medical, technical, and conversational domains, observing a 5.6% improvement in word accuracy over standard CART and 17.3% over a state-of-the-art ASR model. In a user study with 16 DHH participants, CARTGPT captions were rated as significantly more comprehensible, particularly in technical scenarios, while maintaining real-time responsiveness. These findings demonstrate the potential of LLM-assisted captioning to improve accessibility and comprehension for DHH users in real-world settings.
Liang-Yuan Wu, Andrea Kleiver, Dhruv Jain
ASSETS3
2025 Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User
abstract
Current AI sound awareness systems can provide deaf and hard of hearing people with information about sounds, including discrete sound sources and transcriptions. However, synthesizing AI outputs based on DHH people's ever-changing intents in complex auditory environments remains a challenge. In this paper, we describe the co-design process of SoundWeaver, a sound awareness system prototype that dynamically weaves AI outputs from different AI models based on users’ intents and presents synthesized information through a heads-up display. Adopting a Research through Design perspective, we created SoundWeaver with one DHH co-designer, adapting it to his personal contexts and goals (e.g., cooking at home and chatting in a game store). Through this process, we present design implications for the future of “intent-driven” AI systems for sound accessibility.
Jeremy Zhengqi Huang, Jaylin Herskovitz, Liang-Yuan Wu, Cecily Morrison, Dhruv Jain
CHI5
2025 Towards Audio Personalization for Accessible Digital Media
Dhruv Jain
ICMI1
2024 SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
abstract
Mixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually impaired individuals, who rely on sounds and descriptions in their everyday lives. To understand how complex audio information is consumed, we analyzed online forum posts within the blind community, identifying prevailing challenges, needs, and desired solutions. We synthesized the results and propose SoundShift for increasing MR sound awareness, which includes six sound manipulations: Transparency Shift, Envelope Shift, Position Shift, Style Shift, Time Shift, and Sound Append. To evaluate the effectiveness of SoundShift, we conducted a user study with 18 blind participants across three simulated MR scenarios, where participants identified specific sounds within intricate soundscapes. We found that SoundShift increased MR sound awareness and minimized cognitive load. Finally, we developed three real-world example applications to demonstrate the practicality of SoundShift.
Ruei-Che Chang, Chia-Sheng Hung, Bing-Yu Chen 0004, Dhruv Jain, Anhong Guo
Conference on Designing Interactive Systems4
2024 SoundModVR: Sound Modifications in Virtual Reality to Support People who are Deaf and Hard of Hearing
abstract
Previous VR sound accessibility work substituted sounds with visual or haptic output to increase VR accessibility for deaf and hard of hearing (DHH) people. However, deafness occurs on a spectrum, and many DHH people (e.g., those with partial hearing) can also benefit greatly from having more control over the audio instead of substituting it with another modality. In this paper, we explore the possibilities of modifying sounds in VR to support DHH people. To understand the best modification features for this goal, we designed and implemented 18 VR sound modification tools spanning four categories, including prioritizing sounds, modifying sound parameters, providing spatial assistance, and adding additional sounds. We evaluated our tools in five diverse VR scenarios with 10 DHH people, finding that our tool can improve DHH users’ VR experience, but could be further improved by providing more customization options and decreasing distraction. We then compiled a Unity toolkit from select tools and conducted a preliminary evaluation with six Unity VR developers. Findings show that our toolkit is easy to use and debug but could be enhanced through modularization and better documentation. We close by discussing further implications of sound modification in VR.
Xinyun Cao, Dhruv Jain
ASSETS2
2024 Supporting Sound Accessibility by Exploring Sound Augmentations in Virtual Reality
abstract
To increase VR sound accessibility for deaf and hard of hearing users, previous work has substituted sounds with visual or haptic feedback. However, many DHH people (e.g., those with partial hearing) can also benefit from modifying audio (e.g., changing volume based on priorities) instead of fully substituting it with another modality. In this demo paper, we present a toolkit that allows modifying sounds in VR to support DHH people. We designed and implemented 18 VR sound modification tools spanning four categories, including prioritizing sounds, modifying sound parameters, providing spatial assistance, and adding additional sounds. We present five demo scenarios with tools incorporated, covering common VR use cases.
Xinyun Cao, Dhruv Jain
ASSETS2
2024 MaskSound: Exploring Sound Masking Approaches to Support People with Autism in Managing Noise Sensitivity
abstract
Noise sensitivity is a frequently reported characteristic in many autistic individuals. While strategies like sound isolation (e.g., noise-canceling headphones) and avoidance behaviors (e.g., leaving a crowded room) can help, they can reduce situational awareness and limit social engagement. In this paper, we examine an alternate approach to managing noise sensitivity: introducing ambient background sounds to reduce the perception of disruptive noises, i.e., sound masking. Through two studies (with ten and nine autistic individuals respectively), we investigated the autistic individuals’ preferred sound masks (e.g., white noise, brown noise, calming water sounds) for different contexts (e.g., traffic, speech) and elicited reactions for a future interactive tool to deliver effective sound masks. Our findings have implications not just for the accessibility community, but also for designers and researchers working on sound augmentation technology.
Anna Y. Park, Andy Jin, Jeremy Zhengqi Huang, Jesse Carr, Dhruv Jain
ASSETS5
2024 CARTGPT: Improving CART Captioning using Large Language Models
abstract
Communication Access Realtime Translation (CART) is a commonly used real-time captioning technology used by deaf and hard of hearing (DHH) people, due to its accuracy, reliability, and ability to provide a holistic view of the conversational environment (e.g., by displaying speaker names). However, in many real-world situations (e.g., noisy environments, long meetings), the CART captioning accuracy can considerably decline, thereby affecting the comprehension of DHH people. In this work-in-progress paper, we introduce CARTGPT, a system to assist CART captioners in improving their transcription accuracy. CARTGPT takes in errored CART captions and inaccurate automatic speech recognition (ASR) captions as input and uses a large language model to generate corrected captions in real-time. We quantified performance on a noisy speech dataset, showing that our system outperforms both CART (+5.6% accuracy) and a state-of-the-art ASR model (+17.3%). A preliminary evaluation with three DHH users further demonstrates the promise of our approach.
Liang-Yuan Wu, Andrea Kleiver, Dhruv Jain
ASSETS3
2024 A Human-AI Collaborative Approach for Designing Sound Awareness Systems
abstract
Current sound recognition systems for deaf and hard of hearing (DHH) people identify sound sources or discrete events. However, these systems do not distinguish similar sounding events (e.g., a patient monitor beep vs. a microwave beep). In this paper, we introduce HACS, a novel futuristic approach to designing human-AI sound awareness systems. HACS assigns AI models to identify sounds based on their characteristics (e.g., a beep) and prompts DHH users to use this information and their contextual knowledge (e.g., “I am in a kitchen”) to recognize sound events (e.g., a microwave). As a first step for implementing HACS, we articulated a sound taxonomy that classifies sounds based on sound characteristics using insights from a multi-phased research process with people of mixed hearing abilities. We then performed a qualitative (with 9 DHH people) and a quantitative (with a sound recognition model) evaluation. Findings demonstrate the initial promise of HACS for designing accurate and reliable human-AI systems.
Jeremy Zhengqi Huang, Reyna Wood, Hriday Chhabria, Dhruv Jain
CHI4
2024 Hunting imaging biomarkers in pulmonary fibrosis: Benchmarks of the AIIB23 challenge
abstract
• This paper investigates the capacity of AI models for airway modelling on national datasets with paired clinical metadata. • We evaluated AI models against unharmonised, noisy, and out-of-distribution data, as well as the prognostication for FLD. • We found a new biomarker for mortality prediction, outperforming existing clinical measurements (FVC% and fibrosis scores). • In-depth analysis of AI models on airway modelling and prognosis, highlighting challenges and future research directions. Airway-related quantitative imaging biomarkers are crucial for examination, diagnosis, and prognosis in pulmonary diseases. However, the manual delineation of airway structures remains prohibitively time-consuming. While significant efforts have been made towards enhancing automatic airway modelling, current public-available datasets predominantly concentrate on lung diseases with moderate morphological variations. The intricate honeycombing patterns present in the lung tissues of fibrotic lung disease patients exacerbate the challenges, often leading to various prediction errors. To address this issue, the 'Airway-Informed Quantitative CT Imaging Biomarker for Fibrotic Lung Disease 2023′ (AIIB23) competition was organized in conjunction with the official 2023 International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI). The airway structures were meticulously annotated by three experienced radiologists. Competitors were encouraged to develop automatic airway segmentation models with high robustness and generalization abilities, followed by exploring the most correlated QIB of mortality prediction. A training set of 120 high-resolution computerised tomography (HRCT) scans were publicly released with expert annotations and mortality status. The online validation set incorporated 52 HRCT scans from patients with fibrotic lung disease and the offline test set included 140 cases from fibrosis and COVID-19 patients. The results have shown that the capacity of extracting airway trees from patients with fibrotic lung disease could be enhanced by introducing voxel-wise weighted general union loss and continuity loss. In addition to the competitive image biomarkers for mortality prediction, a strong airway-derived biomarker (Hazard ratio>1.5, p < 0.0001) was revealed for survival prognostication compared with existing clinical measurements, clinician assessment and AI-based biomarkers.
Yang Nan 0002, Xiaodan Xing, Zeyu Tang 0001, Federico Felder, Sheng Zhang 0024, Roberta Eufrasia Ledda, Xiaoliu Ding, Feng Shi 0001, Tianyang Sun, Zehong Cao, Yun Gu, Pingyu Wang, Wen Tang 0005, Pengxin Yu, Han Kang, Junqiang Chen, Michail Mamalakis, Francesco Prinzi, Gianluca Carlini, Lisa Cuneo, Abhirup Banerjee, Zhaohu Xing, Lei Zhu 0003, Zacharia Mesbah, Dhruv Jain, Tsiry Mayet, Hongyu Yuan, Qing Lyu 0009, Abdul Qayyum 0002, Moona Mazher, Athol Wells, Simon Walsh, Guang Yang 0006
Medical Image Anal.33
2023 AdaptiveSound: An Interactive Feedback-Loop System to Improve Sound Recognition for Deaf and Hard of Hearing Users
abstract
Sound recognition tools have wide-ranging impacts for deaf and hard of hearing (DHH) people from being informed of safety-critical information (e.g., fire alarms, sirens) to more mundane but still useful information (e.g., door knock, microwave beeps). However, prior sound recognition systems use models that are pre-trained on generic sound datasets and do not adapt well to diverse variations of real-world sounds. We introduce AdaptiveSound, a real-time system for portable devices (e.g., smartphones) that allows DHH users to provide corrective feedback to the sound recognition model to adapt the model to diverse acoustic environments. AdaptiveSound is informed by prior surveys of sound recognition systems, where DHH users strongly desired the ability to provide feedback to a pre-trained sound recognition model to fine-tune it to their environments. Through quantitative experiments and field evaluations with 12 DHH users, we show that AdaptiveSound can achieve a significantly higher accuracy (+14.6%) than prior state-of-the art systems in diverse real-world locations (e.g., homes, parks, streets, and malls) with little end-user effort (about 10 minutes of feedback).
Hang Do, Quan Dang, Jeremy Zhengqi Huang, Dhruv Jain
ASSETS4
2023 "Not There Yet": Feasibility and Challenges of Mobile Sound Recognition to Support Deaf and Hard-of-Hearing People
abstract
While recent advances have enabled mobile sound recognition tools for deaf and hard of hearing (DHH) people, these tools have only been studied in the lab or through short, controlled experiments. To assess the real-world feasibility and guide the future designs of mobile sound awareness systems, we conducted a three-week field study of SoundWatch, a smartwatch-based sound recognition app, with 10 DHH participants. Our findings suggest the app's utility in increasing environmental awareness and facilitating everyday tasks for DHH users. However, several challenges, such as background noises, variability of real-world sounds, and confusion among similar sounding sounds, indicated that mobile sound recognition solutions are “not there yet” for adoption and use in daily life. We close by presenting HCI design opportunities to improve model reliability by increasing contextual awareness, supporting end-user customization, and fostering the collective improvement of sound recognition models.
Jeremy Zhengqi Huang, Hriday Chhabria, Dhruv Jain
ASSETS3
2022 A Workshop on Disability Inclusive Remote Co-Design
abstract
The COVID-19 pandemic forced researchers to find new ways to continue research, as universities and laboratories experienced closure due to nationwide lockdowns in many countries worldwide, including conducting experiments, workshops, and ethnographic work online. While this had a significant impact on the majority of research work across SIGCHI, research relating to disability and ageing was most impacted due to the additional challenges of recruiting participants, finding accessible online platforms, and ensuring seamless participation while juggling platform accessibility issues, facilitation, and supporting participants’ needs. These challenges were more extreme for disabled researchers. In this workshop, we aim to bring together researchers, designers, and practitioners to explore effective strategies and brainstorm actionable guidelines for supporting disability inclusive online research methods and platforms.
Maryam Bandukda, Giulia Barbareschi, Aneesha Singh, Dhruv Jain, Maitraye Das, Tamanna Motahar, Jason Wiese, Lynn Cockburn, David M. Frohlich, Catherine Holloway
ASSETS4
2022 ProtoSound: A Personalized and Scalable Sound Recognition System for Deaf and Hard-of-Hearing Users
abstract
Recent advances have enabled automatic sound recognition systems for deaf and hard of hearing (DHH) users on mobile devices. However, these tools use pre-trained, generic sound recognition models, which do not meet the diverse needs of DHH users. We introduce ProtoSound, an interactive system for customizing sound recognition models by recording a few examples, thereby enabling personalized and fine-grained categories. ProtoSound is motivated by prior work examining sound awareness needs of DHH people and by a survey we conducted with 472 DHH participants. To evaluate ProtoSound, we characterized performance on two real-world sound datasets, showing significant improvement over state-of-the-art (e.g., +9.7% accuracy on the first dataset). We then deployed ProtoSound's end-user training and real-time recognition through a mobile application and recruited 19 hearing participants who listened to the real-world sounds and rated the accuracy across 56 locations (e.g., homes, restaurants, parks). Results show that ProtoSound personalized the model on-device in real-time and accurately learned sounds across diverse acoustic contexts. We close by discussing open challenges in personalizable sound recognition, including the need for better recording interfaces and algorithmic improvements.
Dhruv Jain, Khoa Huynh Anh Nguyen, Steven M. Goodman, Rachel Grossman-Kahn, Hung Ngo, Aditya Kusupati, Ruofei Du, Alex Olwal, Leah Findlater, Jon Froehlich
CHI1
2022 Nonverbal Sound Detection for Disordered Speech
abstract
Voice assistants have become an essential tool for people with various disabilities because they enable complex phone-or tablet-based interactions without the need for fine-grained motor control, such as with touchscreens. However, these systems are not tuned for the unique characteristics of individuals with speech disorders, including many of those who have a motor-speech disorder, are deaf or hard of hearing, have a severe stutter, or are minimally verbal. We introduce an alternative voice-based input system which relies on sound event detection using fifteen nonverbal mouth sounds like "pop", "click", or "eh." This system was designed to work regardless of ones’ speech abilities and allows full access to existing technology. In this paper, we describe the design of a dataset, model considerations for real-world deployment, and efforts towards model personalization. Our fully-supervised model achieves segment-level precision and recall of 88.6% and 88.4% on an internal dataset of 710 adults, while achieving 0.31 false positives per hour on aggressors such as speech. Five-shot personalization enables satisfactory performance in 84.5% of cases where the generic model fails.
Colin Lea, Zifang Huang, Dhruv Jain, Lauren Tooley, Zeinab Liaghat, Shrinath Thelapurath, Leah Findlater, Jeffrey P. Bigham
ICASSP3
2021 A Taxonomy of Sounds in Virtual Reality
abstract
Virtual reality (VR) leverages human sight, hearing and touch senses to convey virtual experiences. For d/Deaf and hard of hearing (DHH) people, information conveyed through sound may not be accessible. To help with future design of accessible VR sound representations for DHH users, this paper contributes a consistent language and structure for representing sounds in VR. Using two studies, we report on the design and evaluation of a novel taxonomy for VR sounds. Study 1 included interviews with 10 VR sound designers to develop our taxonomy along two dimensions: sound source and intent. To evaluate this taxonomy, we conducted another study (Study 2) where eight HCI researchers used our taxonomy to document sounds in 33 VR apps. We found that our taxonomy was able to successfully categorize nearly all sounds (265/267) in these apps. We also uncovered additional insights for designing accessible visual and haptic-based sound substitutes for DHH users.
Dhruv Jain, Sasa Junuzovic, Eyal Ofek, Mike Sinclair, John R. Porter, Chris Yoon, Swetha Machanavajhala, Meredith Ringel Morris
Conference on Designing Interactive Systems1
2021 Mixed Abilities and Varied Experiences: a group autoethnography of a virtual summer internship
abstract
The COVID-19 pandemic forced many people to convert their daily work lives to a “virtual” format where everyone connected remotely from their home. In this new, virtual environment, accessibility barriers changed, in some respects for the better (e.g., more flexibility) and in other aspects, for the worse (e.g., problems including American Sign Language interpreters over video calls). Microsoft Research held its first cohort of all virtual interns in 2020. We the authors, full time and intern members and affiliates of the Ability Team, a research team focused on accessibility, reflect on our virtual work experiences as a team consisting of members with a variety of abilities, positions, and seniority during the summer intern season. Through our autoethnographic method, we provide a nuanced view into the experiences of a mixed-ability, virtual team, and how the virtual setting affected the team’s accessibility. We then reflect on these experiences, noting the successful strategies we used to promote access and the areas in which we could have further improved access. Finally, we present guidelines for future virtual mixed-ability teams looking to improve access.
Kelly Mack, Maitraye Das, Dhruv Jain, Danielle Bragg, John C. Tang, Andrew Begel, Erin Beneteau, Josh Urban Davis, Abraham Glasser, Joon Sung Park 0001, Venkatesh Potluri
ASSETS3
2021 What Do We Mean by "Accessibility Research"?: A Literature Survey of Accessibility Papers in CHI and ASSETS from 1994 to 2019
abstract
Accessibility research has grown substantially in the past few decades, yet there has been no literature review of the field. To understand current and historical trends, we created and analyzed a dataset of accessibility papers appearing at CHI and ASSETS since ASSETS' founding in 1994. We qualitatively coded areas of focus and methodological decisions for the past 10 years (2010-2019, N=506 papers), and analyzed paper counts and keywords over the full 26 years (N=836 papers). Our findings highlight areas that have received disproportionate attention and those that are underserved--for example, over 43% of papers in the past 10 years are on accessibility for blind and low vision people. We also capture common study characteristics, such as the roles of disabled and nondisabled participants as well as sample sizes (e.g., a median of 13 for participant groups with disabilities and older adults). We close by critically reflecting on gaps in the literature and offering guidance for future work in the field.
Kelly Mack, Emma McDonnell, Dhruv Jain, Lucy Lu Wang, Jon Froehlich, Leah Findlater
CHI3
2021 Towards Sound Accessibility in Virtual Reality
abstract
Virtual reality (VR) leverages sight, hearing, and touch senses to convey virtual experiences. For d/Deaf and hard of hearing (DHH) people, however, information conveyed through sound may not be accessible. While prior work has explored making every day sounds accessible to DHH users, the context of VR is, as yet, unexplored. In this paper, we provide a first comprehensive investigation of sound accessibility in VR. Our primary contributions include a design space for developing visual and haptic substitutes of VR sounds to support DHH users and prototypes illustrating several points within the design space. We also characterize sound accessibility in commonly used VR apps and discuss findings from early evaluations of our prototypes with 11 DHH users and 4 VR developers.
Dhruv Jain, Sasa Junuzovic, Eyal Ofek, Mike Sinclair, John R. Porter, Chris Yoon, Swetha Machanavajhala, Meredith Ringel Morris
ICMI1
2020 HoloSound: Combining Speech and Sound Identification for Deaf or Hard of Hearing Users on a Head-mounted Display
abstract
Head-mounted displays can provide private and glanceable speech and sound feedback to deaf and hard of hearing people, yet prior systems have largely focused on speech transcription. We introduce HoloSound, a HoloLens-based augmented reality (AR) prototype that uses deep learning to classify and visualize sound identity and location in addition to providing speech transcription. This poster paper presents a working proof-of-concept prototype, and discusses future opportunities for advancing AR-based sound awareness.
Ru Guo, Yiru Yang, Johnson Kuang, Xue Bin, Dhruv Jain, Steven M. Goodman, Leah Findlater, Jon Froehlich
ASSETS5
2020 SoundWatch: Exploring Smartwatch-based Deep Learning Approaches to Support Sound Awareness for Deaf and Hard of Hearing Users
abstract
Smartwatches have the potential to provide glanceable, always-available sound feedback to people who are deaf or hard of hearing. In this paper, we present a performance evaluation of four low-resource deep learning sound classification models: MobileNet, Inception, ResNet-lite, and VGG-lite across four device architectures: watch-only, watch+phone, watch+phone+cloud, and watch+cloud. While direct comparison with prior work is challenging, our results show that the best model, VGG-lite, performed similar to the state of the art for non-portable devices with an average accuracy of 81.2% (SD=5.8%) across 20 sound classes and 97.6% (SD=1.7%) across the three highest-priority sounds. For device architectures, we found that the watch+phone architecture provided the best balance between CPU, memory, network usage, and classification latency. Based on these experimental results, we built and conducted a qualitative lab evaluation of a smartwatch-based sound awareness app, called SoundWatch (Figure 1), with eight DHH participants. Qualitative findings show support for our sound awareness app but also uncover issues with misclassifications, latency, and privacy concerns. We close by offering design considerations for future wearable sound awareness technology.
Dhruv Jain, Hung Ngo, Pratyush Patel, Steven M. Goodman, Leah Findlater, Jon Froehlich
ASSETS1
2020 Navigating Graduate School with a Disability
abstract
In graduate school, people with disabilities use disability accommodations to learn, network, and do research. However, these accommodations, often scheduled ahead of time, may not work in many situations due to uncertainty and spontaneity of the graduate experience. Through a three-person autoethnography, we present a longitudinal account of our graduate school experiences as people with disabilities, highlighting nuances and tensions of situations when our requested accommodations did not work and the use of alternative coping strategies. We use retrospective journals and field notes to reveal the impact of our self-image, relationships, technologies, and infrastructure on our disabled experience. Using post-hoc reflection on our experiences, we then close with discussing personal and situated ways in which peers, faculty members, universities, and technology designers could improve the graduate school experiences of people with disabilities.
Dhruv Jain, Venkatesh Potluri, Ather Sharif
ASSETS1
2020 Evaluating Smartwatch-based Sound Feedback for Deaf and Hard-of-hearing Users Across Contexts
abstract
We present a qualitative study with 16 deaf and hard of hearing (DHH) participants examining reactions to smartwatch-based visual + haptic sound feedback designs. In Part 1, we conducted a Wizard-of-Oz (WoZ) evaluation of three smartwatch feedback techniques (visual alone, visual + simple vibration, and visual + tacton) and investigated vibrational patterns (tactons) to portray sound loudness, direction, and identity. In Part 2, we visited three public or semi-public locations where we demonstrated sound feedback on the smartwatch in situ to examine contextual influences and explore sound filtering options. Our findings characterize uses for vibration in multimodal sound awareness, both for push notification and for immediately actionable sound information displayed through vibrational patterns (tactons). In situ experiences caused participants to request sound filtering - particularly to limit haptic feedback - as a method for managing soundscape complexity. Additional concerns arose related to learnability, possibility of distraction, and system trust. Our findings have implications for future portable sound awareness systems.
Steven M. Goodman, Susanne Kirchner, Rose Guttman, Dhruv Jain, Jon Froehlich, Leah Findlater
CHI4
2020 HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing Users
abstract
We introduce HomeSound, an in-home sound awareness system for Deaf and hard of hearing (DHH) users. Similar to the Echo Show or Nest Hub, HomeSound consists of a microphone and display, and uses multiple devices installed in each home. We iteratively developed two prototypes, both of which sense and visualize sound information in real-time. Prototype 1 provided a floorplan view of sound occurrences with waveform histories depicting loudness and pitch. A three-week deployment in four DHH homes showed an increase in participants' home- and self-awareness but also uncovered challenges due to lack of line of sight and sound classification. For Prototype 2, we added automatic sound classification and smartwatch support for wearable alerts. A second field deployment in four homes showed further increases in awareness but misclassifications and constant watch vibrations were not well received. We discuss findings related to awareness, privacy, and display placement and implications for future home sound awareness technology.
Dhruv Jain, Kelly Mack, Akli Amrous, Steven M. Goodman, Leah Findlater, Jon Froehlich
CHI1
2019 Autoethnography of a Hard of Hearing Traveler
abstract
Travel experiences offer a diverse view into an individual's interactions with different cultures, societies, and places. In this paper, we present a 2.5-year autoethnographic travel account of a hard of hearing individual-Jain. Through retrospective journals and field notes, we reveal the tensions and nuances in his travel, including the magnified difficulty of social conversations, issues with navigating unfamiliar environments and cultural contexts, and changes in the relationship to personal assistive technologies. By exploring the longitudinal travel experiences of a single individual, we uncover evocative and personal insights rarely available through participant-based research methods. Based on these lived experiences and post hoc reflections, we present two design explorations of personalized technology the autoethnographer created for aiding his travel. Finally, we offer reflections for customized travel technologies for deaf and hard of hearing users, and methodological guidelines for performing first-person research in the context of disability.
Dhruv Jain, Audrey Desjardins, Leah Findlater, Jon Froehlich
ASSETS1
2019 Deaf and Hard-of-hearing Individuals' Preferences for Wearable and Mobile Sound Awareness Technologies
abstract
To investigate preferences for mobile and wearable sound awareness systems, we conducted an online survey with 201 DHH participants. The survey explores how demographic factors affect perceptions of sound awareness technologies, gauges interest in specific sounds and sound characteristics, solicits reactions to three design scenarios (smartphone, smartwatch, head-mounted display) and two output modalities (visual, haptic), and probes issues related to social context of use. While most participants were highly interested in being aware of sounds, this interest was modulated by communication preference--that is, for sign or oral communication or both. Almost all participants wanted both visual and haptic feedback and 75% preferred to have that feedback on separate devices (e.g., haptic on smartwatch, visual on head-mounted display). Other findings related to sound type, full captions vs. keywords, sound filtering, notification styles, and social context provide direct guidance for the design of future mobile and wearable sound awareness systems.
Leah Findlater, Bonnie Chinh, Dhruv Jain, Jon Froehlich, Raja S. Kushalnagar, Angela Carey Lin
CHI3
2019 Exploring Sound Awareness in the Home for People who are Deaf or Hard of Hearing
abstract
The home is filled with a rich diversity of sounds from mundane beeps and whirs to dog barks and children's shouts. In this paper, we examine how deaf and hard of hearing (DHH) people think about and relate to sounds in the home, solicit feedback and reactions to initial domestic sound awareness systems, and explore potential concerns. We present findings from two qualitative studies: in Study 1, 12 DHH participants discussed their perceptions of and experiences with sound in the home and provided feedback on initial sound awareness mockups. Informed by Study 1, we designed three tablet-based sound awareness prototypes, which we evaluated with 10 DHH participants using a Wizard-of-Oz approach. Together, our findings suggest a general interest in smarthome-based sound awareness systems particularly for displaying contextually aware, personalized and glanceable visualizations but key concerns arose related to privacy, activity tracking, cognitive overload, and trust.
Dhruv Jain, Angela Lin, Rose Guttman, Marcus Amalachandran, Aileen Zeng, Leah Findlater, Jon Froehlich
CHI1
2018 Towards Accessible Conversations in a Mobile Context for People who are Deaf and Hard of Hearing
abstract
Prior work has explored communication challenges faced by people who are deaf and hard of hearing (DHH) and the potential role of new captioning and support technologies to address these challenges; however, the focus has been on stationary contexts such as group meetings and lectures. In this paper, we present two studies examining the needs of DHH people in moving contexts (e.g., walking) and the potential for mobile captions on head-mounted displays (HMDs) to support those needs. Our formative study with 12 DHH participants identifies social and environmental challenges unique to or exacerbated by moving contexts. Informed by these findings, we introduce and evaluate a proof-of-concept HMD prototype with 10 DHH participants. Results show that, while walking, HMD captions can support communication access and improve attentional balance between the speakers(s) and navigating the environment. We close by describing open questions in the mobile context space and design guidelines for future technology.
Dhruv Jain, Rachel L. Franz, Leah Findlater, Jackson Cannon, Raja S. Kushalnagar, Jon Froehlich
ASSETS1
2018 Automated GPU Grid Geometry Selection for OPENMP Kernels
abstract
Modern supercomputers are increasingly using GPUs to improve performance per watt. Generating GPU code for target regions in openMP 4.0, or later versions, requires the selection of grid geometry to execute the GPU kernel. Existing industrial-strength compilers use a simple heuristic with arbitrary numbers that are constant for all kernels. After characterizing the relationship between region features, grid geometry and performance, we built a machine-learning model that successfully predicts a suitable geometry for such kernels and results in a performance improvement with a geometric mean of 5% across the benchmarks studied. However, this prediction is impractical because the overhead of the predictor is too high. A careful study of the results of the predictor allowed for the development of a practical low-overhead heuristic that resulted in a performance improvement of up to 7 times with a geometric mean of 25.9%. This paper describes the methodology to build the machine-learning model, and the practical low-overhead heuristic that can be used in industry-strong compilers.
Taylor Lloyd, Artem Chikin, Sanket Kedia, Dhruv Jain, José Nelson Amaral
SBAC-PAD4
2016 Immersive Scuba Diving Simulator Using Virtual Reality
abstract
We present Amphibian, a simulator to experience scuba diving virtually in a terrestrial setting. While existing diving simulators mostly focus on visual and aural displays, Amphibian simulates a wider variety of sensations experienced underwater. Users rest their torso on a motion platform to feel buoyancy. Their outstretched arms and legs are placed in a suspended harness to simulate drag as they swim. An Oculus Rift head-mounted display (HMD) and a pair of headphones delineate the visual and auditory ocean scene. Additional senses simulated in Amphibian are breath motion, temperature changes, and tactile feedback through various sensors. Twelve experienced divers compared Amphibian to real-life scuba diving. We analyzed the system factors that influenced the users' sense of being there while using our simulator. We present future UI improvements for enhancing immersion in VR diving simulators.
Dhruv Jain, Misha Sra, Jingru Guo, Rodrigo Marques, Raymond Wu, Justin Chiu, Chris Schmandt
UIST1
2015 Head-Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of Hearing
abstract
Persons with hearing loss use visual signals such as gestures and lip movement to interpret speech. While hearing aids and cochlear implants can improve sound recognition, they generally do not help the wearer localize sound necessary to leverage these visual cues. In this paper, we design and evaluate visualizations for spatially locating sound on a head-mounted display (HMD). To investigate this design space, we developed eight high-level visual sound feedback dimensions. For each dimension, we created 3-12 example visualizations and evaluated these as a design probe with 24 deaf and hard of hearing participants (Study 1). We then implemented a real-time proof-of-concept HMD prototype and solicited feedback from 4 new participants (Study 2). Study 1 findings reaffirm past work on challenges faced by persons with hearing loss in group conversations, provide support for the general idea of sound awareness visualizations on HMDs, and reveal preferences for specific design options. Although preliminary, Study 2 further contextualizes the design probe and uncovers directions for future work.
Dhruv Jain, Leah Findlater, Jamie Gilkeson, Benjamin Holland, Ramani Duraiswami, Dmitry N. Zotkin, Christian Vogler, Jon Froehlich
CHI1
2014 Path-guided indoor navigation for the visually impaired using minimal building retrofitting
abstract
One of the common problems faced by visually impaired people is of independent path-based mobility in an unfamiliar indoor environment. Existing systems do not provide active guidance or are bulky, expensive and hence are not socially apt. In this paper, we present the design of an omnipresent cellphone based active indoor wayfinding system for the visually impaired. Our system provides step-by-step directions to the destination from any location in the building using minimal additional infrastructure. The carefully calibrated audio, vibration instructions and the small wearable device helps the user to navigate efficiently and unobtrusively. Results from a formative study with five visually impaired individuals informed the design of the system. We then deployed the system in a building and field tested it with ten visually impaired users. The comparison of the quantitative and qualitative results demonstrated that the system is useful and usable, but can still be improved.
Dhruv Jain
ASSETS1
2014 Pilot evaluation of a path-guided indoor navigation system for visually impaired in a public museum
abstract
One of the common problems faced by visually impaired people is of independent path-based mobility in an unfamiliar indoor environment. Existing systems do not provide active guidance or are bulky, expensive and hence are not socially apt. Consequently, no system has found wide scale deployment in a public place. Our system is an omnipresent cellphone based indoor wayfinding system for the visually impaired. It provides step-by-step directions to the destination from any location in the building using minimal additional infrastructure. The carefully calibrated audio, vibration instructions and the small wearable device helps the user to navigate efficiently and unobtrusively. In this paper, we present the results from pilot testing of the system with one visually impaired user in a national science museum.
Dhruv Jain
ASSETS1
2013 A path-guided audio based indoor navigation system for persons with visual impairment
abstract
Independent path-based mobility in an unfamiliar indoor environment is a common problem faced by visually impaired community. We present the design of an infra-red based active wayfinding system for the visually impaired. Our proposed system: downloads the floor plan of the building, locates and tracks the user inside the building, finds the shortest path and provides step-by-step direction to the destination using voice messages. The audio instructions include active guidance for impending turns in the path of travel, distance of each section between turns, obstacle warning instructions and position correction messages when the user gets lost. Results from a needs finding study with visually impaired individuals formed the design of the system. We then deployed the system in a building and field tested it with users using a standardized before-and-after study. The comparison of the results demonstrated that the system is usable and useful.
Dhruv Jain, Akhil Jain, Rohan Paul, Akhila Komarika, M. Balakrishnan
ASSETS1