Liang-Yuan Wu

dblp:377/3812 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0008-3081-1134ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2025 CapTune: Adapting Non-Speech Captions With Anchored Generative Models
abstract
Non-speech captions are essential to the video experience of deaf and hard of hearing (DHH) viewers, yet conventional approaches often overlook the diversity of their preferences. We present CapTune, a system that enables customization of non-speech captions based on DHH viewers' needs while preserving creator intent. CapTune allows caption authors to define safe transformation spaces using concrete examples and empowers viewers to personalize captions across four dimensions: level of detail, expressiveness, sound representation method, and genre alignment. Evaluations with seven caption creators and twelve DHH participants showed that CapTune supported creators' creative control while enhancing viewers' emotional engagement with content. Our findings also reveal trade-offs between information richness and cognitive load, tensions between interpretive and descriptive representations of sound, and the context-dependent nature of caption preferences.
Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu, Dhruv Jain
ASSETS3
2025 Demo of CapTune: Adapting Non-Speech Captions with Anchored Generative Models
Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu, Dhruv Jain
ASSETS3
2025 SoundNarratives: Rich Auditory Scene Descriptions to Support Deaf and Hard of Hearing People
abstract
state-of-the-art audio language model.A user study with 10 DHH participants demonstrated a significant preference for SoundNarratives over a baseline model, along with a potential for improved confidence and situational awareness.
Liang-Yuan Wu, Dhruv Jain
ASSETS1
2025 EvolveCaptions: Real-Time Collaborative ASR Adaptation for DHH Speakers
abstract
Figure 1: Overview of EvolveCaptions.(1) Hearing users correct live captions of the DHH speaker's voice.(2) The DHH speaker records targeted phrases generated from the corrected terms.(3) The Whisper ASR model is fine-tuned with the recordings and adapts to the speaker over time.
Liang-Yuan Wu, Dhruv Jain
ASSETS1
2025 CARTGPT: Real-Time Correction of CART Captions Using Large Language Models
abstract
Communication Access Realtime Translation (CART) is a widely used captioning technology among deaf and hard of hearing (DHH) individuals, valued for its high accuracy and ability to convey speaker cues and contextual sounds in real time. However, CART performance can degrade in challenging conditions such as background noise, technical jargon, or rapid speech—reducing caption quality and impacting comprehension. We introduce CARTGPT, a real-time captioning system that enhances CART transcripts by leveraging large language models (LLMs) and automatic speech recognition (ASR) input to detect and correct transcription errors. To inform the design of CARTGPT, we conducted a formative study with 10 professional CART captioners to identify common sources of error and their perspectives on using AI for caption correction. We evaluated CARTGPT on a 39.7-hour speech dataset spanning medical, technical, and conversational domains, observing a 5.6% improvement in word accuracy over standard CART and 17.3% over a state-of-the-art ASR model. In a user study with 16 DHH participants, CARTGPT captions were rated as significantly more comprehensible, particularly in technical scenarios, while maintaining real-time responsiveness. These findings demonstrate the potential of LLM-assisted captioning to improve accessibility and comprehension for DHH users in real-world settings.
Liang-Yuan Wu, Andrea Kleiver, Dhruv Jain
ASSETS1
2025 Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User
abstract
Current AI sound awareness systems can provide deaf and hard of hearing people with information about sounds, including discrete sound sources and transcriptions. However, synthesizing AI outputs based on DHH people's ever-changing intents in complex auditory environments remains a challenge. In this paper, we describe the co-design process of SoundWeaver, a sound awareness system prototype that dynamically weaves AI outputs from different AI models based on users’ intents and presents synthesized information through a heads-up display. Adopting a Research through Design perspective, we created SoundWeaver with one DHH co-designer, adapting it to his personal contexts and goals (e.g., cooking at home and chatting in a game store). Through this process, we present design implications for the future of “intent-driven” AI systems for sound accessibility.
Jeremy Zhengqi Huang, Jaylin Herskovitz, Liang-Yuan Wu, Cecily Morrison, Dhruv Jain
CHI3
2024 CARTGPT: Improving CART Captioning using Large Language Models
abstract
Communication Access Realtime Translation (CART) is a commonly used real-time captioning technology used by deaf and hard of hearing (DHH) people, due to its accuracy, reliability, and ability to provide a holistic view of the conversational environment (e.g., by displaying speaker names). However, in many real-world situations (e.g., noisy environments, long meetings), the CART captioning accuracy can considerably decline, thereby affecting the comprehension of DHH people. In this work-in-progress paper, we introduce CARTGPT, a system to assist CART captioners in improving their transcription accuracy. CARTGPT takes in errored CART captions and inaccurate automatic speech recognition (ASR) captions as input and uses a large language model to generate corrected captions in real-time. We quantified performance on a noisy speech dataset, showing that our system outperforms both CART (+5.6% accuracy) and a state-of-the-art ASR model (+17.3%). A preliminary evaluation with three DHH users further demonstrates the promise of our approach.
Liang-Yuan Wu, Andrea Kleiver, Dhruv Jain
ASSETS1