VLDB 2026 Research / reviewers in the wild / expert
Caluã de Lacerda Pataca
dblp:259/6374
· DBLP profile ↗
12ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-5046-9884ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASL Educators' Perspectives on AI for Enhancing Student Learning in American Sign Language EducationabstractInterest in learning American Sign Language (ASL) is growing across higher education institutions in North America, as reflected in rising enrollments. Yet this growth is constrained by limited program availability and few opportunities to practice outside the classroom. AI-based technologies show promise for supporting ASL learning, but educators – who bring essential pedagogical, linguistic, and cultural expertise – have been largely absent from conversations on the design of these tools, with prior work focusing primarily on learners. To address this, we conducted formative interviews with eleven Deaf and one hearing ASL instructor, followed by two focus groups with six Deaf educators, to examine how AI tools could support ASL education. Findings revealed priorities for technology design and considerations for integration into existing pedagogical practices, with attention to curricular, linguistic, and access factors. We offer insights for designing and researching technologies aimed at (1) providing adaptive, structured feedback on signing performance and (2) supporting immersive conversational practice with virtual signing partners. Saad Hassan, Laleh Nourian, Caluã de Lacerda Pataca, Michelle M. Olson, Toni D'aurio, Kanupriya Agarwal, Syeda Mah Noor Asad, Garreth W. Tigwell, Matt Huenerfauth |
CHI | 3 |
| 2026 | Fuzzy Feelings: Arousal's Interpretive Noise and the Case for Acoustic-Based HapticsabstractCaptions rarely convey emotional nuances in speech, leaving Deaf and Hard-of-Hearing (dhh) viewers without access to tonal and affective information. We present a two-part mixed-methods study on how haptic feedback can communicate vocal emotion without adding visual load. In Part 1, we replicated an arousal-driven captioning approach using speech-emotion-recognition to modulate typographic weight and vibration intensity. Participants showed divergent mental models and often mapped “more vibration” to loudness rather than emotional arousal, underscoring the construct’s conceptual fuzziness. In Part 2, we evaluated five acoustic-to-haptic mappings that bypass affective inference and translate pitch, rhythm, and waveform cues into vibration patterns. No single pattern dominated, but participants associated options such as pulse or sawtooth with high-arousal emotions, and pitch-normalized signals with calmer states. We derive design guidelines emphasizing contrastive, acoustically grounded mappings and user control for integrating emotional haptics into short-form, captioned media. Caluã de Lacerda Pataca, Stephanie Patterson, Roshan Lalintha Peiris, Matt Huenerfauth |
CHI | 1 |
| 2025 | CapTune: Adapting Non-Speech Captions With Anchored Generative ModelsabstractNon-speech captions are essential to the video experience of deaf and hard of hearing (DHH) viewers, yet conventional approaches often overlook the diversity of their preferences. We present CapTune, a system that enables customization of non-speech captions based on DHH viewers' needs while preserving creator intent. CapTune allows caption authors to define safe transformation spaces using concrete examples and empowers viewers to personalize captions across four dimensions: level of detail, expressiveness, sound representation method, and genre alignment. Evaluations with seven caption creators and twelve DHH participants showed that CapTune supported creators' creative control while enhancing viewers' emotional engagement with content. Our findings also reveal trade-offs between information richness and cognitive load, tensions between interpretive and descriptive representations of sound, and the context-dependent nature of caption preferences. Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu, Dhruv Jain |
ASSETS | 2 |
| 2025 | Demo of CapTune: Adapting Non-Speech Captions with Anchored Generative Models
Jeremy Zhengqi Huang, Caluã de Lacerda Pataca, Liang-Yuan Wu, Dhruv Jain |
ASSETS | 2 |
| 2025 | CuCap: Comparative Analysis of Customized Captioning between North American and South Korean d/Deaf and Hard-of-Hearing UsersabstractAffective and prosodic captions convey not only what a speaker says, but also how they say it-louder words may appear thicker, quieter ones thinner; angry in red, calm in blue.These captions can improve access, satisfaction, and engagement for d/Deaf and Hard-of-Hearing (dhh) users.While prior work has explored their design space, it has focused largely on dhh participants in North America, limiting generalizability beyond English and Latin-based scripts.To uncover the role of culture and language, we ran an exploratory study with 49 dhh participants from North America and South Korea using CuCap, a tool that allowed them to personalize which speech features were displayed, and how.While emotion visualization was a universally favored choice, confirming prior findings, prosody preferences varied across cultures, reflecting linguistic and hearing factors.These findings point to the need for flexible captioning systems that account for cultural, linguistic, and individual differences. Caluã de Lacerda Pataca, Sooyeon Ahn 0001, Suhyeon Yoo, JooYeong Kim, Khai N. Truong, Jin-Hyuk Hong, Roshan Lalintha Peiris, Matt Huenerfauth |
ASSETS | 1 |
| 2025 | Tactile Emotions: Multimodal Affective Captioning with Haptics Improves Narrative Engagement for d/Deaf and Hard-of-Hearing ViewersabstractFigure 1: Multimodal afective captions, combining visual cues and vibrations felt via a wrist-worn device, enrich the viewing experience for d/Deaf or Hard-of-Hearing individuals by portraying speaker emotions, improving engagement. Caluã de Lacerda Pataca, Saad Hassan, Lloyd May, Michelle M. Olson, Toni D'aurio, Roshan Lalintha Peiris, Matt Huenerfauth |
CHI | 1 |
| 2024 | Designing and Evaluating an Advanced Dance Video Comprehension Tool with In-situ Move Identification CapabilitiesabstractAnalyzing dance moves and routines is a foundational step in learning dance. Videos are often utilized at this step, and advancements in machine learning, particularly in human-movement recognition, could further assist dance learners. We developed and evaluated a Wizard-of-Oz prototype of a video comprehension tool that offers automatic in-situ dance move identification functionality. Our system design was informed by an interview study involving 12 dancers to understand the challenges they face when trying to comprehend complex dance videos and taking notes. Subsequently, we conducted a within-subject study with 8 Cuban salsa dancers to identify the benefits of our system compared to an existing traditional feature-based search system. We found that the quality of notes taken by participants improved when using our tool, and they reported a lower workload. Based on participants’ interactions with our system, we offer recommendations on how an AI-powered span-search feature can enhance dance video comprehension tools. Saad Hassan, Caluã de Lacerda Pataca, Laleh Nourian, Garreth W. Tigwell, Briana Davis, Will Zhenya Silver Wagman |
CHI | 2 |
| 2024 | Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing IndividualsabstractAffective captions employ visual typographic modulations to convey a speaker’s emotions, improving speech accessibility for Deaf and Hard-of-Hearing (dhh) individuals. However, the most effective visual modulations for expressing emotions remain uncertain. Bridging this gap, we ran three studies with 39 dhh participants, exploring the design space of affective captions, which include parameters like text color, boldness, size, and so on. Study 1 assessed preferences for nine of these styles, each conveying either valence or arousal separately. Study 2 combined Study 1’s top-performing styles and measured preferences for captions depicting both valence and arousal simultaneously. Participants outlined readability, minimal distraction, intuitiveness, and emotional clarity as key factors behind their choices. In Study 3, these factors and an emotion-recognition task were used to compare how Study 2’s winning styles performed versus a non-styled baseline. Based on our findings, we present the two best-performing styles as design recommendations for applications employing affective captions. Caluã de Lacerda Pataca, Saad Hassan, Nathan Tinker, Roshan Lalintha Peiris, Matt Huenerfauth |
CHI | 1 |
| 2023 | Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing UsersabstractSpeech is expressive in ways that caption text does not capture, with emotion or emphasis information not conveyed. We interviewed eight Deaf and Hard-of-Hearing (dhh) individuals to understand if and how captions’ inexpressiveness impacts them in online meetings with hearing peers. Automatically captioned speech, we found, lacks affective depth, lending it a hard-to-parse ambiguity and general dullness. Interviewees regularly feel excluded, which some understand is an inherent quality of these types of meetings rather than a consequence of current caption text design. Next, we developed three novel captioning models that depicted, beyond words, features from prosody, emotions, and a mix of both. In an empirical study, 16 dhh participants compared these models with conventional captions. The emotion-based model outperformed traditional captions in depicting emotions and emphasis, with only a moderate loss in legibility, suggesting its potential as a more inclusive design for captions. Caluã de Lacerda Pataca, Matthew Watkins, Roshan Lalintha Peiris, Sooyeon Lee, Matt Huenerfauth |
CHI | 1 |
| 2023 | Hidden Bawls, Whispers, and Yelps: Can Text Convey the Sound of Speech, Beyond Words?abstractWhether a word was bawled, whispered, or yelped, captions will typically represent it in the same way. If they are your only way to access what is being said, subjective nuances expressed in the voice will be lost. Since so much of communication is carried by these nuances, we posit that if captions are to be used as an accurate representation of speech, embedding visual representations of paralinguistic qualities into captions could help readers use them to better understand speech beyond its mere textual content. This paper presents a model for processing vocal prosody (its loudness, pitch, and duration) and mapping it into visual dimensions of typography (respectively, font-weight, baseline shift, and letter-spacing), creating a visual representation of these lost vocal subtleties that can be embedded directly into the typographical form of text. An evaluation was carried out where participants were exposed to thisspeech-modulated typographyand asked to match it to its originating audio, presented between similar alternatives. Participants (n=117) were able to correctly identify the original audios with an average accuracy of 65%, with no significant difference when showing them modulations as animated or static text. Additionally, participants’ comments showed their mental models of speech-modulated typography varied widely. Caluã de Lacerda Pataca, Paula Dornhofer Paro Costa |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Support in the Moment: Benefits and use of video-span selection and search for sign-language video comprehension among ASL learnersabstractAs they develop comprehension skills, American Sign Language (ASL) learners often view challenging ASL videos, which may contain unfamiliar signs. Current dictionary tools require students to isolate a single sign they do not understand and input a search query, by selecting linguistic properties or by performing the sign into a webcam. Students may struggle with extracting and re-creating an unfamiliar sign, and they must leave the video-watching task to use an external dictionary tool. We investigate a technology that enables users, in the moment, i.e., while they are viewing a video, to select a span of one or more signs that they do not understand, to view dictionary results. We interviewed 14 American Sign Language (ASL) learners about their challenges in understanding ASL video and workarounds for unfamiliar vocabulary. We then conducted a comparative study and an in-depth analysis with 15 ASL learners to investigate the benefits of using video sub-spans for searching, and their interactions with a Wizard-of-Oz prototype during a video-comprehension task. Our findings revealed benefits of our tool in terms of quality of video translation produced and perceived workload to produce translations. Our in-depth analysis also revealed benefits of an integrated search tool and use of span-selection to constrain video play. These findings inform future designers of such systems, computer vision researchers working on the underlying sign matching technologies, and sign language educators. Saad Hassan, Akhter Al Amin, Caluã de Lacerda Pataca, Diego Navarro, Alexis Gordon, Sooyeon Lee, Matt Huenerfauth |
ASSETS | 3 |
| 2020 | Speech modulated typography: towards an affective representation modelabstractThe transcription of expressive speech into text is a lossy process, since traditional textual resources are typically not capable of fully representing prosody. Speech modulated typography aims to narrow the gap between expressive speech and its textual transcription, with potential applications in affect-sensitive text interfaces, closed-captioning, and automated voice transcriptions. This paper proposes and evaluates two different representation models of prosody-related acoustic features of expressive speech mapped as axes of a variable font. Our experiment tested its participants' preferences for four of these modulations: font weight, letter width, letter slant, and baseline shift. Each of these represented utterances expressed in one of five emotions (anger, happiness, neutrality, sadness, and surprise). Participants preferred font-weight for sentences spoken with intensity and baseline shift for quieter utterances. In both cases, the distance between each syllable's fundamental frequency and centroid frequency was a good predictor of these preferences. Caluã de Lacerda Pataca, Paula Dornhofer Paro Costa |
IUI | 1 |