Rosiana Natalie

dblp:221/1650 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-8912-7627ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 6 first-author · 9 since 2021
YearPublicationVenuePosition
2026 TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
abstract
People who are blind or have low vision regularly use their hands to interact with the physical world to gain access to objects’ shape, size, weight, and texture. However, many rich visual features remain inaccessible through touch alone, making it difficult to distinguish similar objects, interpret visual affordances, and form a complete understanding of objects. In this work, we present TouchScribe, a system that augments hand-object interactions with automated live visual descriptions. We trained a custom egocentric hand interaction model to recognize both common gestures (e.g., grab to inspect, hold side-by-side to compare) and unique ones by blind people (e.g., point to explore color, or swipe to read available texts). Furthermore, TouchScribe provides real-time and adaptive feedback based on hand movement, from hand interaction states, to object labels, and to visual details. Our user study and technical evaluations demonstrate that TouchScribe can provide rich and useful descriptions to support object understanding. Finally, we discuss the implications of making live visual descriptions responsive to users’ physical reach.
Ruei-Che Chang, Rosiana Natalie, Jovan Zheng Feng Yap, Tiange Luo, Venkatesh Potluri, Anhong Guo
CHI2
2026 A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
abstract
Computer Use Agents (CUAs) operate interfaces by pointing, clicking, and typing—mirroring interactions of sighted users (SUs) who can thus monitor CUAs and share control. CUAs do not reflect interactions by blind and low-vision users (BLVUs) who use assistive technology (AT). BLVUs thus cannot easily collaborate with CUAs. To characterize the accessibility gap of CUAs, we present A11y-CUA, a dataset of BLVUs and SUs performing 60 everyday tasks with 40.4 hours and 158,325 events. Our dataset analysis reveals that our collected interaction traces quantitatively confirm distinct interaction styles between SU and BLVU groups (mouse- vs. keyboard-dominant) and demonstrate interaction diversity within each group (sequential vs. shortcut navigation for BLVUs). We then compare collected traces to state-of-the-art CUAs under default and AT conditions (keyboard-only, magnifier). The default CUA executed 78.3% of tasks successfully. But with the AT conditions, CUA’s performance dropped to 41.67% and 28.3% with keyboard-only and magnifier conditions respectively, and did not reflect nuances of real AT use. With our open A11y-CUA dataset, we aim to promote collaborative and accessible CUAs for everyone.
Ananya Gubbi Mohanbabu, Rosiana Natalie, Brandon Kim, Anhong Guo, Amy Pavel
CHI2
2025 Probing the Gaps in ChatGPT's Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
Ruei-Che Chang, Rosiana Natalie, Jovan Zheng Feng Yap, Anhong Guo
ASSETS2
2025 How Well Can Vision Language Models Simulate the Vision Perception of People with Low Vision?
abstract
Advances in Vision Language Models (VLMs) have enabled the simulation of general human behavior through their reasoning and problem solving capabilities.In the accessibility domain, such simulations may support the initial piloting of inclusive design processes, without replacing real human input, and can facilitate the personalization of AI-based application outcomes based on individual profiles.In this work, we conducted a preliminary examination of the extent to which VLMs can simulate the visual perception of people with low vision when interpreting images.We conducted a survey study with 40 low vision participants, collecting their brief and detailed vision information, and both open-ended and multiplechoice image perception and recognition responses to up to 25 images.Using these responses, we constructed prompts for VLMs to create simulated agents of each participant, varying the included information on vision information and example image responses.We evaluated the agreement between LLM-generated responses and participants' original answers.The agreement between the agent' and participants' responses remained low when only either the vision profile (0.59) or example image responses (0.59) were provided, whereas a combination of both significantly increase the agreement (0.70, p < 0.0001).Notably, a single example combining both open-ended and multiple-choice responses, offered significant performance improvements over either alone (p < 0.0001), while additional examples provided minimal benefits (p > 0.05). CCS Concepts• Human-centered computing → Accessibility.
Rosiana Natalie, Ruei-Che Chang, Anhong Guo
ASSETS1
2024 Exploring Conversations between a Practitioner and a Person with Dementia
abstract
In social service centers, practitioners engage in conversations with clients with dementia to facilitate their daily activities and provide support when they are distressed. However, the nature of the care demands the practitioner’s active engagement, which becomes difficult to deliver as the number of people who need care expands. Researchers have been investigating the efficacy of developing agents that assume conversational tasks to alleviate this work. To contribute to the future design of agents for caregiving, we collected and analyzed ten conversations between clients with mild dementia and practitioners who provide care. Our analyses of turn-taking dynamics and dialogue acts with 15k utterances uncovered patterns such as noticeable differences in clients’ and practitioners’ conversational dynamics and the prevalence of neutral-toned, question-oriented utterances by practitioners. We then prototyped a large language model-based script that generates responses to client utterances. We found potential approaches and challenges for making its utterance pattern more similar to that of a practitioner.
Kotaro Hara, Rosiana Natalie, Wei Soon Cheong, Jingjing Gu, Qianli Xu
ASSETS2
2024 Audio Description Customization
abstract
Blind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals’ potentially diverse needs and preferences. This research investigates if customizing AD could improve how BLV individuals consume videos. We conducted an interview study (Study 1) with fifteen BLV participants, which revealed desires for customizing properties like length, emphasis, speed, voice, format, tone, and language. At the same time, concerns like interruptions and increased interaction load due to customization emerged. To examine AD customization’s effectiveness and tradeoffs, we designed CustomAD, a prototype that enables BLV users to customize AD content and presentation. An evaluation study (Study 2) with twelve BLV participants showed using CustomAD significantly enhanced BLV people’s video understanding, immersion, and information navigation efficiency. Our work illustrates the importance of AD customization and offers a design that enhances video accessibility for BLV individuals.
Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, Kotaro Hara
ASSETS1
2023 Supporting Novices Author Audio Descriptions via Automatic Feedback
abstract
Audio descriptions (AD) make videos accessible to those who cannot see them. But many videos lack AD and remain inaccessible as traditional approaches involve expensive professional production. We aim to lower production costs by involving novices in this process. We present an AD authoring system that supports novices to write scene descriptions (SD)-textual descriptions of video scenes-and convert them into AD via text-to-speech. The system combines video scene recognition and natural language processing to review novice-written SD and feeds back what to mention automatically. To assess the effectiveness of this automatic feedback in supporting novices, we recruited 60 participants to author SD with no feedback, human feedback, and automatic feedback. Our study shows that automatic feedback improves SD's descriptiveness, objectiveness, and learning quality, without affecting qualities like sufficiency and clarity. Though human feedback remains more effective, automatic feedback can reduce production costs by 45%.
Rosiana Natalie, Joshua Tseng, Hernisa Kacorri, Kotaro Hara
CHI1
2021 The Efficacy of Collaborative Authoring of Video Scene Descriptions
abstract
The majority of online video contents remain inaccessible to people with visual impairments due to the lack of audio descriptions to depict the video scenes. Content creators have traditionally relied on professionals to author audio descriptions, but their service is costly and not readily-available. We investigate the feasibility of creating more cost-effective audio descriptions that are also of high quality by involving novices. Specifically, we designed, developed, and evaluated ViScene, a web-based collaborative audio description authoring tool that enables a sighted novice author and a reviewer either sighted or blind to interact and contribute to scene descriptions (SDs)—text that can be transformed into audio through text-to-speech. Through a mixed-design study with N = 60 participants, we assessed the quality of SDs created by sighted novices with feedback from both sighted and blind reviewers. Our results showed that with ViScene novices could produce content that is Descriptive, Objective, Referable, and Clear at a cost of i.e., US$2.81pvm to US$5.48pvm, which is 54% to 96% lower than the professional service. However, the descriptions lacked in other quality dimensions (e.g., learning, a measure of how well an SD conveys the video’s intended message). While professional audio describers remain the gold standard, for content creators who cannot afford it, ViScene offers a cost-effective alternative, ultimately leading to a more accessible medium.
Rosiana Natalie, Jolene Loh, Huei Suen Tan, Joshua Tseng, Ian Luke Yi-Ren Chan, Ebrima Jarjue, Hernisa Kacorri, Kotaro Hara
ASSETS1
2021 Uncovering Patterns in Reviewers' Feedback to Scene Description Authors
abstract
Audio descriptions (ADs) can increase access to videos for blind people. Researchers have explored different mechanisms for generating ADs, with some of the most recent studies involving paid novices; to improve the quality of their ADs, novices receive feedback from reviewers. However, reviewer feedback is not instantaneous. To explore the potential for real-time feedback through automation, in this paper, we analyze 1,120 comments that 40 sighted novices received from a sighted or a blind reviewer. We find that feedback patterns tend to fall under four themes: (i) Quality; commenting on different AD quality variables, (ii) Speech Act; the utterance or speech action that the reviewers used, (iii) Required Action; the recommended action that the authors should do to improve the AD, and (iv) Guidance; the additional help that the reviewers gave to help the authors. We discuss which of these patterns could be automated within the review process as design implications for future AD collaborative authoring systems.
Rosiana Natalie, Jolene Loh, Huei Suen Tan, Joshua Tseng, Hernisa Kacorri, Kotaro Hara
ASSETS1
2020 ViScene: A Collaborative Authoring Tool for Scene Descriptions in Videos
abstract
Audio descriptions can make the visual content in videos accessible to people with visual impairments. However, the majority of the online videos lack audio descriptions due in part to the shortage of experts who can create high-quality descriptions. We present ViScene, a web-based authoring tool that taps into the larger pool of sighted non-experts to help them generate high-quality descriptions via two feedback mechanisms—succinct visualizations and comments from an expert. Through a mixed-design study with N = 6 participants, we explore the usability of ViScene and the quality of the descriptions created by sighted non-experts with and without feedback comments. Our results indicate that non-experts can produce better descriptions with feedback comments; preliminary insights also highlight the role that people with visual impairments can play in providing this feedback.
Rosiana Natalie, Ebrima Jarjue, Hernisa Kacorri, Kotaro Hara
ASSETS1