Utku Uckun

dblp:268/8687 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0007-1580-7308ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Screen Reading Enabled by Large Language Models
abstract
Large language models (LLMs), such as the pioneering GPT technology by OpenAI, have undeniably become one of the most significant innovations in recent history. They have achieved phenomenal success across a broad spectrum of applications in numerous industries, transforming how we interact with the digital world. Notwithstanding these remarkable successes, applying LLMs within the realm of accessibility has largely been unexplored. We introduce Savant, as a demonstration of the potential of LLMs for accessibility. Specifically, Savant leverages the impressive text comprehension abilities of LLMs to provide uniform interaction for screen reader users across various applications, mitigating the significant interaction burden imposed by the heterogeneity in user interfaces for blind screen reader users. Savant automates screen reader actions on control elements like buttons, text fields, and drop-down menus via spoken natural language commands (NLCs). Interpreting the NLC, identifying the correct control element, and formulating the action sequence are facilitated by LLMs. Few-shot prompts supply context and guidance for the LLMs to produce appropriate responses, specifically converting the NLC into a correct series of actions on the user interface elements, which are then performed automatically. The demonstration will exhibit Savant’s capability across a variety of exemplar applications, emphasizing its versatility.
Anujay Ghosh, Monalika Padma Reddy, Satwik Ram Kodandaram, Utku Uckun, Vikas Ashok, Xiaojun Bi 0001, I. V. Ramakrishnan
ASSETS4
2024 Enabling Uniform Computer Interaction Experience for Blind Users through Large Language Models
abstract
Blind individuals, who by necessity depend on screen readers to interact with computers, face considerable challenges in navigating the diverse and complex graphical user interfaces of different computer applications. The heterogeneity of various application interfaces often requires blind users to remember different keyboard combinations and navigation methods to use each application effectively. To alleviate this significant interaction burden imposed by heterogeneous application interfaces, we present Savant, a novel assistive technology powered by large language models (LLMs) that allows blind screen reader users to interact uniformly with any application interface through natural language. Novelly, Savant can automate a series of tedious screen reader actions on the control elements of the application when prompted by a natural language command from the user. These commands can be flexible in the sense that the user is not strictly required to specify the exact names of the control elements in the command. A user study evaluation of Savant with 11 blind participants demonstrated significant improvements in interaction efficiency and usability compared to current practices.
Satwik Ram Kodandaram, Utku Uckun, Xiaojun Bi 0001, I. V. Ramakrishnan, Vikas Ashok
ASSETS2
2022 Taming User-Interface Heterogeneity with Uniform Overlays for Blind Users
abstract
For many blind users, interaction with computer applications using screen reader assistive technology is a frustrating and time-consuming affair, mostly due to the complexity and heterogeneity of applications’ user interfaces. An interview study revealed that many applications do not adequately convey their interface structure and controls to blind screen reader users, thereby placing additional burden on these users to acquire this knowledge on their own. This is often an arduous and tedious learning process given the one-dimensional navigation paradigm of screen readers. Moreover, blind users have to repeat this learning process multiple times, i.e., once for each application, since applications differ in their interface designs and implementations. In this paper, we propose a novel push-based approach to make non-visual computer interaction easy, efficient, and uniform across different applications. The key idea is to make screen reader interaction ‘structure-agnostic’, by automatically identifying and extracting all application controls and then instantly ‘pushing’ these controls on demand to the blind user via a custom overlay dashboard interface. Such a custom overlay facilitates uniform and efficient screen reader navigation across all applications. A user study showed significant improvement in user satisfaction and interaction efficiency with our approach compared to a state-of-the-art screen reader.
Utku Uckun, Rohan Tumkur Suresh, Javedul Ferdous, Xiaojun Bi 0001, I. V. Ramakrishnan, Vikas Ashok
UMAP1
2021 Non-Visual Accessibility Assessment of Videos
abstract
Video accessibility is crucial for blind screen-reader users as online videos are increasingly playing an essential role in education, employment, and entertainment. While there exist quite a few techniques and guidelines that focus on creating accessible videos, there is a dearth of research that attempts to characterize the accessibility of existing videos. Therefore in this paper, we define and investigate a diverse set of video and audio-based accessibility features in an effort to characterize accessible and inaccessible videos. As a ground truth for our investigation, we built a custom dataset of 600 videos, in which each video was assigned an accessibility score based on the number of its wins in a Swiss-system tournament, where human annotators performed pairwise accessibility comparisons of videos. In contrast to existing accessibility research where the assessments are typically done by blind users, we recruited sighted users for our effort, since videos comprise a special case where sight could be required to better judge if any particular scene in a video is presently accessible or not. Subsequently, by examining the extent of association between the accessibility features and the accessibility scores, we could determine the features that significantly (positively or negatively) impact video accessibility and therefore serve as good indicators for assessing the accessibility of videos. Using the custom dataset, we also trained machine learning models that leveraged our handcrafted features to either classify an arbitrary video as accessible/inaccessible or predict an accessibility score for the video. Evaluation of our models yielded an F1 score of 0.675 for binary classification and a mean absolute error of 0.53 for score prediction, thereby demonstrating their potential in video accessibility assessment while also illuminating their current limitations and the need for further research in this area.
Ali Selman Aydin, Yu-Jung Ko, Utku Uckun, I. V. Ramakrishnan, Vikas Ashok
CIKM3
2020 Ontology-Driven Transformations for PDF Form Accessibility
abstract
Filling out PDF forms with screen readers has always been a challenge for people who are blind. Many of these forms are not interactive and hence are not accessible; even if they are interactive, the serial reading order of the screen reader makes it difficult to associate the correct labels with the form fields. This demo will present TransPAc[5], an assistive technology that enables blind people to fill out PDF forms. Since blind people are familiar with web browsing, TransPAc leverages this fact by faithfully transforming a PDF document with forms into a HTML page. The blind user fills out the form fields in the HTML page with their screen reader and these filled-in data values are transparently transferred onto the corresponding form fields in the PDF document. TransPAc thus addresses a long standing problem in PDF form accessibility.
Utku Uckun, Ali Selman Aydin, Vikas Ashok, I. V. Ramakrishnan
ASSETS1
2020 Breaking the Accessibility Barrier in Non-Visual Interaction with PDF Forms
abstract
PDF forms are ubiquitous. Businesses big and small, government agencies, health and educational institutions and many others have all embraced PDF forms. People use PDF forms for providing information to these entities. But people who are blind frequently find it very difficult to fill out PDF forms with screen readers, the standard assistive software that they use for interacting with computer applications. Firstly, many of the them are not even accessible as they are non-interactive and hence not editable on a computer. Secondly, even if they are interactive, it is not always easy to associate the correct labels with the form fields, either because the labels are not meaningful or the sequential reading order of the screen reader misses the visual cues that associate the correct labels with the fields. In this paper we present a solution to the accessibility problem of PDF forms. We leverage the fact that many people with visual impairments are familiar with web browsing and are proficient at filling out web forms. Thus, we create a web form layer over the PDF form via a high fidelity transformation process that attempts to preserve all the spatial relationships of the PDF elements including forms, their labels and the textual content. Blind people only interact with the web forms, and the filled out web form fields are transparently transferred to the corresponding fields in the PDF form. An optimization algorithm automatically adjusts the length and width of the PDF fields to accommodate arbitrary size field data. This ensures that the filled out PDF document does not have any truncated form-field values, and additionally, it is readable. A user study with fourteen users with visual impairments revealed that they were able to populate more form fields than the status quo and the self-reported user experience with the proposed interface was superior compared to the status quo.
Utku Uckun, Ali Selman Aydin, Vikas Ashok, I. V. Ramakrishnan
Proc. ACM Hum. Comput. Interact.1