Jason Wu 0001

dblp:78/5374-1 · DBLP profile ↗
← Back
23ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0001-5101-0557ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 21 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving User Interface Generation Models from Designer Feedback
abstract
Despite being trained on vast amounts of data, most LLMs are unable to reliably generate well-designed UIs. Designer feedback is essential to improving performance on UI generation; however, we find that existing RLHF methods based on ratings or rankings are not well-aligned with with designers' workflows and ignore the rich rationale used to critique and improve UI designs. In this paper, we investigate several approaches for designers to give feedback to UI generation models, using familiar interactions such as commenting, sketching and direct manipulation. We first perform an evaluation with 21 designers where they gave feedback using these interactions, which resulted in 1500 design annotations. We then use this data to finetune a series of LLMs to generate higher quality UIs. Finally, we evaluate these models with human judges, and we find that our designer-aligned approaches outperform models trained with traditional ranking feedback and all tested baselines, including GPT-5.
Jason Wu 0001, Amanda Swearngin, Arun Krishnavajjala, Alan Leung, Jeffrey Nichols 0001, Titus Barik
CHI1
2025 CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development
Peya Mowar, Yi-Hao Peng, Jason Wu 0001, Aaron Steinfeld, Jeffrey P. Bigham
CHI3
2025 SQUIRE: Interactive UI Authoring via Slot QUery Intermediate REpresentations
Alan Leung, Ruijia Cheng, Jason Wu 0001, Jeffrey Nichols 0001, Titus Barik
UIST3
2024 DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
Yi-Hao Peng, Faria Huq, Yue Jiang 0002, Jason Wu 0001, Xin Yue Li, Jeffrey P. Bigham, Amy Pavel
ECCV (24)4
2024 FrameKit: A Tool for Authoring Adaptive UIs Using Keyframes
abstract
Adaptive user interfaces (AUIs) can improve user experience by automatically adapting how information and functionality are presented in a user interface. However, the dynamic nature and potentially numerous variations of AUIs make them challenging to author. In this paper, we present a generalized framework for defining adaptation as interpolations between UIs and introduce a computational approach for intelligently generating new variations of a UI from a small set of designs. Based on this approach, we develop FrameKit, an authoring tool with a programming-by-example interface that retains flexibility and control afforded by manual authoring while reducing effort through automatic generation. We demonstrate that FrameKit can support adaptations that typically require domain-specific toolkits, such as those found in context-aware applications, responsive UIs, and ability-based adaptation. We evaluated FrameKit with ten front-end developers, who successfully authored AUIs after a short tutorial session and suggested that FrameKit provides an effective mental model for AUI authoring.
Jason Wu 0001, Kashyap Todi, Joannes Chan, Brad A. Myers, Benjamin J. Lafreniere
IUI1
2024 UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
abstract
Jason Wu, Eldon Schoop, Alan Leung, Titus Barik, Jeffrey Bigham, Jeffrey Nichols. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Wu 0001, Eldon Schoop, Alan Leung, Titus Barik, Jeffrey P. Bigham, Jeffrey Nichols 0001
NAACL-HLT1
2024 UIClip: A Data-driven Model for Assessing User Interface Design
abstract
User interface (UI) design is a difficult yet important task for ensuring the usability, accessibility, and aesthetic qualities of applications. In our paper, we develop a machine-learned model, UIClip, for assessing the design quality and visual relevance of a UI given its screenshot and natural language description. To train UIClip, we used a combination of automated crawling, synthetic augmentation, and human ratings to construct a large-scale dataset of UIs, collated by description and ranked by design quality. Through training on the dataset, UIClip implicitly learns properties of good and bad designs by i) assigning a numerical score that represents a UI design’s relevance and quality and ii) providing design suggestions. In an evaluation that compared the outputs of UIClip and other baselines to UIs rated by 12 human designers, we found that UIClip achieved the highest agreement with ground-truth rankings. Finally, we present three example applications that demonstrate how UIClip can facilitate downstream applications that rely on instantaneous assessment of UI design quality: i) UI code generation, ii) UI design tips generation, and iii) quality-aware UI example search.
Jason Wu 0001, Yi-Hao Peng, Xin Yue Amanda Li, Amanda Swearngin, Jeffrey P. Bigham, Jeffrey Nichols 0001
UIST1
2024 Towards Automated Accessibility Report Generation for Mobile Apps
abstract
Many apps have basic accessibility issues, like missing labels or low contrast. To supplement manual testing, automated tools can help developers and QA testers find basic accessibility issues, but they can be laborious to use or require writing dedicated tests. To motivate our work, we interviewed eight accessibility QA professionals at a large technology company. From these interviews, we synthesized three design goals for accessibility report generation systems. Motivated by these goals, we developed a system to generate whole app accessibility reports by combining varied data collection methods (e.g., app crawling, manual recording) with an existing accessibility scanner. Many such scanners are based on single-screen scanning, and a key problem in whole app accessibility reporting is to effectively de-duplicate and summarize issues collected across an app. To this end, we developed a screen grouping model with 96.9% accuracy (88.8% F1-score) and UI element matching heuristics with 97% accuracy (98.2% F1-score). We combine these technologies in a system to report and summarize unique issues across an app, and enable a unique pixel-based ignore feature to help engineers and testers better manage reported issues across their app’s lifetime. We conducted a user study where 19 accessibility engineers and testers used multiple tools to create lists of prioritized issues in the context of an accessibility audit. Our system helped them create lists they were more satisfied with while addressing key limitations of current accessibility scanning tools.
Amanda Swearngin, Jason Wu 0001, Xiaoyi Zhang 0006, Esteban Gomez, Jen Coughenour, Rachel Stukenborg, Bhavya Garg, Greg Hughes, Adriana Hilliard, Jeffrey P. Bigham, Jeffrey Nichols 0001
ACM Trans. Comput. Hum. Interact.2
2023 WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics
abstract
Modeling user interfaces (UIs) from visual information allows systems to make inferences about the functionality and semantics needed to support use cases in accessibility, app automation, and testing. Current datasets for training machine learning models are limited in size due to the costly and time-consuming process of manually collecting and annotating UIs. We crawled the web to construct WebUI, a large dataset of 400,000 rendered web pages associated with automatically extracted metadata. We analyze the composition of WebUI and show that while automatically extracted data is noisy, most examples meet basic criteria for visual UI modeling. We applied several strategies for incorporating semantics found in web pages to increase the performance of visual UI understanding models in the mobile domain, where less labeled data is available: (i) element detection, (ii) screen classification and (iii) screen similarity.
Jason Wu 0001, Siyan Wang, Siman Shen, Yi-Hao Peng, Jeffrey Nichols 0001, Jeffrey P. Bigham
CHI1
2023 STAR: Smartphone-analogous Typing in Augmented Reality
abstract
While text entry is an essential and frequent task in Augmented Reality (AR) applications, devising an efficient and easy-to-use text entry method for AR remains an open challenge. This research presents STAR, a smartphone-analogous AR text entry technique that leverages a user’s familiarity with smartphone two-thumb typing. With STAR, a user performs thumb typing on a virtual QWERTY keyboard that is overlain on the skin of their hands. During an evaluation study of STAR, participants achieved a mean typing speed of 21.9 WPM (i.e., 56% of their smartphone typing speed), and a mean error rate of 0.3% after 30 minutes of practice. We further analyze the major factors implicated in the performance gap between STAR and smartphone typing, and discuss ways this gap could be narrowed.
Taejun Kim, Amy Karlson, Aakar Gupta, Tovi Grossman, Jason Wu 0001, Parastoo Abtahi, Christopher Collins 0001, Michael Glueck, Hemant Bhaskar Surale
UIST5
2023 Never-ending Learning of User Interfaces
abstract
Machine learning models have been trained to predict semantic information about user interfaces (UIs) to make apps more accessible, easier to test, and to automate. Currently, most models rely on datasets of static screenshots that are labeled by human annotators, a process that is costly and surprisingly error-prone for certain tasks. For example, workers labeling whether a UI element is “tappable” from a screenshot must guess using visual signifiers, and do not have the benefit of tapping on the UI element in the running app and observing the effects. In this paper, we present the Never-ending UI Learner, an app crawler that automatically installs real apps from a mobile app store and crawls them to infer semantic properties of UIs by interacting with UI elements, discovering new and challenging training examples to learn from, and continually updating machine learning models designed to predict these semantics. The Never-ending UI Learner so far has crawled for more than 5,000 device-hours, performing over half a million actions on 6,000 apps to train three computer vision models for i) tappability prediction, ii) draggability prediction, and iii) screen similarity.
Jason Wu 0001, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin, Jeffrey P. Bigham, Jeffrey Nichols 0001
UIST1
2022 Towards Complete Icon Labeling in Mobile Applications
abstract
Accurately recognizing icon types in mobile applications is integral to many tasks, including accessibility improvement, UI design search, and conversational agents. Existing research focuses on recognizing the most frequent icon types, but these technologies fail when encountering an unrecognized low-frequency icon. In this paper, we work towards complete coverage of icons in the wild. After annotating a large-scale icon dataset (327,879 icons) from iPhone apps, we found a highly uneven distribution: 98 common icon types covered 92.8% of icons, while 7.2% of icons were covered by more than 331 long-tail icon types. In order to label icons with widely varying occurrences in apps, our system uses an image classification model to recognize common icon types with an average of 3,000 examples each (96.3% accuracy) and applies a few-shot learning model to classify long-tail icon types with an average of 67 examples each (78.6% accuracy). Our system also detects contextual information that helps characterize icon semantics, including nearby text (95.3% accuracy) and modifier symbols added to the icon (87.4% accuracy). In a validation study with workers (n = 23), we verified the usefulness of our generated icon labels. The icon types supported by our work cover 99.5% of collected icons, improving on the previously highest 78% coverage in icon classification work.
Jieshan Chen, Amanda Swearngin, Jason Wu 0001, Titus Barik, Jeffrey Nichols 0001, Xiaoyi Zhang 0006
CHI3
2022 Understanding Screen Relationships from Screenshots of Smartphone Applications
abstract
All graphical user interfaces are comprised of one or more screens that may be shown to the user depending on their interactions. Identifying different screens of an app and understanding the type of changes that happen on the screens is a challenging task that can be applied in many areas including automatic app crawling, playback of app automation macros and large scale app dataset analysis. For example, an automated app crawler needs to understand if the screen it is currently viewing is the same as any previous screen that it has encountered, so it can focus its efforts on portions of the app that it has not yet explored. Moreover, identifying the type of change on the screen, such as whether any dialogues or keyboards have opened or closed, is useful for an automatic crawler to handle such events while crawling. Understanding screen relationships is a difficult task as instances of the same screen may have visual and structural variation, for example due to different content in a database-backed application, scrolling, dialog boxes opening or closing, or content loading delays. At the same time, instances of different screens from the same app may share some similarities in terms of design, structure, and content. This paper uses a dataset of screenshots from more than 1K iPhone applications to train two ML models that understand similarity in different ways: (1) a screen similarity model that combines a UI object detector with a transformer model architecture to recognize instances of the same screen from a collection of screenshots from a single app, and (2) a screen transition model that uses a siamese network architecture to identify both similarity and three types of events that appear in an interaction trace: the keyboard or a dialog box appearing or disappearing, and scrolling. Our models achieve an F1 score of 0.83 on the screen similarity task, improving on comparable baselines, and an average F1 score of 0.71 across all events in the transition task.
Shirin Feiz, Jason Wu 0001, Xiaoyi Zhang 0006, Amanda Swearngin, Titus Barik, Jeffrey Nichols 0001
IUI2
2022 Diffscriber: Describing Visual Design Changes to Support Mixed-Ability Collaborative Presentation Authoring
abstract
Visual slide-based presentations are ubiquitous, yet slide authoring tools are largely inaccessible to people who are blind or visually impaired (BVI). When authoring presentations, the 9 BVI presenters in our formative study usually work with sighted collaborators to produce visual slides based on the text content they produce. While BVI presenters valued collaborators’ visual design skill, the collaborators often felt they could not fully review and provide feedback on the visual changes that were made. We present Diffscriber, a system that identifies and describes changes to a slide’s content, layout, and style for presentation authoring. Using our system, BVI presentation authors can efficiently review changes to their presentation by navigating either a summary of high-level changes or individual slide elements. To learn more about changes of interest, presenters can use a generated change hierarchy to navigate to lower-level change details and element styles. BVI presenters using Diffscriber were able to identify slide design changes and provide feedback more easily as compared to using only the slides alone. More broadly, Diffscriber illustrates how advances in detecting and describing visual differences can improve mixed-ability collaboration.
Yi-Hao Peng, Jason Wu 0001, Jeffrey P. Bigham, Amy Pavel
UIST2
2021 Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels
abstract
Many accessibility features available on mobile platforms require applications (apps) to provide complete and accurate metadata describing user interface (UI) components. Unfortunately, many apps do not provide sufficient metadata for accessibility features to work as expected. In this paper, we explore inferring accessibility metadata for mobile apps from their pixels, as the visual interfaces often best reflect an app’s full functionality. We trained a robust, fast, memory-efficient, on-device model to detect UI elements using a dataset of 77,637 screens (from 4,068 iPhone apps) that we collected and annotated. To further improve UI detections and add semantic information, we introduced heuristics (e.g., UI grouping and ordering) and additional models (e.g., recognize UI content, state, interactivity). We built Screen Recognition to generate accessibility metadata to augment iOS VoiceOver. In a study with 9 screen reader users, we validated that our approach improves the accessibility of existing mobile apps, enabling even previously inaccessible apps to be used.
Xiaoyi Zhang 0006, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle I. Murray, Lisa Yu, Qi Shan, Jeffrey Nichols 0001, Jason Wu 0001, Chris Fleizach, Aaron Everitt, Jeffrey P. Bigham
CHI9
2021 Screen Parsing: Towards Reverse Engineering of UI Models from Screenshots
abstract
Automated understanding of user interfaces (UIs) from their pixels can improve accessibility, enable task automation, and facilitate interface design without relying on developers to comprehensively provide metadata. A first step is to infer what UI elements exist on a screen, but current approaches are limited in how they infer how those elements are semantically grouped into structured interface definitions. In this paper, we motivate the problem of screen parsing, the task of predicting UI elements and their relationships from a screenshot. We describe our implementation of screen parsing and provide an effective training procedure that optimizes its performance. In an evaluation comparing the accuracy of the generated output, we find that our implementation significantly outperforms current systems (up to 23%). Finally, we show three example applications that are facilitated by screen parsing: (i) UI similarity search, (ii) accessibility enhancement, and (iii) code generation from UI screenshots.
Jason Wu 0001, Xiaoyi Zhang 0006, Jeffrey Nichols 0001, Jeffrey P. Bigham
UIST1
2020 Disability and the COVID-19 Pandemic: Using Twitter to Understand Accessibility during Rapid Societal Transition
abstract
The COVID-19 pandemic has forced institutions to rapidly alter their behavior, which typically has disproportionate negative effects on people with disabilities as accessibility is overlooked. To investigate these issues, we analyzed Twitter data to examine accessibility problems surfaced by the crisis. We identified three key domains at the intersection of accessibility and technology: (i) the allocation of product delivery services, (ii) the transition to remote education, and (iii) the dissemination of public health information. We found that essential retailers expanded their high-risk customer shopping hours and pick-up and delivery services, but individuals with disabilities still lacked necessary access to goods and services. Long-experienced access barriers to online education were exacerbated by the abrupt transition of in-person to remote instruction. Finally, public health messaging has been inconsistent and inaccessible, which is unacceptable during a rapidly-evolving crisis. We argue that organizations should create flexible, accessible technology and policies in calm times to be adaptable in times of crisis to serve individuals with diverse needs.
Cole Gleason, Stephanie Valencia, Lynn Kirabo, Jason Wu 0001, Anhong Guo, Elizabeth J. Carter, Jeffrey P. Bigham, Cynthia L. Bennett, Amy Pavel
ASSETS4
2020 Making Mobile Augmented Reality Applications Accessible
abstract
Augmented Reality (AR) technology creates new immersive experiences in entertainment, games, education, retail, and social media. AR content is often primarily visual and it is challenging to enable access to it non-visually due to the mix of virtual and real-world content. In this paper, we identify common constituent tasks in AR by analyzing existing mobile AR applications for iOS, and characterize the design space of tasks that require accessible alternatives. For each of the major task categories, we create prototype accessible alternatives that we evaluate in a study with 10 blind participants to explore their perceptions of accessible AR. Our study demonstrates that these prototypes make AR possible to use for blind users and reveals a number of insights to move forward. We believe our work sets forth not only exemplars for developers to create accessible AR applications, but also a roadmap for future research to make AR comprehensively accessible.
Jaylin Herskovitz, Jason Wu 0001, Samuel White, Amy Pavel, Gabriel Reyes, Anhong Guo, Jeffrey P. Bigham
ASSETS2
2020 Towards Recommending Accessibility Features on Mobile Devices
abstract
Numerous accessibility features have been developed to increase who and how people can access computing devices. Increasingly, these features are included as part of popular platforms, e.g., Apple iOS, Google Android, and Microsoft Windows. Despite their potential to improve the computing experience, many users are unaware of these features and do not know which combination of them could benefit them. In this work, we first quantified this problem by surveying 100 participants online (including 25 older adults) about their knowledge of accessibility and features that they could benefit from, showing very low awareness. We developed four prototypes spanning numerous accessibility categories (e.g., vision, hearing, motor), that embody signals and detection strategies applicable to accessibility recommendation in general. Preliminary results from a study with 20 older adults show that proactive recommendation is a promising approach for better pairing users with accessibility features they could benefit from.
Jason Wu 0001, Gabriel Reyes, Sam C. White, Xiaoyi Zhang 0006, Jeffrey P. Bigham
ASSETS1
2020 Automated Class Discovery and One-Shot Interactions for Acoustic Activity Recognition
abstract
Acoustic activity recognition has emerged as a foundational element for imbuing devices with context-driven capabilities, enabling richer, more assistive, and more accommodating computational experiences. Traditional approaches rely either on custom models trained in situ, or general models pre-trained on preexisting data, with each approach having accuracy and user burden implications. We present Listen Learner, a technique for activity recognition that gradually learns events specific to a deployed environment while minimizing user burden. Specifically, we built an end-to-end system for self-supervised learning of events labelled through one-shot interaction. We describe and quantify system performance 1) on preexisting audio datasets, 2) on real-world datasets we collected, and 3) through user studies which uncovered system behaviors suitable for this new type of interaction. Our results show that our system can accurately and automatically learn acoustic events across environments (e.g., 97% precision, 87% recall), while adhering to users' preferences for non-intrusive interactive behavior.
Jason Wu 0001, Chris Harrison 0001, Jeffrey P. Bigham, Gierad Laput
CHI1
2019 SelfSync: exploring self-synchronous body-based hotword gestures for initiating interaction
abstract
SelfSync enables rapid, robust initiation of a gesture interface using synchronized movement of different body parts. SelfSync is the gestural equivalent of a hotword such as OK-Google in a speech interface and is enabled by the increasing trend where a user wears two or more wearables, such as a smartwatch, wireless earbuds, or a smartphone. In a user study comparing five potential SelfSync gestures in isolation, our system averages 96%, 98% and 88% for user dependent, user adapted, and user independent accuracy, respectively. For when the user has a phone in a pocket and a smart-watch, we suggest twisting the hand about the wrist while moving the leg with the phone in synchrony left and right. When the user has a head worn device and a smartwatch, we suggest twisting the hand while twisting the head left and right.
Shaurye Aggarwal, Jason Wu 0001, Thad Starner, Woontack Woo
UbiComp3
2018 Seesaw: rapid one-handed synchronous gesture interface for smartwatches
abstract
We present SeeSaw, a synchronous gesture interface for commodity smartwatches to support watch-hand only input with no additional hardware. Our algorithm, which uses correlation to determine whether the user is rotating their wrist in synchrony with a tactile and visual prompt, minimizes false-trigger events while maintaining fast input during situational impairments. Results from a 12 person evaluation of the system, used to respond to notifications on the watch during walking and simulated driving, show interaction speeds of 4.0 s - 5.5 s, which is comparable to the swipe-based interface control condition. SeeSaw is also evaluated as an input interface for watches used in conjunction with a head-worn display. A six subject study showed a 95% success rate in dismissing notifications and a 3.57 s mean dismissal time.
Jason Wu 0001, Cooper Colglazier, Adhithya Ravishankar, Yuyan Duan, Yuanbo Wang 0001, Thomas Plötz, Thad Starner
UbiComp1
2018 NADiA: Neural Network Driven Virtual Human Conversation Agents
abstract
Advances in artificial intelligence and in particular machine learning and neural networks have given rise to a new generation of virtual assistants and chatbots. Within this work, we present NADiA - Neurally Animated Dialog Agent - that leverages both the user's verbal input as well as their facial expressions to respond in a meaningful way. NADiA combines a neural language model that generates appropriate responses to user prompts, a convolutional neural network for facial expression analysis, and virtual human technology that is deployed on a mobile phone. Here, we evaluate NADiA's anthropomorphic characteristics and its ability to understand the human interlocutor using both subjective as well as objective measures. We find that NADiA significantly outperforms state of the art chatbot technology and produces comparable behavior to human generated reference outputs.
Jason Wu 0001, Sayan Ghosh 0004, Mathieu Chollet, Steven Ly, Sharon Mozgai, Stefan Scherer
IVA1