Xiaoyi Zhang 0006

dblp:98/4236-6 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-7185-0470ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 16 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2024 Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
abstract
On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria : a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria.
Fred Hohman, Chaoqun Wang 0002, Jinmook Lee, Jochen Görtler, Dominik Moritz, Jeffrey P. Bigham, Zhile Ren, Cecile Foret, Qi Shan, Xiaoyi Zhang 0006
CHI10
2024 Towards Automated Accessibility Report Generation for Mobile Apps
abstract
Many apps have basic accessibility issues, like missing labels or low contrast. To supplement manual testing, automated tools can help developers and QA testers find basic accessibility issues, but they can be laborious to use or require writing dedicated tests. To motivate our work, we interviewed eight accessibility QA professionals at a large technology company. From these interviews, we synthesized three design goals for accessibility report generation systems. Motivated by these goals, we developed a system to generate whole app accessibility reports by combining varied data collection methods (e.g., app crawling, manual recording) with an existing accessibility scanner. Many such scanners are based on single-screen scanning, and a key problem in whole app accessibility reporting is to effectively de-duplicate and summarize issues collected across an app. To this end, we developed a screen grouping model with 96.9% accuracy (88.8% F1-score) and UI element matching heuristics with 97% accuracy (98.2% F1-score). We combine these technologies in a system to report and summarize unique issues across an app, and enable a unique pixel-based ignore feature to help engineers and testers better manage reported issues across their app’s lifetime. We conducted a user study where 19 accessibility engineers and testers used multiple tools to create lists of prioritized issues in the context of an accessibility audit. Our system helped them create lists they were more satisfied with while addressing key limitations of current accessibility scanning tools.
Amanda Swearngin, Jason Wu 0001, Xiaoyi Zhang 0006, Esteban Gomez, Jen Coughenour, Rachel Stukenborg, Bhavya Garg, Greg Hughes, Adriana Hilliard, Jeffrey P. Bigham, Jeffrey Nichols 0001
ACM Trans. Comput. Hum. Interact.3
2022 Towards Complete Icon Labeling in Mobile Applications
abstract
Accurately recognizing icon types in mobile applications is integral to many tasks, including accessibility improvement, UI design search, and conversational agents. Existing research focuses on recognizing the most frequent icon types, but these technologies fail when encountering an unrecognized low-frequency icon. In this paper, we work towards complete coverage of icons in the wild. After annotating a large-scale icon dataset (327,879 icons) from iPhone apps, we found a highly uneven distribution: 98 common icon types covered 92.8% of icons, while 7.2% of icons were covered by more than 331 long-tail icon types. In order to label icons with widely varying occurrences in apps, our system uses an image classification model to recognize common icon types with an average of 3,000 examples each (96.3% accuracy) and applies a few-shot learning model to classify long-tail icon types with an average of 67 examples each (78.6% accuracy). Our system also detects contextual information that helps characterize icon semantics, including nearby text (95.3% accuracy) and modifier symbols added to the icon (87.4% accuracy). In a validation study with workers (n = 23), we verified the usefulness of our generated icon labels. The icon types supported by our work cover 99.5% of collected icons, improving on the previously highest 78% coverage in icon classification work.
Jieshan Chen, Amanda Swearngin, Jason Wu 0001, Titus Barik, Jeffrey Nichols 0001, Xiaoyi Zhang 0006
CHI6
2022 Understanding Screen Relationships from Screenshots of Smartphone Applications
abstract
All graphical user interfaces are comprised of one or more screens that may be shown to the user depending on their interactions. Identifying different screens of an app and understanding the type of changes that happen on the screens is a challenging task that can be applied in many areas including automatic app crawling, playback of app automation macros and large scale app dataset analysis. For example, an automated app crawler needs to understand if the screen it is currently viewing is the same as any previous screen that it has encountered, so it can focus its efforts on portions of the app that it has not yet explored. Moreover, identifying the type of change on the screen, such as whether any dialogues or keyboards have opened or closed, is useful for an automatic crawler to handle such events while crawling. Understanding screen relationships is a difficult task as instances of the same screen may have visual and structural variation, for example due to different content in a database-backed application, scrolling, dialog boxes opening or closing, or content loading delays. At the same time, instances of different screens from the same app may share some similarities in terms of design, structure, and content. This paper uses a dataset of screenshots from more than 1K iPhone applications to train two ML models that understand similarity in different ways: (1) a screen similarity model that combines a UI object detector with a transformer model architecture to recognize instances of the same screen from a collection of screenshots from a single app, and (2) a screen transition model that uses a siamese network architecture to identify both similarity and three types of events that appear in an interaction trace: the keyboard or a dialog box appearing or disappearing, and scrolling. Our models achieve an F1 score of 0.83 on the screen similarity task, improving on comparable baselines, and an average F1 score of 0.71 across all events in the transition task.
Shirin Feiz, Jason Wu 0001, Xiaoyi Zhang 0006, Amanda Swearngin, Titus Barik, Jeffrey Nichols 0001
IUI3
2021 Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels
abstract
Many accessibility features available on mobile platforms require applications (apps) to provide complete and accurate metadata describing user interface (UI) components. Unfortunately, many apps do not provide sufficient metadata for accessibility features to work as expected. In this paper, we explore inferring accessibility metadata for mobile apps from their pixels, as the visual interfaces often best reflect an app’s full functionality. We trained a robust, fast, memory-efficient, on-device model to detect UI elements using a dataset of 77,637 screens (from 4,068 iPhone apps) that we collected and annotated. To further improve UI detections and add semantic information, we introduced heuristics (e.g., UI grouping and ordering) and additional models (e.g., recognize UI content, state, interactivity). We built Screen Recognition to generate accessibility metadata to augment iOS VoiceOver. In a study with 9 screen reader users, we validated that our approach improves the accessibility of existing mobile apps, enabling even previously inaccessible apps to be used.
Xiaoyi Zhang 0006, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle I. Murray, Lisa Yu, Qi Shan, Jeffrey Nichols 0001, Jason Wu 0001, Chris Fleizach, Aaron Everitt, Jeffrey P. Bigham
CHI1
2021 Screen Parsing: Towards Reverse Engineering of UI Models from Screenshots
abstract
Automated understanding of user interfaces (UIs) from their pixels can improve accessibility, enable task automation, and facilitate interface design without relying on developers to comprehensively provide metadata. A first step is to infer what UI elements exist on a screen, but current approaches are limited in how they infer how those elements are semantically grouped into structured interface definitions. In this paper, we motivate the problem of screen parsing, the task of predicting UI elements and their relationships from a screenshot. We describe our implementation of screen parsing and provide an effective training procedure that optimizes its performance. In an evaluation comparing the accuracy of the generated output, we find that our implementation significantly outperforms current systems (up to 23%). Finally, we show three example applications that are facilitated by screen parsing: (i) UI similarity search, (ii) accessibility enhancement, and (iii) code generation from UI screenshots.
Jason Wu 0001, Xiaoyi Zhang 0006, Jeffrey Nichols 0001, Jeffrey P. Bigham
UIST2
2020 Towards Recommending Accessibility Features on Mobile Devices
abstract
Numerous accessibility features have been developed to increase who and how people can access computing devices. Increasingly, these features are included as part of popular platforms, e.g., Apple iOS, Google Android, and Microsoft Windows. Despite their potential to improve the computing experience, many users are unaware of these features and do not know which combination of them could benefit them. In this work, we first quantified this problem by surveying 100 participants online (including 25 older adults) about their knowledge of accessibility and features that they could benefit from, showing very low awareness. We developed four prototypes spanning numerous accessibility categories (e.g., vision, hearing, motor), that embody signals and detection strategies applicable to accessibility recommendation in general. Preliminary results from a study with 20 older adults show that proactive recommendation is a promising approach for better pairing users with accessibility features they could benefit from.
Jason Wu 0001, Gabriel Reyes, Sam C. White, Xiaoyi Zhang 0006, Jeffrey P. Bigham
ASSETS4
2018 Examining Image-Based Button Labeling for Accessibility in Android Apps through Large-Scale Analysis
abstract
We conduct the first large-scale analysis of the accessibility of mobile apps, examining what unique insights this can provide into the state of mobile app accessibility. We analyzed 5,753 free Android apps for label-based accessibility barriers in three classes of image-based buttons: Clickable Images, Image Buttons, and Floating Action Buttons. An epidemiology-inspired framework was used to structure the investigation. The population of free Android apps was assessed for label-based inaccessible button diseases. Three determinants of the disease were considered: missing labels, duplicate labels, and uninformative labels. The prevalence, or frequency of occurrences of barriers, was examined in apps and in classes of image-based buttons. In the app analysis, 35.9% of analyzed apps had 90% or more of their assessed image-based buttons labeled, 45.9% had less than 10% of assessed image-based buttons labeled, and the remaining apps were relatively uniformly distributed along the proportion of elements that were labeled. In the class analysis, 92.0% of Floating Action Buttons were found to have missing labels, compared to 54.7% of Image Buttons and 86.3% of Clickable Images. We discuss how these accessibility barriers are addressed in existing treatments, including accessibility development guidelines.
Anne Spencer Ross, Xiaoyi Zhang 0006, James Fogarty, Jacob O. Wobbrock
ASSETS2
2018 Interactiles: 3D Printed Tactile Interfaces to Enhance Mobile Touchscreen Accessibility
abstract
The absence of tactile cues such as keys and buttons makes touchscreens difficult to navigate for people with visual impairments. Increasing tactile feedback and tangible interaction on touchscreens can improve their accessibility. However, prior solutions have either required hardware customization or provided limited functionality with static overlays. Prior investigation of tactile solutions for large touchscreens also may not address the challenges on mobile devices. We therefore present Interactiles, a low cost, portable, and unpowered system that enhances tactile interaction on Android touchscreen phones. Interactiles consists of 3D-printed hardware interfaces and software that maps interaction with that hardware to manipulation of a mobile app. The system is compatible with the built-in screen reader without requiring modification of existing mobile apps. We describe the design and implementation of Interactiles, and we evaluate its improvement in task performance and the user experience it enables with people who are blind or have low vision.
Xiaoyi Zhang 0006, Tracy Tran, Yuqian Sun, Ian Culhane, Shobhit Jain, James Fogarty, Jennifer Mankoff
ASSETS1
2018 Robust Annotation of Mobile Application Interfaces in Methods for Accessibility Repair and Enhancement
abstract
Accessibility issues in mobile apps make those apps difficult or impossible to access for many people. Examples include elements that fail to provide alternative text for a screen reader, navigation orders that are difficult, or custom widgets that leave key functionality inaccessible. Social annotation techniques have demonstrated compelling approaches to such accessibility concerns in the web, but have been difficult to apply in mobile apps because of the challenges of robustly annotating interfaces. This research develops methods for robust annotation of mobile app interface elements. Designed for use in runtime interface modification, our methods are based in screen identifiers, element identifiers, and screen equivalence heuristics. We implement initial developer tools for annotating mobile app accessibility metadata, evaluate our current screen equivalence heuristics in a dataset of 2038 screens collected from 50 mobile apps, present three case studies implementing runtime repair of common accessibility issues, and examine repair of real-world accessibility issues in 26 apps. These contributions overall demonstrate strong opportunities for social annotation in mobile accessibility.
Xiaoyi Zhang 0006, Anne Spencer Ross, James Fogarty
UIST1
2018 APPINITE: A Multi-Modal Interface for Specifying Data Descriptions in Programming by Demonstration Using Natural Language Instructions
abstract
A key challenge for generalizing programming-by-demonstration (PBD) scripts is the data description problem - when a user demonstrates performing an action, the system needs to determine features for describing this action and the target object in a way that can reflect the user's intention for the action. However, prior approaches for creating data descriptions in PBD systems have problems with usability, applicability, feasibility, transparency and/or user control. Our APPINITE system introduces a multimodal interface with which users can specify data descriptions verbally using natural language instructions. APPINITE guides users to describe their intentions for the demonstrated actions through mixed-initiative conversations. APPINITE constructs data descriptions for these actions from the natural language instructions. Our evaluation showed that APPINITE is easy-to-use and effective in creating scripts for tasks that would otherwise be difficult to create with prior PBD systems, due to ambiguous data descriptions in demonstrations on GUIs.
Toby Jia-Jun Li, Igor Labutov, Xiaohan Nancy Li, Xiaoyi Zhang 0006, Wenze Shi, Wanling Ding, Tom M. Mitchell, Brad A. Myers
VL/HCC4
2017 Epidemiology as a Framework for Large-Scale Mobile Application Accessibility Assessment
abstract
Mobile accessibility is often a property considered at the level of a single mobile application (app), but rarely on a larger scale of the entire app "ecosystem," such as all apps in an app store, their companies, developers, and user influences. We present a novel conceptual framework for the accessibility of mobile apps inspired by epidemiology. It considers apps within their ecosystems, over time, and at a population level. Under this metaphor, "inaccessibility" is a set of diseases that can be viewed through an epidemiological lens. Accordingly, our framework puts forth notions like risk and protective factors, prevalence, and health indicators found within a population of apps. This new framing offers terminology, motivation, and techniques to reframe how we approach and measure app accessibility. It establishes how app accessibility can benefit from multi-factor, longitudinal, and population-based analyses. Our epidemiology-inspired conceptual framework is the main contribution of this work, intended to provoke thought and inspire new work enhancing app accessibility at a systemic level. In a preliminary exercising of our framework, we perform an analysis of the prevalence of common determinants or accessibility barriers. We assess the health of a stratified sample of 100 popular Android apps using Google's Accessibility Scanner. We find that 100% of apps have at least one of nine accessibility errors and examine which errors are most common. A preliminary analysis of the frequency of co-occurrences of multiple errors in a single app is also presented. We find 72% of apps have five or six errors, suggesting an interaction among different errors or an underlying influence.
Anne Spencer Ross, Xiaoyi Zhang 0006, James Fogarty, Jacob O. Wobbrock
ASSETS2
2017 Smartphone-Based Gaze Gesture Communication for People with Motor Disabilities
abstract
Current eye-tracking input systems for people with ALS or other motor impairments are expensive, not robust under sunlight, and require frequent re-calibration and substantial, relatively immobile setups. Eye-gaze transfer (e-tran) boards, a low-tech alternative, are challenging to master and offer slow communication rates. To mitigate the drawbacks of these two status quo approaches, we created GazeSpeak, an eye gesture communication system that runs on a smartphone, and is designed to be low-cost, robust, portable, and easy-to-learn, with a higher communication bandwidth than an e-tran board. GazeSpeak can interpret eye gestures in real time, decode these gestures into predicted utterances, and facilitate communication, with different user interfaces for speakers and interpreters. Our evaluations demonstrate that GazeSpeak is robust, has good user satisfaction, and provides a speed improvement with respect to an e-tran board; we also identify avenues for further improvement to low-cost, low-effort gaze-based communication technologies.
Xiaoyi Zhang 0006, Harish Kulkarni, Meredith Ringel Morris
CHI1
2017 Interaction Proxies for Runtime Repair and Enhancement of Mobile Application Accessibility
abstract
We introduce interaction proxies as a strategy for runtime repair and enhancement of the accessibility of mobile applications. Conceptually, interaction proxies are inserted between an application's original interface and the manifest interface that a person uses to perceive and manipulate the application. This strategy allows third-party developers and researchers to modify an interaction without an application's source code, without rooting the phone, without otherwise modifying an application, while retaining all capabilities of the system (e.g., Android's full implementation of the TalkBack screen reader). This paper introduces interaction proxies, defines a design space of interaction re-mappings, identifies necessary implementation abstractions, presents details of implementing those abstractions in Android, and demonstrates a set of Android implementations of interaction proxies from throughout our design space. We then present a set of interviews with blind and low-vision people interacting with our prototype interaction proxies, using these interviews to explore the seamlessness of interaction, the perceived usefulness and potential of interaction proxies, and visions of how such enhancements could gain broad usage. By allowing third-party developers and researchers to improve an interaction, interaction proxies offer a new approach to personalizing mobile application accessibility and a new approach to catalyzing development, deployment, and evaluation of mobile accessibility enhancements.
Xiaoyi Zhang 0006, Anne Spencer Ross, Anat Caspi, James Fogarty, Jacob O. Wobbrock
CHI1
2016 Examining Unlock Journaling with Diaries and Reminders for In Situ Self-Report in Health and Wellness
abstract
In situ self-report is widely used in human-computer interaction, ubiquitous computing, and for assessment and intervention in health and wellness. Unfortunately, it remains limited by high burdens. We examine unlock journaling as an alternative. Specifically, we build upon recent work to introduce single-slide unlock journaling gestures appropriate for health and wellness measures. We then present the first field study comparing unlock journaling with traditional diaries and notification-based reminders in self-report of health and wellness measures. We find unlock journaling is less intrusive than reminders, dramatically improves frequency of journaling, and can provide equal or better timeliness. Where appropriate to broader design needs, unlock journaling is thus an overall promising method for in situ self-report.
Xiaoyi Zhang 0006, Laura R. Pina, James Fogarty
CHI1
2015 Leveraging Dual-Observable Input for Fine-Grained Thumb Interaction Using Forearm EMG
abstract
We introduce the first forearm-based EMG input system that can recognize fine-grained thumb gestures, including left swipes, right swipes, taps, long presses, and more complex thumb motions. EMG signals for thumb motions sensed from the forearm are quite weak and require significant training data to classify. We therefore also introduce a novel approach for minimally-intrusive collection of labeled training data for always-available input devices. Our dual-observable input approach is based on the insight that interaction observed by multiple devices allows recognition by a primary device (e.g., phone recognition of a left swipe gesture) to create labeled training examples for another (e.g., forearm-based EMG data labeled as a left swipe). We implement a wearable prototype with dry EMG electrodes, train with labeled demonstrations from participants using their own phones, and show that our prototype can recognize common fine-grained thumb gestures and user-defined complex gestures.
Donny Huang, Xiaoyi Zhang 0006, T. Scott Saponas, James Fogarty, Shyamnath Gollakota
UIST2
2015 BreathSens: A Continuous On-Bed Respiratory Monitoring System With Torso Localization Using an Unobtrusive Pressure Sensing Array
abstract
The ability to continuously monitor respiration rates of patients in homecare or in clinics is an important goal. Past research showed that monitoring patient breathing can lower the associated mortality rates for long-term bedridden patients. Nowadays, in-bed sensors consisting of pressure sensitive arrays are unobtrusive and are suitable for deployment in a wide range of settings. Such systems aim to extract respiratory signals from time-series pressure sequences. However, variance of movements, such as unpredictable extremities activities, affect the quality of the extracted respiratory signals. BreathSens, a high-density pressure sensing system made of e-Textile, profiles the underbody pressure distribution and localizes torso area based on the high-resolution pressure images. With a robust bodyparts localization algorithm, respiratory signals extracted from the localized torso area are insensitive to arbitrary extremities movements. In a study of 12 subjects, BreathSens demonstrated its respiratory monitoring capability with variations of sleep postures, locations, and commonly tilted clinical bed conditions.
Jason J. Liu, Ming-Chun Huang, Wenyao Xu, Xiaoyi Zhang 0006, Luke Stevens, Nabil Alshurafa, Majid Sarrafzadeh
IEEE J. Biomed. Health Informatics4
2014 Using Pressure Map Sequences for Recognition of On Bed Rehabilitation Exercises
abstract
Physical rehabilitation is an important process for patients recovering after surgery. In this paper, we propose and develop a framework to monitor on-bed range of motion exercises that allows physical therapists to evaluate patient adherence to set exercise programs. Using a dense pressure sensitive bedsheet, a sequence of pressure maps are produced and analyzed using manifold learning techniques. We compare two methods, Local Linear Embedding and Isomap, to reduce the dimensionality of the pressure map data. Once the image sequences are converted into a low dimensional manifold, the manifolds can be compared to expected prior data for the rehabilitation exercises. Furthermore, a measure to compare the similarity of manifolds is presented along with experimental results for five on-bed rehabilitation exercises. The evaluation of this framework shows that exercise compliance can be tracked accurately according to prescribed treatment programs.
Ming-Chun Huang, Jason J. Liu, Wenyao Xu, Nabil Alshurafa, Xiaoyi Zhang 0006, Majid Sarrafzadeh
IEEE J. Biomed. Health Informatics5