Zhuohao (Jerry) Zhang

dblp:344/8602 · also Zhuohao Zhang 0001 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-8708-1429ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 14 · 7 first-author · 12 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hesitation and Tolerance in Recommender Systems
abstract
Users’ interactions with recommender systems often involve more than simple acceptance or rejection. We highlight two overlooked states: hesitation, when people deliberate without certainty, and tolerance, when this hesitation escalates into unwanted engagement before ending in disinterest. Across two large-scale surveys (N = 6, 644 and N = 3, 864), hesitation was nearly universal, and tolerance emerged as a recurring source of wasted time, frustration, and diminished trust. Analyses of e-commerce and short-video platforms confirm that tolerance behaviors, such as clicking without purchase or shallow viewing, correlate with decreased activity. Finally, an online field study at scale shows that even lightweight strategies treating tolerance as distinct from interest can improve retention while reducing wasted effort. By surfacing hesitation and tolerance as consequential states, this work reframes how recommender systems should interpret feedback, moving beyond clicks and dwell time toward designs that respect user value, reduce hidden costs, and sustain engagement.
Kuan Zou, Aixin Sun, Yitong Ji, Hao Zhang 0048, Jing Wang 0060, Zhuohao (Jerry) Zhang, Xuemeng Jiang
CHI6
2025 A11yShape: AI-Assisted 3-D Modeling for Blind and Low-Vision Programmers
abstract
Figure 1: With A11yShape, (A) a blind or low-vision (BLV) user can create, interpret, and verify 3-D models through (B) a user interface composed of three parts: Code Editor Panel, AI Assistant Panel, and Model Panel.These panels are linked by a cross-representation highlighting mechanism that connects code, textual descriptions, hierarchical model abstractions, and 3-D visual renderings.The system supports the creation of (C) diverse, customized 3-D models created by BLV users.
Zhuohao (Jerry) Zhang, Haichang Li, Chun Meng Yu, Faraz Faruqi, Junan Xie, Gene S.-H. Kim, Mingming Fan 0001, Angus G. Forbes, Jacob O. Wobbrock, Anhong Guo, Liang He 0005
ASSETS1
2025 VizXpress: Towards Expressive Visual Content by Blind Creators Through AI Support
abstract
From curating the layout of a resume to selecting filters for social media, creating and configuring visual content allows individuals to express identity, communicate intent, and engage socially, yet blind individuals often face significant barriers to such expressive practices.Prior accessibility research primarily addresses functional content configuration, leaving little understanding of blind individuals' expressive visual creation needs.To better understand and support these needs, we conducted a two-stage study: first, we interviewed 10 blind participants to understand their motivations, current practices, and barriers in visual expression, and to ideate on potential visual editing support; second, based on interview insights, we developed an interactive prototype (VizXpress) that provides real-time feedback on visual aesthetics using a vision-language model and supports automated and manual visual editing controls.We used VizXpress as a design probe to further explore accessible design opportunities for visual expression.Our findings highlight many blind users' strong interest in creating visually expressive content, nuanced informational requirements for subjective aesthetics (e.g., color, mood, lighting), and ongoing accessibility challenges with visual creative tools.Grounded in these insights, we propose design implications including richer aesthetic feedback, controlled intelligent editing, and accessible manual editing mechanisms.
Lotus Hanzi Zhang, Zhuohao (Jerry) Zhang, Gina Clepper, Franklin Mingzhe Li, Patrick Carrington, Jacob O. Wobbrock, Leah Findlater
ASSETS2
2025 OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
Jiahao Nick Li, Zhuohao (Jerry) Zhang, Jiaju Ma
CHI2
2025 From Interaction to Impact: Towards Safer AI Agent Through Understanding and Evaluating Mobile UI Operation Impacts
Zhuohao (Jerry) Zhang, Eldon Schoop, Jeffrey Nichols 0001, Anuj Mahajan, Amanda Swearngin
IUI1
2025 SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation Slides
Zhuohao (Jerry) Zhang, Ruiqi Chen 0004, Mingyuan Zhong 0001, Jacob O. Wobbrock
UIST1
2024 ChartA11y: Designing Accessible Touch Experiences of Visualizations with Blind Smartphone Users
abstract
We introduce ChartA11y, an app developed to enable accessible 2-D visualizations on smartphones for blind users through a participatory and iterative design process involving 13 sessions with two blind partners. We also present a design journey for making accessible touch experiences that go beyond simple auditory feedback, incorporating multimodal interactions and multisensory data representations. Together, ChartA11y aimed at providing direct chart accessing and comprehensive chart understanding by applying a two-mode setting: a semantic navigation framework mode and a direct touch mapping mode. By re-designing traditional touch-to-audio interactions, ChartA11y also extends to accessible scatter plots, addressing the under-explored challenges posed by their non-linear data distribution. Our main contributions encompass the detailed participatory design process and the resulting system, ChartA11y, offering a novel approach for blind users to access visualizations on their smartphones.
Zhuohao (Jerry) Zhang, John Thompson 0002, Aditi Shah, Manish Agrawal, Alper Sarikaya 0001, Jacob O. Wobbrock, Edward Cutrell, Bongshin Lee
ASSETS1
2023 Developing and Deploying a Real-World Solution for Accessible Slide Reading and Authoring for Blind Users
abstract
Presentation software like Microsoft PowerPoint and Google Slides remains largely inaccessible for blind users because screen readers are not well suited to 2-D “artboards” that contain different objects in arbitrary arrangements lacking any inherent reading order. To investigate this problem, prior work by Zhang & Wobbrock (2023) developed multimodal interaction techniques in a prototype system called A11yBoard, but their system was limited to a single artboard in a self-contained prototype and was unable to support real-world use. In this work, we present a major extension of A11yBoard that expands upon its initial interaction techniques, addresses numerous real-world issues, and makes it deployable with Google Slides. We describe the new features developed for A11yBoard for Google Slides along with our participatory design process with a blind co-author. We also present two case studies based on real-world deployments showing that participants were able to independently complete slide reading and authoring tasks that were not possible without sighted assistance previously. We conclude with several design guidelines for making accessible digital content creation tools.
Zhuohao (Jerry) Zhang, Gene S.-H. Kim, Jacob O. Wobbrock
ASSETS1
2023 A11yBoard: Making Digital Artboards Accessible to Blind and Low-Vision Users
abstract
Digital artboards, which hold objects rather than pixels (e.g., Microsoft PowerPoint and Google Slides), remain largely inaccessible for blind and low-vision (BLV) users. Building on prior findings about the experiences of BLV users with digital artboards, we present a novel tool called A11yBoard, an interactive multimodal system that makes interpreting and authoring digital artboards accessible. A11yBoard combines a web-based drawing canvas paired with a mobile touch screen device such as a tablet. The mobile device displays the same canvas and enables risk-free spatial exploration of the artboard via touch and gesture. Speech recognition, non-speech audio, and keyboard-based commands are also used for input and output. Through a series of pilot studies and formal task-based user studies with BLV participants, we show that A11yBoard provides (1) intuitive spatial reasoning about two-dimensional objects, (2) multimodal access to objects’ properties and relationships, and (3) eyes-free creating and editing of objects to establish their desired properties and positions.
Zhuohao (Jerry) Zhang, Jacob O. Wobbrock
CHI1
2023 ImageAlly: A Human-AI Hybrid Approach to Support Blind People in Detecting and Redacting Private Image Content
Zhuohao (Jerry) Zhang, Smirity Kaushik, Jooyoung Seo, Haolin Yuan, Sauvik Das, Leah Findlater, Danna Gurari, Abigale Stangl, Yang Wang 0005
SOUPS1
2022 CollabAlly: Accessible Collaboration Awareness in Document Editing
abstract
Collaborative document editing tools are widely used in professional and academic workplaces. While these tools provide basic accessibility support, it is challenging for blind users to gain collaboration awareness that sighted people can easily obtain using visual cues (e.g., who is editing where and what). Through a series of co-design sessions with a blind coauthor, we identified the current practices and challenges in collaborative editing, and iteratively designed CollabAlly, a system that makes collaboration awareness in document editing accessible to blind users. CollabAlly extracts collaborator, comment, and text-change information and their context from a document and presents them in a dialog box to provide easy access and navigation. CollabAlly uses earcons to communicate background events unobtrusively, voice fonts to differentiate collaborators, and spatial audio to convey the location of document activity. In a study with 11 blind participants, we demonstrate that CollabAlly provides improved access to collaboration awareness by centralizing scattered information, sonifying visual information, and simplifying complex operations.
Cheuk Yin Phipson Lee, Zhuohao (Jerry) Zhang, Jaylin Herskovitz, Jooyoung Seo, Anhong Guo
CHI2
2021 CollabAlly: Accessible Collaboration Awareness in Document Editing
abstract
Collaborative document editing tools are widely used in both professional and academic workplaces. While these tools provide some accessibility features, it is still challenging for blind users to gain collaboration awareness that sighted people can easily obtain using visual cues (e.g., who edited or commented where and what in the document). To address this gap, we present CollabAlly, a browser extension that makes extractable collaborative and contextual information in document editing accessible for blind users. With CollabAlly, blind users can easily access collaborators’ information, track real-time or asynchronous content and comment changes, and navigate through these elements. In order to convey this complex information through audio, CollabAlly uses voice fonts and spatial audio to enhance users’ collaboration awareness in shared documents. Through a series of pilot studies with a coauthor who is blind, CollabAlly’s design was refined to include more information and to be more compatible with existing screen readers.
Cheuk Yin Phipson Lee, Zhuohao (Jerry) Zhang, Jaylin Herskovitz, Jooyoung Seo, Anhong Guo
ASSETS2
2019 Designing Interactive 3D Printed Models with Teachers of the Visually Impaired
abstract
Students with visual impairments struggle to learn various concepts in the academic curriculum because diagrams, images, and other visual are not accessible to them. To address this, researchers have design interactive 3D printed models (I3Ms) that provide audio descriptions when a user touches components of a model. In prior work, I3Ms were designed on an ad hoc basis, and it is currently unknown what general guidelines produce effective I3M designs. To address this gap, we conducted two studies with Teachers of the Visually Impaired (TVIs). First, we led two design workshops with 35 TVIs, who modified sample models and added interactive elements to them. Second, we worked with three TVIs to design three I3Ms in an iterative instructional design process. At the end of this process, the TVIs used the I3Ms we designed to teach their students. We conclude that I3Ms should (1) have effective tactile features (e.g., distinctive patterns between components), (2) contain both auditory and visual content (e.g., explanatory animations), and (3) consider pedagogical methods (e.g., overview before details).
Lei Shi 0020, Holly Lawson, Zhuohao (Jerry) Zhang, Shiri Azenkot
CHI3
2018 A Demo of Talkit++: Interacting with 3D Printed Models Using an iOS Device
abstract
Tactile models are important learning materials for visually impaired students. With the adoption of 3D printing technologies, visually impaired students and teachers will have more access to 3D printed tactile models. We designed Talkit++, an iOS application that plays audio and visual content as a user touches parts of a 3D print. With Talkit++, a visually impaired student can explore a printed model tactilely, and use finger gestures and speech commands to get more information about certain elements in the model. Talkit++ detects the model and finger gestures using computer vision algorithms, simple accessories like paper stickers and printable trackers, and the built-in RGB camera on an iOS device. Based on the model's position and the user's input, Talkit++ speaks textual information, plays audio recordings, and displays visual animations.
Lei Shi 0020, Zhuohao (Jerry) Zhang, Shiri Azenkot
ASSETS2