Shengdong Zhao 0001

dblp:23/1695 · DBLP profile ↗
← Back
97ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0001-7971-3107ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 84 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 Spatial Balancing: Designing an LLM-Powered Spatial Externalization Interface for Iterative Science Communication Writing
abstract
Science communication revision requires writers to dynamically balance scientific exposition and narrative engagement - a process where writers often struggle with competing directions. Existing LLM-assisted tools help with co-writing, but offer limited support for navigating this iterative, multi-directional revision process. To address this gap, we designed Spatial Balancing, an exploratory revision environment that maps rhetorical goals and revision strategies onto a two-dimensional spatial canvas for experienced science communication creators with domain expertise but lacking formal professional training. By building a design space of communication strategies and embedding them into a spatial exploratory canvas, our system treats feedback as navigational cues rather than prescriptive judgments. Our findings show that this integrated revision environment helps writers stay focused on writing goals, reason about revision as trajectories, and explore alternatives, which supports greater metacognitive control and confidence without increasing workload. This work highlights the value of spatially externalized revision environments for supporting iterative, reflective thinking during LLM-assisted writing.
Kexue Fu 0002, Jiaye Leng, Jingfei Huang, Yihang Zuo, Runze Cai, Zijian Ding, Ray LC, Shengdong Zhao 0001, Qinyuan Lei
DIS9
2026 Wearable AR for Restorative Breaks: How Interactive Narrative Experiences Support Relaxation for Young People
abstract
Young adults often take breaks from screen-intensive work by consuming digital content on mobile phones, which undermines rest through visual fatigue and inactivity. We introduce a design framework that embeds light break activities into media content on AR smart glasses, balancing engagement and recovery, which employs three strategies: (1) seamlessly guiding users by embedding activity cues aligned with media elements; (2) transitioning to audio-centric formats to reduce visual load while sustaining immersion; and (3) structuring sessions with "rise-peak-closure"pacing for smooth transitions. In a within-subjects study (N=16) comparing passive viewing, reminder-based breaks, and non-narrative activities, InteractiveBreak instantiated from our framework seamlessly guided activities, sustained engagement, and enhanced break quality. These findings demonstrate wearable AR's potential to support restorative relaxation by transforming breaks into engaging, meaningful experiences. © 2026 the owner/author(s).
Jin-Du Wang, Runze Cai, Shuchang Xu, Tianrui Hu, Huamin Qu, Shengdong Zhao 0001, Linping Yuan
CHI6
2026 PersonaMail: Learning and Adapting Personal Communication Preferences for Context-Aware Email Writing
abstract
LLM-assisted writing has seen rapid adoption in interpersonal communication, yet current systems often fail to capture the subtle tones essential for effectiveness. Email writing exemplifies this challenge: effective messages require careful alignment with intent, relationship, and context beyond mere fluency. Through formative studies, we identified three key challenges: articulating nuanced communicative intent, making modifications at multiple levels of granularity, and reusing effective tone strategies across messages. We developed PersonaMail, a system that addresses these gaps through structured communication factor exploration, granular editing controls, and adaptive reuse of successful strategies. Our evaluation compared PersonaMail against standard LLM interfaces, and showed improved efficiency in both immediate and repeated use, alongside higher user satisfaction. We contribute design implications for AI-assisted communication systems that prioritize interpersonal nuance over generic text generation.
Qiuyuan Ren, Felicia Fang-Yi Tan, Yang Chen 0054, Xiaoyu Zhang 0014, Shengdong Zhao 0001
IUI6
2026 FanType: Intention-Inferring Fan-Shaped Thumb Interface for Text Entry on Small XR Keyboards
abstract
In this paper, we present FanType, a text entry technique that enables efficient typing on small virtual keyboards in extended reality (XR), closely resembling thumb-based typing on smartphones and tablets. While virtual environments theoretically provide unlimited space, practical XR usage often occurs in mobile and spatially constrained contexts, such as when using augmented reality (AR) headsets while walking or in confined environments, where large virtual keyboards are impractical, as they occupy substantial visual space and interfere with other virtual or real content. To address this, FanType integrates a fan-shaped interface that groups keys within the thumb's reach and leverages a lift-up gesture for disambiguation, thereby reducing hand movement while preserving natural typing speed. Additionally, an intention-inferring mechanism involving tailored algorithms is implemented to counterbalance touch bias and further ensure input accuracy. A user study with 22 participants evaluated the method on tablet- and smartphone-size portrait keyboards, comparing it with a baseline (touch-only) and a state-of-the-art mid-air technique. Results provide insights into the performance of such a design and highlight the challenges of designing typing methods for small virtual keyboards.
Guanghan Zhao, Louis Teys, Gyeonghwan Yang, Shengdong Zhao 0001, Yoshifumi Kitamura
IEEE Trans. Vis. Comput. Graph.4
2025 AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart Glasses
abstract
Unlike the free exploration of childhood, the demands of daily life reduce our motivation to explore our surroundings, leading to missed opportunities for informal learning. Traditional tools for knowledge acquisition are reactive, relying on user initiative and limiting their ability to uncover hidden interests. Through formative studies, we introduce AiGet, a proactive AI assistant integrated with AR smart glasses, designed to seamlessly embed informal learning into low-demand daily activities (e.g., casual walking and shopping). AiGet analyzes real-time user gaze patterns, environmental context, and user profiles, leveraging large language models to deliver personalized, context-aware knowledge with low disruption to primary tasks. In-lab evaluations and real-world testing, including continued use over multiple days, demonstrate AiGet's effectiveness in uncovering overlooked yet surprising interests, enhancing primary task enjoyment, reviving curiosity, and deepening connections with the environment. We further propose design guidelines for AI-assisted informal learning, focused on transforming everyday moments into enriching learning experiences. © 2025 Copyright held by the owner/author(s).
Runze Cai, Nuwan Janaka, Hyeongcheol Kim 0001, Yang Chen 0054, Shengdong Zhao 0001, Yun Huang 0003, David Hsu
CHI5
2025 ViFeed: Promoting Slow Eating and Food Awareness through Strategic Video Manipulation during Screen-Based Dining
abstract
Given the widespread presence of screens during meals, the notion that digital engagement is inherently incompatible with mindfulness. We demonstrate how the strategic design of digital content can enhance two core aspects of mindful eating: slow eating and food awareness. Our research unfolded in three sequential studies: (1). Zoom Eating Study: Contrary to the assumption that video-watching leads to distraction and overeating, this study revealed that subtle video speed manipulations—can promote slower eating (by 15.31%) and controlled food intake (by 9.65%) while maintaining meal satiation and satisfaction. (2). Co-design workshop: Informed the development of ViFeed, a video playback system strategically incorporating subtle speed adjustments and glanceable visual cues. (3). Field Study: A week-long deployment of ViFeed in daily eating demonstrated its efficacy in fostering food awareness, food appreciation, and sustained engagement. By bridging the gap between ideal mindfulness practices and screen-based behaviors, this work offers insights for designing digital-wellbeing interventions that align with, rather than against, existing habits. © 2025 Copyright held by the owner/author(s).
Yang Chen 0054, Felicia Fang-Yi Tan, Zhuoyu Wang 0001, Jiayi Zhang 0013, Yun Huang 0003, Shengdong Zhao 0001, Ching Chiuan Yen
CHI7
2025 Designing Interactive Multimodal Information Retrieval and Access for Heads Up Computing (DIMIRA-HUC)
abstract
The advancement of wearable intelligent systems presents a unique opportunity to transform how humans interact with digital content. This workshop explores the design of Interactive Multimodal Information Retrieval and Access systems specifically tailored for Heads-Up Computing environments. By leveraging multimodal inputs, such as voice, gaze, and gesture, these systems enable real-time, hands-free access to digital information, facilitating seamless and efficient interaction. The goal is to support tasks requiring rapid information access in dynamic environments while ensuring users remain "heads-up" and engaged with the real world. This half-day workshop will share research outcomes and best practices, foster community building, and facilitate discussions on key challenges. By bringing together researchers and practitioners, it aims to drive further advancements in both research and practical applications within this rapidly evolving field.
Haiming Liu 0002, Shengdong Zhao 0001, Silang Wang, Preben Hansen, Ian Oakley, Khanh-Duy Le
CHIIR2
2025 From Simple to Polychromatic: An Empirical Study on Optimal Color Schemes for Optical See-Through Head-Mounted Displays
abstract
Optical see-through head-mounted displays (OHMDs) blend digital content with the physical world, presenting unique color management challenges. Previous literature suggests using green as the main color, but this severely limits creative freedom. To address this, we conducted an empirical study with 30 participants, evaluating 216 colors under various OHMD usage conditions. Based on the results, we propose color guidelines indicating each hue's clear and comfortable saturation and brightness ranges, along with clarity and comfort scores across hues for different devices and lighting conditions. Our color guidelines expand the usable color palette, offering designers a wider range of color options. These guidelines were used and iteratively refined through feedback in a workshop with 12 designers, integrating them into practical design workflows. The resulting comprehensive color guide provides a valuable resource for OHMD interface designers, enhancing both the aesthetic possibilities and functional effectiveness of augmented reality experiences.
Runze Cai, Ashwin Ram 0002, Haimo Zhang, Shengdong Zhao 0001
IEEE Trans. Vis. Comput. Graph.6
2025 SummonBrush: Enhancing Touch Interaction on Large XR User Interfaces by Augmenting Users' Hands with Virtual Brushes
abstract
Touch interaction is one of the fundamental interaction paradigms in XR, as users have become very familiar with touch interactions on physical touchscreens. However, users typically need to perform extensive arm movements for engaging with XR user interfaces much larger than mobile device touchscreens. We propose the SummonBrush technique to facilitate easy access to hidden windows while interacting with large XR user interfaces, requiring minimal arm movements. The SummonBrush technique adds a virtual brush to the index fingertip of a user's hand. Upon making contact with a virtual user interface, the brush bends and diverges and ink starts to diffuse in it. The more the brush bends and diverges, the more the ink diffuses. The user can summon hidden windows or background applications in situ, which is achieved by firstly pressing the brush against the user interface to make ink fully fill the brush and then perform swipe gestures. Also, the user can press the brush against the thumbtails of background applications in situ to quickly cycle them through. Ecological studies showed that SummonBrush significantly reduced the arm movement time by 39% and 34% in summoning hidden windows and activating/closing background applications, respectively, leading to a significant decrease in reported physical demand.
Yang Tian 0008, Zhao Su, Tianren Luo, Teng Han, Shengdong Zhao 0001, Boyu Gao 0003, Dangxiao Wang
IEEE Trans. Vis. Comput. Graph.5
2025 AmplitudeArrow: On-the-Go AR Menu Selection Using Consecutive Simple Head Gestures and Amplitude Visualization
abstract
Heads-up computing aims to provide synergistic digital assistance that minimally interferes with users' on-the-go daily activities. Currently, the input modalities of heads-up computing are mainly voice and finger gestures. In this work, we propose and evaluate the AmplitudeArrow (AA) technique designed for on-the-go AR menu selection to demonstrate that consecutive simple head gestures can also be an effective input modality for heads-up computing. Specifically, AA arranges menu icons into one/two row(s). To select a target icon, the user first makes their head yaw to pre-select the target icon or the column containing it and then makes their head pitch to make the arrow in the target icon expand until the arrow covers the target icon completely, i.e., the pitch amplitude surpasses the selection confirmation threshold. User studies indicated that AA demonstrated robust resistance to walking-caused head perturbation and external factors such as other people/obstacles, delivering high accuracy (error rate $< $< 5$\%$%) and fast speed ($< $< 1.5s per selection) when there were no more than six icon columns (twelve icons) distributed horizontally and evenly in a menu area with a horizontal visual angle of $43^{\circ }$43∘.
Yang Tian 0008, Yukang Yan, Shengdong Zhao 0001, Xiaojuan Ma, Yuanchun Shi
IEEE Trans. Vis. Comput. Graph.4
2024 GlassMail: Towards Personalised Wearable Assistant for On-the-Go Email Creation on Smart Glasses
abstract
Optical See-through Head-Mounted Displays (OHMDs) offer new opportunities for completing complex information processing tasks on the go. We introduce GlassMail, a Large Language Models (LLMs)-based wearable assistant on OHMDs for mobile email creation. Our formative study identified two challenges of the LLM-based wearable email assistant: (i) achieving efficient and accurate understanding of user intentions, and (ii) ensuring effective information presentation for email processes. Through two empirical studies, we developed a "Single Turn with Optional Clarification " approach for accurate user intention recognition and a "Fade Context with Optional Audio " mode for effective email processing. An observation study then evaluated GlassMail ’s feasibility in composing formal and semi-formal emails, supporting the usefulness and effectiveness of GlassMail in simple scenarios and yielding insights into potential future improvements for complex scenarios. We further discuss the design implications for the future development of wearable AI-enabled assistants.
Ashwin Ram 0002, Can Liu 0003, Yun Huang 0003, Wei Tsang Ooi, Shengdong Zhao 0001
Conference on Designing Interactive Systems9
2024 Heads-Up Multitasker: Simulating Attention Switching On Optical Head-Mounted Displays
abstract
Optical Head-Mounted Displays (OHMDs) allow users to read digital content while walking. A better understanding of how users allocate attention between these two tasks is crucial for improving OHMD interfaces. This paper introduces a computational model for simulating users’ attention switches between reading and walking. We model users’ decision to deploy visual attention as a hierarchical reinforcement learning problem, wherein a supervisory controller optimizes attention allocation while considering both reading activity and walking safety. Our model simulates the control of eye movements and locomotion as an adaptation to the given task priority, design of digital content, and walking speed. The model replicates key multitasking behaviors during OHMD reading while walking, including attention switches, changes in reading and walking speeds, and reading resumptions.
Yunpeng Bai, Aleksi Ikkala, Antti Oulasvirta, Shengdong Zhao 0001, Lucia J. Wang, Pengzhi Yang, Peisen Xu
CHI4
2024 Navigating Real-World Challenges: A Quadruped Robot Guiding System for Visually Impaired People in Diverse Environments
abstract
Blind and Visually Impaired (BVI) people find challenges in navigating unfamiliar environments, even using assistive tools such as white canes or smart devices. Increasingly affordable quadruped robots offer us opportunities to design autonomous guides that could improve how BVI people find ways around unfamiliar environments and maneuver therein. In this work, we designed RDog, a quadruped robot guiding system that supports BVI individuals’ navigation and obstacle avoidance in indoor and outdoor environments. RDog combines an advanced mapping and navigation system to guide users with force feedback and preemptive voice feedback. Using this robot as an evaluation apparatus, we conducted experiments to investigate the difference in BVI people’s ambulatory behaviors using a white cane, a smart cane, and RDog. Results illustrated the benefits of RDog-based ambulation, including faster and smoother navigation with fewer collisions and limitations, and reduced cognitive load. We discuss the implications of our work for multi-terrain assistive guidance systems.
Shaojun Cai, Ashwin Ram 0002, Zhengtai Gou, Mohd Alqama Wasim Shaikh, Yu-An Chen, Yingjia Wan, Kotaro Hara, Shengdong Zhao 0001, David Hsu
CHI8
2024 PANDALens: Towards AI-Assisted In-Context Writing on OHMD During Travels
abstract
While effective for recording and sharing experiences, traditional in-context writing tools are relatively passive and unintelligent, serving more like instruments rather than companions. This reduces primary task (e.g., travel) enjoyment and hinders high-quality writing. Through formative study and iterative development, we introduce PANDALens, a Proactive AI Narrative Documentation Assistant built on an Optical See-Through Head Mounted Display that supports personalized documentation in everyday activities. PANDALens observes multimodal contextual information from user behaviors and environment to confirm interests and elicit contemplation, and employs Large Language Models to transform such multimodal information into coherent narratives with significantly reduced user effort. A real-world travel scenario comparing PANDALens with a smartphone alternative confirmed its effectiveness in improving writing quality and travel enjoyment while minimizing user effort. Accordingly, we propose design guidelines for AI-assisted in-context writing, highlighting the potential of transforming them from tools to intelligent companions.
Runze Cai, Nuwan Janaka, Yang Chen 0054, Lucia J. Wang, Shengdong Zhao 0001, Can Liu 0003
CHI5
2024 Facilitating Virtual Reality Integration in Medical Education: A Case Study of Acceptability and Learning Impact in Childbirth Delivery Training
abstract
Advancements in Virtual Reality (VR) technology have opened new frontiers in medical education, igniting interest among medical educators to incorporate it into mainstream curriculum, complementing traditional training modalities such as manikin training. Despite numerous VR simulators on the market, their uptake in medical education remains limited. This paper explores the acceptability and educational effectiveness of VR in the context of vaginal childbirth delivery training, with the simulator providing a walkthrough for the second and third stages of labour, contrasting it with established manikin-based methods. We conducted a large-scale empirical study with 117 medical students, revealing a significant 24.9% improvement in knowledge scores when using VR as compared to manikin. However, VR received significantly lower self-reported feasibility scores in Confidence, Usability, Enjoyment, Feedback and Presence, indicating low acceptance. The study provides critical insights into the relationship between technological innovation and educational impact, guiding future integration of VR into medical training curricula.
Chang Liu 0157, Felicia Fang-Yi Tan, Shengdong Zhao 0001, Abhiram Kanneganti, Gosavi Arundhati Tushar, Eng Tat Khoo
CHI3
2024 AudioXtend: Assisted Reality Visual Accompaniments for Audiobook Storytelling During Everyday Routine Tasks
abstract
The rise of multitasking in contemporary lifestyles has positioned audio-first content as an essential medium for information consumption. We present AudioXtend, an approach to augment audiobook experiences during daily tasks by integrating glanceable, AI-generated visuals through optical see-through head-mounted displays (OHMDs). Our initial study showed that these visual augmentations not only preserved users’ primary task efficiency but also dramatically enhanced immediate auditory content recall by 33.3% and 7-day recall by 32.7%, alongside a marked improvement in narrative engagement. Through participatory design workshops involving digital arts designers, we crafted a set of design principles for visual augmentations that are attuned to the requirements of multitaskers. Finally, a 3-day take-home field study further revealed new insights for everyday use, underscoring the potential of assisted reality (aR) to enhance heads-up listening and incidental learning experiences.
Felicia Fang-Yi Tan, Peisen Xu, Ashwin Ram 0002, Wei Zhen Suen, Shengdong Zhao 0001, Yun Huang 0003, Christophe Hurter
CHI5
2024 GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
abstract
Virtual assistants have the potential to play an important role in helping users achieves different tasks. However, these systems face challenges in their real-world usability, characterized by inefficiency and struggles in grasping user intentions. Leveraging recent advances in Large Language Models (LLMs), we introduce GptVoiceTasker, a virtual assistant poised to enhance user experiences and task efficiency on mobile devices. GptVoiceTasker excels at intelligently deciphering user commands and executing relevant device interactions to streamline task completion. For unprecedented tasks, GptVoiceTasker utilises the contextual information and on-screen content to continuously explore and execute the tasks. In addition, the system continually learns from historical user commands to automate subsequent task invocations, further enhancing execution efficiency. From our experiments, GptVoiceTasker achieved 84.5% accuracy in parsing human commands into executable actions and 85.7% accuracy in automating multi-step tasks. In our user study, GptVoiceTasker boosted task efficiency in real-world scenarios by 34.85%, accompanied by positive participant feedback. We made GptVoiceTasker open-source, inviting further research into LLMs utilization for diverse tasks through prompt engineering and leveraging user usage data to improve efficiency.
Minh Duc Vu, Han Wang 0023, Jieshan Chen, Zhuang Li 0001, Shengdong Zhao 0001, Zhenchang Xing, Chunyang Chen 0001
UIST5
2024 GestureGPT: Toward Zero-Shot Free-Form Hand Gesture Understanding with Large Language Model Agents
abstract
Existing gesture interfaces only work with a fixed set of gestures defined either by interface designers or by users themselves, which introduces learning or demonstration efforts that diminish their naturalness. Humans, on the other hand, understand free-form gestures by synthesizing the gesture, context, experience, and common sense. In this way, the user does not need to learn, demonstrate, or associate gestures. We introduce GestureGPT, a free-form hand gesture understanding framework that mimics human gesture understanding procedures to enable a natural free-form gestural interface. Our framework leverages multiple Large Language Model agents to manage and synthesize gesture and context information, then infers the interaction intent by associating the gesture with an interface function. More specifically, our triple-agent framework includes a Gesture Description Agent that automatically segments and formulates natural language descriptions of hand poses and movements based on hand landmark coordinates. The description is deciphered by a Gesture Inference Agent through self-reasoning and querying about the interaction context (e.g., interaction history, gaze data), which is managed by a Context Management Agent. Following iterative exchanges, the Gesture Inference Agent discerns the user’s intent by grounding it to an interactive function. We validated our framework offline under two real-world scenarios: smart home control and online video streaming. The average zero-shot Top-1/Top-5 grounding accuracies are 44.79%/83.59% for smart home tasks and 37.50%/73.44% for video streaming tasks. We also provide an extensive discussion that includes rationale for model selection, generalizability, and future research directions for a practical system etc.
Tengxiang Zhang, Chun Yu, Shengdong Zhao 0001, Yiqiang Chen 0001
Proc. ACM Hum. Comput. Interact.5
2024 Kine-Appendage: Enhancing Freehand VR Interaction Through Transformations of Virtual Appendages
abstract
Kinesthetic feedback, the feeling of restriction or resistance when hands contact objects, is essential for natural freehand interaction in VR. However, inducing kinesthetic feedback using mechanical hardware can be cumbersome and hard to control in commodity VR systems. We propose the kine-appendage concept to compensate for the loss of kinesthetic feedback in virtual environments, i.e., a virtual appendage is added to the user's avatar hand; when the appendage contacts a virtual object, it exhibits transformations (rotation and deformation); when it disengages from the contact, it recovers its original appearance. A proof-of-concept kine-appendage technique, BrittleStylus, was designed to enhance isomorphic typing. Our empirical evaluations demonstrated that (i) BrittleStylus significantly reduced the uncorrected error rate of naive isomorphic typing from 6.53% to 1.92% without compromising the typing speed; (ii) BrittleStylus could induce the sense of kinesthetic feedback, the degree of which was parity with that induced by pseudo-haptic (+ visual cue) methods; and (iii) participants preferred BrittleStylus over pseudo-haptic (+ visual cue) methods because of not only good performance but also fluent hand movements.
Yang Tian 0008, Hualong Bai, Shengdong Zhao 0001, Chi-Wing Fu, Chun Yu, Haozhao Qin, Qiong Wang 0001, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.3
2023 Mindful Moments: Exploring On-the-go Mindfulness Practice On Smart-glasses
abstract
Mindfulness technologies have gained research interest in recent years. We explore the use of smart-glasses (Optical Head Mounted Displays or OHMDs) for breath-based mindfulness practice as a well-being technology for everyday users. Since OHMDs do not occlude the wearer’s view, practitioners can access the digital environment while performing daily activities. Through our pilot series, we identified suitable visual and auditory attributes for OHMD mindfulness sessions in casual walking settings, and combined user-preferred features into our proposed Mindful Moments design. Results on physiological, sustained attention and self-reported mindfulness measures suggest that Mindful Moments facilitates higher state mindfulness than the Control. Its results proved comparable to the state-of-the-art Walking Meditation, while also being more accessible, convenient, and easy for novice practitioners to implement in everyday environments. We further evaluate Mindful Moments in a realistic setting, enhancing current understanding of mindfulness practice on OHMDs, thereby contributing a technique for improved health and well-being.
Felicia Fang-Yi Tan, Ashwin Ram 0002, Chloe Dolma Si Ying Haigh, Shengdong Zhao 0001
Conference on Designing Interactive Systems4
2023 ParaGlassMenu: Towards Social-Friendly Subtle Interactions in Conversations
abstract
Interactions with digital devices during social settings can reduce social engagement and interrupt conversations. To overcome these drawbacks, we designed ParaGlassMenu, a semi-transparent circular menu that can be displayed around a conversation partner’s face on Optical See-Through Head-Mounted Display (OHMD) and interacted subtly using a ring mouse. We evaluated ParaGlassMenu with several alternative approaches (Smartphone, Voice assistant, and Linear OHMD menus) by manipulating Internet-of-Things (IoT) devices in a simulated conversation setting with a digital partner. Results indicated that the ParaGlassMenu offered the best overall performance in balancing social engagement and digital interaction needs in conversations. To validate these findings, we conducted a second study in a realistic conversation scenario involving commodity IoT devices. Results confirmed the utility and social acceptance of the ParaGlassMenu. Based on the results, we discuss implications for designing attention-maintaining subtle interaction techniques on OHMDs.
Runze Cai, Nuwan Janaka, Shengdong Zhao 0001, Minghui Sun 0001
CHI3
2023 Can Icons Outperform Text? Understanding the Role of Pictograms in OHMD Notifications
abstract
Optical see-through head-mounted displays (OHMDs) can provide just-in-time digital assistance to users while they are engaged in ongoing tasks. However, given users’ limited attentional resources when multitasking, there is a need to concisely and accurately present information in OHMDs. Existing approaches for digital information presentation involve using either text or pictograms. While pictograms have enabled rapid recognition and easier use in warning messages and traffic signs, most studies using pictograms for digital notifications have exhibited unfavorable results. We thus conducted a series of four iterative studies to understand how we can support effective notification presentation on OHMDs during multitasking scenarios. We find that while icon-augmented notifications can outperform text-only notifications, their effectiveness depends on icon familiarity, encoding density, and environmental brightness. We reveal design implications when using icon-augmented notifications in OHMDs and present plausible reasons for the observed disparity in literature.
Nuwan Janaka, Shengdong Zhao 0001, Shardul Sapkota
CHI2
2023 Not All Spacings are Created Equal: The Effect of Text Spacings in On-the-go Reading Using Optical See-Through Head-Mounted Displays
abstract
The emergent Optical Head-Mounted Display (OHMD) platform has made mobile reading possible by superimposing digital text onto users’ view of the environment. However, mobile reading through OHMD needs to be effectively balanced with the user’s environmental awareness. Hence, a series of studies were conducted to explore how text spacing strategies facilitate such balance. Through these studies, it was found that increasing spacing within the text can significantly enhance mobile reading on OHMDs in both simple and complex navigation scenarios and that such benefits mainly come from increasing the inter-line spacing, but not inter-word spacing. Compared with existing positioning strategies, increasing inter-line spacing improves mobile OHMD information reading in terms of reading speed (11.9% faster), walking speed (3.7% faster), and switching between reading and navigation (106.8% more accurate and 33% faster).
Katherine Fennedy, Felicia Fang-Yi Tan, Shengdong Zhao 0001, Yurui Shao
CHI4
2023 AdaptReview: Towards Effective Video Review Using Text Summaries and Concept Maps
Shan Zhang 0006, Yang Chen 0054, Nuwan Janaka, Chloe Dolma Si Ying Haigh, Shengdong Zhao 0001, Wei Tsang Ooi
INTERACT (2)5
2023 BackTracer: Improving ray-casting 3D target acquisition by backtracking the interaction history
Huagen Wan, Shengdong Zhao 0001, Xutao Liu
Int. J. Hum. Comput. Stud.3
2022 Paracentral and near-peripheral visualizations: Towards attention-maintaining secondary information presentation on OHMDs during in-person social interactions
abstract
Optical see-through Head-Mounted Displays (OST HMDs, OHMDs) are known to facilitate situational awareness while accessing secondary information. However, information displayed on OHMDs can cause attention shifts, which distract users from natural social interactions. We hypothesize that information displayed in paracentral and near-peripheral vision can be better perceived while the user is maintaining eye contact during face-to-face conversations. Leveraging this idea, we designed a circular progress bar to provide progress updates in paracentral and near-peripheral vision. We compared it with textual and linear progress bars under two conversation settings: a simulated one with a digital conversation partner and a realistic one with a real partner. Results show that a circular progress bar can effectively reduce notification distractions without losing eye contact and is more preferred by users. Our findings highlight the potential of utilizing the paracentral and near-peripheral vision for secondary information presentation on OHMDs.
Nuwan Janaka, Chloe Dolma Si Ying Haigh, Hyeongcheol Kim 0001, Shan Zhang 0006, Shengdong Zhao 0001
CHI5
2022 Does Dynamically Drawn Text Improve Learning? Investigating the Effect of Text Presentation Styles in Video Learning
abstract
Dynamically drawn content (e.g., handwritten text) in learning videos is believed to improve users’ engagement and learning over static powerpoint-based ones. However, evidence from existing literature is inconclusive. With the emergence of Optical Head-Mounted Displays (OHMDs), recent work has shown that video learning can be adapted for on-the-go scenarios. To better understand the role of dynamic drawing, we decoupled dynamically drawn text into two factors (font style and motion of appearance) and studied their impact on learning performance under two usage scenarios (while seated with desktop and walking with OHMD). We found that although letter-traced text was more engaging for some users, most preferred learning with typeface text that displayed the entire word at once and achieved better recall (46.7% higher), regardless of the usage scenarios. Insights learned from the studies can better inform designers on how to present text in videos for ubiquitous access.
Ashwin Ram 0002, Shengdong Zhao 0001
CHI2
2021 Exploring Head-based Mode-Switching in Virtual Reality
abstract
Mode-switching supports multilevel operations using a limited number of input methods. In Virtual Reality (VR) head-mounted displays (HMD), common approaches for mode-switching use buttons, controllers, and users’ hands. However, they are inefficient and challenging to do with tasks that require both hands (e.g., when users need to use two hands during drawing operations). Using head gestures for mode-switching can be an efficient and cost-effective way, allowing for a more continuous and smooth transition between modes. In this paper, we explore the use of head gestures for mode-switching especially in scenarios when both users’ hands are performing tasks. We present a first user study that evaluated eight head gestures that could be suitable for VR HMD with a dual-hand line-drawing task. Results show that move forward, move backward, roll left, and roll right led to better performance and are preferred by participants. A second study integrating these four gestures in Tilt Brush, an open-source painting VR application, is conducted to further explore the applicability of these gestures and derive insights. Results show that Tilt Brush with head gestures allowed users to change modes with ease and led to improved interaction and user experience. The paper ends with a discussion on some design recommendations for using head-based mode-switching in VR HMD.
Rongkai Shi, Hai-Ning Liang, Shengdong Zhao 0001
ISMAR4
2021 Ubiquitous Interactions for Heads-Up Computing: Understanding Users' Preferences for Subtle Interaction Techniques in Everyday Settings
abstract
In order to satisfy users’ information needs while incurring minimum interference to their ongoing activities, previous studies have proposed using Optical Head-mounted Displays (OHMDs) with different input techniques. However, it is unclear how these techniques compare against one another in terms of being comfortable and non-intrusive to a user’s everyday tasks. Through a wizard-of-oz study, we thus compared four subtle interaction techniques (feet, arms, thumb-index-fingers, and teeth) in three daily hands-busy tasks under different settings (giving a presentation–sitting, carrying bags–walking, and folding clothes–standing). We found that while each interaction technique has its niche, thumb-index-finger interaction has the best overall balance and is most preferred as a cross-scenario subtle interaction technique for smart glasses. We provide further evaluation of thumb-index-finger interaction with an in-the-wild study with 8 users. Our results contribute to an enhanced understanding of user preferences for subtle interaction techniques with smart glasses for everyday use.
Shardul Sapkota, Ashwin Ram 0002, Shengdong Zhao 0001
MobileHCI3
2020 EYEditor: Towards On-the-Go Heads-Up Text Editing Using Voice and Manual Input
abstract
On-the-go text-editing is difficult, yet frequently done in everyday lives. Using smartphones for editing text forces users into a heads-down posture which can be undesirable and unsafe. We present EYEditor, a heads-up smartglass-based solution that displays the text on a see-through peripheral display and allows text-editing with voice and manual input. The choices of output modality (visual and/or audio) and content presentation were made after a controlled experiment, which showed that sentence-by-sentence visual-only presentation is best for optimizing users' editing and path-navigation capabilities. A second experiment formally evaluated EYEditor against the standard smartphone-based solution for tasks with varied editing complexities and navigation difficulties. The results showed that EYEditor outperformed smartphones as either the path OR the task became more difficult. Yet, the advantage of EYEditor became less salient when both the editing and navigation was difficult. We discuss trade-offs and insights gained for future heads-up text-editing solutions.
Debjyoti Ghosh, Pin Sym Foong, Shengdong Zhao 0001, Can Liu 0003, Nuwan Janaka, Vinitha Erusu
CHI3
2020 Learn with Haptics: Improving Vocabulary Recall with Free-form Digital Annotation on Touchscreen Mobiles
abstract
Mobile vocabulary learning interfaces typically present material only in auditory and visual channels, underutilizing the haptic modality. We explored haptic-integrated learning by adding free-form digital annotation to mobile vocabulary learning interfaces. Through a series of pilot studies, we identified three design factors: annotation mode, presentation sequence, and vibrotactile feedback, that influence recall in haptic-integrated vocabulary interfaces. These factors were then evaluated in a within-subject comparative study using a digital flashcard interface as baseline. Results using a 84-item vocabulary showed that the 'whole word' annotation mode is highly effective, yielding a 24.21% increase in immediate recall scores and a 30.36% increase in the 7-day delayed scores. Effects of presentation sequence and vibrotactile feedback were more transient; they affected the results of immediate tests, but not the delayed tests. We discuss the implications of these factors for designing future mobile learning applications.
Smitha Sheshadri, Shengdong Zhao 0001, Yang Chen 0054, Morten Fjeld
CHI2
2020 Virtually-Extended Proprioception: Providing Spatial Reference in VR through an Appended Virtual Limb
abstract
Selecting targets directly in the virtual world is difficult due to the lack of haptic feedback and inaccurate estimation of egocentric distances. Proprioception, the sense of self-movement and body position, can be utilized to improve virtual target selection by placing targets on or around one's body. However, its effective scope is limited closely around one's body. We explore the concept of virtually-extended proprioception by appending virtual body parts mimicking real body parts to users' avatars, to provide spatial reference to virtual targets. Our studies suggest that our approach facilitates more efficient target selection in VR as compared to no reference or using an everyday object as reference. Besides, by cultivating users' sense of ownership on the appended virtual body part, we can further enhance target selection performance. The effects of transparency and granularity of the virtual body part on target selection performance are also discussed.
Yang Tian 0008, Yuming Bai, Shengdong Zhao 0001, Chi-Wing Fu, Tianpei Yang, Pheng-Ann Heng
CHI3
2020 LiveSnippets: Voice-based Live Authoring of Multimedia Articles about Experiences
abstract
We transform traditional experience writing into in-situ voice-based multimedia authoring. Documenting experiences digitally in blogs and journals is a common activity that allows people to socially connect with others by sharing their experiences (e.g. travelogue). However, documenting such experiences can be time-consuming and cognitively demanding as it is typically done OUT-OF-CONTEXT (after the actual experience). We propose in-situ voice-based multimedia authoring (IVA), an alternative workflow to allow IN-CONTEXT experience documentation. Unlike the traditional approach, IVA encourages in-context content creations using voice-based multimedia input and stores them in multi-modal “snippets”. The snippets can be rearranged to form multimedia articles and can be published with light copy-editing. To improve the output quality from impromptu speech, Q&A scaffolding was introduced to guide the content creation. We implement the IVA workflow in an android application, LiveSnippets - and qualitatively evaluate it under three scenarios (travel writing, recipe creation, product review). Results demonstrated that IVA can effectively lower the barrier of writing with acceptable trade-offs in multitasking.
Hyeongcheol Kim 0001, Shengdong Zhao 0001, Can Liu 0003, Kotaro Hara
MobileHCI2
2020 From Lost to Found: Discover Missing UI Design Semantics through Recovering Missing Tags
abstract
Design sharing sites provide UI designers with a platform to share their works and also an opportunity to get inspiration from others' designs. To facilitate management and search of millions of UI design images, many design sharing sites adopt collaborative tagging systems by distributing the work of categorization to the community. However, designers often do not know how to properly tag one design image with compact textual description, resulting in unclear, incomplete, and inconsistent tags for uploaded examples which impede retrieval, according to our empirical study and interview with four professional designers. Based on a deep neural network, we introduce a novel approach for encoding both the visual and textual information to recover the missing tags for existing UI examples so that they can be more easily found by text queries. We achieve 82.72% accuracy in the tag prediction. Through a simulation test of 5 queries, our system on average returns hundreds more results than the default Dribbble search, leading to better relatedness, diversity and satisfaction.
Chunyang Chen 0001, Sidong Feng, Zhengyang Liu 0004, Zhenchang Xing, Shengdong Zhao 0001
Proc. ACM Hum. Comput. Interact.5
2020 Commanding and Re-Dictation: Developing Eyes-Free Voice-Based Interaction for Editing Dictated Text
abstract
Existing voice-based interfaces have limited support for text editing, especially when seeing the text is difficult, e.g., while walking or cooking. This research develops voice interaction techniques for eyes-free text editing. First, with a Wizard-of-Oz study, we identified two primary user strategies: using commands, e.g., “ replace go with goes ” and re-dictating over an erroneous portion, e.g., correcting “he go there” by saying “he goes there.” To support these user strategies with an actual system implementation, we developed two eyes-free voice interaction techniques, Commanding and Re-dictation , and evaluated them with a controlled experiment. Results showed that while Re-dictation performs significantly better for more semantically complex edits, Commanding is more suitable for making one-word edits, especially deletions. We developed VoiceRev to combine both the techniques in the same interface and evaluated it with realistic tasks. Results showed improved usability of the combined techniques over either of the two techniques used individually.
Debjyoti Ghosh, Can Liu 0003, Shengdong Zhao 0001, Kotaro Hara
ACM Trans. Comput. Hum. Interact.3
2019 Investigating Feedback for Two-Handed Exploration of Digital Maps Without Vision
Sandra Bardot, Marcos Serrano, Simon T. Perrault, Shengdong Zhao 0001, Christophe Jouffrais
INTERACT (1)4
2019 ScaffoMapping: Assisting Concept Mapping for Video Learners
Shan Zhang 0006, Xiaojun Meng, Can Liu 0003, Shengdong Zhao 0001, Vibhor Sehgal, Morten Fjeld
INTERACT (2)4
2019 Gallery D.C.: Design Search and Knowledge Discovery through Auto-created GUI Component Gallery
abstract
Online communities like Dribbble and GraphicBurger allow GUI designers to share their design artwork and learn from each other. These design sharing platforms are important sources for design inspiration, but our survey with GUI designers suggests additional information needs unmet by existing design sharing platforms. First, designers need to see the practical use of certain GUI designs in real applications, rather than just artworks. Second, designers want to see not only the overall designs but also the detailed design of the GUI components. Third, designers need advanced GUI design search abilities (e.g., multi-facets search) and knowledge discovery support (e.g., demographic investigation, cross-company design comparison). This paper presents Gallery D.C. http://mui-collection.herokuapp.com/, a gallery of GUI design components that harness GUI designs crawled from millions of real-world applications using reverse-engineering and computer vision techniques. Through a process of invisible crowdsourcing, Gallery D.C. supports novel ways for designers to collect, analyze, search, summarize and compare GUI designs on a massive scale. We quantitatively evaluate the quality of Gallery D.C. and demonstrate that Gallery D.C. offers additional support for design sharing and knowledge discovery beyond existing platforms.
Chunyang Chen 0001, Sidong Feng, Zhenchang Xing, Linda Liu, Shengdong Zhao 0001, Jinshui Wang
Proc. ACM Hum. Comput. Interact.5
2018 Movespace: on-body athletic interaction for running and cycling
abstract
Wearables are increasingly used during training to quantify performance and provide valuable real-time information. However, interacting with these devices in motion may disrupt the movements of the activity. We propose a method of interaction involving tapping specific locations on the body, identify candidate locations for running and cycling, and compare them in a series of controlled experiments with athletes. A purpose-built prototype measures speed of interaction and gives feedback cues for athletes to report the physical effects on the activity itself. Our results suggest that specific locations are faster and have minimal disruption to movement, even under induced fatigue conditions. The overall method is fast - 1.31s for running and 1.65s for cycling. Preferred locations differ significantly across sports, with stable body parts ranking higher. We effectively demonstrated the use of a single hand for interaction during running with two distinct tap gestures. A set of guidelines inform the design of new sports technologies.
Velko Vechev, Alexandru Dancu, Simon T. Perrault, Quentin Roy, Morten Fjeld, Shengdong Zhao 0001
AVI6
2018 Harvesting Caregiving Knowledge: Design Considerations for Integrating Volunteer Input in Dementia Care
abstract
Improving volunteer performance leads to better caregiving in dementia care settings. However, caregiving knowledge systems have been focused on eliciting and sharing expert, primary caregiver knowledge, rather than volunteer-provided knowledge. Through the use of an experience prototype, we explored the content of volunteer caregiver knowledge and identified ways in which such non-expert knowledge can be useful to dementia care. By using lay language, sharing information specific to the client and collaboratively finding strategies for interaction, volunteers were able to boost the effectiveness of future volunteers. Therapists who reviewed the content affirmed the reliability of volunteer caregiver knowledge and placed value on its recency, variety and its ability to help bridge language and professional barriers. We discuss how future systems designed for eliciting and sharing volunteer caregiver knowledge can be used to promote better dementia care.
Pin Sym Foong, Shengdong Zhao 0001, Felicia Fang-Yi Tan, Joseph Jay Williams
CHI2
2018 EDITalk: Towards Designing Eyes-free Interactions for Mobile Word Processing
abstract
We present EDITalk, a novel voice-based, eyes-free word processing interface. We used a Wizard-of-Oz elicitation study to investigate the viability of eyes-free word processing in the mobile context and to elicit user requirements for such scenarios. Results showed that meta-level operations like highlight and comment, and core operations like insert, delete and replace are desired by users. However, users were challenged by the lack of visual feedback and the cognitive load of remembering text while editing it. We then studied a commercial-grade dictation application and discovered serious limitations that preclude comfortable speak-to-edit interactions. We address these limitations through EDITalk's closed-loop interaction design, enabling eyes-free operation of both meta-level and core word processing operations in the mobile context. Finally, we discuss implications for the design of future mobile, voice-based, eyes-free word processing interface.
Debjyoti Ghosh, Pin Sym Foong, Shengdong Zhao 0001, Morten Fjeld
CHI3
2018 AVEID: Automatic Video System for Measuring Engagement In Dementia
abstract
Engagement in dementia is typically measured using behavior observational scales (BOS) that are tedious and involve intensive manual labor to annotate, and are therefore not easily scalable. We propose AVEID, a low cost and easy-to-use video-based engagement measurement tool to determine the engagement level of a person with dementia (PwD) during digital interaction. We show that the objective behavioral measures computed via AVEID correlate well with subjective expert impressions for the popular MPES and OME BOS, confirming its viability and effectiveness. Moreover, AVEID measures can be obtained for a variety of engagement designs, thereby facilitating large-scale studies with PwD populations.
Viral Parekh, Pin Sym Foong, Shengdong Zhao 0001, Subramanian Ramanathan
IUI3
2017 ReTool: Interactive Microtask and Workflow Design through Demonstration
abstract
In addition to simple form filling, there is an increasing need for crowdsourcing workers to perform freeform interactions directly on content in microtask crowdsourcing (e.g. proofreading articles or specifying object boundary in an image). Such microtasks are often organized within well-designed workflows to optimize task quality and workload distribution. However, designing and implementing the interface and workflow for such microtasks is challenging because it typically requires programming knowledge and tedious manual effort. We present ReTool, a web-based tool for requesters to design and publish interactive microtasks and workflows by demonstrating the microtasks for text and image content. We evaluated ReTool against a task-design tool from a popular crowdsourcing platform and showed the advantages of ReTool over the existing approach.
Xiaojun Meng, Shengdong Zhao 0001, Morten Fjeld
CHI3
2017 VITA: Towards Supporting Volunteer Interactions with Long-Term Care Residents with Dementia
abstract
Volunteers are an important resource at long-term care homes because they can supply services, such as engagement activities, that over-burdened care staff struggle to provide. However, volunteers without sufficient training are often challenged in responding to dementia-linked behaviors, which can lead to frustrating difficulties during interaction. Additionally, short-staffed care homes have difficulties in training and maintaining volunteers. To better support volunteers in providing engagement activities for people with dementia without a high training burden, we created VITA, a tablet-based system that supplies carefully designed profiling and guidance using our dementia-appropriate engagement activity kit. Our evaluation indicated that the instructional guide supplied by VITA significantly improves volunteers' ability to facilitate engagement activities with people with dementia, approaching the level of engagement achievable by professional therapists.
Pin Sym Foong, Shengdong Zhao 0001, Kelsey Carlson, Zhe Liu 0043
CHI2
2017 Follow-My-Lead: Intuitive Indoor Path Creation and Navigation Using Interactive Videos
abstract
We present Follow-My-Lead, an alternative indoor navigation technique that uses visual information recorded on an actual navigation path as a navigational guide. Its design revealed a trade-off between the fidelity of information provided to users and their effort to acquire it. Our first experiment revealed that scrolling through a continuous image stream of the navigation path is highly informative, but it becomes tedious with constant use. Discrete image checkpoints require less effort, but can be confusing. A balance may be struck by adding fast video transitions between image checkpoints, but precise control is required to handle difficult situations. Authoring still image checkpoints is also difficult, and this inspired us to invent a new technique using video checkpoints. We conducted a second experiment on authoring and navigation performance and found video checkpoints plus fast video transitions to be better than both image checkpoints plus fast video transitions and traditional written instructions.
Quentin Roy, Simon T. Perrault, Shengdong Zhao 0001, Richard C. Davis, Anuroop Pattena Vaniyar, Velko Vechev, Youngki Lee 0001, Archan Misra
CHI3
2017 NexP: A Beginner Friendly Toolkit for Designing and Conducting Controlled Experiments
Xiaojun Meng, Pin Sym Foong, Simon T. Perrault, Shengdong Zhao 0001
INTERACT (3)4
2017 Editorial of the Special Issue on Mobile Human-Computer Interaction
abstract
Mobile and wearable devices (e.g., smartphones, FitBit, and Apple Watch) are among the most transformative technologies that have created a rapid worldwide impact on almost every aspect of our soci...
Fiona Fui-Hoon Nah, Dongsong Zhang, John Krogstie, Shengdong Zhao 0001
Int. J. Hum. Comput. Interact.4
2017 Korero: Facilitating Complex Referencing of Visual Materials in Asynchronous Discussion Interface
abstract
In asynchronous online discussions, users actively reference visual materials (e.g., video, document) to provide supporting evidence and additional context. However, creating and comprehending complex references can be challenging, especially when there are multiple referents to refer, or when a referent is highly specific (e.g., specific sentences in a paper rather than the paper as a whole). To identify users' challenges in making references with multiple and specific referents while using existing discussion tools, we conducted an observational study and a preliminary interview. Based on the design lessons, we built Korero, a discussion interface that aims to facilitate complex referencing actions. For evaluation, we compared Korero against conventional interfaces in two user studies with referencing tasks of different referential difficulty. We found that Korero not only significantly reduces the time and effort in making references with multiple and specific referents, but also shows potential in increasing users' engagement with the discussion and referent materials.
Soon Hau Chua, Toni-Jan Keith Palma Monserrat, Dongwook Yoon, Juho Kim 0001, Shengdong Zhao 0001
Proc. ACM Hum. Comput. Interact.5
2016 FusePrint: A DIY 2.5D Printing Technique Embracing Everyday Artifacts
abstract
FusePrint is a Stereolithography-based 2.5D rapid prototyping technique that allows high-precision fabrication without high-end modeling tools, enabling the mixing of everyday physical artifacts and liquid conductive gels with photo-reactive resin during the printing process, facilitating the creation of 2.5D objects that perfectly fit the existing objects. Based on our polynomial model on 2.5D resin printing, we developed the design interface of FusePrint, which allows users to design the printed shapes using physical objects as references, generates projection patterns, and notifies users when to place the objects in the resin during the printing process. Our workshops suggested that FusePrint is easy to learn and use, provides a greater level of interactivity, and could be useful for a wide range of applications domains including: mechanical fabrication, wearable accessory, toys, interactive systems, etc.
Kening Zhu, Alexandru Dancu, Shengdong Zhao 0001
Conference on Designing Interactive Systems3
2016 5-Step Approach to Designing Controlled Experiments
abstract
Controlled experiment, an approach that has been adopted from research methods in psychology, is now widely used in HCI. Design an effective controlled experiment is not necessarily easy, even for experienced researchers. Existing controlled experiment designing tools focused more on the process of designing, but they are often not intuitive and guided enough for less experienced users to use. In this demo paper, we introduce NexP (Next Experiment Tool), a web-based open-source tool for designing controlled experiments. NexP introduces a 5-step approach to guide users through the experimental design process, and helped them to better understand the experimental design process, making it both a useful and educational tool.
Xiaojun Meng, Pin Sym Foong, Simon T. Perrault, Shengdong Zhao 0001
AVI4
2016 HyNote: Integrated Concept Mapping and Notetaking
abstract
Notes can be taken in a linear or nonlinear way. Previous work suggests that nonlinear note taking is advantageous in terms of sense-making and long-term recall. However, previous studies also reveal that the combination of divided attention and time pressure make realtime notetaking a challenge. In this paper, we propose a new hybrid workflow inheriting advantages from both linear and nonlinear notetaking approaches. Our resulting HyNote (Hybrid Notetaking) system uses statistical parsing of linear raw notes to facilitate concepts mapping, allowing users to smoothly switch between linear and nonlinear approaches with low effort and time costs. Results from our preliminary study of HyNote show that users can easily map concepts in realtime and achieve superior understanding of lecture contents in a video learning task compared with using the traditional linear Notepad application.
Xiaojun Meng, Shengdong Zhao 0001, Darren Edge
AVI2
2016 HaptiColor: Interpolating Color Information as Haptic Feedback to Assist the Colorblind
abstract
Most existing colorblind aids help their users to distinguish and recognize colors but not compare them. We present HaptiColor, an assistive wristband that encodes discrete color information into spatiotemporal vibrations to support colorblind users to recognize and compare colors. We ran three experiments: the first found the optimal number and placement of motors around the wrist-worn prototype, and the second tested the optimal way to represent discrete points between the vibration motors. Results suggested that using three vibration motors and pulses of varying duration to encode proximity information in spatiotemporal patterns is the optimal solution. Finally, we evaluated the HaptiColor prototype and encodings with six colorblind participants. Our results show that the participants were able to easily understand the encodings and perform color comparison tasks accurately (94.4% to 100%).
Marta Gonzalez Carcedo, Soon Hau Chua, Simon T. Perrault, Pawel W. Wozniak, Raj Joshi, Mohammad Obaid, Morten Fjeld, Shengdong Zhao 0001
CHI8
2016 Using Crowdsourcing for Scientific Analysis of Industrial Tomographic Images
abstract
In this article, we present a novel application domain for human computation, specifically for crowdsourcing, which can help in understanding particle-tracking problems. Through an interdisciplinary inquiry, we built a crowdsourcing system designed to detect tracer particles in industrial tomographic images, and applied it to the problem of bulk solid flow in silos. As images from silo-sensing systems cannot be adequately analyzed using the currently available computational methods, human intelligence is required. However, limited availability of experts, as well as their high cost, motivates employing additional nonexperts. We report on the results of a study that assesses the task completion time and accuracy of employing nonexpert workers to process large datasets of images in order to generate data for bulk flow research. We prove the feasibility of this approach by comparing results from a user study with data generated from a computational algorithm. The study shows that the crowd is more scalable and more economical than an automatic solution. The system can help analyze and understand the physics of flow phenomena to better inform the future design of silos, and is generalized enough to be applicable to other domains.
Pawel W. Wozniak, Andrzej Romanowski, Mohammad Obaid, Tomasz Jaworski, Jacek Kucharski, Krzysztof Grudzien, Shengdong Zhao 0001, Morten Fjeld
ACM Trans. Intell. Syst. Technol.8
2016 Investigating Expressive Tactile Interaction Design in Artistic Graphical Representations
Maryam Azh, Shengdong Zhao 0001, Sriram Subramanian
ACM Trans. Comput. Hum. Interact.2
2015 OmniVib: Towards Cross-body Spatiotemporal Vibrotactile Notifications for Mobile Phones
abstract
Previous works illustrate that one's palm can reliably recognize 10 or more spatiotemporal vibrotactile patterns. However, recognition of the same patterns on other body parts is unknown. In this paper, we investigate how users perceive spatiotemporal vibrotactile patterns on the arm, palm, thigh, and waist. Results of the first two experiments indicate that precise recognition of either position or orientation is difficult across multiple body parts. Nonetheless, users were able to distinguish whether two vibration pulses were from the same location when played in quick succession. Based on this finding, we designed eight spatiotemporal vibrotactile patterns and evaluated them in two additional experiments. The results demonstrate that these patterns can be reliably recognized (>80%) across the four tested body parts, both in the lab and in a more realistic context.
Jessalyn Alvina, Shengdong Zhao 0001, Simon T. Perrault, Maryam Azh, Thijs Roumen, Morten Fjeld
CHI2
2015 Physical Loci: Leveraging Spatial, Object and Semantic Memory for Command Selection
abstract
Physical Loci, a technique based on an ancient memory technique, allows users to quickly learn a large command set by leveraging spatial, object and verbal/semantic memory to create a cognitive link between individual commands and nearby physical objects in a room (called loci). We first report on an experiment that showed that for learning 25 items Physical Loci outperformed a mid-air Marking Menu baseline. A long-term retention experiment with 48 items then showed that recall was nearly perfect one week later and, surprisingly, independent of whether the command/locus mapping was one's own choice or somebody else's. A final study suggested that recall performance is robust to alterations of the learned mapping, whether systematic or random.
Simon T. Perrault, Eric Lecolinet, Yoann Pascal Bourse, Shengdong Zhao 0001, Yves Guiard
CHI4
2015 NotiRing: A Comparative Study of Notification Channels for Wearable Interactive Rings
abstract
We conducted an empirical investigation of wearable interactive rings on the noticeability of four instantaneous notification channels (light, vibration, sound, poke) and a channel with gradually increased temperature (thermal) during five levels of physical activity (laying down, sitting, standing, walking, and running). Results showed that vibration was the most reliable and fastest channel to convey notification, followed by poke and sound which shared similar noticeability. The noticeability of these three channels was not affected by the level of physical activity. The other two channels, light and thermal, were less noticeable and were affected by the level of physical activity. Our post-experimental survey indicates that while noticeability has a significant influence on user preference, each channel has its own unique advantages that make it suitable for different notification scenarios.
Thijs Roumen, Simon T. Perrault, Shengdong Zhao 0001
CHI3
2015 CoFaçade: A Customizable Assistive Approach for Elders and Their Helpers
abstract
We present CoFaçade, a novel approach to helping elders reach their goals with IT products by working collaboratively with helpers. In this approach, the elder uses an interface with a small number of triggers, where each trigger is a single button (or card) that can execute a procedure. The helper uses a customization interface to link triggers to procedures that accomplish frequently-recurring high-level goals with IT products. Customization can be done either locally or remotely. We conducted an experiment to compare the CoFaçade approach with a baseline approach where helpers taught elders to perform IT tasks. Our results showed that CoFaçade can reduce helpers' time and effort, reduce elders' frustration, and improve elders' success rate in completing IT tasks.
Jason Chen Zhao, Richard C. Davis, Pin Sym Foong, Shengdong Zhao 0001
CHI4
2015 To Risk or Not to Risk?: Improving Financial Risk Taking of Older Adults by Online Social Information
abstract
Increasing number of older adults manage their retirement savings online. A crucial element of better management is to take rational financial risk -- to strike a reasonable balance between expected gain and loss under uncertainty. With the emergence of Web 2.0 technologies, social trading networks can help individuals make better financial decisions by providing information about others' actions. It is, however, unclear whether these resources is beneficial to older adult's own financial decisions, especially because older adults are vulnerable to poor risk management. To address this question, we devise an experiment that improves upon an existing experimental economic task. We find that both peer information (detailed choices by a few individuals) and majority information (aggregated choices of the crowd) help older adults make more risk-neutral decisions. Furthermore, the combination of peer and majority information corrects more mistakes of more risk-averse older adults.
Jason Chen Zhao, Wai-Tat Fu, Hanzhe Zhang, Shengdong Zhao 0001, Henry Been-Lirn Duh
CSCW4
2015 Understanding Learners' General Perception Towards Learning with MOOC Classmates: An Exploratory Study
abstract
In this work-in-progress, we present our preliminary findings from an exploratory study on understanding learners' general behavior and perception towards learning with classmates in MOOCs. One-on-one semi-structured interview designed with grounded theory method was conducted with seven MOOC learners. Initial analysis of the interview data revealed several interesting insights on learners' behavior in working with other learners in MOOCs. We intend to expand the findings in future work to derive design implications for incorporating collaborative features into MOOCs.
Soon Hau Chua, Juho Kim 0001, Toni-Jan Keith Palma Monserrat, Shengdong Zhao 0001
L@S4
2015 Botential: Localizing On-Body Gestures by Measuring Electrical Signatures on the Human Skin
abstract
We present Botential, an on-body interaction method for a wearable input device that can identify the location of on-body tapping gestures, using the entire human body as an interactive surface to expand the usually limited interaction space in the context of mobility. When the sensor is being touched, Botential identifies a body part's unique electric signature, which depends on its physiological and anatomical compositions. This input method exhibits a number of advantages over previous approaches, which include: 1) utilizing the existing signal the human body already emits, to accomplish input with various body parts, 2) the ability to also sense soft and long touches, 3) an increased sensing range that covers the whole body, and 4) the ability to detect taps and hovering through clothes.
Denys J. C. Matthies, Simon T. Perrault, Bodo Urban, Shengdong Zhao 0001
MobileHCI4
2015 ColorBless: Augmenting Visual Information for Colorblind People with Binocular Luster Effect
abstract
Binocular disparity allows interesting visual effects visible only to people with stereoscopic 3D displays. Here, we studied and applied one such effect, binocular luster, to the application of digital colorblind aids with active shutter 3D. We developed two prototype techniques, ColorBless and PatternBless, to investigate the effectiveness of such aids and to explore the potential applications of a luster effect in stereoscopic 3D beyond highlighting. User studies and interviews revealed that luster-based aids were fast and required lower cognitive effort than existing aids and were preferred over other aids by the majority of colorblind participants. We infer design implications of a luster effect from the study and propose potential applications in augmented visualization.
Soon Hau Chua, Haimo Zhang, Muhammad Hammad 0001, Shengdong Zhao 0001, Sahil Goyal, Karan Singh 0004
ACM Trans. Comput. Hum. Interact.4
2014 BezelCopy: an efficient cross-application copy-paste technique for touchscreen smartphones
abstract
Copy-Paste (CP) operations on touchscreen smartphones are not as easy to perform as compared with similar operations on desktop computers. The smaller screen size and input area make both text selection and application switching more difficult to perform. To enable faster copy-paste on touchscreen smartphones, we introduce BezelCopy, a copy-paste technique that uses a bezel-swipe gesture to determine a rough area of interest in the document. Chosen text is magnified in a new panel to enable fast and precise selection. With the new panel, users can perform easy tap-and-drag gestures to select the exact content, and tap the application icon on the bottom of the panel to paste it to the target application. Users can further adjust the location of the pasted text in the target application using drag and drop. We conducted two experiments to compare the performance of BezelCopy with alternative approaches, and our results show that BezelCopy outperform existing copy-paste techniques for a number of commonly performed copy-paste tasks.
Simon T. Perrault, Shengdong Zhao 0001, Wei Tsang Ooi
AVI3
2014 Draco: bringing life to illustrations with kinetic textures
abstract
We present Draco, a sketch-based interface that allows artists and casual users alike to add a rich set of animation effects to their drawings, seemingly bringing illustrations to life. While previous systems have introduced sketch-based animations for individual objects, our contribution is a unified framework of motion controls that allows users to seamlessly add coordinated motions to object collections. We propose a framework built around kinetic textures, which provide continuous animation effects while preserving the unique timeless nature of still illustrations. This enables many dynamic effects difficult or not possible with previous sketch-based tools, such as a school of fish swimming, tree leaves blowing in the wind, or water rippling in a pond. We describe our implementation and illustrate the repertoire of animation effects it supports. A user study with professional animators and casual users demonstrates the variety of animations, applications and creative possibilities our tool provides.
Rubaiat Habib Kazi, Fanny Chevalier, Tovi Grossman, Shengdong Zhao 0001, George W. Fitzmaurice
CHI4
2014 WADE: simplified GUI add-on development for third-party software
abstract
We present the WADE Integrated Development Environment (IDE), which simplifies interface and functionality modification of existing third-party software without access to source code. WADE clones the Graphical User Interface (GUI) of a host program through dynamic-link library (DLL) injection, enabling modifications to (1) the GUI in a WYSIWYG fashion and (2) software functionality. We compare WADE with an alternative state-of-the-art runtime toolkit overloading approach in a user-study, whose results demonstrate that WADE significantly simplifies the task of GUI-based add-on development.
Xiaojun Meng, Shengdong Zhao 0001, James R. Eagan, Subramanian Ramanathan
CHI2
2014 L.IVE: an integrated interactive video-based learning environment
abstract
In this paper, we introduce L.IVE: an online interactive video-based learning environment with an alternative design and architecture that integrates three major interface components: video, comment threads, and assessments. This is in contrast with the approach of existing interfaces which visually separate these components. Our study, which compares L.IVE with existing popular video-based learning environments, suggests advantages in this integrated approach as compared to the separated approach in learning.
Toni-Jan Keith Palma Monserrat, Yawen Li 0004, Shengdong Zhao 0001
CHI3
2014 Food messaging: using edible medium for social messaging
abstract
Food is more than just a means of survival; it is also a form of communication. In this paper, we investigate the potential of food as a social message carrier (a.k.a., food messaging). To investigate how people accept, use, and perceive food messaging, we conducted exploratory interviews, a field study, and follow-up interviews over four weeks in a large information technology (IT) company. We collected 904 messages sent by 343 users. Our results suggest strong acceptance of food messaging as an alternative message channel. Further analysis implies that food messaging embodies characteristics of both text messaging and gifting. It is preferred in close relationships for its evocation of positive emotions. As the first field study on edible social messaging, our empirical findings provide valuable insights into the uniqueness of food as a message carrier and its capabilities to promote greater social bonding.
Xiaojuan Ma, Shengdong Zhao 0001
CHI3
2014 Using Social Media Platforms for Human-Robot Interaction in Domestic Environment
abstract
This article explores the application of existing social media platforms for human–robot interaction. With the increasing popularity of social media platforms that connect humans, we propose to portray domestic robots as buddies on the contact list of family members and present a robot management system that employs complementary social media platforms for humans to interact with the vacuuming robot Roomba and a surveillance robot developed on top of iRobot Create. The social media platforms adopted include short message services (SMS), instant messenger (MSN), an online shared calendar (Google Calendar), and a social networking site (Facebook). Hence, we can provide a rich set of user-familiar, intuitive, and highly accessible interfaces, allowing users to flexibly choose their preferred tools in different situations. An in-lab experiment and a multiday field study are conducted to study the characteristics and strengths of each interface and to investigate users’ perception to the robots and behaviors in choosing the interfaces.
Xiaoning Ma, Xin Yang 0009, Shengdong Zhao 0001, Chi-Wing Fu, Ziquan Lan, Yiming Pu
Int. J. Hum. Comput. Interact.3
2013 NoteVideo: facilitating navigation of blackboard-style lecture videos
abstract
Khan Academy's pre-recorded blackboard-style lecture videos attract millions of online users every month. However, current video navigation tools do not adequately support the kinds of goals that students typically have, like quickly finding a particular concept in a blackboard-style lecture video. This paper reports on the development and evaluation of the new NoteVideo and its improved version, NoteVideo+, systems for identifying the conceptual 'objects' of a blackboard-based video - and then creating a summarized image of the video and using it as an in-scene navigation interface that allows users to directly jump to the video frame where that object first appeared instead of navigating it linearly through time. The research consisted of iteratively implementing the system and then having users perform four different navigation tasks using three different interfaces: Scrubbing, Transcript, and NoteVideo. Results of the study show that participants perform significantly better on all four tasks while using the NoteVideo and its improved version - NoteVideo+ - as compared to others.
Toni-Jan Keith Palma Monserrat, Shengdong Zhao 0001, Kevin McGee, Anshul Vikram Pandey
CHI2
2013 AutoGami: a low-cost rapid prototyping toolkit for automated movable paper craft
abstract
AutoGami is a toolkit for designing automated movable paper craft using the technology of selective inductive power transmission. AutoGami has hardware and software components that allow users to design and implement automated movable paper craft without any prerequisite knowledge of electronics; it also supports rapid prototyping. Apart from developing the toolkit, we have analyzed the design space of movable paper craft and developed a taxonomy to facilitate the design of automated paper craft. AutoGami made consistently strong showings in design workshops, confirming its viability in supporting engagement and creativity as well as its usability in storytelling through paper craft. Additional highlights include rapid prototyping of product design as well as interaction design such as human-robot interactions.
Kening Zhu, Shengdong Zhao 0001
CHI2
2013 Designing an effective vibration-based notification interface for mobile phones
abstract
We conducted an experiment to understand how mobile phone users perceive the urgency of ten simple vibration alerts that were created from four basic signals: short on, short off, long on, and long off. The short and long signals correspond to 200 ms and 600 ms, respectively. To convey the level of urgency of notifications and help users prioritize them, the design of mobile phone vibration alerts should consider that the gap length preceding or succeeding a signal, the number of gaps in the vibration pattern, and the vibration's duration affect an alert's perceived level of urgency. Our study specifically shows that shorter gap lengths between vibrations (200 ms vs. 600 ms), a vibration pattern with one gap instead of two, and shorter vibration all contribute to making the user perceive the alert as more urgent.
Bahador Saket, Chrisnawan Prasojo, Shengdong Zhao 0001
CSCW4
2013 Shared Input Multimodal Mobile Interfaces: Interaction Modality Effects on Menu Selection in Single-Task and Dual-Task Environments
abstract
Journal Article Shared Input Multimodal Mobile Interfaces: Interaction Modality Effects on Menu Selection in Single-Task and Dual-Task Environments Get access Shengdong Zhao, Shengdong Zhao * 1Department of Computer Science, National University of Singapore, 13 Computing Drive, Computing 2, #01-04, Singapore 117417 *Corresponding author: [email protected] Search for other works by this author on: Oxford Academic Google Scholar Duncan P. Brumby, Duncan P. Brumby 2UCL Interaction Centre, University College London, Gower Street, London WC1E 6BT, UK Search for other works by this author on: Oxford Academic Google Scholar Mark Chignell, Mark Chignell 3Knowledge Media Design Institute (KMDI), University of Toronto, 27 King's College Circle, Toronto, Ont., Canada M5S 1A1 Search for other works by this author on: Oxford Academic Google Scholar Dario Salvucci, Dario Salvucci 4Drexel University, 3141 Chestnut Street, Philadelphia, PA 19104, USA Search for other works by this author on: Oxford Academic Google Scholar Sahil Goyal Sahil Goyal 5National University of Singapore, 13 Computing Drive, Computing 2, #01-04, Singapore 117417 Search for other works by this author on: Oxford Academic Google Scholar Interacting with Computers, Volume 25, Issue 5, September 2013, Pages 386–403, https://doi.org/10.1093/iwc/iws021 Published: 06 February 2013 Article history Received: 29 December 2011 Revision received: 11 October 2012 Accepted: 13 November 2012 Published: 06 February 2013
Shengdong Zhao 0001, Duncan P. Brumby, Mark Chignell, Dario D. Salvucci, Sahil Goyal
Interact. Comput.1
2012 AutoComPaste: auto-completing text as an alternative to copy-paste
abstract
The copy-paste command is a fundamental and widely used operation in daily computing. It is generally regarded as a simple task but the process can become tedious when frequent window switching is required to copy-paste across different documents. Auto-completion is another popular operation aimed at reducing users' typing effort. It contrasts to copy-paste by allowing for text completion without switching windows. However, the available content for completion is predefined. We introduce AutoComPaste, an enhanced autocompletion technique for cross-document copy-paste. AutoComPaste allows users to copy-paste different granularity of text from all opened documents without window switching. Our theoretical analysis and empirical study show that AutoComPaste nicely complements traditional copy-paste techniques and outperforms the traditional copy-paste techniques when users have knowledge of the content to be copied.
Shengdong Zhao 0001, Fanny Chevalier, Wei Tsang Ooi, Chee Yuan Lee
AVI1
2012 Vignette: interactive texture design and manipulation with freeform gestures for pen-and-ink illustration
abstract
Vignette is an interactive system that facilitates texture creation in pen-and-ink illustrations. Unlike existing systems, Vignette preserves illustrators' workflow and style: users draw a fraction of a texture and use gestures to automatically fill regions with the texture. We currently support both 1D and 2D synthesis with stitching. Our system also has interactive refinement and editing capabilities to provide a higher level texture control, which helps artists achieve their desired vision. A user study with professional artists shows that Vignette makes the process of illustration more enjoyable and that first time users can create rich textures from scratch within minutes.
Rubaiat Habib Kazi, Takeo Igarashi, Shengdong Zhao 0001, Richard C. Davis
CHI3
2012 Exploring user motivations for eyes-free interaction on mobile devices
abstract
While there is increasing interest in creating eyes-free interaction technologies, a solid analysis of why users need or desire eyes-free interaction has yet to be presented. To gain a better understanding of such user motivations, we conducted an exploratory study with four focus groups, and suggest a classification of motivations for eyes-free interaction under four categories (environmental, social, device features, and personal). Exploring and analyzing these categories, we present early insights pointing to design implications for future eyes-free interactions.
Morten Fjeld, Shengdong Zhao 0001
CHI4
2012 Beyond stereo: an exploration of unconventional binocular presentation for novel visual experience
abstract
Human stereo vision processes the two different images seen by the two eyes to generate depth sensation. While current stereoscopic display technologies look at how to faithfully simulate the stereo viewing experience, we took a look out of this scope, to explore how we may present binocular image pairs that differ in other ways to create novel visual experience. This paper presents several interesting techniques we explored, and discusses their potential applications according to an informal user study.
Haimo Zhang, Shengdong Zhao 0001
CHI3
2012 LUI: lip in multimodal mobile GUI interaction
abstract
Gesture based interactions are commonly used in mobile and ubiquitous environments. Multimodal interaction techniques use lip gestures to enhance speech recognition or control mouse movement on the screen. In this paper we extend the previous work to explore LUI: lip gestures as an alternative input technique for controlling the user interface elements in a ubiquitous environment. In addition to use lips to control cursor movement, we use lip gestures to control music players and activate menus. A LUI Motion-Action library is also provided to guide future interaction design using lip gestures.
Maryam Azh, Shengdong Zhao 0001
ICMI2
2012 ICMI'12 grand challenge: haptic voice recognition
abstract
This paper describes the Haptic Voice Recognition (HVR) Grand Challenge 2012 and its datasets. The HVR Grand Challenge 2012 is a research oriented competition designed to bring together researchers across multiple disciplines to work on novel multimodal text entry methods involving speech and touch inputs. Annotated datasets were collected and released for this grand challenge as well as future research purposes. A simple recipe for building an HVR system using the Hidden Markov Model Toolkit (HTK) was also provided. In this paper, detailed analyses of the datasets will be given. Experimental results obtained using these data will also be presented.
Khe Chai Sim, Shengdong Zhao 0001, Kai Yu 0004, Hank Liao
ICMI2
2012 Tracing Tuples Across Dimensions: A Comparison of Scatterplots and Parallel Coordinate Plots
abstract
Abstract One of the fundamental tasks for analytic activity is retrieving (i.e., reading) the value of a particular quantity in an information visualization. However, few previous studies have compared user performance in such value retrieval tasks for different visualizations. We present an experimental comparison of user performance (time and error distance) across four multivariate data visualizations. Three variants of scatterplot (SCP) visualizations, namely SCPs with common vertical axes (SCP‐common), SCPs with a staircase layout (SCP‐staircase), and SCPs with rotated axes between neighboring cells (SCP‐rotated), and a baseline parallel coordinate plots (PCP) were compared. Results show that the baseline PCP is better than SCP‐rotated and SCP‐staircase under all conditions, while the difference between SCP‐common and PCP depends on the dimensionality and density of the dataset. PCP shows advantages over SCP‐common when the dimensionality and density of the dataset are low, but SCP‐common eventually outperforms PCP as data dimensionality and density increase. The results suggest guidelines for the use of SCPs and PCPs that can benefit future researchers and practitioners.
Xiaole Kuang, Haimo Zhang, Shengdong Zhao 0001, Michael J. McGuffin
Comput. Graph. Forum3
2011 Farmer's tale: a facebook game to promote volunteerism
abstract
Volunteering is an important activity that brings great benefits to societies. However, encouraging volunteerism is difficult due to the altruistic nature of volunteer activities and the high resource demand in carrying them out. We have created a Facebook game called "Farmer's Tale" to attract and make it easier for people to volunteer. We evaluated people's acceptance to this novel idea and the results revealed great potential in such type of games.
Don Sim Jianqiang, Xiaojuan Ma, Shengdong Zhao 0001, Jing Ting Khoo, Swee Ling Bay, Zhenhui Jiang
CHI3
2011 SandCanvas: a multi-touch art medium inspired by sand animation
abstract
Sand animation is a performance art technique in which an artist tells stories by creating animated images with sand. Inspired by this medium, we have developed a new multi-touch digital artistic medium named SandCanvas that simplifies the creation of sand animations. SandCanvas also goes beyond traditional sand animation with tools for mixing sand animation with video and replicating recorded free-form hand gestures. In this paper, we analyze common sand animation hand gestures, present SandCanvas's intuitive UI, and describe implementation challenges we encountered. We also present an evaluation with professional and novice artists that shows the importance and unique affordances of this new medium.
Rubaiat Habib Kazi, Kien Chuan Chua, Shengdong Zhao 0001, Richard C. Davis, Kok-Lim Low
CHI3
2011 Measuring web page revisitation in tabbed browsing
abstract
Browsing the web has been shown to be a highly recurrent activity. Aimed to optimize the browsing experience, extensive previous research has been carried out on users' revisitation behavior. However, the conventional definition for revisitation, which only considers page loading activities by monitoring http requests initiated by the browser, largely underestimates users' intended revisitation activities with tabbed browsers. Thus, we introduce a goal-oriented definition and a refined revisitation measurement based on page viewings in tabbed browsers. An empirical analysis of statistics taken from a client-side log study showed that although the overall revisitation rate remained relatively constant, tabbed browsing has introduced new behaviors warrant future investigations.
Haimo Zhang, Shengdong Zhao 0001
CHI2
2011 MOGCLASS: evaluation of a collaborative system of mobile devices for classroom music education of young children
abstract
Composition, listening, and performance are essential activities in classroom music education, yet conventional music classes impose unnecessary limitations on students' ability to develop these skills. Based on in-depth fieldwork and a user-centered design approach, we created MOGCLASS, a multimodal collaborative music environment that enhances students' musical experience and improves teachers' management of the classroom.
Yinsheng Zhou, Graham Percival, Xinxi Wang, Ye Wang 0007, Shengdong Zhao 0001
CHI5
2010 MOGCLASS: a collaborative system of mobile devices forclassroom music education
abstract
We introduce MOGCLASS: a system of networked mobile devices to amplify and extend children's capabilities to perceive, perform and produce music collaboratively in classroom context. MOGCLASS includes various features for students to enhance their motivation, interest, and collaboration in music class. It provides a wide-ranging palette of easy-to-use musical instruments for students to choose from, and supports both collaborative silent practice with headphones, and collaborative performance with loudspeakers. To facilitate classroom management, the teacher's interface is used to control students' activities. Our evaluation results indicate that MOGCLASS is effective in increasing students' motivation in learning music and in supporting teachers' classroom management
Yinsheng Zhou, Graham Percival, Xinxi Wang, Ye Wang 0007, Shengdong Zhao 0001
ACM Multimedia5
2010 Towards understanding user tolerance to network latency and data rate in remote viewing of progressive meshes
abstract
We conducted experiments with 38 users who interacted with 3 progressively streamed and rendered 3D meshes in order to study their tolerance levels for network data rate and delay. Our study shows that over 90% of users can tolerate a data rate of 80 KBps and above (when the delay is 400ms) and over 95% of users can tolerate delay up to 1 second (when the data rate is 100 KBps). Our study shows that data rate and delay tolerance levels do not vary significantly among the three meshes we used.
Ransi Nilaksha De Silva, Wei Tsang Ooi, Shengdong Zhao 0001
NOSSDAV4
2010 MusicFlow: an interactive music composition system
abstract
Music notation has evolved to the point in which music scores can be digitalized to give composers a different dimension of music composition. Traditional method of music notation using pen and paper requires much time and effort especially during reviewing and editing of hand written music scores. On the other hand of the spectrum, computerizing the entire music composition process can potentially reduce the workload; however, previous approaches in digitizing music notation suffer from the overwhelming functions and the lack of human touch. In this paper, we designed, implemented, and evaluated a multi-touch application called Musicflow, which allows for automatic transcription of composers' music into digital music scores through one's fingertips. To facilitate natural, efficient interaction, MusicFlow supports many multi-touch gestures such as music notation, editing and playing back for reviewing. In addition, Musicflow includes a collaborative teaching tool which further enhances music education by engaging both the teacher and students actively on a multi-touch table. Our initial evaluation indicates that MusicFlow is intuitive to use and effective for music composition.
Sharon Yee Ping Tan, Zhijia Hu, Alan Yih Lun Koh, Felicia Tan 0001, Shengdong Zhao 0001
VCIP5
2009 Magic cards: a paper tag interface for implicit robot control
abstract
Typical Human Robot Interaction (HRI) assumes that the user explicitly interacts with robots. However, explicit control with robots can be unnecessary or even undesirable in certain cases, such as dealing with domestic services (or housework). In this paper, we propose an alternative strategy of interaction: the user implicitly controls a robot by issuing commands on corresponding real world objects and the environment. Robots then discover these commands and complete them in the background. We implemented a paper-tag-based interface to support such implicit robot control in a sensor-augmented home environment. Our initial user studies indicated that the paper-tag-based interface is particularly simple to use and provides users with flexibility in planning and controlling their housework tasks in a simulated home environment.
Shengdong Zhao 0001, Koichi Nakamura, Kentaro Ishii, Takeo Igarashi
CHI1
2009 Designing Laser Gesture Interface for Robot Control
Kentaro Ishii, Shengdong Zhao 0001, Masahiko Inami, Takeo Igarashi, Michita Imai
INTERACT (2)2
2009 Towards characterizing user interaction with progressively transmitted 3D meshes
abstract
10.1145/1631272.1631438
Ransi Nilaksha De Silva, Wei Tsang Ooi, Shengdong Zhao 0001
ACM Multimedia5
2007 InkSeine: In Situ search for active note taking
abstract
Using a notebook to sketch designs, reflect on a topic, or capture and extend creative ideas are examples of active note taking tasks. Optimal experience for such tasks demands concentration without interruption. Yet active note taking may also require reference documents or emails from team members. InkSeine is a Tablet PC application that supports active note taking by coupling a pen-and-ink interface with an in situ search facility that flows directly from a user's ink notes (Fig. 1). InkSeine integrates four key concepts: it leverages preexisting ink to initiate a search; it provides tight coupling of search queries with application content; it persists search queries as first class objects that can be commingled with ink notes; and it enables a quick and flexible workflow where the user may freely interleave inking, searching, and gathering content. InkSeine offers these capabilities in an interface that is tailored to the unique demands of pen input, and that maintains the primacy of inking above all other tasks.
Ken Hinckley, Shengdong Zhao 0001, Raman Sarin, Patrick Baudisch, Edward Cutrell, Michael Shilman, Desney S. Tan
CHI2
2007 Earpod: eyes-free menu selection using touch input and reactive audio feedback
abstract
We present the design and evaluation of earPod: an eyes-free menu technique using touch input and reactive auditory feedback. Studies comparing earPod with an iPod-like visual menu technique on reasonably-sized static menus indicate that they are comparable in accuracy. In terms of efficiency (speed), earPod is initially slower, but outperforms the visual technique within 30 minutes of practice. Our results indicate that earPod is potentially a reasonable eyes-free menu technique for general use, and is a particularly exciting technique for use in mobile device interfaces.
Shengdong Zhao 0001, Pierre Dragicevic, Mark Chignell, Ravin Balakrishnan, Patrick Baudisch
CHI1
2007 The Adaptive Hybrid Cursor: A Pressure-Based Target Selection Technique for Pen-Based User Interfaces
Xiangshi Ren, Jibin Yin, Shengdong Zhao 0001
INTERACT (1)3
2006 Zone and polygon menus: using relative position to increase the breadth of multi-stroke marking menus
abstract
We present Zone and Polygon menus, two new variants of multi-stroke marking menus that consider both the relative position and orientation of strokes. Our menus are designed to increase menu breadth over the 8 item limit of status quo orientation-based marking menus. An experiment shows that Zone and Polygon menus can successfully increase breadth by a factor of 2 or more over orientation-based marking menus, while maintaining high selection speed and accuracy. We also discuss hybrid techniques that may further increase menu breadth and performance. Our techniques offer UI designers new options for balancing menu breadth and depth against selection speed and accuracy.
Shengdong Zhao 0001, Maneesh Agrawala, Ken Hinckley
CHI1
2006 Phosphor: explaining transitions in the user interface using afterglow effects
abstract
Sometimes users fail to notice a change that just took place on their display. For example, the user may have accidentally deleted an icon or a remote collaborator may have changed settings in a control panel. Animated transitions can help, but they force users to wait for the animation to complete. This can be cumbersome, especially in situations where users did not need an explanation. We propose a different approach. Phosphor objects show the outcome of their transition instantly; at the same time they explain their change in retrospect. Manipulating a phosphor slider, for example, leaves an afterglow that illustrates how the knob moved. The parallelism of instant outcome and explanation supports both types of users. Users who already understood the transition can continue interacting without delay, while those who are inexperienced or may have been distracted can take time to view the effects at their own pace. We present a framework of transition designs for widgets, icons, and objects in drawing programs. We evaluate phosphor objects in two user studies and report significant performance benefits for phosphor objects.
Patrick Baudisch, Desney S. Tan, Maxime Collomb, Daniel C. Robbins, Ken Hinckley, Maneesh Agrawala, Shengdong Zhao 0001, Gonzalo A. Ramos
UIST7
2004 Simple vs. compound mark hierarchical marking menus
abstract
We present a variant of hierarchical marking menus where items are selected using a series of inflection-free simple marks, rather than the single "zig-zag" compound mark used in the traditional design. Theoretical analysis indicates that this simple mark approach has the potential to significantly increase the number of items in a marking menu that can be selected efficiently and accurately. A user experiment is presented that compares the simple and compound mark techniques. Results show that the simple mark technique allows for significantly more accurate and faster menu selections overall, but most importantly also in menus with a large number of items where performance of the compound mark technique is particularly poor. The simple mark technique also requires significantly less physical input space to perform the selections, making it particularly suitable for small footprint pen-based input devices. Visual design alternatives are also discussed.
Shengdong Zhao 0001, Ravin Balakrishnan
UIST1
2003 Listen to the Music: Audio Preview Cues for Exploration of Online Music
m. c. schraefel, Maria Karam, Shengdong Zhao 0001
INTERACT3
2002 Hunter gatherer: interaction support for the creation and management of within-web-page collections
abstract
Hunter Gatherer is an interface that lets Web users carry out three main tasks: (1) collect components from within Web pages; (2) represent those components in a collection; (3) edit those component collections. Our research shows that while the practice of making collections of content from within Web pages is common, it is not frequent, due in large part to poor interaction support in existing tools. We engaged with users in task analysis as well as iterative design reviews in order to understand the interaction issues that are part of within-Web-page collection making and to design an interaction that would support that process.We report here on that design development, as well as on the evaluations of the tool that evolved from that process, and the future work stemming from these results, in which our critical question is: what happens to users perceptions and expectations of web-based information (their web-based information management practices) when they can treat this information as harvestable, recontextualizable data, rather than as fixed pages?
m. c. schraefel, Yuxiang Zhu, David Modjeska, Daniel J. Wigdor, Shengdong Zhao 0001
WWW5