EDBT 2026 Demo / reviewers in the wild / expert
Shumin Zhai
dblp:z/ShuminZhai
· DBLP profile ↗
110ranked-venue papers
23as first author
15since 2021 · last 2025
0000-0003-0752-2090ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 106 · 23 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorArtificial intelligence and machine learning · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Simulating Errors in Touchscreen Typingabstract| openaire: EC/HE/101141916/EU//Artificial User Danqing Shi, Yujun Zhu, Francisco Erivaldo Fernandes Junior, Shumin Zhai, Antti Oulasvirta |
CHI | 4 |
| 2025 | Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphonesabstractlayer that integrates the tap location into the LLM's attention mechanism, enabling it to utilize the tap location for text correction. We fine-tuned the touch location-informed LLM on synthetic touch locations and correction commands, achieving significantly higher correction accuracy than the state-of-the-art method VT [45]. A 16-person user study demonstrated that Tap&Say outperforms VT [45] with 16.4% shorter task completion time and 47.5% fewer keyboard clicks and is preferred by users. Maozheng Zhao, Michael Xuelin Huang, Nathan G. Huang, Shanqing Cai, Henry Huang, Michael G. Huang, Shumin Zhai, I. V. Ramakrishnan, Xiaojun Bi 0001 |
CHI | 7 |
| 2024 | Rambler: Supporting Writing With Speech via LLM-Assisted Gist ManipulationabstractDictation enables efficient text input on mobile devices. However, writing with speech can produce disfluent, wordy, and incoherent text and thus requires heavy post-processing. This paper presents Rambler, an LLM-powered graphical user interface that supports gist-level manipulation of dictated text with two main sets of functions: gist extraction and macro revision. Gist extraction generates keywords and summaries as anchors to support the review and interaction with spoken text. LLM-assisted macro revisions allow users to respeak, split, merge, and transform dictated text without specifying precise editing locations. Together they pave the way for interactive dictation and revision that help close gaps between spontaneously spoken words and well-structured writing. In a comparative study with 12 participants performing verbal composition tasks, Rambler outperformed the baseline of a speech-to-text editor + ChatGPT, as it better facilitates iterative revisions with enhanced user control over the content while supporting surprisingly diverse user strategies. Susan Lin, Jeremy Warner, J. D. Zamfirescu-Pereira, Matthew G. Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Björn Hartmann, Can Liu 0003 |
CHI | 9 |
| 2024 | CRTypist: Simulating Touchscreen Typing Behavior via Computational RationalityabstractTouchscreen typing requires coordinating the fingers and visual attention for button-pressing, proofreading, and error correction. Computational models need to account for the associated fast pace, coordination issues, and closed-loop nature of this control problem, which is further complicated by the immense variety of keyboards and users. The paper introduces CRTypist, which generates human-like typing behavior. Its key feature is a reformulation of the supervisory control problem, with the visual attention and motor system being controlled with reference to a working memory representation tracking the text typed thus far. Movement policy is assumed to asymptotically approach optimal performance in line with cognitive and design-related bounds. This flexible model works directly from pixels, without requiring hand-crafted feature engineering for keyboards. It aligns with human data in terms of movements and performance, covers individual differences, and can generalize to diverse keyboard designs. Though limited to skilled typists, the model generates useful estimates of the typing performance achievable under various conditions. Danqing Shi, Yujun Zhu, Jussi P. P. Jokinen, Aditya Acharya, Aini Putkonen, Shumin Zhai, Antti Oulasvirta |
CHI | 6 |
| 2024 | Can Capacitive Touch Images Enhance Mobile Keyboard Decoding?abstractCapacitive touch sensors capture the two-dimensional spatial profile (referred to as a touch heatmap) of a finger’s contact with a mobile touchscreen. However, the research and design of touchscreen mobile keyboards – one of the most speed and accuracy demanding touch interfaces – has focused on the location of the touch centroid derived from the touch image heatmap as the input, discarding the rest of the raw spatial signals. In this paper, we investigate whether touch heatmaps can be leveraged to further improve the tap decoding accuracy for mobile touchscreen keyboards. Specifically, we developed and evaluated machine-learning models that interpret user taps by using the centroids and/or the heatmaps as their input and studied the contribution of the heatmaps to model performance. The results show that adding the heatmap into the input feature set led to 21.4% relative reduction of character error rates on average, compared to using the centroid alone. Furthermore, we conducted a live user study with the centroid-based and heatmap-based decoders built into Pixel 6 Pro devices and observed lower error rate, faster typing speed, and higher self-reported satisfaction score based on the heatmap-based decoder than the centroid-based decoder. These findings underline the promise of utilizing touch heatmaps for improving typing experience in mobile keyboards. Piyawat Lertvittayakumjorn, Shanqing Cai, Billy Dou, Cedric Ho, Shumin Zhai |
UIST | 5 |
| 2024 | SkipWriter: LLM-Powered Abbreviated Writing on TabletsabstractLarge Language Models (LLMs) may offer transformative opportunities for text input, especially for physically demanding modalities like handwriting. We studied a form of abbreviated handwriting by designing, developing, and evaluating a prototype, named SkipWriter, that converts handwritten strokes of a variable-length prefix-based abbreviation (e.g., "ho a y" as handwritten strokes) into the intended full phrase (e.g., "how are you" in the digital format) based on the preceding context. SkipWriter consists of an in-production handwriting recognizer and an LLM fine-tuned on this task. With flexible pen input, SkipWriter allows the user to add and revise prefix strokes when predictions do not match the user’s intent. An user evaluation demonstrated a 60% reduction in motor movements with an average speed of 25.78 WPM. We also showed that this reduction is close to the ceiling of our model in an offline simulation. Zheer Xu, Shanqing Cai, Mukund Varma T., Subhashini Venugopalan, Shumin Zhai |
UIST | 5 |
| 2023 | WordGesture-GAN: Modeling Word-Gesture Movement with Generative Adversarial NetworkabstractWord-gesture production models that can synthesize word-gestures are critical to the training and evaluation of word-gesture keyboard decoders. We propose WordGesture-GAN, a conditional generative adversarial network that takes arbitrary text as input to generate realistic word-gesture movements in both spatial (i.e., (x, y) coordinates of touch points) and temporal (i.e., timestamps of touch points) dimensions. WordGesture-GAN introduces a Variational Auto-Encoder to extract and embed variations of user-drawn gestures into a Gaussian distribution which can be sampled to control variation in generated gestures. Our experiments on a dataset with 38k gesture samples show that WordGesture-GAN outperforms existing gesture production models including the minimum jerk model [37] and the style-transfer GAN [31, 32] in generating realistic gestures. Overall, our research demonstrates that the proposed GAN structure can learn variations in user-drawn gestures, and the resulting WordGesture-GAN can generate word-gesture movement and predict the distribution of gestures. WordGesture-GAN can serve as a valuable tool for designing and evaluating gestural input systems. Jeremy Chu, Dongsheng An, Yan Ma 0006, Wenzhe Cui, Shumin Zhai, Xianfeng Gu, Xiaojun Bi 0001 |
CHI | 5 |
| 2023 | TouchType-GAN: Modeling Touch Typing with Generative Adversarial NetworkabstractModels that can generate touch typing tasks are important to the development of touch typing keyboards. We propose TouchType-GAN, a Conditional Generative Adversarial Network that can simulate locations and time stamps of touch points in touch typing. TouchType-GAN takes arbitrary text as input to generate realistic touch typing both spatially (i.e., (x, y) coordinates of touch points) and temporally (i.e., timestamps of touch points). TouchType-GAN introduces a variational generator that estimates Gaussian Distributions for every target letter to prevent mode collapse. Our experiments on a dataset with 3k typed sentences show that TouchType-GAN outperforms existing touch typing models, including the Rotational Dual Gaussian model [36] for simulating the distribution of touch points, and the Finger-Fitts Euclidean Model [30] for simulating typing time. Overall, our research demonstrates that the proposed GAN structure can learn the distribution of user typed touch points, and the resulting TouchType-GAN can also estimate typing movements. TouchType-GAN can serve as a valuable tool for designing and evaluating touch typing input systems. Jeremy Chu, Yan Ma 0006, Shumin Zhai, Xianfeng Gu, Xiaojun Bi 0001 |
UIST | 3 |
| 2023 | C-PAK: Correcting and Completing Variable-Length Prefix-Based Abbreviated KeystrokesabstractImproving keystroke savings is a long-term goal of text input research. We present a study into the design space of an abbreviated style of text input called C-PAK (Correcting and completing variable-length Prefix-based Abbreviated Keystrokes) for text entry on mobile devices. Given a variable length and potentially inaccurate input string (e.g., “li g t m”), C-PAK aims to expand it into a complete phrase (e.g., “looks good to me”). We develop a C-PAK prototype keyboard, PhraseWriter , based on a current state-of-the-art mobile keyboard consisting of 1.3 million n -grams and 164,000 words. Using computational simulations on a large dataset of realistic input text, we found that, in comparison to conventional single-word suggestions, PhraseWriter improves the maximum keystroke savings rate by 6.7% (from 46.3% to 49.4,), reduces the word error rate by 14.7%, and is particularly advantageous for common phrases. We conducted a lab study of novice user behavior and performance which found that users could quickly utilize the C-PAK style abbreviations implemented in PhraseWriter, achieving a higher keystroke savings rate than forward suggestions (25% vs. 16%). Furthermore, they intuitively and successfully abbreviated more with common phrases. However, users had a lower overall text entry rate due to their limited experience with the system (28.5 words per minute vs. 37.7). We outline future technical directions to improve C-PAK over the PhraseWriter baseline, and further opportunities to study the perceptual, cognitive, and physical action trade-offs that underlie the learning curve of C-PAK systems. Tianshi Li 0001, Philip Quinn, Shumin Zhai |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2022 | TypeAnywhere: A QWERTY-Based Text Entry Solution for Ubiquitous ComputingabstractWe present a QWERTY-based text entry system, TypeAnywhere, for use in off-desktop computing environments. Using a wearable device that can detect finger taps, users can leverage their touch-typing skills from physical keyboards to perform text entry on any surface. TypeAnywhere decodes typing sequences based only on finger-tap sequences without relying on tap locations. To achieve optimal decoding performance, we trained a neural language model and achieved a 1.6% character error rate (CER) in an offline evaluation, compared to a 5.3% CER from a traditional n-gram language model. Our user study showed that participants achieved an average performance of 70.6 WPM, or 80.4% of their physical keyboard speed, and 1.50% CER after 2.5 hours of practice over five days on a table surface. They also achieved 43.9 WPM and 1.37% CER when typing on their laps. Our results demonstrate the strong potential of QWERTY typing as a ubiquitous text entry solution. Mingrui Ray Zhang, Shumin Zhai, Jacob O. Wobbrock |
CHI | 2 |
| 2022 | EyeSayCorrect: Eye Gaze and Voice Based Hands-free Text Correction for Mobile DevicesabstractText correction on mobile devices usually requires precise and repetitive manual control. In this paper, we present EyeSayCorrect, an eye gaze and voice based hands-free text correction method for mobile devices. To correct text with EyeSayCorrect, the user first utilizes the gaze location on the screen to select a word, then speaks the new phrase. EyeSayCorrect would then infer the user’s correction intention based on the inputs and the text context. We used a Bayesian approach for determining the selected word given an eye-gaze trajectory. Given each sampling point in an eye-gaze trajectory, the posterior probability of selecting a word is calculated and accumulated. The target word would be selected when its accumulated interest is larger than a threshold. The misspelt words have higher priors. Our user studies showed that using priors for misspelt words reduced the task completion time up to 23.79% and the text selection time up to 40.35%, and EyeSayCorrect is a feasible hands-free text correction method on mobile devices. Maozheng Zhao, Henry Huang, Zhi Li 0052, Wenzhe Cui, Kajal Toshniwal, Ananya Goel, Sina Rashidian, Furqan Baig, Khiem Phi, Shumin Zhai, I. V. Ramakrishnan, Fusheng Wang 0001, Xiaojun Bi 0001 |
IUI | 13 |
| 2021 | TapNet: The Design, Training, Implementation, and Applications of a Multi-Task Learning CNN for Off-Screen Mobile InputabstractTo make off-screen interaction without specialized hardware practical, we investigate using deep learning methods to process the common built-in IMU sensor (accelerometers and gyroscopes) on mobile phones into a useful set of one-handed interaction events. We present the design, training, implementation and applications of TapNet, a multi-task network that detects tapping on the smartphone. With phone form factor as auxiliary information, TapNet can jointly learn from data across devices and simultaneously recognize multiple tap properties, including tap direction and tap location. We developed two datasets consisting of over 135K training samples, 38K testing samples, and 32 participants in total. Experimental evaluation demonstrated the effectiveness of the TapNet design and its significant improvement over the state of the art. Along with the datasets, codebase1, and extensive experiments, TapNet establishes a new technical foundation for off-screen mobile input. Michael Xuelin Huang, Yang Li 0058, Nazneen Nazneen, Alexander Chao, Shumin Zhai |
CHI | 5 |
| 2021 | PhraseFlow: Designs and Empirical Studies of Phrase-Level InputabstractDecoding on phrase-level may afford more correction accuracy than on word-level according to previous research. However, how phrase-level input affects the user typing behavior, and how to design the interaction to make it practical remain under explored. We present PhraseFlow, a phrase-level input keyboard that is able to correct previous text based on the subsequently input sequences. Computational studies show that phrase-level input reduces the error rate of autocorrection by over 16%. We found that phrase-level input introduced extra cognitive load to the user that hindered their performance. Through an iterative design-implement-research process, we optimized the design of PhraseFlow that alleviated the cognitive load. An in-lab study shows that users could adopt PhraseFlow quickly, resulting in 19% fewer error without losing speed. In real-life settings, we conducted a six-day deployment study with 42 participants, showing that 78.6% of the users would like to have the phrase-level input feature in future keyboards. Mingrui Ray Zhang, Shumin Zhai |
CHI | 2 |
| 2021 | Modeling Touch Point Distribution with Rotational Dual Gaussian ModelabstractTouch point distribution models are important tools for designing touchscreen interfaces. In this paper, we investigate how the finger movement direction affects the touch point distribution, and how to account for it in modeling. We propose the Rotational Dual Gaussian model, a refinement and generalization of the Dual Gaussian model, to account for the finger movement direction in predicting touch point distribution. In this model, the major axis of the prediction ellipse of the touch point distribution is along the finger movement direction, and the minor axis is perpendicular to the finger movement direction. We also propose using projected target width and height, in lieu of nominal target width and height to model touch point distribution. Evaluation on three empirical datasets shows that the new model reflects the observation that the touch point distribution is elongated along the finger movement direction, and outperforms the original Dual Gaussian Model in all prediction tests. Compared with the original Dual Gaussian model, the Rotational Dual Gaussian model reduces the RMSE of touch error rate prediction from 8.49% to 4.95%, and more accurately predicts the touch point distribution in target acquisition. Using the Rotational Dual Gaussian model can also improve the soft keyboard decoding accuracy on smartwatches. Yan Ma 0006, Shumin Zhai, I. V. Ramakrishnan, Xiaojun Bi 0001 |
UIST | 2 |
| 2021 | Voice and Touch Based Error-tolerant Multimodal Text Editing and Correction for SmartphonesabstractEditing operations such as cut, copy, paste, and correcting errors in typed text are often tedious and challenging to perform on smartphones. In this paper, we present VT, a voice and touch-based multi-modal text editing and correction method for smartphones. To edit text with VT, the user glides over a text fragment with a finger and dictates a command, such as "bold" to change the format of the fragment, or the user can tap inside a text area and speak a command such as "highlight this paragraph" to edit the text. For text correcting, the user taps approximately at the area of erroneous text fragment and dictates the new content for substitution or insertion. VT combines touch and voice inputs with language context such as language model and phrase similarity to infer a user's editing intention, which can handle ambiguities and noisy input signals. It is a great advantage over the existing error correction methods (e.g., iOS's Voice Control) which require precise cursor control or text selection. Our evaluation shows that VT significantly improves the efficiency of text editing and text correcting on smartphones over the touch-only method and the iOS's Voice Control method. Our user studies showed that VT reduced the text editing time by 30.80%, and text correcting time by 29.97% over the touch-only method. VT reduced the text editing time by 30.81%, and text correcting time by 47.96% over the iOS's Voice Control method. Maozheng Zhao, Wenzhe Cui, I. V. Ramakrishnan, Shumin Zhai, Xiaojun Bi 0001 |
UIST | 4 |
| 2020 | Modeling Two Dimensional Touch PointingabstractModeling touch pointing is essential to touchscreen interface development and research, as pointing is one of the most basic and common touch actions users perform on touchscreen devices. Finger-Fitts Law [4] revised the conventional Fitts' law into a 1D (one-dimensional) pointing model for finger touch by explicitly accounting for the fat finger ambiguity (absolute error) problem which was unaccounted for in the original Fitts' law. We generalize Finger-Fitts law to 2D touch pointing by solving two critical problems. First, we extend two of the most successful 2D Fitts law forms to accommodate finger ambiguity. Second, we discovered that using nominal target width and height is a conceptually simple yet effective approach for defining amplitude and directional constraints for 2D touch pointing across different movement directions. The evaluation shows our derived 2D Finger-Fitts law models can be both principled and powerful. Specifically, they outperformed the existing 2D Fitts' laws, as measured by the regression coefficient and model selection information criteria (e.g., Akaike Information Criterion) considering the number of parameters. Finally, 2D Finger-Fitts laws also advance our understanding of touch pointing and thereby serve as the basis for touch interface designs. Yu-Jung Ko, Hang Zhao 0005, Yoonsang Kim, I. V. Ramakrishnan, Shumin Zhai, Xiaojun Bi 0001 |
UIST | 5 |
| 2019 | Touchscreen Haptic Augmentation Effects on Tapping, Drag and Drop, and Path FollowingabstractWe study the effects of haptic augmentation on tapping, path following, and drag & drop tasks based on a recent flagship smartphone with refined touch sensing and haptic actuator technologies. Results show actuated haptic confirmation on tapping targets was subjectively appreciated by some users but did not improve tapping speed or accuracy. For drag & drop, a clear performance improvement was measured when haptic feedback is applied to target boundary crossing, particularly when the targets are small. For path following tasks, virtual haptic feedback improved accuracy at a reduced speed in a sitting condition. Stronger results were achieved in a physical haptic mock-up. Overall, we found actuated touchscreen haptic feedback particularly effective when the touched object was visually interfered by the finger. Participants subjective experience of haptic feedback in all tasks tended to be more positive than their time or accuracy performance suggests. We compare and discuss these findings with previous results on early generations of devices. The work provides an empirical foundation to product design and future research of touch input and haptic systems. Mitchell L. Gordon, Shumin Zhai |
CHI | 2 |
| 2019 | Active Edge: Designing Squeeze Gestures for the Google Pixel 2abstractActive Edge is a feature of Google Pixel 2 smartphone devices that creates a force-sensitive interaction surface along their sides, allowing users to perform gestures by holding and squeezing their device. Supported by strain gauge elements adhered to the inner sidewalls of the device chassis, these gestures can be more natural and ergonomic than on-screen (touch) counterparts. Developing these interactions is an integration of several components: (1) an insight and understanding of the user experiences that benefit from squeeze gestures; (2) hardware with the sensitivity and reliability to sense a user's squeeze in any operating environment; (3) a gesture design that discriminates intentional squeezes from innocuous handling; and (4) an interaction design to promote a discoverable and satisfying user experience. This paper describes the design and evaluation of Active Edge in these areas as part of the product's development and engineering. Philip Quinn, Seungyon Claire Lee, Melissa Barnhart, Shumin Zhai |
CHI | 4 |
| 2019 | Text Entry Throughput: Towards Unifying Speed and Accuracy in a Single Performance MetricabstractHuman-computer input performance inherently involves speed-accuracy tradeoffs---the faster users act, the more inaccurate those actions are. Therefore, comparing speeds and accuracies separately can result in ambiguous outcomes: Does a fast but inaccurate technique perform better or worse overall than a slow but accurate one? For pointing, speed and accuracy has been unified for over 60 years as throughput (bits/s) (Crossman 1957, Welford 1968), but to date, no similar metric has been established for text entry. In this paper, we introduce a text entry method-independent throughput metric based on Shannon information theory (1948). To explore the practical usability of the metric, we conducted an experiment in which 16 participants typed with a laptop keyboard using different cognitive sets, i.e., speed-accuracy biases. Our results show that as a performance metric, text entry throughput remains relatively stable under different speed-accuracy conditions. We also evaluated a smartphone keyboard with 12 participants, finding that throughput varied least compared to other text entry metrics. This work allows researchers to characterize text entry performance with a single unified measure of input efficiency. Mingrui Ray Zhang, Shumin Zhai, Jacob O. Wobbrock |
CHI | 2 |
| 2019 | i'sFree: Eyes-Free Gesture Typing via a Touch-Enabled Remote ControlabstractEntering text without having to pay attention to the keyboard is compelling but challenging due to the lack of visual guidance. We propose i'sFree to enable eyes-free gesture typing on a distant display from a touch-enabled remote control. i'sFree does not display the keyboard or gesture trace but decodes gestures drawn on the remote control into text according to an invisible and shifting Qwerty layout. i'sFree decodes gestures similar to a general gesture typing decoder, but learns from the instantaneous and historical input gestures to dynamically adjust the keyboard location. We designed it based on the understanding of how users perform eyes-free gesture typing. Our evaluation shows eyes-free gesture typing is feasible: reducing visual guidance on the distant display hardly affects the typing speed. Results also show that the i'sFree gesture decoding algorithm is effective, enabling an input speed of 23 WPM, 46% faster than the baseline eyes-free condition built on a general gesture decoder. Finally, i'sFree is easy to learn: participants reached 22 WPM in the first ten minutes, even though 40% of them were first-time gesture typing users. Suwen Zhu, Jingjie Zheng, Shumin Zhai, Xiaojun Bi 0001 |
CHI | 3 |
| 2018 | Understanding the Uncertainty in 1D Unidirectional Moving Target SelectionabstractIn contrast to the extensive studies on static target pointing, much less formal understanding of moving target acquisition can be found in the HCI literature. We designed a set of experiments to identify regularities in 1D unidirectional moving target selection, and found a Ternary-Gaussian model to be descriptive of the endpoint distribution in such tasks. The shape of the distribution as characterized by μ and σ in the Gaussian model were primarily determined by the speed and size of the moving target. The model fits the empirical data well with 0.95 and 0.94 R2 values for μ and σ , respectively. We also demonstrated two extensions of the model, including 1) predicting error rates in moving target selection; and 2) a novel interaction technique to implicitly aid moving target selection. By applying them in a game interface design, we observed good performances in both predicting error rates (e.g., 2.7% mean absolute error) and assisting moving target selection (e.g., 33% or a greater increase in pointing accuracy). Jin Huang 0009, Feng Tian 0001, Xiangmin Fan, Xiaolong Zhang 0001, Shumin Zhai |
CHI | 5 |
| 2018 | M3 Gesture Menu: Design and Experimental Analyses of Marking Menus for Touchscreen Mobile InteractionabstractDespite their learning advantages in theory, marking menus have faced adoption challenges in practice, even on today's touchscreen-based mobile devices. We address these challenges by designing, implementing, and evaluating multiple versions of M3 Gesture Menu (M3), a reimagination of marking menus targeted at mobile interfaces. M3 is defined on a grid rather than in a radial space, relies on gestural shapes rather than directional marks, and has constant and stationary space use. Our first controlled experiment on expert performance showed M3 was faster and less error-prone by a factor of two than traditional marking menus. A second experiment on learning demonstrated for the first time that users could successfully transition to recall-based execution of a dozen commands after three ten-minute practice sessions with both M3 and Multi-Stroke Marking Menu. Together, M3, with its demonstrated resolution, learning, and space use benefits, contributes to the design and understanding of menu selection in the mobile-first era of end-user computing. Jingjie Zheng, Xiaojun Bi 0001, Yang Li 0058, Shumin Zhai |
CHI | 5 |
| 2018 | Typing on an Invisible KeyboardabstractA virtual keyboard takes a large portion of precious screen real estate. We have investigated whether an invisible keyboard is a feasible design option, how to support it, and how well it performs. Our study showed users could correctly recall relative key positions even when keys were invisible, although with greater absolute errors and overlaps between neighboring keys. Our research also showed adapting the spatial model in decoding improved the invisible keyboard performance. This method increased the input speed by 11.5% over simply hiding the keyboard and using the default spatial model. Our 3-day multi-session user study showed typing on an invisible keyboard could reach a practical level of performance after only a few sessions of practice: the input speed increased from 31.3 WPM to 37.9 WPM after 20 - 25 minutes practice on each day in 3 days, approaching that of a regular visible keyboard (41.6 WPM). Overall, our investigation shows an invisible keyboard with adapted spatial model is a practical and promising interface option for the mobile text entry systems. Suwen Zhu, Tianyao Luo, Xiaojun Bi 0001, Shumin Zhai |
CHI | 4 |
| 2018 | Modeling Gesture-Typing MovementsabstractWord–Gesture keyboards allow users to enter text using continuous input strokes (also known as gesture typing or shape writing). We developed a production model of gesture typing input based on a human motor control theory of optimal control (specifically, modeling human drawing movements as a minimization of jerk—the third derivative of position). In contrast to existing models, which consider gestural input as a series of concatenated aiming movements and predict a user’s time performance, this descriptive theory of human motor control predicts the shapes and trajectories that users will draw. The theory is supported by an analysis of user-produced gestures that found qualitative and quantitative agreement between the shapes users drew and the minimum jerk theory of motor control. Furthermore, by using a small number of statistical via-points whose distributions reflect the sensorimotor noise and speed–accuracy trade-off in gesture typing, we developed a model of gesture production that can predict realistic gesture trajectories for arbitrary text input tasks. The model accurately reflects features in the figural shapes and dynamics observed from users and can be used to improve the design and evaluation of gestural input systems. Philip Quinn, Shumin Zhai |
Hum. Comput. Interact. | 2 |
| 2017 | Modern Touchscreen Keyboards as Intelligent User Interfaces: A Research ReviewabstractEssential to mobile communication, the touchscreen keyboard is the most ubiquitous intelligent user interface on modern mobile phones. Developing smarter, more efficient, easy to learn, and fun to use keyboards has presented many fascinating IUI research and design questions. Some have been addressed by academic research and practitioners in industry, while others remain significant ongoing research challenges. In this IUI 2017 keynote address I will review and synthesize the progress and open research questions of the past 15 years in text input, focusing on those my co-authors and I have directly dealt with through publications, such as the cost-benefit equations of automation and prediction [9], the power of machine/statistical intelligence [4, 7, 12], the human performance models fundamental to the design of error-correction algorithms [1, 2, 8], spatial scaling from a phone to a watch and the implications on human-machine labor division [5], user behavior and learning innovation [7, 11, 12, 13], and the challenges of evaluating the longitudinal effects of personalization and adaptation [4]. Through this research program review, I will illustrate why intelligent user interfaces, or the combination of machine intelligence and human factors, holds the future of human-computer interaction, and information technology at large. Shumin Zhai |
IUI | 1 |
| 2016 | IJQwerty: What Difference Does One Key Change Make? Gesture Typing Keyboard Optimization Bounded by One Key Position Change from QwertyabstractDespite of a significant body of research in optimizing the virtual keyboard layout, none of them has gained large adoption, primarily due to the steep learning curve. To address this learning problem, we introduced three types of Qwerty constraints, Qwerty1, QwertyH1, and One-Swap bounds in layout optimization, and investigated their effects on layout learnability and performance. This bounded optimization process leads to IJQwerty, which has only one pair of keys different from Qwerty. Our theoretical analysis and user study show that IJQwerty improves the accuracy and input speed of gesture typing over Qwerty once a user reaches the expert mode. IJQwerty is also extremely easy to learn. The initial upon-use text entry speed is the same with Qwerty. Given the high performance and learnability, such a layout will more likely gain large adoption than any of previously obtained layouts. Our research also shows the disparity from Qwerty substantially affects layout learning. To minimize the learning effort, a new layout needs to hold a strong resemblance to Qwerty. Xiaojun Bi 0001, Shumin Zhai |
CHI | 2 |
| 2016 | WatchWriter: Tap and Gesture Typing on a Smartwatch Miniature Keyboard with Statistical DecodingabstractWe present WatchWriter, a finger operated keyboard that supports both touch and gesture typing with statistical decoding on a smartwatch. Just like on modern smartphones, users type one letter per tap or one word per gesture stroke on WatchWriter but in a much smaller spatial scale. WatchWriter demonstrates that human motor control adaptability, coupled with modern statistical decoding and error correction technologies developed for smartphones, can enable a surprisingly effective typing performance despite the small watch size. In a user performance experiment entirely run on a smartwatch, 36 participants reached a speed of 22-24 WPM with near zero error rate. Mitchell L. Gordon, Tom Ouyang, Shumin Zhai |
CHI | 3 |
| 2016 | A Cost-Benefit Study of Text Entry Suggestion InteractionabstractMobile keyboards often present error corrections and word completions (suggestions) as candidates for anticipated user input. However, these suggestions are not cognitively free: they require users to attend, evaluate, and act upon them. To understand this trade-off between suggestion savings and interaction costs, we conducted a text transcription experiment that controlled interface assertiveness: the tendency for an interface to present itself. Suggestions were either always present (extraverted), never present (introverted), or gated by a probability threshold (ambiverted). Results showed that although increasing the assertiveness of suggestions reduced the number of keyboard actions to enter text and was subjectively preferred, the costs of attending to and using the suggestions impaired average time performance. Philip Quinn, Shumin Zhai |
CHI | 2 |
| 2016 | Predicting Finger-Touch Accuracy Based on the Dual Gaussian Distribution ModelabstractAccurately predicting the accuracy of finger-touch target acquisition is crucial for designing touchscreen UI and for modeling complex and higher level touch interaction behaviors. Despite its importance, there has been little theoretical work on creating such models. Building on the Dual Gaussian Distribution Model[3], we derived an accuracy model that predicts the success rate of target acquisition based on the target size. We evaluated the model by comparing the predicted success rates with empirical measures for three types of targets including 1-dimensional vertical and horizontal, and 2-dimensional circular targets. The predictions matched the empirical data very well: the differences between predicted and observed success rates were under 5% for 4.8 mm and 7.2 mm targets, and under 10% for 2.4 mm targets. The evaluation results suggest that our simple model can reliably predict touch accuracy. Xiaojun Bi 0001, Shumin Zhai |
UIST | 2 |
| 2015 | Effects of Language Modeling and its Personalization on Touchscreen Typing PerformanceabstractModern smartphones correct typing errors and learn user-specific words (such as proper names). Both techniques are useful, yet little has been published about their technical specifics and concrete benefits. One reason is that typing accuracy is difficult to measure empirically on a large scale. We describe a closed-loop, smart touch keyboard (STK) evaluation system that we have implemented to solve this problem. It includes a principled typing simulator for generating human-like noisy touch input, a simple-yet-effective decoder for reconstructing typed words from such spatial data, a large web-scale background language model (LM), and a method for incorporating LM personalization. Using the Enron email corpus as a personalization test set, we show for the first time at this scale that a combined spatial-language model reduces word error rate from a pre-model baseline of 38.4% down to 5.7%, and that LM personalization can improve this further to 4.6%. Andrew Fowler, Kurt Partridge, Ciprian Chelba, Xiaojun Bi 0001, Tom Ouyang, Shumin Zhai |
CHI | 6 |
| 2015 | Performance and User Experience of Touchscreen and Gesture Keyboards in a Lab Setting and in the WildabstractWe study the performance and user experience of two popular mainstream mobile text entry methods: the Smart Touch Keyboard (STK) and the Smart Gesture Keyboard (SGK). Our first study is a lab-based ten-session text entry experiment. In our second study we use a new text entry evaluation methodology based on the experience sampling method (ESM). In the ESM study, participants installed an Android app on their own mobile phones that periodically sampled their text entry performance and user experience amid their everyday activities for four weeks. The studies show that text can be entered at an average speed of 28 to 39 WPM, depending on the method and the user's experience, with 1.0% to 3.6% character error rates remaining. Error rates of touchscreen input, particularly with SGK, are a major challenge; and reducing out-of-vocabulary errors is particularly important. Both SGK and STK have strengths, weaknesses, and different individual awareness and preferences. Two-thumb touch typing in a focused setting is particularly effective on STK, whereas one-handed SGK typing with the thumb is particularly effective in more mobile situations. When exposed to both, users tend to migrate from STK to SGK. We also conclude that studies in the lab and in the wild can both be informative to reveal different aspects of keyboard experience, but used in conjunction is more reliable in comprehensively assessing input technologies of current and future generations. Shyam Reyal, Shumin Zhai, Per Ola Kristensson |
CHI | 2 |
| 2015 | Optimizing Touchscreen Keyboards for Gesture TypingabstractDespite its growing popularity, gesture typing suffers from a major problem not present in touch typing: gesture ambiguity on the Qwerty keyboard. By applying rigorous mathematical optimization methods, this paper systematically investigates the optimization space related to the accuracy, speed, and Qwerty similarity of a gesture typing keyboard. Our investigation shows that optimizing the layout for gesture clarity (a metric measuring how unique word gestures are on a keyboard) drastically improves the accuracy of gesture typing. Moreover, if we also accommodate gesture speed, or both gesture speed and Qwerty similarity, we can still reduce error rates by 52% and 37% over Qwerty, respectively. In addition to investigating the optimization space, this work contributes a set of optimized layouts such as GK-D and GK-T that can immediately benefit mobile device users. Brian A. Smith 0001, Xiaojun Bi 0001, Shumin Zhai |
CHI | 3 |
| 2015 | Long short term memory neural network for keyboard gesture decodingabstractGesture typing is an efficient input method for phones and tablets using continuous traces created by a pointed object (e.g., finger or stylus). Translating such continuous gestures into textual input is a challenging task as gesture inputs exhibit many features found in speech and handwriting such as high variability, co-articulation and elision. In this work, we address these challenges with a hybrid approach, combining a variant of recurrent networks, namely Long Short Term Memories [1] with conventional Finite State Transducer decoding [2]. Results using our approach show considerable improvement relative to a baseline shape-matching-based system, amounting to 4% and 22% absolute improvement respectively for small and large lexicon decoding on real datasets and 2% on a synthetic large scale dataset. Ouais Alsharif, Tom Ouyang, Françoise Beaufays, Shumin Zhai, Thomas M. Breuel, Johan Schalkwyk |
ICASSP | 4 |
| 2015 | Differences and Similarities between Finger and Pen Stroke Gestures on Stationary and Mobile devicesabstractThis study investigated differences and similarities between finger and pen gestures on stationary devices (sitting posture) and mobile devices (sitting and walking postures). The recorded gestures were analyzed according to multiple gesture features. We found (1) pen and index finger gestures were different in features like size ratio but similar in features like angle difference ; (2) implement (pen vs. index finger vs. thumb) interacted with gesture complexity and size in features like articulation time ; (3) features like time and shape distance , were different between the pen and index finger on mobile devices (walking) but similar on stationary devices; (4) one-handed thumb gestures had worse performances than index finger gestures by time and accuracy in sitting but similar performances in walking; and (5) for the three implements, gesture drawing time and accuracy on mobile devices reduced from sitting to walking condition. We discuss these findings with implications for future gesture design and research. Huawei Tu, Xiangshi Ren, Shumin Zhai |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2015 | TOCHI Editor-in-Chief Transition: Farewell from Shumin Zhai, Welcome Ken HinckleyabstractNo abstract available. Shumin Zhai |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2014 | Both complete and correct?: multi-objective optimization of touchscreen keyboardabstractCorrecting erroneous input (i.e., correction) and completing a word based on partial input (i.e., completion) are two important "smart" capabilities of a modern intelligent touchscreen keyboard. However little is known whether these two capabilities are conflicting or compatible with each other in the keyboard parameter tuning. Applying computational optimization methods, this work explores the optimality issues related to them. The work demonstrates that it is possible to simultaneously optimize a keyboard algorithm for both correction and completion. The keyboard simultaneously optimized for both introduces no compromise to correction and only a slight compromise to completion when compared to the keyboards exclusively optimized for one objective. Our research also demonstrates the effectiveness of the proposed optimization method in keyboard algorithm design, which is based on the Pareto multi-objective optimization and the Metropolis algorithm. For the development and test datasets used in our experiments, computational optimization improved the correction accuracy rate by 8.3% and completion power by 17.7%. Xiaojun Bi 0001, Tom Ouyang, Shumin Zhai |
CHI | 3 |
| 2014 | Editorial: TOCHI turns twentyabstracteditorial Free Access Share on Editorial: TOCHI turns twenty Editor: Shumin Zhai View Profile Authors Info & Claims ACM Transactions on Computer-Human InteractionVolume 21Issue 1February 2014 Article No.: 1pp 1–4https://doi.org/10.1145/2568193Published:01 February 2014Publication History 1citation635DownloadsMetricsTotal Citations1Total Downloads635Last 12 Months34Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Shumin Zhai |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2013 | Octopus: evaluating touchscreen keyboard correction and recognition algorithms viaabstractThe time and labor demanded by a typical laboratory-based keyboard evaluation are limiting resources for algorithmic adjustment and optimization. We propose Remulation, a complementary method for evaluating touchscreen keyboard correction and recognition algorithms. It replicates prior user study data through real-time, on-device simulation. We have developed Octopus, a Remulation-based evaluation tool that enables keyboard developers to efficiently measure and inspect the impact of algorithmic changes without conducting resource-intensive user studies. It can also be used to evaluate third-party keyboards in a "black box" fashion, without access to their algorithms or source code. Octopus can evaluate both touch keyboards and word-gesture keyboards. Two empirical examples show that Remulation can efficiently and effectively measure many aspects of touch screen keyboards at both macro and micro levels. Additionally, we contribute two new metrics to measure keyboard accuracy at the word level: the Ratio of Error Reduction (RER) and the Word Score. Xiaojun Bi 0001, Shiri Azenkot, Kurt Partridge, Shumin Zhai |
CHI | 4 |
| 2013 | FFitts law: modeling finger touch with fitts' lawabstractFitts' law has proven to be a strong predictor of pointing performance under a wide range of conditions. However, it has been insufficient in modeling small-target acquisition with finger-touch based input on screens. We propose a dual-distribution hypothesis to interpret the distribution of the endpoints in finger touch input. We hypothesize the movement endpoint distribution as a sum of two independent normal distributions. One distribution reflects the relative precision governed by the speed-accuracy tradeoff rule in the human motor system, and the other captures the absolute precision of finger touch independent of the speed-accuracy tradeoff effect. Based on this hypothesis, we derived the FFitts model - an expansion of Fitts' law for finger touch input. We present three experiments in 1D target acquisition, 2D target acquisition and touchscreen keyboard typing tasks respectively. The results showed that FFitts law is more accurate than Fitts' law in modeling finger input on touchscreens. At 0.91 or a greater R2 value, FFitts' index of difficulty is able to account for significantly more variance than conventional Fitts' index of difficulty based on either a nominal target width or an effective target width in all the three experiments. Xiaojun Bi 0001, Yang Li 0058, Shumin Zhai |
CHI | 3 |
| 2013 | Making touchscreen keyboards adaptive to keys, hand postures, and individuals: a hierarchical spatial backoff model approachabstractWe propose a new approach for improving text entry accuracy on touchscreen keyboards by adapting the underlying spatial model to factors such as input hand postures, individuals, and target key positions. To combine these factors together, we introduce a hierarchical spatial backoff model (SBM) that consists of submodels with different levels of complexity. The most general model includes no adaptive factors, whereas the most specific model includes all three. Considering that in practice people may switch hand postures (e.g., from two-thumb to one-finger) to better suit a situation, and that the specific submodels may take time to train for each user, a specific submodel should be applied only if its corresponding input posture can be identified with confidence, and if the submodel has enough training data from the user. We introduce the backoff mechanism to fall back to a simpler model if either of these conditions are not met. We implemented a prototype system capable of reducing the language-model-independent error rate by 13.2% using an online posture classifier with 86.4% accuracy. Further improvements in error rate may be possible with even better posture classification. Tom Ouyang, Kurt Partridge, Shumin Zhai |
CHI | 4 |
| 2013 | Bayesian touch: a statistical criterion of target selection with finger touchabstractTo improve the accuracy of target selection for finger touch, we conceptualize finger touch input as an uncertain process, and derive a statistical target selection criterion, Bayesian Touch Criterion, by combining the basic Bayes' rule of probability with the generalized dual Gaussian distribution hypothesis of finger touch. The Bayesian Touch Criterion selects the intended target as the candidate with the shortest Bayesian Touch Distance to the touch point, which is computed from the touch point to the target center distance and the target size. We give the derivation of the Bayesian Touch Criterion and its empirical evaluation with two experiments. The results showed that for 2-dimensional circular target selection, the Bayesian Touch Criterion is significantly more accurate than the commonly used Visual Boundary Criterion (i.e., a target is selected if and only if the touch point falls within its boundary) and its two variants. Xiaojun Bi 0001, Shumin Zhai |
UIST | 2 |
| 2013 | The impact of candidate display styles for Japanese and Chinese characters on input efficiency
Xiangshi Ren, Shumin Zhai |
Int. J. Hum. Comput. Stud. | 3 |
| 2012 | A comparative evaluation of finger and pen stroke gesturesabstractThis paper reports an empirical investigation in which participants produced a set of stroke gestures with varying degrees of complexity and in different target sizes using both the finger and the pen. The recorded gestures were then analyzed according to multiple measures characterizing many aspects of stroke gestures. Our findings were as follows: (1) Finger drawn gestures were quite different to pen drawn gestures in basic measures including size ratio and average speed. Finger drawn gestures tended to be larger and faster than pen drawn gestures. They also differed in shape geometry as measured by, for example, aperture of closed gestures, corner shape distance and intersecting points deviation; (2) Pen drawn gestures and finger drawn gestures were similar in several measures including articulation time, indicative angle difference, axial symmetry and proportional shape distance; (3) There were interaction effects between gesture implement (finger vs. pen) and target gesture size and gesture complexity. Our findings show that half of the features we tested were performed well enough by the finger. This finding suggests that "finger friendly" systems should exploit these features when designing finger interfaces and avoid using the other features in which the finger does not perform as well as the pen. Huawei Tu, Xiangshi Ren, Shumin Zhai |
CHI | 3 |
| 2012 | Touch behavior with different postures on soft smartphone keyboardsabstractText entry on smartphones is far slower and more error-prone than on traditional desktop keyboards, despite sophisticated detection and auto-correct algorithms. To strengthen the empirical and modeling foundation of smartphone text input improvements, we explore touch behavior on soft QWERTY keyboards when used with two thumbs, an index finger, and one thumb. We collected text entry data from 32 participants in a lab study and describe touch accuracy and precision for different keys. We found that distinct patterns exist for input among the three hand postures, suggesting that keyboards should adapt to different postures. We also discovered that participants' touch precision was relatively high given typical key dimensions, but there were pronounced and consistent touch offsets that can be leveraged by keyboard algorithms to correct errors. We identify patterns in our empirical findings and discuss implications for design and improvements of soft keyboards. Shiri Azenkot, Shumin Zhai |
Mobile HCI | 2 |
| 2012 | Bimanual gesture keyboardabstractGesture keyboards represent an increasingly popular way to input text on mobile devices today. However, current gesture keyboards are exclusively unimanual. To take advantage of the capability of modern multi-touch screens, we created a novel bimanual gesture text entry system, extending the gesture keyboard paradigm from one finger to multiple fingers. To address the complexity of recognizing bimanual gesture, we designed and implemented two related interaction methods, finger-release and space-required, both based on a new multi-stroke gesture recognition algorithm. A formal experiment showed that bimanual gesture behaviors were easy to learn. They improved comfort and reduced the physical demand relative to unimanual gestures on tablets. The results indicated that these new gesture keyboards were valuable complements to unimanual gesture and regular typing keyboards. Xiaojun Bi 0001, Ciprian Chelba, Tom Ouyang, Kurt Partridge, Shumin Zhai |
UIST | 5 |
| 2012 | Multilingual Touchscreen Keyboard Design and OptimizationabstractA keyboard design, once adopted, tends to have a longlasting and worldwide impact on daily user experience. There is a substantial body of research on touch-screen stylus keyboard optimization. Most of it has focused on English only. Applying rigorous mathematical optimization methods and addressing diacritic character design issues, this article expands this body of work to French, Spanish, German, and Chinese. More important and counter to the intuition that optimization by nature is necessarily specific to each language, this article demonstrates that it is possible to find common layouts that are highly optimized across multiple languages for stylus (or single finger) typing. We first obtained a layout that is highly optimized for both English and French input. We then obtained a layout that is optimized for English, French, Spanish, German, and Chinese pinyin simultaneously, reducing its stylus travel distance to about half of QWERTY's for all of the five languages. In comparison to QWERTY's 3.31, 3.... Xiaojun Bi 0001, Barton A. Smith, Shumin Zhai |
Hum. Comput. Interact. | 3 |
| 2011 | Smart phone use by non-mobile business usersabstractThe rapid increase in smart phone capabilities has introduced new opportunities for mobile information access and computing. However, smart phone use may still be constrained by both device affordances and work environments. To understand how current business users employ smart phones and to identify opportunities for improving business smart phone use, we conducted two studies of actual and perceived performance of standard work tasks. Our studies involved 243 smart phone users from a large corporation. We intentionally chose users who primarily work with desktops and laptops, as these "non-mobile" users represent the largest population of business users. Our results go beyond the general intuition that smart phones are better for consuming than producing information: we provide concrete measurements that show how fast reading is on phones and how much slower and more effortful text entry is on phones than on computers. We also demonstrate that security mechanisms are a significant barrier to wider business smart phone use. We offer design suggestions to overcome these barriers. Patti Bao, Jeffrey S. Pierce, Steve Whittaker 0001, Shumin Zhai |
Mobile HCI | 4 |
| 2011 | Understanding information preview in mobile email processingabstractBrowsing a collection of information on a mobile device is a common task, yet it can be difficult due to the small size of mobile displays. A common trade-off offered by many current mobile interfaces is to allow users to switch between an overview and detailed views of particular items. An open question is how much preview of each item to include in the overview. Using a mobile email processing task, we attempted to answer that question. We investigated participants' email processing behaviors under differing preview conditions in a semi-controlled, naturalistic study. We collected log data of participants' actual behaviors as well as their subjective impressions of different conditions. Our results suggest that a moderate level of two to three lines of preview should be the default. The overall benefit of a moderate amount of preview was supported by both positive subjective ratings and fewer transitions between the overview and individual items. Kimberly Weaver, Huahai Yang, Shumin Zhai, Jeffrey S. Pierce |
Mobile HCI | 3 |
| 2010 | Quasi-qwerty soft keyboard optimizationabstractIt has been well understood that optimized soft keyboard layouts improve motor movement efficiency over the standard Qwerty layouts, but have the drawback of long initial visual search time for novice users. To ease the initial searching time on optimized soft keyboards, we explored "Quasi-Qwerty optimization" so that the resulting layouts are close to Qwerty. Our results show that a middle ground between the optimized but new, and the familiar (Qwerty) but inefficient does exist. We show that by allowing letters to move at most one step (key) away from their original positions on Qwerty in an optimization process, one can achieve about half of what free optimization could gain in movement efficiency. An experiment shows that due to users' familiarity with Qwerty, a layout with quasi Qwerty optimization could significantly reduce novice user's visual search time to a level between those of Qwerty and a freely optimized layout. The results in this work provide designers with a new quantitative understanding of the soft keyboard design space. Xiaojun Bi 0001, Barton A. Smith, Shumin Zhai |
CHI | 3 |
| 2010 | SHRIMP: solving collision and out of vocabulary problems in mobile predictive input with motion gestureabstractDictionary-based disambiguation (DBD) is a very popular solution for text entry on mobile phone keypads but suffers from two problems: 1. the resolution of encoding collision (two or more words sharing the same numeric key sequence) and 2. entering out-of-vocabulary (OOV) words. In this paper, we present SHRIMP, a system and method that addresses these two problems by integrating DBD with camera based motion sensing that enables the user to express preference through a tilting or movement gesture. SHRIMP (Small Handheld Rapid Input with Motion and Prediction) runs on camera phones equipped with a standard 12-key keypad. SHRIMP maintains the speed advantage of DBD driven predictive text input while enabling the user to overcome DBD collision and OOV problems seamlessly without even a mode switch. An initial empirical study demonstrates that SHRIMP can be learned very quickly, performed immediately faster than MultiTap and handled OOV words more efficiently than DBD. Shumin Zhai, John F. Canny |
CHI | 2 |
| 2010 | Pen pressure control in trajectory-based interactionabstractThis study presents a series of three experiments that evaluate human capabilities and limitations in using pen-tip pressure as an additional channel of control information in carrying out trajectory tasks such as drawing, writing and gesturing on computer screen. The first experiment measured the natural range of force used in regular drawing and writing tasks. The second experiment tested human performance of maintaining pen-tip pressure at different levels with and without a visual display of the pen pressure. The third experiment, using the steering law paradigm, studied path steering performance as a function of the steering law index of difficulty, steering path type (linear and circular) and pressure precision tolerance interval. The main conclusions of our investigation are as follows. The natural range of pressure used in drawing and writing is concentrated in the 0.82 N (SI force unit Newton (N) is used in this article) to 3.16 N region. The resting force of the pen tip on the screen is between 0.78 N and 1.58 N. Pressure near or below the resting force is markedly more difficult to control. Visual feedback improves pressure-modulated trajectory tasks. Up to six layers of pressure can be controlled in steering tasks, but the error rate changed from 4.9% for one layer of pressure to 35% for six layers. The steering law holds for pressure steering tasks, which enables systematic prediction of successful steering time for a given path's length, width and pressure precision criterion. The steering time can also be modelled as a logarithmic function of pressure control precision ratio σ. Taken together, the current work provides a systematic body of empirical knowledge as basis for future research and design of digital pen applications. Jibin Yin, Xiangshi Ren, Shumin Zhai |
Behav. Inf. Technol. | 3 |
| 2010 | Phone n' Computer: teaming up an information appliance with a PC
Min Yin, Jeffrey S. Pierce, Shumin Zhai |
Pers. Ubiquitous Comput. | 3 |
| 2010 | "Writing with music": Exploring the use of auditory feedback in gesture interfacesabstractWe investigate the use of auditory feedback in pen-gesture interfaces in a series of informal and formal experiments. Initial iterative exploration showed that gaining performance advantage with auditory feedback was possible using absolute cues and state feedback after the gesture was produced and recognized. However, gaining learning or performance advantage from auditory feedback tightly coupled with the pen-gesture articulation and recognition process was more difficult. To establish a systematic baseline, Experiment 1 formally evaluated gesture production accuracy as a function of auditory and visual feedback. Size of gestures and the aperture of the closed gestures were influenced by the visual or auditory feedback, while other measures such as shape distance and directional difference were not, supporting the theory that feedback is too slow to strongly influence the production of pen stroke gestures. Experiment 2 focused on the subjective aspects of auditory feedback in pen-gesture interfaces. Participants' rating on the dimensions of being wonderful and stimulating was significantly higher with musical auditory feedback. Several lessons regarding pen gestures and auditory feedback are drawn from our exploration: a few simple functions such as indicating the pen-gesture recognition results can be achieved, gaining performance and learning advantage through tightly coupled process-based auditory feedback is difficult, pen-gesture sets and their recognizers can be designed to minimize visual dependence, and people's subjective experience of gesture interaction can be influenced using musical auditory feedback. These lessons may serve as references and stepping stones toward future research and development in pen-gesture interfaces with auditory feedback. Tue Haste Andersen, Shumin Zhai |
ACM Trans. Appl. Percept. | 2 |
| 2010 | Foundations for designing and evaluating user interfaces based on the crossing paradigmabstractTraditional graphical user interfaces have been designed with the desktop mouse in mind, a device well characterized by Fitts' law. Yet in recent years, hand-held devices and tablet personal computers using a pen (or fingers) as the primary mean of interaction have become more and more popular. These new interaction modalities have pushed the traditional focus on pointing to its limit. In this paper we explore whether a different paradigm—goal crossing-based on pen strokes—may substitute or complement pointing as another fundamental interaction method. First we describe a study in which we establish that goal crossing is dependent on an index of difficulty analogous to Fitts' law, and that in some settings, goal crossing completion time is shorter or comparable to pointing performance under the same index of difficulty. We then demonstrate the expressiveness of the crossing-based interaction paradigm by implementing CrossY, an application which only uses crossing for selecting commands. CrossY demonstrates that crossing-based interactions can be more expressive than the standard point and click approach. We also show how crossing-based interactions encourage the fluid composition of commands. Finally after observing that users' performance could be influenced by the general direction of travel, we report on the results of a study characterizing this effect. These latter results led us to propose a general guideline for dialog box interaction. Together, these results provide the foundation for the design of effective crossing-based interactions. Georg Apitz, François Guimbretière, Shumin Zhai |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2009 | Using strokes as command shortcuts: cognitive benefits and toolkit supportabstractThis paper investigates using stroke gestures as shortcuts to menu selection. We first experimentally measured the performance and ease of learning of stroke shortcuts in comparison to keyboard shortcuts when there is no mnemonic link between the shortcut and the command. While both types of shortcuts had the same level of performance with enough practice, stroke shortcuts had substantial cognitive advantages in learning and recall. With the same amount of practice, users could successfully recall more shortcuts and make fewer errors with stroke shortcuts than with keyboard shortcuts. The second half of the paper focuses on UI development support and articulates guidelines for toolkits to implement stroke shortcuts in a wide range of software applications. We illustrate how to apply these guidelines by introducing the Stroke Shortcuts Toolkit (SST) which is a library for adding stroke shortcuts to Java Swing applications with just a few lines of code. Caroline Appert, Shumin Zhai |
CHI | 2 |
| 2009 | The performance of touch screen soft buttonsabstractThe introduction of a new generation of attractive touch screen-based devices raises many basic usability questions whose answers may influence future design and market direction. With a set of current mobile devices, we conducted three experiments focusing on one of the most basic interaction actions on touch screens: the operation of soft buttons. Issues investigated in this set of experiments include: a comparison of soft button and hard button performance; the impact of audio and vibrato-tactile feedback; the impact of different types of touch sensors on use, behavior, and performance; a quantitative comparison of finger and stylus operation; and an assessment of the impact of soft button sizes below the traditional 22 mm recommendation as well as below finger width. Seungyon Claire Lee, Shumin Zhai |
CHI | 2 |
| 2008 | Interlaced QWERTY: accommodating ease of visual search and input flexibility in shape writingabstractShape writing is an input technology for touch-screen mobile phones and pen-tablets. To shape write text, the user spells out word patterns by sliding a finger or stylus over a graphical keyboard. The user's trace is then recognized by a pattern recognizer. In this paper we analyze and evaluate various keyboard layouts, including alphabetic, optimized (ATOMIK), QWERTY, and interlaced QWERTY for shape writing. The goodness of a layout for shape writing has two aspects. For users' initial ease of use the letters should be easy to visually locate. For long term use, however, the layout should maximize the imprecision tolerance and writing flexibility for all words. We present empirical studies for the former and mathematical analyses for the latter. Our results led to a new layout, interlaced QWERTY, which offers excellent separation of word shapes, while still maintaining a low visual search time. Many of the findings in our study also apply to traditional soft keyboards tapped with a stylus or one finger. Shumin Zhai, Per Ola Kristensson |
CHI | 1 |
| 2008 | On the ease and efficiency of human-computer interfacesabstractEase and efficiency are two critical qualities of human-computer interfaces. In this address, I will first examine various contributing factors to these two qualities, including recognition and recall, open and closed loop control, controlled and automatic process [Schneider and Shiffrin, 1977; Shiffrin and Schneider, 1977; Schneider and Chein, 2003], mapping directness [Hutchins et al., 1985], unit of operation, and chunking. Shumin Zhai |
ETRA | 1 |
| 2008 | Improving word-recognizers using an interactive lexicon with active and passive wordsabstractThe words a user is likely to write comprise the user's active vocabulary. This vocabulary is considerably smaller than the passive vocabulary of words a user reads. We explore an interactive adaptive lexicon method that separates a large lexicon into active and passive sets, and gradually expands and adapts the active set to reflect the user's active vocabulary. The adaptation is achieved through lightweight interaction as a by product of actual use. The effectiveness of the technique is demonstrated through a computational experiment and a user study. Per Ola Kristensson, Shumin Zhai |
IUI | 2 |
| 2007 | Modeling human performance of pen stroke gesturesabstractThis paper presents a quantitative human performance model of making single-stroke pen gestures within certain error constraints in terms of production time. Computed from the properties of Curves, Line segments, and Corners (CLC) in a gesture stroke, the model may serve as a foundation for the design and evaluation of existing and future gesture-based user interfaces at the basic motor control efficiency level, similar to the role of previous "laws of action" played to pointing, crossing or steering-based user interfaces. We report and discuss our experimental results on establishing and validating the CLC model, together with other basic empirical findings in stroke gesture production. Shumin Zhai |
CHI | 2 |
| 2007 | Hard lessons: effort-inducing interfaces benefit spatial learningabstractInterface designers normally strive for a design that minimises the user's effort. However, when the design's objective is to train users to interact with interfaces that are highly dependent on spatial properties (e.g. keypad layout or gesture shapes) we contend that designers should consider explicitly increasing the mental effort of interaction. To test the hypothesis that effort aids spatial memory, we designed a "frost-brushing" interface that forces the user to mentally retrieve spatial information, or to physically brush away the frost to obtain visual guidance. We report results from two experiments using virtual keypad interfaces -- the first concerns spatial location learning of buttons on the keypad, and the second concerns both location and trajectory learning of gesture shape. The results support our hypothesis, showing that the frost-brushing design improved spatial learning. The participants' subjective responses emphasised the connections between effort, engagement, boredom, frustration, and enjoyment, suggesting that effort requires careful parameterisation to maximise its effectiveness. Andy Cockburn, Per Ola Kristensson, Jason Alexander, Shumin Zhai |
CHI | 4 |
| 2007 | Command strokes with and without preview: using pen gestures on keyboard for command selectionabstractThis paper presents a new command selection method that provides an alternative to pull-down menus in pen-based mobile interfaces. Its primary advantage is the ability forusers to directly select commands from a very large set without the need to traverse menu hierarchies. The proposed method maps the character strings representing the commands onto continuous pen-traces on a stylus keyboard. The user enters a command by stroking part of its character string. We call this method "command strokes." We present the results of three experiments assessing the usefulness of the technique. The first experiment shows that command strokes are 1.6 times faster than the de-facto standard pull-down menus and that users find command strokes more fun to use. The second and third experiments investigate the effect of displaying a visual preview of the currently recognized command while the user is still articulating the command stroke. These experiments show that visual preview does not slow users down and leads to significantly lower error rates and shorter gestures when users enter new unpracticed commands. Per Ola Kristensson, Shumin Zhai |
CHI | 2 |
| 2006 | The benefits of augmenting telephone voice menu navigation with visual browsing and searchabstractAutomatic interactive voice response (IVR) based telephone routing has long been recognized as a frustrating interaction experience. This paper presents a series of experiments examining the benefits of augmenting telephone voice menus with coordinated visual displays and keyword search. The first experiment qualitatively studied callers' experience of having a visual menu on a screen in synchronization with the telephone voice menu tree navigation. The second experiment quantitatively measured callers' performance in time and accuracy with and without visual display augmentation. The third experiment tested keyword search in comparison to visual browsing of telephone menu trees. Study participants uniformly and enthusiastically liked the visual augmentation of voice menus. On average with visual augmentation callers could navigate phone trees 36% faster with 75% fewer errors, and made choices ahead of the voice menu over 60% of the time. Search vs. browsing had similar navigation performance but offered different and complementary user experiences. Overall our studies conclude that telephone voice menu navigation can be significantly improved with a visual channel augmentation, resulting in both business cost reduction and user experience satisfaction. Min Yin, Shumin Zhai |
CHI | 2 |
| 2006 | Camera phone based motion sensing: interaction techniques, applications and performance studyabstractThis paper presents TinyMotion, a pure software approach for detecting a mobile phone user's hand movement in real time by analyzing image sequences captured by the built-in camera. We present the design and implementation of TinyMotion and several interactive applications based on TinyMotion. Through both an informal evaluation and a formal 17-participant user study, we found that 1. TinyMotion can detect camera movement reliably under most background and illumination conditions. 2. Target acquisition tasks based on TinyMotion follow Fitts' law and Fitts law parameters can be used for TinyMotion based pointing performance measurement. 3. The users can use Vision TiltText, a TinyMotion enabled input method, to enter sentences faster than MultiTap with a few minutes of practicing. 4. Using camera phone as a handwriting capture device and performing large vocabulary, multilingual real time handwriting recognition on the cell phone are feasible. 5. TinyMotion based gaming is enjoyable and immediately available for the current generation camera phones. We also report user experiences and problems with TinyMotion based interaction as resources for future design and development of mobile interfaces. Shumin Zhai, John F. Canny |
UIST | 2 |
| 2005 | Conversing with the user based on eye-gaze patternsabstractMotivated by and grounded in observations of eye-gaze patterns in human-human dialogue, this study explores using eye-gaze patterns in managing human-computer dialogue. We developed an interactive system, iTourist, for city trip planning, which encapsulated knowledge of eye-gaze patterns gained from studies of human-human collaboration systems. User study results show that it was possible to sense users' interest based on eye-gaze patterns and manage computer information output accordingly. Study participants could successfully plan their trip with iTourist and positively rated their experience of using it. We demonstrate that eye-gaze could play an important role in managing future multimodal human-computer dialogues. Pernilla Qvarfordt, Shumin Zhai |
CHI | 2 |
| 2005 | RealTourist - A Study of Augmenting Human-Human and Human-Computer Dialogue with Eye-Gaze Overlay
Pernilla Qvarfordt, David Beymer, Shumin Zhai |
INTERACT | 3 |
| 2005 | Relaxing stylus typing precision by geometric pattern matchingabstractFitts' law models the inherent speed-accuracy trade-off constraint in stylus typing. Users attempting to go beyond the Fitts' law speed ceiling will tend to land the stylus outside the targeted key, resulting in erroneous words and increasing users' frustration. We propose a geometric pattern matching technique to overcome this problem. Our solution can be used either as an enhanced spell checker or as a way to enable users to escape the Fitts' law constraint in stylus typing, potentially resulting in higher text entry speeds than what is currently theoretically modeled. We view the hit points on a stylus keyboard as a high resolution geometric pattern. This pattern can be matched against patterns formed by the letter key center positions of legitimate words in a lexicon. We present the development and evaluation of an "elastic" stylus keyboard capable of correcting words even if the user misses all the intended keys, as long as the user's tapping pattern is close enough to the intended word. Per Ola Kristensson, Shumin Zhai |
IUI | 2 |
| 2005 | Dial and see: tackling the voice menu navigation problem with cross-device user experience integrationabstractIVR (interactive voice response) menu navigation has long been recognized as a frustrating interaction experience. We propose an IM-based system that sends a coordinated visual IVR menu to the caller's computer screen. The visual menu is updated in real time in response to the caller's actions. With this automatically opened supplementary channel, callers can take advantages of different modalities over different devices and interact with the IVR system with the ease of graphical menu selection. Our approach of utilizing existing network infrastructure to pinpoint the caller's virtual location and coordinating multiple devices and multiple channels based on users' ID registration can also be more generally applied to create integrated user experiences across a group of devices. Min Yin, Shumin Zhai |
UIST | 2 |
| 2005 | In search of effective text input interfaces for off the desktop computingabstractIt is generally recognized that today's frontier of HCI research lies beyond the traditional desktop computers whose GUI interfaces were built on the foundation of display-pointing device-full keyboard.Many interface challenges arise without such a physical UI foundation.Text writingranging from entering URLs and search queries, filling forms, typing commands, to taking notes and writing emails and chat messages-is one of the hard problems awaiting for solutions in off-desktop computing.This paper summarizes and synthesizes a research program on this topic at the IBM Almaden Research Center.It analyzes various dimensions that constitute a good text input interface; briefly reviews related literature; discusses the evaluation methodology issues of text input; presents the major ideas and results of two systems, ATOMIK and SHARK; and points out current and future directions in the area from our current vantage point. Shumin Zhai, Per Ola Kristensson, Barton A. Smith |
Interact. Comput. | 1 |
| 2005 | Introduction to sensing-based interactionabstractNo abstract available. Shumin Zhai, Victoria Bellotti |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2004 | View size and pointing difficulty in multi-scale navigationabstractUsing a new taxonomy of pointing tasks which includes view pointing beside traditional cursor pointing, we introduce the concept of multi-scale pointing. Analyzing the impact of view size, we demonstrate theoretically and experimentally that (1) the time needed to reach a remotely located target in a multi-scale interface still obeys Fitts' law and (2) the bandwidth of the interaction (i.e., the inverse of Fitts' law slope) is proportional to view size, a relationship bounded by an early ceiling effect. We discuss these results with special reference to navigation in miniaturized and enlarged interfaces. Yves Guiard, Michel Beaudouin-Lafon, Julien Bastin, Dennis Pasveer, Shumin Zhai |
AVI | 5 |
| 2004 | TNT: a numeric keypad based text input methodabstractWith the evolving functionality in television-based (TV-based) information and entertainment appliances, there is an increased need to enable users input text through remote control devices. We present a novel text input method, The Numpad Typer (TNT), for interactive TV, multimedia home terminals or other similar applications. Embodied in a TV remote control and guided by a visual map on the TV screen, TNT was designed for consistent spatial Stimuli-Response (S-R) compatibility and consistency of use. Five users tested TNT in ten sessions of 45-minutes. This initial investigation showed that users on average could type 9.3 and 17.7 correct words per minute with TNT doing the slowest and the fastest session respectively. The study also showed that the users found the TNT method easy to grasp and fun to use. Subjectively the participants felt they mastered the method rather quickly in comparison to their actual speed improvement. Magnus Ingmarsson, David Dinka, Shumin Zhai |
CHI | 3 |
| 2004 | SHARK2: a large vocabulary shorthand writing system for pen-based computersabstractZhai and Kristensson (2003) presented a method of speed-writing for pen-based computing which utilizes gesturing on a stylus keyboard for familiar words and tapping for others. In SHARK2:, we eliminated the necessity to alternate between the two modes of writing, allowing any word in a large vocabulary (e.g. 10,000-20,000 words) to be entered as a shorthand gesture. This new paradigm supports a gradual and seamless transition from visually guided tracing to recall-based gesturing. Based on the use characteristics and human performance observations, we designed and implemented the architecture, algorithms and interfaces of a high-capacity multi-channel pen-gesture recognition system. The system's key components and performance are also reported. Per Ola Kristensson, Shumin Zhai |
UIST | 2 |
| 2004 | Top-down learning strategies: can they facilitate stylus keyboard learning?
Paul Ung-Joon Lee, Shumin Zhai |
Int. J. Hum. Comput. Stud. | 2 |
| 2004 | Characterizing computer input with Fitts' law parameters-the information and non-information aspects of pointing
Shumin Zhai |
Int. J. Hum. Comput. Stud. | 1 |
| 2004 | Speed-accuracy tradeoff in Fitts' law tasks-on the equivalency of actual and nominal pointing precision
Shumin Zhai, Xiangshi Ren |
Int. J. Hum. Comput. Stud. | 1 |
| 2003 | Refining Fitts' law models for bivariate pointingabstractWe investigate bivariate pointing in light of the recent progress in the modeling of univariate pointing. Unlike previous studies, we focus on the effect of target shape (width and height ratio) on pointing performance, particularly when such a ratio is between 1 and 2. Results showed unequal impact of amplitude and directional constraints, with the former dominating the latter. Investigating models based on the notion of weighted Lp norm, we found that our empirical findings were best captured by an Euclidean model with one free weight. This model significantly outperforms the best model to date. Johnny Accot, Shumin Zhai |
CHI | 2 |
| 2003 | High precision touch screen interactionabstractBare hand pointing on touch screens both benefits and suffers from the nature of direct input. This work explores techniques to overcome its limitations. Our goal is to design interaction tools allowing pixel level pointing in a fast and efficient manner. Based on several cycles of iterative design and testing, we propose two techniques: Cross-Keys that uses discrete taps on virtual keys integrated with a crosshair cursor, and an analog Precision-Handle that uses a leverage (gain) effect to amplify movement precision from the user's finger tip to the end cursor. We conducted a formal experiment with these two techniques, in addition to the previously known Zoom-Pointing and Take-Off as baseline anchors. Both subjective and performance measurements indicate that Precision-Handle and Cross-Keys complement existing techniques for touch screen interaction. Pär-Anders Albinsson, Shumin Zhai |
CHI | 2 |
| 2003 | Human on-line response to target expansionabstractMcGuffin and Balakrishnan (M&B) have recently reported evidence that target expansion during a reaching movement reduces pointing time even if the expansion occurs as late as in the last 10% of the distance to be covered by the cursor. While M&B massed their static and expanding targets in separate blocks of trials, thus making expansion predictable for participants, we replicated their experiment with one new condition in which the target could unpredictably expand, shrink, or stay unchanged. Our results show that target expansion occurring as late as in M&B's experiment enhances pointing performance in the absence of expectation. We discuss these findings in terms of the basic human processes that underlie target-acquisition movements, and we address the implications for user interface design by introducing a revised design for the Mac OS X Dock. Shumin Zhai, Stéphane Conversy, Michel Beaudouin-Lafon, Yves Guiard |
CHI | 1 |
| 2003 | Shorthand writing on stylus keyboardabstractWe propose a method for computer-based speed writing, SHARK (shorthand aided rapid keyboarding), which augments stylus keyboarding with shorthand gesturing. SHARK defines a shorthand symbol for each word according to its movement pattern on an optimized stylus keyboard. The key principles for the SHARK design include high efficiency stemmed from layout optimization, duality of gesturing and stylus tapping, scale and location independent writing, Zipf's law, and skill transfer from tapping to shorthand writing due to pattern consistency. We developed a SHARK system based on a classic handwriting recognition algorithm. A user study demonstrated the feasibility of the SHARK method. Shumin Zhai, Per Ola Kristensson |
CHI | 1 |
| 2003 | Candidate Display Styles in Japanese Input
Xiangshi Ren, Kinya Tamura, Shumin Zhai |
INTERACT | 4 |
| 2003 | Collaboration Meets Fitts' Law: Passing Virtual Objects with and without Haptic Force Feedback
Eva-Lotta Sallnäs, Shumin Zhai |
INTERACT | 2 |
| 2003 | Human Movement Performance in Relation to Path Constraint - The Law of Steering in LocomotionabstractWe examine the law of steering - a quantitative model of human movement time in relation to path width and length previously established in hand drawing movement - in a VR locomotion paradigm. Participants drove a simulated vehicle in a virtual environment on paths whose shape and width were manipulated Results showed that the law of steering also applies to locomotion. Participants' mean trial completion times linearly correlated (r/sup 2/ between 0.985 and 0.999) with an index of difficulty quantified as path distance to width ratio for the straight and circular paths used in this experiment. Their average mean and maximum speed was linearly proportional to path width. Such human performance regularity provides a quantitative tool for 3D human machine interface design and evaluation. Shumin Zhai, Rogier Woltjer |
VR | 1 |
| 2002 | More than dotting the i's - foundations for crossing-based interfacesabstractToday's graphical interactive systems largely depend upon pointing actions, i.e. entering an object and selecting it. In this paper we explore whether an alternate paradigm --- crossing boundaries --- may substitute or complement pointing as another fundamental interaction method. We describe an experiment in which we systematically evaluate two target-pointing tasks and four goal-crossing tasks, which differ by the direction of the movement variability constraint (collinear vs. orthogonal) and by the nature of the action (pointing vs. crossing, discrete vs. continuous). We found that participants' temporal performance in each of the six tasks was dependent on the index of difficulty formulated in the same way as in Fitts' law, but that the parameters differ by task. We also found that goal crossing completion time was shorter or no longer than pointing performance under the same index of difficulty. These regularities, as well as qualitative characterizations of crossing actions and their application in HCI, lay the foundation for designing crossing-based user interfaces Johnny Accot, Shumin Zhai |
CHI | 2 |
| 2002 | Movement model, hits distribution and learning in virtual keyboardingabstractIn a ten-session experiment, six participants practiced typing with an expanding rehearsal method on an optimized virtual keyboard. Based on a large amount of in-situ performance data, this paper reports the following findings. First, the Fitts-digraph movement efficiency model of virtual keyboards is revised. The format and parameters of Fitts' law used previously in virtual keyboards research were incorrect. Second, performance limit predictions of various layouts are calculated with the new model. Third, learning with expanding rehearsal intervals for maximum memory benefits is effective, but many improvements of the training algorithm used can be made in the future. Finally, increased visual load when typing previously practiced text did not significantly change users' performance at this stage of learning, but typing unpracticed text did have a performance effect, suggesting a certain degree of text specific learning when typing on virtual keyboards Shumin Zhai, Alison E. Sue, Johnny Accot |
CHI | 1 |
| 2002 | Performance Optimization of Virtual Keyboards
Shumin Zhai, Michael A. Hunter, Barton A. Smith |
Hum. Comput. Interact. | 1 |
| 2001 | Scale effects in steering law tasksabstractInteraction tasks on a computer screen can technically be scaled to a much larger or much smaller sized input control area by adjusting the input device's control gain or the control-display (C-D) ratio. However, human performance as a function of movement scale is not a well concluded topic. This study introduces a new task paradigm to study the scale effect in the framework of the steering law. The results confirmed a U-shaped performance-scale function and rejected straight-line or no-effect hypotheses in the literature. We found a significant scale effect in path steering performance, although its impact was less than that of the steering law's index of difficulty. We analyzed the scale effects in two plausible causes: movement joints shift and motor precision limitation. The theoretical implications of the scale effects to the validity of the steering law, and the practical implications of input device size and zooming functions are discussed in the paper. Johnny Accot, Shumin Zhai |
CHI | 2 |
| 2001 | Chinese input with keyboard and eye-tracking: an anatomical studyabstractChinese input presents unique challenges to the field of human computer interaction. This study provides an anatomical analysis of today's standard Chinese input process, which is based on pinyin, a phonetic spelling system in Roman characters. Through a combination of human performance modeling and experimentation, our study decomposed the Chinese input process into sub-tasks and found that choice reaction time and numeric keying, two component resulted from the large number of homophones in Chinese, were the major usability bottlenecks. Choice reaction alone took 36% of the total input time in our experiment. Numeric keying for multiple candidates selection tends to take the user's attention away from the computer visual screen. We designed and implemented the EASE (Eye Assisted Selection and Entry) system to help maintaining complete touch-typing experience without diverting visual (spacebar) and implicit eye-tracking to replace the numeric keystrokes. Our experiment showed that such a system could indeed work, even with today's imperfecteye-tracking technology. Shumin Zhai, Hui Su |
CHI | 2 |
| 2001 | Telepresence under Exceptional Circumstances: Enriching the Connection to School for Sick Children
Deborah I. Fels, Judith K. Waalen, Shumin Zhai, Patrice L. (Tamar) Weiss |
INTERACT | 3 |
| 2001 | Designing Feedback for an Attentive Office
Teenie Matlock, Christopher S. Campbell, Paul P. Maglio, Shumin Zhai, Barton A. Smith |
INTERACT | 4 |
| 2001 | Optimised Virtual Keyboards with and without Alphabetical Ordering - A Novice User Study
Barton A. Smith, Shumin Zhai |
INTERACT | 2 |
| 2000 | Hand eye coordination patterns in target selectionabstractIn this paper, we describe the use of eye gaze tracking and trajectory analysis in the testing of the performance of input devices for cursor control in Graphical User Interfaces (GUIs). By closely studying the behavior of test subjects performing pointing tasks, we can gain a more detailed understanding of the device design factors that may influence the overall performance with these devices. Our Results show them are many patterns of hand eye coordination at the computer interface which differ from patterns found in direct hand pointing at physical targets (Byrne, Anderson, Douglass, & Matessa, 1999). Barton A. Smith, Janet Ho, Wendy S. Ark, Shumin Zhai |
ETRA | 4 |
| 2000 | Gaze and Speech in Attentive User Interfaces
Paul P. Maglio, Teenie Matlock, Christopher S. Campbell, Shumin Zhai, Barton A. Smith |
ICMI | 4 |
| 2000 | The metropolis keyboard - an exploration of quantitative techniques for virtual keyboard designabstractText entry user interfaces have been a bottleneck of nontraditional computing devices.One of the promising methods is the virtual keyboard on touch screens.Various layouts have been manually designed to replace the dominant QWERTY layout.This paper presents two computerized quantitative design techniques to search for the optimal virtual keyboard.The first technique simulated the dynamics of a keyboard with "digraph springs" between keys, which produced a "Hooke's" keyboard with 41.6 wpm performance.The second technique used a Metropolis random walk algorithm guided by a "Fitts energy" objective function, which produced a "Metropolis" keyboard with 43.1 wpm performance.The paper also models and evaluates the performance of four existing keyboard layouts.We corrected erroneous estimates in the literature and predicted the performance of QWERTY, CHUBON, FITALY, OPTI to be in the neighborhood of 30, 33, 36 and 38 wpm respectively.Our best design was 40% faster than QWERTY and 10% faster than OPTI, illustrating the advantage of quantitative user interface design techniques based on models of human performance over traditional trial and error designs guided by heuristics. Shumin Zhai, Michael A. Hunter, Barton A. Smith |
UIST | 1 |
| 1999 | Performance Evaluation of Input Devices in Trajectory-Based Tasks: An Application of the Steering LawabstractChoosing input devices for interactive systems that best suit users needs remains a challenge, especially consid- ering the increasing number of devices available. The choice often has to be made through empirical evalua- tions. The most frequently used evaluation task hitherto is target acquisition, a task that can be accurately modeled by Fitts law. However, todays use of computer input devices has gone beyond target acquisition alone. In particular, we often need to perform trajectory-based tasks, such as drawing, writing, and navigation. This paper illustrates how a recently discovered model, the steering law, can be applied as an evaluation paradigm complementary to Fitts law. We tested five commonly used computer input devices in two steering tasks, one linear and one circular. Results showed that subjects performance with the five devices could be generally classified into three groups in the following order: 1. the tablet and the mouse, 2. the trackpoint, 3. the touch- pad and the trackball. The steering law proved to hold for all five devices with greater than 0.98 correlation. The ability to generalize the experimental results and the limitations of the steering law are also discussed. Johnny Accot, Shumin Zhai |
CHI | 2 |
| 1999 | Manual and Gaze Input Cascaded (MAGIC) PointingabstractThis work explores a new direction in utilizing eye gaze for computer input. Gaze tracking has long been considered as an alternative or potentially superior pointing method for computer input. We believe that many fundamental limitations exist with traditional gaze pointing. In particular, it is unnatural to overload a perceptual channel such as vision with a motor control task. We therefore propose an alternative approach, dubbed MAGIC (Manual And Gaze Input Cascaded) pointing. With such an approach, pointing appears to the user to be a manual task, used for fine manipulation and selection. However, a large portion of the cursor movement is eliminated by warping the cursor to the eye gaze area, which encompasses the target. Two specific MAGIC pointing techniques, one conservative and one liberal, were designed, analyzed, and implemented with an eye tracker we developed. They were then tested in a pilot study. This early- stage exploration showed that the MAGIC pointing techniques might offer many advantages, including reduced physical effort and fatigue as compared to traditional manual pointing, greater accuracy and naturalness than traditional gaze pointing, and possibly faster speed than manual pointing. The pros and cons of the two techniques are discussed in light of both performance data and subjective reports. Shumin Zhai, Carlos Hitoshi Morimoto, Steven Ihde |
CHI | 1 |
| 1999 | What You Feel Must Be What You See: Adding Tactile Feedback to the Trackpoint
Christopher S. Campbell, Shumin Zhai, Kim M. May, Paul P. Maglio |
INTERACT | 2 |
| 1998 | Quantifying Coordination in Multiple DOF Movement and Its Application to Evaluating 6 DOF Input DevicesabstractStudy of computer input devices has primarily focused on trial completion time and target acquisition errors.To deepen our understanding of input devices, particularly those with high degrees of freedom @OF), this paper explores device influence on the user's ability to coordinate controlled movements in a 3D interface.After reviewing various existing methods, a new measure of quantifying coordination in multiple degrees of freedom, based on movement efficiency, is proposed and applied to the evaluation of two 6 DOF devices: a free-moving position-control device and a desk-top rate-controlled hand controller.Results showed that while the users of the free moving device had shorter completion time than the users of an elastic rate controller, their movement trajectories were less coordinated.These new findings should better inform system designers on development and selection of input devices.Issues such as mental rotation and isomorphism vs. tools operation as means of computer input are also discussed. Shumin Zhai, Paul Milgram |
CHI | 1 |
| 1998 | Manual and Cognitive Benefits of Two-Handed Input: An Experimental StudyabstractOne of the recent trends in computer input is to utilize users' natural bimanual motor skills. This article further explores the potential benefits of such two-handed input. We have observed that bimanual manipulation may bring two types of advantages to human-computer interaction: manual and cognitive. Manual benefits come from increased time-motion efficiency, due to the twice as many degrees of freedom simultaneously available to the user. Cognitive benefits arise as a result of reducing the load of mentally composing and visualizing the task at an unnaturally low level which is imposed by traditional unimanual techniques. Area sweeping was selected as our experimental task. It is representative of what one encounters, for example, when sweeping out the bounding box surrounding a set of objects in a graphics program. Such tasks cannot be modeled by Fitts' Law alone and have not been previously studied in the literature. In our experiments, two bimanual techniques were compared with the conventional one-handed GUI approach. Both bimanual techniques employed the two-handed “stretchy” technique first demonstrated by Krueger in 1983. We also incorporated the “Toolglass” technique introduced by Bier et al. in 1993. Overall, the bimanual techniques resulted in significantly faster performance than thestatus quoone-handed technique, and these benefits increased with the difficulty of mentally visualizing the task, supporting our bimanual cognitive advantage hypothesis. There was no significant difference between the two bimanual techniques. This study makes two types of contributions to the literature. First, practically we studied yet another class of transaction where significant benefits can be realized by applying bimanual techniques. Furthermore, we have done so using easily available commercial hardware in the context to our understanding of why bimanual interaction techniques have an advantage over unimanual techniques. A literature review on two-handed computer input and some of the relevant bimanual human mototr control studies is also included. Andrea Leganchuk, Shumin Zhai, William Buxton |
ACM Trans. Comput. Hum. Interact. | 2 |
| 1997 | Beyond Fitts' Law: Models for Trajectory-Based HCI TasksabstractArticle Free Access Share on Beyond Fitts' law: models for trajectory-based HCI tasks Authors: Johnny Accot Centre d'Études de la, Navigation Aérienne, 7 avenue Edouard Belin, 31055 Toulouse cedex, France and Input Research Group, CSRI, University of Toronto, Toronto, ON M5S 1A4, Canada Centre d'Études de la, Navigation Aérienne, 7 avenue Edouard Belin, 31055 Toulouse cedex, France and Input Research Group, CSRI, University of Toronto, Toronto, ON M5S 1A4, CanadaView Profile , Shumin Zhai Input Research Group, CSRI, University of Toronto, Toronto, ON M5S 1A4, Canada and IBM Almaden, Research Center, 650 Harry Road, San Jose, CA Input Research Group, CSRI, University of Toronto, Toronto, ON M5S 1A4, Canada and IBM Almaden, Research Center, 650 Harry Road, San Jose, CAView Profile Authors Info & Claims CHI '97: Proceedings of the ACM SIGCHI Conference on Human factors in computing systemsMarch 1997 Pages 295–302https://doi.org/10.1145/258549.258760Published:27 March 1997Publication History 328citation6,971DownloadsMetricsTotal Citations328Total Downloads6,971Last 12 Months602Last 6 weeks96 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Johnny Accot, Shumin Zhai |
CHI | 2 |
| 1997 | An Isometric Tongue Pointing DeviceabstractNo abstract available. Chris Salem, Shumin Zhai |
CHI | 2 |
| 1997 | Improving Browsing Performance: A study of four input devices for scrolling and pointing tasks
Shumin Zhai, Barton A. Smith, Ted Selker |
INTERACT | 1 |
| 1997 | Graphical Means of Directing User's Attention in the Visual Interface
Shumin Zhai, Julie Wright, Ted Selker, Sabra-Anne Kelin |
INTERACT | 1 |
| 1997 | Anisotropic human performance in six degree-of-freedom tracking: an evaluation of three-dimensional display and control interfacesabstractMotivated by the need for human performance evaluations of advanced interface technologies, this paper presents an empirical evaluation of a 3D interface, from the point of view of both display and control, in a pursuit tracking experiment. The paper derives methods for decomposing tracking performance into six dimensions (three in translation and three in rotation). This dimensional decomposition approach has the advantage of revealing overall performance levels in the depth dimension relative to performance in the horizontal and vertical dimensions. With interposition, linear perspective, stereoscopic disparity and partial occlusion cues incorporated into a single 3D display system, subjects' tracking errors in the depth dimension were about 35% (with no practice) to 35% (with practice) larger than those in the horizontal and vertical dimensions. It was also found that subjects initially had larger tracking errors along the vertical axis than along the horizontal axis, likely due to their attention allocation strategy. Analysis of rotation errors generated a similar anisotropic pattern. Shumin Zhai, Paul Milgram, Anu Rastogi |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 1996 | The Influence of Muscle Groups on Performance of Multiple Degree-of-Freedom InputabstractThe literature has long suggested that the design of computer input devices should make use of the fine, smaller muscle groups and joints in the fingers, since they are richly represented in the human motor and sensory cortex and they have higher information processing bandwidth than other body parts.This hypothesis, however, has not been conclusively verified with empirical research.The present work studied such a hypothesis in the context of designing 6 degree-of-freedom (DOF) input devices.The work attempts to address both a practical need designing efficient 6 DOF input devices -and the theoretical issue of muscle group differences in input control.Two alternative 6 DOF input devices, one including and the other excluding the fingers from the 6 DOF manipulation, were designed and tested in a 3D object docking experiment.Users' task completion times were significantly shorter with the device that utilised the fingers.The results of this study strongly suggest that the shape and size of future input device designs should constitute affordances that invite finger participation in input control. K e y w o r d sInput devices, 3-D interface, 6 DOF input, motor control, muscle group differences, hand, fingers, arm, homunculus model. Shumin Zhai, Paul Milgram, William Buxton |
CHI | 1 |
| 1996 | The Partial-Occlusion Effect: Utilizing Semitransparency in 3D Human-Computer InteractionabstractThis study investigates human performance when using semitransparent tools in interactive 3D computer graphics environments. The article briefly reviews techniques for presenting depth information and examples of applying semitransparency in computer interface design. We hypothesize that when the user moves a semitransparent surface in a 3D environment, the “partial-occlusion” effect introduced through semitransparency acts as an effective cue in target localization—an essential component in many 3D interaction tasks. This hypothesis was tested in an experiment in which subjects were asked to capture dynamic targets (virtual fish) with two versions of a 3D box cursor, one with and one without semitransparent surfaces. Results showed that the partial-occlusion effect through semitransparency significantly improved users' performance in terms of trial completion time, error rate, and error magnitude in both monoscopic and stereoscopic displays. Subjective evaluations supported the conclusions drawn from performance measures. The experimental results and their implications are discussed, with emphasis on the relative, discrete nature of the partial-occlusion effect and on interactions between different depth cues. The article concludes with proposals of a few future research issues and applications of semitransparency in human-computer interaction. Shumin Zhai, William Buxton, Paul Milgram |
ACM Trans. Comput. Hum. Interact. | 1 |
| 1994 | The "Silk Cursor": investigating transparency for 3D target acquisitionabstractThis study investigates dynamic 3D target acquisition.The focus is on the relative effect of specific perceptual cues.A novel technique is introduced and we report on an experiment that evaluates its effectiveness.There are two aspects to the new technique.First, in contrast to normal practice, the tracking symbol is a volume rather than a point.Second, the surface of this volume is semi-transparent, thereby affording occlusion cues during target acquisition.The experiment shows that the volumehcclusion cues were effective in both monocular and stereoscopic conditions.For some tasks where stereoscopic presentation is unavailable or infeasible, the new techniaue offers an effective alternative. Shumin Zhai, William Buxton, Paul Milgram |
CHI | 1 |
| 1993 | Applications of augmented reality for human-robot communicationabstractThe director/agent (D/A) metaphor of telerobotic interaction is discussed as a potential means for achieving human-robot synergy. In order for the human operator to communicate spatial information to the robot during D/A operations, the medium of augmented reality through overlaid virtual stereographics is proposed, leading to what is referred to as virtual control. An overview is given of the ARGOS (Augmented Reality through Graphic Overlays on Stereovideo) system. In particular, the uses of overlaid virtual pointers for enhancing absolute depth judgement tasks, virtual tape measures for real-world quantification, virtual tethers for perceptual enhancement in manual teleoperation, virtual landmarks for enhancing depth scaling, and virtual object overlays for on-object edge enhancement and display superposition are all presented and discussed. Paul Milgram, Shumin Zhai, David Drascic, Julius Grodski |
IROS | 2 |
| 1993 | From Icons to Interface Models: Designing Hypermedia from the Bottom Up
John A. Waterworth, Mark Chignell, Shumin Zhai |
Int. J. Man Mach. Stud. | 3 |
| 1993 | Virtual Reality for Palmtop ComputersabstractWe are exploring how virtual reahty theories can be applied toward palmtop computers. In our prototype, called the Chameleon, a small 4-inch hand-held monitor acts as a palmtop computer with the capabihties of a Silicon graphics workstation. A 6D input device and a response button are attached to tbe small monitor to detect user gestures and input selections for issuing commands. An experiment was conducted to evaluate our design and to see how well depth could be perceived in the small screen compared to a large 21-inch screen, and the extent to which movement of the small display (in a palmtop virtual reality condition) could improve depth perception, Results show that with very little training, perception of depth in the palmtop virtual reality condition is about as good as corresponding depth perception in a large (but static) display. Variations to the initial design are also discussed, along with issues to be explored in future research, Our research suggests that palmtop virtual reality may support effective navigation and search and retrieval, in rich and portable information spaces. George W. Fitzmaurice, Shumin Zhai, Mark Chignell |
ACM Trans. Inf. Syst. | 2 |