Per Ola Kristensson

dblp:k/PerOlaKristensson · DBLP profile ↗
← Back
139ranked-venue papers
16as first author
71since 2021 · last 2026
0000-0002-7139-871XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 99 · 13 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 4 first-author · 37 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning
abstract
Reinforcement learning (RL)-based biomechanical simulations have the potential to revolutionize HCI research and interaction design, but currently lack usability and interpretability. Using the Human Action Cycle as a design lens, we identify key limitations of biomechanical RL frameworks and develop MyoInteract, a novel framework for fast prototyping of biomechanical HCI tasks. MyoInteract allows designers to setup tasks, user models, and training parameters from an easy-to-use GUI within minutes. It trains and evaluates muscle-actuated simulated users within minutes, reducing training times by up to 98%. A workshop study with 12 interaction designers revealed that MyoInteract allowed novices in biomechanical RL to successfully setup, train, and assess goal-directed user movements within a single session. By transforming biomechanical RL from a days-long expert task into an accessible hour-long workflow, this work significantly lowers barriers to entry and accelerates iteration cycles in HCI biomechanics research.
Ankit Bhattarai, Hannah Selder, Florian Fischer 0001, Arthur Fleig, Per Ola Kristensson
DIS5
2026 Cost-Aware Bayesian Optimization for Interactive Devices
abstract
Deciding which idea is worth prototyping is a central concern in iterative design. A prototype should be produced when the expected improvement is high and the cost is low. However, this is hard to decide, because costs can vary drastically: a simple parameter tweak may take seconds, while fabricating hardware consumes material and energy. Such asymmetries, can discourage a designer from exploring the design space. In this paper, we present an extension of cost-aware Bayesian optimization to account for diverse prototyping costs. The method builds on the power of Bayesian optimization and requires only a minimal modification to the acquisition function. The key idea is to use designer-estimated costs to guide sampling toward more cost-effective prototypes. In technical evaluations, the method achieved comparable utility to a cost-agnostic baseline while requiring only \({\approx }70\%\) of the cost; under strict budgets, it outperformed the baseline threefold. A within-subjects study with 12 participants in a realistic joystick design task demonstrated similar benefits. These results show that accounting for prototyping costs can make Bayesian optimization more compatible with real-world design projects.
Thomas Langerak, Renate Zhang, Per Ola Kristensson, Antti Oulasvirta
CHI4
2026 How Do We Evaluate Experiences in Immersive Environments?
abstract
How do we evaluate experiences in immersive environments? Despite decades of research in immersive technologies such as virtual reality, the field remains fragmented. Studies rely on overlapping constructs, heterogeneous instruments, and little agreement on what counts as immersive experience. To better understand this landscape, we conducted a bottom-up scoping review of 375 papers published in ACM CHI, UIST, VRST, SUI, IEEE VR, ISMAR, and TVCG. Our analysis reveals that evaluation practices are often domain- and purpose-specific, shaped more by local choices than by shared standards. Yet this diversity also points to new directions. Instead of multiplying instruments, researchers benefit from integrating and refining them into smarter measures. Rather than focusing only on system outputs, evaluations must center the user’s lived experience. Computational modeling offers opportunities to bridge signals across methods, but lasting progress requires open and sustainable evaluation practices that support comparability and reuse. Ultimately, our contribution is to map current practices and outline a forward-looking agenda for immersive experience research.
Xiang Li 0101, Wei He 0028, Per Ola Kristensson
CHI3
2026 Unbounded: Object-Boundary Interaction in Mixed Reality
abstract
Boundaries such as walls, windows, and doors are ubiquitous in the physical world, yet their potential in mixed reality (MR) remains underexplored. We present Unbounded, a Research through Design inquiry into object--boundary interaction (OBI). Building on prior work, we articulate a design space aimed at providing a shared language for OBI. To demonstrate its potential, we design and implement eight examples across productivity and art exploration scenarios, showcasing how OBIs can enrich and reframe everyday interactions. We further engage with six MR experts in one-on-one feedback sessions, using the design space and examples as design probes. Their reflections broaden the conceptual scope of OBI, reveal new possibilities for how the framework may be applied, and highlight implications for future MR interaction design. https://www.zhuoyuelyu.com/unbounded
Zhuoyue Lyu, Per Ola Kristensson
CHI2
2026 Objestures: Everyday Objects Meet Mid-Air Gestures for Expressive Interaction
abstract
Everyday object-based interactions (EOIs) and mid-air gesture interactions (MAIs) have been widely explored, yet prior work on their integration often targets narrow use cases or specific technologies, leaving designers and developers with limited guidance that generalizes across diverse EOIs and MAIs. We introduce Objestures (“Obj” + “Gestures”)—five interaction types spanning EOIs and MAIs, forming a design space for expressive uni- and bimanual interaction. To evaluate the usefulness of Objestures, we conducted an exploratory user study (N = 12) on basic 3D tasks (rotation and scaling), which showed performance comparable to the headset’s native freehand manipulation. To understand the user experience, we conducted case studies with the same participants across three applications (Sound, Draw, and Shadow), where participants found the interactions intuitive, engaging, and expressive, and indicated interest in everyday use. We further demonstrate the potential of Objestures across diverse contexts through 30 examples, and discuss limitations and implications.
Zhuoyue Lyu, Per Ola Kristensson
CHI2
2026 Grand Challenges around Designing Computers' Control Over Our Bodies
abstract
Advances in emerging technologies, such as on-body mechanical actuators and electrical muscle stimulation, have allowed computers to take control over our bodies. This presents opportunities as well as challenges, raising fundamental questions about agency and the role of our bodies when interacting with technology. To advance this research field as a whole, we brought together expert perspectives in a week-long seminar to articulate the grand challenges that should be tackled when it comes to the design of computers’ control over our bodies. These grand challenges span technical, design, user, and ethical aspects. By articulating these grand challenges, we aim to begin initiating a research agenda that positions bodily control not only as a technical feature but as a central, experiential, and ethical concern for future human–computer interaction endeavors.
Florian 'Floyd' Mueller, Nadia Bianchi-Berthouze, Misha Sra, Mar González-Franco, Henning Pohl, Susanne Boll, Richard Byrne 0001, Arthur Pitzer Caetano, Masahiko Inami, Jarrod Knibbe, Per Ola Kristensson, Xiang Li 0101, Zhuying Li 0001, Joe Marshall, Louise Petersen Matjeka, Minna Orvokki Nygren, Rakesh Patibanda, Sara Price, Harald Reiterer, Aryan Saini, Oliver Schneider 0006, Ambika Shahu, Phoebe O. Toups Dugas, Samitha Elvitigala
CHI11
2026 Optimal Explanations: A Quantitative Model of Human Error in Causal Graph Interpretation
abstract
When Artificial Intelligence (AI) reasoning is explained via causal graphs for human oversight, the human-computer interface is the performance bottleneck for decision-supported actions. As explanations grow more complex, humans’ interpretation ability degrades, resulting in ineffective oversight. This paper contributes a quantitative model of human causal reasoning bounds and demonstrates their utility for interpretable AI explanations.
Paul-David Joshua Zuercher, Thomas Bohné, Per Ola Kristensson
IUI3
2026 Human-Inspired Perspectives: A Survey on AI Long-Term Memory
abstract
With the rapid advancement of AI systems, their abilities to store, retrieve, and utilize information over the long term - referred to as long-term memory - have become increasingly significant. These capabilities are crucial for enhancing the performance of AI systems across a wide range of tasks. However, there is currently no comprehensive survey that systematically investigates AI's long-term memory capabilities, formulates a theoretical framework, and inspires the development of next-generation AI long-term memory systems. This paper begins by introducing the mechanisms of human long-term memory, then explores AI long-term memory mechanisms, establishing a mapping between the two. Based on the mapping relationships identified, we extend the current cognitive architectures and propose the Cognitive Architecture of Self-Adaptive Long-term Memory (SALM). SALM provides a theoretical framework for the practice of AI long-term memory and holds potential for guiding the creation of next-generation long-term memory driven AI systems. Finally, we delve into the future directions and application prospects of AI long-term memory.
Zihong He, Weizhe Lin, Fan Zhang 0017, Matt W. Jones, Laurence Aitchison, Xuhai Xu, Miao Liu 0007, Hai-Ning Liang, Per Ola Kristensson, Junxiao Shen
Proc. IEEE10
2026 Sensorimotor Regularities as Alignment between Humans and Large Language Models
abstract
Large Language Models (LLMs) do not construct conceptual representations in ways that align with human cognition, posing risks for human–AI interaction. While LLMs solely rely on linguistic distributional knowledge, humans leverage both linguistic and sensorimotor knowledge. To systematically assess human-LLM alignment in concept representations, we propose a novel evaluation framework based on sensorimotor regularities, operationalized as image schemas—multimodal gestalts derived from repeated sensorimotor experiences. Investigating linguistic manifestations of such schemas, we systematically identify human-LLM alignments and misalignments in the encoding of sensorimotor regularities. Results indicate that three contemporary disembodied LLMs encode highly human-like sensorimotor gestalts. However, these models exhibit reduced alignment when mapping such gestalts to concepts, and they do not systematically combine these gestalts in ways consistent with human patterns. We identify each LLM’s misalignments with human patterns in image schema distribution, conceptual associations, and image schema co-occurrences. Building on these findings, we augment gpt-4-1106 with targeted sensorimotor priors derived from its identified misalignments with human patterns. In a downstream user study, this augmentation yields sentence continuations rated by humans as significantly more conceptually clear, contextually contingent, and human-like than baseline outputs. Our work establishes a foundation for evaluating and improving human-LLM alignment at the conceptual level.
Jinghui Hu, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.3
2026 Preemptive, Buffered, or Guided? Empirical Studies on Human-AI Interaction Strategies for Software Test Case Development
abstract
Applications of large language models (LLMs) in software engineering are soaring quickly. Despite ample evidence supporting LLMs’ capabilities of generating high-quality and creative suggestions, human–LLM interactions in software testing remain under-explored, and there lacks empirical evidence supporting efficacious human–LLM system designs. In software testing, the industry’s pursuit of “autonomous” systems neglects intensive human involvement in prior-design and post-review processes. This article discusses two empirical user studies for a test case brainstorming task: (1) exploring user behaviors in human–LLM interactions compared to web search ( \(N_{1}=16\) ) and (2) investigating three modified interaction strategies—preemptive prompting, buffered response, and guided input ( \(N_{2}=24\) ). We consolidate nine cross-disciplinary metrics to quantitatively evaluate the holistic performance of the human–LLM system, covering three perspectives: test quality, creativity, and attention span. Our findings reveal that users spend 126% more time interacting with LLMs compared to Google search. Moreover, preemptively prompting the LLM system significantly improves test quality and task creativity by over 30%, while simultaneously reducing user idle time by up to 49%. Based on the results, this article discusses three interaction design principles—mixed initiative, acceptability, and appropriation—as guidance for future iterations of an efficacious LLM-assisted software testing system.
Billy Shi, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.2
2026 EnVisionVR: A Scene Interpretation Tool for Visual Accessibility in Virtual Reality
abstract
Effective visual accessibility in Virtual Reality (VR) is crucial for Blind and Low Vision (BLV) users. However, designing visual accessibility systems is challenging due to the complexity of 3D VR environments and the need for techniques that can be easily retrofitted into existing applications. While prior work has studied how to enhance or translate visual information, the advancement of Vision Language Models (VLMs) provides an exciting opportunity to advance the scene interpretation capability of current systems. This paper presents EnVisionVR, an accessibility tool for VR scene interpretation. Through a formative study of usability barriers, we confirmed the lack of visual accessibility features as a key barrier for BLV users of VR content and applications. In response, we used our findings from the formative study to inform the design and development of EnVisionVR, a novel visual accessibility system leveraging a VLM, voice input and multimodal feedback for scene interpretation and virtual object interaction in VR. An evaluation with 12 BLV users demonstrated that EnVisionVR significantly improved their ability to locate virtual objects, effectively supporting scene understanding and object interaction.
Rosella P. Galindo Esparza, Vanja Garaj, Per Ola Kristensson, John J. Dudley
IEEE Trans. Vis. Comput. Graph.4
2026 LocoScooter: Designing a Stationary Scooter-Based Locomotion System for Navigation in Virtual Reality
abstract
Virtual locomotion remains a challenge in VR, especially in space-limited environments where room-scale walking is impractical. We present LocoScooter, a low-cost, deployable locomotion interface combining foot-sliding on a compact treadmill with handlebar steering inspired by scooter riding. Built from commodity hardware, it supports embodied navigation through familiar, physically engaging movement. In a within-subject study $(N = 14)$, LocoScooter significantly improved immersion, enjoyment, and bodily involvement over joystick navigation, while maintaining comparable efficiency and usability. Despite higher physical demand, users did not report increased fatigue, suggesting familiar movements can enrich VR navigation.
Wei He 0028, Xiang Li 0101, Per Ola Kristensson, Ge Lin 0001
IEEE Trans. Vis. Comput. Graph.3
2026 Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
abstract
As virtual reality (VR) continues to evolve, traditional input methods such as handheld controllers and gesture systems often face challenges with precision, social accessibility, and user fatigue. These limitations motivate the exploration of microgestures, which promise more subtle, ergonomic, and device-free interactions. We introduce microGEXT, a lightweight microgesture-based system designed for text editing in VR without external sensors, which utilizes small, subtle hand movements to reduce physical strain compared to standard gestures. We evaluated microGEXT in three user studies. In Study 1 ($N=20$N=20), microGEXT reduced overall edit time and fatigue compared to a ray-casting + pinch menu baseline, the default text editing approach in commercial VR systems. Study 2 ($N=20$N=20) found that microGEXT performed well in short text selection tasks but was slower for longer text ranges. In Study 3 ($N=10$N=10), participants found microGEXT intuitive for open-ended information-gathering tasks. Across all studies, microGEXT demonstrated enhanced user experience and reduced physical effort, offering a promising alternative to traditional VR text editing techniques.
Xiang Li 0101, Wei He 0028, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.3
2026 Assessing the Readiness of Augmented Reality for Industrial Assembly: A Deployment Study Comparing Immersive and Non-Immersive Solutions
abstract
Integrating augmented reality (AR) into industrial assembly processes has great potential, yet most studies use simplified lab tasks and rarely benchmark AR against commercial systems, limiting their industrial relevance. To address this gap, we present a comparative study of immersive AR and non-immersive touchscreen-based assembly support systems on a manufacturing shop floor, with 16 participants assembling two industry-grade assets of different complexity. The results reveal that while AR's immersive capabilities excelled in spatial guidance for complex tasks, the simplicity and reliability of the touchscreen interface proved more effective for simple assemblies. We derive three design implications for deployment practice.
Xinyi Tu 0001, Benedikt Hartmann, Per Ola Kristensson, John Liu, Thomas Bohné, Slawomir Konrad Tadeja
IEEE Trans. Vis. Comput. Graph.3
2026 Stiffness Discrimination Thresholds of Force-Feedback Gloves for Volumetric Data Exploration in Virtual Reality
abstract
Can force-feedback gloves deliver the haptic fidelity required for medical training or engineering design in VR? We systematically measured stiffness discrimination thresholds using Dexmo force-feedback gloves through psychophysical experiments with 23 participants exploring virtual volumetric data. Our dual-method approach-adaptive staircase procedures and population-level psychometric functions-revealed Just Noticeable Difference (JND) thresholds of 48.1% (Low Reference) and 26.3% (High Reference), far exceeding the 10-20% thresholds achievable with natural touch. This performance gap, compounded by a 1.8:1 Weber fraction ratio indicating perceptual scaling violations, raises questions about current haptic VR capabilities. However, spatial analysis uncovered actionable insights: positioning interactions at mid-reach positions (avoiding near-body and extended-arm extremes) improves accuracy by 7%. Our findings establish critical benchmarks for haptic VR development and provide evidence-based guidelines for designing applications that work within current technological constraints rather than assuming real-world haptic equivalence.
Qisong Wang, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.2
2025 Human-Computer Interaction for AI Systems Design: Reflections on an Online Course on Human-AI Interaction for Professionals
abstract
Human–Computer Interaction for AI Systems Design is an eight-week short online course aimed at professional students. It is part of an online course platform called Cambridge Advance Online, which is a joint effort between Cambridge University Press & Assessment and the University of Cambridge. This course launched in July 2023 amidst a massive increase in interest in AI and its applications, and quickly became one of the platform's highest-enrolling courses, attracting about 50 students per quarterly course run. To date, more than 200 students have completed the course, and more than 90 percent have rated their experience `good' or `excellent'. This paper reports on our experiences in designing and teaching this course.
Per Ola Kristensson, Emily Patterson
AAAI1
2025 Exclusion Rates among Disabled and Older Users of Virtual and Augmented Reality
abstract
This paper examines the levels of exclusion encountered by disabled and older users of consumer-level VR and AR technology and identifies methods formed by people with diverse access needs to circumvent encountered barriers to use. First, we estimate exclusion rates for a selection of nine immersive experiences of VR and AR, computed using population statistics data for the United Kingdom (UK). We then present an empirical lab-based study evaluating the usability of the same VR and AR experiences. The study involved 60 UK-based participants with varying access needs and the study results were used to calculate the empirical exclusion rates. Both the estimated and empirical exclusion rates display high levels of exclusion, which for the more complex experiences in the study reached 100 %. However, multiple participants overcame usability barriers and completed experiences through provided assistance and self-initiated adaptations, suggesting that future VR and AR can become more inclusive if designed to counter these barriers.
Rosella P. Galindo Esparza, John J. Dudley, Vanja Garaj, Per Ola Kristensson
CHI4
2025 Seeing and Touching the Air: Unraveling Eye-Hand Coordination in Mid-Air Gesture Typing for Mixed Reality
abstract
Mid-air text entry in mixed reality (MR) headsets has shown promise but remains less efficient than traditional input methods. While research has focused on improving typing performance, the mechanics of mid-air gesture typing, especially eye-hand coordination, are less understood. This paper investigates visuomotor coordination of mid-air gesture keyboards through a user study (n = 16) comparing gesture typing on a tablet and in mid-air. Through an expert task we demonstrate that users were able to achieve a comparable text input performance. Our in-depth analysis of eye-hand coordination reveals significant differences in the eye-hand coordination patterns between gesture typing on a tablet and in-air. The mid-air gesture typing necessitates almost all of the visual attention on the keyboard area and a more consistent synchronization in eye-hand coordination to compensate for the increased motor and cognitive demands without physical boundaries. These insights provide important implications for the design of more efficient text input methods.
Jinghui Hu, John J. Dudley, Per Ola Kristensson
CHI3
2025 Making Hardware Devices at Scale is Still Hard: Challenges and Opportunities for the HCI Community
Bo Kang, Steve Hodges 0001, Per Ola Kristensson
CHI3
2025 AlphaPIG: The Nicest Way to Prolong Interactive Gestures in Extended Reality
abstract
Mid-air gestures serve as a common interaction modality across Extended Reality (XR) applications, enhancing engagement and ownership through intuitive body movements. However, prolonged arm movements induce shoulder fatigue—known as "Gorilla Arm Syndrome"—degrading user experience and reducing interaction duration. Although existing ergonomic techniques derived from Fitts’ law (such as reducing target distance, increasing target width, and modifying control-display gain) provide some fatigue mitigation, their implementation in XR applications remains challenging due to the complex balance between user engagement and physical exertion. We present AlphaPIG, a meta-technique designed to Prolong Interactive Gestures by leveraging real-time fatigue predictions. AlphaPIG assists designers in extending and improving XR interactions by enabling automated fatigue-based interventions. Through adjustment of intervention timing and intensity decay rate, designers can explore and control the trade-off between fatigue reduction and potential effects such as decreased body ownership. We validated AlphaPIG’s effectiveness through a study (N=22) implementing the widely-used Go-Go technique. Results demonstrated that AlphaPIG significantly reduces shoulder fatigue compared to non-adaptive Go-Go, while maintaining comparable perceived body ownership and agency. Based on these findings, we discuss positive and negative perceptions of the intervention. By integrating real-time fatigue prediction with adaptive intervention mechanisms, AlphaPIG constitutes a critical first step towards creating fatigue-aware applications in XR.
Yi Li 0058, Florian Fischer 0001, Tim Dwyer, Barrett Ens, Robert George Crowther, Per Ola Kristensson, Benjamin Tag
CHI6
2025 ReachVox: Clutter-Free Reachability Visualization for Robot Motion Planning in Virtual Reality
abstract
Human-Robot-Collaboration can enhance workflows by leveraging the mutual strengths of human operators and robots. Planning and understanding robot movements remain major challenges in this domain. This problem is prevalent in dynamic environments that might need constant robot motion path adaptation. In this paper, we investigate whether a minimalistic encoding of the reachability of a point near an object of interest, which we call ReachVox, can aid the collaboration between a remote operator and a robotic arm in VR. Through a user study ($\mathrm{n}=20$), we indicate the strength of the visualization relative to a point-based reachability check-up.
Steffen Hauck, Diar Abdlkarim, John J. Dudley, Per Ola Kristensson, Eyal Ofek, Jens Grubert
ISMAR4
2025 Demystifying Reward Design in Reinforcement Learning for Upper Extremity Interaction: Practical Guidelines for Biomechanical Simulations in HCI
abstract
Designing effective reward functions is critical for reinforcement learning-based biomechanical simulations, yet HCI researchers and practitioners often waste (computation) time with unintuitive trial-and-error tuning. This paper demystifies reward function design by systematically analyzing the impact of effort minimization, task completion bonuses, and target proximity incentives on typical HCI tasks such as pointing, tracking, and choice reaction. We show that proximity incentives are essential for guiding movement, while completion bonuses ensure task success. Effort terms, though optional, help refine motion regularity when appropriately scaled. We perform an extensive analysis of how sensitive task success and completion time depend on the weights of these three reward components. From these results we derive practical guidelines to create plausible biomechanical simulations without the need for reinforcement learning expertise, which we then validate on remote control and keyboard typing tasks. This paper advances simulation-based interaction design and evaluation in HCI by improving the efficiency and applicability of biomechanical user modeling for real-world interface development.
Hannah Selder, Florian Fischer 0001, Per Ola Kristensson, Arthur Fleig
UIST3
2025 Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
abstract
As more applications of large language models (LLMs) for 3D content in immersive environments emerge, it is crucial to study user behavior to identify interaction patterns and potential barriers to guide the future design of immersive content creation and editing systems which involve LLMs. In an empirical user study with 12 participants, we combine quantitative usage data with post-experience questionnaire feedback to reveal common interaction patterns and key barriers in LLM-assisted 3D scene editing systems. We identify opportunities for improving natural language interfaces in 3D design tools and propose design recommendations. Through an empirical study, we demonstrate that LLM-assisted interactive systems can be used productively in immersive environments.
Jens Grubert, Per Ola Kristensson
VR3
2025 On the Benefits of Sensorimotor Regularities as Design Constraints for Superpower Interactions in Mixed Reality
abstract
Mixed Reality (MR) systems enable users to perform augmented superpowers that transcend real-world limitations. However, it remains unclear what types of action-outcome mappings can enable users to easily learn, control, and feel a sense of ownership of these augmented superpowers. Humans develop a set of sensorimotor regularities (i.e., image schemas and lawful relations between them) from recurring bodily experiences since early infancy, and use them to predict the outcome of our actions, or choose actions based on the desired outcome. We investigate whether sensorimotor regularities (SRs) can serve as effective design constraints for superpower interactions, by comparing three temporal manipulation methods in MR games: (1) mid-air button control; (2) gestures incongruent with SRs embedded in the concept of temporal manipulation; and (3) gestures congruent with these SRs. A within-subject study with 18 participants reveals that the SRs-congruent method enables significantly improved task performance, lower overall workload, and a greater sense of agency and presence compared to both an SRs-incongruent method and a mid-air button-based method. The SRs-congruent method also enabled faster mastery of the augmented superpower. No significant difference was observed in any of the above-mentioned metrics between the SRs-incongruent and mid-air button-based methods. These results empirically demonstrate multiple benefits of SRs as design constraints for superpower controls in MR, and encourage future research to explore their wider applicability in superpower interaction design.
Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.2
2025 Working in Extended Reality in the Wild: Worker and Bystander Experiences of XR Virtual Displays in Public Real-World Settings
abstract
Although access to sufficient screen space is crucial to knowledge work, workers often find themselves with limited access to display infrastructure in remote or public settings. While virtual displays can be used to extend the available screen space through extended reality (XR) head-worn displays (HWD), we must better understand the implications of working with them in public settings from both users' and bystanders' viewpoints. To this end, we conducted two user studies. We first explored the usage of a hybrid AR display across real-world settings and tasks. We focused on how users take advantage of virtual displays and what social and environmental factors impact their usage of the system. A second study investigated the differences between working with a laptop, an AR system, or a VR system in public. We focused on a single location and participants performed a predefined task to enable direct comparisons between the conditions while also gathering data from bystanders. The combined results suggest a positive acceptance of XR technology in public settings and show that virtual displays can be used to accompany existing devices. We highlighted some environmental and social factors. We saw that previous XR experience and personality can influence how people perceive the use of XR in public. In addition, we confirmed that using XR in public still makes users stand out and that bystanders are curious about the devices, yet have no clear understanding of how they can be used.
Leonardo Pavanatto, Verena Biener, Jennifer Chandran, Snehanjali Kalamkar, Feiyu Lu 0001, John J. Dudley, Jinghui Hu, Gabriella N. Ramirez, Per Ola Kristensson, Alexander Giovannelli, Luke Schlueter, Jörg Müller 0001, Jens Grubert, Doug A. Bowman
IEEE Trans. Vis. Comput. Graph.9
2025 Handows: A Palm-Based Interactive Multi-Window Management System in Virtual Reality
abstract
Window management in virtual reality (VR) remains a challenging task due to the spatial complexity and physical demands of current interaction methods. We introduce Handows, a palm-based interface that enables direct manipulation of spatial windows through familiar smartphone-inspired gestures on the user's non-dominant hand. Combining ergonomic layout design with body-centric input and passive haptics, Handows supports four core operations: window selection, closure, positioning, and scaling. We evaluate Handows in a user study (N = 15) against two common VR techniques (virtual hand and controller) across four core window operations. Results show that Handows significantly reduces physical effort and head movement while improving task efficiency and interaction precision. A follow-up case study (N = 8) demonstrates Handows' usability in realistic multitasking scenarios, highlighting user-adapted workflows and spontaneous layout strategies. Our findings also suggest the potential of embedding mobile-inspired metaphors into proprioceptive body-centric interfaces to support low-effort and spatially coherent interaction in VR.
Jin-Du Wang, Per Ola Kristensson, Xiang Li 0101
IEEE Trans. Vis. Comput. Graph.4
2025 Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality
abstract
The rising popularity of 360-degree images and virtual reality (VR) has spurred a growing interest among creators in producing visually appealing content through effective color grading processes. Although existing computational approaches have simplified the global color adjustment for entire images with Preferential Bayesian Optimization (PBO), they neglect local colors for points of interest and are not optimized for the immersive nature of VR. In response, we propose a dual-level PBO framework that integrates global and local color adjustments tailored for VR environments. We design and evaluate a novel context-aware preferential Gaussian Process (GP) to learn contextual preferences for local colors, taking into account the dynamic contexts of previously established global colors. Additionally, recognizing the limitations of desktop-based interfaces for comparing 360-degree images, we design three VR interfaces for color comparison. We conduct a controlled user study to investigate the effectiveness of the three VR interface designs and find that users prefer to be enveloped by one 360-degree image at a time and to compare two rather than four color-graded options.
Linping Yuan, John J. Dudley, Per Ola Kristensson, Huamin Qu
IEEE Trans. Vis. Comput. Graph.3
2024 On the Benefits of Image-Schematic Metaphors when Designing Mixed Reality Systems
abstract
A Mixed Reality (MR) system encompasses various aspects, such as visualization and spatial registration of user interface elements, user interactions and interaction feedback. Image-schematic metaphors (ISMs) are universal knowledge structures shared by a wide range of users. They hold a theoretical promise of facilitating greater ease of learning and use for interactive systems without costly adaptations. This paper investigates whether image-schematic metaphors (ISMs) can improve user learning, by comparing an existing MR instruction authoring system with or without ISM enhancements. In a user study with 32 participants, we found that the ISM-enhanced system significantly improved task performance, learnability and mental efficiency compared to the baseline. Participants also rated the ISM-enhanced system significantly higher in terms of perspicuity, efficiency, and novelty. These results empirically demonstrate multiple benefits of ISMs when integrated into the design of this MR system and encourage further studies to explore the wider applicability of ISMs in user interface design.
Per Ola Kristensson
CHI2
2024 Efficient Mid-Air Text Input Correction in Virtual Reality
abstract
The task of inputting text within virtual reality has attracted significant research attention over the last five years. Less well explored is the related task of correcting inputted text when errors are made. This is despite the fact that considerable time and frustration stems from efforts to correct text. In this paper, we bridge this gap in prior research and explore efficient methods for supporting text input correction in virtual reality. We present a characterization of the types and frequencies of errors encountered when inputting text in virtual reality and an analysis of effective editing strategies. We also present the results of a user study evaluating the performance and usability trade-offs for several interaction methods leveraging the unique capabilities of modern head-mounted displays.
John J. Dudley, Amy Karlson, Kashyap Todi, Hrvoje Benko, Matt Longest, Robert Wang 0002, Per Ola Kristensson
ISMAR7
2024 LookUP: Command Search Using Dwell-free Eye Typing in Mixed Reality
abstract
We introduce LookUP, a novel general purpose command search system for mixed reality headsets, offering a hands-free experience through dwell-free eye typing. With LookUP, users can trigger the display of a virtual keyboard with a simple upward head motion. The keyboard then uses a statistical decoder to interpret users’ intended text based on their eye movements. This approach diverges from traditional dwell-time methods, significantly enhancing typing speed and efficiency. Our research involved deploying LookUP on a HoloLens 2, and benchmarking it against a dwell-based command search baseline and the native HoloLens system menu. Our user study indicated that participants spent a significantly shorter time using LookUP with dwell-free eye typing in command search and entry, demonstrating LookUP’s potential to be a complementary command input for mixed reality headsets.
Jinghui Hu, John J. Dudley, Per Ola Kristensson
ISMAR3
2024 Accented Character Entry Using Physical Keyboards in Virtual Reality
abstract
Research on text entry in Virtual Reality (VR) has gained popularity but the efficient entry of accented characters, characters with diacritical marks, in VR remains underexplored. Entering accented characters is supported on most capacitive touch keyboards through a long press on a base character and a subsequent selection of the accented character. However, entering those characters on physical keyboards is still challenging, as they require a recall and an entry of respective numeric codes. To address this issue this paper investigates three techniques to support accented character entry on physical keyboards in VR. Specifically, we compare a contextaware numeric code technique that does not require users to recall a code, a key-press-only condition in which the accented characters are dynamically remapped to physical keys next to a base character, and a multimodal technique, in which eye gaze is used to select the accented version of a base character previously selected by keypress on the keyboard. The results from our user study $(n=18)$ reveal that both the key-press-only and the multimodal technique outperform the baseline technique in terms of text entry speed.
Snehanjali Kalamkar, Verena Biener, Daniel Pauls, Leon Lindlein, Morteza Izadifar, Per Ola Kristensson, Jens Grubert
ISMAR6
2024 Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
abstract
We depend on our own memory to encode, store, and retrieve our experiences. However, memory lapses can occur. One promising avenue for achieving memory augmentation is through the use of augmented reality head-mounted displays to capture and preserve egocentric videos, a practice commonly referred to as lifelogging. However, a significant challenge arises from the sheer volume of video data generated through lifelogging, as the current technology lacks the capability to encode and store such large amounts of data efficiently. Further, retrieving specific information from extensive video archives requires substantial computational power, further complicating the task of quickly accessing desired content. To address these challenges, we propose a memory augmentation agent that involves leveraging natural language encoding for video data and storing them in a vector database. This approach harnesses the power of large vision language models to perform the language encoding process. Additionally, we propose using large language models to facilitate natural language querying. Our agent underwent extensive evaluation using the QA-Ego4D dataset and achieved state-of-the-art results with a BLEU score of 8.3, outperforming conventional machine learning models that scored between 3.4 and 5.8. Additionally, we conducted a user study in which participants interacted with the human memory augmentation agent through episodic memory and open-ended questions. The results of this study show that the agent results in significantly better recall performance on episodic memory tasks compared to human participants. The results also highlight the agent’s practical applicability and user acceptance.
Junxiao Shen, John J. Dudley, Per Ola Kristensson
ISMAR3
2024 Towards Open-World Gesture Recognition
abstract
Providing users with accurate gestural interfaces, such as gesture recognition based on wrist-worn devices, is a key challenge in mixed reality. However, static machine learning processes in gesture recognition assume that training and test data come from the same underlying distribution. Unfortunately, in real-world applications involving gesture recognition, such as gesture recognition based on wrist-worn devices, the data distribution may change over time. We formulate this problem of adapting recognition models to new tasks, where new data patterns emerge, as open-world gesture recognition (OWGR). We propose the use of continual learning to enable machine learning models to be adaptive to new tasks without degrading performance on previously learned tasks. However, the process of exploring parameters for questions around when, and how, to train and deploy recognition models requires resource-intensive user studies may be impractical. To address this challenge, we propose a design engineering approach that enables offline analysis on a collected large-scale dataset by systematically examining various parameters and comparing different continual learning methods. Finally, we provide design guidelines to enhance the development of an open-world wrist-worn gesture recognition process.
Junxiao Shen, Matthias De Lange, Xuhai Xu, Enmin Zhou, Ran Tan, Naveen Suda, Maciej Lazarewicz, Per Ola Kristensson, Amy Karlson, Evan Strasnick
ISMAR8
2024 SkiMR: Dwell-free Eye Typing in Mixed Reality
abstract
We present SkiMR: a dwell-free eye typing system that enables fast and accurate hands-free text entry on mixed reality headsets. SkiMR uses a statistical decoder to infer users’ intended text based on users’ eye movements on a virtual keyboard. It does not rely on dwell timeouts for key selections, which enables it to be faster than traditional eye typing. We study this dwell-free eye typing system, deployed on a HoloLens 2, in two studies. In the first study (n = 12) we show that dwell-free eye typing results in a significantly faster text entry rate compared to traditional dwell-based eye typing with word prediction support, and a hybrid dwell-free method that uses dwell timeouts to delimit word entry. Based on the insights from the first study we evaluate the feasibility of a refined system in a more realistic composition task and use an interaction mechanism that provides real-time predictions during dwell-free eye typing. The second study (n = 16) demonstrates that this final system allow users to compose original text at 12 words per minute with a corrected character error rate of 1.1%. Overall, this work demonstrates the high potential for fast and accurate hands-free text entry using dwell-free eye typing for mixed reality headsets.
Jinghui Hu, John J. Dudley, Per Ola Kristensson
VR3
2024 Swarm manipulation: An efficient and accurate technique for multi-object manipulation in virtual reality
abstract
The theory of swarm control shows promise for controlling multiple objects, however, scalability is hindered by cost constraints, such as hardware and infrastructure. Virtual Reality (VR) can overcome these limitations, but research on swarm interaction in VR is limited. This paper introduces a novel swarm manipulation technique and compares it with two baseline techniques: Virtual Hand and Controller (ray-casting). We evaluated these techniques in a user study ( N = 12) in three tasks (selection, rotation, and resizing) across five conditions. Our results indicate that swarm manipulation yielded superior performance, with significantly faster speeds in most conditions across the three tasks. It notably reduced resizing size deviations but introduced a trade-off between speed and accuracy in the rotation task. Additionally, we conducted a follow-up user study ( N = 6) using swarm manipulation in two complex VR scenarios and obtained insights through semi-structured interviews, shedding light on optimized swarm control mechanisms and perceptual changes induced by this interaction paradigm. These results demonstrate the potential of the swarm manipulation technique to enhance the usability and user experience in VR compared to conventional manipulation techniques. In future studies, we aim to understand and improve swarm interaction via internal swarm particle cooperation. • We define swarm manipulation in VR and evaluate its usability and utility in two user studies. • Study 1 compares swarm manipulation to hand and controller techniques across five conditions. • Study 2 explores swarm manipulation in complex scenarios with detailed user feedback. • We analyze strengths and limitations of swarm manipulation in usability and workload. • We provide design insights for integrating swarm manipulation into VR applications.
Xiang Li 0101, Jin-Du Wang, John J. Dudley, Per Ola Kristensson
Comput. Graph.4
2024 Investigating Creation Perspectives and Icon Placement Preferences for On-Body Menus in Virtual Reality
abstract
On-body menus present a novel interaction paradigm within Virtual Reality (VR) environments by embedding virtual interfaces directly onto the user’s body. Unlike traditional screen-based interfaces, on-body menus enable users to interact with virtual options or icons visually attached to their physical form. In this paper, We investigated the impact of the creation process on the effectiveness of on-body menus, comparing first-person, third-person, and mirror perspectives. Our first study ( N = 12) revealed that the mirror perspective led to faster creation times and more accurate recall compared to the other two perspectives. To further explore user preferences, we conducted a second study ( N = 18) utilizing a VR system with integrated body tracking. By combining distributions of icons from both studies ( N = 30), we confirmed significant preferences in on-body menu placement based on icon category ( e.g. , Social Media icons were consistently placed on forearms). We also discovered associations between categories, such as Leisure and Social Media icons frequently co-occurring. Our findings highlight the importance of the creation process, uncover user preferences for on-body menu organization, and provide insights to guide the development of intuitive and effective on-body interactions within virtual environments.
Xiang Li 0101, Wei He 0028, Shan Jin 0002, Jan Gugenheimer, Pan Hui 0001, Hai-Ning Liang, Per Ola Kristensson
Proc. ACM Hum. Comput. Interact.7
2024 40 Years of Eye Typing: Challenges, Gaps, and Emergent Strategies
abstract
Gaze interaction enables users to communicate through eye tracking, and is often the only channel of effective and efficient communication for individuals with severe motor disabilities. While there has been significant research and development of eye typing systems, in the context of augmentative and alternative communication (AAC), there is no comprehensive review that integrates the key findings from the variety of aspects that constitute the complex landscape of gaze communication. This paper presents a detailed review and characterization of the literature and aims to consolidate the disparate efforts to provide eye typing solutions for AAC users. We provide a systematic understanding of the components and functionalities that underpin eye typing solutions, and analyze the interplay of the different facets and their role in shaping the user-experience, accessibility, performance, and overall effectiveness of eye typing technology. We also identify the major challenges and highlight several areas that require further research attention.
Aleesha Hamid, Per Ola Kristensson
Proc. ACM Hum. Comput. Interact.2
2024 Cooperative Multi-Objective Bayesian Design Optimization
abstract
Computational methods can potentially facilitate user interface design by complementing designer intuition, prior experience, and personal preference. Framing a user interface design task as a multi-objective optimization problem can help with operationalizing and structuring this process at the expense of designer agency and experience. While offering a systematic means of exploring the design space, the optimization process cannot typically leverage the designer’s expertise in quickly identifying that a given “bad” design is not worth evaluating. We here examine a cooperative approach where both the designer and optimization process share a common goal and work in partnership by establishing a shared understanding of the design space. We tackle the research question: How can we foster cooperation between the designer and a systematic optimization process in order to best leverage their combined strength? We introduce and present an evaluation of a cooperative approach that allows the user to express their design insight and work in concert with a multi-objective design process. We find that the cooperative approach successfully encourages designers to explore more widely in the design space than when they are working without assistance from an optimization process. The cooperative approach also delivers design outcomes that are comparable to an optimization process run without any direct designer input but achieves this with greater efficiency and substantially higher designer engagement levels.
George B. Mo, John J. Dudley, Li-Wei Chan 0001, Yi-Chi Liao 0001, Antti Oulasvirta, Per Ola Kristensson
ACM Trans. Interact. Intell. Syst.6
2024 Guiding the Design of Inclusive Interactive Systems: Do Younger and Older Adults Use the Same Image-schematic Metaphors?
abstract
The use of image-schematic metaphors is often promoted for being near-universal across user groups, suggesting that these metaphors have the potential to make novel interactive systems easy to use by both younger and older adults. This study empirically investigates this by eliciting image-schematic metaphors from the spoken language and interaction behaviors of 12 younger adults and 12 older adults undertaking tasks in a technology learning domain. For the first time, we reveal an almost-perfect overlap between image-schematic metaphors used by the younger and older groups, despite the two groups showing significant differences in prior technological knowledge. This finding provides empirical evidence for the near-universality of image-schematic metaphor use across age groups. The study also identifies 37 image-schematic metaphors shared between the two age groups in the technology learning domain to support future design of age-inclusive interactive systems.
Nathan Crilly, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.3
2024 Hold Tight: Identifying Behavioral Patterns During Prolonged Work in VR Through Video Analysis
abstract
VR devices have recently been actively promoted as tools for knowledge workers and prior work has demonstrated that VR can support some knowledge worker tasks. However, only a few studies have explored the effects of prolonged use of VR such as a study observing 16 participants working in VR and a physical environment for one work-week each and reporting mainly on subjective feedback. As a nuanced understanding of participants' behavior in VR and how it evolves over time is still missing, we report on the results from an analysis of 559 hours of video material obtained in this prior study. Among other findings, we report that (1) the frequency of actions related to adjusting the headset reduced by 46% and the frequency of actions related to supporting the headset reduced by 42% over the five days; (2) the HMD was removed 31% less frequently over the five days but for 41% longer periods; (3) wearing an HMD is disruptive to normal patterns of eating and drinking, but not to social interactions, such as talking. The combined findings in this work demonstrate the value of long-term studies of deployed VR systems and can be used to inform the design of better, more ergonomic VR systems as tools for knowledge workers.
Verena Biener, Forouzan Farzinnejad, Rinaldo Schuster, Seyedmasih Tabaei, Leon Lindlein, Jinghui Hu, Negar Nouri, John J. Dudley, Per Ola Kristensson, Jörg Müller 0001, Jens Grubert
IEEE Trans. Vis. Comput. Graph.9
2024 Text Entry Performance and Situation Awareness of a Joint Optical See-Through Head-Mounted Display and Smartphone System
abstract
Optical see-through head-mounted displays (OST HMDs) are a popular output medium for mobile Augmented Reality (AR) applications. To date, they lack efficient text entry techniques. Smartphones are a major text entry medium in mobile contexts but attentional demands can contribute to accidents while typing on the go. Mobile multi-display ecologies, such as combined OST HMD-smartphone systems, promise performance and situation awareness benefits over single-device use. We study the joint performance of text entry on mobile phones with text output on optical see-through head-mounted displays. A series of five experiments with a total of 86 participants indicate that, as of today, the challenges in such a joint interactive system outweigh the potential benefits.
Jens Grubert, Lukas Witzani, Alexander Otte, Travis Gesslein, Matthias Kranz, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.6
2024 Design and Evaluation of Controller-Based Raycasting Methods for Efficient Alphanumeric and Special Character Entry in Virtual Reality
abstract
Alphanumeric and special characters are essential during text entry. Text entry in virtual reality (VR) is usually performed on a virtual Qwerty keyboard to minimize the need to learn new layouts. As such, entering capitals, symbols, and numbers in VR is often a direct migration from a physical/touchscreen Qwerty keyboard-that is, using the mode-switching keys to switch between different types of characters and symbols. However, there are inherent differences between a keyboard in VR and a physical/touchscreen keyboard, and as such, a direct adaptation of mode-switching via switch keys may not be suitable for VR. The high flexibility afforded by VR opens up more possibilities for entering alphanumeric and special characters using the Qwerty layout. In this work, we designed two controller-based raycasting text entry methods for alphanumeric and special characters input (Layer-ButtonSwitch and Key-ButtonSwitch) and compared them with two other methods (Standard Qwerty Keyboard and Layer-PointSwitch) that were derived from physical and soft Qwerty keyboards. We explored the performance and user preference of these four methods via two user studies (one short-term and one prolonged use), where participants were instructed to input text containing alphanumeric and special characters. Our results show that Layer-ButtonSwitch led to the highest statistically significant performance, followed by Key-ButtonSwitch and Standard Qwerty Keyboard, while Layer-PointSwitch had the slowest speed. With continuous practice, participants' performance using Key-ButtonSwitch reached that of Layer-ButtonSwitch. Further, the results show that the key-level layout used in Key-ButtonSwitch led users to parallel mode switching and character input operations because this layout showed all characters on one layer. We distill three recommendations from the results that can help guide the design of text entry techniques for alphanumeric and special characters in VR.
Tingjie Wan, Yushi Wei, Rongkai Shi, Junxiao Shen, Per Ola Kristensson, Katie Atkinson, Hai-Ning Liang
IEEE Trans. Vis. Comput. Graph.5
2024 Should I Evaluate My Augmented Reality System in an Industrial Environment? Investigating the Effects of Classroom and Shop Floor Settings on Guided Assembly
abstract
Numerous prior studies have investigated real-time assembly instructions using Augmented Reality (AR). However, most such experiments were conducted in laboratory settings with simplistic assembly tasks, failing to represent real-world industrial conditions. To ascertain to what extent results obtained in a laboratory environment may differ from studies in actual industrial environments, we carried out a user study with 32 manufacturing apprentices. We compared assembly task execution results in two settings, a classroom and an industrial workshop environment. To facilitate the experiments, we developed AR-guided manual assembly systems for simple and more complex assets. Our findings reveal a significantly improved task performance in the industrial workshop, reflected in faster task completion times, fewer errors, and subjectively perceived higher flow. This contradicted participants' subjective ratings, as they expected to perform better in the classroom environment. Our results suggest that the actual manufacturing environment is critical in evaluating AR systems for real-world industrial applications.
Vicky Zhang, Alexander Albers, Christine Saeedi-Givi, Per Ola Kristensson, Thomas Bohné, Slawomir Konrad Tadeja
IEEE Trans. Vis. Comput. Graph.4
2023 Relative Design Acquisition: A Computational Approach for Creating Visual Interfaces to Steer User Choices
abstract
A central objective in computational design is that an optimal design is desired which optimizes a performance metric. We explore a different problem class with a computational approach we call relative design acquisition. As a motivational example, consider a user prompted to make a choice using buttons. One button may have a more visually appealing design and hence is visually optimal to steer users to click it more often than the second button. In such a design case, a relative design is acquired of a certain quality with respect to a reference design to guide a user decision. After mathematically formalizing this problem, we report the results of three experiments that demonstrate the approach’s efficacy in generating relative designs in a visual interface preference setting. The relative designs are controllable by a quality factor, which affects both comparative ratings and human decision time between the reference and relative designs.
George B. Mo, Per Ola Kristensson
CHI2
2023 Understanding Adoption Barriers to Dwell-Free Eye-Typing: Design Implications from a Qualitative Deployment Study and Computational Simulations
abstract
Eye-typing is a slow and cumbersome text entry method typically used by individuals with no other practical means of communication. As an alternative, prior HCI research has proposed dwell-free eye-typing as a potential improvement that eliminates time-consuming and distracting dwell-timeouts. However, it is rare that such research ideas are translated into working products. This paper reports on a qualitative deployment study of a product that was developed to allow users access to a dwell-free eye-typing research solution. This allowed us to understand how such a research solution would work in practice, as part of users’ current communication solutions in their own homes. Based on interviews and observations, we discuss a number of design issues that currently act as barriers preventing widespread adoption of dwell-free eye-typing. The study findings are complemented with computational simulations in a range of conditions that were inspired by the findings in the deployment study. These simulations serve to both contextualize the qualitative findings and to explore quantitative implications of possible interface redesigns. The combined analysis gives rise to a set of design implications for enabling wider adoption of dwell-free eye-typing in practice.
Per Ola Kristensson, Morten Mjelde, Keith Vertanen
IUI1
2023 Imperfect Surrogate Users: Understanding Performance Implications of Augmentative and Alternative Communication Systems through Bounded Rationality, Human Error, and Interruption Modeling
abstract
Nonspeaking individuals with motor disabilities frequently rely on augmentative and alternative communication (AAC) systems that allow users to communicate through a text entry interface coupled with a speech synthesizer. Such systems are notoriously difficult to evaluate with end-users. However, recent research has proposed envelope analysis as a method to estimate text entry rates and keystroke savings by simulating the interaction of an expert surrogate user entering sentences on a conceptual word-predictive text entry system. While only a part of the evaluation process of an AAC system, this method enables AAC designers to benefit from quantitative insights early on in the design process. This paper extends prior work by (1) demonstrating how to incorporate natural language generation, such as sentence generation, in such analyses; (2) presenting a model of an imperfect surrogate user that incorporates bounded rationality, human error, and interruptions to provide a more realistic simulation of text entry behavior; and (3) demonstrating how to estimate model parameters by observing users' actual typing behavior. We validate the model with data collected from eight participants using an AAC system on a touchscreen.
Boyin Yang, Per Ola Kristensson
Proc. ACM Hum. Comput. Interact.2
2023 Evaluating the Performance of Hand-Based Probabilistic Text Input Methods on a Mid-Air Virtual Qwerty Keyboard
abstract
Integrated hand-tracking on modern virtual reality (VR) headsets can be readily exploited to deliver mid-air virtual input surfaces for text entry. These virtual input surfaces can closely replicate the experience of typing on a Qwerty keyboard on a physical touchscreen, thereby allowing users to leverage their pre-existing typing skills. However, the lack of passive haptic feedback, unconstrained user motion, and potential tracking inaccuracies or observability issues encountered in this interaction setting typically degrades the accuracy of user articulations. We present a comprehensive exploration of error-tolerant probabilistic hand-based input methods to support effective text input on a mid-air virtual Qwerty keyboard. Over three user studies we examine the performance potential of hand-based text input under both gesture and touch typing paradigms. We demonstrate typical entry rates in the range of 20 to 30 wpm and average peak entry rates of 40 to 45 wpm.
John J. Dudley, Jingyao Zheng, Aakar Gupta, Hrvoje Benko, Matt Longest, Robert Wang 0002, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.7
2023 Fast and Robust Mid-Air Gesture Typing for AR Headsets using 3D Trajectory Decoding
abstract
We present a fast mid-air gesture keyboard for head-mounted optical see-through augmented reality (OST AR) that supports users in articulating word patterns by merely moving their own physical index finger in relation to a virtual keyboard plane without a need to indirectly control a visual 2D cursor on a keyboard plane. To realize this, we introduce a novel decoding method that directly translates users' three-dimensional fingertip gestural trajectories into their intended text. We evaluate the efficacy of the system in three studies that investigate various design aspects, such as immediate efficacy, accelerated learning, and whether it is possible to maintain performance without providing visual feedback. We find that the new 3D trajectory decoding design results in significant improvements in entry rates while maintaining low error rates. In addition, we demonstrate that users can maintain their performance even without fingertip and gesture trace visualization.
Junxiao Shen, John J. Dudley, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.3
2023 HotGestures: Complementing Command Selection and Use with Delimiter-Free Gesture-Based Shortcuts in Virtual Reality
abstract
Conventional desktop applications provide users with hotkeys as shortcuts for triggering different functionality. In this paper we consider what constitutes an effective parallel to hotkeys in a 3D interaction space where the input modality is no longer limited to the use of a keyboard. We propose HotGestures: a gesture-based interaction system for rapid tool selection and usage. Hand gestures are frequently used during human communication to convey information and provide natural associations with meaning. HotGestures provide shortcuts for users to seamlessly activate and use virtual tools by performing hand gestures. This approach naturally complements conventional menu interactions. We evaluate the potential of HotGestures in a set of two user studies and observe that our gesture-based technique provides fast and effective shortcuts for tool selection and usage. Participants found HotGestures to be distinctive, fast, and easy to use while also complementing conventional menu-based interaction.
Zhaomou Song, John J. Dudley, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.3
2022 Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
abstract
Designers reportedly struggle with design optimization tasks where they are asked to find a combination of design parameters that maximizes a given set of objectives. In HCI, design optimization problems are often exceedingly complex, involving multiple objectives and expensive empirical evaluations. Model-based computational design algorithms assist designers by generating design examples during design, however they assume a model of the interaction domain. Black box methods for assistance, on the other hand, can work with any design problem. However, virtually all empirical studies of this human-in-the-loop approach have been carried out by either researchers or end-users. The question stands out if such methods can help designers in realistic tasks. In this paper, we study Bayesian optimization as an algorithmic method to guide the design optimization process. It operates by proposing to a designer which design candidate to try next, given previous observations. We report observations from a comparative study with 40 novice designers who were tasked to optimize a complex 3D touch interaction technique. The optimizer helped designers explore larger proportions of the design space and arrive at a better solution, however they reported lower agency and expressiveness. Designers guided by an optimizer reported lower mental effort but also felt less creative and less in charge of the progress. We conclude that human-in-the-loop optimization can support novice designers in cases where agency is not critical.
Li-Wei Chan 0001, Yi-Chi Liao 0001, George B. Mo, John J. Dudley, Chun-Lien Cheng, Per Ola Kristensson, Antti Oulasvirta
CHI6
2022 Personalization of a Mid-Air Gesture Keyboard using Multi-Objective Bayesian Optimization
abstract
We present AdaptiKeyboard, a mid-air gesture keyboard that uses multi-objective Bayesian optimization to adaptively change layout size to simultaneously optimize speed and accuracy. Gesture keyboards are well suited for enabling mid-air text entry in augmented reality (AR) due to their relative robustness to articulation inaccuracy. However, transplanting gesture keyboards to AR involves a larger design and operational space compared to touchscreen interactions. One potential advantage of this larger design and operational space is that mid-air keyboards presented in AR can be more versatile than their touchscreen equivalents. A key component of a mid-air gesture keyboard is the layout size, which can be made adaptive in order to optimize text entry speed and accuracy at the individual user level. This adaptive personalization can refine the keyboard design to reflect the differences users exhibit in motor behaviors and personal preferences. In this paper, we propose a multi-objective Bayesian optimization approach for adapting the layout size of a mid-air gesture keyboard to individual users. We show that this process can deliver a 14.4% improvement in speed and a 13.8% improvement in accuracy relative to a baseline design with a constant size derived from the default system keyboard on the HoloLens 2.
Junxiao Shen, Jinghui Hu, John J. Dudley, Per Ola Kristensson
ISMAR4
2022 Efficient Special Character Entry on a Virtual Keyboard by Hand Gesture-Based Mode Switching
abstract
The need to support efficient text input in virtual reality continues to attract significant research attention. However, much of this research understandably focuses exclusively on core text input tasks involving the entry of standard alphabetic characters. The less common though still critical task of entering special characters is often ignored. In this paper we focus on this niche use case as chiefly encountered when entering passwords. Current commercial virtual keyboards allow users to switch between different layers of the keyboard in order to access capital letters, numerals and special characters by pressing an explicit mode-switch button. We propose a new method of switching between layers of a virtual keyboard using hand gestures. Critically, these hand gestures are seamlessly performed in conjunction with key selections to deliver an efficient and intuitive interaction. We report on a user study with 16 participants entering standard passwords comparing our gesture-based mode switching approach to a conventional button-based baseline. We find that with practice our proposed method results in significantly faster entry rates without any deterioration in accuracy. Feedback from users also indicates that our technique is considered efficient and comfortable to use.
Zhaomou Song, John J. Dudley, Per Ola Kristensson
ISMAR3
2022 KWickChat: A Multi-Turn Dialogue System for AAC Using Context-Aware Sentence Generation by Bag-of-Keywords
abstract
We present KWickChat (Keyword Quick Chat): a multi-turn augmentative and alternative communication (AAC) dialogue system for nonspeaking individuals with motor disabilities. The central objective of KWickChat is to reduce the communication gap between nonspeaking and speaking partners by exploring a sentence-based text entry system that automatically generates suitable sentences for the nonspeaking partner based on keyword entry. The system is underpinned by a GPT-2 language model and leverages context information, including dialogue history and persona tags, to improve the quality of the generated responses. We evaluate the system by analyzing the functional design and decomposing it into key functions and parameters that are systematically investigated using envelope analysis. We pursue this methodology as a necessary precursor to evaluation with AAC users. Our results show that with word prediction and with a threshold word error rate of 0.65, the keystroke savings of the KWickChat system is around 71%. To complement the envelope analysis, we also recruited two human judges to evaluate the semantic consistency between 400 sentences generated by KWickChat and reference sentences. Both judges reported a median rating of 4 on a scale from 1 (very bad) to 5 (very good) for the best generated sentence in each exchange and achieved an inter-rater reliability of 0.92 across all 400 sentences judged.
Junxiao Shen, Boyin Yang, John J. Dudley, Per Ola Kristensson
IUI4
2022 Supporting Playful Rehabilitation in the Home using Virtual Reality Headsets and Force Feedback Gloves
abstract
Virtual Reality (VR) is a promising platform for home rehabilitation with the potential to completely immerse users within a playful experience. To explore this area we design, implement, and evaluate a system that uses a VR headset in conjunction with force feed-back gloves to present users with a playful experience for home rehabilitation. The system immerses the user within a virtual cat bathing simulation that allows users to practice fine motor skills by progressively completing three cat-care tasks. The study results demonstrate the positive role that playfulness may play in the user experience of VR rehabilitation.
Qisong Wang, Bo Kang, Per Ola Kristensson
VR3
2022 Understanding user performance of acquiring targets with motion-in-depth in virtual reality
Jin Huang 0009, John J. Dudley, Stephen Uzor, Per Ola Kristensson, Feng Tian 0001
Int. J. Hum. Comput. Stud.5
2022 PoVRPoint: Authoring Presentations in Mobile Virtual Reality
abstract
Virtual Reality (VR) has the potential to support mobile knowledge workers by complementing traditional input devices with a large three-dimensional output space and spatial input. Previous research on supporting VR knowledge work explored domains such as text entry using physical keyboards and spreadsheet interaction using combined pen and touch input. Inspired by such work, this paper probes the VR design space for authoring presentations in mobile settings. We propose PoVRPoint-a set of tools coupling pen- and touch-based editing of presentations on mobile devices, such as tablets, with the interaction capabilities afforded by VR. We study the utility of extended display space to, for example, assist users in identifying target slides, supporting spatial manipulation of objects on a slide, creating animations, and facilitating arrangements of multiple, possibly occluded shapes or objects. Among other things, our results indicate that 1) the wide field of view afforded by VR results in significantly faster target slide identification times compared to a tablet-only interface for visually salient targets; and 2) the three-dimensional view in VR enables significantly faster object reordering in the presence of occlusion compared to two baseline interfaces. A user study further confirmed that the interaction techniques were found to be usable and enjoyable.
Verena Biener, Travis Gesslein, Daniel Schneider 0006, Felix Kawala, Alexander Otte, Per Ola Kristensson, Michel Pahud, Eyal Ofek, Cuauhtli Campos, Matjaz Kljun, Klen Copic Pucihar, Jens Grubert
IEEE Trans. Vis. Comput. Graph.6
2022 Quantifying the Effects of Working in VR for One Week
abstract
Virtual Reality (VR) provides new possibilities for modern knowledge work. However, the potential advantages of virtual work environments can only be used if it is feasible to work in them for an extended period of time. Until now, there are limited studies of long-term effects when working in VR. This paper addresses the need for understanding such long-term effects. Specifically, we report on a comparative study $i$, in which participants were working in VR for an entire week-for five days, eight hours each day-as well as in a baseline physical desktop environment. This study aims to quantify the effects of exchanging a desktop-based work environment with a VR-based environment. Hence, during this study, we do not present the participants with the best possible VR system but rather a setup delivering a comparable experience to working in the physical desktop environment. The study reveals that, as expected, VR results in significantly worse ratings across most measures. Among other results, we found concerning levels of simulator sickness, below average usability ratings and two participants dropped out on the first day using VR, due to migraine, nausea and anxiety. Nevertheless, there is some indication that participants gradually overcame negative first impressions and initial discomfort. Overall, this study helps lay the groundwork for subsequent research, by clearly highlighting current shortcomings and identifying opportunities for improving the experience of working in VR.
Verena Biener, Snehanjali Kalamkar, Negar Nouri, Eyal Ofek, Michel Pahud, John J. Dudley, Jinghui Hu, Per Ola Kristensson, Maheshya Weerasinghe, Klen Copic Pucihar, Matjaz Kljun, Stephan Streuber, Jens Grubert
IEEE Trans. Vis. Comput. Graph.8
2022 Gesture Spotter: A Rapid Prototyping Tool for Key Gesture Spotting in Virtual and Augmented Reality Applications
abstract
In this paper we examine the task of key gesture spotting: accurate and timely online recognition of hand gestures. We specifically seek to address two key challenges faced by developers when integrating key gesture spotting functionality into their applications. These are: i) achieving high accuracy and zero or negative activation lag with single-time activation; and ii) avoiding the requirement for deep domain expertise in machine learning. We address the first challenge by proposing a key gesture spotting architecture consisting of a novel gesture classifier model and a novel single-time activation algorithm. This key gesture spotting architecture was evaluated on four separate hand skeleton gesture datasets, and achieved high recognition accuracy with early detection. We address the second challenge by encapsulating different data processing and augmentation strategies, as well as the proposed key gesture spotting architecture, into a graphical user interface and an application programming interface. Two user studies demonstrate that developers are able to efficiently construct custom recognizers using both the graphical user interface and the application programming interface.
Junxiao Shen, John J. Dudley, George B. Mo, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.4
2021 Crowdsourcing Design Guidance for Contextual Adaptation of Text Content in Augmented Reality
abstract
Augmented Reality (AR) can deliver engaging user experiences that seamlessly meld virtual content with the physical environment. However, building such experiences is challenging due to the developer’s inability to assess how uncontrolled deployment contexts may influence the user experience. To address this issue, we demonstrate a method for rapidly conducting AR experiments and real-world data collection in the user’s own physical environment using a privacy-conscious mobile web application. The approach leverages the large number of distinct user contexts accessible through crowdsourcing to efficiently source diverse context and perceptual preference data. The insights gathered through this method complement emerging design guidance and sample-limited lab-based studies. The utility of the method is illustrated by re-examining the design challenge of adapting AR text content to the user’s environment. Finally, we demonstrate how gathered design insight can be operationalized to provide adaptive text content functionality in an AR headset.
John J. Dudley, Jason T. Jacques, Per Ola Kristensson
CHI3
2021 Understanding, Detecting and Mitigating the Effects of Coactivations in Ten-Finger Mid-Air Typing in Virtual Reality
abstract
Typing with ten fingers on a virtual keyboard in virtual or augmented reality exposes a challenging input interpretation problem. There are many sources of noise in this interaction context and these exacerbate the challenge of accurately translating human actions into text. A particularly challenging input noise source arises from the physiology of the hand. Intentional finger movements can produce unintentional coactivations in other fingers. On a physical keyboard, the resistance of the keys alleviates this issue. On a virtual keyboard, coactivations are likely to introduce spurious input events under a naïve solution to input detection. In this paper we examine the features that discriminate intentional activations from coactivations. Based on this analysis, we demonstrate three alternative coactivation detection strategies with high discrimination power. Finally, we integrate coactivation detection into a probabilistic decoder and demonstrate its ability to further reduce uncorrected character error rates by approximately 10% relative and 0.9% absolute.
Conor R. Foy, John J. Dudley, Aakar Gupta, Hrvoje Benko, Per Ola Kristensson
CHI5
2021 Enhancing the Composition Task in Text Entry Studies: Eliciting Difficult Text and Improving Error Rate Calculation
abstract
Participants in text entry studies usually copy phrases or compose novel messages. A composition task mimics actual user behavior and can allow researchers to better understand how a system might perform in reality. A problem with composition is that participants may gravitate towards writing simple text, that is, text containing only common words. Such simple text is insufficient to explore all factors governing a text entry method, such as its error correction features. We contribute to enhancing composition tasks in two ways. First, we show participants can modulate the difficulty of their compositions based on simple instructions. While it took more time to compose difficult messages, they were longer, had more difficult words, and resulted in more use of error correction features. Second, we compare two methods for obtaining a participant’s intended text, comparing both methods with a previously proposed crowdsourced judging procedure. We found participant-supplied references were more accurate.
Dylan Gaines, Per Ola Kristensson, Keith Vertanen
CHI2
2021 Design and Analysis of Intelligent Text Entry Systems with Function Structure Models and Envelope Analysis
abstract
Designing intelligent interactive text entry systems often relies on factors that are difficult to estimate or assess using traditional HCI design and evaluation methods. We introduce a complementary approach by adapting function structure models from engineering design. We extend their use by extracting controllable and uncontrollable parameters from function structure models and visualizing their impact using envelope analysis. Function structure models allow designers to understand a system in terms of its functions and flows between functions and decouple functions from function carriers. Envelope analysis allows the designer to further study how parameters affect variables of interest, for example, accuracy, keystroke savings and other dependent variables. We provide examples of function structure models and illustrate a complete envelope analysis by investigating a parameterized function structure model of predictive text entry. We discuss the implications of this design approach for both text entry system design and for critique of system contributions.
Per Ola Kristensson, Thomas Müllners
CHI1
2021 Gesture Knitter: A Hand Gesture Design Tool for Head-Mounted Mixed Reality Applications
abstract
Hand gestures are a natural and expressive input method enabled by modern mixed reality headsets. However, it remains challenging for developers to create custom gestures for their applications. Conventional strategies to bespoke gesture recognition involve either hand-crafting or data-intensive deep-learning. Neither approach is well suited for rapid prototyping of new interactions. This paper introduces a flexible and efficient alternative approach for constructing hand gestures. We present Gesture Knitter: a design tool for creating custom gesture recognizers with minimal training data. Gesture Knitter allows the specification of gesture primitives that can then be combined to create more complex gestures using a visual declarative script. Designers can build custom recognizers by declaring them from scratch or by providing a demonstration that is automatically decoded into its primitive components. Our developer study shows that Gesture Knitter achieves high recognition accuracy despite minimal training data and delivers an expressive and creative design experience.
George B. Mo, John J. Dudley, Per Ola Kristensson
CHI3
2021 Investigating the Accessibility of Crowdwork Tasks on Mechanical Turk
abstract
Crowdwork can enable invaluable opportunities for people with disabilities, not least the work flexibility and the ability to work from home, especially during the current Covid-19 pandemic. This paper investigates how engagement in crowdwork tasks is affected by individual disabilities and the resulting implications for HCI. We first surveyed 1,000 Amazon Mechanical Turk (AMT) workers to identify demographics of crowdworkers who identify as having various disabilities within the AMT ecosystem—including vision, hearing, cognition/mental, mobility, reading and motor impairments. Through a second focused survey and follow-up interviews, we provide insights into how respondents cope with crowdwork tasks. We found that standard task factors, such as task completion time and presentation, often do not account for the needs of users with disabilities, resulting in anxiety and a feeling of depression on occasion. We discuss how to alleviate barriers to enable effective interaction for crowdworkers with disabilities.
Stephen Uzor, Jason T. Jacques, John J. Dudley, Per Ola Kristensson
CHI4
2021 The Imaginative Generative Adversarial Network: Automatic Data Augmentation for Dynamic Skeleton-Based Hand Gesture and Human Action Recognition
abstract
Deep learning approaches deliver state-of-the-art performance in recognition of spatiotemporal human motion data. However, one of the main challenges in these recognition tasks is limited available training data. Insufficient training data results in over-fitting and data augmentation is one approach to address this challenge. Existing data augmentation strategies based on scaling, shifting and interpolating offer limited generalizability and typically require detailed inspection of the dataset as well as hundreds of GPU hours for hyperparameter optimization. In this paper, we present a novel automatic data augmentation model, the Imaginative Generative Adversarial Network (GAN), that approximates the distribution of the input data and samples new data from this distribution. It is automatic in that it requires no data inspection and little hyperparameter tuning and therefore it is a low-cost and low-effort approach to generate synthetic data. We demonstrate our approach on small-scale skeleton-based datasets with a comprehensive experimental analysis. Our results show that the augmentation strategy is fast to train and can improve classification accuracy for both conventional neural networks and state-of-the-art methods.
Junxiao Shen, John J. Dudley, Per Ola Kristensson
FG3
2021 Simulating Realistic Human Motion Trajectories of Mid-Air Gesture Typing
abstract
The eventual success of many AR and VR intelligent interactive systems relies on the ability to collect user motion data at large scale. Realistic simulation of human motion trajectories is a potential solution to this problem. Simulated user motion data can facilitate prototyping and speed up the design process. There are also potential benefits in augmenting training data for deep learning-based AR/VR applications to improve performance. However, the generation of realistic motion data is nontrivial. In this paper, we examine the specific challenge of simulating index finger movement data to inform mid-air gesture keyboard design. The mid-air gesture keyboard is deployed on an optical see-through display that allows the user to enter text by articulating word gesture patterns with their physical index finger in the vicinity of a visualized keyboard layout. We propose and compare four different approaches to simulating this type of motion data, including a Jerk-Minimization model, a Recurrent Neural Network (RNN)-based generative model, and a Generative Adversarial Network (GAN)-based model with two modes: style transfer and data alteration. We also introduce a procedure for validating the quality of the generated trajectories in terms of realism and diversity. The GAN-based model shows significant potential for generating synthetic motion trajectories to facilitate design and deep learning for advanced gesture keyboards deployed in AR and VR.
Junxiao Shen, John J. Dudley, Per Ola Kristensson
ISMAR3
2021 Supporting Iterative Virtual Reality Analytics Design and Evaluation by Systematic Generation of Surrogate Clustered Datasets
abstract
Virtual Reality (VR) is a promising technology platform for immersive visual analytics. However, the design space of VR analytics interface design is vast and difficult to explore using traditional A/B comparisons in formal or informal controlled experiments— a fundamental part of an iterative design process. A key factor that complicates such comparisons is the dataset. Exposing participants to the same dataset in all conditions introduces an unavoidable learning effect. On the other hand, using different datasets for all experimental conditions introduces the dataset itself as an uncontrolled variable, which reduces internal validity to an unacceptable degree. In this paper, we propose to rectify this problem by introducing a generative process for synthesizing clustered datasets for VR analytics experiments. This process generates datasets that are distinct while simultaneously allowing systematic comparisons in experiments. A key advantage is that these datasets can then be used in iterative design processes. In a two-part experiment, we show the validity of the generative process and demonstrate how new insights in VR-based visual analytics can be gained using synthetic datasets.
Slawomir Konrad Tadeja, Patrick Langdon, Per Ola Kristensson
ISMAR3
2021 Exploring gestural input for engineering surveys of real-life structures in virtual reality using photogrammetric 3D models
abstract
Abstract Photogrammetry is a promising set of methods for generating photorealistic 3D models of physical objects and structures. Such methods may rely solely on camera-captured photographs or include additional sensor data. Digital twins are digital replicas of physical objects and structures. Photogrammetry is an opportune approach for generating 3D models for the purpose of preparing digital twins. At a sufficiently high level of quality, digital twins provide effective archival representations of physical objects and structures and become effective substitutes for engineering inspections and surveying. While photogrammetric techniques are well-established, insights about effective methods for interacting with such models in virtual reality remain underexplored. We report the results of a qualitative engineering case study in which we asked six domain experts to carry out engineering measurement tasks in an immersive environment using bimanual gestural input coupled with gaze-tracking. The qualitative case study revealed that gaze-supported bimanual interaction of photogrammetric 3D models is a promising modality for domain experts. It allows the experts to efficiently manipulate and measure elements of the 3D model. To better allow designers to support this modality, we report design implications distilled from the feedback from the domain experts.
Slawomir Konrad Tadeja, Yupu Lu, Maciej Rydlewicz, Wojciech Rydlewicz, Tomasz Bubas, Per Ola Kristensson
Multim. Tools Appl.6
2021 Mining, analyzing, and modeling text written on mobile devices
abstract
Abstract We present a method for mining the web for text entered on mobile devices. Using searching, crawling, and parsing techniques, we locate text that can be reliably identified as originating from 300 mobile devices. This includes 341,000 sentences written on iPhones alone. Our data enables a richer understanding of how users type “in the wild” on their mobile devices. We compare text and error characteristics of different device types, such as touchscreen phones, phones with physical keyboards, and tablet computers. Using our mined data, we train language models and evaluate these models on mobile test data. A mixture model trained on our mined data, Twitter, blog, and forum data predicts mobile text better than baseline models. Using phone and smartwatch typing data from 135 users, we demonstrate our models improve the recognition accuracy and word predictions of a state-of-the-art touchscreen virtual keyboard decoder. Finally, we make our language models and mined dataset available to other researchers.
Keith Vertanen, Per Ola Kristensson
Nat. Lang. Eng.2
2021 An Exploration of Freehand Crossing Selection in Head-Mounted Augmented Reality
abstract
Crossing, or goal crossing, has proven useful in various selection scenarios, including pen, mouse, touch, and virtual reality (VR). However, crossing has not been exploited for freehand selection using augmented reality head-mounted displays (AR HMDs). Using the HoloLens, we explore freehand crossing for selection and compare it to the state-of-the-art “gaze and commit” (head gaze) method. We report on three studies investigating freehand crossing in multiple use cases. The first study shows that crossing outperforms head gaze in selection scenarios of varying target arrangements. The second explores crossing, head gaze, and hand pointing in radial menu and dynamic interface scenarios. The third explores crossing as a function carrier for a variety of basic interaction functions in a drawing application. This work builds on existing knowledge on the goal-crossing paradigm by demonstrating its potential as a useful interaction method in 3D AR HMD interfaces.
Stephen Uzor, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.2
2021 Complex Interaction as Emergent Behaviour: Simulating Mid-Air Virtual Keyboard Typing using Reinforcement Learning
abstract
Accurately modelling user behaviour has the potential to significantly improve the quality of human-computer interaction. Traditionally, these models are carefully hand-crafted to approximate specific aspects of well-documented user behaviour. This limits their availability in virtual and augmented reality where user behaviour is often not yet well understood. Recent efforts have demonstrated that reinforcement learning can approximate human behaviour during simple goal-oriented reaching tasks. We build on these efforts and demonstrate that reinforcement learning can also approximate user behaviour in a complex mid-air interaction task: typing on a virtual keyboard. We present the first reinforcement learning-based user model for mid-air and surface-aligned typing on a virtual keyboard. Our model is shown to replicate high-level human typing behaviour. We demonstrate that this approach may be used to augment or replace human testing during the validation and development of virtual keyboards.
Lorenz Hetzel, John J. Dudley, Anna Maria Feit, Per Ola Kristensson
IEEE Trans. Vis. Comput. Graph.4
2020 A Design Engineering Approach for Quantitatively Exploring Context-Aware Sentence Retrieval for Nonspeaking Individuals with Motor Disabilities
abstract
Nonspeaking individuals with motor disabilities typically have very low communication rates. This paper proposes a design engineering approach for quantitatively exploring context-aware sentence retrieval as a promising complementary input interface, working in tandem with a word-prediction keyboard. We motivate the need for complementary design engineering methodology in the design of augmentative and alternative communication and explain how such methods can be used to gain additional design insights. We then study the theoretical performance envelopes of a context-aware sentence retrieval system, identifying potential keystroke savings as a function of the parameters of the subsystems, such as the accuracy of the underlying auto-complete word prediction algorithm and the accuracy of sensed context information under varying assumptions. We find that context-aware sentence retrieval has the potential to provide users with considerable improvements in keystroke savings under reasonable parameter assumptions of the underlying subsystems. This highlights how complementary design engineering methods can reveal additional insights into design for augmentative and alternative communication.
Per Ola Kristensson, James Lilley, Rolf Black, Annalu Waller
CHI1
2020 Pen-based Interaction with Spreadsheets in Mobile Virtual Reality
abstract
Virtual Reality (VR) can enhance the display and interaction of mobile knowledge work and in particular, spreadsheet applications. While spreadsheets are widely used yet are challenging to interact with, especially on mobile devices, using them in VR has not been explored in depth. A special uniqueness of the domain is the contrast between the immersive and large display space afforded by VR, contrasted by the very limited interaction space that may be afforded for the information worker on the go, such as an airplane seat or a small work-space. To close this gap, we present a tool-set for enhancing spreadsheet interaction on tablets using immersive VR headsets and pen-based input. This combination opens up many possibilities for enhancing the productivity for spreadsheet interaction. We propose to use the space around and in front of the tablet for enhanced visualization of spreadsheet data and meta-data. For example, extending sheet display beyond the bounds of the physical screen, or easier debugging by uncovering hidden dependencies between sheet's cells. Combining the precise on-screen input of a pen with spatial sensing around the tablet, we propose tools for the efficient creation and editing of spreadsheets functions such as off-the-screen layered menus, visualization of sheets dependencies, and gaze-and-touch-based switching between spreadsheet tabs. We study the feasibility of the proposed tool-set using a video-based online survey and an expert-based assessment of indicative human performance potential.
Travis Gesslein, Verena Biener, Philipp Gagel, Daniel Schneider 0006, Per Ola Kristensson, Eyal Ofek, Michel Pahud, Jens Grubert
ISMAR5
2020 Investigating Remote Tactile Feedback for Mid-Air Text-Entry in Virtual Reality
abstract
In this paper, we investigate the utility of remote tactile feedback for freehand text-entry on a mid-air Qwerty keyboard in VR. To that end, we use insights from prior work to design a virtual keyboard along with different forms of tactile feedback, both spatial and non-spatial, for fingers and for wrists. We report on a multi-session text-entry study with 24 participants where we investigated four vibrotactile feedback conditions: on-fingers, on-wrist spatialized, on-wrist non-spatialized, and audio-visual only. We use micro-metrics analyses and participant interviews to analyze the mechanisms underpinning the observed performance and user experience. The results show comparable performance across feedback types. However, participants overwhelmingly prefer the tactile feedback conditions and rate on-fingers feedback as significantly lower in mental demand, frustration, and effort. Results also show that spatialization of vibrotactile feedback on the wrist as a way to provide finger-specific feedback is comparable in performance and preference to a single vibration location. The micro-metrics analyses suggest that users compensated for the lack of tactile feedback with higher visual and cognitive attention, which ensured similar performance but higher user effort.
Aakar Gupta, Majed Samad, Kenrick Kin, Per Ola Kristensson, Hrvoje Benko
ISMAR4
2020 Breaking the Screen: Interaction Across Touchscreen Boundaries in Virtual Reality for Mobile Knowledge Workers
abstract
Virtual Reality (VR) has the potential to transform knowledge work. One advantage of VR knowledge work is that it allows extending 2D displays into the third dimension, enabling new operations, such as selecting overlapping objects or displaying additional layers of information. On the other hand, mobile knowledge workers often work on established mobile devices, such as tablets, limiting interaction with those devices to a small input space. This challenge of a constrained input space is intensified in situations when VR knowledge work is situated in cramped environments, such as airplanes and touchdown spaces. In this paper, we investigate the feasibility of interacting jointly between an immersive VR head-mounted display and a tablet within the context of knowledge work. Specifically, we 1) design, implement and study how to interact with information that reaches beyond a single physical touchscreen in VR; 2) design and evaluate a set of interaction concepts; and 3) build example applications and gather user feedback on those applications.
Verena Biener, Daniel Schneider 0006, Travis Gesslein, Alexander Otte, Bastian Kuth, Per Ola Kristensson, Eyal Ofek, Michel Pahud, Jens Grubert
IEEE Trans. Vis. Comput. Graph.6
2019 Crowdsourcing Interface Feature Design with Bayesian Optimization
abstract
Designing novel interfaces is challenging. Designers typically rely on experience or subjective judgment in the absence of analytical or objective means for selecting interface parameters. We demonstrate Bayesian optimization as an efficient tool for objective interface feature refinement. Specifically, we show that crowdsourcing paired with Bayesian optimization can rapidly and effectively assist interface design across diverse deployment environments. Experiment 1 evaluates the approach on a familiar 2D interface design problem: a map search and review use case. Adding a degree of complexity, Experiment 2 extends Experiment 1 by switching the deployment environment to mobile-based virtual reality. The approach is then demonstrated as a case study for a fundamentally new and unfamiliar interaction design problem: web-based augmented reality. Finally, we show how the model generated as an outcome of the refinement process can be used for user simulation and queried to deliver various design insights.
John J. Dudley, Jason T. Jacques, Per Ola Kristensson
CHI3
2019 Crowdworker Economics in the Gig Economy
abstract
The nature of work is changing. As labor increasingly trends to casual work in the emerging gig economy, understanding the broader economic context is crucial to effective engagement with a contingent workforce. Crowdsourcing represents an early manifestation of this fluid, laisser-faire, on-demand workforce. This work analyzes the results of four large-scale surveys of US-based Amazon Mechanical Turk workers recorded over a six-year period, providing comparable measures to national statistics. Our results show that despite unemployment far higher than national levels, crowdworkers are seeing positive shifts in employment status and household income. Our most recent surveys indicate a trend away from full-time-equivalent crowdwork, coupled with a reduction in estimated poverty levels to below national figures. These trends are indicative of an increasingly flexible workforce, able to maximize their opportunities in a rapidly changing national labor market, which may have material impacts on existing models of crowdworker behavior.
Jason T. Jacques, Per Ola Kristensson
CHI2
2019 VelociWatch: Designing and Evaluating a Virtual Keyboard for the Input of Challenging Text
abstract
Virtual keyboard typing is typically aided by an auto-correct method that decodes a user's noisy taps into their intended text. This decoding process can reduce error rates and possibly increase entry rates by allowing users to type faster but less precisely. However, virtual keyboard decoders sometimes make mistakes that change a user's desired word into another. This is particularly problematic for challenging text such as proper names. We investigate whether users can guess words that are likely to cause auto-correct problems and whether users can adjust their behavior to assist the decoder. We conduct computational experiments to decide what predictions to offer in a virtual keyboard and design a smartwatch keyboard named VelociWatch. Novice users were able to use the features of VelociWatch to enter challenging text at 17 words-per-minute with a corrected error rate of 3%. Interestingly, they wrote slightly faster and just as accurately on a simpler keyboard with limited correction options. Our finding suggest users may be able to type difficult words on a smartwatch simply by tapping precisely without the use of auto-correct.
Keith Vertanen, Dylan Gaines, Crystal Fletcher, Alex M. Stanage, Robbie Watling, Per Ola Kristensson
CHI6
2019 Performance Envelopes of Virtual Keyboard Text Input Strategies in Virtual Reality
abstract
Virtual and Augmented Reality deliver engaging interaction experiences that can transport and extend the capabilities of the user. To ensure these paradigms are more broadly usable and effective, however, it is necessary to also deliver many of the conventional functions of a smartphone or personal computer. It remains unclear how conventional input tasks, such as text entry, can best be translated into virtual and augmented reality. In this paper we examine the performance potential of four alternative text entry strategies in virtual reality (VR). These four strategies are selected to provide full coverage of two fundamental design dimensions: i) physical surface association; and ii) number of engaged fingers. Specifically, we examine typing with index fingers on a surface and in mid-air and typing using all ten fingers on a surface and in mid-air. The central objective is to evaluate the human performance potential of these four typing strategies without being constrained by current tracking and statistical text decoding limitations. To this end we introduce an auto-correction simulator that uses knowledge of the stimulus to emulate statistical text decoding within constrained experimental parameters and use high-precision motion tracking hardware to visualise and detect fingertip interactions. We find that alignment of the virtual keyboard with a physical surface delivers significantly faster entry rates over a mid-air keyboard. Also, users overwhelmingly fail to effectively engage all ten fingers in mid-air typing, resulting in slower entry rates and higher error rates compared to just using two index fingers. In addition to identifying the envelopes of human performance for the four strategies investigated, we also provide a detailed analysis of the underlying features that distinguish each strategy in terms of its performance and behaviour.
John J. Dudley, Hrvoje Benko, Daniel J. Wigdor, Per Ola Kristensson
ISMAR4
2019 An Evaluation of Discrete and Continuous Mid-Air Loop and Marking Menu Selection in Optical See-Through HMDs
abstract
This paper investigates discrete and continuous hand-drawn loops and marks in mid-air as a selection input for gesture-based menu systems on optical see-through head-mounted displays (OST HMDs). We explore two fundamental methods of providing menu selection: the marking menu and the loop menu, and a hybrid method which combines the two. The loop menu design uses a selection mechanism with loops to approximate directional selections in a menu system. We evaluate the merits of loop and marking menu selection in an experiment with two phases and report that 1) the loop-based selection mechanism provides smooth and effective interaction; 2) users prioritize accuracy and comfort over speed for mid-air gestures; 3) users can exploit the flexibility of a final hybrid marking/loop menu design; and, finally, 4) users tend to chunk gestures depending on the selection task and their level of familiarity with the menu layout.
Zhi Han Lim, Per Ola Kristensson
MobileHCI2
2019 How do People Type on Mobile Devices?: Observations from a Study with 37, 000 Volunteers
abstract
This paper presents a large-scale dataset on mobile text entry collected via a web-based transcription task performed by 37,370 volunteers. The average typing speed was 36.2 WPM with 2.3% uncorrected errors. The scale of the data enables powerful statistical analyses on the correlation between typing performance and various factors, such as demographics, finger usage, and use of intelligent text entry techniques. We report effects of age and finger usage on performance that correspond to previous studies. We also find evidence of relationships between performance and use of intelligent text entry techniques: auto-correct usage correlates positively with entry rates, whereas word prediction usage has a negative correlation. To aid further work on modeling, machine learning and design improvements in mobile text entry, we make the code and dataset openly available.
Kseniia Palin, Anna Maria Feit, Sunjun Kim, Per Ola Kristensson, Antti Oulasvirta
MobileHCI4
2019 Ticker: An Adaptive Single-Switch Text Entry Method for Visually Impaired Users
abstract
Ticker is a probabilistic stereophonic single-switch text entry method for visually-impaired users with motor disabilities who rely on single-switch scanning systems to communicate. Such scanning systems are sensitive to a variety of noise sources, which are inevitably introduced in practical use of single-switch systems. Ticker uses a novel interaction model based on stereophonic sound coupled with statistical models for robust inference of the user's intended text in the presence of noise. As a consequence of its design, Ticker is resilient to noise and therefore a practical solution for single-switch scanning systems. Ticker's performance is validated using a combination of simulations and empirical user studies.
Emli-Mari Nel, Per Ola Kristensson, David J. C. MacKay
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 ReconViguRation: Reconfiguring Physical Keyboards in Virtual Reality
abstract
Physical keyboards are common peripherals for personal computers and are efficient standard text entry devices. Recent research has investigated how physical keyboards can be used in immersive head-mounted display-based Virtual Reality (VR). So far, the physical layout of keyboards has typically been transplanted into VR for replicating typing experiences in a standard desktop environment. In this paper, we explore how to fully leverage the immersiveness of VR to change the input and output characteristics of physical keyboard interaction within a VR environment. This allows individual physical keys to be reconfigured to the same or different actions and visual output to be distributed in various ways across the VR representation of the keyboard. We explore a set of input and output mappings for reconfiguring the virtual presentation of physical keyboards and probe the resulting design space by specifically designing, implementing and evaluating nine VR-relevant applications: emojis, languages and special characters, application shortcuts, virtual text processing macros, a window manager, a photo browser, a whack-a-mole game, secure password entry and a virtual touch bar. We investigate the feasibility of the applications in a user study with 20 participants and find that, among other things, they are usable in VR. We discuss the limitations and possibilities of remapping the input and output characteristics of physical keyboards in VR based on empirical findings and analysis and suggest future research directions in this area.
Daniel Schneider 0006, Alexander Otte, Travis Gesslein, Philipp Gagel, Bastian Kuth, Mohamad Shahm Damlakhi, Oliver Dietz, Eyal Ofek, Michel Pahud, Per Ola Kristensson, Jörg Müller 0001, Jens Grubert
IEEE Trans. Vis. Comput. Graph.10
2018 Bare-Handed 3D Drawing in Augmented Reality
abstract
Head-mounted augmented reality (AR) enables embodied in situ drawing in three dimensions (3D). We explore 3D drawing interactions based on uninstrumented, unencumbered (bare) hands that preserve the user's ability to freely navigate and interact with the physical environment. We derive three alternative interaction techniques supporting bare-handed drawing in AR from the literature and by analysing several envisaged use cases. The three interaction techniques are evaluated in a controlled user study examining three distinct drawing tasks: planar drawing, path description, and 3D object reconstruction. The results indicate that continuous freehand drawing supports faster line creation than the control point based alternatives, although with reduced accuracy. User preferences for the different techniques are mixed and vary considerably between the different tasks, highlighting the value of diverse and flexible interactions. The combined effectiveness of these three drawing techniques is illustrated in an example application of 3D AR drawing.
John J. Dudley, Hendrik Schuff, Per Ola Kristensson
Conference on Designing Interactive Systems3
2018 Change Blindness in Proximity-Aware Mobile Interfaces
abstract
Interface designs on both small and large displays can encourage people to alter their physical distance to the display. Mobile devices support this form of interaction naturally, as the user can move the device closer or further away as needed. The current generation of mobile devices can employ computer vision, depth sensing and other inference methods to determine the distance between the user and the display. Once this distance is known, a system can adapt the rendering of display content accordingly and enable proximity-aware mobile interfaces. The dominant method of exploiting proximity-aware interfaces is to remove or superimpose visual information. In this paper, we investigate change blindness in such interfaces. We present the results of two experiments. In our first experiment we show that a proximity-aware mobile interface results in significantly more change blindness errors than a non-moving interface. The absolute difference in error rates was 13.7%. In our second experiment we show that within a proximity-aware mobile interface, gradual changes induce significantly more change blindness errors than instant changes---confirming expected change blindness behavior. Based on our results we discuss the implications of either exploiting change blindness effects or mitigating them when designing mobile proximity-aware interfaces.
Michael Oliver Brock, Aaron J. Quigley, Per Ola Kristensson
CHI3
2018 eystrokes
abstract
We report on typing behaviour and performance of 168,000 volunteers in an online study. The large dataset allows detailed statistical analyses of keystroking patterns, linking them to typing performance. Besides reporting distributions and confirming some earlier findings, we report two new findings. First, letter pairs that are typed by different hands or fingers are more predictive of typing speed than, for example, letter repetitions. Second, rollover-typing, wherein the next key is pressed before the previous one is released, is sur- prisingly prevalent. Notwithstanding considerable variation in typing patterns, unsupervised clustering using normalised inter-key intervals reveals that most users can be divided into eight groups of typists that differ in performance, accuracy, hand and finger usage, and rollover. The code and dataset are released for scientific use.
Vivek Dhakal, Anna Maria Feit, Per Ola Kristensson, Antti Oulasvirta
CHI3
2018 The Impact of Word, Multiple Word, and Sentence Input on Virtual Keyboard Decoding Performance
abstract
Entering text on non-desktop computing devices is often done via an onscreen virtual keyboard. Input on such keyboards normally consists of a sequence of noisy tap events that specify some amount of text, most commonly a single word. But is single word-at-a-time entry the best choice? This paper compares user performance and recognition accuracy of word-at-a-time, phrase-at-a-time, and sentence-at-a-time text entry on a smartwatch keyboard. We evaluate the impact of differing amounts of input in both text copy and free composition tasks. We found providing input of an entire sentence significantly improved entry rates from 26 wpm to 32 wpm while keeping character error rates below 4%. In offline experiments with more processing power and memory, sentence input was recognized with a much lower 2.0% error rate. Our findings suggest virtual keyboards can enhance performance by encouraging users to provide more input per recognition event.
Keith Vertanen, Crystal Fletcher, Dylan Gaines, Jacob Gould, Per Ola Kristensson
CHI5
2018 Effects of Hand Representations for Typing in Virtual Reality
abstract
Alphanumeric text entry is a challenge for Virtual Reality (VR) applications. VR enables new capabilities, impossible in the real world, such as an unobstructed view of the keyboard, without occlusion by the user's physical hands. Several hand representations have been proposed for typing in VR on standard physical keyboards. However, to date, these hand representations have not been compared regarding their performance and effects on presence for VR text entry. Our work addresses this gap by comparing existing hand representations with minimalistic fingertip visualization. We study the effects of four hand representations (no hand representation, inverse kinematic model, fingertip visualization using spheres and video inlay) on typing in VR using a standard physical keyboard with 24 participants. We found that the fingertip visualization and video inlay both resulted in statistically significant lower text entry error rates compared to no hand or inverse kinematic model representations. We found no statistical differences in text entry speed.
Jens Grubert, Lukas Witzani, Eyal Ofek, Michel Pahud, Matthias Kranz, Per Ola Kristensson
VR6
2018 Text Entry in Immersive Head-Mounted Display-Based Virtual Reality Using Standard Keyboards
abstract
We study the performance and user experience of two popular mainstream text entry devices, desktop keyboards and touchscreen keyboards, for use in Virtual Reality (VR) applications. We discuss the limitations arising from limited visual feedback, and examine the efficiency of different strategies of use. We analyze a total of 24 hours of typing data in VR from 24 participants and find that novice users are able to retain about 60% of their typing speed on a desktop keyboard and about 40-45% of their typing speed on a touchscreen keyboard. We also find no significant learning effects, indicating that users can transfer their typing skills fast into VR. Besides investigating baseline performances, we study the position in which keyboards and hands are rendered in space. We find that this does not adversely affect performance for desktop keyboard typing and results in a performance trade-off for touchscreen keyboard typing.
Jens Grubert, Lukas Witzani, Eyal Ofek, Michel Pahud, Matthias Kranz, Per Ola Kristensson
VR6
2018 Performance Envelopes of in-Air Direct and Smartwatch Indirect Control for Head-Mounted Augmented Reality
abstract
The scarcity of established input methods for augmented reality (AR) head-mounted displays (HMD) motivates us to investigate the performance envelopes of two easily realisable solutions: indirect cursor control via a smartwatch and direct control by in-air touch. Indirect cursor control via a smartwatch has not been previously investigated for AR HMDs. We evaluate these two techniques for carrying out three fundamental user interface actions: target acquisition, goal crossing, and circular steering. We find that in-air is faster than smartwatch (p <; 0.001) for target acquisition and circular steering. We observe, however, that in-air selection can lead to discomfort after extended use and suggest that smartwatch control offers a complementary alternative.
Dennis Wolf 0002, John J. Dudley, Per Ola Kristensson
VR3
2018 A Review of User Interface Design for Interactive Machine Learning
abstract
Interactive Machine Learning (IML) seeks to complement human perception and intelligence by tightly integrating these strengths with the computational power and speed of computers. The interactive process is designed to involve input from the user but does not require the background knowledge or experience that might be necessary to work with more traditional machine learning techniques. Under the IML process, non-experts can apply their domain knowledge and insight over otherwise unwieldy datasets to find patterns of interest or develop complex data-driven applications. This process is co-adaptive in nature and relies on careful management of the interaction between human and machine. User interface design is fundamental to the success of this approach, yet there is a lack of consolidated principles on how such an interface should be implemented. This article presents a detailed review and characterisation of Interactive Machine Learning from an interactive systems perspective. We propose and describe a structural and behavioural model of a generalised IML system and identify solution principles for building effective interfaces for IML. Where possible, these emergent solution principles are contextualised by reference to the broader human-computer interaction literature. Finally, we identify strands of user interface research key to unlocking more efficient and productive non-expert interactive machine learning applications.
John J. Dudley, Per Ola Kristensson
ACM Trans. Interact. Intell. Syst.2
2018 Fast and Precise Touch-Based Text Entry for Head-Mounted Augmented Reality with Variable Occlusion
abstract
We present the VISAR keyboard: An augmented reality (AR) head-mounted display (HMD) system that supports text entry via a virtualised input surface. Users select keys on the virtual keyboard by imitating the process of single-hand typing on a physical touchscreen display. Our system uses a statistical decoder to infer users’ intended text and to provide error-tolerant predictions. There is also a high-precision fall-back mechanism to support users in indicating which keys should be unmodified by the auto-correction process. A unique advantage of leveraging the well-established touch input paradigm is that our system enables text entry with minimal visual clutter on the see-through display, thus preserving the user’s field-of-view. We iteratively designed and evaluated our system and show that the final iteration of the system supports a mean entry rate of 17.75wpm with a mean character error rate less than 1%. This performance represents a 19.6% improvement relative to the state-of-the-art baseline investigated: A gaze-then-gesture text entry technique derived from the system keyboard on the Microsoft HoloLens. Finally, we validate that the system is effective in supporting text entry in a fully mobile usage scenario likely to be encountered in industrial applications of AR HMDs.
John J. Dudley, Keith Vertanen, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.3
2017 Investigating Tilt-based Gesture Keyboard Entry for Single-Handed Text Entry on Large Devices
abstract
The popularity of mobile devices with large screens is making single-handed interaction difficult. We propose and evaluate a novel design point around a tilt-based text entry technique which supports single handed usage. Our technique is based on the gesture keyboard (shape writing). However, instead of drawing gestures with a finger or stylus, users articulate a gesture by tilting the device. This can be especially useful when the user's other hand is otherwise encumbered or unavailable. We show that novice users achieve an entry rate of 15 words-per-minute (wpm) after minimal practice. A pilot longitudinal study reveals that a single participant achieved an entry rate of 32 wpm after approximate 90 minutes of practice. Our data indicate that tilt-based gesture keyboard entry enables walk-up use and provides a suitable text entry rate for occasional use and can act as a promising alternative to single-handed typing in certain situations.
Hui-Shyong Yeo, Xiao-Shen Phang, Steven J. Castellucci, Per Ola Kristensson, Aaron J. Quigley
CHI4
2016 What's Hot in Intelligent User Interfaces
abstract
The ACM Conference on Intelligent User Interfaces (IUI) is the annual meeting of the intelligent user interface community and serves as a premier international forum for reporting outstanding research and development on intelligent user interfaces. ACM IUI is where the Human-Computer Interaction (HCI) community meets the Artificial Intelligence (AI) community. Here we summarize the latest trends in IUI based on our experience organizing the 20th ACM IUI Conference in Atlanta in 2015.
Shimei Pan, Oliver Brdiczka, Giuseppe Carenini, Polo Chau, Per Ola Kristensson
AAAI5
2016 Dexmo: An Inexpensive and Lightweight Mechanical Exoskeleton for Motion Capture and Force Feedback in VR
abstract
We present Dexmo: an inexpensive and lightweight mechanical exoskeleton system for motion capturing and force feedback in virtual reality applications. Dexmo combines multiple types of sensors, actuation units and link rod structures to provide users with a pleasant virtual reality experience. The device tracks the user's motion and uniquely provides passive force feedback. In combination with a 3D graphics rendered environment, Dexmo provides the user with a realistic sensation of interaction when a user is for example grasping an object. An initial evaluation with 20 participants demonstrate that the device is working reliably and that the addition of force feedback resulted in a significant reduction in error rate. Informal comments by the participants were overwhelmingly positive.
Xiaochi Gu, Weize Sun, Yuanzhe Bian, Dao Zhou, Per Ola Kristensson
CHI6
2015 Performance and User Experience of Touchscreen and Gesture Keyboards in a Lab Setting and in the Wild
abstract
We study the performance and user experience of two popular mainstream mobile text entry methods: the Smart Touch Keyboard (STK) and the Smart Gesture Keyboard (SGK). Our first study is a lab-based ten-session text entry experiment. In our second study we use a new text entry evaluation methodology based on the experience sampling method (ESM). In the ESM study, participants installed an Android app on their own mobile phones that periodically sampled their text entry performance and user experience amid their everyday activities for four weeks. The studies show that text can be entered at an average speed of 28 to 39 WPM, depending on the method and the user's experience, with 1.0% to 3.6% character error rates remaining. Error rates of touchscreen input, particularly with SGK, are a major challenge; and reducing out-of-vocabulary errors is particularly important. Both SGK and STK have strengths, weaknesses, and different individual awareness and preferences. Two-thumb touch typing in a focused setting is particularly effective on STK, whereas one-handed SGK typing with the thumb is particularly effective in more mobile situations. When exposed to both, users tend to migrate from STK to SGK. We also conclude that studies in the lab and in the wild can both be informative to reveal different aspects of keyboard experience, but used in conjunction is more reliable in comprehensively assessing input technologies of current and future generations.
Shyam Reyal, Shumin Zhai, Per Ola Kristensson
CHI3
2015 VelociTap: Investigating Fast Mobile Text Entry using Sentence-Based Decoding of Touchscreen Keyboard Input
abstract
We present VelociTap: a state-of-the-art touchscreen keyboard decoder that supports a sentence-based text entry approach. VelociTap enables users to seamlessly choose from three word-delimiter actions: pushing a space key, swiping to the right, or simply omitting the space key and letting the decoder infer spaces automatically. We demonstrate that VelociTap has a significantly lower error rate than Google's keyboard while retaining the same entry rate. We show that intermediate visual feedback does not significantly affect entry or error rates and we find that using the space key results in the most accurate results. We also demonstrate that enabling flexible word-delimiter options does not incur an error rate penalty. Finally, we investigate how small we can make the keyboard when using VelociTap. We show that novice users can reach a mean entry rate of 41 wpm on a 40 mm wide smartwatch-sized keyboard at a 3% character error rate.
Keith Vertanen, Haythem Memmi, Justin Emge, Shyam Reyal, Per Ola Kristensson
CHI5
2014 Phoneme-based predictive text entry interface
abstract
Phoneme-based text entry provides an alternative typing method for nonspeaking individuals who often experience difficulties in orthographic spelling. In this paper, we investigate the application of rate enhancement strategies to improve the user performance of phoneme-based text entry systems. We have developed a phoneme-based predictive typing system, which employs statistical language modeling techniques to dynamically reduce the phoneme search space and offer accurate word predictions. Results of a case study with a nonspeaking participant demonstrated that our rate enhancement strategies led to improved text entry speed and error rates.
Ha Trinh, Annalu Waller, Keith Vertanen, Per Ola Kristensson, Vicki L. Hanson
ASSETS4
2014 AwToolkit: attention-aware user interface widgets
abstract
Increasing screen real-estate allows for the development of applications where a single user can manage a large amount of data and related tasks through a distributed user interface. However, such users can easily become overloaded and become unaware of display changes as they alternate their attention towards different displays. We propose AwToolkit, a novel widget set for developers that supports users in maintaing awareness in multi-display systems. The AwToolkit widgets automatically determine which display a user is looking at and provide users with notifications with different levels of subtlety to make the user aware of any unattended display changes. The toolkit uses four notification levels (unnoticeable, subtle, intrusive and disruptive), ranging from an almost imperceptible visual change to a clear and visually saliant change. We describe AwToolkit's six widgets, which have been designed for C# developers, and the design of a user study with an application oriented towards healthcare environments. The evaluation results reveal a marked increase in user awareness in comparison to the same application implemented without AwToolkit.
Juan Enrique Garrido, Victor M. Ruiz Penichet, María Dolores Lozano 0001, Aaron J. Quigley, Per Ola Kristensson
AVI5
2014 An evaluation of Dasher with a high-performance language model as a gaze communication method
abstract
Dasher is a promising fast assistive gaze communication method. However, previous evaluations of Dasher have been inconclusive. Either the studies have been too short, involved too few participants, suffered from sampling bias, lacked a control condition, used an inappropriate language model, or a combination of the above. To rectify this, we report results from two new evaluations of Dasher carried out using a Tobii P10 assistive eye-tracker machine. We also present a method of modifying Dasher so that it can use a state-of-the-art long-span statistical language model. Our experimental results show that compared to a baseline eye-typing method, Dasher resulted in significantly faster entry rates (12.6 wpm versus 6.0 wpm in Experiment 1, and 14.2 wpm versus 7.0 wpm in Experiment 2). These faster entry rates were possible while maintaining error rates comparable to the baseline eye-typing method. Participants' perceived physical demand, mental demand, effort and frustration were all significantly lower for Dasher. Finally, participants significantly rated Dasher as being more likeable, requiring less concentration and being more fun.
Daniel J. Rough, Keith Vertanen, Per Ola Kristensson
AVI3
2014 Modeling the perception of user performance
abstract
This paper studies how users perceive their own performance in two alternative user interfaces. We extend methodology from psychophysics to the study of interactive performance and conduct two experiments in order to create a model of users' perception of their own performance. In our studies, two interfaces are sequentially used in a pointing task, and users are asked to rate in which interface their performance was higher. We first differentiate the effects of objective performance (speed and accuracy) versus interface qualities (distance between elements and width of elements) on perceived performance. We then derive a model that predicts the amount of change required in an interface for users to reliably detect a difference. The model is useful as a heuristic for predicting if a new interface design is better enough for users to reliably appreciate the obtained gain in user performance. We validate the model via a separate user study, and conclude by discussing how to apply our findings to design problems.
Max Nicosia, Antti Oulasvirta, Per Ola Kristensson
CHI3
2014 Uncertain text entry on mobile devices
abstract
Users often struggle to enter text accurately on touchscreen keyboards. To address this, we present a flexible decoder for touchscreen text entry that combines probabilistic touch models with a language model. We investigate two different touch models. The first touch model is based on a Gaussian Process regression approach and implicitly models the inherent uncertainty of the touching process. The second touch model allows users to explicitly control the uncertainty via touch pressure. Using the first model we show that the character error rate can be reduced by up to 7% over a baseline method, and by up to 1.3% over a leading commercial keyboard. Using the second model we demonstrate that providing users with control over input certainty reduces the amount of text users have to correct manually and increases the text entry rate.
Daryl Weir, Henning Pohl, Simon Rogers, Keith Vertanen, Per Ola Kristensson
CHI5
2014 SpiderEyes: designing attention- and proximity-aware collaborative interfaces for wall-sized displays
abstract
With the proliferation of large multi-faceted datasets, a critical question is how to design collaborative environments, in which this data can be analysed in an efficient and insightful manner. Exploiting people's movements and distance to the data display and to collaborators, proxemic interactions can potentially support such scenarios in a fluid and seamless way, supporting both tightly coupled collaboration as well as parallel explorations. In this paper we introduce the concept of collaborative proxemics: enabling groups of people to collaboratively use attention- and proximity-aware applications. To help designers create such applications we have developed SpiderEyes: a system and toolkit for designing attention- and proximity-aware collaborative interfaces for wall-sized displays. SpiderEyes is based on low-cost technology and allows accurate markerless attention-aware tracking of multiple people interacting in front of a display in real-time. We discuss how this toolkit can be applied to design attention- and proximity-aware collaborative scenarios around large wall-sized displays, and how the information visualisation pipeline can be extended to incorporate proxemic interactions.
Jakub Dostal, Uta Hinrichs, Per Ola Kristensson, Aaron J. Quigley
IUI3
2014 The inviscid text entry rate and its application as a grand goal for mobile text entry
abstract
We introduce the concept of the inviscid text entry rate: the point when the user's creativity is the bottleneck rather than the text entry method. We then apply the inviscid text entry rate to define a grand goal for mobile text entry. Via a proxy measure we estimate the population mean of the sufficiently inviscid entry rate to be 67 wpm. We then compare existing mobile text entry methods against this estimate and find that the vast majority of text entry methods in the literature are substantially slower. This analysis suggests the mobile text entry field needs to focus on methods that can viably approach the inviscid entry rate.
Per Ola Kristensson, Keith Vertanen
Mobile HCI1
2014 Estimating and using absolute and relative viewing distance in interactive systems
Jakub Dostal, Per Ola Kristensson, Aaron J. Quigley
Pervasive Mob. Comput.2
2014 Complementing text entry evaluations with a composition task
abstract
A common methodology for evaluating text entry methods is to ask participants to transcribe a predefined set of memorable sentences or phrases. In this article, we explore if we can complement the conventional transcription task with a more externally valid composition task. In a series of large-scale crowdsourced experiments, we found that participants could consistently and rapidly invent high quality and creative compositions with only modest reductions in entry rates. Based on our series of experiments, we provide a best-practice procedure for using composition tasks in text entry evaluations. This includes a judging protocol which can be performed either by the experimenters or by crowdsourced workers on a microtask market. We evaluated our composition task procedure using a text entry method unfamiliar to participants. Our empirical results show that the composition task can serve as a valid complementary text entry evaluation method.
Keith Vertanen, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.2
2013 The feasibility of eyes-free touchscreen keyboard typing
abstract
Typing on a touchscreen keyboard is very difficult without being able to see the keyboard. We propose a new approach in which users imagine a Qwerty keyboard somewhere on the device and tap out an entire sentence without any visual reference to the keyboard and without intermediate feedback about the letters or words typed. To demonstrate the feasibility of our approach, we developed an algorithm that decodes blind touchscreen typing with a character error rate of 18.5%. Our decoder currently uses three components: a model of the keyboard topology and tap variability, a point transformation algorithm, and a long-span statistical language model. Our initial results demonstrate that our proposed method provides fast entry rates and promising error rates. On one-third of the sentences, novices' highly noisy input was successfully decoded with no errors.
Keith Vertanen, Haythem Memmi, Per Ola Kristensson
ASSETS3
2013 Multi-touch rotation gestures: performance and ergonomics
abstract
Rotations performed with the index finger and thumb involve some of the most complex motor action among common multi-touch gestures, yet little is known about the factors affecting performance and ergonomics. This note presents results from a study where the angle, direction, diameter, and position of rotations were systematically manipulated. Subjects were asked to perform the rotations as quickly as possible without losing contact with the display, and were allowed to skip rotations that were too uncomfortable. The data show surprising interaction effects among the variables, and help us identify whole categories of rotations that are slow and cumbersome for users.
Eve E. Hoggan, John Williamson 0001, Antti Oulasvirta, Miguel A. Nacenta, Per Ola Kristensson, Anu Lehtiö
CHI5
2013 Memorability of pre-designed and user-defined gesture sets
abstract
We studied the memorability of free-form gesture sets for invoking actions. We compared three types of gesture sets: user-defined gesture sets, gesture sets designed by the authors, and random gesture sets in three studies with 33 participants in total. We found that user-defined gestures are easier to remember, both immediately after creation and on the next day (up to a 24% difference in recall rate compared to pre-designed gestures). We also discovered that the differences between gesture sets are mostly due to association errors (rather than gesture form errors), that participants prefer user-defined sets, and that they think user-defined gestures take less time to learn. Finally, we contribute a qualitative analysis of the tradeoffs involved in gesture type selection and share our data and a video corpus of 66 gestures for replicability and further analysis.
Miguel A. Nacenta, Yemliha Kamber, Yizhou Qiang, Per Ola Kristensson
CHI4
2013 Improving two-thumb text entry on touchscreen devices
abstract
We study the design of split keyboards for fast text entry with two thumbs on mobile touchscreen devices. The layout of KALQ was determined through first studying how users should grip a device with two hands. We then assigned letters to keys computationally, using a model of two-thumb tapping. KALQ minimizes thumb travel distance and maximizes alternation between thumbs. An error correction algorithm was added to help address linguistic and motor errors. Users reached a rate of 37 words per minute (with a 5% error rate) after a training program.
Antti Oulasvirta, Anna Reichel, Wenbin Li 0003, Yan Zhang 0001, Myroslav Bachynskyi, Keith Vertanen, Per Ola Kristensson
CHI7
2013 Crowdsourcing a HIT: Measuring Workers' Pre-Task Interactions on Microtask Markets
abstract
The ability to entice and engage crowd workers to participate in human intelligence tasks (HITs) is critical for many human computation systems and large-scale experiments. While various metrics have been devised to measure and improve the quality of worker output via task designs, effective recruitment of crowd workers is often overlooked. To help us gain a better understanding of crowd recruitment strategies we propose three new metrics for measuring crowd workers' willingness to participate in advertised HITs: conversion rate, conversion rate over time, and nominal conversion rate. We discuss how the conversion rate of workers—the number of potential workers aware of a task that choose to accept the task—can affect the quantity, quality, and validity of any data collected via crowdsourcing. We also contribute a tool — turkmill — that enables requesters on Amazon Mechanical Turk to easily measure the conversion rate of HITs. We then present the results of two experiments that demonstrate how conversion rate metrics can be used to evaluate the effect of different HIT designs. We investigate how four HIT design features (value proposition, branding, quality of presentation, and intrinsic motivation) affect conversion rates. Among other things, we find that including a clear value proposition has a strong significant, positive effect on the nominal conversion rate. We also find that crowd workers prefer commercial entities to non-profit or university requesters.
Jason T. Jacques, Per Ola Kristensson
HCOMP2
2013 Subtle gaze-dependent techniques for visualising display changes in multi-display environments
abstract
This paper explores techniques for visualising display changes in multi-display environments. We present four subtle gaze-dependent techniques for visualising change on unattended displays called FreezeFrame, PixMap, WindowMap and Aura. To enable the techniques to be directly deployed to workstations, we also present a system that automatically identifies the user's eyes using computer vision and a set of web cameras mounted on the displays. An evaluation confirms this system can detect which display the user is attending to with high accuracy. We studied the efficacy of the visualisation techniques in a five-day case study with a working professional. This individual used our system eight hours per day for five consecutive days. The results of the study show that the participant found the system and the techniques useful, subtle, calm and non-intrusive. We conclude by discussing the challenges in evaluating intelligent subtle interaction techniques using traditional experimental paradigms.
Jakub Dostal, Per Ola Kristensson, Aaron J. Quigley
IUI2
2012 iSCAN: a phoneme-based predictive communication aid for nonspeaking individuals
abstract
The high incidence of literacy deficits among people with severe speech impairments (SSI) has been well documented. Without literacy skills, people with SSI are unable to effectively use orthographic-based communication systems to generate novel linguistic items in spontaneous conversation. To address this problem, phoneme-based communication systems have been proposed which enable users to create spoken output from phoneme sequences. In this paper, we investigate whether prediction techniques can be employed to improve the usability of such systems. We have developed iSCAN, a phoneme-based predictive communication system, which offers phoneme prediction and phoneme-based word prediction. A pilot study with 16 able-bodied participants showed that our predictive methods led to a 108.4% increase in phoneme entry speed and a 79.0% reduction in phoneme error rate. The benefits of the predictive methods were also demonstrated in a case study with a cerebral palsied participant. Moreover, results of a comparative evaluation conducted with the same participant after 16 sessions using iSCAN indicated that our system outperformed an orthographic-based predictive communication device that the participant has used for over 4 years.
Ha Trinh, Annalu Waller, Keith Vertanen, Per Ola Kristensson, Vicki L. Hanson
ASSETS4
2012 I did that!: measuring users' experience of agency in their own actions
abstract
Cognitive neuroscience defines the sense of agency as the experience of controlling one's own actions and, through this control, affecting the external world. We believe that the sense of personal agency is a key factor in how people experience interactions with technology. This paper draws on theoretical perspectives in cognitive neuroscience and describes two implicit methods through which personal agency can be empirically investigated. We report two experiments applying these methods to HCI problems. One shows that a new input modality - skin-based interaction - can substantially increase users' sense of agency. The second demonstrates that variations in the parameters of assistance techniques such as predictive mouse acceleration can have a significant impact on users' sense of agency. The methods presented provide designers with new ways of evaluating and refining empowering interaction techniques and interfaces, in which users experience an instinctive sense of control and ownership over their actions.
David Coyle, James W. Moore, Per Ola Kristensson, Paul C. Fletcher, Alan F. Blackwell
CHI3
2012 The potential of dwell-free eye-typing for fast assistive gaze communication
abstract
We propose a new research direction for eye-typing which is potentially much faster: dwell-free eye-typing. Dwell-free eye-typing is in principle possible because we can exploit the high redundancy of natural languages to allow users to simply look at or near their desired letters without stopping to dwell on each letter. As a first step we created a system that simulated a perfect recognizer for dwell-free eye-typing. We used this system to investigate how fast users can potentially write using a dwell-free eye-typing interface. We found that after 40 minutes of practice, users reached a mean entry rate of 46 wpm. This indicates that dwell-free eye-typing may be more than twice as fast as the current state-of-the-art methods for writing by gaze. A human performance model further demonstrates that it is highly unlikely traditional eye-typing systems will ever surpass our dwell-free eye-typing performance estimate.
Per Ola Kristensson, Keith Vertanen
ETRA1
2012 Spelling as a Complementary Strategy for Speech Recognition
abstract
We compare a variety of strategies for incorporating spelling to create more robust voice-only speech in-terfaces. These strategies use different combinations of speaking the word, spelling the word, and spelling the word using a phonetic alphabet. For correcting a single recognition error, spelling the word or speaking and spelling the word reduced error rates substantially. Phonetic-spelling was very accurate with error rates on a 5K task approaching zero. Most importantly, multiple input strategies can be used simultaneously with only a modest degradation in performance compared to allow-ing only a single input strategy. Thus our work shows that spelling-based input strategies offer the potential of a simple, natural and effective way for users to both avoid and correct recognition errors. Index Terms: speech recognition, error correction 1.
Keith Vertanen, Per Ola Kristensson
INTERSPEECH2
2012 Continuous recognition of one-handed and two-handed gestures using 3D full-body motion tracking sensors
abstract
In this paper we present a new bimanual markerless gesture interface for 3D full-body motion tracking sensors, such as the Kinect. Our interface uses a probabilistic algorithm to incrementally predict users' intended one-handed and twohanded gestures while they are still being articulated. It supports scale and translation invariant recognition of arbitrarily defined gesture templates in real-time. The interface supports two ways of gesturing commands in thin air to displays at a distance. First, users can use one-handed and two-handed gestures to directly issue commands. Second, users can use their non-dominant hand to modulate single-hand gestures. Our evaluation shows that the system recognizes one-handed and two-handed gestures with an accuracy of 92.7%--96.2%.
Per Ola Kristensson, Thomas Nicholson, Aaron J. Quigley
IUI1
2012 Performance comparisons of phrase sets and presentation styles for text entry evaluations
abstract
We empirically compare five different publicly-available phrase sets in two large-scale (N = 225 and N = 150) crowdsourced text entry experiments. We also investigate the impact of asking participants to memorize phrases before writing them versus allowing participants to see the phrase during text entry. We find that asking participants to memorize phrases increases entry rates at the cost of slightly increased error rates. This holds for both a familiar and for an unfamiliar text entry method. We find statistically significant differences between some of the phrase sets in terms of both entry and error rates. Based on our data, we arrive at a set of recommendations for choosing suitable phrase sets for text entry evaluations.
Per Ola Kristensson, Keith Vertanen
IUI1
2011 Design Dimensions of Intelligent Text Entry Tutors
Per Ola Kristensson
AIED1
2011 The Imagination of Crowds: Conversational AAC Language Modeling using Crowdsourcing and Large Data Sources
Keith Vertanen, Per Ola Kristensson
EMNLP2
2011 Asynchronous Multimodal Text Entry Using Speech and Gesture Keyboards
abstract
We propose reducing errors in text entry by combining speech and gesture keyboard input. We describe a merge model that combines recognition results in an asynchronous and flexible manner. We collected speech and gesture data of users entering both short email sentences and web search queries. By merging recognition results from both modalities, word error rate was reduced by 53% relative for email sentences and 29% relative for web searches. For email utterances with speech errors, we investigated providing gesture keyboard corrections of only the erroneous words. Without the user explicitly indicating the incorrect words, our model was able to reduce the word error rate by 44% relative. Copyright © 2011 ISCA.
Per Ola Kristensson, Keith Vertanen
INTERSPEECH1
2011 A versatile dataset for text entry evaluations based on genuine mobile emails
abstract
Mobile text entry methods are typically evaluated by having study participants copy phrases. However, currently there is no available phrase set that has been composed by mobile users. Instead researchers have resorted to using invented phrases that probably suffer from low external validity. Further, there is no available phrase set whose phrases have been verified to be memorable. In this paper we present a collection of mobile email sentences written by actual users on actual mobile devices. We obtained our sentences from emails written by Enron employees on their BlackBerry mobile devices. We provide empirical data on how easy the sentences were to remember and how quickly and accurately users could type these sentences on a full-sized keyboard. Using this empirical data, we construct a series of phrase sets we suggest for use in text entry evaluations.
Keith Vertanen, Per Ola Kristensson
Mobile HCI2
2010 Intelligently Aiding Human-Guided Correction of Speech Recognition
abstract
Correcting recognition errors is often necessary in a speech interface. These errors not only reduce users' overall entry rate, but can also lead to frustration. While making fewer recognition errors is undoubtedly helpful, facilities for supporting user-guided correction are also critical. We explore how to better support user corrections using Parakeet — a continuous speech recognition system for mobile touch-screen devices. Parakeet's interface is designed for easy error correction on a handheld device. Users correct errors by selecting alternative words from a word confusion network and by typing on a predictive software keyboard. Our interface design was guided by computational experiments and used a variety of information sources to aid the correction process. In user studies, participants were able to write text effectively despite sometimes high initial recognition error rates. Using Parakeet as an example, we discuss principles we found were important for building an effective speech correction interface.
Keith Vertanen, Per Ola Kristensson
AAAI2
2010 Getting it right the second time: Recognition of spoken corrections
abstract
We investigate ways to improve recognition accuracy on spoken corrections. We show that a variety of simple techniques can greatly improve the accuracy on corrections. We further develop a flexible merge model that improves accuracy by combining information from the original recognition and the spoken correction. Our merge model operates on word confusion networks and can easily incorporate prior beliefs about the recognition events (e.g. which words are likely correct or incorrect). By combining all of our techniques, the percentage of correctly recognized spoken corrections increased from 21% to 53%.
Keith Vertanen, Per Ola Kristensson
SLT2
2009 Automatic selection of recognition errors by respeaking the intended text
abstract
We investigate how to automatically align spoken corrections with an initial speech recognition result. Such automatic alignment would enable one-step voice-only correction in which users simply respeak their intended text. We present three new models for automatically aligning corrections: a 1-best model, a word confusion network model, and a revision model. The revision model allows users to alter what they intended to write even when the initial recognition was completely correct. We evaluate our models with data gathered from two user studies. We show that providing just a single correct word of context dramatically improves alignment success from 65% to 84%. We find that a majority of users provide such context without being explicitly instructed to do so. We find that the revision model is superior when users modify words in their initial recognition, improving alignment success from 73% to 83%. We show how our models can easily incorporate prior information about correction location and we show that such information aids alignment success. Last, we observe that users speak their intended text faster and with fewer re-recordings than if they are forced to speak misrecognized text.
Keith Vertanen, Per Ola Kristensson
ASRU2
2009 Text entry performance of state of the art unconstrained handwriting recognition: a longitudinal user study
abstract
We report on a longitudinal study of unconstrained handwriting recognition performance. After 250 minutes of practice, participants had a mean text entry rate of 24.1 wpm. For the first four hours of usage, entry and error rates of handwriting recognition are about the same as for a baseline QWERTY software keyboard. Our results reveal that unconstrained handwriting is faster than what was previously assumed in the text entry literature.
Per Ola Kristensson, Leif C. Denby
CHI1
2009 Recognition and correction of voice web search queries
abstract
In this work we investigate how to recognize and correct voice web search queries. We describe our corpus of web search queries and show how it was used to improve recognition accuracy. We show that using a search-specific vocabulary with automatically generated pronunciations is superior to using a vocabulary limited to a fixed pronunciation dictionary. We conducted a formative user study to investigate recognition and correction aspects of voice search in a mobile context. In the user study, we found that despite a word error rate of 48%, users were able to speak and correct search queries in about 18 seconds. Users did this while walking around using a mobile touch-screen device. Index Terms: speech recognition, voice search, error correction, mobile web search
Keith Vertanen, Per Ola Kristensson
INTERSPEECH2
2009 Parakeet: a continuous speech recognition system for mobile touch-screen devices
abstract
We present Parakeet, a system for continuous speech recognition on mobile touch-screen devices. The design of Parakeet was guided by computational experiments and validated by a user study. Participants had an average text entry rate of 18 words-per-minute (WPM) while seated indoors and 13 WPM while walking outdoors. In an expert pilot study, we found that speech recognition has the potential to be a highly competitive mobile text entry method, particularly in an actual mobile setting where users are walking around while entering text.
Keith Vertanen, Per Ola Kristensson
IUI2
2009 Parakeet: a demonstration of speech recognition on a mobile touch-screen device
abstract
We demonstrate Parakeet -- a continuous speech recognition system for mobile touch-screen devices. Parakeet's interface is designed to make correcting errors easy on a handheld device while on the move. Users correct errors using a touch-screen to either select alternative words from a word confusion network or by typing on a predictive software keyboard. Our interface design was guided by computational experiments. We conducted a user study to validate our design. We found novices entered text at 18 WPM while seated indoors and 13 WPM while walking outdoors.
Keith Vertanen, Per Ola Kristensson
IUI2
2009 An Evaluation of Space Time Cube Representation of Spatiotemporal Patterns
abstract
Space time cube representation is an information visualization technique where spatiotemporal data points are mapped into a cube. Information visualization researchers have previously argued that space time cube representation is beneficial in revealing complex spatiotemporal patterns in a data set to users. The argument is based on the fact that both time and spatial information are displayed simultaneously to users, an effect difficult to achieve in other representations. However, to our knowledge the actual usefulness of space time cube representation in conveying complex spatiotemporal patterns to users has not been empirically validated. To fill this gap, we report on a between-subjects experiment comparing novice users' error rates and response times when answering a set of questions using either space time cube or a baseline 2D representation. For some simple questions, the error rates were lower when using the baseline representation. For complex questions where the participants needed an overall understanding of the spatiotemporal structure of the data set, the space time cube representation resulted in on average twice as fast response times with no difference in error rates compared to the baseline. These results provide an empirical foundation for the hypothesis that space time cube representation benefits users analyzing complex spatiotemporal patterns.
Per Ola Kristensson, Nils Dahlbäck, Daniel Anundi, Marius Björnstad, Hanna Gillberg, Jonas Haraldsson, Ingrid Mårtensson, Mathias Nordvall, Josefine Ståhl
IEEE Trans. Vis. Comput. Graph.1
2008 On the benefits of confidence visualization in speech recognition
abstract
In a typical speech dictation interface, the recognizer's best-guess is displayed as normal, unannotated text. This ignores potentially useful information about the recognizer's confidence in its recognition hypothesis. Using a confidence measure (which itself may sometimes be inaccurate), we investigated providing visual feedback about low-confidence portions of the recognition using shaded, red underlining. An evaluation showed, compared to a baseline without underlining, underlining low-confidence areas did not increase user's speed or accuracy in detecting errors. However, we found that when recognition errors were correctly underlined, they were discovered significantly more often than baseline. Conversely, when errors failed to be underlined, they were discovered less often. Our results indicate confidence visualization can be effective --- but only if the confidence measure has high accuracy. Further, since our results show that users tend to trust confidence visualization, designers should be careful in its application if a high accuracy confidence measure is not available.
Keith Vertanen, Per Ola Kristensson
CHI2
2008 Interlaced QWERTY: accommodating ease of visual search and input flexibility in shape writing
abstract
Shape writing is an input technology for touch-screen mobile phones and pen-tablets. To shape write text, the user spells out word patterns by sliding a finger or stylus over a graphical keyboard. The user's trace is then recognized by a pattern recognizer. In this paper we analyze and evaluate various keyboard layouts, including alphabetic, optimized (ATOMIK), QWERTY, and interlaced QWERTY for shape writing. The goodness of a layout for shape writing has two aspects. For users' initial ease of use the letters should be easy to visually locate. For long term use, however, the layout should maximize the imprecision tolerance and writing flexibility for all words. We present empirical studies for the former and mathematical analyses for the latter. Our results led to a new layout, interlaced QWERTY, which offers excellent separation of word shapes, while still maintaining a low visual search time. Many of the findings in our study also apply to traditional soft keyboards tapped with a stylus or one finger.
Shumin Zhai, Per Ola Kristensson
CHI2
2008 Improving word-recognizers using an interactive lexicon with active and passive words
abstract
The words a user is likely to write comprise the user's active vocabulary. This vocabulary is considerably smaller than the passive vocabulary of words a user reads. We explore an interactive adaptive lexicon method that separates a large lexicon into active and passive sets, and gradually expands and adapts the active set to reflect the user's active vocabulary. The adaptation is achieved through lightweight interaction as a by product of actual use. The effectiveness of the technique is demonstrated through a computational experiment and a user study.
Per Ola Kristensson, Shumin Zhai
IUI1
2007 Hard lessons: effort-inducing interfaces benefit spatial learning
abstract
Interface designers normally strive for a design that minimises the user's effort. However, when the design's objective is to train users to interact with interfaces that are highly dependent on spatial properties (e.g. keypad layout or gesture shapes) we contend that designers should consider explicitly increasing the mental effort of interaction. To test the hypothesis that effort aids spatial memory, we designed a "frost-brushing" interface that forces the user to mentally retrieve spatial information, or to physically brush away the frost to obtain visual guidance. We report results from two experiments using virtual keypad interfaces -- the first concerns spatial location learning of buttons on the keypad, and the second concerns both location and trajectory learning of gesture shape. The results support our hypothesis, showing that the frost-brushing design improved spatial learning. The participants' subjective responses emphasised the connections between effort, engagement, boredom, frustration, and enjoyment, suggesting that effort requires careful parameterisation to maximise its effectiveness.
Andy Cockburn, Per Ola Kristensson, Jason Alexander, Shumin Zhai
CHI2
2007 Command strokes with and without preview: using pen gestures on keyboard for command selection
abstract
This paper presents a new command selection method that provides an alternative to pull-down menus in pen-based mobile interfaces. Its primary advantage is the ability forusers to directly select commands from a very large set without the need to traverse menu hierarchies. The proposed method maps the character strings representing the commands onto continuous pen-traces on a stylus keyboard. The user enters a command by stroking part of its character string. We call this method "command strokes." We present the results of three experiments assessing the usefulness of the technique. The first experiment shows that command strokes are 1.6 times faster than the de-facto standard pull-down menus and that users find command strokes more fun to use. The second and third experiments investigate the effect of displaying a visual preview of the currently recognized command while the user is still articulating the command stroke. These experiments show that visual preview does not slow users down and leads to significantly lower error rates and shorter gestures when users enter new unpracticed commands.
Per Ola Kristensson, Shumin Zhai
CHI1
2005 Relaxing stylus typing precision by geometric pattern matching
abstract
Fitts' law models the inherent speed-accuracy trade-off constraint in stylus typing. Users attempting to go beyond the Fitts' law speed ceiling will tend to land the stylus outside the targeted key, resulting in erroneous words and increasing users' frustration. We propose a geometric pattern matching technique to overcome this problem. Our solution can be used either as an enhanced spell checker or as a way to enable users to escape the Fitts' law constraint in stylus typing, potentially resulting in higher text entry speeds than what is currently theoretically modeled. We view the hit points on a stylus keyboard as a high resolution geometric pattern. This pattern can be matched against patterns formed by the letter key center positions of legitimate words in a lexicon. We present the development and evaluation of an "elastic" stylus keyboard capable of correcting words even if the user misses all the intended keys, as long as the user's tapping pattern is close enough to the intended word.
Per Ola Kristensson, Shumin Zhai
IUI1
2005 In search of effective text input interfaces for off the desktop computing
abstract
It is generally recognized that today's frontier of HCI research lies beyond the traditional desktop computers whose GUI interfaces were built on the foundation of display-pointing device-full keyboard.Many interface challenges arise without such a physical UI foundation.Text writingranging from entering URLs and search queries, filling forms, typing commands, to taking notes and writing emails and chat messages-is one of the hard problems awaiting for solutions in off-desktop computing.This paper summarizes and synthesizes a research program on this topic at the IBM Almaden Research Center.It analyzes various dimensions that constitute a good text input interface; briefly reviews related literature; discusses the evaluation methodology issues of text input; presents the major ideas and results of two systems, ATOMIK and SHARK; and points out current and future directions in the area from our current vantage point.
Shumin Zhai, Per Ola Kristensson, Barton A. Smith
Interact. Comput.2
2004 SHARK2: a large vocabulary shorthand writing system for pen-based computers
abstract
Zhai and Kristensson (2003) presented a method of speed-writing for pen-based computing which utilizes gesturing on a stylus keyboard for familiar words and tapping for others. In SHARK2:, we eliminated the necessity to alternate between the two modes of writing, allowing any word in a large vocabulary (e.g. 10,000-20,000 words) to be entered as a shorthand gesture. This new paradigm supports a gradual and seamless transition from visually guided tracing to recall-based gesturing. Based on the use characteristics and human performance observations, we designed and implemented the architecture, algorithms and interfaces of a high-capacity multi-channel pen-gesture recognition system. The system's key components and performance are also reported.
Per Ola Kristensson, Shumin Zhai
UIST1
2003 Shorthand writing on stylus keyboard
abstract
We propose a method for computer-based speed writing, SHARK (shorthand aided rapid keyboarding), which augments stylus keyboarding with shorthand gesturing. SHARK defines a shorthand symbol for each word according to its movement pattern on an optimized stylus keyboard. The key principles for the SHARK design include high efficiency stemmed from layout optimization, duality of gesturing and stylus tapping, scale and location independent writing, Zipf's law, and skill transfer from tapping to shorthand writing due to pattern consistency. We developed a SHARK system based on a classic handwriting recognition algorithm. A user study demonstrated the feasibility of the SHARK method.
Shumin Zhai, Per Ola Kristensson
CHI2