VLDB 2026 Research / reviewers in the wild / expert
John J. Dudley
dblp:211/7174
· DBLP profile ↗
37ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0001-6692-4853ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 23 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-AI Interaction for Time-Critical Sensemaking in Missing Persons InvestigationsabstractEvery year an estimated 200,000 people go missing in the UK alone. Missing persons investigations involve challenging time-critical sensemaking tasks based on fragmented data sources. This paper describes a mixed-methods participatory study evaluating data science and AI-driven techniques (summarisation, fact extraction, and data visualisation) for supporting these investigations as part of a human-centered workflow. A series of human-AI interfaces were iteratively designed and tested with search officers and domain experts at Police Scotland. Based on findings, we describe: (1) user and information needs for missing persons investigations; (2) insights on the benefits and challenges of applying LLM-based techniques in high-risk contexts; and (3) lessons for integrating AI for sensemaking tasks in policing more broadly. We highlight that in high-stakes contexts, where accuracy and context-sensitivity are paramount, AI techniques must be balanced with other approaches and designed in close partnership with end-users. Pola Zuzanna Labedzka, Dorian Peters, John J. Dudley, Miri Zilka |
CHI | 3 |
| 2026 | EnVisionVR: A Scene Interpretation Tool for Visual Accessibility in Virtual RealityabstractEffective visual accessibility in Virtual Reality (VR) is crucial for Blind and Low Vision (BLV) users. However, designing visual accessibility systems is challenging due to the complexity of 3D VR environments and the need for techniques that can be easily retrofitted into existing applications. While prior work has studied how to enhance or translate visual information, the advancement of Vision Language Models (VLMs) provides an exciting opportunity to advance the scene interpretation capability of current systems. This paper presents EnVisionVR, an accessibility tool for VR scene interpretation. Through a formative study of usability barriers, we confirmed the lack of visual accessibility features as a key barrier for BLV users of VR content and applications. In response, we used our findings from the formative study to inform the design and development of EnVisionVR, a novel visual accessibility system leveraging a VLM, voice input and multimodal feedback for scene interpretation and virtual object interaction in VR. An evaluation with 12 BLV users demonstrated that EnVisionVR significantly improved their ability to locate virtual objects, effectively supporting scene understanding and object interaction. Rosella P. Galindo Esparza, Vanja Garaj, Per Ola Kristensson, John J. Dudley |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Exclusion Rates among Disabled and Older Users of Virtual and Augmented RealityabstractThis paper examines the levels of exclusion encountered by disabled and older users of consumer-level VR and AR technology and identifies methods formed by people with diverse access needs to circumvent encountered barriers to use. First, we estimate exclusion rates for a selection of nine immersive experiences of VR and AR, computed using population statistics data for the United Kingdom (UK). We then present an empirical lab-based study evaluating the usability of the same VR and AR experiences. The study involved 60 UK-based participants with varying access needs and the study results were used to calculate the empirical exclusion rates. Both the estimated and empirical exclusion rates display high levels of exclusion, which for the more complex experiences in the study reached 100 %. However, multiple participants overcame usability barriers and completed experiences through provided assistance and self-initiated adaptations, suggesting that future VR and AR can become more inclusive if designed to counter these barriers. Rosella P. Galindo Esparza, John J. Dudley, Vanja Garaj, Per Ola Kristensson |
CHI | 2 |
| 2025 | Seeing and Touching the Air: Unraveling Eye-Hand Coordination in Mid-Air Gesture Typing for Mixed RealityabstractMid-air text entry in mixed reality (MR) headsets has shown promise but remains less efficient than traditional input methods. While research has focused on improving typing performance, the mechanics of mid-air gesture typing, especially eye-hand coordination, are less understood. This paper investigates visuomotor coordination of mid-air gesture keyboards through a user study (n = 16) comparing gesture typing on a tablet and in mid-air. Through an expert task we demonstrate that users were able to achieve a comparable text input performance. Our in-depth analysis of eye-hand coordination reveals significant differences in the eye-hand coordination patterns between gesture typing on a tablet and in-air. The mid-air gesture typing necessitates almost all of the visual attention on the keyboard area and a more consistent synchronization in eye-hand coordination to compensate for the increased motor and cognitive demands without physical boundaries. These insights provide important implications for the design of more efficient text input methods. Jinghui Hu, John J. Dudley, Per Ola Kristensson |
CHI | 2 |
| 2025 | ReachVox: Clutter-Free Reachability Visualization for Robot Motion Planning in Virtual RealityabstractHuman-Robot-Collaboration can enhance workflows by leveraging the mutual strengths of human operators and robots. Planning and understanding robot movements remain major challenges in this domain. This problem is prevalent in dynamic environments that might need constant robot motion path adaptation. In this paper, we investigate whether a minimalistic encoding of the reachability of a point near an object of interest, which we call ReachVox, can aid the collaboration between a remote operator and a robotic arm in VR. Through a user study ($\mathrm{n}=20$), we indicate the strength of the visualization relative to a point-based reachability check-up. Steffen Hauck, Diar Abdlkarim, John J. Dudley, Per Ola Kristensson, Eyal Ofek, Jens Grubert |
ISMAR | 3 |
| 2025 | Working in Extended Reality in the Wild: Worker and Bystander Experiences of XR Virtual Displays in Public Real-World SettingsabstractAlthough access to sufficient screen space is crucial to knowledge work, workers often find themselves with limited access to display infrastructure in remote or public settings. While virtual displays can be used to extend the available screen space through extended reality (XR) head-worn displays (HWD), we must better understand the implications of working with them in public settings from both users' and bystanders' viewpoints. To this end, we conducted two user studies. We first explored the usage of a hybrid AR display across real-world settings and tasks. We focused on how users take advantage of virtual displays and what social and environmental factors impact their usage of the system. A second study investigated the differences between working with a laptop, an AR system, or a VR system in public. We focused on a single location and participants performed a predefined task to enable direct comparisons between the conditions while also gathering data from bystanders. The combined results suggest a positive acceptance of XR technology in public settings and show that virtual displays can be used to accompany existing devices. We highlighted some environmental and social factors. We saw that previous XR experience and personality can influence how people perceive the use of XR in public. In addition, we confirmed that using XR in public still makes users stand out and that bystanders are curious about the devices, yet have no clear understanding of how they can be used. Leonardo Pavanatto, Verena Biener, Jennifer Chandran, Snehanjali Kalamkar, Feiyu Lu 0001, John J. Dudley, Jinghui Hu, Gabriella N. Ramirez, Per Ola Kristensson, Alexander Giovannelli, Luke Schlueter, Jörg Müller 0001, Jens Grubert, Doug A. Bowman |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Personalized Dual-Level Color Grading for 360-degree Images in Virtual RealityabstractThe rising popularity of 360-degree images and virtual reality (VR) has spurred a growing interest among creators in producing visually appealing content through effective color grading processes. Although existing computational approaches have simplified the global color adjustment for entire images with Preferential Bayesian Optimization (PBO), they neglect local colors for points of interest and are not optimized for the immersive nature of VR. In response, we propose a dual-level PBO framework that integrates global and local color adjustments tailored for VR environments. We design and evaluate a novel context-aware preferential Gaussian Process (GP) to learn contextual preferences for local colors, taking into account the dynamic contexts of previously established global colors. Additionally, recognizing the limitations of desktop-based interfaces for comparing 360-degree images, we design three VR interfaces for color comparison. We conduct a controlled user study to investigate the effectiveness of the three VR interface designs and find that users prefer to be enveloped by one 360-degree image at a time and to compare two rather than four color-graded options. Linping Yuan, John J. Dudley, Per Ola Kristensson, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Efficient Mid-Air Text Input Correction in Virtual RealityabstractThe task of inputting text within virtual reality has attracted significant research attention over the last five years. Less well explored is the related task of correcting inputted text when errors are made. This is despite the fact that considerable time and frustration stems from efforts to correct text. In this paper, we bridge this gap in prior research and explore efficient methods for supporting text input correction in virtual reality. We present a characterization of the types and frequencies of errors encountered when inputting text in virtual reality and an analysis of effective editing strategies. We also present the results of a user study evaluating the performance and usability trade-offs for several interaction methods leveraging the unique capabilities of modern head-mounted displays. John J. Dudley, Amy Karlson, Kashyap Todi, Hrvoje Benko, Matt Longest, Robert Wang 0002, Per Ola Kristensson |
ISMAR | 1 |
| 2024 | LookUP: Command Search Using Dwell-free Eye Typing in Mixed RealityabstractWe introduce LookUP, a novel general purpose command search system for mixed reality headsets, offering a hands-free experience through dwell-free eye typing. With LookUP, users can trigger the display of a virtual keyboard with a simple upward head motion. The keyboard then uses a statistical decoder to interpret users’ intended text based on their eye movements. This approach diverges from traditional dwell-time methods, significantly enhancing typing speed and efficiency. Our research involved deploying LookUP on a HoloLens 2, and benchmarking it against a dwell-based command search baseline and the native HoloLens system menu. Our user study indicated that participants spent a significantly shorter time using LookUP with dwell-free eye typing in command search and entry, demonstrating LookUP’s potential to be a complementary command input for mixed reality headsets. Jinghui Hu, John J. Dudley, Per Ola Kristensson |
ISMAR | 2 |
| 2024 | Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric PerceptionabstractWe depend on our own memory to encode, store, and retrieve our experiences. However, memory lapses can occur. One promising avenue for achieving memory augmentation is through the use of augmented reality head-mounted displays to capture and preserve egocentric videos, a practice commonly referred to as lifelogging. However, a significant challenge arises from the sheer volume of video data generated through lifelogging, as the current technology lacks the capability to encode and store such large amounts of data efficiently. Further, retrieving specific information from extensive video archives requires substantial computational power, further complicating the task of quickly accessing desired content. To address these challenges, we propose a memory augmentation agent that involves leveraging natural language encoding for video data and storing them in a vector database. This approach harnesses the power of large vision language models to perform the language encoding process. Additionally, we propose using large language models to facilitate natural language querying. Our agent underwent extensive evaluation using the QA-Ego4D dataset and achieved state-of-the-art results with a BLEU score of 8.3, outperforming conventional machine learning models that scored between 3.4 and 5.8. Additionally, we conducted a user study in which participants interacted with the human memory augmentation agent through episodic memory and open-ended questions. The results of this study show that the agent results in significantly better recall performance on episodic memory tasks compared to human participants. The results also highlight the agent’s practical applicability and user acceptance. Junxiao Shen, John J. Dudley, Per Ola Kristensson |
ISMAR | 2 |
| 2024 | SkiMR: Dwell-free Eye Typing in Mixed RealityabstractWe present SkiMR: a dwell-free eye typing system that enables fast and accurate hands-free text entry on mixed reality headsets. SkiMR uses a statistical decoder to infer users’ intended text based on users’ eye movements on a virtual keyboard. It does not rely on dwell timeouts for key selections, which enables it to be faster than traditional eye typing. We study this dwell-free eye typing system, deployed on a HoloLens 2, in two studies. In the first study (n = 12) we show that dwell-free eye typing results in a significantly faster text entry rate compared to traditional dwell-based eye typing with word prediction support, and a hybrid dwell-free method that uses dwell timeouts to delimit word entry. Based on the insights from the first study we evaluate the feasibility of a refined system in a more realistic composition task and use an interaction mechanism that provides real-time predictions during dwell-free eye typing. The second study (n = 16) demonstrates that this final system allow users to compose original text at 12 words per minute with a corrected character error rate of 1.1%. Overall, this work demonstrates the high potential for fast and accurate hands-free text entry using dwell-free eye typing for mixed reality headsets. Jinghui Hu, John J. Dudley, Per Ola Kristensson |
VR | 2 |
| 2024 | Swarm manipulation: An efficient and accurate technique for multi-object manipulation in virtual realityabstractThe theory of swarm control shows promise for controlling multiple objects, however, scalability is hindered by cost constraints, such as hardware and infrastructure. Virtual Reality (VR) can overcome these limitations, but research on swarm interaction in VR is limited. This paper introduces a novel swarm manipulation technique and compares it with two baseline techniques: Virtual Hand and Controller (ray-casting). We evaluated these techniques in a user study ( N = 12) in three tasks (selection, rotation, and resizing) across five conditions. Our results indicate that swarm manipulation yielded superior performance, with significantly faster speeds in most conditions across the three tasks. It notably reduced resizing size deviations but introduced a trade-off between speed and accuracy in the rotation task. Additionally, we conducted a follow-up user study ( N = 6) using swarm manipulation in two complex VR scenarios and obtained insights through semi-structured interviews, shedding light on optimized swarm control mechanisms and perceptual changes induced by this interaction paradigm. These results demonstrate the potential of the swarm manipulation technique to enhance the usability and user experience in VR compared to conventional manipulation techniques. In future studies, we aim to understand and improve swarm interaction via internal swarm particle cooperation. • We define swarm manipulation in VR and evaluate its usability and utility in two user studies. • Study 1 compares swarm manipulation to hand and controller techniques across five conditions. • Study 2 explores swarm manipulation in complex scenarios with detailed user feedback. • We analyze strengths and limitations of swarm manipulation in usability and workload. • We provide design insights for integrating swarm manipulation into VR applications. Xiang Li 0101, Jin-Du Wang, John J. Dudley, Per Ola Kristensson |
Comput. Graph. | 3 |
| 2024 | Cooperative Multi-Objective Bayesian Design OptimizationabstractComputational methods can potentially facilitate user interface design by complementing designer intuition, prior experience, and personal preference. Framing a user interface design task as a multi-objective optimization problem can help with operationalizing and structuring this process at the expense of designer agency and experience. While offering a systematic means of exploring the design space, the optimization process cannot typically leverage the designer’s expertise in quickly identifying that a given “bad” design is not worth evaluating. We here examine a cooperative approach where both the designer and optimization process share a common goal and work in partnership by establishing a shared understanding of the design space. We tackle the research question: How can we foster cooperation between the designer and a systematic optimization process in order to best leverage their combined strength? We introduce and present an evaluation of a cooperative approach that allows the user to express their design insight and work in concert with a multi-objective design process. We find that the cooperative approach successfully encourages designers to explore more widely in the design space than when they are working without assistance from an optimization process. The cooperative approach also delivers design outcomes that are comparable to an optimization process run without any direct designer input but achieves this with greater efficiency and substantially higher designer engagement levels. George B. Mo, John J. Dudley, Li-Wei Chan 0001, Yi-Chi Liao 0001, Antti Oulasvirta, Per Ola Kristensson |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2024 | Hold Tight: Identifying Behavioral Patterns During Prolonged Work in VR Through Video AnalysisabstractVR devices have recently been actively promoted as tools for knowledge workers and prior work has demonstrated that VR can support some knowledge worker tasks. However, only a few studies have explored the effects of prolonged use of VR such as a study observing 16 participants working in VR and a physical environment for one work-week each and reporting mainly on subjective feedback. As a nuanced understanding of participants' behavior in VR and how it evolves over time is still missing, we report on the results from an analysis of 559 hours of video material obtained in this prior study. Among other findings, we report that (1) the frequency of actions related to adjusting the headset reduced by 46% and the frequency of actions related to supporting the headset reduced by 42% over the five days; (2) the HMD was removed 31% less frequently over the five days but for 41% longer periods; (3) wearing an HMD is disruptive to normal patterns of eating and drinking, but not to social interactions, such as talking. The combined findings in this work demonstrate the value of long-term studies of deployed VR systems and can be used to inform the design of better, more ergonomic VR systems as tools for knowledge workers. Verena Biener, Forouzan Farzinnejad, Rinaldo Schuster, Seyedmasih Tabaei, Leon Lindlein, Jinghui Hu, Negar Nouri, John J. Dudley, Per Ola Kristensson, Jörg Müller 0001, Jens Grubert |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | Evaluating the Performance of Hand-Based Probabilistic Text Input Methods on a Mid-Air Virtual Qwerty KeyboardabstractIntegrated hand-tracking on modern virtual reality (VR) headsets can be readily exploited to deliver mid-air virtual input surfaces for text entry. These virtual input surfaces can closely replicate the experience of typing on a Qwerty keyboard on a physical touchscreen, thereby allowing users to leverage their pre-existing typing skills. However, the lack of passive haptic feedback, unconstrained user motion, and potential tracking inaccuracies or observability issues encountered in this interaction setting typically degrades the accuracy of user articulations. We present a comprehensive exploration of error-tolerant probabilistic hand-based input methods to support effective text input on a mid-air virtual Qwerty keyboard. Over three user studies we examine the performance potential of hand-based text input under both gesture and touch typing paradigms. We demonstrate typical entry rates in the range of 20 to 30 wpm and average peak entry rates of 40 to 45 wpm. John J. Dudley, Jingyao Zheng, Aakar Gupta, Hrvoje Benko, Matt Longest, Robert Wang 0002, Per Ola Kristensson |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Fast and Robust Mid-Air Gesture Typing for AR Headsets using 3D Trajectory DecodingabstractWe present a fast mid-air gesture keyboard for head-mounted optical see-through augmented reality (OST AR) that supports users in articulating word patterns by merely moving their own physical index finger in relation to a virtual keyboard plane without a need to indirectly control a visual 2D cursor on a keyboard plane. To realize this, we introduce a novel decoding method that directly translates users' three-dimensional fingertip gestural trajectories into their intended text. We evaluate the efficacy of the system in three studies that investigate various design aspects, such as immediate efficacy, accelerated learning, and whether it is possible to maintain performance without providing visual feedback. We find that the new 3D trajectory decoding design results in significant improvements in entry rates while maintaining low error rates. In addition, we demonstrate that users can maintain their performance even without fingertip and gesture trace visualization. Junxiao Shen, John J. Dudley, Per Ola Kristensson |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | HotGestures: Complementing Command Selection and Use with Delimiter-Free Gesture-Based Shortcuts in Virtual RealityabstractConventional desktop applications provide users with hotkeys as shortcuts for triggering different functionality. In this paper we consider what constitutes an effective parallel to hotkeys in a 3D interaction space where the input modality is no longer limited to the use of a keyboard. We propose HotGestures: a gesture-based interaction system for rapid tool selection and usage. Hand gestures are frequently used during human communication to convey information and provide natural associations with meaning. HotGestures provide shortcuts for users to seamlessly activate and use virtual tools by performing hand gestures. This approach naturally complements conventional menu interactions. We evaluate the potential of HotGestures in a set of two user studies and observe that our gesture-based technique provides fast and effective shortcuts for tool selection and usage. Participants found HotGestures to be distinctive, fast, and easy to use while also complementing conventional menu-based interaction. Zhaomou Song, John J. Dudley, Per Ola Kristensson |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction TechniquesabstractDesigners reportedly struggle with design optimization tasks where they are asked to find a combination of design parameters that maximizes a given set of objectives. In HCI, design optimization problems are often exceedingly complex, involving multiple objectives and expensive empirical evaluations. Model-based computational design algorithms assist designers by generating design examples during design, however they assume a model of the interaction domain. Black box methods for assistance, on the other hand, can work with any design problem. However, virtually all empirical studies of this human-in-the-loop approach have been carried out by either researchers or end-users. The question stands out if such methods can help designers in realistic tasks. In this paper, we study Bayesian optimization as an algorithmic method to guide the design optimization process. It operates by proposing to a designer which design candidate to try next, given previous observations. We report observations from a comparative study with 40 novice designers who were tasked to optimize a complex 3D touch interaction technique. The optimizer helped designers explore larger proportions of the design space and arrive at a better solution, however they reported lower agency and expressiveness. Designers guided by an optimizer reported lower mental effort but also felt less creative and less in charge of the progress. We conclude that human-in-the-loop optimization can support novice designers in cases where agency is not critical. Li-Wei Chan 0001, Yi-Chi Liao 0001, George B. Mo, John J. Dudley, Chun-Lien Cheng, Per Ola Kristensson, Antti Oulasvirta |
CHI | 4 |
| 2022 | Personalization of a Mid-Air Gesture Keyboard using Multi-Objective Bayesian OptimizationabstractWe present AdaptiKeyboard, a mid-air gesture keyboard that uses multi-objective Bayesian optimization to adaptively change layout size to simultaneously optimize speed and accuracy. Gesture keyboards are well suited for enabling mid-air text entry in augmented reality (AR) due to their relative robustness to articulation inaccuracy. However, transplanting gesture keyboards to AR involves a larger design and operational space compared to touchscreen interactions. One potential advantage of this larger design and operational space is that mid-air keyboards presented in AR can be more versatile than their touchscreen equivalents. A key component of a mid-air gesture keyboard is the layout size, which can be made adaptive in order to optimize text entry speed and accuracy at the individual user level. This adaptive personalization can refine the keyboard design to reflect the differences users exhibit in motor behaviors and personal preferences. In this paper, we propose a multi-objective Bayesian optimization approach for adapting the layout size of a mid-air gesture keyboard to individual users. We show that this process can deliver a 14.4% improvement in speed and a 13.8% improvement in accuracy relative to a baseline design with a constant size derived from the default system keyboard on the HoloLens 2. Junxiao Shen, Jinghui Hu, John J. Dudley, Per Ola Kristensson |
ISMAR | 3 |
| 2022 | Efficient Special Character Entry on a Virtual Keyboard by Hand Gesture-Based Mode SwitchingabstractThe need to support efficient text input in virtual reality continues to attract significant research attention. However, much of this research understandably focuses exclusively on core text input tasks involving the entry of standard alphabetic characters. The less common though still critical task of entering special characters is often ignored. In this paper we focus on this niche use case as chiefly encountered when entering passwords. Current commercial virtual keyboards allow users to switch between different layers of the keyboard in order to access capital letters, numerals and special characters by pressing an explicit mode-switch button. We propose a new method of switching between layers of a virtual keyboard using hand gestures. Critically, these hand gestures are seamlessly performed in conjunction with key selections to deliver an efficient and intuitive interaction. We report on a user study with 16 participants entering standard passwords comparing our gesture-based mode switching approach to a conventional button-based baseline. We find that with practice our proposed method results in significantly faster entry rates without any deterioration in accuracy. Feedback from users also indicates that our technique is considered efficient and comfortable to use. Zhaomou Song, John J. Dudley, Per Ola Kristensson |
ISMAR | 2 |
| 2022 | KWickChat: A Multi-Turn Dialogue System for AAC Using Context-Aware Sentence Generation by Bag-of-KeywordsabstractWe present KWickChat (Keyword Quick Chat): a multi-turn augmentative and alternative communication (AAC) dialogue system for nonspeaking individuals with motor disabilities. The central objective of KWickChat is to reduce the communication gap between nonspeaking and speaking partners by exploring a sentence-based text entry system that automatically generates suitable sentences for the nonspeaking partner based on keyword entry. The system is underpinned by a GPT-2 language model and leverages context information, including dialogue history and persona tags, to improve the quality of the generated responses. We evaluate the system by analyzing the functional design and decomposing it into key functions and parameters that are systematically investigated using envelope analysis. We pursue this methodology as a necessary precursor to evaluation with AAC users. Our results show that with word prediction and with a threshold word error rate of 0.65, the keystroke savings of the KWickChat system is around 71%. To complement the envelope analysis, we also recruited two human judges to evaluate the semantic consistency between 400 sentences generated by KWickChat and reference sentences. Both judges reported a median rating of 4 on a scale from 1 (very bad) to 5 (very good) for the best generated sentence in each exchange and achieved an inter-rater reliability of 0.92 across all 400 sentences judged. Junxiao Shen, Boyin Yang, John J. Dudley, Per Ola Kristensson |
IUI | 3 |
| 2022 | Understanding user performance of acquiring targets with motion-in-depth in virtual reality
Jin Huang 0009, John J. Dudley, Stephen Uzor, Per Ola Kristensson, Feng Tian 0001 |
Int. J. Hum. Comput. Stud. | 2 |
| 2022 | Quantifying the Effects of Working in VR for One WeekabstractVirtual Reality (VR) provides new possibilities for modern knowledge work. However, the potential advantages of virtual work environments can only be used if it is feasible to work in them for an extended period of time. Until now, there are limited studies of long-term effects when working in VR. This paper addresses the need for understanding such long-term effects. Specifically, we report on a comparative study $i$, in which participants were working in VR for an entire week-for five days, eight hours each day-as well as in a baseline physical desktop environment. This study aims to quantify the effects of exchanging a desktop-based work environment with a VR-based environment. Hence, during this study, we do not present the participants with the best possible VR system but rather a setup delivering a comparable experience to working in the physical desktop environment. The study reveals that, as expected, VR results in significantly worse ratings across most measures. Among other results, we found concerning levels of simulator sickness, below average usability ratings and two participants dropped out on the first day using VR, due to migraine, nausea and anxiety. Nevertheless, there is some indication that participants gradually overcame negative first impressions and initial discomfort. Overall, this study helps lay the groundwork for subsequent research, by clearly highlighting current shortcomings and identifying opportunities for improving the experience of working in VR. Verena Biener, Snehanjali Kalamkar, Negar Nouri, Eyal Ofek, Michel Pahud, John J. Dudley, Jinghui Hu, Per Ola Kristensson, Maheshya Weerasinghe, Klen Copic Pucihar, Matjaz Kljun, Stephan Streuber, Jens Grubert |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Gesture Spotter: A Rapid Prototyping Tool for Key Gesture Spotting in Virtual and Augmented Reality ApplicationsabstractIn this paper we examine the task of key gesture spotting: accurate and timely online recognition of hand gestures. We specifically seek to address two key challenges faced by developers when integrating key gesture spotting functionality into their applications. These are: i) achieving high accuracy and zero or negative activation lag with single-time activation; and ii) avoiding the requirement for deep domain expertise in machine learning. We address the first challenge by proposing a key gesture spotting architecture consisting of a novel gesture classifier model and a novel single-time activation algorithm. This key gesture spotting architecture was evaluated on four separate hand skeleton gesture datasets, and achieved high recognition accuracy with early detection. We address the second challenge by encapsulating different data processing and augmentation strategies, as well as the proposed key gesture spotting architecture, into a graphical user interface and an application programming interface. Two user studies demonstrate that developers are able to efficiently construct custom recognizers using both the graphical user interface and the application programming interface. Junxiao Shen, John J. Dudley, George B. Mo, Per Ola Kristensson |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Crowdsourcing Design Guidance for Contextual Adaptation of Text Content in Augmented RealityabstractAugmented Reality (AR) can deliver engaging user experiences that seamlessly meld virtual content with the physical environment. However, building such experiences is challenging due to the developer’s inability to assess how uncontrolled deployment contexts may influence the user experience. To address this issue, we demonstrate a method for rapidly conducting AR experiments and real-world data collection in the user’s own physical environment using a privacy-conscious mobile web application. The approach leverages the large number of distinct user contexts accessible through crowdsourcing to efficiently source diverse context and perceptual preference data. The insights gathered through this method complement emerging design guidance and sample-limited lab-based studies. The utility of the method is illustrated by re-examining the design challenge of adapting AR text content to the user’s environment. Finally, we demonstrate how gathered design insight can be operationalized to provide adaptive text content functionality in an AR headset. John J. Dudley, Jason T. Jacques, Per Ola Kristensson |
CHI | 1 |
| 2021 | Understanding, Detecting and Mitigating the Effects of Coactivations in Ten-Finger Mid-Air Typing in Virtual RealityabstractTyping with ten fingers on a virtual keyboard in virtual or augmented reality exposes a challenging input interpretation problem. There are many sources of noise in this interaction context and these exacerbate the challenge of accurately translating human actions into text. A particularly challenging input noise source arises from the physiology of the hand. Intentional finger movements can produce unintentional coactivations in other fingers. On a physical keyboard, the resistance of the keys alleviates this issue. On a virtual keyboard, coactivations are likely to introduce spurious input events under a naïve solution to input detection. In this paper we examine the features that discriminate intentional activations from coactivations. Based on this analysis, we demonstrate three alternative coactivation detection strategies with high discrimination power. Finally, we integrate coactivation detection into a probabilistic decoder and demonstrate its ability to further reduce uncorrected character error rates by approximately 10% relative and 0.9% absolute. Conor R. Foy, John J. Dudley, Aakar Gupta, Hrvoje Benko, Per Ola Kristensson |
CHI | 2 |
| 2021 | Gesture Knitter: A Hand Gesture Design Tool for Head-Mounted Mixed Reality ApplicationsabstractHand gestures are a natural and expressive input method enabled by modern mixed reality headsets. However, it remains challenging for developers to create custom gestures for their applications. Conventional strategies to bespoke gesture recognition involve either hand-crafting or data-intensive deep-learning. Neither approach is well suited for rapid prototyping of new interactions. This paper introduces a flexible and efficient alternative approach for constructing hand gestures. We present Gesture Knitter: a design tool for creating custom gesture recognizers with minimal training data. Gesture Knitter allows the specification of gesture primitives that can then be combined to create more complex gestures using a visual declarative script. Designers can build custom recognizers by declaring them from scratch or by providing a demonstration that is automatically decoded into its primitive components. Our developer study shows that Gesture Knitter achieves high recognition accuracy despite minimal training data and delivers an expressive and creative design experience. George B. Mo, John J. Dudley, Per Ola Kristensson |
CHI | 2 |
| 2021 | Investigating the Accessibility of Crowdwork Tasks on Mechanical TurkabstractCrowdwork can enable invaluable opportunities for people with disabilities, not least the work flexibility and the ability to work from home, especially during the current Covid-19 pandemic. This paper investigates how engagement in crowdwork tasks is affected by individual disabilities and the resulting implications for HCI. We first surveyed 1,000 Amazon Mechanical Turk (AMT) workers to identify demographics of crowdworkers who identify as having various disabilities within the AMT ecosystem—including vision, hearing, cognition/mental, mobility, reading and motor impairments. Through a second focused survey and follow-up interviews, we provide insights into how respondents cope with crowdwork tasks. We found that standard task factors, such as task completion time and presentation, often do not account for the needs of users with disabilities, resulting in anxiety and a feeling of depression on occasion. We discuss how to alleviate barriers to enable effective interaction for crowdworkers with disabilities. Stephen Uzor, Jason T. Jacques, John J. Dudley, Per Ola Kristensson |
CHI | 3 |
| 2021 | The Imaginative Generative Adversarial Network: Automatic Data Augmentation for Dynamic Skeleton-Based Hand Gesture and Human Action RecognitionabstractDeep learning approaches deliver state-of-the-art performance in recognition of spatiotemporal human motion data. However, one of the main challenges in these recognition tasks is limited available training data. Insufficient training data results in over-fitting and data augmentation is one approach to address this challenge. Existing data augmentation strategies based on scaling, shifting and interpolating offer limited generalizability and typically require detailed inspection of the dataset as well as hundreds of GPU hours for hyperparameter optimization. In this paper, we present a novel automatic data augmentation model, the Imaginative Generative Adversarial Network (GAN), that approximates the distribution of the input data and samples new data from this distribution. It is automatic in that it requires no data inspection and little hyperparameter tuning and therefore it is a low-cost and low-effort approach to generate synthetic data. We demonstrate our approach on small-scale skeleton-based datasets with a comprehensive experimental analysis. Our results show that the augmentation strategy is fast to train and can improve classification accuracy for both conventional neural networks and state-of-the-art methods. Junxiao Shen, John J. Dudley, Per Ola Kristensson |
FG | 2 |
| 2021 | Simulating Realistic Human Motion Trajectories of Mid-Air Gesture TypingabstractThe eventual success of many AR and VR intelligent interactive systems relies on the ability to collect user motion data at large scale. Realistic simulation of human motion trajectories is a potential solution to this problem. Simulated user motion data can facilitate prototyping and speed up the design process. There are also potential benefits in augmenting training data for deep learning-based AR/VR applications to improve performance. However, the generation of realistic motion data is nontrivial. In this paper, we examine the specific challenge of simulating index finger movement data to inform mid-air gesture keyboard design. The mid-air gesture keyboard is deployed on an optical see-through display that allows the user to enter text by articulating word gesture patterns with their physical index finger in the vicinity of a visualized keyboard layout. We propose and compare four different approaches to simulating this type of motion data, including a Jerk-Minimization model, a Recurrent Neural Network (RNN)-based generative model, and a Generative Adversarial Network (GAN)-based model with two modes: style transfer and data alteration. We also introduce a procedure for validating the quality of the generated trajectories in terms of realism and diversity. The GAN-based model shows significant potential for generating synthetic motion trajectories to facilitate design and deep learning for advanced gesture keyboards deployed in AR and VR. Junxiao Shen, John J. Dudley, Per Ola Kristensson |
ISMAR | 2 |
| 2021 | Complex Interaction as Emergent Behaviour: Simulating Mid-Air Virtual Keyboard Typing using Reinforcement LearningabstractAccurately modelling user behaviour has the potential to significantly improve the quality of human-computer interaction. Traditionally, these models are carefully hand-crafted to approximate specific aspects of well-documented user behaviour. This limits their availability in virtual and augmented reality where user behaviour is often not yet well understood. Recent efforts have demonstrated that reinforcement learning can approximate human behaviour during simple goal-oriented reaching tasks. We build on these efforts and demonstrate that reinforcement learning can also approximate user behaviour in a complex mid-air interaction task: typing on a virtual keyboard. We present the first reinforcement learning-based user model for mid-air and surface-aligned typing on a virtual keyboard. Our model is shown to replicate high-level human typing behaviour. We demonstrate that this approach may be used to augment or replace human testing during the validation and development of virtual keyboards. Lorenz Hetzel, John J. Dudley, Anna Maria Feit, Per Ola Kristensson |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Crowdsourcing Interface Feature Design with Bayesian OptimizationabstractDesigning novel interfaces is challenging. Designers typically rely on experience or subjective judgment in the absence of analytical or objective means for selecting interface parameters. We demonstrate Bayesian optimization as an efficient tool for objective interface feature refinement. Specifically, we show that crowdsourcing paired with Bayesian optimization can rapidly and effectively assist interface design across diverse deployment environments. Experiment 1 evaluates the approach on a familiar 2D interface design problem: a map search and review use case. Adding a degree of complexity, Experiment 2 extends Experiment 1 by switching the deployment environment to mobile-based virtual reality. The approach is then demonstrated as a case study for a fundamentally new and unfamiliar interaction design problem: web-based augmented reality. Finally, we show how the model generated as an outcome of the refinement process can be used for user simulation and queried to deliver various design insights. John J. Dudley, Jason T. Jacques, Per Ola Kristensson |
CHI | 1 |
| 2019 | Performance Envelopes of Virtual Keyboard Text Input Strategies in Virtual RealityabstractVirtual and Augmented Reality deliver engaging interaction experiences that can transport and extend the capabilities of the user. To ensure these paradigms are more broadly usable and effective, however, it is necessary to also deliver many of the conventional functions of a smartphone or personal computer. It remains unclear how conventional input tasks, such as text entry, can best be translated into virtual and augmented reality. In this paper we examine the performance potential of four alternative text entry strategies in virtual reality (VR). These four strategies are selected to provide full coverage of two fundamental design dimensions: i) physical surface association; and ii) number of engaged fingers. Specifically, we examine typing with index fingers on a surface and in mid-air and typing using all ten fingers on a surface and in mid-air. The central objective is to evaluate the human performance potential of these four typing strategies without being constrained by current tracking and statistical text decoding limitations. To this end we introduce an auto-correction simulator that uses knowledge of the stimulus to emulate statistical text decoding within constrained experimental parameters and use high-precision motion tracking hardware to visualise and detect fingertip interactions. We find that alignment of the virtual keyboard with a physical surface delivers significantly faster entry rates over a mid-air keyboard. Also, users overwhelmingly fail to effectively engage all ten fingers in mid-air typing, resulting in slower entry rates and higher error rates compared to just using two index fingers. In addition to identifying the envelopes of human performance for the four strategies investigated, we also provide a detailed analysis of the underlying features that distinguish each strategy in terms of its performance and behaviour. John J. Dudley, Hrvoje Benko, Daniel J. Wigdor, Per Ola Kristensson |
ISMAR | 1 |
| 2018 | Bare-Handed 3D Drawing in Augmented RealityabstractHead-mounted augmented reality (AR) enables embodied in situ drawing in three dimensions (3D). We explore 3D drawing interactions based on uninstrumented, unencumbered (bare) hands that preserve the user's ability to freely navigate and interact with the physical environment. We derive three alternative interaction techniques supporting bare-handed drawing in AR from the literature and by analysing several envisaged use cases. The three interaction techniques are evaluated in a controlled user study examining three distinct drawing tasks: planar drawing, path description, and 3D object reconstruction. The results indicate that continuous freehand drawing supports faster line creation than the control point based alternatives, although with reduced accuracy. User preferences for the different techniques are mixed and vary considerably between the different tasks, highlighting the value of diverse and flexible interactions. The combined effectiveness of these three drawing techniques is illustrated in an example application of 3D AR drawing. John J. Dudley, Hendrik Schuff, Per Ola Kristensson |
Conference on Designing Interactive Systems | 1 |
| 2018 | Performance Envelopes of in-Air Direct and Smartwatch Indirect Control for Head-Mounted Augmented RealityabstractThe scarcity of established input methods for augmented reality (AR) head-mounted displays (HMD) motivates us to investigate the performance envelopes of two easily realisable solutions: indirect cursor control via a smartwatch and direct control by in-air touch. Indirect cursor control via a smartwatch has not been previously investigated for AR HMDs. We evaluate these two techniques for carrying out three fundamental user interface actions: target acquisition, goal crossing, and circular steering. We find that in-air is faster than smartwatch (p <; 0.001) for target acquisition and circular steering. We observe, however, that in-air selection can lead to discomfort after extended use and suggest that smartwatch control offers a complementary alternative. Dennis Wolf 0002, John J. Dudley, Per Ola Kristensson |
VR | 2 |
| 2018 | A Review of User Interface Design for Interactive Machine LearningabstractInteractive Machine Learning (IML) seeks to complement human perception and intelligence by tightly integrating these strengths with the computational power and speed of computers. The interactive process is designed to involve input from the user but does not require the background knowledge or experience that might be necessary to work with more traditional machine learning techniques. Under the IML process, non-experts can apply their domain knowledge and insight over otherwise unwieldy datasets to find patterns of interest or develop complex data-driven applications. This process is co-adaptive in nature and relies on careful management of the interaction between human and machine. User interface design is fundamental to the success of this approach, yet there is a lack of consolidated principles on how such an interface should be implemented. This article presents a detailed review and characterisation of Interactive Machine Learning from an interactive systems perspective. We propose and describe a structural and behavioural model of a generalised IML system and identify solution principles for building effective interfaces for IML. Where possible, these emergent solution principles are contextualised by reference to the broader human-computer interaction literature. Finally, we identify strands of user interface research key to unlocking more efficient and productive non-expert interactive machine learning applications. John J. Dudley, Per Ola Kristensson |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2018 | Fast and Precise Touch-Based Text Entry for Head-Mounted Augmented Reality with Variable OcclusionabstractWe present the VISAR keyboard: An augmented reality (AR) head-mounted display (HMD) system that supports text entry via a virtualised input surface. Users select keys on the virtual keyboard by imitating the process of single-hand typing on a physical touchscreen display. Our system uses a statistical decoder to infer users’ intended text and to provide error-tolerant predictions. There is also a high-precision fall-back mechanism to support users in indicating which keys should be unmodified by the auto-correction process. A unique advantage of leveraging the well-established touch input paradigm is that our system enables text entry with minimal visual clutter on the see-through display, thus preserving the user’s field-of-view. We iteratively designed and evaluated our system and show that the final iteration of the system supports a mean entry rate of 17.75wpm with a mean character error rate less than 1%. This performance represents a 19.6% improvement relative to the state-of-the-art baseline investigated: A gaze-then-gesture text entry technique derived from the system keyboard on the Microsoft HoloLens. Finally, we validate that the system is effective in supporting text entry in a fully mobile usage scenario likely to be encountered in industrial applications of AR HMDs. John J. Dudley, Keith Vertanen, Per Ola Kristensson |
ACM Trans. Comput. Hum. Interact. | 1 |