VLDB 2026 Research / reviewers in the wild / expert
Amy Karlson
dblp:277/5645
· DBLP profile ↗
10ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0001-8934-7761ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Viago: Exploring Visual-Audio Modality Transitions for Social Media Consumption on the Go
Ruei-Che Chang, Tovi Grossman, Carine Rognon, Michael Glueck, Christopher Collins 0001, Amy Karlson, Hemant Bhaskar Surale |
UIST | 6 |
| 2024 | Boosting Gesture Recognition with an Automatic Gesture Annotation FrameworkabstractTraining a real-time gesture recognition model heavily relies on annotated data. However, manual data annotation is costly and demands substantial human effort. In order to address this challenge, we propose a framework that can automatically annotate gesture classes and identify their temporal ranges. Our framework consists of two key components: (1) a novel annotation model that leverages the Connectionist Temporal Classification (CTC) loss, and (2) a semi-supervised learning pipeline that enables the model to improve its performance by training on its own predictions, known as pseudo labels. These high-quality pseudo labels can also be used to enhance the accuracy of other downstream gesture recognition models. To evaluate our framework, we conducted experiments using two publicly available gesture datasets. Our ablation study demonstrates that our annotation model design surpasses the baseline in terms of both gesture classification accuracy (3–4 % improvement) and localization accuracy (71-75% improvement). Additionally, we illustrate that the pseudo-labeled dataset produced from the proposed framework significantly boosts the accuracy of a pre-trained downstream gesture recognition model by 11-18%. We believe that this annotation framework has immense potential to improve the training of downstream gesture recognition models using unlabeled datasets. Junxiao Shen, Xuhai Xu, Ran Tan, Amy Karlson, Evan Strasnick |
FG | 4 |
| 2024 | Efficient Mid-Air Text Input Correction in Virtual RealityabstractThe task of inputting text within virtual reality has attracted significant research attention over the last five years. Less well explored is the related task of correcting inputted text when errors are made. This is despite the fact that considerable time and frustration stems from efforts to correct text. In this paper, we bridge this gap in prior research and explore efficient methods for supporting text input correction in virtual reality. We present a characterization of the types and frequencies of errors encountered when inputting text in virtual reality and an analysis of effective editing strategies. We also present the results of a user study evaluating the performance and usability trade-offs for several interaction methods leveraging the unique capabilities of modern head-mounted displays. John J. Dudley, Amy Karlson, Kashyap Todi, Hrvoje Benko, Matt Longest, Robert Wang 0002, Per Ola Kristensson |
ISMAR | 2 |
| 2024 | Towards Open-World Gesture RecognitionabstractProviding users with accurate gestural interfaces, such as gesture recognition based on wrist-worn devices, is a key challenge in mixed reality. However, static machine learning processes in gesture recognition assume that training and test data come from the same underlying distribution. Unfortunately, in real-world applications involving gesture recognition, such as gesture recognition based on wrist-worn devices, the data distribution may change over time. We formulate this problem of adapting recognition models to new tasks, where new data patterns emerge, as open-world gesture recognition (OWGR). We propose the use of continual learning to enable machine learning models to be adaptive to new tasks without degrading performance on previously learned tasks. However, the process of exploring parameters for questions around when, and how, to train and deploy recognition models requires resource-intensive user studies may be impractical. To address this challenge, we propose a design engineering approach that enables offline analysis on a collected large-scale dataset by systematically examining various parameters and comparing different continual learning methods. Finally, we provide design guidelines to enhance the development of an open-world wrist-worn gesture recognition process. Junxiao Shen, Matthias De Lange, Xuhai Xu, Enmin Zhou, Ran Tan, Naveen Suda, Maciej Lazarewicz, Per Ola Kristensson, Amy Karlson, Evan Strasnick |
ISMAR | 9 |
| 2024 | RingGesture: A Ring-Based Mid-Air Gesture Typing System Powered by a Deep-Learning Word Prediction FrameworkabstractText entry is a critical capability for any modern computing experience, with lightweight augmented reality (AR) glasses being no exception. Designed for all-day wearability, a limitation of lightweight AR glass is the restriction to the inclusion of multiple cameras for extensive field of view in hand tracking. This constraint underscores the need for an additional input device. We propose a system to address this gap: a ring-based mid-air gesture typing technique, RingGesture, utilizing electrodes to mark the start and end of gesture trajectories and inertial measurement units (IMU) sensors for hand tracking. This method offers an intuitive experience similar to raycast-based mid-air gesture typing found in VR headsets, allowing for a seamless translation of hand movements into cursor navigation. To enhance both accuracy and input speed, we propose a novel deep-learning word prediction framework, Score Fusion, comprised of three key components: a) a word-gesture decoding model, b) a spatial spelling correction model, and c) a lightweight contextual language model. In contrast, this framework fuses the scores from the three models to predict the most likely words with higher precision. We conduct comparative and longitudinal studies to demonstrate two key findings: firstly, the overall effectiveness of RingGesture, which achieves an average text entry speed of 27.3 words per minute (WPM) and a peak performance of 47.9 WPM. Secondly, we highlight the superior performance of the Score Fusion framework, which offers a 28.2% improvement in uncorrected Character Error Rate over a conventional word prediction framework, Naive Correction, leading to a 55.2% improvement in text entry speed for RingGesture. Additionally, RingGesture received a System Usability Score of 83 signifying its excellent usability. Junxiao Shen, Roger Boldu, Arpit Kalla, Michael Glueck, Hemant Bhaskar Surale, Amy Karlson |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-TrainingabstractText entry with word-gesture keyboards (WGK) is emerging as a popular method and becoming a key interaction for Extended Reality (XR). However, the diversity of interaction modes, keyboard sizes, and visual feedback in these environments introduces divergent word-gesture trajectory data patterns, thus leading to complexity in decoding trajectories into text. Template-matching decoding methods, such as SHARK2 [32], are commonly used for these WGK systems because they are easy to implement and configure. However, these methods are susceptible to decoding inaccuracies for noisy trajectories. While conventional neural-network-based decoders (neural decoders) trained on word-gesture trajectory data have been proposed to improve accuracy, they have their own limitations: they require extensive data for training and deep-learning expertise for implementation. To address these challenges, we propose a novel solution that combines ease of implementation with high decoding accuracy: a generalizable neural decoder enabled by pre-training on large-scale coarsely discretized word-gesture trajectories. This approach produces a ready-to-use WGK decoder that is generalizable across mid-air and on-surface WGK systems in augmented reality (AR) and virtual reality (VR), which is evident by a robust average Top-4 accuracy of 90.4% on four diverse datasets. It significantly outperforms SHARK2 with a 37.2% enhancement and surpasses the conventional neural decoder by 7.4%. Moreover, the Pre-trained Neural Decoder's size is only 4 MB after quantization, without sacrificing accuracy, and it can operate in real-time, executing in just 97 milliseconds on Quest 3. Junxiao Shen, Khadija Khaldi, Enmin Zhou, Hemant Bhaskar Surale, Amy Karlson |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Investigating Wrist Deflection Scrolling Techniques for Extended RealityabstractScrolling in extended reality (XR) is currently performed using handheld controllers or vision-based arm-in-front gestures, which have the limitations of encumbering the user’s hands or requiring a specific arm posture, respectively. To address these limitations, we investigate freehand, posture-independent scrolling driven by wrist deflection. We propose two novel techniques: Wrist Joystick, which uses rate control, and Wrist Drag, which uses position control. In an empirical study of a rapid item acquisition task and a casual browsing task, both Wrist Drag and Wrist Joystick performed on par with a comparable state-of-the-art technique on one of the two tasks. Further, using a relaxed arm-at-side posture, participants retained their arm-in-front performance for both wrist techniques. Finally, we analyze behavioral and ergonomic data to provide design insights for wrist deflection scrolling. Our results demonstrate that wrist deflection provides a promising method for performant scrolling controls while offering additional benefits over existing XR interaction techniques. Jacqui Fashimpaur, Amy Karlson, Tanya R. Jonker, Hrvoje Benko, Aakar Gupta |
CHI | 2 |
| 2023 | How Visualising Emotions Affects Interpersonal Trust and Task Collaboration in a Shared Virtual SpaceabstractEmotion is dynamic. Changes in emotion can be hard to process during face-to-face interaction, yet transferring them into a shared virtual space becomes more challenging. This research first explores nine visual representations to amplify emotions in a virtual space, leading to a bi-directional emotion-sharing system (FeelMoji i/o). The second study investigates the effect of explicit emotion-sharing in interpersonal trust and task collaboration through three conditions - verbal only, verbal+positive visual, and verbal+honest visual using FeelMoji through the proposal of a framework of four factors (usability, integrity, behaviour, and collaboration). The results indicate that FeelMoji yields frequent emotion consensus as task milestones and positive interdependent behaviours between collaborators, which help develop conversations, affirm decision-making, and build familiarity and trust between strangers. Moreover, we discuss how our study can inspire future investigation in human-AI agent behaviours and large-scale multi-user virtual environments. Allison Jing, Michael Frederick, Monica Sewell, Amy Karlson, Brian Simpson, Missie Smith |
ISMAR | 4 |
| 2023 | STAR: Smartphone-analogous Typing in Augmented RealityabstractWhile text entry is an essential and frequent task in Augmented Reality (AR) applications, devising an efficient and easy-to-use text entry method for AR remains an open challenge. This research presents STAR, a smartphone-analogous AR text entry technique that leverages a user’s familiarity with smartphone two-thumb typing. With STAR, a user performs thumb typing on a virtual QWERTY keyboard that is overlain on the skin of their hands. During an evaluation study of STAR, participants achieved a mean typing speed of 21.9 WPM (i.e., 56% of their smartphone typing speed), and a mean error rate of 0.3% after 30 minutes of practice. We further analyze the major factors implicated in the performance gap between STAR and smartphone typing, and discuss ways this gap could be narrowed. Taejun Kim, Amy Karlson, Aakar Gupta, Tovi Grossman, Jason Wu 0001, Parastoo Abtahi, Christopher Collins 0001, Michael Glueck, Hemant Bhaskar Surale |
UIST | 2 |
| 2020 | Understanding In-Situ Use of Commonly Available Navigation Technologies by People with Visual ImpairmentsabstractDespite the large body of work in accessibility concerning the design of novel navigation technologies, little is known about commonly available technologies that people with visual impairments currently use for navigation. We address this gap with a qualitative study consisting of interviews with 23 people with visual impairments, ten of whom also participated in a follow-up diary study. We develop the idea of complementarity first introduced by Williams et al. [53] and find that in addition to using apps to complement mobility aids, technologies and apps complemented each other and filled in for the gaps inherent in one another. Furthermore, the complementarity between apps and other apps/aids was primarily the result of the differences in information and modalities in which this information is communicated by apps, technology and mobility aids. We propose design recommendations to enhance this complementarity and guide the development of improved navigation experiences for people with visual impairments. Vaishnav Kameswaran, Alexander Fiannaca, Melanie Kneitmix, Amy Karlson, Edward Cutrell, Meredith Ringel Morris |
ASSETS | 4 |