VLDB 2026 Research / reviewers in the wild / expert
Stephanie Santosa
dblp:58/2593
· DBLP profile ↗
16ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-6010-005XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 14 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating Aggregated vs. Sequential Command Recommendation in Graphical User InterfacesabstractAdvances in artificial intelligence open the possibility of predicting and recommending sequences of GUI commands to a user. An interesting question raised by this capability is how to present such recommendations to the user – as a sequential set of individual command recommendations, or as one aggregated recommendation consisting of multiple commands. In this paper we propose an interface for aggregated command recommendation and conduct controlled studies to compare sequential versus aggregated command recommendation across a range of simulated utility conditions. Our results indicate that aggregated command recommendation can improve overall task performance over sequential recommendation, and that this benefit comes from enabling users to rapidly recognize and use high-utility aggregated recommendations. The aggregated command recommendation approach also reduced deliberation time when evaluating and correcting imperfect sets of recommended commands. Benjamin J. Lafreniere, Zachary J. Davis 0003, Michelle Li, Junmeng Andrew Han, Tovi Grossman, Stephanie Santosa, Daniel J. Wigdor |
Graphics Interface | 6 |
| 2025 | An Investigation of Multimodal Kinematic Template Matching for Ray Pointing Prediction for Target Selection in VRabstractWe explore the use of multimodal input to predict the landing position of a ray pointer while selecting targets in a virtual reality (VR) environment. We first extend a prior 2D Kinematic Template Matching technique to include head movements. This new technique, Head-Coupled Kinematic Template Matching, was found to improve upon the existing 2D approach, with an angular error of 10.0° when a user was 40% of the way through their movement. We then investigate two additional models that incorporated eye gaze, which were both found to further improve the predicted landing positions. The first model, Gaze-Coupled Kinematic Template Matching resulted in angular error of 6.8° for reciprocal target layouts and 9.1° for random target layouts, when a user was 40% of the way through their movement. The second model, Hybrid Kinematic Template Matching, resulted in angular error of 5.2° for reciprocal target layouts and 7.2° for random target layouts when a user was 40% of the way through their movement. We also found that using just the current gaze location resulted in sufficient predictions in many conditions. We reflect on our results by discussing the broader implications of utilizing multimodal input to inform selection predictions in VR. Marcello Giordano, Tovi Grossman, Aakar Gupta, Rorik Henrikson, Sean Trowbridge, Stephanie Santosa, Michael Glueck, Tanya R. Jonker, Hrvoje Benko, Daniel J. Wigdor |
ACM Trans. Comput. Hum. Interact. | 7 |
| 2024 | Fidgets: Building Blocks for a Predictive UI ToolkitabstractThe rapid growth of AR platforms, combined with the rising predictive power of intelligent systems, will fundamentally change interactive computing. Interaction will increasingly happen on the go, causing I/O to become constrained, ultimately leading to reliance on user intent prediction for aid. In this pictorial, we argue that to support the development of such systems, new predictive UI toolkits are required. We place the reader in the shoes of an App designer and outline the challenges that will be faced. We then describe a new predictive toolkit, leveraging Fuzzy Widgets, or “Fidgets” as the main UI building block. Fidgets extend Responsive Design into the realm of intelligent systems, to adapt not only to spatial constraints, but to system predictions as well. We then describe a working implementation of a predictive music application, built using our described framework, showcasing its benefits and range of adaptive abilities. Joannes Chan, Chris De Paoli, Michelle Li, Tovi Grossman, Stephanie Santosa, Daniel J. Wigdor, Michael Glueck |
Conference on Designing Interactive Systems | 5 |
| 2024 | GraspUI: Seamlessly Integrating Object-Centric Gestures within the Seven Phases of GraspingabstractObjects are indispensable tools in our daily lives. Recent research has demonstrated their potential to act as conduits for digital interactions with microgestures, however, the primary focus was on situations where the hand firmly grasps an object. We introduce GraspUI, an exploratory design space of object-centric gestures within the seven distinct phases of the grasping process, spanning pre-, during, and post-grasp movements. We conducted ideation sessions with mixed-reality designers from industry and academia to explore gesture integration throughout the entire grasping process. The outcome was 38 storyboards envisioning practical applications. To evaluate the design space’s utility, we performed a video-based assessment with end-users. We then implemented an interactive prototype and quantified the overhead cost of performing proposed gestures through a secondary study. Participants reacted positively to gestures and could integrate them into existing usage of objects. To conclude, we highlight technical and usability guidelines for implementing and extending GraspUI systems. Adwait Sharma, Alexander Ivanov 0004, Frances Lai, Tovi Grossman, Stephanie Santosa |
Conference on Designing Interactive Systems | 5 |
| 2024 | Body Language for VUIs: Exploring Gestures to Enhance Interactions with Voice User InterfacesabstractWith the progress in Large Language Models (LLMs) and rapid development of wearable smart devices like smart glasses, there is a growing opportunity for users to interact with on-device virtual assistants through voice and gestures with ease. Although voice user interfaces (VUIs) have been widely studied, the potential uses of full-body gestures in VUIs that can fully understand users’ surroundings and gestures are relatively unexplored. In this two-phase research using a Wizard-of-Oz approach, we aim to investigate the role of gestures in VUI interactions and explore their design space. In an initial exploratory user study with six participants, we identify influential factors for VUI gestures and establish an initial design space. In the second phase, we conducted a user study with 12 participants to validate and refine our initial findings. Our results showed that users are open and ready to adopt and utilize gestures to interact with multi-modal VUIs, especially in scenarios with poor voice capture quality. The study also highlighted three key categories of gesture functions for enhancing multi-modal VUI interactions: context reference, alternative input, and flow control. Finally, we present a design space for multi-modal VUI gestures along with demonstrations to enlighten future design for coupling multi-modal VUIs with gestures. Liwei Wu 0002, Benjamin J. Lafreniere, Tovi Grossman, Thomas White, Stephanie Santosa |
Conference on Designing Interactive Systems | 5 |
| 2024 | OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMsabstractThe progression to “Pervasive Augmented Reality” envisions easy access to multimodal information continuously. However, in many everyday scenarios, users are occupied physically, cognitively or socially. This may increase the friction to act upon the multimodal information that users encounter in the world. To reduce such friction, future interactive interfaces should intelligently provide quick access to digital actions based on users’ context. To explore the range of possible digital actions, we conducted a diary study that required participants to capture and share the media that they intended to perform actions on (e.g., images or audio), along with their desired actions and other contextual information. Using this data, we generated a holistic design space of digital follow-up actions that could be performed in response to different types of multimodal sensory inputs. We then designed OmniActions, a pipeline powered by large language models (LLMs) that processes multimodal sensory inputs and predicts follow-up actions on the target information grounded in the derived design space. Using the empirical data collected in the diary study, we performed quantitative evaluations on three variations of LLM techniques (intent classification, in-context learning and finetuning) and identified the most effective technique for our task. Additionally, as an instantiation of the pipeline, we developed an interactive prototype and reported preliminary user feedback about how people perceive and react to the action predictions and its errors. Jiahao Nick Li, Tovi Grossman, Stephanie Santosa, Michelle Li |
CHI | 4 |
| 2023 | Affordance-Based and User-Defined Gestures for Spatial Tangible InteractionabstractAlthough mid-air hand gestures have been widely adopted by VR/AR products (e.g., Quest 2 and HoloLens), some drawbacks remain due to their lack of tangibility and tactile feedback. Opportunistic Tangible User Interfaces could address these shortcomings by repurposing existing objects in one's physical environment. However, there has yet to be a systematic investigation of the gestures that would be desirable when using opportunistic objects or how such gestures would be impacted by such objects. In this work, we conducted an elicitation study to investigate the desirability of object and gesture combinations across a variety of interactions. The results contribute (1) an opportunistic tangible UI gesture set for spatial interfaces, and (2) an Affordance-Based Object Selector Scheme that identifies ideal objects for tangible input given a desired input gesture, based on that object's physical affordances. Arising from these findings is the vision of the Adaptive Tangible User Interface, which supports the on-the-fly composition of tangible interfaces based on the affordances found in the physical environment and a user's input task. Valentin Weilun Gong, Stephanie Santosa, Tovi Grossman, Michael Glueck, Frances Lai |
Conference on Designing Interactive Systems | 2 |
| 2023 | GazeRayCursor: Facilitating Virtual Reality Target Selection by Blending Gaze and Controller RaycastingabstractRaycasting is a common method for target selection in virtual reality (VR). However, it results in selection ambiguity whenever a ray intersects multiple targets that are located at different depths. To resolve these ambiguities, we estimate object depth by projecting the closest intersection between the gaze and controller rays onto the controller ray. An evaluation of this method found that it significantly outperformed a previous eye convergence depth estimation technique. Based on these results, we developed GazeRayCursor, a novel selection technique that enhances Raycasting, by leveraging gaze for object depth estimation. In a second study, we compared two variations of GazeRayCursor with RayCursor, a recent technique developed for a similar purpose, in a dense target environment. The results indicated that GazeRayCursor decreased selection time by 45.0% and reduced manual depth adjustments by a factor of 10 in a dense target environment. Our findings showed that GazeRayCursor is an effective method for target disambiguation in VR selection without incurring extra effort. Di Laura Chen, Marcello Giordano, Hrvoje Benko, Tovi Grossman, Stephanie Santosa |
VRST | 5 |
| 2022 | Gaze as an Indicator of Input Recognition ErrorsabstractInput recognition errors are common in gesture- and touch-based recognition systems, and negatively affect user experience and performance. When errors occur, systems are unaware of them, but the user's gaze following an error may provide valuable cues for error detection. A study was conducted using a manual serial selection task to investigate whether gaze could be used to discriminate user-initiated selections from injected false positive selection errors. Logistic regression models of gaze dynamics could successfully identify injected selection errors as early as 50 milliseconds following a selection, with performance peaking at 550 milliseconds. A two-phase gaze pattern was observed in which users exhibited high gaze motion immediately following errors, and then decreased gaze motion as the error was noticed. Together, these results provide the first demonstration that gaze dynamics can be used to detect input recognition errors, and open new possibilities for systems that can assist with error recovery. Candace E. Peacock, Benjamin J. Lafreniere, Ting Zhang 0013, Stephanie Santosa, Hrvoje Benko, Tanya R. Jonker |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2021 | StickyPie: A Gaze-Based, Scale-Invariant Marking Menu Optimized for AR/VRabstractThis work explores the design of marking menus for gaze-based AR/VR menu selection by expert and novice users. It first identifies and explains the challenges inherent in ocular motor control and current eye tracking hardware, including overshooting, incorrect selections, and false activations. Through three empirical studies, we optimized and validated design parameters to mitigate these errors while reducing completion time, task load, and eye fatigue. Based on the findings from these studies, we derived a set of design guidelines to support gaze-based marking menus in AR/VR. To overcome the overshoot errors found with eye-based expert marking menu behaviour, we developed StickyPie, a marking menu technique that enables scale-independent marking input by estimating saccade landing positions. An evaluation of StickyPie revealed that StickyPie was easier to learn than the traditional technique (i.e., RegularPie) and was 10% more efficient after 3 sessions. Sunggeun Ahn, Stephanie Santosa, Mark Parent, Daniel J. Wigdor, Tovi Grossman, Marcello Giordano |
CHI | 2 |
| 2021 | False Positives vs. False Negatives: The Effects of Recovery Time and Cognitive Costs on Input Error PreferenceabstractExisting approaches to trading off false positive versus false negative errors in input recognition are based on imprecise ideas of how these errors affect user experience that are unlikely to hold for all situations. To inform dynamic approaches to setting such a tradeoff, two user studies were conducted on how relative preference for false positive versus false negative errors is influenced by differences in the temporal cost of error recovery, and high-level task factors (time pressure, multi-tasking). Participants completed a tile selection task in which false positive and false negative errors were injected at a fixed rate, and the temporal cost to recover from each of the two types of error was varied, and then indicated a preference for one error type or the other, and a frustration rating for the task. Responses indicate that the temporal costs of error recovery can drive both frustration and relative error type preference, and that participants exhibit a bias against false positive errors, equivalent to ∼1.5 seconds or more of added temporal recovery time. Several explanations for this bias were revealed, including that false positive errors impose a greater attentional demand on the user, and that recovering from false positive errors imposes a task switching cost. Benjamin J. Lafreniere, Tanya R. Jonker, Stephanie Santosa, Mark Parent, Michael Glueck, Tovi Grossman, Hrvoje Benko, Daniel J. Wigdor |
UIST | 3 |
| 2017 | Argument Mapper: Countering Cognitive Biases in Analysis with Critical (Visual) ThinkingabstractHumans are vulnerable to cognitive biases such as neglect of probability, framing effect, confirmation bias, conservatism (belief revision) and anchoring. Argument Mapper addresses these biases in intelligence analysis by providing an easy-to-use, theoretically sound, web-based interactive software tool that enables the application of evidence-based reasoning to analytic questions. Designed in collaboration with analytic methodologists, this tool combines structured argument mapping methodology with visualization techniques to help analysts make sense of complex problems and overcome cognitive biases. The tool uses Baconian probability and conjunctive logic to automatically calculate the inferential force on the upper level hypothesis. Evaluations with 16 analysts showed the tool was easy to use and easy to understand. William Wright, David Sheffield, Stephanie Santosa |
IV | 3 |
| 2014 | LACES: live authoring through compositing and editing of streaming videoabstractVideo authoring activity typically consists of three phases: planning (pre-production), capture (production) and processing (post-production). The status quo is that these phases occur separately, and the latter two have a significant amount of "slack time", where the camera operator is watching the scene unfold during capture, and the editor is re-watching and navigating through recorded footage during post-production. While this process is well suited to creating polished or professional video, video clips produced by casual video makers as seen in online forums could benefit from some editing without the overhead of current authoring tools. We introduce LACES, a tablet-based system enabling simple video manipulations in the midst of filming. Seamless in-situ integration of video capture and manipulation forms a novel workflow, allowing greater spontaneity and exploration of video creation. Dustin Freeman, Stephanie Santosa, Fanny Chevalier, Ravin Balakrishnan, Karan Singh 0004 |
CHI | 2 |
| 2013 | Direct space-time trajectory control for visual media editingabstractWe explore the design space for using object motion trajectories to create and edit visual elements in various media across space and time. We introduce a suite of pen-based techniques that facilitate fluid stylization, annotation and editing of space-time content such as video, slide presentations and 2D animation, utilizing pressure and multi-touch input. We implemented and evaluated these techniques in DirectPaint, a system for creating free-hand painting and annotation over video. Stephanie Santosa, Fanny Chevalier, Ravin Balakrishnan, Karan Singh 0004 |
CHI | 1 |
| 2013 | A field study of multi-device workflows in distributed workspacesabstractWith the selection of devices encompassing a wider range of computing surfaces, along with the near ubiquity of wireless networks, the nature of the workspace has become distributed over multiple locations and digital artifacts. We interviewed 22 professionals across a wide range of industries about their use of artifacts in their workflows to dis-cover new cross-device interaction paradigms and issues. We explore the impact of today's cloud services and app-based computing on how devices are used together. Gaps in data management and cross-device interactions were identified as the main obstacles and opportunities for improvement for multi-device interaction. Stephanie Santosa, Daniel J. Wigdor |
UbiComp | 1 |
| 2013 | MAV-Vis: a notation for model uncertaintyabstractWe apply the “Physics of Notations” theory to design MAV-Vis, a concrete syntax for partial models, i.e., models where design uncertainty is explicitly captured. To validate our implementation of this theory in creating MAV-Vis, we designed and executed an empirical user study comparing the cognitive effectiveness of MAV-Vis with the existing, ad-hoc notation, MAV-Text. We measured the ease, speed, and accuracy of each notation for reading and writing partial models. Michalis Famelis, Stephanie Santosa |
MiSE | 2 |