Jaime Ruiz 0002

dblp:41/2122-2 · DBLP profile ↗
← Back
52ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-9139-6172ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 44 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Comparison of Text-Based Inputs for Human-in-the-Loop Feedback in Vision-Language Models
abstract
Human-in-the-loop methods leverage human feedback to enhance machine learning and AI. Manual review of outputs can correct errors, identify model weaknesses, or expand labels to broaden model capabilities. Feedback collection methods range from simple flagging of outputs as correct or incorrect to more complex feature-level adjustments or natural language interpretations. This article presents a user study evaluating changes in user performance over time and explores the tradeoff between feedback quality and human effort. We compare four interactive input methods for reviewing and correcting outcomes in object detection and activity recognition in videos. Our findings indicate that while some complex input methods, such as free-text, require more time, the quality and impact of their feedback on model accuracy often surpass those of simpler methods that require less effort. However, more effort does not always lead to better-quality feedback, especially when aiming to improve the model. Our VLM experiments show that the most accurate models were trained using detailed natural language feedback or precise word-level corrections, while simple yes/no judgments also led to solid performance at a much lower annotation cost.
Reza Shahriari, Amal Hashky, Shivvrat Arya, Tyler Audino, Eric D. Ragan, Vibhav Gogate, Jaime Ruiz 0002
ACM Trans. Interact. Intell. Syst.7
2025 Challenges of Precueing Instructions for Compound Task Procedures in Mixed Reality
abstract
Augmented reality (AR) and virtual reality (VR) can enhance task guidance by overlaying visual information to improve efficiency and reduce errors. However, challenges remain in designing the appropriate presentation format and amount of information for real-time assistance. Prior research has shown benefits of visual cues in procedural tasks, but these findings are limited to simplified scenarios, highlighting a gap in understanding their effectiveness for complex, real-world applications. Therefore, we study visual design and cue effectiveness in the context of compound procedures encompassing subtasks and heterogeneous instructions. We present an experiment assessing different visual cues in VR to test a user’s ability to harness distinct information streams for different tasks, separating cues for object search and object placement for multi-step procedures. The results show that even for compound tasks requiring processing of multiple types of information, the addition of simple interaction cues for individual subtasks did significantly improved task performance for both time and errors. However, in contrast to prior studies showing successful precueing of future steps in more simplistic tasks, the study did not find evidence of precueing with the more complex tasks.
Ahmed Rageeb Ahsan, Andrew W. Tompkins, Eric D. Ragan, Jaime Ruiz 0002, Ryan P. McMahan
Graphics Interface4
2025 How Hand Constraints Influence User Defined Gestures in Mixed Reality
abstract
How do user-defined gestures for mixed reality change when users’ hands are engaged in tasks? To address this question, we conducted a gesture elicitation study to understand user preferences and the characteristics of gestures conceptualized in three scenarios with varying levels of hand constraints, namely: "both hands free", "one hand fixed", and "both hands busy". We analyzed these gestures across multiple dimensions and compared our findings with those from prior research. Our results indicate that when both hands are occupied, users tend to favor head gestures over those involving other body parts, such as the eyes or legs. Additionally, we found that most of the proposed gestures were metaphorical, with many influenced by legacy bias. These insights enhance our understanding of how hand constraints influence gesture choices in mixed reality scenarios.
Alexander Barquero, Niriksha Regmi, Oluwatomisin Obajemu, Rohith Venkatakrishnan, Christina Boucher 0001, Lisa Anthony, Jaime Ruiz 0002
Graphics Interface7
2025 Exploring Users' Perceptions on Position, Gaze Direction, and Gender of Virtual Agents in Augmented Reality
abstract
Prior research has highlighted users’ preferences for embodiment when interacting with virtual agents in augmented reality headsets. However, open questions remain regarding users’ preferences towards agent placement and gaze direction. In our study, we asked 48 adults to wear the Microsoft HoloLens 2 and find objects in a hidden object game with the help of embodied agents. We examined four distinct agent configurations for both male and female agents: a human-size agent standing beside participants, a human-size agent sitting beside participants, a small desk agent facing the screen, and a small desk agent facing the participant. Overall, participants preferred male over female virtual agents when receiving assistance, and no consistent preference emerged regarding the agents’ position or gaze direction. From our results, we build upon existing guidelines for designing better virtual agents for AR with headsets.
Rodrigo Luis Calvo, Heting Wang, Alexander Barquero, Jaime Ruiz 0002
Graphics Interface4
2025 Exploring Interactions with Companion Virtual Agents
abstract
Although companion virtual agents (CVAs) are increasingly adopted to support well-being, interaction patterns between adults and CVAs remain underexplored. This study examines how users engage with CVAs over time, exploring conversation patterns, interaction frequency, attitudes, and emotional responses over a seven-day period. Twenty-four adults engaged with a GPT-4-powered embodied CVA daily, discussing topics from personal interests to emotional reflections. Quantitative measures, including loneliness and affect scales, revealed no significant reduction in loneliness but noted decreases in positive affect and nervousness. Qualitative analysis highlighted evolving conversational dynamics, with participants shifting from exploratory questions to more reflective and personal discussions. Participants appreciated the agent’s ability to engage in fluid and meaningful conversations. However, participants also noted shortcomings, including limited recall and occasional conversational unnaturalness. These findings inform the design of CVAs, emphasizing the need for adaptive conversational strategies, enhanced emotional responsiveness, and improved memory systems to foster meaningful connections.
Rodrigo Luis Calvo, Heting Wang, Alexander Barquero, Xuanpu Zhang, Rohith Venkatakrishnan, Jaime Ruiz 0002
HAI6
2025 When is Self-Gaze Helpful? Examining Uni- vs Bi-Directional Gaze Visualization in Collocated AR Tasks
abstract
Shared-gaze visualizations (SGVs) in augmented reality enable collaborators to share focus and intentions through gaze interactions. Most prior research has examined bi-directional visualizations, where both users see their own and their partner's gaze, to provide feedback on how their gaze is communicated to their partner. However, bi-directional SGV approaches are largely based on research for remote collaboration. In collocated settings, bi-directional SGVs can obstruct views and cause distractions. Additionally, collocated applications differ from remote ones. We propose that if eye-tracking is well-calibrated, bi-directional visualizations may be unnecessary in collocated settings. To explore this, we conducted a user study comparing perceptions of uni- and bi-directional gaze visualizations in a virtual collaborative sorting task. Our results suggest that self-gaze may not always be necessary for users; however, there are cases in which self-gaze helps them feel more confident in the task. We offer a deeper understanding for future collaborative gaze interaction systems.
Daniel Alexander Delgado, Christopher Bowers, Rodrigo Luis Calvo, Jaime Ruiz 0002
ISMAR4
2025 Natural Language Interaction for Editing Visual Knowledge Graphs
abstract
Knowledge graphs are often visualized using node-link diagrams that reveal relationships and structure. In many applications using graphs, it is desirable to allow users to edit graphs to ensure data accuracy or provides updates. Commonly in graph visualization, users can interact directly with the visual elements by clicking and typing updates to specific items through traditional interaction methods in the graphical user interface. However, it can become tedious to make many updates due to the need to individually select and change numerous items in a graph. Our research investigates natural language input as an alternative method for editing network graphs. We present a user study comparing GUI graph editing with two natural language alternatives to contribute novel empirical data of the trade-offs of the different interaction methods. The findings show natural language methods to be significantly more effective than traditional GUI interaction.
Reza Shahriari, Eric D. Ragan, Jaime Ruiz 0002
K-CAP3
2025 CReLeRI: Explainable, Concept-centric, Representation, Learning, Reasoning, and Interaction Video Analysis System
abstract
Existing video analysis models often lack explainability, perform poorly on long videos, and frequently hallucinate. Commercial solutions are closed-source and costly. We introduce CReLeRI, an open-source system for action detection in untrimmed videos. CReLeRI segments videos using scene and action transitions, detects actions and their arguments and grounds them in 3D space to improve interpretability and reduce hallucinations. The system promotes transparency and trust in AI-driven analysis of complex, real-world videos. A demonstration video is also available.
Michael Francis Perez, Yichi Yang, Yuheng Zha, Enze Ma, Danish Nisar Ahmed Tamboli, Haodi Ma, Reza Shahriari, Vyom Pathak, Dzmitry Kasinets, Rohith Venkatakrishnan, Daisy Zhe Wang, Jaime Ruiz 0002, Eric D. Ragan, Zhiting Hu, Eric P. Xing, Jun-Yan Zhu
ACM Multimedia12
2025 May The Force be With You: Cloning Distant Objects to Improve Medium-Field Interactions in Augmented Reality
abstract
Augmented Reality (AR) interactions feature users interacting with virtual objects registered in the physical world. With contemporary AR experiences increasingly featuring interactions at distances, we conceptualized The Force, a technique that allows users to clone distant objects and manipulate their replicas. An empirical evaluation was conducted, comparing it against two well-established techniques including controller-based ray-casting and a gaze-based pinching technique in a pick-and-place task. We employed a within-subjects design, collecting data on both objective performance and subjective user experience. Results suggest that The Force allows for higher levels of accuracy and efficiency in medium-field tasks that require precision and fine motor control. Furthermore, we discovered avenues towards iteratively refining this technique. We go on to discuss the implications of our findings in an effort to facilitate better interactions in augmented reality.
Danish Nisar Ahmed Tamboli, Rohith Venkatakrishnan, Roshan Venkatakrishnan, Balagopal Raveendranath, Julia Woodward, Isaac Wang, Jesse Smith, Jaime Ruiz 0002
VR8
2024 Understanding User Needs for Task Guidance Systems Through the Lens of Cooking
abstract
To design intuitive and effective context-aware task guidance systems, we must understand users’ thought processes and the obstacles they experience when they perform tasks. Though task guidance systems have proven beneficial in many domains for improving task performance and reducing user frustration, there is a lack of general guidelines and design principles for their development. Prior work has shown that recipe-based cooking is a strong medium for studying task planning and execution. In response, we conducted a contextual inquiry study in home kitchens, observing eight different participants’ cooking sessions. We used affinity diagramming of our notes and transcripts to identify common obstacles faced by participants and establish user needs in the areas of object interaction, safety, knowledge base, and task coordination. We discuss how these findings can inform the design of technology-driven solutions for task guidance systems beyond cooking.
Alexander Barquero, Rodrigo Luis Calvo, Daniel Alexander Delgado, Isaac Wang, Lisa Anthony, Jaime Ruiz 0002
Conference on Designing Interactive Systems6
2024 Investigating Contextual Notifications to Drive Self-Monitoring in mHealth Apps for Weight Maintenance
abstract
Mobile health applications for weight maintenance offer self-monitoring as a tool to empower users to achieve health goals (e.g., losing weight); yet maintaining consistent self-monitoring over time proves challenging for users. These apps use push notifications to help increase users’ app engagement and reduce long-term attrition, but they are often ignored by users due to appearing at inopportune moments. Therefore, we analyzed whether delivering push notifications based on time alone or also considering user context (e.g., current activity) affected users’ engagement in a weight maintenance app, in a 4-week in-the-wild study with 30 participants. We found no difference in participants’ overall (across the day) self-monitoring frequency between the two conditions, but in the context-based condition, participants responded faster and more frequently to notifications, and logged their data more timely (as eating/exercising occurs). Our work informs the design of notifications in weight maintenance apps to improve their efficacy in promoting self-monitoring.
Julia Woodward, Dinank Bista, Xuanpu Zhang, Ishvina Singh, Oluwatomisin Obajemu, Meena N. Shankar, Kathryn M. Ross, Jaime Ruiz 0002, Lisa Anthony
CHI9
2024 Human-Centered Evaluation of EMG-Based Upper-Limb Prosthetic Control Modes
abstract
The aim of this study was to experimentally test the effects of different electromyographic-based prosthetic control modes on user task performance, cognitive workload, and perceived usability to inform further human-centered design and application of these prosthetic control interfaces. We recruited 30 able-bodied participants for a between-subjects comparison of three control modes: direct control (DC), pattern recognition (PR), and continuous control (CC). Multiple human-centered evaluations were used, including task performance, cognitive workload, and usability assessments. To ensure that the results were not task-dependent, this study used two different test tasks, including the clothespin relocation task and Southampton hand assessment procedure-door handle task. Results revealed performance with each control mode to vary among tasks. When the task had high-angle adjustment accuracy requirements, the PR control outperformed DC. For cognitive workload, the CC mode was superior to DC in reducing user load across tasks. Both CC and PR control appear to be effective alternatives to DC in terms of task performance and cognitive load. Furthermore, we observed that, when comparing control modes, multitask testing and multifaceted evaluations are critical to avoid task-induced or method-induced evaluation bias. Hence, future studies with larger samples and different designs will be needed to expand the understanding of prosthetic device features and workload relationships.
Yunmei Liu, Joseph Berman, Albert Dodson, Maryam Zahabi, He Huang 0002, Jaime Ruiz 0002, David B. Kaber
IEEE Trans. Hum. Mach. Syst.7
2023 Designing Textual Information in AR Headsets to Aid in Adults' and Children's Task Performance
abstract
Augmented reality (AR) headsets are being utilized in different task-based domains (e.g., healthcare, education) for both adults and children. However, prior work has mainly examined the applicability of AR headsets instead of how to design the visual information being displayed. It is essential to study how visual information should be presented in AR headsets to maximize task performance for both adults and children. Therefore, we conducted two studies (adults vs. children) analyzing distinct design combinations of critical and secondary textual information during a procedural assembly task. We found that while the design of information did not affect adults' task performance, the location of information had a direct effect on children's task performance. Our work contributes new understanding on how to design textual information in AR headsets to aid in adults’ and children's task performance. In addition, we identify specific differences on how to design textual information between adults and children.
Julia Woodward, Jaime Ruiz 0002
IDC2
2023 Stop Copying Me: Evaluating nonverbal mimicry in embodied motivational agents
abstract
Motivational agents are virtual agents that seek to motivate users by providing feedback and guidance. Prior work has shown how certain factors of an agent, such as the type of feedback given or the agent's appearance, can influence user motivation when completing tasks. However, it is not known how nonverbal mirroring affects an agent's ability to motivate users. Specifically, would an agent that mirrors be more motivating than an agent that does not? Would an agent trained on real human behaviors be better? We conducted a within-subjects study asking 30 participants to play a "find-the-hidden-object" game while interacting with a motivational agent that would provide hints and feedback on the user's performance. We created three agents: a Control agent that did not respond to the user's movements, a simple Mimic agent that mirrored the user's movements on a delay, and a Complex agent that used a machine-learned behavior model. We asked participants to complete a questionnaire asking them to rate their levels of motivation and perceptions of the agent and its feedback. Our results showed that the Mimic agent was more motivating than the Control agent and more helpful than the Complex agent. We also found that when participants became aware of the mimicking behavior, it can feel weird or creepy; therefore, it is important to consider the detection of mimicry when designing virtual agents.
Isaac Wang, Rodrigo Luis Calvo, Heting Wang, Jaime Ruiz 0002
IVA4
2023 Cognitive Workload and Usability of Virtual Reality Simulation for Prosthesis Training
abstract
Amputees use prosthetic devices to perform activities of daily living. However, some users reject their devices due to the lack of usability or high cognitive workload. Although virtual reality has been studied in this domain for training purposes, there has not been any investigation on usability and cognitive workload of using virtual reality simulations for training of prosthetic devices. The objective of this study was to compare cognitive workload and usability of using virtual reality-based simulation of electromyography based prosthetic devices and physical devices. The findings suggested that using virtual reality simulations were helpful in reducing cognitive workload and increasing perceived usability of prosthetic devices.
Austin Music, Daniel Delgado, Joseph Berman, Albert Dodson, Yunmei Liu, Jaime Ruiz 0002, He Huang 0002, David B. Kaber, Maryam Zahabi
SMC7
2023 Analytic Review of Using Augmented Reality for Situational Awareness
abstract
Situational awareness is the perception and understanding of the surrounding environment. Maintaining situational awareness is vital for performance and error prevention in safety critical domains. Prior work has examined applying augmented reality (AR) to the context of improving situational awareness, but has mainly focused on the applicability of using AR rather than on information design. Hence, there is a need to investigate how to design the presentation of information, especially in AR headsets, to increase users' situational awareness. We conducted a Systematic Literature Review to research how information is currently presented in AR, especially in systems that are being utilized for situational awareness. Comparing current presentations of information to existing design recommendations aided in identifying future areas of design. In addition, this survey further discusses opportunities and challenges in applying AR to increasing users' situational awareness.
Julia Woodward, Jaime Ruiz 0002
IEEE Trans. Vis. Comput. Graph.2
2022 "It Would Be Cool to Get Stampeded by Dinosaurs": Analyzing Children's Conceptual Model of AR Headsets Through Co-Design
abstract
Children are being presented with augmented reality (AR) in different contexts, such as education and gaming. However, little is known about how children conceptualize AR, especially AR headsets. Prior work has shown that children's interaction behaviors and expectations of technological devices can be quite different from adults’. It is important to understand children's mental models of AR headsets to design more effective experiences for them. To elicit children's perceptions, we conducted four participatory design sessions with ten children on designing content for imaginary AR headsets. We found that children expect AR systems to be highly intelligent and to recognize and virtually transform surroundings to create immersive environments. Also, children are in favor of using these devices for difficult tasks but prefer to work on their own for easy tasks. Our work contributes new understanding on how children comprehend AR headsets and provides recommendations for designing future headsets for children.
Julia Woodward, Feben Alemu, Natalia E. López Adames, Lisa Anthony, Jason C. Yip 0001, Jaime Ruiz 0002
CHI6
2021 AmbiTeam: Providing Team Awareness Through Ambient Displays
abstract
Due to the COVID-19 pandemic, research is increasingly conducted remotely without the benefit of informal interactions that help maintain awareness of each collaborator's work progress. We developed AmbiTeam, an ambient display that shows activity related to the files of a team project, to help collaborations preserve a sense of the team's involvement while working remotely. We found that using AmbiTeam did have a quantifiable effect on researchers' perceptions of their collaborators' project prioritization. We also found that the use of the system motivated researchers to work on their collaborative projects. This effect is known as "the motivational presences of others," one of the key challenges that make distance work difficult. We discuss how ambient displays can support remote collaborative work by recreating the motivational presence of others.
Sarah Morrison-Smith, Lydia B. Chilton, Jaime Ruiz 0002
Graphics Interface3
2021 Examining the Use of Nonverbal Communication in Virtual Agents
abstract
Virtual agents are systems that add a social dimension to computing, often featuring not only natural language input but also an embodiment or avatar. This allows them to take on a more social role and leverage the use of nonverbal communication (NVC). In humans, NVC is used for many purposes, including communicating intent, directing attention, and conveying emotion. As a result, researchers have developed agents that emulate these behaviors. However, challenges pervade the design and development of NVC in agents. Some articles reveal inconsistencies in the benefits of agent NVC; others show signs of difficulties in the process of analyzing and implementing behaviors. Thus, it is unclear what the specific outcomes and effects of incorporating NVC in agents and what outstanding challenges underlie development. This survey seeks to review the uses, outcomes, and development of NVC in virtual agents to identify challenges and themes to improve and motivate the design of future virtual agents.
Isaac Wang, Jaime Ruiz 0002
Int. J. Hum. Comput. Interact.2
2021 A Survey on Applying Automated Recognition of Touchscreen Stroke Gestures to Children's Input
abstract
Abstract Gesture recognition algorithms help designers create intelligent user interfaces for a number of application areas. However, these recognition algorithms are usually designed to recognize the gestures of adults, not children, and as such they generally do not perform as well for children as adults. Recognition of younger children’s gestures is particularly poor when compared to recognition of older children’s and adults’ gestures. Researchers have begun to examine the aspects of children’s gesture articulation patterns that make recognition difficult. This paper extends the initial work examining child-specific recognition approaches by considering general-purpose approaches and how they might apply to the problem of recognizing children’s touchscreen gestures. This paper presents a survey of existing recognition and analysis techniques for gestures of both adults and children from a human-centered perspective, highlighting ways in which improved recognition can lead to a better experience for children using touchscreen gestures in a variety of contexts.
Jaime Ruiz 0002, Lisa Anthony
Interact. Comput.2
2020 Diana's World: A Situated Multimodal Interactive Agent
abstract
State of the art unimodal dialogue agents lack some core aspects of peer-to-peer communication—the nonverbal and visual cues that are a fundamental aspect of human interaction. To facilitate true peer-to-peer communication with a computer, we present Diana, a situated multimodal agent who exists in a mixed-reality environment with a human interlocutor, is situation- and context-aware, and responds to the human's language, gesture, and affect to complete collaborative tasks.
Nikhil Krishnaswamy, Pradyumna Narayana, Rahul Bangar, Kyeongmin Rim, Dhruva Patil, David G. McNeely-White, Jaime Ruiz 0002, Bruce A. Draper, J. Ross Beveridge, James Pustejovsky
AAAI7
2020 Evaluating the Scalability of Non-Preferred Hand Mode Switching in Augmented Reality
abstract
Mode switching allows applications to support a wide range of operations (e.g. selection, manipulation, and navigation) using a limited input space. While the performance of different mode switching techniques has been extensively examined for pen- and touch-based interfaces, investigating mode switching in augmented reality (AR) is still relatively new. Prior work found that using non-preferred hand is an efficient mode switching technique in AR. However, it is unclear how the technique performs when increasing the number of modes, which is more indicative of real-world applications. Therefore, we examined the scalability of non-preferred hand mode switching in AR with two, four, six, and eight modes. We found that as the number of modes increase, performance plateaus after the four-mode condition. We also found that counting gestures have varying effects on mode switching performance in AR. Our findings suggest that modeling mode switching performance in AR is more complex than simply counting the number of available modes. Our work lays a foundation for understanding the costs associated with scaling interaction techniques in AR.
Jesse Smith, Isaac Wang, Winston Wei, Julia Woodward, Jaime Ruiz 0002
AVI5
2020 Examining Fitts' and FFitts' Law Models for Children's Pointing Tasks on Touchscreens
abstract
Fitts' law has accurately modeled both children's and adults' pointing movements, but it is not as precise for modeling movement to small targets. To address this issue, prior work presented FFitts' law, which is more exact than Fitts' law for modeling adults' finger input on touchscreens. Since children's touch interactions are more variable than adults, it is unclear if FFitts' law should be applied to children. We conducted a 2D target acquisition task with 54 children (ages 5-10) to examine if FFitts' law can accurately model children's touchscreen movement time. We found that Fitts' law using nominal target widths is more accurate, with a R2 value of 0.93, than FFitts' law for modeling children's finger input on touchscreens. Our work contributes new understanding of how to accurately predict children's finger touch performance on touchscreens.
Julia Woodward, Jahelle Cato, Jesse Smith, Isaac Wang, Brett Benda, Lisa Anthony, Jaime Ruiz 0002
AVI7
2020 Examining the Presentation of Information in Augmented Reality Headsets for Situational Awareness
abstract
Augmented Reality (AR) headsets are being employed in industrial settings (e.g., the oil industry); however, there has been little work on how information should be presented in these headsets, especially in the context of situational awareness. We present a study examining three different presentation styles (Display, Environment, Mixed Environment) for textual secondary information in AR headsets. We found that the Display and Environment presentation styles assisted in perception and comprehension. Our work contributes a first step to understanding how to design visual information in AR headsets to support situational awareness.
Julia Woodward, Jesse Smith, Isaac Wang, Sofia Cuenca, Jaime Ruiz 0002
AVI5
2020 Examining the Link between Children's Cognitive Development and Touchscreen Interaction Patterns
abstract
It is well established that children's touch and gesture interactions on touchscreen devices are different from those of adults, with much prior work showing that children's input is recognized more poorly than adults? input. In addition, researchers have shown that recognition of touchscreen input is poorest for young children and improves for older children when simply considering their age; however, individual differences in cognitive and motor development could also affect children's input. An understanding of how cognitive and motor skill influence touchscreen interactions, as opposed to only coarser measurements like age and grade level, could help in developing personalized and tailored touchscreen interfaces for each child. To investigate how cognitive and motor development may be related to children's touchscreen interactions, we conducted a study of 28 participants ages 4 to 7 that included validated assessments of the children's motor and cognitive skills as well as typical touchscreen target acquisition and gesture tasks. We correlated participants? touchscreen behaviors to their cognitive development level, including both fine motor skills and executive function. We compare our analysis of touchscreen interactions based on cognitive and motor development to prior work based on children's age. We show that all four factors (age, grade level, motor skill, and executive function) show similar correlations with target miss rates and gesture recognition rates. Thus, we conclude that age and grade level are sufficiently sensitive when considering children's touchscreen behaviors.
Aishat Aloba, Pavlo D. Antonenko, Jaime Ruiz 0002, Lisa Anthony
ICMI6
2020 MMGatorAuth: A Novel Multimodal Dataset for Authentication Interactions in Gesture and Voice
abstract
The future of smart environments is likely to involve both passive and active interactions on the part of users. Depending on what sensors are available in the space, users may make use of multimodal interaction modalities such as hand gestures or voice commands. There is a shortage of robust yet controlled multimodal interaction datasets for smart environment applications. One application domain of interest based on current state-of-the-art is authentication for sensitive or private tasks, such as banking and email. We present a novel, large multimodal dataset for authentication interactions in both gesture and voice, collected from 106 volunteers who each performed 10 examples of each of a set of hand gesture and spoken voice commands chosen from prior literature (10,600 gesture samples and 13,780 voice samples). We present the data collection method, raw data and common features extracted, and a case study illustrating how this dataset could be useful to researchers. Our goal is to provide a benchmark dataset for testing future multimodal authentication solutions, enabling comparison across approaches.
Sarah Morrison-Smith, Aishat Aloba, Hangwei Lu, Brett Benda, Shaghayegh Esmaeili, Gianne Flores, Jesse Smith, Nikita Soni 0001, Isaac Wang, Rejin Joy, Damon L. Woodard, Jaime Ruiz 0002, Lisa Anthony
ICMI12
2020 Truly Visual Caller ID? An Analysis of Anti-Robocall Applications and their Accessibility to Visually Impaired Users
abstract
Robocalls interrupt daily activity, cause financial harm, and influence users to ignore calls from unfamiliar numbers. Service providers and developers have created Anti-Robocall applications to attempt to restore trust in the phone and decrease the impact of robocalls on daily life. However, whether or not such applications meet accessibility standards and are therefore usable by vulnerable populations, particularly the visually impaired, is unknown. In this paper, we use a combination of the W3C's Mobile Web Content Accessibility Guidelines (MWCAG) and interviews with 11 visually impaired users to establish accessibility metrics for Anti-Robocall applications. We then evaluate 56 Anti-Robocall applications for Android to assess whether they met the needs of the visually impaired community. Our results indicate that 100% of the applications fail to meet all basic accessibility guidelines including minimum color contrast, button labels (to assist screen readers), and automatic audible alerts. As a result, we show that despite the availability of a variety of tools to help developers identify and correct these problems, this important class of applications does not meet basic accessibility requirements. We conclude by suggesting viable paths forward that ensure inclusion and protection for the visually impaired community.
Imani N. S. Munyaka, Jasmine D. Bowers, Liz-Laure Laborde, Juan E. Gilbert, Jaime Ruiz 0002, Patrick Traynor
ISTAS5
2020 Wow, You Are Terrible at This!: An Intercultural Study on Virtual Agents Giving Mixed Feedback
abstract
While the effects of virtual agents in terms of likeability, uncanniness, etc. are well explored, it is unclear how their appearance and the feedback they give affects people's reactions. Is critical feedback from an agent embodied as a mouse or a robot taken less serious than from a human agent? In an intercultural study with 120 participants from Germany and the US, participants had to find hidden objects in a game and received feedback on their performance by virtual agents with different appearances. As some levels were designed to be unsolvable, critical feedback was unavoidable. We hypothesized that feedback would be taken more serious, the more human the agent looked. Also, we expected the subjects from the US to react more sensitively to criticism. Surprisingly, our results showed that the agents' appearance did not significantly change the participants' perception. Also, while we found highly significant differences in inspirational and motivational effects as well as in perceived task load between the two cultures, the reactions to criticism were contrary to expectations based on established cultural models. This work improves our understanding on how affective virtual agents are to be designed, both with respect to culture and to dialogue strategies.
Isaac Wang, Lea Buchweitz, Jesse Smith, Lara-Sophie Bornholdt, Jonas Grund, Jaime Ruiz 0002, Oliver Korn
IVA6
2020 Are You Going to Answer That? Measuring User Responses to Anti-Robocall Application Indicators
Imani N. S. Munyaka, Jasmine D. Bowers, Keith McNamara Jr., Juan E. Gilbert, Jaime Ruiz 0002, Patrick Traynor
NDSS5
2019 Exploring Virtual Agents for Augmented Reality
abstract
Prior work has shown that embodiment can benefit virtual agents, such as increasing rapport and conveying non-verbal information. However, it is unclear if users prefer an embodied to a speech-only agent for augmented reality (AR) headsets that are designed to assist users in completing real-world tasks. We conducted a study to examine users' perceptions and behaviors when interacting with virtual agents in AR. We asked 24 adults to wear the Microsoft HoloLens and find objects in a hidden object game while interacting with an agent that would offer assistance. We presented participants with four different agents: voice-only, non-human, full-size embodied, and a miniature embodied agent. Overall, users preferred the miniature embodied agent due to the novelty of his size and reduced uncanniness as opposed to the larger agent. From our results, we draw conclusions about how agent representation matters and derive guidelines on designing agents for AR headsets.
Isaac Wang, Jesse Smith, Jaime Ruiz 0002
CHI3
2019 Experimental Analysis of Single Mode Switching Techniques in Augmented Reality
Jesse Smith, Isaac Wang, Julia Woodward, Jaime Ruiz 0002
Graphics Interface4
2018 EASEL: Easy Automatic Segmentation Event Labeler
abstract
Video annotation is a vital part of research examining gestural and multimodal interaction as well as computer vision, machine learning, and interface design. However, annotation is a difficult, time-consuming task that requires high cognitive effort. Existing tools for labeling and annotation still require users to manually label most of the data, limiting the tools helpfulness. In this paper, we present the Easy Automatic Segmentation Event Labeler (EASEL), a tool supporting gesture analysis. EASEL streamlines the annotation process by introducing assisted annotation, using automatic gesture segmentation and recognition to automatically annotate gestures. To evaluate the efficacy of assisted annotation, we conducted a user study with 24 participants and found that assisted annotation decreased the time needed to annotate videos with no difference in accuracy compared with manual annotation. The results of our study demonstrate the benefit of adding computational intelligence to video and audio annotation tasks.
Isaac Wang, Pradyumna Narayana, Jesse Smith, Bruce A. Draper, J. Ross Beveridge, Jaime Ruiz 0002
IUI6
2018 Investigating Separation of Territories and Activity Roles in Children's Collaboration around Tabletops
abstract
Prior work has shown that children exhibit negative collaborative behaviors, such as blocking others' access to objects, when collaborating on interactive tabletop computers. We implemented previous design recommendations, namely separate physical territories and activity roles, which had been recommended to decrease these negative collaborative behaviors. We developed a multi-touch "I-Spy" picture searching application with separate territory partitions and activity roles. We conducted a deep qualitative analysis of how six pairs of children, ages 6 to 10, interacted with the application. Our analysis revealed that the collaboration styles differed for each pair, both in regards to the interaction with the task and with each other. Several pairs exhibited negative physical and verbal collaborative behaviors, such as nudging each other out of the way. Based on our analysis, we suggest that it is important for a collaborative task to offer equal opportunities for interaction, but it may not be necessary to strive for complete equity of collaboration. We examine the applicability of prior design guidelines and suggest open questions for future research to inform the design of tabletop applications to support collaboration for children.
Julia Woodward, Shaghayegh Esmaeili, Ayushi Jain, John Bell, Jaime Ruiz 0002, Lisa Anthony
Proc. ACM Hum. Comput. Interact.5
2017 Comparing human and machine recognition of children's touchscreen stroke gestures
abstract
Children's touchscreen stroke gestures are poorly recognized by existing recognition algorithms, especially compared to adults' gestures. It seems clear that improved recognition is necessary, but how much is realistic? Human recognition rates may be a good starting point, but no prior work exists establishing an empirical threshold for a target accuracy in recognizing children's gestures based on human recognition. To this end, we present a crowdsourcing study in which naïve adult viewers recruited via Amazon Mechanical Turk were asked to classify gestures produced by 5- to 10-year-old children. We found a significant difference between human (90.60%) and machine (84.14%) recognition accuracy, over all ages. We also found significant differences between human and machine recognition of gestures of different types: humans perform much better than machines do on letters and numbers versus symbols and shapes. We provide an empirical measure of the accuracy that future machine recognition should aim for, as well as a guide for which categories of gestures have the most room for improvement in automated recognition. Our findings will inform future work on recognition of children's gestures and improving applications for children.
Jaime Ruiz 0002, Lisa Anthony
ICMI2
2017 Tablets, tabletops, and smartphones: cross-platform comparisons of children's touchscreen interactions
abstract
The proliferation of smartphones and tablets has increased children’s access to and usage of touchscreen devices. Prior work on smartphones has shown that children’s touch interactions differ from adults’. However, larger screen devices like tablets and tabletops have not been studied at the same granularity for children as smaller devices. We present two studies: one of 13 children using tablets with pen and touch, and one of 18 children using a touchscreen tabletop device. Participants completed target touching and gesture drawing tasks. We found significant differences in performance by modality for tablet: children responded faster and slipped less with touch than pen. In the tabletop study, children responded more accurately to changing target locations (fewer holdovers), and were more accurate touching targets around the screen. Gesture recognition rates were consistent across devices. We provide design guidelines for children’s touchscreen interactions across screen sizes to inform the design of future touchscreen applications for children.
Julia Woodward, Aishat Aloba, Ayushi Jain, Jaime Ruiz 0002, Lisa Anthony
ICMI5
2016 Exploring Non-touchscreen Gestures for Smartwatches
abstract
Although smartwatches are gaining popularity among mainstream consumers, the input space is limited due to their small form factor. The goal of this work is to explore how to design non-touchscreen gestures to extend the input space of smartwatches. We conducted an elicitation study eliciting gestures for 31 smartwatch tasks. From this study, we demonstrate that a consensus exists among the participants on the mapping of gesture to command and use this consensus to specify a user-defined gesture set. Using gestures collected during our study, we define a taxonomy describing the mapping and physical characteristics of the gestures. Lastly, we provide insights to inform the design of non-touchscreen gestures for smartwatch interaction.
Shaikh Shawon Arefin Shimon, Courtney Lutton, Zichun Xu, Sarah Morrison-Smith, Christina Boucher 0001, Jaime Ruiz 0002
CHI6
2016 Using Audio Cues to Support Motion Gesture Interaction on Mobile Devices
abstract
Motion gestures are an underutilized input modality for mobile interaction despite numerous potential advantages. Negulescu et al. found that the lack of feedback on attempted motion gestures made it difficult for participants to diagnose and correct errors, resulting in poor recognition performance and user frustration. In this article, we describe and evaluate a training and feedback technique, Glissando , which uses audio characteristics to provide feedback on the system’s interpretation of user input. This technique enables feedback by verbally confirming correct gestures and notifying users of errors in addition to providing continuous feedback by manipulating the pitch of distinct musical notes mapped to each of three dimensional axes in order to provide both spatial and temporal information.
Sarah Morrison-Smith, Megan Hofmann, Yang Li 0058, Jaime Ruiz 0002
ACM Trans. Appl. Percept.4
2015 Soft-Constraints to Reduce Legacy and Performance Bias to Elicit Whole-body Gestures with Low Arm Fatigue
abstract
Participant biases can influence proposed gestures in elicitation studies. There is a legacy bias from previous experience with, or even knowledge of, existing input devices, interfaces, and technologies. There is also a performance bias, where the artificial study setting does not encourage consideration of long-term aspects such as fatigue. These biases make it especially difficult to uncover gestures appropriate for whole-body gestural input. We propose using soft constraints to correct for legacy and performance biases by penalizing physical movements. We use wrist weights as a soft constraint to elicit whole-body gestures with low arm fatigue. We show soft constraints encourage a wider range of gestures using subtler arm movements or alternate body parts and lower consumed endurance for arm movements.
Jaime Ruiz 0002, Daniel Vogel 0001
CHI1
2015 Exploring User-Defined Back-Of-Device Gestures for Mobile Devices
abstract
Many studies have highlighted the advantages of expanding the input space of mobile devices by utilizing the back of the device. We extend this work by performing an elicitation study to explore users' mapping of gestures to smartphone commands and identify their criteria for using back-of-device gestures. Using the data collected from our study, we present elicited gestures and highlight common user motivations, both of which inform the design of back-of-device gestures for mobile interaction.
Shaikh Shawon Arefin Shimon, Sarah Morrison-Smith, Noah John, Ghazal Fahimi, Jaime Ruiz 0002
MobileHCI5
2015 Student Response to Teaching of Memory Cues and Resumption Strategies in Computer Science Classes
abstract
Programming is a creative process that requires the ability to concentrate and juggle multiple concepts simultaneously in one's mind. Existing research shows there is a tangible cost when a programmer is interrupted as the programmer must recover the context of his work and refocus on the task at hand. However, CS students are rarely taught about interruptions and how to manage them. Instead, teaching tends to focus only on technical concepts. In addition, there is little research on interruptions with respect to CS students. Therefore, our research examines what happens when CS students are taught about interruptions and how to cope with them.
Noah John, Jaime Ruiz 0002
SIGCSE2
2014 Analyzing intended use effects in target acquisition
abstract
Recent work by Mandryk and Lough demonstrated that the movement time of Fitts-style pointing tasks varies based on intended use of a target, suggesting major implications for HCI research that models pointing using Fitts' Law. We replicate the study of Mandryk and Lough to determine exactly how and why observed movement times vary. We demonstrate that any variation in movement time is the result of differences in additive factors (a in Fitts' equation) and can be attributed to changes in the time a user spends over their primary target.
Jaime Ruiz 0002, Edward Lank
AVI1
2012 Territoriality and behaviour on and around large vertical publicly-shared displays
abstract
We investigate behaviours on, and around, large vertical displays during concurrent usage. Using an observational field study, we identify fundamental patterns of how people use existing public displays: their orientation, positioning, group identification, and behaviour within and between social groups just-before, during, and just-after usage. These results are then used to motivate a controlled experiment where two individuals or two pairs of individuals complete tasks concurrently on a simulated large vertical display. Results from our controlled study demonstrates that vertical surface territories are similar to those found in horizontal tabletops in function, but their definitions and social conventions are different. In addition, the nature of use-while-standing systems results in more complex and dynamic physical territories around the display. We show that the anthropological notion of personal space must be slightly refined for application to vertical displays.
Alec Azad, Jaime Ruiz 0002, Daniel Vogel 0001, Mark S. Hancock, Edward Lank
Conference on Designing Interactive Systems2
2012 Tap, swipe, or move: attentional demands for distracted smartphone input
abstract
Smartphones are frequently used in environments where the user is distracted by another task, for example by walking or by driving. While the typical interface for smartphones involves hardware and software buttons and surface gestures, researchers have recently posited that, for distracted environments, benefits may exist in using motion gestures to execute commands. In this paper, we examine the relative cognitive demands of motion gestures and surface taps and gestures in two specific distracted scenarios: a walking scenario, and an eyes-free seated scenario. We show, first, that there is no significant difference in reaction time for motion gestures, taps, or surface gestures on smartphones. We further show that motion gestures result in significantly less time looking at the smartphone during walking than does tapping on the screen, even with interfaces optimized for eyes-free input. Taken together, these results show that, despite somewhat lower throughput, there may be benefits to making use of motion gestures as a modality for distracted input on smartphones.
Matei Negulescu, Jaime Ruiz 0002, Yang Li 0058, Edward Lank
AVI2
2012 A recognition safety net: bi-level threshold recognition for mobile motion gestures
abstract
Designers of motion gestures for mobile devices face the difficult challenge of building a recognizer that can separate gestural input from motion noise. A threshold value is often used to classify motion and effectively balances the rates of false positives and false negatives. We present a bi-level threshold recognition technique designed to lower the rate of recognition failures by accepting either a tightly thresholded gesture or two consecutive possible gestures recognized by a relaxed model. Evaluation of the technique demonstrates that the technique can aid in recognition for users who have trouble performing motion gestures. Lastly, we suggest the use of bi-level thresholding to scaffold the learning of gestures.
Matei Negulescu, Jaime Ruiz 0002, Edward Lank
Mobile HCI2
2011 DoubleFlip: a motion gesture delimiter for mobile interaction
abstract
To make motion gestures more widely adopted on mobile devices it is important that devices be able to distinguish between motion intended for mobile interaction and every-day motion. In this paper, we present DoubleFlip, a unique motion gesture designed as an input delimiter for mobile motion-based interaction. The DoubleFlip gesture is distinct from regular motion of a mobile device. Based on a collection of 2,100 hours of motion data captured from 99 users, we found that our DoubleFlip recognizer is extremely resistant to false positive conditions, while still achieving a high recognition rate. Since DoubleFlip is easy to perform and unlikely to be accidentally invoked, it provides an always-active input event for mobile interaction.
Jaime Ruiz 0002, Yang Li 0058
CHI1
2011 User-defined motion gestures for mobile interaction
abstract
Modern smartphones contain sophisticated sensors to monitor three-dimensional movement of the device. These sensors permit devices to recognize motion gestures - deliberate movements of the device by end-users to invoke commands. However, little is known about best-practices in motion gesture design for the mobile computing paradigm. To address this issue, we present the results of a guessability study that elicits end-user motion gestures to invoke commands on a smartphone device. We demonstrate that consensus exists among our participants on parameters of movement and on mappings of motion gestures onto commands. We use this consensus to develop a taxonomy for motion gestures and to specify an end-user inspired motion gesture set. We highlight the implications of this work to the design of smartphone applications and hardware. Finally, we argue that our results influence best practices in design for all gestural interfaces.
Jaime Ruiz 0002, Yang Li 0058, Edward Lank
CHI1
2010 Speeding pointing in tiled widgets: understanding the effects of target expansion and misprediction
abstract
Target expansion is a pointing facilitation technique where the user's target, typically an interface widget, is dynamically enlarged to speed pointing in interfaces. However, with densely packed (tiled) arrangements of widgets, interfaces cannot expand all potential targets; they must, instead, predict the user's desired target. As a result, mispredictions will occur which may disrupt the pointing task. In this paper, we present a model describing the cost/benefit of expanding multiple targets using the probability distribution of a given predictor. Using our model, we demonstrate how the model can be used to infer the accuracy required by target prediction techniques. The results of this work are another step toward pointing facilitation techniques that allow users to outperform Fitts' Law in realistic pointing tasks.
Jaime Ruiz 0002, Edward Lank
IUI1
2008 A model of non-preferred hand mode switching
Jaime Ruiz 0002, Andrea Bunt, Edward Lank
Graphics Interface1
2008 Analyzing the kinematics of bivariate pointing
Jaime Ruiz 0002, David Tausky, Andrea Bunt, Edward Lank, Richard Mann
Graphics Interface1
2007 Endpoint prediction using motion kinematics
abstract
Recently proposed novel interaction techniques such as cursor jumping [1] and target expansion for tiled arrangements [13] are predicated on an ability to effectively estimate the endpoint of an input gesture prior to its completion. However, current endpoint estimation techniques lack the precision to make these interaction techniques possible. To address a recognized lack of effective endpoint prediction mechanisms, we propose a new technique for endpoint prediction that applies established laws of motion kinematics in a novel way to the identification of motion endpoint. The technique derives a model of speed over distance that permits extrapolation. We verify our model experimentally using stylus targeting tasks, and demonstrate that our endpoint prediction is almost twice as accurate as the previously tested technique [13] at points more than twice as distant from motion endpoint.
Edward Lank, Yi-Chun Nikko Cheng, Jaime Ruiz 0002
CHI3
2007 A study on the scalability of non-preferred hand mode manipulation
abstract
In pen-tablet input devices modes allow overloading of the electronic stylus. In the case of two modes, switching modes with the non-preferred hand is most effective [12]. Further, allowing temporal overlap of mode switch and pen action boosts speed [11]. We examine the effect of increasing the number of interface modes accessible via non-preferred hand mode switching on task performance in pen-tablet interfaces. We demonstrate that the temporal benefit of overlapping mode-selection and pen action for the two mode case is preserved as the number of modes increases. This benefit is the result of both concurrent action of the hands, and reduced planning time for the overall task. Finally, while allowing bimanual overlap is still faster it takes longer to switch modes as the number of modes increases. Improved understanding of the temporal costs presented assists in the design of pen-tablet interfaces with larger sets of interface modes.
Jaime Ruiz 0002, Edward Lank
ICMI1
2006 Concurrent bimanual stylus interaction: a study of non-preferred hand mode manipulation
Edward Lank, Jaime Ruiz 0002, William B. Cowan
Graphics Interface2