VLDB 2026 Research / reviewers in the wild / expert
Daniel Szafir
dblp:97/11299 · also Daniel J. Szafir
· DBLP profile ↗
50ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0003-1848-7884ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 36 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 29 · 3 first-author · 15 since 2021Systems, architecture and hardware · 13 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting ApproachabstractUnstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named "external facelist", to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems. Spiros Tsalikis, William J. Schroeder, Daniel Szafir, Kenneth Moreland |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | MARCER: Multimodal Augmented Reality for Composing and Executing Robot TasksabstractIn this work, we combine the strengths of humans and robots by developing MARCER, a novel interactive and multimodal end-user robot programming system. MARCER utilizes a Large Language Model to translate users' natural language task descriptions and environmental context into Action Plans for robot execution, based on a trigger-action programming paradigm that facilitates authoring reactive robot behaviors. MARCER also affords interaction via augmented reality to help users parameterize and validate robot programs and provide real-time, visual previews and feedback directly in the context of the robot's operating environment. We present the design, implementation, and evaluation of MARCER to explore the usability of such systems and demonstrate how trigger-action programming, Large Language Models, and augmented reality hold deep-seated synergies that, when combined, empower users to program general-purpose robots to perform everyday tasks. Bryce Ikeda, Maitrey Gramopadhye, LillyAnn Nekervis, Daniel Szafir |
HRI | 4 |
| 2025 | Supporting Long-Horizon Tasks in Human-Robot Collaboration by Aligning Intentions via Augmented RealityabstractHuman-involved robot learning has made significant strides in performing everyday tasks. However, long-horizon human-robot collaborative tasks remain challenging due to ambiguous subtask goals, which we refer to as intention misalignment between the user and the robot. To address this, we propose a novel human-robot collaboration (HRC) paradigm that utilizes augmented reality (AR) as a bidirectional communication channel. This channel enables the robot to communicate its intentions for ambiguous subtasks to users and adapt based on their feedback. To operationalize these aligned intentions as actionable goals, we employ a goal-conditioned reinforcement learning model, where the goal reflects the user-aligned intention. We validate our approach in a classic pick-and-place task involving three distinct objects, where the user specifies the desired goal object and its target pose. Yue Yang 0024, Bryce Ikeda, Daniel Szafir |
HRI | 4 |
| 2025 | ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video SynthesisabstractVision-language-action (VLA) models present a promising paradigm by training policies directly on real robot datasets like Open X-Embodiment. However, the high cost of real-world data collection hinders further data scaling, thereby restricting the generalizability of VLAs. In this paper, we introduce ReBot, a novel real-to-sim-to-real approach for scaling real robot datasets and adapting VLA models to target domains, which is the last-mile deployment challenge in robot manipulation. Specifically, ReBot replays real-world robot trajectories in simulation to diversify manipulated objects (real-to-sim), and integrates the simulated movements with inpainted real-world background to synthesize physically realistic and temporally consistent robot videos (sim-to-real). Our approach has several advantages: 1) it enjoys the benefit of real data to minimize the sim-to-real gap; 2) it leverages the scalability of simulation; and 3) it can generalize a pretrained VLA to a target domain with fully automated data pipelines. Extensive experiments in both simulation and real-world environments show that ReBot significantly enhances the performance and robustness of VLAs. For example, in SimplerEnv with the WidowX robot, ReBot improved the in-domain performance of Octo by 7.2% and OpenVLA by 21.8%, and out-of-domain generalization by 19.9% and 9.4%, respectively. For real-world evaluation with a Franka robot, ReBot increased the success rates of Octo by 17% and OpenVLA by 20%. More information can be found at our project page. Yue Yang 0024, Xinghao Zhu, Gedas Bertasius, Daniel Szafir, Mingyu Ding |
IROS | 6 |
| 2025 | Investigating Encoding and Perspective for Augmented Reality Motion GuidanceabstractAugmented reality (AR) offers promising opportunities to support movement-based activities, such as personal training or physical therapy, with real-time, spatially-situated visual cues. While many approaches leverage AR to guide motion, existing design guidelines focus on simple, upper-body movements within the user's field of view. We lack evidence-based design recommendations for guiding more diverse scenarios involving movements with varying levels of visibility and direction. We conducted an experiment to investigate how different visual encodings and perspectives affect motion guidance performance and usability, using three exercises that varied in visibility and planes of motion. Our findings reveal significant differences in preference and performance across designs. Notably, the best perspective varied depending on motion visibility and showing more information about the overall motion did not necessarily improve motion execution. We provide empirically-grounded guidelines for designing immersive, interactive visualizations for motion guidance to support more effective AR systems. Jade Kandel, Sriya Kasumarthi, Spiros Tsalikis, Chelsea Duppen, Daniel Szafir, Michael Lewek, Henry Fuchs, Danielle Albers Szafir |
ISMAR | 5 |
| 2024 | PD-Insighter: A Visual Analytics System to Monitor Daily Actions for Parkinson's Disease TreatmentabstractPeople with Parkinson's Disease (PD) can slow the progression of their symptoms with physical therapy. However, clinicians lack insight into patients' motor function during daily life, preventing them from tailoring treatment protocols to patient needs. This paper introduces PD-Insighter, a system for comprehensive analysis of a person's daily movements for clinical review and decision-making. PD-Insighter provides an overview dashboard for discovering motor patterns and identifying critical deficits during activities of daily living and an immersive replay for closely studying the patient's body movements with environmental context. Developed using an iterative design study methodology in consultation with clinicians, we found that PD-Insighter's ability to aggregate and display data with respect to time, actions, and local environment enabled clinicians to assess a person's overall functioning during daily life outside the clinic. PD-Insighter's design offers future guidance for generalized multiperspective body motion analytics, which may significantly improve clinical decision-making and slow the functional decline of PD and other medical conditions. Jade Kandel, Chelsea Duppen, Qian Zhang 0066, Howard Jiang, Angelos Angelopoulos, Ashley Paula-Ann Neall, Pranav Wagh, Daniel Szafir, Henry Fuchs, Michael Lewek, Danielle Albers Szafir |
CHI | 8 |
| 2024 | The Cyber-Physical Control Room: A Mixed Reality Interface for Mobile Robot Teleoperation and Human-Robot TeamingabstractIn this work, we present the design and evaluation of an immersive Cyber-Physical Control Room interface for remote mobile robots that provides users with both robot-egocentric and robot-exocentric 3D perspectives. We evaluate the Cyber-Physical Control room against a traditional robot interface in a mock disaster response scenario that features a mixed human-robot field team. In our evaluation, we found that the Cyber-Physical Control Room improved robot operator effectiveness by 28% while navigating a complex warehouse environment and performing a visual search. The Cyber-Physical Control Room also enhanced various aspects of human-robot teaming, including social engagement, the ability of a remote robot teleoperator to track their human partner in the field, and opinions of human teammate leadership qualities. Michael E. Walker, Maitrey Gramopadhye, Bryce Ikeda, Jack Burns, Daniel Szafir |
HRI | 5 |
| 2024 | ARCADE: Scalable Demonstration Collection and Generation via Augmented Reality for Imitation LearningabstractRobot Imitation Learning (IL) is a crucial technique in robot learning, where agents learn by mimicking human demonstrations. However, IL encounters scalability challenges stemming from both non-user-friendly demonstration collection methods and the extensive time required to amass a sufficient number of demonstrations for effective training. In response, we introduce the Augmented Reality for Collection and generAtion of DEmonstrations (ARCADE) framework, designed to scale up demonstration collection for robot manipulation tasks. Our framework combines two key capabilities: 1) it leverages AR to make demonstration collection as simple as users performing daily tasks using their hands, and 2) it enables the automatic generation of additional synthetic demonstrations from a single human-derived demonstration, significantly reducing user effort and time. We assess ARCADE’s performance on a real Fetch robot across three robotics tasks: 3-Waypoints-Reach, Push, and Pick-And-Place. Using our framework, we were able to rapidly train a policy using vanilla Behavioral Cloning (BC), a classic IL algorithm, which excelled across these three tasks. We also deploy ARCADE on a real household task, Pouring-Water, achieving an 80% success rate. Yue Yang 0024, Bryce Ikeda, Gedas Bertasius, Daniel Szafir |
IROS | 4 |
| 2024 | Incorporating Retakes in a Robotics Class with LabsabstractLab assignments, in which students build and program robots to accomplish tasks in various environments, are a central component in many undergraduate robotics classes. Such activities require that students operationalize concepts learned in class. However, physical robots are prone to uncertain real-world behavior, making debugging challenging and causing many students to feel stressed about being graded based on their robot's performance. Therefore, we incorporated retakes into our undergraduate robotics class, allowing students to learn from their mistakes and master class content while improving their robots. Initial results show that students widely embrace retakes, use the opportunity to improve, and feel less stressed about the assignments. Janine Hoelscher, Bryce Ikeda, Daniel Szafir, Ron Alterovitz |
SIGCSE (2) | 3 |
| 2024 | Assistance in Teleoperation of Redundant Robots through Predictive Joint ManeuveringabstractIn teleoperation of redundant robotic manipulators, translating an operator’s end effector motion command to joint space can be a tool for maintaining feasible and precise robot motion. Through optimizing redundancy resolution, the control system can ensure the end effector maintains maneuverability by avoiding joint limits and kinematic singularities. In autonomous motion planning, this optimization can be done over an entire trajectory to improve performance over local optimization. However, teleoperation involves a human-in-the-loop who determines the trajectory to be executed through a dynamic sequence of motion commands. We present two systems, Predictive Kinematic Control Tree and Predictive Kinematic Control Search, for utilizing a predictive model of operator commands to accomplish this redundancy resolution in a manner that considers future expected motion during teleoperation. Using a probabilistic model of operator commands allows optimization over an expected trajectory of future motion rather than consideration of local motion alone. Evaluation through a user study demonstrates improved control outcomes from this predictive redundancy resolution over minimum joint velocity solutions and inverse kinematics-based motion controllers. Connor Brooks, Wyatt Rees, Daniel Szafir |
ACM Trans. Hum. Robot Interact. | 3 |
| 2024 | PRogramAR: Augmented Reality End-User Robot ProgrammingabstractThe field of end-user robot programming seeks to develop methods that empower non-expert programmers to task and modify robot operations. In doing so, researchers may enhance robot flexibility and broaden the scope of robot deployments into the real world. We introduce PRogramAR (Programming Robots using Augmented Reality), a novel end-user robot programming system that combines the intuitive visual feedback of augmented reality (AR) with the simplistic and responsive paradigm of trigger-action programming (TAP) to facilitate human-robot collaboration. Through PRogramAR, users are able to rapidly author task rules and desired reactive robot behaviors, while specifying task constraints and observing program feedback contextualized directly in the real world. PRogramAR provides feedback by simulating the robot’s intended behavior and providing instant evaluation of TAP rule executability to help end users better understand and debug their programs during development. In a system validation, 17 end users ranging from ages 18 to 83 used PRogramAR to program a robot to assist them in completing three collaborative tasks. Our results demonstrate how merging the benefits of AR and TAP using elements from prior robot programming research into a single novel system can successfully enhance the robot programming process for non-expert users. Bryce Ikeda, Daniel Szafir |
ACM Trans. Hum. Robot Interact. | 2 |
| 2023 | Identifying the Focus of Attention in Human-Robot Conversational GroupsabstractWe propose a method for detecting the group’s focus of attention: the visual point at which a majority of participants direct their gaze in a conversation. This information enables a robot to infer important conversational cues and adjust its behavior to support more natural conversational interactions. Our approach uses a Hidden Markov Model based on mimicry, where the robot observes the head orientation of participants and infers their gaze direction to identify the group’s focus of attention. We demonstrate our method by replicating the gaze patterns of the group members, showing that the robot can accurately determine the focal point. We evaluated our algorithm using a combination of datasets and real-world scenarios with a Fetch robot, demonstrating an accuracy of 81% compared to a baseline of 54%. Our proposed method has the potential to significantly improve group-oriented human-robot interaction. Hooman Hedayati, Annika Muehlbradt, James Kennedy 0001, Daniel Szafir |
HAI | 4 |
| 2023 | Generating Executable Action Plans with Environmentally-Aware Language ModelsabstractLarge Language Models (LLMs) trained using massive text datasets have recently shown promise in generating action plans for robotic agents from high-level text queries. However, these models typically do not consider the robot's environment, resulting in generated plans that may not actually be executable, due to ambiguities in the planned actions or environmental constraints. In this paper, we propose an approach to generate environmentally-aware action plans that agents are better able to execute. Our approach involves integrating environmental objects and object relations as additional inputs into LLM action plan generation to provide the system with an awareness of its surroundings, resulting in plans where each generated action is mapped to objects present in the scene. We also design a novel scoring function that, along with generating the action steps and associating them with objects, helps the system disambiguate among object instances and take into account their states. We evaluated our approach using the VirtualHome simulator and the ActivityPrograms knowledge base and found that action plans generated from our system had a 310% improvement in executability and a 147% improvement in correctness over prior work. The complete code and a demo of our method is publicly available at https://github.com/hri-ironlab/scene_aware_language_planner. Maitrey Gramopadhye, Daniel Szafir |
IROS | 2 |
| 2023 | Virtual, Augmented, and Mixed Reality for Human-robot Interaction: A Survey and Virtual Design Element TaxonomyabstractVirtual, Augmented, and Mixed Reality for Human-Robot Interaction (VAM-HRI) has been gaining considerable attention in HRI research in recent years. However, the HRI community lacks a set of shared terminology and framework for characterizing aspects of mixed reality interfaces, presenting serious problems for future research. Therefore, it is important to have a common set of terms and concepts that can be used to precisely describe and organize the diverse array of work being done within the field. In this article, we present a novel taxonomic framework for different types of VAM-HRI interfaces, composed of four main categories of virtual design elements (VDEs). We present and justify our taxonomy and explain how its elements have been developed over the past 30 years as well as the current directions VAM-HRI is headed in the coming decade. Michael E. Walker, Thao Phung, Tathagata Chakraborti, Tom Williams 0001, Daniel Szafir |
ACM Trans. Hum. Robot Interact. | 5 |
| 2022 | Drone Brush: Mixed Reality Drone Path PlanningabstractIn this paper we present Drone Brush, a prototype mixed reality interface for immersive planning of drone paths for tasks such as collaborative photogrammetry and inspection. This interface employs Microsoft's HoloLens 2 to allow users to draw paths for drone navigation in 3D using hand gestures. Users can place waypoints with a simple pinch gesture, and similarly, delete and move existing waypoints. To validate paths, we leverage the HoloLens spatial map to check for potential collisions ahead of time, greatly reducing the likelihood of a collision during drone navigation. Paths are simplified and cleaned up using density-based clustering to prevent complex or redundant drone movement. In this Late-Breaking Report, we present the design and implementation of our system that integrates mixed reality, natural hand gestures, and drone path planning, which we plan to evaluate in a user study in the near future. Angelos Angelopoulos, Austin Hale, Husam Shaik, Akshay Paruchuri, Ziyu Liu 0002, Randal Tuggle, Daniel Szafir |
HRI | 7 |
| 2022 | Predicting Positions of People in Human-Robot Conversational GroupsabstractRobots that operate in social settings must be able to recognize, understand, and reason about human conversational groups (i.e., F-formations). While several algorithms have been developed for identifying such groups, there has been little research on how robots might reason about inaccuracies following group classification (e.g., recognizing only 4 of 5 group members). We address this gap through a data-driven approach that builds knowledge of human group positioning. By analyzing multiple conversational group data sets, we have developed a system for identifying high probability regions that indicate areas where people are likely to stand in a group relative to a single anchor participant. We use knowledge of these regions to train two models, which we implement on a social robot. The first model can estimate the true size of a partially-observed conversational group (i.e., a group where only some of the participants were detected). Our second model can predict the locations where any undetected participants are likely to reside. Together, these mod-els may improve F-formation detection algorithms by increasing robustness to noisy input data. Hooman Hedayati, Daniel Szafir |
HRI | 2 |
| 2022 | Advancing the Design of Visual Debugging Tools for RoboticistsabstractProgramming robots is a challenging task exacer-bated by software bugs, faulty hardware, and environmental fac-tors. When coding issues arise, traditional debugging techniques such as output logs or print statements that may help in typical computer applications are not always useful for roboticists. As a result, roboticists often leverage visualizations that depict various aspects of robot, sensor, and environment states. In this paper, we explore various design approaches towards such visualizations for robotics debugging support, including 3D visualizations presented on 2D displays, as in the popular RViz tool within the ROS ecosystem, visualizations in a two-dimensional graphical user interfaces (2D GUI), and emerging immersive three-dimensional (3D) augmented reality (AR). We present a qualitative evaluation of feedback gathered from 24 roboticists across two universities who used one of these debugging tools and synthesize design guidelines for advancing robotics debugging interfaces. Bryce Ikeda, Daniel Szafir |
HRI | 2 |
| 2022 | The Predictive Kinematic Control Tree: Enhancing Teleoperation of Redundant Robots through Probabilistic User ModelsabstractWhen teleoperating complex robotic manipula-tors, operators often find it most natural to issue commands that dictate end effector movements in task space. If the robot has redundant degrees of freedom, the translation of this com-mand from task space into configuration space can affect the robot's maneuverability, smoothness of motion, and the general precision of the teleoperated system. In this paper, we propose a novel method for performing this translation that predicts future operator commands in order to choose joint motions that maintain maneuverability in future timesteps. We introduce a Predictive Kinematic Control Tree (PrediKCT) that optimizes joint movement in the nullspace of the Jacobian over multiple future timesteps by reasoning over probabilistic models of the human operator. In essence, PrediKCT builds out and evaluates a tree of possible future commands. We implement this system on two simulated and one physical 7 -degree-of-freedom robotic arms and characterize performance by analyzing robot motions produced through multiple command trajectories with differing user model accuracies and tree parameters, demonstrating benefits to path accuracy over both a minimum-norm joint velocity solution and local optimization of joint movement. Connor Brooks, Daniel Szafir |
IROS | 2 |
| 2022 | "I'mConfident This Will End Poorly": Robot Proficiency Self-Assessment in Human-Robot TeamingabstractHuman-robot teams are expected to accomplish complex tasks in high-risk and uncertain environments. In domains such as space exploration or search & rescue, a human operator may not be a robotics expert, but will need to establish a baseline understanding of the robot's capabilities with respect to a given task in order to appropriately utilize and rely on the robot. This willingness to rely, also known as trust, is based partly on the operator's belief in the robot's task proficiency. If trust is too high, the operator may unknowingly push the robot beyond its capabilities. If trust is too low, the operator may not utilize it when they otherwise could have, wasting precious time and resources. In this work, we discuss results from an online human-subjects study investigating how a robot communicated report of its task proficiency with respect to an operator's expectations affects trust and performance in a navigation task. Our results show that communication of a robot self-assessment helped operators understand when reliance on the robot was appropriate given the task and conditions. This led to improvements in task performance, informed choices of autonomy level, and increased trust. Nicholas Conlon, Daniel Szafir, Nisar R. Ahmed |
IROS | 2 |
| 2021 | Connecting Human-Robot Interaction and Data VisualizationabstractHuman-robot interaction (HRI) research frequently explores how to design interfaces that enable humans to effectively teleoperate and supervise robots. One of the principle goals of such systems is to support data collection, analysis, and human decision making, which requires representing robot data in ways that support fast and accurate analyses by humans. However, the interfaces for these systems do not always use best-practice principles for effectively visualizing data. We present a new framework to scaffold reasoning about robot interface design that emphasizes the need to consider data visualization for supporting analysis and decision making processes, detail several data visualization best practices relevant to HRI, identify a set of core data tasks that commonly occur in HRI, and highlight several promising opportunities for further synergistic activities at the intersection of these two research areas. Daniel Szafir, Danielle Albers Szafir |
HRI | 1 |
| 2021 | ARC-LfD: Using Augmented Reality for Interactive Long-Term Robot Skill Maintenance via Constrained Learning from DemonstrationabstractLearning from Demonstration (LfD) enables novice users to teach robots new skills. However, many LfD methods do not facilitate skill maintenance and adaptation. Changes in task requirements or in the environment often reveal the lack of resiliency and adaptability in the skill model. To overcome these limitations, we introduce ARC-LfD: an Augmented Reality (AR) interface for constrained Learning from Demonstration that allows users to maintain, update, and adapt learned skills. This is accomplished through in-situ visualizations of learned skills and constraint-based editing of existing skills without requiring further demonstration. We describe the existing algorithmic basis for this system as well as our Augmented Reality interface and the novel capabilities it provides. Finally, we provide three case studies that demonstrate how ARC-LfD enables users to adapt to changes in the environment or task which require a skill to be altered after initial teaching has taken place. Matthew B. Luebbers, Connor Brooks, Carl L. Mueller, Daniel Szafir, Bradley Hayes |
ICRA | 4 |
| 2021 | What Information Should a Robot Convey?abstractRobotic technologies are becoming pervasive within industrial and domestic settings, resulting in more frequent interactions between humans and robots. To ensure these interactions are effective, Human-Robot Interaction (HRI) researchers have argued that robots and humans must establish a shared common ground by communicating fundamental pieces of information to each other, such as their intentions, goals, plans, status, etc. Although a large body of work has explored how robots might signal individual aspects of such information to users, we still know relatively little regarding the importance of such information overall (e.g., is communicating robot status more important than communicating robot goals?). Such information is necessary for robots acting in the wild to create prioritized lists of communicative goals as, at any given time, it is unlikely that a robot will be able to convey all possibly relevant or important aspects of information to users. Prioritizing information for users is a complex problem as many factors might influence information priority, including task context, user expertise, and robot capability. In this work, we first address the current state-of-the-art signaling methods for non-humanoid robots. Second, we take an initial step towards understanding prioritization by exploring what types of information users request, and how the rankings of informational importance that users assign change, in a prototypical shared-environment interaction with three different types of robots. Our results, collected from 150 participants on Amazon’s Mechanical Turk, generally show that users value information related to the robot’s battery, capabilities, task, safety, navigation, communication, and privacy, with user priorities of these items varying across a small ground robot, a large ground robot, and an aerial robot. Hooman Hedayati, Mark D. Gross, Daniel Szafir |
IROS | 3 |
| 2021 | A Mixed Reality Supervision and Telepresence Interface for Outdoor Field RoboticsabstractCollaborative human-robot field operations rely on timely decision-making and coordination, which can be challenging for heterogeneous teams operating in large-scale deployments. In this work, we present the design of an immersive, mixed reality (MR) interface to support sense-making and situational awareness based on the data collection capabilities of both human and robotic team members. Our solution integrates state-of-the-art methods in environment mapping and MR so that users may gain rapid insights regarding the working environment, the current and previous locations of human and robot team members, and the environment data such team members have collected. We describe the implementation of our system, share lessons learned in collaborating with emergency responders throughout our design process, and offer a vision for the use of immersive displays for human-robot field team deployments in large-scale outdoor environments. Michael E. Walker, Zhaozhong Chen, Matt Whitlock, David Blair, Danielle Albers Szafir, Christoffer R. Heckman, Daniel Szafir |
IROS | 7 |
| 2020 | RoomShift: Room-scale Dynamic Haptics for VR with Furniture-moving Swarm RobotsabstractRoomShift is a room-scale dynamic haptic environment for virtual reality, using a small swarm of robots that can move furniture. RoomShift consists of nine shape-changing robots: Roombas with mechanical scissor lifts. These robots drive beneath a piece of furniture to lift, move and place it. By augmenting virtual scenes with physical objects, users can sit on, lean against, place and otherwise interact with furniture with their whole body; just as in the real world. When the virtual scene changes or users navigate within it, the swarm of robots dynamically reconfigures the physical environment to match the virtual content. We describe the hardware and software implementation, applications in virtual tours and architectural design and interaction techniques. Ryo Suzuki 0001, Hooman Hedayati, Clement Zheng, James L. Bohn, Daniel Szafir, Ellen Yi-Luen Do, Mark D. Gross, Daniel Leithinger |
CHI | 5 |
| 2020 | Visualization of Intended Assistance for Acceptance of Shared ControlabstractIn shared control, advances in autonomous robotics are applied to help empower a human user in operating a robotic system. While these systems have been shown to improve efficiency and operation success, users are not always accepting of the new control paradigm produced by working with an assistive controller. This mismatch between performance and acceptance can prevent users from taking advantage of the benefits of shared control systems for robotic operation. To address this mismatch, we develop multiple types of visualizations for improving both the legibility and perceived predictability of assistive controllers, then conduct a user study to evaluate the impact that these visualizations have on user acceptance of shared control systems. Our results demonstrate that shared control visualizations must be designed carefully to be effective, with users requiring visualizations that improve both legibility and predictability of the assistive controller in order to voluntarily relinquish control. Connor Brooks, Daniel Szafir |
IROS | 2 |
| 2020 | REFORM: Recognizing F-formations for Social RobotsabstractRecognizing and understanding conversational groups, or F-formations, is a critical task for situated agents designed to interact with humans. F-formations contain complex structures and dynamics, yet are used intuitively by people in everyday face-to-face conversations. Prior research exploring ways of identifying F-formations has largely relied on heuristic algorithms that may not capture the rich dynamic behaviors employed by humans. We introduce REFORM (REcognize F-FORmations with Machine learning), a data-driven approach for detecting F-formations given human and agent positions and orientations. REFORM decomposes the scene into all possible pairs and then reconstructs F-formations with a voting-based scheme. We evaluated our approach across three datasets: the SALSA dataset, a newly collected human-only dataset, and a new set of acted human-robot scenarios, and found that REFORM yielded improved accuracy over a state-of-the-art F-formation detection algorithm. We also introduce symmetry and tightness as quantitative measures to characterize F-formations. Hooman Hedayati, Annika Muehlbradt, Daniel Szafir, Sean Andrist |
IROS | 3 |
| 2020 | PufferBot: Actuated Expandable Structures for Aerial RobotsabstractWe present PufferBot, an aerial robot with an expandable structure that may expand to protect a drone's propellers when the robot is close to obstacles or collocated humans. PufferBot is made of a custom 3D-printed expandable scissor structure, which utilizes a one degree of freedom actuator with rack and pinion mechanism. We propose four designs for the expandable structure, each with unique characterizations for different situations. Finally, we present three motivating scenarios in which PufferBot may extend the utility of existing static propeller guard structures. Hooman Hedayati, Ryo Suzuki 0001, Daniel Leithinger, Daniel Szafir |
IROS | 4 |
| 2019 | HugBot: A soft robot designed to give human-like hugsabstractAs robots increasingly enter our daily lives, there is a need to understand how to design robots capable of emotional interaction with humans, especially children, due to their sensitivity and vulnerability. For example, robots that provide children with social and emotional support might be more effective at also helping children develop cognitive abilities, rather than designing robots that focus solely on helping children acquire cognitive skill. In this paper, we examine the design of robots that can provide human-like hugs as a particular form of social and emotional support. We first discuss the need to design robots that can interact emotionally with children. Then, we present the development of a shirt augmented with pressure sensors used to collect data on how humans hug each other. Finally, we detail the design of "Hugbot", a soft robot that could use this data to give human-like hugs, and discuss our planned future work on this system. Hooman Hedayati, Srinjita Bhaduri, Tamara Sumner, Daniel Szafir, Mark D. Gross |
IDC | 4 |
| 2019 | RoboGraphics: Dynamic Tactile Graphics Powered by Mobile RobotsabstractTactile graphics are a common way to present information to people with vision impairments. Tactile graphics can be used to explore a broad range of static visual content but aren't well suited to representing animation or interactivity. We introduce a new approach to creating dynamic tactile graphics that combines a touch screen tablet, static tactile overlays, and small mobile robots. We introduce a prototype system called RoboGraphics and several proof-of-concept applications. We evaluated our prototype with seven participants with varying levels of vision, comparing the RoboGraphics approach to a flat screen, audio-tactile interface. Our results show that dynamic tactile graphics can help visually impaired participants explore data quickly and accurately. Darren Guinness, Annika Muehlbradt, Daniel Szafir, Shaun K. Kane |
ASSETS | 3 |
| 2019 | RoboGraphics: Using Mobile Robots to Create Dynamic Tactile GraphicsabstractTactile graphics are a common way to present information to people with vision impairments. Tactile graphics can be used to explore a broad range of content including presenting data, telling stories, and conveying map information, but aren't well suited to representing dynamic or moving information. We introduce and demonstrate a new approach RoboGraphics for creating dynamic tactile graphics by combining static tactile overlays, touch screen tablets, and off-the-shelf tangible robots. To evaluate the RoboGraphics approach for conveying information to blind users, we created a set of reference applications, including an interface for exploring graph data, braille characters, the story of the tortoise and the hare, the analog clock, and the ruminant digestive process. A design probe demonstrated that haptic graphics can be used to help people with vision impairments explore data using the audio-tactile display. Darren Guinness, Annika Muehlbradt, Daniel Szafir, Shaun K. Kane |
ASSETS | 3 |
| 2019 | The Reality-Virtuality Interaction Cube: A Framework for Conceptualizing Mixed-Reality Interaction Design Elements for HRIabstractThere has recently been an explosion of work in the human-robot interaction (HRI) community on the use of mixed, augmented, and virtual reality. We present a novel conceptual framework to characterize and cluster work in this new area and identify gaps for future research. We begin by introducing the Plane of Interaction: a framework for characterizing interactive technologies in a 2D space informed by the Model-View-Controller design pattern. We then describe how Interactive Design Elements that contribute to the interactivity of a technology can be characterized within this space and present a taxonomy of mixed-reality interactive design elements. We then discuss how these elements may be rendered onto both reality- and virtuality-based environments using a variety of hardware devices and introduce the Reality-Virtuality Interaction Cube: a three-dimensional continuum representing the design space of interactive technologies formed by combining the Plane of Interaction with the Reality-Virtuality Continuum. Finally, we demonstrate the feasibility and utility of this framework by clustering and analyzing the set of papers presented at the 2018 VAM-HRI workshop. Tom Williams 0001, Daniel Szafir, Tathagata Chakraborti |
HRI | 2 |
| 2019 | Virtual, Augmented, and Mixed Reality for Human-Robot Interaction (VAM-HRI)abstractThe 2ndInternational Workshop on Virtual, Augmented, and Mixed Reality for Human-Robot Interactions (VAM-HRI) will bring together HRI, Robotics, and Mixed Reality researchers to identify challenges in mixed reality interactions between humans and robots. Topics relevant to the workshop include development of robots that can interact with humans in mixed reality, use of virtual reality for developing interactive robots, the design of new augmented reality interfaces that mediate communication between humans and robots, comparisons of the capabilities and perceptions of robots and virtual agents, and best design practices. VAM-HRI was held for the first time at HRI 2018, where it served as the first workshop of its kind at an academic AI or Robotics conference, and served as a timely call to arms to the academic community in response to the growing promise of this emerging field. VAM-HRI 2019 will follow on the success of VAM-HRI 2018, and present new opportunities for expanding this nascent research community. Tom Williams 0001, Daniel Szafir, Tathagata Chakraborti, Elizabeth Phillips |
HRI | 2 |
| 2019 | Balanced Information Gathering and Goal-Oriented Actions in Shared AutonomyabstractRobotic teleoperation can be a complex task due to factors such as high degree-of-freedom manipulators, operator inexperience, and limited operator situational awareness. To reduce teleoperation complexity, researchers have developed the shared autonomy control paradigm that involves joint control of a robot by a human user and an autonomous control system. We introduce the concept of active learning into shared autonomy by developing a method for systems to leverage information gathering: minimizing the system's uncertainty about user goals by moving to information-rich states to observe user input. We create a framework for balancing information gathering actions, which help the system gain information about user goals, with goal-oriented actions, which move the robot towards the goal the system has inferred from the user. We conduct an evaluation within the context of users who are multitasking that compares pure teleoperation with two forms of shared autonomy: our balanced system and a traditional goal-oriented system. Our results show significant improvements for both shared autonomy systems over pure teleoperation in terms of belief convergence about the user's goal and task completion speed and reveal trade-offs across shared autonomy strategies that may inform future investigations in this space. Connor Brooks, Daniel Szafir |
HRI | 2 |
| 2019 | Recognizing F-Formations in the Open WorldabstractA key skill for social robots in the wild will be to understand the structure and dynamics of conversational groups in order to fluidly participate in them. Social scientists have long studied the rich complexity underlying such focused encounters, or F-formations. However, current state-of-the-art algorithms that robots might use to recognize F-formations are highly heuristic and quite brittle. In this report, we explore a data-driven approach to detect F-formations from sets of tracked human positions and orientations, trained and evaluated on two openly available human-only datasets and a small human-robot dataset that we collected. We also discuss the potential for further computational characterization of F-formations beyond simply detecting their occurrence. Hooman Hedayati, Daniel Szafir, Sean Andrist |
HRI | 2 |
| 2019 | Robot Teleoperation with Augmented Reality Virtual SurrogatesabstractTeleoperation remains a dominant control paradigm for human interaction with robotic systems. However, teleoperation can be quite challenging, especially for novice users. Even experienced users may face difficulties or inefficiencies when operating a robot with unfamiliar and/or complex dynamics, such as industrial manipulators or aerial robots, as teleoperation forces users to focus on low-level aspects of robot control, rather than higher level goals regarding task completion, data analysis, and problem solving. We explore how advances in augmented reality (AR) may enable the design of novel teleoperation interfaces that increase operation effectiveness, support the user in conducting concurrent work, and decrease stress. Our key insight is that AR may be used in conjunction with prior work on predictive graphical interfaces such that a teleoperator controls a virtual robot surrogate, rather than directly operating the robot itself, providing the user with foresight regarding where the physical robot will end up and how it will get there. We present the design of two AR interfaces using such a surrogate: one focused on real-time control and one inspired by waypoint delegation. We compare these designs against a baseline teleoperation system in a laboratory experiment in which novice and expert users piloted an aerial robot to inspect an environment and analyze data. Our results revealed that the augmented reality prototypes provided several objective and subjective improvements, demonstrating the promise of leveraging AR to improve human-robot interactions. Michael E. Walker, Hooman Hedayati, Daniel Szafir |
HRI | 3 |
| 2019 | The Influence of Size in Augmented Reality Telepresence AvatarsabstractIn this work, we explore how advances in augmented reality technologies are creating a new design space for long-distance telepresence communication through virtual avatars. Studies have shown that the relative size of a speaker has a significant impact on many aspects of human communication including perceived dominance and persuasiveness. Our system synchronizes the body pose of a remote user with a realistic, virtual human avatar visible to a local user wearing an augmented reality head-mounted display. We conducted a two-by-two (relative system size: equivalent vs. small; leader vs. follower), between participants study (N = 40) to investigate the effect of avatar size on the interactions between remote and local user. We found the equal-sized avatars to be significantly more influential than the small-sized avatars and that the small avatars commanded significantly less attention than the equal-sized avatars. Additionally, we found the assigned leadership role to significantly impact participant subjective satisfaction of the task outcome. Michael E. Walker, Daniel Szafir, Irene Rae |
VR | 2 |
| 2018 | Improving Collocated Robot Teleoperation with Augmented RealityabstractRobot teleoperation can be a challenging task, often requiring a great deal of user training and expertise, especially for platforms with high degrees-of-freedom (e.g., industrial manipulators and aerial robots). Users often struggle to synthesize information robots collect (e.g., a camera stream) with contextual knowledge of how the robot is moving in the environment. We explore how advances in augmented reality (AR) technologies are creating a new design space for mediating robot teleoperation by enabling novel forms of intuitive, visual feedback. We prototype several aerial robot teleoperation interfaces using AR, which we evaluate in a 48-participant user study where participants completed an environmental inspection task. Our new interface designs provided several objective and subjective performance benefits over existing systems, which often force users into an undesirable paradigm that divides user attention between monitoring the robot and monitoring the robot»s camera feed(s). Hooman Hedayati, Michael E. Walker, Daniel Szafir |
HRI | 3 |
| 2018 | Communicating Robot Motion Intent with Augmented RealityabstractHumans coordinate teamwork by conveying intent through social cues, such as gestures and gaze behaviors. However, these methods may not be possible for appearance-constrained robots that lack anthropomorphic or zoomorphic features, such as aerial robots. We explore a new design space for communicating robot motion intent by investigating how augmented reality (AR) might mediate human-robot interactions. We develop a series of explicit and implicit designs for visually signaling robot motion intent using AR, which we evaluate in a user study. We found that several of our AR designs significantly improved objective task efficiency over a baseline in which users only received physically-embodied orientation cues. In addition, our designs offer several trade-offs in terms of intent clarity and user perceptions of the robot as a teammate. Michael E. Walker, Hooman Hedayati, Jennifer Lee, Daniel Szafir |
HRI | 4 |
| 2018 | Improving Object Disambiguation from Natural Language using Empirical ModelsabstractRobots, virtual assistants, and other intelligent agents need to effectively interpret verbal references to environmental objects in order to successfully interact and collaborate with humans in complex tasks. However, object disambiguation can be a challenging task due to ambiguities in natural language. To reduce uncertainty when describing an object, humans often use a combination of unique object features and locative prepositions --prepositional phrases that describe where an object is located relative to other features (i.e., reference objects) in a scene. We present a new system for object disambiguation in cluttered environments based on probabilistic models of unique object features and spatial relationships. Our work extends prior models of spatial relationship semantics by collecting and encoding empirical data from a series of crowdsourced studies to better understand how and when people use locative prepositions, how reference objects are chosen, and how to model prepositional geometry in 3D space (e.g., capturing distinctions between "next to" and "beside"). Our approach also introduces new techniques for responding to compound locative phrases of arbitrary complexity and proposes a new metric for disambiguation confidence. An experimental validation revealed our method can improve object disambiguation accuracy and performance over past approaches. Daniel Prendergast, Daniel Szafir |
ICMI | 2 |
| 2018 | Proactive Robot Assistants for Freeform Collaborative Tasks Through Multimodal Recognition of Generic SubtasksabstractSuccessful human-robot collaboration depends on a shared understanding of task state and current goals. In nonlinear or freeform tasks without an explicit task model, robot partners are unable to provide assistance without the ability to translate perception into meaningful task knowledge. In this paper, we explore the utility of multimodal recurrent neural networks (RNNs) with long short-term memory (LSTM) units for real-time subtask recognition in order to provide context-aware assistance during generic assembly tasks. We train RNNs to recognize specific subtasks in individual modalities, then combine the high-level representations of these networks through a nonlinear connection layer to create a multimodal subtask recognition system. We report results from implementing the system on a robot that uses the subtask recognition system to provide predictive assistance to a human partner during a laboratory experiment involving a human-robot team completing an assembly task. Generalizability of the system is evaluated through training and testing on separate tasks with some similar subtasks. Our results demonstrate the value of such a system in providing assistance to human partners during a freeform assembly scenario and increasing humans' perception of the robot's agency and usefulness. Connor Brooks, Madhur Atreya, Daniel Szafir |
IROS | 3 |
| 2018 | Virtual-to-Real-World Transfer Learning for Robots on Wilderness TrailsabstractRobots hold promise in many scenarios involving outdoor use, such as search-and-rescue, wildlife management, and collecting data to improve environment, climate, and weather forecasting. However, autonomous navigation of outdoor trails remains a challenging problem. Recent work has sought to address this issue using deep learning. Although this approach has achieved state-of-the-art results, the deep learning paradigm may be limited due to a reliance on large amounts of annotated training data. Collecting and curating training datasets may not be feasible or practical in many situations, especially as trail conditions may change due to seasonal weather variations, storms, and natural erosion. In this paper, we explore an approach to address this issue through virtual-to-real-world transfer learning using a variety of deep learning models trained to classify the direction of a trail in an image. Our approach utilizes synthetic data gathered from virtual environments for model training, bypassing the need to collect a large amount of real images of the outdoors. We validate our approach in three main ways. First, we demonstrate that our models achieve classification accuracies upwards of 95% on our synthetic data set. Next, we utilize our classification models in the control system of a simulated robot to demonstrate feasibility. Finally, we evaluate our models on real-world trail data and demonstrate the potential of virtual-to-real-world transfer learning. Michael L. Iuzzolino, Michael E. Walker, Daniel Szafir |
IROS | 3 |
| 2018 | The Haptic Video Player: Using Mobile Robots to Create Tangible Video AnnotationsabstractVideo and animation are common ways of delivering concepts that cannot be easily communicated through text. This visual information is often inaccessible to blind and visually impaired people, and alternative representations such as Braille and audio may leave out important details. Audio-haptic displays with along with supplemental descriptions allow for the presentation of complex spatial information, along with accompanying description. We introduce the Haptic Video Player, a system for authoring and presenting audio-haptic content from videos. The Haptic Video Player presents video using mobile robots that can be touched as they move over a touch screen. We describe the design of the Haptic Video Player system, and present user studies with educators and blind individuals that demonstrate the ability of this system to render dynamic visual content non-visually. Darren Guinness, Annika Muehlbradt, Daniel Szafir, Shaun K. Kane |
ISS | 3 |
| 2017 | GUI Robots: Using Off-the-Shelf Robots as Tangible Input and Output Devices for Unmodified GUI ApplicationsabstractTraditional GUI applications provide limited support for tangible interaction, as most applications are not programmed to support tangible input, and most input devices do not provide haptic feedback. To address this limitation, we introduce GUI Robots, a software framework that enables developers to repurpose off-the-shelf robots as tangible input and haptic output devices, and to connect them to unmodified desktop applications. We introduce the GUI Robots framework and present several proof-of-concept applications, including a haptic scroll wheel, force feedback game controllers, a 3D mouse, and a self-driving notification robot. To evaluate whether GUI Robots can be used to prototype tangible interfaces for existing applications, we conducted a user study in which developers created customized tangible interfaces for two applications. Study participants were able to create tangible user interfaces for these applications in less than an hour. GUI Robots allows developers to easily extend applications with tangible input and haptic output. Darren Guinness, Daniel Szafir, Shaun K. Kane |
Conference on Designing Interactive Systems | 2 |
| 2017 | Designing for Depth Perceptions in Augmented RealityabstractAugmented reality technologies allow people to view and interact with virtual objects that appear alongside physical objects in the real world. For augmented reality applications to be effective, users must be able to accurately perceive the intended real world location of virtual objects. However, when creating augmented reality applications, developers are faced with a variety of design decisions that may affect user perceptions regarding the real world depth of virtual objects. In this paper, we conducted two experiments using a perceptual matching task to understand how shading, cast shadows, aerial perspective, texture, dimensionality (i.e., 2D vs. 3D shapes) and billboarding affected participant perceptions of virtual object depth relative to real world targets. The results of these studies quantify trade-offs across virtual object designs to inform the development of applications that take advantage of users' visual abilities to better blend the physical and virtual world. Catherine Diaz, Michael E. Walker, Danielle Albers Szafir, Daniel Szafir |
ISMAR | 4 |
| 2015 | Communicating Directionality in Flying RobotsabstractSmall flying robots represent a rapidly emerging family of robotic technologies with aerial capabilities that enable unique forms of assistance in a variety of collaborative tasks. Such tasks will necessitate interaction with humans in close proximity, requiring that designers consider human perceptions regarding robots flying and acting within human environments. We explore the design space regarding explicit robot communication of flight intentions to nearby viewers. We apply design constraints to robot flight behaviors, using biological and airplane flight as inspiration, and develop a set of signaling mechanisms for visually communicating directionality while operating under such constraints. We implement our designs on two commercial flyers, requiring little modification to the base platforms, and evaluate each signaling mechanism, as well as a no-signaling baseline, in a user study in which participants were asked to predict robot intent. We found that three of our designs significantly improved viewer response time and accuracy over the baseline and that the form of the signal offered tradeoffs in precision, generalizability, and perceived robot usability. Daniel Szafir, Bilge Mutlu, Terrence Fong |
HRI | 1 |
| 2015 | From 9 to 90: Engaging Learners of All AgesabstractThis paper details the creation of a two-day computer science and robotics outreach course aimed at simultaneously engaging youth (children, ages 9-14) and senior (their grandparents, ages 55+) students. Our goal is to encourage enthusiasm for science and technology in students of all ages as well as provide practical instruction regarding common computer science concepts, including variables, loops, and boolean logic. To this end, we ground our course in the emerging field of social robotics, which enables the design of several multidisciplinary hands-on activities for students. We report on a four-year experience in the development of our course, which has been offered twelve times and involved over 210 youth and senior students. Our work presents a discussion regarding the challenges in designing a course for students from diverse ages, guidelines for creating similar courses, and a reflection on how we might improve our own class. The activities and project code developed for our course are available online as open-source resources. Allison Sauppé, Daniel Szafir, Chien-Ming Huang 0001, Bilge Mutlu |
SIGCSE | 2 |
| 2014 | Communication of intent in assistive free flyersabstractAssistive free-flyers (AFFs) are an emerging robotic platform with unparalleled flight capabilities that appear uniquely suited to exploration, surveillance, inspection, and telepresence tasks. However, unconstrained aerial movements may make it difficult for colocated operators, collaborators, and observers to understand AFF intentions, potentially leading to difficulties understanding whether operator instructions are being executed properly or to safety concerns if future AFF motions are unknown or difficult to predict. To increase AFF usability when working in close proximity to users, we explore the design of natural and intuitive flight motions that may improve AFF abilities to communicate intent while simultaneously accomplishing task goals. We propose a formalism for representing AFF flight paths as a series of motion primitives and present two studies examining the effects of modifying the trajectories and velocities of these flight primitives based on natural motion principles. Our first study found that modified flight motions might allow AFFs to more effectively communicate intent and, in our second study, participants preferred interacting with an AFF that used a manipulated flight path, rated modified flight motions as more natural, and felt safer around an AFF with modified motion. Our proposed formalism and findings highlight the importance of robot motion in achieving effective human-robot interactions. Daniel Szafir, Bilge Mutlu, Terrence Fong |
HRI | 1 |
| 2013 | ARTFul: adaptive review technology for flipped learningabstractInternet technology is revolutionizing education. Teachers are developing massive open online courses (MOOCs) and using innovative practices such as flipped learning in which students watch lectures at home and engage in hands-on, problem solving activities in class. This work seeks to explore the design space afforded by these novel educational paradigms and to develop technology for improving student learning. Our design, based on the technique of adaptive content review, monitors student attention during educational presentations and determines which lecture topic students might benefit the most from reviewing. An evaluation of our technology within the context of an online art history lesson demonstrated that adaptively reviewing lesson content improved student recall abilities 29% over a baseline system and was able to match recall gains achieved by a full lesson review in less time. Our findings offer guidelines for a novel design space in dynamic educational technology that might support both teachers and online tutoring systems. Daniel Szafir, Bilge Mutlu |
CHI | 1 |
| 2012 | Pay attention!: designing adaptive agents that monitor and improve user engagementabstractEmbodied agents hold great promise as educational assistants, exercise coaches, and team members in collaborative work. These roles require agents to closely monitor the behavioral, emotional, and mental states of their users and provide appropriate, effective responses. Educational agents, for example, will have to monitor student attention and seek to improve it when student engagement decreases. In this paper, we draw on techniques from brain-computer interfaces (BCI) and knowledge from educational psychology to design adaptive agents that monitor student attention in real time using measurements from electroencephalography (EEG) and recapture diminishing attention levels using verbal and nonverbal cues. An experimental evaluation of our approach showed that an adaptive robotic agent employing behavioral techniques to regain attention during drops in engagement improved student recall abilities 43% over the baseline regardless of student gender and significantly improved female motivation and rapport. Our findings offer guidelines for developing effective adaptive agents, particularly for educational settings. Daniel Szafir, Bilge Mutlu |
CHI | 1 |
| 2011 | An Exploration of the Utilization of Electroencephalography and Neural Nets to Control Robots
Daniel Szafir, Robert Signorile |
INTERACT (4) | 1 |