VLDB 2026 Research / reviewers in the wild / expert
Justin W. Hart
dblp:29/5724 · also Justin Wildrick Hart
· DBLP profile ↗
20ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-4317-4129ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perceived Social Intelligence and Human Compliance in Incidental Human-Robot EncountersabstractABSTRACT This study examines how socially compliant robot behaviours, operationalized as body language and verbal cues, affect human perceptions of social intelligence and compliance with a quadruped robot during incidental human–robot encounters. Extending prior work that primarily focused on direct human–robot interactions, this study addresses the understudied context of incidental encounters, which are brief, unplanned interactions between bystanders and autonomous robots in public environments. In an online video‐based study with 385 participants using a within‐subject design, we found that both verbal and body language behaviours significantly improved the robot's perceived social intelligence (PSI), which in turn positively correlated with human compliance. Verbal communication alone increased compliance likelihood more than body language, and combining both yielded the highest PSI ratings and compliance scores. To validate these findings beyond controlled video scenarios, a follow‐up real‐world study with 26 participants replicated the results: 11 out of 13 of participants complied in the Body Language + Verbal condition versus only 1 in the baseline condition. Free‐text responses revealed the importance of clearly stated intentions, politeness, and concerns about robot legitimacy and safety. Together, these findings highlight the critical role of perceived social intelligence in fostering human assistance for robots during incidental encounters and offer design strategies to support robot deployment in public environments. Yao-Cheng Chan, Elliott Hauser, Sadanand Modak, Joydeep Biswas, Justin W. Hart |
Expert Syst. J. Knowl. Eng. | 5 |
| 2025 | Principles and Guidelines for Evaluating Social Robot Navigation AlgorithmsabstractA major challenge to deploying robots widely is navigation in human-populated environments, commonly referred to as social robot navigation . While the field of social navigation has advanced tremendously in recent years, the fair evaluation of algorithms that tackle social navigation remains hard because it involves not just robotic agents moving in static environments but also dynamic human agents and their perceptions of the appropriateness of robot behavior. In contrast, clear, repeatable, and accessible benchmarks have accelerated progress in fields like computer vision, natural language processing and traditional robot navigation by enabling researchers to fairly compare algorithms, revealing limitations of existing solutions and illuminating promising new directions. We believe the same approach can benefit social navigation. In this article, we pave the road toward common, widely accessible, and repeatable benchmarking criteria to evaluate social robot navigation. Our contributions include (a) a definition of a socially navigating robot as one that respects the principles of safety, comfort, legibility, politeness, social competency, agent understanding, proactivity, and responsiveness to context, (b) guidelines for the use of metrics, development of scenarios, benchmarks, datasets, and simulators to evaluate social navigation, and (c) a design of a social navigation metrics framework to make it easier to compare results from different simulators, robots, and datasets. Anthony G. Francis, Claudia Pérez-D'Arpino, Chengshu Li 0002, Fei Xia 0002, Alexandre Alahi, Rachid Alami 0001, Aniket Bera, Abhijat Biswas, Joydeep Biswas, Rohan Chandra, Hao-Tien Chiang, Michael Everett, Sehoon Ha, Justin W. Hart, Jonathan P. How, Haresh Karnan, Tsang-Wei Edward Lee, Luis Manso, Reuth Mirsky, Sören Pirk, Phani-Teja Singamaneni, Peter Stone 0001, Ada V. Taylor, Pete Trautman, Nathan Tsoi, Marynel Vázquez, Xuesu Xiao, Peng Xu 0010, Naoki Yokoyama, Alexander Toshev, Roberto Martin Martin |
ACM Trans. Hum. Robot Interact. | 14 |
| 2024 | Recovering Missed Detections in an Elevator Button Segmentation TaskabstractOne obstacle that mobile service robots face is operating elevators. Reading elevator control panel buttons involves both an instance segmentation of buttons and labels and associating buttons with their respective metal labels in the elevator. Segmentation algorithms, however, can miss detections. This paper presents a segmentation model specifically designed to solve the problem of missed detections. This can be used to recover detections that the initial model misses. This work presents: 1) a new elevator button dataset containing both 108 images sampled from the internet and 292 images imaged from 24 buildings from the University of Texas at Austin campus and the surrounding neighborhood, along with their segmentation boundaries and associated labels; 2) a vision pipeline based on Mask-RCNN for solving the initial image segmentation and labeling task; and 3) a novel method for identifying missed detections, using a Mask-RCNN network trained on expected button locations. Results show that the missed detections model, specifically developed to recover buttons and labels that were missed by the initial pass, is accurate on up to 99.33% of its predicted missed features on a synthetic missed-detection dataset and 97.14% of its predictions for features missed on a non-synthetic dataset. In the case of the average accuracy of successful button and label detections of a specifically-trained "weak" initial detector at a standard IoU threshold of 0.5, the missed detection model improves the detector’s success rate from 80.38% on the button recognition task with the initial segmentation model only to an average accuracy of 90.7% with the missed detections model enabled. The overall accuracy of the best-performing pipeline implementing the missed detections model is 91.73% and 98.27% on our Internet subset and Campus subset of our dataset, respectively. Nicholas Verzic, Abhinav Chadaga, Justin W. Hart |
IROS | 3 |
| 2024 | Vid2Real HRI: Align video-based HRI study designs with real-world settingsabstractHRI research using autonomous robots in real-world settings can produce results with the highest ecological validity of any study modality, but many difficulties limit such studies’ feasibility and effectiveness. We propose Vid2Real HRI, a research framework to maximize real-world insights offered by video-based studies. The Vid2Real HRI framework was used to design an online study using first-person videos of robots as real-world encounter surrogates. The online study (n=385) distinguished the within-subjects effects of four robot behavioral conditions on perceived social intelligence and human willingness to help the robot enter an exterior door. A real-world, between-subjects replication (n=26) using two conditions confirmed the validity of the online study’s findings and the sufficiency of the participant recruitment target (n=22) based on a power analysis of online study results. The Vid2Real HRI framework offers HRI researchers a principled way to take advantage of the efficiency of video-based study modalities while generating directly transferable knowledge of real-world HRI. Code and data from the study are provided at vid2real.github.io/vid2realHRI. Elliott Hauser, Yao-Cheng Chan, Sadanand Modak, Joydeep Biswas, Justin W. Hart |
RO-MAN | 5 |
| 2024 | Dobby: A Conversational Service Robot Driven by GPT-4abstractThis work introduces a robotics platform which comprehensively integrates multi-step action execution, natural language understanding, and memory to interactively perform service tasks in accordance with variable needs and intentions of users. The proposed architecture is built around an AI agent, derived from GPT-4, which is embedded in an embodied system. Our approach utilizes semantic matching, plan validation, and state messages to ground the agent in the physical world, enabling a seamless merger between communication and behavior. We demonstrate the advantages of this system with an HRI study comparing mobile robots with and without conversational AI capabilities in a free-form tour-guide scenario. The increased adaptability of the system is measured along five dimensions: flexible task planning, interactive exploration of information, emotional-friendliness, personalization, and increased overall user satisfaction. Carson Stark, Bohkyung Chun, Casey Charleston, Varsha Ravi, Luis Pabon, Surya Sunkari, Tarun Mohan, Peter Stone 0001, Justin W. Hart |
RO-MAN | 9 |
| 2024 | Conflict Avoidance in Social Navigation - a SurveyabstractA major goal in robotics is to enable intelligent mobile robots to operate smoothly in shared human-robot environments. One of the most fundamental capabilities in service of this goal is competent navigation in this “social” context. As a result, there has been a recent surge of research on social navigation; and especially as it relates to the handling of conflicts between agents during social navigation. These developments introduce a variety of models and algorithms, however as this research area is inherently interdisciplinary, many of the relevant papers are not comparable and there is no shared standard vocabulary. This survey aims at bridging this gap by introducing such a common language, using it to survey existing work, and highlighting open problems. It starts by defining the boundaries of this survey to a limited, yet highly common type of social navigation—conflict avoidance. Within this proposed scope, this survey introduces a detailed taxonomy of the conflict avoidance components. This survey then maps existing work into this taxonomy, while discussing papers using its framing. Finally, this article proposes some future research directions and open problems that are currently on the frontier of social navigation to aid ongoing and future research. Reuth Mirsky, Xuesu Xiao, Justin W. Hart, Peter Stone 0001 |
ACM Trans. Hum. Robot Interact. | 3 |
| 2022 | Longitudinal Social Impacts of HRI over Long-Term DeploymentsabstractThe Longitudinal Social Impacts of HRI over Long-Term Deployments Workshop seeks to bring together researchers working on all aspects of thoroughly understanding such deployments. This includes researchers working in contributing areas, such as longitudinal studies of human-robot interaction, long-term autonomy, and real-world reployments. This workshop seeks to grow the study of how real-world, deployed robot systems impact the people who interact with them and the social structure of the places that they inhabit. Historically, research in this area has been high-impact. As the world sees robots begin to inhabit places designed for people - delivery robots on city streets, and robots with jobs in airports, shopping malls, and in the home - we expect the importance of understanding these impacts to grow. Justin W. Hart, Elliott Hauser, Samuel Baker, Joydeep Biswas, Junfeng Jiao, Luis Sentis |
HRI | 1 |
| 2022 | Human-Interactive Robot Learning (HIRL)abstractWith robots poised to enter our daily environments, we conjecture that they will not only need to work for people, but also learn from them. An active area of investigation in the robotics, machine learning, and human-robot interaction communities is the design of teachable robotic agents that can learn interactively from human input. To refer to these research efforts, we use the umbrella term Human-Interactive Robot Learning (HIRL). While algorithmic solutions for robots learning from people have been investigated in a variety of ways, HIRL, as a fairly new research area, is still lacking: 1) a formal set of definitions to classify related but distinct research problems or solutions, 2) benchmark tasks, interactions, and metrics to evaluate the performance of HIRL algorithms and interactions, and 3) clear long-term research challenges to be addressed by different communities. The main goal of this workshop will be to consolidate relevant recent work falling under the HIRL umbrella into a coherent set of long, medium, and short-term research problems, and identify the most pressing future research goals in this area. As HIRL is a developing research area, this workshop is an opportunity to break the existing boundaries between relevant research communities by developing and sharing a diverse set of benchmark tasks and metrics for HIRL, inspired by other fields including neuroscience, biology, and ethics research. Reuth Mirsky, Kim Baraka, Taylor Kessler Faulkner, Justin W. Hart, Harel Yedidsion, Xuesu Xiao |
HRI | 4 |
| 2021 | Watch Where You're Going! Gaze and Head Orientation as Predictors for Social Robot NavigationabstractMobile robots deployed in human-populated environments must be able to safely and comfortably navigate in close proximity to people. Head orientation and gaze are both mechanisms which help people to interpret where other people intend to walk, which in turn enables them to coordinate their movement. Head orientation has previously been leveraged to develop classifiers which are able to predict the goal of a person’s walking motion. Gaze is believed to generally precede head orientation, with a person quickly moving their eyes to a target and then following it with a turn of their head. This study leverages state-of-the-art virtual reality technology to place participants into a simulated environment in which their gaze and motion can be observed. The results of this study indicate that position, velocity, head orientation, and gaze can all be used as predictive features of the goal of a person’s walking motion. The results also indicate that gaze both precedes head orientation and can be used to predict the goal of a person’s walking motion at a higher level of accuracy earlier in their walking trajectory. These findings can be leveraged in the design of social navigation systems for mobile robots. Blake Holman, Abrar Anwar, Mauricio Tec, Justin W. Hart, Peter Stone 0001 |
ICRA | 5 |
| 2020 | Deep R-Learning for Continual Area SweepingabstractCoverage path planning is a well-studied problem in robotics in which a robot must plan a path that passes through every point in a given area repeatedly, usually with a uniform frequency. To address the scenario in which some points need to be visited more frequently than others, this problem has been extended to non-uniform coverage planning. This paper considers the variant of non-uniform coverage in which the robot does not know the distribution of relevant events beforehand and must nevertheless learn to maximize the rate of detecting events of interest. This continual area sweeping problem has been previously formalized in a way that makes strong assumptions about the environment, and to date only a greedy approach has been proposed. We generalize the continual area sweeping formulation to include fewer environmental constraints, and propose a novel approach based on reinforcement learning in a Semi-Markov Decision Process. This approach is evaluated in an abstract simulation and in a high fidelity Gazebo simulation. These evaluations show significant improvement upon the existing approach in general settings, which is especially relevant in the growing area of service robotics. We also present a video demonstration on a real service robot. Rishi Shah, Yuqian Jiang, Justin W. Hart, Peter Stone 0001 |
IROS | 3 |
| 2020 | Jointly Improving Parsing and Perception for Natural Language Commands through Human-Robot DialogabstractIn this work, we present methods for using human-robot dialog to improve language understanding for a mobile robot agent. The agent parses natural language to underlying semantic meanings and uses robotic sensors to create multi-modal models of perceptual concepts like red and heavy. The agent can be used for showing navigation routes, delivering objects to people, and relocating objects from one location to another. We use dialog clari_cation questions both to understand commands and to generate additional parsing training data. The agent employs opportunistic active learning to select questions about how words relate to objects, improving its understanding of perceptual concepts. We evaluated this agent on Amazon Mechanical Turk. After training on data induced from conversations, the agent reduced the number of dialog questions it asked while receiving higher usability ratings. Additionally, we demonstrated the agent on a robotic platform, where it learned new perceptual concepts on the y while completing a real-world task. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
J. Artif. Intell. Res. | 7 |
| 2019 | Improving Grounded Natural Language Understanding through Human-Robot DialogabstractNatural language understanding for robotics can require substantial domain- and platform-specific engineering. For example, for mobile robots to pick-and-place objects in an environment to satisfy human commands, we can specify the language humans use to issue such commands, and connect concept words like red can to physical object properties. One way to alleviate this engineering for a new domain is to enable robots in human environments to adapt dynamically-continually learning new language constructions and perceptual concepts. In this work, we present an end-to-end pipeline for translating natural language commands to discrete robot actions, and use clarification dialogs to jointly improve language parsing and concept grounding. We train and evaluate this agent in a virtual setting on Amazon Mechanical Turk, and we transfer the learned agent to a physical robot platform to demonstrate it in the real world. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
ICRA | 7 |
| 2018 | PRISM: Pose Registration for Integrated Semantic MappingabstractMany robotics applications involve navigating to positions specified in terms of their semantic significance. A robot operating in a hotel may need to deliver room service to a named room. In a hospital, it may need to deliver medication to a patient's room. The Building-Wide Intelligence Project at UT Austin has been developing a fleet of autonomous mobile robots, called BWIBots, which perform tasks in the computer science department. Tasks include guiding a person, delivering a message, or bringing an object to a location such as an office, lecture hall, or classroom. The process of constructing a map that a robot can use for navigation has been simplified by modern SLAM algorithms. The attachment of semantics to map data, however, remains a tedious manual process of labeling locations in otherwise automatically generated maps. This paper introduces a system called PRISM to automate a step in this process by enabling a robot to localize door signs - a semantic markup intended to aid the human occupants of a building - and to annotate these locations in its map. Justin W. Hart, Rishi Shah, Sean Kirmani, Nick Walker 0001, Kathryn Baldauf, Nathan John, Peter Stone 0001 |
IROS | 1 |
| 2018 | Passive Demonstrations of Light-Based Robot Signals for Improved Human InterpretabilityabstractWhen mobile robots navigate crowded, human-populated environments, the potential for conflict arises in the form of intersecting trajectories. This study investigates the use of light-emitting diodes (LEDs) arranged along the chassis of a robot in an arrangement similar to a turn signal on a car as a non-anthropomorphic, yet familiar signal to convey the intended path of a mobile service robot. We study the scenario of a human and a robot heading directly toward each other in a hallway, which may give rise to the familiar human experience in which both parties step to the right, then the left, then the right, continuing to block each other's paths until they are able to coordinate their movements and pass each other. We conducted a pilot study which revealed that people do not always interpret this signal as one may expect, which would be similar to how a car uses its turn signal. This motivated a 2 × 2 experiment in which the robot either does or does not use LEDs to indicate its intended direction of travel, and in which study participants either are able to or unable to witness the robot's “lane-changing” behavior further down the hallway prior to coming into direct proximal contact with the robot. The results demonstrate that exposing participants to the robot's use of the LED signal only once prior to passing each other in the hallway is sufficient to disambiguate its meaning to the user, and thus greatly enhances its utility in-situ, with no direct instruction or training to the user. These findings suggest a paradigm of passive demonstration of such signals in future applications. Rolando Fernandez, Nathan John, Sean Kirmani, Justin W. Hart, Jivko Sinapov, Peter Stone 0001 |
RO-MAN | 4 |
| 2012 | Mirror Perspective-Taking with a Humanoid RobotabstractThe ability to use a mirror as an instrument for spatial reasoning enables an agent to make meaningful inferences about the positions of objects in space based on the appearance of their reflections in mirrors. The model presented in this paper enables a robot to infer the perspective from which objects reflected in a mirror appear to be observed, allowing the robot to use this perspective as a virtual camera. Prior work by our group presented an architecture through which a robot learns the spatial relationship between its body and visual sense, mimicking an early form of self-knowledge in which infants learn about their bodies and senses through their interactions with each other. In this work, this self-knowledge is utilized in order to determine the mirror's perspective. Witnessing the position of its end-effector in a mirror in several distinct poses, the robot determines a perspective that is consistent with these observations. The system is evaluated by measuring how well the robot's predictions of its end-effector's position in 3D, relative to the robot's egocentric coordinate system, and in 2D, as projected onto it's cameras, match measurements of a marker tracked by its stereo vision system. Reconstructions of the 3D position end-effector, as computed from the perspective of the mirror, are found to agree with the forward kinematic model within a mean of 31.55mm. When observed directly by the robot's cameras, reconstructions agree within 5.12mm. Predictions of the 2D position of the end-effector in the visual field agree with visual measurements within a mean of 18.47 pixels, when observed in the mirror, or 5.66 pixels, when observed directly by the robot's cameras. Justin W. Hart, Brian Scassellati |
AAAI | 1 |
| 2011 | Effects related to synchrony and repertoire in perceptions of robot danceabstractIn this work we identify low-level aspects of robot motion that can be exploited to create impressions of agency and lifelikeness. In two experiments, participants view split-screen videos of multiple robots set to music and rate the robots on their dance ability, lifelikeness, and entertainment value. The first experiment tests the impact of the correspondence (or lack thereof) of the robot's motion to the underlying rhythm of the music, and the effect of matching changes in the robot's movement to changes in the music, such as a phrase of vocals or drumming. This motivates a second experiment which more deeply explores the relationships of asynchrony and changes in motion repertoire to participants' perceptions of the lifelikeness of the robot's motion. Findings indicate that perceptions of the lifelikeness of the robot and the quality of the dance can be manipulated by simple changes, such as variation in the repertoire of motions, coordination of changes in behavior with events in the music, and the addition of flaws to the robot's synchrony with the music. Eleanor R. Avrunin, Justin W. Hart, Ashley Douglas, Brian Scassellati |
HRI | 2 |
| 2010 | No fair!!: an interaction with a cheating robotabstractUsing a humanoid robot and a simple children's game, we examine the degree to which variations in behavior result in attributions of mental state and intentionality. Participants play the well-known children's game "rock-paper-scissors" against a robot that either plays fairly, or that cheats in one of two ways. In the "verbal cheat" condition, the robot announces the wrong outcome on several rounds which it loses, declaring itself the winner. In the "action cheat"' condition, the robot changes its gesture after seeing its opponent's play. We find that participants display a greater level of social engagement and make greater attributions of mental state when playing against the robot in the conditions in which it cheats. Elaine Short, Justin W. Hart, Michelle Vu, Brian Scassellati |
HRI | 2 |
| 2009 | Incorporating active vision into the body schemaabstractNo abstract available. Justin W. Hart, Eleanor R. Avrunin, David Golub, Brian Scassellati, Steven W. Zucker |
HRI | 1 |
| 2008 | The effect of presence on human-robot interactionabstractThis study explores how a robotpsilas physical or virtual presence affects unconscious human perception of the robot as a social partner. Subjects collaborated on simple book-moving tasks with either a physically present humanoid robot or a video-displayed robot. Each task examined a single aspect of interaction: greetings, cooperation, trust, and personal space. Subjects readily greeted and cooperated with the robot in both conditions. However, subjects were more likely to fulfill an unusual instruction and to afford greater personal space to the robot in the physical condition than in the video-displayed condition. The same tendencies occurred when the virtual robot was supplemented by disambiguating 3-D information. Wilma A. Bainbridge, Justin W. Hart, Elizabeth S. Kim, Brian Scassellati |
RO-MAN | 2 |
| 2006 | QBF Modeling: Exploiting Player Symmetry for Simplicity and Efficiency
Ashish Sabharwal, Carlos Ansótegui, Carla P. Gomes, Justin W. Hart, Bart Selman |
SAT | 4 |