Andrea Thomaz

dblp:l/AndreaLockerd · also Andrea Lockerd, Andrea Lockerd Thomaz · DBLP profile ↗
← Back
79ranked-venue papers
11as first author
10since 2021 · last 2023
0000-0002-3734-964XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 74 · 10 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 37 · 7 first-author · 2 since 2021Systems, architecture and hardware · 26 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 Robots in Real Life: Putting HRI to Work
abstract
This talk will be focused on the unique challenges in deploying a mobile manipulation robot into an environment where the robot is working closely with people on a daily basis. Diligent Robotics' first product, Moxi, is a mobile manipulation service robot that is at work in hospitals today assisting nurses and other front line staff with materials management tasks. This talk will dive into the computational complexity of developing a mobile manipulator with social intelligence. Dr. Thomaz will focus on how human-robot interaction theories and algorithms translate into the real-world and the impact on functionality and perception of robots that perform delivery tasks in a busy human environment. The talk will include many examples and data from the field, with commentary and discussion around both the expected and unexpected hard problems in building robots operating 24/7 as reliable teammates.
Andrea Thomaz
HRI1
2023 Using Learning Curve Predictions to Learn from Incorrect Feedback
abstract
Robots can incorporate data from human teachers when learning new tasks. However, this data can often be noisy, which can cause robots to learn slowly or not at all. One method for learning from human teachers is Human-in-the-loop Reinforcement Learning (HRL), which can combine information from both an environmental reward and external feedback from human teachers. However, many HRL methods assume near-perfect information from teachers or must know the skill level of each teacher before starting the learning process. Our algorithm, Classification for Learning Erroneous Assessments using Rewards (CLEAR), is a feedback filter for Reinforcement Learning (RL) algorithms, enabling learning agents to learn from imperfect teachers without prior modeling. CLEAR is able to determine whether human feedback is correct based on observations of the RL learning curve. Our results suggest that CLEAR improves the quality of human feedback - from 57.5% to 65% correct in a human study - and performs more reliably than baselines by matching or outperforming RL without human teachers in all tested cases.
Taylor Kessler Faulkner, Andrea Thomaz
ICRA2
2022 Abstraction in Data-Sparse Task Transfer (Extended Abstract)
abstract
When a robot adapts a learned task for a novel environment, any changes to objects in the novel environment have an unknown effect on its task execution. For example, replacing an object in a pick-and-place task affects where the robot should target its actions, but does not necessarily affect the underlying action model. In contrast, replacing a tool that the robot will use to complete a task will effectively alter its end-effector pose with respect to the robot's base coordinate system, and thus the robot's motion must be replanned accordingly. These examples highlight the relationship among (i) differences between the source and target environments, (ii) the level of abstraction at which a robot's task model should be represented to enable transfer to the target environment, and (iii) the information needed to ground the abstracted task representation in the target environment. In this abstract, summarizing our full article [Fitzgerald et al., 2021], we present our taxonomy of transfer problems based on this relationship. We also describe a knowledge representation called the Tiered Task Abstraction (TTA) and demonstrate its applicability to a variety of transfer problems in the taxonomy. Our experimental results indicate a trade-off between the generality and data requirements of a task representation, and reinforce the need for multiple transfer methods that operate at different levels of abstraction.
Tesca Fitzgerald, Ashok K. Goel 0001, Andrea Thomaz
IJCAI3
2022 Understanding Acoustic Patterns of Human Teachers Demonstrating Manipulation Tasks to Robots
abstract
Humans use audio signals in the form of spoken language or verbal reactions effectively when teaching new skills or tasks to other humans. While demonstrations allow humans to teach robots in a natural way, learning from trajectories alone does not leverage other available modalities including audio from human teachers. To effectively utilize audio cues accompanying human demonstrations, first it is important to understand what kind of information is present and conveyed by such cues. This work characterizes audio from human teachers demonstrating multi-step manipulation tasks to a situated Sawyer robot along three dimensions: (1) duration of speech used, (2) expressiveness in speech or prosody, and (3) semantic content of speech. We analyze these features for four different independent variables and find that teachers convey similar semantic content via spoken words for different conditions of (1) demonstration types, (2) audio usage instructions, (3) subtasks, and (4) errors during demonstrations. However, differentiating properties of speech in terms of duration and expressiveness are present for the four independent variables, highlighting that human audio carries rich information, potentially beneficial for technological advancement of robot learning from demonstration methods.
Akanksha Saran, Kush Desai, Mai Lee Chang, Rudolf Lioutikov, Andrea Thomaz, Scott Niekum
IROS5
2021 Extending Policy Shaping to Continuous State Spaces (Student Abstract)
abstract
Policy Shaping is a Human-in-the-loop Reinforcement Learning (HRL) algorithm. We extend this work to continuous states with our algorithm, Deep Policy Shaping (DPS). DPS uses a feedback neural network that learns the optimality of actions from noisy feedback combined with an RL algorithm. In simulation, we find that DPS outperforms or matches baselines averaged over multiple hyperparameter settings and varying feedback correctness.
Thomas Benjamin Wei, Taylor Kessler Faulkner, Andrea Thomaz
AAAI3
2021 Towards Safe Motion Planning in Human Workspaces: A Robust Multi-agent Approach
abstract
It is becoming increasingly feasible for robots to share a workspace with humans. However, for them to do so safely while maintaining agile performance, they need the ability to smoothly handle the dynamics and uncertainty caused by human motions. Markov Decision Processes (MDPs) serve as a common framework to formulate robot planning problems. However, because of its single-agent formulation, such planner cannot account for human reaction when evaluating robot actions. The robot can thus suffer from unsafe motions and move in ways that are hard for nearby humans to understand. To resolve this, we instead model robot planning in human workspaces as a Stochastic Game, and contribute a robust planning algorithm, which enables the robot to account for its prediction errors in human responses to prevent collision, while not losing agility, opposed to traditional maximin optimization techniques, by applying maximin operation only at "critical states". We validate the approach under partial knowledge of pedestrian behaviors, and show that our approach encounters zero collision despite imperfect prediction, while improving path efficiency, compared to baselines.
Shih-Yun Lo, Benito Fernandez, Peter Stone 0001, Andrea Thomaz
ICRA4
2021 Robust Planning with Emergent Human-like Behavior for Agents Traveling in Groups
abstract
To enable robots to smoothly interact with humans during their travels together as a group, robots need the ability to adapt their motions under environmental changes and ensure all group members’ routes are feasible. To achieve this ability, robots require knowledge of the final destination and the subgoals in between. In practice, such information is seldom shared explicitly among group members, and may be frequently updated. Under this uncertain setting, maintaining travel efficiency and behavior appropriateness becomes a challenge. Previous literature approached the problem by generating compliant coordinating motions inspired by human groups, with subgoal uncertainty remaining isolated from the plan evaluation process. We show that such coordination can lead the robot to "bad" transient states where inefficient planning and lost tracking may incur. We propose to resolve the problem by formulating the coordinating motion as a Bayesian stochastic game, to plan for the robot as a group member, in the meanwhile considering the long-term effect of uncertainty during path coordination. We show that the approach improves travel efficiency and partner tracking robustness, by preventing assertive decisions during the inference update process. Moreover, the approach presents "agency", in the sense that it can generate human-like motions, which can be applied and contribute to the pedestrian simulation literature; the approach also affords variants from the human-like motions to generate robot behaviors based on sensing capabilities, contributing to the methodology of robot behavior design.
Shih-Yun Lo, Elaine Short, Andrea Thomaz
ICRA3
2021 Communication Strategy for Efficient Guidance Providing : Domain-structure Awareness, Performance Trade-offs, and Value of Future Observations
Shih-Yun Lo, Andrea Thomaz
ICRA2
2021 Unfair! Perceptions of Fairness in Human-Robot Teams
abstract
How team members are treated influences their performance in the team and their desire to be a part of the team in the future. Prior research in human-robot teamwork proposes fairness definitions for human-robot teaming that are based on the work completed by each team member. However, metrics that properly capture people’s perception of fairness in human-robot teaming remains a research gap. We present work on assessing how well objective metrics capture people’s perception of fairness. First, we extend prior fairness metrics based on team members’ capabilities and workload to a bigger team. We also develop a new metric to quantify the amount of time that the robot spends working on the same task as each person. We conduct an online user study (n=95) and show that these metrics align with perceived fairness. Importantly, we discover that there are bleed-over effects in people’s assessment of fairness. When asked to rate fairness based on the amount of time that the robot spends working with each person, participants used two factors (fairness based on the robot’s time and teammates’ capabilities). This bleed-over effect is stronger when people are asked to assess fairness based on capability. From these insights, we propose design guidelines for algorithms to enable robotic teammates to consider fairness in its decision-making to maintain positive team social dynamics and team task performance.
Mai Lee Chang, J. Gregory Trafton, J. Malcolm McCurry, Andrea Thomaz
RO-MAN4
2021 Abstraction in data-sparse task transfer
Tesca Fitzgerald, Ashok K. Goel 0001, Andrea Thomaz
Artif. Intell.3
2020 Planning with Partner Uncertainty Modeling for Efficient Information Revealing in Teamwork
abstract
Communication among team members is important for efficient teamwork, to coordinate behavior and ensure that all team members have the information they need to complete the task. To enable effective communication and thus efficient teamwork, we propose a multi-agent planning approach to revealing information based on its benefit to joint team performance. By explicitly modeling the partner's knowledge and behavior, our approach allows a robot in a team to reason about when information is useful, how the communication is effective, and to communicate through efficient actions. That is, the robot provides only the necessary information for task completion, provides the information at the time that it is needed, and through the action(s) that optimizes team performance. We validated this approach in a human study in which participants walk together with a robot to a destination that is known only to the robot. We compared to a legible motion generation approach, and showed that users perceived our approach as more natural, socially appropriate, and fluent to team with, while being both more predictable and intent-clear. The ratings of our approach are equal or higher than legible motion across all 18 survey items.
Shih-Yun Lo, Elaine Short, Andrea Thomaz
HRI3
2020 Interactive Reinforcement Learning with Inaccurate Feedback
abstract
Interactive Reinforcement Learning (RL) enables agents to learn from two sources: rewards taken from observations of the environment, and feedback or advice from a secondary critic source, such as human teachers or sensor feedback. The addition of information from a critic during the learning process allows the agents to learn more quickly than non-interactive RL. There are many methods that allow policy feedback or advice to be combined with RL. However, critics can often give imperfect information. In this work, we introduce a framework for characterizing Interactive RL methods with imperfect teachers and propose an algorithm, Revision Estimation from Partially Incorrect Resources (REPaIR), which can estimate corrections to imperfect feedback over time. We run experiments both in simulations and demonstrate performance on a physical robot, and find that when baseline algorithms do not have prior information on the exact quality of a feedback source, using REPaIR matches or improves the expected performance of these algorithms.
Taylor Kessler Faulkner, Elaine Short, Andrea Thomaz
ICRA3
2020 TASC: Teammate Algorithm for Shared Cooperation
abstract
For robots to be perceived as full-fledged team members, they must display intelligent behavior along multiple dimensions. One challenge is that even when the robot and human are on the same team, the interaction may not feel like teamwork to the human. We present a novel algorithm, Teammate Algorithm for Shared Cooperation (TASC). TASC is motivated by the concept of shared cooperative activity (SCA) for human-human teamwork, developed in prior work by Bratman. We focus on enabling the robot to prioritize certain SCA facets in its action selection depending on the task. We evaluated TASC in three experiments using different tasks with human users on Amazon Mechanical Turk. Our results show that TASC enabled participants to predict the robot's goal earlier by one robot move and with greater confidence. The robot also helped reduce participants' energy usage in a simulated block-moving task. Altogether, these results show that considering the SCA facets in the robot's action selection improves teamwork.
Mai Lee Chang, Taylor Kessler Faulkner, Thomas Benjamin Wei, Elaine Short, Gokul Anandaraman, Andrea Thomaz
IROS6
2020 Defining Fairness in Human-Robot Teams
abstract
We seek to understand the human teammate's perception of fairness during a human-robot physical collaborative task where certain subtasks leverage the robot's strengths and others leverage the human's. We conduct a user study (n=30) to investigate the effects of fluency (absent vs. present) and effort (absent vs. present) on participants' perception of fairness. Fluency controls if the robot minimizes the idle time between the human's action and robot's action. Effort controls if the robot performs tasks that it is least skilled at, i.e., most time-consuming tasks, as quickly as possible. We evaluated four human-robot teaming algorithms that consider different levels of fluency and effort. Our results show that effort and fluency help improve fairness without making a trade-off with efficiency. When the robot displays effort, this significantly increased participants' perceived fairness. Participants' perception of fairness is also influenced by team members' skill levels and task type. To that end, we propose three notions of fairness for effective human-robot teamwork: equality of workload, equality of capability, and equality of task type.
Mai Lee Chang, Zachary Pope, Elaine Short, Andrea Thomaz
RO-MAN4
2019 Learning from Corrective Demonstrations
abstract
Robots deployed in human environments will inevitably encounter unmodeled scenarios which are likely to result in execution failures. To address this issue, we would like to allow co-present naive users to correct and improve the robot's behavior as these edge cases are encountered over time.
Reymundo Gutierrez, Elaine Short, Scott Niekum, Andrea Thomaz
HRI4
2019 Enhancing Robot Learning with Human Social Cues
abstract
Imagine a learning scenario between two humans: a teacher demonstrating how to play a new musical instrument or a craftsman teaching a new skill like pottery or knitting to a novice. Even though learning a skill has a learning curve to get the nuances of the technique right, some basic social principles are followed between the teacher and the student to make the learning process eventually succeed. There are several assumptions or social priors in this communication for teaching: mutual eye contact to draw attention to instructions, following the gaze of the teacher to understand the skill, the teacher following the student's gaze during imitation to give feedback, the teacher demonstrating by pointing towards something she is going to approach or manipulate and verbal interruptions or corrections during the learning process [1], [2]. In prior research, verbal and non-verbal social cues such as eye gaze and gestures have been shown to make human-human interactions seamless and augment verbal, collaborative behavior [3], [4]. They serve as an indicator of engagement, interest and attention when people interact face-to-face with one another [5], [6].
Akanksha Saran, Elaine Short, Andrea Thomaz, Scott Niekum
HRI3
2019 SAIL: Simulation-Informed Active In-the-Wild Learning
abstract
Robots in real-world environments may need to adapt context-specific behaviors learned in one environment to new environments with new constraints. In many cases, copresent humans can provide the robot with information, but it may not be safe for them to provide hands-on demonstrations and there may not be a dedicated supervisor to provide constant feedback. In this work we present the SAIL (Simulation-Informed Active In-the-Wild Learning) algorithm for learning new approaches to manipulation skills starting from a single demonstration. In this three-step algorithm, the robot simulates task execution to choose new potential approaches; collects unsupervised data on task execution in the target environment; and finally, chooses informative actions to show to co-present humans and obtain labels. Our approach enables a robot to learn new ways of executing two different tasks by using success/failure labels obtained from naïve users in a public space, performing 496 manipulation actions and collecting 163 labels from users in the wild over six 45-minute to 1-hour deployments. We show that classifiers based low-level sensor data can be used to accurately distinguish between successful and unsuccessful motions in a multi-step task ( ), even when trained in the wild. We also show that using the sensor data to choose which actions to sample is more effective than choosing the least-sampled action.
Elaine Short, Adam Allevato, Andrea Thomaz
HRI3
2019 Real-time Multisensory Affordance-based Control for Adaptive Object Manipulation
abstract
We address the challenge of how a robot can adapt its actions to successfully manipulate objects it has not previously encountered. We introduce Real-time Multisensory Affordance-based Control (RMAC), which enables a robot to adapt existing affordance models using multisensory inputs. We show that using the combination of haptic, audio, and visual information with RMAC allows the robot to learn afforance models and adaptively manipulate two very different objects (drawer, lamp), in multiple novel configurations. Offline evaluations and real-time online evaluations show that RMAC allows the robot to accurately open different drawer configurations and turn-on novel lamps with an average accuracy of 75%.
Vivian Chu, Reymundo Gutierrez, Sonia Chernova, Andrea Thomaz
ICRA4
2018 Detecting Contingency for HRI in Open-World Environments
abstract
This paper presents a novel algorithm for detecting contingent reactions to robot behavior in noisy real-world environments with naive users. Prior work has established that one way to detect contingency is by calculating a difference metric between sensor data before and after a robot probe of the environment. Our algorithm, CIRCLE (Contingency for Interactive Real-time CLassification of Engagement) provides a new approach to calculating this difference and detecting contingency, improving the running time for the difference calculation from 2.5 seconds to approximately 0.001 seconds on an 1100-sample vector, and effectively enabling real-time detection of contingent events. We show accuracy comparable to the best offline results for detecting contingency in this way (89.5% vs 91% in prior work), and demonstrate the utility of the real-time contingency detection in a field study of a survey-administering robot in a noisy open-world environment with naive users, showing that the robot can decrease the number of requests it makes (from 38 to 13) while more efficiently collecting survey responses (30% response rate rather than 26.3%).
Elaine Short, Mai Lee Chang, Andrea Thomaz
HRI3
2018 Human-Driven Feature Selection for a Robotic Agent Learning Classification Tasks from Demonstration
abstract
The state features available to a robot define the variables on which the learning computation depends. However, little prior work considers feature selection in the context of deploying a general-purpose robot able to learn new tasks. In this work, we explore human-driven feature selection in which a robotic agent can identify useful features with the aid of a human user, by extracting information from users about which features are most informative for discriminating between classes of objects needed for a given task (e.g. sorting groceries). The research questions examine (a) whether a domain expert is able to identify a subset of informative task features, (b) whether human selected features will enable the agent to classify unseen examples as accurately as using computational feature selection, and (c) if the interaction strategy used to elicit the information from the user impacts the quality of resulting feature selection. Toward that end, we conducted a user study with 30 participants on campus, given a multi-class classification task and one of five different approaches for conveying information about informative features to a robot learner. Our findings show that when features are semantically interpretable, human feature selection is effective in LfD scenarios because it is able to outperform computational methods when there is limited training data, yet still remains on-par with computational methods as the training sample size increases.
Kalesha Bullard, Sonia Chernova, Andrea Thomaz
ICRA3
2018 Incremental Task Modification via Corrective Demonstrations
abstract
In realistic environments, fully specifying a task model such that a robot can perform a task in all situations is impractical. In this work, we present Incremental Task Modification via Corrective Demonstrations (ITMCD), a novel algorithm that allows a robot to update a learned model by making use of corrective demonstrations from an end-user in its environment. We propose three different types of model updates that make structural changes to a finite state automaton (FSA) representation of the task by first converting the FSA into a state transition auto-regressive hidden Markov model (STARHMM). The STARHMM's probabilistic properties are then used to perform approximate Bayesian model selection to choose the best model update, if any. We evaluate ITMCD Model Selection in a simulated block sorting domain and the full algorithm on a real-world pouring task. The simulation results show our approach can choose new task models that sufficiently incorporate new demonstrations while remaining as simple as possible. The results from the pouring task show that ITMCD performs well when the modeled segments of the corrective demonstrations closely comply with the original task model.
Reymundo Gutierrez, Vivian Chu, Andrea Thomaz, Scott Niekum
ICRA3
2018 Towards Intelligent Arbitration of Diverse Active Learning Queries
abstract
Active learning literature has explored the selection of optimal queries by a learning agent with respect to given criteria, but prior work in classification has focused only on obtaining labels for queried samples. In contrast, proficient learners, like humans, integrate multiple forms of information during learning. This work seeks to enable an active learner to reason about multiple query types concurrently, aimed at soliciting both instance and feature information from the teacher, and to autonomously arbitrate between queries of different types. We contribute the design of rule-based and decision-theoretic arbitration strategies and evaluate all against baselines of more traditional passive and active learning. Our findings show that all arbitration strategies lead to more efficient learning, compared to the baselines. Moreover, given a dynamically changing environment and constrained questioning budget (typical in human settings), the decision-theoretic strategy statistically outperforms all other methods since it reasons about both what query to make and when to make a query, in order to most effectively utilize its questioning budget.
Kalesha Bullard, Andrea Thomaz, Sonia Chernova
IROS2
2018 Effects of Integrated Intent Recognition and Communication on Human-Robot Collaboration
abstract
Human-robot interaction research to date has investigated intent recognition and communication separately. In this paper, we explore the effects of integrating both the robot's ability to generate intentional motion and predict the human's motion in a collaborative physical task. We implemented an intent recognition system to recognize the human partner's hand motion intent and a motion planner system to enable the robot to communicate its intent by using legible and predictable motion. We tested this bi-directional intent system in a 2-way within-subjects user study. Results suggest that an integrated intent recognition and communication system may facilitate more collaborative behavior among team members.
Mai Lee Chang, Reymundo Gutierrez, Priyanka Khante, Elaine Short, Andrea Thomaz
IROS5
2018 Policy Shaping with Supervisory Attention Driven Exploration
abstract
Robots deployed for long periods of time need to be able to explore and learn from their environment. One approach to this problem has been reinforcement learning (RL), in which robots receive rewards from the environment that allow them to choose optimal actions. To speed learning when human supervision is available, interactive reinforcement learning solicits feedback from a human teacher. However, this approach typically assumes that learning takes place under continuous supervision, which is unlikely to hold in long-term scenarios. We propose an extension to a method of interactive reinforcement learning, policy shaping, that takes into account human attention. Our approach enables better performance while unattended by favoring information-gathering actions when attended and actions that have received positive feedback when unattended. We test our approach in both simulation and on a robot, finding that our method learns faster than policy shaping and performs more safely than policy shaping while no one is paying attention to the robot.
Taylor Kessler Faulkner, Elaine Short, Andrea Thomaz
IROS3
2018 Human Gaze Following for Human-Robot Interaction
abstract
Gaze provides subtle informative cues to aid fluent interactions among people. Incorporating human gaze predictions can signify how engaged a person is while interacting with a robot and allow the robot to predict a human's intentions or goals. We propose a novel approach to predict human gaze fixations relevant for human-robot interaction tasks-both referential and mutual gaze-in real time on a robot. We use a deep learning approach which tracks a human's gaze from a robot's perspective in real time. The approach builds on prior work which uses a deep network to predict the referential gaze of a person from a single 2D image. Our work uses an interpretable part of the network, a gaze heat map, and incorporates contextual task knowledge such as location of relevant objects, to predict referential gaze. We find that the gaze heat map statistics also capture differences between mutual and referential gaze conditions, which we use to predict whether a person is facing the robot's camera or not. We highlight the challenges of following a person's gaze on a robot in real time and show improved performance for referential gaze and mutual gaze prediction.
Akanksha Saran, Srinjoy Majumdar, Elaine Short, Andrea Thomaz, Scott Niekum
IROS4
2018 Human-Guided Object Mapping for Task Transfer
abstract
When transferring a learned task to an environment containing new objects, a core problem is identifying the mapping between objects in the old and new environments. This object mapping is dependent on the task being performed and the roles objects play in that task. Prior work assumes (i) the robot has access to multiple new demonstrations of the task or (ii) the primary features for object mapping have been specified. We introduce an approach that is not constrained by either assumption but rather uses structured interaction with a human teacher to infer an object mapping for task transfer. We describe three experiments: an extensive evaluation of assisted object mapping in simulation, an interactive evaluation incorporating demonstration and assistance data from a user study involving 10 participants, and an offline evaluation of the robot’s confidence during object mapping. Our results indicate that human-guided object mapping provided a balance between mapping performance and autonomy, resulting in (i) up to 2.25× as many correct object mappings as mapping without human interaction, and (ii) more efficient transfer than requiring the human teacher to re-demonstrate the task in the new environment, correctly inferring the object mapping across 93.3% of the tasks and requiring at most one interactive assist in the typical case.
Tesca Fitzgerald, Ashok K. Goel 0001, Andrea Thomaz
ACM Trans. Hum. Robot Interact.3
2017 Human-Robot Co-Creativity: Task Transfer on a Spectrum of Similarity
Tesca Fitzgerald, Ashok K. Goel 0001, Andrea Thomaz
ICCC3
2017 Situated Bayesian Reasoning Framework for Robots Operating in Diverse Everyday Environments
Sonia Chernova, Vivian Chu, Angel Andres Daruna, Haley Garrison, Meera Hahn, Priyanka Khante, Andrea Thomaz
ISRR8
2016 Learning Object Affordances by Leveraging the Combination of Human-Guidance and Self-Exploration
abstract
Our work focuses on robots to be deployed in human environments. These robots, which will need specialized object manipulation skills, should leverage end-users to efficiently learn the affordances of objects in their environment. This approach is promising because people naturally focus on showing salient aspects of the objects [1]. We replicate prior results and build on them to create a combination of self and supervised learning. We present experimental results with a robot learning 5 affordances on 4 objects using 1219 interactions. We compare three conditions: (1) learning through self-exploration, (2) learning from supervised examples provided by 10 naïve users, and (3) self-exploration biased by the user input. Our results characterize the benefits of self and supervised affordance learning and show that a combined approach is the most efficient and successful.
Vivian Chu, Tesca Fitzgerald, Andrea Thomaz
HRI3
2016 Learning and Grounding Haptic Affordances Using Demonstration and Human-Guided Exploration
abstract
We present a system for learning haptic affordance models of complex manipulation skills. The goal of a haptic affordance model is to improve task completion by characterizing the feel of a particular object-action pair. We use learning from demonstration to provide the robot with an example of a successful interaction with a given object. We then use environmental scaffolding and a wrist-mounted force/torque (F/T) sensor to collect grounded examples (successes and unsuccessful “near misses”) of the haptic data for the object-action pair. From this, we build one “success” Hidden Markov Model (HMM) and one “near-miss” HMM for each object-action pair. We evaluate this approach with five different actions on seven different objects to learn two specific affordances (open-able and scoop-able). We show that by building a library of object-action pairs for each affordance, we can successfully monitor a trajectory of haptic data to determine if the robot finds an affordance.
Vivian Chu, Andrea Thomaz
HRI2
2016 Work those arms: Toward dynamic and stable humanoid walking that optimizes full-body motion
abstract
Humanoid robots are designed with dozens of actuated joints to suit a variety of tasks, but walking controllers rarely make the best use of all of this freedom. We present a framework for maximizing the use of the full humanoid body for the purpose of stable dynamic locomotion, which requires no restriction to a planning template (e.g. LIPM). Using a hybrid zero dynamics (HZD) framework, this approach optimizes a set of outputs which provides requirements for the motion for all actuated links, including arms. These output equations are then rapidly solved by a whole-body inverse-kinematic (IK) solver, providing a set of joint trajectories to the robot. We apply this procedure to a simulation of the humanoid robot, DRC-HUBO, which has over 27 actuators. As a consequence, the resulting gaits swing their arms, not by a user defining swinging motions a priori or superimposing them on gaits post hoc, but as an emergent behavior from optimizing the dynamic gait. We also present preliminary dynamic walking experiments with DRC-HUBO in hardware, thereby building a case that hybrid zero dynamics as augmented by inverse kinematics (HZD+IK) is becoming a viable approach for controlling the full complexity of humanoid locomotion.
Christian Hubicki, Ayonga Hereid, Michael X. Grey, Andrea Thomaz, Aaron D. Ames
ICRA4
2016 Hierarchical rejection sampling for informed kinodynamic planning in high-dimensional spaces
abstract
We present hierarchical rejection sampling (HRS) to improve the efficiency of asymptotically optimal sampling-based planners for high-dimensional problems with differential constraints. Pruning nodes and rejecting samples that cannot improve the currently best solution have been shown to improve performance for certain problems. We show that in high-dimensional domains this improvement can be so large that rejecting samples becomes the bottleneck of the algorithm because almost all samples are rejected. This contradicts general wisdom that collision checking is always the bottleneck of sampling-based planners. Only samples in the informed subset of the state space can potentially improve the current solution. For systems without differential constraints the informed subset forms an ellipsoid, which can be parameterized and sampled directly. For systems with differential constraints the informed subset is more complicated and no such direct sampling methods exist. HRS improves the efficiency of finding samples within the informed subset without parameterizing it explicitly. Thus, it can also be applied to systems with differential constraints for which a steering method is available. In our experiments we demonstrate efficiency improvements of an RRT* planner of up to two orders of magnitude.
Tobias Kunz, Andrea Thomaz, Henrik I. Christensen
ICRA2
2016 Humanoid manipulation planning using backward-forward search
abstract
This paper explores combining task and manipulation planning for humanoid robots. Existing methods tend to either take prohibitively long to compute for humanoids or artificially limit the physical capabilities of the humanoid platform by restricting the robot's actions to predetermined trajectories. We present a hybrid planning system which is able to scale well for complex tasks without relying on predetermined robot actions. Our system utilizes the hybrid backward-forward planning algorithm for high-level task planning combined with humanoid primitives for standing and walking motion planning. These primitives are designed to be efficiently computable during planning, despite the large amount of complexity present in humanoid robots, while still informing the task planner of the geometric constraints present in the problem. Our experiments apply our method to simulated pick-and-place problems with additional gate constraints impacting navigation using the DRC-HUBO1 robot. Our system is able to solve puzzle-like problems on a humanoid within a matter of minutes.
Michael X. Grey, Caelan Reed Garrett, C. Karen Liu, Aaron D. Ames, Andrea Thomaz
IROS5
2016 Grounding action parameters from demonstration
abstract
When a robot is deployed to a new setting, it must reason about how to accomplish the goals of domain-appropriate tasks within the environment it is situated. We investigate the problem of enabling robots to interactively learn how to perform known tasks in new environments. Each task is composed of a sequence of parameterized actions, which we assume are given to the robot in the form of a task recipe. In order to learn how to ground the task in a new environment, our learner builds classifiers to model each of the parameters (i.e. all unique objects and semantic locations) associated with the task. In evaluation for two tasks across three different environments, our results show that these groundings are both (1) capable of being learned efficiently from demonstrations, and (2) necessary to learn for each new environment.
Kalesha Bullard, Baris Akgün, Sonia Chernova, Andrea Thomaz
RO-MAN4
2015 Visual Case Retrieval for Interpreting Skill Demonstrations
Tesca Fitzgerald, Keith McGreggor, Baris Akgün, Andrea Thomaz, Ashok K. Goel 0001
ICCBR4
2015 Policy Shaping with Human Teachers
Thomas Cederborg, Ishaan Grover, Charles L. Isbell Jr., Andrea Thomaz
IJCAI4
2015 Self-improvement of learned action models with learned goal models
abstract
We introduce a new method for robots to further improve upon skills acquired through Learning from Demonstration. Previously, we have introduced a method to learn both an action model to execute the skill and a goal model to monitor the execution of the skill. In this paper we show how to use the learned goal models to improve the learned action models autonomously, without further user interaction. Trajectories are sampled from the action model and executed on the robot. The goal model then labels them as success or failure and the successful ones are used to update the action model. We introduce an adaptive sampling method to speed up convergence. We show through both simulation and real robot experiments that our method can fix a failed action model.
Baris Akgün, Andrea Thomaz
IROS2
2015 An evaluation of GUI and kinesthetic teaching methods for constrained-keyframe skills
abstract
Keyframe-based Learning from Demonstration has been shown to be an effective method for allowing end-users to teach robots skills. We propose a method for using multiple keyframe demonstrations to learn skills as sequences of positional constraints (c-keyframes) which can be planned between for skill execution. We also introduce an interactive GUI which can be used for displaying the learned c-keyframes to the teacher, for altering aspects of the skill after it has been taught, or for specifying a skill directly without providing kinesthetic demonstrations. We compare 3 methods of teaching c-keyframe skills: kinesthetic teaching, GUI teaching, and kinesthetic teaching followed by GUI editing of the learned skill (K-GUI teaching). Based on user evaluation, the K-GUI method of teaching is found to be the most preferred, and the GUI to be the least preferred. Kinesthetic teaching is also shown to result in more robust constraints than GUI teaching, and several use cases of K-GUI teaching are discussed to show how the GUI can be used to improve the results of kinesthetic teaching.
Andrey Kurenkov, Baris Akgün, Andrea Thomaz
IROS3
2015 Real-time changes to social dynamics in human-robot turn-taking
abstract
In order for robots to work alongside humans in a range of domains, they will need to operate with a variety of social dynamics that each context will require. This paper builds on previous work with a parameterized turn-taking model, CADENCE, in which different parameter settings resulted in different social dynamics. In contrast to the static parameter settings of previous work, we now investigate the problem of changing these turn-taking parameter sets dynamically within a single interaction session. This ability is necessary for successful peer-to-peer collaborations, in which balance of control between leading and following must be maintained. We present our dynamic switching approach and an experiment with 15 participants. Our results confirm that it is possible to achieve the same changes in social dynamics within a single interaction session that were previously seen only between independent sessions of different parameter settings. Moreover, we show that such a change in social dynamics is contingent upon changing parameters at socially appropriate turn boundaries.
Justin S. Smith, Crystal Chao, Andrea Thomaz
IROS3
2014 Multimodal real-time contingency detection for HRI
abstract
Our goal is to develop robots that naturally engage people in social exchanges. In this paper, we focus on the problem of recognizing that a person is responsive to a robot's request for interaction. Inspired by human cognition, our approach is to treat this as a contingency detection problem. We present a simple discriminative Support Vector Machine (SVM) classifier to compare against previous generative methods introduced in prior work by Lee et al. [1]. We evaluate these methods in two ways. First, by training three separate SVMs with multi-modal sensory input on a set of batch data collected in a controlled setting, where we obtain an average F1score of 0.82. Second, in an open-ended experiment setting with seven participants, we show that our model is able to perform contingency detection in real-time and generalize to new people with a best F1score of 0.72.
Vivian Chu, Kalesha Bullard, Andrea Thomaz
IROS3
2014 Eliciting good teaching from humans for machine learners
Maya Cakmak, Andrea Thomaz
Artif. Intell.2
2014 Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains
Luis C. Cobo, Kaushik Subramanian, Charles L. Isbell Jr., Aaron D. Lanterman, Andrea Thomaz
Artif. Intell.5
2013 Collaborative manipulation: new challenges for robotics and HRI
Anca D. Dragan, Andrea Thomaz, Siddhartha S. Srinivasa
HRI2
2013 Revisioning HRI given exponential technological growth
Peter H. Kahn Jr., Gerhard Sagerer, Andrea Thomaz, Takayuki Kanda 0001
HRI3
2013 Policy Shaping: Integrating Human Feedback with Reinforcement Learning
abstract
A long term goal of Interactive Reinforcement Learning is to incorporate non-expert human feedback to solve complex tasks. State-of-the-art methods have approached this problem by mapping human information to reward and value signals to indicate preferences and then iterating over them to compute the necessary control policy. In this paper we argue for an alternate, more effective characterization of human feedback: Policy Shaping. We introduce Advise, a Bayesian approach that attempts to maximize the information gained from human feedback by utilizing it as direct labels on the policy. We compare Advise to state-of-the-art approaches and highlight scenarios where it outperforms them and importantly is robust to infrequent and inconsistent human feedback.
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L. Isbell Jr., Andrea Thomaz
NIPS5
2013 Controlling social dynamics with a parametrized model of floor regulation
abstract
Turn-taking is ubiquitous in human communication, yet turn-taking between humans and robots continues to be stilted and awkward for human users. The goal of our work is to build autonomous robot controllers for successfully engaging in human-like turn-taking interactions. Towards this end, we present CADENCE, a novel computational model and architecture that explicitly reasons about the four components of floor regulation: seizing the floor, yielding the floor, holding the floor, and auditing the owner of the floor. The model is parametrized to enable the robot to achieve a range of social dynamics for the human-robot dyad. In a between-groups experiment with 30 participants, our humanoid robot uses this turn-taking system at two contrasting parametrizations to engage users in autonomous object play interactions. Our results from the study show that: (1) manipulating these turn-taking parameters results in significantly different robot behavior; (2) people perceive the robot's behavioral differences and consequently attribute different personalities to the robot; and (3) changing the robot's personality results in different behavior from the human, manipulating the social dynamics of the dyad. We discuss the implications of this work for various contextual applications as well as the key limitations of the system to be addressed in future work.
Crystal Chao, Andrea Thomaz
J. Hum. Robot Interact.2
2012 Trajectories and keyframes for kinesthetic teaching: a human-robot interaction perspective
abstract
Kinesthetic teaching is an approach to providing demonstrations to a robot in Learning from Demonstration whereby a human physically guides a robot to perform a skill. In the common usage of kinesthetic teaching, the robot's trajectory during a demonstration is recorded from start to end. In this paper we consider an alternative, keyframe demonstrations, in which the human provides a sparse set of consecutive keyframes that can be connected to perform the skill. We present a user-study (n=34) comparing the two approaches and highlighting their complementary nature. The study also tests and shows the potential benefits of iterative and adaptive versions of keyframe demonstrations. Finally, we introduce a hybrid method that combines trajectories and keyframes in a single demonstration.
Baris Akgün, Maya Cakmak, Jae Wook Yoo, Andrea Thomaz
HRI4
2012 Designing robot learners that ask good questions
abstract
Programming new skills on a robot should take minimal time and effort. One approach to achieve this goal is to allow the robot to ask questions. This idea, called Active Learning, has recently caught a lot of attention in the robotics community. However, it has not been explored from a human-robot interaction perspective. In this paper, we identify three types of questions (label, demonstration and feature queries) and discuss how a robot can use these while learning new skills. Then, we present an experiment on human question asking which characterizes the extent to which humans use these question types. Finally, we evaluate the three question types within a human-robot teaching interaction. We investigate the ease with which different types of questions are answered and whether or not there is a general preference of one type of question over another. Based on our findings from both experiments we provide guidelines for designing question asking behaviors on a robot learner.
Maya Cakmak, Andrea Thomaz
HRI2
2012 Enhancing interaction through exaggerated motion synthesis
abstract
Other than eye gaze and referential gestures (e.g. pointing), the relationship between robot motion and observer attention is not well understood. We explore this relationship to achieve social goals, such as influencing human partner behavior or directing attention. We present an algorithm that creates exaggerated variants of a motion in real-time. Through two experiments we confirm that exaggerated motion is perceptibly different than the input motion, provided that the motion is sufficiently exaggerated. We found that different levels of exaggeration correlate to human expectations of robot-like, human-like, and cartoon-like motion. We present empirical evidence that use of exaggerated motion in experiments enhances the interaction through the benefits of increased engagement and perceived entertainment value. Finally, we provide statistical evidence that exaggerated motion causes predictable human partner gaze direction and better retention of interaction details.
Michael J. Gielniak, Andrea Thomaz
HRI2
2012 Timing in multimodal turn-taking interactions: control and analysis using timed Petri nets
abstract
Turn-taking interactions with humans are multimodal and reciprocal in nature. In addition, the timing of actions is of great importance, as it influences both social and task strategies. To enable the precise control and analysis of timed discrete events for a robot, we develop a system for multimodal collaboration based on a timed Petri net (TPN) representation. We also argue for action interruptions in reciprocal interaction and describe its implementation within our system. Using the system, our autonomously operating humanoid robot Simon collaborates with humans through both speech and physical action to solve the Towers of Hanoi, during which the human and the robot take turns manipulating objects in a shared physical workspace. We hypothesize that action interruptions have a positive impact on turn-taking and evaluate this in the Towers of Hanoi domain through two experimental methods. One is a between-groups user study with 16 participants. The other is a simulation experiment using 200 simulated users of varying speed, initiative, compliance, and correctness. In these experiments, action interruptions are either present or absent in the system. Our collective results show that action interruptions lead to increased task efficiency through increased user initiative, improved interaction balance, and higher sense of fluency. In arriving at these results, we demonstrate how these evaluation methods can be highly complementary in the analysis of interaction dynamics.
Crystal Chao, Andrea Thomaz
J. Hum. Robot Interact.2
2011 Learning Tasks and Skills Together From a Human Teacher
abstract
We are interested in developing Learning from Demonstration (LfD) systems that are tailored to be used by everyday people. We highlight and tackle the issues of skill learning, task learning and interaction in the context of LfD As part of the AAAI 2011 LfD Challenge, we will demonstrate some of our most recent Socially Guided-Machine Learning work, in which the PR2 robot learns both low-level skills and high-level tasks through an ongoing social dialog with a human partner
Baris Akgün, Kaushik Subramanian, Jaeeun Shim, Andrea Thomaz
AAAI4
2011 Touched by a robot: an investigation of subjective responses to robot-initiated touch
abstract
By initiating physical contact with people, robots can be more useful. For example, a robotic caregiver might make contact to provide physical assistance or facilitate communication. So as to better understand how people respond to robot-initiated touch, we conducted a 2x2 between-subjects experiment with 56 people in which a robotic nurse autonomously touched and wiped the subject's forearm. Our independent variables were whether or not the robot verbally warned the person before contact, and whether the robot verbally indicated that the touch was intended to clean the person's skin (instrumental touch) or to provide comfort (affective touch). On average, regardless of the treatment, participants had a generally positive subjective response. However, with instrumental touch people responded significantly more favorably. Since the physical behavior of the robot was the same for all trials, our results demonstrate that the perceived intent of the robot can significantly influence a person's subjective response to robot-initiated touch. Our results suggest that roboticists should consider this factor in addition to the mechanics of physical interaction. Unexpectedly, we found that participants tended to respond more favorably without a verbal warning. Although inconclusive, our results suggest that verbal warnings prior to contact should be carefully designed, if used at all.
Tiffany L. Chen, Chih-Hung King, Andrea Thomaz, Charles C. Kemp
HRI3
2011 Spatiotemporal correspondence as a metric for human-like robot motion
abstract
Coupled degrees-of-freedom exhibit correspondence, in that their trajectories influence each other. In this paper we add evidence to the hypothesis that spatiotemporal correspondence (STC) of distributed actuators is a component of human-like motion. We demonstrate a method for making robot motion more human-like, by optimizing with respect to a nonlinear STC metric. Quantitative evaluation of STC between coordinated robot motion, human motion capture data, and retargeted human motion capture data projected onto an anthropomorphic robot suggests that coordinating robot motion with respect to the STC metric makes the motion more human-like. A user study based on mimicking shows that STC-optimized motion is (1) more often recognized as a common human motion, (2) more accurately identified as the originally intended motion, and (3) mimicked more accurately than a non-optimized version. We conclude that coordinating robot motion with respect to the STC metric makes the motion more human-like. Finally, we present and discuss data on potential reasons why coordinating motion increases recognition and ability to mimic.
Michael J. Gielniak, Andrea Thomaz
HRI2
2011 Vision-based contingency detection
abstract
We present a novel method for the visual detection of a contingent response by a human to the stimulus of a robot action. Contingency is defined as a change in an agent's behavior within a specific time window in direct response to a signal from another agent; detection of such responses is essential to assess the willingness and interest of a human in interacting with the robot. Using motion-based features to describe the possible contingent action, our approach assesses the visual self-similarity of video subsequences captured before the robot exhibits its signaling behavior and statistically models the typical graph-partitioning cost of separating an arbitrary subsequence of frames from the others. After the behavioral signal, the video is similarly analyzed and the cost of separating the after-signal frames from the before-signal sequences is computed; a lower than typical cost indicates likely contingent reaction. We present a preliminary study in which data were captured and analyzed for algorithmic performance.
Jinhan Lee, Jeffrey F. Kiser, Aaron F. Bobick, Andrea Thomaz
HRI4
2011 Task-aware variations in robot motion
abstract
Social robots can benefit from motion variance because non-repetitive gestures will be more natural and intuitive for human partners. We introduce a new approach for synthesizing variance, both with and without constraints, using a stochastic process. Based on optimal control theory and operational space control, our method can generate an infinite number of variations in real-time that resemble the kinematic and dynamic characteristics from the single input motion sequence. We also introduce a stochastic method to generate smooth but nondeterministic transitions between arbitrary motion variants. Furthermore, we quantitatively evaluate task-aware variance against random white torque noise, operational space control, style-based inverse kinematics, and retargeted human motion to prove that task-aware variance generates human-like motion. Finally, we demonstrate the ability of task-aware variance to maintain velocity and time-dependent features that exist in the input motion.
Michael J. Gielniak, C. Karen Liu, Andrea Thomaz
ICRA3
2011 Automatic State Abstraction from Demonstration
Luis C. Cobo, Peng Zang, Charles L. Isbell Jr., Andrea Thomaz
IJCAI4
2011 Simon plays Simon says: The timing of turn-taking in an imitation game
abstract
Turn-taking is fundamental to the way humans engage in information exchange, but robots currently lack the turn-taking skills required for natural communication. In order to bring effective turn-taking to robots, we must first understand the underlying processes in the context of what is possible to implement. We describe a data collection experiment with an interaction format inspired by “Simon says,” a turn-taking imitation game that engages the channels of gaze, speech, and motion. We analyze data from 23 human subjects interacting with a humanoid social robot and propose the principle of minimum necessary information (MNI) as a factor in determining the timing of the human response.We also describe the other observed phenomena of channel exclusion, efficiency, and adaptation. We discuss the implications of these principles and propose some ways to incorporate our findings into a computational model of turn-taking.
Crystal Chao, Jinhan Lee, Momotaz Begum, Andrea Thomaz
RO-MAN4
2011 Generating anticipation in robot motion
abstract
Robots that display anticipatory motion provide their human partners with greater time to respond in interactive tasks because human partners are aware of robot intent earlier. We create anticipatory motion autonomously from a single motion exemplar by extracting hand and body symbols that communicate motion intent and moving them earlier in the motion. We validate that our algorithm extracts the most salient frame (i.e. the correct symbol) which is the most informative about motion intent to human observers. Furthermore, we show that anticipatory variants allow humans to discern motion intent sooner than motions without anticipation, and that humans are able to reliably predict motion intent prior to the symbol frame when motion is anticipatory. Finally, we quantified the time range for robot motion when humans can perceive intent more accurately and the collaborative social benefits of anticipatory motion are greatest.
Michael J. Gielniak, Andrea Thomaz
RO-MAN2
2011 Effects of responding to, initiating and ensuring joint attention in human-robot interaction
abstract
Inspired by the developmental timeline of joint attention in humans, we propose a conceptual model of joint attention with three parts: responding to joint attention, initiating joint attention, and ensuring joint attention.We conduct two experiments to investigate effects of joint attention in human-robot interaction. The first experiment explores the effects of responding to joint attention. We show that a robot responding to joint attention improves task performance and is perceived as more competent and socially interactive. The second experiment studies the importance of ensuring joint attention in human-robot interaction.We find that a robot's ensuring joint attention behavior is judged as having better performance in human-robot interactive tasks and is perceived as a natural behavior.
Chien-Ming Huang 0001, Andrea Thomaz
RO-MAN2
2011 Human-like action segmentation for option learning
abstract
Robots learning interactively with a human partner has several open questions, one of which is increasing the efficiency of learning. One approach to this problem in the Reinforcement Learning domain is to use options, temporally extended actions, instead of primitive actions. In this paper, we aim to develop a robot system that can discriminate meaningful options from observations of human use of low-level primitive actions. Our approach is inspired by psychological findings about human action parsing, which posits that we attend to low-level statistical regularities to determine action boundary choices. We implement a human-like action segmentation system for automatic option discovery and evaluate our approach and show that option-based learning converges to the optimal solutions faster compared with primitive-action-based learning.
Jaeeun Shim, Andrea Thomaz
RO-MAN2
2010 Transparent active learning for robots
abstract
This research aims to enable robots to learn from human teachers. Motivated by human social learning, we believe that a transparent learning process can help guide the human teacher to provide the most informative instruction. We believe active learning is an inherently transparent machine learning approach because the learner formulates queries to the oracle that reveal information about areas of uncertainty in the underlying model. In this work, we implement active learning on the Simon robot in the form of nonverbal gestures that query a human teacher about a demonstration within the context of a social dialogue. Our preliminary pilot study data show potential for transparency through active learning to improve the accuracy and efficiency of the teaching process. However, our data also seem to indicate possible undesirable effects from the human teacher's perspective regarding balance of the interaction. These preliminary results argue for control strategies that balance leading and following during a social learning interaction.
Crystal Chao, Maya Cakmak, Andrea Thomaz
HRI3
2010 Stylized motion generalization through adaptation of velocity profiles
abstract
Stylized motion is prevalent in the field of Human-Robot Interaction (HRI). Robot designers typically hand craft or work with professional animators to design behaviors for a robot that will be communicative or life-like when interacting with a human partner. A challenge is to apply this stylized trajectory in varied contexts (e.g. performing a stylized gesture with different end-effector constraints). The goal of this research is to create useful, task-based motion with variance that spans the reachable space of the robot and satisfies constraints, while preserving the “style” of the original motion. We claim the appropriate representation for adapting and generalizing a trajectory is not in Cartesian or joint angle space, but rather in joint velocity space, which allows for unspecified initial conditions to be supplied by interaction with the dynamic environment. The benefit of this representation is that a single trajectory can be extended to accomplish similar tasks in the world given constraints in the environment. We present quantitative data using a continuity metric to prove that, given a stylized initial trajectory, we can create smoother generalized motion than with traditional techniques such as cyclic-coordinate descent.
Michael J. Gielniak, C. Karen Liu, Andrea Thomaz
RO-MAN3
2010 Secondary action in robot motion
abstract
Secondary action, a concept borrowed from character animation, improves the animation realism by augmenting natural, passive motion to primary action. We use dynamic simulation to induce three techniques of secondary motion for robot hardware, which exploit actuation passivity to overcome hardware constraints and change the dynamic perception of the robot and its motion characteristics. Results of secondary motion due to internal and external forces are presented including discussion on how to choose the appropriate technique for a particular application.
Michael J. Gielniak, C. Karen Liu, Andrea Thomaz
RO-MAN3
2010 The interplay of context and emotion for non-anthropomorphic robots
abstract
Household robots are becoming commonplace. The application of social cues, such as emotion, has the potential to make such robots easier to use and understand. However, it remains unclear how household robots can or should display emotion, and what considerations should be given to emotive behavior regarding the expected set of contexts in which the robot will operate. In this paper, we report the results of our systematic evaluation of context and emotion recognition of a non-anthropomorphic robot, the iRobot Roomba. Considerations, implications, and future work are discussed.
Bryan Wiltgen, Jenay M. Beer, Keith McGreggor, Karl Jiang, Andrea Thomaz
RO-MAN5
2009 Learning about objects with human teachers
abstract
A general learning task for a robot in a new environment is to learn about objects and what actions/effects they afford. To approach this, we look at ways that a human partner can intuitively help the robot learn, Socially Guided Machine Learning. We present experiments conducted with our robot, Junior, and make six observations characterizing how people approached teaching about objects. We show that Junior successfully used transparency to mitigate errors. Finally, we present the impact of "social" versus "non-social" data sets when training SVM classifiers.
Andrea Thomaz, Maya Cakmak
HRI1
2009 Effective Robot Task Learning by focusing on Task-relevant objects
abstract
In a Robot Learning from Demonstration framework involving environments with many objects, one of the key problems is to decide which objects are relevant to a given task. In this paper, we analyze this problem and propose a biologically-inspired computational model that enables the robot to focus on the task-relevant objects. To filter out incompatible task models, we compute a Task Relevance Value (TRV) for each object, which shows a human demonstrator's implicit indication of the relevance to the task. By combining an intentional action representation with ‘motionese’ [2], our model exhibits recognition capabilities compatible with the way that humans demonstrate. We evaluate the system on demonstrations from five different human subjects, showing its ability to correctly focus on the appropriate objects in these demonstrations.
Kyuhwa Lee, Jinhan Lee, Andrea Thomaz, Aaron F. Bobick
IROS3
2009 Effects of social exploration mechanisms on robot learning
abstract
Social learning in robotics has largely focused on imitation learning. Here we take a broader view and are interested in the multifaceted ways that a social partner can influence the learning process. We implement four social learning mechanisms on a robot: stimulus enhancement, emulation, mimicking, and imitation, and illustrate the computational benefits of each. In particular, we illustrate that some strategies are about directing the attention of the learner to objects and others are about actions. Taken together these strategies form a rich repertoire allowing social learners to use a social partner to greatly impact their learning process. We demonstrate these results in simulation and with physical robot `playmates'.
Maya Cakmak, Nick DePalma, Andrea Thomaz, Rosa I. Arriaga
RO-MAN3
2008 Learning from human teachers with Socially Guided Exploration
abstract
We present a learning mechanism, socially guided exploration, in which a robot learns new tasks through a combination of self-exploration and social interaction. The system's motivational drives (novelty, mastery), along with social scaffolding from a human partner, bias behavior to create learning opportunities for a reinforcement learning mechanism. The system is able to learn on its own, but can flexibly use the guidance of a human partner to improve performance. An experiment with non-expert human subjects shows a human is able to shape the learning process through suggesting actions and drawing attention to goal states. Human guidance results in a task set that is significantly more focused and efficient, while self exploration results in a broader set.
Cynthia Breazeal, Andrea Thomaz
ICRA2
2008 Teachable robots: Understanding human teaching behavior to build more effective robot learners
Andrea Thomaz, Cynthia Breazeal
Artif. Intell.1
2008 Experiments in socially guided exploration: lessons learned in building robots that learn with and without human teachers
abstract
We present a learning system, socially guided exploration, in which a social robot learns new tasks through a combination of self-exploration and social interaction. The system's motivational drives, along with social scaffolding from a human partner, bias behaviour to create learning opportunities for a hierarchical reinforcement learning mechanism. The robot is able to learn on its own, but can flexibly take advantage of the guidance of a human teacher. We report the results of an experiment that analyses what the robot learns on its own as compared to being taught by human subjects. We also analyse the video of these interactions to understand human teaching behaviour and the social dynamics of the human-teacher/robot-learner system. With respect to learning performance, human guidance results in a task set that is significantly more focused and efficient at the tasks the human was trying to teach, whereas self-exploration results in a more diverse set. Analysis of human teaching behaviour reveals insights of social coupling between the human teacher and robot learner, different teaching styles, strong consistency in the kinds and frequency of scaffolding acts across teachers and nuances in the communicative intent behind positive and negative feedback.
Andrea Thomaz, Cynthia Breazeal
Connect. Sci.1
2007 Asymmetric Interpretations of Positive and Negative Human Feedback for a Social Learning Agent
abstract
The ability for people to interact with robots and teach them new skills will be crucial to the successful application of robots in everyday human environments. In order to design agents that learn efficiently and effectively from their instruction, it is important to understand how people, that are not experts in Machine Learning or robotics, will try to teach social robots. In prior work we have shown that human trainers use positive and negative feedback differentially when interacting with a reinforcement learning agent. In this paper we present experiments and implementations on two platforms, a robotic and a computer game platform, that explore the asymmetric communicative intents of positive and negative feedback from a human partner, in particular that negative feedback is both about the past and about intentions for future action.
Andrea Thomaz, Cynthia Breazeal
RO-MAN1
2006 Perspective Taking: An Organizing Principle for Learning in Human-Robot Interaction
Matt Berlin, Jesse Gray, Andrea Thomaz, Cynthia Breazeal
AAAI3
2006 Reinforcement Learning with Human Teachers: Evidence of Feedback and Guidance with Implications for Learning Performance
Andrea Thomaz, Cynthia Breazeal
AAAI1
2006 Experiments in socially guided machine learning: understanding how humans geach
abstract
In Socially Guided Machine Learning we explore the ways in which machine learning can more fully take advantage of natural human interaction. In this work we are studying the role real-time human interaction plays in training assistive robots to perform new tasks. We describe an experimental platform, Sophie's World, and present descriptive analysis of human teaching behavior found in a user study. We report three important observations of how people administer reward and punishment to teach a simulated robot a new task through Reinforcement Learning. People adjust their behavior as they develop a model of the learner, they use the reward channel for guidance as well as feedback, and they may also use it as a motivational channel.
Andrea Thomaz, Guy Hoffman, Cynthia Breazeal
HRI1
2006 Teachable Characters: User Studies, Design Principles, and Learning Performance
Andrea Thomaz, Cynthia Breazeal
IVA1
2006 Reinforcement Learning with Human Teachers: Understanding How People Want to Teach Robots
abstract
While reinforcement learning (RL) is not traditionally designed for interactive supervisory input from a human teacher, several works in both robot and software agents have adapted it for human input by letting a human trainer control the reward signal. In this work, we experimentally examine the assumption underlying these works, namely that the human-given reward is compatible with the traditional RL reward signal. We describe an experimental platform with a simulated RL robot and present an analysis of real-time human teaching behavior found in a study in which untrained subjects taught the robot to perform a new task. We report three main observations on how people administer feedback when teaching a robot a task through reinforcement learning: (a) they use the reward channel not only for feedback, but also for future-directed guidance; (b) they have a positive bias to their feedback -possibly using the signal as a motivational channel; and (c) they change their behavior as they develop a mental model of the robotic learner. In conclusion, we discuss future extensions to RL to accommodate these lessons
Andrea Thomaz, Guy Hoffman, Cynthia Breazeal
RO-MAN1
2005 Effects of nonverbal communication on efficiency and robustness in human-robot teamwork
abstract
Nonverbal communication plays an important role in coordinating teammates' actions for collaborative activities. In this paper, we explore the impact of non-verbal social cues and behavior on task performance by a human-robot team. We report our results from an experiment where naive human subjects guide a robot to perform a physical task using speech and gesture. Both self-report via questionnaire and behavioral analysis of video offer evidence to support our hypothesis that implicit non-verbal communication positively impacts human-robot task performance with respect to understandability of the robot, efficiency of task performance, and robustness to errors that arise from miscommunication.
Cynthia Breazeal, Cory D. Kidd, Andrea Thomaz, Guy Hoffman, Matt Berlin
IROS3
2004 Tutelage and socially guided robot learning
abstract
We view the problem of machine learning as a collaboration between the human and the machine. Inspired by human-style tutelage, we situate the learning problem within a dialog in which social interaction structures the learning experience, providing instruction, directing attention, and controlling the complexity of the task. We present a learning mechanism, implemented on a humanoid robot, to demonstrate that a collaborative dialog framework allows a robot to efficiently learn a task from a human, generalize this ability to a new task configuration, and show commitment to the overall goal of the learned task. We also compare this approach to traditional machine learning approaches.
Andrea Thomaz, Cynthia Breazeal
IROS1
2003 DriftCatcher: The Implicit Social Context of Email
Andrea Thomaz, Ted Selker
INTERACT1