VLDB 2026 Research / reviewers in the wild / expert
Dylan P. Losey
dblp:175/5609
· DBLP profile ↗
28ranked-venue papers
5as first author
19since 2021 · last 2025
0000-0002-8787-5293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 14 since 2021Systems, architecture and hardware · 18 · 3 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Personalizing Interfaces to Humans with User-Friendly PriorsabstractRobots often need to convey information to human users. For example, robots can leverage visual, auditory, and haptic interfaces to display their intent or express their internal state. In some scenarios there are socially agreed upon conventions for what these signals mean: e.g., a red light indicates an autonomous car is slowing down. But as robots develop new capabilities and seek to convey more complex data, the meaning behind their signals is not always mutually understood: one user might think a flashing light indicates the autonomous car is an aggressive driver, while another user might think the same signal means the autonomous car is defensive. In this paper we enable robots to adapt their interfaces to the current user so that the human's personalized interpretation is aligned with the robot's meaning. We start with an information theoretic end-to-end approach, which automatically tunes the interface policy to optimize the correlation between human and robot. But to ensure that this learning policy is intuitive - and to accelerate how quickly the interface adapts to the human - we recognize that humans have priors over how interfaces should function. For instance, humans expect interface signals to be proportional and convex. Our approach biases the robot's interface towards these priors, resulting in signals that are adapted to the current user while still following social expectations. Our simulations and user study results across 15 participants suggest that these priors improve robot-to-human communication. See videos here: https://youtu.be/mO_bz5updDc Benjamin A. Christie, Heramb Nemlekar, Dylan P. Losey |
ICRA | 3 |
| 2025 | Using High-Level Patterns to Estimate How Humans Predict a Robot will BehaveabstractHumans interacting with robots often form predictions of what the robot will do next. For instance, based on the recent behavior of an autonomous car, a nearby human driver might predict that the car is going to remain in the same lane. It is important for the robot to understand the human’s prediction for safe and seamless interaction: e.g., if the autonomous car knows the human thinks it is not merging — but the autonomous car actually intends to merge — then the car can adjust its behavior to prevent an accident. Prior works typically assume that humans make precise predictions of robot behavior. However, recent research on human-human prediction suggests the opposite: humans tend to approximate other agents by predicting their high-level behaviors. We apply this finding to develop a second-order theory of mind approach that enables robots to estimate how humans predict they will behave. To extract these high-level predictions directly from data, we embed the recent human and robot trajectories into a discrete latent space. Each element of this latent space captures a different type of behavior (e.g., merging in front of the human, remaining in the same lane) and decodes into a vector field across the state space that is consistent with the underlying behavior type. We hypothesize that our resulting high-level and course predictions of robot behavior will correspond to actual human predictions. We provide initial evidence in support of this hypothesis through proof-of-concept simulations, testing our method’s predictions against those of real users, and experiments on a real-world interactive driving dataset. Sagar Parekh, Lauren Bramblett, Nicola Bezzo, Dylan P. Losey |
IROS | 4 |
| 2025 | RECON: Reducing Causal Confusion with Human-Placed MarkersabstractImitation learning enables robots to learn new tasks from human examples. One fundamental limitation while learning from humans is causal confusion. Causal confusion occurs when the robot’s observations include both task-relevant and extraneous information: for instance, a robot’s camera might see not only the intended goal, but also clutter and changes in lighting within its environment. Because the robot does not know which aspects of its observations are important a priori, it often misinterprets the human’s examples and fails to learn the desired task. To address this issue, we highlight that — while the robot learner may not know what to focus on — the human teacher does. In this paper we propose that the human proactively marks key parts of their task with small, lightweight beacons. Under our framework (RECON) the human attaches these beacons to task-relevant objects before providing demonstrations: as the human shows examples of the task, beacons track the position of marked objects. We then harness this offline beacon data to train a task-relevant state embedding. Specifically, we embed the robot’s observations to a latent state that is correlated with the measured beacon readings: in practice, this causes the robot to autonomously filter out extraneous observations and make decisions based on features learned from the beacon data. Our simulations and a real robot experiment suggest that this framework for human-placed beacons mitigates causal confusion. Indeed, we find that using RECON significantly reduces the number of demonstrations needed to convey the task, lowering the overall time required for human teaching. See videos here: https://youtu.be/oy85xJvtLSU Robert Ramirez Sanchez, Heramb Nemlekar, Shahabedin Sagheb, Cara M. Nunez, Dylan P. Losey |
IROS | 5 |
| 2025 | PECAN: Personalizing Robot Behaviors through a Learned Canonical SpaceabstractRobots should personalize how they perform tasks to match the needs of individual human users. Today’s robots achieve this personalization by asking for the human’s feedback in the task space. For example, an autonomous car might show the human two different ways to decelerate at stoplights, and ask the human which of these motions they prefer. This current approach to personalization is indirect : Based on the behaviors the human selects (e.g., decelerating slowly), the robot tries to infer their underlying preference (e.g., defensive driving). By contrast, our article develops a learning and interface-based approach that enables humans to directly indicate their desired style. We do this by learning an abstract, low-dimensional, and continuous canonical space from human demonstration data. Each point in the canonical space corresponds to a different style (e.g., defensive or aggressive driving), and users can directly personalize the robot’s behavior by simply clicking on a point. Given the human’s selection, the robot then decodes this canonical style across each task in the dataset—e.g., if the human selects a defensive style, the autonomous car personalizes its behavior to drive defensively when decelerating, passing other cars, or merging onto highways. We refer to our resulting approach as PECAN: Pe rsonalizing Robot Behaviors through a Learned Can onical Space. Our simulations and user studies suggest that humans prefer using PECAN to directly personalize robot behavior (particularly when those users become familiar with PECAN), and that users find the learned canonical space to be intuitive and consistent. See videos here: https://youtu.be/wRJpyr23PKI . Heramb Nemlekar, Robert Ramirez Sanchez, Dylan P. Losey |
ACM Trans. Hum. Robot Interact. | 3 |
| 2024 | Aligning Learning with Communication in Shared AutonomyabstractAssistive robot arms can help humans by partially automating their desired tasks. Consider an adult with motor impairments controlling an assistive robot arm to eat dinner. The robot can reduce the number of human inputs — and how precise those inputs need to be — by recognizing what the human wants (e.g., a fork) and assisting for that task (e.g., moving towards the fork). Prior research has largely focused on learning the human’s task and providing meaningful assistance. But as the robot learns and assists, we also need to ensure that the human understands the robot’s intent (e.g., does the human know the robot is reaching for a fork?). In this paper, we study the effects of communicating learned assistance from the robot back to the human operator. We do not focus on the specific interfaces used for communication. Instead, we develop experimental and theoretical models of a) how communication changes the way humans interact with assistive robot arms, and b) how robots can harness these changes to better align with the human’s intent. We first conduct online and in-person user studies where participants operate robots that provide partial assistance, and we measure how the human’s inputs change with and without communication. With communication, we find that humans are more likely to intervene when the robot incorrectly predicts their intent, and more likely to release control when the robot correctly understands their task. We then use these findings to modify an established robot learning algorithm so that the robot can correctly interpret the human’s inputs when communication is present. Our results from a second in-person user study suggest that this combination of communication and learning outperforms assistive systems that isolate either learning or communication. See videos here: https://youtu.be/BET9yuVTVU4 Joshua Hoegerman, Shahabedin Sagheb, Benjamin A. Christie, Dylan P. Losey |
IROS | 4 |
| 2024 | Kiri-Spoon: A Soft Shape-Changing Utensil for Robot-Assisted FeedingabstractAssistive robot arms have the potential to help disabled or elderly adults eat everyday meals without relying on a caregiver. To provide meaningful assistance, these robots must reach for food items, pick them up, and then carry them to the human’s mouth. Current work equips robot arms with standard utensils (e.g., forks and spoons). But — although these utensils are intuitive for humans — they are not easy for robots to control. If the robot arm does not carefully and precisely orchestrate its motion, food items may fall out of a spoon or slide off of the fork. Accordingly, in this paper we design, model, and test Kiri-Spoon, a novel utensil specifically intended for robot-assisted feeding. Kiri-Spoon combines the familiar shape of traditional utensils with the capabilities of soft grippers. By actuating a kirigami structure the robot can rapidly adjust the curvature of Kiri-Spoon: at one extreme the utensil wraps around food items to make them easier for the robot to pick up and carry, and at the other extreme the utensil returns to a typical spoon shape so that human users can easily take a bite of food. Our studies with able-bodied human operators suggest that robot arms equipped with Kiri-Spoon carry foods more robustly than when leveraging traditional utensils. See videos here: https://youtu.be/nddAniZLFPk Maya N. Keely, Heramb Nemlekar, Dylan P. Losey |
IROS | 3 |
| 2024 | Waypoint-Based Reinforcement Learning for Robot Manipulation TasksabstractRobot arms should be able to learn new tasks. One framework here is reinforcement learning, where the robot is given a reward function that encodes the task, and the robot autonomously learns actions to maximize its reward. Existing approaches to reinforcement learning often frame this problem as a Markov decision process, and learn a policy (or a hierarchy of policies) to complete the task. These policies reason over hundreds of fine-grained actions that the robot arm needs to take: e.g., moving slightly to the right or rotating the end-effector a few degrees. But the manipulation tasks that we want robots to perform can often be broken down into a small number of high-level motions: e.g., reaching an object or turning a handle. In this paper we therefore propose a waypoint-based approach for model-free reinforcement learning. Instead of learning a low-level policy, the robot now learns a trajectory of waypoints, and then interpolates between those waypoints using existing controllers. Our key novelty is framing this waypoint-based setting as a sequence of multi-armed bandits: each bandit problem corresponds to one waypoint along the robot’s motion. We theoretically show that an ideal solution to this reformulation has lower regret bounds than standard frameworks. We also introduce an approximate posterior sampling solution that builds the robot’s motion one waypoint at a time. Results across benchmark simulations and two real-world experiments suggest that this proposed approach learns new tasks more quickly than state-of-the-art baselines. See our website here: https://collab.me.vt.edu/rl-waypoints/ Shaunak A. Mehta, Soheil Habibian, Dylan P. Losey |
IROS | 3 |
| 2024 | LIMIT: Learning Interfaces to Maximize Information TransferabstractRobots can use auditory, visual, or haptic interfaces to convey information to human users. The way these interfaces select signals is typically pre-defined by the designer: for instance, a haptic wristband might vibrate when the robot is moving and squeeze when the robot stops. But different people interpret the same signals in different ways, so that what makes sense to one person might be confusing or unintuitive to another. In this article, we introduce a unified algorithmic formalism for learning co-adaptive interfaces from scratch . Our method does not need to know the human’s task (i.e., what the human is using these signals for). Instead, our insight is that interpretable interfaces should select signals that maximize correlation between the human’s actions and the information the interface is trying to convey. Applying this insight we develop Learning Interfaces to Maximize Information Transfer (LIMIT). LIMIT optimizes a tractable, real-time proxy of information gain in continuous spaces. The first time a person works with our system the signals may appear random; but over repeated interactions, the interface learns a one-to-one mapping between displayed signals and human responses. Our resulting approach is both personalized to the current user and not tied to any specific interface modality. We compare LIMIT to state-of-the-art baselines across controlled simulations, an online survey, and an in-person user study with auditory, visual, and haptic interfaces. Overall, our results suggest that LIMIT learns interfaces that enable users to complete the task more quickly and efficiently, and users subjectively prefer LIMIT to the alternatives. See videos here: https://youtu.be/IvQ3TM1_2fA . Benjamin A. Christie, Dylan P. Losey |
ACM Trans. Hum. Robot Interact. | 2 |
| 2024 | SARI: Shared Autonomy across Repeated InteractionabstractAssistive robot arms try to help their users perform everyday tasks. One way robots can provide this assistance is shared autonomy . Within shared autonomy, both the human and robot maintain control over the robot’s motion: as the robot becomes confident it understands what the human wants, it intervenes to automate the task. But how does the robot know these tasks in the first place? State-of-the-art approaches to shared autonomy often rely on prior knowledge. For instance, the robot may need to know the human’s potential goals beforehand. During long-term interaction these methods will inevitably break down—sooner or later the human will attempt to perform a task that the robot does not expect. Accordingly, in this article we formulate an alternate approach to shared autonomy that learns assistance from scratch. Our insight is that operators repeat important tasks on a daily basis (e.g., opening the fridge, making coffee). Instead of relying on prior knowledge, we therefore take advantage of these repeated interactions to learn assistive policies. We introduce SARI, an algorithm that recognizes the human’s task, replicates similar demonstrations, and returns control when unsure. We then combine learning with control to demonstrate that the error of our approach is uniformly ultimately bounded. We perform simulations to support this error bound, compare our approach to imitation learning baselines, and explore its capacity to assist for an increasing number of tasks. Finally, we conduct three user studies with industry-standard methods and shared autonomy baselines, including a pilot test with a disabled user. Our results indicate that learning shared autonomy across repeated interactions matches existing approaches for known tasks and outperforms baselines on new tasks. See videos of our user studies here: https://youtu.be/3vE4omSvLvc . Ananth Jonnavittula, Shaunak A. Mehta, Dylan P. Losey |
ACM Trans. Hum. Robot Interact. | 3 |
| 2024 | Unified Learning from Demonstrations, Corrections, and Preferences during Physical Human-Robot InteractionabstractHumans can leverage physical interaction to teach robot arms. This physical interaction takes multiple forms depending on the task, the user, and what the robot has learned so far. State-of-the-art approaches focus on learning from a single modality, or combine some interaction types. Some methods do so by assuming that the robot has prior information about the features of the task and the reward structure. By contrast, in this article, we introduce an algorithmic formalism that unites learning from demonstrations, corrections, and preferences. Our approach makes no assumptions about the tasks the human wants to teach the robot; instead, we learn a reward model from scratch by comparing the human’s input to nearby alternatives, i.e., trajectories close to the human’s feedback. We first derive a loss function that trains an ensemble of reward models to match the human’s demonstrations, corrections, and preferences. The type and order of feedback is up to the human teacher: We enable the robot to collect this feedback passively or actively. We then apply constrained optimization to convert our learned reward into a desired robot trajectory. Through simulations and a user study, we demonstrate that our proposed approach more accurately learns manipulation tasks from physical human interaction than existing baselines, particularly when the robot is faced with new or unexpected objectives. Videos of our user study are available at https://youtu.be/FSUJsTYvEKU . Shaunak A. Mehta, Dylan P. Losey |
ACM Trans. Hum. Robot Interact. | 2 |
| 2023 | Towards Robots that Influence Humans over Long-Term InteractionabstractWhen humans interact with robots influence is inevitable. Consider an autonomous car driving near a human: the speed and steering of the autonomous car will affect how the human drives. Prior works have developed frameworks that enable robots to influence humans towards desired behaviors. But while these approaches are effective in the short-term (i.e., the first few human-robot interactions), here we explore long-term influence (i.e., repeated interactions between the same human and robot). Our central insight is that humans are dynamic: people adapt to robots, and behaviors which are influential now may fall short once the human learns to anticipate the robot's actions. With this insight, we experimentally demonstrate that a prevalent game-theoretic formalism for generating influential robot behaviors becomes less effective over repeated interactions. Next, we propose three modifications to Stackelberg games that make the robot's policy both influential and unpredictable. We finally test these modifications across simulations and user studies: our results suggest that robots which purposely make their actions harder to anticipate are better able to maintain influence over long-term interaction. See videos here: https://youtu.be/ydO83cgjZ2Q Shahabedin Sagheb, Ye-Ji Mun, Neema Ahmadian, Benjamin A. Christie, Andrea Bajcsy, Katherine Rose Driggs-Campbell, Dylan P. Losey |
ICRA | 7 |
| 2022 | Communicating Robot Conventions through Shared AutonomyabstractWhen humans control robot arms these robots often need to infer the human's desired task. Prior research on assistive teleoperation and shared autonomy explores how robots can determine the desired task based on the human's joystick inputs. In order to perform this inference the robot relies on an internal mapping between joystick inputs and discrete tasks: e.g., pressing the joystick left indicates that the human wants a plate, while pressing the joystick right indicates a cup. This approach works well after the human understands how the robot interprets their inputs - but inexperienced users still have to learn these mappings through trial and error! Here we recognize that the robot's mapping between tasks and inputs is a convention. There are multiple, equally efficient conventions that the robot could use: rather than passively waiting for the human, we introduce a shared autonomy approach where the robot actively reveals its chosen convention. Across repeated interactions the robot intervenes and exaggerates the arm's motion to demonstrate more efficient inputs while also assisting for the current task. We compare this approach to a state-of-the-art baseline - where users must identify the convention by themselves - as well as written instructions. Our user study results indicate that modifying the robot's behavior to reveal its convention outperforms the baselines and reduces the amount of time that humans spend controlling the robot. See videos of our user study here: https://youtu.be/jROTVOp469I Ananth Jonnavittula, Dylan P. Losey |
ICRA | 2 |
| 2022 | Learning Latent Actions without Human DemonstrationsabstractWe can make it easier for disabled users to control assistive robots by mapping the user's low-dimensional joystick inputs to high-dimensional, complex actions. Prior works learn these mappings from human demonstrations: a non-disabled human either teleoperates or kinesthetically guides the robot arm through a variety of motions, and the robot learns to reproduce the demonstrated behaviors. But this framework is often impractical — disabled users will not always have access to external demonstrations! Here we instead learn diverse teleoperation mappings without either human demonstrations or pre-defined tasks. Under our unsupervised approach the robot first optimizes for object state entropy: i.e., the robot autonomously learns to push, pull, open, close, or otherwise change the state of nearby objects. We then embed these diverse, object-oriented behaviors into a latent space for real-time control: now pressing the joystick causes the robot to perform dexterous motions like pushing or opening. We experimentally show that — with a best-case human operator — our unsupervised approach actually outperforms the teleoperation mappings learned from human demonstrations, particularly if those demonstrations are noisy or imperfect. But our user study results were less clear-cut: although our approach enabled participants to complete tasks more quickly and with fewer changes of direction, users were confused when the unsupervised robot learned unexpected behaviors. See videos of the user study here: https://youtu.be/BkqHQjsUKDg Shaunak A. Mehta, Sagar Parekh, Dylan P. Losey |
ICRA | 3 |
| 2022 | Assisting Operators of Articulated Machinery with Optimal Planning and Goal InferenceabstractOperating an articulated machine is a complex and hierarchical task, involving several levels of decision making. Motivated by the timber-harvesting applications of these machines, we are interested in developing a collaborative framework for operating an articulated machine/robot in order to increase its level of autonomy. In this paper, we consider two problems in the context of collaborative operation of a feller-buncher: first, the problem of planning a sequence of cut/grasp/bunch tasks for the trees in the vicinity of the machine. Here we propose a human-inspired planning algorithm based on our observations of the operators in the field. Then, a Markov Decision Process (MDP) framework is provided, which enables us to obtain an optimal sequence of tasks. We provide numerical illustrations of how our MDP framework works. Second is the problem of inferring the operator's goal from the motions of the machine. The goal inference algorithm presented here enables the robot equipped with the planning intelligence to perceive the human's intent in real-time. We evaluate the performance of our goal inference algorithm through a user-study with a feller-buncher simulator. The results show the benefits of our algorithm over a robot that assumes the human is moving to the closest target. Ehsan Yousefi, Dylan P. Losey, Inna Sharf |
ICRA | 2 |
| 2022 | RILI: Robustly Influencing Latent IntentabstractWhen robots interact with human partners, often these partners change their behavior in response to the robot. On the one hand this is challenging because the robot must learn to coordinate with a dynamic partner. But on the other hand - if the robot understands these dynamics - it can harness its own behavior, influence the human, and guide the team towards effective collaboration. Prior research enables robots to learn to influence other robots or simulated agents. In this paper we extend these learning approaches to now influence humans. What makes humans especially hard to influence is that - not only do humans react to the robot - but the way a single user reacts to the robot may change over time, and different humans will respond to the same robot behavior in different ways. We therefore propose a robust approach that learns to influence changing partner dynamics. Our method first trains with a set of partners across repeated interactions, and learns to predict the current partner's behavior based on the previous states, actions, and rewards. Next, we rapidly adapt to new partners by sampling trajectories the robot learned with the original partners, and then leveraging those existing behaviors to influence the new partner dynamics. We compare our resulting algorithm to state-of-the-art baselines across simulated environments and a user study where the robot and participants collaborate to build towers. We find that our approach outperforms the alternatives, even when the partner follows new or unexpected dynamics. Videos of the user study are available here: https://youtu.be/1YsWM8An18g Sagar Parekh, Soheil Habibian, Dylan P. Losey |
IROS | 3 |
| 2022 | Here's What I've Learned: Asking Questions that Reveal Reward LearningabstractRobots can learn from humans by asking questions. In these questions, the robot demonstrates a few different behaviors and asks the human for their favorite. But how should robots choose which questions to ask? Today’s robots optimize for informative questions that actively probe the human’s preferences as efficiently as possible. But while informative questions make sense from the robot’s perspective, human onlookers may find them arbitrary and misleading . For example, consider an assistive robot learning to put away the dishes. Based on your answers to previous questions this robot knows where it should stack each dish; however, the robot is unsure about right height to carry these dishes. A robot optimizing only for informative questions focuses purely on this height: it shows trajectories that carry the plates near or far from the table, regardless of whether or not they stack the dishes correctly. As a result, when we see this question, we mistakenly think that the robot is still confused about where to stack the dishes! In this article, we formalize active preference-based learning from the human’s perspective. We hypothesize that—from the human’s point-of-view —the robot’s questions reveal what the robot has and has not learned. Our insight enables robots to use questions to make their learning process transparent to the human operator. We develop and test a model that robots can leverage to relate the questions they ask to the information these questions reveal. We then introduce a tradeoff between informative and revealing questions that considers both human and robot perspectives: a robot that optimizes for this tradeoff actively gathers information from the human while simultaneously keeping the human up to date with what it has learned. We evaluate our approach across simulations, online surveys, and in-person user studies. We find that robots, which consider the human’s point of view learn just as quickly as state-of-the-art baselines while also communicating what they have learned to the human operator. Videos of our user studies and results are available here: https://youtu.be/tC6y_jHN7Vw. Soheil Habibian, Ananth Jonnavittula, Dylan P. Losey |
ACM Trans. Hum. Robot Interact. | 3 |
| 2021 | I Know What You Meant: Learning Human Objectives by (Under)estimating Their Choice SetabstractAssistive robots have the potential to help people perform everyday tasks. However, these robots first need to learn what it is their user wants them to do. Teaching assistive robots is hard for inexperienced users, elderly users, and users living with physical disabilities, since often these individuals are unable to show the robot their desired behavior. We know that inclusive learners should give human teachers credit for what they cannot demonstrate. But today’s robots do the opposite: they assume every user is capable of providing any demonstration. As a result, these robots learn to mimic the demonstrated behavior, even when that behavior is not what the human really meant! Here we propose a different approach to reward learning: robots that reason about the user’s demonstrations in the context of similar or simpler alternatives. Unlike prior works — which err towards overestimating the human’s capabilities — here we err towards underestimating what the human can input (i.e., their choice set). Our theoretical analysis proves that underestimating the human’s choice set is riskaverse, with better worst-case performance than overestimating. We formalize three properties to generate similar and simpler alternatives. Across simulations and a user study, our resulting algorithm better extrapolates the human’s objective. See the user study here: https://youtu.be/RgbH2YULVRo. Ananth Jonnavittula, Dylan P. Losey |
ICRA | 2 |
| 2021 | Learning Human Objectives from Sequences of Physical CorrectionsabstractWhen personal, assistive, and interactive robots make mistakes, humans naturally and intuitively correct those mistakes through physical interaction. In simple situations, one correction is sufficient to convey what the human wants. But when humans are working with multiple robots or the robot is performing an intricate task often the human must make several corrections to fix the robot’s behavior. Prior research assumes each of these physical corrections are independent events, and learns from them one-at-a-time. However, this misses out on crucial information: each of these interactions are interconnected, and may only make sense if viewed together. Alternatively, other work reasons over the final trajectory produced by all of the human’s corrections. But this method must wait until the end of the task to learn from corrections, as opposed to inferring from the corrections in an online fashion. In this paper we formalize an approach for learning from sequences of physical corrections during the current task. To do this we introduce an auxiliary reward that captures the human’s trade-off between making corrections which improve the robot’s immediate reward and long-term performance. We evaluate the resulting algorithm in remote and in-person human-robot experiments, and compare to both independent and final baselines. Our results indicate that users are best able to convey their objective when the robot reasons over their sequence of corrections. Mengxi Li, Alper Canberk, Dylan P. Losey, Dorsa Sadigh |
ICRA | 3 |
| 2021 | Learning to Share Autonomy Across Repeated InteractionabstractWheelchair-mounted robotic arms (and other assistive robots) should help their users perform everyday tasks. One way robots can provide this assistance is shared autonomy. Within shared autonomy, both the human and robot maintain control over the robot’s motion: as the robot becomes confident it understands what the human wants, it increasingly intervenes to automate the task. But how does the robot know what tasks the human may want to perform in the first place? Today’s shared autonomy approaches often rely on prior knowledge: for example, the robot must know the set of possible human goals a priori. In the long-term, however, this prior knowledge will inevitably break down — sooner or later the human will reach for a goal that the robot did not expect. In this paper we propose a learning approach to shared autonomy that takes advantage of repeated interactions. Learning to assist humans would be impossible if they performed completely different tasks at every interaction: but our insight is that users living with physical disabilities repeat important tasks on a daily basis (e.g., opening the fridge, making coffee, and having dinner). We introduce an algorithm that exploits these repeated interactions to recognize the human’s task, replicate similar demonstrations, and return control when unsure. As the human repeatedly works with this robot, our approach continually learns to assist tasks that were never specified beforehand: these tasks include both discrete goals (e.g., reaching a cup) and continuous skills (e.g., opening a drawer). Across simulations and an in-person user study, we demonstrate that robots leveraging our approach match existing shared autonomy methods for known goals, and outperform imitation learning baselines on new tasks. See videos here: https://youtu.be/NazeLVbQ2og Ananth Jonnavittula, Dylan P. Losey |
IROS | 2 |
| 2020 | When Humans Aren't Optimal: Robots that Collaborate with Risk-Aware HumansabstractIn order to collaborate safely and efficiently, robots need to anticipate how their human partners will behave. Some of today's robots model humans as if they were also robots, and assume users are always optimal. Other robots account for human limitations, and relax this assumption so that the human is noisily rational. Both of these models make sense when the human receives deterministic rewards: i.e., gaining either $100 or $130 with certainty. But in real-world scenarios, rewards are rarely deterministic. Instead, we must make choices subject to risk and uncertainty-and in these settings, humans exhibit a cognitive bias towards suboptimal behavior. For example, when deciding between gaining $100 with certainty or $130 only 80% of the time, people tend to make the risk-averse choice-even though it leads to a lower expected gain! In this paper, we adopt a well-known Risk-Aware human model from behavioral economics called Cumulative Prospect Theory and enable robots to leverage this model during human-robot interaction (HRI). In our user studies, we offer supporting evidence that the Risk-Aware model more accurately predicts suboptimal human behavior. We find that this increased modeling accuracy results in safer and more efficient human-robot collaboration. Overall, we extend existing rational human models so that collaborative robots can anticipate and plan around suboptimal human behavior during HRI. Minae Kwon, Erdem Biyik, Aditi Talati, Karan Bhasin, Dylan P. Losey, Dorsa Sadigh |
HRI | 5 |
| 2020 | Controlling Assistive Robots with Learned Latent ActionsabstractAssistive robotic arms enable users with physical disabilities to perform everyday tasks without relying on a caregiver. Unfortunately, the very dexterity that makes these arms useful also makes them challenging to teleoperate: the robot has more degrees-of-freedom than the human can directly coordinate with a handheld joystick. Our insight is that we can make assistive robots easier for humans to control by leveraging latent actions. Latent actions provide a lowdimensional embedding of high-dimensional robot behavior: for example, one latent dimension might guide the assistive arm along a pouring motion. In this paper, we design a teleoperation algorithm for assistive robots that learns latent actions from task demonstrations. We formulate the controllability, consistency, and scaling properties that user-friendly latent actions should have, and evaluate how different lowdimensional embeddings capture these properties. Finally, we conduct two user studies on a robotic arm to compare our latent action approach to both state-of-the-art shared autonomy baselines and a teleoperation strategy currently used by assistive arms. Participants completed assistive eating and cooking tasks more efficiently when leveraging our latent actions, and also subjectively reported that latent actions made the task easier to perform. The video accompanying this paper can be found at: https://youtu.be/wjnhrzugBj4. Dylan P. Losey, Krishnan Srinivasan, Ajay Mandlekar, Animesh Garg, Dorsa Sadigh |
ICRA | 1 |
| 2020 | Learning User-Preferred Mappings for Intuitive Robot ControlabstractWhen humans control drones, cars, and robots, we often have some preconceived notion of how our inputs should make the system behave. Existing approaches to teleoperation typically assume a one-size-fits-all approach, where the designers pre-define a mapping between human inputs and robot actions, and every user must adapt to this mapping over repeated interactions. Instead, we propose a personalized method for learning the human's preferred or preconceived mapping from a few robot queries. Given a robot controller, we identify an alignment model that transforms the human's inputs so that the controller's output matches their expectations. We make this approach data-efficient by recognizing that human mappings have strong priors: we expect the input space to be proportional, reversable, and consistent. Incorporating these priors ensures that the robot learns an intuitive mapping from few examples. We test our learning approach in robot manipulation tasks inspired by assistive settings, where each user has different personal preferences and physical capabilities for teleoperating the robot arm. Our simulated and experimental results suggest that learning the mapping between inputs and robot actions improves objective and subjective performance when compared to manually defined alignments or learned alignments without intuitive priors. The supplementary video showing these user studies can be found at: https://youtu.be/rKHka0_48-Q. Mengxi Li, Dylan P. Losey, Jeannette Bohg, Dorsa Sadigh |
IROS | 2 |
| 2020 | Learning the Correct Robot Trajectory in Real-Time from Physical Human InteractionsabstractWe present a learning and control strategy that enables robots to harness physical human interventions to update their trajectory and goal during autonomous tasks. Within the state of the art, the robot typically reacts to physical interactions by modifying a local segment of its trajectory, or by searching for the global trajectory offline, using either replanning or previous demonstrations. Instead, we explore a one-shot approach: here, the robot updates its entire trajectory and goal in real time without relying on multiple iterations, offline demonstrations, or replanning. Our solution is grounded in optimal control and gradient descent, and extends linear-quadratic regulator controllers to generalize across methods that locally or globally modify the robot’s underlying trajectory. In the best case, this Linear-quadratic regulator + Learning approach matches the optimal offline response to physical interactions, and—in more challenging cases—our strategy is robust to noisy and unexpected human corrections. We compare the proposed solution against other real-time strategies in a user study and demonstrate its efficacy in terms of both objective and subjective measures. Dylan P. Losey, Marcia Kilchenman O'Malley |
ACM Trans. Hum. Robot Interact. | 1 |
| 2019 | Robots that Take Advantage of Human TrustabstractHumans often assume that robots are rational. We believe robots take optimal actions given their objective; hence, when we are uncertain about what the robot's objective is, we interpret the robot's actions as optimal with respect to our estimate of its objective. This approach makes sense when robots straightforwardly optimize their objective, and enables humans to learn what the robot is trying to achieve. However, our insight is that-when robots are aware that humans learn by trusting that the robot actions are rational-intelligent robots do not act as the human expects; instead, they take advantage of the human's trust, and exploit this trust to more efficiently optimize their own objective. In this paper, we formally model instances of human-robot interaction (HRI) where the human does not know the robot's objective using a two-player game. We formulate different ways in which the robot can model the uncertain human, and compare solutions of this game when the robot has conservative, optimistic, rational, and trusting human models. In an offline linear-quadratic case study and a real-time user study, we show that trusting human models can naturally lead to communicative robot behavior, which influences end-users and increases their involvement. Dylan P. Losey, Dorsa Sadigh |
IROS | 1 |
| 2018 | Learning from Physical Human Corrections, One Feature at a TimeabstractWe focus on learning robot objective functions from human guidance: specifically, from physical corrections provided by the person while the robot is acting. Objective functions are typically parametrized in terms of features, which capture aspects of the task that might be important. When the person intervenes to correct the robot»s behavior, the robot should update its understanding of which features matter, how much, and in what way. Unfortunately, real users do not provide optimal corrections that isolate exactly what the robot was doing wrong. Thus, when receiving a correction, it is difficult for the robot to determine which features the person meant to correct, and which features were changed unintentionally. In this paper, we propose to improve the efficiency of robot learning during physical interactions by reducing unintended learning. Our approach allows the human-robot team to focus on learning one feature at a time, unlike state-of-the-art techniques that update all features at once. We derive an online method for identifying the single feature which the human is trying to change during physical interaction, and experimentally compare this one-at-a-time approach to the all-at-once baseline in a user study. Our results suggest that users teaching one-at-a-time perform better, especially in tasks that require changing multiple features. Andrea Bajcsy, Dylan P. Losey, Marcia Kilchenman O'Malley, Anca D. Dragan |
HRI | 2 |
| 2018 | Trajectory Deformations From Physical Human-Robot InteractionabstractRobots are finding new applications where physical interaction with a human is necessary, such as manufacturing, healthcare, and social tasks. Accordingly, the field of physical human-robot interaction (pHRI) has leveraged impedance control approaches, which support compliant interactions between human and robot. However, a limitation of traditional impedance control is that-despite provisions for the human to modify the robot's current trajectory-the human cannot affect the robot's future desired trajectory through pHRI. In this paper, we present an algorithm for physically interactive trajectory deformations which, when combined with impedance control, allows the human to modulate both the actual and desired trajectories of the robot. Unlike related works, our method explicitly deforms the future desired trajectory based on forces applied during pHRI, but does not require constant human guidance. We present our approach and verify that this method is compatible with traditional impedance control. Next, we use constrained optimization to derive the deformation shape. Finally, we describe an algorithm for real-time implementation, and perform simulations to test the arbitration parameters. Experimental results demonstrate reduction in the human's effort and improvement in the movement quality when compared to pHRI with impedance control alone. Dylan P. Losey, Marcia Kilchenman O'Malley |
IEEE Trans. Robotics | 1 |
| 2017 | Effects of discretization on the K-width of series elastic actuatorsabstractRigid haptic devices enable humans to physically interact with virtual environments, and the range of impedances that can be safely rendered using these rigid devices is quantified by the Z-Width metric. Series elastic actuators (SEAs) similarly modulate the impedance felt by the human operator when interacting with a robotic device, and, in particular, the robot's perceived stiffness can be controlled by changing the elastic element's equilibrium position. In this paper, we explore the K-Width of SEAs, while specifically focusing on how discretization inherent in the computer-control architecture affects the system's passivity. We first propose a hybrid model for a single degree-of-freedom (DoF) SEA based on prior hybrid models for rigid haptic systems. Next, we derive a closed-form bound on the K-Width of SEAs that is a generalization of known constraints for both rigid haptic systems and continuous time SEA models. This bound is first derived under a continuous time approximation, and is then numerically supported with discrete time analysis. Finally, experimental results validate our finding that large pure masses are the most destabilizing operator in human-SEA interactions, and demonstrate the accuracy of our theoretical K-Width bound. Dylan P. Losey, Marcia Kilchenman O'Malley |
ICRA | 1 |
| 2016 | Minimal Assist-as-Needed Controller for Upper Limb Robotic RehabilitationabstractRobotic rehabilitation of the upper limb following neurological injury is most successful when subjects are engaged in the rehabilitation protocol. Developing assistive control strategies that maximize subject participation is accordingly an active area of research, with aims to promote neural plasticity and, in turn, increase the potential for recovery of motor coordination. Unfortunately, state-of-the-art control strategies either ignore more complex subject capabilities or assume underlying patterns govern subject behavior and may therefore intervene suboptimally. In this paper, we present a minimal assist-as-needed (mAAN) controller for upper limb rehabilitation robots. The controller employs sensorless force estimation to dynamically determine subject inputs without any underlying assumptions as to the nature of subject capabilities and computes a corresponding assistance torque with adjustable ultimate bounds on position error. Our adaptive input estimation scheme is shown to yield fast, stable, and accurate measurements regardless of subject interaction and exceeds the performance of current approaches that estimate only position-dependent force inputs from the user. Two additional algorithms are introduced in this paper to further promote active participation of subjects with varying degrees of impairment. First, a bound modification algorithm is described, which alters allowable error. Second, a decayed disturbance rejection algorithm is presented, which encourages subjects who are capable of leading the reference trajectory. The mAAN controller and accompanying algorithms are demonstrated experimentally with healthy subjects in the RiceWrist-S exoskeleton. Ali Utku Pehlivan, Dylan P. Losey, Marcia Kilchenman O'Malley |
IEEE Trans. Robotics | 2 |