EDBT 2026 Demo / reviewers in the wild / expert
Daniel S. Brown
dblp:141/7769
· DBLP profile ↗
30ranked-venue papers
11as first author
18since 2021 · last 2025
0000-0002-9570-1832ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 9 first-author · 18 since 2021Systems, architecture and hardware · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Human Input to Enable Robust, Interactive, and Aligned AI SystemsabstractEnsuring that AI systems do what we, as humans, actually want them to do, is one of the biggest open research challenges in AI alignment and safety. My research seeks to directly address this challenge by enabling AI systems to interact with humans to learn aligned and robust behaviors. The way in which robots and other AI systems behave is often the result of optimizing a reward function. However, manually designing good reward functions is highly challenging and error prone, even for domain experts. Consider trying to write down a reward function that describes good driving behavior or how you like your bed made in the morning. While reward functions for these tasks are difficult to manually specify, human feedback in the form of demonstrations or preferences are often much easier to obtain. However, human data is often difficult to interpret, due to ambiguity and noise. Thus, it is critical that AI systems take into account epistemic uncertainty over the human's true intent. My talk will give an overview of my lab's progress along the following fundamental research areas: (1) efficiently maintaining uncertainty over human intent, (2) directly optimizing behavior to be robust to uncertainty over human intent, and (3) actively querying for additional human input to reduce uncertainty over human intent. Daniel S. Brown |
AAAI | 1 |
| 2025 | Toward Zero-Shot User Intent Recognition in Shared AutonomyabstractA fundamental challenge of shared autonomy is to use high-DoF robots to assist, rather than hinder, humans by first inferring user intent and then empowering the user to achieve their intent. Although successful, prior methods either rely heavily on a priori knowledge of all possible human intents or require many demonstrations and interactions with the human to learn these intents before being able to assist the user. We propose and study a zero-shot, vision-only shared autonomy (VOSA) frame-work designed to allow robots to use end-effector vision to estimate zero-shot human intents in conjunction with blended control to help humans accomplish manipulation tasks with unknown and dynamically changing object locations. To demonstrate the effectiveness of our VOSA framework, we instantiate a simple version of VOSA on a Kinova Gen3 manipulator and evaluate our system by conducting a user study on three tabletop manipulation tasks. The performance of VOSA matches that of an oracle baseline model that receives privileged knowledge of possible human intents while also requiring significantly less effort than unassisted teleoperation. In more realistic settings, where the set of possible human intents is fully or partially unknown, we demonstrate that VOSA requires less human effort and time than baseline approaches while being preferred by a majority of the participants. Our results demonstrate the efficacy and efficiency of using off-the-shelf vision algorithms to enable flexible and beneficial shared control of a robot manipulator. Code and videos available here: https://sites.google.com/view/zeroshot-sharedautonomy/home Atharv Belsare, Zohre Karimi, Connor Mattson, Daniel S. Brown |
HRI | 4 |
| 2025 | Discovery and Deployment of Emergent Robot Swarm Behaviors via Representation Learning and Real2Sim2Real Transfer
Connor Mattson, Varun Raveendra, Ricardo Vega, Cameron Nowzari, Daniel S. Drew, Daniel S. Brown |
AAMAS | 6 |
| 2024 | Autonomous Assessment of Demonstration Sufficiency via Bayesian Inverse Reinforcement LearningabstractWe examine the problem of determining demonstration sufficiency: how can a robot self-assess whether it has received enough demonstrations from an expert to ensure a desired level of performance? To address this problem, we propose a novel self-assessment approach based on Bayesian inverse reinforcement learning and value-at-risk, enabling learning-from-demonstration ("LfD") robots to compute high-confidence bounds on their performance and use these bounds to determine when they have a sufficient number of demonstrations. We propose and evaluate two definitions of sufficiency: (1) normalized expected value difference, which measures regret with respect to the human's unobserved reward function, and (2) percent improvement over a baseline policy. We demonstrate how to formulate high-confidence bounds on both of these metrics. We evaluate our approach in simulation for both discrete and continuous state-space domains and illustrate the feasibility of developing a robotic system that can accurately evaluate demonstration sufficiency. We also show that the robot can utilize active learning in asking for demonstrations from specific states which results in fewer demos needed for the robot to still maintain high confidence in its policy. Finally, via a user study, we show that our approach successfully enables robots to perform at users' desired performance levels, without needing too many or perfectly optimal demonstrations. Tu Trinh, Daniel S. Brown |
HRI | 3 |
| 2024 | Bayesian Constraint Inference from User Demonstrations Based on Margin-Respecting Preference ModelsabstractIt is crucial for robots to be aware of the presence of constraints in order to acquire safe policies. However, explicitly specifying all constraints in an environment can be a challenging task. State-of-the-art constraint inference algorithms learn constraints from demonstrations, but tend to be computationally expensive and prone to instability issues. In this paper, we propose a novel Bayesian method that infers constraints based on preferences over demonstrations. The main advantages of our proposed approach are that it 1) infers constraints without calculating a new policy at each iteration, 2) uses a simple and more realistic ranking of groups of demonstrations, without requiring pairwise comparisons over all demonstrations, and 3) adapts to cases where there are varying levels of constraint violation. Our empirical results demonstrate that our proposed Bayesian approach infers constraints of varying severity, more accurately than state-of-the-art constraint inference methods. Code and videos: https://sites.google.com/berkeley.edu/pbicrl. Dimitris Papadimitriou, Daniel S. Brown |
ICRA | 2 |
| 2023 | The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback TypesabstractWhen inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a "rationality coefficient" capturing how much noise or entropy we expect to see in the human behavior. Prior work typically sets the rationality level to a constant value, regardless of the type, or quality, of human feedback. However, in many settings, giving one type of feedback (e.g. a demonstration) may be much more difficult than a different type of feedback (e.g. answering a comparison query). Thus, we expect to see more or less noise depending on the type of human feedback. In this work, we advocate that grounding the rationality coefficient in real data for each feedback type, rather than assuming a default value, has a significant positive effect on reward learning. We test this in both simulated experiments and in a user study with real human feedback. We find that overestimating human rationality can have dire effects on reward learning accuracy and regret. We also find that fitting the rationality coefficient to human data enables better reward learning, even when the human deviates significantly from the noisy-rational choice model due to systematic biases. Further, we find that the rationality level affects the informativeness of each feedback type: surprisingly, demonstrations are not always the most informative---when the human acts very suboptimally, comparisons actually become more informative, even when the rationality level is the same for both. Ultimately, our results emphasize the importance and advantage of paying attention to the assumed human-rationality-level, especially when agents actively learn from multiple types of human feedback. Gaurav R. Ghosal, Matthew Zurek, Daniel S. Brown, Anca D. Dragan |
AAAI | 3 |
| 2023 | Leveraging Human Feedback to Evolve and Discover Novel Emergent Behaviors in Robot SwarmsabstractRobot swarms often exhibit emergent behaviors that are fascinating to observe; however, it is often difficult to predict what swarm behaviors can emerge under a given set of agent capabilities. We seek to efficiently leverage human input to automatically discover a taxonomy of collective behaviors that can emerge from a particular multi-agent system, without requiring the human to know beforehand what behaviors are interesting or even possible. Our proposed approach adapts to user preferences by learning a similarity space over swarm collective behaviors using self-supervised learning and human-in-the-loop queries. We combine our learned similarity metric with novelty search and clustering to explore and categorize the space of possible swarm behaviors. We also propose several general-purpose heuristics that improve the efficiency of our novelty search by prioritizing robot controllers that are likely to lead to interesting emergent behaviors. We test our approach in simulation on two robot capability models and show that our methods consistently discover a richer set of emergent behaviors than prior work. Code, videos, and datasets are available at https://sites.google.com/view/evolving-novel-swarms. Connor Mattson, Daniel S. Brown |
GECCO | 2 |
| 2023 | SIRL: Similarity-based Implicit Representation LearningabstractWhen robots learn reward functions using high capacity models that take raw state directly as input, they need to both learn a representation for what matters in the task --- the task "features" --- as well as how to combine these features into a single objective. If they try to do both at once from input designed to teach the full reward function, it is easy to end up with a representation that contains spurious correlations in the data, which fails to generalize to new settings. Instead, our ultimate goal is to enable robots to identify and isolate the causal features that people actually care about and use when they represent states and behavior. Our idea is that we can tune into this representation by asking users what behaviors they consider similar: behaviors will be similar if the features that matter are similar, even if low-level behavior is different; conversely, behaviors will be different if even one of the features that matter differs. This, in turn, is what enables the robot to disambiguate between what needs to go into the representation versus what is spurious, as well as what aspects of behavior can be compressed together versus not. The notion of learning representations based on similarity has a nice parallel in contrastive learning, a self-supervised representation learning technique that maps visually similar data points to similar embeddings, where similarity is defined by a designer through data augmentation heuristics. By contrast, in order to learn the representations that people use, so we can learn their preferences and objectives, we use their definition of similarity. In simulation as well as in a user study, we show that learning through such similarity queries leads to representations that, while far from perfect, are indeed more generalizable than self-supervised and task-input alternatives. Andreea Bobu, Rohin Shah, Daniel S. Brown, Anca D. Dragan |
HRI | 4 |
| 2023 | Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D. Dragan, Daniel S. Brown |
ICLR | 5 |
| 2023 | Contextual Reliability: When Different Features Matter in Different ContextsabstractDeep neural networks often fail catastrophically by relying on spurious correlations. Most prior work assumes a clear dichotomy into spurious and reliable features; however, this is often unrealistic. For example, most of the time we do not want an autonomous car to simply copy the speed of surrounding cars---we don't want our car to run a red light if a neighboring car does so. However, we cannot simply enforce invariance to next-lane speed, since it could provide valuable information about an unobservable pedestrian at a crosswalk. Thus, universally ignoring features that are sometimes (but not always) reliable can lead to non-robust performance. We formalize a new setting called contextual reliability which accounts for the fact that the "right" features to use may vary depending on the context. We propose and analyze a two-stage framework called Explicit Non-spurious feature Prediction (ENP) which first identifies the relevant features to use for a given context, then trains a model to rely exclusively on these features. Our work theoretically and empirically demonstrates the advantages of ENP over existing methods and provides new benchmarks for contextual reliability. Gaurav R. Ghosal, Amrith Setlur, Daniel S. Brown, Anca D. Dragan, Aditi Raghunathan |
ICML | 3 |
| 2023 | Efficient Preference-Based Reinforcement Learning Using Learned Dynamics ModelsabstractPreference-based reinforcement learning (PbRL) can enable robots to learn to perform tasks based on an individual's preferences without requiring a hand-crafted re-ward function. However, existing approaches either assume access to a high-fidelity simulator or analytic model or take a model-free approach that requires extensive, possibly unsafe online environment interactions. In this paper, we study the benefits and challenges of using a learned dynamics model when performing PbRL. In particular, we provide evidence that a learned dynamics model offers the following benefits when performing PbRL: (1) preference elicitation and policy optimization require significantly fewer environment interactions than model-free PbRL, (2) diverse preference queries can be synthesized safely and efficiently as a byproduct of standard model-based RL, and (3) reward pre-training based on suboptimal demonstrations can be performed without any environmental interaction. Our paper provides empirical ev-idence that learned dynamics models enable robots to learn customized policies based on user preferences in ways that are safer and more sample efficient than prior preference learning approaches. Supplementary materials and code are available at https://sites.google.com/berkeley.edu/mop-rl. Gaurav Datta, Ellen R. Novoseller, Daniel S. Brown |
ICRA | 4 |
| 2022 | LEGS: Learning Efficient Grasp Sets for Exploratory GraspingabstractWhile deep learning has enabled significant progress in designing general purpose robot grasping systems, there remain objects which still pose challenges for these systems. Recent work on Exploratory Grasping has formalized the problem of systematically exploring grasps on these adversarial objects and explored a multi-armed bandit model for identifying high-quality grasps on each object stable pose. However, these systems are still limited to exploring a small number or grasps on each object. We present Learned Efficient Grasp Sets (LEGS), an algorithm that efficiently explores thousands of possible grasps by maintaining small active sets of promising grasps and determining when it can stop exploring the object with high confidence. Experiments suggest that LEGS can identify a high-quality grasp more efficiently than prior algorithms which do not use active sets. In simulation experiments, we measure the gap between the success probability of the best grasp identified by LEGS, baselines, and the most-robust grasp (verified ground truth). After 3000 exploration steps, LEGS outperforms baseline algorithms on 10/14 and 25/39 objects on the Dex-Net Adversarial and EGAD! datasets respectively. We then evaluate LEGS in physical experiments; trials on 3 challenging objects suggest that LEGS converges to high-performing grasps significantly faster than baselines. See https://sites.google.com/view/LEGS-exp-grasping for supplemental material and videos. Letian Fu, Michael Danielczuk, Ashwin Balakrishna, Daniel S. Brown, Jeffrey Ichnowski, Eugen Solowjow, Kenneth Y. Goldberg |
ICRA | 4 |
| 2022 | Teaching Robots to Span the Space of Functional Expressive MotionabstractOur goal is to enable robots to perform functional tasks in emotive ways, be it in response to their users' emotional states, or expressive of their confidence levels. Prior work has proposed learning independent cost functions from user feedback for each target emotion, so that the robot may optimize it alongside task and environment specific objectives for any situation it encounters. However, this approach is inefficient when modeling multiple emotions and unable to generalize to new ones. In this work, we leverage the fact that emotions are not independent of each other: they are related through a latent space of Valence-Arousal-Dominance (VAD). Our key idea is to learn a model for how trajectories map onto VAD with user labels. Considering the distance between a trajectory's mapping and a target VAD allows this single model to represent cost functions for all emotions. As a result 1) all user feedback can contribute to learning about every emotion; 2) the robot can generate trajectories for any emotion in the space instead of only a few predefined ones; and 3) the robot can respond emotively to user-generated natural language by mapping it to a target VAD. We introduce a method that interactively learns to map trajectories to this latent space and test it in simulation and in a user study. In experiments, we use a simple vacuum robot as well as the Cassie biped. Arjun Sripathy, Andreea Bobu, Zhongyu Li 0003, Koushil Sreenath, Daniel S. Brown, Anca D. Dragan |
IROS | 5 |
| 2022 | Monte Carlo Augmented Actor-Critic for Sparse Reward Deep Reinforcement Learning from Suboptimal DemonstrationsabstractProviding densely shaped reward functions for RL algorithms is often exceedingly challenging, motivating the development of RL algorithms that can learn from easier-to-specify sparse reward functions. This sparsity poses new exploration challenges. One common way to address this problem is using demonstrations to provide initial signal about regions of the state space with high rewards. However, prior RL from demonstrations algorithms introduce significant complexity and many hyperparameters, making them hard to implement and tune. We introduce Monte Carlo Actor-Critic (MCAC), a parameter free modification to standard actor-critic algorithms which initializes the replay buffer with demonstrations and computes a modified $Q$-value by taking the maximum of the standard temporal distance (TD) target and a Monte Carlo estimate of the reward-to-go. This encourages exploration in the neighborhood of high-performing trajectories by encouraging high $Q$-values in corresponding regions of the state space. Experiments across $5$ continuous control domains suggest that MCAC can be used to significantly increase learning efficiency across $6$ commonly used RL and RL-from-demonstrations algorithms. See https://sites.google.com/view/mcac-rl for code and supplementary material. Albert Wilcox, Ashwin Balakrishna, Jules Dedieu, Wyame Benslimane, Daniel S. Brown, Kenneth Y. Goldberg |
NeurIPS | 5 |
| 2021 | Value Alignment VerificationabstractAs humans interact with autonomous agents to perform increasingly complicated, potentially risky tasks, it is important to be able to efficiently evaluate an agent’s performance and correctness. In this paper we formalize and theoretically analyze the problem of efficient value alignment verification: how to efficiently test whether the behavior of another agent is aligned with a human’s values? The goal is to construct a kind of "driver’s test" that a human can give to any agent which will verify value alignment via a minimal number of queries. We study alignment verification problems with both idealized humans that have an explicit reward function as well as problems where they have implicit values. We analyze verification of exact value alignment for rational agents, propose and test heuristics for value alignment verification in gridworlds and a continuous autonomous driving domain, and prove that there exist sufficient conditions such that we can verify epsilon-alignment in any environment via a constant-query-complexity alignment test. Daniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott Niekum |
ICML | 1 |
| 2021 | Policy Gradient Bayesian Robust Optimization for Imitation LearningabstractThe difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are often many different reward functions that explain the human feedback, leaving agents with uncertainty over what the true reward function is. While most policy optimization approaches handle this uncertainty by optimizing for expected performance, many applications demand risk-averse behavior. We derive a novel policy gradient-style robust optimization approach, PG-BROIL, that optimizes a soft-robust objective that balances expected performance and risk. To the best of our knowledge, PG-BROIL is the first policy optimization algorithm robust to a distribution of reward hypotheses which can scale to continuous MDPs. Results suggest that PG-BROIL can produce a family of behaviors ranging from risk-neutral to risk-averse and outperforms state-of-the-art imitation learning algorithms when learning from ambiguous demonstrations by hedging against uncertainty, rather than seeking to uniquely identify the demonstrator’s reward function. Zaynah Javed, Daniel S. Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca D. Dragan, Kenneth Y. Goldberg |
ICML | 2 |
| 2021 | Dynamically Switching Human Prediction Models for Efficient PlanningabstractAs environments involving both robots and humans become increasingly common, so does the need to account for people during planning. To plan effectively, robots must be able to respond to and sometimes influence what humans do. This requires a human model which predicts future human actions. A simple model may assume the human will continue what they did previously; a more complex one might predict that the human will act optimally, disregarding the robot; whereas an even more complex one might capture the robot’s ability to influence the human. These models make different trade-offs between computational time and performance of the resulting robot plan. Using only one model of the human either wastes computational resources or is unable to handle critical situations. In this work, we give the robot access to a suite of human models and enable it to assess the performance-computation trade-off online. By estimating how an alternate model could improve human prediction and how that may translate to performance gain, the robot can dynamically switch human models whenever the additional computation is justified. Our experiments in a driving simulator showcase how the robot can achieve performance comparable to always using the best human model, but with greatly reduced computation. Arjun Sripathy, Andreea Bobu, Daniel S. Brown, Anca D. Dragan |
ICRA | 3 |
| 2021 | Situational Confidence Assistance for Lifelong Shared AutonomyabstractShared autonomy enables robots to infer user intent and assist in accomplishing it. But when the user wants to do a new task that the robot does not know about, shared autonomy will hinder their performance by attempting to assist them with something that is not their intent. Our key idea is that the robot can detect when its repertoire of intents is insufficient to explain the user's input, and give them back control. This then enables the robot to observe unhindered task execution, learn the new intent behind it, and add it to this repertoire. We demonstrate with both a case study and a user study that our proposed method maintains good performance when the human's intent is in the robot's repertoire, outperforms prior shared autonomy approaches when it isn't, and successfully learns new skills, enabling efficient lifelong learning for confidence-based shared autonomy. Matthew Zurek, Andreea Bobu, Daniel S. Brown, Anca D. Dragan |
ICRA | 3 |
| 2020 | Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesabstractBayesian reward learning from demonstrations enables rigorous safety and uncertainty analysis when performing imitation learning. However, Bayesian reward learning methods are typically computationally intractable for complex control problems. We propose Bayesian Reward Extrapolation (Bayesian REX), a highly efficient Bayesian reward learning algorithm that scales to high-dimensional imitation learning problems by pre-training a low-dimensional feature encoding via self-supervised tasks and then leveraging preferences over demonstrations to perform fast Bayesian inference. Bayesian REX can learn to play Atari games from demonstrations, without access to the game score and can generate 100,000 samples from the posterior over reward functions in only 5 minutes on a personal laptop. Bayesian REX also results in imitation learning performance that is competitive with or better than state-of-the-art methods that only learn point estimates of the reward function. Finally, Bayesian REX enables efficient high-confidence policy evaluation without having access to samples of the reward function. These high-confidence performance bounds can be used to rank the performance and risk of a variety of evaluation policies and provide a way to detect reward hacking behaviors. Daniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott Niekum |
ICML | 1 |
| 2020 | Bayesian Robust Optimization for Imitation LearningabstractOne of the main challenges in imitation learning is determining what action an agent should take when outside the state distribution of the demonstrations. Inverse reinforcement learning (IRL) can enable generalization to new states by learning a parameterized reward function, but these approaches still face uncertainty over the true reward function and corresponding optimal policy. Existing safe imitation learning approaches based on IRL deal with this uncertainty using a maxmin framework that optimizes a policy under the assumption of an adversarial reward function, whereas risk-neutral IRL approaches either optimize a policy for the mean or MAP reward function. While completely ignoring risk can lead to overly aggressive and unsafe policies, optimizing in a fully adversarial sense is also problematic as it can lead to overly conservative policies that perform poorly in practice. To provide a bridge between these two extremes, we propose Bayesian Robust Optimization for Imitation Learning (BROIL). BROIL leverages Bayesian reward function inference and a user specific risk tolerance to efficiently optimize a robust policy that balances expected return and conditional value at risk. Our empirical results show that BROIL provides a natural way to interpolate between return-maximizing and risk-minimizing behaviors and outperforms existing risk-sensitive and risk-neutral inverse reinforcement learning algorithms. Daniel S. Brown, Scott Niekum, Marek Petrik |
NeurIPS | 1 |
| 2019 | Machine Teaching for Inverse Reinforcement Learning: Algorithms and ApplicationsabstractInverse reinforcement learning (IRL) infers a reward function from demonstrations, allowing for policy improvement and generalization. However, despite much recent interest in IRL, little work has been done to understand the minimum set of demonstrations needed to teach a specific sequential decisionmaking task. We formalize the problem of finding maximally informative demonstrations for IRL as a machine teaching problem where the goal is to find the minimum number of demonstrations needed to specify the reward equivalence class of the demonstrator. We extend previous work on algorithmic teaching for sequential decision-making tasks by showing a reduction to the set cover problem which enables an efficient approximation algorithm for determining the set of maximallyinformative demonstrations. We apply our proposed machine teaching algorithm to two novel applications: providing a lower bound on the number of queries needed to learn a policy using active IRL and developing a novel IRL algorithm that can learn more efficiently from informative demonstrations than a standard IRL approach. Daniel S. Brown, Scott Niekum |
AAAI | 1 |
| 2019 | Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from ObservationsabstractA critical flaw of existing inverse reinforcement learning (IRL) methods is their inability to significantly outperform the demonstrator. This is because IRL typically seeks a reward function that makes the demonstrator appear near-optimal, rather than inferring the underlying intentions of the demonstrator that may have been poorly executed in practice. In this paper, we introduce a novel reward-learning-from-observation algorithm, Trajectory-ranked Reward EXtrapolation (T-REX), that extrapolates beyond a set of (approximately) ranked demonstrations in order to infer high-quality reward functions from a set of potentially poor demonstrations. When combined with deep reinforcement learning, T-REX outperforms state-of-the-art imitation learning and IRL methods on multiple Atari and MuJoCo benchmark tasks and achieves performance that is often more than twice the performance of the best demonstration. We also demonstrate that T-REX is robust to ranking noise and can accurately extrapolate intention by simply watching a learner noisily improve at a task over time. Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, Scott Niekum |
ICML | 1 |
| 2018 | Efficient Probabilistic Performance Bounds for Inverse Reinforcement LearningabstractIn the field of reinforcement learning there has been recent progress towards safety and high-confidence bounds on policy performance. However, to our knowledge, no practical methods exist for determining high-confidence policy performance bounds in the inverse reinforcement learning setting---where the true reward function is unknown and only samples of expert behavior are given. We propose a sampling method based on Bayesian inverse reinforcement learning that uses demonstrations to determine practical high-confidence upper bounds on the alpha-worst-case difference in expected return between any evaluation policy and the optimal policy under the expert's unknown reward function. We evaluate our proposed bound on both a standard grid navigation task and a simulated driving task and achieve tighter and more accurate bounds than a feature count-based baseline. We also give examples of how our proposed bound can be utilized to perform risk-aware policy selection and risk-aware policy improvement. Because our proposed bound requires several orders of magnitude fewer demonstrations than existing high-confidence bounds, it is the first practical method that allows agents that learn from demonstration to express confidence in the quality of their learned policy. Daniel S. Brown, Scott Niekum |
AAAI | 1 |
| 2017 | Exact and Heuristic Algorithms for Risk-Aware Stochastic Physical SearchabstractWe consider an intelligent agent seeking to obtain an item from one of several physical locations, where the cost to obtain the item at each location is stochastic. We study risk‐aware stochastic physical search (RA‐SPS), where both the cost to travel and the cost to obtain the item are taken from the same budget and where the objective is to maximize the probability of success while minimizing the required budget. This type of problem models many task‐planning scenarios, such as space exploration, shopping, or surveillance. In these types of scenarios, the actual cost of completing an objective at a location may only be revealed when an agent physically arrives at the location, and the agent may need to use a single resource to both search for and acquire the item of interest. We present exact and heuristic algorithms for solving RA‐SPS problems on complete metric graphs. We first formulate the problem as mixed integer linear programming problem. We then develop custom branch and bound algorithms that result in a dramatic reduction in computation time. Using these algorithms, we generate empirical insights into the hardness landscape of the RA‐SPS problem and compare the performance of several heuristics. Daniel S. Brown, Jeffrey Hudack, Nathaniel Gemelli, Bikramjit Banerjee |
Comput. Intell. | 1 |
| 2016 | Classifying swarm behavior via compressive subspace learningabstractBio-inspired robot swarms encompass a rich space of dynamics and collective behaviors.Given some agent measurements of a swarm at a particular time instance, an important problem is the classification of the swarm behavior.This is challenging in practical scenarios where information from only a small number of agents may be available, resulting in limited agent samples for classification.Another challenge is recognizing emerging behavior: the prediction of swarm behavior prior to convergence of the attracting state.In this paper we address these challenges by modeling a swarm's collective motion as a low-dimensional linear subspace.We illustrate that for both synthetic and real data, these behaviors manifest as low-dimensional subspaces, and that these subspaces are highly discriminative.We also show that these subspaces generalize well to predicting emerging behavior, highlighting that there exists low-dimensional structure in transient agent behavior.In order to learn distinct behavior subspaces, we extend previous work on subspace estimation and identification from missing data to that of compressive measurements, where compressive measurements arise due to agent positions scattered throughout the domain.We demonstrate improvement in performance over prior works with respect to limited agent samples over a wide range of agent models and scenarios. Matthew Berger, Lee M. Seversky, Daniel S. Brown |
ICRA | 3 |
| 2016 | Two invariants of human-swarm interactionabstractThe search for invariants is a fundamental aim of scientific endeavors. These invariants, such as Newton's laws of motion, allow us to model and predict the behavior of systems across many different problems. In the nascent field of Human-Swarm Interaction (HSI), a systematic identification of fundamental invariants is still lacking. Discovering and formalizing these invariants will provide a foundation for developing, and better understanding, effective methods for HSI. We propose two invariants underlying HSI for geometric-based swarms: (1) collective state is the fundamental percept associated with a bio-inspired swarm, and (2) a human's ability to influence and understand the collective state of a swarm is determined by the balance between the span and persistence. We provide evidence of these invariants by synthesizing much of our previous work in the area of HSI with several new results, including a novel user study where users manage multiple swarms simultaneously. We also discuss how these invariants can be applied to enable more efficient and successful teaming between humans and bio-inspired collectives and identify several promising directions for future research into the invariants of HSI. Daniel S. Brown, Michael A. Goodrich, Shin-Young Jung, Sean Kerman |
J. Hum. Robot Interact. | 1 |
| 2015 | Multiobjective Optimization for the Stochastic Physical Search Problem
Jeffrey Hudack, Nathaniel Gemelli, Daniel S. Brown, Steven Loscalzo, Jae C. Oh |
IEA/AIE | 3 |
| 2014 | Human-swarm interactions based on managing attractorsabstractLeveraging the abilities of multiple affordable robots as a swarm is enticing because of the resulting robustness and emergent behaviors of a swarm. However, because swarms are composed of many different agents, it is difficult for a human to influence the swarm by managing individual agents. Instead, we propose that human influence should focus on (a)~managing the higher level attractors of the swarm system and (b)~managing trade-offs that appear in mission-relevant performance. We claim that managing attractors theoretically allows a human to abstract the details of individual agents and focus on managing the collective as a whole. Using a swarm model with two attractors, we demonstrate this concept by showing how limited human influence can cause the swarm to switch between attractors. We further claim that using quorum sensing allows a human to manage trade-offs between the scalability of interactions and mitigating the vulnerability of the swarm to agent failures. Daniel S. Brown, Sean Kerman, Michael A. Goodrich |
HRI | 1 |
| 2014 | Balancing human and inter-agent influences for shared control of bio-inspired collectivesabstractHuman interaction with bio-inspired collectives provides an interesting setting for studying shared control. A human will often have knowledge of global objectives and high-level plans, but the collective will often have more detailed lower-level knowledge about the particulars of the situation at hand. Thus it is important to understand how control can be appropriately shared between the human and the collective. We analyze human interaction with bio-inspired collectives using graph theory, and propose that there are two human-side elements that determine how well control is shared: span and persistence. We additionally propose that there is a collective-side element that determines how well control is shared: connectivity. We study two examples of shared-control between a human and a bio-inspired collective: shaping a spatial formation and causing a collective to switch between stable collective states. Our empirical results show that span, persistence, and connectivity combine to affect (1) how influence is shared between the human and the collective and (2) the resulting success of human-collective interactions. Daniel S. Brown, Shin-Young Jun, Michael A. Goodrich |
SMC | 1 |
| 2013 | Shaping Couzin-Like Torus Swarms through Coordinated MediationabstractHuman-swarm interaction methods often allow a human to influence a swarm through either leadership or predation. These methods of influence have two main limitations: (1) although leaders sustain influence over nominal agents for a long period of time, they tend to cause all collective structures to turn in to flocks (negating the benefit of other swarm formations) and (2) predators tend to cause collective structures to fragment. We introduce the use of mediators as a novel shared control method for human-swarm influence and use mediators to shape Couzin-like tori [1]. The mediator method uses special agents that operate from within the spatial center of a swarm. This approach allows a human operator to transform and move a dynamic torus formation while sustaining influence over the torus, avoiding fragmentation, and maintaining the torus' connectivity. The use of mediators allows a human to mold and adapt the torus' behavior and structure to a wide range of spatio-temporal tasks such as military protection and decontamination tasks. Shin-Young Jun, Daniel S. Brown, Michael A. Goodrich |
SMC | 2 |