EDBT 2026 Demo / reviewers in the wild / expert
Peter Stone 0001
dblp:s/PeterStone
· DBLP profile ↗
338ranked-venue papers
26as first author
92since 2021 · last 2026
0000-0002-6795-420XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 324 · 24 first-author · 90 since 2021Graphics, computer vision, multimedia, augmented reality and games · 89 · 3 first-author · 18 since 2021Systems, architecture and hardware · 62 · 32 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single PolicyabstractGeneralization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act within environments with varying contexts, such as self-driving cars or quadrupedal robots that need to operate in different terrains or weather conditions than they were trained for. We tackle the critical task of generalizing to out-of-distribution (OOD) settings, without access to explicit context information at test time. Recent work has addressed this problem by training a context encoder and a history adaptation module in separate stages. While promising, this two-phase approach is cumbersome to implement and train. We simplify the methodology and introduce SPARC: single-phase adaptation for robust control. We test SPARC on varying contexts within the high-fidelity racing simulator Gran Turismo 7 and wind-perturbed MuJoCo environments, and find that it achieves reliable and robust OOD generalization. Bram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone 0001, Peter R. Wurman |
AAAI | 4 |
| 2026 | The Essentials of AI for Life and Society: A Full-Scale AI Literacy Course Accessible to AllabstractIn Fall 2023, we introduced a new AI Literacy class called The Essentials of AI for Life and Society (CS 109), a one-credit, seminar course consisting mainly of guest lectures, which was open to the entire university, including students, staff, and faculty. Building on its success and popularity, this paper describes our significant expansion of the course into a full-scale three-credit undergraduate course (CS 309), with an expanded emphasis on student engagement, interactivity, and ethics-related components. To knit together content from the guest lecturers, we implemented a flipped classroom. This model used weekly asynchronous learning modules---integrating pre-recorded expert lectures, collaborative readings, and ethical reflections---which were then unified by the course instructor during a live, interactive discussion session. To maintain the broad accessibility of the material (no prerequisites), the course introduced substantive, non-programming homework assignments in which students applied AI concepts to grounded, real-world problems. This work culminated in a final project analyzing the ethical and societal implications of a chosen AI tool. The redesigned course received overwhelmingly positive student feedback, highlighting its interactivity, coherence, and accessible and engaging assignments. This paper details the course's evolution, its pedagogical structure, and the lessons learned in developing a core AI literacy course. All course materials are freely available for others to use and build upon. Zifan Xu, Kristen Procko, Michael J. Munje, Kristin Patterson, Lea Sabatini, Joydeep Biswas, Peter Stone 0001 |
AAAI | 7 |
| 2025 | The Essentials of AI for Life and Society: An AI Literacy Course for the University CommunityabstractWe describe the development of a one-credit course to promote AI literacy at the University of Texas at Austin. In response to a call for the rapid deployment of class that would serve a broad audience in Fall of 2023, we designed a 14-week seminar-style course that incorporated an interdisciplinary group of speakers who lectured on topics ranging from the fundamentals of AI to societal concerns including disinformation and employment. University students, faculty, and staff, and even community members outside of the University were invited to enroll in this online offering: The Essentials of AI for Life and Society. We collected feedback from course participants through weekly reflections and a final survey. Satisfyingly, we found that attendees reported gains in their AI literacy. We sought critical feedback through quantitative and qualitative analysis, which uncovered challenges in designing a course for this general audience. We utilized the course feedback to design a three-credit version of the course that is being offered in Fall of 2024. The lessons we learned and our plans for this new iteration may serve as a guide to instructors designing AI courses for a broad audience. Joydeep Biswas, Donald S. Fussell, Peter Stone 0001, Kristin Patterson, Kristen Procko, Lea Sabatini, Zifan Xu |
AAAI | 3 |
| 2025 | Deep Reinforcement Learning for Robotics: A Survey of Real-World SuccessesabstractReinforcement learning (RL), particularly its combination with deep neural networks referred to as deep RL (DRL), has shown tremendous promise across a wide range of applications, suggesting its potential for enabling the development of sophisticated robotic behaviors. Robotics problems, however, pose fundamental difficulties for the application of RL, stemming from the complexity and cost of interacting with the physical world. These challenges notwithstanding, recent advances have enabled DRL to succeed at some real-world robotic tasks. However, state-of-the-art DRL solutions’ maturity varies significantly across robotic applications. In this talk, I will review the current progress of DRL in real-world robotic applications based on our recent survey paper (with Tang, Abbatematteo, Hu, Chandra, and Martı́n-Martı́n), with a particular focus on evaluating the real-world successes achieved with DRL in realizing several key robotic competencies, including locomotion, navigation, stationary manipulation, mobile manipulation, human-robot interaction, and multi-robot interaction. The analysis aims to identify the key factors underlying those exciting successes, reveal underexplored areas, and provide an overall characterization of the status of DRL in robotics. I will also highlight several important avenues for future work, emphasizing the need for stable and sample-efficient real-world RL paradigms, holistic approaches for discovering and integrating various competencies to tackle complex long-horizon, open-world tasks, and principled development and evaluation procedures. The talk is designed to offer insights for RL practitioners and roboticists toward harnessing RL’s power to create generally capable real-world robotic systems. Chen Tang 0001, Ben Abbatematteo, Jiaheng Hu, Rohan Chandra, Roberto Martin Martin, Peter Stone 0001 |
AAAI | 6 |
| 2025 | Argus: A Compact and Versatile Foundation Model for VisionabstractWhile existing vision and multi-modal foundation models can handle multiple computer vision tasks, they often suffer from significant limitations, including huge demand for data and computational resources during training and inconsistent performance across vision tasks at deployment time. To address these challenges, we introduce Argus1, a compact and versatile vision foundation model designed to support a wide range of vision tasks through a unified multitask architecture. Argus employs a two-stage training strategy: (i) multitask pretraining over core vision tasks with a shared backbone that includes a lightweight adapter to inject task-specific inductive biases, and (ii) scalable and efficient adaptation to new tasks by fine-tuning only the task-specific decoders. Extensive evaluations demonstrate that Argus, despite its relatively compact and training-efficient design of merely 100M backbone parameters (only 13.6% of which are trained using 1.6M images), competes with and even surpasses much larger models. Compared to state-of-the-art foundation models, Argus not only covers a broader set of vision tasks but also matches or outperforms the models with similar sizes on 12 tasks. We expect that Argus will accelerate the real-world adoption of vision foundation models in resource-constrained scenarios. Weiming Zhuang, Chen Chen 0043, Sina Sajadmanesh, Jiabo Huang, Vikash Sehwag, Vivek Sharma 0001, Hirotaka Shinozaki, Felan Carlo Garcia, Yihao Zhan, Naohiro Adachi, Ryoji Eki, Michael Spranger, Peter Stone 0001, Lingjuan Lyu |
CVPR | 15 |
| 2025 | Learning a Fast Mixing Exogenous Block MDP using a Single TrajectoryabstractIn order to train agents that can quickly adapt to new objectives or reward functions, efficient unsupervised representation learning in sequential decision-making environments can be important. Frameworks such as the Exogenous Block Markov Decision Process (Ex-BMDP) have been proposed to formalize this representation-learning problem (Efroni et al., 2022b). In the Ex-BMDP framework, the agent's high-dimensional observations of the environment have two latent factors: a controllable factor, which evolves deterministically within a small state space according to the agent's actions, and an exogenous factor, which represents time-correlated noise, and can be highly complex. The goal of the representation learning problem is to learn an encoder that maps from observations into the controllable latent space, as well as the dynamics of this space. Efroni et al. (2022b) has shown that this is possible with a sample complexity that depends only on the size of the controllable latent space, and not on the size of the noise factor. However, this prior work has focused on the episodic setting, where the controllable latent state resets to a specific start state after a finite horizon.
By contrast, if the agent can only interact with the environment in a single continuous trajectory, prior works have not established sample-complexity bounds. We propose STEEL, the first provably sample-efficient algorithm for learning the controllable dynamics of an Ex-BMDP from a single trajectory, in the function approximation setting. STEEL has a sample complexity that depends only on the sizes of the controllable latent space and the encoder function class, and (at worst linearly) on the mixing time of the exogenous noise factor. We prove that STEEL is correct and sample-efficient, and demonstrate STEEL on two toy problems. Code is available at: https://github.com/midi-lab/steel. Alexander Levine 0001, Peter Stone 0001, Amy Zhang 0001 |
ICLR | 2 |
| 2025 | Longhorn: State Space Models are Amortized Online LearnersabstractThe most fundamental capability of modern AI methods such as Large Language Models (LLMs) is the ability to predict the next token in a long sequence of tokens, known as “sequence modeling.” Although the Transformers model is the current dominant approach to sequence modeling, its quadratic computational cost with respect to sequence length is a significant drawback. State-space models (SSMs) offer a promising alternative due to their linear decoding efficiency and high parallelizability during training. However, existing SSMs often rely on seemingly ad hoc linear recurrence designs.
In this work, we explore SSM design through the lens of online learning, conceptualizing SSMs as meta-modules for specific online learning problems. This approach links SSM design to formulating precise online learning objectives, with state transition rules derived from optimizing these objectives.
Based on this insight, we introduce a novel deep SSM architecture based on the implicit update for optimizing an online regression objective. Our experimental results show that our models outperform state-of-the-art SSMs, including the Mamba model, on standard sequence modeling benchmarks and language modeling tasks. Bo Liu 0042, Lemeng Wu, Yihao Feng, Peter Stone 0001, Qiang Liu 0001 |
ICLR | 5 |
| 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningabstractRecent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to overfitting.
These large networks avoid overfitting by integrating components that induce a simplicity bias, guiding models toward simple and generalizable solutions.
However, in deep RL, designing and scaling up networks have been less explored.
Motivated by this opportunity, we present SimBa, an architecture designed to scale up parameters in deep RL by injecting a simplicity bias. SimBa consists of three components: (i) an observation normalization layer that standardizes inputs with running statistics, (ii) a residual feedforward block to provide a linear pathway from the input to output, and (iii) a layer normalization to control feature magnitudes.
By scaling up parameters with SimBa, the sample efficiency of various deep RL algorithms—including off-policy, on-policy, and unsupervised methods—is consistently improved.
Moreover, solely by integrating SimBa architecture into SAC, it matches or surpasses state-of-the-art deep RL methods with high computational efficiency across DMC, MyoSuite, and HumanoidBench.
These results demonstrate SimBa's broad applicability and effectiveness across diverse RL algorithms and environments. Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R. Wurman, Jaegul Choo, Peter Stone 0001, Takuma Seno |
ICLR | 9 |
| 2025 | Proto Successor Measure: Representing the Behavior Space of an RL AgentabstractHaving explored an environment, intelligent agents should be able to transfer their knowledge to most downstream tasks within that environment without additional interactions. Referred to as "zero-shot learning", this ability remains elusive for general-purpose reinforcement learning algorithms. While recent works have attempted to produce zero-shot RL agents, they make assumptions about the nature of the tasks or the structure of the MDP. We present Proto Successor Measure: the basis set for all possible behaviors of a Reinforcement Learning Agent in a dynamical system. We prove that any possible behavior (represented using visitation distributions) can be represented using an affine combination of these policy-independent basis functions. Given a reward function at test time, we simply need to find the right set of linear weights to combine these bases corresponding to the optimal policy. We derive a practical algorithm to learn these basis functions using reward-free interaction data from the environment and show that our approach can produce the near-optimal policy at test time for any given reward function without additional environmental interactions. Project page: agarwalsiddhant10.github.io/projects/psm.html. Siddhant Agarwal, Harshit Sikchi, Peter Stone 0001, Amy Zhang 0001 |
ICML | 3 |
| 2025 | Hyperspherical Normalization for Scalable Deep Reinforcement LearningabstractScaling up the model size and computation has brought consistent performance improvements in supervised learning. However, this lesson often fails to apply to reinforcement learning (RL) because training the model on non-stationary data easily leads to overfitting and unstable optimization. In response, we introduce SimbaV2, a novel RL architecture designed to stabilize optimization by (i) constraining the growth of weight and feature norm by hyperspherical normalization; and (ii) using a distributional value estimation with reward scaling to maintain stable gradients under varying reward magnitudes. Using the soft actor-critic as a base algorithm, SimbaV2 scales up effectively with larger models and greater compute, achieving state-of-the-art performance on 57 continuous control tasks across 4 domains. Youngdo Lee, Takuma Seno, Donghu Kim, Peter Stone 0001, Jaegul Choo |
ICML | 5 |
| 2025 | FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-TuningabstractIn recent years, the Robotics field has initiated several efforts toward building generalist robot policies through large-scale multi-task Behavior Cloning. However, direct deployments of these policies have led to unsatisfactory performance, where the policy struggles with unseen states and tasks. How can we break through the performance plateau of these models and elevate their capabilities to new heights? In this paper, we propose FLaRe, a large-scale Reinforcement Learning fine-tuning framework that integrates robust pre-trained representations, large-scale training, and gradient stabilization techniques. Our method aligns pre-trained policies towards task completion, achieving state-of-the-art (SoTA) performance both on previously demonstrated and on entirely novel tasks and embodiments. Specifically, on a set of long-horizon mobile manipulation tasks, FLaRe achieves an average success rate of 79.5% in unseen environments, with absolute improvements of$+23.6 \%$in simulation and$+30.7 \%$on real robots over prior SoTA methods. By utilizing only sparse rewards, our approach can enable generalizing to new capabilities beyond the pretraining data with minimal human effort. Moreover, we demonstrate rapid adaptation to new embodiments and behaviors with less than a day of fine-tuning. Videos, code, and appendix can be found on the project website at robot-flare.github.io Jiaheng Hu, Rose Hendrix, Ali Farhadi, Aniruddha Kembhavi, Roberto Martin Martin, Peter Stone 0001, Kuo-Hao Zeng, Kiana Ehsani |
ICRA | 6 |
| 2025 | Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot SoccerabstractRobot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) is a promising approach to learning decisionmaking in such domains, however, end-to-end RL in complex environments is often intractable. To address this challenge in the RoboCup Standard Platform League (SPL) domain, we developed a novel architecture integrating RL within a classical robotics stack, while employing a multi-fidelity sim2real approach and decomposing behavior into learned sub-behaviors with heuristic selection. Our architecture led to victory in the 2024 RoboCup SPL Challenge Shield Division. In this work, we fully describe our system's architecture and empirically analyze key design decisions that contributed to its success. Our approach demonstrates how RL-based behaviors can be integrated into complete robot behavior architectures. Adam Labiosa, Zhihan Wang, Siddhant Agarwal, William Cong, Geethika Hemkumar, Abhinav Narayan Harish, Benjamin Hong, Josh Kelle, Zisen Shao, Peter Stone 0001, Josiah Hanna |
ICRA | 12 |
| 2025 | PRESTO: Fast Motion Planning Using Diffusion Models Based on Key-Configuration Environment RepresentationabstractWe introduce a learning-guided motion planning framework that generates seed trajectories using a diffusion model for trajectory optimization. Given a workspace, our method approximates the configuration space (C-space) obstacles through an environment representation consisting of a sparse set of task-related key configurations, which is then used as a conditioning input to the diffusion model. The diffusion model integrates regularization terms that encourage smooth, collision-free trajectories during training, and trajectory optimization refines the generated seed trajectories to correct any colliding segments. Our experimental results demonstrate that high-quality trajectory priors, learned through our C-space-grounded diffusion model, enable the efficient generation of collision-free trajectories in narrow-passage environments, outperforming previous learning- and planning-based baselines. Videos and additional materials can be found on the project page: https://kiwi-sherbet.github.io/PRESTO. Mingyo Seo, Yoonyoung Cho, Yoonchang Sung, Peter Stone 0001, Yuke Zhu |
ICRA | 4 |
| 2025 | L3M+P: Lifelong Planning with Large Language ModelsabstractBy combining classical planning methods with large language models (LLMs), recent research such as LLM+P has enabled agents to plan for general tasks given in natural language. However, scaling these methods to general-purpose service robots remains challenging: (1) classical planning algorithms generally require a detailed and consistent specification of the environment, which is not always readily available; and (2) existing frameworks mainly focus on isolated planning tasks, whereas robots are often meant to serve in long-term continuous deployments, and therefore must maintain a dynamic memory of the environment which can be updated with multi-modal inputs and extracted as planning knowledge for future tasks. To address these two issues, this paper introduces L3M+P (Lifelong LLM+P), a framework that uses an external knowledge graph as a representation of the world state. The graph can be updated from multiple sources of information, including sensory input and natural language interactions with humans. L3M+P enforces rules for the expected format of the absolute world state graph to maintain consistency between graph updates. At planning time, given a natural language description of a task, L3M+P retrieves context from the knowledge graph and generates a problem definition for classical planners. Evaluated on household robot simulators and on a real-world service robot, L3M+P achieves significant improvement over baseline methods both on accurately registering natural language state changes and on correctly generating plans, thanks to the knowledge graph retrieval and verification. Krish Agarwal, Yuqian Jiang, Jiaheng Hu, Bo Liu 0042, Peter Stone 0001 |
IROS | 5 |
| 2025 | Multi-Agent Inverse Reinforcement Learning in Real World Unstructured Pedestrian CrowdsabstractSocial robot navigation in crowded public spaces such as university campuses, restaurants, grocery stores, and hospitals, is an increasingly important area of research. One of the core strategies for achieving this goal is to understand humans’ intent–underlying psychological factors that govern their motion–by learning how humans assign rewards to their actions, typically via inverse reinforcement learning (IRL). Despite significant progress in IRL, learning reward functions of multiple agents simultaneously in dense unstructured pedestrian crowds has remained intractable due to the nature of the tightly coupled social interactions that occur in these scenarios e.g. passing, intersections, swerving, weaving, etc. In this paper, we present a new multi-agent maximum entropy inverse reinforcement learning algorithm for real world unstructured pedestrian crowds. Key to our approach is a simple, but effective, mathematical trick which we name the so-called "tractability-rationality trade-off" trick that achieves tractability at the cost of a slight reduction in accuracy. We compare our approach to the classical single-agent MaxEnt IRL as well as state-of-the-art trajectory prediction methods on several datasets including the ETH, UCY, SCAND, JRDB, and a new dataset, called Speedway, collected at a busy intersection on a University campus focusing on dense, complex agent interactions. Our key findings show that, on the dense Speedway dataset, our approach ranks 1stamong top 7 baselines with > 2× improvement over single-agent IRL, and is competitive with state-of-the-art large transformer-based encoder-decoder models on sparser datasets such as ETH/UCY (ranks 3rdamong top 7 baselines). Rohan Chandra, Haresh Karnan, Negar Mehr, Peter Stone 0001, Joydeep Biswas |
IROS | 4 |
| 2025 | Dyna-LfLH: Learning Agile Navigation in Dynamic Environments from Learned HallucinationabstractThis paper introduces Dynamic Learning from Learned Hallucination (Dyna-LfLH), a self-supervised method for training motion planners to navigate environments with dense and dynamic obstacles. Classical planners struggle with dense, unpredictable obstacles due to limited computation, while learning-based planners face challenges in acquiring high-quality demonstrations for imitation learning or dealing with exploration inefficiencies in reinforcement learning. Building on Learning from Hallucination (LfH), which synthesizes training data from past successful navigation experiences in simpler environments, Dyna-LfLH incorporates dynamic obstacles by generating them through a learned latent distribution. This enables efficient and safe motion planner training. We evaluate Dyna-LfLH on a ground robot in both simulated and real environments, achieving up to a 25% improvement in success rate compared to baselines. Saad Abdul Ghani, Peter Stone 0001, Xuesu Xiao |
IROS | 3 |
| 2025 | GACL: Grounded Adaptive Curriculum Learning with Active Task and Performance MonitoringabstractCurriculum learning has emerged as a promising approach for training complex robotics tasks, yet current applications predominantly rely on manually designed curricula, which demand significant engineering effort and can suffer from subjective and suboptimal human design choices. While automated curriculum learning has shown success in simple domains like grid worlds and games where task distributions can be easily specified, robotics tasks present unique challenges: they require handling complex task spaces while maintaining relevance to target domain distributions that are only partially known through limited samples. To this end, we propose Grounded Adaptive Curriculum Learning (GACL1), a framework specifically designed for robotics curriculum learning with three key innovations: (1) a task representation that consistently handles complex robot task design, (2) an active performance tracking mechanism that allows adaptive curriculum generation appropriate for the robot’s current capabilities, and (3) a grounding approach that maintains target domain relevance through alternating sampling between reference and synthetic tasks. We validate GACL on wheeled navigation in constrained environments and quadruped locomotion in challenging 3D confined spaces, achieving 6.8% and 6.1% higher success rates, respectively, than state-of-the-art methods in each domain. Linji Wang, Zifan Xu, Peter Stone 0001, Xuesu Xiao |
IROS | 3 |
| 2025 | RLZero: Direct Policy Inference from Language Without In-Domain SupervisionabstractThe reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward signal is notoriously difficult, as humans are often unable to predict the optimal behavior corresponding to a reward function. Natural language offers an intuitive alternative for instructing reinforcement learning (RL) agents, yet previous language-conditioned approaches either require costly supervision or test-time training given a language instruction. In this work, we present a new approach that uses a pretrained RL agent trained using only unlabeled, offline interactions—without task-specific supervision or labeled trajectories—to get zero-shot test-time policy inference from arbitrary natural language instructions. We introduce a framework comprising three steps: *imagine*, *project*, and *imitate*. First, the agent imagines a sequence of observations corresponding to the provided language description using video generative models. Next, these imagined observations are projected into the target environment domain. Finally, an agent pretrained in the target environment with unsupervised RL instantly imitates the projected observation sequence through a closed-form solution. To the best of our knowledge, our method, RLZero, is the first approach to show direct language-to-behavior generation abilities on a variety of tasks and environments without any in-domain supervision. We further show that components of RLZero can be used to generate policies zero-shot from cross-embodied videos, such as those available on YouTube, even for complex embodiments like humanoids. Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli, Caleb Chuck, Max Rudolph, Peter Stone 0001, Amy Zhang 0001, Scott Niekum |
NeurIPS | 7 |
| 2025 | Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using ConcordiaabstractLarge Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In this paper, we introduce a method for evaluating the ability of LLM-based agents to cooperate in zero-shot, mixed-motive environments using Concordia, a natural language multi-agent simulation environment. Our method measures general cooperative intelligence by testing an agent's ability to identify and exploit opportunities for mutual gain across diverse partners and contexts. We present empirical results from the NeurIPS 2024 Concordia Contest, where agents were evaluated on their ability to achieve mutual gains across a suite of diverse scenarios ranging from negotiation to collective action problems. Our findings reveal significant gaps between current agent capabilities and the robust generalization required for reliable cooperation, particularly in scenarios demanding persuasion and norm enforcement. Chandler Smith, Marwa Abdulhai, Manfred Diaz, Marko Tesic, Rakshit S. Trivedi, Alexander Vezhnevets, Lewis Hammond, Jesse Clifton, Minsuk Chang, Edgar A. Duéñez-Guzmán, John P. Agapiou, Jayd Matyas, Danny Karmon, Beining Zhang, Jim Dilkes, Akash Kundu, Emanuel Tewolde, Jebish Purbey, Ram Mohan Rao Kadiyala, Siddhant Gupta, Aliaksei Korshuk, Buyantuev Alexander, Ilya Makarov, Rolando Fernandez, Zhihan Wang, Caroline Wang, Jiaxun Cui, Lingyun Xiao, Yoonchang Sung, Muhammad Arrasy Rahman, Peter Stone 0001, Yipeng Kang, Hyeonggeun Yun, Ananya, Taehun Cha, Elizaveta Tennant, Olivia Macmillan-Scott, Marta Segura, Diana Riazi, Fuyang Cui, Sriram Ganapathi, Toryn Q. Klassen, Nico Schiavone, Mogtaba Alim, Sheila A. McIlraith, Manuel Ríos, Oswaldo Peña, Manuela Chacon-Chamorro, Rubén Manrique, Luis Felipe Giraldo, Nicanor Quijano, Fangwei Zhong, Wenming Tu, Zhaowei Zhang 0001, Zixia Jia, Zilong Zheng, Chichen Lin, Weijian Fan, Chenao Liu, Sneheel Sarangi, Shuqing Shi, Yali Du 0001, Avinaash Anand Kulandaivel, Yang Liu 0266, Ruiyang Wu 0007, Chetan Talele, Sunjia Lu, Gema Parreno, Shamika Dhuri, Bain McHale, Tim Baarslag, Dylan Hadfield-Menell, Natasha Jaques, José Hernández-Orallo, Joel Z. Leibo |
NeurIPS | 34 |
| 2025 | Dyn-O: Building Structured World Models with Object-Centric RepresentationsabstractWorld models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered on interactions among objects within the environment. This motivates the development of world models that operate on object-centric rather than monolithic representations, with the goal of more effectively capturing environment dynamics and enhancing compositional generalization. However, the development of object-centric world models has largely been explored in environments with limited visual complexity (such as basic geometries). It remains underexplored whether such models can be effective in more challenging settings. In this paper, we fill this gap by introducing Dyn-O, an enhanced structured world model built upon object-centric representations. Compared to prior work in object-centric representations, Dyn-O improves in both learning representations and modeling dynamics. On the challenging Procgen games, we demonstrate that our method can learn object-centric world models directly from pixel observations, outperforming DreamerV3 in rollout prediction accuracy. Furthermore, by decoupling object centric features into dynamic-agnostic and dynamic-aware components, we enable finer-grained manipulation of these features and generate more diverse imagined trajectories. The code of Dyn-O can be found at: https://github.com/wangzizhao/dyn-O. Li Zhao 0007, Peter Stone 0001, Jiang Bian 0002 |
NeurIPS | 4 |
| 2025 | Learning a robust multiagent driving policy for traffic congestion reduction
Yulin Zhang 0001, William Macke, Jiaxun Cui, Sharon Hornstein, Daniel Urieli, Peter Stone 0001 |
Neural Comput. Appl. | 6 |
| 2025 | Principles and Guidelines for Evaluating Social Robot Navigation AlgorithmsabstractA major challenge to deploying robots widely is navigation in human-populated environments, commonly referred to as social robot navigation . While the field of social navigation has advanced tremendously in recent years, the fair evaluation of algorithms that tackle social navigation remains hard because it involves not just robotic agents moving in static environments but also dynamic human agents and their perceptions of the appropriateness of robot behavior. In contrast, clear, repeatable, and accessible benchmarks have accelerated progress in fields like computer vision, natural language processing and traditional robot navigation by enabling researchers to fairly compare algorithms, revealing limitations of existing solutions and illuminating promising new directions. We believe the same approach can benefit social navigation. In this article, we pave the road toward common, widely accessible, and repeatable benchmarking criteria to evaluate social robot navigation. Our contributions include (a) a definition of a socially navigating robot as one that respects the principles of safety, comfort, legibility, politeness, social competency, agent understanding, proactivity, and responsiveness to context, (b) guidelines for the use of metrics, development of scenarios, benchmarks, datasets, and simulators to evaluate social navigation, and (c) a design of a social navigation metrics framework to make it easier to compare results from different simulators, robots, and datasets. Anthony G. Francis, Claudia Pérez-D'Arpino, Chengshu Li 0002, Fei Xia 0002, Alexandre Alahi, Rachid Alami 0001, Aniket Bera, Abhijat Biswas, Joydeep Biswas, Rohan Chandra, Hao-Tien Chiang, Michael Everett, Sehoon Ha, Justin W. Hart, Jonathan P. How, Haresh Karnan, Tsang-Wei Edward Lee, Luis Manso, Reuth Mirsky, Sören Pirk, Phani-Teja Singamaneni, Peter Stone 0001, Ada V. Taylor, Pete Trautman, Nathan Tsoi, Marynel Vázquez, Xuesu Xiao, Peng Xu 0010, Naoki Yokoyama, Alexander Toshev, Roberto Martin Martin |
ACM Trans. Hum. Robot Interact. | 22 |
| 2024 | Reward (Mis)design for Autonomous Driving (Abstract Reprint)abstractThis article considers the problem of diagnosing certain common errors in reward design. Its insights are also applicable to the design of cost functions and performance metrics more generally. To diagnose common errors, we develop 8 simple sanity checks for identifying flaws in reward functions. We survey research that is published in top-tier venues and focuses on reinforcement learning (RL) for autonomous driving (AD). Specifically, we closely examine the reported reward function in each publication and present these reward functions in a complete and standardized format in the appendix. Wherever we have sufficient information, we apply the 8 sanity checks to each surveyed reward function, revealing near-universal flaws in reward design for AD that might also exist pervasively across reward design for other tasks. Lastly, we explore promising directions that may aid the design of reward functions for AD in subsequent research, following a process of inquiry that can be adapted to other domains. W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt 0001, Peter Stone 0001 |
AAAI | 5 |
| 2024 | Learning Optimal Advantage from Preferences and Mistaking It for RewardabstractWe consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most recent work assumes that human preferences are generated based only upon the reward accrued within those segments, or their partial return. Recent work casts doubt on the validity of this assumption, proposing an alternative preference model based upon regret. We investigate the consequences of assuming preferences are based upon partial return when they actually arise from regret. We argue that the learned function is an approximation of the optimal advantage function, not a reward function. We find that if a specific pitfall is addressed, this incorrect assumption is not particularly harmful, resulting in a highly shaped reward function. Nonetheless, this incorrect usage of the approximation of the optimal advantage function is less desirable than the appropriate and simpler approach of greedy maximization of it. From the perspective of the regret preference model, we also provide a clearer interpretation of fine tuning contemporary large language models with RLHF. This paper overall provides insight regarding why learning under the partial return preference model tends to work so well in practice, despite it conforming poorly to how humans give preferences. W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur O. Adalgeirsson, Serena Booth, Anca D. Dragan, Peter Stone 0001, Scott Niekum |
AAAI | 6 |
| 2024 | Minimum Coverage Sets for Training Robust Ad Hoc Teamwork AgentsabstractRobustly cooperating with unseen agents and human partners presents significant challenges due to the diverse cooperative conventions these partners may adopt. Existing Ad Hoc Teamwork (AHT) methods address this challenge by training an agent with a population of diverse teammate policies obtained through maximizing specific diversity metrics. However, prior heuristic-based diversity metrics do not always maximize the agent's robustness in all cooperative problems. In this work, we first propose that maximizing an AHT agent's robustness requires it to emulate policies in the minimum coverage set (MCS), the set of best-response policies to any partner policies in the environment. We then introduce the L-BRDiv algorithm that generates a set of teammate policies that, when used for AHT training, encourage agents to emulate policies from the MCS. L-BRDiv works by solving a constrained optimization problem to jointly train teammate policies for AHT training and approximating AHT agent policies that are members of the MCS. We empirically demonstrate that L-BRDiv produces more robust AHT agents than state-of-the-art methods in a broader range of two-player cooperative problems without the need for extensive hyperparameter tuning for its objectives. Our study shows that L-BRDiv outperforms the baseline methods by prioritizing discovering distinct members of the MCS instead of repeatedly finding redundant policies. Muhammad Rahman 0008, Jiaxun Cui, Peter Stone 0001 |
AAAI | 3 |
| 2024 | Building Minimal and Reusable Causal State Abstractions for Reinforcement LearningabstractTwo desiderata of reinforcement learning (RL) algorithms are the ability to learn from relatively little experience and the ability to learn policies that generalize to a range of problem specifications. In factored state spaces, one approach towards achieving both goals is to learn state abstractions, which only keep the necessary variables for learning the tasks at hand. This paper introduces Causal Bisimulation Modeling (CBM), a method that learns the causal relationships in the dynamics and reward functions for each task to derive a minimal, task-specific abstraction. CBM leverages and improves implicit modeling to train a high-fidelity causal dynamics model that can be reused for all tasks in the same environment. Empirical validation on two manipulation environments and four tasks reveals that CBM's learned implicit dynamics models identify the underlying causal relationships and state abstractions more accurately than explicit ones. Furthermore, the derived state abstractions allow a task learner to achieve near-oracle levels of sample efficiency and outperform baselines on all tasks. Caroline Wang, Xuesu Xiao, Yuke Zhu, Peter Stone 0001 |
AAAI | 5 |
| 2024 | Sample Efficient Myopic Exploration Through Multitask Reinforcement Learning with Diverse TasksabstractMultitask Reinforcement Learning (MTRL) approaches have gained increasing attention for its wide applications in many important Reinforcement Learning (RL) tasks. However, while recent advancements in MTRL theory have focused on the improved statistical efficiency by assuming a shared structure across tasks, exploration--a crucial aspect of RL--has been largely overlooked. This paper addresses this gap by showing that when an agent is trained on a sufficiently diverse set of tasks, a generic policy-sharing algorithm with myopic exploration design like $\epsilon$-greedy that are inefficient in general can be sample-efficient for MTRL. To the best of our knowledge, this is the first theoretical demonstration of the "exploration benefits" of MTRL. It may also shed light on the enigmatic success of the wide applications of myopic exploration in practice. To validate the role of diversity, we conduct experiments on synthetic robotic control environments, where the diverse task set aligns with the task selection by automatic curriculum learning, which is empirically shown to improve sample-efficiency. Ziping Xu, Zifan Xu, Runxuan Jiang, Peter Stone 0001, Ambuj Tewari |
ICLR | 4 |
| 2024 | Wait, That Feels Familiar: Learning to Extrapolate Human Preferences for Preference-Aligned Path PlanningabstractAutonomous mobility tasks such as last-mile delivery require reasoning about operator-indicated preferences over terrains on which the robot should navigate to ensure both robot safety and mission success. However, coping with out of distribution data from novel terrains or appearance changes due to lighting variations remains a fundamental problem in visual terrain-adaptive navigation. Existing solutions either require labor-intensive manual data re-collection and labeling or use hand-coded reward functions that may not align with operator preferences. In this work, we posit that operator preferences for visually novel terrains, which the robot should adhere to, can often be extrapolated from established terrain preferences within the inertial-proprioceptive-tactile domain. Leveraging this insight, we introduce Preference extrApolation for Terrain-awarE Robot Navigation (PATERN), a novel framework for extrapolating operator terrain preferences for visual navigation. PATERN learns to map inertial-proprioceptive-tactile measurements from the robot’s observations to a representation space and performs nearest-neighbor search in this space to estimate operator preferences over novel terrains. Through physical robot experiments in outdoor environments, we assess PATERN’s capability to extrapolate preferences and generalize to novel terrains and challenging lighting conditions. Compared to baseline approaches, our findings indicate that PATERN1robustly generalizes to diverse terrains and varied lighting conditions, while navigating in a preference-aligned manner. Haresh Karnan, Elvin Yang, Garrett Warnell, Joydeep Biswas, Peter Stone 0001 |
ICRA | 5 |
| 2024 | Rethinking Social Robot Navigation: Leveraging the Best of Two WorldsabstractEmpowering robots to navigate in a socially compliant manner is essential for the acceptance of robots moving in human-inhabited environments. Previously, roboticists have developed geometric navigation systems with decades of empirical validation to achieve safety and efficiency. However, the many complex factors of social compliance make geometric navigation systems hard to adapt to social situations, where no amount of tuning enables them to be both safe (people are too unpredictable) and efficient (the frozen robot problem). With recent advances in deep learning approaches, the common reaction has been to entirely discard these classical navigation systems and start from scratch, building a completely new learning-based social navigation planner. In this work, we find that this reaction is unnecessarily extreme: using a large-scale real-world social navigation dataset, SCAND, we find that geometric systems can produce trajectory plans that align with the human demonstrations in a large number of social situations. We, therefore, ask if we can rethink the social robot navigation problem by leveraging the advantages of both geometric and learning-based methods. We validate this hybrid paradigm through a proof-of-concept experiment, in which we develop a hybrid planner that switches between geometric and learning-based planning. Our experiments on both SCAND and two physical robots show that the hybrid planner can achieve better social compliance compared to using either the geometric or learning-based approach alone. Amir Hossain Raj, Zichao Hu, Haresh Karnan, Rohan Chandra, Amirreza Payandeh, Luisa Mao, Peter Stone 0001, Joydeep Biswas, Xuesu Xiao |
ICRA | 7 |
| 2024 | Asynchronous Task Plan Refinement for Multi-Robot Task and Motion PlanningabstractThis paper explores general multi-robot task and motion planning, where multiple robots in close proximity manipulate objects while satisfying constraints and a given goal. In particular, we formulate the plan refinement problem—which, given a task plan, finds valid assignments of variables corresponding to solution trajectories—as a hybrid constraint satisfaction problem. The proposed algorithm follows several design principles that yield the following features: (1) efficient solution finding due to sequential heuristics and implicit time and roadmap representations, and (2) maximized feasible solution space obtained by introducing minimally necessary coordination-induced constraints and not relying on prevalent simplifications that exist in the literature. The evaluation results demonstrate the planning efficiency of the proposed algorithm, outperforming the synchronous approach in terms of makespan. Yoonchang Sung, Rahul Shome, Peter Stone 0001 |
ICRA | 3 |
| 2024 | Dexterous Legged Locomotion in Confined 3D Spaces with Reinforcement LearningabstractRecent advances of locomotion controllers utilizing deep reinforcement learning (RL) have yielded impressive results in terms of achieving rapid and robust locomotion across challenging terrain, such as rugged rocks, non-rigid ground, and slippery surfaces. However, while these controllers primarily address challenges underneath the robot, relatively little research has investigated legged mobility through confined 3D spaces, such as narrow tunnels or irregular voids, which impose all-around constraints. The cyclic gait patterns resulted from existing RL-based methods to learn parameterized locomotion skills characterized by motion parameters, such as velocity and body height, may not be adequate to navigate robots through challenging confined 3D spaces, requiring both agile 3D obstacle avoidance and robust legged locomotion. Instead, we propose to learn locomotion skills end-to-end from goal-oriented navigation in confined 3D spaces. To address the inefficiency of tracking distant navigation goals, we introduce a hierarchical locomotion controller that combines a classical planner tasked with planning waypoints to reach a faraway global goal location, and an RL-based policy trained to follow these waypoints by generating low-level motion commands. This approach allows the policy to explore its own locomotion skills within the entire solution space and facilitates smooth transitions between local goals, enabling long-term navigation towards distant goals. In simulation, our hierarchical approach succeeds at navigating through demanding confined 3D environments, outperforming both pure end-to-end learning approaches and parameterized locomotion skills. We further demonstrate the successful real-world deployment of our simulation-trained controller on a real robot. Zifan Xu, Amir Hossain Raj, Xuesu Xiao, Peter Stone 0001 |
ICRA | 4 |
| 2024 | Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement LearningabstractA hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery methods often learn entangled skills where one skill variable simultaneously influences many entities in the environment, making downstream skill chaining extremely challenging. We propose Disentangled Unsupervised Skill Discovery (DUSDi), a method for learning disentangled skills that can be efficiently reused to solve downstream tasks. DUSDi decomposes skills into disentangled components, where each skill component only affects one factor of the state space. Importantly, these skill components can be concurrently composed to generate low-level actions, and efficiently chained to tackle downstream tasks through hierarchical Reinforcement Learning. DUSDi defines a novel mutual-information-based objective to enforce disentanglement between the influences of different skill components, and utilizes value factorization to optimize this objective efficiently. Evaluated in a set of challenging environments, DUSDi successfully learns disentangled skills, and significantly outperforms previous skill discovery methods when it comes to applying the learned skills to solve downstream tasks. Jiaheng Hu, Peter Stone 0001, Roberto Martin Martin |
NeurIPS | 3 |
| 2024 | Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy ExplorationabstractThe ability to approach the same problem from different angles is a cornerstone of human intelligence that leads to robust solutions and effective adaptation to problem variations. In contrast, current RL methodologies tend to lead to policies that settle on a single solution to a given problem, making them brittle to problem variations. Replicating human flexibility in reinforcement learning agents is the challenge that we explore in this work. We tackle this challenge by extending state-of-the-art approaches to introduce DUPLEX, a method that explicitly defines a diversity objective with constraints and makes robust estimates of policies’ expected behavior through successor features. The trained agents can (i) learn a diverse set of near-optimal policies in complex highly-dynamic environments and (ii) exhibit competitive and diverse skills in out-of-distribution (OOD) contexts. Empirical results indicate that DUPLEX improves over previous methods and successfully learns competitive driving styles in a hyper-realistic simulator (i.e., GranTurismo ™ 7) as well as diverse and effective policies in several multi-context robotics MuJoCo simulations with OOD gravity forces and height limits. To the best of our knowledge, our method is the first to achieve diverse solutions in complex driving simulators and OOD robotic contexts. DUPLEX agents demonstrating diverse behaviors can be found at https://ai.sony/publications/Discovering-Creative-Behaviors-through-DUPLEX-Diverse-Universal-Features-for-Policy-Exploration/. Borja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman, Peter Stone 0001 |
NeurIPS | 5 |
| 2024 | SkiLD: Unsupervised Skill Discovery Guided by Factor InteractionsabstractUnsupervised skill discovery carries the promise that an intelligent agent can learn reusable skills through autonomous, reward-free interactions with environments. Existing unsupervised skill discovery methods learn skills by encouraging distinguishable behaviors that cover diverse states. However, in complex environments with many state factors (e.g., household environments with many objects), learning skills that cover all possible states is impossible, and naively encouraging state diversity often leads to simple skills that are not ideal for solving downstream tasks. This work introduces Skill Discovery from Local Dependencies (SkiLD), which leverages state factorization as a natural inductive bias to guide the skill learning process. The key intuition guiding SkiLD is that skills that induce \textbf{diverse interactions} between state factors are often more valuable for solving downstream tasks. To this end, SkiLD develops a novel skill learning objective that explicitly encourages the mastering of skills that effectively induce different interactions within an environment. We evaluate SkiLD in several domains with challenging, long-horizon sparse reward tasks including a realistic simulated household robot domain, where SkiLD successfully learns skills with clear semantic meaning and shows superior performance compared to existing unsupervised reinforcement learning methods that only maximize state coverage. Jiaheng Hu, Caleb Chuck, Roberto Martin Martin, Amy Zhang 0001, Scott Niekum, Peter Stone 0001 |
NeurIPS | 8 |
| 2024 | N-agent Ad Hoc TeamworkabstractCurrent approaches to learning cooperative multi-agent behaviors assume relatively restrictive settings. In standard fully cooperative multi-agent reinforcement learning, the learning algorithm controls *all* agents in the scenario, while in ad hoc teamwork, the learning algorithm usually assumes control over only a *single* agent in the scenario. However, many cooperative settings in the real world are much less restrictive. For example, in an autonomous driving scenario, a company might train its cars with the same learning algorithm, yet once on the road, these cars must cooperate with cars from another company. Towards expanding the class of scenarios that cooperative learning methods may optimally address, we introduce $N$*-agent ad hoc teamwork* (NAHT), where a set of autonomous agents must interact and cooperate with dynamically varying numbers and types of teammates. This paper formalizes the problem, and proposes the *Policy Optimization with Agent Modelling* (POAM) algorithm. POAM is a policy gradient, multi-agent reinforcement learning approach to the NAHT problem, that enables adaptation to diverse teammate behaviors by learning representations of teammate behaviors. Empirical evaluation on tasks from the multi-agent particle environment and StarCraft II shows that POAM improves cooperative task returns compared to baseline approaches, and enables out-of-distribution generalization to unseen teammates. Caroline Wang, Arrasy Rahman, Ishan Durugkar, Elad Liebman, Peter Stone 0001 |
NeurIPS | 5 |
| 2024 | Dobby: A Conversational Service Robot Driven by GPT-4abstractThis work introduces a robotics platform which comprehensively integrates multi-step action execution, natural language understanding, and memory to interactively perform service tasks in accordance with variable needs and intentions of users. The proposed architecture is built around an AI agent, derived from GPT-4, which is embedded in an embodied system. Our approach utilizes semantic matching, plan validation, and state messages to ground the agent in the physical world, enabling a seamless merger between communication and behavior. We demonstrate the advantages of this system with an HRI study comparing mobile robots with and without conversational AI capabilities in a free-form tour-guide scenario. The increased adaptability of the system is measured along five dimensions: flexible task planning, interactive exploration of information, emotional-friendliness, personalization, and increased overall user satisfaction. Carson Stark, Bohkyung Chun, Casey Charleston, Varsha Ravi, Luis Pabon, Surya Sunkari, Tarun Mohan, Peter Stone 0001, Justin W. Hart |
RO-MAN | 8 |
| 2024 | Data-Efficient Policy Evaluation Through Behavior Policy SearchabstractWe consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its performance. We show that the data collected from deploying a different policy, commonly called the behavior policy, can be used to produce unbiased estimates with lower mean squared error than this standard technique. We derive an analytic expression for a minimal variance behavior policy -- a behavior policy that minimizes the mean squared error of the resulting estimates. Because this expression depends on terms that are unknown in practice, we propose a novel policy evaluation sub-problem, behavior policy search: searching for a behavior policy that reduces mean squared error. We present two behavior policy search algorithms and empirically demonstrate their effectiveness in lowering the mean squared error of policy performance estimates. Josiah Hanna, Yash Chandak, Philip S. Thomas, Martha White, Peter Stone 0001, Scott Niekum |
J. Mach. Learn. Res. | 5 |
| 2024 | Conflict Avoidance in Social Navigation - a SurveyabstractA major goal in robotics is to enable intelligent mobile robots to operate smoothly in shared human-robot environments. One of the most fundamental capabilities in service of this goal is competent navigation in this “social” context. As a result, there has been a recent surge of research on social navigation; and especially as it relates to the handling of conflicts between agents during social navigation. These developments introduce a variety of models and algorithms, however as this research area is inherently interdisciplinary, many of the relevant papers are not comparable and there is no shared standard vocabulary. This survey aims at bridging this gap by introducing such a common language, using it to survey existing work, and highlighting open problems. It starts by defining the boundaries of this survey to a limited, yet highly common type of social navigation—conflict avoidance. Within this proposed scope, this survey introduces a detailed taxonomy of the conflict avoidance components. This survey then maps existing work into this taxonomy, while discussing papers using its framing. Finally, this article proposes some future research directions and open problems that are currently on the frontier of social navigation to aid ongoing and future research. Reuth Mirsky, Xuesu Xiao, Justin W. Hart, Peter Stone 0001 |
ACM Trans. Hum. Robot Interact. | 4 |
| 2023 | Metric Residual Network for Sample Efficient Goal-Conditioned Reinforcement LearningabstractGoal-conditioned reinforcement learning (GCRL) has a wide range of potential real-world applications, including manipulation and navigation problems in robotics. Especially in such robotics tasks, sample efficiency is of the utmost importance for GCRL since, by default, the agent is only rewarded when it reaches its goal. While several methods have been proposed to improve the sample efficiency of GCRL, one relatively under-studied approach is the design of neural architectures to support sample efficiency. In this work, we introduce a novel neural architecture for GCRL that achieves significantly better sample efficiency than the commonly-used monolithic network architecture. The key insight is that the optimal action-value function must satisfy the triangle inequality in a specific sense. Furthermore, we introduce the metric residual network (MRN) that deliberately decomposes the action-value function into the negated summation of a metric plus a residual asymmetric component. MRN provably approximates any optimal action-value function, thus making it a fitting neural architecture for GCRL. We conduct comprehensive experiments across 12 standard benchmark environments in GCRL. The empirical results demonstrate that MRN uniformly outperforms other state-of-the-art GCRL neural architectures in terms of sample efficiency. The code is available at https://github.com/Cranial-XIX/metric-residual-network. Bo Liu 0042, Yihao Feng, Qiang Liu 0001, Peter Stone 0001 |
AAAI | 4 |
| 2023 | The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsabstractIn reinforcement learning (RL), a reward function that aligns exactly with a task's true performance metric is often necessarily sparse. For example, a true task metric might encode a reward of 1 upon success and 0 otherwise. The sparsity of these true task metrics can make them hard to learn from, so in practice they are often replaced with alternative dense reward functions. These dense reward functions are typically designed by experts through an ad hoc process of trial and error. In this process, experts manually search for a reward function that improves performance with respect to the task metric while also enabling an RL algorithm to learn faster. This process raises the question of whether the same reward function is optimal for all algorithms, i.e., whether the reward function can be overfit to a particular algorithm. In this paper, we study the consequences of this wide yet unexamined practice of trial-and-error reward design. We first conduct computational experiments that confirm that reward functions can be overfit to learning algorithms and their hyperparameters. We then conduct a controlled observation study which emulates expert practitioners' typical experiences of reward design, in which we similarly find evidence of reward function overfitting. We also find that experts' typical approach to reward design---of adopting a myopic strategy and weighing the relative goodness of each state-action pair---leads to misdesign through invalid task specifications, since RL algorithms use cumulative reward rather than rewards for individual state-action pairs as an optimization target. Code, data: github.com/serenabooth/reward-design-perils Serena Booth, W. Bradley Knox, Julie A. Shah, Scott Niekum, Peter Stone 0001, Alessandro Allievi |
AAAI | 5 |
| 2023 | DM²: Decentralized Multi-Agent Reinforcement Learning via Distribution MatchingabstractCurrent approaches to multi-agent cooperation rely heavily on centralized mechanisms or explicit communication protocols to ensure convergence. This paper studies the problem of distributed multi-agent learning without resorting to centralized components or explicit communication. It examines the use of distribution matching to facilitate the coordination of independent agents. In the proposed scheme, each agent independently minimizes the distribution mismatch to the corresponding component of a target visitation distribution. The theoretical analysis shows that under certain conditions, each agent minimizing its individual distribution mismatch allows the convergence to the joint policy that generated the target distribution. Further, if the target distribution is from a joint policy that optimizes a cooperative task, the optimal policy for a combination of this task reward and the distribution matching reward is the same joint policy. This insight is used to formulate a practical algorithm (DM^2), in which each individual agent matches a target distribution derived from concurrently sampled trajectories from a joint expert policy. Experimental validation on the StarCraft domain shows that combining (1) a task reward, and (2) a distribution matching reward for expert demonstrations for the same task, allows agents to outperform a naive distributed baseline. Additional experiments probe the conditions under which expert demonstrations need to be sampled to obtain the learning benefits. Caroline Wang, Ishan Durugkar, Elad Liebman, Peter Stone 0001 |
AAAI | 4 |
| 2023 | MACTA: A Multi-agent Reinforcement Learning Approach for Cache Timing Attacks and Detection
Jiaxun Cui, Mulong Luo, Gary Geunbae Lee, Peter Stone 0001, Hsien-Hsin S. Lee, G. Edward Suh, Wenjie Xiong 0001, Yuandong Tian |
ICLR | 5 |
| 2023 | Learning Perceptual Hallucination for Multi-Robot Navigation in Narrow HallwaysabstractWhile current systems for autonomous robot navigation can produce safe and efficient motion plans in static environments, they usually generate suboptimal behaviors when multiple robots must navigate together in confined spaces. For example, when two robots meet each other in a narrow hallway, they may either turn around to find an alternative route or collide with each other. This paper presents a new approach to navigation that allows two robots to pass each other in a narrow hallway without colliding, stopping, or waiting. Our approach, Perceptual Hallucination for Hallway Passing (PHHP), learns to synthetically generate virtual obstacles (i.e., perceptual hallucination) to facilitate passing in narrow hallways by multiple robots that utilize otherwise standard autonomous navigation systems. Our experiments on physical robots in a variety of hallways show improved performance compared to multiple baselines. Jin Soo Park, Xuesu Xiao, Garrett Warnell, Harel Yedidsion, Peter Stone 0001 |
ICRA | 5 |
| 2023 | Benchmarking Reinforcement Learning Techniques for Autonomous NavigationabstractDeep reinforcement learning (RL) has brought many successes for autonomous robot navigation. However, there still exists important limitations that prevent real-world use of RL-based navigation systems. For example, most learning approaches lack safety guarantees; and learned navigation systems may not generalize well to unseen environments. Despite a variety of recent learning techniques to tackle these challenges in general, a lack of an open-source benchmark and reproducible learning methods specifically for autonomous navigation makes it difficult for roboticists to choose what learning methods to use for their mobile robots and for learning researchers to identify current shortcomings of general learning methods for autonomous navigation. In this paper, we identify four major desiderata of applying deep RL approaches for autonomous navigation: (D1) reasoning under uncertainty, (D2) safety, (D3) learning from limited trial-and-error data, and (D4) generalization to diverse and novel environments. Then, we explore four major classes of learning techniques with the purpose of achieving one or more of the four desiderata: memory-based neural network architectures (D1), safe RL (D2), model-based RL (D2, D3), and domain randomization (D4). By deploying these learning techniques in a new open-source large-scale navigation benchmark and real-world environments, we perform a comprehensive study aimed at establishing to what extent can these techniques achieve these desiderata for RL-based navigation systems. Zifan Xu, Bo Liu 0042, Xuesu Xiao, Anirudh Nair, Peter Stone 0001 |
ICRA | 5 |
| 2023 | Symbolic State Space Optimization for Long Horizon Mobile Manipulation PlanningabstractIn existing task and motion planning (TAMP) research, it is a common assumption that experts manually specify the state space for task-level planning. A well-developed state space enables the desirable distribution of limited computational resources between task planning and motion planning. However, developing such task-level state spaces can be non-trivial in practice. In this paper, we consider a long horizon mobile manipulation domain including repeated navigation and manipulation. We propose Symbolic State Space Optimization (S3O) for computing a set of abstracted locations and their 2D geometric groundings for generating task-motion plans in such domains. Our approach has been extensively evaluated in simulation and demonstrated on a real mobile manipulator working on clearing up dining tables. Results show the superiority of the proposed method over TAMP baselines in task completion rate and execution time. Xiaohan Zhang 0002, Yan Ding 0002, Yuqian Jiang, Yuke Zhu, Peter Stone 0001, Shiqi Zhang 0001 |
IROS | 6 |
| 2023 | A Novel Control Law for Multi-Joint Human-Robot Interaction Tasks While Maintaining Postural CoordinationabstractExoskeleton robots are capable of safe torque-controlled interactions with a wearer while moving their limbs through predefined trajectories. However, affecting and assisting the wearer's movements while incorporating their inputs (effort and movements) effectively during an interaction re-mains an open problem due to the complex and variable nature of human motion. In this paper, we present a control algorithm that leverages task-specific movement behaviors to control robot torques during unstructured interactions by implementing a force field that imposes a desired joint angle coordination behavior. This control law, built by using principal component analysis (PCA), is implemented and tested with the Harmony exoskeleton. We show that the proposed control law is versatile enough to allow for the imposition of different coordination behaviors with varying levels of impedance stiffness. We also test the feasibility of our method for unstructured human-robot interaction. Specifically, we demonstrate that participants in a human-subject experiment are able to effectively perform reaching tasks while the exoskeleton imposes the desired joint coordination under different movement speeds and interaction modes. Survey results further suggest that the proposed control law may offer a reduction in cognitive or motor effort. This control law opens up the possibility of using the exoskeleton for training the participating in accomplishing complex multi-joint motor tasks while maintaining postural coordination. Keya Ghonasgi, Reuth Mirsky, Adrian M. Haith, Peter Stone 0001, Ashish D. Deshpande |
IROS | 4 |
| 2023 | f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-DivergencesabstractGoal-Conditioned Reinforcement Learning (RL) problems often have access to sparse rewards where the agent receives a reward signal only when it has achieved the goal, making policy optimization a difficult problem.
Several works augment this sparse reward with a learned dense reward function, but this can lead to sub-optimal policies if the reward is misaligned.
Moreover, recent works have demonstrated that effective shaping rewards for a particular problem can depend on the underlying learning algorithm.
This paper introduces a novel way to encourage exploration called
$f$-Policy Gradients, or $f$-PG. $f$-PG minimizes the f-divergence between the agent's state visitation distribution and the goal, which we show can lead to an optimal policy. We derive gradients for various f-divergences to optimize this objective. Our learning paradigm provides dense learning signals for exploration in sparse reward settings. We further introduce an entropy-regularized policy optimization objective, that we call $state$-MaxEnt RL (or $s$-MaxEnt RL) as a special case of our objective. We show that several metric-based shaping rewards like L2 can be used with $s$-MaxEnt RL, providing a common ground to study such metric-based shaping rewards with efficient exploration. We find that $f$-PG has better performance compared to standard policy gradient methods on a challenging gridworld as well as the Point Maze and FetchReach environments. More information on our website https://agarwalsiddhant10.github.io/projects/fpg.html. Siddhant Agarwal, Ishan Durugkar, Peter Stone 0001, Amy Zhang 0001 |
NeurIPS | 3 |
| 2023 | FAMO: Fast Adaptive Multitask OptimizationabstractOne of the grand enduring goals of AI is to create generalist agents that can learn multiple different tasks from diverse data via multitask learning (MTL). However, in practice, applying gradient descent (GD) on the average loss across all tasks may yield poor multitask performance due to severe under-optimization of certain tasks. Previous approaches that manipulate task gradients for a more balanced loss decrease require storing and computing all task gradients ($\mathcal{O}(k)$ space and time where $k$ is the number of tasks), limiting their use in large-scale scenarios. In this work, we introduce Fast Adaptive Multitask Optimization (FAMO), a dynamic weighting method that decreases task losses in a balanced way using $\mathcal{O}(1)$ space and time. We conduct an extensive set of experiments covering multi-task supervised and reinforcement learning problems. Our results indicate that FAMO achieves comparable or superior performance to state-of-the-art gradient manipulation techniques while offering significant improvements in space and computational efficiency. Code is available at \url{https://github.com/Cranial-XIX/FAMO}. Bo Liu 0042, Yihao Feng, Peter Stone 0001, Qiang Liu 0001 |
NeurIPS | 3 |
| 2023 | LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot LearningabstractLifelong learning offers a promising paradigm of building a generalist agent that learns and adapts over its lifespan. Unlike traditional lifelong learning problems in image and text domains, which primarily involve the transfer of declarative knowledge of entities and concepts, lifelong learning in decision-making (LLDM) also necessitates the transfer of procedural knowledge, such as actions and behaviors. To advance research in LLDM, we introduce LIBERO, a novel benchmark of lifelong learning for robot manipulation. Specifically, LIBERO highlights five key research topics in LLDM: 1) how to efficiently transfer declarative knowledge, procedural knowledge, or the mixture of both; 2) how to design effective policy architectures and 3) effective algorithms for LLDM; 4) the robustness of a lifelong learner with respect to task ordering; and 5) the effect of model pretraining for LLDM. We develop an extendible procedural generation pipeline that can in principle generate infinitely many tasks. For benchmarking purpose, we create four task suites (130 tasks in total) that we use to investigate the above-mentioned research topics. To support sample-efficient learning, we provide high-quality human-teleoperated demonstration data for all tasks. Our extensive experiments present several insightful or even unexpected discoveries: sequential finetuning outperforms existing lifelong learning methods in forward transfer, no single visual encoder architecture excels at all types of knowledge transfer, and naive supervised pretraining can hinder agents' performance in the subsequent LLDM. Bo Liu 0042, Chongkai Gao, Yihao Feng, Qiang Liu 0001, Yuke Zhu, Peter Stone 0001 |
NeurIPS | 7 |
| 2023 | ELDEN: Exploration via Local DependenciesabstractTasks with large state space and sparse rewards present a longstanding challenge to reinforcement learning. In these tasks, an agent needs to explore the state space efficiently until it finds a reward. To deal with this problem, the community has proposed to augment the reward function with intrinsic reward, a bonus signal that encourages the agent to visit interesting states. In this work, we propose a new way of defining interesting states for environments with factored state spaces and complex chained dependencies, where an agent's actions may change the value of one entity that, in order, may affect the value of another entity. Our insight is that, in these environments, interesting states for exploration are states where the agent is uncertain whether (as opposed to how) entities such as the agent or objects have some influence on each other. We present ELDEN, Exploration via Local DepENdencies, a novel intrinsic reward that encourages the discovery of new interactions between entities. ELDEN utilizes a novel scheme --- the partial derivative of the learned dynamics to model the local dependencies between entities accurately and computationally efficiently. The uncertainty of the predicted dependencies is then used as an intrinsic reward to encourage exploration toward new interactions. We evaluate the performance of ELDEN on four different domains with complex dependencies, ranging from 2D grid worlds to 3D robotic tasks. In all domains, ELDEN correctly identifies local dependencies and learns successful policies, significantly outperforming previous state-of-the-art exploration methods. Jiaheng Hu, Peter Stone 0001, Roberto Martin Martin |
NeurIPS | 3 |
| 2023 | Composing Efficient, Robust Tests for Policy SelectionabstractModern reinforcement learning systems produce many high-quality policies throughout the learning process. However, to choose which policy to actually deploy in the real world, they must be tested under an intractable number of environmental conditions. We introduce RPOSST, an algorithm to select a small set of test cases from a larger pool based on a relatively small number of sample evaluations. RPOSST treats the test case selection problem as a two-player game and optimizes a solution with provable $k$-of-$N$ robustness, bounding the error relative to a test that used all the test cases in the pool. Empirical results demonstrate that RPOSST finds a small set of test cases that identify high quality policies in a toy one-shot game, poker datasets, and a high-fidelity racing simulator. Dustin Morrill, Thomas J. Walsh 0001, Daniel Hernandez, Peter R. Wurman, Peter Stone 0001 |
UAI | 5 |
| 2023 | Reward (Mis)design for autonomous drivingabstractThis article considers the problem of diagnosing certain common errors in reward design. Its insights are also applicable to the design of cost functions and performance metrics more generally. To diagnose common errors, we develop 8 simple sanity checks for identifying flaws in reward functions. We survey research that is published in top-tier venues and focuses on reinforcement learning (RL) for autonomous driving (AD). Specifically, we closely examine the reported reward function in each publication and present these reward functions in a complete and standardized format in the appendix. Wherever we have sufficient information, we apply the 8 sanity checks to each surveyed reward function, revealing near-universal flaws in reward design for AD that might also exist pervasively across reward design for other tasks. Lastly, we explore promising directions that may aid the design of reward functions for AD in subsequent research, following a process of inquiry that can be adapted to other domains. W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt 0001, Peter Stone 0001 |
Artif. Intell. | 5 |
| 2023 | A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha |
Neural Networks | 37 |
| 2022 | Coopernaut: End-to-End Driving with Cooperative Perception for Networked VehiclesabstractOptical sensors and learning algorithms for autonomous vehicles have dramatically advanced in the past few years. Nonetheless, the reliability of today's autonomous vehicles is hindered by the limited line-of-sight sensing capability and the brittleness of data-driven methods in handling extreme situations. With recent developments of telecommunication technologies, cooperative perception with vehicle-to-vehicle communications has become a promising paradigm to enhance autonomous driving in dangerous or emergency situations. We introduce Coopernaut,an end-to-end learning model that uses cross-vehicle perception for vision-based cooperative driving. Our model encodes Li-DAR information into compact point-based representations that can be transmitted as messages between vehicles via realistic wireless channels. To evaluate our model, we develop Autocastsim,a network-augmented driving simulation framework with example accident-prone scenarios. Our experiments on Autocastsim suggest that our cooperative perception driving models lead to a 40% improvement in average success rate over egocentric driving mod-els in these challenging driving situations and a$5\times$smaller bandwidth requirement than prior work V2VNet. Cooper-nautand Autocastsim are available at https://ut-austin-rpl.github.io/Coopernaut/. Jiaxun Cui, Hang Qiu 0001, Dian Chen 0005, Peter Stone 0001, Yuke Zhu |
CVPR | 4 |
| 2022 | A Survey of Ad Hoc Teamwork Research
Reuth Mirsky, Ignacio Carlucho, Arrasy Rahman, Elliot Fosong, William Macke, Mohan Sridharan, Peter Stone 0001, Stefano V. Albrecht |
EUMAS | 7 |
| 2022 | Effective mutation rate adaptation through group elite selectionabstractEvolutionary algorithms are sensitive to the mutation rate (MR); no single value of this parameter works well across domains. Self-adaptive MR approaches have been proposed but they tend to be brittle: Sometimes they decay the MR to zero, thus halting evolution. To make self-adaptive MR robust, this paper introduces the Group Elite Selection of Mutation Rates (GESMR) algorithm. GESMR co-evolves a population of solutions and a population of MRs, such that each MR is assigned to a group of solutions. The resulting best mutational change in the group, instead of average mutational change, is used for MR selection during evolution, thus avoiding the vanishing MR problem. With the same number of function evaluations and with almost no overhead, GESMR converges faster and to better solutions than previous approaches on a wide range of continuous test optimization problems. GESMR also scales well to high-dimensional neuroevolution for supervised image-classification tasks and for reinforcement learning control tasks. Remarkably, GESMR produces MRs that are optimal in the long-term, as demonstrated through a comprehensive look-ahead grid search. Thus, GESMR and its theoretical and empirical analysis demonstrate how self-adaptation can be harnessed to improve performance in several applications of evolutionary computation. Akarsh Kumar, Bo Liu 0042, Risto Miikkulainen, Peter Stone 0001 |
GECCO | 4 |
| 2022 | Causal Dynamics Learning for Task-Independent State AbstractionabstractLearning dynamics models accurately is an important goal for Model-Based Reinforcement Learning (MBRL), but most MBRL methods learn a dense dynamics model which is vulnerable to spurious correlations and therefore generalizes poorly to unseen states. In this paper, we introduce Causal Dynamics Learning for Task-Independent State Abstraction (CDL), which first learns a theoretically proved causal dynamics model that removes unnecessary dependencies between state variables and the action, thus generalizing well to unseen states. A state abstraction can then be derived from the learned dynamics, which not only improves sample efficiency but also applies to a wider range of tasks than existing state abstraction methods. Evaluated on two simulated environments and downstream tasks, both the dynamics model and policies learned by the proposed method generalize well to unseen states and the derived state abstraction improves sample efficiency compared to learning without it. Xuesu Xiao, Zifan Xu, Yuke Zhu, Peter Stone 0001 |
ICML | 5 |
| 2022 | Skeletal Feature Compensation for Imitation Learning with Embodiment MismatchabstractLearning from demonstrations in the wild (e.g. YouTube videos) is a tantalizing goal in imitation learning. However, for this goal to be achieved, imitation learning algorithms must deal with the fact that the demonstrators and learners may have bodies that differ from one another. This condition — “embodiment mismatch” — is ignored by many recent imitation learning algorithms. Our proposed imitation learning technique, SILEM (Skeletal feature compensation for Imitation Learning with Embodiment Mismatch), addresses a particular type of embodiment mismatch by introducing a learned affine transform to compensate for differences in the skeletal features obtained from the learner and expert. We create toy domains based on PyBullet's HalfCheetah and Ant to assess SILEM's benefits for this type of embodiment mismatch. We also provide qualitative and quantitative results on more realistic problems — teaching simulated humanoid agents, including Atlas from Boston Dynamics, to walk by observing human demonstrations. Eddy Hudson, Garrett Warnell, Faraz Torabi, Peter Stone 0001 |
ICRA | 4 |
| 2022 | Adversarial Imitation Learning from Video Using a State ObserverabstractThe imitation learning research community has recently made significant progress towards the goal of enabling artificial agents to imitate behaviors from video demonstrations alone. However, current state-of-the-art approaches developed for this problem exhibit high sample complexity due, in part, to the high-dimensional nature of video observations. Towards addressing this issue, we introduce here a new algorithm called Visual Generative Adversarial Imitation from Observation using a State Observer (VGAIfO-SO). At its core, VGAIfO-SO seeks to address sample inefficiency using a novel, self-supervised state observer, which provides estimates of lower-dimensional proprioceptive state representations from high-dimensional images. We show experimentally in several continuous control environments that VGAIfO-SO is more sample efficient than other IfO algorithms at learning from video-only demonstrations and can sometimes even achieve performance close to the Generative Adversarial Imitation from Observation (GAIfO) algorithm that has privileged access to the demonstrator's proprioceptive state information. Haresh Karnan, Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
ICRA | 4 |
| 2022 | VOILA: Visual-Observation-Only Imitation Learning for Autonomous NavigationabstractWhile imitation learning for vision-based au-tonomous mobile robot navigation has recently received a great deal of attention in the research community, existing approaches typically require state-action demonstrations that were gathered using the deployment platform. However, what if one cannot easily outfit their platform to record these demonstration signals or-worse yet-the demonstrator does not have access to the platform at all? Is imitation learning for vision-based autonomous navigation even possible in such scenarios? In this work, we hypothesize that the answer is yes and that recent ideas from the Imitation from Observation (IfO) literature can be brought to bear such that a robot can learn to navigate using only ego-centric video collected by a demonstrator, even in the presence of viewpoint mismatch. To this end, we introduce a new algorithm, Visual-Observation-only Imitation Learning for Autonomous navigation (VOILA), that can successfully learn navigation policies from a single video demonstration collected from a physically different agent. We evaluate VOILA in the AirSim simulator and show that VOILA not only successfully imitates the expert, but that it also learns navigation policies that can generalize to novel environments. Further, we demonstrate the effectiveness of VOILA in a real-world setting by showing that it allows a wheeled Jackal robot to successfully imitate a human walking in an environment while recording video with a handheld mobile phone camera. Haresh Karnan, Garrett Warnell, Xuesu Xiao, Peter Stone 0001 |
ICRA | 4 |
| 2022 | Visually Grounded Task and Motion Planning for Mobile ManipulationabstractTask and motion planning (TAMP) algorithms aim to help robots achieve task-level goals, while maintaining motion-level feasibility. This paper focuses on TAMP domains that involve robot behaviors that take extended periods of time (e.g., long-distance navigation). In this paper, we develop a visual grounding approach to help robots probabilistically evaluate action feasibility, and introduce a TAMP algorithm, called GROP, that optimizes both feasibility and efficiency. We have collected a dataset that includes 96, 000 simulated trials of a robot conducting mobile manipulation tasks, and then used the dataset to learn to ground symbolic spatial relationships for action feasibility evaluation. Compared with competitive TAMP baselines, GROP exhibited a higher task-completion rate while maintaining lower or comparable action costs. In addition to these extensive experiments in simulation, GROP is fully implemented and tested on a real robot system. Xiaohan Zhang 0002, Yan Ding 0002, Yuke Zhu, Peter Stone 0001, Shiqi Zhang 0001 |
ICRA | 5 |
| 2022 | Dynamic Sparse Training for Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) agents are trained through trial-and-error interactions with the environment. This leads to a long training time for dense neural networks to achieve good performance. Hence, prohibitive computation and memory resources are consumed. Recently, learning efficient DRL agents has received increasing attention. Yet, current methods focus on accelerating inference time. In this paper, we introduce for the first time a dynamic sparse training approach for deep reinforcement learning to accelerate the training process. The proposed approach trains a sparse neural network from scratch and dynamically adapts its topology to the changing data distribution during training. Experiments on continuous control tasks show that our dynamic sparse agents achieve higher performance than the equivalent dense methods, reduce the parameter count and floating-point operations (FLOPs) by 50%, and have a faster learning speed that enables reaching the performance of dense agents with 40−50% reduction in the training steps. Ghada Sokar, Elena Mocanu, Decebal Constantin Mocanu, Mykola Pechenizkiy, Peter Stone 0001 |
IJCAI | 5 |
| 2022 | Quantifying Changes in Kinematic Behavior of a Human-Exoskeleton Interactive SystemabstractWhile human-robot interaction studies are becoming more common, quantification of the effects of repeated interaction with an exoskeleton remains unexplored. We draw upon existing literature in human skill assessment and present extrinsic and intrinsic performance metrics that quantify how the human-exoskeleton system's behavior changes over time. Specifically, in this paper, we present a new performance metric that provides insight into the system's kinematics associated with ‘successful’ movements resulting in a richer characterization of changes in the system's behavior. A human subject study is carried out wherein participants learn to play a challenging and dynamic reaching game over multiple attempts, while donning an upper-body exoskeleton. The results demonstrate that repeated practice results in learning over time as identified through the improvement of extrinsic performance. Changes in the newly developed kinematics-based measure further illumi-nate how the participant's intrinsic behavior is altered over the training period. Thus, we are able to quantify the changes in the human-exoskeleton system's behavior observed in relation with learning. Keya Ghonasgi, Reuth Mirsky, Adrian M. Haith, Peter Stone 0001, Ashish D. Deshpande |
IROS | 4 |
| 2022 | VI-IKD: High-Speed Accurate Off-Road Navigation using Learned Visual-Inertial Inverse KinodynamicsabstractOne of the key challenges in high-speed off-road navigation on ground vehicles is that the kinodynamics of the vehicle-terrain interaction can differ dramatically depending on the terrain. Previous approaches to addressing this challenge have considered learning an inverse kinodynamics (IKD) model, conditioned on inertial information of the vehicle to sense the kinodynamic interactions. In this paper, we hypothesize that to enable accurate high-speed off-road navigation using a learned IKD model, in addition to inertial information from the past, one must also anticipate the kinodynamic interactions of the vehicle with the terrain in the future. To this end, we introduce Visual-Inertial Inverse Kinodynamics (VI-IKD), a novel learning based IKD model that is conditioned on visual information from a terrain patch ahead of the robot in addition to past inertial information, enabling it to anticipate kinodynamic interactions in the future. We validate the effectiveness of VI-IKD in accurate high-speed off-road navigation experimentally on a scale 1/5 UT-AlphaTruck off-road autonomous vehicle in both indoor and outdoor environments and show that compared to other state-of-the-art approaches, VI-IKD enables more accurate and robust off-road navigation on a variety of different terrains at speeds of up to 3.5m/s. Haresh Karnan, Kavan Singh Sikand, Pranav Atreya, Sadegh Rabiee, Xuesu Xiao, Garrett Warnell, Peter Stone 0001, Joydeep Biswas |
IROS | 7 |
| 2022 | BOME! Bilevel Optimization Made Easy: A Simple First-Order ApproachabstractBilevel optimization (BO) is useful for solving a variety of important machine learning problems including but not limited to hyperparameter optimization, meta-learning, continual learning, and reinforcement learning.Conventional BO methods need to differentiate through the low-level optimization process with implicit differentiation, which requires expensive calculations related to the Hessian matrix. There has been a recent quest for first-order methods for BO, but the methods proposed to date tend to be complicated and impractical for large-scale deep learning applications. In this work, we propose a simple first-order BO algorithm that depends only on first-order gradient information, requires no implicit differentiation, and is practical and efficient for large-scale non-convex functions in deep learning. We provide non-asymptotic convergence analysis of the proposed method to stationary points for non-convex objectives and present empirical results that show its superior practical performance. Bo Liu 0042, Mao Ye 0006, Stephen Wright, Peter Stone 0001, Qiang Liu 0001 |
NeurIPS | 4 |
| 2022 | Value Function Decomposition for Iterative Design of Reinforcement Learning AgentsabstractDesigning reinforcement learning (RL) agents is typically a difficult process that requires numerous design iterations. Learning can fail for a multitude of reasons and standard RL methods provide too few tools to provide insight into the exact cause. In this paper, we show how to integrate \textit{value decomposition} into a broad class of actor-critic algorithms and use it to assist in the iterative agent-design process. Value decomposition separates a reward function into distinct components and learns value estimates for each. These value estimates provide insight into an agent's learning and decision-making process and enable new training methods to mitigate common problems. As a demonstration, we introduce SAC-D, a variant of soft actor-critic (SAC) adapted for value decomposition. SAC-D maintains similar performance to SAC, while learning a larger set of value predictions. We also introduce decomposition-based tools that exploit this information, including a new reward \textit{influence} metric, which measures each reward component's effect on agent decision-making. Using these tools, we provide several demonstrations of decomposition's use in identifying and addressing problems in the design of both environments and agents. Value decomposition is broadly applicable and easy to incorporate into existing algorithms and workflows, making it a powerful tool in an RL practitioner's toolbox. James MacGlashan, Evan Archer, Alisa Devlic, Takuma Seno, Craig Sherstan, Peter R. Wurman, Peter Stone 0001 |
NeurIPS | 7 |
| 2022 | Towards a Real-Time, Low-Resource, End-to-End Object Detection Pipeline for Robot Soccer
Sai Kiran Narayanaswami, Mauricio Tec, Ishan Durugkar, Siddharth Desai, Bharath Masetty, Sanmit Narvekar, Peter Stone 0001 |
RoboCup | 7 |
| 2022 | Lucid dreaming for experience replay: refreshing past states with the current policy
Yunshu Du, Garrett Warnell, Assefaw Hadish Gebremedhin, Peter Stone 0001, Matthew E. Taylor |
Neural Comput. Appl. | 4 |
| 2021 | Demonstration of the EMPATHIC Framework for Task Learning from Implicit Human FeedbackabstractReactions such as gestures, facial expressions, and vocalizations are an abundant, naturally occurring channel of information that humans provide during interactions. An agent could leverage an understanding of such implicit human feedback to improve its task performance at no cost to the human. This approach contrasts with common agent teaching methods based on demonstrations, critiques, or other guidance that need to be attentively and intentionally provided. In this work, we demonstrate a novel data-driven framework for learning from implicit human feedback, EMPATHIC. This two-stage method consists of (1) mapping implicit human feedback to relevant task statistics such as reward, optimality, and advantage; and (2) using such a mapping to learn a task. We instantiate the first stage and three second-stage evaluations of the learned mapping. To do so, we collect a dataset of human facial reactions while participants observe an agent execute a sub-optimal policy for a prescribed training task. We train a deep neural network on this data and demonstrate its ability to (1) infer relative reward ranking of events in the training task from prerecorded human facial reactions; (2) improve the policy of an agent in the training task using live human facial reactions; and (3) transfer to a novel domain in which it evaluates robot manipulation trajectories. In the video, we focus on demonstrating the online learning capability of our instantiation of EMPATHIC. Yuchen Cui, Qiping Zhang, Sahil Jain, Alessandro Allievi, Peter Stone 0001, Scott Niekum, W. Bradley Knox |
AAAI | 5 |
| 2021 | Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksabstractIn continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an optimal policy in this setting typically requires a large amount of training experiences. Reward shaping is a common approach for incorporating domain knowledge into reinforcement learning in order to speed up convergence to an optimal policy. However, to the best of our knowledge, the theoretical properties of reward shaping have thus far only been established in the discounted setting. This paper presents the first reward shaping framework for average-reward learning and proves that, under standard assumptions, the optimal policy under the original reward function can be recovered. In order to avoid the need for manual construction of the shaping function, we introduce a method for utilizing domain knowledge expressed as a temporal logic formula. The formula is automatically translated to a shaping function that provides additional reward throughout the learning process. We evaluate the proposed method on three continuing tasks. In all cases, shaping speeds up the average-reward learning rate without any reduction in the performance of the learned policy compared to relevant baselines. Yuqian Jiang, Suda Bharadwaj, Bo Wu 0005, Rishi Shah, Ufuk Topcu, Peter Stone 0001 |
AAAI | 6 |
| 2021 | Goal Blending for Responsive Shared Autonomy in a Navigating VehicleabstractHuman-robot shared autonomy techniques for vehicle navigation hold promise for reducing a human driver’s workload, ensuring safety, and improving navigation efficiency. However, because typical techniques achieve these improvements by effectively removing human control at critical moments, these approaches often exhibit poor responsiveness to human commands—especially in cluttered environments. In this paper, we propose a novel goal-blending shared autonomy (GBSA) system, which aims to improve responsiveness in shared autonomy systems by blending human and robot input during the selection of local navigation goals as opposed to low-level motor (servo-level) commands. We validate the proposed approach by performing a human study involving an intelligent wheelchair and compare GBSA to a representative servo-level shared control system that uses a policy-blending approach. The results of both quantitative performance analysis and a subjective survey show that GBSA exhibits significantly better system responsiveness and induces higher user satisfaction than the existing approach. Yu-Sian Jiang, Garrett Warnell, Peter Stone 0001 |
AAAI | 3 |
| 2021 | Expected Value of Communication for Planning in Ad Hoc TeamworkabstractA desirable goal for autonomous agents is to be able to coordinate on the fly with previously unknown teammates. Known as “ad hoc teamwork”, enabling such a capability has been receiving increasing attention in the research community. One of the central challenges in ad hoc teamwork is quickly recognizing the current plans of other agents and planning accordingly. In this paper, we focus on the scenario in which teammates can communicate with one another, but only at a cost. Thus, they must carefully balance plan recognition based on observations vs. that based on communication. This paper proposes a new metric for evaluating how similar are two policies that a teammate may be following - the Expected Divergence Point (EDP). We then present a novel planning algorithm for ad hoc teamwork, determining which query to ask and planning accordingly. We demonstrate the effectiveness of this algorithm in a range of increasingly general communication in ad hoc teamwork problems. William Macke, Reuth Mirsky, Peter Stone 0001 |
AAAI | 3 |
| 2021 | Coach-Player Multi-agent Reinforcement Learning for Dynamic Team CompositionabstractIn real-world multi-agent systems, agents with different capabilities may join or leave without altering the team’s overarching goals. Coordinating teams with such dynamic composition is challenging: the optimal team strategy varies with the composition. We propose COPA, a coach-player framework to tackle this problem. We assume the coach has a global view of the environment and coordinates the players, who only have partial views, by distributing individual strategies. Specifically, we 1) adopt the attention mechanism for both the coach and the players; 2) propose a variational objective to regularize learning; and 3) design an adaptive communication method to let the coach decide when to communicate with the players. We validate our methods on a resource collection task, a rescue game, and the StarCraft micromanagement tasks. We demonstrate zero-shot generalization to new team compositions. Our method achieves comparable or better performance than the setting where all players have a full view of the environment. Moreover, we see that the performance remains high even when the coach communicates as little as 13% of the time using the adaptive communication strategy. Bo Liu 0042, Qiang Liu 0001, Peter Stone 0001, Animesh Garg, Yuke Zhu, Anima Anandkumar |
ICML | 3 |
| 2021 | Watch Where You're Going! Gaze and Head Orientation as Predictors for Social Robot NavigationabstractMobile robots deployed in human-populated environments must be able to safely and comfortably navigate in close proximity to people. Head orientation and gaze are both mechanisms which help people to interpret where other people intend to walk, which in turn enables them to coordinate their movement. Head orientation has previously been leveraged to develop classifiers which are able to predict the goal of a person’s walking motion. Gaze is believed to generally precede head orientation, with a person quickly moving their eyes to a target and then following it with a turn of their head. This study leverages state-of-the-art virtual reality technology to place participants into a simulated environment in which their gaze and motion can be observed. The results of this study indicate that position, velocity, head orientation, and gaze can all be used as predictive features of the goal of a person’s walking motion. The results also indicate that gaze both precedes head orientation and can be used to predict the goal of a person’s walking motion at a higher level of accuracy earlier in their walking trajectory. These findings can be leveraged in the design of social navigation systems for mobile robots. Blake Holman, Abrar Anwar, Mauricio Tec, Justin W. Hart, Peter Stone 0001 |
ICRA | 6 |
| 2021 | Efficient Real-Time Inference in Temporal Convolution NetworksabstractIt has been recently demonstrated that Temporal Convolution Networks (TCNs) provide state-of-the-art results in many problem domains where the input data is a time-series. TCNs typically incorporate information from a long history of inputs (the receptive field) into a single output using many convolution layers. Real-time inference using a trained TCN can be challenging on devices with limited compute and memory, especially if the receptive field is large. This paper introduces the RT-TCN algorithm that reuses the output of prior convolution operations to minimize the computational requirements and persistent memory footprint of a TCN during real-time inference. We also show that when a TCN is trained using time slices of the input time-series, it can be executed in realtime continually using RT-TCN. In addition, we provide TCN architecture guidelines that ensure that real-time inference can be performed within memory and computational constraints. Piyush Khandelwal, James MacGlashan, Peter R. Wurman, Peter Stone 0001 |
ICRA | 4 |
| 2021 | Towards Safe Motion Planning in Human Workspaces: A Robust Multi-agent ApproachabstractIt is becoming increasingly feasible for robots to share a workspace with humans. However, for them to do so safely while maintaining agile performance, they need the ability to smoothly handle the dynamics and uncertainty caused by human motions. Markov Decision Processes (MDPs) serve as a common framework to formulate robot planning problems. However, because of its single-agent formulation, such planner cannot account for human reaction when evaluating robot actions. The robot can thus suffer from unsafe motions and move in ways that are hard for nearby humans to understand. To resolve this, we instead model robot planning in human workspaces as a Stochastic Game, and contribute a robust planning algorithm, which enables the robot to account for its prediction errors in human responses to prevent collision, while not losing agility, opposed to traditional maximin optimization techniques, by applying maximin operation only at "critical states". We validate the approach under partial knowledge of pedestrian behaviors, and show that our approach encounters zero collision despite imperfect prediction, while improving path efficiency, compared to baselines. Shih-Yun Lo, Benito Fernandez, Peter Stone 0001, Andrea Thomaz |
ICRA | 3 |
| 2021 | APPLI: Adaptive Planner Parameter Learning From InterventionsabstractWhile classical autonomous navigation systems can typically move robots from one point to another safely and in a collision-free manner, these systems may fail or produce suboptimal behavior in certain scenarios. The current practice in such scenarios is to manually re-tune the system’s parameters, e.g. max speed, sampling rate, inflation radius, to optimize performance. This practice requires expert knowledge and may jeopardize performance in the originally good scenarios. Meanwhile, it is relatively easy for a human to identify those failure or suboptimal cases and provide a teleoperated intervention to correct the failure or suboptimal behavior. In this work, we seek to learn from those human interventions to improve navigation performance. In particular, we propose Adaptive Planner Parameter Learning from Interventions (APPLI), in which multiple sets of navigation parameters are learned during training and applied based on a confidence measure to the underlying navigation system during deployment. In our physical experiments, the robot achieves better performance compared to the planner with static default parameters, and even dynamic parameters learned from a full human demonstration. We also show APPLI’s generalizability in another unseen physical test course, and a suite of 300 simulated navigation environments. Xuesu Xiao, Bo Liu 0042, Garrett Warnell, Peter Stone 0001 |
ICRA | 5 |
| 2021 | Agile Robot Navigation through Hallucinated Learning and Sober DeploymentabstractLearning from Hallucination (LfH) is a recent machine learning paradigm for autonomous navigation, which uses training data collected in completely safe environments and adds numerous imaginary obstacles to make the environment densely constrained, to learn navigation planners that produce feasible navigation even in highly constrained (more dangerous) spaces. However, LfH requires hallucinating the robot perception during deployment to match with the hallucinated training data, which creates a need for sometimes-infeasible prior knowledge and tends to generate very conservative planning. In this work, we propose a new LfH paradigm that does not require runtime hallucination—a feature we call "sober deployment"—and can therefore adapt to more realistic navigation scenarios. This novel Hallucinated Learning and Sober Deployment (HLSD) paradigm is tested in a benchmark testbed of 300 simulated navigation environments with a wide range of difficulty levels, and in the real-world. In most cases, HLSD outperforms both the original LfH method and a classical navigation planner. Xuesu Xiao, Bo Liu 0042, Peter Stone 0001 |
ICRA | 3 |
| 2021 | APPLR: Adaptive Planner Parameter Learning from ReinforcementabstractClassical navigation systems typically operate using a fixed set of hand-picked parameters (e.g. maximum speed, sampling rate, inflation radius, etc.) and require heavy expert re-tuning in order to work in new environments. To mitigate this requirement, it has been proposed to learn parameters for different contexts in a new environment using human demonstrations collected via teleoperation. However, learning from human demonstration limits deployment to the training environment, and limits overall performance to that of a potentially-suboptimal demonstrator. In this paper, we introduce APPLR, Adaptive Planner Parameter Learning from Reinforcement, which allows existing navigation systems to adapt to new scenarios by using a parameter selection scheme discovered via reinforcement learning (RL) in a wide variety of simulation environments. We evaluate APPLR on a robot in both simulated and physical experiments, and show that it can outperform both a fixed set of hand-tuned parameters and also a dynamic parameter tuning scheme learned from human demonstration. Zifan Xu, Gauraang Dhamankar, Anirudh Nair, Xuesu Xiao, Garrett Warnell, Bo Liu 0042, Peter Stone 0001 |
ICRA | 8 |
| 2021 | A Scavenger Hunt for Service RobotsabstractCreating robots that can perform general-purpose service tasks in a human-populated environment has been a longstanding grand challenge for AI and Robotics research. One particularly valuable skill that is relevant to a wide variety of tasks is the ability to locate and retrieve objects upon request. This paper models this skill as a Scavenger Hunt (SH) game, which we formulate as a variation of the NP-hard stochastic traveling purchaser problem. In this problem, the goal is to find a set of objects as quickly as possible, given probability distributions of where they may be found. We investigate the performance of several solution algorithms for the SH problem, both in simulation and on a real mobile robot. We use Reinforcement Learning (RL) to train an agent to plan a minimal cost path, and show that the RL agent can outperform a range of heuristic algorithms, achieving near optimal performance. In order to stimulate research on this problem, we introduce a publicly available software stack and associated website that enable users to upload scavenger hunts which robots can download, perform, and learn from to continually improve their performance on future hunts. Harel Yedidsion, Jennifer Suriadinata, Zifan Xu, Stefan Debruyn, Peter Stone 0001 |
ICRA | 5 |
| 2021 | Capturing Skill State in Curriculum Learning for Human Skill AcquisitionabstractHumans learn complex motor skills with practice and training. Though the learning process is not fully understood, several theories from motor learning, neuroscience, education, and game design suggest that curriculum-based training may be the key to efficient skill acquisition. However, designing such a curriculum and understanding its effects on learning are challenging problems. In this paper, we define the Human-skill Curriculum Markov Decision Process (H-CMDP) to systematize the design of training protocols. We also identify a vocabulary of performance features to enable the approximation for a human’s skill level across a variety of cognitive and motor tasks. A novel task domain is introduced as a testbed to evaluate the effectiveness of our approach. Human subject experiments show that (1) participants can learn to improve their performance in tasks within this domain, (2) the learning is quantifiable via our performance features, and (3) the domain is flexible enough to create distinct levels of difficulty. The long-term goal of this work is to systematize the process of curriculum-based training toward the design of protocols for robot-mediated rehabilitation. Keya Ghonasgi, Reuth Mirsky, Sanmit Narvekar, Bharath Masetty, Adrian M. Haith, Peter Stone 0001, Ashish D. Deshpande |
IROS | 6 |
| 2021 | Team Orienteering Coverage Planning with Uncertain RewardabstractMany municipalities and large organizations have fleets of vehicles that need to be coordinated for tasks such as garbage collection or infrastructure inspection. Motivated by this need, this paper focuses on the common subproblem in which a team of vehicles needs to plan coordinated routes to patrol an area over iterations while minimizing temporally and spatially dependent costs. In particular, at a specific location (e.g., a vertex on a graph), we assume the cost accumulates over time and its growth rate is a random variable with a fixed but unknown mean, and the cost is reset to zero whenever any vehicle visits the vertex (representing the robot "servicing" the vertex). We formulate this problem in graph terminology and call it Team Orienteering Coverage Planning with Uncertain Reward (TOCPUR). We propose to solve TOCPUR by simultaneously estimating the accumulated cost at every vertex on the graph and solving a novel variant of the Team Orienteering Problem (TOP) iteratively, which we call the Team Orienteering Coverage Problem (TOCP). We provide the first mixed integer programming formulation for the TOCP, as a significant adaptation of the original TOP. We introduce a new benchmark consisting of hundreds of randomly generated graphs for comparing different methods. We show the proposed solution outperforms both the exact TOP solution and a greedy algorithm. In addition, we provide a demo of our method on a team of three physical robots in a real-world environment. The code is publicly available at https://github.com/Cranial-XIX/TOCPUR.git. Bo Liu 0042, Xuesu Xiao, Peter Stone 0001 |
IROS | 3 |
| 2021 | DEALIO: Data-Efficient Adversarial Learning for Imitation from ObservationabstractIn imitation learning from observation (IfO), a learning agent seeks to imitate a demonstrating agent using only observations of the demonstrated behavior without access to the control signals generated by the demonstrator. Recent methods based on adversarial imitation learning have led to state-of-the-art performance on IfO problems, but they typically suffer from high sample complexity due to a reliance on data-inefficient, model-free reinforcement learning algorithms. This issue makes them impractical to deploy in real-world settings, where gathering samples can incur high costs in terms of time, energy, and risk. In this work, we hypothesize that we can incorporate ideas from model-based reinforcement learning with adversarial methods for IfO in order to increase the data efficiency of these methods without sacrificing performance. Specifically, we consider time-varying linear Gaussian policies, and propose a method that integrates the linear-quadratic regulator with path integral policy improvement into an existing adversarial IfO framework. The result is a more data-efficient IfO algorithm with better performance, which we show empirically in four simulation domains: using far fewer interactions with the environment, the proposed method exhibits similar or better performance than the existing technique. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IROS | 3 |
| 2021 | From Agile Ground to Aerial Navigation: Learning from Learned HallucinationabstractThis paper presents a self-supervised Learning from Learned Hallucination (LfLH) method to learn fast and reactive motion planners for ground and aerial robots to navigate through highly constrained environments. The recent Learning from Hallucination (LfH) paradigm for autonomous navigation executes motion plans by random exploration in completely safe obstacle-free spaces, uses hand-crafted hallucination techniques to add imaginary obstacles to the robot’s perception, and then learns motion planners to navigate in realistic, highly-constrained, dangerous spaces. However, current hand-crafted hallucination techniques need to be tailored for specific robot types (e.g., a differential drive ground vehicle), and use approximations heavily dependent on certain assumptions (e.g., a short planning horizon). In this work, instead of manually designing hallucination functions, LfLH learns to hallucinate obstacle configurations, where the motion plans from random exploration in open space are optimal, in a self-supervised manner. LfLH is robust to different robot types and does not make assumptions about the planning horizon. Evaluated in both simulated and physical environments with a ground and an aerial robot, LfLH outperforms or performs comparably to previous hallucination approaches, along with sampling- and optimization-based classical methods. Xuesu Xiao, Alexander J. Nettekoven, Kadhiravan Umasankar, Anika Singh, Sriram Bommakanti, Ufuk Topcu, Peter Stone 0001 |
IROS | 8 |
| 2021 | Adversarial Intrinsic Motivation for Reinforcement LearningabstractLearning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning. In this paper, we investigate whether one such objective, the Wasserstein-1 distance between a policy's state visitation distribution and a target distribution, can be utilized effectively for reinforcement learning (RL) tasks. Specifically, this paper focuses on goal-conditioned reinforcement learning where the idealized (unachievable) target distribution has full measure at the goal. This paper introduces a quasimetric specific to Markov Decision Processes (MDPs) and uses this quasimetric to estimate the above Wasserstein-1 distance. It further shows that the policy that minimizes this Wasserstein-1 distance is the policy that reaches the goal in as few steps as possible. Our approach, termed Adversarial Intrinsic Motivation (AIM), estimates this Wasserstein-1 distance through its dual objective and uses it to compute a supplemental reward function. Our experiments show that this reward function changes smoothly with respect to transitions in the MDP and directs the agent's exploration to find the goal efficiently. Additionally, we combine AIM with Hindsight Experience Replay (HER) and show that the resulting algorithm accelerates learning significantly on several simulated robotics tasks when compared to other rewards that encourage exploration or accelerate learning. Ishan Durugkar, Mauricio Tec, Scott Niekum, Peter Stone 0001 |
NeurIPS | 4 |
| 2021 | Machine versus Human Attention in Deep Reinforcement Learning TasksabstractDeep reinforcement learning (RL) algorithms are powerful tools for solving visuomotor decision tasks. However, the trained models are often difficult to interpret, because they are represented as end-to-end deep neural networks. In this paper, we shed light on the inner workings of such trained models by analyzing the pixels that they attend to during task execution, and comparing them with the pixels attended to by humans executing the same tasks. To this end, we investigate the following two questions that, to the best of our knowledge, have not been previously studied. 1) How similar are the visual representations learned by RL agents and humans when performing the same task? and, 2) How do similarities and differences in these learned representations explain RL agents' performance on these tasks? Specifically, we compare the saliency maps of RL agents against visual attention models of human experts when learning to play Atari games. Further, we analyze how hyperparameters of the deep RL algorithm affect the learned representations and saliency maps of the trained agents. The insights provided have the potential to inform novel algorithms for closing the performance gap between human experts and RL agents. Sihang Guo, Bo Liu 0042, Dana H. Ballard, Mary M. Hayhoe, Peter Stone 0001 |
NeurIPS | 7 |
| 2021 | Conflict-Averse Gradient Descent for Multi-task learningabstractThe goal of multi-task learning is to enable more efficient learning than single task learning by sharing model structures for a diverse set of tasks. A standard multi-task learning objective is to minimize the average loss across all tasks. While straightforward, using this objective often results in much worse final performance for each task than learning them independently. A major challenge in optimizing a multi-task model is the conflicting gradients, where gradients of different task objectives are not well aligned so that following the average gradient direction can be detrimental to specific tasks' performance. Previous work has proposed several heuristics to manipulate the task gradients for mitigating this problem. But most of them lack convergence guarantee and/or could converge to any Pareto-stationary point.In this paper, we introduce Conflict-Averse Gradient descent (CAGrad) which minimizes the average loss function, while leveraging the worst local improvement of individual tasks to regularize the algorithm trajectory. CAGrad balances the objectives automatically and still provably converges to a minimum over the average loss. It includes the regular gradient descent (GD) and the multiple gradient descent algorithm (MGDA) in the multi-objective optimization (MOO) literature as special cases. On a series of challenging multi-task supervised learning and reinforcement learning tasks, CAGrad achieves improved performance over prior state-of-the-art multi-objective gradient manipulation methods. Bo Liu 0042, Xingchao Liu, Peter Stone 0001, Qiang Liu 0001 |
NeurIPS | 4 |
| 2021 | UT Austin Villa: RoboCup 2021 3D Simulation League Competition Champions
Patrick MacAlpine, Bo Liu 0042, William Macke, Caroline Wang, Peter Stone 0001 |
RoboCup | 5 |
| 2021 | Recent advances in leveraging human guidance for sequential decision-making tasks
Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 4 |
| 2021 | Agent-Based Markov Modeling for Improved COVID-19 Mitigation PoliciesabstractThe year 2020 saw the covid-19 virus lead to one of the worst global pandemics in history. As a result, governments around the world have been faced with the challenge of protecting public health while keeping the economy running to the greatest extent possible. Epidemiological models provide insight into the spread of these types of diseases and predict the effects of possible intervention policies. However, to date, even the most data-driven intervention policies rely on heuristics. In this paper, we study how reinforcement learning (RL) and Bayesian inference can be used to optimize mitigation policies that minimize economic impact without overwhelming hospital capacity. Our main contributions are (1) a novel agent-based pandemic simulator which, unlike traditional models, is able to model fine-grained interactions among people at specific locations in a community; (2) an RLbased methodology for optimizing fine-grained mitigation policies within this simulator; and (3) a Hidden Markov Model for predicting infected individuals based on partial observations regarding test results, presence of symptoms, and past physical contacts. This article is part of the special track on AI and COVID-19. Roberto Capobianco, Varun Raj Kompella, James Ault, Guni Sharon, Stacy Jong, Spencer J. Fox, Lauren Ancel Meyers, Peter R. Wurman, Peter Stone 0001 |
J. Artif. Intell. Res. | 9 |
| 2021 | Grounded action transformation for sim-to-real reinforcement learningabstractAbstract Reinforcement learning in simulation is a promising alternative to the prohibitive sample cost of reinforcement learning in the physical world. Unfortunately, policies learned in simulation often perform worse than hand-coded policies when applied on the target, physical system. Grounded simulation learning (gsl) is a general framework that promises to address this issue by altering the simulator to better match the real world (Farchy et al. 2013 in Proceedings of the 12th international conference on autonomous agents and multiagent systems (AAMAS)). This article introduces a new algorithm for gsl—Grounded Action Transformation (GAT)—and applies it to learning control policies for a humanoid robot. We evaluate our algorithm in controlled experiments where we show it to allow policies learned in simulation to transfer to the real world. We then apply our algorithm to learning a fast bipedal walk on a humanoid robot and demonstrate a 43.27% improvement in forward walk velocity compared to a state-of-the art hand-coded walk. This striking empirical success notwithstanding, further empirical analysis shows that gat may struggle when the real world has stochastic state transitions. To address this limitation we generalize gat to the stochasticgat (sgat) algorithm and empirically show that sgat leads to successful real world transfer in situations where gat may fail to find a good policy. Our results contribute to a deeper understanding of grounded simulation learning and demonstrate its effectiveness for applying reinforcement learning to learn robot control policies entirely in simulation. Josiah Hanna, Siddharth Desai, Haresh Karnan, Garrett Warnell, Peter Stone 0001 |
Mach. Learn. | 5 |
| 2021 | Importance sampling in reinforcement learning with an estimated behavior policyabstractAbstract In reinforcement learning, importance sampling is a widely used method for evaluating an expectation under the distribution of data of one policy when the data has in fact been generated by a different policy. Importance sampling requires computing the likelihood ratio between the action probabilities of a target policy and those of the data-producing behavior policy. In this article, we study importance sampling where the behavior policy action probabilities are replaced by their maximum likelihood estimate of these probabilities under the observed data. We show this general technique reduces variance due to sampling error in Monte Carlo style estimators. We introduce two novel estimators that use this technique to estimate expected values that arise in the RL literature. We find that these general estimators reduce the variance of Monte Carlo sampling methods, leading to faster learning for policy gradient algorithms and more accurate off-policy policy evaluation. We also provide theoretical analysis showing that our new estimators are consistent and have asymptotically lower variance than Monte Carlo estimators. Josiah Hanna, Scott Niekum, Peter Stone 0001 |
Mach. Learn. | 3 |
| 2020 | Reducing Sampling Error in Batch Temporal Difference LearningabstractTemporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This paper studies the use of TD(0), a canonical TD algorithm, to estimate the value function of a given policy from a batch of data. In this batch setting, we show that TD(0) may converge to an inaccurate value function because the update following an action is weighted according to the number of times that action occurred in the batch – not the true probability of the action under the given policy. To address this limitation, we introduce \emph{policy sampling error corrected}-TD(0) (PSEC-TD(0)). PSEC-TD(0) first estimates the empirical distribution of actions in each state in the batch and then uses importance sampling to correct for the mismatch between the empirical weighting and the correct weighting for updates following each action. We refine the concept of a certainty-equivalence estimate and argue that PSEC-TD(0) is a more data efficient estimator than TD(0) for a fixed batch of data. Finally, we conduct an empirical evaluation of PSEC-TD(0) on three batch value function learning tasks, with a hyperparameter sensitivity analysis, and show that PSEC-TD(0) produces value function estimates with lower mean squared error than TD(0). Brahma S. Pavse, Ishan Durugkar, Josiah Hanna, Peter Stone 0001 |
ICML | 4 |
| 2020 | Balancing Individual Preferences and Shared Objectives in Multiagent Reinforcement LearningabstractIn multiagent reinforcement learning scenarios, it is often the case that independent agents must jointly learn to perform a cooperative task. This paper focuses on such a scenario in which agents have individual preferences regarding how to accomplish the shared task. We consider a framework for this setting which balances individual preferences against task rewards using a linear mixing scheme. In our theoretical analysis we establish that agents can reach an equilibrium that leads to optimal shared task reward even when they consider individual preferences which aren't fully aligned with this task. We then empirically show, somewhat counter-intuitively, that there exist mixing schemes that outperform a purely task-oriented baseline. We further consider empirically how to optimize the mixing scheme. Ishan Durugkar, Elad Liebman, Peter Stone 0001 |
IJCAI | 3 |
| 2020 | A Penny for Your Thoughts: The Value of Communication in Ad Hoc TeamworkabstractIn ad hoc teamwork, multiple agents need to collaborate without having knowledge about their teammates or their plans a priori. A common assumption in this research area is that the agents cannot communicate. However, just as two random people may speak the same language, autonomous teammates may also happen to share a communication protocol. This paper considers how such a shared protocol can be leveraged, introducing a means to reason about Communication in Ad Hoc Teamwork (CAT). The goal of this work is enabling improved ad hoc teamwork by judiciously leveraging the ability of the team to communicate. We situate our study within a novel CAT scenario, involving tasks with multiple steps, where teammates' plans are unveiled over time. In this context, the paper proposes methods to reason about the timing and value of communication and introduces an algorithm for an ad hoc agent to leverage these methods. Finally, we introduces a new multiagent domain, the tool fetching domain, and we study how varying this domain's properties affects the usefulness of communication. Empirical results show the benefits of explicit reasoning about communication content and timing in ad hoc teamwork. Reuth Mirsky, William Macke, Harel Yedidsion, Peter Stone 0001 |
IJCAI | 5 |
| 2020 | Stochastic Grounded Action Transformation for Robot Learning in SimulationabstractRobot control policies learned in simulation do not often transfer well to the real world. Many existing solutions to this sim-to-real problem, such as the Grounded Action Transformation (GAT) algorithm, seek to correct for- or ground-these differences by matching the simulator to the real world. However, the efficacy of these approaches is limited if they do not explicitly account for stochasticity in the target environment. In this work, we analyze the problems associated with grounding a deterministic simulator in a stochastic real world environment, and we present examples where GAT fails to transfer a good policy due to stochastic transitions in the target domain. In response, we introduce the Stochastic Grounded Action Transformation (SGAT) algorithm, which models this stochasticity when grounding the simulator. We find experimentally-for both simulated and physical target domains-that SGAT can find policies that are robust to stochasticity in the target domain. Siddharth Desai, Haresh Karnan, Josiah Hanna, Garrett Warnell, Peter Stone 0001 |
IROS | 5 |
| 2020 | Reinforced Grounded Action Transformation for Sim-to-Real TransferabstractRobots can learn to do complex tasks in simulation, but often, learned behaviors fail to transfer well to the real world due to simulator imperfections (the "reality gap"). Some existing solutions to this sim-to-real problem, such as Grounded Action Transformation (gat), use a small amount of real-world experience to minimize the reality gap by "grounding" the simulator. While very effective in certain scenarios, gat is not robust on problems that use complex function approximation techniques to model a policy. In this paper, we introduce Reinforced Grounded Action Transformation (rgat), a new sim-to-real technique that uses Reinforcement Learning (RL) not only to update the target policy in simulation, but also to perform the grounding step itself. This novel formulation allows for end-to-end training during the grounding step, which, compared to gat, produces a better grounded simulator. Moreover, we show experimentally in several MuJoCo domains that our approach leads to successful transfer for policies modeled using neural networks. Haresh Karnan, Siddharth Desai, Josiah Hanna, Garrett Warnell, Peter Stone 0001 |
IROS | 5 |
| 2020 | Deep R-Learning for Continual Area SweepingabstractCoverage path planning is a well-studied problem in robotics in which a robot must plan a path that passes through every point in a given area repeatedly, usually with a uniform frequency. To address the scenario in which some points need to be visited more frequently than others, this problem has been extended to non-uniform coverage planning. This paper considers the variant of non-uniform coverage in which the robot does not know the distribution of relevant events beforehand and must nevertheless learn to maximize the rate of detecting events of interest. This continual area sweeping problem has been previously formalized in a way that makes strong assumptions about the environment, and to date only a greedy approach has been proposed. We generalize the continual area sweeping formulation to include fewer environmental constraints, and propose a novel approach based on reinforcement learning in a Semi-Markov Decision Process. This approach is evaluated in an abstract simulation and in a high fidelity Gazebo simulation. These evaluations show significant improvement upon the existing approach in general settings, which is especially relevant in the growing area of service robotics. We also present a video demonstration on a real service robot. Rishi Shah, Yuqian Jiang, Justin W. Hart, Peter Stone 0001 |
IROS | 4 |
| 2020 | An Imitation from Observation Approach to Transfer Learning with Dynamics MismatchabstractWe examine the problem of transferring a policy learned in a source environment to a target environment with different dynamics, particularly in the case where it is critical to reduce the amount of interaction with the target environment during learning. This problem is particularly important in sim-to-real transfer because simulators inevitably model real-world dynamics imperfectly. In this paper, we show that one existing solution to this transfer problem-- grounded action transformation --is closely related to the problem of imitation from observation (IfO): learning behaviors that mimic the observations of behavior demonstrations. After establishing this relationship, we hypothesize that recent state-of-the-art approaches from the IfO literature can be effectively repurposed for grounded transfer learning. To validate our hypothesis we derive a new algorithm -- generative adversarial reinforced action transformation (GARAT) -- based on adversarial imitation from observation techniques. We run experiments in several domains with mismatched dynamics, and find that agents trained with GARAT achieve higher returns in the target environment compared to existing black-box transfer methods. Siddharth Desai, Ishan Durugkar, Haresh Karnan, Garrett Warnell, Josiah Hanna, Peter Stone 0001 |
NeurIPS | 6 |
| 2020 | Firefly Neural Architecture Descent: a General Approach for Growing Neural NetworksabstractWe propose firefly neural architecture descent, a general framework for progressively and dynamically growing neural networks to jointly optimize the networks' parameters and architectures. Our method works in a steepest descent fashion, which iteratively finds the best network within a functional neighborhood of the original network that includes a diverse set of candidate network structures. By using Taylor approximation, the optimal network structure in the neighborhood can be found with a greedy selection procedure. We show that firefly descent can flexibly grow networks both wider and deeper, and can be applied to learn accurate but resource-efficient neural architectures that avoid catastrophic forgetting in continual learning. Empirically, firefly descent achieves promising results on both neural architecture search and continual learning. In particular, on a challenging continual image classification task, it learns networks that are smaller in size but have higher average accuracy than those learned by the state-of-the-art methods. Lemeng Wu, Bo Liu 0042, Peter Stone 0001, Qiang Liu 0001 |
NeurIPS | 3 |
| 2020 | Learning and Reasoning for Robot Dialog and Navigation TasksabstractReinforcement learning and probabilistic reasoning algorithms aim at learning from interaction experiences and reasoning with probabilistic contextual knowledge respectively.In this research, we develop algorithms for robot task completions, while looking into the complementary strengths of reinforcement learning and probabilistic reasoning techniques.The robots learn from trial-and-error experiences to augment their declarative knowledge base, and the augmented knowledge can be used for speeding up the learning process in potentially different tasks.We have implemented and evaluated the developed algorithms using mobile robots conducting dialog and navigation tasks.From the results, we see that our robot's performance can be improved by both reasoning with human knowledge and learning from task-completion experience.More interestingly, the robot was able to learn from navigation tasks to improve its dialog strategies. Keting Lu, Shiqi Zhang 0001, Peter Stone 0001 |
SIGdial | 3 |
| 2020 | Agents teaching agents: a survey on inter-agent transfer learning
Felipe Leno da Silva, Garrett Warnell, Anna Helena Reali Costa, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 4 |
| 2020 | Special issue on autonomous agents modelling other agents: Guest editorial
Stefano V. Albrecht, Peter Stone 0001, Michael P. Wellman |
Artif. Intell. | 2 |
| 2020 | The PETLON Algorithm to Plan Efficiently for Task-Level-Optimal NavigationabstractIntelligent mobile robots have recently become able to operate autonomously in large-scale indoor environments for extended periods of time. In this process, mobile robots need the capabilities of both task and motion planning. Task planning in such environments involves sequencing the robot’s high-level goals and subgoals, and typically requires reasoning about the locations of people, rooms, and objects in the environment, and their interactions to achieve a goal. One of the prerequisites for optimal task planning that is often overlooked is having an accurate estimate of the actual distance (or time) a robot needs to navigate from one location to another. State-of-the-art motion planning algorithms, though often computationally complex, are designed exactly for this purpose of finding routes through constrained spaces. In this article, we focus on integrating task and motion planning (TMP) to achieve task-level-optimal planning for robot navigation while maintaining manageable computational efficiency. To this end, we introduce TMP algorithm PETLON (Planning Efficiently for Task-Level-Optimal Navigation), including two configurations with different trade-offs over computational expenses between task and motion planning, for everyday service tasks using a mobile robot. Experiments have been conducted both in simulation and on a mobile robot using object delivery tasks in an indoor office environment. The key observation from the results is that PETLON is more efficient than a baseline approach that pre-computes motion costs of all possible navigation actions, while still producing plans that are optimal at the task level. We provide results with two different task planning paradigms in the implementation of PETLON, and offer TMP practitioners guidelines for the selection of task planners from an engineering perspective. Shih-Yun Lo, Shiqi Zhang 0001, Peter Stone 0001 |
J. Artif. Intell. Res. | 3 |
| 2020 | Jointly Improving Parsing and Perception for Natural Language Commands through Human-Robot DialogabstractIn this work, we present methods for using human-robot dialog to improve language understanding for a mobile robot agent. The agent parses natural language to underlying semantic meanings and uses robotic sensors to create multi-modal models of perceptual concepts like red and heavy. The agent can be used for showing navigation routes, delivering objects to people, and relocating objects from one location to another. We use dialog clari_cation questions both to understand commands and to generate additional parsing training data. The agent employs opportunistic active learning to select questions about how words relate to objects, improving its understanding of perceptual concepts. We evaluated this agent on Amazon Mechanical Turk. After training on data induced from conversations, the agent reduced the number of dialog questions it asked while receiving higher usability ratings. Additionally, we demonstrated the agent on a robotic platform, where it learned new perceptual concepts on the y while completing a real-world task. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
J. Artif. Intell. Res. | 8 |
| 2020 | Curriculum Learning for Reinforcement Learning Domains: A Framework and SurveyabstractReinforcement learning (RL) is a popular paradigm for addressing sequential decision tasks in which the agent has only limited environmental feedback. Despite many advances over the past three decades, learning in many domains still requires a large amount of interaction with the environment, which can be prohibitively expensive in realistic scenarios. To address this problem, transfer learning has been applied to reinforcement learning such that experience gained in one task can be leveraged when starting to learn the next, harder task. More recently, several lines of research have explored how tasks, or data samples themselves, can be sequenced into a curriculum for the purpose of learning a problem that may otherwise be too difficult to learn from scratch. In this article, we present a framework for curriculum learning (CL) in reinforcement learning, and use it to survey and classify existing CL methods in terms of their assumptions, capabilities, and goals. Finally, we use our framework to find open problems and suggest directions for future RL curriculum learning research. Sanmit Narvekar, Bei Peng 0001, Matteo Leonetti, Jivko Sinapov, Matthew E. Taylor, Peter Stone 0001 |
J. Mach. Learn. Res. | 6 |
| 2019 | Selecting Compliant Agents for Opt-in Micro-TollingabstractThis paper examines the impact of tolls on social welfare in the context of a transportation network in which only a portion of the agents are subject to tolls. More specifically, this paper addresses the question: which subset of agents provides the most system benefit if they are compliant with an approximate marginal cost tolling scheme? Since previous work suggests this problem is NP-hard, we examine a heuristic approach. Our experimental results on three real-world traffic scenarios suggest that evaluating the marginal impact of a given agent serves as a particularly strong heuristic for selecting an agent to be compliant. Results from using this heuristic for selecting 7.6% of the agents to be compliant achieved an increase of up to 10.9% in social welfare over not tolling at all. The presented heuristic approach and conclusions can help practitioners target specific agents to participate in an opt-in tolling scheme. Josiah Hanna, Guni Sharon, Stephen D. Boyles, Peter Stone 0001 |
AAAI | 4 |
| 2019 | Importance Sampling Policy Evaluation with an Estimated Behavior PolicyabstractWe consider the problem of off-policy evaluation in Markov decision processes. Off-policy evaluation is the task of evaluating the expected return of one policy with data generated by a different, behavior policy. Importance sampling is a technique for off-policy evaluation that re-weights off-policy returns to account for differences in the likelihood of the returns between the two policies. In this paper, we study importance sampling with an estimated behavior policy where the behavior policy estimate comes from the same set of data used to compute the importance sampling estimate. We find that this estimator often lowers the mean squared error of off-policy evaluation compared to importance sampling with the true behavior policy or using a behavior policy that is estimated from a separate data set. Intuitively, estimating the behavior policy in this way corrects for error due to sampling in the action-space. Our empirical results also extend to other popular variants of importance sampling and show that estimating a non-Markovian behavior policy can further lower large-sample mean squared error even when the true behavior policy is Markovian. Josiah Hanna, Scott Niekum, Peter Stone 0001 |
ICML | 3 |
| 2019 | Improving Grounded Natural Language Understanding through Human-Robot DialogabstractNatural language understanding for robotics can require substantial domain- and platform-specific engineering. For example, for mobile robots to pick-and-place objects in an environment to satisfy human commands, we can specify the language humans use to issue such commands, and connect concept words like red can to physical object properties. One way to alleviate this engineering for a new domain is to enable robots in human environments to adapt dynamically-continually learning new language constructions and perceptual concepts. In this work, we present an end-to-end pipeline for translating natural language commands to discrete robot actions, and use clarification dialogs to jointly improve language parsing and concept grounding. We train and evaluate this agent in a virtual setting on Amazon Mechanical Turk, and we transfer the learned agent to a physical robot platform to demonstrate it in the real world. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
ICRA | 8 |
| 2019 | Ad Hoc Teamwork With Behavior Switching AgentsabstractAs autonomous AI agents proliferate in the real world, they will increasingly need to cooperate with each other to achieve complex goals without always being able to coordinate in advance. This kind of cooperation, in which agents have to learn to cooperate on the fly, is called ad hoc teamwork. Many previous works investigating this setting assumed that teammates behave according to one of many predefined types that is fixed throughout the task. This assumption of stationarity in behaviors, is a strong assumption which cannot be guaranteed in many real-world settings. In this work, we relax this assumption and investigate settings in which teammates can change their types during the course of the task. This adds complexity to the planning problem as now an agent needs to recognize that a change has occurred in addition to figuring out what is the new type of the teammate it is interacting with. In this paper, we present a novel Convolutional-Neural-Network-based Change point Detection (CPD) algorithm for ad hoc teamwork. When evaluating our algorithm on the modified predator prey domain, we find that it outperforms existing Bayesian CPD algorithms. Manish Ravula, Shani Alkoby, Peter Stone 0001 |
IJCAI | 3 |
| 2019 | Imitation Learning from Video by Leveraging ProprioceptionabstractClassically, imitation learning algorithms have been developed for idealized situations, e.g., the demonstrations are often required to be collected in the exact same environment and usually include the demonstrator's actions. Recently, however, the research community has begun to address some of these shortcomings by offering algorithmic solutions that enable imitation learning from observation (IfO), e.g., learning to perform a task from visual demonstrations that may be in a different environment and do not include actions. Motivated by the fact that agents often also have access to their own internal states (i.e., proprioception), we propose and study an IfO algorithm that leverages this information in the policy learning process. The proposed architecture learns policies over proprioceptive state representations and compares the resulting trajectories visually to the demonstration data. We experimentally test the proposed technique on several MuJoCo domains and show that it outperforms other imitation from observation algorithms by a large margin. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IJCAI | 3 |
| 2019 | Recent Advances in Imitation Learning from ObservationabstractImitation learning is the process by which one agent tries to learn how to perform a certain task using information generated by another, often more-expert agent performing that same task. Conventionally, the imitator has access to both state and action information generated by an expert performing the task (e.g., the expert may provide a kinesthetic demonstration of object placement using a robotic arm). However, requiring the action information prevents imitation learning from a large number of existing valuable learning resources such as online videos of humans performing tasks. To overcome this issue, the specific problem of imitation from observation (IfO) has recently garnered a great deal of attention, in which the imitator only has access to the state information (e.g., video frames) generated by the expert. In this paper, we provide a literature review of methods developed for IfO, and then point out some open research problems and potential future work. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IJCAI | 3 |
| 2019 | Leveraging Human Guidance for Deep Reinforcement Learning TasksabstractReinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated using imitation learning, where the agent learns to imitate human demonstrated decisions. However, human guidance is not limited to the demonstrations. Other types of guidance could be more suitable for certain tasks and require less human effort. This survey provides a high-level overview of five recent learning frameworks that primarily rely on human guidance other than conventional, step-by-step action demonstrations. We review the motivation, assumption, and implementation of each framework. We then discuss possible future research directions. Faraz Torabi, Lin Guan 0003, Dana H. Ballard, Peter Stone 0001 |
IJCAI | 5 |
| 2019 | Task-Motion Planning with Reinforcement Learning for Adaptable Mobile Service RobotsabstractTask-motion planning (TMP) addresses the problem of efficiently generating executable and low-cost task plans in a discrete space such that the (initially unknown) action costs are determined by motion plans in a corresponding continuous space. A task-motion plan for a mobile service robot that behaves in a highly dynamic domain can be sensitive to domain uncertainty and changes, leading to suboptimal behaviors or execution failures. In this paper, we propose a novel framework, TMP-RL, which is an integration of TMP and reinforcement learning (RL), to solve the problem of robust TMP in dynamic and uncertain domains. The robot first generates a low-cost, feasible task-motion plan by iteratively planning in the discrete space and updating relevant action costs evaluated by the motion planner in continuous space. During execution, the robot learns via model-free RL to further improve its task-motion plans. RL enables adaptability to the current domain, but can be costly with regards to experience; using TMP, which does not rely on experience, can jump-start the learning process before executing in the real world. TMP-RL is evaluated in a mobile service robot domain where the robot navigates in an office area, showing significantly improved adaptability to unseen domain dynamics over TMP and task planning (TP)-RL methods. Yuqian Jiang, Fangkai Yang, Shiqi Zhang 0001, Peter Stone 0001 |
IROS | 4 |
| 2019 | UT Austin Villa: RoboCup 2019 3D Simulation League Competition and Technical Challenge Champions
Patrick MacAlpine, Faraz Torabi, Brahma S. Pavse, Peter Stone 0001 |
RoboCup | 4 |
| 2019 | Task planning in robotics: an empirical comparison of PDDL- and ASP-based systemsabstractRobots need task planning algorithms to sequence actions toward accomplishing goals that are impossible through individual actions. Off-the-shelf task planners can be used by intelligent robotics practitioners to solve a variety of planning problems. However, many different planners exist, each with different strengths and weaknesses, and there are no general rules for which planner would be best to apply to a given problem. In this study, we empirically compare the performance of state-of-the-art planners that use either the planning domain description language (PDDL) or answer set programming (ASP) as the underlying action language. PDDL is designed for task planning, and PDDL-based planners are widely used for a variety of planning problems. ASP is designed for knowledge-intensive reasoning, but can also be used to solve task planning problems. Given domain encodings that are as similar as possible, we find that PDDL-based planners perform better on problems with longer solutions, and ASP-based planners are better on tasks with a large number of objects or tasks in which complex reasoning is required to reason about action preconditions and effects. The resulting analysis can inform selection among general-purpose planning systems for particular robot task planning domains. Yuqian Jiang, Shiqi Zhang 0001, Piyush Khandelwal, Peter Stone 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | DyETC: Dynamic Electronic Toll Collection for Traffic Congestion AlleviationabstractTo alleviate traffic congestion in urban areas, electronic toll collection (ETC) systems are deployed all over the world. Despite the merits, tolls are usually pre-determined and fixed from day to day, which fail to consider traffic dynamics and thus have limited regulation effect when traffic conditions are abnormal. In this paper, we propose a novel dynamic ETC (DyETC) scheme which adjusts tolls to traffic conditions in realtime. The DyETC problem is formulated as a Markov decision process (MDP), the solution of which is very challenging due to its 1) multi-dimensional state space, 2) multi-dimensional, continuous and bounded action space, and 3) time-dependent state and action values. Due to the complexity of the formulated MDP, existing methods cannot be applied to our problem. Therefore, we develop a novel algorithm, PG-beta, which makes three improvements to traditional policy gradient method by proposing 1) time-dependent value and policy functions, 2) Beta distribution policy function and 3) state abstraction. Experimental results show that, compared with existing ETC schemes, DyETC increases traffic volume by around 8%, and reduces travel time by around 14:6% during rush hour. Considering the total traffic volume in a traffic network, this contributes to a substantial increase to social welfare. Haipeng Chen 0001, Bo An 0001, Guni Sharon, Josiah Hanna, Peter Stone 0001, Chunyan Miao, Yeng Chai Soh |
AAAI | 5 |
| 2018 | Adversarial Goal Generation for Intrinsic Motivation
Ishan Durugkar, Peter Stone 0001 |
AAAI | 2 |
| 2018 | Traffic Optimization for a Mixture of Self-Interested and Compliant AgentsabstractThis paper focuses on two commonly used path assignment policies for agents traversing a congested network: self-interested routing, and system-optimum routing. In the self-interested routing policy each agent selects a path that optimizes its own utility, while in the system-optimum routing, agents are assigned paths with the goal of maximizing system performance. This paper considers a scenario where a centralized network manager wishes to optimize utilities over all agents, i.e., implement a system-optimum routing policy. In many real-life scenarios, however, the system manager is unable to influence the route assignment of all agents due to limited influence on route choice decisions. Motivated by such scenarios, a computationally tractable method is presented that computes the minimal amount of agents that the system manager needs to influence (compliant agents) in order to achieve system optimal performance. Moreover, this methodology can also determine whether a given set of compliant agents is sufficient to achieve system optimum and compute the optimal route assignment for the compliant agents to do so. Experimental results are presented showing that in several large-scale, realistic traffic networks optimal flow can be achieved with as low as 13% of the agent being compliant and up to 54%. Guni Sharon, Michael Albert 0002, Tarun Rambha, Stephen D. Boyles, Peter Stone 0001 |
AAAI | 5 |
| 2018 | Guiding Exploratory Behaviors for Multi-Modal Grounding of Linguistic DescriptionsabstractA major goal of grounded language learning research is to enable robots to connect language predicates to a robot's physical interactive perception of the world. Coupling object exploratory behaviors such as grasping, lifting, and looking with multiple sensory modalities (e.g., audio, haptics, and vision) enables a robot to ground non-visual words like ``heavy'' as well as visual words like ``red''. A major limitation of existing approaches to multi-modal language grounding is that a robot has to exhaustively explore training objects with a variety of actions when learning a new such language predicate. This paper proposes a method for guiding a robot's behavioral exploration policy when learning a novel predicate based on known grounded predicates and the novel predicate's linguistic relationship to them. We demonstrate our approach on two datasets in which a robot explored large sets of objects and was tasked with learning to recognize whether novel words applied to those objects. Jesse Thomason, Jivko Sinapov, Raymond J. Mooney, Peter Stone 0001 |
AAAI | 4 |
| 2018 | Deep TAMER: Interactive Agent Shaping in High-Dimensional State SpacesabstractWhile recent advances in deep reinforcement learning have allowed autonomous learning agents to succeed at a variety of complex tasks, existing algorithms generally require a lot oftraining data. One way to increase the speed at which agent sare able to learn to perform tasks is by leveraging the input of human trainers. Although such input can take many forms, real-time, scalar-valued feedback is especially useful in situations where it proves difficult or impossible for humans to provide expert demonstrations. Previous approaches have shown the usefulness of human input provided in this fashion (e.g., the TAMER framework), but they have thus far not considered high-dimensional state spaces or employed the use of deep learning. In this paper, we do both: we propose DeepTAMER, an extension of the TAMER framework that leverages the representational power of deep neural networks inorder to learn complex tasks in just a short amount of time with a human trainer. We demonstrate Deep TAMER’s success by using it and just 15 minutes of human-provided feedback to train an agent that performs better than humans on the Atari game of Bowling - a task that has proven difficult for even state-of-the-art reinforcement learning methods. Garrett Warnell, Nicholas R. Waytowich, Vernon Lawhern, Peter Stone 0001 |
AAAI | 4 |
| 2018 | Learning a Policy for Opportunistic Active LearningabstractActive learning identifies data points to label that are expected to be the most useful in improving a supervised model.Opportunistic active learning incorporates active learning into interactive tasks that constrain possible queries during interactions.Prior work has shown that opportunistic active learning can be used to improve grounding of natural language descriptions in an interactive object retrieval task.In this work, we use reinforcement learning for such an object retrieval task, to learn a policy that effectively trades off task completion with model improvement that would benefit future tasks. Aishwarya Padmakumar, Peter Stone 0001, Raymond J. Mooney |
EMNLP | 2 |
| 2018 | Inferring User Intention using Gaze in VehiclesabstractMotivated by the desire to give vehicles better information about their drivers, we explore human intent inference in the setting of a human driver riding in a moving vehicle. Specifically, we consider scenarios in which the driver intends to go to or learn about a specific point of interest along the vehicle's route, and an autonomous system is tasked with inferring this point of interest using gaze cues. Because the scene under observation is highly dynamic --- both the background and objects in the scene move independently relative to the driver --- such scenarios are significantly different from the static scenes considered by most literature in the eye tracking community. In this paper, we provide a formulation for this new problem of determining a point of interest in a dynamic scenario. We design an experimental framework to systematically evaluate initial solutions to this novel problem, and we propose our own solution called dynamic interest point detection (DIPD). We experimentally demonstrate the success of DIPD when compared to baseline nearest-neighbor or filtering approaches. Yu-Sian Jiang, Garrett Warnell, Peter Stone 0001 |
ICMI | 3 |
| 2018 | Multi-modal Predicate Identification using Dynamically Learned Robot ControllersabstractIntelligent robots frequently need to explore the objects in their working environments. Modern sensors have enabled robots to learn object properties via perception of multiple modalities. However, object exploration in the real world poses a challenging trade-off between information gains and exploration action costs. Mixed observability Markov decision process (MOMDP) is a framework for planning under uncertainty, while accounting for both fully and partially observable components of the state. Robot perception frequently has to face such mixed observability. This work enables a robot equipped with an arm to dynamically construct query-oriented MOMDPs for multi-modal predicate identification (MPI) of objects. The robot's behavioral policy is learned from two datasets collected using real robots. Our approach enables a robot to explore object properties in a way that is significantly faster while improving accuracies in comparison to existing methods that rely on hand-coded exploration strategies. Saeid Amiri, Suhua Wei, Shiqi Zhang 0001, Jivko Sinapov, Jesse Thomason, Peter Stone 0001 |
IJCAI | 6 |
| 2018 | Behavioral Cloning from ObservationabstractHumans often learn how to perform tasks via imitation: they observe others perform a task, and then very quickly infer the appropriate actions to take based on their observations. While extending this paradigm to autonomous agents is a well-studied problem in general, there are two particular aspects that have largely been overlooked: (1) that the learning is done from observation only (i.e., without explicit action information), and (2) that the learning is typically done very quickly. In this work, we propose a two-phase, autonomous imitation learning technique called behavioral cloning from observation (BCO), that aims to provide improved performance with respect to both of these aspects. First, we allow the agent to acquire experience in a self-supervised fashion. This experience is used to develop a model which is then utilized to learn a particular task by observing an expert perform that task without the knowledge of the specific actions taken. We experimentally compare BCO to imitation learning methods, including the state-of-the-art, generative adversarial imitation learning (GAIL) technique, and we show comparable task performance in several different simulation domains while exhibiting increased learning speed after expert trajectories become available. Faraz Torabi, Garrett Warnell, Peter Stone 0001 |
IJCAI | 3 |
| 2018 | PRISM: Pose Registration for Integrated Semantic MappingabstractMany robotics applications involve navigating to positions specified in terms of their semantic significance. A robot operating in a hotel may need to deliver room service to a named room. In a hospital, it may need to deliver medication to a patient's room. The Building-Wide Intelligence Project at UT Austin has been developing a fleet of autonomous mobile robots, called BWIBots, which perform tasks in the computer science department. Tasks include guiding a person, delivering a message, or bringing an object to a location such as an office, lecture hall, or classroom. The process of constructing a map that a robot can use for navigation has been simplified by modern SLAM algorithms. The attachment of semantics to map data, however, remains a tedious manual process of labeling locations in otherwise automatically generated maps. This paper introduces a system called PRISM to automate a step in this process by enabling a robot to localize door signs - a semantic markup intended to aid the human occupants of a building - and to annotate these locations in its map. Justin W. Hart, Rishi Shah, Sean Kirmani, Nick Walker 0001, Kathryn Baldauf, Nathan John, Peter Stone 0001 |
IROS | 7 |
| 2018 | Passive Demonstrations of Light-Based Robot Signals for Improved Human InterpretabilityabstractWhen mobile robots navigate crowded, human-populated environments, the potential for conflict arises in the form of intersecting trajectories. This study investigates the use of light-emitting diodes (LEDs) arranged along the chassis of a robot in an arrangement similar to a turn signal on a car as a non-anthropomorphic, yet familiar signal to convey the intended path of a mobile service robot. We study the scenario of a human and a robot heading directly toward each other in a hallway, which may give rise to the familiar human experience in which both parties step to the right, then the left, then the right, continuing to block each other's paths until they are able to coordinate their movements and pass each other. We conducted a pilot study which revealed that people do not always interpret this signal as one may expect, which would be similar to how a car uses its turn signal. This motivated a 2 × 2 experiment in which the robot either does or does not use LEDs to indicate its intended direction of travel, and in which study participants either are able to or unable to witness the robot's “lane-changing” behavior further down the hallway prior to coming into direct proximal contact with the robot. The results demonstrate that exposing participants to the robot's use of the LED signal only once prior to passing each other in the hallway is sufficient to disambiguate its meaning to the user, and thus greatly enhances its utility in-situ, with no direct instruction or training to the user. These findings suggest a paradigm of passive demonstration of such signals in future applications. Rolando Fernandez, Nathan John, Sean Kirmani, Justin W. Hart, Jivko Sinapov, Peter Stone 0001 |
RO-MAN | 6 |
| 2018 | A Study of Human-Robot Copilot Systems for En-route Destination ChangingabstractIn this paper, we introduce the problem of enroute destination changing for a self-driving car, and we study the effectiveness of human-robot copilot systems as a solution. The copilot system is one in which the autonomous vehicle not only handles low-level vehicle control, but also continually monitors the intent of the human passenger in order to respond to dynamic changes in desired destination. We specifically consider a vehicle parking task, where the vehicle must respond to the user's intent to drive to and park next to a particular roadside sign board, and we study a copilot system that detects the passenger's intended destination based on gaze. We conduct a human study to investigate, in the context of our parking task, (a) if there is benefit in using a copilot system over manual driving, and (b) if copilot systems that use eye tracking to detect the intended destination have any benefit compared to those that use a more traditional, keyboard-based system. We find that the answers to both of these questions are affirmative: our copilot systems can complete the autonomous parking task more efficiently than human drivers can, and our copilot system that utilizes gaze information enjoys an increased success rate over one that utilizes typed input. Yu-Sian Jiang, Garrett Warnell, Eduardo Munera Sánchez, Peter Stone 0001 |
RO-MAN | 4 |
| 2018 | UT Austin Villa: RoboCup 2018 3D Simulation League Champions
Patrick MacAlpine, Faraz Torabi, Brahma S. Pavse, John Sigmon, Peter Stone 0001 |
RoboCup | 5 |
| 2018 | Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V. Albrecht, Peter Stone 0001 |
Artif. Intell. | 2 |
| 2018 | Overlapping layered learning
Patrick MacAlpine, Peter Stone 0001 |
Artif. Intell. | 2 |
| 2017 | Automated Design of Robust MechanismsabstractWe introduce a new class of mechanisms, robust mechanisms, that is an intermediary between ex-post mechanisms and Bayesian mechanisms. This new class of mechanisms allows the mechanism designer to incorporate imprecise estimates of the distribution over bidder valuations in a way that provides strong guarantees that the mechanism will perform at least as well as ex-post mechanisms, while in many cases performing better. We further extend this class to mechanisms that are with high probability incentive compatible and individually rational, ε-robust mechanisms. Using techniques from automated mechanism design and robust optimization, we provide an algorithm polynomial in the number of bidder types to design robust and ε-robust mechanisms. We show experimentally that this new class of mechanisms can significantly outperform traditional mechanism design techniques when the mechanism designer has an estimate of the distribution and the bidder’s valuation is correlated with an externally verifiable signal. Michael Albert 0002, Vincent Conitzer, Peter Stone 0001 |
AAAI | 3 |
| 2017 | Grounded Action Transformation for Robot Learning in SimulationabstractRobot learning in simulation is a promising alternative to the prohibitive sample cost of learning in the physical world. Unfortunately, policies learned in simulation often perform worse than hand-coded policies when applied on the physical robot. Grounded simulation learning (GSL) promises to address this issue by altering the simulator to better match the real world. This paper proposes a new algorithm for GSL -- Grounded Action Transformation -- and applies it to learning of humanoid bipedal locomotion. Our approach results in a 43.27% improvement in forward walk velocity compared to a state-of-the art hand-coded walk. We further evaluate our methodology in controlled experiments using a second, higher-fidelity simulator in place of the real world. Our results contribute to a deeper understanding of grounded simulation learning and demonstrate its effectiveness for learning robot control policies. Josiah Hanna, Peter Stone 0001 |
AAAI | 2 |
| 2017 | Grounded Action Transformation for Robot Learning in SimulationabstractRobot learning in simulation is a promising alternative to the prohibitive sample cost of learning in the physical world. Unfortunately, policies learned in simulation often perform worse than hand-coded policies when applied on the physical robot. This paper proposes a new algorithm for learning in simulation — Grounded Action Transformation — and applies it to learning of humanoid bipedal locomotion. Our approach results in a 43.27% improvement in forward walk velocity compared to a state-of-the art hand-coded walk. Josiah Hanna, Peter Stone 0001 |
AAAI | 2 |
| 2017 | Bootstrapping with Models: Confidence Intervals for Off-Policy EvaluationabstractFor an autonomous agent, executing a poor policy may be costly or even dangerous. For such agents, it is desirable to determine confidence interval lower bounds on the performance of any given policy without executing said policy. Current methods for exact high confidence off-policy evaluation that use importance sampling require a substantial amount of data to achieve a tight lower bound. Existing model-based methods only address the problem in discrete state spaces. Since exact bounds are intractable for many domains we trade off strict guarantees of safety for more data-efficient approximate bounds. In this context, we propose two bootstrapping off-policy evaluation methods which use learned MDP transition models in order to estimate lower confidence bounds on policy performance with limited data in both continuous and discrete state spaces. Since direct use of a model may introduce bias, we derive a theoretical upper bound on model bias for when the model transition function is estimated with i.i.d.~trajectories. This bound broadens our understanding of the conditions under which model-based methods have high bias. Finally, we empirically evaluate our proposed methods and analyze the settings in which different bootstrapping off-policy confidence interval methods succeed and fail. Josiah Hanna, Peter Stone 0001, Scott Niekum |
AAAI | 2 |
| 2017 | Designing Better Playlists with Monte Carlo Tree Search
Elad Liebman, Piyush Khandelwal, Maytal Saar-Tsechansky, Peter Stone 0001 |
AAAI | 4 |
| 2017 | Automatic Curriculum Graph Generation for Reinforcement Learning AgentsabstractIn recent years, research has shown that transfer learning methods can be leveraged to construct curricula that sequence a series of simpler tasks such that performance on a final target task is improved. A major limitation of existing approaches is that such curricula are handcrafted by humans that are typically domain experts. To address this limitation, we introduce a method to generate a curriculum based on task descriptors and a novel metric of transfer potential. Our method automatically generates a curriculum as a directed acyclic graph (as opposed to a linear sequence as done in existing work). Experiments in both discrete and continuous domains show that our method produces curricula that improve the agent's learning performance when compared to the baseline condition of learning on the target task from scratch. Maxwell Svetlik, Matteo Leonetti, Jivko Sinapov, Rishi Shah, Nick Walker 0001, Peter Stone 0001 |
AAAI | 6 |
| 2017 | Dynamically Constructed (PO)MDPs for Adaptive Robot PlanningabstractTo operate in human-robot coexisting environments, intelligent robots need to simultaneously reason with commonsense knowledge and plan under uncertainty. Markov decision processes (MDPs) and partially observable MDPs (POMDPs), are good at planning under uncertainty toward maximizing long-term rewards; P-LOG, a declarative programming language under Answer Set semantics, is strong in commonsense reasoning. In this paper, we present a novel algorithm called iCORPP to dynamically reason about, and construct (PO)MDPs using P-LOG. iCORPP successfully shields exogenous domain attributes from (PO)MDPs, which limits computational complexity and enables (PO)MDPs to adapt to the value changes these attributes produce. We conduct a number of experimental trials using two example problems in simulation and demonstrate iCORPP on a real robot. Results show significant improvements compared to competitive baselines. Shiqi Zhang 0001, Piyush Khandelwal, Peter Stone 0001 |
AAAI | 3 |
| 2017 | CC-Log: Drastically Reducing Storage Requirements for Robots Using Classification and Compression
Santiago Gonzalez, Vijay Chidambaram, Jivko Sinapov, Peter Stone 0001 |
HotStorage | 4 |
| 2017 | Data-Efficient Policy Evaluation Through Behavior Policy SearchabstractWe consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its performance. We show that the data collected from deploying a different policy, commonly called the behavior policy, can be used to produce unbiased estimates with lower mean squared error than this standard technique. We derive an analytic expression for the optimal behavior policy — the behavior policy that minimizes the mean squared error of the resulting estimates. Because this expression depends on terms that are unknown in practice, we propose a novel policy evaluation sub-problem, behavior policy search: searching for a behavior policy that reduces mean squared error. We present a behavior policy search algorithm and empirically demonstrate its effectiveness in lowering the mean squared error of policy performance estimates. Josiah Hanna, Philip S. Thomas, Peter Stone 0001, Scott Niekum |
ICML | 3 |
| 2017 | Autonomous Task Sequencing for Customized Curriculum Design in Reinforcement LearningabstractTransfer learning is a method where an agent reuses knowledge learned in a source task to improve learning on a target task. Recent work has shown that transfer learning can be extended to the idea of curriculum learning, where the agent incrementally accumulates knowledge over a sequence of tasks (i.e. a curriculum). In most existing work, such curricula have been constructed manually. Furthermore, they are fixed ahead of time, and do not adapt to the progress or abilities of the agent. In this paper, we formulate the design of a curriculum as a Markov Decision Process, which directly models the accumulation of knowledge as an agent interacts with tasks, and propose a method that approximates an execution of an optimal policy in this MDP to produce an agent-specific curriculum. We use our approach to automatically sequence tasks for 3 agents with varying sensing and action capabilities in an experimental domain, and show that our method produces curricula customized for each agent that improve performance relative to learning from scratch or using a different agent's curriculum. Sanmit Narvekar, Jivko Sinapov, Peter Stone 0001 |
IJCAI | 3 |
| 2017 | Leveraging commonsense reasoning and multimodal perception for robot spoken dialog systemsabstractProbabilistic graphical models, such as partially observable Markov decision processes (POMDPs), have been used in stochastic spoken dialog systems to handle the inherent uncertainty in speech recognition and language understanding. Such dialog systems suffer from the fact that only a relatively small number of domain variables are allowed in the model, so as to ensure the generation of good-quality dialog policies. At the same time, the non-language perception modalities on robots, such as vision-based facial expression recognition and Lidar-based distance detection, can hardly be integrated into this process. In this paper, we use a probabilistic commonsense reasoner to “guide” our POMDP-based dialog manager, and present a principled, multimodal dialog management (MDM) framework that allows the robot's dialog belief state to be seamlessly updated by both observations of human spoken language, and exogenous events such as the change of human facial expressions. The MDM approach has been implemented and evaluated both in simulation and on a real mobile robot using guidance tasks. Dongcai Lu, Shiqi Zhang 0001, Peter Stone 0001 |
IROS | 3 |
| 2017 | UT Austin Villa: RoboCup 2017 3D Simulation League Competition and Technical Challenges Champions
Patrick MacAlpine, Peter Stone 0001 |
RoboCup | 2 |
| 2017 | Fast and Precise Black and White Ball Detection for RoboCup Soccer
Jacob Menashe, Josh Kelle, Katie Genter, Josiah Hanna, Elad Liebman, Sanmit Narvekar, Peter Stone 0001 |
RoboCup | 8 |
| 2017 | Special issue on multiagent interaction without prior coordination: guest editorial
Stefano V. Albrecht, Somchaya Liemhetcharat, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2017 | Three years of the RoboCup standard platform league drop-in player competition - Creating and maintaining a large scale ad hoc teamwork robotics competition
Katie Genter, Tim Laue 0001, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2017 | Making friends on the fly: Cooperating with new teammates
Samuel Barrett, Avi Rosenfeld, Sarit Kraus, Peter Stone 0001 |
Artif. Intell. | 4 |
| 2017 | Intrinsically motivated model learning for developing curious robots
Todd Hester, Peter Stone 0001 |
Artif. Intell. | 2 |
| 2017 | Machine Learning Capabilities of a Simulated CerebellumabstractThis paper describes the learning and control capabilities of a biologically constrained bottom-up model of the mammalian cerebellum. Results are presented from six tasks: 1) eyelid conditioning; 2) pendulum balancing; 3) proportional-integral-derivative control; 4) robot balancing; 5) pattern recognition; and 6) MNIST handwritten digit recognition. These tasks span several paradigms of machine learning, including supervised learning, reinforcement learning, control, and pattern recognition. Results over these six domains indicate that the cerebellar simulation is capable of robustly identifying static input patterns even when randomized across the sensory apparatus. This capability allows the simulated cerebellum to perform several different supervised learning and control tasks. On the other hand, both reinforcement learning and temporal pattern recognition prove problematic due to the delayed nature of error signals and the simulator's inability to solve the credit assignment problem. These results are consistent with previous findings which hypothesize that in the human brain, the basal ganglia is responsible for reinforcement learning, while the cerebellum handles supervised learning. Matthew J. Hausknecht, Wen-Ke Li, Michael D. Mauk, Peter Stone 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | What's Hot at RoboCupabstractThe aim of this paper is to give an overview of the latest and most innovative developments at RoboCup, as well as highlighting some of the current and future challenges upon which today's RoboCup participants are focused. Peter Stone 0001 |
AAAI | 1 |
| 2016 | Autonomous Electricity Trading Using Time-of-Use Tariffs in a Competitive MarketabstractThis paper studies the impact of Time-Of-Use (TOU) tariffs in a competitive electricity market place. Specifically, it focuses on the question of how should an autonomous broker agent optimize TOU tariffs in a competitive retail market, and what is the impact of such tariffs on the economy. We formalize the problem of TOU tariff optimization and propose an algorithm for approximating its solution. We extensively experiment with our algorithm in a large-scale, detailed electricity retail markets simulation of the Power Trading Agent Competition (Power TAC) and: 1) find that our algorithm results in 15% peak-demand reduction, 2) find that its peak-flattening results in greater profit and/or profit-share for the broker and allows it to win against the 1st and 2nd place brokers from the Power TAC 2014 finals, and 3) analyze several economic implications of using TOU tariffs in competitive retail markets. Daniel Urieli, Peter Stone 0001 |
AAAI | 2 |
| 2016 | On the Analysis of Complex Backup Strategies in Monte Carlo Tree SearchabstractOver the past decade, Monte Carlo Tree Search (MCTS) and specifically Upper Confidence Bound in Trees (UCT) have proven to be quite effective in large probabilistic planning domains. In this paper, we focus on how values are backpropagated in the MCTS tree, and apply complex return strategies from the Reinforcement Learning (RL) literature to MCTS, producing 4 new MCTS variants. We demonstrate that in some probabilistic planning benchmarks from the International Planning Competition (IPC), selecting a MCTS variant with a backup strategy different from Monte Carlo averaging can lead to substantially better results. We also propose a hypothesis for why different backup strategies lead to different performance in particular environments, and manipulate a carefully structured grid-world domain to provide empirical evidence supporting our hypothesis. Piyush Khandelwal, Elad Liebman, Scott Niekum, Peter Stone 0001 |
ICML | 4 |
| 2016 | Learning to Order Objects Using Haptic and Proprioceptive Exploratory Behaviors
Jivko Sinapov, Priyanka Khante, Maxwell Svetlik, Peter Stone 0001 |
IJCAI | 4 |
| 2016 | Learning Multi-Modal Grounded Linguistic Semantics by Playing "I Spy"
Jesse Thomason, Jivko Sinapov, Maxwell Svetlik, Peter Stone 0001, Raymond J. Mooney |
IJCAI | 4 |
| 2016 | Robot Scavenger Hunt: A Standardized Framework for Evaluating Intelligent Mobile Robots
Shiqi Zhang 0001, Dongcai Lu, Peter Stone 0001 |
IJCAI | 4 |
| 2016 | UT Austin Villa RoboCup 3D Simulation Base Code Release
Patrick MacAlpine, Peter Stone 0001 |
RoboCup | 2 |
| 2016 | Prioritized Role Assignment for Marking
Patrick MacAlpine, Peter Stone 0001 |
RoboCup | 2 |
| 2016 | UT Austin Villa: RoboCup 2016 3D Simulation League Competition and Technical Challenges Champions
Patrick MacAlpine, Peter Stone 0001 |
RoboCup | 2 |
| 2016 | A synthesis of automated planning and reinforcement learning for efficient, robust decision-making
Matteo Leonetti, Luca Iocchi, Peter Stone 0001 |
Artif. Intell. | 3 |
| 2015 | Cooperating with Unknown Teammates in Complex Domains: A Robot Soccer Case Study of Ad Hoc TeamworkabstractMany scenarios require that robots work together as a team in order to effectively accomplish their tasks. However, pre-coordinating these teams may not always be possible given the growing number of companies and research labs creating these robots. Therefore, it is desirable for robots to be able to reason about ad hoc teamwork and adapt to new teammates on the fly. Past research on ad hoc teamwork has focused on relatively simple domains, but this paper demonstrates that agents can reason about ad hoc teamwork in complex scenarios. To handle these complex scenarios, we introduce a new algorithm, PLASTIC–Policy, that builds on an existing ad hoc teamwork approach. Specifically, PLASTIC– Policy learns policies to cooperate with past teammates and reuses these policies to quickly adapt to new teammates. This approach is tested in the 2D simulation soccer league of RoboCup using the half field offense task. Samuel Barrett, Peter Stone 0001 |
AAAI | 2 |
| 2015 | Placing Influencing Agents in a FlockabstractFlocking is a emergent behavior exhibited by many different animal species, including birds and fish. In our work we consider adding a small set of influencing agents, that are under our control, into a flock. Following ad hoc teamwork methodology, we assume that we are given knowledge of, but no direct control over, the rest of the flock. In our ongoing work highlighted in this abstract, we are specifically considering the problem of where to initially place influencing agents that we add to such a flock. We use these influencing agents to influence the flock to behave in a particular way - for example, to fly in a particular orientation or fly in a particular pattern such as to avoid an obstacle. Katie Genter, Peter Stone 0001 |
AAAI | 2 |
| 2015 | UT Austin Villa 2014: RoboCup 3D Simulation League Champion via Overlapping Layered LearningabstractLayered learning is a hierarchical machine learning paradigm that enables learning of complex behaviors by incrementally learning a series of sub-behaviors. A key feature of layered learning is that higher layers directly depend on the learned lower layers. In its original formulation, lower layers were frozen prior to learning higher layers. This paper considers an extension to the paradigm that allows learning certain behaviors independently, and then later stitching them together by learning at the "seams" where their influences overlap. The UT Austin Villa 2014 RoboCup 3D simulation team, using such overlapping layered learning, learned a total of 19 layered behaviors for a simulated soccer-playing robot, organized both in series and in parallel. To the best of our knowledge this is more than three times the number of layered behaviors in any prior layered learning system. Furthermore, the complete learning process is repeated on four different robot body types, showcasing its generality as a paradigm for efficient behavior learning. The resulting team won the RoboCup 2014 championship with an undefeated record, scoring 52 goals and conceding none. This paper includes a detailed experimental analysis of the team's performance and the overlapping layered learning approach that led to its success. Patrick MacAlpine, Mike Depinet, Peter Stone 0001 |
AAAI | 3 |
| 2015 | SCRAM: Scalable Collision-avoiding Role Assignment with Minimal-Makespan for Formational PositioningabstractTeams of mobile robots often need to divide up subtasks efficiently. In spatial domains, a key criterion for doing so may depend on distances between robots and the subtasks' locations. This paper considers a specific such criterion, namely how to assign interchangeable robots, represented as point masses, to a set of target goal locations within an open two dimensional space such that the makespan (time for all robots to reach their target locations) is minimized while also preventing collisions among robots. We present scaleable (computable in polynomial time) role assignment algorithms that we classify as being SCRAM (Scalable Collision-avoiding Role Assignment with Minimal-makespan). SCRAM role assignment algorithms use a graph theoretic approach to map agents to target goal locations such that our objectives for both minimizing the makespan and avoiding agent collisions are met. A system using SCRAM role assignment was originally designed to allow for decentralized coordination among physically realistic simulated humanoid soccer playing robots in the partially observable, non-deterministic, noisy, dynamic, and limited communication setting of the RoboCup 3D simulation league. In its current form, SCRAM role assignment generalizes well to many realistic and real-world multiagent systems, and scales to thousands of agents. Patrick MacAlpine, Eric Price 0001, Peter Stone 0001 |
AAAI | 3 |
| 2015 | CORPP: Commonsense Reasoning and Probabilistic Planning, as Applied to Dialog with a Mobile RobotabstractIn order to be fully robust and responsive to a dynamically changing real-world environment, intelligent robots will need to engage in a variety of simultaneous reasoning modalities. In particular, in this paper we consider their needs to i) reason with commonsense knowledge, ii) model their nondeterministic action outcomes and partial observability, and iii) plan toward maximizing long-term rewards. On one hand, Answer Set Programming (ASP) is good at representing and reasoning with commonsense and default knowledge, but is ill-equipped to plan under probabilistic uncertainty. On the other hand, Partially Observable Markov Decision Processes(POMDPs) are strong at planning under uncertainty toward maximizing long-term rewards, but are not designed to incorporate commonsense knowledge and inference. This paper introduces the CORPP algorithm which combines P-log,a probabilistic extension of ASP, with POMDPs to integrate commonsense reasoning with planning under uncertainty.Our approach is fully implemented and tested on a shopping request identification problem both in simulation and on a real robot. Compared with existing approaches using P-log or POMDPs individually, we observe significant improvements in both efficiency and accuracy. Shiqi Zhang 0001, Peter Stone 0001 |
AAAI | 2 |
| 2015 | When Security Games Go Green: Designing Defender Strategies to Prevent Poaching and Illegal Fishing
Fei Fang 0001, Peter Stone 0001, Milind Tambe |
IJCAI | 2 |
| 2015 | Learning to Interpret Natural Language Commands through Human-Robot Dialog
Jesse Thomason, Shiqi Zhang 0001, Raymond J. Mooney, Peter Stone 0001 |
IJCAI | 4 |
| 2015 | Benchmarking robot cooperation without pre-coordination in the RoboCup Standard Platform League drop-in player competitionabstractThe Standard Platform League is one of the main competitions of the annual RoboCup world championships. In this competition, teams of five humanoid robots play soccer against each other. In 2014, the league added a new sub-competition which serves as a testbed for cooperation without pre-coordination: the Drop-in Player Competition. Instead of homogeneous robot teams that are each programmed by the same people and hence implicitly pre-coordinated, this competition features ad hoc teams, i. e. teams that consist of robots originating from different RoboCup teams and that are each running different software. In this paper, we provide an overview of this competition, including its motivation and rules. We then present and analyze the results of the 2014 competition, which gathered robots from 23 teams, involved at least 50 human participants, and consisted of fifteen 20-minute games for a total playing time of 300 minutes. We also suggest improvements for future iterations, many of which will be evaluated at RoboCup 2015. Katie Genter, Tim Laue 0001, Peter Stone 0001 |
IROS | 3 |
| 2015 | Mobile Robot Planning Using Action Language BC with an Abstraction Hierarchy
Shiqi Zhang 0001, Fangkai Yang, Piyush Khandelwal, Peter Stone 0001 |
LPNMR | 4 |
| 2015 | A Study of Layered Learning Strategies Applied to Individual Behaviors in Robot SoccerabstractHierarchical task decomposition strategies allow robots and agents in general to address complex decision-making tasks. Layered learning is a hierarchical machine learning paradigm where a complex behavior is learned from a series of incrementally trained sub-tasks. This paper describes how layered learning can be applied to design individual behaviors in the context of soccer robotics. Three different layered learning strategies are implemented and analyzed using a ball-dribbling behavior as a case study. Performance indices for evaluating dribbling speed and ball-control are defined and measured. Experimental results validate the usefulness of the implemented layered learning strategies showing a trade-off between performance and learning speed. David Leonardo Leottau, Javier Ruiz-del-Solar, Patrick MacAlpine, Peter Stone 0001 |
RoboCup | 4 |
| 2015 | UT Austin Villa: RoboCup 2015 3D Simulation League Competition and Technical Challenges ChampionsabstractThe UT Austin Villa team, from the University of Texas at Austin, won the 2015 RoboCup 3D Simulation League, winning all 19 games that the team played. During the course of the competition the team scored 87 goals and conceded only 1. Additionally the team won the RoboCup 3D Simulation League technical challenge by winning each of a series of three league challenges: drop-in player, kick accuracy, and free challenge. This paper describes the changes and improvements made to the team between 2014 and 2015 that allowed it to win both the main competition and each of the league technical challenges. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Patrick MacAlpine, Josiah Hanna, Jason Liang, Peter Stone 0001 |
RoboCup | 4 |
| 2015 | Framing reinforcement learning from human reward: Reward positivity, temporal discounting, episodicity, and performance
W. Bradley Knox, Peter Stone 0001 |
Artif. Intell. | 2 |
| 2014 | TacTex'13: A Champion Adaptive Power Trading AgentabstractSustainable energy systems of the future will no longer be able to rely on the current paradigm that energy supply follows demand. Many of the renewable energy resources do not produce power on demand, and therefore there is a need for new market structures that motivate sustainable behaviors by participants. The Power Trading Agent Competition (Power TAC) is a new annual competition that focuses on the design and operation of future retail power markets, specifically in smart grid environments with renewable energy production, smart metering, and autonomous agents acting on behalf of customers and retailers. It uses a rich, open-source simulation platform that is based on real-world data and state-of-the-art customer models. Its purpose is to help researchers understand the dynamics of customer and retailer decision-making, as well as the robustness of proposed market designs. This paper introduces TacTex'13, the champion agent from the inaugural competition in 2013. TacTex'13 learns and adapts to the environment in which it operates, by heavily relying on reinforcement learning and prediction methods. This paper describes the constituent components of TacTex'13 and examines its success through analysis of competition results and subsequent controlled experiments. Daniel Urieli, Peter Stone 0001 |
AAAI | 2 |
| 2014 | Communicating with Unknown TeammatesabstractPast research has investigated a number of methods for coordinating teams of agents, but with the growing number of sources of agents, it is likely that agents will encounter teammates that do not share their coordination methods. Therefore, it is desirable for agents to adapt to these teammates, forming an effective ad hoc team. Past ad hoc teamwork research has focused on cases where the agents do not directly communicate. However when teammates do communicate, it can provide a valuable channel for coordination. Therefore, this paper tackles the problem of communication in ad hoc teams, introducing a minimal version of the multiagent, multiarmed bandit problem with limited communication between the agents. The theoretical results in this paper prove that this problem setting can be solved in polynomial time when the agent knows the set of possible teammates. Furthermore, the empirical results show that an agent can cooperate with a variety of teammates following unknown behaviors even when its models of these teammates are imperfect. Samuel Barrett, Noa Agmon, Noam Hazon, Sarit Kraus, Peter Stone 0001 |
ECAI | 5 |
| 2014 | The RoboCup 2013 drop-in player challenges: Experiments in ad hoc teamworkabstractAs the prevalence of autonomous agents grows, so does the number of interactions between these agents. Therefore, it is desirable for these agents to be capable of banding together with previously unknown teammates towards a common goal: to collaborate without pre-coordination. While past research on ad hoc teamwork has focused mainly on theoretical treatments and empirical studies in relatively simple domains, the long-term vision has been to enable robots and other autonomous agents to exhibit the sort of flexibility and adaptability on complex tasks that people do, for example when they play games of “pick-up” basketball or soccer. This paper introduces a series of pick-up robot soccer experiments that were carried out in three different leagues at the international RoboCup competition in 2013. In all cases, agents from different labs were put on teams with no pre-coordination. This paper introduces the structure of these experiments, describes the strategies used by UT Austin Villa in each challenge, and analyzes the results. The paper's main contribution is the introduction of a new large-scale ad hoc teamwork testbed that can serve as a starting point for future experimental ad hoc teamwork research. Patrick MacAlpine, Katie Genter, Samuel Barrett, Peter Stone 0001 |
IROS | 4 |
| 2014 | Keyframe Sampling, Optimization, and Behavior Integration: Towards Long-Distance Kicking in the RoboCup 3D Simulation League
Mike Depinet, Patrick MacAlpine, Peter Stone 0001 |
RoboCup | 3 |
| 2014 | UT Austin Villa: RoboCup 2014 3D Simulation League Competition and Technical Challenge Champions
Patrick MacAlpine, Mike Depinet, Jason Liang, Peter Stone 0001 |
RoboCup | 4 |
| 2014 | Multiagent learning in the presence of memory-bounded agents
Doran Chakraborty, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 2 |
| 2014 | A Neuroevolution Approach to General Atari Game PlayingabstractThis paper addresses the challenge of learning to play many different video games with little domain-specific knowledge. Specifically, it introduces a neuroevolution approach to general Atari 2600 game playing. Four neuroevolution algorithms were paired with three different state representations and evaluated on a set of 61 Atari games. The neuroevolution agents represent different points along the spectrum of algorithmic sophistication - including weight evolution on topologically fixed neural networks (conventional neuroevolution), covariance matrix adaptation evolution strategy (CMA-ES), neuroevolution of augmenting topologies (NEAT), and indirect network encoding (HyperNEAT). State representations include an object representation of the game screen, the raw pixels of the game screen, and seeded noise (a comparative baseline). Results indicate that direct-encoding methods work best on compact state representations while indirect-encoding methods (i.e., HyperNEAT) allow scaling to higher dimensional representations (i.e., the raw game screen). Previous approaches based on temporal-difference (TD) learning had trouble dealing with the large state spaces and sparse reward gradients often found in Atari games. Neuroevolution ameliorates these problems and evolved policies achieve state-of-the-art results, even surpassing human high scores on three games. These results suggest that neuroevolution is a promising approach to general video game playing (GVGP). Matthew J. Hausknecht, Joel Lehman, Risto Miikkulainen, Peter Stone 0001 |
IEEE Trans. Comput. Intell. AI Games | 4 |
| 2013 | Teamwork with Limited Knowledge of TeammatesabstractWhile great strides have been made in multiagent teamwork, existing approaches typically assume extensive information exists about teammates and how to coordinate actions. This paper addresses how robust teamwork can still be created even if limited or no information exists about a specific group of teammates, as in the ad hoc teamwork scenario. The main contribution of this paper is the first empirical evaluation of an agent cooperating with teammates not created by the authors, where the agent is not provided expert knowledge of its teammates. For this purpose, we develop a general-purpose teammate modeling method and test the resulting ad hoc team agent's ability to collaborate with more than 40 unknown teams of agents to accomplish a benchmark task. These agents were designed by people other than the authors without these designers planning for the ad hoc teamwork setting. A secondary contribution of the paper is a new transfer learning algorithm, TwoStageTransfer, that can improve results when the ad hoc team agent does have some limited observations of its current teammates. Samuel Barrett, Peter Stone 0001, Sarit Kraus, Avi Rosenfeld |
AAAI | 2 |
| 2013 | Learning non-myopically from human-generated rewardabstractRecent research has demonstrated that human-generated reward signals can be effectively used to train agents to perform a range of reinforcement learning tasks. Such tasks are either episodic - i.e., conducted in unconnected episodes of activity that often end in either goal or failure states - or continuing - i.e., indefinitely ongoing. Another point of difference is whether the learning agent highly discounts the value of future reward - a myopic agent - or conversely values future reward appreciably. In recent work, we found that previous approaches to learning from human reward all used myopic valuation [7]. This study additionally provided evidence for the desirability of myopic valuation in task domains that are both goal-based and episodic. W. Bradley Knox, Peter Stone 0001 |
IUI | 2 |
| 2013 | Model-Selection for Non-parametric Function Approximation in Continuous Control Problems: A Case Study in a Smart Energy System
Daniel Urieli, Peter Stone 0001 |
ECML/PKDD (1) | 2 |
| 2013 | The 2012 UT Austin Villa Code Release
Samuel Barrett, Katie Genter, Yuchen He 0003, Todd Hester, Piyush Khandelwal, Jacob Menashe, Peter Stone 0001 |
RoboCup | 7 |
| 2013 | The Open-Source TEXPLORE Code Release for Reinforcement Learning on Robots
Todd Hester, Peter Stone 0001 |
RoboCup | 2 |
| 2013 | Teaching and leading an ad hoc teammate: Collaboration without pre-coordination
Peter Stone 0001, Gal A. Kaminka, Sarit Kraus, Jeffrey S. Rosenschein, Noa Agmon |
Artif. Intell. | 1 |
| 2013 | TEXPLORE: real-time sample-efficient reinforcement learning for robots
Todd Hester, Peter Stone 0001 |
Mach. Learn. | 2 |
| 2013 | Using a million cell simulation of the cerebellum: Network scaling and task generality
Wen-Ke Li, Matthew J. Hausknecht, Peter Stone 0001, Michael D. Mauk |
Neural Networks | 3 |
| 2012 | Design and Optimization of an Omnidirectional Humanoid Walk: A Winning Approach at the RoboCup 2011 3D Simulation CompetitionabstractThis paper presents the design and learning architecture for an omnidirectional walk used by a humanoid robot soccer agent acting in the RoboCup 3D simulation environment. The walk, which was originally designed for and tested on an actual Nao robot before being employed in the 2011 RoboCup 3D simulation competition, was the crucial component in the UT Austin Villa team winning the competition in 2011. To the best of our knowledge, this is the first time that robot behavior has been conceived and constructed on a real robot for the end purpose of being used in simulation. The walk is based on a double linear inverted pendulum model, and multiple sets of its parameters are optimized via a novel framework. The framework optimizes parameters for different tasks in conjunction with one another, a little-understood problem with substantial practical significance. Detailed experiments show that the UT Austin Villa agent significantly outperforms all the other agents in the competition with the optimized walk being the key to its success. Patrick MacAlpine, Samuel Barrett, Daniel Urieli, Victor Vu, Peter Stone 0001 |
AAAI | 5 |
| 2012 | HyperNEAT-GGP: a hyperNEAT-based atari general game playerabstractThis paper considers the challenge of enabling agents to learn with as little domain-specific knowledge as possible. The main contribution is HyperNEAT-GGP, a HyperNEAT-based General Game Playing approach to Atari games. By leveraging the geometric regularities present in the Atari game screen, HyperNEAT effectively evolves policies for playing two different Atari games, Asterix and Freeway. Results show that HyperNEAT-GGP outperforms existing benchmarks on these games. HyperNEAT-GGP represents a step towards the ambitious goal of creating an agent capable of learning and seamlessly transitioning between many different tasks. Matthew J. Hausknecht, Piyush Khandelwal, Risto Miikkulainen, Peter Stone 0001 |
GECCO | 4 |
| 2012 | PAC Subset Selection in Stochastic Multi-armed Bandits
Shivaram Kalyanakrishnan, Ambuj Tewari, Peter Auer, Peter Stone 0001 |
ICML | 4 |
| 2012 | On coordination in practical multi-robot patrolabstractMulti-robot patrol is a fundamental application of multi-robot systems. While much theoretical work exists providing an understanding of the optimal patrol strategy for teams of coordinated homogeneous robots, little work exists on building and evaluating the performance of such systems for real. In this paper, we evaluate the performance of multirobot patrol in a practical outdoor distributed robotic system, and evaluate the effect of different coordination schemes on the performance of the robotic team. The multi-robot patrol algorithms evaluated vary in the level of robot coordination: no coordination, loose coordination, and tight coordination. In addition, we evaluate versions of these algorithms that distribute state information-either individual state, or entire team state (global-view state). Our experiments show that while tight coordination is theoretically optimal, it is not practical in practice. Instead, uncoordinated patrol performs best in terms of average waypoint visitation frequency, though loosely coordinated patrol that shares only individual state performed best in terms of worst-case frequency. Both are significantly better than a loosely coordinated algorithm based on sharing global-view state. We respond to this discrepancy between theory and practice, caused primarily by robot heterogeneity, by extending the theory to account for such heterogeneity, and find that the new theory accounts for the empirical results. Noa Agmon, Chien-Liang Fok, Yehuda Elmaliach, Peter Stone 0001, Christine Julien 0001, Sriram Vishwanath |
ICRA | 4 |
| 2012 | Setpoint scheduling for autonomous vehicle controllersabstractThis paper considers the problem of controlling an autonomous vehicle to arrive at a specific position on a road at a given time and velocity. This ability is particularly useful for a recently introduced autonomous intersection management protocol, called AIM, which has been shown to lead to lower delays than traffic signals and stop signs. Specifically, we introduce a setpoint scheduling algorithm for generating setpoints for the PID controllers for the brake and throttle actuators of an autonomous vehicle. The algorithm constructs a feasible setpoint schedule such that the vehicle arrives at the position at the correct time and velocity. Our experimental results show that the algorithm outperforms a heuristic-based setpoint scheduler that does not provide any guarantee about the arrival time and velocity. Tsz-Chiu Au, Michael J. Quinlan, Peter Stone 0001 |
ICRA | 3 |
| 2012 | RTMBA: A Real-Time Model-Based Reinforcement Learning Architecture for robot controlabstractReinforcement Learning (RL) is a paradigm for learning decision-making tasks that could enable robots to learn and adapt to their situation on-line. For an RL algorithm to be practical for robotic control tasks, it must learn in very few samples, while continually taking actions in real-time. Existing model-based RL methods learn in relatively few samples, but typically take too much time between each action for practical on-line learning. In this paper, we present a novel parallel architecture for model-based RL that runs in real-time by 1) taking advantage of sample-based approximate planning methods and 2) parallelizing the acting, model learning, and planning processes in a novel way such that the acting process is sufficiently fast for typical robot control cycles. We demonstrate that algorithms using this architecture perform nearly as well as methods using the typical sequential architecture when both are given unlimited time, and greatly out-perform these methods on tasks that require real-time actions such as controlling an autonomous vehicle. Todd Hester, Michael J. Quinlan, Peter Stone 0001 |
ICRA | 3 |
| 2012 | Evasion planning for autonomous vehicles at intersectionsabstractAutonomous intersection management (AIM) is a new intersection control protocol that exploits the capabilities of autonomous vehicles to control traffic at intersections in a way better than traffic signals and stop signs. A key assumption of this protocol is that vehicles can always follow their trajectories. But mechanical failures can occur in real life, causing vehicles to deviate from their trajectories. A previous approach for handling mechanical failure was to prevent vehicles from entering the intersection after the failure. However, this approach cannot prevent collisions among vehicles already in the intersection or too close to stop because (1) the lack of coordination among vehicles can cause collisions during the execution of evasive actions; and (2) the intersection may not have enough room for evasive actions. In this paper, we propose a preemptive approach that pre-computes evasion plans for several common types of mechanical failures before vehicles enter an intersection. This preemptive approach is necessary because there are situations in which vehicles cannot evade without pre-allocation of space for evasion. We present a modified AIM protocol and demonstrate the effectiveness of evasion plan execution on a miniature autonomous intersection testbed. Tsz-Chiu Au, Chien-Liang Fok, Sriram Vishwanath, Christine Julien 0001, Peter Stone 0001 |
IROS | 5 |
| 2012 | Video: RoboCup robot soccer history 1997 - 2011abstractRoboCup is an international initiative to foster inter-disciplinary research and education in robotics, artificial intelligence, computer science, and engineering. We focus on the challenges of multi-robot systems, where robots cooperate with each other and when needed with humans to achieve goals in complex and uncertain environments, such as robot soccer, as RoboCupSoccer, robot rescue, as RoboCupRescue, and the wide spectrum of robot applications in daily life, as RoboCup@Home. We also include sponsored demonstrations that explore possible new scientific challenges, such as collaborative logistics. Furthermore, we are committed to contribute to the education of children in robotics: RoboCupJunior provides an exciting introduction to science and engineering for children. Overall, RoboCup is a large vibrant community, composed of university faculty and student researchers and engineers, school teachers, children, and parents. RoboCup serves as a substrate to a wide variety of academic entreprises, ranging from courses and class projects to undergraduate, Masters, and PhD research theses. RoboCup has an international annual event consisting of robot competitions and a symposium. RoboCup has consistently grown, from a few hundred participants in 1997 to close to 3,000 in 2011. Manuela M. Veloso, Peter Stone 0001 |
IROS | 2 |
| 2012 | Reinforcement learning from human reward: Discounting in episodic tasksabstractSeveral studies have demonstrated that teaching agents by human-generated reward can be a powerful technique. However, the algorithmic space for learning from human reward has hitherto not been explored systematically. Using model-based reinforcement learning from human reward in goal-based, episodic tasks, we investigate how anticipated future rewards should be discounted to create behavior that performs well on the task that the human trainer intends to teach. We identify a “positive circuits” problem with low discounting (i.e., high discount factors) that arises from an observed bias among humans towards giving positive reward. Empirical analyses indicate that high discounting (i.e., low discount factors) of human reward is necessary in goal-based, episodic tasks and lend credence to the existence of the positive circuits problem. W. Bradley Knox, Peter Stone 0001 |
RO-MAN | 2 |
| 2012 | UT Austin Villa 2012: Standard Platform League World Champions
Samuel Barrett, Katie Genter, Yuchen He 0003, Todd Hester, Piyush Khandelwal, Jacob Menashe, Peter Stone 0001 |
RoboCup | 7 |
| 2012 | Positioning to Win: A Dynamic Role Assignment and Formation Positioning System
Patrick MacAlpine, Francisco Barrera, Peter Stone 0001 |
RoboCup | 3 |
| 2012 | UT Austin Villa: RoboCup 2012 3D Simulation League Champion
Patrick MacAlpine, Nick Collins, Adrian Lopez-Mobilia, Peter Stone 0001 |
RoboCup | 4 |
| 2011 | Multiagent Patrol Generalized to Complex Environmental ConditionsabstractThe problem of multiagent patrol has gained considerable attention during the past decade, with the immediate applicability of the problem being one of its main sources of interest. In this paper we concentrate on frequency-based patrol, in which the agents' goal is to optimize a frequency criterion, namely, minimizing the time between visits to a set of interest points. We consider multiagent patrol in environments with complex environmental conditions that affect the cost of traveling from one point to another. For example, in marine environments, the travel time of ships depends on parameters such as wind, water currents, and waves. We demonstrate that in such environments there is a need to consider a new multiagent patrol strategy which divides the given area into parts in which more than one agent is active, for improving frequency. We show that in general graphs this problem is intractable, therefore we focus on simplified (yet realistic) cyclic graphs with possible inner edges. Although the problem remains generally intractable in such graphs, we provide a heuristic algorithm that is shown to significantly improve point-visit frequency compared to other patrol strategies. For evaluation of our work we used a custom developed ship simulator that realistically models ship movement constraints such as engine force and drag and reaction of the ship to environmental changes. Noa Agmon, Daniel Urieli, Peter Stone 0001 |
AAAI | 3 |
| 2011 | Enforcing Liveness in Autonomous Traffic ManagementabstractLooking ahead to the time when autonomous cars will be common, Dresner and Stone proposed a multiagent systems-based intersection control protocol called Autonomous Intersection Management (AIM). They showed that by leveraging the capacities of autonomous vehicles it is possible to dramatically reduce the time wasted in traffic, and therefore also fuel consumption and air pollution. The proposed protocol, however, handles reservation requests one at a time and does not prioritize reservations according to their relative priorities and waiting times, causing potentially large inequalities in granting reservations. For example, at an intersection between a main street and an alley, vehicles from the alley can take an excessively long time to get reservations to enter the intersection, causing a waste of time and fuel. The same is true in a network of intersections, in which gridlock may occur and cause traffic congestion. In this paper, we introduce the batch processing of reservations in AIM to enforce liveness properties in intersections and analyze the conditions under which no vehicle will get stuck in traffic. Our experimental results show that our prioritizing schemes outperform previous intersection control protocols in unbalanced traffic. Tsz-Chiu Au, Neda Shahidi, Peter Stone 0001 |
AAAI | 3 |
| 2011 | Ad Hoc Teamwork in Variations of the Pursuit DomainabstractIn multiagent team settings, the agents are often given a protocol for coordinating their actions. When such a protocol is not available, agents must engage in ad hoc teamwork to effectively cooperate with one another. A fully general ad hoc team agent needs to be capable of collaborating with a wide range of potential teammates on a varying set of joint tasks. This paper extends previous research in a new direction with the introduction of an efficient method for reasoning about the value of information. Then, we show how previous theoretical results can aid ad hoc agents in a set of testbed pursuit domains. Samuel Barrett, Peter Stone 0001 |
AAAI | 2 |
| 2011 | Role-Based Ad Hoc TeamworkabstractAn ad hoc team setting is one in which teammates must work together to obtain a common goal, but without any prior agreement regarding how to work together. In this abstract we present a role-based approach for ad hoc teamwork, in which each teammate is inferred to be following a specialized role that accomplishes a specific task or exhibits a particular behavior. In such cases, the role an ad hoc agent should select depends both on its own capabilities and on the roles currently selected by the other team members. We present methods for evaluating the influence of the ad hoc agent's role selection on the team's utility and we examine empirically how to select the best suited method for role assignment in a complex environment. Finally, we show that an appropriate assignment method can be determined from a limited amount of data and used successfully in similar new tasks that the team has not encountered before. Katie Genter, Noa Agmon, Peter Stone 0001 |
AAAI | 3 |
| 2011 | Comparing Agents' Success against People in Security DomainsabstractThe interaction of people with autonomous agents has become increasingly prevalent. Some of these settings include security domains, where people can be characterized as uncooperative, hostile, manipulative, and tending to take advantage of the situation for their own needs. This makes it challenging to design proficient agents to interact with people in such environments. Evaluating the success of the agents automatically before evaluating them with people or deploying them could alleviate this challenge and result in better designed agents. In this paper we show how Peer Designed Agents (PDAs) -- computer agents developed by human subjects -- can be used as a method for evaluating autonomous agents in security domains. Such evaluation can reduce the effort and costs involved in evaluating autonomous agents interacting with people to validate their efficacy. Our experiments included more than 70 human subjects and 40 PDAs developed by students. The study provides empirical support that PDAs can be used to compare the proficiency of autonomous agents when matched with people in security domains. Raz Lin, Sarit Kraus, Noa Agmon, Samuel Barrett, Peter Stone 0001 |
AAAI | 5 |
| 2011 | On learning with imperfect representationsabstractIn this paper we present a perspective on the relationship between learning and representation in sequential decision making tasks. We undertake a brief survey of existing real-world applications, which demonstrates that the classical “tabular” representation seldom applies in practice. Specifically, several practical tasks suffer from state aliasing, and most demand some form of generalization and function approximation. Coping with these representational aspects thus becomes an important direction for furthering the advent of reinforcement learning in practice. The central thesis we present in this position paper is that in practice, learning methods specifically developed to work with imperfect representations are likely to perform better than those developed for perfect representations and then applied in imperfect-representation settings. We specify an evaluation criterion for learning methods in practice, and propose a framework for their synthesis. In particular, we highlight the degrees of “representational bias” prevalent in different learning methods. We reference a variety of relevant literature as a background for this introspective essay. Shivaram Kalyanakrishnan, Peter Stone 0001 |
ADPRL | 2 |
| 2011 | Protecting against evaluation overfitting in empirical reinforcement learningabstractEmpirical evaluations play an important role in machine learning. However, the usefulness of any evaluation depends on the empirical methodology employed. Designing good empirical methodologies is difficult in part because agents can overfit test evaluations and thereby obtain misleadingly high scores. We argue that reinforcement learning is particularly vulnerable to environment overfitting and propose as a remedy generalized methodologies, in which evaluations are based on multiple environments sampled from a distribution. In addition, we consider how to summarize performance when scores from different environments may not have commensurate values. Finally, we present proof-of-concept results demonstrating how these methodologies can validate an intuitively useful range-adaptive tile coding method. Shimon Whiteson, Brian Tanner, Matthew E. Taylor, Peter Stone 0001 |
ADPRL | 4 |
| 2011 | Structure Learning in Ergodic Factored MDPs without Knowledge of the Transition Function's In-Degree
Doran Chakraborty, Peter Stone 0001 |
ICML | 2 |
| 2011 | Autonomous Intersection Management: Multi-intersection optimizationabstractAdvances in autonomous vehicles and intelligent transportation systems indicate a rapidly approaching future in which intelligent vehicles will automatically handle the process of driving. However, increasing the efficiency of today's transportation infrastructure will require intelligent traffic control mechanisms that work hand in hand with intelligent vehicles. To this end, Dresner and Stone proposed a new intersection control mechanism called Autonomous Intersection Management (AIM) and showed in simulation that by studying the problem from a multiagent perspective, intersection control can be made more efficient than existing control mechanisms such as traffic signals and stop signs. We extend their study beyond the case of an individual intersection and examine the unique implications and abilities afforded by using AIM-based agents to control a network of interconnected intersections. We examine different navigation policies by which autonomous vehicles can dynamically alter their planned paths, observe an instance of Braess' paradox, and explore the new possibility of dynamically reversing the flow of traffic along lanes in response to minute-by-minute traffic conditions. Studying this multiagent system in simulation, we quantify the substantial improvements in efficiency imparted by these agent-based traffic control methods. Matthew J. Hausknecht, Tsz-Chiu Au, Peter Stone 0001 |
IROS | 3 |
| 2011 | WrightEagle and UT Austin Villa: RoboCup 2011 Simulation League Champions
Aijun Bai, Patrick MacAlpine, Daniel Urieli, Samuel Barrett, Peter Stone 0001 |
RoboCup | 6 |
| 2011 | A Low Cost Ground Truth Detection System for RoboCup Using the Kinect
Piyush Khandelwal, Peter Stone 0001 |
RoboCup | 2 |
| 2011 | Characterizing reinforcement learning methods through parameterized learning problems
Shivaram Kalyanakrishnan, Peter Stone 0001 |
Mach. Learn. | 2 |
| 2010 | Ad Hoc Autonomous Agent Teams: Collaboration without Pre-CoordinationabstractAs autonomous agents proliferate in the real world, both in software and robotic settings, they will increasingly need to band together for cooperative activities with previously unfamiliar teammates. In such ad hoc team settings, team strategies cannot be developed a priori. Rather, an agent must be prepared to cooperate with many types of teammates: it must collaborate without pre-coordination. This paper challenges the AI community to develop theory and to implement prototypes of ad hoc team agents. It defines the concept of ad hoc team agents, specifies an evaluation paradigm, and provides examples of possible theoretical and empirical approaches to challenge. The goal is to encourage progress towards this ambitious, newly realistic, and increasingly important research goal. Peter Stone 0001, Gal A. Kaminka, Sarit Kraus, Jeffrey S. Rosenschein |
AAAI | 1 |
| 2010 | Convergence, Targeted Optimality, and Safety in Multiagent Learning
Doran Chakraborty, Peter Stone 0001 |
ICML | 2 |
| 2010 | Efficient Selection of Multiple Bandit Arms: Theory and Practice
Shivaram Kalyanakrishnan, Peter Stone 0001 |
ICML | 2 |
| 2010 | Boosting for Regression Transfer
David Pardoe, Peter Stone 0001 |
ICML | 2 |
| 2010 | Generalized model learning for Reinforcement Learning on a humanoid robotabstractReinforcement learning (RL) algorithms have long been promising methods for enabling an autonomous robot to improve its behavior on sequential decision-making tasks. The obvious enticement is that the robot should be able to improve its own behavior without the need for detailed step-by-step programming. However, for RL to reach its full potential, the algorithms must be sample efficient: they must learn competent behavior from very few real-world trials. From this perspective, model-based methods, which use experiential data more efficiently than model-free approaches, are appealing. But they often require exhaustive exploration to learn an accurate model of the domain. In this paper, we present an algorithm, Reinforcement Learning with Decision Trees (RL-DT), that uses decision trees to learn the model by generalizing the relative effect of actions across states. The agent explores the environment until it believes it has a reasonable policy. The combination of the learning approach with the targeted exploration policy enables fast learning of the model. We compare RL-DT against standard model-free and model-based learning methods, and demonstrate its effectiveness on an Aldebaran Nao humanoid robot scoring goals in a penalty kick scenario. Todd Hester, Michael J. Quinlan, Peter Stone 0001 |
ICRA | 3 |
| 2010 | Bringing simulation to life: A mixed reality autonomous intersectionabstractFully autonomous vehicles are technologically feasible with the current generation of hardware, as demonstrated by recent robot car competitions. Dresner and Stone proposed a new intersection control protocol called Autonomous Intersection Management (AIM) and showed that with autonomous vehicles it is possible to make intersection control much more efficient than the traditional control mechanisms such as traffic signals and stop signs. The protocol, however, has only been tested in simulation and has not been evaluated with real autonomous vehicles. To realistically test the protocol, we implemented a mixed reality platform on which an autonomous vehicle can interact with multiple virtual vehicles in a simulation at a real intersection in real time. From this platform we validated realistic parameters for our autonomous vehicle to safely traverse an intersection in AIM. We present several techniques to improve efficiency and show that the AIM protocol can still outperform traffic signals and stop signs even if the cars are not as precisely controllable as has been assumed in previous studies. Michael J. Quinlan, Tsz-Chiu Au, Jesse Zhu, Nicolae Stiurca, Peter Stone 0001 |
IROS | 5 |
| 2010 | Gaussian Processes for Sample Efficient Reinforcement Learning with RMAX-Like Exploration
Tobias Jung 0001, Peter Stone 0001 |
ECML/PKDD (1) | 2 |
| 2010 | Learning Powerful Kicks on the Aibo ERS-7: The Quest for a Striker
Matthew J. Hausknecht, Peter Stone 0001 |
RoboCup | 2 |
| 2010 | Critical factors in the empirical performance of temporal difference and evolutionary methods for reinforcement learningabstractTemporal difference and evolutionary methods are two of the most common approaches to solving reinforcement learning problems. However, there is little consensus on their relative merits and there have been few empirical studies that directly compare their performance. This article aims to address this shortcoming by presenting results of empirical comparisons between Sarsa and NEAT, two representative methods, in mountain car and keepaway, two benchmark reinforcement learning tasks. In each task, the methods are evaluated in combination with both linear and nonlinear representations to determine their best configurations. In addition, this article tests two specific hypotheses about the critical factors contributing to these methods’ relative performance: (1) that sensor noise reduces the final performance of Sarsa more than that of NEAT, because Sarsa’s learning updates are not reliable in the absence of the Markov property and (2) that stochasticity, by introducing noise in fitness estimates, reduces the learning speed of NEAT more than that of Sarsa. Experiments in variations of mountain car and keepaway designed to isolate these factors confirm both these hypotheses. Shimon Whiteson, Matthew E. Taylor, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2010 | Adaptive Auction Mechanism Design and the Incorporation of Prior KnowledgeabstractElectronic auction markets are economic information systems that facilitate transactions between buyers and sellers. Whereas auction design has traditionally been an analytic process that relies on theory-driven assumptions such as bidders' rationality, bidders often exhibit unknown and variable behaviors. In this paper we present a data-driven adaptive auction mechanism that capitalizes on key properties of electronic auction markets, such as the large transaction volume, access to information, and the ability to dynamically alter the mechanism's design to acquire information about the benefits from different designs and adapt the auction mechanism online in response to actual bidders' behaviors. Our auction mechanism does not require an explicit representation of bidder behavior to infer about design profitability—a key limitation of prior approaches when they address complex auction settings. Our adaptive mechanism can also incorporate prior general knowledge of bidder behavior to enhance the search for effective designs. The data-driven adaptation and the capacity to use prior knowledge render our mechanisms particularly useful when there is uncertainty regarding bidders' behaviors or when bidders' behaviors change over time. Extensive empirical evaluations demonstrate that the adaptive mechanism outperforms any single fixed mechanism design under a variety of settings, including when bidders' strategies evolve in response to the seller's adaptation; our mechanism's performance is also more robust than that of alternatives when prior general information about bidders' behaviors differs from the encountered behaviors. David Pardoe, Peter Stone 0001, Maytal Saar-Tsechansky, Tayfun Keskin, Kerem Tomak |
INFORMS J. Comput. | 2 |
| 2009 | Improving particle filter performance using SSE instructionsabstractRobotics researchers are often faced with real-time constraints, and for that reason algorithmic and implementation-level optimization can dramatically increase the overall performance of a robot. In this paper we illustrate how a substantial run-time gain can be achieved by taking advantage of the extended instruction sets found in modern processors, in particular the SSE1 and SSE2 instruction sets. We present an SSE version of Monte Carlo Localization that results in an impressive 9x speedup over an optimized scalar implementation. In the process, we discuss SSE implementations of atan, atan2 and exp that achieve up to a 4x speedup in these mathematical operations alone. Peter Djeu, Michael J. Quinlan, Peter Stone 0001 |
IROS | 3 |
| 2009 | Interactively shaping agents via human reinforcement: the TAMER frameworkabstractAs computational learning agents move into domains that incur real costs (e.g., autonomous driving or financial investment), it will be necessary to learn good policies without numerous high-cost learning trials. One promising approach to reducing sample complexity of learning a task is knowledge transfer from humans to agents. Ideally, methods of transfer should be accessible to anyone with task knowledge, regardless of that person's expertise in programming and AI. This paper focuses on allowing a human trainer to interactively shape an agent's policy via reinforcement signals. Specifically, the paper introduces "Training an Agent Manually via Evaluative Reinforcement," or TAMER, a framework that enables such shaping. Differing from previous approaches to interactive shaping, a TAMER agent models the human's reinforcement and exploits its model by choosing actions expected to be most highly reinforced. Results from two domains demonstrate that lay users can train TAMER agents without defining an environmental reward function (as in an MDP) and indicate that human training within the TAMER framework can reduce sample complexity over autonomous learning algorithms. W. Bradley Knox, Peter Stone 0001 |
K-CAP | 2 |
| 2009 | Compositional Models for Reinforcement Learning
Nicholas K. Jong, Peter Stone 0001 |
ECML/PKDD (1) | 2 |
| 2009 | Feature Selection for Value Function Approximation Using Bayesian Model Selection
Tobias Jung 0001, Peter Stone 0001 |
ECML/PKDD (1) | 2 |
| 2009 | Three Humanoid Soccer Platforms: Comparison and Synthesis
Shivaram Kalyanakrishnan, Todd Hester, Michael J. Quinlan, Yinon Bentor, Peter Stone 0001 |
RoboCup | 5 |
| 2009 | Learning Complementary Multiagent Behaviors: A Case Study
Shivaram Kalyanakrishnan, Peter Stone 0001 |
RoboCup | 2 |
| 2009 | Transfer Learning for Reinforcement Learning Domains: A Survey
Matthew E. Taylor, Peter Stone 0001 |
J. Mach. Learn. Res. | 2 |
| 2008 | Online kernel selection for Bayesian reinforcement learningabstractKernel-based Bayesian methods for Reinforcement Learning (RL) such as Gaussian Process Temporal Difference (GPTD) are particularly promising because they rigorously treat uncertainty in the value function and make it easy to specify prior knowledge. However, the choice of prior distribution significantly affects the empirical performance of the learning agent, and little work has been done extending existing methods for prior model selection to the online setting. This paper develops Replacing-Kernel RL, an online model selection method for GPTD using sequential Monte-Carlo methods. Replacing-Kernel RL is compared to standard GPTD and tile-coding on several RL domains, and is shown to yield significantly better asymptotic performance for many different kernel families. Furthermore, the resulting kernels capture an intuitively useful notion of prior state covariance that may nevertheless be difficult to capture manually. Joseph Reisinger, Peter Stone 0001, Risto Miikkulainen |
ICML | 2 |
| 2008 | Negative information and line observations for Monte Carlo localizationabstractLocalization is a very important problem in robotics and is critical to many tasks performed on a mobile robot. In order to localize well in environments with few landmarks, a robot must make full use of all the information provided to it. This paper moves towards this goal by studying the effects of incorporating line observations and negative information into the localization algorithm. We extend the general Monte Carlo localization algorithm to utilize observations of lines such as carpet edges. We also make use of the information available when the robot expects to see a landmark but does not, by incorporating negative information into the algorithm. We compare our implementations of these ideas to previous similar approaches and demonstrate the effectiveness of these improvements through localization experiments performed both on a Sony AIBO ERS-7 robot and in simulation. Todd Hester, Peter Stone 0001 |
ICRA | 2 |
| 2008 | Person recognition on a Segway Robot: A video of UT Austin Villa Robocup@Home 2007 finals demonstrationabstractThis video shows a Segway robot from the University of Texas at Austin competing in the finals of the 2007 Robocup @Home competition, which featured home assistant robots performing various challenging tasks. This demonstration combines a few tasks which will likely be performed by a future home assistant robot. The robot learns a human's appearance, follows the human with his back turned, distinguishes the human from a similarly clothed stranger, and adapts when it notices that the human has changed his clothing. For this task, we introduce a novel two-classifier architecture, using the subject's face as a primary identifying characteristic and his shirt as a secondary characteristic. W. Bradley Knox, Peter Stone 0001 |
ICRA | 3 |
| 2008 | Person tracking on a mobile robot with heterogeneous inter-characteristic feedbackabstractFor a mobile robot that interacts with humans such as a home assistant or a tour guide robot, tracking a particular person among multiple persons is a fundamental, yet challenging task. Uniquely identifying characteristics such as a person's face, may not be visible consistently enough to be used as the sole form of identification. Rather, it may be useful to also track more frequently visible, but perhaps less uniquely identifying characteristics such as a person's clothes. After learning various characteristics of a person, the tracking system is required to autonomously update itself with additional training data, since the learned features may change over space and time due to the mobile nature of the robot. In this paper, we introduce a novel algorithm for merging multiple, heterogeneous sub-classifiers designed to track and associate different characteristics of a person being tracked. These heterogeneous classifiers give feedback to each other by identifying additional online training data for one another, thus improving the performance of each classifier and the accuracy of the overall system. Our algorithm has been fully implemented and tested on a Segway base. Peter Stone 0001 |
ICRA | 2 |
| 2008 | Maximum likelihood estimation of sensor and action model functions on a mobile robotabstractIn order for a mobile robot to accurately interpret its sensations and predict the effects of its actions, it must have accurate models of its sensors and actuators. These models are typically tuned manually, a brittle and laborious process. Autonomous model learning is a promising alternative to manual calibration, but previous work has assumed the presence of an accurate action or sensor model in order to train the other model. This paper presents an adaptation of the Expectation-Maximization (EM) algorithm to enable a mobile robot to learn both its action and sensor model functions, starting without an accurate version of either. The resulting algorithm is validated experimentally both on a Sony Aibo ERS-7 robot and in simulation. Daniel Stronger, Peter Stone 0001 |
ICRA | 2 |
| 2008 | Online Multiagent Learning against Memory Bounded Adversaries
Doran Chakraborty, Peter Stone 0001 |
ECML/PKDD (1) | 2 |
| 2008 | Transferring Instances for Model-Based Reinforcement Learning
Matthew E. Taylor, Nicholas K. Jong, Peter Stone 0001 |
ECML/PKDD (2) | 3 |
| 2008 | Domestic Interaction on a Segway Base
W. Bradley Knox, Peter Stone 0001 |
RoboCup | 3 |
| 2008 | A Multiagent Approach to Autonomous Intersection ManagementabstractArtificial intelligence research is ushering in a new era of sophisticated, mass-market transportation technology. While computers can already fly a passenger jet better than a trained human pilot, people are still faced with the dangerous yet tedious task of driving automobiles. Intelligent Transportation Systems (ITS) is the field that focuses on integrating information technology with vehicles and transportation infrastructure to make transportation safer, cheaper, and more efficient. Recent advances in ITS point to a future in which vehicles themselves handle the vast majority of the driving task. Once autonomous vehicles become popular, autonomous interactions amongst multiple vehicles will be possible. Current methods of vehicle coordination, which are all designed to work with human drivers, will be outdated. The bottleneck for roadway efficiency will no longer be the drivers, but rather the mechanism by which those drivers' actions are coordinated. While open-road driving is a well-studied and more-or-less-solved problem, urban traffic scenarios, especially intersections, are much more challenging. We believe current methods for controlling traffic, specifically at intersections, will not be able to take advantage of the increased sensitivity and precision of autonomous vehicles as compared to human drivers. In this article, we suggest an alternative mechanism for coordinating the movement of autonomous vehicles through intersections. Drivers and intersections in this mechanism are treated as autonomous agents in a multiagent system. In this multiagent system, intersections use a new reservation-based approach built around a detailed communication protocol, which we also present. We demonstrate in simulation that our new mechanism has the potential to significantly outperform current intersection control technology -- traffic lights and stop signs. Because our mechanism can emulate a traffic light or stop sign, it subsumes the most popular current methods of intersection control. This article also presents two extensions to the mechanism. The first extension allows the system to control human-driven vehicles in addition to autonomous vehicles. The second gives priority to emergency vehicles without significant cost to civilian vehicles. The mechanism, including both extensions, is implemented and tested in simulation, and we present experimental results that strongly attest to the efficacy of this approach. Kurt M. Dresner, Peter Stone 0001 |
J. Artif. Intell. Res. | 2 |
| 2007 | Representation Transfer via Elaboration
Matthew E. Taylor, Peter Stone 0001 |
AAAI | 2 |
| 2007 | Temporal Difference and Policy Search Methods for Reinforcement Learning: An Empirical Comparison
Matthew E. Taylor, Shimon Whiteson, Peter Stone 0001 |
AAAI | 3 |
| 2007 | Graph-Based Domain Mapping for Transfer Learning in General Games
Gregory Kuhlmann, Peter Stone 0001 |
ECML | 2 |
| 2007 | Cross-domain transfer for reinforcement learningabstractA typical goal for transfer learning algorithms is to utilize knowledge gained in a source task to learn a target task faster. Recently introduced transfer methods in reinforcement learning settings have shown considerable promise, but they typically transfer between pairs of very similar tasks. This work introduces Rule Transfer, a transfer algorithm that first learns rules to summarize a source task policy and then leverages those rules to learn faster in a target task. This paper demonstrates that Rule Transfer can effectively speed up learning in Keepaway, a benchmark RL problem in the robot soccer domain, based on experience from source tasks in the gridworld domain. We empirically show, through the use of three distinct transfer metrics, that Rule Transfer is effective across these domains. Matthew E. Taylor, Peter Stone 0001 |
ICML | 2 |
| 2007 | A Comparison of Two Approaches for Vision and Self-Localization on a Mobile RobotabstractThis paper considers two approaches to the problem of vision and self-localization on a mobile robot. In the first approach, the perceptual processing is primarily bottom-up, with visual object recognition entirely preceding localization. In the second, significant top-down information is incorporated, with vision and localization being intertwined. That is, the processing of vision is highly dependent on the robot's estimate of its location. The two approaches are implemented and tested on a Sony Aibo ERS-7 robot, localizing as it walks through a color-coded test-bed domain. This paper's contributions are an exposition of two different approaches to vision and localization on a mobile robot, an empirical comparison of the two methods, and a discussion of the relative advantages of each method. Daniel Stronger, Peter Stone 0001 |
ICRA | 2 |
| 2007 | General Game Learning Using Knowledge Transfer
Bikramjit Banerjee, Peter Stone 0001 |
IJCAI | 2 |
| 2007 | Sharing the Road: Autonomous Vehicles Meet Human Drivers
Kurt M. Dresner, Peter Stone 0001 |
IJCAI | 2 |
| 2007 | Color Learning on a Mobile Robot: Towards Full Autonomy under Changing Illumination
Mohan Sridharan, Peter Stone 0001 |
IJCAI | 2 |
| 2007 | Learning and Multiagent Reasoning for Autonomous Agents
Peter Stone 0001 |
IJCAI | 1 |
| 2007 | Machine Learning for On-Line Hardware Reconfiguration
Jonathan Wildstrom, Peter Stone 0001, Emmett Witchel, Michael Dahlin |
IJCAI | 2 |
| 2007 | Global action selection for illumination invariant color modelingabstractA major challenge in the path of widespread use of mobile robots is the ability to function autonomously, learning useful features from the environment and using them to adapt to environmental changes. We propose an algorithm for mobile robots equipped with color cameras that allows for smooth operation under illumination changes. The robot uses image statistics and the environmental structure to autonomously detect and adapt to both major and minor illumination changes. Furthermore, the robot autonomously plans an action sequence that maximizes color learning opportunities while minimizing localization errors. Our approach is fully implemented and tested on the Sony AIBO robots. Mohan Sridharan, Peter Stone 0001 |
IROS | 2 |
| 2007 | Instance-Based Action Models for Fast Action Planning
Mazda Ahmadi, Peter Stone 0001 |
RoboCup | 2 |
| 2007 | A Neural Network-Based Approach to Robot Motion Control
Uli Grasemann, Daniel Stronger, Peter Stone 0001 |
RoboCup | 3 |
| 2007 | Model-Based Reinforcement Learning in a Complex Domain
Shivaram Kalyanakrishnan, Peter Stone 0001 |
RoboCup | 2 |
| 2007 | Multiagent learning is not the answer. It is the question
Peter Stone 0001 |
Artif. Intell. | 1 |
| 2007 | Transfer Learning via Inter-Task Mappings for Temporal Difference Learning
Matthew E. Taylor, Peter Stone 0001 |
J. Mach. Learn. Res. | 2 |
| 2006 | Adaptive mechanism design: a metalearning approachabstractAuction mechanism design has traditionally been a largely analytic process, relying on assumptions such as fully rational bidders. In practice, however, bidders often exhibit unknown and variable behavior, making them difficult to model and complicating the design process. To address this challenge, we explore the use of an adaptive auction mechanism: one that learns to adjust its parameters in response to past empirical bidder behavior so as to maximize an objective function such as auctioneer revenue. In this paper, we give an overview of our general approach and then present an instantiation in a specific auction scenario. In addition, we show how predictions of possible bidder behavior can be incorporated into the adaptive mechanism through a metalearning process. The approach is fully implemented and tested. Results indicate that the adaptive mechanism is able to outperform any single fixed mechanism, and that the addition of metalearning improves performance substantially. David Pardoe, Peter Stone 0001, Maytal Saar-Tsechansky, Kerem Tomak |
ICEC | 2 |
| 2006 | Keeping in Touch: Maintaining Biconnected Structure by Homogeneous Robots
Mazda Ahmadi, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Biconnected Structure for Multi-Robot Systems
Mazda Ahmadi, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Traffic Intersections of the Future
Kurt M. Dresner, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Making Autonomous Intersection Management Backwards-Compatible
Kurt M. Dresner, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Know Thine Enemy: A Champion RoboCup Coach Agent
Gregory Kuhlmann, W. Bradley Knox, Peter Stone 0001 |
AAAI | 3 |
| 2006 | Automatic Heuristic Construction in a Complete General Game Player
Gregory Kuhlmann, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Automatic Heuristic Construction for General Game Playing
Gregory Kuhlmann, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Value-Function-Based Transfer for Reinforcement Learning Using Structure Mapping
Peter Stone 0001 |
AAAI | 2 |
| 2006 | TacTex-05: A Champion Supply Chain Management Agent
David Pardoe, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Expectation-Based Vision for Self-Localization on a Legged Robot
Daniel Stronger, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Inter-Task Action Correlation for Reinforcement Learning Tasks
Matthew E. Taylor, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Sample-Efficient Evolutionary Function Approximation for Reinforcement Learning
Shimon Whiteson, Peter Stone 0001 |
AAAI | 2 |
| 2006 | Designing safe, profitable automated stock trading agents using evolutionary algorithmsabstractTrading rules are widely used by practitioners as an effective means to mechanize aspects of their reasoning about stock price trends. However, due to the simplicity of these rules, each rule is susceptible to poor behavior in specific types of adverse market conditions. Naive combinations of such rules are not very effective in mitigating the weaknesses of component rules. We demonstrate that sophisticated approaches to combining these trading rules enable us to overcome these problems and gainfully utilize them in autonomous agents. We achieve this combination through the use of genetic algorithms and genetic programs. Further, we show that it is possible to use qualitative characterizations of stochastic dynamics to improve the performance of these agents by delineating safe, or feasible, regions. We present the results of experiments conducted within the Penn-Lehman Automated Trading project. In this way we are able to demonstrate that autonomous agents can achieve consistent profitability in a variety of market conditions, in ways that are human competitive. Categories and Subject Descriptors [Real World Applications]: finance, intelligent agents, automated Harish Subramanian, Subramanian Ramamoorthy, Peter Stone 0001, Benjamin Kuipers |
GECCO | 3 |
| 2006 | Comparing evolutionary and temporal difference methods in a reinforcement learning domainabstractBoth genetic algorithms (GAs) and temporal difference (TD) methods have proven effective at solving reinforcement learning (RL) problems. However, since few rigorous empirical comparisons have been conducted, there are no general guidelines describing the methods' relative strengths and weaknesses. This paper presents the results of a detailed empirical comparison between a GA and a TD method in Keepaway, a standard RL benchmark domain based on robot soccer. In particular, we compare the performance of NEAT [19], a GA that evolves neural networks, with Sarsa [16, 17], a popular TD method. The results demonstrate that NEAT can learn better policies in this task, though it requires more evaluations to do so. Additional experiments in two variations of Keepaway demonstrate that Sarsa learns better policies when the task is fully observable and NEAT learns faster when the task is deterministic. Together, these results help isolate the factors critical to the performance of each method and yield insights into their general strengths and weaknesses. Matthew E. Taylor, Shimon Whiteson, Peter Stone 0001 |
GECCO | 3 |
| 2006 | On-line evolutionary computation for reinforcement learning in stochastic domainsabstractIn reinforcement learning, an agent interacting with its environment strives to learn a policy that specifies, for each state it may encounter, what action to take. Evolutionary computation is one of the most promising approaches to reinforcement learning but its success is largely restricted to off-line scenarios. In on-line scenarios, an agent must strive to maximize the reward it accrues while it is learning. Temporal difference (TD) methods, another approach to reinforcement learning, naturally excel in on-line scenarios because they have selection mechanisms for balancing the need to search for better policies exploration) with the need to accrue maximal reward (exploitation). This paper presents a novel way to strike this balance in evolutionary methods by borrowing the selection mechanisms used by TD methods to choose individual actions and using them in evolution to choose policies for evaluation. Empirical results in the mountain car and server job scheduling domains demonstrate that these techniques can substantially improve evolution's on-line performance in stochastic domains. Shimon Whiteson, Peter Stone 0001 |
GECCO | 2 |
| 2006 | Autonomous Planned Color Learning on a Mobile Robot Without Labeled DataabstractColor segmentation is a challenging yet integral subtask of mobile robot systems that use visual sensors, especially since such systems typically have limited computational and memory resources. We present an online approach for a mobile robot to autonomously learn the colors in its environment without any explicitly labeled training data, thereby making it robust to re-colorings in the environment. The robot plans its motion and extracts structure from a color-coded environment to learn colors autonomously and incrementally, with the knowledge acquired at any stage of the learning process being used as a bootstrap mechanism to aid the robot in planning its motion during subsequent stages. With our novel representation, the robot is able to use the same algorithm both within the constrained setting of our lab and in much more uncontrolled settings such as indoor corridors. The segmentation and localization accuracies are comparable to that obtained by a time-consuming offline training process. The algorithm is fully implemented and tested on SONY Aibo robots Mohan Sridharan, Peter Stone 0001 |
ICARCV | 2 |
| 2006 | A Multi-robot System for Continuous Area Sweeping TasksabstractAs mobile robots become increasingly autonomous over extended periods of time, opportunities arise for their use on repetitive tasks. We define and implement behaviors for a class of such tasks that we call continuous area sweeping tasks. A continuous area sweeping task is one in which a group of robots must repeatedly visit all points in a fixed area, possibly with nonuniform frequency, as specified by a task-dependent cost function. Examples of problems that need continuous area sweeping are trash removal in a large building and routine surveillance. In our previous work we have introduced a single-robot approach to this problem. In this paper, we extend that approach to multi-robot scenarios. The focus of this paper is adaptive and decentralized task assignment in continuous area sweeping problems, with the aim of ensuring stability in environments with dynamic factors, such as robot malfunctions or the addition of new robots to the team. Our proposed negotiation-based approach is fully implemented and tested both in simulation and on physical robots Mazda Ahmadi, Peter Stone 0001 |
ICRA | 2 |
| 2006 | Polynomial Regression with Automated Degree: A Function Approximator for Autonomous AgentsabstractIn order for an autonomous agent to behave robustly in a variety of environments, it must have the ability to learn approximations to many different functions. The function approximator used by such an agent is subject to a number of constraints that may not apply in a traditional supervised learning setting. Many different function approximators exist and are appropriate for different problems. This paper proposes a set of criteria for function approximators for autonomous agents. Additionally, for those problems on which polynomial regression is a candidate technique, the paper presents an enhancement that meets these criteria. In particular, using polynomial regression typically requires a manual choice of the polynomial's degree, trading off between function accuracy and computational and memory efficiency. Polynomial regression with automated degree (PRAD) is a novel function approximation method that uses training data to automatically identify an appropriate degree for the polynomial. PRAD is fully implemented. Empirical tests demonstrate its ability to efficiently and accurately approximate both a wide variety of synthetic functions and real-world data gathered by a mobile robot Daniel Stronger, Peter Stone 0001 |
ICTAI | 2 |
| 2006 | The Chin Pinch: A Case Study in Skill Learning on a Legged Robot
Peggy Fidelman, Peter Stone 0001 |
RoboCup | 2 |
| 2006 | Half Field Offense in RoboCup Soccer: A Multiagent Reinforcement Learning Case Study
Shivaram Kalyanakrishnan, Peter Stone 0001 |
RoboCup | 3 |
| 2006 | Autonomous Learning of Stable Quadruped Locomotion
Manish Saggar, Thomas D'Silva, Nate Kohl, Peter Stone 0001 |
RoboCup | 4 |
| 2006 | Autonomous Planned Color Learning on a Legged Robot
Mohan Sridharan, Peter Stone 0001 |
RoboCup | 2 |
| 2006 | Selective Visual Attention for Object Detection on a Legged Robot
Daniel Stronger, Peter Stone 0001 |
RoboCup | 2 |
| 2006 | Cobot in LambdaMOO: An Adaptive Social Statistics Agent
Charles L. Isbell Jr., Michael Kearns, Satinder Singh 0001, Christian R. Shelton, Peter Stone 0001, David P. Kormann |
Auton. Agents Multi Agent Syst. | 5 |
| 2006 | Towards autonomous sensor and actuator model induction on a mobile robotabstractThis article presents a novel methodology for a robot to autonomously induce models of its actions and sensors called ASAMI (autonomous sensor and actuator model induction). While previous approaches to model learning rely on an independent source of training data, we show how a robot can induce action and sensor models without any well-calibrated feedback. Specifically, the only inputs to the ASAMI learning process are the data the robot would naturally have access to: its raw sensations and knowledge of its own action selections. From the perspective of developmental robotics, our robot’s goal is to obtain self-consistent internal models, rather than to perform any externally defined tasks. Furthermore, the target function of each model-learning process comes from within the system, namely the most current version of another internal system model. Concretely realizing this model-learning methodology presents a number of challenges, and we introduce a broad class of settings in which solutions to these challenges are presented. ASAMI is fully implemented and tested, and empirical results validate our approach in a robotic testbed domain using a Sony Aibo ERS-7 robot. Daniel Stronger, Peter Stone 0001 |
Connect. Sci. | 2 |
| 2006 | Evolutionary Function Approximation for Reinforcement LearningabstractTemporal difference methods are theoretically grounded and empirically effective methods for addressing reinforcement learning problems. In most real-world reinforcement learning tasks, TD methods require a function approximator to represent the value function. However, using function approximators requires manually making crucial representational decisions. This paper investigates evolutionary function approximation, a novel approach to automatically selecting function approximator representations that enable efficient individual learning. This method evolves individuals that are better able to learn. We present a fully implemented instantiation of evolutionary function approximation which combines NEAT, a neuroevolutionary optimization technique, with Q-learning, a popular TD method. The resulting NEAT+Q algorithm automatically discovers effective representations for neural network function approximators. This paper also presents on-line evolutionary computation, which improves the on-line performance of evolutionary computation by borrowing selection mechanisms used in TD methods to choose individual actions and using them in evolutionary computation to select policies for evaluation. We evaluate these contributions with extended empirical studies in two domains: 1) the mountain car task, a standard reinforcement learning benchmark on which neural network function approximators have previously performed poorly and 2) server job scheduling, a large probabilistic domain drawn from the field of autonomic computing. The results demonstrate that evolutionary function approximation can significantly improve the performance of TD methods and on-line evolutionary computation can significantly improve evolutionary methods. This paper also presents additional tests that offer insight into what factors can make neural network function approximation difficult in practice. Shimon Whiteson, Peter Stone 0001 |
J. Mach. Learn. Res. | 2 |
| 2005 | Improving Action Selection in MDP's via Knowledge Transfer
Alexander A. Sherstov, Peter Stone 0001 |
AAAI | 2 |
| 2005 | Autonomous Color Learning on a Mobile Robot
Mohan Sridharan, Peter Stone 0001 |
AAAI | 2 |
| 2005 | Value Functions for RL-Based Behavior Transfer: A Comparative Study
Matthew E. Taylor, Peter Stone 0001 |
AAAI | 2 |
| 2005 | Automatic feature selection in neuroevolutionabstractFeature selection is the process of finding the set of inputs to a machine learning algorithm that will yield the best performance. Developing a way to solve this problem automatically would make current machine learning methods much more useful. Previous efforts to automate feature selection rely on expensive meta-learning or are applicable only when labeled training data is available. This paper presents a novel method called FS-NEAT which extends the NEAT neuroevolution method to automatically determine an appropriate set of inputs for the networks it evolves. By learning the network's inputs, topology, and weights simultaneously, FS-NEAT addresses the feature selection problem without relying on meta-learning or labeled data. Initial experiments in an autonomous car racing simulation demonstrate that FS-NEAT can learn better and faster than regular NEAT. In addition, the networks it evolves are smaller and require fewer inputs. Furthermore, FS-NEAT's performance remains robust even as the feature selection task it faces is made increasingly difficult. Shimon Whiteson, Peter Stone 0001, Kenneth O. Stanley, Risto Miikkulainen, Nate Kohl |
GECCO | 2 |
| 2005 | Practical Vision-Based Monte Carlo Localization on a Legged RobotabstractMobile robot localization, the ability of a robot to determine its global position and orientation, continues to be a major research focus in robotics. In most past cases, such localization has been studied on wheeled robots with range finding sensors such as sonar or lasers. In this paper, we consider the more challenging scenario of a legged robot localizing with a limited field-of-view camera as its primary sensory input. We begin with a baseline implementation adapted from the literature that provides a reasonable level of competence, but that exhibits some weaknesses in real-world tests. We propose a series of practical enhancements designed to improve the robot’s sensory and actuator models that enable our robots to achieve a 50% improvement in localization accuracy over the baseline implementation. We go on to demonstrate how the accuracy improvement is even more dramatic when the robot is subjected to large unmodeled movements. These enhancements are each individually straightforward, but together they provide a roadmap for avoiding potential pitfalls when implementing Monte Carlo Localization on vision-based and/or legged robots. Mohan Sridharan, Gregory Kuhlmann, Peter Stone 0001 |
ICRA | 3 |
| 2005 | Simultaneous Calibration of Action and Sensor Models on a Mobile RobotabstractThis paper presents a technique for the Simultaneous Calibration of Action and Sensor Models (SCASM) on a mobile robot. While previous approaches to calibration make use of an independent source of feedback, SCASM is unsupervised, in that it does not receive any well-calibrated feedback about its location. Starting with only an inaccurate action model, it learns accurate relative action and sensor models. Furthermore, SCASM is fully autonomous, in that it operates with no human supervision. SCASM is fully implemented and tested on a Sony Aibo ERS-7 robot. Daniel Stronger, Peter Stone 0001 |
ICRA | 2 |
| 2005 | State Abstraction Discovery from Irrelevant State Variables
Nicholas K. Jong, Peter Stone 0001 |
IJCAI | 2 |
| 2005 | Real-time vision on a mobile robot platformabstractComputer vision is a broad and significant ongoing research challenge, even when performed on an individual image or on streaming video from a high-quality stationary camera with abundant computational resources. When faced with streaming video from a lower-quality, rapidly moving camera and limited computational resources, the challenge increases. We present our implementation of a vision system on a mobile robot platform that uses a camera image as the primary sensory input. Having to perform all processing, including segmentation and object detection, in real-time on-board the robot, eliminates the possibility of using some state-of-the-art methods that otherwise might apply. We describe the methods that we developed to achieve a practical vision system within these constraints. Our approach is fully implemented and tested on a team of Sony AIBO robots. Mohan Sridharan, Peter Stone 0001 |
IROS | 2 |
| 2005 | Towards Eliminating Manual Color Calibration at RoboCup
Mohan Sridharan, Peter Stone 0001 |
RoboCup | 2 |
| 2005 | Keepaway Soccer: From Machine Learning Testbed to Benchmark
Peter Stone 0001, Gregory Kuhlmann, Matthew E. Taylor |
RoboCup | 1 |
| 2005 | A polynomial-time Nash equilibrium algorithm for repeated games
Michael L. Littman, Peter Stone 0001 |
Decis. Support Syst. | 2 |
| 2005 | Evolving Soccer Keepaway Players Through Task Decomposition
Shimon Whiteson, Nate Kohl, Risto Miikkulainen, Peter Stone 0001 |
Mach. Learn. | 4 |
| 2004 | Machine Learning for Fast Quadrupedal Locomotion
Nate Kohl, Peter Stone 0001 |
AAAI | 2 |
| 2004 | Towards Autonomic Computing: Adaptive Job Routing and Scheduling
Shimon Whiteson, Peter Stone 0001 |
AAAI | 2 |
| 2004 | Policy Gradient Reinforcement Learning for Fast Quadrupedal LocomotionabstractThis paper presents a machine learning approach to optimizing a quadrupedal trot gait for forward speed. Given a parameterized walk designed for a specific robot, we propose using a form of policy gradient reinforcement learning to automatically search the set of possible parameters with the goal of finding the fastest possible walk. We implement and test our approach on a commercially available quadrupedal robot platform, namely the Sony Aibo robot. After about three hours of learning, all on the physical robots and with no human intervention other than to change the batteries, the robots achieved a gait faster than any previously known gait known for the Aibo, significantly outperforming a variety of existing hand-coded and learned solutions. Nate Kohl, Peter Stone 0001 |
ICRA | 2 |
| 2004 | The UT Austin Villa 2003 Champion Simulator Coach: A Machine Learning Approach
Gregory Kuhlmann, Peter Stone 0001, Justin Lallinger |
RoboCup | 2 |
| 2004 | Towards Illumination Invariance in the Legged League
Mohan Sridharan, Peter Stone 0001 |
RoboCup | 2 |
| 2004 | A Model-Based Approach to Robot Joint Control
Daniel Stronger, Peter Stone 0001 |
RoboCup | 2 |
| 2004 | Adaptive job routing and scheduling
Shimon Whiteson, Peter Stone 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2004 | Using RoboCup in university-level computer science educationabstractIn the education literature, team-based projects have proven to be an effective pedagogical methodology. We have been using RoboCup challenges as the basis for class projects in undergraduate and masters level courses. This article discusses several independent efforts in this direction and presents our work in the development of shared resources and evaluation instruments. We outline three courses and describe related class projects in order to make the context of our investigation clear and make it possible for others to replicate and extend our work as well as contribute to the shared resource. Elizabeth Sklar, Simon Parsons, Peter Stone 0001 |
ACM J. Educ. Resour. Comput. | 3 |
| 2003 | Performance analysis of a counter-intuitive automated stock-trading agentabstractAutonomous trading in stock markets is an area of great interest in both academic and commercial circles. A lot of trading strategies have been proposed and practiced from the perspectives of Artificial Intelligence, market making, external data indication, technical analysis, etc. This paper examines some properties of a counter-intuitive automated stock-trading strategy in the context of the Penn-Lehman Automated Trading (PLAT) simulator [1], which is a realtime, real-data market simulator. While it might seem natural to buy when the market is on the rise and sell when it is on the decline, our strategy does exactly the opposite. As a result, we call it the reverse strategy. The reverse strategy was the winner strategy in the first and second PLAT live competitions. In this paper, we analyze the performance of the reverse strategy. Also, we suggest ways to control the risk of using the reverse strategy in certain kinds of markets. 1. Ronggang Yu, Peter Stone 0001 |
ICEC | 2 |
| 2003 | Evolving Keepaway Soccer Players through Task Decomposition
Shimon Whiteson, Nate Kohl, Risto Miikkulainen, Peter Stone 0001 |
GECCO | 4 |
| 2003 | Learning Predictive State Representations
Satinder Singh 0001, Michael L. Littman, Nicholas K. Jong, David Pardoe, Peter Stone 0001 |
ICML | 5 |
| 2003 | Progress in Learning 3 vs. 2 Keepaway
Gregory Kuhlmann, Peter Stone 0001 |
RoboCup | 2 |
| 2003 | RoboCup in Higher Education: A Preliminary Report
Elizabeth Sklar, Simon Parsons, Peter Stone 0001 |
RoboCup | 3 |
| 2003 | RoboCup as an Introduction to CS Research
Peter Stone 0001 |
RoboCup | 1 |
| 2003 | A polynomial-time nash equilibrium algorithm for repeated gamesabstractWith the increasing reliance on game theory as a foundation for auctions and electronic commerce, efficient algorithms for computing equilibria in multiplayer general-sum games are of great theoretical and practical interest. The computational complexity of finding a Nash equilibrium for a one-shot bimatrix game is a well known open problem. This paper treats a closely related problem, that of finding a Nash equilibrium for an average-payoff phrepeated bimatrix game, and presents a polynomial-time algorithm. Our approach draws on the "folk theorem" from game theory and shows how finite-state equilibrium strategies can be found efficiently and expressed succinctly. Michael L. Littman, Peter Stone 0001 |
EC | 2 |
| 2003 | Progress in learning 3 vs. 2 keepawayabstractReinforcement learning has been successfully applied to several subtasks in the RoboCup simulated soccer domain. Keepaway is one such task. One notable success in the keepaway domain has been the application of SMDP Sarsa(/spl lambda/) with tile-coding function approximation. However, this success was achieved with the help of some significant task simplifications, including the delivery of complete, noise-free world-state information to the agents. Here we demonstrate that this task simplification was unnecessary: the agents are able to learn even in the presence of noisy, incomplete information. We also scale up to larger problems than have been previously tried. The main contribution of this paper is a deeper understanding of the difficulties of scaling up reinforcement learning to RoboCup soccer. We address several focused questions about the previous results with detailed experiments. Gregory Kuhlmann, Peter Stone 0001 |
SMC | 2 |
| 2003 | The RoboCup Soccer Server and CMUnited Clients: Implemented Infrastructure for MAS Research
Itsuki Noda, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 2 |
| 2003 | Decision-Theoretic Bidding Based on Learned Density Models in Simultaneous, Interacting AuctionsabstractAuctions are becoming an increasingly popular method for transacting business, especially over the Internet. This article presents a general approach to building autonomous bidding agents to bid in multiple simultaneous auctions for interacting goods. A core component of our approach learns a model of the empirical price dynamics based on past data and uses the model to analytically calculate, to the greatest extent possible, optimal bids. We introduce a new and general boosting-based algorithm for conditional density estimation problems of this kind, i.e., supervised learning problems in which the goal is to estimate the entire conditional distribution of the real-valued label. This approach is fully implemented as ATTac-2001, a top-scoring agent in the second Trading Agent Competition (TAC-01). We present experiments demonstrating the effectiveness of our boosting-based price predictor relative to several reasonable alternatives. Peter Stone 0001, Robert E. Schapire, Michael L. Littman, János A. Csirik, David A. McAllester |
J. Artif. Intell. Res. | 1 |
| 2002 | Modeling Auction Price Uncertainty Using Boosting-based Conditional Density Estimation
Robert E. Schapire, Peter Stone 0001, David A. McAllester, Michael L. Littman, János A. Csirik |
ICML | 2 |
| 2002 | Multiagent Competitions and Research: Lessons from RoboCup and TAC
Peter Stone 0001 |
RoboCup | 1 |
| 2001 | Scaling Reinforcement Learning toward RoboCup Soccer
Peter Stone 0001, Richard S. Sutton |
ICML | 1 |
| 2001 | Cobot: A Social Reinforcement Learning AgentabstractWe report on the use of reinforcement learning with Cobot, a software agent residing in the well-known online community LambdaMOO. Our initial work on Cobot (Isbell et al.2000) provided him with the ability to collect social statistics and report them to users. Here we describe an application of RL allowing Cobot to take proactive actions in this complex social environment, and adapt behavior from multiple sources of human reward. After 5 months of training, and 3171 reward and punishment events from 254 different LambdaMOO users, Cobot learned nontrivial preferences for a number of users, modifing his behavior based on his current state. Here we describe LambdaMOO and the state and action spaces of Cobot, and report the statistical results of the learning experiment. Charles L. Isbell Jr., Christian R. Shelton, Michael Kearns, Satinder Singh 0001, Peter Stone 0001 |
NIPS | 5 |
| 2001 | ATTUnited-2001: Using Heterogeneous Players
Peter Stone 0001 |
RoboCup | 1 |
| 2001 | Keepaway Soccer: A Machine Learning Testbed
Peter Stone 0001, Richard S. Sutton |
RoboCup | 1 |
| 2001 | ATTac-2000: An Adaptive Autonomous Bidding AgentabstractThe First Trading Agent Competition (TAC) was held from June 22nd to July 8th, 2000. TAC was designed to create a benchmark problem in the complex domain of e-marketplaces and to motivate researchers to apply unique approaches to a common task. This article describes ATTac-2000, the first-place finisher in TAC. ATTac-2000 uses a principled bidding strategy that includes several elements of adaptivity. In addition to the success at the competition, isolated empirical results are presented indicating the robustness and effectiveness of ATTac-2000's adaptive strategy. Peter Stone 0001, Michael L. Littman, Satinder Singh 0001, Michael Kearns |
J. Artif. Intell. Res. | 1 |
| 2000 | Layered Learning
Peter Stone 0001, Manuela M. Veloso |
ECML | 1 |
| 2000 | TPOT-RL Applied to Network Routing
Peter Stone 0001 |
ICML | 1 |
| 2000 | Keeping the Ball from CMUnited-99
David A. McAllester, Peter Stone 0001 |
RoboCup | 2 |
| 2000 | ATT-CMUnited-2000: Third Place Finisher in the RoboCup-2000 Simulator League
Patrick F. Riley, Peter Stone 0001, David A. McAllester, Manuela M. Veloso |
RoboCup | 2 |
| 2000 | Overview of RoboCup-2000
Peter Stone 0001, Minoru Asada, Tucker R. Balch, Masahiro Fujita 0002, Gerhard K. Kraetzschmar, Henrik Hautop Lund, Paul Scerri, Satoshi Tadokoro, Gordon F. Wyeth |
RoboCup | 1 |
| 2000 | Reinforcement Learning for 3 vs. 2 Keepaway
Peter Stone 0001, Richard S. Sutton, Satinder Singh 0001 |
RoboCup | 1 |
| 1999 | The CMUnited-99 Champion Simulator Team
Peter Stone 0001, Patrick F. Riley, Manuela M. Veloso |
RoboCup | 1 |
| 1999 | Layered Learning and Flexible Teamwork in RoboCup Simulation Agents
Peter Stone 0001, Manuela M. Veloso |
RoboCup | 1 |
| 1999 | Overview of RoboCup-99
Manuela M. Veloso, Hiroaki Kitano, Enrico Pagello, Gerhard K. Kraetzschmar, Peter Stone 0001, Tucker R. Balch, Minoru Asada, Silvia Coradeschi, Lars Karlsson, Masahiro Fujita 0002 |
RoboCup | 5 |
| 1999 | Task Decomposition, Dynamic Role Assignment, and Low-Bandwidth Communication for Real-Time Strategic Teamwork
Peter Stone 0001, Manuela M. Veloso |
Artif. Intell. | 1 |
| 1998 | Team-Partitioned, Opaque-Transition Reinforced Learning
Peter Stone 0001, Manuela M. Veloso |
RoboCup | 1 |
| 1998 | The CMUnited-98 Champion Simulator Team
Peter Stone 0001, Manuela M. Veloso, Patrick F. Riley |
RoboCup | 1 |
| 1998 | The CMUnited-98 Small-Robot Team
Manuela M. Veloso, Michael H. Bowling, Sorin Achim, Kwun Han, Peter Stone 0001 |
RoboCup | 5 |
| 1998 | Towards collaborative and adversarial learning: a case study in robotic soccer
Peter Stone 0001, Manuela M. Veloso |
Int. J. Hum. Comput. Stud. | 1 |
| 1997 | The RoboCup Synthetic Agent Challenge 97
Hiroaki Kitano, Milind Tambe, Peter Stone 0001, Manuela M. Veloso, Silvia Coradeschi, Eiichi Osawa, Hitoshi Matsubara, Itsuki Noda, Minoru Asada |
IJCAI (1) | 3 |
| 1997 | The RoboCup Physical Agent Challenge: Goals and Protocols for Phase 1
Minoru Asada, Peter Stone 0001, Hiroaki Kitano, Alexis Drogoul, Dominique Duhaut, Manuela M. Veloso, Hajime Asama, Sho'ji Suzuki |
RoboCup | 2 |
| 1997 | The RoboCup Synthetic Agent Challenge 97
Hiroaki Kitano, Milind Tambe, Peter Stone 0001, Manuela M. Veloso, Silvia Coradeschi, Eiichi Osawa, Hitoshi Matsubara, Itsuki Noda, Minoru Asada |
RoboCup | 3 |
| 1997 | Using Decision Tree Confidence Factors for Multiagent Control
Peter Stone 0001, Manuela M. Veloso |
RoboCup | 1 |
| 1997 | The CMUnited-97 Simulator Team
Peter Stone 0001, Manuela M. Veloso |
RoboCup | 1 |
| 1997 | The CMUnited-97 Small Robot Team
Manuela M. Veloso, Peter Stone 0001, Kwun Han, Sorin Achim |
RoboCup | 2 |
| 1995 | Beating a Defender in Robotic Soccer: Memory-Based Learning of a Continuous Function
Peter Stone 0001, Manuela M. Veloso |
NIPS | 1 |
| 1995 | FLECS: Planning with a Flexible Commitment StrategyabstractThere has been evidence that least-commitment planners can efficiently handle planning problems that involve difficult goal interactions. This evidence has led to the common belief that delayed-commitment is the "best" possible planning strategy. However, we recently found evidence that eager-commitment planners can handle a variety of planning problems more efficiently, in particular those with difficult operator choices. Resigned to the futility of trying to find a universally successful planning strategy, we devised a planner that can be used to study which domains and problems are best for which planning strategies. In this article we introduce this new planning algorithm, FLECS, which uses a FLExible Commitment Strategy with respect to plan-step orderings. It is able to use any strategy from delayed-commitment to eager-commitment. The combination of delayed and eager operator-ordering commitments allows FLECS to take advantage of the benefits of explicitly using a simulated execution state and reasoning about planning constraints. FLECS can vary its commitment strategy across different problems and domains, and also during the course of a single planning problem. FLECS represents a novel contribution to planning in that it explicitly provides the choice of which commitment strategy to use while planning. FLECS provides a framework to investigate the mapping from planning domains and problems to efficient planning strategies. Manuela M. Veloso, Peter Stone 0001 |
J. Artif. Intell. Res. | 2 |