EDBT 2026 Demo / reviewers in the wild / expert
Matthew E. Taylor
dblp:46/4287 · also Matthew Edmund Taylor
· DBLP profile ↗
91ranked-venue papers
17as first author
38since 2021 · last 2026
0000-0001-8946-0211ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 80 · 16 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 8 first-author · 11 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-Interactive Robot Learning: Definition, Challenges, and RecommendationsabstractRobot learning from humans has been proposed and researched for several decades as a means to enable robots to learn new skills or adapt existing ones to new situations. Recent advances in AI, including learning approaches like reinforcement learning and architectures like transformers and foundation models, combined with access to massive datasets, have created attractive opportunities to apply those data-hungry techniques to this problem. We argue that the focus on massive amounts of pre-collected data, and the resulting learning paradigm, where humans demonstrate and robots learn in isolation, is overshadowing a specialized area of work we term Human-Interactive Robot Learning (HIRL). This paradigm, wherein robots and humans interact during the learning process , is at the intersection of multiple fields (AI, robotics, human–computer interaction, design and others) and holds unique promise. Using HIRL, robots can achieve greater sample efficiency (as humans can provide task knowledge through interaction), align with human preferences (as humans can guide the robot behavior toward their expectations), and explore more meaningfully and safely (as humans can utilize domain knowledge to guide learning and prevent catastrophic failures). This can result in robotic systems that can more quickly and easily adapt to new tasks in human environments. The objective of this article is to provide a broad and consistent overview of HIRL research and to guide researchers toward understanding the scope of HIRL, and current open or underexplored challenges related to four themes—namely, human, robot learning, interaction, and broader context. The article includes concrete use cases to illustrate the interaction between these challenges and inspire further research according to broad recommendations and a call for action for the growing HIRL community. Kim Baraka, Ifrah Idrees, Taylor Kessler Faulkner, Erdem Biyik, Serena Booth, Mohamed Chetouani, Daniel H. Grollman, Akanksha Saran, Emmanuel Senft, Silvia Tulli, Anna-Lisa Vollmer, Antonio Andriella, Helen Beierling, Tiffany Horter, Jens Kober, Isaac S. Sheidlower, Matthew E. Taylor, Sanne van Waveren, Xuesu Xiao |
ACM Trans. Hum. Robot Interact. | 17 |
| 2025 | An LLM-Guided Tutoring System for Social Skills TrainingabstractSocial skills training targets behaviors necessary for success in social interactions. However, traditional classroom training for such skills is often insufficient to teach effective communication — one-to-one interaction in real-world scenarios is preferred to lecture-style information delivery. This paper introduces a framework that allows instructors to collaborate with large language models to dynamically design realistic scenarios for students to communicate. Our framework uses these scenarios to enable student rehearsal, provide immediate feedback and visualize performance for both students and instructors. Unlike traditional intelligent tutoring systems, instructors can easily co-create scenarios with a large language model without technical skills. Additionally, the system generates new scenario branches in real time when existing options don't fit the student's response. Michael Guevarra, Indronil Bhattacharjee, Srijita Das 0001, Christabel Wayllace, Carrie Demmans Epp, Matthew E. Taylor, Alan Tay |
AAAI | 6 |
| 2025 | Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningabstractTo create useful reinforcement learning (RL) agents, step zero is to design a suitable reward function that captures the nuances of the task. However, reward engineering can be a difficult and time-consuming process.
Instead, human-in-the-loop RL methods hold the promise of learning reward functions from human feedback. Despite recent successes, many of the human-in-the-loop RL methods still require numerous human interactions to learn successful reward functions.
To improve the feedback efficiency of human-in-the-loop RL methods (i.e., require less human interaction), this paper introduces Sub-optimal Data Pre-training, SDP, an approach that leverages reward-free, sub-optimal data to improve scalar- and preference-based RL algorithms. In SDP, we start by pseudo-labeling all low-quality data with the minimum environment reward. Through this process, we obtain reward labels
to pre-train our reward model without requiring human labeling or preferences.
This pre-training phase provides the reward model a head start in learning, enabling it to recognize that low-quality transitions should be assigned low rewards. Through extensive experiments with both simulated and human teachers, we find that SDP can at least meet, but often significantly improve, state of the art human-in-the-loop RL performance across a variety of simulated robotic tasks. Calarina Muslimani, Matthew E. Taylor |
ICLR | 2 |
| 2025 | Model-Based Exploration in Monitored Markov Decision ProcessesabstractA tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or malfunctioning, or rewards may be inaccessible during deployment. Monitored Markov decision processes (Mon-MDPs) have recently been proposed to model such settings. However, existing Mon-MDP algorithms have several limitations: they do not fully exploit the problem structure, cannot leverage a known monitor, lack worst-case guarantees for "unsolvable" Mon-MDPs without specific initialization, and offer only asymptotic convergence proofs. This paper makes three contributions. First, we introduce a model-based algorithm for Mon-MDPs that addresses these shortcomings. The algorithm employs two instances of model-based interval estimation: one to ensure that observable rewards are reliably captured, and another to learn the minimax-optimal policy. Second, we empirically demonstrate the advantages. We show faster convergence than prior algorithms in more than four dozen benchmarks, and even more dramatic improvements when the monitoring process is known. Third, we present the first finite-sample bound on performance. We show convergence to a minimax-optimal policy even when some rewards are never observable. Alireza Kazemipour, Matthew E. Taylor, Michael H. Bowling |
ICML | 2 |
| 2025 | Taming Multi-Agent Reinforcement Learning with Estimator Variance Reduction
Taher Jafferjee, Juliusz Krysztof Ziomek, Tianpei Yang, Zipeng Dai, Matthew E. Taylor, Kun Shao, Jun Wang 0012, David Mguni |
AAMAS | 6 |
| 2025 | Boosting Robustness in Preference-Based Reinforcement Learning with Dynamic Sparsity
Calarina Muslimani, Bram Grooten, Deepak Ranganatha Sastry Mamillapalli, Mykola Pechenizkiy, Decebal Constantin Mocanu, Matthew E. Taylor |
AAMAS | 6 |
| 2025 | Empowering Generalization for Deep Reinforcement Learning via Symbolic Planning
Tianpei Yang, Srijita Das 0001, Christabel Wayllace, Matthew E. Taylor |
AAMAS | 4 |
| 2025 | The Evolving Landscape of LLM- and VLM-Integrated Reinforcement LearningabstractReinforcement learning (RL) has shown impressive results in sequential decision-making tasks. Large Language Models (LLMs) and Vision-Language Models (VLMs) have recently emerged, exhibiting impressive capabilities in multimodal understanding and reasoning. These advances have led to a surge of research integrating LLMs and VLMs into RL. This survey reviews representative works in which LLMs and VLMs are used to overcome key challenges in RL, such as lack of prior knowledge, long-horizon planning, and reward design. We present a taxonomy that categorizes these LLM/VLM-assisted RL approaches into three roles: agent, planner, and reward. We conclude by exploring open problems, including grounding, bias mitigation, improved representations, and action advice. By consolidating existing research and identifying future directions, this survey establishes a framework for integrating LLMs and VLMs into RL, advancing approaches that unify natural language and visual understanding with sequential decision-making. Sheila Schoepp, Masoud Jafaripour, Yingyue Cao, Tianpei Yang, Fatemeh Abdollahi, Shadan Golestan, Zahin Sufiyan, Osmar R. Zaïane, Matthew E. Taylor |
IJCAI | 9 |
| 2025 | Pilot Trainees Benefit from Modelling and Adaptive FeedbackabstractLimited training capacity has contributed to a critical shortage of licensed commercial pilots.Adaptive educational technologies and simulators could alleviate current training bottlenecks if these technologies could assess trainee performance and provide appropriate feedback.Agents can be used to assess trainee performance, but there is insufficient guidance on how to provide concurrent feedback in simulation-based learning environments.So, we designed 4 feedback conditions that provide varying degrees of elaboration and used a within-subject study (𝑛 = 20) to compare feedback approaches.Trainee performance was best when they received highly-elaborative feedback that modeled expert behaviour.Variability in participant performance and preferences indicates a need to adapt the feedback type to individual learners and provides insight into the use of concurrent feedback in simulation-based learning environments.Specifically, learners appreciated the expert model because it facilitated a sense of control which was associated with lower negative affect and lower extraneous cognitive load. Yalmaz Ali Abdullah, Michael Guevarra, Minghao Cai, Matthew E. Taylor, Carrie Demmans Epp |
UMAP | 5 |
| 2025 | hammer: Multi-level coordination of reinforcement learning agents via learned messaging
Nikunj Gupta, G. Srinivasaraghavan 0001, Swarup Mohalik, Matthew E. Taylor |
Neural Comput. Appl. | 5 |
| 2025 | Human-AI collaboration in real-world complex environment with reinforcement learning
Md. Saiful Islam 0007, Srijita Das 0001, Sai Krishna Gottipati, William Duguay, Clodéric Mars, Jalal Arabneydi, Antoine Fagette, Matthew Guzdial, Matthew E. Taylor |
Neural Comput. Appl. | 9 |
| 2025 | Do as you teach: a multi-teacher approach to self-play in deep reinforcement learning
Chaitanya Kharyal, Sai Krishna Gottipati, Tanmay Kumar Sinha, Fatemeh Abdollahi, Srijita Das 0001, Matthew E. Taylor |
Neural Comput. Appl. | 6 |
| 2024 | PORTAL: Automatic Curricula Generation for Multiagent Reinforcement LearningabstractDespite many breakthroughs in recent years, it is still hard for MultiAgent Reinforcement Learning (MARL) algorithms to directly solve complex tasks in MultiAgent Systems (MASs) from scratch. In this work, we study how to use Automatic Curriculum Learning (ACL) to reduce the number of environmental interactions required to learn a good policy. In order to solve a difficult task, ACL methods automatically select a sequence of tasks (i.e., curricula). The idea is to obtain maximum learning progress towards the final task by continuously learning on tasks that match the current capabilities of the learners. The key question is how to measure the learning progress of the learner for better curriculum selection. We propose a novel ACL framework, PrOgRessive mulTiagent Automatic curricuLum (PORTAL), for MASs. PORTAL selects curricula according to two critera: 1) How difficult is a task, relative to the learners’ current abilities? 2) How similar is a task, relative to the final task? By learning a shared feature space between tasks, PORTAL is able to characterize different tasks based on the distribution of features and select those that are similar to the final task. Also, the shared feature space can effectively facilitate the policy transfer between curricula. Experimental results show that PORTAL can train agents to master extremely hard cooperative tasks, which can not be achieved with previous state-of-the-art MARL algorithms. Jizhou Wu, Jianye Hao, Tianpei Yang, Xiaotian Hao, Yan Zheng 0002, Weixun Wang, Matthew E. Taylor |
AAAI | 7 |
| 2024 | A Transfer Approach Using Graph Neural Networks in Deep Reinforcement LearningabstractTransfer learning (TL) has shown great potential to improve Reinforcement Learning (RL) efficiency by leveraging prior knowledge in new tasks. However, much of the existing TL research focuses on transferring knowledge between tasks that share the same state-action spaces. Further, transfer from multiple source tasks that have different state-action spaces is more challenging and needs to be solved urgently to improve the generalization and practicality of the method in real-world scenarios. This paper proposes TURRET (Transfer Using gRaph neuRal nETworks), to utilize the generalization capabilities of Graph Neural Networks (GNNs) to facilitate efficient and effective multi-source policy transfer learning in the state-action mismatch setting. TURRET learns a semantic representation by accounting for the intrinsic property of the agent through GNNs, which leads to a unified state embedding space for all tasks. As a result, TURRET achieves more efficient transfer with strong generalization ability between different tasks and can be easily combined with existing Deep RL algorithms. Experimental results show that TURRET significantly outperforms other TL methods on multiple continuous action control tasks, successfully transferring across robots with different state-action spaces. Tianpei Yang, Heng You, Jianye Hao, Yan Zheng 0002, Matthew E. Taylor |
AAAI | 5 |
| 2024 | Local Linearity is All You Need (in Data-Driven Teleoperation)abstractOne of the critical aspects of assistive robotics is to provide a control system of a high-dimensional robot from a low-dimensional user input (i.e. a 2D joystick). Data-driven teleoperation seeks to provide an intuitive user interface called an action map to map the low dimensional input to robot velocities from human demonstrations. Action maps are machine learning models trained on robotic demonstration data to map user input directly to desired movements as opposed to aspects of robot pose ("move to cup or pour content" vs. "move along x- or y-axis"). Many works have investigated nonlinear action maps with multi-layer perceptrons, but recent work suggests that local-linear neural approximations provide better control of the system. However, local linear models assume actions exist on a linear subspace and may not capture nuanced motions in training data. In this work, we hypothesize that local-linear neural networks are effective because they make the action map odd w.r.t. the user input, enhancing the intuitiveness of the controller. Based on this assumption, we propose two nonlinear means of encoding odd behavior that do not constrain the action map to a local linear function. However, our analysis reveals that these models effectively behave like local linear models for relevant mappings between user joysticks and robot movements. We support this claim in simulation, and show on a realworld use case that there is no statistical benefit of using non-linear maps, according to the users experience. These negative results suggest that further investigation into model architectures beyond local linear models may offer diminishing returns for improving user experience in data-driven teleoperation systems. Michael Przystupa, Gauthier Gidel, Matthew E. Taylor, Martin Jägersand, Justus H. Piater, Samuele Tosatto |
IROS | 3 |
| 2024 | Human-in-the-Loop Reinforcement Learning: A Survey and Position on Requirements, Challenges, and OpportunitiesabstractArtificial intelligence (AI) and especially reinforcement learning (RL) have the potential to enable agents to learn and perform tasks autonomously with superhuman performance. However, we consider RL as fundamentally a Human-in-the-Loop (HITL) paradigm, even when an agent eventually performs its task autonomously. In cases where the reward function is challenging or impossible to define, HITL approaches are considered particularly advantageous. The application of Reinforcement Learning from Human Feedback (RLHF) in systems such as ChatGPT demonstrates the effectiveness of optimizing for user experience and integrating their feedback into the training loop. In HITL RL, human input is integrated during the agent’s learning process, allowing iterative updates and fine-tuning based on human feedback, thus enhancing the agent’s performance. Since the human is an essential part of this process, we argue that human-centric approaches are the key to successful RL, a fact that has not been adequately considered in the existing literature. This paper aims to inform readers about current explainability methods in HITL RL. It also shows how the application of explainable AI (xAI) and specific improvements to existing explainability approaches can enable a better human-agent interaction in HITL RL for all types of users, whether for lay people, domain experts, or machine learning specialists. Accounting for the workflow in HITL RL and based on software and machine learning methodologies, this article identifies four phases for human involvement for creating HITL RL systems: (1) Agent Development, (2) Agent Learning, (3) Agent Evaluation, and (4) Agent Deployment. We highlight human involvement, explanation requirements, new challenges, and goals for each phase. We furthermore identify low-risk, high-return opportunities for explainability research in HITL RL and present long-term research goals to advance the field. Finally, we propose a vision of human-robot collaboration that allows both parties to reach their full potential and cooperate effectively. Carl Orge Retzlaff, Srijita Das 0001, Christabel Wayllace, Payam Mousavi, Mohammad Afshari, Tianpei Yang, Anna Saranti, Alessa Angerschmid, Matthew E. Taylor, Andreas Holzinger |
J. Artif. Intell. Res. | 9 |
| 2024 | Comparing explanations in RL
Brittany Davis Pierson, Dustin Arendt, Matthew E. Taylor |
Neural Comput. Appl. | 4 |
| 2024 | Applying reinforcement learning to learn best net to rip and re-route in global routingabstractPhysical designers typically employ heuristics to solve challenging problems in global routing. However, these heuristic solutions are not adaptable to the ever-changing fabrication demands, and the experience and creativity of designers can limit their effectiveness. Reinforcement learning (RL) is an effective method to tackle sequential optimization problems due to its ability to adapt and learn through trial and error. Hence, RL can create policies that can handle complex tasks. This work presents an RL framework for global routing that incorporates a self-learning model called RL-Ripper. The primary function of RL-Ripper is to identify the best nets that need to be ripped and rerouted in order to decrease the number of total short violations. In this work, we show that the proposed RL-Ripper framework’s approach can reduce the number of short violations for ISPD 2018 benchmarks when compared to the state-of-the-art global router CUGR. Moreover, RL-Ripper reduced the total number of short violations after the first iteration of detailed routing over the baseline while being on par with the wirelength, VIA, and runtime. The proposed framework’s major impact is providing a novel learning-based approach to global routing that can be replicated for newer technologies. Upma Gandhi, Erfan Aghaeekiasaraee, Sahir, Payam Mousavi, Ismail Bustany, Matthew E. Taylor, Laleh Behjat |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2023 | Augmenting Flight Training with AI to Efficiently Train PilotsabstractWe propose an AI-based pilot trainer to help students learn how to fly aircraft. First, an AI agent uses behavioral cloning to learn flying maneuvers from qualified flight instructors. Later, the system uses the agent's decisions to detect errors made by students and provide feedback to help students correct their errors. This paper presents an instantiation of the pilot trainer. We focus on teaching straight and level flying maneuvers by automatically providing formative feedback to the human student. Michael Guevarra, Srijita Das 0001, Christabel Wayllace, Carrie Demmans Epp, Matthew E. Taylor, Alan Tay |
AAAI | 5 |
| 2023 | Learning to Shape Rewards Using a Game of Two PartnersabstractReward shaping (RS) is a powerful method in reinforcement learning (RL) for overcoming the problem of sparse or uninformative rewards. However, RS typically relies on manually engineered shaping-reward functions whose construc- tion is time-consuming and error-prone. It also requires domain knowledge which runs contrary to the goal of autonomous learning. We introduce Reinforcement Learning Optimising Shaping Algorithm (ROSA), an automated reward shaping framework in which the shaping-reward function is constructed in a Markov game between two agents. A reward-shaping agent (Shaper) uses switching controls to determine which states to add shaping rewards for more efficient learning while the other agent (Controller) learns the optimal policy for the task using these shaped rewards. We prove that ROSA, which adopts existing RL algorithms, learns to construct a shaping-reward function that is beneficial to the task thus ensuring efficient convergence to high performance policies. We demonstrate ROSA’s properties in three didactic experiments and show its superior performance against state-of-the-art RS algorithms in challenging sparse reward environments. David Mguni, Taher Jafferjee, Nicolas Perez Nieves, Wenbin Song, Feifei Tong, Matthew E. Taylor, Tianpei Yang, Zipeng Dai, Jiangcheng Zhu, Kun Shao, Jun Wang 0012, Yaodong Yang 0001 |
AAAI | 7 |
| 2023 | Model AI Assignments 2023abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of six AI assignments from the 2023 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http://modelai.gettysburg.edu . Todd W. Neller, Raechel Walker, Olivia Dias, Zeynep Yalcin, Cynthia Breazeal, Matthew E. Taylor, Michele Donini, Erin Talvitie, Charlie Pilgrim, Paolo Turrini, James Maher, Matthew Boutell, Justin Wilson, Narges Norouzi, Jonathan Scott |
AAAI | 6 |
| 2023 | C2Tutor: Helping People Learn to Avoid Present Bias During Decision Making
Calarina Muslimani, Saba Gul, Matthew E. Taylor, Carrie Demmans Epp, Christabel Wayllace |
AIED | 3 |
| 2023 | Innovating AI Leadership EducationabstractThis research to practice full paper explores a new educational framework for AI-informed leadership and evaluates its curriculum and pedagogical approach through a novel, tailored, research instrument. Artificial Intelligence continues to rapidly transform many aspects of markets, solutions, and organizational culture across companies, agencies, and institutions in the public and private sectors. Within complex organizations, AI tools, technologies, and applications inform how leaders engage in strategy-making, management, operations, human resources, and professional education. Non-technical managers and executives are increasingly expected to lead teams to implement responsible AI solutions with the promise to improve efficiency, effectiveness, productivity, profitability, and more. AI is rapidly transforming organizational culture, requiring non-technical leaders to develop AI literacy and essential skills to lead teams in implementing responsible AI solutions. In the face of AI-driven change, business leaders need to be AI literate and develop their own essential skills, knowledge, procedures, and perspectives to successfully set vision and strategy to lead teams that can leverage AI to achieve inward-facing and outward-facing business goals. This presents challenges and opportunities to develop new pedagogical approaches and measures to prepare and assess business leaders' AI leadership skills - including understanding human-AI systems in the workplace and their responsible development and ethical use. There are also cultural and organizational behavior challenges in successfully adopting these new capabilities into a global and diverse human-AI workforce at scale. To advance these, we present an innovative hands-on AI leadership curriculum, where participants learn by making and team problem-solving, for United States Air Force (USAF) leaders to learn about AI and its responsible use in human-robot teaming with autonomous robots. We contribute new measures to assess their attitudinal shifts in AI leadership with respect to culture, mindsets, and ethics. We present a pilot study to evaluate our curriculum design and pedagogical approach to foster positive shifts in our AI leadership measures. Xiaoxue Du, Sharifa Alghowinem, Matthew E. Taylor, Kate Darling, Cynthia Breazeal |
FIE | 3 |
| 2023 | RL-Ripper: : A Framework for Global Routing Using Reinforcement Learning and Smart Net Ripping TechniquesabstractPhysical designers have been using heuristics to solve challenging problems in routing. However, these heuristic solutions are not adaptable to the ever-changing fabrication demands and their effectiveness is limited by the experience and creativity of the designer. Reinforcement learning is an effective method to tackle sequential optimization problems due to its ability to adapt and learn through trial and error, creating policies that can handle complex tasks. This study presents an RL framework for global routing that incorporates a self-learning model called RL-Ripper. The primary function of RL-Ripper is to identify the best nets to rip to decrease the number of total short violations. In this work, the final global routing results are evaluated against CUGR, a state-of-the-art global router, using the ISPD 2018 benchmarks. The proposed RL-Ripper framework's approach can reduce the short violations compared to CUGR. Moreover, the RL-Ripper reduced the total number of short violations after the first iteration of detailed routing over the baseline while being on par with the wirelength, VIA, and runtime. The major impact of the proposed framework is to provide a novel learning-based approach to global routing that can be replicated for newer technologies. Upma Gandhi, Erfan Aghaeekiasaraee, Ismail Bustany, Payam Mousavi, Matthew E. Taylor, Laleh Behjat |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | Can You Improve My Code? Optimizing Programs with Local SearchabstractThis paper introduces a local search method for improving an existing program with respect to a measurable objective. Program Optimization with Locally Improving Search (POLIS) exploits the structure of a program, defined by its lines. POLIS improves a single line of the program while keeping the remaining lines fixed, using existing brute-force synthesis algorithms, and continues iterating until it is unable to improve the program's performance. POLIS was evaluated with a 27-person user study, where participants wrote programs attempting to maximize the score of two single-agent games: Lunar Lander and Highway. POLIS was able to substantially improve the participants' programs with respect to the game scores. A proof-of-concept demonstration on existing Stack Overflow code measures applicability in real-world problems. These results suggest that POLIS could be used as a helpful programming assistant for programming problems with measurable objectives. Fatemeh Abdollahi, Saqib Ameen, Matthew E. Taylor, Levi Lelis |
IJCAI | 3 |
| 2023 | Multi-Agent Advisor Q-Learning (Extended Abstract)abstractIn the last decade, there have been significant advances in multi-agent reinforcement learning (MARL) but there are still numerous challenges, such as high sample complexity and slow convergence to stable policies, that need to be overcome before wide-spread deployment is possible. However, many real-world environments already, in practice, deploy sub-optimal or heuristic approaches for generating policies. An interesting question that arises is how to best use such approaches as advisors to help improve reinforcement learning in multi-agent domains. We provide a principled framework for incorporating action recommendations from online sub-optimal advisors in multi-agent settings. We describe the problem of ADvising Multiple Intelligent Reinforcement Agents (ADMIRAL) in nonrestrictive general-sum stochastic game environments and present two novel Q-learning-based algorithms: ADMIRAL - Decision Making (ADMIRAL-DM) and ADMIRAL - Advisor Evaluation (ADMIRAL-AE), which allow us to improve learning by appropriately incorporating advice from an advisor (ADMIRAL-DM), and evaluate the effectiveness of an advisor (ADMIRAL-AE). We analyze the algorithms theoretically and provide fixed point guarantees regarding their learning in general-sum stochastic games. Furthermore, extensive experiments illustrate that these algorithms: can be used in a variety of environments, have performances that compare favourably to other related baselines, can scale to large state-action spaces, and are robust to poor advice from advisors. Sriram Ganapathi Subramanian, Matthew E. Taylor, Kate Larson, Mark Crowley 0001 |
IJCAI | 2 |
| 2023 | Ignorance is Bliss: Robust Control via Information GatingabstractInformational parsimony provides a useful inductive bias for learning representations that achieve better generalization by being robust to noise and spurious correlations. We propose *information gating* as a way to learn parsimonious representations that identify the minimal information required for a task. When gating information, we can learn to reveal as little information as possible so that a task remains solvable, or hide as little information as possible so that a task becomes unsolvable. We gate information using a differentiable parameterization of the signal-to-noise ratio, which can be applied to arbitrary values in a network, e.g., erasing pixels at the input layer or activations in some intermediate layer. When gating at the input layer, our models learn which visual cues matter for a given task. When gating intermediate layers, our models learn which activations are needed for subsequent stages of computation. We call our approach *InfoGating*. We apply InfoGating to various objectives such as multi-step forward and inverse dynamics models, Q-learning, and behavior cloning, highlighting how InfoGating can naturally help in discarding information not relevant for control. Results show that learning to identify and use minimal information can improve generalization in downstream tasks. Policies based on InfoGating are considerably more robust to irrelevant visual features, leading to improved pretraining and finetuning of RL models. Manan Tomar, Riashat Islam, Matthew E. Taylor, Sergey Levine, Philip Bachman |
NeurIPS | 3 |
| 2023 | ASN: action semantics network for multiagent reinforcement learning
Tianpei Yang, Weixun Wang, Jianye Hao, Matthew E. Taylor, Yong Liu 0007, Xiaotian Hao, Yujing Hu, Changjie Fan, Chunxu Ren, Jiangcheng Zhu, Yang Gao 0001 |
Auton. Agents Multi Agent Syst. | 4 |
| 2023 | Improving reinforcement learning with human assistance: an argument for human subject studies with HIPPO Gym
Matthew E. Taylor, Nicholas Nissen, Neda Navidi |
Neural Comput. Appl. | 1 |
| 2022 | Decentralized Mean Field GamesabstractMultiagent reinforcement learning algorithms have not been widely adopted in large scale environments with many agents as they often scale poorly with the number of agents. Using mean field theory to aggregate agents has been proposed as a solution to this problem. However, almost all previous methods in this area make a strong assumption of a centralized system where all the agents in the environment learn the same policy and are effectively indistinguishable from each other. In this paper, we relax this assumption about indistinguishable agents and propose a new mean field system known as Decentralized Mean Field Games, where each agent can be quite different from others. All agents learn independent policies in a decentralized fashion, based on their local observations. We define a theoretical solution concept for this system and provide a fixed point guarantee for a Q-learning based algorithm in this system. A practical consequence of our approach is that we can address a `chicken-and-egg' problem in empirical mean field reinforcement learning algorithms. Further, we provide Q-learning and actor-critic algorithms that use the decentralized mean field learning approach and give stronger performances compared to common baselines in this area. In our setting, agents do not need to be clones of each other and learn in a fully decentralized fashion. Hence, for the first time, we show the application of mean field learning methods in fully competitive environments, large-scale continuous action space environments, and other environments with heterogeneous agents. Importantly, we also apply the mean field method in a ride-sharing problem using a real-world dataset. We propose a decentralized solution to this problem, which is more practical than existing centralized training methods. Sriram Ganapathi Subramanian, Matthew E. Taylor, Mark Crowley 0001, Pascal Poupart |
AAAI | 2 |
| 2022 | PMIC: Improving Multi-Agent Reinforcement Learning with Progressive Mutual Information CollaborationabstractLearning to collaborate is critical in Multi-Agent Reinforcement Learning (MARL). Previous works promote collaboration by maximizing the correlation of agents’ behaviors, which is typically characterized by Mutual Information (MI) in different forms. However, we reveal sub-optimal collaborative behaviors also emerge with strong correlations, and simply maximizing the MI can, surprisingly, hinder the learning towards better collaboration. To address this issue, we propose a novel MARL framework, called Progressive Mutual Information Collaboration (PMIC), for more effective MI-driven collaboration. PMIC uses a new collaboration criterion measured by the MI between global states and joint actions. Based on this criterion, the key idea of PMIC is maximizing the MI associated with superior collaborative behaviors and minimizing the MI associated with inferior ones. The two MI objectives play complementary roles by facilitating better collaborations while avoiding falling into sub-optimal ones. Experiments on a wide range of MARL benchmarks show the superior performance of PMIC compared with other algorithms. Pengyi Li 0001, Hongyao Tang, Tianpei Yang, Xiaotian Hao, Tong Sang, Yan Zheng 0002, Jianye Hao, Matthew E. Taylor, Wenyuan Tao, Zhen Wang 0004 |
ICML | 8 |
| 2022 | Multiagent Q-learning with Sub-Team CoordinationabstractIn many real-world cooperative multiagent reinforcement learning (MARL) tasks, teams of agents can rehearse together before deployment, but then communication constraints may force individual agents to execute independently when deployed. Centralized training and decentralized execution (CTDE) is increasingly popular in recent years, focusing mainly on this setting. In the value-based MARL branch, credit assignment mechanism is typically used to factorize the team reward into each individual’s reward — individual-global-max (IGM) is a condition on the factorization ensuring that agents’ action choices coincide with team’s optimal joint action. However, current architectures fail to consider local coordination within sub-teams that should be exploited for more effective factorization, leading to faster learning. We propose a novel value factorization framework, called multiagent Q-learning with sub-team coordination (QSCAN), to flexibly represent sub-team coordination while honoring the IGM condition. QSCAN encompasses the full spectrum of sub-team coordination according to sub-team size, ranging from the monotonic value function class to the entire IGM function class, with familiar methods such as QMIX and QPLEX located at the respective extremes of the spectrum. Experimental results show that QSCAN’s performance dominates state-of-the-art methods in matrix games, predator-prey tasks, the Switch challenge in MA-Gym. Additionally, QSCAN achieves comparable performances to those methods in a selection of StarCraft II micro-management tasks. Wenhan Huang, Kai Li 0022, Kun Shao, Tianze Zhou, Matthew E. Taylor, Jun Luo 0009, Dongge Wang 0001, Hangyu Mao, Jianye Hao, Jun Wang 0012, Xiaotie Deng |
NeurIPS | 5 |
| 2022 | Cross-domain adaptive transfer reinforcement learning based on state-action correspondenceabstractDespite the impressive success achieved in various domains, deep reinforcement learning (DRL) is still faced with the sample inefficiency problem. Transfer learning (TL), which leverages prior knowledge from different but related tasks to accelerate the target task learning, has emerged as a promising direction to improve RL efficiency. The majority of prior work considers TL across tasks with the same state-action spaces, while transferring across domains with different state-action spaces is relatively unexplored. Furthermore, such existing cross-domain transfer approaches only enable transfer from a single source policy, leaving open the important question of how to best transfer from multiple source policies. This paper proposes a novel framework called Cross-domain Adaptive Transfer (CAT) to accelerate DRL. CAT learns the state-action correspondence from each source task to the target task and adaptively transfers knowledge from multiple source task policies to the target policy. CAT can be easily combined with existing DRL algorithms and experimental results show that CAT significantly accelerates learning and outperforms other cross-domain transfer methods on multiple continuous action control tasks. Heng You, Tianpei Yang, Yan Zheng 0002, Jianye Hao, Matthew E. Taylor |
UAI | 5 |
| 2022 | Multi-Agent Advisor Q-LearningabstractIn the last decade, there have been significant advances in multi-agent reinforcement learning (MARL) but there are still numerous challenges, such as high sample complexity and slow convergence to stable policies, that need to be overcome before wide-spread deployment is possible. However, many real-world environments already, in practice, deploy sub-optimal or heuristic approaches for generating policies. An interesting question that arises is how to best use such approaches as advisors to help improve reinforcement learning in multi-agent domains. In this paper, we provide a principled framework for incorporating action recommendations from online suboptimal advisors in multi-agent settings. We describe the problem of ADvising Multiple Intelligent Reinforcement Agents (ADMIRAL) in nonrestrictive general-sum stochastic game environments and present two novel Q-learning based algorithms: ADMIRAL - Decision Making (ADMIRAL-DM) and ADMIRAL - Advisor Evaluation (ADMIRAL-AE), which allow us to improve learning by appropriately incorporating advice from an advisor (ADMIRAL-DM), and evaluate the effectiveness of an advisor (ADMIRAL-AE). We analyze the algorithms theoretically and provide fixed point guarantees regarding their learning in general-sum stochastic games. Furthermore, extensive experiments illustrate that these algorithms: can be used in a variety of environments, have performances that compare favourably to other related baselines, can scale to large state-action spaces, and are robust to poor advice from advisors. Sriram Ganapathi Subramanian, Matthew E. Taylor, Kate Larson, Mark Crowley 0001 |
J. Artif. Intell. Res. | 2 |
| 2022 | Policy invariant explicit shaping: an efficient alternative to reward shapingabstractAbstract Reinforcement learning(RL) is a powerful learning paradigm in which agents can learn to maximize sparse and delayed reward signals. Although RL has had many impressive successes in complex domains, learning can take hours, days, or even years of training data. A major challenge of contemporary RL research is to discover how to learn with less data. Previous work has shown that domain information can be successfully used to shape the reward; by adding additional reward information, the agent can learn with much less data. Furthermore, if the reward is constructed from a potential function, the optimal policy is guaranteed to be unaltered. While suchpotential-based reward shaping(PBRS) holds promise, it is limited by the need for a well-defined potential function. Ideally, we would like to be able to take arbitrary advice from a human or other agent and improve performance without affecting the optimal policy. The recently introduceddynamic potential-based advice(DPBA) was proposed to tackle this challenge by predicting the potential function values as part of the learning process. However, this article demonstrates theoretically and empirically that, while DPBA can facilitate learning with good advice, it does in fact alter the optimal policy. We further show that when adding the correction term to “fix” DPBA it no longer shows effective shaping with good advice. We then present a simple method calledpolicy invariant explicit shaping(PIES) and show theoretically and empirically that PIES can use arbitrary advice, speed-up learning, and leave the optimal policy unchanged. Paniz Behboudian, Yash Satsangi, Matthew E. Taylor, Anna Harutyunyan, Michael H. Bowling |
Neural Comput. Appl. | 3 |
| 2022 | Lucid dreaming for experience replay: refreshing past states with the current policy
Yunshu Du, Garrett Warnell, Assefaw Hadish Gebremedhin, Peter Stone 0001, Matthew E. Taylor |
Neural Comput. Appl. | 5 |
| 2021 | Towered Actor Critic For Handling Multiple Action Types In Reinforcement Learning For Drug DiscoveryabstractReinforcement learning (RL) has made significant progress in both abstract and real-world domains, but the majority of state-of-the-art algorithms deal only with monotonic actions. However, some applications require agents to reason over different types of actions. Our application simulates reaction-based molecule generation, used as part of the drug discovery pipeline, and includes both uni-molecular and bi-molecular reactions. This paper introduces a novel framework, towered actor critic (TAC), to handle multiple action types. The TAC framework is general in that it is designed to be combined with any existing RL algorithms for continuous action space. We combine it with TD3 to empirically obtain significantly better results than existing methods in the drug discovery setting. TAC is also applied to RL benchmarks in OpenAI Gym and results show that our framework can improve, or at least does not hurt, performance relative to standard TD3. Sai Krishna Gottipati, Yashaswi Pathak, Boris Sattarov, Sahir, Rohan Nuttall, Matthew E. Taylor, Sarath Chandar |
AAAI | 7 |
| 2021 | Reinforcement Learning for Electronic Design Automation: Successes and OpportunitiesabstractReinforcement learning is a machine learning technique that has been applied in many domains, including robotics, game playing, and finance. This talk will briefly introduce reinforcement learning with two use cases related to compiler optimization and chip design. Interested participants will also have materials suggested to learn a more at a technical or non-technical level about this exciting tool. Matthew E. Taylor |
ISPD | 1 |
| 2020 | Providing Uncertainty-Based Advice for Deep Reinforcement Learning Agents (Student Abstract)
Felipe Leno da Silva, Pablo Hernandez-Leal, Bilal Kartal, Matthew E. Taylor |
AAAI | 4 |
| 2020 | Uncertainty-Aware Action Advising for Deep Reinforcement Learning AgentsabstractAlthough Reinforcement Learning (RL) has been one of the most successful approaches for learning in sequential decision making problems, the sample-complexity of RL techniques still represents a major challenge for practical applications. To combat this challenge, whenever a competent policy (e.g., either a legacy system or a human demonstrator) is available, the agent could leverage samples from this policy (advice) to improve sample-efficiency. However, advice is normally limited, hence it should ideally be directed to states where the agent is uncertain on the best action to execute. In this work, we propose Requesting Confidence-Moderated Policy advice (RCMP), an action-advising framework where the agent asks for advice when its epistemic uncertainty is high for a certain state. RCMP takes into account that the advice is limited and might be suboptimal. We also describe a technique to estimate the agent uncertainty by performing minor modifications in standard value-function-based RL methods. Our empirical evaluations show that RCMP performs better than Importance Advising, not receiving advice, and receiving it at random states in Gridworld and Atari Pong scenarios. Felipe Leno da Silva, Pablo Hernandez-Leal, Bilal Kartal, Matthew E. Taylor |
AAAI | 4 |
| 2020 | Curriculum Learning for Reinforcement Learning Domains: A Framework and SurveyabstractReinforcement learning (RL) is a popular paradigm for addressing sequential decision tasks in which the agent has only limited environmental feedback. Despite many advances over the past three decades, learning in many domains still requires a large amount of interaction with the environment, which can be prohibitively expensive in realistic scenarios. To address this problem, transfer learning has been applied to reinforcement learning such that experience gained in one task can be leveraged when starting to learn the next, harder task. More recently, several lines of research have explored how tasks, or data samples themselves, can be sequenced into a curriculum for the purpose of learning a problem that may otherwise be too difficult to learn from scratch. In this article, we present a framework for curriculum learning (CL) in reinforcement learning, and use it to survey and classify existing CL methods in terms of their assumptions, capabilities, and goals. Finally, we use our framework to find open problems and suggest directions for future RL curriculum learning research. Sanmit Narvekar, Bei Peng 0001, Matteo Leonetti, Jivko Sinapov, Matthew E. Taylor, Peter Stone 0001 |
J. Mach. Learn. Res. | 5 |
| 2019 | Achieving cooperation through deep multiagent reinforcement learning in sequential prisoner's dilemmasabstractThe Iterated Prisoner's Dilemma has guided research on social dilemmas for decades. However, it distinguishes between only two atomic actions: cooperate and defect. In real-world prisoner's dilemmas, these choices are temporally extended and different strategies may correspond to sequences of actions, reflecting grades of cooperation. We introduce a Sequential Prisoner's Dilemma (SPD) game to better capture the aforementioned characteristics. In this work, we propose a deep multiagent reinforcement-learning approach that investigates the evolution of mutual cooperation in SPD games. Our approach consists of two phases. The first phase is offline: it synthesizes policies with different cooperation degrees and then trains a cooperation degree detection network. The second phase is online: an agent adaptively selects its policy based on the detected degree of opponent cooperation. The effectiveness of our approach is demonstrated in two representative SPD 2D games: the Apple-Pear game and the Fruit Gathering game. Experimental results show that our strategy can avoid being exploited by exploitative opponents and achieve cooperation with cooperative opponents. Weixun Wang, Jianye Hao, Yixi Wang 0003, Matthew E. Taylor |
DAI | 4 |
| 2019 | Interactive Reinforcement Learning with Dynamic Reuse of Prior Knowledge from Human and Agent DemonstrationsabstractReinforcement learning has enjoyed multiple impressive successes in recent years. However, these successes typically require very large amounts of data before an agent achieves acceptable performance. This paper focuses on a novel way of combating such requirements by leveraging existing (human or agent) knowledge. In particular, this paper leverages demonstrations, allowing an agent to quickly achieve high performance. This paper introduces the Dynamic Reuse of Prior (DRoP) algorithm, which combines the offline knowledge (demonstrations recorded before learning) with online confidence-based performance analysis. DRoP leverages the demonstrator's knowledge by automatically balancing between reusing the prior knowledge and the current learned policy, allowing the agent to outperform the original demonstrations. We compare with multiple state-of-the-art learning algorithms and empirically show that DRoP can achieve superior performance in two domains. Additionally, we show that this confidence measure can be used to selectively request additional demonstrations, significantly improving the learning performance of the agent. Zhaodong Wang, Matthew E. Taylor |
IJCAI | 2 |
| 2019 | Metatrace Actor-Critic: Online Step-Size Tuning by Meta-gradient Descent for Reinforcement Learning ControlabstractReinforcement learning (RL) has had many successes, but significant hyperparameter tuning is commonly required to achieve good performance. Furthermore, when nonlinear function approximation is used, non-stationarity in the state representation can lead to learning instability. A variety of techniques exist to combat this --- most notably experience replay or the use of parallel actors. These techniques stabilize learning by making the RL problem more similar to the supervised setting. However, they come at the cost of moving away from the RL problem as it is typically formulated, that is, a single agent learning online without maintaining a large database of training examples. To address these issues, we propose Metatrace, a meta-gradient descent based algorithm to tune the step-size online. Metatrace leverages the structure of eligibility traces, and works for both tuning a scalar step-size and a respective step-size for each parameter. We empirically evaluate Metatrace for actor-critic on the Arcade Learning Environment. Results show Metatrace can speed up learning, and improve performance in non-stationary settings. Kenny Young, Baoxiang Wang 0001, Matthew E. Taylor |
IJCAI | 3 |
| 2019 | Towers of Saliency: A Reinforcement Learning Visualization Using Immersive EnvironmentsabstractDeep reinforcement learning (DRL) has had many successes on complex tasks, but is typically considered a black box. Opening this black box would enable better understanding and trust of the model which can be helpful for researchers and end users to better interact with the learner. In this paper, we propose a new visualization to better analyze DRL agents and present a case study using the Pommerman benchmark domain. This visualization combines two previously proven methods for improving human understanding of systems: saliency mapping and immersive visualization. Nathan Douglas, Dianna Yim, Bilal Kartal, Pablo Hernandez-Leal, Frank Maurer, Matthew E. Taylor |
ISS | 6 |
| 2019 | A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, Matthew E. Taylor |
Auton. Agents Multi Agent Syst. | 3 |
| 2019 | Analysis of University Fitness Center Data Uncovers Interesting Patterns, Enables PredictionabstractData is increasingly being used to make everyday life easier and better. Applications such as waiting time estimation, traffic prediction, and parking search are good examples of how data from different sources can be used to facilitate our daily life. In this study, we consider an under-utilized data source: university ID cards. Such cards are used on many campuses to purchase food, allow access to different areas, and even take attendance in classes. In this article, we use data from our university to analyze usage of the university fitness center and build a predictor for future visit volume. The work makes several contributions: it demonstrates the richness of the data source, shows how the data can be leveraged to improve student services, discovers interesting trends and behavior, and serves as a case study illustrating the entire data science process. Yunshu Du, Assefaw Hadish Gebremedhin, Matthew E. Taylor |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | Autonomously Reusing Knowledge in Multiagent Reinforcement LearningabstractAutonomous agents are increasingly required to solve complex tasks; hard-coding behaviors has become infeasible. Hence, agents must learn how to solve tasks via interactions with the environment. In many cases, knowledge reuse will be a core technology to keep training times reasonable, and for that, agents must be able to autonomously and consistently reuse knowledge from multiple sources, including both their own previous internal knowledge and from other agents. In this paper, we provide a literature review of methods for knowledge reuse in Multiagent Reinforcement Learning. We define an important challenge problem for the AI community, survey the existent methods, and discuss how they can all contribute to this challenging problem. Moreover, we highlight gaps in the current literature, motivating "low-hanging fruit'' for those interested in the area. Our ambition is that this paper will encourage the community to work on this difficult and relevant research challenge. Felipe Leno da Silva, Matthew E. Taylor, Anna Helena Reali Costa |
IJCAI | 2 |
| 2018 | Improving Reinforcement Learning with Human InputabstractReinforcement learning (RL) has had many successes when learning autonomously. This paper and accompanying talk consider how to make use of a non-technical human participant, when available. In particular, we consider the case where a human could 1) provide demonstrations of good behavior, 2) provide online evaluative feedback, or 3) define a curriculum of tasks for the agent to learn on. In all cases, our work has shown such information can be effectively leveraged. After giving a high-level overview of this work, we will highlight a set of open questions and suggest where future work could be usefully focused. Matthew E. Taylor |
IJCAI | 1 |
| 2017 | Scalable Multitask Policy Gradient Reinforcement LearningabstractPolicy search reinforcement learning (RL) allows agents to learn autonomously with limited feedback. However, such methods typically require extensive experience for successful behavior due to their tabula rasa nature. Multitask RL is an approach, which aims to reduce data requirements by allowing knowledge transfer between tasks. Although successful, current multitask learning methods suffer from scalability issues when considering large number of tasks. The main reasons behind this limitation is the reliance on centralized solutions. This paper proposes to a novel distributed multitask RL framework, improving the scalability across many different types of tasks. Our framework maps multitask RL to an instance of general consensus and develops an efficient decentralized solver. We justify the correctness of the algorithm both theoretically and empirically: we first proof an improvement of convergence speed to an order of O(1/k) with k being the number of iterations, and then show our algorithm surpassing others on multiple dynamical system benchmarks. Salam El Bsat, Haitham Bou-Ammar, Matthew E. Taylor |
AAAI | 3 |
| 2017 | AI Projects for Computer Science Capstone Classes (Extended Abstract)abstractCapstone senior design projects provide students with a collaborative software design and development experience to reinforce learned material while allowing students latitude in developing real-world applications. Our two-semester capstone classes are required for all computer science majors. Students must have completed a software engineering course — capstone classes are typically taken during their last two semesters. Project proposals come from a variety of sources, including industry, WSU faculty (from our own and other departments), local agencies, and entrepreneurs. We have recently targeted projects in AI — although students typically have little background, they find the ideas and methods compelling. This paper outlines our instructional approach and reports our experiences with three projects. Matthew E. Taylor, Sakire Arslan Ay |
AAAI | 1 |
| 2017 | Interactive Learning from Policy-Dependent Human FeedbackabstractThis paper investigates the problem of interactively learning behaviors communicated by a human teacher using positive and negative feedback. Much previous work on this problem has made the assumption that people provide feedback for decisions that is dependent on the behavior they are teaching and is independent from the learner’s current policy. We present empirical results that show this assumption to be false—whether human trainers give a positive or negative feedback for a decision is influenced by the learner’s current policy. Based on this insight, we introduce Convergent Actor-Critic by Humans (COACH), an algorithm for learning from policy-dependent feedback that converges to a local optimum. Finally, we demonstrate that COACH can successfully learn multiple behaviors on a physical robot. James MacGlashan, Mark K. Ho, Robert Tyler Loftin, Bei Peng 0001, David L. Roberts 0001, Matthew E. Taylor, Michael L. Littman |
ICML | 7 |
| 2017 | Leveraging Human Knowledge in Tabular Reinforcement Learning: A Study of Human Subjects
Ariel Rosenfeld, Matthew E. Taylor, Sarit Kraus |
IJCAI | 2 |
| 2017 | Improving Reinforcement Learning with Confidence-Based DemonstrationsabstractReinforcement learning has had many successes, but in practice it often requires significant amounts of data to learn high-performing policies. One common way to improve learning is to allow a trained (source) agent to assist a new (target) agent. The goals in this setting are to 1) improve the target agent's performance, relative to learning unaided, and 2) allow the target agent to outperform the source agent. Our approach leverages source agent demonstrations, removing any requirements on the source agent's learning algorithm or representation. The target agent then estimates the source agent's policy and improves upon it. The key contribution of this work is to show that leveraging the target agent's uncertainty in the source agent's policy can significantly improve learning in two complex simulated domains, Keepaway and Mario. Zhaodong Wang, Matthew E. Taylor |
IJCAI | 2 |
| 2017 | Efficiently detecting switches against non-stationary opponents
Pablo Hernandez-Leal, Yusen Zhan, Matthew E. Taylor, Luis Enrique Sucar, Enrique Munoz de Cote |
Auton. Agents Multi Agent Syst. | 3 |
| 2017 | An exploration strategy for non-stationary opponents
Pablo Hernandez-Leal, Yusen Zhan, Matthew E. Taylor, Luis Enrique Sucar, Enrique Munoz de Cote |
Auton. Agents Multi Agent Syst. | 3 |
| 2017 | Multi-objectivization and ensembles of shapings in reinforcement learning
Tim Brys, Anna Harutyunyan, Peter Vrancx, Ann Nowé, Matthew E. Taylor |
Neurocomputing | 5 |
| 2017 | Nonconvex Policy Search Using Variational InequalitiesabstractPolicy search is a class of reinforcement learning algorithms for finding optimal policies in control problems with limited feedback. These methods have been shown to be successful in high-dimensional problems such as robotics control. Though successful, current methods can lead to unsafe policy parameters that potentially could damage hardware units. Motivated by such constraints, we propose projection-based methods for safe policies. These methods, however, can handle only convex policy constraints. In this letter, we propose the first safe policy search reinforcement learner capable of operating under nonconvex policy constraints. This is achieved by observing, for the first time, a connection between nonconvex variational inequalities and policy search problems. We provide two algorithms, Mann and two-step iteration, to solve the above problems and prove convergence in the nonconvex stochastic setting. Finally, we demonstrate the performance of the algorithms on six benchmark dynamical systems and show that our new method is capable of outperforming previous methods under a variety of settings. Yusen Zhan, Haitham Bou-Ammar, Matthew E. Taylor |
Neural Comput. | 3 |
| 2017 | Scalable lifelong reinforcement learning
Yusen Zhan, Haitham Bou-Ammar, Matthew E. Taylor |
Pattern Recognit. | 3 |
| 2016 | Theoretically-Grounded Policy Advice from Multiple Teachers in Reinforcement Learning Settings with Applications to Negative Transfer
Yusen Zhan, Haitham Bou-Ammar, Matthew E. Taylor |
IJCAI | 3 |
| 2016 | Lifelong learning for disturbance rejection on mobile robotsabstractNo two robots are exactly the same-even for a given model of robot, different units will require slightly different controllers. Furthermore, because robots change and degrade over time, a controller will need to change over time to remain optimal. This paper leverages lifelong learning in order to learn controllers for different robots. In particular, we show that by learning a set of control policies over robots with different (unknown) motion models, we can quickly adapt to changes in the robot, or learn a controller for a new robot with a unique set of disturbances. Furthermore, the approach is completely model-free, allowing us to apply this method to robots that have not, or cannot, be fully modeled. David Isele, José-Marcio Luna, Eric Eaton, Gabriel Victor de la Cruz, James Irwin, Brandon Kallaher, Matthew E. Taylor |
IROS | 7 |
| 2016 | Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
Robert Tyler Loftin, Bei Peng 0001, James MacGlashan, Michael L. Littman, Matthew E. Taylor, Jeff Huang 0002, David L. Roberts 0001 |
Auton. Agents Multi Agent Syst. | 5 |
| 2015 | Unsupervised Cross-Domain Transfer in Policy Gradient Reinforcement Learning via Manifold AlignmentabstractThe success of applying policy gradient reinforcement learning (RL) to difficult control tasks hinges crucially on the ability to determine a sensible initialization for the policy. Transfer learning methods tackle this problem by reusing knowledge gleaned from solving other related tasks. In the case of multiple task domains, these algorithms require an inter-task mapping to facilitate knowledge transfer across domains. However, there are currently no general methods to learn an inter-task mapping without requiring either background knowledge that is not typically present in RL settings, or an expensive analysis of an exponential number of inter-task mappings in the size of the state and action spaces. This paper introduces an autonomous framework that uses unsupervised manifold alignment to learn inter-task mappings and effectively transfer samples between different task domains. Empirical results on diverse dynamical systems, including an application to quadrotor control, demonstrate its effectiveness for cross-domain transfer in the context of policy gradient RL. Haitham Bou-Ammar, Eric Eaton, Paul Ruvolo, Matthew E. Taylor |
AAAI | 4 |
| 2015 | Reinforcement Learning from Demonstration through Shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E. Taylor, Ann Nowé |
IJCAI | 5 |
| 2014 | Combining Multiple Correlated Reward and Shaping Signals by Measuring ConfidenceabstractMulti-objective problems with correlated objectives are a class of problems that deserve specific attention. In contrast to typical multi-objective problems, they do not require the identification of trade-offs between the objectives, as (near-) optimal solutions for any objective are (near-) optimal for every objective. Intelligently combining the feedback from these objectives, instead of only looking at a single one, can improve optimization. This class of problems is very relevant in reinforcement learning, as any single-objective reinforcement learning problem can be framed as such a multi-objective problem using multiple reward shaping functions. After discussing this problem class, we propose a solution technique for such reinforcement learning problems, called adaptive objective selection. This technique makes a temporal difference learner estimate the Q-function for each objective in parallel, and introduces a way of measuring confidence in these estimates. This confidence metric is then used to choose which objective's estimates to use for action selection. We show significant improvements in performance over other plausible techniques on two problem domains. Finally, we provide an intuitive analysis of the technique's decisions, yielding insights into the nature of the problems being solved. Tim Brys, Ann Nowé, Daniel Kudenko, Matthew E. Taylor |
AAAI | 4 |
| 2014 | A Strategy-Aware Technique for Learning Behaviors from Discrete Human FeedbackabstractThis paper introduces two novel algorithms for learning behaviors from human-provided rewards. The primary novelty of these algorithms is that instead of treating the feedback as a numeric reward signal, they interpret feedback as a form of discrete communication that depends on both the behavior the trainer is trying to teach and the teaching strategy used by the trainer. For example, some human trainers use a lack of feedback to indicate whether actions are correct or incorrect, and interpreting this lack of feedback accurately can significantly improve learning speed. Results from user studies show that humans use a variety of training strategies in practice and both algorithms can learn a contextual bandit task faster than algorithms that treat the feedback as numeric. Simulated trainers are also employed to evaluate the algorithms in both contextual bandit and sequential decision-making tasks with similar results. Robert Tyler Loftin, James MacGlashan, Bei Peng 0001, Matthew E. Taylor, Michael L. Littman, Jeff Huang 0002, David L. Roberts 0001 |
AAAI | 4 |
| 2014 | Using Ensemble Techniques and Multi-Objectivization to Solve Reinforcement Learning ProblemsabstractRecent work on multi-objectivization has shown how a single-objective reinforcement learning problem can be turned into a multi-objective problem with correlated objectives, by providing multiple reward shaping functions. The information contained in these correlated objectives can be exploited to solve the base, single-objective problem faster and better, given techniques specifically aimed at handling such correlated objectives. In this paper, we identify ensemble techniques as a set of methods that is suitable to solve multi-objectivized reinforcement learning problems. We empirically demonstrate their use on the Pursuit domain. Tim Brys, Matthew E. Taylor, Ann Nowé |
ECAI | 2 |
| 2014 | Online Multi-Task Learning for Policy Gradient MethodsabstractPolicy gradient algorithms have shown considerable recent success in solving high-dimensional sequential decision making tasks, particularly in robotics. However, these methods often require extensive experience in a domain to achieve high performance. To make agents more sample-efficient, we developed a multi-task policy gradient method to learn decision making tasks consecutively, transferring knowledge between tasks to accelerate learning. Our approach provides robust theoretical guarantees, and we show empirically that it dramatically accelerates learning on a variety of dynamical systems, including an application to quadrotor control. Haitham Bou-Ammar, Eric Eaton, Paul Ruvolo, Matthew E. Taylor |
ICML | 4 |
| 2014 | Multi-objectivization of reinforcement learning problems by reward shapingabstractMulti-objectivization is the process of transforming a single objective problem into a multi-objective problem. Research in evolutionary optimization has demonstrated that the addition of objectives that are correlated with the original objective can make the resulting problem easier to solve compared to the original single-objective problem. In this paper we investigate the multi-objectivization of reinforcement learning problems. We propose a novel method for the multi-objectivization of Markov Decision problems through the use of multiple reward shaping functions. Reward shaping is a technique to speed up reinforcement learning by including additional heuristic knowledge in the reward signal. The resulting composite reward signal is expected to be more informative during learning, leading the learner to identify good actions more quickly. Good reward shaping functions are by definition correlated with the target value function for the base reward signal, and we show in this paper that adding several correlated signals can help to solve the basic single objective problem faster and better. We prove that the total ordering of solutions, and by consequence the optimality of solutions, is preserved in this process, and empirically demonstrate the usefulness of this approach on two reinforcement learning tasks: a pathfinding problem and the Mario domain. Tim Brys, Anna Harutyunyan, Peter Vrancx, Matthew E. Taylor, Daniel Kudenko, Ann Nowé |
IJCNN | 4 |
| 2014 | Agents Teaching Agents in Reinforcement Learning (Nectar Abstract)
Matthew E. Taylor, Lisa Torrey |
ECML/PKDD (3) | 1 |
| 2014 | Learning something from nothing: Leveraging implicit human feedback strategiesabstractIn order to be useful in real-world situations, it is critical to allow non-technical users to train robots. Existing work has considered the problem of a robot or virtual agent learning behaviors from evaluative feedback provided by a human trainer. That work, however, has treated feedback as a numeric reward that the agent seeks to maximize, and has assumed that all trainers will provide feedback in the same way when teaching the same behavior. We report the results of a series of user studies that indicate human trainers use a variety of approaches to providing feedback in practice, which we describe as different “training strategies.” For example, users may not always give explicit feedback in response to an action, and may be more likely to provide explicit reward than explicit punishment, or vice versa. If the trainer is consistent in their strategy, then it may be possible to infer knowledge about the desired behavior from cases where no explicit feedback is provided. We discuss a probabilistic model of human-provided feedback that can be used to classify these different training strategies based on when the trainer chooses to provide explicit reward and/or explicit punishment, and when they choose to provide no feedback. Additionally, we investigate how training strategies may change in response to the appearance of the learning agent. Ultimately, based on this work, we argue that learning agents designed to understand and adapt to different users' training strategies will allow more efficient and intuitive learning experiences. Robert Tyler Loftin, Bei Peng 0001, James MacGlashan, Michael L. Littman, Matthew E. Taylor, Jeff Huang 0002, David L. Roberts 0001 |
RO-MAN | 5 |
| 2014 | Distributed learning and multi-objectivity in traffic light controlabstractTraffic jams and suboptimal traffic flows are ubiquitous in modern societies, and they create enormous economic losses each year. Delays at traffic lights alone account for roughly 10% of all delays in US traffic. As most traffic light scheduling systems currently in use are static, set up by human experts rather than being adaptive, the interest in machine learning approaches to this problem has increased in recent years. Reinforcement learning (RL) approaches are often used in these studies, as they require little pre-existing knowledge about traffic flows. Distributed constraint optimisation approaches (DCOP) have also been shown to be successful, but are limited to cases where the traffic flows are known. The distributed coordination of exploration and exploitation (DCEE) framework was recently proposed to introduce learning in the DCOP framework. In this paper, we present a study of DCEE and RL techniques in a complex simulator, illustrating the particular advantages of each, comparing them against standard isolated traffic actuated signals. We analyse how learning and coordination behave under different traffic conditions, and discuss the multi-objective nature of the problem. Finally we evaluate several alternative reward signals in the best performing approach, some of these taking advantage of the correlation between the problem-inherent objectives to improve performance. Tim Brys, Tong T. Pham, Matthew E. Taylor |
Connect. Sci. | 3 |
| 2014 | Reinforcement learning agents providing advice in complex video gamesabstractThis article introduces a teacher–student framework for reinforcement learning, synthesising and extending material that appeared in conference proceedings [Torrey, L., & Taylor, M. E. (2013)]. Teaching on a budget: Agents advising agents in reinforcement learning. {Proceedings of the international conference on autonomous agents and multiagent systems}] and in a non-archival workshop paper [Carboni, N., &Taylor, M. E. (2013, May)]. Preliminary results for 1 vs. 1 tactics in StarCraft. {Proceedings of the adaptive and learning agents workshop (at AAMAS-13)}]. In this framework, a teacher agent instructs a student agent by suggesting actions the student should take as it learns. However, the teacher may only give such advice a limited number of times. We present several novel algorithms that teachers can use to budget their advice effectively, and we evaluate them in two complex video games: StarCraft and Pac-Man. Our results show that the same amount of advice, given at different moments, can have different effects on student learning, and that teachers can significantly affect student learning even when students use different learning methods and state representations. Matthew E. Taylor, Nicholas Carboni, Anestis Fachantidis, Ioannis P. Vlahavas, Lisa Torrey |
Connect. Sci. | 1 |
| 2013 | Automatically Mapped Transfer between Reinforcement Learning Tasks via Three-Way Restricted Boltzmann Machines
Haitham Bou-Ammar, Decebal Constantin Mocanu, Matthew E. Taylor, Kurt Driessens, Karl Tuyls, Gerhard Weiss 0001 |
ECML/PKDD (2) | 3 |
| 2013 | Mitigating multi-path fading in a mobile mesh network
Marcos A. M. Vieira, Matthew E. Taylor, Prateek Tandon 0002, Ramesh Govindan, Gaurav S. Sukhatme, Milind Tambe |
Ad Hoc Networks | 2 |
| 2011 | Protecting against evaluation overfitting in empirical reinforcement learningabstractEmpirical evaluations play an important role in machine learning. However, the usefulness of any evaluation depends on the empirical methodology employed. Designing good empirical methodologies is difficult in part because agents can overfit test evaluations and thereby obtain misleadingly high scores. We argue that reinforcement learning is particularly vulnerable to environment overfitting and propose as a remedy generalized methodologies, in which evaluations are based on multiple environments sampled from a distribution. In addition, we consider how to summarize performance when scores from different environments may not have commensurate values. Finally, we present proof-of-concept results demonstrating how these methodologies can validate an intuitively useful range-adaptive tile coding method. Shimon Whiteson, Brian Tanner, Matthew E. Taylor, Peter Stone 0001 |
ADPRL | 3 |
| 2010 | Evolving Compiler Heuristics to Manage Communication and ContentionabstractAs computer architectures become increasingly complex, hand-tuning compiler heuristics becomes increasingly tedious and time consuming for compiler developers. This paper presents a case study that uses a genetic algorithm to learn a compiler policy. The target policy implicitly balances communication and contention among processing elements of the TRIPS processor, a physically realized prototype chip. We learn specialized policies for individual programs as well as general policies that work well across all programs. We also employ a two-stage method that first classifies the code being compiled based on salient characteristics, and then chooses a specialized policy based on that classification.This work is particularly interesting for the AI community because it 1) emphasizes the need for increased collaboration between AI researchers and researchers from other branches of computer science and 2) discusses a machine learning setup where training on the custom hardware requires weeks of training, rather than the more typical minutes or hours. Matthew E. Taylor, Katherine E. Coons, Behnam Robatmili, Bertrand A. Maher, Doug Burger, Kathryn S. McKinley |
AAAI | 1 |
| 2010 | Critical factors in the empirical performance of temporal difference and evolutionary methods for reinforcement learningabstractTemporal difference and evolutionary methods are two of the most common approaches to solving reinforcement learning problems. However, there is little consensus on their relative merits and there have been few empirical studies that directly compare their performance. This article aims to address this shortcoming by presenting results of empirical comparisons between Sarsa and NEAT, two representative methods, in mountain car and keepaway, two benchmark reinforcement learning tasks. In each task, the methods are evaluated in combination with both linear and nonlinear representations to determine their best configurations. In addition, this article tests two specific hypotheses about the critical factors contributing to these methods’ relative performance: (1) that sensor noise reduces the final performance of Sarsa more than that of NEAT, because Sarsa’s learning updates are not reliable in the absence of the Markov property and (2) that stochasticity, by introducing noise in fitness estimates, reduces the learning speed of NEAT more than that of Sarsa. Experiments in variations of mountain car and keepaway designed to isolate these factors confirm both these hypotheses. Shimon Whiteson, Matthew E. Taylor, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 2 |
| 2009 | DCOPs Meet the Real World: Exploring Unknown Reward Matrices with Applications to Mobile Sensor Networks
Matthew E. Taylor, Milind Tambe, Makoto Yokoo |
IJCAI | 2 |
| 2009 | Transfer Learning for Reinforcement Learning Domains: A Survey
Matthew E. Taylor, Peter Stone 0001 |
J. Mach. Learn. Res. | 1 |
| 2008 | Feature selection and policy optimization for distributed instruction placement using reinforcement learningabstractCommunication overheads are one of the fundamental challenges in a multiprocessor system. As the number of processors on a chip increases, communication overheads and the distribution of computation and data become increasingly important performance factors. Explicit Dataflow Graph Execution (EDGE) processors, in which instructions communicate with one another directly on a distributed substrate, give the compiler control over communication overheads at a fine granularity. Prior work shows that compilers can effectively reduce fine-grained communication overheads in EDGE architectures using a spatial instruction placement algorithm with a heuristic-based cost function. While this algorithm is effective, the cost function must be painstakingly tuned. Heuristics tuned to perform well across a variety of applications leave users with little ability to tune performance-critical applications, yet we find that the best placement heuristics vary significantly with the application. Katherine E. Coons, Behnam Robatmili, Matthew E. Taylor, Bertrand A. Maher, Doug Burger, Kathryn S. McKinley |
PACT | 3 |
| 2008 | Transferring Instances for Model-Based Reinforcement Learning
Matthew E. Taylor, Nicholas K. Jong, Peter Stone 0001 |
ECML/PKDD (2) | 1 |
| 2007 | Autonomous Inter-Task Transfer in Reinforcement Learning Domains
Matthew E. Taylor |
AAAI | 1 |
| 2007 | Representation Transfer via Elaboration
Matthew E. Taylor, Peter Stone 0001 |
AAAI | 1 |
| 2007 | Temporal Difference and Policy Search Methods for Reinforcement Learning: An Empirical Comparison
Matthew E. Taylor, Shimon Whiteson, Peter Stone 0001 |
AAAI | 1 |
| 2007 | Cross-domain transfer for reinforcement learningabstractA typical goal for transfer learning algorithms is to utilize knowledge gained in a source task to learn a target task faster. Recently introduced transfer methods in reinforcement learning settings have shown considerable promise, but they typically transfer between pairs of very similar tasks. This work introduces Rule Transfer, a transfer algorithm that first learns rules to summarize a source task policy and then leverages those rules to learn faster in a target task. This paper demonstrates that Rule Transfer can effectively speed up learning in Keepaway, a benchmark RL problem in the robot soccer domain, based on experience from source tasks in the gridworld domain. We empirically show, through the use of three distinct transfer metrics, that Rule Transfer is effective across these domains. Matthew E. Taylor, Peter Stone 0001 |
ICML | 1 |
| 2007 | Transfer Learning via Inter-Task Mappings for Temporal Difference Learning
Matthew E. Taylor, Peter Stone 0001 |
J. Mach. Learn. Res. | 1 |
| 2006 | Inter-Task Action Correlation for Reinforcement Learning Tasks
Matthew E. Taylor, Peter Stone 0001 |
AAAI | 1 |
| 2006 | Comparing evolutionary and temporal difference methods in a reinforcement learning domainabstractBoth genetic algorithms (GAs) and temporal difference (TD) methods have proven effective at solving reinforcement learning (RL) problems. However, since few rigorous empirical comparisons have been conducted, there are no general guidelines describing the methods' relative strengths and weaknesses. This paper presents the results of a detailed empirical comparison between a GA and a TD method in Keepaway, a standard RL benchmark domain based on robot soccer. In particular, we compare the performance of NEAT [19], a GA that evolves neural networks, with Sarsa [16, 17], a popular TD method. The results demonstrate that NEAT can learn better policies in this task, though it requires more evaluations to do so. Additional experiments in two variations of Keepaway demonstrate that Sarsa learns better policies when the task is fully observable and NEAT learns faster when the task is deterministic. Together, these results help isolate the factors critical to the performance of each method and yield insights into their general strengths and weaknesses. Matthew E. Taylor, Shimon Whiteson, Peter Stone 0001 |
GECCO | 1 |
| 2005 | Value Functions for RL-Based Behavior Transfer: A Comparative Study
Matthew E. Taylor, Peter Stone 0001 |
AAAI | 1 |
| 2005 | Keepaway Soccer: From Machine Learning Testbed to Benchmark
Peter Stone 0001, Gregory Kuhlmann, Matthew E. Taylor |
RoboCup | 3 |