EDBT 2026 Demo / reviewers in the wild / expert
Kagan Tumer
dblp:22/378
· DBLP profile ↗
77ranked-venue papers
12as first author
19since 2021 · last 2025
0009-0007-3809-7257ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 12 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 3 since 2021Systems, architecture and hardware · 5 · 1 since 2021Computer networks · 3Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multiagent Quality-Diversity for Effective AdaptationabstractRobust adaptation in multiagent settings requires learning not just a single optimal behavior, but a repertoire of high-performing and diverse team behaviors that can succeed under environmental contingencies. Traditional multiagent reinforcement learning methods typically converge to a single specialized team behavior, limiting their adaptability. Recent approaches like Mix-ME promote behavioral diversity but rely solely on evolutionary operators, often resulting in sample-inefficiency and uncoordinated team composition. This work introduces Multiagent Sample-Efficient Quality-Diversity (MASQD), a learning framework that produces an archive of diverse, high-performing multiagent teams. MASQD builds on the Cross-Entropy Method Reinforcement Learning algorithm and extends it to the multiagent setting by representing teams as parameter-shared neural networks, directing exploration from previously discovered behaviors, and guiding refinement through a descriptor-conditioned critic. Through this coupling of anchored exploration and targeted exploitation, MASQD produces functional diversity: teams that are not only behaviorally distinct but also robust and effective under varied conditions. Experiments across four Multiagent MuJoCo tasks show that MASQD outperforms state-of-the-art baselines in both team fitness and functional diversity. Siddarth Iyer, Ayhan Alp Aydeniz, Kagan Tumer |
ECAI | 4 |
| 2025 | Dynamic Influence For Coevolutionary AgentsabstractMultiagent settings are naturally characterized by coevolutionary dynamics, where agents must adapt and learn in the context of their teammates. A key challenge in such domains is determining how to credit an individual agent for their contribution to team performance. Fitness shaping approaches partially address this by identifying and isolating an agent's direct contribution to the team's success. However, when an agent's contribution is indirect—such as influencing other teammates to succeed—existing methods fail to account for its influence on the team. This paper introduces Dynamic Influence, a fitness shaping method for heterogeneous teams that isolates both direct and indirect contributions by evaluating how agents influence others over time. By considering inter-agent influence at a high temporal resolution, Dynamic Influence-Based Fitness Shaping allows agents to distill and extract direct credit from indirect interactions. Results in an autonomous aerial and terrestrial vehicle coordination problem demonstrate the efficacy of Dynamic Influence-Based Fitness Shaping, achieving superior cooperative behaviors compared to several static fitness shaping baselines. Everardo Gonzalez, Kagan Tumer |
GECCO | 3 |
| 2025 | Multiagent Credit Assignment for Multi-Objective CoordinationabstractMany real-world coordination tasks—such as environmental monitoring, traffic management, and underwater exploration—are best modelled as multiagent problems with multiple, often conflicting objectives. Achieving effective coordination in these settings requires addressing two main challenges: 1) balancing multiple objectives and 2) resolving the credit assignment problem to isolate each agent's contribution from team-level feedback. Existing multiagent credit assignment methods collapse multi-objective reward vectors into a single scalar—potentially overlooking nuanced trade-offs. In this paper, we introduce the Multi-Objective Difference Evaluation (DMO) operator to assign agent-level credit without a priori scalarisation. DMO measures the change in hypervolume when an agent's policy is replaced by a counterfactual default, capturing how much that policy contributes to each objective and to the Pareto front. We embed DMO into the popular NSGA-II algorithm to evolve a population of joint policies with distinct trade-offs. Empirical results on the Multi-Objective Beach Problem and the Multi-Objective Rover Exploration domain show that our approach matches or surpasses existing baselines, delivering up to a 33% performance improvement. Raghav Thakar, Siddarth Iyer, Kagan Tumer |
GECCO | 4 |
| 2025 | Safe Entropic Agents under Team Constraints
Ayhan Alp Aydeniz, Enrico Marchesini, Robert Tyler Loftin, Christopher Amato, Kagan Tumer |
AAMAS | 5 |
| 2025 | Multi-objective reinforcement learning framework for beneficent artificial intelligence
Anna Nickelson, Russell Perkins, Alex John London, Paul Robinette, Kagan Tumer |
Neural Comput. Appl. | 5 |
| 2024 | Objective-Informed Diversity for Multi-Objective Multiagent CoordinationabstractTo coordinate in multiagent settings characterized by multiple objectives, asymmetric agents (agents with distinct capabilities and preferences) must learn diverse behaviors to balance trade-offs between agent-specific and team objectives. Hierarchical methods partially address this by leveraging a combination of Quality-Diversity methods that illuminate the behavior space and evolutionary algorithms that use non-dominated sorting over the explored behaviors to improve coverage in the objective space. However, optimizing diverse behaviors and trade-offs in isolation is susceptible to producing egocentric behaviors that favor agent-specific objectives at the cost of team objectives. This work introduces the Multi-Objective Informed Island Model (MOI-IM), an asymmetric multiagent learning framework that fosters diverse behaviors and rich inter-agent relationships, necessary to balance potentially conflicting and misaligned objectives. An evolutionary algorithm improves coverage in the objective space by evolving a population of teams, while a gradient-based optimization infers and progressively explores the behavior space by fluidly adapting search to regions that produce policies with non-dominated trade-offs. The two processes are coupled via shared replay buffers to ensure alignment between coverage in the behavior and objective space. Empirical results on an asymmetric multi-objective coordination problem highlight MOI-IM’s ability to produce teams that can express diverse trade-offs and robust relationships required to balance misaligned objectives. Kagan Tumer |
ECAI | 2 |
| 2024 | Learning Aligned Local Evaluations For Better Credit Assignment In Cooperative CoevolutionabstractCooperative coevolutionary algorithms prove effective in solving tasks that can be easily decoupled into subproblems. When applied to problems with high coupling (where the fitness depends heavily on specific joint actions), evolution is often stifled by the credit assignment problem. This is due to each agent evolving their policy using a shared evaluation function that is sensitive to the "noise" of all other agents' actions. Using fitness critics alleviates this problem by approximating a local model of an agent's contribution and using that signal as a fitness function. However, fitness critics suffer when the quality of the local approximation degrades. In this work, we present Global Aligned Local Error (GALE), a loss function that generates better credit-assigning local evaluations that aim to maximize the alignment of the local and global evaluations. In a multiagent exploration domain, we show GALE's ability to learn better credit assignment, which leads to improved teaming behavior. Joshua Cook, Kagan Tumer |
GECCO | 2 |
| 2024 | Informed Diversity Search for Learning in Asymmetric Multiagent SystemsabstractTo coordinate in multiagent settings, asymmetric agents (agents with distinct objectives and capabilities) must learn diverse behaviors that allow them to maximize their individual and team objectives. Hierarchical learning techniques partially address this by leveraging a combination of Quality-Diversity to learn diverse agent-specific behaviors and evolutionary optimization to maximize team objectives. However, isolating diversity search from team optimization is prone to producing egocentric behaviors that have misaligned objectives. This work introduces Diversity Aligned Island Model (DA-IM), a coevolutionary framework that fluidly adapts diversity search to focus on behaviors that yield high fitness teams. An evolutionary algorithm evolves a population of teams to optimize the team objective. Concurrently, a combination of gradient-based optimizers utilize experiences collected by the teams to reinforce agent-specific behaviors and selectively mutate them based on their fitness on the team objective. Periodically, the mutated policies are added to the evolutionary population to inject diversity and to ensure alignment between the two processes. Empirical evaluations on two asymmetric coordination problems with varying degrees of alignment highlight DA-IM's ability to produce diverse behaviors that outperform existing population-based diversity search methods. Kagan Tumer |
GECCO | 2 |
| 2024 | Reinforcing Inter-Class Dependencies in the Asymmetric Island ModelabstractMultiagent learning allows agents to learn cooperative behaviors necessary to accomplish team objectives. However, coordination requires agents to learn diverse behaviors that work well as part of a team, a task made more difficult by all agents simultaneously learning their own individual behaviors. This is made more challenging when there are multiple classes of asymmetric agents in the system with differing capabilities that work together as a team. The Asymmetric Island Model alleviates these difficulties by simultaneously optimizing for class-specific and team-wide behaviors as independent processes that enable agents to discover and refine optimal joint-behaviors. However, agents learn to optimize agent-specific behaviors in isolation from other agent classes, leading them to learn egocentric behaviors that are potentially sub-optimal when paired with other agent classes. This work introduces Reinforced Asymmetric Island Model (RAIM), a framework for explicitly reinforcing closely dependent inter-class agent behaviors. When optimizing the class-specific behaviors, agents learn alongside stationary representations of other classes, allowing them to efficiently optimize class-specific behaviors that are conditioned on the expectation of the behaviors of the complementary agent classes. Experiments in an asymmetric harvest environment highlight the effectiveness of our method in learning robust inter-agent behaviors that can adapt to diverse environment dynamics. Andrew Festa, Kagan Tumer |
GECCO | 3 |
| 2024 | Influence Based Fitness Shaping for Coevolutionary AgentsabstractCoevolving cooperative teams creates a challenging joint-action discovery problem because fitness functions generally evaluate team performance rather than individual agent performance. Feedback "sparsity" where agents only receive feedback when they jointly stumble upon a valuable action compounds this problem. Fitness shaping techniques alleviate this problem by extracting agent contributions and providing stepping stone incentives in sparse feedback settings. However, such techniques require agents to make direct and measurable impacts to system performance. If agents have indirect impacts, such as influencing other agents to accomplish tasks, existing shaping methods fail to provide adequate feedback. In this work, we introduce Influence Based Fitness Shaping (IBFS) to capture and incentivize indirect impacts. IBFS extracts an agent's impact based on how it influences other agents and guides exploration towards influencing actions in sparse feedback settings. Our results in a multiagent shepherding problem show that IBFS outperforms standard fitness shaping, and the gains increase when feedback becomes sparser. Everardo Gonzalez, Siddarth Viswanathan, Kagan Tumer |
GECCO | 3 |
| 2023 | Knowledge Injection for Multiagent Systems via Counterfactual Perception ShapingabstractReward shaping can be used to train coordinated agent teams, but most learning approaches optimize for training conditions and by design, are limited by knowledge directly captured by the reward function. Advances in adaptive systems (e.g., transfer learning) may enable agents to quickly learn new policies in response to changing conditions, but retraining agents is both difficult and risks losing team coordination altogether. In this work we introduce Counterfactual Knowledge Injection (CKI), a novel approach to injecting high-level information into a multiagent system outside of the learning process. CKI encodes knowledge into counterfactual state representations to shape agent perceptions of the system so that their current policies better match the current system conditions. We demonstrate CKI in a multiagent exploration task where agents must collaborate to observe various Points of Interest (POI). We show that CKI successfully imparts high-level system knowledge to agents in response to imperceptible changes. We also show that CKI enables agents to adjust their level of agent-to-agent coordination ranging from tasks individuals can complete up to tasks that require the entire team. Nicholas Zerbel, Kagan Tumer |
ECAI | 2 |
| 2023 | Novelty Seeking Multiagent Evolutionary Reinforcement LearningabstractCoevolving teams of agents promises effective solutions for many coordination tasks such as search and rescue missions or deep ocean exploration. Good team performance in such domains generally relies on agents discovering complex joint policies, which is particularly difficult when the fitness functions are sparse (where many joint policies return the same or even zero fitness values). In this paper, we introduce Novelty Seeking Multiagent Evolutionary Reinforcement Learning (NS-MERL), which enables agents to more efficiently explore their joint strategy space. The key insight of NS-MERL is to promote good exploratory behaviors for individual agents using a dense, novelty-based fitness function. Though the overall team-level performance is still evaluated via a sparse fitness function, agents using NS-MERL more efficiently explore their joint action space and more readily discover good joint policies. Our results in complex coordination tasks show that teams of agents trained with NS-MERL perform significantly better than agents trained solely with task-specific fitnesses. Ayhan Alp Aydeniz, Robert Tyler Loftin, Kagan Tumer |
GECCO | 3 |
| 2023 | Leveraging Fitness Critics To Learn Robust TeamworkabstractCo-evolutionary algorithms have successfully trained agent teams for tasks such as autonomous exploration or robot soccer. However generally, such approaches seek a single strong team, whereas many real-world applications require agents to effectively cooperate across multiple teams. To adapt to different teammates, agents need to learn more general teamwork skills rather than a single team-specific role. Previous work primarily frames this as a fitness-shaping problem, providing high-quality but expensive evaluation methods to isolate an agent's contribution. In this work, we introduce Learned Evaluations for Robust Teaming (LERT), an approach that provides a local evaluation that leverages state trajectories of agents to better quantify their impact across multiple teams. The key insight of this work is that agent state trajectories and previous experiences carry sufficient information to map agent abilities to team performance. As a result, LERT cooperatively co-evolves agents to work together across arbitrary teams. While only using local information and significantly fewer team evaluations, LERT performs as well as-if not better than-current methods. Joshua Cook, Kagan Tumer, Tristan Scheiner |
GECCO | 2 |
| 2023 | Learning Synergies for Multi-Objective Optimization in Asymmetric Multiagent SystemsabstractAgents in a multiagent system must learn diverse policies that allow them to express complex inter-agent relationships required to optimize a single team objective. Multiagent Quality Diversity methods partially address this by transforming the agents' large joint policy space to a tractable sub-space that can produce synergistic agent policies. However, a majority of real-world problems are inherently multi-objective and require asymmetric agents (agents with different capabilities and objectives) to learn policies that represent diverse trade-offs between agent-specific and team objectives. This work introduces Multi-objective Asymmetric Island Model (MO-AIM), a multi-objective multiagent learning framework for the discovery of generalizable agent synergies and trade-offs that is based on adapting the population dynamics over a spectrum of tasks. The key insight is that the competitive pressure arising from the changing populations on the team tasks forces agents to acquire robust synergies required to balance their individual and team objectives in response to the nature of their teams and task dynamics. Results on several variations of a multi-objective habitat problem highlight the potential of MO-AIM in producing teams with diverse specializations and trade-offs that readily adapt to unseen tasks. Kagan Tumer |
GECCO | 2 |
| 2023 | Contextual Multi-Objective Path PlanningabstractMany critical robot environments, such as healthcare and security, require robots to account for contextdependent criteria when performing their functions (e.g., navigation). Such domains require decisions that balance multiple factors, making it difficult for robots to make contextually appropriate decisions. Multi-Objective Optimization (MOO) methods offer a potential solution by trading off between objectives; however concepts like Pareto fronts are not only expensive to compute but struggle with differentiating among solutions on the Pareto front. This work introduces the Contextual Multi-Objective Path Planning (CMOPP) algorithm, which enables the robot to trade off different complex costs dependent on context. The key insight of this work is to separate the path planning and path cost estimation into two independent steps, thus significantly reducing computation cost without impacting the quality of the resulting path. As a result, CMOPP is able to accurately model path costs, which provide meaningful trade-offs when choosing a path that best fits the context. We show the benefits of CMOPP on case studies that demonstrate its contextual path planning capabilities. CMOPP finds contextually appropriate paths by first reducing the search space up to 99.9% to a near-optimal set of paths. This reduction enables the generation of accurate path cost models, using up to 90% less computation than similar methods. Anna Nickelson, Kagan Tumer, William D. Smart |
ICRA | 2 |
| 2023 | Shaping the Behavior Space with Counterfactual Agents in Multi-Objective Map Elites
Anna Nickelson, Nicholas Zerbel, Kagan Tumer |
IJCCI | 4 |
| 2022 | Fitness shaping for multiple teamsabstractCoevolutionary algorithms have effectively trained multiagent teams to collectively solve complex problems. However, in many real-world applications, changes to the environment or agent functionality require agents to function well with multiple different teams. In this paper, we provide a counterfactual-state-based shaped fitness evaluation that provides an agent-specific signal that promotes effective cooperation across a variety of teams. The key insight leading to this result is that the shaped fitnesses across multiple teams can be aggregated because those performances are independent of each other. As a result, this approach leads to a single signal that captures an agent's performance across multiple teams. We show that this method provides significant improvement over standard multiagent fitness-shaped methods in learning robust cooperative behavior. Joshua Cook, Kagan Tumer |
GECCO | 2 |
| 2022 | Diversifying behaviors for learning in asymmetric multiagent systemsabstractTo achieve coordination in multiagent systems such as air traffic control or search and rescue, agents must not only evolve their policies, but also adapt to the behaviors of other agents. However, extending coevolutionary algorithms to complex domains is difficult because agents evolve in the dynamic environment created by the changing policies of other agents. This problem is exacerbated when the teams consist of diverse asymmetric agents (agents with different capabilities and objectives), making it difficult for agents to evolve complementary policies. Quality-Diversity methods solve part of the problem by allowing agents to discover not just optimal, but diverse behaviors, but are computationally intractable in multiagent settings. This paper introduces a multiagent learning framework to allow asymmetric agents to specialize and explore diverse behaviors needed for coordination in a shared environment. The key insight of this work is that a hierarchical decomposition of diversity search, fitness optimization, and team composition modeling allows the fitness on the team-wide objective to direct the diversity search in a dynamic environment. Experimental results in multiagent environments with temporal and spatial coupling requirements demonstrate the diversity of acquired agent synergies in response to a changing environment and team compositions. Everardo Gonzalez, Kagan Tumer |
GECCO | 3 |
| 2021 | MAEDyS: multiagent evolution via dynamic skill selectionabstractEvolving effective coordination strategies in tightly coupled multi-agent settings with sparse team fitness evaluations is challenging. It relies on multiple agents simultaneously stumbling upon the goal state to generate a learnable feedback signal. In such settings, estimating an agent's contribution to the overall team performance is extremely difficult, leading to a well-known structural credit assignment problem. This problem is further exacerbated when agents must complete sub-tasks with added spatial and temporal constraints, and different sub-tasks may require different local skills. We introduce MAEDyS, Multiagent Evolution via Dynamic Skill Selection, a hybrid bi-level optimization framework that augments evolutionary methods with policy gradient methods to generate effective coordination policies. MAEDyS learns to dynamically switch between multiple local skills towards optimizing the team fitness. It adopts fast policy gradients to learn several local skills using dense local rewards. It utilizes an evolutionary process to optimize the delayed team fitness by recruiting the most optimal skill at any given time. The ability to switch between various local skills during an episode eliminates the need for designing heuristic mixing functions. We evaluate MAEDyS in complex multiagent coordination environments with spatial and temporal constraints and show that it outperforms prior methods. Enna Sachdeva, Shauharda Khadka, Somdeb Majumdar, Kagan Tumer |
GECCO | 4 |
| 2020 | Multi-fitness learning for behavior-driven cooperationabstractEvolutionary learning algorithms have been successfully applied to multiagent problems where the desired system behavior can be captured by a single fitness signal. However, the complexity of many real world applications cannot be reduced to a single number, particularly when the fitness (i) arrives after a lengthy sequence of actions, and (ii) depends on the joint-action of multiple teammates. In this paper, we introduce the multi-fitness learning paradigm to enable multiagent teams to identify which fitness matters when in domains that require long-term, complex coordination. We demonstrate that multi-fitness learning efficiently solves a cooperative exploration task where teams of rovers must coordinate to observe various points of interest in a specific but unknown order. Connor Yates, Reid Christopher, Kagan Tumer |
GECCO | 3 |
| 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationabstractMany cooperative multiagent reinforcement learning environments provide agents with a sparse team-based reward, as well as a dense agent-specific reward that incentivizes learning basic skills. Training policies solely on the team-based reward is often difficult due to its sparsity. Also, relying solely on the agent-specific reward is sub-optimal because it usually does not capture the team coordination objective. A common approach is to use reward shaping to construct a proxy reward by combining the individual rewards. However, this requires manual tuning for each environment. We introduce Multiagent Evolutionary Reinforcement Learning (MERL), a split-level training platform that handles the two objectives separately through two optimization processes. An evolutionary algorithm maximizes the sparse team-based objective through neuroevolution on a population of teams. Concurrently, a gradient-based optimizer trains policies to only maximize the dense agent-specific rewards. The gradient-based policies are periodically added to the evolutionary population as a way of information transfer between the two optimization processes. This enables the evolutionary algorithm to use skills learned via the agent-specific rewards toward optimizing the global objective. Results demonstrate that MERL significantly outperforms state-of-the-art methods, such as MADDPG, on a number of difficult coordination benchmarks. Somdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer, Kagan Tumer |
ICML | 5 |
| 2020 | The impact of agent definitions and interactions on multiagent learning for coordination in traffic management domains
Jen Jen Chung, Damjan Miklic, Lorenzo Sabattini, Kagan Tumer, Roland Siegwart |
Auton. Agents Multi Agent Syst. | 4 |
| 2019 | Collaborative Evolutionary Reinforcement LearningabstractDeep reinforcement learning algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically struggle with achieving effective exploration and are extremely sensitive to the choice of hyperparameters. One reason is that most approaches use a noisy version of their operating policy to explore - thereby limiting the range of exploration. In this paper, we introduce Collaborative Evolutionary Reinforcement Learning (CERL), a scalable framework that comprises a portfolio of policies that simultaneously explore and exploit diverse regions of the solution space. A collection of learners - typically proven algorithms like TD3 - optimize over varying time-horizons leading to this diverse portfolio. All learners contribute to and use a shared replay buffer to achieve greater sample efficiency. Computational resources are dynamically distributed to favor the best learners as a form of online algorithm selection. Neuroevolution binds this entire process to generate a single emergent learner that exceeds the capabilities of any individual learner. Experiments in a range of continuous control benchmarks demonstrate that the emergent learner significantly outperforms its composite learners while remaining overall more sample-efficient - notably solving the Mujoco Humanoid benchmark where all of its composite learners (TD3) fail entirely in isolation. Shauharda Khadka, Somdeb Majumdar, Tarek Nassar, Zach Dwiel, Evren Tumer, Santiago Miret, Yinyin Liu, Kagan Tumer |
ICML | 8 |
| 2019 | Neuroevolution of a Modular Memory-Augmented Neural Network for Deep Memory ProblemsabstractWe present Modular Memory Units (MMUs), a new class of memory-augmented neural network. MMU builds on the gated neural architecture of Gated Recurrent Units (GRUs) and Long Short Term Memory (LSTMs), to incorporate an external memory block, similar to a Neural Turing Machine (NTM). MMU interacts with the memory block using independent read and write gates that serve to decouple the memory from the central feedforward operation. This allows for regimented memory access and update, giving our network the ability to choose when to read from memory, update it, or simply ignore it. This capacity to act in detachment allows the network to shield the memory from noise and other distractions, while simultaneously using it to effectively retain and propagate information over an extended period of time. We train MMU using both neuroevolution and gradient descent, and perform experiments on two deep memory benchmarks. Results demonstrate that MMU performs significantly faster and more accurately than traditional LSTM-based methods, and is robust to dramatic increases in the sequence depth of these memory benchmarks. Shauharda Khadka, Jen Jen Chung, Kagan Tumer |
Evol. Comput. | 3 |
| 2018 | Evolution-Guided Policy Gradient in Reinforcement LearningabstractDeep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically suffer from three core difficulties: temporal credit assignment with sparse rewards, lack of effective exploration, and brittle convergence properties that are extremely sensitive to hyperparameters. Collectively, these challenges severely limit the applicability of these approaches to real world problems. Evolutionary Algorithms (EAs), a class of black box optimization techniques inspired by natural evolution, are well suited to address each of these three challenges. However, EAs typically suffer from high sample complexity and struggle to solve problems that require optimization of a large number of parameters. In this paper, we introduce Evolutionary Reinforcement Learning (ERL), a hybrid algorithm that leverages the population of an EA to provide diversified data to train an RL agent, and reinserts the RL agent into the EA population periodically to inject gradient information into the EA. ERL inherits EA's ability of temporal credit assignment with a fitness metric, effective exploration with a diverse set of policies, and stability of a population-based approach and complements it with off-policy DRL's ability to leverage gradients for higher sample efficiency and faster learning. Experiments in a range of challenging continuous control benchmarks demonstrate that ERL significantly outperforms prior DRL and EA methods. Shauharda Khadka, Kagan Tumer |
NeurIPS | 2 |
| 2017 | Evolving memory-augmented neural architecture for deep memory problemsabstractIn this paper, we present a new memory-augmented neural network called Gated Recurrent Unit with Memory Block (GRU-MB). Our architecture builds on the gated neural architecture of a Gated Recurrent Unit (GRU) and integrates an external memory block, similar to a Neural Turing Machine (NTM). GRU-MB interacts with the memory block using independent read and write gates that serve to decouple the memory from the central feedforward operation. This allows for regimented memory access and update, administering our network the ability to choose when to read from memory, update it, or simply ignore it. This capacity to act in detachment allows the network to shield the memory from noise and other distractions, while simultaneously using it to effectively retain and propagate information over an extended period of time. We evolve GRU-MB using neuroevolution and perform experiments on two different deep memory tasks. Results demonstrate that GRU-MB performs significantly faster and more accurately than traditional memory-based methods, and is robust to dramatic increases in the depth of these tasks. Shauharda Khadka, Jen Jen Chung, Kagan Tumer |
GECCO | 3 |
| 2017 | Fitness function shaping in multiagent cooperative coevolutionary algorithms
Mitchell K. Colby, Kagan Tumer |
Auton. Agents Multi Agent Syst. | 2 |
| 2016 | Multiobjective Neuroevolutionary Control for a Fuel Cell Turbine Hybrid Energy SystemabstractIncreased energy demands are driving the development of new power generation technologies with high efficient. Direct fired fuel cell turbine hybrid systems are one such development, which have the potential to dramatically increase power generation efficiency, quickly respond to transient loads (and are generally flexible), and offer fast start up times. However, traditional control techniques are often inadequate in these systems because of extremely high nonlinearities and coupling between system parameters. In this work, we develop multi-objective neural network controller via neuroevolution and the Pareto Concavity Elimination Transformation (PaCcET). In order for the training process to be computationally tractable, we develop a computationally efficient plant simulator based on physical plant data, allowing for rapid fitness assignment. Results demonstrate that the multi-objective algorithm is able to develop a Pareto front of control policies which represent tradeoffs between tracking desired turbine speed profiles and minimizing transient operation of the fuel cell. Mitchell K. Colby, Logan Michael Yliniemi, Paolo Pezzini, David Tucker, Kenneth Mark Bryden, Kagan Tumer |
GECCO | 6 |
| 2016 | Neuroevolution of a Hybrid Power Plant SimulatorabstractEver increasing energy demands are driving the development of high-efficiency power generation technologies such as direct-fired fuel cell turbine hybrid systems. Due to lack of an accurate system model, high nonlinearities and high coupling between system parameters, traditional control strategies are often inadequate. To resolve this problem, learning based controllers trained using neuroevolution are currently being developed. In order for the neuroevolution of these controllers to be computationally tractable, a computationally efficient simulator of the plant is required. Despite the availability of real-time sensor data from a physical plant, supervised learning techniques such as backpropagation are deficient as minute errors at each step tend to propagate over time. In this paper, we implement a neuroevolutionary method in conjunction with backpropagation to ameliorate this problem. Furthermore, a novelty search method is implemented which is shown to diversify our neural network based-simulator, making it more robust to local optima. Results show that our simulator is able to achieve an overall average error of 0.39% and a maximum error of 1.26% for any state variable averaged over the time-domain simulation of the hybrid power plant. Shauharda Khadka, Kagan Tumer, Mitchell K. Colby, Dave Tucker, Paolo Pezzini, Kenneth Mark Bryden |
GECCO | 2 |
| 2016 | D++: Structural credit assignment in tightly coupled multiagent domainsabstractAutonomous multi-robot teams can be used in complex coordinated exploration tasks to improve exploration performance in terms of both speed and effectiveness. However, use of multi-robot systems presents additional challenges. Specifically, in domains where the robots' actions are tightly coupled, coordinating multiple robots to achieve cooperative behavior at the group level is difficult. In this paper, we demonstrate that reward shaping can greatly benefit learning in multi-robot exploration tasks. We propose a novel reward framework based on the idea of counterfactuals to tackle the coordination problem in tightly coupled domains. We show that the proposed algorithm provides superior performance (166% performance improvement and a quadruple convergence speed up) compared to policies learned using either the global reward or the difference reward [1]. Aida Rahmattalabi, Jen Jen Chung, Mitchell K. Colby, Kagan Tumer |
IROS | 4 |
| 2016 | Multi-objective multiagent credit assignment in reinforcement learning and NSGA-II
Logan Michael Yliniemi, Kagan Tumer |
Soft Comput. | 2 |
| 2015 | An Evolutionary Game Theoretic Analysis of Difference Evaluation FunctionsabstractOne of the key difficulties in cooperative coevolutionary algorithms is solving the credit assignment problem. Given the performance of a team of agents, it is difficult to determine the effectiveness of each agent in the system. One solution to solving the credit assignment problem is the difference evaluation function, which has produced excellent results in many multiagent coordination domains, and exhibits the desirable theoretical properties of alignment and sensitivity. However, to date, there has been no prescriptive theoretical analysis deriving conditions under which difference evaluations improve the probability of selecting optimal actions. In this paper, we derive such conditions. Further, we prove that difference evaluations do not alter the Nash equilibria locations or the relative ordering of fitness values for each action, meaning that difference evaluations do not typically harm converged system performance in cases where the conditions are not met. We then demonstrate the theoretical findings using an empirical basins of attraction analysis. Mitchell K. Colby, Kagan Tumer |
GECCO | 2 |
| 2015 | Implicit adaptive multi-robot coordination in dynamic environmentsabstractMulti-robot teams offer key advantages over single robots in exploration missions by increasing efficiency (explore larger areas), reducing risk (partial mission failure with robot failures), and enabling new data collection modes (multi-modal observations). However, coordinating multiple robots to achieve a system-level task is difficult, particularly if the task may change during the mission. In this work, we demonstrate how multiagent cooperative coevolutionary algorithms can develop successful control policies for dynamic and stochastic multi-robot exploration missions. We find that agents using difference evaluation functions (a technique that quantifies each individual agent's contribution to the team) provides superior system performance (up to 15%) compared to global evaluation functions and a hand-coded algorithm. Mitchell K. Colby, Jen Jen Chung, Kagan Tumer |
IROS | 3 |
| 2015 | Learning to trick cost-based planners into cooperative behaviorabstractIn this paper we consider the problem of routing autonomously guided robots by manipulating the cost space to induce safe trajectories in the work space. Specifically, we examine the domain of UAV traffic management in urban airspaces. Each robot does not explicitly coordinate with other vehicles in the airspace. Instead, the robots execute their own individual internal cost-based planner to travel between locations. Given this structure, our goal is to develop a high-level UAV traffic management (UTM) system that can dynamically adapt the cost space to reduce the number of conflict incidents in the airspace without knowing the internal planners of each robot. We propose a decentralized and distributed system of high-level traffic controllers that each learn appropriate costing strategies via a neuro-evolutionary algorithm. The policies learned by our algorithm demonstrated a 16.4% reduction in the total number of conflict incidents experienced in the airspace while maintaining throughput performance. Carrie Rebhuhn, Ryan Skeele, Jen Jen Chung, Geoffrey A. Hollinger, Kagan Tumer |
IROS | 5 |
| 2015 | Learning Tensegrity Locomotion Using Open-Loop Control Signals and Coevolutionary AlgorithmsabstractSoft robots offer many advantages over traditional rigid robots. However, soft robots can be difficult to control with standard control methods. Fortunately, evolutionary algorithms can offer an elegant solution to this problem. Instead of creating controls to handle the intricate dynamics of these robots, we can simply evolve the controls using a simulation to provide an evaluation function. In this article, we show how such a control paradigm can be applied to an emerging field within soft robotics: robots based on tensegrity structures. We take the model of the Spherical Underactuated Planetary Exploration Robot ball (SUPERball), an icosahedron tensegrity robot under production at NASA Ames Research Center, develop a rolling locomotion algorithm, and study the learned behavior using an accurate model of the SUPERball simulated in the NASA Tensegrity Robotics Toolkit. We first present the historical-average fitness-shaping algorithm for coevolutionary algorithms to speed up learning while favoring robustness over optimality. Second, we use a distributed control approach by coevolving open-loop control signals for each controller. Being simple and distributed, open-loop controllers can be readily implemented on SUPERball hardware without the need for sensor information or precise coordination. We analyze signals of different complexities and frequencies. Among the learned policies, we take one of the best and use it to analyze different aspects of the rolling gait, such as lengths, tensions, and energy consumption. We also discuss the correlation between the signals controlling different parts of the tensegrity robot. Atil Iscen, Ken Caluwaerts, Jonathan Bruce, Adrian K. Agogino, Vytas SunSpiral, Kagan Tumer |
Artif. Life | 6 |
| 2015 | Simulation of the introduction of new technologies in air traffic managementabstractAccurate simulation of the effects of integrating new technologies into a complex system is critical to the modernisation of large infrastructure problems. This is especially true in the modernisation of our antiquated air traffic system, where there exist many layers of interacting procedures, controls, and automation all designed to cooperate with human operators. Additions of even simple new technologies may result in unexpected emergent behaviour due to complex human/machine interactions. One approach is to create high-fidelity human models coming from the field of human factors that can simulate a rich set of behaviours. However, such models are difficult to produce, especially to show unexpected emergent behaviour coming from many human operators interacting simultaneously within a complex system. Instead, we introduce an alternate approach. Instead of engineering complex human models, we directly model the emergent behaviour with relatively simple goal-directed agents. In this model, each autonomous agent in a system pursues individual goals, and the high-level behaviour of the system emerges from the interactions, foreseen or unforeseen, between the agents/actors. We show that this method is capable of reflecting the integration of new technologies in a historical case, and apply the same methodology for a possible future technology. Finally, we show how these high-level simulated behaviours compare to actual deployed air traffic control mechanisms in use today. Logan Michael Yliniemi, Adrian K. Agogino, Kagan Tumer |
Connect. Sci. | 3 |
| 2014 | Hierarchical simulation for complex domains: air traffic flow managementabstractA key element in the continuing growth of air traffic is the increased use of automation. The Next Generation (Next-Gen) Air Traffic System will include automated decision support systems and satellite navigation that will let pilots know the precise locations of other aircraft around them. This Next-Gen suggestion system can assist pilots in making good decisions when they have to direct the aircraft themselves. However, effective automation is critical in achieving the capacity and safety goals of the Next-Gen Air Traffic System. In this paper we show that evolutionary algorithms can be used to achieve this effective automation. William J. Curran, Adrian K. Agogino, Kagan Tumer |
GECCO | 3 |
| 2014 | Evolutionary agent-based simulation of the introduction of new technologies in air traffic managementabstractAccurate simulation of the effects of integrating new technologies into a complex system is critical to the modernization of our antiquated air traffic system, where there exist many layers of interacting procedures, controls, and automation all designed to cooperate with human operators. Additions of even simple new technologies may result in unexpected emergent behavior due to complex human/machine interactions. One approach is to create high-fidelity human models coming from the field of human factors that can simulate a rich set of behaviors. However, such models are difficult to produce, especially to show unexpected emergent behavior coming from many human operators interacting simultaneously within a complex system. Instead of engineering complex human models, we directly model the emergent behavior by evolving goal directed agents, representing human users. Using evolution we can predict how the agent representing the human user reacts given his/her goals. In this paradigm, each autonomous agent in a system pursues individual goals, and the behavior of the system emerges from the interactions, foreseen or unforeseen, between the agents/actors. We show that this method reflects the integration of new technologies in a historical case, and apply the same methodology for a possible future technology. Logan Michael Yliniemi, Adrian K. Agogino, Kagan Tumer |
GECCO | 3 |
| 2014 | Flop and roll: Learning robust goal-directed locomotion for a Tensegrity RobotabstractTensegrity robots are composed of compression elements (rods) that are connected via a network of tension elements (cables). Tensegrity robots provide many advantages over standard robots, such as compliance, robustness, and flexibility. Moreover, sphere-shaped tensegrity robots can provide non-traditional modes of locomotion, such as rolling. While they have advantageous physical properties, tensegrity robots are hard to control because of their nonlinear dynamics and oscillatory nature. In this paper, we present a robust, distributed, and directional rolling algorithm, “flop and roll”. The algorithm uses coevolution and exploits the distributed nature and symmetry of the tensegrity structure. We validate this algorithm using the NASA Tensegrity Robotics Toolkit (NTRT) simulator, as well as the highly accurate model of the physical SUPERBall being developped under the NASA Innovative and Advanced Concepts (NIAC) program. Flop and roll improves upon previous approaches in that it provides rolling to a desired location. It is also robust to both unexpected external forces and partial hardware failures. Additionally, it handles variable terrain (hills up to 33% grade). Finally, results are compatible with the hardware since the algorithm relies on realistic sensing and actuation capabilities of the SUPERBall. Atil Iscen, Adrian K. Agogino, Vytas SunSpiral, Kagan Tumer |
IROS | 4 |
| 2013 | Multiagent Learning with a Noisy Global Reward SignalabstractScaling multiagent reinforcement learning to domains with many agents is a complex problem. In particular, multiagent credit assignment becomes a key issue as the system size increases. Some multiagent systems suffer from a global reward signal that is very noisy or difficult to analyze. This makes deriving a learnable local reward signal very difficult. Difference rewards (a particular instance of reward shaping) have been used to alleviate this concern, but they remain difficult to compute in many domains. In this paper we present an approach to modeling the global reward using function approximation that allows the quick computation of local rewards. We demonstrate how this model can result in significant improvements in behavior for three congestion problems: a multiagent ``bar problem'', a complex simulation of the United States airspace, and a generic air traffic domain. We show how the model of the global reward may be either learned on- or off-line using either linear functions or neural networks. For the bar problem, we show an increase in reward of nearly 200% over learning using the global reward directly. For the air traffic problem, we show a decrease in costs of 25% over learning using the global reward directly. Scott Proper, Kagan Tumer |
AAAI | 2 |
| 2013 | Controlling tensegrity robots through evolutionabstractTensegrity structures (built from interconnected rods and cables) have the potential to offer a revolutionary new robotic design that is light-weight, energy-efficient, robust to failures, capable of unique modes of locomotion, impact tolerant, and compliant (reducing damage between the robot and its environment). Unfortunately robots built from tensegrity structures are difficult to control with traditional methods due to their oscillatory nature, nonlinear coupling between components and overall complexity. Fortunately this formidable control challenge can be overcome through the use of evolutionary algorithms. In this paper we show that evolutionary algorithms can be used to efficiently control a ball shaped tensegrity robot. Experimental results performed with a variety of evolutionary algorithms in a detailed soft-body physics simulator show that a centralized evolutionary algorithm performs 400% better than a hand-coded solution, while the multiagent evolution performs 800% better. In addition, evolution is able to discover diverse control solutions (both crawling and rolling) that are robust against structural failures and can be adapted to a wide range of energy and actuation constraints. These successful controls will form the basis for building high-performance tensegrity robots in the near future. Atil Iscen, Adrian K. Agogino, Vytas SunSpiral, Kagan Tumer |
GECCO | 4 |
| 2013 | Coordinating actions in congestion games: impact of top-down and bottom-up utilities
Kagan Tumer, Scott Proper |
Auton. Agents Multi Agent Syst. | 1 |
| 2013 | Efficient Objective Functions for Coordinated Learning in Large-Scale Distributed OSA SystemsabstractIn this paper, we derive and evaluate private objective functions for large-scale, distributed opportunistic spectrum access (OSA) systems. By means of any learning algorithms, these derived objective functions enable OSA users to assess, locate, and exploit unused spectrum opportunities effectively by maximizing the users' average received rewards. We consider the elastic traffic model, suitable for elastic applications such as file transfer and web browsing, and in which an SU's received reward increases proportionally to the amount of received service when the amount is higher than a certain threshold. But when this amount is below the threshold, the reward decreases exponentially with the amount of received service. In this model, SUs are assumed to be treated fairly in that the SUs using the same band will roughly receive an equal share of the total amount of service offered by the band. We show that the proposed objective functions are: near-optimal, as they achieve high performances in terms of average received rewards; highly scalable, as they perform well for small- as well as large-scale systems; highly learnable, as they reach up near-optimal values very quickly; and distributive, as they require information sharing only among OSA users belonging to the same band. MohammadJavad NoroozOliaee, Bechir Hamdaoui, Kagan Tumer |
IEEE Trans. Mob. Comput. | 3 |
| 2012 | Evolving distributed resource sharing for cubesat constellationsabstractAdvances in miniaturization will allow for the commoditization of large numbers of tiny satellites, known as "CubeSats." However, current algorithms made for small tightly-managed space missions are ill-designed to take advantage of the huge amount of resources available in a decentralized collection of these CubeSats. We believe that multiagent evolutionary algorithms are ideally suited to exploit the distributed nature of this new problem. This paper presents a solution where a customer in need of satellite observations can reliably obtain these observations at low cost, through the help of a multiagent system as an intermediary. Each agent in this system is assigned to a single CubeSat. Given a set of the customer's observational needs, and models of the CubeSats' salient properties, the agents evolve policies that attempt to purchase an appropriate set of observations at a low price. This system is especially flexible as it demands no centralized resource broker, contracts or commitments of resources. We perform a series of experiments on an Earth-observition domain. The results show that the evolutionary methods combined with multiagent techniques have three times the performance of a simple hand-coded allocation algorithm, and twice the performance of simple evolving agents. Adrian K. Agogino, Chris HolmesParker, Kagan Tumer |
GECCO | 3 |
| 2012 | Evolving large scale UAV communication systemabstractUnmanned Aerial Vehicles (UAVs) have traditionally been used for short duration missions involving surveillance or military operations. Advances in batteries, photovoltaics and electric motors though, will soon allow large numbers of small, cheap, solar powered unmanned aerial vehicles (UAVs) to fly long term missions at high altitudes. This will revolutionize the way UAVs are used, allowing them to form vast communication networks. However, to make effective use of thousands (and perhaps millions) of UAVs owned by numerous disparate institutions, intelligent and robust coordination algorithms are needed, as this domain introduces unique congestion and signal-to-noise issues. In this paper, we present a solution based on evolutionary algorithms to a specific ad-hoc communication problem, where UAVs communicate to ground-based customers over a single wide-spectrum communication channel. To maximize their bandwidth, UAVs need to optimally control their output power levels and orientation. Experimental results show that UAVs using evolutionary algorithms in combination with appropriately shaped evaluation functions can form a robust communication network and perform 180% better than a fixed baseline algorithm as well as 90% better than a basic evolutionary algorithm. Adrian K. Agogino, Chris HolmesParker, Kagan Tumer |
GECCO | 3 |
| 2012 | A multiagent approach to managing air traffic flow
Adrian K. Agogino, Kagan Tumer |
Auton. Agents Multi Agent Syst. | 2 |
| 2012 | Coordinating Secondary-User Behaviors for Inelastic Traffic Reward Maximization in Large-Scale \osa NetworksabstractWe develop efficient coordination techniques that support inelastic traffic in large-scale distributed dynamic spectrum access (DSA) networks. By means of any learning algorithm, the proposed techniques enable DSA users to locate and exploit spectrum opportunities effectively, thereby increasing their achieved throughput (or “rewards” to be more general). Basically, learning algorithms allow DSA users to learn by interacting with the environment, and use their acquired knowledge to select the proper actions that maximize their own objectives, thereby “hopefully” maximizing their long-term cumulative received reward. However, when DSA users' objectives are not carefully coordinated, learning algorithms can lead to poor overall system performance, resulting in lesser per-user average achieved rewards. In this paper, we derive efficient objective functions that DSA users can aim to maximize, and that by doing so, users' collective behavior also leads to good overall system performance, thus maximizing each user's long-term cumulative received rewards. We show that the proposed techniques are: (i) efficient by enabling users to achieve high rewards, (ii) scalable by performing well in systems with a small as well as a large number of users, (iii) learnable by allowing users to reach up high rewards very quickly, and (iv) distributive by being implementable in a decentralized manner. Bechir Hamdaoui, MohammadJavad NoroozOliaee, Kagan Tumer, Ammar Rayes |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2012 | Evolving a Multiagent Controller for Micro Aerial VehiclesabstractMicro aerial vehicles (MAVs) are notoriously difficult to control as they are light, susceptible to minor fluctuations in the environment, and obey highly nonlinear dynamics. Indeed, traditional control methods, particularly those relying on difficult to obtain models of the interaction between an MAV and its environment, have been unable to provide adequate control beyond simple maneuvers. In this paper, we address the problem of controlling an MAV (which has segmented control surfaces) by evolving a neurocontroller and fine tuning it using multiagent coordination techniques. This approach is based on a control strategy that learns to map MAV states (position and velocity) to MAV actions (e.g., actuator position) to achieve good performance (e.g., flight time) by maximizing an objective function. The main difficulty with this approach is defining the objective functions at the MAV level that allow good performance. In addition, to provide added robustness, we investigate a multiagent approach to control where each control surface aims to optimize a local objective. Our results show that this approach not only provides good MAV control, but provides robustness to: 1) wind gusts by a factor of 6; 2) turbulence by a factor of 4; and 3) hardware failures by a factor of 8 over a traditional control method. Max Salichon, Kagan Tumer |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2011 | Optimizing ballast design of wave energy converters using evolutionary algorithmsabstractWave energy converters promise to be a viable alternative to current electrical generation methods. However, these generators must become more efficient before wide-scale industrial use can become cost-effective. The efficiency of these devices is primarily dependent upon their geometry and ballast configuration which are both difficult to evaluate, due to slow computation time and high computation cost of current models. In this paper, we use evolutionary algorithms to optimize the ballast geometry of a wave energy generator using a two step process. First, we generate a function approximator (neural network) to predict wave energy converter power output with respect to key geometric design variables. This is a critical step as the computation time of using a full model (e.g., AQWA) to predict energy output prohibits the use of an evolutionary algorithm for design optimization. The function approximator reduced the computation time by over 99% while having an average error of only 1.5%. The evolutionary algorithm then optimized the weight distribution of a wave energy generator, resulting in an 84% improvement in power output over a ballast-free wave energy converter. Mitchell K. Colby, Ehsan M. Nasroullahi, Kagan Tumer |
GECCO | 3 |
| 2011 | Agent fitness functions for evolving coordinated sensor networksabstractDistributed sensor networks are an attractive area for research in agent systems. This is due primarily to the level of information available in applications where sensing technology has improved dramatically. These include energy systems and area coverage where it is desirable for sensor networks to have the ability to self-organize and be robust to changes in network structure. The challenges presented when investigating distributed sensor networks for such applications include the need for small sensor packages that are still capable of making good decisions to cover areas where multiple types of information may be present. For example in energy systems, singular areas in power plants may produce several types of valuable information, such as temperature, pressure, or chemical indicators. Matt Knudson, Kagan Tumer |
GECCO | 3 |
| 2011 | Aligning Spectrum-User Objectives for Maximum Inelastic-Traffic RewardabstractWe develop objective functions for large-scale distributed dynamic spectrum access (DSA) networks that, by means of any learning algorithm, enable DSA users to locate and exploit spectrum opportunities effectively, thereby increasing their achieved throughput (or "rewards" to be more general). We show that the proposed functions are: (i) optimal by enabling users to achieve high rewards, (ii) scalable by performing well in systems with a small as well as a large number of users, (iii) learnable by allowing users to reach up high rewards very quickly, and (iv) distributed by being implementable in a decentralized manner. Bechir Hamdaoui, MohammadJavad NoroozOliaee, Kagan Tumer, Ammar Rayes |
ICCCN | 3 |
| 2010 | Coevolution of heterogeneous multi-robot teamsabstractEvolving multiple robots so that each robot acting independently can contribute to the maximization of a system level objective presents significant scientific challenges. For example, evolving multiple robots to maximize aggregate information in exploration domains (e.g., planetary exploration, search and rescue) requires coordination, which in turn requires the careful design of the evaluation functions. Additionally, where communication among robots is expensive (e.g., limited power or computation), the coordination must be achieved passively, without robots explicitly informing others of their states/intended actions. Coevolving robots in these situations is a potential solution to producing coordinated behavior, where the robots are coupled through their evaluation functions. In this work, we investigate coevolution in three types of domains: (i) where precisely n homogeneous robots need to perform a task; (ii) where n is the optimal number of homogeneous robots for the task; and (iii) where n is the optimal number of heterogeneous robots for the task. Our results show that coevolving robots with evaluation functions that are locally aligned with the system evaluation significantly improve performance over robots evolving using the system evaluation function directly, particularly in dynamic environments. Matt Knudson, Kagan Tumer |
GECCO | 2 |
| 2010 | A neuro-evolutionary approach to micro aerial vehicle controlabstractApplying classical control methods to Micro Aerial Vehicles (MAVs) is a difficult process due to the complexity of the control laws with fast and highly non-linear dynamics. Such methods rely heavily on difficult to obtain models and are particularly ill-suited to the stochastic and dynamic environments in which MAVs operate. Instead, in this paper, we focus on a neuro-evolutionary method that learns to map MAV states (position, velocity) to MAV actions (e.g., actuator position). Our results show significant improvements in response times to minor altitude and heading corrections over a traditional PID controller. In addition, we show that the MAV response to maintaining altitude in the presence of wind gusts improves by a factor of five. Similarly, we show that the MAV response to maintaining heading in the presence of turbulence improves by factors of three. Max Salichon, Kagan Tumer |
GECCO | 2 |
| 2010 | Robust neuro-control for a micro quadrotorabstractQuadrotors are unique among Micro Aerial Vehicles in providing excellent maneuverability (as opposed to winged flight), while maintaing a simple mechanical construction (as opposed to helicopters). This mechanical simplicity comes at a cost of increased controller complexity. Quadrotors are inherently unstable, and micro quadrotors are particularly difficult to control. In this paper, we evolve a hierarchical neuro-controller for micro (0.5 kg) quadrotor control. The first stage of control aims to stabilize the craft and outputs rotor speeds based on a requested attitude (pitch, roll, yaw, and vertical velocity). This controller is developed in four parts around each of the variables, and then combined and trained further to increase robustness. The second stage of control aims to achieve a requested (x, y, z) position by providing the first stage with the appropriate attitude. The results show that stable quadrotor control is achieved through this architecture. In addition, the results show that neuro-evolutionary control recovers from disturbances over an order of magnitude faster than a basic PID controller. Finally, the neuro-evolutionary controller provides stable flight in the presence of 5 times more sensor noise and 8 times more actuator noise as compared to the PID controller. Jack F. Shepherd III, Kagan Tumer |
GECCO | 2 |
| 2008 | Adaptive Management of Air Traffic Flow: A Multiagent Coordination Approach
Kagan Tumer, Adrian K. Agogino |
AAAI | 1 |
| 2008 | Analyzing and visualizing multiagent rewards in dynamic and stochastic domains
Adrian K. Agogino, Kagan Tumer |
Auton. Agents Multi Agent Syst. | 2 |
| 2008 | Efficient Evaluation Functions for Evolving CoordinationabstractAbstract This paper presents fitness evaluation functions that efficiently evolve coordination in large multi-component systems. In particular, we focus on evolving distributed control policies that are applicable to dynamic and stochastic environments. While it is appealing to evolve such policies directly for an entire system, the search space is prohibitively large in most cases to allow such an approach to provide satisfactory results. Instead, we present an approach based on evolving system components individually where each component aims to maximize its own fitness function. Though this approach sidesteps the exploding state space concern, it introduces two new issues: (1) how to create component evaluation functions that are aligned with the global evaluation function; and (2) how to create component evaluation functions that are sensitive to the fitness changes of that component, while remaining relatively insensitive to the fitness changes of other components in the system. If the first issue is not addressed, the resulting system becomes uncoordinated; if the second issue is not addressed, the evolutionary process becomes either slow to converge or worse, incapable of converging to good solutions. This paper shows how to construct evaluation functions that promote coordination by satisfying these two properties. We apply these evaluation functions to the distributed control problem of coordinating multiple rovers to maximize aggregate information collected. We focus on environments that are highly dynamic (changing points of interest), noisy (sensor and actuator faults), and communication limited (both for observation of other rovers and points of interest) forcing the rovers to evolve generalized solutions. On this difficult coordination problem, the control policy evolved using aligned and component-sensitive evaluation functions outperforms global evaluation functions by up to 400%. More notably, the performance improvements increase when the problems become more difficult (larger, noisier, less communication). In addition we provide an analysis of the results by quantifying the two characteristics (alignment and sensitivity discussed above) leading to a systematic study of the presented fitness functions. Adrian K. Agogino, Kagan Tumer |
Evol. Comput. | 2 |
| 2008 | Ensemble clustering with voting active clusters
Kagan Tumer, Adrian K. Agogino |
Pattern Recognit. Lett. | 1 |
| 2007 | Evolving distributed agents for managing air trafficabstractAir traffic management offers an intriguing real world challenge to designing large scale distributed systems using evolutionary computation. The ability to evolve effective air traffic flow strategies depends not only on evolving good local strategies, but also on ensuring that those local strategies result in good global solutions. While traditional, direct evolutionary strategies can be highly effective in certain combinatorial domains, they are not well-suited to complex air traffic flow problems because of the large interdependencies among the local subsystems. In this paper, we propose an evolutionary agent-based solution to the air traffic flow problem. In this approach, we evolve agents both to learn the right local flow strategies to alleviate congestion in their immediate surroundings, and to prevent the creation of congestion "downstream" from their local areas. The agent-based approach leads to better and more fault-tolerant solutions. To validate this approach, we use FACET, an air traffic simulator developed at NASA and used extensively by the FAA and industry. On a scenario composed of three hundred aircraft and two points of congestion, our results show that an agent based evolutionary computation method, where each agent uses the system evaluation function, achieves 40% improvement over a direct evolutionary algorithm. In addition by creating agent-specific "difference evaluation functions" we achieve an additional 30% improvement over agents using the system evaluation. Adrian K. Agogino, Kagan Tumer |
GECCO | 2 |
| 2006 | QUICR-Learning for Multi-Agent Coordination
Adrian K. Agogino, Kagan Tumer |
AAAI | 2 |
| 2006 | Distributed evaluation functions for fault tolerant multi-rover systemsabstractThe ability to evolve fault tolerant control strategies for large collections of agents is critical to the successful application of evolutionary strategies to domains where failures are common. Furthermore, while evolutionary algorithms have been highly successful in discovering single-agent control strategies, extending such algorithms to multi-agent domains has proven to be difficult. In this paper we present a method for shaping evaluation functions for agents that provide control strategies that are both tolerant to different types of failures and lead to coordinated behavior in a multi-agent setting. This method neither relies on a centralized strategy (susceptible to single points of failures) nor a distributed strategy where each agent uses a system wide evaluation function (severe credit assignment problem). In a multi-rover problem, we show that agents using our agent-specific evaluation perform up to 500% better than agents using the system evaluation. In addition we show that agents are still able to maintain a high level of performance when up to 60% of the agents fail due to actuator, communication or controller faults. Adrian K. Agogino, Kagan Tumer |
GECCO | 2 |
| 2006 | Handling Communication Restrictions and Team Formation in Congestion Games
Adrian K. Agogino, Kagan Tumer |
Auton. Agents Multi Agent Syst. | 2 |
| 2005 | Efficient credit assignment through evaluation function decompositionabstractEvolutionary methods are powerful tools in discovering solutions for difficult continuous tasks.When such a solution is encoded over multiple genes, a genetic algorithm faces the difficult credit assignment problem of evaluating how a single gene in a chromosome contributes to the full solution.Typically a single evaluation function is used for the entire chromosome, implicitly giving each gene in the chromosome the same evaluation.This method is inefficient because a gene will get credit for the contribution of all the other genes as well.Accurately measuring the fitness of individual genes in such a large search space requires many trials.This paper instead proposes turning this single complex search problem into a multi-agent search problem, where each agent has the simpler task of discovering a suitable gene.Gene-specific evaluation functions can then be created that have better theoreticaal properties than a single evaluation function over all genes.This method is tested in the difficult double-pole balancing problem, showing that agents using gene-specific evaluation functions can create a successful control policy in 20% fewer trials than the best existing genetic.algorithms.The method is extended to more distributed problems, achieving 95% performance gains over tradition methods in the multi-rover domain. Adrian K. Agogino, Kagan Tumer, Risto Miikkulainen |
GECCO | 2 |
| 2005 | Coordinating multi-rover systems: evaluation functions for dynamic and noisy environmentsabstractThis paper addresses the evolution of control strategies for a collective: a set of entities that collectively strives to maximize a global evaluation function that rates the performance of the full system. Directly addressing such problems by having a population of collectives and applying the evolutionary algorithm to that population is appealing, but the search space is prohibitively large in most cases. Instead, we focus on evolving control policies for each member of the collective. The main difficulty with this approach is creating an evaluation function for each member of the collective that is both aligned with the global evaluation function and sensitive to the fitness changes of the member. We show how to construct evaluation functions in dynamic, noisy and communication-limited collective environments. On a rover coordination problem, a control policy evolved using aligned and member-sensitive evaluations outperforms global evaluation methods by up to 400%. More notably, in the presence of a larger number of rovers or rovers with noisy and communication limited sensors, the improvements due to the proposed method become significantly more pronounced. Kagan Tumer, Adrian K. Agogino |
GECCO | 1 |
| 2004 | Efficient Evaluation Functions for Multi-rover Systems
Adrian K. Agogino, Kagan Tumer |
GECCO (1) | 2 |
| 2004 | Overcoming communication restrictions in collectivesabstractThe performance of distributed systems generally depend on the actions and interactions of a large number of independent components (e.g., agents, neurons). Such "collectives" are often subject to communication restrictions, making it difficult for the components to coordinate their actions to provide good system level performance. In this article, we address that coordination problem and derive four agent utility functions that make different tradeoffs between alignedness between agent and system utilities and the signal-to-noise each agent encounters. The results show that these utility functions outperform both traditional methods and previous collective-based methods by up to 75% in systems with communication restrictions. Kagan Tumer, Adrian K. Agogino |
IJCNN | 1 |
| 2003 | Input decimated ensembles
Kagan Tumer, Nikunj C. Oza |
Pattern Anal. Appl. | 1 |
| 2002 | Collective Intelligence, Data Routing and Braess' ParadoxabstractWe consider the problem of designing the the utility functions of the utility-maximizing agents in a multi-agent system so that they work synergistically to maximize a global utility. The particular problem domain we explore is the control of network routing by placing agents on all the routers in the network. Conventional approaches to this task have the agents all use the Ideal Shortest Path routing Algorithm (ISPA). We demonstrate that in many cases, due to the side-effects of one agent's actions on another agent's performance, having agents use ISPA's is suboptimal as far as global aggregate cost is concerned, even when they are only used to route infinitesimally small amounts of traffic. The utility functions of the individual agents are not ``aligned'' with the global utility, intuitively speaking. As a particular example of this we present an instance of Braess' paradox in which adding new links to a network whose agents all use the ISPA results in a decrease in overall throughput. We also demonstrate that load-balancing, in which the agents' decisions are collectively made to optimize the global cost incurred by all traffic currently being routed, is suboptimal as far as global cost averaged across time is concerned. This is also due to `side-effects', in this case of current routing decision on future traffic. The mathematics of Collective Intelligence (COIN) is concerned precisely with the issue of avoiding such deleterious side-effects in multi-agent systems, both over time and space. We present key concepts from that mathematics and use them to derive an algorithm whose ideal version should have better performance than that of having all agents use the ISPA, even in the infinitesimal limit. We present experiments verifying this, and also showing that a machine-learning-based version of this COIN algorithm in which costs are only imprecisely estimated via empirical means (a version potentially applicable in the real world) also outperforms the ISPA, despite having access to less information than does the ISPA. In particular, this COIN algorithm almost always avoids Braess' paradox. David H. Wolpert, Kagan Tumer |
J. Artif. Intell. Res. | 2 |
| 2002 | Robust Combining of Disparate Classifiers through Order Statistics
Kagan Tumer, Joydeep Ghosh |
Pattern Anal. Appl. | 1 |
| 2001 | Reinforcement Learning in Distributed Domains: Beyond Team Games
David H. Wolpert, Joseph Sill, Kagan Tumer |
IJCAI | 3 |
| 1998 | Using Collective Intelligence to Route Internet Traffic
David H. Wolpert, Kagan Tumer, Jeremy Frank |
NIPS | 2 |
| 1996 | Estimating the Bayes error rate through classifier combiningabstractThe Bayes error provides the lowest achievable error rate for a given pattern classification problem. There are several classical approaches for estimating or finding bounds for the Bayes error. One type of approach focuses on obtaining analytical bounds, which are both difficult to calculate and dependent on distribution parameters that may not be known. Another strategy is to estimate the class densities through non-parametric methods, and use these estimates to obtain bounds on the Bayes error. This article presents a novel approach to estimating the Bayes error based on classifier combining techniques. For an artificial data set where the Bayes error is known, the combiner-based estimate outperforms the classical methods. Kagan Tumer, Joydeep Ghosh |
ICPR | 1 |
| 1996 | Spectroscopic Detection of Cervical Pre-Cancer through Radial Basis Function Networks
Kagan Tumer, Nirmala Ramanujam, Rebecca R. Richards-Kortum, Joydeep Ghosh |
NIPS | 1 |
| 1996 | Error Correlation and Error Reduction in Ensemble ClassifiersabstractUsing an ensemble of classifiers, instead of a single classifier, can lead to improved generalization. The gains obtained by combining, however, are often affected more by the selection of what is presented to the combiner than by the actual combining method that is chosen. In this paper, we focus on data selection and classifier training methods, in order to 'prepare' classifiers for combining. We review a combining framework for classification problems that quantifies the need for reducing the correlation among individual classifiers. Then, we discuss several methods that make the classifiers in an ensemble more complementary. Experimental results are provided to illustrate the benefits and pitfalls of reducing the correlation among classifiers, especially when the training data are in limited supply. Kagan Tumer, Joydeep Ghosh |
Connect. Sci. | 1 |
| 1996 | Analysis of decision boundaries in linearly combined neural classifiers
Kagan Tumer, Joydeep Ghosh |
Pattern Recognit. | 1 |
| 1995 | Designing genetic algorithms for the state assignment problemabstractFinding the best state assignment for implementing a synchronous sequential circuit is important for reducing silicon area or chip count in many digital designs. This state assignment problem (SAP) belongs to a broader class of combinatorial optimization problems than the well studied traveling salesman problem, which can be formulated as a special case of SAP. The search for a good solution is considerably involved for the SAP due to a large number of equivalent solutions, and no effective heuristic has been found so far to cater to all types of circuits. In this paper, a matrix representation is used as the genotype for a genetic algorithm (GA) approach to this problem. A novel selection mechanism is introduced, and suitable genetic operators for crossover and mutation, are constructed. The properties of each of these elements of the GA are discussed and an analysis of parameters that influence the algorithm is given. A canonical form for a solution is defined to significantly reduce the search space and number of local minima. Experiments with several examples show that the GA approach yields results that are often comparable to, or better than those obtained using established heuristics that embody extensive domain knowledge.> José Nelson Amaral, Kagan Tumer, Joydeep Ghosh |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1994 | Sequence Recognition by Input Anticipation
Kagan Tumer, Joydeep Ghosh |
IEA/AIE | 1 |