EDBT 2026 Demo / reviewers in the wild / expert
Shauharda Khadka
dblp:183/9233
· DBLP profile ↗
9ranked-venue papers
6as first author
3since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 97% Deep learning architectures and training · 3% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › population-based learning › evolutionary learning › population-based reinforcement learning
evolutionary reinforcement learning |
1.1 | 3 | 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020 Collaborative Evolutionary Reinforcement Learning · ICML 2019 Evolution-Guided Policy Gradient in Reinforcement Learning · NeurIPS 2018 |
Machine learning › Reinforcement learning › relational reinforcement learning
graph reinforcement learning |
0.5 | 1 | 2021 | Optimizing Memory Placement using Evolutionary Graph Reinforcement Learning · ICLR 2021 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.4 | 1 | 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.4 | 1 | 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020 |
Machine learning › Reinforcement learning
sparse reward |
0.4 | 1 | 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2019 | Collaborative Evolutionary Reinforcement Learning · ICML 2019 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2019 | Collaborative Evolutionary Reinforcement Learning · ICML 2019 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.3 | 1 | 2018 | Evolution-Guided Policy Gradient in Reinforcement Learning · NeurIPS 2018 |
Mathematical optimization › multi-objective optimization
evolutionary algorithm |
0.3 | 1 | 2018 | Evolution-Guided Policy Gradient in Reinforcement Learning · NeurIPS 2018 |
Mathematical optimization › metaheuristic optimization
population-based optimization |
0.3 | 1 | 2018 | Evolution-Guided Policy Gradient in Reinforcement Learning · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › neural network training
neuroevolution |
0.1 | 1 | 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020 |
Machine learning › Reinforcement learning
continuous control |
0.1 | 1 | 2019 | Collaborative Evolutionary Reinforcement Learning · ICML 2019 |
Methods — techniques the papers use, named apart from their topics
evolutionary algorithm · 1.1evolutionary graph reinforcement learning · 1.0neuroevolution · 0.8policy gradient · 0.7off-policy learning · 0.7gradient-based optimization · 0.4replay buffer · 0.4TD3 · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Learning Intrinsic Symbolic Rewards in Reinforcement LearningabstractLearning effective policies for sparse objectives is a key challenge in Deep Reinforcement Learning (RL). A common approach is to design task-related dense rewards to improve task learnability. While such rewards are easily interpreted, they rely on heuristics and domain expertise. Alternate approaches that train neural networks to discover dense surrogate rewards avoid heuristics, but are high-dimensional, black-box solutions offering little interpretability. In this paper, we present a method that discovers dense rewards in the form of low-dimensional symbolic trees - thus making them more tractable for analysis. The trees use simple functional operators to map an agent's observations to a scalar reward, which then supervises the policy gradient learning of a neural network policy. We test our method on continuous action spaces in Mujoco and discrete action spaces in Atari and Pygame environments. We show that the discovered dense rewards are an effective signal for an RL policy to solve the benchmark tasks. Notably, we significantly outperform a widely used, contemporary neural-network based reward-discovery algorithm in all environments considered. Hassam Ullah Sheikh, Shauharda Khadka, Santiago Miret, Somdeb Majumdar, Mariano Phielipp |
IJCNN | 2 |
| 2021 | MAEDyS: multiagent evolution via dynamic skill selectionabstractEvolving effective coordination strategies in tightly coupled multi-agent settings with sparse team fitness evaluations is challenging. It relies on multiple agents simultaneously stumbling upon the goal state to generate a learnable feedback signal. In such settings, estimating an agent's contribution to the overall team performance is extremely difficult, leading to a well-known structural credit assignment problem. This problem is further exacerbated when agents must complete sub-tasks with added spatial and temporal constraints, and different sub-tasks may require different local skills. We introduce MAEDyS, Multiagent Evolution via Dynamic Skill Selection, a hybrid bi-level optimization framework that augments evolutionary methods with policy gradient methods to generate effective coordination policies. MAEDyS learns to dynamically switch between multiple local skills towards optimizing the team fitness. It adopts fast policy gradients to learn several local skills using dense local rewards. It utilizes an evolutionary process to optimize the delayed team fitness by recruiting the most optimal skill at any given time. The ability to switch between various local skills during an episode eliminates the need for designing heuristic mixing functions. We evaluate MAEDyS in complex multiagent coordination environments with spatial and temporal constraints and show that it outperforms prior methods. Enna Sachdeva, Shauharda Khadka, Somdeb Majumdar, Kagan Tumer |
GECCO | 2 |
| 2021 | Optimizing Memory Placement using Evolutionary Graph Reinforcement Learning
Shauharda Khadka, Estelle Aflalo, Mattias Marder, Avrech Ben-David, Santiago Miret, Shie Mannor, Tamir Hazan, Somdeb Majumdar |
ICLR | 1 |
| 2020 | Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationabstractMany cooperative multiagent reinforcement learning environments provide agents with a sparse team-based reward, as well as a dense agent-specific reward that incentivizes learning basic skills. Training policies solely on the team-based reward is often difficult due to its sparsity. Also, relying solely on the agent-specific reward is sub-optimal because it usually does not capture the team coordination objective. A common approach is to use reward shaping to construct a proxy reward by combining the individual rewards. However, this requires manual tuning for each environment. We introduce Multiagent Evolutionary Reinforcement Learning (MERL), a split-level training platform that handles the two objectives separately through two optimization processes. An evolutionary algorithm maximizes the sparse team-based objective through neuroevolution on a population of teams. Concurrently, a gradient-based optimizer trains policies to only maximize the dense agent-specific rewards. The gradient-based policies are periodically added to the evolutionary population as a way of information transfer between the two optimization processes. This enables the evolutionary algorithm to use skills learned via the agent-specific rewards toward optimizing the global objective. Results demonstrate that MERL significantly outperforms state-of-the-art methods, such as MADDPG, on a number of difficult coordination benchmarks. Somdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer, Kagan Tumer |
ICML | 2 |
| 2019 | Collaborative Evolutionary Reinforcement LearningabstractDeep reinforcement learning algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically struggle with achieving effective exploration and are extremely sensitive to the choice of hyperparameters. One reason is that most approaches use a noisy version of their operating policy to explore - thereby limiting the range of exploration. In this paper, we introduce Collaborative Evolutionary Reinforcement Learning (CERL), a scalable framework that comprises a portfolio of policies that simultaneously explore and exploit diverse regions of the solution space. A collection of learners - typically proven algorithms like TD3 - optimize over varying time-horizons leading to this diverse portfolio. All learners contribute to and use a shared replay buffer to achieve greater sample efficiency. Computational resources are dynamically distributed to favor the best learners as a form of online algorithm selection. Neuroevolution binds this entire process to generate a single emergent learner that exceeds the capabilities of any individual learner. Experiments in a range of continuous control benchmarks demonstrate that the emergent learner significantly outperforms its composite learners while remaining overall more sample-efficient - notably solving the Mujoco Humanoid benchmark where all of its composite learners (TD3) fail entirely in isolation. Shauharda Khadka, Somdeb Majumdar, Tarek Nassar, Zach Dwiel, Evren Tumer, Santiago Miret, Yinyin Liu, Kagan Tumer |
ICML | 1 |
| 2019 | Neuroevolution of a Modular Memory-Augmented Neural Network for Deep Memory ProblemsabstractWe present Modular Memory Units (MMUs), a new class of memory-augmented neural network. MMU builds on the gated neural architecture of Gated Recurrent Units (GRUs) and Long Short Term Memory (LSTMs), to incorporate an external memory block, similar to a Neural Turing Machine (NTM). MMU interacts with the memory block using independent read and write gates that serve to decouple the memory from the central feedforward operation. This allows for regimented memory access and update, giving our network the ability to choose when to read from memory, update it, or simply ignore it. This capacity to act in detachment allows the network to shield the memory from noise and other distractions, while simultaneously using it to effectively retain and propagate information over an extended period of time. We train MMU using both neuroevolution and gradient descent, and perform experiments on two deep memory benchmarks. Results demonstrate that MMU performs significantly faster and more accurately than traditional LSTM-based methods, and is robust to dramatic increases in the sequence depth of these memory benchmarks. Shauharda Khadka, Jen Jen Chung, Kagan Tumer |
Evol. Comput. | 1 |
| 2018 | Evolution-Guided Policy Gradient in Reinforcement LearningabstractDeep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically suffer from three core difficulties: temporal credit assignment with sparse rewards, lack of effective exploration, and brittle convergence properties that are extremely sensitive to hyperparameters. Collectively, these challenges severely limit the applicability of these approaches to real world problems. Evolutionary Algorithms (EAs), a class of black box optimization techniques inspired by natural evolution, are well suited to address each of these three challenges. However, EAs typically suffer from high sample complexity and struggle to solve problems that require optimization of a large number of parameters. In this paper, we introduce Evolutionary Reinforcement Learning (ERL), a hybrid algorithm that leverages the population of an EA to provide diversified data to train an RL agent, and reinserts the RL agent into the EA population periodically to inject gradient information into the EA. ERL inherits EA's ability of temporal credit assignment with a fitness metric, effective exploration with a diverse set of policies, and stability of a population-based approach and complements it with off-policy DRL's ability to leverage gradients for higher sample efficiency and faster learning. Experiments in a range of challenging continuous control benchmarks demonstrate that ERL significantly outperforms prior DRL and EA methods. Shauharda Khadka, Kagan Tumer |
NeurIPS | 1 |
| 2017 | Evolving memory-augmented neural architecture for deep memory problemsabstractIn this paper, we present a new memory-augmented neural network called Gated Recurrent Unit with Memory Block (GRU-MB). Our architecture builds on the gated neural architecture of a Gated Recurrent Unit (GRU) and integrates an external memory block, similar to a Neural Turing Machine (NTM). GRU-MB interacts with the memory block using independent read and write gates that serve to decouple the memory from the central feedforward operation. This allows for regimented memory access and update, administering our network the ability to choose when to read from memory, update it, or simply ignore it. This capacity to act in detachment allows the network to shield the memory from noise and other distractions, while simultaneously using it to effectively retain and propagate information over an extended period of time. We evolve GRU-MB using neuroevolution and perform experiments on two different deep memory tasks. Results demonstrate that GRU-MB performs significantly faster and more accurately than traditional memory-based methods, and is robust to dramatic increases in the depth of these tasks. Shauharda Khadka, Jen Jen Chung, Kagan Tumer |
GECCO | 1 |
| 2016 | Neuroevolution of a Hybrid Power Plant SimulatorabstractEver increasing energy demands are driving the development of high-efficiency power generation technologies such as direct-fired fuel cell turbine hybrid systems. Due to lack of an accurate system model, high nonlinearities and high coupling between system parameters, traditional control strategies are often inadequate. To resolve this problem, learning based controllers trained using neuroevolution are currently being developed. In order for the neuroevolution of these controllers to be computationally tractable, a computationally efficient simulator of the plant is required. Despite the availability of real-time sensor data from a physical plant, supervised learning techniques such as backpropagation are deficient as minute errors at each step tend to propagate over time. In this paper, we implement a neuroevolutionary method in conjunction with backpropagation to ameliorate this problem. Furthermore, a novelty search method is implemented which is shown to diversify our neural network based-simulator, making it more robust to local optima. Results show that our simulator is able to achieve an overall average error of 0.39% and a maximum error of 1.26% for any state variable averaged over the time-domain simulation of the hybrid power plant. Shauharda Khadka, Kagan Tumer, Mitchell K. Colby, Dave Tucker, Paolo Pezzini, Kenneth Mark Bryden |
GECCO | 1 |