EDBT 2026 Demo / reviewers in the wild / expert
Kaushik Subramanian
dblp:19/8399
· DBLP profile ↗
10ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Reinforcement learning · 46% Motion planning and robot control · 22% Trustworthy machine learning · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% |
Topics — the 20 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
1.2 | 2 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024 |
Machine learning › Reinforcement learning
deep reinforcement learning |
1.2 | 2 | 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025 Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018 |
Machine learning › Reinforcement learning › multi-task reinforcement learning
contextual reinforcement learning |
1.0 | 1 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 |
Robotics › Motion planning and robot control
robot control |
1.0 | 1 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 |
Robotics › Motion planning and robot control › robot control
robust control |
1.0 | 1 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 |
Machine learning › Learning theory › inductive bias
simplicity bias |
0.9 | 1 | 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning › policy optimization
diverse policy learning |
0.8 | 1 | 2024 | Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024 |
Robotics › Autonomous driving › autonomous vehicle navigation
intersection navigation |
0.3 | 1 | 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.2 | 2 | 2011 | Apprenticeship Learning About Multiple Intentions · ICML 2011 Generalizing Apprenticeship Learning across Hypothesis Classes · ICML 2010 |
Machine learning › Reinforcement learning
efficient reinforcement learning |
0.2 | 1 | 2014 | Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains · Artif. Intell. 2014 |
Machine learning › Reinforcement learning › policy search
bayesian policy search |
0.2 | 1 | 2013 | Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013 |
Machine learning › Reinforcement learning
human feedback |
0.2 | 1 | 2013 | Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013 |
Machine learning › Reinforcement learning › human-in-the-loop reinforcement learning
interactive reinforcement learning |
0.2 | 1 | 2013 | Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013 |
Robotics › Robot manipulation
learning from demonstration |
0.1 | 1 | 2011 | Learning Tasks and Skills Together From a Human Teacher · AAAI 2011 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning |
0.1 | 1 | 2011 | Learning Tasks and Skills Together From a Human Teacher · AAAI 2011 |
Robotics › Motion planning and robot control › robot learning
task learning |
0.1 | 1 | 2011 | Learning Tasks and Skills Together From a Human Teacher · AAAI 2011 |
Machine learning › Reinforcement learning
behavior learning |
0.1 | 1 | 2010 | Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models · AAAI 2010 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.1 | 1 | 2010 | Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models · AAAI 2010 |
Machine learning › Reinforcement learning
imitation learning |
0.1 | 1 | 2010 | Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models · AAAI 2010 |
Robotics › Robot navigation and mapping › active perception
active sensing |
0.1 | 1 | 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018 |
Methods — techniques the papers use, named apart from their topics
single-phase adaptation · 1.0context encoder · 1.0residual connections · 0.9observation normalization · 0.9layer normalization · 0.9successor features · 0.8diversity objective with constraints · 0.8deep reinforcement learning · 0.3demonstration · 0.2abstraction · 0.2social dialog · 0.1learning from demonstration · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single PolicyabstractGeneralization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act within environments with varying contexts, such as self-driving cars or quadrupedal robots that need to operate in different terrains or weather conditions than they were trained for. We tackle the critical task of generalizing to out-of-distribution (OOD) settings, without access to explicit context information at test time. Recent work has addressed this problem by training a context encoder and a history adaptation module in separate stages. While promising, this two-phase approach is cumbersome to implement and train. We simplify the methodology and introduce SPARC: single-phase adaptation for robust control. We test SPARC on varying contexts within the high-fidelity racing simulator Gran Turismo 7 and wind-perturbed MuJoCo environments, and find that it achieves reliable and robust OOD generalization. Bram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone 0001, Peter R. Wurman |
AAAI | 3 |
| 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningabstractRecent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to overfitting.
These large networks avoid overfitting by integrating components that induce a simplicity bias, guiding models toward simple and generalizable solutions.
However, in deep RL, designing and scaling up networks have been less explored.
Motivated by this opportunity, we present SimBa, an architecture designed to scale up parameters in deep RL by injecting a simplicity bias. SimBa consists of three components: (i) an observation normalization layer that standardizes inputs with running statistics, (ii) a residual feedforward block to provide a linear pathway from the input to output, and (iii) a layer normalization to control feature magnitudes.
By scaling up parameters with SimBa, the sample efficiency of various deep RL algorithms—including off-policy, on-policy, and unsupervised methods—is consistently improved.
Moreover, solely by integrating SimBa architecture into SAC, it matches or surpasses state-of-the-art deep RL methods with high computational efficiency across DMC, MyoSuite, and HumanoidBench.
These results demonstrate SimBa's broad applicability and effectiveness across diverse RL algorithms and environments. Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R. Wurman, Jaegul Choo, Peter Stone 0001, Takuma Seno |
ICLR | 6 |
| 2024 | Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy ExplorationabstractThe ability to approach the same problem from different angles is a cornerstone of human intelligence that leads to robust solutions and effective adaptation to problem variations. In contrast, current RL methodologies tend to lead to policies that settle on a single solution to a given problem, making them brittle to problem variations. Replicating human flexibility in reinforcement learning agents is the challenge that we explore in this work. We tackle this challenge by extending state-of-the-art approaches to introduce DUPLEX, a method that explicitly defines a diversity objective with constraints and makes robust estimates of policies’ expected behavior through successor features. The trained agents can (i) learn a diverse set of near-optimal policies in complex highly-dynamic environments and (ii) exhibit competitive and diverse skills in out-of-distribution (OOD) contexts. Empirical results indicate that DUPLEX improves over previous methods and successfully learns competitive driving styles in a hyper-realistic simulator (i.e., GranTurismo ™ 7) as well as diverse and effective policies in several multi-context robotics MuJoCo simulations with OOD gravity forces and height limits. To the best of our knowledge, our method is the first to achieve diverse solutions in complex driving simulators and OOD robotic contexts. DUPLEX agents demonstrating diverse behaviors can be found at https://ai.sony/publications/Discovering-Creative-Behaviors-through-DUPLEX-Diverse-Universal-Features-for-Policy-Exploration/. Borja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman, Peter Stone 0001 |
NeurIPS | 3 |
| 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement LearningabstractProviding an efficient strategy to navigate safely through unsignaled intersections is a difficult task that requires determining the intent of other drivers. We explore the effectiveness of Deep Reinforcement Learning to handle intersection problems. Using recent advances in Deep RL, we are able to learn policies that surpass the performance of a commonly-used heuristic approach in several metrics including task completion time and goal success rate and have limited ability to generalize. We then explore a system's ability to learn active sensing behaviors to enable navigating safely in the case of occlusions. Our analysis, provides insight into the intersection handling problem, the solutions learned by the network point out several shortcomings of current rule-based methods, and the failures of our current deep reinforcement learning system point to future research directions. David Isele, Reza Rahimi, Akansel Cosgun, Kaushik Subramanian, Kikuo Fujimura |
ICRA | 4 |
| 2014 | Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains
Luis C. Cobo, Kaushik Subramanian, Charles L. Isbell Jr., Aaron D. Lanterman, Andrea Thomaz |
Artif. Intell. | 2 |
| 2013 | Policy Shaping: Integrating Human Feedback with Reinforcement LearningabstractA long term goal of Interactive Reinforcement Learning is to incorporate non-expert human feedback to solve complex tasks. State-of-the-art methods have approached this problem by mapping human information to reward and value signals to indicate preferences and then iterating over them to compute the necessary control policy. In this paper we argue for an alternate, more effective characterization of human feedback: Policy Shaping. We introduce Advise, a Bayesian approach that attempts to maximize the information gained from human feedback by utilizing it as direct labels on the policy. We compare Advise to state-of-the-art approaches and highlight scenarios where it outperforms them and importantly is robust to infrequent and inconsistent human feedback. Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L. Isbell Jr., Andrea Thomaz |
NIPS | 2 |
| 2011 | Learning Tasks and Skills Together From a Human TeacherabstractWe are interested in developing Learning from Demonstration (LfD) systems that are tailored to be used by everyday people. We highlight and tackle the issues of skill learning, task learning and interaction in the context of LfD As part of the AAAI 2011 LfD Challenge, we will demonstrate some of our most recent Socially Guided-Machine Learning work, in which the PR2 robot learns both low-level skills and high-level tasks through an ongoing social dialog with a human partner Baris Akgün, Kaushik Subramanian, Jaeeun Shim, Andrea Thomaz |
AAAI | 2 |
| 2011 | Apprenticeship Learning About Multiple Intentions
Monica Babes-Vroman, Vukosi Marivate, Kaushik Subramanian, Michael L. Littman |
ICML | 3 |
| 2010 | Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture ModelsabstractIn this paper a system was developed for robot behavior acquisition using kinesthetic demonstrations. It enables a humanoid robot to imitate constrained reaching gestures directed towards a target using a learning algorithm based on Gaussian Mixture Models. The imitation trajectory can be reshaped in order to satisfy the constraints of the task and it can adapt to changes in the initial conditions and to target displacements occurring during movement execution. The potential of this method was evaluated using experiments with the Nao, Aldebaran’s humanoid robot. Kaushik Subramanian |
AAAI | 1 |
| 2010 | Generalizing Apprenticeship Learning across Hypothesis Classes
Thomas J. Walsh 0001, Kaushik Subramanian, Michael L. Littman, Carlos Diuk |
ICML | 2 |