Kaushik Subramanian

dblp:19/8399 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Reinforcement learning · 46% Motion planning and robot control · 22% Trustworthy machine learning · 13%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 20 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
out-of-distribution generalization
1.222026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024
Machine learning › Reinforcement learning
deep reinforcement learning
1.222025
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025
Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018
Machine learning › Reinforcement learning › multi-task reinforcement learning
contextual reinforcement learning
1.012026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Robotics › Motion planning and robot control
robot control
1.012026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Robotics › Motion planning and robot control › robot control
robust control
1.012026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Machine learning › Learning theory › inductive bias
simplicity bias
0.912025
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning › policy optimization
diverse policy learning
0.812024
Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024
Robotics › Autonomous driving › autonomous vehicle navigation
intersection navigation
0.312018
Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.222011
Apprenticeship Learning About Multiple Intentions · ICML 2011
Generalizing Apprenticeship Learning across Hypothesis Classes · ICML 2010
Machine learning › Reinforcement learning
efficient reinforcement learning
0.212014
Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains · Artif. Intell. 2014
Machine learning › Reinforcement learning › policy search
bayesian policy search
0.212013
Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013
Machine learning › Reinforcement learning
human feedback
0.212013
Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013
Machine learning › Reinforcement learning › human-in-the-loop reinforcement learning
interactive reinforcement learning
0.212013
Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013
Robotics › Robot manipulation
learning from demonstration
0.112011
Learning Tasks and Skills Together From a Human Teacher · AAAI 2011
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning
0.112011
Learning Tasks and Skills Together From a Human Teacher · AAAI 2011
Robotics › Motion planning and robot control › robot learning
task learning
0.112011
Learning Tasks and Skills Together From a Human Teacher · AAAI 2011
Machine learning › Reinforcement learning
behavior learning
0.112010
Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models · AAAI 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model
0.112010
Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models · AAAI 2010
Machine learning › Reinforcement learning
imitation learning
0.112010
Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models · AAAI 2010
Robotics › Robot navigation and mapping › active perception
active sensing
0.112018
Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018

Methods — techniques the papers use, named apart from their topics

single-phase adaptation · 1.0context encoder · 1.0residual connections · 0.9observation normalization · 0.9layer normalization · 0.9successor features · 0.8diversity objective with constraints · 0.8deep reinforcement learning · 0.3demonstration · 0.2abstraction · 0.2social dialog · 0.1learning from demonstration · 0.1
YearPublicationVenuePosition
2026 Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy
abstract
Generalization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act within environments with varying contexts, such as self-driving cars or quadrupedal robots that need to operate in different terrains or weather conditions than they were trained for. We tackle the critical task of generalizing to out-of-distribution (OOD) settings, without access to explicit context information at test time. Recent work has addressed this problem by training a context encoder and a history adaptation module in separate stages. While promising, this two-phase approach is cumbersome to implement and train. We simplify the methodology and introduce SPARC: single-phase adaptation for robust control. We test SPARC on varying contexts within the high-fidelity racing simulator Gran Turismo 7 and wind-perturbed MuJoCo environments, and find that it achieves reliable and robust OOD generalization.
Bram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone 0001, Peter R. Wurman
AAAI3
2025 SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
abstract
Recent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to overfitting. These large networks avoid overfitting by integrating components that induce a simplicity bias, guiding models toward simple and generalizable solutions. However, in deep RL, designing and scaling up networks have been less explored. Motivated by this opportunity, we present SimBa, an architecture designed to scale up parameters in deep RL by injecting a simplicity bias. SimBa consists of three components: (i) an observation normalization layer that standardizes inputs with running statistics, (ii) a residual feedforward block to provide a linear pathway from the input to output, and (iii) a layer normalization to control feature magnitudes. By scaling up parameters with SimBa, the sample efficiency of various deep RL algorithms—including off-policy, on-policy, and unsupervised methods—is consistently improved. Moreover, solely by integrating SimBa architecture into SAC, it matches or surpasses state-of-the-art deep RL methods with high computational efficiency across DMC, MyoSuite, and HumanoidBench. These results demonstrate SimBa's broad applicability and effectiveness across diverse RL algorithms and environments.
Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R. Wurman, Jaegul Choo, Peter Stone 0001, Takuma Seno
ICLR6
2024 Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration
abstract
The ability to approach the same problem from different angles is a cornerstone of human intelligence that leads to robust solutions and effective adaptation to problem variations. In contrast, current RL methodologies tend to lead to policies that settle on a single solution to a given problem, making them brittle to problem variations. Replicating human flexibility in reinforcement learning agents is the challenge that we explore in this work. We tackle this challenge by extending state-of-the-art approaches to introduce DUPLEX, a method that explicitly defines a diversity objective with constraints and makes robust estimates of policies’ expected behavior through successor features. The trained agents can (i) learn a diverse set of near-optimal policies in complex highly-dynamic environments and (ii) exhibit competitive and diverse skills in out-of-distribution (OOD) contexts. Empirical results indicate that DUPLEX improves over previous methods and successfully learns competitive driving styles in a hyper-realistic simulator (i.e., GranTurismo ™ 7) as well as diverse and effective policies in several multi-context robotics MuJoCo simulations with OOD gravity forces and height limits. To the best of our knowledge, our method is the first to achieve diverse solutions in complex driving simulators and OOD robotic contexts. DUPLEX agents demonstrating diverse behaviors can be found at https://ai.sony/publications/Discovering-Creative-Behaviors-through-DUPLEX-Diverse-Universal-Features-for-Policy-Exploration/.
Borja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman, Peter Stone 0001
NeurIPS3
2018 Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning
abstract
Providing an efficient strategy to navigate safely through unsignaled intersections is a difficult task that requires determining the intent of other drivers. We explore the effectiveness of Deep Reinforcement Learning to handle intersection problems. Using recent advances in Deep RL, we are able to learn policies that surpass the performance of a commonly-used heuristic approach in several metrics including task completion time and goal success rate and have limited ability to generalize. We then explore a system's ability to learn active sensing behaviors to enable navigating safely in the case of occlusions. Our analysis, provides insight into the intersection handling problem, the solutions learned by the network point out several shortcomings of current rule-based methods, and the failures of our current deep reinforcement learning system point to future research directions.
David Isele, Reza Rahimi, Akansel Cosgun, Kaushik Subramanian, Kikuo Fujimura
ICRA4
2014 Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains
Luis C. Cobo, Kaushik Subramanian, Charles L. Isbell Jr., Aaron D. Lanterman, Andrea Thomaz
Artif. Intell.2
2013 Policy Shaping: Integrating Human Feedback with Reinforcement Learning
abstract
A long term goal of Interactive Reinforcement Learning is to incorporate non-expert human feedback to solve complex tasks. State-of-the-art methods have approached this problem by mapping human information to reward and value signals to indicate preferences and then iterating over them to compute the necessary control policy. In this paper we argue for an alternate, more effective characterization of human feedback: Policy Shaping. We introduce Advise, a Bayesian approach that attempts to maximize the information gained from human feedback by utilizing it as direct labels on the policy. We compare Advise to state-of-the-art approaches and highlight scenarios where it outperforms them and importantly is robust to infrequent and inconsistent human feedback.
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L. Isbell Jr., Andrea Thomaz
NIPS2
2011 Learning Tasks and Skills Together From a Human Teacher
abstract
We are interested in developing Learning from Demonstration (LfD) systems that are tailored to be used by everyday people. We highlight and tackle the issues of skill learning, task learning and interaction in the context of LfD As part of the AAAI 2011 LfD Challenge, we will demonstrate some of our most recent Socially Guided-Machine Learning work, in which the PR2 robot learns both low-level skills and high-level tasks through an ongoing social dialog with a human partner
Baris Akgün, Kaushik Subramanian, Jaeeun Shim, Andrea Thomaz
AAAI2
2011 Apprenticeship Learning About Multiple Intentions
Monica Babes-Vroman, Vukosi Marivate, Kaushik Subramanian, Michael L. Littman
ICML3
2010 Task Space Behavior Learning for Humanoid Robots using Gaussian Mixture Models
abstract
In this paper a system was developed for robot behavior acquisition using kinesthetic demonstrations. It enables a humanoid robot to imitate constrained reaching gestures directed towards a target using a learning algorithm based on Gaussian Mixture Models. The imitation trajectory can be reshaped in order to satisfy the constraints of the task and it can adapt to changes in the initial conditions and to target displacements occurring during movement execution. The potential of this method was evaluated using experiments with the Nao, Aldebaran’s humanoid robot.
Kaushik Subramanian
AAAI1
2010 Generalizing Apprenticeship Learning across Hypothesis Classes
Thomas J. Walsh 0001, Kaushik Subramanian, Michael L. Littman, Carlos Diuk
ICML2