Srinath Mahankali

dblp:321/0657 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 81% Legged, aerial and field robots · 19%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
constrained reinforcement learning
0.812024
Maximizing Quadruped Velocity by Minimizing Energy · ICRA 2024
Machine learning › Reinforcement learning
exploration
0.812024
Random Latent Exploration for Deep Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › exploration
exploration bonus
0.812024
Random Latent Exploration for Deep Reinforcement Learning · ICML 2024
Robotics › Legged, aerial and field robots › legged robots › legged robot locomotion
quadruped locomotion
0.812024
Maximizing Quadruped Velocity by Minimizing Energy · ICRA 2024
Machine learning › Reinforcement learning › exploration
randomized value functions
0.812024
Random Latent Exploration for Deep Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning
deep reinforcement learning
0.212024
Random Latent Exploration for Deep Reinforcement Learning · ICML 2024

Methods — techniques the papers use, named apart from their topics

reward perturbation · 0.8random latent exploration · 0.8proximal policy optimization · 0.8extrinsic-intrinsic policy optimization · 0.8
YearPublicationVenuePosition
2024 Random Latent Exploration for Deep Reinforcement Learning
abstract
The ability to efficiently explore high-dimensional state spaces is essential for the practical success of deep Reinforcement Learning (RL). This paper introduces a new exploration technique called Random Latent Exploration (RLE), that combines the strengths of exploration bonuses and randomized value functions (two popular approaches for effective exploration in deep RL). RLE leverages the idea of perturbing rewards by adding structured random rewards to the original task rewards in certain (random) states of the environment, to encourage the agent to explore the environment during training. RLE is straightforward to implement and performs well in practice. To demonstrate the practical effectiveness of RLE, we evaluate it on the challenging Atari and IsaacGym benchmarks and show that RLE exhibits higher overall scores across all the tasks than other approaches, including action-noise and randomized value function exploration.
Srinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin, Pulkit Agrawal 0001
ICML1
2024 Maximizing Quadruped Velocity by Minimizing Energy
abstract
Reinforcement Learning (RL) has been a powerful tool for training robots to acquire agile locomotion skills. To learn locomotion, it is commonly necessary to introduce additional reward-shaping terms, such as an energy minimization term, to guide an algorithm like Proximal Policy Optimization (PPO) to good performance. Prior works rely on hyper-parameter tuning on the weight of the reward shaping terms to obtain satisfactory task performance. To save the efforts of tuning these weights, we adopt the Extrinsic-Intrinsic Policy Optimization (EIPO) framework. The key idea of EIPO is to establish a constrained optimization framework for the primary objective of enhancing task performance and the secondary objective of minimizing energy consumption. It seeks a policy that minimizes the energy consumption objective within the optimal policy space for task performance. This guarantees that the learned policy excels in task performance while conserving energy, all without requiring manual weight adjustments for both objectives. Our experiments evaluate EIPO on various quadruped locomotion tasks, revealing that policies trained with EIPO consistently achieve higher task performance than PPO comparisons while maintaining comparable energy consumption levels. Furthermore, EIPO exhibits superior task performance in real-world evaluations compared to PPO.
Srinath Mahankali, Chi-Chang Lee, Gabriel B. Margolis, Zhang-Wei Hong, Pulkit Agrawal 0001
ICRA1