VLDB 2026 Research / reviewers in the wild / expert
Srinath Mahankali
dblp:321/0657
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 81% Legged, aerial and field robots · 19% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.8 | 1 | 2024 | Maximizing Quadruped Velocity by Minimizing Energy · ICRA 2024 |
Machine learning › Reinforcement learning
exploration |
0.8 | 1 | 2024 | Random Latent Exploration for Deep Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning › exploration
exploration bonus |
0.8 | 1 | 2024 | Random Latent Exploration for Deep Reinforcement Learning · ICML 2024 |
Robotics › Legged, aerial and field robots › legged robots › legged robot locomotion
quadruped locomotion |
0.8 | 1 | 2024 | Maximizing Quadruped Velocity by Minimizing Energy · ICRA 2024 |
Machine learning › Reinforcement learning › exploration
randomized value functions |
0.8 | 1 | 2024 | Random Latent Exploration for Deep Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.2 | 1 | 2024 | Random Latent Exploration for Deep Reinforcement Learning · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
reward perturbation · 0.8random latent exploration · 0.8proximal policy optimization · 0.8extrinsic-intrinsic policy optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Random Latent Exploration for Deep Reinforcement LearningabstractThe ability to efficiently explore high-dimensional state spaces is essential for the practical success of deep Reinforcement Learning (RL). This paper introduces a new exploration technique called Random Latent Exploration (RLE), that combines the strengths of exploration bonuses and randomized value functions (two popular approaches for effective exploration in deep RL). RLE leverages the idea of perturbing rewards by adding structured random rewards to the original task rewards in certain (random) states of the environment, to encourage the agent to explore the environment during training. RLE is straightforward to implement and performs well in practice. To demonstrate the practical effectiveness of RLE, we evaluate it on the challenging Atari and IsaacGym benchmarks and show that RLE exhibits higher overall scores across all the tasks than other approaches, including action-noise and randomized value function exploration. Srinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin, Pulkit Agrawal 0001 |
ICML | 1 |
| 2024 | Maximizing Quadruped Velocity by Minimizing EnergyabstractReinforcement Learning (RL) has been a powerful tool for training robots to acquire agile locomotion skills. To learn locomotion, it is commonly necessary to introduce additional reward-shaping terms, such as an energy minimization term, to guide an algorithm like Proximal Policy Optimization (PPO) to good performance. Prior works rely on hyper-parameter tuning on the weight of the reward shaping terms to obtain satisfactory task performance. To save the efforts of tuning these weights, we adopt the Extrinsic-Intrinsic Policy Optimization (EIPO) framework. The key idea of EIPO is to establish a constrained optimization framework for the primary objective of enhancing task performance and the secondary objective of minimizing energy consumption. It seeks a policy that minimizes the energy consumption objective within the optimal policy space for task performance. This guarantees that the learned policy excels in task performance while conserving energy, all without requiring manual weight adjustments for both objectives. Our experiments evaluate EIPO on various quadruped locomotion tasks, revealing that policies trained with EIPO consistently achieve higher task performance than PPO comparisons while maintaining comparable energy consumption levels. Furthermore, EIPO exhibits superior task performance in real-world evaluations compared to PPO. Srinath Mahankali, Chi-Chang Lee, Gabriel B. Margolis, Zhang-Wei Hong, Pulkit Agrawal 0001 |
ICRA | 1 |