Joe Eappen

dblp:267/5377 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0001-9386-5545ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 60% Robot navigation and mapping · 20% Motion planning and robot control · 20%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
mobile robot navigation
0.812024
Co-learning Planning and Control Policies Constrained by Differentiable Logic Specifications · ICRA 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Information-Directed Pessimism for Offline Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › offline reinforcement learning
pessimism
0.812024
Information-Directed Pessimism for Offline Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning
policy optimization
0.812024
Information-Directed Pessimism for Offline Reinforcement Learning · ICML 2024
Robotics › Motion planning and robot control
robot control
0.812024
Co-learning Planning and Control Policies Constrained by Differentiable Logic Specifications · ICRA 2024

Methods — techniques the papers use, named apart from their topics

stein discrepancy · 0.8reinforcement learning · 0.8differentiable logic specifications · 0.8concentration bounds · 0.8
YearPublicationVenuePosition
2024 Information-Directed Pessimism for Offline Reinforcement Learning
abstract
Policy optimization from batch data, i.e., offline reinforcement learning (RL) is important when collecting data from a current policy is not possible. This setting incurs distribution mismatch between batch training data and trajectories from the current policy. Pessimistic offsets estimate mismatch using concentration bounds, which possess strong theoretical guarantees and simplicity of implementation. Mismatch may be conservative in sparse data regions and less so otherwise, which can result in under-performing their no-penalty variants in practice. We derive a new pessimistic penalty as the distance between the data and the true distribution using an evaluable one-sample test known as Stein Discrepancy that requires minimal smoothness conditions, and noticeably, allows a mixture family representation of distribution over next states. This entity forms a quantifier of information in offline data, which justifies calling this approach *information-directed pessimism* (IDP) for offline RL. We further establish that this new penalty based on discrete Stein discrepancy yields practical gains in performance while generalizing the regret of prior art to multimodal distributions.
Alec Koppel, Sujay Bhatt, Jiacheng Guo, Joe Eappen, Mengdi Wang 0001, Sumitra Ganesh
ICML4
2024 Co-learning Planning and Control Policies Constrained by Differentiable Logic Specifications
abstract
Synthesizing planning and control policies in robotics is a fundamental task, further complicated by factors such as complex logic specifications and high-dimensional robot dynamics. This paper presents a novel reinforcement learning approach to solving high-dimensional robot navigation tasks with complex logic specifications by co-learning planning and control policies. Notably, this approach significantly reduces the sample complexity in training, allowing us to train high-quality policies with much fewer samples compared to existing reinforcement learning algorithms. In addition, our methodology streamlines complex specification extraction from map images and enables the efficient generation of long-horizon robot motion paths across different map layouts. Moreover, our approach also demonstrates capabilities for high-dimensional control and avoiding suboptimal policies via policy alignment. The efficacy of our approach is demonstrated through experiments involving simulated high-dimensional quadruped robot dynamics and a real-world differential drive robot (TurtleBot3) under different types of task specifications.
Zikang Xiong, Daniel Lawson, Joe Eappen, Ahmed H. Qureshi, Suresh Jagannathan
ICRA3
2022 Model-free Neural Lyapunov Control for Safe Robot Navigation
abstract
Model-free Deep Reinforcement Learning (DRL) controllers have demonstrated promising results on various challenging non-linear control tasks. While a model-free DRL algorithm can solve unknown dynamics and high-dimensional problems, it lacks safety assurance. Although safety constraints can be encoded as part of a reward function, there still exists a large gap between an RL controller trained with this modified reward and a safe controller. In contrast, instead of implicitly encoding safety constraints with rewards, we explicitly colearn a Twin Neural Lyapunov Function (TNLF) with the control policy in the DRL training loop and use the learned TNLF to build a runtime monitor. Combined with the path generated from a planner, the monitor chooses appropriate waypoints that guide the learned controller to provide collision-free control trajectories. Our approach inherits the scalability advantages from DRL while enhancing safety guarantees. Our experimental evaluation demonstrates the effectiveness of our approach compared to DRL with augmented rewards and constrained DRL methods over a range of high-dimensional safety-sensitive navigation tasks.
Zikang Xiong, Joe Eappen, Ahmed H. Qureshi, Suresh Jagannathan
IROS2
2022 DistSPECTRL: Distributing Specifications in Multi-Agent Reinforcement Learning Systems
Joe Eappen, Suresh Jagannathan
ECML/PKDD (4)1
2022 Defending Observation Attacks in Deep Reinforcement Learning via Detection and Denoising
Zikang Xiong, Joe Eappen, He Zhu 0001, Suresh Jagannathan
ECML/PKDD (3)2