Jeongmo Kim

dblp:396/8344 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 79% Trustworthy machine learning · 21%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
meta-reinforcement learning
0.912025
Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.912025
Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
task representation learning
0.912025
Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Exclusively Penalized Q-learning for Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › value-based reinforcement learning
underestimation bias
0.812024
Exclusively Penalized Q-learning for Offline Reinforcement Learning · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

state regularization · 0.9metric-based representation learning · 0.9value function penalization · 0.8exclusively penalized q-learning · 0.8
YearPublicationVenuePosition
2025 Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks
abstract
Meta reinforcement learning aims to develop policies that generalize to unseen tasks sampled from a task distribution. While context-based meta-RL methods improve task representation using task latents, they often struggle with out-of-distribution (OOD) tasks. To address this, we propose Task-Aware Virtual Training (TAVT), a novel algorithm that accurately captures task characteristics for both training and OOD scenarios using metric-based representation learning. Our method successfully preserves task characteristics in virtual tasks and employs a state regularization technique to mitigate overestimation errors in state-varying environments. Numerical results demonstrate that TAVT significantly enhances generalization to OOD tasks across various MuJoCo and MetaWorld environments. Our code is available at https://github.com/JM-Kim-94/tavt.git.
Jeongmo Kim, Yisak Park, Minung Kim, Seungyul Han
ICML1
2024 Exclusively Penalized Q-learning for Offline Reinforcement Learning
abstract
Constraint-based offline reinforcement learning (RL) involves policy constraints or imposing penalties on the value function to mitigate overestimation errors caused by distributional shift. This paper focuses on a limitation in existing offline RL methods with penalized value function, indicating the potential for underestimation bias due to unnecessary bias introduced in the value function. To address this concern, we propose Exclusively Penalized Q-learning (EPQ), which reduces estimation bias in the value function by selectively penalizing states that are prone to inducing estimation errors. Numerical results show that our method significantly reduces underestimation bias and improves performance in various offline control tasks compared to other offline RL methods.
Junghyuk Yeom, Yonghyeon Jo, Jeongmo Kim, Seungyul Han
NeurIPS3