EDBT 2026 Demo / reviewers in the wild / expert
Jeongmo Kim
dblp:396/8344
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 79% Trustworthy machine learning · 21% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
meta-reinforcement learning |
0.9 | 1 | 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.9 | 1 | 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
task representation learning |
0.9 | 1 | 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Exclusively Penalized Q-learning for Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
underestimation bias |
0.8 | 1 | 2024 | Exclusively Penalized Q-learning for Offline Reinforcement Learning · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
state regularization · 0.9metric-based representation learning · 0.9value function penalization · 0.8exclusively penalized q-learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution TasksabstractMeta reinforcement learning aims to develop policies that generalize to unseen tasks sampled from a task distribution. While context-based meta-RL methods improve task representation using task latents, they often struggle with out-of-distribution (OOD) tasks. To address this, we propose Task-Aware Virtual Training (TAVT), a novel algorithm that accurately captures task characteristics for both training and OOD scenarios using metric-based representation learning. Our method successfully preserves task characteristics in virtual tasks and employs a state regularization technique to mitigate overestimation errors in state-varying environments. Numerical results demonstrate that TAVT significantly enhances generalization to OOD tasks across various MuJoCo and MetaWorld environments. Our code is available at https://github.com/JM-Kim-94/tavt.git. Jeongmo Kim, Yisak Park, Minung Kim, Seungyul Han |
ICML | 1 |
| 2024 | Exclusively Penalized Q-learning for Offline Reinforcement LearningabstractConstraint-based offline reinforcement learning (RL) involves policy constraints or imposing penalties on the value function to mitigate overestimation errors caused by distributional shift. This paper focuses on a limitation in existing offline RL methods with penalized value function, indicating the potential for underestimation bias due to unnecessary bias introduced in the value function. To address this concern, we propose Exclusively Penalized Q-learning (EPQ), which reduces estimation bias in the value function by selectively penalizing states that are prone to inducing estimation errors. Numerical results show that our method significantly reduces underestimation bias and improves performance in various offline control tasks compared to other offline RL methods. Junghyuk Yeom, Yonghyeon Jo, Jeongmo Kim, Seungyul Han |
NeurIPS | 3 |