EDBT 2026 Demo / reviewers in the wild / expert
Jeonghye Kim
dblp:172/6718
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 76% Transfer learning and domain adaptation · 7% Language models and text generation · 6% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
2.4 | 3 | 2025 | Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025 Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning · NeurIPS 2024 Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making · ICLR 2024 |
Machine learning › Reinforcement learning › reward design
reward scaling |
1.7 | 2 | 2025 | Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025 ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
1.0 | 1 | 2026 | RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation · AAAI 2026 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection · EMNLP 2025 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.9 | 1 | 2025 | ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning |
0.9 | 1 | 2025 | Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › value function estimation
value estimation bias |
0.9 | 1 | 2025 | Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.8 | 1 | 2024 | Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › offline reinforcement learning
return-conditioned supervised learning |
0.8 | 1 | 2024 | Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework · ICML 2023 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
0.7 | 1 | 2023 | LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework · ICML 2023 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework |
0.7 | 1 | 2023 | LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
phase orchestration · 1.0parameter transfer · 1.0reflection · 0.9react · 0.9q-learning · 0.9periodic network reset · 0.9layer normalization · 0.9large language model · 0.9adaptive reward scaling · 0.9SPOT · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RL-Studio: A System for Multi-Phase Reinforcement Learning ExperimentationabstractReinforcement learning (RL) has evolved beyond monolithic training, yet existing frameworks remain limited to single algorithms or simple offline-to-online transitions. We present multi-phase RL, a framework that orchestrates multiple learning phases for continual policy improvement. It enables efficient fine-tuning of pretrained policies with new data and smooth adaptation from simulation to real-world environments. To support this paradigm, we introduce RL-Studio, a platform that addresses key implementation barriers, including neural architecture mismatches, parameter transfer complexities, and experiment management overhead. It provides phase orchestration, transition-point monitoring, and full experiment lineage tracking. We demonstrate the effectiveness of multi-phase RL through representative scenarios and highlight RL-Studio’s capabilities. Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Jeonghye Kim, Yongjae Shin, Suhyun Jung, Hyundam Yoo, Chanwoo Moon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee |
AAAI | 4 |
| 2025 | ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State ReflectionabstractRecent advances in LLM agents have largely built on reasoning backbones like ReAct (Yao et al., 2023), which interleaves thought and action in complex environments.However, Re-Act often produces ungrounded or incoherent reasoning steps, leading to misalignment between the agent's actual state and goal.Our analysis finds that this stems from ReAct's inability to maintain consistent internal beliefs and goal alignment, causing compounding errors and hallucinations.To address this, we introduce ReflAct, a novel backbone that shifts reasoning from merely planning next actions to continuously reflecting on the agent's state relative to its goal.By explicitly grounding decisions in states and enforcing ongoing goal alignment, ReflAct dramatically improves strategic reliability.This design delivers substantial empirical gains: ReflAct surpasses Re-Act by 27.7% on average, achieving a 93.3% success rate in ALFWorld.Notably, ReflAct even outperforms ReAct with added enhancement modules (e.g., Reflexion, WKM), showing that strengthening the core reasoning backbone is key to reliable agent performance.Observation: You pick up the apple 1 from the countertop 1. Thought: Now I take an apple 1.Next, I need to go to a microwave 1 and heat it.Action: go to microwave 1 Observation: The microwave 1 is closed.Thought: The microwave 1 is closed, so I need to open it first.Action: open microwave 1 Observation: You open the microwave 1.The microwave 1 is open.In it, you see nothing. Jeonghye Kim, Sojeong Rhee, Minbeom Kim, Sangmook Lee, Youngchul Sung, Kyomin Jung |
EMNLP | 1 |
| 2025 | ARS: Adaptive Reward Scaling for Multi-Task Reinforcement LearningabstractMulti-task reinforcement learning (RL) encounters significant challenges due to varying task complexities and their reward distributions from the environment. To address these issues, in this paper, we propose Adaptive Reward Scaling (ARS), a novel framework that dynamically adjusts reward magnitudes and leverages a periodic network reset mechanism. ARS introduces a history-based reward scaling strategy that ensures balanced reward distributions across tasks, enabling stable and efficient training. The reset mechanism complements this approach by mitigating overfitting and ensuring robust convergence. Empirical evaluations on the Meta-World benchmark demonstrate that ARS significantly outperforms baseline methods, achieving superior performance on challenging tasks while maintaining overall learning efficiency. These results validate ARS’s effectiveness in tackling diverse multi-task RL problems, paving the way for scalable solutions in complex real-world applications. Myungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul Sung |
ICML | 3 |
| 2025 | Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline DataabstractReinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond the data range is particularly problematic. To mitigate this, we propose guiding the gradual decrease of Q-values outside the data range, which is achieved through reward scaling with layer normalization (RS-LN) and a penalization mechanism for infeasible actions (PA). By combining RS-LN and PA, we develop a new algorithm called PARS. We evaluate PARS across a range of tasks, demonstrating superior performance compared to state-of-the-art algorithms in both offline training and online fine-tuning on the D4RL benchmark, with notable success in the challenging AntMaze Ultra task. Jeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngchul Sung, Kanghoon Lee, Woohyung Lim |
ICML | 1 |
| 2025 | Online Pre-Training for Offline-to-Online Reinforcement LearningabstractOffline-to-online reinforcement learning (RL) aims to integrate the complementary strengths of offline and online RL by pre-training an agent offline and subsequently fine-tuning it through online interactions. However, recent studies reveal that offline pre-trained agents often underperform during online fine-tuning due to inaccurate value estimation caused by distribution shift, with random initialization proving more effective in certain cases. In this work, we propose a novel method, Online Pre-Training for Offline-to-Online RL (OPT), explicitly designed to address the issue of inaccurate value estimation in offline pre-trained agents. OPT introduces a new learning phase, Online Pre-Training, which allows the training of a new value function tailored specifically for effective online fine-tuning. Implementation of OPT on TD3 and SPOT demonstrates an average 30% improvement in performance across a wide range of D4RL environments, including MuJoCo, Antmaze, and Adroit. Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geon-Hyeong Kim, Jongseong Chae, Youngchul Sung, Kanghoon Lee, Woohyung Lim |
ICML | 2 |
| 2024 | Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision MakingabstractThe recent success of Transformer in natural language processing has sparked its use in various domains. In offline reinforcement learning (RL), Decision Transformer (DT) is emerging as a promising model based on Transformer. However, we discovered that the attention module of DT is not appropriate to capture the inherent local dependence pattern in trajectories of RL modeled as a Markov decision process. To overcome the limitations of DT, we propose a novel action sequence predictor, named Decision ConvFormer (DC), based on the architecture of MetaFormer, which is a general structure to process multiple entities in parallel and understand the interrelationship among the multiple entities. DC employs local convolution filtering as the token mixer and can effectively capture the inherent local associations of the RL dataset. In extensive experiments, DC achieved state-of-the-art performance across various standard RL benchmarks while requiring fewer resources. Furthermore, we show that DC better understands the underlying meaning in data and exhibits enhanced generalization capability. Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung |
ICLR | 1 |
| 2024 | Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement LearningabstractOffline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability of $Q$-functions. By analyzing $Q$-function over-generalization, which impairs stable stitching, QCS adaptively integrates $Q$-aid into RCSL's loss function based on trajectory return. Empirical results show that QCS significantly outperforms RCSL and value-based methods, consistently achieving or exceeding the highest trajectory returns across diverse offline RL benchmarks. QCS represents a breakthrough in offline RL, pushing the limits of what can be achieved and fostering further innovations. Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung |
NeurIPS | 1 |
| 2023 | LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option FrameworkabstractIn this paper, a unified framework for exploration in reinforcement learning (RL) is proposed based on an option-critic architecture. The proposed framework learns to integrate a set of diverse exploration strategies so that the agent can adaptively select the most effective exploration strategy to realize an effective exploration-exploitation trade-off for each given task. The effectiveness of the proposed exploration framework is demonstrated by various experiments in the MiniGrid and Atari environments. Woojun Kim, Jeonghye Kim, Youngchul Sung |
ICML | 2 |