Jeonghye Kim

dblp:172/6718 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 76% Transfer learning and domain adaptation · 7% Language models and text generation · 6%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
offline reinforcement learning
2.432025
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025
Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning · NeurIPS 2024
Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making · ICLR 2024
Machine learning › Reinforcement learning › reward design
reward scaling
1.722025
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025
ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning · ICML 2025
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
1.012026
RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation · AAAI 2026
Natural language and speech › Language models and text generation
LLM agents
0.912025
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection · EMNLP 2025
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.912025
ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.912025
Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › value function estimation
value estimation bias
0.912025
Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.812024
Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
return-conditioned supervised learning
0.812024
Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
exploration
0.712023
LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework · ICML 2023
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.712023
LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework · ICML 2023
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework
0.712023
LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework · ICML 2023

Methods — techniques the papers use, named apart from their topics

phase orchestration · 1.0parameter transfer · 1.0reflection · 0.9react · 0.9q-learning · 0.9periodic network reset · 0.9layer normalization · 0.9large language model · 0.9adaptive reward scaling · 0.9SPOT · 0.9
YearPublicationVenuePosition
2026 RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation
abstract
Reinforcement learning (RL) has evolved beyond monolithic training, yet existing frameworks remain limited to single algorithms or simple offline-to-online transitions. We present multi-phase RL, a framework that orchestrates multiple learning phases for continual policy improvement. It enables efficient fine-tuning of pretrained policies with new data and smooth adaptation from simulation to real-world environments. To support this paradigm, we introduce RL-Studio, a platform that addresses key implementation barriers, including neural architecture mismatches, parameter transfer complexities, and experiment management overhead. It provides phase orchestration, transition-point monitoring, and full experiment lineage tracking. We demonstrate the effectiveness of multi-phase RL through representative scenarios and highlight RL-Studio’s capabilities.
Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Jeonghye Kim, Yongjae Shin, Suhyun Jung, Hyundam Yoo, Chanwoo Moon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee
AAAI4
2025 ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
abstract
Recent advances in LLM agents have largely built on reasoning backbones like ReAct (Yao et al., 2023), which interleaves thought and action in complex environments.However, Re-Act often produces ungrounded or incoherent reasoning steps, leading to misalignment between the agent's actual state and goal.Our analysis finds that this stems from ReAct's inability to maintain consistent internal beliefs and goal alignment, causing compounding errors and hallucinations.To address this, we introduce ReflAct, a novel backbone that shifts reasoning from merely planning next actions to continuously reflecting on the agent's state relative to its goal.By explicitly grounding decisions in states and enforcing ongoing goal alignment, ReflAct dramatically improves strategic reliability.This design delivers substantial empirical gains: ReflAct surpasses Re-Act by 27.7% on average, achieving a 93.3% success rate in ALFWorld.Notably, ReflAct even outperforms ReAct with added enhancement modules (e.g., Reflexion, WKM), showing that strengthening the core reasoning backbone is key to reliable agent performance.Observation: You pick up the apple 1 from the countertop 1. Thought: Now I take an apple 1.Next, I need to go to a microwave 1 and heat it.Action: go to microwave 1 Observation: The microwave 1 is closed.Thought: The microwave 1 is closed, so I need to open it first.Action: open microwave 1 Observation: You open the microwave 1.The microwave 1 is open.In it, you see nothing.
Jeonghye Kim, Sojeong Rhee, Minbeom Kim, Sangmook Lee, Youngchul Sung, Kyomin Jung
EMNLP1
2025 ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning
abstract
Multi-task reinforcement learning (RL) encounters significant challenges due to varying task complexities and their reward distributions from the environment. To address these issues, in this paper, we propose Adaptive Reward Scaling (ARS), a novel framework that dynamically adjusts reward magnitudes and leverages a periodic network reset mechanism. ARS introduces a history-based reward scaling strategy that ensures balanced reward distributions across tasks, enabling stable and efficient training. The reset mechanism complements this approach by mitigating overfitting and ensuring robust convergence. Empirical evaluations on the Meta-World benchmark demonstrate that ARS significantly outperforms baseline methods, achieving superior performance on challenging tasks while maintaining overall learning efficiency. These results validate ARS’s effectiveness in tackling diverse multi-task RL problems, paving the way for scalable solutions in complex real-world applications.
Myungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul Sung
ICML3
2025 Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data
abstract
Reinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond the data range is particularly problematic. To mitigate this, we propose guiding the gradual decrease of Q-values outside the data range, which is achieved through reward scaling with layer normalization (RS-LN) and a penalization mechanism for infeasible actions (PA). By combining RS-LN and PA, we develop a new algorithm called PARS. We evaluate PARS across a range of tasks, demonstrating superior performance compared to state-of-the-art algorithms in both offline training and online fine-tuning on the D4RL benchmark, with notable success in the challenging AntMaze Ultra task.
Jeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngchul Sung, Kanghoon Lee, Woohyung Lim
ICML1
2025 Online Pre-Training for Offline-to-Online Reinforcement Learning
abstract
Offline-to-online reinforcement learning (RL) aims to integrate the complementary strengths of offline and online RL by pre-training an agent offline and subsequently fine-tuning it through online interactions. However, recent studies reveal that offline pre-trained agents often underperform during online fine-tuning due to inaccurate value estimation caused by distribution shift, with random initialization proving more effective in certain cases. In this work, we propose a novel method, Online Pre-Training for Offline-to-Online RL (OPT), explicitly designed to address the issue of inaccurate value estimation in offline pre-trained agents. OPT introduces a new learning phase, Online Pre-Training, which allows the training of a new value function tailored specifically for effective online fine-tuning. Implementation of OPT on TD3 and SPOT demonstrates an average 30% improvement in performance across a wide range of D4RL environments, including MuJoCo, Antmaze, and Adroit.
Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geon-Hyeong Kim, Jongseong Chae, Youngchul Sung, Kanghoon Lee, Woohyung Lim
ICML2
2024 Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making
abstract
The recent success of Transformer in natural language processing has sparked its use in various domains. In offline reinforcement learning (RL), Decision Transformer (DT) is emerging as a promising model based on Transformer. However, we discovered that the attention module of DT is not appropriate to capture the inherent local dependence pattern in trajectories of RL modeled as a Markov decision process. To overcome the limitations of DT, we propose a novel action sequence predictor, named Decision ConvFormer (DC), based on the architecture of MetaFormer, which is a general structure to process multiple entities in parallel and understand the interrelationship among the multiple entities. DC employs local convolution filtering as the token mixer and can effectively capture the inherent local associations of the RL dataset. In extensive experiments, DC achieved state-of-the-art performance across various standard RL benchmarks while requiring fewer resources. Furthermore, we show that DC better understands the underlying meaning in data and exhibits enhanced generalization capability.
Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung
ICLR1
2024 Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability of $Q$-functions. By analyzing $Q$-function over-generalization, which impairs stable stitching, QCS adaptively integrates $Q$-aid into RCSL's loss function based on trajectory return. Empirical results show that QCS significantly outperforms RCSL and value-based methods, consistently achieving or exceeding the highest trajectory returns across diverse offline RL benchmarks. QCS represents a breakthrough in offline RL, pushing the limits of what can be achieved and fostering further innovations.
Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung
NeurIPS1
2023 LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework
abstract
In this paper, a unified framework for exploration in reinforcement learning (RL) is proposed based on an option-critic architecture. The proposed framework learns to integrate a set of diverse exploration strategies so that the agent can adaptively select the most effective exploration strategy to realize an effective exploration-exploitation trade-off for each given task. The effectiveness of the proposed exploration framework is demonstrated by various experiments in the MiniGrid and Atari environments.
Woojun Kim, Jeonghye Kim, Youngchul Sung
ICML2