Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Deunsol Yoon

dblp:225/5388 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 81% Deep learning architectures and training · 10% Transfer learning and domain adaptation · 9%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 67% Energy systems and smart grids · 33%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.922026
RAPID: A Rapid Prototyping Platform for Industrial Automation · AAAI 2026
Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning
actor-critic methods
1.422025
Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning · ICML 2025
Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic · ICLR 2021
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
1.012026
RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation · AAAI 2026
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.912025
Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › reward design
reward scaling
0.912025
Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025
Machine learning › Reinforcement learning › value function estimation
value estimation bias
0.912025
Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.612022
Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022
Machine learning › Deep learning architectures and training › transformer
structure-aware transformer
0.612022
Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning › deep reinforcement learning
transformer-based policy
0.612022
Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.312026
RAPID: A Rapid Prototyping Platform for Industrial Automation · AAAI 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.312025
Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning · ICML 2025
Energy systems and smart grids › power system planning and operation
power grid management
0.112021
Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic · ICLR 2021

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 3.0behavior simulation · 2.0phase orchestration · 1.0parameter transfer · 1.0q-learning · 0.9layer normalization · 0.9generalized advantage estimation · 0.9attention-based aggregation · 0.9actor-critic · 0.9PPO · 0.9
YearPublicationVenuePosition
2026 RAPID: A Rapid Prototyping Platform for Industrial Automation
abstract
Industrial automation in smart logistics and factories requires simulation platforms that support rapid environment building before costly physical deployment. Yet existing tools often require substantial expertise, complex setup, and long configuration times, hindering agile prototyping. We present RAPID, a simulation platform with two components: layout design, which enables intuitive visual configuration of factory layouts, and behavior simulation and validation, which allows users to attach behavior models and evaluate system performance. RAPID lowers the entry barrier to industrial simulation, letting users apply existing behavior models or trained reinforcement learning (RL) agents to new layouts with minimal effort. This approach lets practitioners prototype facilities in minutes rather than weeks and gives researchers a standardized environment for benchmarking multi-agent RL and coordination algorithms. By combining rapid design with simulation-based validation, RAPID accelerates automation development from concept to implementation.
Sunghoon Hong, Whiyoung Jung, Deunsol Yoon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee
AAAI4
2026 RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation
abstract
Reinforcement learning (RL) has evolved beyond monolithic training, yet existing frameworks remain limited to single algorithms or simple offline-to-online transitions. We present multi-phase RL, a framework that orchestrates multiple learning phases for continual policy improvement. It enables efficient fine-tuning of pretrained policies with new data and smooth adaptation from simulation to real-world environments. To support this paradigm, we introduce RL-Studio, a platform that addresses key implementation barriers, including neural architecture mismatches, parameter transfer complexities, and experiment management overhead. It provides phase orchestration, transition-point monitoring, and full experiment lineage tracking. We demonstrate the effectiveness of multi-phase RL through representative scenarios and highlight RL-Studio’s capabilities.
Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Jeonghye Kim, Yongjae Shin, Suhyun Jung, Hyundam Yoo, Chanwoo Moon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee
AAAI3
2025 Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning
abstract
Multi-Agent Reinforcement Learning (MARL) struggles with coordination in sparse reward environments. Macro-actions —sequences of actions executed as single decisions— facilitate long-term planning but introduce asynchrony, complicating Centralized Training with Decentralized Execution (CTDE). Existing CTDE methods use padding to handle asynchrony, risking misaligned asynchronous experiences and spurious correlations. We propose the Agent-Centric Actor-Critic (ACAC) algorithm to manage asynchrony without padding. ACAC uses agent-centric encoders for independent trajectory processing, with an attention-based aggregation module integrating these histories into a centralized critic for improved temporal abstractions. The proposed structure is trained via a PPO-based algorithm with a modified Generalized Advantage Estimation for asynchronous environments. Experiments show ACAC accelerates convergence and enhances performance over baselines in complex MARL tasks.
Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Kanghoon Lee, Woohyung Lim
ICML3
2025 Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data
abstract
Reinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond the data range is particularly problematic. To mitigate this, we propose guiding the gradual decrease of Q-values outside the data range, which is achieved through reward scaling with layer normalization (RS-LN) and a penalization mechanism for infeasible actions (PA). By combining RS-LN and PA, we develop a new algorithm called PARS. We evaluate PARS across a range of tasks, demonstrating superior performance compared to state-of-the-art algorithms in both offline training and online fine-tuning on the D4RL benchmark, with notable success in the challenging AntMaze Ultra task.
Jeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngchul Sung, Kanghoon Lee, Woohyung Lim
ICML5
2025 Online Pre-Training for Offline-to-Online Reinforcement Learning
abstract
Offline-to-online reinforcement learning (RL) aims to integrate the complementary strengths of offline and online RL by pre-training an agent offline and subsequently fine-tuning it through online interactions. However, recent studies reveal that offline pre-trained agents often underperform during online fine-tuning due to inaccurate value estimation caused by distribution shift, with random initialization proving more effective in certain cases. In this work, we propose a novel method, Online Pre-Training for Offline-to-Online RL (OPT), explicitly designed to address the issue of inaccurate value estimation in offline pre-trained agents. OPT introduces a new learning phase, Online Pre-Training, which allows the training of a new value function tailored specifically for effective online fine-tuning. Implementation of OPT on TD3 and SPOT demonstrates an average 30% improvement in performance across a wide range of D4RL environments, including MuJoCo, Antmaze, and Adroit.
Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geon-Hyeong Kim, Jongseong Chae, Youngchul Sung, Kanghoon Lee, Woohyung Lim
ICML5
2025 Hierarchical Decomposition Framework for Steiner Tree Packing Problem
Hanbum Ko, Minu Kim 0001, Han-Seul Jeong, Sunghoon Hong, Deunsol Yoon, Youngjoon Park, Woohyung Lim, Honglak Lee, Moontae Lee, Kanghoon Lee, Sungbin Lim, Sungryull Sohn
ICORES5
2022 Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning
Sunghoon Hong, Deunsol Yoon, Kee-Eung Kim
ICLR2
2021 Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic
Deunsol Yoon, Sunghoon Hong, Byung-Jun Lee 0001, Kee-Eung Kim
ICLR1