EDBT 2026 Demo / reviewers in the wild / expert
Deunsol Yoon
dblp:225/5388
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 81% Deep learning architectures and training · 10% Transfer learning and domain adaptation · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 67% Energy systems and smart grids · 33% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.9 | 2 | 2026 | RAPID: A Rapid Prototyping Platform for Industrial Automation · AAAI 2026 Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning
actor-critic methods |
1.4 | 2 | 2025 | Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning · ICML 2025 Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic · ICLR 2021 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
1.0 | 1 | 2026 | RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation · AAAI 2026 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.9 | 1 | 2025 | Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025 |
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning |
0.9 | 1 | 2025 | Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › reward design
reward scaling |
0.9 | 1 | 2025 | Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data · ICML 2025 |
Machine learning › Reinforcement learning › value function estimation
value estimation bias |
0.9 | 1 | 2025 | Online Pre-Training for Offline-to-Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.6 | 1 | 2022 | Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022 |
Machine learning › Deep learning architectures and training › transformer
structure-aware transformer |
0.6 | 1 | 2022 | Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.6 | 1 | 2022 | Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022 |
Machine learning › Reinforcement learning › deep reinforcement learning
transformer-based policy |
0.6 | 1 | 2022 | Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning · ICLR 2022 |
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments |
0.3 | 1 | 2026 | RAPID: A Rapid Prototyping Platform for Industrial Automation · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution |
0.3 | 1 | 2025 | Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement Learning · ICML 2025 |
Energy systems and smart grids › power system planning and operation
power grid management |
0.1 | 1 | 2021 | Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic · ICLR 2021 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 3.0behavior simulation · 2.0phase orchestration · 1.0parameter transfer · 1.0q-learning · 0.9layer normalization · 0.9generalized advantage estimation · 0.9attention-based aggregation · 0.9actor-critic · 0.9PPO · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAPID: A Rapid Prototyping Platform for Industrial AutomationabstractIndustrial automation in smart logistics and factories requires simulation platforms that support rapid environment building before costly physical deployment. Yet existing tools often require substantial expertise, complex setup, and long configuration times, hindering agile prototyping. We present RAPID, a simulation platform with two components: layout design, which enables intuitive visual configuration of factory layouts, and behavior simulation and validation, which allows users to attach behavior models and evaluate system performance. RAPID lowers the entry barrier to industrial simulation, letting users apply existing behavior models or trained reinforcement learning (RL) agents to new layouts with minimal effort. This approach lets practitioners prototype facilities in minutes rather than weeks and gives researchers a standardized environment for benchmarking multi-agent RL and coordination algorithms. By combining rapid design with simulation-based validation, RAPID accelerates automation development from concept to implementation. Sunghoon Hong, Whiyoung Jung, Deunsol Yoon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee |
AAAI | 4 |
| 2026 | RL-Studio: A System for Multi-Phase Reinforcement Learning ExperimentationabstractReinforcement learning (RL) has evolved beyond monolithic training, yet existing frameworks remain limited to single algorithms or simple offline-to-online transitions. We present multi-phase RL, a framework that orchestrates multiple learning phases for continual policy improvement. It enables efficient fine-tuning of pretrained policies with new data and smooth adaptation from simulation to real-world environments. To support this paradigm, we introduce RL-Studio, a platform that addresses key implementation barriers, including neural architecture mismatches, parameter transfer complexities, and experiment management overhead. It provides phase orchestration, transition-point monitoring, and full experiment lineage tracking. We demonstrate the effectiveness of multi-phase RL through representative scenarios and highlight RL-Studio’s capabilities. Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Jeonghye Kim, Yongjae Shin, Suhyun Jung, Hyundam Yoo, Chanwoo Moon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee |
AAAI | 3 |
| 2025 | Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) struggles with coordination in sparse reward environments. Macro-actions —sequences of actions executed as single decisions— facilitate long-term planning but introduce asynchrony, complicating Centralized Training with Decentralized Execution (CTDE). Existing CTDE methods use padding to handle asynchrony, risking misaligned asynchronous experiences and spurious correlations. We propose the Agent-Centric Actor-Critic (ACAC) algorithm to manage asynchrony without padding. ACAC uses agent-centric encoders for independent trajectory processing, with an attention-based aggregation module integrating these histories into a centralized critic for improved temporal abstractions. The proposed structure is trained via a PPO-based algorithm with a modified Generalized Advantage Estimation for asynchronous environments. Experiments show ACAC accelerates convergence and enhances performance over baselines in complex MARL tasks. Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Kanghoon Lee, Woohyung Lim |
ICML | 3 |
| 2025 | Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline DataabstractReinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond the data range is particularly problematic. To mitigate this, we propose guiding the gradual decrease of Q-values outside the data range, which is achieved through reward scaling with layer normalization (RS-LN) and a penalization mechanism for infeasible actions (PA). By combining RS-LN and PA, we develop a new algorithm called PARS. We evaluate PARS across a range of tasks, demonstrating superior performance compared to state-of-the-art algorithms in both offline training and online fine-tuning on the D4RL benchmark, with notable success in the challenging AntMaze Ultra task. Jeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngchul Sung, Kanghoon Lee, Woohyung Lim |
ICML | 5 |
| 2025 | Online Pre-Training for Offline-to-Online Reinforcement LearningabstractOffline-to-online reinforcement learning (RL) aims to integrate the complementary strengths of offline and online RL by pre-training an agent offline and subsequently fine-tuning it through online interactions. However, recent studies reveal that offline pre-trained agents often underperform during online fine-tuning due to inaccurate value estimation caused by distribution shift, with random initialization proving more effective in certain cases. In this work, we propose a novel method, Online Pre-Training for Offline-to-Online RL (OPT), explicitly designed to address the issue of inaccurate value estimation in offline pre-trained agents. OPT introduces a new learning phase, Online Pre-Training, which allows the training of a new value function tailored specifically for effective online fine-tuning. Implementation of OPT on TD3 and SPOT demonstrates an average 30% improvement in performance across a wide range of D4RL environments, including MuJoCo, Antmaze, and Adroit. Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geon-Hyeong Kim, Jongseong Chae, Youngchul Sung, Kanghoon Lee, Woohyung Lim |
ICML | 5 |
| 2025 | Hierarchical Decomposition Framework for Steiner Tree Packing Problem
Hanbum Ko, Minu Kim 0001, Han-Seul Jeong, Sunghoon Hong, Deunsol Yoon, Youngjoon Park, Woohyung Lim, Honglak Lee, Moontae Lee, Kanghoon Lee, Sungbin Lim, Sungryull Sohn |
ICORES | 5 |
| 2022 | Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning
Sunghoon Hong, Deunsol Yoon, Kee-Eung Kim |
ICLR | 2 |
| 2021 | Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic
Deunsol Yoon, Sunghoon Hong, Byung-Jun Lee 0001, Kee-Eung Kim |
ICLR | 1 |