EDBT 2026 Demo / reviewers in the wild / expert
Seungyul Han
dblp:183/6417
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-5376-1976ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Reinforcement learning · 92% Trustworthy machine learning · 8% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
1.8 | 3 | 2024 | FoX: Formation-Aware Exploration in Multi-Agent Reinforcement Learning · AAAI 2024 A Max-Min Entropy Framework for Reinforcement Learning · NeurIPS 2021 Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration · ICML 2021 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.6 | 2 | 2025 | Wolfpack Adversarial Attack for Robust Multi-Agent Reinforcement Learning · ICML 2025 FoX: Formation-Aware Exploration in Multi-Agent Reinforcement Learning · AAAI 2024 |
Machine learning › Reinforcement learning
imitation learning |
1.2 | 2 | 2023 | Domain Adaptive Imitation Learning with Visual Observation · NeurIPS 2023 Robust Imitation Learning against Variations in Environment Dynamics · ICML 2022 |
Machine learning › Reinforcement learning
actor-critic methods |
1.0 | 2 | 2021 | A Max-Min Entropy Framework for Reinforcement Learning · NeurIPS 2021 Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration · ICML 2021 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
0.9 | 1 | 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.9 | 1 | 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
robust multi-agent reinforcement learning |
0.9 | 1 | 2025 | Wolfpack Adversarial Attack for Robust Multi-Agent Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
task representation learning |
0.9 | 1 | 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks · ICML 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Exclusively Penalized Q-learning for Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
underestimation bias |
0.8 | 1 | 2024 | Exclusively Penalized Q-learning for Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › imitation learning › transfer imitation learning
cross-domain imitation learning |
0.7 | 1 | 2023 | Domain Adaptive Imitation Learning with Visual Observation · NeurIPS 2023 |
Machine learning › Reinforcement learning › imitation learning › transfer imitation learning
domain adaptive imitation learning |
0.7 | 1 | 2023 | Domain Adaptive Imitation Learning with Visual Observation · NeurIPS 2023 |
Machine learning › Reinforcement learning › imitation learning
robust imitation learning |
0.6 | 1 | 2022 | Robust Imitation Learning against Variations in Environment Dynamics · ICML 2022 |
Machine learning › Reinforcement learning › regularization for reinforcement learning
entropy regularization |
0.5 | 1 | 2021 | Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration · ICML 2021 |
Machine learning › Reinforcement learning
maximum entropy reinforcement learning |
0.5 | 1 | 2021 | A Max-Min Entropy Framework for Reinforcement Learning · NeurIPS 2021 |
Machine learning › Reinforcement learning › actor-critic methods
soft actor-critic |
0.5 | 1 | 2021 | A Max-Min Entropy Framework for Reinforcement Learning · NeurIPS 2021 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2019 | Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning · ICML 2019 |
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization |
0.4 | 1 | 2019 | Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning · ICML 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | Wolfpack Adversarial Attack for Robust Multi-Agent Reinforcement Learning · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
state regularization · 0.9metric-based representation learning · 0.9adversarial training · 0.9adversarial attack · 0.9value function penalization · 0.8formation-based equivalence relation · 0.8exclusively penalized q-learning · 0.8deep multi-agent reinforcement learning · 0.8image reconstruction · 0.7dual feature extraction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wolfpack Adversarial Attack for Robust Multi-Agent Reinforcement LearningabstractTraditional robust methods in multi-agent reinforcement learning (MARL) often struggle against coordinated adversarial attacks in cooperative scenarios. To address this limitation, we propose the Wolfpack Adversarial Attack framework, inspired by wolf hunting strategies, which targets an initial agent and its assisting agents to disrupt cooperation. Additionally, we introduce the Wolfpack-Adversarial Learning for MARL (WALL) framework, which trains robust MARL policies to defend against the proposed Wolfpack attack by fostering system-wide collaboration. Experimental results underscore the devastating impact of the Wolfpack attack and the significant robustness improvements achieved by WALL. Our code is available at https://github.com/sunwoolee0504/WALL. Sunwoo Lee 0006, Jaebak Hwang, Yonghyeon Jo, Seungyul Han |
ICML | 4 |
| 2025 | Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution TasksabstractMeta reinforcement learning aims to develop policies that generalize to unseen tasks sampled from a task distribution. While context-based meta-RL methods improve task representation using task latents, they often struggle with out-of-distribution (OOD) tasks. To address this, we propose Task-Aware Virtual Training (TAVT), a novel algorithm that accurately captures task characteristics for both training and OOD scenarios using metric-based representation learning. Our method successfully preserves task characteristics in virtual tasks and employs a state regularization technique to mitigate overestimation errors in state-varying environments. Numerical results demonstrate that TAVT significantly enhances generalization to OOD tasks across various MuJoCo and MetaWorld environments. Our code is available at https://github.com/JM-Kim-94/tavt.git. Jeongmo Kim, Yisak Park, Minung Kim, Seungyul Han |
ICML | 4 |
| 2025 | Adaptive multi-model fusion learning for sparse-reward reinforcement learning
Giseung Park, Whiyoung Jung, Seungyul Han, Sungho Choi, Youngchul Sung |
Neurocomputing | 3 |
| 2024 | FoX: Formation-Aware Exploration in Multi-Agent Reinforcement LearningabstractRecently, deep multi-agent reinforcement learning (MARL) has gained significant popularity due to its success in various cooperative multi-agent tasks. However, exploration still remains a challenging problem in MARL due to the partial observability of the agents and the exploration space that can grow exponentially as the number of agents increases. Firstly, in order to address the scalability issue of the exploration space, we define a formation-based equivalence relation on the exploration space and aim to reduce the search space by exploring only meaningful states in different formations. Then, we propose a novel formation-aware exploration (FoX) framework that encourages partially observable agents to visit the states in diverse formations by guiding them to be well aware of their current formation solely based on their own observations. Numerical results show that the proposed FoX framework significantly outperforms the state-of-the-art MARL algorithms on Google Research Football (GRF) and sparse Starcraft II multi-agent challenge (SMAC) tasks. Yonghyeon Jo, Sunwoo Lee 0006, Junghyuk Yeom, Seungyul Han |
AAAI | 4 |
| 2024 | Exclusively Penalized Q-learning for Offline Reinforcement LearningabstractConstraint-based offline reinforcement learning (RL) involves policy constraints or imposing penalties on the value function to mitigate overestimation errors caused by distributional shift. This paper focuses on a limitation in existing offline RL methods with penalized value function, indicating the potential for underestimation bias due to unnecessary bias introduced in the value function. To address this concern, we propose Exclusively Penalized Q-learning (EPQ), which reduces estimation bias in the value function by selectively penalizing states that are prone to inducing estimation errors. Numerical results show that our method significantly reduces underestimation bias and improves performance in various offline control tasks compared to other offline RL methods. Junghyuk Yeom, Yonghyeon Jo, Jeongmo Kim, Seungyul Han |
NeurIPS | 5 |
| 2023 | Domain Adaptive Imitation Learning with Visual ObservationabstractIn this paper, we consider domain-adaptive imitation learning with visual observation, where an agent in a target domain learns to perform a task by observing expert demonstrations in a source domain. Domain adaptive imitation learning arises in practical scenarios where a robot, receiving visual sensory data, needs to mimic movements by visually observing other robots from different angles or observing robots of different shapes. To overcome the domain shift in cross-domain imitation learning with visual observation, we propose a novel framework for extracting domain-independent behavioral features from input observations that can be used to train the learner, based on dual feature extraction and image reconstruction. Empirical results demonstrate that our approach outperforms previous algorithms for imitation learning from visual observation with domain shift. Sungho Choi, Seungyul Han, Woojun Kim, Jongseong Chae, Whiyoung Jung, Youngchul Sung |
NeurIPS | 2 |
| 2022 | Robust Imitation Learning against Variations in Environment DynamicsabstractIn this paper, we propose a robust imitation learning (IL) framework that improves the robustness of IL when environment dynamics are perturbed. The existing IL framework trained in a single environment can catastrophically fail with perturbations in environment dynamics because it does not capture the situation that underlying environment dynamics can be changed. Our framework effectively deals with environments with varying dynamics by imitating multiple experts in sampled environment dynamics to enhance the robustness in general variations in environment dynamics. In order to robustly imitate the multiple sample experts, we minimize the risk with respect to the Jensen-Shannon divergence between the agent’s policy and each of the sample experts. Numerical results show that our algorithm significantly improves robustness against dynamics perturbations compared to conventional IL baselines. Jongseong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho, Sungho Choi, Youngchul Sung |
ICML | 2 |
| 2021 | Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient ExplorationabstractIn this paper, sample-aware policy entropy regularization is proposed to enhance the conventional policy entropy regularization for better exploration. Exploiting the sample distribution obtainable from the replay buffer, the proposed sample-aware entropy regularization maximizes the entropy of the weighted sum of the policy action distribution and the sample action distribution from the replay buffer for sample-efficient exploration. A practical algorithm named diversity actor-critic (DAC) is developed by applying policy iteration to the objective function with the proposed sample-aware entropy regularization. Numerical results show that DAC significantly outperforms existing recent algorithms for reinforcement learning. Seungyul Han, Youngchul Sung |
ICML | 1 |
| 2021 | A Max-Min Entropy Framework for Reinforcement LearningabstractIn this paper, we propose a max-min entropy framework for reinforcement learning (RL) to overcome the limitation of the soft actor-critic (SAC) algorithm implementing the maximum entropy RL in model-free sample-based learning. Whereas the maximum entropy RL guides learning for policies to reach states with high entropy in the future, the proposed max-min entropy framework aims to learn to visit states with low entropy and maximize the entropy of these low-entropy states to promote better exploration. For general Markov decision processes (MDPs), an efficient algorithm is constructed under the proposed max-min entropy framework based on disentanglement of exploration and exploitation. Numerical results show that the proposed algorithm yields drastic performance improvement over the current state-of-the-art RL algorithms. Seungyul Han, Youngchul Sung |
NeurIPS | 1 |
| 2019 | Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement LearningabstractIn importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning. However, policy update from clipped statistics induces large bias in tasks with high action dimensions, and bias from clipping makes it difficult to reuse old samples with large IS weights. In this paper, we consider PPO, a representative on-policy algorithm, and propose its improvement by dimension-wise IS weight clipping which separately clips the IS weight of each action dimension to avoid large bias and adaptively controls the IS weight to bound policy update from the current policy. This new technique enables efficient learning for high action-dimensional tasks and reusing of old samples like in off-policy learning to increase the sample efficiency. Numerical results show that the proposed new algorithm outperforms PPO and other RL algorithms in various Open AI Gym tasks. Seungyul Han, Youngchul Sung |
ICML | 1 |