Myungsik Cho

dblp:233/3959 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 84% Planning, search and constraint satisfaction · 16%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-task reinforcement learning
1.622025
ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning · ICML 2025
Hard Tasks First: Multi-Task Reinforcement Learning Through Task Scheduling · ICML 2024
Machine learning › Reinforcement learning › reward design
reward scaling
0.912025
ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › scheduling
task scheduling
0.812024
Hard Tasks First: Multi-Task Reinforcement Learning Through Task Scheduling · ICML 2024
Machine learning › Reinforcement learning
meta-reinforcement learning
0.712023
Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask Decomposition · NeurIPS 2023
Machine learning › Reinforcement learning › hierarchical reinforcement learning
sub-task decomposition
0.712023
Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask Decomposition · NeurIPS 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › hierarchical problem solving
task decomposition
0.712023
Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask Decomposition · NeurIPS 2023
Machine learning › Reinforcement learning
constrained reinforcement learning
0.612022
Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability · NeurIPS 2022
Machine learning › Reinforcement learning
imitation learning
0.612022
Robust Imitation Learning against Variations in Environment Dynamics · ICML 2022
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.612022
Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability · NeurIPS 2022
Machine learning › Reinforcement learning › imitation learning
robust imitation learning
0.612022
Robust Imitation Learning against Variations in Environment Dynamics · ICML 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.412019
Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning · AAAI 2019
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication
0.412019
Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning · AAAI 2019
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412019
Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning · AAAI 2019
Machine learning › Reinforcement learning › value-based reinforcement learning
distributional reinforcement learning
0.212022
Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

periodic network reset · 0.9adaptive reward scaling · 0.9network reset · 0.8dynamic task prioritization · 0.8virtual training · 0.7Gaussian mixture VAE · 0.7quantile estimation · 0.6large deviation principle · 0.6lagrange multipliers · 0.6jensen-shannon divergence minimization · 0.6
YearPublicationVenuePosition
2025 ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning
abstract
Multi-task reinforcement learning (RL) encounters significant challenges due to varying task complexities and their reward distributions from the environment. To address these issues, in this paper, we propose Adaptive Reward Scaling (ARS), a novel framework that dynamically adjusts reward magnitudes and leverages a periodic network reset mechanism. ARS introduces a history-based reward scaling strategy that ensures balanced reward distributions across tasks, enabling stable and efficient training. The reset mechanism complements this approach by mitigating overfitting and ensuring robust convergence. Empirical evaluations on the Meta-World benchmark demonstrate that ARS significantly outperforms baseline methods, achieving superior performance on challenging tasks while maintaining overall learning efficiency. These results validate ARS’s effectiveness in tackling diverse multi-task RL problems, paving the way for scalable solutions in complex real-world applications.
Myungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul Sung
ICML1
2024 Hard Tasks First: Multi-Task Reinforcement Learning Through Task Scheduling
abstract
Multi-task reinforcement learning (RL) faces the significant challenge of varying task difficulties, often leading to negative transfer when simpler tasks overshadow the learning of more complex ones. To overcome this challenge, we propose a novel algorithm, Scheduled Multi-Task Training (SMT), that strategically prioritizes more challenging tasks, thereby enhancing overall learning efficiency. SMT introduces a dynamic task prioritization strategy, underpinned by an effective metric for assessing task difficulty. This metric ensures an efficient and targeted allocation of training resources, significantly improving learning outcomes. Additionally, SMT incorporates a reset mechanism that periodically reinitializes key network parameters to mitigate the simplicity bias, further enhancing the adaptability and robustness of the learning process across diverse tasks. The efficacy of SMT's scheduling method is validated by significantly improving performance on challenging Meta-World benchmarks.
Myungsik Cho, Jongeui Park, Suyoung Lee, Youngchul Sung
ICML1
2023 Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask Decomposition
abstract
Meta-reinforcement learning (meta-RL) techniques have demonstrated remarkable success in generalizing deep reinforcement learning across a range of tasks. Nevertheless, these methods often struggle to generalize beyond tasks with parametric variations. To overcome this challenge, we propose Subtask Decomposition and Virtual Training (SDVT), a novel meta-RL approach that decomposes each non-parametric task into a collection of elementary subtasks and parameterizes the task based on its decomposition. We employ a Gaussian mixture VAE to meta-learn the decomposition process, enabling the agent to reuse policies acquired from common subtasks. Additionally, we propose a virtual training procedure, specifically designed for non-parametric task variability, which generates hypothetical subtask compositions, thereby enhancing generalization to previously unseen subtask compositions. Our method significantly improves performance on the Meta-World ML-10 and ML-45 benchmarks, surpassing current state-of-the-art techniques.
Suyoung Lee, Myungsik Cho, Youngchul Sung
NeurIPS2
2022 Robust Imitation Learning against Variations in Environment Dynamics
abstract
In this paper, we propose a robust imitation learning (IL) framework that improves the robustness of IL when environment dynamics are perturbed. The existing IL framework trained in a single environment can catastrophically fail with perturbations in environment dynamics because it does not capture the situation that underlying environment dynamics can be changed. Our framework effectively deals with environments with varying dynamics by imitating multiple experts in sampled environment dynamics to enhance the robustness in general variations in environment dynamics. In order to robustly imitate the multiple sample experts, we minimize the risk with respect to the Jensen-Shannon divergence between the agent’s policy and each of the sample experts. Numerical results show that our algorithm significantly improves robustness against dynamics perturbations compared to conventional IL baselines.
Jongseong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho, Sungho Choi, Youngchul Sung
ICML4
2022 Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability
abstract
Constrained reinforcement learning (RL) is an area of RL whose objective is to find an optimal policy that maximizes expected cumulative return while satisfying a given constraint. Most of the previous constrained RL works consider expected cumulative sum cost as the constraint. However, optimization with this constraint cannot guarantee a target probability of outage event that the cumulative sum cost exceeds a given threshold. This paper proposes a framework, named Quantile Constrained RL (QCRL), to constrain the quantile of the distribution of the cumulative sum cost that is a necessary and sufficient condition to satisfy the outage constraint. This is the first work that tackles the issue of applying the policy gradient theorem to the quantile and provides theoretical results for approximating the gradient of the quantile. Based on the derived theoretical results and the technique of the Lagrange multiplier, we construct a constrained RL algorithm named Quantile Constrained Policy Optimization (QCPO). We use distributional RL with the Large Deviation Principle (LDP) to estimate quantiles and tail probability of the cumulative sum cost for the implementation of QCPO. The implemented algorithm satisfies the outage probability constraint after the training period.
Whiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul Sung
NeurIPS2
2019 Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning
abstract
In this paper, we propose a new learning technique named message-dropout to improve the performance for multi-agent deep reinforcement learning under two application scenarios: 1) classical multi-agent reinforcement learning with direct message communication among agents and 2) centralized training with decentralized execution. In the first application scenario of multi-agent systems in which direct message communication among agents is allowed, the messagedropout technique drops out the received messages from other agents in a block-wise manner with a certain probability in the training phase and compensates for this effect by multiplying the weights of the dropped-out block units with a correction probability. The applied message-dropout technique effectively handles the increased input dimension in multi-agent reinforcement learning with communication and makes learning robust against communication errors in the execution phase. In the second application scenario of centralized training with decentralized execution, we particularly consider the application of the proposed messagedropout to Multi-Agent Deep Deterministic Policy Gradient (MADDPG), which uses a centralized critic to train a decentralized actor for each agent. We evaluate the proposed message-dropout technique for several games, and numerical results show that the proposed message-dropout technique with proper dropout rate improves the reinforcement learning performance significantly in terms of the training speed and the steady-state performance in the execution phase.
Woojun Kim, Myungsik Cho, Youngchul Sung
AAAI2