Kyunghwan Son

dblp:206/9135 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Computer networks · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 82% Planning, search and constraint satisfaction · 12% Efficient and distributed learning · 7%
Databases, data mining, and information retrieval
2 papers
Web and social media mining · 100%
Computer networks
1 paper
Network measurement and analytics · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
1.022022
Disentangling Sources of Risk for Distributional Multi-Agent Reinforcement Learning · ICML 2022
QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning · ICML 2019
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.022022
Disentangling Sources of Risk for Distributional Multi-Agent Reinforcement Learning · ICML 2022
Learning to Schedule Communication in Multi-agent Reinforcement Learning · ICLR (Poster) 2019
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.712023
Imitating Graph-Based Planning with Goal-Conditioned Policies · ICLR 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search-based planning
graph-based planning
0.712023
Imitating Graph-Based Planning with Goal-Conditioned Policies · ICLR 2023
Machine learning › Reinforcement learning
risk-sensitive policy
0.612022
Disentangling Sources of Risk for Distributional Multi-Agent Reinforcement Learning · ICML 2022
Network measurement and analytics › social network analysis
information diffusion
0.412020
Information Source Finding in Networks: Querying With Budgets · IEEE/ACM Trans. Netw. 2020
Network measurement and analytics › network diffusion
source detection
0.412020
Information Source Finding in Networks: Querying With Budgets · IEEE/ACM Trans. Netw. 2020
Machine learning › Efficient and distributed learning
communication scheduling
0.412019
Learning to Schedule Communication in Multi-agent Reinforcement Learning · ICLR (Poster) 2019
Machine learning › Reinforcement learning
deep reinforcement learning
0.412019
Solving Continual Combinatorial Selection via Deep Reinforcement Learning · IJCAI 2019
Machine learning › Reinforcement learning
large action space
0.412019
Solving Continual Combinatorial Selection via Deep Reinforcement Learning · IJCAI 2019
Machine learning › Reinforcement learning
markov decision process
0.412019
Solving Continual Combinatorial Selection via Deep Reinforcement Learning · IJCAI 2019
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
0.412019
QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning · ICML 2019
Web and social media mining
information diffusion
0.312017
Rumor source detection under querying with untruthful answers · INFOCOM 2017
Web and social media mining › information diffusion
rumor source detection
0.312017
Rumor source detection under querying with untruthful answers · INFOCOM 2017
Web and social media mining › social network analysis
social network
0.312017
Rumor source detection under querying with untruthful answers · INFOCOM 2017

Methods — techniques the papers use, named apart from their topics

information theory · 0.9adaptive querying · 0.9imitation learning · 0.7goal-conditioned policy · 0.7distributional reinforcement learning · 0.6weight-shared q-networks · 0.4value-based reinforcement learning · 0.4reinforcement learning · 0.4iterative select-MDP · 0.4centralized training with decentralized execution · 0.4simulation · 0.3maximum likelihood estimation · 0.3
YearPublicationVenuePosition
2023 Imitating Graph-Based Planning with Goal-Conditioned Policies
Younggyo Seo, Sungsoo Ahn, Kyunghwan Son, Jinwoo Shin
ICLR4
2022 Disentangling Sources of Risk for Distributional Multi-Agent Reinforcement Learning
abstract
In cooperative multi-agent reinforcement learning, the outcomes of agent-wise policies are highly stochastic due to the two sources of risk: (a) random actions taken by teammates and (b) random transition and rewards. Although the two sources have very distinct characteristics, existing frameworks are insufficient to control the risk-sensitivity of agent-wise policies in a disentangled manner. To this end, we propose Disentangled RIsk-sensitive Multi-Agent reinforcement learning (DRIMA) to separately access the risk sources. For example, our framework allows an agent to be optimistic with respect to teammates (who can prosocially adapt) but more risk-neutral with respect to the environment (which does not adapt). Our experiments demonstrate that DRIMA significantly outperforms prior state-of-the-art methods across various scenarios in the StarCraft Multi-agent Challenge environment. Notably, DRIMA shows robust performance where prior methods learn only a highly suboptimal policy, regardless of reward shaping, exploration scheduling, and noisy (random or adversarial) agents.
Kyunghwan Son, Sungsoo Ahn, Roben Delos Reyes, Yung Yi, Jinwoo Shin
ICML1
2021 Neuro-DCF: Design of Wireless MAC via Multi-Agent Reinforcement Learning Approach
abstract
The carrier sense multiple access (CSMA) algorithm has been used in the wireless medium access control (MAC) under standard 802.11 implementation due to its simplicity and generality. An extensive body of research on CSMA has long been made not only in the context of practical protocols, but also in a distributed way of optimal MAC scheduling. However, the current state-of-the-art CSMA (or its extensions) still suffers from poor performance, especially in multi-hop scenarios, and often requires patch-based solutions rather than a universal solution. In this paper, we propose an algorithm which adopts an experience-driven approach and train CSMA-based wireless MAC by using deep reinforcement learning. We name our protocol, Neuro-DCF. Two key challenges are: (i) a stable training method for distributed execution and (ii) a unified training method for embracing various interference patterns and configurations. For (i), we adopt a multi-agent reinforcement learning framework, and for (ii) we introduce a novel graph neural network (GNN) based training structure. We provide extensive simulation results which demonstrate that our protocol, Neuro-DCF, significantly outperforms 802.11 DCF and O-DCF, a recent theory-based MAC protocol, especially in terms of improving delay performance while preserving optimal utility. We believe our multi-agent reinforcement learning based approach would get broad interest from other learning-based network controllers in different layers that require distributed operation.
Sumyeong Ahn, Kyunghwan Son, Yung Yi
MobiHoc3
2020 Information Source Finding in Networks: Querying With Budgets
abstract
In this paper, we study a problem of detecting the source of diffused information by querying individuals, given a sample snapshot of the information diffusion graph, where two queries are asked: (i) whether the respondent is the source or not, and (ii) if not, which neighbor spreads the information to the respondent. We consider the case when respondents may not always be truthful and some cost is taken for each query. Our goal is to quantify the necessary and sufficient budgets to achieve the detection probability 1- δ for any given 0 <; δ <; 1. To this end, we study two types of algorithms: adaptive and non-adaptive ones, each of which corresponds to whether we adaptively select the next respondents based on the answers of the previous respondents or not. We first provide the information theoretic lower bounds for the necessary budgets in both algorithm types. In terms of the sufficient budgets, we propose two practical estimation algorithms, each of non-adaptive and adaptive types, and for each algorithm, we quantitatively analyze the budget which ensures 1- δ detection accuracy. This theoretical analysis not only quantifies the budgets needed by practical estimation algorithms achieving a given target detection accuracy in finding the diffusion source, but also enables us to quantitatively characterize the amount of extra budget required in non-adaptive type of estimation, referred to as adaptivity gap. We validate our theoretical findings over synthetic and real-world social network topologies.
Jaeyoung Choi 0001, Jiin Woo, Kyunghwan Son, Jinwoo Shin, Yung Yi
IEEE/ACM Trans. Netw.4
2019 Learning to Schedule Communication in Multi-agent Reinforcement Learning
Daewoo Kim, David Hostallero, Wan Ju Kang, Kyunghwan Son, Yung Yi
ICLR (Poster)6
2019 QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning
abstract
We explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently. However, VDN and QMIX are representative examples that use the idea of factorization of the joint action-value function into individual ones for decentralized execution. VDN and QMIX address only a fraction of factorizable MARL tasks due to their structural constraint in factorization such as additivity and monotonicity. In this paper, we propose a new factorization method for MARL, QTRAN, which is free from such structural constraints and takes on a new approach to transforming the original joint action-value function into an easily factorizable one, with the same optimal actions. QTRAN guarantees more general factorization than VDN or QMIX, thus covering a much wider class of MARL tasks than does previous methods. Our experiments for the tasks of multi-domain Gaussian-squeeze and modified predator-prey demonstrate QTRAN’s superior performance with especially larger margins in games whose payoffs penalize non-cooperative behavior more aggressively.
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Hostallero, Yung Yi
ICML1
2019 Solving Continual Combinatorial Selection via Deep Reinforcement Learning
abstract
We consider the Markov Decision Process (MDP) of selecting a subset of items at each step, termed the Select-MDP (S-MDP). The large state and action spaces of S-MDPs make them intractable to solve with typical reinforcement learning (RL) algorithms especially when the number of items is huge. In this paper, we present a deep RL algorithm to solve this issue by adopting the following key ideas. First, we convert the original S-MDP into an Iterative Select-MDP (IS-MDP), which is equivalent to the S-MDP in terms of optimal actions. IS-MDP decomposes a joint action of selecting K items simultaneously into K iterative selections resulting in the decrease of actions at the expense of an exponential increase of states. Second, we overcome this state space explosion by exploiting a special symmetry in IS-MDPs with novel weight shared Q-networks, which provably maintain sufficient expressive power. Various experiments demonstrate that our approach works well even when the item space is large and that it scales to environments with item spaces different from those used in training.
HyungSeok Song, Hyeryung Jang, Hai H. Tran, Se-eun Yoon, Kyunghwan Son, Donggyu Yun, Hyoju Chung, Yung Yi
IJCAI5
2017 Rumor source detection under querying with untruthful answers
abstract
Social networks are the major routes for most individuals to exchange their opinions about new products, social trends and political issues via their interactions. It is often of significant importance to figure out who initially diffuses the information, i.e., finding a rumor source or a trend setter. It is known that such a task is highly challenging and the source detection probability cannot be beyond 31% for regular trees, if we just estimate the source from a given diffusion snapshot. In practice, finding the source often entails the process of querying that asks “Are you the rumor source?” or “Who tells you the rumor?” that would increase the chance of detecting the source. In this paper, we consider two kinds of querying: (a) simple batch querying and (b) interactive querying with direction under the assumption that queriees can be untruthful with some probability. We propose estimation algorithms for those queries, and quantify their detection performance and the amount of extra budget due to untruthfulness, analytically showing that querying significantly improves the detection performance. We perform extensive simulations to validate our theoretical findings over synthetic and real-world social network topologies.
Jaeyoung Choi 0001, Jiin Woo, Kyunghwan Son, Jinwoo Shin, Yung Yi
INFOCOM4