Hongjoon Ahn

dblp:236/5812 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 48% Graph learning · 17% Learning paradigms · 13%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.022025
Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple Baseline · ICLR 2025
SS-IL: Separated Softmax for Incremental Learning · ICCV 2021
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning
0.912025
Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple Baseline · ICLR 2025
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
negative transfer
0.912025
Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple Baseline · ICLR 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Learning paradigms
continual learning
0.822020
Continual Learning with Node-Importance based Adaptive Group Sparse Regularization · NeurIPS 2020
Uncertainty-based Continual Learning with Adaptive Regularization · NeurIPS 2019
Machine learning › Reinforcement learning › offline reinforcement learning
offline preference-based reinforcement learning
0.812024
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning
0.812024
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
0.812024
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024
Machine learning › Graph learning
graph neural network
0.612022
Descent Steps of a Relation-Aware Energy Produce Heterogeneous Graph Neural Networks · NeurIPS 2022
Machine learning › Graph learning › graph neural network
heterogeneous graph neural network
0.612022
Descent Steps of a Relation-Aware Energy Produce Heterogeneous Graph Neural Networks · NeurIPS 2022
Machine learning › Graph learning › graph neural network
node classification
0.612022
Descent Steps of a Relation-Aware Energy Produce Heterogeneous Graph Neural Networks · NeurIPS 2022
Machine learning › Graph learning › graph neural network › deep graph neural network
over-smoothing
0.612022
Descent Steps of a Relation-Aware Energy Produce Heterogeneous Graph Neural Networks · NeurIPS 2022
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.512021
SS-IL: Separated Softmax for Incremental Learning · ICCV 2021
Machine learning › Learning paradigms › continual learning
class-incremental learning
0.512021
SS-IL: Separated Softmax for Incremental Learning · ICCV 2021
Machine learning › Optimization for machine learning › sparse learning
group-sparse regularization
0.412020
Continual Learning with Node-Importance based Adaptive Group Sparse Regularization · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › recursive bayesian estimation
bayesian online learning
0.412019
Uncertainty-based Continual Learning with Adaptive Regularization · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.412019
Uncertainty-based Continual Learning with Adaptive Regularization · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 0.9network resetting · 0.9knowledge distillation · 0.9advantage estimation · 0.9ternary feedback · 0.8ranked list of trajectories · 0.8energy function descent · 0.6bi-level optimization · 0.6separated softmax · 0.5exemplar memory · 0.5
YearPublicationVenuePosition
2025 Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple Baseline
abstract
We argue that the negative transfer problem occurring when the new task to learn arrives is an important problem that needs not be overlooked when developing effective Continual Reinforcement Learning (CRL) algorithms. Through comprehensive experimental validation, we demonstrate that such issue frequently exists in CRL and cannot be effectively addressed by several recent work on either mitigating plasticity loss of RL agents or enhancing the positive transfer in CRL scenario. To that end, we develop Reset & Distill (R&D), a simple yet highly effective baseline method, to overcome the negative transfer problem in CRL. R&D combines a strategy of resetting the agent's online actor and critic networks to learn a new task and an offline learning step for distilling the knowledge from the online actor and previous expert's action probabilities. We carried out extensive experiments on long sequence of Meta World tasks and show that our simple baseline method consistently outperforms recent approaches, achieving significantly higher success rates across a range of tasks. Our findings highlight the importance of considering negative transfer in CRL and emphasize the need for robust strategies like R&D to mitigate its detrimental effects.
Hongjoon Ahn, Jinu Hyeon, Bosun Hwang, Taesup Moon
ICLR1
2025 Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning
abstract
Offline goal-conditioned reinforcement learning (GCRL) offers a practical learning paradigm in which goal-reaching policies are trained from abundant state–action trajectory datasets without additional environment interaction. However, offline GCRL still struggles with long-horizon tasks, even with recent advances that employ hierarchical policy structures, such as HIQL. Identifying the root cause of this challenge, we observe the following insight. Firstly, performance bottlenecks mainly stem from the high-level policy’s inability to generate appropriate subgoals. Secondly, when learning the high-level policy in the long-horizon regime, the sign of the advantage estimate frequently becomes incorrect. Thus, we argue that improving the value function to produce a clear advantage estimate for learning the high-level policy is essential. In this paper, we propose a simple yet effective solution: _**Option-aware Temporally Abstracted**_ value learning, dubbed **OTA**, which incorporates temporal abstraction into the temporal-difference learning process. By modifying the value update to be _option-aware_, our approach contracts the effective horizon length, enabling better advantage estimates even in long-horizon regimes. We experimentally show that the high-level policy learned using the OTA value function achieves strong performance on complex tasks from OGBench, a recently proposed offline GCRL benchmark, including maze navigation and visual robotic manipulation environments. Our code is available at https://github.com/ota-v/ota-v
Hongjoon Ahn, Heewoong Choi, Jisu Han, Taesup Moon
NeurIPS1
2024 Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
abstract
In Reinforcement Learning (RL), designing precise reward functions remains to be a challenge, particularly when aligning with human intent. Preference-based RL (PbRL) was introduced to address this problem by learning reward models from human feedback. However, existing PbRL methods have limitations as they often overlook the second-order preference that indicates the relative strength of preference. In this paper, we propose Listwise Reward Estimation (LiRE), a novel approach for offline PbRL that leverages second-order preference information by constructing a Ranked List of Trajectories (RLT), which can be efficiently built by using the same ternary feedback type as traditional methods. To validate the effectiveness of LiRE, we propose a new offline PbRL dataset that objectively reflects the effect of the estimated rewards. Our extensive experiments on the dataset demonstrate the superiority of LiRE, i.e., outperforming state-of-the-art baselines even with modest feedback budgets and enjoying robustness with respect to the number of feedbacks and feedback noise. Our code is available at https://github.com/chwoong/LiRE
Heewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup Moon
ICML3
2022 Descent Steps of a Relation-Aware Energy Produce Heterogeneous Graph Neural Networks
abstract
Heterogeneous graph neural networks (GNNs) achieve strong performance on node classification tasks in a semi-supervised learning setting. However, as in the simpler homogeneous GNN case, message-passing-based heterogeneous GNNs may struggle to balance between resisting the oversmoothing that may occur in deep models, and capturing long-range dependencies of graph structured data. Moreover, the complexity of this trade-off is compounded in the heterogeneous graph case due to the disparate heterophily relationships between nodes of different types. To address these issues, we propose a novel heterogeneous GNN architecture in which layers are derived from optimization steps that descend a novel relation-aware energy function. The corresponding minimizer is fully differentiable with respect to the energy function parameters, such that bilevel optimization can be applied to effectively learn a functional form whose minimum provides optimal node representations for subsequent classification tasks. In particular, this methodology allows us to model diverse heterophily relationships between different node types while avoiding oversmoothing effects. Experimental results on 8 heterogeneous graph benchmarks demonstrates that our proposed method can achieve competitive node classification accuracy.
Hongjoon Ahn, Yongyi Yang, Taesup Moon, David P. Wipf
NeurIPS1
2021 SS-IL: Separated Softmax for Incremental Learning
abstract
We consider class incremental learning (CIL) problem, in which a learning agent continuously learns new classes from incrementally arriving training data batches and aims to predict well on all the classes learned so far. The main challenge of the problem is the catastrophic forgetting, and for the exemplar-memory based CIL methods, it is generally known that the forgetting is commonly caused by the classification score bias that is injected due to the data imbalance between the new classes and the old classes (in the exemplar-memory). While several methods have been proposed to correct such score bias by some additional post-processing, e.g., score re-scaling or balanced fine-tuning, no systematic analysis on the root cause of such bias has been done. To that end, we analyze that computing the softmax probabilities by combining the output scores for all old and new classes could be the main cause of the bias. Then, we propose a new method, dubbed as Separated Softmax for Incremental Learning (SS-IL), that consists of separated softmax (SS) output layer combined with task-wise knowledge distillation (TKD) to resolve such bias. Throughout our extensive experimental results on several large-scale CIL benchmark datasets, we show our SS-IL achieves strong state-of-the-art accuracy through attaining much more balanced prediction scores across old and new classes, without any additional post-processing.
Hongjoon Ahn, Jihwan Kwak, Subin Lim, Hyeonsu Bang, Hyojun Kim, Taesup Moon
ICCV1
2020 Continual Learning with Node-Importance based Adaptive Group Sparse Regularization
abstract
We propose a novel regularization-based continual learning method, dubbed as Adaptive Group Sparsity based Continual Learning (AGS-CL), using two group sparsity-based penalties. Our method selectively employs the two penalties when learning each neural network node based on its the importance, which is adaptively updated after learning each task. By utilizing the proximal gradient descent method, the exact sparsity and freezing of the model is guaranteed during the learning process, and thus, the learner explicitly controls the model capacity. Furthermore, as a critical detail, we re-initialize the weights associated with unimportant nodes after learning each task in order to facilitate efficient learning and prevent the negative transfer. Throughout the extensive experimental results, we show that our AGS-CL uses orders of magnitude less memory space for storing the regularization parameters, and it significantly outperforms several state-of-the-art baselines on representative benchmarks for both supervised and reinforcement learning.
Sangwon Jung, Hongjoon Ahn, Sungmin Cha, Taesup Moon
NeurIPS2
2020 Iterative Channel Estimation for Discrete Denoising under Channel Uncertainty
abstract
We propose a novel iterative channel estimation (ICE) algorithm that essentially removes the critical known noisy channel assumption for universal discrete denoising problem. Our algorithm is based on Neural DUDE (N-DUDE), a recently proposed neural network-based discrete denoiser, and it estimates the channel transition matrix as well as the neural network parameters in an alternating manner until convergence. While we do not make any probabilistic assumption on the underlying clean data, our ICE resembles Expectation-Maximization (EM) with variational approximation, and it takes advantage of the property that N-DUDE can always induce a marginal posterior distribution of the clean data. We carefully validate the channel estimation quality of ICE, and with extensive experiments on several radically different types of data, we show the ICE equipped neural network-based denoisers can perform \emph{universally} well regardless of the uncertainties in both the channel and the clean source. Moreover, we show ICE becomes extremely robust to its hyperparameters, and show the denoisers with ICE significantly outperform the strong baseline that can handle the channel uncertainties for denoising, the widely used Baum-Welch (BW) algorithm for hidden Markov models (HMM).
Hongjoon Ahn, Taesup Moon
UAI1
2019 Uncertainty-based Continual Learning with Adaptive Regularization
abstract
We introduce a new neural network-based continual learning algorithm, dubbed as Uncertainty-regularized Continual Learning (UCL), which builds on traditional Bayesian online learning framework with variational inference. We focus on two significant drawbacks of the recently proposed regularization-based methods: a) considerable additional memory cost for determining the per-weight regularization strengths and b) the absence of gracefully forgetting scheme, which can prevent performance degradation in learning new tasks. In this paper, we show UCL can solve these two problems by introducing a fresh interpretation on the Kullback-Leibler (KL) divergence term of the variational lower bound for Gaussian mean-field approximation. Based on the interpretation, we propose the notion of node-wise uncertainty, which drastically reduces the number of additional parameters for implementing per-weight regularization. Moreover, we devise two additional regularization terms that enforce \emph{stability} by freezing important parameters for past tasks and allow \emph{plasticity} by controlling the actively learning parameters for a new task. Through extensive experiments, we show UCL convincingly outperforms most of recent state-of-the-art baselines not only on popular supervised learning benchmarks, but also on challenging lifelong reinforcement learning tasks. The source code of our algorithm is available at https://github.com/csm9493/UCL.
Hongjoon Ahn, Sungmin Cha, Donggyu Lee, Taesup Moon
NeurIPS1