Heewoong Choi

dblp:373/2138 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction
0.912025
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › offline reinforcement learning
offline preference-based reinforcement learning
0.812024
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning
0.812024
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
0.812024
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 0.9advantage estimation · 0.9ternary feedback · 0.8ranked list of trajectories · 0.8
YearPublicationVenuePosition
2025 Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning
abstract
Offline goal-conditioned reinforcement learning (GCRL) offers a practical learning paradigm in which goal-reaching policies are trained from abundant state–action trajectory datasets without additional environment interaction. However, offline GCRL still struggles with long-horizon tasks, even with recent advances that employ hierarchical policy structures, such as HIQL. Identifying the root cause of this challenge, we observe the following insight. Firstly, performance bottlenecks mainly stem from the high-level policy’s inability to generate appropriate subgoals. Secondly, when learning the high-level policy in the long-horizon regime, the sign of the advantage estimate frequently becomes incorrect. Thus, we argue that improving the value function to produce a clear advantage estimate for learning the high-level policy is essential. In this paper, we propose a simple yet effective solution: _**Option-aware Temporally Abstracted**_ value learning, dubbed **OTA**, which incorporates temporal abstraction into the temporal-difference learning process. By modifying the value update to be _option-aware_, our approach contracts the effective horizon length, enabling better advantage estimates even in long-horizon regimes. We experimentally show that the high-level policy learned using the OTA value function achieves strong performance on complex tasks from OGBench, a recently proposed offline GCRL benchmark, including maze navigation and visual robotic manipulation environments. Our code is available at https://github.com/ota-v/ota-v
Hongjoon Ahn, Heewoong Choi, Jisu Han, Taesup Moon
NeurIPS2
2024 Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
abstract
In Reinforcement Learning (RL), designing precise reward functions remains to be a challenge, particularly when aligning with human intent. Preference-based RL (PbRL) was introduced to address this problem by learning reward models from human feedback. However, existing PbRL methods have limitations as they often overlook the second-order preference that indicates the relative strength of preference. In this paper, we propose Listwise Reward Estimation (LiRE), a novel approach for offline PbRL that leverages second-order preference information by constructing a Ranked List of Trajectories (RLT), which can be efficiently built by using the same ternary feedback type as traditional methods. To validate the effectiveness of LiRE, we propose a new offline PbRL dataset that objectively reflects the effect of the estimated rewards. Our extensive experiments on the dataset demonstrate the superiority of LiRE, i.e., outperforming state-of-the-art baselines even with modest feedback budgets and enjoying robustness with respect to the number of feedbacks and feedback noise. Our code is available at https://github.com/chwoong/LiRE
Heewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup Moon
ICML1
2024 NCIS: Neural Contextual Iterative Smoothing for Purifying Adversarial Perturbations
abstract
We propose a novel and effective purification-based adversarial defense method against pre-processor blind white-and black-box attacks, without requiring any adversarial training or retraining of the classification model. Based on the observation of the adversarial noise, we propose a simple iterative Gaussian Smoothing (GS) that smoothes out adversarial noise and achieves substantially high robust accuracy. To further improve the method, we propose Neural Contextual Iterative Smoothing (NCIS), which trains a blind-spot network (BSN) in a self-supervised manner to reconstruct the discriminative features of the smoothed original image. From the extensive experiments on the large-scale ImageNet, we show that our method achieves both competitive standard accuracy and state-of-the-art robust accuracy against most strong purifier-blind white- and black-box attacks. Also, we propose a new evaluation benchmark based on commercial image classification APIs, including AWS, Azure, Clarifai, and Google, and demonstrate that users can use our method to increase the adversarial robustness of APIs.
Sungmin Cha, Naeun Ko, Heewoong Choi, Young Joon Yoo, Taesup Moon
WACV3