EDBT 2026 Demo / reviewers in the wild / expert
Heewoong Choi
dblp:373/2138
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning |
0.9 | 1 | 2025 | Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.9 | 1 | 2025 | Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.9 | 1 | 2025 | Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction |
0.9 | 1 | 2025 | Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning › offline reinforcement learning
offline preference-based reinforcement learning |
0.8 | 1 | 2024 | Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.8 | 1 | 2024 | Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.8 | 1 | 2024 | Listwise Reward Estimation for Offline Preference-based Reinforcement Learning · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
temporal difference learning · 0.9advantage estimation · 0.9ternary feedback · 0.8ranked list of trajectories · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningabstractOffline goal-conditioned reinforcement learning (GCRL) offers a practical learning paradigm in which goal-reaching policies are trained from abundant state–action trajectory datasets without additional environment interaction. However, offline GCRL still struggles with long-horizon tasks, even with recent advances that employ hierarchical policy structures, such as HIQL. Identifying the root cause of this challenge, we observe the following insight. Firstly, performance bottlenecks mainly stem from the high-level policy’s inability to generate appropriate subgoals. Secondly, when learning the high-level policy in the long-horizon regime, the sign of the advantage estimate frequently becomes incorrect. Thus, we argue that improving the value function to produce a clear advantage estimate for learning the high-level policy is essential. In this paper, we propose a simple yet effective solution: _**Option-aware Temporally Abstracted**_ value learning, dubbed **OTA**, which incorporates temporal abstraction into the temporal-difference learning process. By modifying the value update to be _option-aware_, our approach contracts the effective horizon length, enabling better advantage estimates even in long-horizon regimes. We experimentally show that the high-level policy learned using the OTA value function achieves strong performance on complex tasks from OGBench, a recently proposed offline GCRL benchmark, including maze navigation and visual robotic manipulation environments. Our code is available at https://github.com/ota-v/ota-v Hongjoon Ahn, Heewoong Choi, Jisu Han, Taesup Moon |
NeurIPS | 2 |
| 2024 | Listwise Reward Estimation for Offline Preference-based Reinforcement LearningabstractIn Reinforcement Learning (RL), designing precise reward functions remains to be a challenge, particularly when aligning with human intent. Preference-based RL (PbRL) was introduced to address this problem by learning reward models from human feedback. However, existing PbRL methods have limitations as they often overlook the second-order preference that indicates the relative strength of preference. In this paper, we propose Listwise Reward Estimation (LiRE), a novel approach for offline PbRL that leverages second-order preference information by constructing a Ranked List of Trajectories (RLT), which can be efficiently built by using the same ternary feedback type as traditional methods. To validate the effectiveness of LiRE, we propose a new offline PbRL dataset that objectively reflects the effect of the estimated rewards. Our extensive experiments on the dataset demonstrate the superiority of LiRE, i.e., outperforming state-of-the-art baselines even with modest feedback budgets and enjoying robustness with respect to the number of feedbacks and feedback noise. Our code is available at https://github.com/chwoong/LiRE Heewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup Moon |
ICML | 1 |
| 2024 | NCIS: Neural Contextual Iterative Smoothing for Purifying Adversarial PerturbationsabstractWe propose a novel and effective purification-based adversarial defense method against pre-processor blind white-and black-box attacks, without requiring any adversarial training or retraining of the classification model. Based on the observation of the adversarial noise, we propose a simple iterative Gaussian Smoothing (GS) that smoothes out adversarial noise and achieves substantially high robust accuracy. To further improve the method, we propose Neural Contextual Iterative Smoothing (NCIS), which trains a blind-spot network (BSN) in a self-supervised manner to reconstruct the discriminative features of the smoothed original image. From the extensive experiments on the large-scale ImageNet, we show that our method achieves both competitive standard accuracy and state-of-the-art robust accuracy against most strong purifier-blind white- and black-box attacks. Also, we propose a new evaluation benchmark based on commercial image classification APIs, including AWS, Azure, Clarifai, and Google, and demonstrate that users can use our method to increase the adversarial robustness of APIs. Sungmin Cha, Naeun Ko, Heewoong Choi, Young Joon Yoo, Taesup Moon |
WACV | 3 |