EDBT 2026 Demo / reviewers in the wild / expert
Kee-Eung Kim
dblp:35/6703
· DBLP profile ↗
88ranked-venue papers
8as first author
32since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 8 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing Preferential Rate in Retail Lending with Causal Inference and Domain Adaptation
Jimyung Choi, Hyeryeong Oh, Sumin Shin, Kee-Eung Kim |
AAAI | 7 |
| 2025 | Monet: Mixture of Monosemantic Experts for TransformersabstractUnderstanding the internal computations of large language models (LLMs) is crucial for aligning them with human values and preventing undesirable behaviors like toxic content generation. However, mechanistic interpretability is hindered by *polysemanticity*—where individual neurons respond to multiple, unrelated concepts. While Sparse Autoencoders (SAEs) have attempted to disentangle these features through sparse dictionary learning, they have compromised LLM performance due to reliance on post-hoc reconstruction loss. To address this issue, we introduce **Mixture of Monosemantic Experts for Transformers (Monet)** architecture, which incorporates sparse dictionary learning directly into end-to-end Mixture-of-Experts pretraining. Our novel expert decomposition method enables scaling the expert count to 262,144 per layer while total parameters scale proportionally to the square root of the number of experts. Our analyses demonstrate mutual exclusivity of knowledge across experts and showcase the parametric knowledge encapsulated within individual experts. Moreover, **Monet** allows knowledge manipulation over domains, languages, and toxicity mitigation without degrading general performance. Our pursuit of transparent LLMs highlights the potential of scaling expert counts to enhance mechanistic interpretability and directly resect the internal knowledge to fundamentally adjust model behavior. Jungwoo Park, Ahn Young Jin, Kee-Eung Kim, Jaewoo Kang |
ICLR | 3 |
| 2025 | Goal-Conditioned DPO: Prioritizing Safety in Misaligned InstructionsabstractJoo Bon Maeng, Seongmin Lee, Seokin Seo, Kee-Eung Kim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Joo Bon Maeng, Seongmin Lee 0011, Seokin Seo, Kee-Eung Kim |
NAACL (Long Papers) | 4 |
| 2025 | DPAIL: Training Diffusion Policy for Adversarial Imitation Learning without Policy OptimizationabstractHuman experts employ diverse strategies to complete a task, producing to multi-modal demonstration data. Although traditional Adversarial Imitation Learning (AIL) methods have achieved notable success, they often collapse theses multi-modal behaviors into a single strategy, failing to replicate expert behaviors. To overcome this limitation, we propose DPAIL, an adversarial IL framework that leverages diffusion models as a policy class to enhance expressiveness. Building on the Adversarial Soft Advantage Fitting (ASAF) framework, which removes the need for policy optimization steps, DPAIL trains a diffusion policy using a binary cross-entropy objective to distinguish expert trajectories from generated ones. To enable optimization of the diffusion policy, we introduce a novel, tractable lower bound on the policy's likelihood. Through comprehensive quantitative and qualitative evaluations against various baselines, we demonstrate that our method not only captures diverse behaviors but also remains robust as the number of behavior modes increases. Yunseon Choi, Minchan Jeong, Soobin Um, Kee-Eung Kim |
NeurIPS | 4 |
| 2024 | Stitching Sub-trajectories with Conditional Diffusion Model for Goal-Conditioned Offline RLabstractOffline Goal-Conditioned Reinforcement Learning (Offline GCRL) is an important problem in RL that focuses on acquiring diverse goal-oriented skills solely from pre-collected behavior datasets. In this setting, the reward feedback is typically absent except when the goal is achieved, which makes it difficult to learn policies especially from a finite dataset of suboptimal behaviors. In addition, realistic scenarios involve long-horizon planning, which necessitates the extraction of useful skills within sub-trajectories. Recently, the conditional diffusion model has been shown to be a promising approach to generate high-quality long-horizon plans for RL. However, their practicality for the goal-conditioned setting is still limited due to a number of technical assumptions made by the methods. In this paper, we propose SSD (Sub-trajectory Stitching with Diffusion), a model-based offline GCRL method that leverages the conditional diffusion model to address these limitations. In summary, we use the diffusion model that generates future plans conditioned on the target goal and value, with the target value estimated from the goal-relabeled offline dataset. We report state-of-the-art performance in the standard benchmark set of GCRL tasks, and demonstrate the capability to successfully stitch the segments of suboptimal trajectories in the offline data to generate high-quality plans. Sungyoon Kim, Yunseon Choi, Daiki E. Matsunaga, Kee-Eung Kim |
AAAI | 4 |
| 2024 | A Submodular Optimization Approach to Accountable Loan ApprovalabstractIn the field of finance, the underwriting process is an essential step in evaluating every loan application. During this stage, the borrowers' creditworthiness and ability to repay the loan are assessed to ultimately decide whether to approve the loan application. One of the core components of underwriting is credit scoring, in which the probability of default is estimated. As such, there has been significant progress in enhancing the predictive accuracy of credit scoring models through the use of machine learning, but there still exists a need to ultimately construct an approval rule that takes into consideration additional criteria beyond the score itself. This construction process is traditionally done manually to ensure that the approval rule remains interpretable to humans. In this paper, we outline an automated system for optimizing a rule-based system for approving loan applications, which has been deployed at Hyundai Capital Services (HCS). The main challenge lay in creating a high-quality rule base that is simultaneously simple enough to be interpretable by risk analysts as well as customers, since the approval decision should be accountable. We addressed this challenge through principled submodular optimization. The deployment of our system has led to a 14% annual growth in the volume of loan services at HCS, while maintaining the target bad rate, and has resulted in the approval of customers who might have otherwise been rejected. Kyungsik Lee, Hana Yoo, Sumin Shin, Yeonung Baek, Hyunjin Kang, Kee-Eung Kim |
AAAI | 8 |
| 2024 | Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RLabstractYunseon Choi, Sangmin Bae, Seonghyun Ban, Minchan Jeong, Chuheng Zhang, Lei Song, Li Zhao, Jiang Bian, Kee-Eung Kim. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yunseon Choi, Sangmin Bae, Seonghyun Ban, Minchan Jeong, Chuheng Zhang, Lei Song 0001, Li Zhao 0007, Jiang Bian 0002, Kee-Eung Kim |
ACL (1) | 9 |
| 2024 | GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNetsabstractA critical component of the current generation of language models is preference alignment, which aims to precisely control the model's behavior to meet human needs and values.The most notable among such methods is Reinforcement Learning with Human Feedback (RLHF) and its offline variant Direct Preference Optimization (DPO), both of which seek to maximize a reward model based on human preferences.In particular, DPO derives reward signals directly from the offline preference data, but in doing so overfits the reward signals and generates suboptimal responses that may contain human biases in the dataset.In this work, we propose a practical application of a diversity-seeking RL algorithm called GFlowNet-DPO (GDPO) in an offline preference alignment setting to curtail such challenges.Empirical results show GDPO can generate far more diverse responses than the baseline methods that are still relatively aligned with human values in dialog generation and summarization tasks. Oh Joon Kwon, Daiki E. Matsunaga, Kee-Eung Kim |
EMNLP | 3 |
| 2024 | Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL PoliciesabstractWe consider off-policy evaluation (OPE) of deterministic target policies for reinforcement learning (RL) in environments with continuous action spaces. While it is common to use importance sampling for OPE, it suffers from high variance when the behavior policy deviates significantly from the target policy. In order to address this issue, some recent works on OPE proposed in-sample learning with importance resampling. Yet, these approaches are not applicable to deterministic target policies for continuous action spaces. To address this limitation, we propose to relax the deterministic target policy using a kernel and learn the kernel metrics that minimize the overall mean squared error of the estimated temporal difference update vector of an action value function, where the action value function is used for policy evaluation. We derive the bias and variance of the estimation error due to this relaxation and provide analytic solutions for the optimal kernel metric. In empirical studies using various test domains, we show that the OPE with in-sample learning using the kernel with optimized metric achieves significantly improved accuracy than other baselines. Haanvid Lee, Tri Wahyu Guntara, Jongmin Lee 0004, Yung-Kyun Noh, Kee-Eung Kim |
ICLR | 5 |
| 2024 | Diversification of Adaptive Policy for Effective Offline Reinforcement Learning
Yunseon Choi, Li Zhao 0007, Chuheng Zhang, Lei Song 0001, Jiang Bian 0002, Kee-Eung Kim |
IJCAI | 6 |
| 2024 | SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token SynchronizationabstractVisual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip gestures that represent different phonemes. Prior approaches have sought to distinguish fine-grained visemes by aligning visual and auditory semantics, but often fell short of full synchronization. To address this, we present SyncVSR, an end-to-end learning framework that leverages quantized audio for frame-level crossmodal supervision. By integrating a projection layer that synchronizes visual representation with acoustic data, our encoder learns to generate discrete audio tokens from a video sequence in a non-autoregressive manner. SyncVSR shows versatility across tasks, languages, and modalities at the cost of a forward pass. Our empirical evaluations show that it not only achieves state-of-the-art results but also reduces data usage by up to ninefold. Youngjin Ahn, Jungwoo Park, Sangha Park, Kee-Eung Kim |
INTERSPEECH | 5 |
| 2024 | Data Augmentation with Diffusion for Open-Set Semi-Supervised LearningabstractSemi-supervised learning (SSL) seeks to utilize unlabeled data to overcome the limited amount of labeled data and improve model performance. However, many SSL methods typically struggle in real-world scenarios, particularly when there is a large number of irrelevant instances in the unlabeled data that do not belong to any class in the labeled data. Previous approaches often downweight instances from irrelevant classes to mitigate the negative impact of class distribution mismatch on model training. However, by discarding irrelevant instances, they may result in the loss of valuable information such as invariance, regularity, and diversity within the data. In this paper, we propose a data-centric generative augmentation approach that leverages a diffusion model to enrich labeled data using both labeled and unlabeled samples. A key challenge is extracting the diversity inherent in the unlabeled data while mitigating the generation of samples irrelevant to the labeled data. To tackle this issue, we combine diffusion model training with a discriminator that identifies and reduces the impact of irrelevant instances. We also demonstrate that such a trained diffusion model can even convert an irrelevant instance into a relevant one, yielding highly effective synthetic data for training. Through a comprehensive suite of experiments, we show that our data augmentation approach significantly enhances the performance of SSL methods, especially in the presence of class distribution mismatch. Seonghyun Ban, Heesan Kong, Kee-Eung Kim |
NeurIPS | 3 |
| 2024 | Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionabstractWe consider offline imitation learning (IL), which aims to train an agent to imitate from the dataset of expert demonstrations without online interaction with the environment. Behavioral Cloning (BC) has been a simple yet effective approach to offline IL, but it is also well-known to be vulnerable to the covariate shift resulting from the mismatch between the state distributions induced by the learned policy and the expert policy. Moreover, as often occurs in practice, when expert datasets are collected from an arbitrary state distribution instead of a stationary one, these shifts become more pronounced, potentially leading to substantial failures in existing IL methods. Specifically, we focus on covariate shift resulting from arbitrary state data distributions, such as biased data collection or incomplete trajectories, rather than shifts induced by changes in dynamics or noisy expert actions. In this paper, to mitigate the effect of the covariate shifts in BC, we propose DrilDICE, which utilizes a distributionally robust BC objective by employing a stationary distribution correction ratio estimation (DICE) to derive a feasible solution. We evaluate the effectiveness of our method through an extensive set of experiments covering diverse covariate shift scenarios. The results demonstrate the efficacy of the proposed approach in improving the robustness against the shifts, outperforming existing offline IL methods in such scenarios. Seokin Seo, Byung-Jun Lee 0001, Jongmin Lee 0004, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim |
NeurIPS | 6 |
| 2023 | Trustworthy Residual Vehicle Value Prediction for Auto FinanceabstractThe residual value (RV) of a vehicle refers to its estimated worth at some point in the future. It is a core component in every auto financial product, used to determine the credit lines and the leasing rates. As such, an accurate prediction of RV is critical for the auto finance industry, since it can pose a risk of revenue loss by over-prediction or make the financial product incompetent by under-prediction. Although there are a number of prior studies on training machine learning models on a large amount of used car sales data, we had to cope with real-world operational requirements such as compliance with regulations (i.e. monotonicity of output with respect to a subset of features) and generalization to unseen input (i.e. new and rare car models). In this paper, we describe how we coped with these practical challenges and created value for our business at Hyundai Capital Services, the top auto financial service provider in Korea. Mihye Kim, Jimyung Choi, Yeonung Baek, Gisuk Bang, Kwangwoon Son, Yeonman Ryou, Kee-Eung Kim |
AAAI | 9 |
| 2023 | Information-Theoretic State Space Model for Multi-View Reinforcement LearningabstractMulti-View Reinforcement Learning (MVRL) seeks to find an optimal control for an agent given multi-view observations from various sources. Despite recent advances in multi-view learning that aim to extract the latent representation from multi-view data, it is not straightforward to apply them to control tasks, especially when the observations are temporally dependent on one another. The problem can be even more challenging if the observations are intermittently missing for a subset of views. In this paper, we introduce Fuse2Control (F2C), an information-theoretic approach to capturing the underlying state space model from the sequences of multi-view observations. We conduct an extensive set of experiments in various control tasks showing that our method is highly effective in aggregating task-relevant information across many views, that scales linearly with the number of views while retaining robustness to arbitrary missing view scenarios. HyeongJoo Hwang, Seokin Seo, Youngsoo Jang, Sungyoon Kim, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung Kim |
ICML | 7 |
| 2023 | AlberDICE: Addressing Out-Of-Distribution Joint Actions in Offline Multi-Agent RL via Alternating Stationary Distribution Correction EstimationabstractOne of the main challenges in offline Reinforcement Learning (RL) is the distribution shift that arises from the learned policy deviating from the data collection policy. This is often addressed by avoiding out-of-distribution (OOD) actions during policy improvement as their presence can lead to substantial performance degradation. This challenge is amplified in the offline Multi-Agent RL (MARL) setting since the joint action space grows exponentially with the number of agents.
To avoid this curse of dimensionality, existing MARL methods adopt either value decomposition methods or fully decentralized training of individual agents. However, even when combined with standard conservatism principles, these methods can still result in the selection of OOD joint actions in offline MARL. To this end, we introduce AlberDICE,
an offline MARL algorithm that alternatively performs centralized training of individual agents based on stationary distribution optimization. AlberDICE circumvents the exponential complexity of MARL by computing the best response of one agent at a time while effectively avoiding OOD joint action selection. Theoretically, we show that the alternating optimization procedure converges to Nash policies. In the experiments, we demonstrate that AlberDICE significantly outperforms baseline algorithms on a standard suite of MARL benchmarks. Daiki E. Matsunaga, Jongmin Lee 0004, Jaeseok Yoon, Stefanos Leonardos, Pieter Abbeel, Kee-Eung Kim |
NeurIPS | 6 |
| 2023 | Regularized Behavior Cloning for Blocking the Leakage of Past Action InformationabstractFor partially observable environments, imitation learning with observation histories (ILOH) assumes that control-relevant information is sufficiently captured in the observation histories for imitating the expert actions. In the offline setting wherethe agent is required to learn to imitate without interaction with the environment, behavior cloning (BC) has been shown to be a simple yet effective method for imitation learning. However, when the information about the actions executed in the past timesteps leaks into the observation histories, ILOH via BC often ends up imitating its own past actions. In this paper, we address this catastrophic failure by proposing a principled regularization for BC, which we name Past Action Leakage Regularization (PALR). The main idea behind our approach is to leverage the classical notion of conditional independence to mitigate the leakage. We compare different instances of our framework with natural choices of conditional independence metric and its estimator. The result of our comparison advocates the use of a particular kernel-based estimator for the conditional independence metric. We conduct an extensive set of experiments on benchmark datasets in order to assess the effectiveness of our regularization method. The experimental results show that our method significantly outperforms prior related approaches, highlighting its potential to successfully imitate expert actions when the past action information leaks into the observation histories. Seokin Seo, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim |
NeurIPS | 4 |
| 2022 | COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
Jongmin Lee 0004, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess, Doina Precup, Kee-Eung Kim, Arthur Guez |
ICLR | 6 |
| 2022 | Structure-Aware Transformer Policy for Inhomogeneous Multi-Task Reinforcement Learning
Sunghoon Hong, Deunsol Yoon, Kee-Eung Kim |
ICLR | 3 |
| 2022 | GPT-Critic: Offline Reinforcement Learning for End-to-End Task-Oriented Dialogue Systems
Youngsoo Jang, Jongmin Lee 0004, Kee-Eung Kim |
ICLR | 3 |
| 2022 | DemoDICE: Offline Imitation Learning with Supplementary Imperfect Demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee 0004, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim |
ICLR | 7 |
| 2022 | PAC-Net: A Model Pruning Approach to Inductive Transfer LearningabstractInductive transfer learning aims to learn from a small amount of training data for the target task by utilizing a pre-trained model from the source task. Most strategies that involve large-scale deep learning models adopt initialization with the pre-trained model and fine-tuning for the target task. However, when using over-parameterized models, we can often prune the model without sacrificing the accuracy of the source task. This motivates us to adopt model pruning for transfer learning with deep learning models. In this paper, we propose PAC-Net, a simple yet effective approach for transfer learning based on pruning. PAC-Net consists of three steps: Prune, Allocate, and Calibrate (PAC). The main idea behind these steps is to identify essential weights for the source task, fine-tune on the source task by updating the essential weights, and then calibrate on the target task by updating the remaining redundant weights. Under the various and extensive set of inductive transfer learning experiments, we show that our method achieves state-of-the-art performance by a large margin. Sanghoon Myung, In Huh, Wonik Jang, Jae Myung Choe, Jisu Ryu, Daesin Kim, Kee-Eung Kim, Changwook Jeong |
ICML | 7 |
| 2022 | Data Augmentation for Learning to Play in Text-Based GamesabstractImproving generalization in text-based games serves as a useful stepping-stone towards reinforcement learning (RL) agents with generic linguistic ability. Data augmentation for generalization in RL has shown to be very successful in classic control and visual tasks, but there is no prior work for text-based games. We propose Transition-Matching Permutation, a novel data augmentation technique for text-based games, where we identify phrase permutations that match as many transitions in the trajectory data. We show that applying this technique results in state-of-the-art performance in the Cooking Game benchmark suite for text-based games. Jinhyeon Kim, Kee-Eung Kim |
IJCAI | 2 |
| 2022 | LobsDICE: Offline Learning from Observation via Stationary Distribution Correction EstimationabstractWe consider the problem of learning from observation (LfO), in which the agent aims to mimic the expert's behavior from the state-only demonstrations by experts. We additionally assume that the agent cannot interact with the environment but has access to the action-labeled transition data collected by some agents with unknown qualities. This offline setting for LfO is appealing in many real-world scenarios where the ground-truth expert actions are inaccessible and the arbitrary environment interactions are costly or risky. In this paper, we present LobsDICE, an offline LfO algorithm that learns to imitate the expert policy via optimization in the space of stationary distributions. Our algorithm solves a single convex minimization problem, which minimizes the divergence between the two state-transition distributions induced by the expert and the agent policy. Through an extensive set of offline LfO tasks, we show that LobsDICE outperforms strong baseline methods. Geon-Hyeong Kim, Jongmin Lee 0004, Youngsoo Jang, Hongseok Yang, Kee-Eung Kim |
NeurIPS | 5 |
| 2022 | Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsabstractWe consider local kernel metric learning for off-policy evaluation (OPE) of deterministic policies in contextual bandits with continuous action spaces. Our work is motivated by practical scenarios where the target policy needs to be deterministic due to domain requirements, such as prescription of treatment dosage and duration in medicine. Although importance sampling (IS) provides a basic principle for OPE, it is ill-posed for the deterministic target policy with continuous actions. Our main idea is to relax the target policy and pose the problem as kernel-based estimation, where we learn the kernel metric in order to minimize the overall mean squared error (MSE). We present an analytic solution for the optimal metric, based on the analysis of bias and variance. Whereas prior work has been limited to scalar action spaces or kernel bandwidth selection, our work takes a step further being capable of vector action spaces and metric optimization. We show that our estimator is consistent, and significantly reduces the MSE compared to baseline OPE methods through experiments on various domains. Haanvid Lee, Jongmin Lee 0004, Yunseon Choi, Wonseok Jeon, Byung-Jun Lee 0001, Yung-Kyun Noh, Kee-Eung Kim |
NeurIPS | 7 |
| 2021 | Dual Correction Strategy for Ranking Distillation in Top-N Recommender SystemabstractKnowledge Distillation (KD), which transfers the knowledge of a well-trained large model (teacher) to a small model (student), has become an important area of research for practical deployment of recommender systems. Recently, Relaxed Ranking Distillation (RRD) has shown that distilling the ranking information in the recommendation list significantly improves the performance. However, the method still has limitations in that 1) it does not fully utilize the prediction errors of the student model, which makes the training not fully efficient, and 2) it only distills the user-side ranking information, which provides an insufficient view under the sparse implicit feedback. This paper presents Dual Correction strategy for Distillation (DCD), which transfers the ranking information from the teacher model to the student model in a more efficient manner. Most importantly, DCD uses the discrepancy between the teacher model and the student model predictions to decide which knowledge to be distilled. By doing so, DCD essentially provides the learning guidance tailored to "correcting" what the student model has failed to accurately predict. This process is applied for transferring the ranking information from the user-side as well as the item-side to address sparse implicit user feedback. Our experiments show that the proposed method outperforms the state-of-the-art baselines, and ablation studies validate the effectiveness of each component. Youngjune Lee, Kee-Eung Kim |
CIKM | 2 |
| 2021 | Representation Balancing Offline Model-based Reinforcement Learning
Byung-Jun Lee 0001, Jongmin Lee 0004, Kee-Eung Kim |
ICLR | 3 |
| 2021 | Monte-Carlo Planning and Learning with Language Action Value Estimates
Youngsoo Jang, Seokin Seo, Jongmin Lee 0004, Kee-Eung Kim |
ICLR | 4 |
| 2021 | Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic
Deunsol Yoon, Sunghoon Hong, Byung-Jun Lee 0001, Kee-Eung Kim |
ICLR | 4 |
| 2021 | OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationabstractWe consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL, the distributional shift becomes the primary source of difficulty, which arises from the deviation of the target policy being optimized from the behavior policy used for data collection. This typically causes overestimation of action values, which poses severe problems for model-free algorithms that use bootstrapping. To mitigate the problem, prior offline RL algorithms often used sophisticated techniques that encourage underestimation of action values, which introduces an additional set of hyperparameters that need to be tuned properly. In this paper, we present an offline RL algorithm that prevents overestimation in a more principled way. Our algorithm, OptiDICE, directly estimates the stationary distribution corrections of the optimal policy and does not rely on policy-gradients, unlike previous offline RL algorithms. Using an extensive set of benchmark datasets for offline RL, we show that OptiDICE performs competitively with the state-of-the-art methods. Jongmin Lee 0004, Wonseok Jeon, Byung-Jun Lee 0001, Joelle Pineau, Kee-Eung Kim |
ICML | 5 |
| 2021 | Multi-View Representation Learning via Total Correlation ObjectiveabstractMulti-View Representation Learning (MVRL) aims to discover a shared representation of observations from different views with the complex underlying correlation. In this paper, we propose a variational approach which casts MVRL as maximizing the amount of total correlation reduced by the representation, aiming to learn a shared latent representation that is informative yet succinct to capture the correlation among multiple views. To this end, we introduce a tractable surrogate objective function under the proposed framework, which allows our method to fuse and calibrate the observations in the representation space. From the information-theoretic perspective, we show that our framework subsumes existing multi-view generative models. Lastly, we show that our approach straightforwardly extends to the Partial MVRL (PMVRL) setting, where the observations are missing without any regular pattern. We demonstrate the effectiveness of our approach in the multi-view translation and classification tasks, outperforming strong baseline methods. HyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung Kim |
NeurIPS | 4 |
| 2021 | Corrigendum to 'Extensions to Hybrid Code Networks for FAIR Dialog Data' Computer Speech & Language volume 53 (2019) Pages 80-91
Jiyeon Ham, Soohyun Lim, Kyeng-Hun Lee, Kee-Eung Kim |
Comput. Speech Lang. | 4 |
| 2020 | Bayes-Adaptive Monte-Carlo Planning and Learning for Goal-Oriented DialoguesabstractWe consider a strategic dialogue task, where the ability to infer the other agent's goal is critical to the success of the conversational agent. While this problem can be naturally formulated as Bayesian planning, it is known to be a very difficult problem due to its enormous search space consisting of all possible utterances. In this paper, we introduce an efficient Bayes-adaptive planning algorithm for goal-oriented dialogues, which combines RNN-based dialogue generation and MCTS-based Bayesian planning in a novel way, leading to robust decision-making under the uncertainty of the other agent's goal. We then introduce reinforcement learning for the dialogue agent that uses MCTS as a strong policy improvement operator, casting reinforcement learning as iterative alternation of planning and supervised-learning of self-generated dialogues. In the experiments, we demonstrate that our Bayes-adaptive dialogue planning agent significantly outperforms the state-of-the-art in a negotiation dialogue domain. We also show that reinforcement learning via MCTS further improves end-task performance without diverging from human language. Youngsoo Jang, Jongmin Lee 0004, Kee-Eung Kim |
AAAI | 3 |
| 2020 | Residual Neural ProcessesabstractA Neural Process (NP) is a map from a set of observed input-output pairs to a predictive distribution over functions, which is designed to mimic other stochastic processes' inference mechanisms. NPs are shown to work effectively in tasks that require complex distributions, where traditional stochastic processes struggle, e.g. image completion tasks. This paper concerns the practical capacity of set function approximators despite their universality. By delving deeper into the relationship between an NP and a Bayesian last layer (BLL), it is possible to see that NPs may struggle in simple examples, which other stochastic processes can easily solve. In this paper, we propose a simple yet effective remedy; the Residual Neural Process (RNP) that leverages traditional BLL for faster training and better prediction. We demonstrate that the RNP shows faster convergence and better performance, both qualitatively and quantitatively. Byung-Jun Lee 0001, Seunghoon Hong, Kee-Eung Kim |
AAAI | 3 |
| 2020 | Monte-Carlo Tree Search in Continuous Action Spaces with Value GradientsabstractMonte-Carlo Tree Search (MCTS) is the state-of-the-art online planning algorithm for large problems with discrete action spaces. However, many real-world problems involve continuous action spaces, where MCTS is not as effective as in discrete action spaces. This is mainly due to common practices such as coarse discretization of the entire action space and failure to exploit local smoothness. In this paper, we introduce Value-Gradient UCT (VG-UCT), which combines traditional MCTS with gradient-based optimization of action particles. VG-UCT simultaneously performs a global search via UCT with respect to the finitely sampled set of actions and performs a local improvement via action value gradients. In the experiments, we demonstrate that our approach outperforms existing MCTS methods and other strong baseline algorithms for continuous action spaces. Jongmin Lee 0004, Wonseok Jeon, Geon-Hyeong Kim, Kee-Eung Kim |
AAAI | 4 |
| 2020 | End-to-End Neural Pipeline for Goal-Oriented Dialogue Systems using GPT-2abstractThe goal-oriented dialogue system needs to be optimized for tracking the dialogue flow and carrying out an effective conversation under various situations to meet the user goal.The traditional approach to building such a dialogue system is to take a pipelined modular architecture, where its modules are optimized individually.However, such an optimization scheme does not necessarily yield an overall performance improvement of the whole system.On the other hand, end-to-end dialogue systems with monolithic neural architecture are often trained only with input-output utterances, without taking into account the entire annotations available in the corpus.This scheme makes it difficult for goal-oriented dialogues where the system needs to be integrated with external systems or to provide interpretable information about why the system generated a particular response.In this paper, we present an end-to-end neural architecture for dialogue systems that addresses both challenges above.Our dialogue system achieved the success rate of 68.32%, the language understanding score of 4.149, and the response appropriateness score of 4.287 in human evaluations, which ranked the system at the top position in the end-to-end multi-domain dialogue system task in the 8th dialogue systems technology challenge (DSTC8). DongHoon Ham, Jeong-Gwan Lee, Youngsoo Jang, Kee-Eung Kim |
ACL | 4 |
| 2020 | Variational Inference for Sequential Data with Future Likelihood EstimatesabstractThe recent development of flexible and scalable variational inference algorithms has popularized the use of deep probabilistic models in a wide range of applications. However, learning and reasoning about high-dimensional models with nondifferentiable densities are still a challenge. For such a model, inference algorithms struggle to estimate the gradients of variational objectives accurately, due to high variance in their estimates. To tackle this challenge, we present a novel variational inference algorithm for sequential data, which performs well even when the density from the model is not differentiable, for instance, due to the use of discrete random variables. The key feature of our algorithm is that it estimates future likelihoods at all time steps. The estimated future likelihoods form the core of our new low-variance gradient estimator. We formally analyze our gradient estimator from the perspective of variational objective, and show the effectiveness of our algorithm with synthetic and real datasets. Geon-Hyeong Kim, Youngsoo Jang, Hongseok Yang, Kee-Eung Kim |
ICML | 4 |
| 2020 | Batch Reinforcement Learning with Hyperparameter GradientsabstractWe consider the batch reinforcement learning problem where the agent needs to learn only from a fixed batch of data, without further interaction with the environment. In such a scenario, we want to prevent the optimized policy from deviating too much from the data collection policy since the estimation becomes highly unstable otherwise due to the off-policy nature of the problem. However, imposing this requirement too strongly will result in a policy that merely follows the data collection policy. Unlike prior work where this trade-off is controlled by hand-tuned hyperparameters, we propose a novel batch reinforcement learning approach, batch optimization of policy and hyperparameter (BOPAH), that uses a gradient-based optimization of the hyperparameter using held-out data. We show that BOPAH outperforms other batch reinforcement learning algorithms in tabular and continuous control tasks, by finding a good balance to the trade-off between adhering to the data collection policy and pursuing the possible policy improvement. Byung-Jun Lee 0001, Jongmin Lee 0004, Peter Vrancx, Kee-Eung Kim |
ICML | 5 |
| 2020 | Reinforcement Learning for Control with Multiple FrequenciesabstractMany real-world sequential decision problems involve multiple action variables whose control frequencies are different, such that actions take their effects at different periods. While these problems can be formulated with the notion of multiple action persistences in factored-action MDP (FA-MDP), it is non-trivial to solve them efficiently since an action-persistent policy constructed from a stationary policy can be arbitrarily suboptimal, rendering solution methods for the standard FA-MDPs hardly applicable. In this paper, we formalize the problem of multiple control frequencies in RL and provide its efficient solution method. Our proposed method, Action-Persistent Policy Iteration (AP-PI), provides a theoretical guarantee on the convergence to an optimal solution while incurring only a factor of $|A|$ increase in time complexity during policy improvement step, compared to the standard policy iteration for FA-MDPs. Extending this result, we present Action-Persistent Actor-Critic (AP-AC), a scalable RL algorithm for high-dimensional control tasks. In the experiments, we demonstrate that AP-AC significantly outperforms the baselines on several continuous control tasks and a traffic control simulation, which highlights the effectiveness of our method that directly optimizes the periodic non-stationary policy for tasks with multiple control frequencies. Jongmin Lee 0004, Byung-Jun Lee 0001, Kee-Eung Kim |
NeurIPS | 3 |
| 2020 | Variational Interaction Information Maximization for Cross-domain DisentanglementabstractCross-domain disentanglement is the problem of learning representations partitioned into domain-invariant and domain-specific representations, which is a key to successful domain transfer or measuring semantic distance between two domains. Grounded in information theory, we cast the simultaneous learning of domain-invariant and domain-specific representations as a joint objective of multiple information constraints, which does not require adversarial training or gradient reversal layers. We derive a tractable bound of the objective and propose a generative model named Interaction Information Auto-Encoder (IIAE). Our approach reveals insights on the desirable representation for cross-domain disentanglement and its connection to Variational Auto-Encoder (VAE). We demonstrate the validity of our model in the image-to-image translation and the cross-domain retrieval tasks. We further show that our model achieves the state-of-the-art performance in the zero-shot sketch based image retrieval task, even without external knowledge. HyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung Kim |
NeurIPS | 4 |
| 2020 | Foreword: special issue for the journal track of the 12th Asian conference on machine learning (ACML 2020)
Kee-Eung Kim, Vineeth N. Balasubramanian |
Mach. Learn. | 1 |
| 2020 | Foreword: special issue for the journal track of the 11th Asian Conference on Machine Learning (ACML 2019)
Kee-Eung Kim |
Mach. Learn. | 1 |
| 2020 | Layered Behavior Modeling via Combining Descriptive and Prescriptive Approaches: A Case Study of Infantry Company EngagementabstractDefense modeling and simulation (DM&S) has brought insights into how to efficiently operate combat entities, such as soldiers and weapon systems. Most DM&S works have been developed to reflect accurate descriptions of military doctrines, yet these doctrines provide only guidelines of military operations, not details about how the combat entities should behave. Because such vague parts are often fulfilled with the appropriate behavior of combat entities in a battlefield, one part argues that DM&S should consider individual combat behaviors as well. However, it is known as an infeasible problem discovering best individual actions from infinite searching space, such as the battlefield. This paper proposes a layered behavior modeling to practically resolve this issue. The proposed method applies descriptive modeling to reduce the searching space by employing domain-specific knowledge; and prescriptive modeling to discover best individual actions in the reduced space. For the generalization, the proposed method adapts both modeling methods being modularized, and then the proposed method suggested an interface between them that is based on their semantic analogies. Both modeling methods are modularized, so they are interacted through an interface defined in the proposed method. This paper presents a realization of the proposed method through a case study of infantry company-level operations. In the case study, the proposed method is implemented with discrete event system specification formalism as the descriptive part and Markov decision process as the prescriptive part. The experimental results illustrated that the combat effectiveness resulted from the proposed method is statistically better than that from the descriptive-only modeling, and the difference would be guided by the objective of the combat behavior. Through the presented experimental results and the discussion, this paper argues that future DM&S should consider a broad spectrum from the battlefield incorporating the rational behavior of military individuals. Jang Won Bae, Junseok Lee 0001, Kanghoon Lee, Jongmin Lee 0004, Kee-Eung Kim, Il-Chul Moon |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2019 | Trust Region Sequential Variational InferenceabstractStochastic variational inference has emerged as an effective method for performing inference on or learning complex models for data. Yet, one of the challenges in stochastic variational inference is handling high-dimensional data, such as sequential data, and models with non-differentiable densities caused by, for instance, the use of discrete latent variables. In such cases, it is challenging to control the variance of the gradient estimator used in stochastic variational inference, while low variance is often one of the key properties needed for successful inference. In this work, we present a new algorithm for stochastic variational inference of sequential models which trades off bias for variance to tackle this challenge effectively. Our algorithm is inspired by variance reduction techniques in reinforcement learning, yet it uniquely adopts their key ideas in the context of stochastic variational inference. We demonstrate the effectiveness of our approach through formal analysis and experiments on synthetic and real-world datasets. Geon-Hyeong Kim, Youngsoo Jang, Jongmin Lee 0004, Wonseok Jeon, Hongseok Yang, Kee-Eung Kim |
ACML | 6 |
| 2019 | Extensions to hybrid code networks for FAIR dialog dataset
Jiyeon Ham, Soohyun Lim, Kyeng-Hun Lee, Kee-Eung Kim |
Comput. Speech Lang. | 4 |
| 2019 | Bayesian optimistic Kullback-Leibler exploration
Kanghoon Lee, Geon-Hyeong Kim, Pedro A. Ortega, Daniel D. Lee, Kee-Eung Kim |
Mach. Learn. | 5 |
| 2018 | Imitation Learning via Kernel Mean EmbeddingabstractImitation learning refers to the problem where an agent learns a policy that mimics the demonstration provided by the expert, without any information on the cost function of the environment. Classical approaches to imitation learning usually rely on a restrictive class of cost functions that best explains the expert's demonstration, exemplified by linear functions of pre-defined features on states and actions. We show that the kernelization of a classical algorithm naturally reduces the imitation learning to a distribution learning problem, where the imitation policy tries to match the state-action visitation distribution of the expert. Closely related to our approach is the recent work on leveraging generative adversarial networks (GANs) for imitation learning, but our reduction to distribution learning is much simpler, robust to scarce expert demonstration, and sample efficient. We demonstrate the effectiveness of our approach on a wide range of high-dimensional control tasks. Kee-Eung Kim, Hyun Soo Park |
AAAI | 1 |
| 2018 | A Bayesian Approach to Generative Adversarial Imitation LearningabstractGenerative adversarial training for imitation learning has shown promising results on high-dimensional and continuous control tasks. This paradigm is based on reducing the imitation learning problem to the density matching problem, where the agent iteratively refines the policy to match the empirical state-action visitation frequency of the expert demonstration. Although this approach has shown to robustly learn to imitate even with scarce demonstration, one must still address the inherent challenge that collecting trajectory samples in each iteration is a costly operation. To address this issue, we first propose a Bayesian formulation of generative adversarial imitation learning (GAIL), where the imitation policy and the cost function are represented as stochastic neural networks. Then, we show that we can significantly enhance the sample efficiency of GAIL leveraging the predictive density of the cost, on an extensive set of imitation learning tasks with high-dimensional states and actions. Wonseok Jeon, Seokin Seo, Kee-Eung Kim |
NeurIPS | 3 |
| 2018 | Monte-Carlo Tree Search for Constrained POMDPsabstractMonte-Carlo Tree Search (MCTS) has been successfully applied to very large POMDPs, a standard model for stochastic sequential decision-making problems. However, many real-world problems inherently have multiple goals, where multi-objective formulations are more natural. The constrained POMDP (CPOMDP) is such a model that maximizes the reward while constraining the cost, extending the standard POMDP model. To date, solution methods for CPOMDPs assume an explicit model of the environment, and thus are hardly applicable to large-scale real-world problems. In this paper, we present CC-POMCP (Cost-Constrained POMCP), an online MCTS algorithm for large CPOMDPs that leverages the optimization of LP-induced parameters and only requires a black-box simulator of the environment. In the experiments, we demonstrate that CC-POMCP converges to the optimal stochastic action selection in CPOMDP and pushes the state-of-the-art by being able to scale to very large problems. Jongmin Lee 0004, Geon-Hyeong Kim, Pascal Poupart, Kee-Eung Kim |
NeurIPS | 4 |
| 2018 | Cross-Language Neural Dialog State Tracker for Large Ontologies Using Hierarchical AttentionabstractDialog state tracking, which refers to identifying the user intent from utterances, is one of the most important tasks in dialog management. In this paper, we present our dialog state tracker developed for the fifth dialog state tracking challenge, which focused on cross-language adaptation using a very scarce machine-translated training data when compared to the size of the ontology. Our dialog state tracker is based on the bi-directional long short-term memory network with a hierarchical attention mechanism in order to spot important words in user utterances. The user intent is predicted by finding the closest keyword in the ontology to the attention-weighted word vector. With the suggested methodology, our tracker can overcome various difficulties due to the scarce training data that existing machine learning-based trackers had, such as predicting user intents they have not seen before. We show that our tracker outperforms other trackers submitted to the challenge with respect to most of the performance measures. Youngsoo Jang, Jiyeon Ham, Byung-Jun Lee 0001, Kee-Eung Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Hierarchically-partitioned Gaussian Process ApproximationabstractThe Gaussian process (GP) is a simple yet powerful probabilistic framework for various machine learning tasks. However, exact algorithms for learning and prediction are prohibitive to be applied to large datasets due to inherent computational complexity. To overcome this main limitation, various techniques have been proposed, and in particular, local GP algorithms that scales “truly linearly” with respect to the dataset size. In this paper, we introduce a hierarchical model based on local GP for large-scale datasets, which stacks inducing points over inducing points in layers. By using different kernels in each layer, the overall model becomes multi-scale and is able to capture both long- and short-range dependencies. We demonstrate the effectiveness of our model by speed-accuracy performance on challenging real-world datasets. Byung-Jun Lee 0001, Jongmin Lee 0004, Kee-Eung Kim |
AISTATS | 3 |
| 2017 | Constrained Bayesian Reinforcement Learning via Approximate Linear ProgrammingabstractIn this paper, we consider the safe learning scenario where we need to restrict the exploratory behavior of a reinforcement learning agent. Specifically, we treat the problem as a form of Bayesian reinforcement learning in an environment that is modeled as a constrained MDP (CMDP) where the cost function penalizes undesirable situations. We propose a model-based Bayesian reinforcement learning (BRL) algorithm for such an environment, eliciting risk-sensitive exploration in a principled way. Our algorithm efficiently solves the constrained BRL problem by approximate linear programming, and generates a finite state controller in an off-line manner. We provide theoretical guarantees and demonstrate empirically that our approach outperforms the state of the art. Jongmin Lee 0004, Youngsoo Jang, Pascal Poupart, Kee-Eung Kim |
IJCAI | 4 |
| 2017 | Generative Local Metric Learning for Kernel RegressionabstractThis paper shows how metric learning can be used with Nadaraya-Watson (NW) kernel regression. Compared with standard approaches, such as bandwidth selection, we show how metric learning can significantly reduce the mean square error (MSE) in kernel regression, particularly for high-dimensional data. We propose a method for efficiently learning a good metric function based upon analyzing the performance of the NW estimator for Gaussian-distributed data. A key feature of our approach is that the NW estimator with a learned metric uses information from both the global and local structure of the training data. Theoretical and empirical results confirm that the learned metric can considerably reduce the bias and MSE for kernel regression even when the data are not confined to Gaussian. Yung-Kyun Noh, Masashi Sugiyama, Kee-Eung Kim, Frank C. Park 0001, Daniel D. Lee |
NIPS | 3 |
| 2017 | Hybrid modeling and simulation of tactical maneuvers in computer generated forceabstractDefense modeling and simulation (DM&S) offers insights into the efficient operations of combat entities, e.g., soldiers and weapon systems. Most DM&S aim at exact description of military doctrines, but often the doctrines fails to provide detail action procedures about how the combat entities conduct military operations. Such unspecified descriptions are filled with the rational behaviors of the combat entities in a battlefield, and thereby the combat effectiveness from these combat entities would differ. Also, by incorporating such rational factors, this could provide the insights that cannot be captured from the traditional works. To examine this postulation, this paper developed a computer generated force where the tactical maneuver of combat entities are realized by the combination of descriptive and prescriptive modeling. Specifically, the descriptive models describe the explicit action rules in military doctrines, and they are modeled using discrete event system specification (DEVS) formalism; the predictive models denoted the rational behavior of the combat entities under the military doctrines, and they are modeled using partially observable Markov decision process (POMDP). The provided results illustrated that the proposed approach helps to maintain a team formation effectively, and this formation maintenance lead to the better combat efficiency. Jang Won Bae, Bowon Nam, Kee-Eung Kim, Junseok Lee 0001, Il-Chul Moon |
SMC | 3 |
| 2017 | Foreword: special issue for the journal track of the 8th Asian conference on machine learning (ACML 2016)
Robert J. Durrant, Kee-Eung Kim, Geoff Holmes 0001, Stephen R. Marsland, Masashi Sugiyama, Zhi-Hua Zhou |
Mach. Learn. | 2 |
| 2016 | Bayesian Reinforcement Learning with Behavioral Feedback
Teakgyu Hong, Jongmin Lee 0004, Kee-Eung Kim, Pedro A. Ortega, Daniel D. Lee |
IJCAI | 3 |
| 2016 | Neural dialog state tracker for large ontologies by attention mechanismabstractThis paper presents a dialog state tracker submitted to Dialog State Tracking Challenge 5 (DSTC 5) with details. To tackle the challenging cross-language human-human dialog state tracking task with limited training data, we propose a tracker that focuses on words with meaningful context based on attention mechanism and bi-directional long short term memory (LSTM). The vocabulary including a plenty of proper nouns is vectorized with a sufficient amount of related texts crawled from web to learn a good embedding for words not existent in training dialogs. Despite its simplicity, our proposed tracker succeeded to achieve high accuracy without sophisticated pre- and post-processing. Youngsoo Jang, Jiyeon Ham, Byung-Jun Lee 0001, Youngjae Chang 0001, Kee-Eung Kim |
SLT | 5 |
| 2015 | Reward Shaping for Model-Based Bayesian Reinforcement LearningabstractBayesian reinforcement learning (BRL) provides a formal framework for optimal exploration-exploitation tradeoff in reinforcement learning. Unfortunately, it is generally intractable to find the Bayes-optimal behavior except for restricted cases. As a consequence, many BRL algorithms, model-based approaches in particular, rely on approximated models or real-time search methods. In this paper, we present potential-based shaping for improving the learning performance in model-based BRL. We propose a number of potential functions that are particularly well suited for BRL, and are domain-independent in the sense that they do not require any prior knowledge about the actual environment. By incorporating the potential function into real-time heuristic search, we show that we can significantly improve the learning performance in standard benchmark domains. Hyeoneun Kim, Woosang Lim, Kanghoon Lee, Yung-Kyun Noh, Kee-Eung Kim |
AAAI | 5 |
| 2015 | Tighter Value Function Bounds for Bayesian Reinforcement LearningabstractBayesian reinforcement learning (BRL) provides a principled framework for optimal exploration-exploitation tradeoff in reinforcement learning. We focus on model based BRL, which involves a compact formulation of the optimal tradeoff from the Bayesian perspective. However, it still remains a computational challenge to compute the Bayes-optimal policy. In this paper, we propose a novel approach to compute tighter value function bounds of the Bayes-optimal value function, which is crucial for improving the performance of many model-based BRL algorithms. We then present how our bounds can be integrated into real-time AO* heuristic search, and provide a theoretical analysis on the impact of improved bounds on the search efficiency. We also provide empirical results on standard BRL domains that demonstrate the effectiveness of our approach. Kanghoon Lee, Kee-Eung Kim |
AAAI | 2 |
| 2015 | Approximate Linear Programming for Constrained Partially Observable Markov Decision ProcessesabstractIn many situations, it is desirable to optimize a sequence of decisions by maximizing a primary objective while respecting some constraints with respect to secondary objectives. Such problems can be naturally modeled as constrained partially observable Markov decision processes (CPOMDPs) when the environment is partially observable. In this work, we describe a technique based on approximate linear programming to optimize policies in CPOMDPs. The optimization is performed offline and produces a finite state controller with desirable performance guarantees. The approach outperforms a constrained version of point-based value iteration on a suite of benchmark problems. Pascal Poupart, Aarti Malhotra, Pei Pei, Kee-Eung Kim, Bongseok Goh, Michael H. Bowling |
AAAI | 4 |
| 2015 | Reactive bandits with attitudeabstractWe consider a general class of K-armed bandits that adapt to the actions of the player. A single continuous parameter characterizes the “attitude” of the bandit, ranging from stochastic to cooperative or to fully adversarial in nature. The player seeks to maximize the expected return from the adaptive bandit, and the associated optimization problem is related to the free energy of a statistical mechanical system under an external field. When the underlying stochastic distribution is Gaussian, we derive an analytic solution for the long run optimal player strategy for different regimes of the bandit. In the fully adversarial limit, this solution is equivalent to the Nash equilibrium of a two-player, zero-sum semi-infinite game. We show how optimal strategies can be learned from sequential draws and reward observations in these adaptive bandits using Bayesian filtering and Thompson sampling. Results show the qualitative difference in policy pseudo-regret between our proposed strategy and other well-known bandit algorithms. Pedro A. Ortega, Kee-Eung Kim, Daniel D. Lee |
AISTATS | 2 |
| 2015 | Hierarchical Bayesian Inverse Reinforcement LearningabstractInverse reinforcement learning (IRL) is the problem of inferring the underlying reward function from the expert's behavior data. The difficulty in IRL mainly arises in choosing the best reward function since there are typically an infinite number of reward functions that yield the given behavior data as optimal. Another difficulty comes from the noisy behavior data due to sub-optimal experts. We propose a hierarchical Bayesian framework, which subsumes most of the previous IRL algorithms as well as models the sub-optimality of the expert's behavior. Using a number of experiments on a synthetic problem, we demonstrate the effectiveness of our approach including the robustness of our hierarchical Bayesian framework to the sub-optimal expert behavior data. Using a real dataset from taxi GPS traces, we additionally show that our approach predicts the driving behavior with a high accuracy. Jaedeug Choi, Kee-Eung Kim |
IEEE Trans. Cybern. | 2 |
| 2014 | Optimizing Generative Dialog State Tracker via Cascading Gradient DescentabstractFor robust spoken dialog management, various dialog state tracking methods have been proposed. Although discriminative models are gaining popularity due to their superior performance, generative models based on the Partially Observable Markov Decision Process model still remain at-tractive since they provide an integrated framework for dialog state tracking and dialog policy optimization. Although a straightforward way to fit a generative model is to independently train the com-ponent probability models, we present a gradient descent algorithm that simultane-ously train all the component models. We show that the resulting tracker performs competitively with other top-performing trackers that participated in DSTC2. 1 Byung-Jun Lee 0001, Woosang Lim, Daejoong Kim, Kee-Eung Kim |
SIGDIAL Conference | 4 |
| 2013 | Bayesian Nonparametric Feature Construction for Inverse Reinforcement Learning
Jaedeug Choi, Kee-Eung Kim |
IJCAI | 2 |
| 2013 | Engineering Statistical Dialog State Trackers: A Case Study on DSTC
Daejoong Kim, Jaedeug Choi, Kee-Eung Kim, Jungsu Lee, Jinho Sohn |
SIGDIAL Conference | 3 |
| 2012 | Nonparametric Bayesian Inverse Reinforcement Learning for Multiple Reward FunctionsabstractWe present a nonparametric Bayesian approach to inverse reinforcement learning (IRL) for multiple reward functions. Most previous IRL algorithms assume that the behaviour data is obtained from an agent who is optimizing a single reward function, but this assumption is hard to be met in practice. Our approach is based on integrating the Dirichlet process mixture model into Bayesian IRL. We provide an efficient Metropolis-Hastings sampling algorithm utilizing the gradient of the posterior to estimate the underlying reward functions, and demonstrate that our approach outperforms the previous ones via experiments on a number of problem domains. Jaedeug Choi, Kee-Eung Kim |
NIPS | 2 |
| 2012 | Cost-Sensitive Exploration in Bayesian Reinforcement LearningabstractIn this paper, we consider Bayesian reinforcement learning (BRL) where actions incur costs in addition to rewards, and thus exploration has to be constrained in terms of the expected total cost while learning to maximize the expected long-term total reward. In order to formalize cost-sensitive exploration, we use the constrained Markov decision process (CMDP) as the model of the environment, in which we can naturally encode exploration requirements using the cost function. We extend BEETLE, a model-based BRL method, for learning in the environment with cost constraints. We demonstrate the cost-sensitive exploration behaviour in a number of simulated problems. Kee-Eung Kim, Pascal Poupart |
NIPS | 2 |
| 2012 | Exploiting symmetries for single- and multi-agent Partially Observable Stochastic Domains
Byung Kon Kang, Kee-Eung Kim |
Artif. Intell. | 2 |
| 2011 | A POMDP-Based Optimal Control of P300-Based Brain-Computer InterfacesabstractMost of the previous work on brain-computer interfaces (BCIs) exploiting the P300 in electroencephalography (EEG) has focused on low-level signal processing algorithms such as feature extraction and classification methods. Although a significant improvement has been made in the past, the accuracy of detecting P300 is limited by the inherently low signal-to-noise ratio in EEGs. In this paper, we present a systematic approach to optimize the interface using partially observable Markov decision processes (POMDPs). Through experiments involving human subjects, we show the P300 speller system that is optimized using the POMDP achieves a significant performance improvement in terms of the communication bandwidth in the interaction. Kee-Eung Kim, Yoon-Kyu Song |
AAAI | 2 |
| 2011 | Point-Based Value Iteration for Constrained POMDPs
Jaesong Lee, Kee-Eung Kim, Pascal Poupart |
IJCAI | 3 |
| 2011 | MAP Inference for Bayesian Inverse Reinforcement LearningabstractThe difficulty in inverse reinforcement learning (IRL) arises in choosing the best reward function since there are typically an infinite number of reward functions that yield the given behaviour data as optimal. Using a Bayesian framework, we address this challenge by using the maximum a posteriori (MAP) estimation for the reward function, and show that most of the previous IRL algorithms can be modeled into our framework. We also present a gradient method for the MAP estimation based on the (sub)differentiability of the posterior distribution. We show the effectiveness of our approach by comparing the performance of the proposed method to those of the previous algorithms. Jaedeug Choi, Kee-Eung Kim |
NIPS | 2 |
| 2011 | A Geometric Traversal Algorithm for Reward-Uncertain MDPs
Eunsoo Oh, Kee-Eung Kim |
UAI | 2 |
| 2011 | Inverse Reinforcement Learning in Partially Observable Environments
Jaedeug Choi, Kee-Eung Kim |
J. Mach. Learn. Res. | 2 |
| 2011 | Robust Performance Evaluation of POMDP-Based Dialogue SystemsabstractPartially observable Markov decision processes (POMDPs) have received significant interest in research on spoken dialogue systems, due to among many benefits its ability to naturally model the dialogue strategy selection problem under unreliable automated speech recognition. However, the POMDP approaches are essentially model-based, and as a result, the dialogue strategy computed from POMDP is still subject to the correctness of the model. In this paper, we extend some of the previous MDP user models to POMDPs, and evaluate the effects of user models on the dialogue strategy computed from POMDPs. We experimentally show that the strategies computed from POMDPs perform better than those from MDPs, and the strategies computed from poor user models fail severely when tested on different user models. This paper further investigates the evaluation methods for dialogue strategies, and proposes a method based on the bias-variance analysis for reliably estimating the dialogue performance. Jin H. Kim, Kee-Eung Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2010 | A POMDP approach to P300-based brain-computer interfacesabstractMost of the previous work on non-invasive brain-computer interfaces (BCIs) has been focused on feature extraction and classification algorithms to achieve high performance for the communication between the brain and the computer. While significant progress has been made in the lower layer of the BCI system, the issues in the higher layer have not been sufficiently addressed. Existing P300-based BCI systems, for example the P300 speller, use a random order of stimulus sequence for eliciting P300 signal for identifying users' intentions. This paper is about computing an optimal sequence of stimulus in order to minimize the number of stimuli, hence improving the performance. To accomplish this, we model the problem as a partially observable Markov decision process (POMDP), which is a model for planning in partially observable stochastic environments. Through simulation and human subject experiments, we show that our approach achieves a significant performance improvement in terms of the success rate and the bit rate. Kee-Eung Kim, Sungho Jo |
IUI | 2 |
| 2010 | Point-Based Bounded Policy Iteration for Decentralized POMDPs
Kee-Eung Kim |
PRICAI | 2 |
| 2009 | Inverse Reinforcement Learning in Partially Observable Environments
Jaedeug Choi, Kee-Eung Kim |
IJCAI | 2 |
| 2008 | Exploiting Symmetries in POMDPs for Point-Based Algorithms
Kee-Eung Kim |
AAAI | 1 |
| 2008 | Symbolic Heuristic Search Value Iteration for Factored POMDPs
Hyeong Seop Sim, Kee-Eung Kim, Jin Hyung Kim, Du-Seong Chang, Myoung-Wan Koo |
AAAI | 2 |
| 2008 | Effects of user modeling on POMDP-based dialogue systems
Hyeong Seop Sim, Kee-Eung Kim, Jin Hyung Kim, Hyunjeong Kim, Joo Won Sung |
INTERSPEECH | 3 |
| 2006 | Hand Grip Pattern Recognition for Mobile User Interfaces
Kee-Eung Kim, Wook Chang, Sung-Jung Cho, Junghyun Shim, Hyunjeong Lee, Joonah Park, Youngbeom Lee, Sangryoung Kim |
AAAI | 1 |
| 2005 | Variable bandwidth allocation scheme for energy efficient wireless sensor networkabstractIncreasing the lifetime of wireless sensors is essential for the proliferation of wireless sensor networks in various environments. In this paper, the relationship between bandwidth and energy consumption is exploited to increase the lifetime of the sensors. A variable bandwidth allocation scheme that uses time-frequency slot assignment is proposed to reduce the energy consumption of a collaborative sensor network which has large spatial variation in node density and event rates. To assign the time-frequency slots to the sensor network, a novel algorithm is presented, which results in significant energy savings over the conventional constant bandwidth allocation scheme. SeongHwan Cho, Kee-Eung Kim |
ICC | 2 |
| 2003 | Solving factored MDPs using non-homogeneous partitions
Kee-Eung Kim, Thomas L. Dean |
Artif. Intell. | 1 |
| 2002 | Solving Factored MDPs with Large Action Space Using Algebraic Decision Diagrams
Kee-Eung Kim, Thomas L. Dean |
PRICAI | 1 |
| 2001 | Solving Factored MDPs via Non-Homogeneous Partitioning
Kee-Eung Kim, Thomas L. Dean |
IJCAI | 1 |
| 2000 | Learning to Cooperate via Policy Search
Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau, Leslie Pack Kaelbling |
UAI | 2 |
| 1999 | Solving POMDPs by Searching the Space of Finite Policies
Nicolas Meuleau, Kee-Eung Kim, Leslie Pack Kaelbling, Anthony R. Cassandra |
UAI | 2 |
| 1999 | Learning Finite-State Controllers for Partially Observable Environments
Nicolas Meuleau, Leonid Peshkin, Kee-Eung Kim, Leslie Pack Kaelbling |
UAI | 3 |