EDBT 2026 Demo / reviewers in the wild / expert
Kyungjae Lee 0001
dblp:13/7265-1
· DBLP profile ↗
36ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0003-0147-2715ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 6 first-author · 18 since 2021Systems, architecture and hardware · 19 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function ApproximationabstractDistributional reinforcement learning improves performance by capturing environmental stochasticity, but a comprehensive theoretical understanding of its effectiveness remains elusive.
In addition, the intractable element of the infinite dimensionality of distributions has been overlooked.
In this paper, we present a regret analysis of distributional reinforcement learning with general value function approximation in a finite episodic Markov decision process setting.
We first introduce a key notion of Bellman unbiasedness which is essential for exactly learnable and provably efficient distributional updates in an online manner.
Among all types of statistical functionals for representing infinite-dimensional return distributions, our theoretical results demonstrate that only moment functionals can exactly capture the statistical information.
Secondly, we propose a provably efficient algorithm, SF-LSVI, that achieves a tight regret bound of $\tilde{O}(d_E H^{\frac{3}{2}}\sqrt{K})$ where $H$ is the horizon, $K$ is the number of episodes, and $d_E$ is the eluder dimension of a function class. Taehyun Cho, Seungyub Han, Seokhun Ju, Dohyeong Kim, Kyungjae Lee 0001, Jungwoo Lee 0001 |
ICML | 5 |
| 2025 | Policy-labeled Preference Learning: Is Preference Enough for RLHF?abstractTo design reward that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing models using reinforcement learning algorithms.
However, existing RLHF methods often misinterpret trajectories as being generated by an optimal policy, causing inaccurate likelihood estimation and suboptimal learning. To address this, we propose Policy-labeled Preference Learning (PPL) within the Direct Preference Optimization (DPO) framework, which resolves these likelihood mismatch problems by modeling human preferences with regret, reflecting the efficiency of executed policies. Additionally, we introduce a contrastive KL regularization term derived from regret-based principles to enhance sequential contrastive learning. Experiments in high-dimensional continuous control environments demonstrate PPL's significant improvements in offline RLHF performance and its effectiveness in online settings. Taehyun Cho, Seokhun Ju, Seungyub Han, Dohyeong Kim, Kyungjae Lee 0001, Jungwoo Lee 0001 |
ICML | 5 |
| 2025 | Learning-Based Dynamic Robot-to-Human HandoverabstractThis paper presents a novel learning-based approach to dynamic robot-to-human handover, addressing the challenges of delivering objects to a moving receiver. We hypothesize that dynamic handover, where the robot adjusts to the receiver's movements, results in more efficient and comfortable interaction compared to static handover, where the receiver is assumed to be stationary. To validate this, we developed a nonparametric method for generating continuous handover motion, conditioned on the receiver's movements, and trained the model using a dataset of 1,000 human-to-human handover demonstrations. We integrated preference learning for improved handover effectiveness and applied impedance control to ensure user safety and adaptiveness. The approach was evaluated in both simulation and real-world settings, with user studies demonstrating that dynamic handover significantly reduces handover time and improves user comfort compared to static methods. Videos and demonstrations of our approach are available at https://zerotohero7886.github.io/dyn-r2h-handover/. Hyeonseong Kim, Matthew K. X. J. Pan, Kyungjae Lee 0001 |
ICRA | 4 |
| 2025 | Self-Corrective Task Planning by Inverse Prompting with Large Language ModelsabstractIn robot task planning, large language models (LLMs) have shown significant promise in generating complex and long-horizon action sequences. However, it is observed that LLMs often produce responses that sound plausible but are not accurate. To address these problems, existing methods typically employ predefined error sets or external knowledge sources, requiring human efforts and computation resources. Recently, self-correction approaches have emerged, where LLM generates and refines plans, identifying errors by itself. Despite their effectiveness, they are more prone to failures in correction due to insufficient reasoning. In this paper, we introduce InversePrompt, a novel self-corrective task planning approach that leverages inverse prompting to enhance interpretability. Our method incorporates reasoning steps to provide clear, interpretable feedback. It generates inverse actions corresponding to the initially generated actions and verifies whether these inverse actions can restore the system to its original state, explicitly validating the logical coherence of the generated plans. The results on benchmark datasets show an average 16.3% higher success rate over existing LLM-based task planning methods. Our approach offers clearer justifications for feedback in real-world environments, resulting in more successful task completion than existing self-correction approaches across various scenarios. Hayun Lee, Jonghyeon Kim, Kyungjae Lee 0001, Eunwoo Kim |
ICRA | 4 |
| 2025 | Pareto Optimal Risk-Agnostic Distributional Bandits with Heavy-Tail RewardsabstractThis paper addresses the problem of multi-risk measure agnostic multi-armed bandits in heavy-tailed reward settings.
We propose a framework that leverages novel deviation inequalities for the $1$-Wasserstein distance to construct confidence intervals for Lipschitz risk measures.
The distributional LCB (DistLCB) algorithm is introduced, which achieves asymptotic optimality by deriving the first lower bounds for risk measure aware bandits with explicit sub-optimality gap dependencies.
The DistLCB is further extended to multi-risk objectives, which enables Pareto-optimal solutions that consider multiple aspects of reward distributions.
Additionally, we provide a regret analysis that includes both gap-dependent and gap-independent bounds for multi-risk settings.
Experiments validate the effectiveness of the proposed methods in synthetic and real-world applications. Kyungjae Lee 0001, Dohyeong Kim, Taehyun Cho, Chaeyeon Kim, Yunkyung Ko, Seungyub Han, Seokhun Ju, Dohyeok Lee, Sungbin Lim |
NeurIPS | 1 |
| 2024 | Placement Aware Grasp Planning for Efficient Sequential ManipulationabstractIn this paper, we address the problem of sequential pick-and-place with multiple objects when a specific goal configuration is given. The sequence of pick-and-place can generally be optimized using search-based methods. However, when the number of objects increases, there remains a challenge due to the exponential growth in computational complexity. Especially, the most challenging aspect is that a significant number of sequences are infeasible, leading to extended search times. In this regard, we propose an approach that efficiently addresses this issue by considering both pick and place aspects simultaneously while searching grasp poses with learning-based techniques to find feasible solutions expediently. In our experiments, the proposed method showed an enhancement of up to 90% in the average quality of trajectory for rearrangement benchmarks and about a 50% improvement in computational time in certain scenarios. Juhan Park, Daejong Jin, Kyungjae Lee 0001 |
ECAI | 3 |
| 2024 | Contextually Adaptive Algorithms for Gaussian Process Bandit Optimization Under Heavy-Tailed NoiseabstractWe consider a Gaussian process (GP) bandit optimization problem when the objective function lives in a reproducing kernel Hilbert space (RKHS), assuming that the payoffs follow a heavy-tailed distribution with a bounded (1+ϵ)-th moment for some ϵ∈(0,1]. Existing algorithms for this setting face practical challenges due to their significant computational demands and inconsistent theoretical guarantee to translation of noise distribution. To address these issues, we introduce two robust algorithms. The first algorithm utilizes a truncation estimator, achieving the same regret bound as that of the existing algorithm up to logarithmic terms with reduced time complexity. The second algorithm employs a median-of-means estimator and achieves more stable regret bound to alteration of noise distribution with lower time and space complexities compared to existing methods. Finally, we empirically validate the performance of our proposed algorithms against previous methods in both synthetic and real-world datasets. Hyeonjun Park 0001, Kyungjae Lee 0001 |
ECAI | 2 |
| 2024 | SPOTS: Stable Placement of Objects with Reasoning in Semi-Autonomous Teleoperation SystemsabstractPick-and-place is one of the fundamental tasks in robotics research. However, the attention has been mostly focused on the "pick" task, leaving the "place" task relatively unexplored. In this paper, we address the problem of placing objects in the context of a teleoperation framework. Particularly, we focus on two aspects of the place task: stability robustness and contextual reasonableness of object placements. Our proposed method combines simulation-driven physical stability verification via real-to-sim and the semantic reasoning capability of large language models. In other words, given place context information (e.g., user preferences, object to place, and current scene information), our proposed method outputs a probability distribution over the possible placement candidates, considering the robustness and reasonableness of the place task. Our proposed method is extensively evaluated in two simulation and one real world environments and we show that our method can greatly increase the physical plausibility of the placement as well as contextual soundness while considering user preferences. Code, video, and details are available at: https://joonhyunglee.github.io/spots/ Joonhyung Lee, Jeongeun Park 0002, Kyungjae Lee 0001 |
ICRA | 4 |
| 2024 | Spectral-Risk Safe Reinforcement Learning with Convergence GuaranteesabstractThe field of risk-constrained reinforcement learning (RCRL) has been developed to effectively reduce the likelihood of worst-case scenarios by explicitly handling risk-measure-based constraints.
However, the nonlinearity of risk measures makes it challenging to achieve convergence and optimality.
To overcome the difficulties posed by the nonlinearity, we propose a spectral risk measure-constrained RL algorithm, spectral-risk-constrained policy optimization (SRCPO), a bilevel optimization approach that utilizes the duality of spectral risk measures.
In the bilevel optimization structure, the outer problem involves optimizing dual variables derived from the risk measures, while the inner problem involves finding an optimal policy given these dual variables.
The proposed method, to the best of our knowledge, is the first to guarantee convergence to an optimum in the tabular setting.
Furthermore, the proposed method has been evaluated on continuous control tasks and showed the best performance among other RCRL algorithms satisfying the constraints.
Our code is available at https://github.com/rllab-snu/Spectral-Risk-Constrained-RL. Dohyeong Kim, Taehyun Cho, Seungyub Han, Hojun Chung, Kyungjae Lee 0001, Songhwai Oh |
NeurIPS | 5 |
| 2024 | Minimax Optimal Bandits for Heavy Tail RewardsabstractStochastic multiarmed bandits (stochastic MABs) are a problem of sequential decision-making with noisy rewards, where an agent sequentially chooses actions under unknown reward distributions to minimize cumulative regret. The majority of prior works on stochastic MABs assume that the reward distribution of each action has bounded supports or follows light-tailed distribution, i.e., sub-Gaussian distribution. However, in a variety of decision-making problems, the reward distributions follow a heavy-tailed distribution. In this regard, we consider stochastic MABs with heavy-tailed rewards, whose$p$th moment is bounded by a constant$\nu_{p}$for$1 Kyungjae Lee 0001, Sungbin Lim |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Perturbation-Based Best Arm Identification for Efficient Task Planning with Monte-Carlo Tree SearchabstractCombining task and motion planning (TAMP) is crucial for intelligent robots to perform complex and long-horizon tasks. In TAMP, many approaches generally employ Monte-Carlo tree search (MCTS) with upper confidence bound (UCB) for task planning to handle exploration-exploitation trade-off and find globally optimal solutions. However, since UCB basically considers the estimation error caused by noise, the error caused by insufficient optimization of the sub-tree is not represented. Hence, UCB-based approaches have the disadvantage of not exploring underestimated sub-trees. To alleviate this issue, we propose a novel tree search method using perturbation-based best-arm identification (PBAI). We theoretically prove the bound of the simple regret of our method and empirically verify that PBAI finds the optimal task plans faster and more efficiently than the existing algorithms. The source code of our proposed algorithm is available at https://github.com/jdj2261/pytamp. Daejong Jin, Juhan Park, Kyungjae Lee 0001 |
ICRA | 3 |
| 2023 | Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk CriterionabstractDistributional reinforcement learning algorithms have attempted to utilize estimated uncertainty for exploration, such as optimism in the face of uncertainty. However, using the estimated variance for optimistic exploration may cause biased data collection and hinder convergence or performance. In this paper, we present a novel distributional reinforcement learning that selects actions by randomizing risk criterion without losing the risk-neutral objective. We provide a perturbed distributional Bellman optimality operator by distorting the risk measure. Also,we prove the convergence and optimality of the proposed method with the weaker contraction property. Our theoretical results support that the proposed method does not fall into biased exploration and is guaranteed to converge to an optimal return. Finally, we empirically show that our method outperforms other existing distribution-based algorithms in various environments including Atari 55 games. Taehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee 0001, Jungwoo Lee 0001 |
NeurIPS | 4 |
| 2023 | Sequential Preference Ranking for Efficient Reinforcement Learning from Human FeedbackabstractReinforcement learning from human feedback (RLHF) alleviates the problem of designing a task-specific reward function in reinforcement learning by learning it from human preference. However, existing RLHF models are considered inefficient as they produce only a single preference data from each human feedback. To tackle this problem, we propose a novel RLHF framework called SeqRank, that uses sequential preference ranking to enhance the feedback efficiency. Our method samples trajectories in a sequential manner by iteratively selecting a defender from the set of previously chosen trajectories $\mathcal{K}$ and a challenger from the set of unchosen trajectories $\mathcal{U}\setminus\mathcal{K}$, where $\mathcal{U}$ is the replay buffer. We propose two trajectory comparison methods with different defender sampling strategies: (1) sequential pairwise comparison that selects the most recent trajectory and (2) root pairwise comparison that selects the most preferred trajectory from $\mathcal{K}$. We construct a data structure and rank trajectories by preference to augment additional queries. The proposed method results in at least 39.2% higher average feedback efficiency than the baseline and also achieves a balance between feedback efficiency and data dependency. We examine the convergence of the empirical risk and the generalization bound of the reward model with Rademacher complexity. While both trajectory comparison methods outperform conventional pairwise comparison, root pairwise comparison improves the average reward in locomotion tasks and the average success rate in manipulation tasks by 29.0% and 25.0%, respectively. The source code and the videos are provided in the supplementary material. Minyoung Hwang, Gunmin Lee, Hogun Kee, Kyungjae Lee 0001, Songhwai Oh |
NeurIPS | 5 |
| 2023 | Trust Region-Based Safe Distributional Reinforcement Learning for Multiple ConstraintsabstractIn safety-critical robotic tasks, potential failures must be reduced, and multiple constraints must be met, such as avoiding collisions, limiting energy consumption, and maintaining balance.
Thus, applying safe reinforcement learning (RL) in such robotic tasks requires to handle multiple constraints and use risk-averse constraints rather than risk-neutral constraints.
To this end, we propose a trust region-based safe RL algorithm for multiple constraints called a safe distributional actor-critic (SDAC).
Our main contributions are as follows: 1) introducing a gradient integration method to manage infeasibility issues in multi-constrained problems, ensuring theoretical convergence, and 2) developing a TD($\lambda$) target distribution to estimate risk-averse constraints with low biases.
We evaluate SDAC through extensive experiments involving multi- and single-constrained robotic tasks.
While maintaining high scores, SDAC shows 1.93 times fewer steps to satisfy all constraints in multi-constrained tasks and 1.78 times fewer constraint violations in single-constrained tasks compared to safe RL baselines.
Code is available at: https://github.com/rllab-snu/Safe-Distributional-Actor-Critic. Dohyeong Kim, Kyungjae Lee 0001, Songhwai Oh |
NeurIPS | 2 |
| 2023 | Score-based Generative Modeling through Stochastic Evolution Equations in Hilbert SpacesabstractContinuous-time score-based generative models consist of a pair of stochastic differential equations (SDEs)—a forward SDE that smoothly transitions data into a noise space and a reverse SDE that incrementally eliminates noise from a Gaussian prior distribution to generate data distribution samples—are intrinsically connected by the time-reversal theory on diffusion processes. In this paper, we investigate the use of stochastic evolution equations in Hilbert spaces, which expand the applicability of SDEs in two aspects: sample space and evolution operator, so they enable encompassing recent variations of diffusion models, such as generating functional data or replacing drift coefficients with image transformation. To this end, we derive a generalized time-reversal formula to build a bridge between probabilistic diffusion models and stochastic evolution equations and propose a score-based generative model called Hilbert Diffusion Model (HDM). Combining with Fourier neural operator, we verify the superiority of HDM for sampling functions from functional datasets with a power of kernel two-sample test of 4.2 on Quadratic, 0.2 on Melbourne, and 3.6 on Gridwatch, which outperforms existing diffusion models formulated in function spaces. Furthermore, the proposed method shows its strength in motion synthesis tasks by utilizing the Wiener process with values in Hilbert space. Finally, our empirical results on image datasets also validate a connection between HDM and diffusion models using heat dissipation, revealing the potential for exploring evolution operators and sample spaces. Sungbin Lim, Eun-Bi Yoon, Taehyun Byun, Taewon Kang, Seungwoo Kim, Kyungjae Lee 0001 |
NeurIPS | 6 |
| 2022 | Domain Generalization by Mutual-Information Regularization with Pre-trained Models
Junbum Cha, Kyungjae Lee 0001, Sungrae Park, Sanghyuk Chun |
ECCV (23) | 2 |
| 2022 | Semi-Autonomous Teleoperation via Learning Non-Prehensile Manipulation SkillsabstractIn this paper, we present a semi-autonomous teleoperation framework for a pick-and-place task using an RGB-D sensor. In particular, we assume that the target object is located in a cluttered environment where both prehensile grasping and non-prehensile manipulation are combined for efficient teleoperation. A trajectory-based reinforcement learning is utilized for learning the non-prehensile manipulation to rearrange the objects for enabling direct grasping. From the depth image of the cluttered environment and the location of the goal object, the learned policy can provide multiple options of non-prehensile manipulation to the human operator. We carefully design a reward function for the rearranging task where the policy is trained in a simulational environment. Then, the trained policy is transferred to a real-world and evaluated in a number of real-world experiments with the varying number of objects where we show that the proposed method outperforms manual keyboard control in terms of the time duration for the grasping. Yoonbyung Chai, Jeongeun Park 0002, Kyungjae Lee 0001 |
ICRA | 5 |
| 2022 | Safety Guided Policy OptimizationabstractIn reinforcement learning (RL), exploration is essential to achieve a globally optimal policy but unconstrained exploration can cause damages to robots and nearby people. To handle this safety issue in exploration, safe RL has been proposed to keep the agent under the specified safety constraints while maximizing cumulative rewards. This paper introduces a new safe RL method which can be applied to robots to operate under the safety constraints while learning. The key component of the proposed method is the safeguard module. The safeguard predicts the constraints in the near future and corrects actions such that the predicted constraints are not violated. Since actions are safely modified by the safeguard during exploration and policies are trained to imitate the corrected actions, the agent can safely explore. Additionally, the safeguard is sample efficient as it does not require long horizontal trajectories for training, so constraints can be satisfied within short time steps. The proposed method is extensively evaluated in simulation and experiments using a real robot. The results show that the proposed method achieves the best performance while satisfying safety constraints with minimal interaction with environments in all experiments. Dohyeong Kim, Kyungjae Lee 0001, Songhwai Oh |
IROS | 3 |
| 2020 | Monte Carlo Tree Search in Continuous Spaces Using Voronoi Optimistic Optimization with Regret BoundsabstractMany important applications, including robotics, data-center management, and process control, require planning action sequences in domains with continuous state and action spaces and discontinuous objective functions. Monte Carlo tree search (MCTS) is an effective strategy for planning in discrete action spaces. We provide a novel MCTS algorithm (voot) for deterministic environments with continuous action spaces, which, in turn, is based on a novel black-box function-optimization algorithm (voo) to efficiently sample actions. The voo algorithm uses Voronoi partitioning to guide sampling, and is particularly efficient in high-dimensional spaces. The voot algorithm has an instance of voo at each node in the tree. We provide regret bounds for both algorithms and demonstrate their empirical effectiveness in several high-dimensional problems including two difficult robotics planning problems. Kyungjae Lee 0001, Sungbin Lim, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 2 |
| 2020 | Task Agnostic Robust Learning on Corrupt Outputs by Correlation-Guided Mixture Density NetworksabstractIn this paper, we focus on weakly supervised learning with noisy training data for both classification and regression problems. We assume that the training outputs are collected from a mixture of a target and correlated noise distributions. Our proposed method simultaneously estimates the target distribution and the quality of each data which is defined as the correlation between the target and data generating distributions. The cornerstone of the proposed method is a Cholesky Block that enables modeling dependencies among mixture distributions in a differentiable manner where we maintain the distribution over the network weights. We first provide illustrative examples in both regression and classification tasks to show the effectiveness of the proposed method. Then, the proposed method is extensively evaluated in a number of experiments where we show that it constantly shows comparable or superior performances compared to existing baseline methods in the handling of noisy data. Kyungjae Lee 0001, Sungbin Lim |
CVPR | 3 |
| 2020 | Hierarchical 6-DoF Grasping with Approaching Direction SelectionabstractIn this paper, we tackle the problem of 6-DoF grasp detection which is crucial for robot grasping in cluttered real-world scenes. Unlike existing approaches which synthesize 6-DoF grasp data sets and train grasp quality networks with input grasp representations based on point clouds, we rather take a novel hierarchical approach which does not use any 6-DoF grasp data. We cast the 6-DoF grasp detection problem as a robot arm approaching direction selection problem using the existing 4-DoF grasp detection algorithm, by exploiting a fully convolutional grasp quality network for evaluating the quality of an approaching direction. To select the best approaching direction with the highest grasp quality, we propose an approaching direction selection method which leverages a geometry-based prior and a derivative-free optimization method. Specifically, we optimize the direction iteratively using the cross entropy method with initial samples of surface normal directions. Our algorithm efficiently finds diverse 6-DoF grasps by the novel way of evaluating and optimizing approaching directions. We validate that the proposed method outperforms other selection methods in scenarios with cluttered objects in a physics-based simulator. Finally, we show that our method outperforms the state-of-the-art grasp detection method in real-world experiments with robots. Hogun Kee, Kyungjae Lee 0001, Jaegoo Choy, Junhong Min, Sohee Lee, Songhwai Oh |
ICRA | 3 |
| 2020 | No-Regret Shannon Entropy Regularized Neural Contextual Bandit Online Learning for Robotic GraspingabstractIn this paper, we propose a novel contextual bandit algorithm that employs a neural network as a reward estimator and utilizes Shannon entropy regularization to encourage exploration, which is called Shannon entropy regularized neural contextual bandits (SERN). In many learning-based algorithms for robotic grasping, the lack of the real-world data hampers the generalization performance of a model and makes it difficult to apply a trained model to real-world problems. To handle this issue, the proposed method utilizes the benefit of an online learning. The proposed method trains a neural network to predict the success probability of a given grasp pose based on a depth image, which is called a grasp quality. We theoretically show that the SERN has a no regret property. We empirically demonstrate that the SERN outperforms ε-greedy in terms of sample efficiency. Kyungjae Lee 0001, Jaegu Choy, Hogun Kee, Songhwai Oh |
IROS | 1 |
| 2020 | MixGAIL: Autonomous Driving Using Demonstrations with Mixed QualitiesabstractIn this paper, we consider autonomous driving of a vehicle using imitation learning. Generative adversarial imitation learning (GAIL) is a widely used algorithm for imitation learning. This algorithm leverages positive demonstrations to imitate the behavior of an expert. In this paper, we propose a novel method, called mixed generative adversarial imitation learning (MixGAIL), which incorporates both of expert demonstrations and negative demonstrations, such as vehicle collisions. To this end, the proposed method utilizes an occupancy measure and a constraint function. The occupancy measure is used to follow expert demonstrations and provides a positive feedback. On the other hand, the constraint function is used for negative demonstrations to assert a negative feedback. Experimental results show that the proposed algorithm converges faster than the other baseline methods. Also, hardware experiments using a real-world RC car shows an outstanding performance and faster convergence compared with existing methods. Gunmin Lee, Dohyeong Kim, Wooseok Oh, Kyungjae Lee 0001, Songhwai Oh |
IROS | 4 |
| 2020 | Optimal Algorithms for Stochastic Multi-Armed Bandits with Heavy Tailed RewardsabstractIn this paper, we consider stochastic multi-armed bandits (MABs) with heavy-tailed rewards, whose p-th moment is bounded by a constant nu_p for 1 Kyungjae Lee 0001, Hongjun Yang, Sungbin Lim, Songhwai Oh |
NeurIPS | 1 |
| 2019 | Distributional Deep Reinforcement Learning with a Mixture of GaussiansabstractIn this paper, we propose a novel distributional reinforcement learning (RL) method which models the distribution of the sum of rewards using a mixture density network. Recently, it has been shown that modeling the randomness of the return distribution leads to better performance in Atari games and control tasks. Despite the success of the prior work, it has limitations which come from the use of a discrete distribution. First, it needs a projection step and softmax parametrization for the distribution, since it minimizes the KL divergence loss. Secondly, its performance depends on discretization hyperparameters such as the number of atoms and bounds of the support which require domain knowledge. We mitigate these problems with the proposed parameterization, a mixture of Gaussians. Furthermore, we propose a new distance metric called the Jensen-Tsallis distance, which allows the computation of the distance between two mixtures of Gaussians in a closed form. We have conducted various experiments to validate the proposed method, including Atari games and autonomous vehicle driving. Kyungjae Lee 0001, Songhwai Oh |
ICRA | 2 |
| 2019 | Soft Action Particle Deep Reinforcement Learning for a Continuous Action SpaceabstractRecent advances of actor-critic methods in deep reinforcement learning have enabled performing several continuous control problems. However, existing actor-critic algorithms require a large number of parameters to model policy and value functions where it can lead to overfitting issue and is difficult to tune hyperparameter. In this paper, we introduce a new off-policy actor-critic algorithm, which can reduce a significant number of parameters compared to existing actorcritic algorithms without any performance loss. The proposed method replaces the actor network with a set of action particles that employ few parameters. Then, the policy distribution is represented using state action value network with action particles. During the learning phase, to improve the performance of policy distribution, the location of action particles is updated to maximize state action values. To enhance the exploration and stable convergence, we add perturbation to action particles during training. In the experiment, we validate the proposed method in MuJoCo environments and empirically show that our method shows similar or better performance than the state-of-the-art actor-critic method with a smaller number of parameters. The experimental video can be found at http: //rllab.snu.ac.kr/multimedia. Minjae Kang 0002, Kyungjae Lee 0001, Songhwai Oh |
IROS | 2 |
| 2019 | Robust Learning From Demonstrations With Mixed Qualities Using Leveraged Gaussian ProcessesabstractIn this paper, we focus on the problem of learning from demonstration (LfD) where demonstrations with different proficiencies are provided without labeling. To this end, we model multiple policies with different qualities as correlated Gaussian processes and present a leverage optimization method that estimates the leverage of each policy where the difference between two leverages defines the correlation between the corresponding policies. To recover a single policy function of an expert, we present a sparsity constraint on the leverage parameters. We first show that the proposed leverage optimization method can recover the correlations between sensory fields where the fields are realized from correlated Gaussian processes and sensor measurements are collected from the fields. Furthermore, we applied the proposed method to autonomous driving experiments, where demonstrations are collected from three different driving modes. While the driving policies are not realized from correlated processes, the proposed method assigns reasonable leverages to the driving demonstrations. The estimated driving policy of an expert, which incorporates the optimized leverages, outperforms previous LfD methods in terms of both safety and driving quality. Kyungjae Lee 0001, Songhwai Oh |
IEEE Trans. Robotics | 2 |
| 2018 | Uncertainty-Aware Learning from Demonstration Using Mixture Density Networks with Sampling-Free Variance ModelingabstractIn this paper, we propose an uncertainty-aware learning from demonstration method by presenting a novel uncertainty estimation method utilizing a mixture density network appropriate for modeling complex and noisy human behaviors. The proposed uncertainty acquisition can be done with a single forward path without Monte Carlo sampling and is suitable for real-time robotics applications. Then, we show that it can be decomposed into explained variance and unexplained variance where the connections between aleatoric and epistemic uncertainties are addressed. The properties of the proposed uncertainty measure are analyzed through three different synthetic examples, absence of data, heavy measurement noise, and composition of functions scenarios. We show that each case can be distinguished using the proposed uncertainty measure and presented an uncertainty-aware learning from demonstration method for autonomous driving using this property. The proposed uncertainty-aware learning from demonstration method outperforms other compared methods in terms of safety using a complex real-world driving dataset. Kyungjae Lee 0001, Sungbin Lim, Songhwai Oh |
ICRA | 2 |
| 2018 | A Nonparametric Motion Flow Model for Human Robot CooperationabstractIn this paper, we present a novel nonparametric motion flow model that effectively describes a motion trajectory of a human and its application to human robot cooperation. To this end, motion flow similarity measure which considers both spatial and temporal properties of a trajectory is proposed by utilizing the mean and variance functions of a Gaussian process. We also present a human robot cooperation method using the proposed motion flow model. Given a set of interacting trajectories of two workers, the underlying reward function of cooperating behaviors is optimized by using the learned motion description as an input to the reward function where a stochastic trajectory optimization method is used to control a robot. The presented human robot cooperation method is compared with the state-of-the-art algorithm, which utilizes a mixture of interaction primitives (MIP), in terms of the RMS error between generated and target trajectories. While the proposed method shows comparable performance with the MIP when the full observation of human demonstrations is given, it shows superior performance when partial trajectory information is given. Kyungjae Lee 0001, Hyungju Andy Park, Songhwai Oh |
ICRA | 2 |
| 2018 | Maximum Causal Tsallis Entropy Imitation LearningabstractIn this paper, we propose a novel maximum causal Tsallis entropy (MCTE) framework for imitation learning which can efficiently learn a sparse multi-modal policy distribution from demonstrations. We provide the full mathematical analysis of the proposed framework. First, the optimal solution of an MCTE problem is shown to be a sparsemax distribution, whose supporting set can be adjusted. The proposed method has advantages over a softmax distribution in that it can exclude unnecessary actions by assigning zero probability. Second, we prove that an MCTE problem is equivalent to robust Bayes estimation in the sense of the Brier score. Third, we propose a maximum causal Tsallis entropy imitation learning (MCTEIL) algorithm with a sparse mixture density network (sparse MDN) by modeling mixture weights using a sparsemax distribution. In particular, we show that the causal Tsallis entropy of an MDN encourages exploration and efficient mixture utilization while Boltzmann Gibbs entropy is less effective. We validate the proposed method in two simulation studies and MCTEIL outperforms existing imitation learning methods in terms of average returns and learning multi-modal policies. Kyungjae Lee 0001, Songhwai Oh |
NeurIPS | 1 |
| 2017 | Scalable robust learning from demonstration with leveraged deep neural networksabstractIn this paper, we propose a novel algorithm for learning from demonstration, which can learn a policy function robustly from a large number of demonstrations with mixed qualities. While most of the existing approaches assume that demonstrations are collected from skillful experts, the proposed method alleviates such restrictions by estimating the proficiency level of each demonstration using the proposed leverage optimization. Furthermore, a novel leveraged cost function is proposed to represent a policy function using deep neural networks by reformulating the objective function of leveraged Gaussian process regression using the representer theorem. The proposed method is successfully applied to autonomous track driving tasks, where a large number of demonstrations with mixed qualities are given as training data without labels. Kyungjae Lee 0001, Songhwai Oh |
IROS | 2 |
| 2016 | Robust learning from demonstration using leveraged Gaussian processes and sparse-constrained optimizationabstractIn this paper, we propose a novel method for robust learning from demonstration using leveraged Gaussian process regression. While existing learning from demonstration (LfD) algorithms assume that demonstrations are given from skillful experts, the proposed method alleviates such assumption by allowing demonstrations from casual or novice users. To learn from demonstrations of mixed quality, we present a sparse-constrained leveraged optimization algorithm using proximal linearized minimization. The proposed sparse constrained leverage optimization algorithm is successfully applied to sensory field reconstruction and direct policy learning for planar navigation problems. In experiments, the proposed sparse-constrained method outperforms existing LfD methods. Kyungjae Lee 0001, Songhwai Oh |
ICRA | 2 |
| 2016 | Gaussian random paths for real-time motion planningabstractIn this paper, we propose Gaussian random paths by defining a probability distribution over continuous paths interpolating a finite set of anchoring points using Gaussian process regression. By utilizing the generative property of Gaussian random paths, a Gaussian random path planner is developed to safely steer a robot to a goal position. The Gaussian random path planner can be used in a number of applications, including local path planning for a mobile robot and trajectory optimization for whole body motion planning. We have conducted an extensive set of simulations and experiments, showing that the proposed planner outperforms look-ahead planners which use a pre-defined subset of egocentric trajectories in terms of collision rates and trajectory lengths. Furthermore, we apply the proposed method to existing trajectory optimization methods as an initialization step and demonstrate that it can help produce more cost-efficient trajectories. Kyungjae Lee 0001, Songhwai Oh |
IROS | 2 |
| 2016 | Robust modeling and prediction in dynamic environments using recurrent flow networksabstractTo enable safe motion planning in a dynamic environment, it is vital to anticipate and predict object movements. In practice, however, an accurate object identification among multiple moving objects is extremely challenging, making it infeasible to accurately track and predict individual objects. Furthermore, even for a single object, its appearance can vary significantly due to external effects, such as occlusions, varying perspectives, or illumination changes. In this paper, we propose a novel recurrent network architecture called a recurrent flow network that can infer the velocity of each cell and the probability of future occupancy from a sequence of occupancy grids which we refer to as an occupancy flow. The parameters of the recurrent flow network are optimized using Bayesian optimization. The proposed method outperforms three baseline optical flow methods, Lucas-Kanade, Lucas-Kanade with Tikhonov regularization, and HornSchunck methods, and a Bayesian occupancy grid filter in terms of both prediction accuracy and robustness to noise. Kyungjae Lee 0001, Songhwai Oh |
IROS | 2 |
| 2016 | Inverse reinforcement learning with leveraged Gaussian processesabstractIn this paper, we propose a novel inverse reinforcement learning algorithm with leveraged Gaussian processes that can learn from both positive and negative demonstrations. While most existing inverse reinforcement learning (IRL) methods suffer from the lack of information near low reward regions, the proposed method alleviates this issue by incorporating (negative) demonstrations of what not to do. To mathematically formulate negative demonstrations, we introduce a novel generative model which can generate both positive and negative demonstrations using a parameter, called proficiency. Moreover, since we represent a reward function using a leveraged Gaussian process which can model a nonlinear function, the proposed method can effectively estimate the structure of a nonlinear reward function. Kyungjae Lee 0001, Songhwai Oh |
IROS | 1 |
| 2015 | Leveraged non-stationary Gaussian process regression for autonomous robot navigationabstractIn this paper, we propose a novel regression method that can incorporate both positive and negative training data into a single regression framework. In detail, a leveraged kernel function for non-stationary Gaussian process regression is proposed. With this new kernel function, we can vary the correlation betwen two inputs in both positive and negative directions by adjusting leverage parameters. By using this property, the resulting leveraged non-stationary Gaussian process regression can anchor the regressor to the positive data while avoiding the negative data. We first prove the positive semi-definiteness of the leveraged kernel function using Bochner's theorem. Then, we apply the leveraged non-stationary Gaussian process regression to a real-time motion control problem. In this case, the positive data refer to what to do and the negative data indicate what not to do. The results show that the controller using both positive and negative data outperforms the controller using positive data only in terms of the collision rate given training sets of the same size. Eunwoo Kim, Kyungjae Lee 0001, Songhwai Oh |
ICRA | 3 |