EDBT 2026 Demo / reviewers in the wild / expert
Woo Kyung Kim
dblp:306/0140
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-6214-4171ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 45% Generative modeling · 10% Planning, search and constraint satisfaction · 10% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill-based reinforcement learning |
1.6 | 2 | 2025 | In-Context Policy Adaptation via Cross-Domain Skill Diffusion · AAAI 2025 Robust Policy Learning via Offline Skill Diffusion · AAAI 2024 |
Machine learning › Reinforcement learning
policy adaptation |
1.5 | 2 | 2025 | In-Context Policy Adaptation via Cross-Domain Skill Diffusion · AAAI 2025 Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents · NeurIPS 2023 |
Machine learning › Generative modeling
diffusion model |
1.5 | 2 | 2024 | LLM-based Skill Diffusion for Zero-shot Policy Adaptation · NeurIPS 2024 Robust Policy Learning via Offline Skill Diffusion · AAAI 2024 |
Machine learning › Reinforcement learning
imitation learning |
1.4 | 2 | 2024 | Incremental Learning of Retrievable Skills For Efficient Continual Task Adaptation · NeurIPS 2024 One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill · ICML 2023 |
Machine learning › Learning paradigms
continual learning |
0.8 | 1 | 2024 | Incremental Learning of Retrievable Skills For Efficient Continual Task Adaptation · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning |
0.8 | 1 | 2024 | Embodied CoT Distillation From LLM To Off-the-shelf Agents · ICML 2024 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.8 | 1 | 2024 | Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation · IJCAI 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
language-based planning |
0.8 | 1 | 2024 | Embodied CoT Distillation From LLM To Off-the-shelf Agents · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
LLM distillation |
0.8 | 1 | 2024 | Embodied CoT Distillation From LLM To Off-the-shelf Agents · ICML 2024 |
Machine learning › Reinforcement learning
policy learning |
0.8 | 1 | 2024 | Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation · IJCAI 2024 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.7 | 1 | 2023 | One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill · ICML 2023 |
Machine learning › Reinforcement learning › imitation learning › few-shot imitation learning
one-shot imitation learning |
0.7 | 1 | 2023 | One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill · ICML 2023 |
Computer vision › Vision and language
vision-language model |
0.7 | 1 | 2023 | Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
zero-shot domain adaptation |
0.7 | 1 | 2023 | One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill · ICML 2023 |
Robotics › Motion planning and robot control › robot learning
manipulation skill learning |
0.3 | 1 | 2025 | In-Context Policy Adaptation via Cross-Domain Skill Diffusion · AAAI 2025 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.2 | 1 | 2024 | Robust Policy Learning via Offline Skill Diffusion · AAAI 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2024 | Embodied CoT Distillation From LLM To Off-the-shelf Agents · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models |
0.2 | 1 | 2024 | Embodied CoT Distillation From LLM To Off-the-shelf Agents · ICML 2024 |
Robotics › Motion planning and robot control › robot learning
embodied manipulation |
0.2 | 1 | 2023 | Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents · NeurIPS 2023 |
Robotics › Robot navigation and mapping
embodied navigation |
0.2 | 1 | 2023 | Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.6offline dataset · 0.9dynamic domain prompting · 0.9pareto optimization · 0.8knowledge graph · 0.8in-context learning · 0.8imitation learning · 0.8hierarchical encoding · 0.8guided diffusion · 0.8chain-of-thought distillation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aspect-augmented distillation of task-oriented dialogues to small language modelsabstract• Considering user aspects improves task-oriented dialogue performance • Large language models adapt to user aspects; small models lack aspectawareness • Large language models generate synthetic aspect-specific dialogues for distillation • Aspect-aware capabilities distilled from large to small language models Research on developing dialogue systems with large language models (LLMs) has been extensive, relying heavily on LLMs’ capabilities to generate contextually nuanced responses. Yet, these approaches are not easily transferable to smaller language models (sLMs), particularly in task-oriented dialogue (ToD) scenarios, where dialogue systems are required to engage in personalized interactions with humans. In this paper, we investigate LLM distillation approaches for sLM-based ToD systems and present an Aspect-Augmented Dialogue Distillation (A2D2) framework, aiming to compress the human aspect-aware capabilities of an LLM into an sLM while ensuring the fulfillment on task specific requirements. The framework incorporates a set of human aspects individually into LLM-based ToD data generation to improve the effectiveness and efficiency of the LLM-to-sLM distillation process, thereby establishing robust sLM-based ToD systems that are adaptable to diverse users and achieving higher task success rates. We demonstrate that the sLM-based ToD systems derived through A2D2 yield competitive performance on various ToD scenarios including unseen task settings, adapting to a wide range of synthetic users characterized by multiple aspects. Jongmoon Jun, Woo Kyung Kim, Hyunseong Na, Honguk Woo, Jeehyeong Kim |
Expert Syst. Appl. | 2 |
| 2025 | In-Context Policy Adaptation via Cross-Domain Skill DiffusionabstractIn this work, we present an in-context policy adaptation (ICPAD) framework designed for long-horizon multi-task environments, exploring diffusion-based skill learning techniques in cross-domain settings. The framework enables rapid adaptation of skill-based reinforcement learning policies to diverse target domains, especially under stringent constraints on no model updates and only limited target domain data. Specifically, the framework employs a cross-domain skill diffusion scheme, where domain-agnostic prototype skills and a domain-grounded skill adapter are learned jointly and effectively from an offline dataset through cross-domain consistent diffusion processes. The prototype skills act as primitives for common behavior representations of long-horizon policies, serving as a lingua franca to bridge different domains. Furthermore, to enhance the in-context adaptation performance, we develop a dynamic domain prompting scheme that guides the diffusion-based skill adapter toward better alignment with the target domain. Through experiments with robotic manipulation in Metaworld and autonomous driving in CARLA, we show that our ICPAD framework achieves superior policy adaptation performance under limited target domain data conditions for various cross-domain configurations including differences in environment dynamics, agent embodiment, and task horizon. Minjong Yoo, Woo Kyung Kim, Honguk Woo |
AAAI | 2 |
| 2024 | Robust Policy Learning via Offline Skill DiffusionabstractSkill-based reinforcement learning (RL) approaches have shown considerable promise, especially in solving long-horizon tasks via hierarchical structures. These skills, learned task-agnostically from offline datasets, can accelerate the policy learning process for new tasks. Yet, the application of these skills in different domains remains restricted due to their inherent dependency on the datasets, which poses a challenge when attempting to learn a skill-based policy via RL for a target domain different from the datasets' domains. In this paper, we present a novel offline skill learning framework DuSkill which employs a guided Diffusion model to generate versatile skills extended from the limited skills in datasets, thereby enhancing the robustness of policy learning for tasks in different domains. Specifically, we devise a guided diffusion-based skill decoder in conjunction with the hierarchical encoding to disentangle the skill embedding space into two distinct representations, one for encapsulating domain-invariant behaviors and the other for delineating the factors that induce domain variations in the behaviors. Our DuSkill framework enhances the diversity of skills learned offline, thus enabling to accelerate the learning procedure of high-level policies for different domains. Through experiments, we show that DuSkill outperforms other skill-based imitation learning and RL algorithms for several long-horizon tasks, demonstrating its benefits in few-shot imitation and online RL. Woo Kyung Kim, Minjong Yoo, Honguk Woo |
AAAI | 1 |
| 2024 | Embodied CoT Distillation From LLM To Off-the-shelf AgentsabstractWe address the challenge of utilizing large language models (LLMs) for complex embodied tasks, in the environment where decision-making systems operate timely on capacity-limited, off-the-shelf devices. We present DeDer, a framework for decomposing and distilling the embodied reasoning capabilities from LLMs to efficient, small language model (sLM)-based policies. In DeDer, the decision-making process of LLM-based strategies is restructured into a hierarchy with a reasoning-policy and planning-policy. The reasoning-policy is distilled from the data that is generated through the embodied in-context learning and self-verification of an LLM, so it can produce effective rationales. The planning-policy, guided by the rationales, can render optimized plans efficiently. In turn, DeDer allows for adopting sLMs for both policies, deployed on off-the-shelf devices. Furthermore, to enhance the quality of intermediate rationales, specific to embodied tasks, we devise the embodied knowledge graph, and to generate multiple rationales timely through a single inference, we also use the contrastively prompted attention model. Our experiments with the ALFRED benchmark demonstrate that DeDer surpasses leading language planning and distillation approaches, indicating the applicability and efficiency of sLM-based embodied policies derived through DeDer. Wonje Choi 0003, Woo Kyung Kim, Minjong Yoo, Honguk Woo |
ICML | 2 |
| 2024 | Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation
Woo Kyung Kim, Minjong Yoo, Honguk Woo |
IJCAI | 1 |
| 2024 | LLM-based Skill Diffusion for Zero-shot Policy AdaptationabstractRecent advances in data-driven imitation learning and offline reinforcement learning have highlighted the use of expert data for skill acquisition and the development of hierarchical policies based on these skills. However, these approaches have not significantly advanced in adapting these skills to unseen contexts, which may involve changing environmental conditions or different user requirements. In this paper, we present a novel LLM-based policy adaptation framework LDuS which leverages an LLM to guide the generation process of a skill diffusion model upon contexts specified in language, facilitating zero-shot skill-based policy adaptation to different contexts. To implement the skill diffusion model, we adapt the loss-guided diffusion with a sequential in-painting technique, where target trajectories are conditioned by masking them with past state-action sequences, thereby enabling the robust and controlled generation of skill trajectories in test-time. To have a loss function for a given context, we employ the LLM-based code generation with iterative refinement, by which the code and controlled trajectory are validated to align with the context in a closed-loop manner. Through experiments, we demonstrate the zero-shot adaptability of LDuS to various context types including different specification levels, multi-modality, and varied temporal conditions for several robotic manipulation tasks, outperforming other language-conditioned imitation and planning methods. Woo Kyung Kim, Jooyoung Kim 0001, Honguk Woo |
NeurIPS | 1 |
| 2024 | Incremental Learning of Retrievable Skills For Efficient Continual Task AdaptationabstractContinual Imitation Learning (CiL) involves extracting and accumulating task knowledge from demonstrations across multiple stages and tasks to achieve a multi-task policy. With recent advancements in foundation models, there has been a growing interest in adapter-based CiL approaches, where adapters are established parameter-efficiently for tasks newly demonstrated. While these approaches isolate parameters for specific tasks and tend to mitigate catastrophic forgetting, they limit knowledge sharing among different demonstrations. We introduce IsCiL, an adapter-based CiL framework that addresses this limitation of knowledge sharing by incrementally learning shareable skills from different demonstrations, thus enabling sample-efficient task adaptation using the skills particularly in non-stationary CiL environments. In IsCiL, demonstrations are mapped into the state embedding space, where proper skills can be retrieved upon input states through prototype-based memory. These retrievable skills are incrementally learned on their corresponding adapters. Our CiL experiments with complex tasks in the Franka-Kitchen and Meta-World demonstrate the robust performance of IsCiL in both task adaptation and sample-efficiency. We also show a simple extension of IsCiL for task unlearning scenarios. Daehee Lee 0001, Minjong Yoo, Woo Kyung Kim, Wonje Choi 0003, Honguk Woo |
NeurIPS | 3 |
| 2023 | One-shot Imitation in a Non-Stationary Environment via Multi-Modal SkillabstractOne-shot imitation is to learn a new task from a single demonstration, yet it is a challenging problem to adopt it for complex tasks with the high domain diversity inherent in a non-stationary environment. To tackle the problem, we explore the compositionality of complex tasks, and present a novel skill-based imitation learning framework enabling one-shot imitation and zero-shot adaptation; from a single demonstration for a complex unseen task, a semantic skill sequence is inferred and then each skill in the sequence is converted into an action sequence optimized for environmental hidden dynamics that can vary over time. Specifically, we leverage a vision-language model to learn a semantic skill set from offline video datasets, where each skill is represented on the vision-language embedding space, and adapt meta-learning with dynamics inference to enable zero-shot skill adaptation. We evaluate our framework with various one-shot imitation scenarios for extended multi-stage Meta-world tasks, showing its superiority in learning complex tasks, generalizing to dynamics changes, and extending to different demonstration conditions and modalities, compared to other baselines. Sangwoo Shin, Daehee Lee 0001, Minjong Yoo, Woo Kyung Kim, Honguk Woo |
ICML | 4 |
| 2023 | Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied AgentsabstractFor embodied reinforcement learning (RL) agents interacting with the environment, it is desirable to have rapid policy adaptation to unseen visual observations, but achieving zero-shot adaptation capability is considered as a challenging problem in the RL context. To address the problem, we present a novel contrastive prompt ensemble (ConPE) framework which utilizes a pretrained vision-language model and a set of visual prompts, thus enables efficient policy learning and adaptation upon a wide range of environmental and physical changes encountered by embodied agents. Specifically, we devise a guided-attention-based ensemble approach with multiple visual prompts on the vision-language model to construct robust state representations. Each prompt is contrastively learned in terms of an individual domain factors that significantly affects the agent's egocentric perception and observation. For a given task, the attention-based ensemble and policy are jointly learned so that the resulting state representations not only generalize to various domains but are also optimized for learning the task. Through experiments, we show that ConPE outperforms other state-of-the-art algorithms for several embodied agent tasks including navigation in AI2THOR, manipulation in Metaworld, and autonomous driving in CARLA, while also improving the sample efficiency of policy learning and adaptation. Wonje Choi 0003, Woo Kyung Kim, Honguk Woo |
NeurIPS | 2 |