Honguk Woo

dblp:63/6072 · DBLP profile ↗
← Back
41ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0001-6948-3440ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 2 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorSystems, architecture and hardware · 3 · 1 since 2021Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Aspect-augmented distillation of task-oriented dialogues to small language models
abstract
• Considering user aspects improves task-oriented dialogue performance • Large language models adapt to user aspects; small models lack aspectawareness • Large language models generate synthetic aspect-specific dialogues for distillation • Aspect-aware capabilities distilled from large to small language models Research on developing dialogue systems with large language models (LLMs) has been extensive, relying heavily on LLMs’ capabilities to generate contextually nuanced responses. Yet, these approaches are not easily transferable to smaller language models (sLMs), particularly in task-oriented dialogue (ToD) scenarios, where dialogue systems are required to engage in personalized interactions with humans. In this paper, we investigate LLM distillation approaches for sLM-based ToD systems and present an Aspect-Augmented Dialogue Distillation (A2D2) framework, aiming to compress the human aspect-aware capabilities of an LLM into an sLM while ensuring the fulfillment on task specific requirements. The framework incorporates a set of human aspects individually into LLM-based ToD data generation to improve the effectiveness and efficiency of the LLM-to-sLM distillation process, thereby establishing robust sLM-based ToD systems that are adaptable to diverse users and achieving higher task success rates. We demonstrate that the sLM-based ToD systems derived through A2D2 yield competitive performance on various ToD scenarios including unseen task settings, adapting to a wide range of synthetic users characterized by multiple aspects.
Jongmoon Jun, Woo Kyung Kim, Hyunseong Na, Honguk Woo, Jeehyeong Kim
Expert Syst. Appl.4
2025 In-Context Policy Adaptation via Cross-Domain Skill Diffusion
abstract
In this work, we present an in-context policy adaptation (ICPAD) framework designed for long-horizon multi-task environments, exploring diffusion-based skill learning techniques in cross-domain settings. The framework enables rapid adaptation of skill-based reinforcement learning policies to diverse target domains, especially under stringent constraints on no model updates and only limited target domain data. Specifically, the framework employs a cross-domain skill diffusion scheme, where domain-agnostic prototype skills and a domain-grounded skill adapter are learned jointly and effectively from an offline dataset through cross-domain consistent diffusion processes. The prototype skills act as primitives for common behavior representations of long-horizon policies, serving as a lingua franca to bridge different domains. Furthermore, to enhance the in-context adaptation performance, we develop a dynamic domain prompting scheme that guides the diffusion-based skill adapter toward better alignment with the target domain. Through experiments with robotic manipulation in Metaworld and autonomous driving in CARLA, we show that our ICPAD framework achieves superior policy adaptation performance under limited target domain data conditions for various cross-domain configurations including differences in environment dynamics, agent embodiment, and task horizon.
Minjong Yoo, Woo Kyung Kim, Honguk Woo
AAAI3
2025 NFVAgent: A Retrieval-Augmented LLM Agent for Resilient NFV Failure Recovery
abstract
As modern networks continue to grow in scale, heterogeneity, and complexity, automation in Network Function Virtualization (NFV) management has become increasingly critical. Traditional NFV management systems, which rely on static policies and manual interventions, often fail to respond effectively to dynamic network conditions. This leads to delayed or suboptimal failure recovery. Moreover, the increasing diversity of NFV deployments, spanning heterogeneous configurations, tenants, and operational standards, restricts the flexibility and scalability of existing policy-based solutions. Considering these limitations, this work explores Large Language Model (LLM) agent approaches by harnessing the advanced reasoning capabilities of LLMs within NFV operational context. While LLMbased approaches are structurally promising, LLMs often lack explicit exposure to the diverse and domain-specific characteristics of NFV environments, resulting in limited generalization and reliability in real-world scenarios. We introduce NFVAgent, an LLM-driven NFV recovery framework built on RetrievalAugmented Generation (RAG). It overcomes adaptation limitations by continuously updating its knowledge base through experiences gathered from both testbed and deployment environments. This evolving knowledge, validated across diverse NFV environments, is seamlessly integrated into the agent’s decisionmaking loop, enabling it to generate appropriate recovery actions. The framework supports dynamic reasoning over up-to-date knowledge, allowing the agent not only to react to failures but to continually refine its understanding of NFV failure patterns and effective recovery strategies. We evaluate NFVAgent on an alarm dataset comprising over 500 failure events with diverse NFV environments, each configured based on industry standards (e.g., ONAP, ETSI, O-RAN). Experimental results show that NFVAgent achieves up to $99.8 \%$ recovery accuracy, outperforming existing policy-based methods by an average margin of $37.7 \%$ in multitenant environments. This highlights the practical viability and performance benefits of integrating LLM agents with retrieval mechanisms that leverage NFV-specific operational knowledge in real-world recovery tasks.
Yeunjin Woo, Honguk Woo
CNSM2
2025 NeSyC: A Neuro-symbolic Continual Learner For Complex Embodied Tasks in Open Domains
abstract
We explore neuro-symbolic approaches to generalize actionable knowledge, enabling embodied agents to tackle complex tasks more effectively in open-domain environments. A key challenge for embodied agents is the generalization of knowledge across diverse environments and situations, as limited experiences often confine them to their prior knowledge. To address this issue, we introduce a novel framework, NeSyC, a neuro-symbolic continual learner that emulates the hypothetico-deductive model by continually formulating and validating knowledge from limited experiences through the combined use of Large Language Models (LLMs) and symbolic tools. Specifically, we devise a contrastive generality improvement scheme within NeSyC, which iteratively generates hypotheses using LLMs and conducts contrastive validation via symbolic tools. This scheme reinforces the justification for admissible actions while minimizing the inference of inadmissible ones. Additionally, we incorporate a memory-based monitoring scheme that efficiently detects action errors and triggers the knowledge refinement process across domains. Experiments conducted on diverse embodied task benchmarks—including ALFWorld, VirtualHome, Minecraft, RLBench, and a real-world robotic scenario—demonstrate that NeSyC is highly effective in solving complex embodied tasks across a range of open-domain environments.
Wonje Choi 0003, Sanghyun Ahn, Daehee Lee 0001, Honguk Woo
ICLR5
2025 Model Risk-sensitive Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) is becoming critical in risk-sensitive areas such as finance and autonomous driving, where incorrect decisions can lead to substantial financial loss or compromised safety. However, traditional risk-sensitive offline RL methods often struggle with accurately assessing risk, with minor errors in the estimated return potentially causing significant inaccuracies of risk estimation. These challenges are intensified by distribution shifts inherent in offline RL. To mitigate these issues, we propose a model risk-sensitive offline RL framework designed to minimize the worst-case of risks across a set of plausible alternative scenarios rather than solely focusing on minimizing estimated risk. We present a critic-ensemble criterion method that identifies the plausible alternative scenarios without introducing additional hyperparameters. We also incorporate the learned Fourier feature framework and the IQN framework to address spectral bias in neural networks, which can otherwise lead to severe errors in calculating model risk. Our experiments in finance and self-driving scenarios demonstrate that the proposed framework significantly reduces risk, by $11.2\%$ to $18.5\%$, compared to the most outperforming risk-sensitive offline RL baseline, particularly in highly uncertain environments.
Gwangpyo Yoo, Honguk Woo
ICLR2
2025 World Model Implanting for Test-time Adaptation of Embodied Agents
abstract
In embodied AI, a persistent challenge is enabling agents to robustly adapt to novel domains without requiring extensive data collection or retraining. To address this, we present a world model implanting framework (WorMI) that combines the reasoning capabilities of large language models (LLMs) with independently learned, domain-specific world models through test-time composition. By allowing seamless implantation and removal of the world models, the embodied agent’s policy achieves and maintains cross-domain adaptability. In the WorMI framework, we employ a prototype-based world model retrieval approach, utilizing efficient trajectory-based abstract representation matching, to incorporate relevant models into test-time composition. We also develop a world-wise compound attention method that not only integrates the knowledge from the retrieved world models but also aligns their intermediate representations with the reasoning model’s representation within the agent’s policy. This framework design effectively fuses domain-specific knowledge from multiple world models, ensuring robust adaptation to unseen domains. We evaluate our WorMI on the VirtualHome and ALFWorld benchmarks, demonstrating superior zero-shot and few-shot performance compared to several LLM-based approaches across a range of unseen domains. These results highlight the framework’s potential for scalable, real-world deployment in embodied agent scenarios where adaptability and data efficiency are essential.
Minjong Yoo, Jinwoo Jang, Sihyung Yoon, Honguk Woo
ICML4
2025 Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning
abstract
Recent advances in large language models (LLMs) have enabled the automatic generation of executable code for task planning and control in embodied agents such as robots, demonstrating the potential of LLM-based embodied intelligence. However, these LLM-based code-as-policies approaches often suffer from limited environmental grounding, particularly in dynamic or partially observable settings, leading to suboptimal task success rates due to incorrect or incomplete code generation. In this work, we propose a neuro-symbolic embodied task planning framework that incorporates explicit symbolic verification and interactive validation processes during code generation. In the validation phase, the framework generates exploratory code that actively interacts with the environment to acquire missing observations while preserving task-relevant states. This integrated process enhances the grounding of generated code, resulting in improved task reliability and success rates in complex environments. We evaluate our framework on RLBench and in real-world settings across dynamic, partially observable scenarios. Experimental results demonstrate that our framework improves task success rates by 46.2\% over Code as Policies baselines and attains over 86.8\% executability of task-relevant actions, thereby enhancing the reliability of task planning in dynamic environments.
Sanghyun Ahn, Wonje Choi 0003, Honguk Woo
NeurIPS5
2025 NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
abstract
We address the challenge of adopting language models (LMs) for embodied tasks in dynamic environments, where online access to large-scale inference engines or symbolic planners is constrained due to latency, connectivity, and resource limitations. To this end, we present NeSyPr, a novel embodied reasoning framework that compiles knowledge via neurosymbolic proceduralization, thereby equipping LM-based agents with structured, adaptive, and timely reasoning capabilities. In NeSyPr, task-specific plans are first explicitly generated by a symbolic tool leveraging its declarative knowledge. These plans are then transformed into composable procedural representations that encode the plans' implicit production rules, enabling the resulting composed procedures to be seamlessly integrated into the LM's inference process. This neurosymbolic proceduralization abstracts and generalizes multi-step symbolic structured path-finding and reasoning into single-step LM inference, akin to human knowledge compilation. It supports efficient test-time inference without relying on external symbolic guidance, making it well suited for deployment in latency-sensitive and resource-constrained physical systems. We evaluate NeSyPr on the embodied benchmarks PDDLGym, VirtualHome, and ALFWorld, demonstrating its efficient reasoning capabilities over large-scale reasoning models and a symbolic planner, while using more compact LMs.
Wonje Choi 0003, Honguk Woo
NeurIPS3
2025 Policy Compatible Skill Incremental Learning via Lazy Learning Interface
abstract
Skill Incremental Learning (SIL) is the process by which an embodied agent expands and refines its skill set over time by leveraging experience gained through interaction with its environment or by the integration of additional data. SIL facilitates efficient acquisition of hierarchical policies grounded in reusable skills for downstream tasks. However, as the skill repertoire evolves, it can disrupt compatibility with existing skill-based policies, limiting their reusability and generalization. In this work, we propose SIL-C, a novel framework that ensures skill-policy compatibility, allowing improvements in incrementally learned skills to enhance the performance of downstream policies without requiring policy re-training or structural adaptation. SIL-C employs a bilateral lazy learning-based mapping technique to dynamically align the subtask space referenced by policies with the skill space decoded into agent behaviors. This enables each subtask, derived from the policy's decomposition of a complex task, to be executed by selecting an appropriate skill based on trajectory distribution similarity. We evaluate SIL-C across diverse SIL scenarios and demonstrate that it maintains compatibility between evolving skills and downstream policies while ensuring efficiency throughout the learning process.
Daehee Lee 0001, Dongsu Lee, TaeYoon Kwack, Wonje Choi 0003, Honguk Woo
NeurIPS5
2024 Robust Policy Learning via Offline Skill Diffusion
abstract
Skill-based reinforcement learning (RL) approaches have shown considerable promise, especially in solving long-horizon tasks via hierarchical structures. These skills, learned task-agnostically from offline datasets, can accelerate the policy learning process for new tasks. Yet, the application of these skills in different domains remains restricted due to their inherent dependency on the datasets, which poses a challenge when attempting to learn a skill-based policy via RL for a target domain different from the datasets' domains. In this paper, we present a novel offline skill learning framework DuSkill which employs a guided Diffusion model to generate versatile skills extended from the limited skills in datasets, thereby enhancing the robustness of policy learning for tasks in different domains. Specifically, we devise a guided diffusion-based skill decoder in conjunction with the hierarchical encoding to disentangle the skill embedding space into two distinct representations, one for encapsulating domain-invariant behaviors and the other for delineating the factors that induce domain variations in the behaviors. Our DuSkill framework enhances the diversity of skills learned offline, thus enabling to accelerate the learning procedure of high-level policies for different domains. Through experiments, we show that DuSkill outperforms other skill-based imitation learning and RL algorithms for several long-horizon tasks, demonstrating its benefits in few-shot imitation and online RL.
Woo Kyung Kim, Minjong Yoo, Honguk Woo
AAAI3
2024 SemTra: A Semantic Skill Translator for Cross-Domain Zero-Shot Policy Adaptation
abstract
This work explores the zero-shot adaptation capability of semantic skills, semantically interpretable experts' behavior patterns, in cross-domain settings, where a user input in interleaved multi-modal snippets can prompt a new long-horizon task for different domains. In these cross-domain settings, we present a semantic skill translator framework SemTra which utilizes a set of multi-modal models to extract skills from the snippets, and leverages the reasoning capabilities of a pretrained language model to adapt these extracted skills to the target domain. The framework employs a two-level hierarchy for adaptation: task adaptation and skill adaptation. During task adaptation, seq-to-seq translation by the language model transforms the extracted skills into a semantic skill sequence, which is tailored to fit the cross-domain contexts. Skill adaptation focuses on optimizing each semantic skill for the target domain context, through parametric instantiations that are facilitated by language prompting and contrastive learning-based context inferences. This hierarchical adaptation empowers the framework to not only infer a complex task specification in one-shot from the interleaved multi-modal snippets, but also adapt it to new domains with zero-shot learning abilities. We evaluate our framework with Meta-World, Franka Kitchen, RLBench, and CARLA environments. The results clarify the framework's superiority in performing long-horizon tasks and adapting to different domains, showing its broad applicability in practical use cases, such as cognitive robots interpreting abstract instructions and autonomous vehicles operating under varied configurations.
Sangwoo Shin, Minjong Yoo, Honguk Woo
AAAI4
2024 Risk-Conditioned Reinforcement Learning: A Generalized Approach for Adapting to Varying Risk Measures
abstract
In application domains requiring mission-critical decision making, such as finance and robotics, the optimal policy derived by reinforcement learning (RL) often hinges on a preference for risk management. Yet, the dynamic nature of risk measures poses considerable challenges to achieving generalization and adaptation of risk-sensitive policies in the context of RL. In this paper, we propose a risk-conditioned RL model that enables rapid policy adaptation to varying risk measures via a unified risk representation, the Weighted Value-at-Risk (WV@R). To sample risk measures that avoid undue optimism, we construct a risk proposal network employing a conditional adversarial auto-encoder and a normalizing flow. This network establishes coherent representations for risk measures, preserving the continuity in terms of the Wasserstein distance on the risk measures. The normalizing flow is used to support non-crossing quantile regression that obtains valid samples for risk measures, and it is also applied to the agent’s critic to ascertain the preservation of monotonicity in quantile estimations. Through experiments with locomotion, finance, and self-driving scenarios, we show that our model is capable of adapting to a range of risk measures, achieving comparable performance to the baseline models individually trained for each measure. Our model often outperforms the baselines, especially in the cases when exploration is required during training but risk-aversion is favored during evaluation.
Gwangpyo Yoo, Honguk Woo
AAAI3
2024 Model Adaptation for Time Constrained Embodied Control
abstract
When adopting a deep learning model for embodied agents, it is required that the model structure be optimized for specific tasks and operational conditions. Such optimization can be static such as model compression or dynamic such as adaptive inference. Yet, these techniques have not been fully investigated for embodied control systems subject to time constraints, which necessitate sequential decision-making for multiple tasks, each with distinct inference latency limitations. In this paper, we present MoDeC, a time constraint-aware embodied control framework using the modular model adaptation. We formulate model adaptation to varying operational conditions on resource and time restrictions as dynamic routing on a modular network, incorporating these conditions as part of multi-task objectives. Our evaluation across several vision-based embodied environments demonstrates the robustness of MoDeC, showing that it outperforms other model adaptation methods in both performance and adherence to time constraints in robotic manipulation and autonomous driving applications.
Jaehyun Song, Minjong Yoo, Honguk Woo
CVPR3
2024 Embodied CoT Distillation From LLM To Off-the-shelf Agents
abstract
We address the challenge of utilizing large language models (LLMs) for complex embodied tasks, in the environment where decision-making systems operate timely on capacity-limited, off-the-shelf devices. We present DeDer, a framework for decomposing and distilling the embodied reasoning capabilities from LLMs to efficient, small language model (sLM)-based policies. In DeDer, the decision-making process of LLM-based strategies is restructured into a hierarchy with a reasoning-policy and planning-policy. The reasoning-policy is distilled from the data that is generated through the embodied in-context learning and self-verification of an LLM, so it can produce effective rationales. The planning-policy, guided by the rationales, can render optimized plans efficiently. In turn, DeDer allows for adopting sLMs for both policies, deployed on off-the-shelf devices. Furthermore, to enhance the quality of intermediate rationales, specific to embodied tasks, we devise the embodied knowledge graph, and to generate multiple rationales timely through a single inference, we also use the contrastively prompted attention model. Our experiments with the ALFRED benchmark demonstrate that DeDer surpasses leading language planning and distillation approaches, indicating the applicability and efficiency of sLM-based embodied policies derived through DeDer.
Wonje Choi 0003, Woo Kyung Kim, Minjong Yoo, Honguk Woo
ICML4
2024 Offline Policy Learning via Skill-step Abstraction for Long-horizon Goal-Conditioned Tasks
Minjong Yoo, Honguk Woo
IJCAI3
2024 Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation
Woo Kyung Kim, Minjong Yoo, Honguk Woo
IJCAI3
2024 LLM-based Skill Diffusion for Zero-shot Policy Adaptation
abstract
Recent advances in data-driven imitation learning and offline reinforcement learning have highlighted the use of expert data for skill acquisition and the development of hierarchical policies based on these skills. However, these approaches have not significantly advanced in adapting these skills to unseen contexts, which may involve changing environmental conditions or different user requirements. In this paper, we present a novel LLM-based policy adaptation framework LDuS which leverages an LLM to guide the generation process of a skill diffusion model upon contexts specified in language, facilitating zero-shot skill-based policy adaptation to different contexts. To implement the skill diffusion model, we adapt the loss-guided diffusion with a sequential in-painting technique, where target trajectories are conditioned by masking them with past state-action sequences, thereby enabling the robust and controlled generation of skill trajectories in test-time. To have a loss function for a given context, we employ the LLM-based code generation with iterative refinement, by which the code and controlled trajectory are validated to align with the context in a closed-loop manner. Through experiments, we demonstrate the zero-shot adaptability of LDuS to various context types including different specification levels, multi-modality, and varied temporal conditions for several robotic manipulation tasks, outperforming other language-conditioned imitation and planning methods.
Woo Kyung Kim, Jooyoung Kim 0001, Honguk Woo
NeurIPS4
2024 Incremental Learning of Retrievable Skills For Efficient Continual Task Adaptation
abstract
Continual Imitation Learning (CiL) involves extracting and accumulating task knowledge from demonstrations across multiple stages and tasks to achieve a multi-task policy. With recent advancements in foundation models, there has been a growing interest in adapter-based CiL approaches, where adapters are established parameter-efficiently for tasks newly demonstrated. While these approaches isolate parameters for specific tasks and tend to mitigate catastrophic forgetting, they limit knowledge sharing among different demonstrations. We introduce IsCiL, an adapter-based CiL framework that addresses this limitation of knowledge sharing by incrementally learning shareable skills from different demonstrations, thus enabling sample-efficient task adaptation using the skills particularly in non-stationary CiL environments. In IsCiL, demonstrations are mapped into the state embedding space, where proper skills can be retrieved upon input states through prototype-based memory. These retrievable skills are incrementally learned on their corresponding adapters. Our CiL experiments with complex tasks in the Franka-Kitchen and Meta-World demonstrate the robust performance of IsCiL in both task adaptation and sample-efficiency. We also show a simple extension of IsCiL for task unlearning scenarios.
Daehee Lee 0001, Minjong Yoo, Woo Kyung Kim, Wonje Choi 0003, Honguk Woo
NeurIPS5
2024 Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following
abstract
This study presents an Exploratory Retrieval-Augmented Planning (ExRAP) framework, designed to tackle continual instruction following tasks of embodied agents in dynamic, non-stationary environments. The framework enhances Large Language Models' (LLMs) embodied reasoning capabilities by efficiently exploring the physical environment and establishing the environmental context memory, thereby effectively grounding the task planning process in time-varying environment contexts. In ExRAP, given multiple continual instruction following tasks, each instruction is decomposed into queries on the environmental context memory and task executions conditioned on the query results. To efficiently handle these multiple tasks that are performed continuously and simultaneously, we implement an exploration-integrated task planning scheme by incorporating the information-based exploration into the LLM-based planning process. Combined with memory-augmented query evaluation, this integrated scheme not only allows for a better balance between the validity of the environmental context memory and the load of environment exploration, but also improves overall task performance. Furthermore, we devise a temporal consistency refinement scheme for query evaluation to address the inherent decay of knowledge in the memory. Through experiments with VirtualHome, ALFRED, and CARLA, our approach demonstrates robustness against a variety of embodied instruction following scenarios involving different instruction scales and types, and non-stationarity degrees, and it consistently outperforms other state-of-the-art LLM-based task planning approaches in terms of both goal success rate and execution efficiency.
Minjong Yoo, Jinwoo Jang, Wei-Jin Park, Honguk Woo
NeurIPS4
2023 One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill
abstract
One-shot imitation is to learn a new task from a single demonstration, yet it is a challenging problem to adopt it for complex tasks with the high domain diversity inherent in a non-stationary environment. To tackle the problem, we explore the compositionality of complex tasks, and present a novel skill-based imitation learning framework enabling one-shot imitation and zero-shot adaptation; from a single demonstration for a complex unseen task, a semantic skill sequence is inferred and then each skill in the sequence is converted into an action sequence optimized for environmental hidden dynamics that can vary over time. Specifically, we leverage a vision-language model to learn a semantic skill set from offline video datasets, where each skill is represented on the vision-language embedding space, and adapt meta-learning with dynamics inference to enable zero-shot skill adaptation. We evaluate our framework with various one-shot imitation scenarios for extended multi-stage Meta-world tasks, showing its superiority in learning complex tasks, generalizing to dynamics changes, and extending to different demonstration conditions and modalities, compared to other baselines.
Sangwoo Shin, Daehee Lee 0001, Minjong Yoo, Woo Kyung Kim, Honguk Woo
ICML5
2023 Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
abstract
For embodied reinforcement learning (RL) agents interacting with the environment, it is desirable to have rapid policy adaptation to unseen visual observations, but achieving zero-shot adaptation capability is considered as a challenging problem in the RL context. To address the problem, we present a novel contrastive prompt ensemble (ConPE) framework which utilizes a pretrained vision-language model and a set of visual prompts, thus enables efficient policy learning and adaptation upon a wide range of environmental and physical changes encountered by embodied agents. Specifically, we devise a guided-attention-based ensemble approach with multiple visual prompts on the vision-language model to construct robust state representations. Each prompt is contrastively learned in terms of an individual domain factors that significantly affects the agent's egocentric perception and observation. For a given task, the attention-based ensemble and policy are jointly learned so that the resulting state representations not only generalize to various domains but are also optimized for learning the task. Through experiments, we show that ConPE outperforms other state-of-the-art algorithms for several embodied agent tasks including navigation in AI2THOR, manipulation in Metaworld, and autonomous driving in CARLA, while also improving the sample efficiency of policy learning and adaptation.
Wonje Choi 0003, Woo Kyung Kim, Honguk Woo
NeurIPS4
2022 An Efficient Combinatorial Optimization Model Using Learning-to-Rank Distillation
abstract
Recently, deep reinforcement learning (RL) has proven its feasibility in solving combinatorial optimization problems (COPs). The learning-to-rank techniques have been studied in the field of information retrieval. While several COPs can be formulated as the prioritization of input items, as is common in the information retrieval, it has not been fully explored how the learning-to-rank techniques can be incorporated into deep RL for COPs. In this paper, we present the learning-to-rank distillation-based COP framework, where a high-performance ranking policy obtained by RL for a COP can be distilled into a non-iterative, simple model, thereby achieving a low-latency COP solver. Specifically, we employ the approximated ranking distillation to render a score-based ranking model learnable via gradient descent. Furthermore, we use the efficient sequence sampling to improve the inference performance with a limited delay. With the framework, we demonstrate that a distilled model not only achieves comparable performance to its respective, high-performance RL, but also provides several times faster inferences. We evaluate the framework with several COPs such as priority-based task scheduling and multidimensional knapsack, demonstrating the benefits of the framework in terms of inference latency and performance.
Honguk Woo, Hyunsung Lee, Sangwoo Cho
AAAI1
2022 Structure Learning-Based Task Decomposition for Reinforcement Learning in Non-stationary Environments
abstract
Reinforcement learning (RL) agents empowered by deep neural networks have been considered a feasible solution to automate control functions in a cyber-physical system. In this work, we consider an RL-based agent and address the issue of learning via continual interaction with a time-varying dynamic system modeled as a non-stationary Markov decision process (MDP). We view such a non-stationary MDP as a time series of conventional MDPs that can be parameterized by hidden variables. To infer the hidden parameters, we present a task decomposition method that exploits CycleGAN-based structure learning. This method enables the separation of time-variant tasks from a non-stationary MDP, establishing the task decomposition embedding specific to time-varying information. To mitigate the adverse effect due to inherent noises of task embedding, we also leverage continual learning on sequential tasks by adapting the orthogonal gradient descent scheme with a sliding window. Through various experiments, we demonstrate that our approach renders the RL agent adaptable to time-varying dynamic environment conditions, outperforming other methods including state-of-the-art non-stationary MDP algorithms.
Honguk Woo, Gwangpyo Yoo, Minjong Yoo
AAAI1
2022 Skills Regularized Task Decomposition for Multi-task Offline Reinforcement Learning
abstract
Reinforcement learning (RL) with diverse offline datasets can have the advantage of leveraging the relation of multiple tasks and the common skills learned across those tasks, hence allowing us to deal with real-world complex problems efficiently in a data-driven way. In offline RL where only offline data is used and online interaction with the environment is restricted, it is yet difficult to achieve the optimal policy for multiple tasks, especially when the data quality varies for the tasks. In this paper, we present a skill-based multi-task RL technique on heterogeneous datasets that are generated by behavior policies of different quality. To learn the shareable knowledge across those datasets effectively, we employ a task decomposition method for which common skills are jointly learned and used as guidance to reformulate a task in shared and achievable subtasks. In this joint learning, we use Wasserstein Auto-Encoder (WAE) to represent both skills and tasks on the same latent space and use the quality-weighted loss as a regularization term to induce tasks to be decomposed into subtasks that are more consistent with high-quality skills than others. To improve the performance of offline RL agents learned on the latent space, we also augment datasets with imaginary trajectories relevant to high-quality skills for each task. Through experiments, we show that our multi-task offline RL approach is robust to different-quality datasets and it outperforms other state-of-the-art algorithms for several robotic manipulation tasks and drone navigation tasks.
Minjong Yoo, Sangwoo Cho, Honguk Woo
NeurIPS3
2021 GPU-Ether: GPU-native Packet I/O for GPU Applications on Commodity Ethernet
abstract
Despite the advent of various network enhancement technologies, it is yet a challenge to provide high-performance networking for GPU-accelerated applications on commodity Ethernet. Kernel-bypass I/O, such as DPDK or netmap, which is normally optimized for host memory-based CPU applications, has limitations on improving the performance of GPU-accelerated applications due to the data transfer overhead between host and GPU. In this paper, we propose GPU-Ether, GPU-native packet I/O on commodity Ethernet, which enables direct network access from GPU via dedicated persistent kernel threads. We implement GPU-Ether prototype on a commodity Ethernet NIC and perform extensive testing to evaluate it. The results show that GPU-Ether can provide high throughput and low latency for GPU applications.
Changue Jung, Suhwan Kim 0008, Ikjun Yeom, Honguk Woo
INFOCOM4
2021 ML for RT: Priority Assignment Using Machine Learning
abstract
As machine learning (ML) has been proven effective in solving various problems, researchers in the real-time systems (RT) community have recently paid increasing attention to ML. While most of them focused on timing issues for ML applications (i.e., RT for ML), only a little has been done on the use of ML for solving fundamental RT problems. In this paper, we aim at utilizing ML to solve a fundamental RT problem of priority assignment for global fixed-priority preemptive (gFP) scheduling on a multiprocessor platform. This problem is known to be challenging in the case of a large number (n) of tasks in a task set because exhaustive testing of all priority assignments (as many as n!) is intractable and existing heuristics cannot find a schedulable priority assignment, even if exists, for a number of task sets. We systematically incorporate RT domain knowledge into ML and develop an ML framework tailored to the problem, called PAL. First, raising and addressing technical issues including neural architecture selection and training sample regulation, we enable PAL to infer a schedulable priority assignment of a set of n tasks, by training PAL with same-size (i.e., with n tasks) samples each of whose schedulable priority assignment has already been identified. Second, considering the exhaustive testing of all priority assignments of each task set with large n makes it intractable to provide training samples to PAL, we derive inductive properties that can generate training samples with large n from those with small n, through empirical observation of PAL and mathematical analysis of the target gFP schedulability test. Finally, utilizing the inductive properties and additional techniques, we propose how to systematically implement PAL whose training sample generation process not only yields unbiased samples but also is tractable even for large n. Our experimental results demonstrate PAL covers a number of additional task sets, each of which has never been proven schedulable by any existing approaches for gFP.
Seunghoon Lee 0002, Hyeongboo Baek, Honguk Woo, Kang G. Shin, Jinkyu Lee 0001
RTAS3
2017 A Revisit to Web Browsing on Wearable Devices
Jinwoo Song, Hyunjune Kim, Honguk Woo
WEBIST4
2016 NDN-Based Pub/Sub System for Scalable IoT Cloud
abstract
The Internet of Things (IoT) has recently become mainstream, different from the incomplete realization of ubiquitous computing, sensor networks, and others in past that commonly shared the notion of visionary hyper-connectivity, i.e., everything is connected. Among many, one significant reason for this achievement of IoT is the nowadays availability of cloud computing which can cost effectively deal with the sheer number of globally distributed devices and data generated from those devices. In this paper, we address the issue of IoT scalability in cloud environments by proposing the unique integration scheme between pub/sub systems and Named Data Networking (NDN). While the pub/sub architecture has been applied in modern IoT cloud platforms, its implementation and deployment practice have not been fully studied for large scale IoT applications. Our approach concentrates on how to leverage the essential scalability of NDN for building pub/sub systems, thereby achieving scalable IoT cloud services.
Sungwon Han 0002, Honguk Woo
CloudCom2
2014 The resource optimization of heterogeneous network interfaces in wireless mobile devices
abstract
Although using high-quality streaming services becomes more common on mobile devices, the increment of streaming quality, which generally requires the larger playback bandwidth, frequently incurs undesirable delays or jitters that hamper user experiences. Recent mobile devices usually equip multiple network interfaces and their concurrent utilization can be beneficial to retain the necessary bandwidth. This, however, requires the more energy consumption of mobile devices whose resource is limited due to battery-based operations. For this reason, streaming performance and energy consumption should not be independently considered. Under multiple network interface utilization, we investigate the characteristics of each component, based on two-level convex optimization, and propose a solution to optimize them together. The algorithm is tested on Android phone where LTE and Wi-Fi networks are concurrently enabled. Compared to LTE only method, it shows that the proposed approach can obtain 1.8 times good-put increment only with the half amount of energy consumption. Our approach can be utilized for almost all mobile devices without server-side modifications.
Daehyun Ban, Jeongsik In, Honguk Woo
ICC3
2014 Buffering in proxy mobile IPv6: implementation and analysis
Changyong Park, NamYeong Kwon, Honguk Woo, Hyunseung Choo
J. Supercomput.3
2013 GreenBag: Energy-Efficient Bandwidth Aggregation for Real-Time Streaming in Heterogeneous Mobile Wireless Networks
abstract
Modern mobile devices are equipped with multiple network interfaces, including 3G/LTE and WiFi. Bandwidth aggregation over LTE and WiFi links offers an attractive opportunity of supporting bandwidth-intensive services, such as high-quality video streaming, on mobile devices. However, achieving effective bandwidth aggregation in mobile environments raises several challenges related to deployment, link heterogeneity, network fluctuation, and energy consumption. We present GreenBag, an energy-efficient bandwidth aggregation middleware that supports real-time data-streaming services over asymmetric wireless links, requiring no modifications to the existing Internet infrastructure and servers. GreenBag employs several techniques, including medium load balancing, efficient segment management, and energy-aware mode control, to resolve such challenges. We implement a prototype of GreenBag on Android-based mobile devices which hosts, to the best knowledge of the authors, the first LTE-enabled bandwidth aggregation prototype for energy-efficient real-time video streaming. Our experiment results in both emulated and real-world environments show that GreenBag not only achieves good bandwidth aggregation to provide QoS in bandwidth-scarce environments but also efficiently saves energy on mobile devices. Moreover, energy-aware GreenBag can minimize video interruption while consuming 14-25% less energy than the non-energy-aware counterpart in real-world experiments.
Duc Hoang Bui, Kilho Lee, Sangeun Oh, Insik Shin, Hyojeong Shin, Honguk Woo, Daehyun Ban
RTSS6
2013 Enhanced cooperative communication MAC for mobile wireless networks
Donghyeok An, Honguk Woo, Hyunsoo Yoon, Ikjun Yeom
Comput. Networks2
2008 Incorporating Resource Safety Verification to Executable Model-based Development for Embedded Systems
abstract
This paper formulates and illustrates the integration of resource safety verification into a design methodology for development of verified and robust real-time embedded systems. Resource-related concerns are not closely linked with current xUML model-based software development although they are critical for embedded systems. We describe how to integrate resource analysis techniques into the early phase of an xUML-based development cycle. Our hybrid framework for resource safety verification combines static resource analysis and runtime monitoring. A case study based on an embedded controller for satellite simulation, TableSat, illustrates the benefits obtained by incorporating resource verification into design and combining static analysis and runtime monitoring.
Jianliang Yi, Honguk Woo, James C. Browne, Aloysius K. Mok, Ella M. Atkins, Chan-Gun Lee
IEEE Real-Time and Embedded Technology and Applications Symposium2
2007 Real-Time Monitoring of Uncertain Data Streams Using Probabilistic Similarity
abstract
Data uncertainty is a common problem for the real-time monitoring of data streams. In this paper, we address the issue of efficiently monitoring the satisfaction/violation of user-defined constraints over data streams where the data uncertainty can be probabilistically characterized. We propose a monitoring architecture SPMON that can incorporate probabilistic models of uncertainty in constraint monitoring. We adapt the concept of data similarity in real-time databases to the processing of uncertain data streams. In doing so, we generalize the data similarity by a new concept psr (probabilistic similarity region) that allows us to define similarity relations for probabilistic data with respect to the set of constraints being monitored. This enables the construction of lightweight filters for saving bandwidth. We also show how to efficiently update the filter conditions at run-time.
Honguk Woo, Aloysius K. Mok
RTSS1
2006 Probabilistic Timing Join over Uncertain Event Streams
abstract
This paper addresses the problem of processing eventtiming queries over event streams where the uncertainty in the values of the timestamps is characterizable by histograms. We describe a stream-partitioning technique for checking the satisfaction of a probabilistic timing constraint upon event arrivals in a systematic way in order to delimit the "probing range" in event streams. This technique can be formalized as a probabilistic timing join (PTJoin) operator where the join condition is specified by a time window and a confidence threshold in our model. We present efficient PTJoin algorithms that tightly delimit the probing range and efficiently invalidate events in event streams.
Aloysius K. Mok, Honguk Woo, Chan-Gun Lee
RTCSA2
2006 A Generic Framework for Monitoring Timing Constraints over Uncertain Events
abstract
This paper provides a comprehensive approach to the problem of monitoring timing constraints over event streams for which the timestamp values are inherently uncertain. We first propose a generic framework for capturing the early detection of the violation of timing constraints, based on the notion of probabilistic violation time. In doing so, we provide a systemic approach for deriving a set of necessary constraints at compilation time. Our work is innovative in that the framework is formulated to be "modular" with respect to the probability distributions on timestamp values. We demonstrate the applicability of the framework for two different timestamp models, Gaussian and histogram. The Gaussian model is appropriate for representing event timing from a wide variety of sensors with well-modelled physical noise characteristics; we show how we can efficiently derive the probabilistic violation time of timing constraints by exploiting the relation between the Gaussian distribution parameters. The histogram model can be used where the timestamps of events are available from measurements only as arbitrary probability distributions: we show how to derive an efficient timing constraint monitoring method for the histogram model
Honguk Woo, Aloysius K. Mok, Chan-Gun Lee
RTSS1
2004 Specifying Timing Constraints and Composite Events: An Application in the Design of Electronic Brokerages
abstract
Increasingly, business applications need to capture consumers' complex preferences interactively and monitor those preferences by translating them into event-condition-action (ECA) rules and syntactically correct processing specification. An expressive event model to specify primitive and composite events that may involve timing constraints among events is critical to such applications. Relying on the work done in active databases and real-time systems, this research proposes a new composite event model based on real-time logic (RTL). The proposed event model does not require fixed event consumption policies and allows the users to represent the exact correlation of event instances in defining composite events. It also supports a wide-range of domain-specific temporal events and constraints, such as future events, time-constrained events, and relative events. This event model is validated within an electronic brokerage architecture that unbundles the required functionalities into three separable components - business rule manager, ECA rule manager, and event monitor - with well-defined interfaces. A proof-of-concept prototype was implemented in the Java programming language to demonstrate the expressiveness of the event model and the feasibility of the architecture. The performance of the composite event monitor was evaluated by varying the number of rules, event arrival rates, and type of composite events.
Aloysius K. Mok, Prabhudev Konana, Guangtian Liu, Chan-Gun Lee, Honguk Woo
IEEE Trans. Software Eng.5
2002 The Monitoring of Timing Constraints on Time Intervals
abstract
Efficient algorithms have been developed by a number of authors to detect constraint violation or satisfaction of timed events. In extant work, the time of every event occurrence is assumed to be known exactly. However there are practical situations where we are not sure about the exact time of occurrence of an event but we may be able to capture the uncertainty by a time interval. In this paper we propose new types of timing constraints: possible and certain constraints that are pertinent to an event model where timestamps are given by time intervals. We extend previous work in timed event monitoring that is time-point based to our interval-based model. We give an efficient algorithm for monitoring timing constraints under event timing uncertainty, and sketch its proof of correctness by extending the pruning algorithm on the constraint graph to cover interval timestamps.
Aloysius K. Mok, Chan-Gun Lee, Honguk Woo, Prabhudev Konana
RTSS3
2002 DBCache: database caching for web application servers
abstract
Many e-Business applications today are being developed and deployed on multi-tier environments involving browser-based clients, web application servers and backend databases. The dynamic nature of these applications necessitates generating web pages on-demand, making middle-tier database caching an effective approach to achieve high scalability and performance [3]. In the DBCache project, we are incorporating a database cache feature in DB2 UDB by modifying the engine code and leveraging existing federated database functionality. This allows us to take advantage of DB2's sophisticated distributed query processing power for database caching. As a result, the user queries can be executed at either the local database cache or the remote backend server, or more importantly, the query can be partitioned and then distributed to both databases for cost optimum execution.DBCache also includes a cache initialization component that takes a backend database schema and SQL queries in the workload, and generates a middle-tier database schema for the cache. We have implemented an initial prototype of the system that supports table level caching. As DB2's functionality is extended, we will be able to support subtable level caching, XML data caching and caching of execution results of web services.
Mehmet Altinel, Qiong Luo 0001, Sailesh Krishnamurthy, C. Mohan 0001, Hamid Pirahesh, Bruce G. Lindsay 0001, Honguk Woo, Larry Brown
SIGMOD Conference7
2002 Middle-tier database caching for e-business
abstract
While scaling up to the enormous and growing Internet population with unpredictable usage patterns, E-commerce applications face severe challenges in cost and manageability, especially for database servers that are deployed as those applications' backends in a multi-tier configuration. Middle-tier database caching is one solution to this problem. In this paper, we present a simple extension to the existing federated features in DB2 UDB, which enables a regular DB2 instance to become a DBCache without any application modification. On deployment of a DBCache at an application server, arbitrary SQL statements generated from the unchanged application that are intended for a backend database server, can be answered: at the cache, at the backend database server, or at both locations in a distributed manner. The factors that determine the distribution of workload include the SQL statement type, the cache content, the application requirement on data freshness, and cost-based optimization at the cache. We have developed a research prototype of DBCache, and conducted an extensive set of experiments with an E-Commerce benchmark to show the benefits of this approach and illustrate tradeoffs in caching considerations.
Qiong Luo 0001, Sailesh Krishnamurthy, C. Mohan 0001, Hamid Pirahesh, Honguk Woo, Bruce G. Lindsay 0001, Jeffrey F. Naughton
SIGMOD Conference5
2000 Implementation and Performance Evaluation of a Real-Time E-Brokerage System
abstract
Timeliness is an important attribute for e-brokerage both in the detection of opportunities in a narrow time window and also in facilitating differentiation of end-to-end quality of service (QoS). In this paper, we demonstrate how real-time event monitoring techniques can be applied to time-critical e-brokerage. We start with a formal timed event model which provides the semantics for specifying complex timing correlation rules in composite events. The formal event model of the system is based on RTL (Real Time Logic), which is important for disambiguating informal specifications when multiple instances of the same event type may appear in a timing correlation rule involving composite events. This leads to our design of an e-brokerage, online stock monitoring and alerting system which enables users to express their complex preferences and provides an alert service to users in a timely manner. The user's preferences are monitored by a real-time event monitor. We report the design of this system and some performance data characterizing some key timing parameters. Our system is being used by a class of MBA students in a field test.
Prabhudev Konana, Aloysius K. Mok, Chan-Gun Lee, Honguk Woo, Guangtian Liu
RTSS4