VLDB 2026 Research / reviewers in the wild / expert
Kimin Lee
dblp:183/6849
· DBLP profile ↗
59ranked-venue papers
9as first author
43since 2021 · last 2026
0000-0001-9017-3084ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 8 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unintended Misalignment from Agentic Fine-Tuning: Risks and MitigationabstractBeyond simple text generation, Large Language Models (LLMs) have evolved into agentic systems capable of planning and interacting with external tools to solve complex tasks. This evolution involves fine-tuning LLMs on agent-specific tasks to enhance their proficiency. However, safety concerns are frequently overlooked during this fine-tuning process. In this work, we show that aligned LLMs can become unintentionally misaligned, leading to a higher likelihood of executing harmful tasks and a reduced tendency to refuse them when fine-tuned to execute agentic tasks. To address these safety challenges, we propose Prefix INjection Guard (PING), a simple yet effective method that prepends automatically generated natural language prefixes to agent responses, guiding them to refuse harmful requests while preserving performance on benign tasks. Specifically, we introduce an iterative approach that alternates between (1) generating candidate prefixes and (2) selecting those that optimize both task performance and refusal behavior. Experimental results demonstrate that PING significantly enhances the safety of fine-tuned LLM agents without sacrificing their effectiveness. PING consistently outperforms existing prompting approaches across diverse benchmarks in both web navigation and code generation tasks. Our analysis of internal hidden states via linear probes reveals that prefix tokens are crucial for behavior modification, explaining the performance gains. Dongyoon Hahm, Taywon Min, Woogyeol Jin, Kimin Lee |
AAAI | 4 |
| 2026 | MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device ControlabstractAutonomous agents powered by large language models (LLMs) show promising potential in assistive tasks across various domains, including mobile device control. As these agents interact directly with personal information and device settings, ensuring their safe and reliable behavior is crucial to prevent undesirable outcomes. However, no benchmark exists for standardized evaluation of the safety of mobile device-control agents. In this work, we introduce MobileSafetyBench, a benchmark designed to evaluate the safety of device-control agents within a realistic mobile environment based on Android emulators. We develop a diverse set of tasks involving interactions with various mobile applications, including messaging and banking applications, challenging agents with managing risks encompassing the misuse and negative side effects. These tasks include tests to evaluate the safety of agents in daily scenarios as well as their robustness against indirect prompt injection attacks. Our experiments demonstrate that baseline agents, based on state-of-the-art LLMs, often fail to effectively prevent harm while performing the tasks. To mitigate these safety concerns, we propose a prompting method that encourages agents to prioritize safety considerations. While this method shows promise in promoting safer behaviors, there is still considerable room for improvement to fully earn user trust. This highlights the urgent need for continued research to develop more robust safety mechanisms in mobile environments. Dongyoon Hahm, June Suk Choi, W. Bradley Knox, Kimin Lee |
AAAI | 5 |
| 2026 | State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital LivingabstractWhen working on digital devices, people often face distractions that can lead to a decline in productivity and efficiency, as well as negative psychological and emotional impacts. To address this challenge, we introduce a novel Artificial Intelligence (AI) assistant that elicits a user’s intention, assesses whether ongoing activities are in line with that intention, and provides gentle nudges when deviations occur. The system leverages a large language model to analyze screenshots, application titles, and URLs, issuing notifications when behavior diverges from the stated goal. Its detection accuracy is refined through initial clarification dialogues and continuous user feedback. In a three-week, within-subjects field deployment with 22 participants, we compared our assistant to both a rule-based intent reminder system and a passive baseline that only logged activity. Results indicate that our AI assistant effectively supports users in maintaining focus and aligning their digital behavior with their intentions. Our source code is publicly available at https://intentassistant.github.io Juheon Choi, Jian Kim, Taywon Min, W. Bradley Knox, Min Kyung Lee, Kimin Lee |
CHI | 8 |
| 2025 | DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion ModelsabstractFine-tuning text-to-image diffusion models to maximize rewards has proven effective for enhancing model performance. However, reward fine-tuning methods often suffer from slow convergence due to online sample generation. Therefore, obtaining diverse samples with strong reward signals is crucial for improving sample efficiency and overall performance. In this work, we introduce DiffExp, a simple yet effective exploration strategy for reward fine-tuning of text-to-image models. Our approach employs two key strategies: (a) dynamically adjusting the scale of classifier-free guidance to enhance sample diversity, and (b) randomly weighting phrases of the text prompt to exploit high-quality reward signals. We demonstrate that these strategies significantly enhance exploration during online sample generation, improving the sample efficiency of recent reward fine-tuning methods, such as DDPO and AlignProp. Daewon Chae, June Suk Choi, Kimin Lee |
AAAI | 4 |
| 2025 | What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMsabstractWe investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ).Our experiments utilize context length of up to 128K tokens.Through comprehensive analysis with various many-shot attack settings with different instruction styles, shot density, topic, and format, we reveal that context length is the primary factor determining attack effectiveness.Critically, we find that successful attacks do not require carefully crafted harmful content.Even repetitive shots or random dummy text can circumvent model safety measures, suggesting fundamental limitations in long-context processing capabilities of LLMs.The safety behavior of well-aligned models becomes increasingly inconsistent with longer contexts.These findings highlight significant safety gaps in context expansion capabilities of LLMs, emphasizing the need for new safety mechanisms. Sangyeop Kim 0001, Yongwoo Song, Kimin Lee |
ACL (1) | 4 |
| 2025 | Understanding Impact of Human Feedback via Influence FunctionsabstractIn Reinforcement Learning from Human Feedback (RLHF), it is crucial to learn suitable reward models from human feedback to align large language models (LLMs) with human intentions.However, human feedback can often be noisy, inconsistent, or biased, especially when evaluating complex responses.Such feedback can lead to misaligned reward signals, potentially causing unintended side effects during the RLHF process.To address these challenges, we explore the use of influence functions to measure the impact of human feedback on the performance of reward models.We propose a compute-efficient approximation method that enables the application of influence functions to LLM-based reward models and large-scale preference datasets.Our experiments showcase two key applications of influence functions: (1) detecting common labeler biases in human feedback datasets and (2) guiding labelers in refining their strategies to better align with expert feedback.By quantifying the impact of human feedback, we believe that influence functions can enhance feedback interpretability and contribute to scalable oversight in RLHF, helping labelers provide more accurate and consistent feedback.Source code is available at https://github.com/mintaywon/IF_RLHF. Taywon Min, Haeone Lee, Yongchan Kwon, Kimin Lee |
ACL (1) | 4 |
| 2025 | Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsabstractText-to-image diffusion models have achieved remarkable success in generating high-quality contents from text prompts. However, their reliance on publicly available data and the growing trend of data sharing for fine-tuning make these models particularly vulnerable to data poisoning attacks. In this work, we introduce the Silent Branding Attack, a novel data poisoning method that manipulates text-to-image diffusion models to generate images containing specific brand logos or symbols without any text triggers. We find that when certain visual patterns are repeatedly in the training data, the model learns to reproduce them naturally in its outputs, even without prompt mentions. Leveraging this, we develop an automated data poisoning algorithm that unobtrusively injects logos into original images, ensuring they blend naturally and remain undetected. Models trained on this poisoned dataset generate images containing logos without degrading image quality or text alignment. We experimentally validate our silent branding attack across two realistic settings on large-scale high-quality image datasets and style personalization datasets, achieving high success rates even without a specific text trigger. Human evaluation and quantitative metrics including logo detection show that our method can stealthily embed logos. Our project page is at https://silent-branding.github.io/. Sangwon Jang, June Suk Choi, Jaehyeong Jo, Kimin Lee, Sung Ju Hwang |
CVPR | 4 |
| 2025 | Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level TokenizersabstractTokenization is a crucial step that bridges human-readable text with model-readable discrete tokens.However, recent studies have revealed that tokenizers can be exploited to elicit unwanted model behaviors.In this work, we investigate incomplete tokens, i.e., undecodable tokens with stray bytes resulting from bytelevel byte-pair encoding (BPE) tokenization.We hypothesize that such tokens are heavily reliant on their adjacent tokens and are fragile when paired with unfamiliar tokens.To demonstrate this vulnerability, we introduce improbable bigrams: out-of-distribution combinations of incomplete tokens designed to exploit their dependency.Our experiments show that improbable bigrams are significantly prone to hallucinatory behaviors.Surprisingly, the same phrases have drastically lower rates of hallucination (90% reduction in Llama3.1)when an alternative tokenization is used.We caution against the potential vulnerabilities introduced by byte-level BPE tokenizers, which may introduce blind spots to language models. Eugene Jang, Kimin Lee, Jin-Woo Chung, Keuntae Park, Seungwon Shin 0001 |
EMNLP | 2 |
| 2025 | DiffusionGuard: A Robust Defense Against Malicious Diffusion-based Image EditingabstractRecent advances in diffusion models have introduced a new era of text-guided image manipulation, enabling users to create realistic edited images with simple textual prompts. However, there is significant concern about the potential misuse of these methods, especially in creating misleading or harmful content. Although recent defense strategies, which introduce imperceptible adversarial noise to induce model failure, have shown promise, they remain ineffective against more sophisticated manipulations, such as editing with a mask. In this work, we propose DiffusionGuard, a robust and effective defense method against unauthorized edits by diffusion-based image editing models, even in challenging setups. Through a detailed analysis of these models, we introduce a novel objective that generates adversarial noise targeting the early stage of the diffusion process. This approach significantly improves the efficiency and effectiveness of adversarial noises. We also introduce a mask-augmentation technique to enhance robustness against various masks during test time. Finally, we introduce a comprehensive benchmark designed to evaluate the effectiveness and robustness of methods in protecting against privacy threats in realistic scenarios. Through extensive experiments, we show that our method achieves stronger protection and improved mask robustness with lower computational costs compared to the strongest baseline. Additionally, our method exhibits superior transferability and better resilience to noise removal techniques compared to all baseline methods. Our source code is publicly available at https://choi403.github.io/diffusionguard. June Suk Choi, Kyungmin Lee, Jongheon Jeong, Saining Xie, Jinwoo Shin, Kimin Lee |
ICLR | 6 |
| 2025 | Subtask-Aware Visual Reward Learning from Segmented DemonstrationsabstractReinforcement Learning (RL) agents have demonstrated their potential across various robotic tasks. However, they still heavily rely on human-engineered reward functions, requiring extensive trial-and-error and access to target behavior information, often unavailable in real-world settings. This paper introduces REDS: REward learning from Demonstration with Segmentations, a novel reward learning framework that leverages action-free videos with minimal supervision. Specifically, REDS employs video demonstrations segmented into subtasks from diverse sources and treats these segments as ground-truth rewards. We train a dense reward function conditioned on video segments and their corresponding subtasks to ensure alignment with ground-truth reward signals by minimizing the Equivalent-Policy Invariant Comparison distance. Additionally, we employ contrastive learning objectives to align video representations with subtasks, ensuring precise subtask inference during online interactions. Our experiments show that REDS significantly outperforms baseline methods on complex robotic manipulation tasks in Meta-World and more challenging real-world tasks, such as furniture assembly in FurnitureBench, with minimal human intervention. Moreover, REDS facilitates generalization to unseen tasks and robot embodiments, highlighting its potential for scalable deployment in diverse environments. Changyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee, Jinwoo Shin, Joseph J. Lim, Kimin Lee |
ICLR | 7 |
| 2025 | Spread Preference Annotation: Direct Preference Judgment for Efficient LLM AlignmentabstractAligning large language models (LLMs) with human preferences becomes a key component to obtaining state-of-the-art performance, but it yields a huge cost to construct a large human-annotated preference dataset. To tackle this problem, we propose a new framework, Spread Preference Annotation with direct preference judgment (SPA), that boosts the alignment of LLMs using only a very small amount of human-annotated preference data.
Our key idea is leveraging the human prior knowledge within the small (seed) data and progressively improving the alignment of LLM, by iteratively generating the responses and learning from them with the self-annotated preference data.
To be specific, we propose to derive the preference label from the logits of LLM to explicitly extract the model's inherent preference.
Compared to the previous approaches using external reward models or implicit in-context learning, we observe that the proposed approach is significantly more effective.
In addition, we introduce a noise-aware preference learning algorithm to mitigate the risk of low quality within generated preference data.
Our experimental results demonstrate that the proposed framework significantly boosts the alignment of LLMs.
For example, we achieve superior alignment performance on AlpacaEval 2.0 with only 3.3% of the ground-truth preference labels in the Ultrafeedback data compared to the cases using the entire data or state-of-the-art baselines. Dongyoung Kim, Kimin Lee, Jinwoo Shin, Jaehyung Kim 0001 |
ICLR | 2 |
| 2025 | Learning to Contextualize Web Pages for Enhanced Decision Making by LLM AgentsabstractRecent advances in large language models (LLMs) have led to a growing interest in developing LLM-based agents for automating web tasks. However, these agents often struggle with even simple tasks on real-world websites due to their limited capability to understand and process complex web page structures. In this work, we introduce LCoW, a framework for Learning language models to Contextualize complex Web pages into a more comprehensible form, thereby enhancing decision making by LLM agents. LCoW decouples web page understanding from decision making by training a separate contextualization module to transform complex web pages into comprehensible format, which are then utilized by the decision-making agent. We demonstrate that our contextualization module effectively integrates with LLM agents of various scales to significantly enhance their decision-making capabilities in web automation tasks. Notably, LCoW improves the success rates of closed-source LLMs (e.g., Gemini-1.5-flash, GPT-4o, Claude-3.5-Sonnet) by an average of 15.6%, and demonstrates a 23.7% average improvement in success rates for open-source LMs (e.g., Llama-3.1-8B, Llama-3.1-70B) on the WorkArena benchmark.
Moreover, the Gemini-1.5-flash agent with LCoW achieves state-of-the-art results on the WebShop benchmark, outperforming human experts. The relevant code materials are available at our project page: https://lcowiclr2025.github.io. Kyuyoung Kim, Jihoon Tack, Jinwoo Shin, Yee Whye Teh, Kimin Lee |
ICLR | 7 |
| 2025 | Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion ModelsabstractFine-tuning text-to-image diffusion models with human feedback is an effective method for aligning model behavior with human intentions. However, this alignment process often suffers from slow convergence due to the large size and noise present in human feedback datasets. In this work, we propose FiFA, a novel automated data filtering algorithm designed to enhance the fine-tuning of diffusion models using human feedback datasets with direct preference optimization (DPO). Specifically, our approach selects data by solving an optimization problem to maximize three components: preference margin, text quality, and text diversity. The concept of preference margin is used to identify samples that are highly informative in addressing the noisy nature of feedback dataset, which is calculated using a proxy reward model. Additionally, we incorporate text quality, assessed by large language models to prevent harmful contents, and consider text diversity through a k-nearest neighbor entropy estimator to improve generalization. Finally, we integrate all these components into an optimization process, with approximating the solution by assigning importance score to each data pair and selecting the most important ones. As a result, our method efficiently filters data automatically, without the need for manual intervention, and can be applied to any large-scale dataset. Experimental results show that FiFA significantly enhances training stability and achieves better performance, being preferred by humans 17% more, while using less than 0.5% of the full data and thus 1% of the GPU hours compared to utilizing full human feedback datasets. Yongjin Yang, Sihyeon Kim, Hojung Jung, Sangmin Bae, Se-Young Yun, Kimin Lee |
ICLR | 7 |
| 2025 | Latent Action Pretraining from VideosabstractWe introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators during pretraining, which significantly limits possible data sources and scale. In this work, we propose a method to learn from internet-scale videos that do not have robot action labels. We first train an action quantization model leveraging VQ-VAE-based objective to learn discrete latent actions between image frames, then pretrain a latent VLA model to predict these latent actions from observations and task descriptions, and finally finetune the VLA on small-scale robot manipulation data to map from latent to robot actions. Experimental results demonstrate that our method significantly outperforms existing techniques that train robot manipulation policies from large-scale videos. Furthermore, it outperforms the state-of-the-art VLA model trained with robotic action labels on real-world manipulation tasks that require language conditioning, generalization to unseen objects, and semantic generalization to unseen instructions. Training only on human manipulation videos also shows positive transfer, opening up the potential for leveraging web-scale data for robotics foundation models. Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Y. Lin, Lars Liden, Kimin Lee, Jianfeng Gao 0001, Luke Zettlemoyer, Dieter Fox, Minjoon Seo |
ICLR | 12 |
| 2025 | SelfReplay: Adapting Self-Supervised Sensory Models via Adaptive Meta-Task ReplayabstractSelf-supervised learning enables effective model pre-training on large-scale unlabeled data, which is crucial for user-specific fine-tuning in mobile sensing applications. However, pre-trained models often face significant domain shifts during fine-tuning due to user diversity, leading to performance degradation. To address this, we propose SelfReplay, an adaptive approach designed to align self-supervised models to different domains. SelfReplay consists of two stages: MetaSSL, which leverages meta-learning with self-supervised learning to pre-train domain-adaptive weights, and ReplaySSL, which further adapts the pre-trained model to each user's domain by replaying the meta-learned self-supervised task with a few user-specific samples. This produces a personalized model tailored to each user. Evaluations on mobile sensing benchmarks demonstrate that SelfReplay outperforms existing baselines, improving the F1-score by 9.4%p on average. On-device analyses on a commodity smartphone show the efficiency of SelfReplay's adaptation step, required just once after deployment, with SimCLR completing in only 10 seconds while using less than 100MB of memory. Hyungjun Yoon, Jae Hyun Kwak, Biniyam Aschalew Tolera, Gaole Dai, Mo Li 0001, Taesik Gong, Kimin Lee, Sung-Ju Lee 0001 |
SenSys | 7 |
| 2025 | When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
Hanna Kim, Minkyoo Song, Seung Ho Na, Seungwon Shin 0001, Kimin Lee |
USENIX Security Symposium | 5 |
| 2024 | Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image GenerationabstractHuman feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate reward function has not been conclusively established. This paper investigates the effectiveness of fine-grained feedback which captures nuanced distinctions in image quality and prompt-alignment, compared to traditional coarse-grained feedback (for example, thumbs up/down or ranking between a set of options). While fine-grained feedback holds promise, particularly for systems catering to diverse societal preferences, we show that demonstrating its superiority to coarse-grained feedback is not automatic. Through experiments on real and synthetic preference data, we surface the complexities of building effective models due to the interplay of model choice, feedback type, and the alignment between human judgment and computational interpretation. We identify key challenges in eliciting and utilizing fine-grained feedback, prompting a reassessment of its assumed benefits and practicality. Our findings -- e.g., that fine-grained feedback can lead to worse models for a fixed budget, in some settings; however, in controlled settings with known attributes, fine grained rewards can indeed be more helpful -- call for careful consideration of feedback attributes and potentially beckon novel modeling approaches to appropriately unlock the potential value of fine-grained feedback in-the-wild. Katie Collins, Najoung Kim, Yonatan Bitton, Verena Rieser, Shayegan Omidshafiei, Yushi Hu, Sherol Chen, Senjuti Dutta, Minsuk Chang, Kimin Lee, Youwei Liang, Georgina Evans, Sahil Singla 0005, Gang Li 0021, Adrian Weller, Junfeng He, Deepak Ramachandran, Krishnamurthy Dvijotham |
AIES (1) | 10 |
| 2024 | Promptable Behaviors: Personalizing Multi-Objective Rewards from Human PreferencesabstractCustomizing robotic behaviors to be aligned with di-verse human preferences is an underexplored challenge in the field of embodied AI. In this paper, we present Prompt-able Behaviors, a novel framework that facilitates efficient personalization of robotic agents to diverse human prefer-ences in complex environments. We use multi-objective re-inforcement learning to train a single policy adaptable to a broad spectrum of preferences. We introduce three distinct methods to infer human preferences by leveraging different types of interactions: (1) human demonstrations, (2) prefer-ence feedback on trajectory comparisons, and (3) language instructions. We evaluate the proposed method in person-alized object-goal navigation and flee navigation tasks in ProcTHOR [18] and RoboTHOR [17], demonstrating the ability to prompt agent behaviors to satisfy human prefer-ences in various scenarios. Project page: https://promptable-behaviors.github.io Minyoung Hwang, Luca Weihs, Chanwoo Park, Kimin Lee, Aniruddha Kembhavi, Kiana Ehsani |
CVPR | 4 |
| 2024 | By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual PromptingabstractLarge language models (LLMs) have demonstrated exceptional abilities across various domains.However, utilizing LLMs for ubiquitous sensing applications remains challenging as existing text-prompt methods show significant performance degradation when handling long sensor data sequences.We propose a visual prompting approach for sensor data using multimodal LLMs (MLLMs).We design a visual prompt that directs MLLMs to utilize visualized sensor data alongside the target sensory task descriptions.Additionally, we introduce a visualization generator that automates the creation of optimal visualizations tailored to a given sensory task, eliminating the need for prior task-specific knowledge.We evaluated our approach on nine sensory tasks involving four sensing modalities, achieving an average of 10% higher accuracy than text-based prompts and reducing token costs by 15.8×.Our findings highlight the effectiveness and cost-efficiency of visual prompts with MLLMs for various sensory tasks.The source code is available at https://github. com/diamond264/ByMyEyes. ### InstructionYou are an expert in sensor data analysis.Given the sensor data, determine the correct answer from the options listed in the question.Provide the answer with the format of ANSWER , where ANSWER corresponds to one of the options listed in the question.If the answer is not in the options, choose the most possible option.The ECG data is collected from a lead II ECG sensor.The ECG data is recorded over 10 seconds.The data is normalized with the statistics of the user's data.Please refer to the provided examples and use them to answer the following question for the target data.### Examples *Example of normal*: Average heartbeat in the ECG signal (list of ['lead II']): [-0.34, -0.34, -0.35, -0.35, -0.36, … ECG_P_Peaks in the ECG signal (list of (index, value)): [(21, 0.08), (129, 0.19), (239, 0.22), … ECG_Q_Peaks in the ECG signal (list of (index, value)): [(35, -0.61), (137, -0.46), (246, -0.48), … ECG_S_Peaks in the ECG signal (list of (index, value)): [(42, -1.4), (149, -1.35), (260, -1.33), … ECG_T_Peaks in the ECG signal (list of (index, value)): [(63, 2.18), (171, 2.11), (282, 2.31), … *Example of conduction disturbance*: Average heartbeat in the ECG signal (list of ['lead II']): [-0.15, -0.24, -0.29, -0.31, -0.28, … ECG_P_Peaks in the ECG signal (list of (index, value)): [(4, 0.14), (57, 0.3), (103, 0.22), … ECG_Q_Peaks in the ECG signal (list of (index, value)): [(14, -0.05), (65, -0.05), (109, -0.07), … ECG_S_Peaks in the ECG signal (list of (index, value)): [(22, -2.28), (73, -2.1), (124, -2.35), … ECG_T_Peaks in the ECG signal (list of (index, value)): [(82, 0.03), (142, 0.64), (245, 0.25), … ### Question Average heartbeat in the ECG signal (list of ['lead II']): [-0.39, -0.39, -0.39, -0.39, -0.4,… ECG_P_Peaks in the ECG signal (list of (index, value)): [(15, 0.14), (94, -0.26), (173, -0.23), … ECG_Q_Peaks in the ECG signal (list of (index, value)): [(23, -0.27), (102, -0.81), (182, -0.55), … ECG_S_Peaks in the ECG signal (list of (index, value)): [(34, 0.15), (116, -0.51), (192, -0.45), … ECG_T_Peaks in the ECG signal (list of (index, value)): [(50, 1.39), (130, 1.1), (209, 1.31), … *Question*: When the sensor data is used for a task for classifying ECG data into 2 categories: conduction disturbance, normal, what is the most likely answer among ['conduction disturbance', 'normal']?*Answer*: Hyungjun Yoon, Biniyam Aschalew Tolera, Taesik Gong, Kimin Lee, Sung-Ju Lee 0001 |
EMNLP | 4 |
| 2024 | Confidence-aware Reward Optimization for Fine-tuning Text-to-Image ModelsabstractFine-tuning text-to-image models with reward functions trained on human feedback data has proven effective for aligning model behavior with human intent. However, excessive optimization with such reward models, which serve as mere proxy objectives, can compromise the performance of fine-tuned models, a phenomenon known as reward overoptimization. To investigate this issue in depth, we introduce the Text-Image Alignment Assessment (TIA2) benchmark, which comprises a diverse collection of text prompts, images, and human annotations. Our evaluation of several state-of-the-art reward models on this benchmark reveals their frequent misalignment with human assessment. We empirically demonstrate that overoptimization occurs notably when a poorly aligned reward model is used as the fine-tuning objective. To address this, we propose TextNorm, a simple method that enhances alignment based on a measure of reward model confidence estimated across a set of semantically contrastive text prompts. We demonstrate that incorporating the confidence-calibrated rewards in fine-tuning effectively reduces overoptimization, resulting in twice as many wins in human evaluation for text-image alignment compared against the baseline reward models. Kyuyoung Kim, Jongheon Jeong, Minyong An, Mohammad Ghavamzadeh, Krishnamurthy Dvijotham, Jinwoo Shin, Kimin Lee |
ICLR | 7 |
| 2024 | Identity Decoupling for Multi-Subject Personalization of Text-to-Image ModelsabstractText-to-image diffusion models have shown remarkable success in generating personalized subjects based on a few reference images. However, current methods often fail when generating multiple subjects simultaneously, resulting in mixed
identities with combined attributes from different subjects. In this work, we present MuDI, a novel framework that enables multi-subject personalization by effectively decoupling identities from multiple subjects. Our main idea is to utilize segmented subjects generated by a foundation model for segmentation (Segment Anything) for both training and inference, as a form of data augmentation for training and initialization for the generation process. Moreover, we further introduce a new metric to better evaluate the performance of our method on multi-subject personalization. Experimental results show that our MuDI can produce high-quality personalized images without identity mixing, even for highly similar subjects as shown in Figure 1. Specifically, in human evaluation, MuDI obtains twice the success rate for personalizing multiple subjects without identity mixing over existing baselines and is preferred over 70% against the strongest baseline. Sangwon Jang, Jaehyeong Jo, Kimin Lee, Sung Ju Hwang |
NeurIPS | 3 |
| 2023 | Preference Transformer: Modeling Human Preferences using Transformers for RL
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, Kimin Lee |
ICLR | 6 |
| 2023 | Controllability-Aware Unsupervised Skill DiscoveryabstractOne of the key capabilities of intelligent agents is the ability to discover useful skills without external supervision. However, the current unsupervised skill discovery methods are often limited to acquiring simple, easy-to-learn skills due to the lack of incentives to discover more complex, challenging behaviors. We introduce a novel unsupervised skill discovery method, Controllability-aware Skill Discovery (CSD), which actively seeks complex, hard-to-control skills without supervision. The key component of CSD is a controllability-aware distance function, which assigns larger values to state transitions that are harder to achieve with the current skills. Combined with distance-maximizing skill discovery, CSD progressively learns more challenging skills over the course of training as our jointly trained distance function reduces rewards for easy-to-achieve skills. Our experimental results in six robotic manipulation and locomotion environments demonstrate that CSD can discover diverse complex skills including object manipulation and locomotion skills with no supervision, significantly outperforming prior unsupervised skill discovery methods. Videos and code are available at https://seohong.me/projects/csd/ Seohong Park, Kimin Lee, Youngwoon Lee, Pieter Abbeel |
ICML | 2 |
| 2023 | Multi-View Masked World Models for Visual Robotic ManipulationabstractVisual robotic manipulation research and applications often use multiple cameras, or views, to better perceive the world. How else can we utilize the richness of multi-view data? In this paper, we investigate how to learn good representations with multi-view data and utilize them for visual robotic manipulation. Specifically, we train a multi-view masked autoencoder which reconstructs pixels of randomly masked viewpoints and then learn a world model operating on the representations from the autoencoder. We demonstrate the effectiveness of our method in a range of scenarios, including multi-view control and single-view control with auxiliary cameras for representation learning. We also show that the multi-view masked autoencoder trained with multiple randomized viewpoints enables training a policy with strong viewpoint randomization and transferring the policy to solve real-robot tasks without camera calibration and an adaptation procedure. Video demonstrations are available at: https://sites.google.com/view/mv-mwm. Younggyo Seo, Stephen James, Kimin Lee, Jinwoo Shin, Pieter Abbeel |
ICML | 4 |
| 2023 | Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Olivia Watkins, Hao Liu 0055, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee 0001, Kimin Lee |
NeurIPS | 10 |
| 2023 | Guide Your Agent with Adaptive Multimodal RewardsabstractDeveloping an agent capable of adapting to unseen environments remains a difficult challenge in imitation learning. This work presents Adaptive Return-conditioned Policy (ARP), an efficient framework designed to enhance the agent's generalization ability using natural language task descriptions and pre-trained multimodal encoders. Our key idea is to calculate a similarity between visual observations and natural language instructions in the pre-trained multimodal embedding space (such as CLIP) and use it as a reward signal. We then train a return-conditioned policy using expert demonstrations labeled with multimodal rewards. Because the multimodal rewards provide adaptive signals at each timestep, our ARP effectively mitigates the goal misgeneralization. This results in superior generalization performances even when faced with unseen text instructions, compared to existing text-conditioned policies. To improve the quality of rewards, we also introduce a fine-tuning method for pre-trained multimodal encoders, further enhancing the performance. Video demonstrations and source code are available on the project website: \url{https://sites.google.com/view/2023arp}. Changyeon Kim, Younggyo Seo, Lisa Lee, Jinwoo Shin, Honglak Lee, Kimin Lee |
NeurIPS | 7 |
| 2023 | StyleDrop: Text-to-Image Synthesis of Any StyleabstractPre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language, and out-of-distribution effects make it hard to synthesize arbitrary image styles, leveraging a specific design pattern, texture or material. In this paper, we introduce *StyleDrop*, a method that enables the synthesis of images that faithfully follow a specific style using a text-to-image model. StyleDrop is extremely versatile and captures nuances and details of a user-provided style, such as color schemes, shading, design patterns, and local and global effects. StyleDrop works by efficiently learning a new style by fine-tuning very few trainable parameters (less than 1\% of total model parameters), and improving the quality via iterative training with either human or automated feedback. Better yet, StyleDrop is able to deliver impressive results even when the user supplies only a *single* image specifying the desired style. An extensive study shows that, for the task of style tuning text-to-image models, StyleDrop on Muse convincingly outperforms other methods, including DreamBooth and textual inversion on Imagen or Stable Diffusion. More results are available at our project website: [https://styledrop.github.io](https://styledrop.github.io). Kihyuk Sohn, Lu Jiang 0004, Jarred Barber, Kimin Lee, Nataniel Ruiz, Dilip Krishnan, Huiwen Chang, Yuanzhen Li, Irfan A. Essa, Michael Rubinstein, Yuan Hao, Glenn Entis, Irina Blok, Daniel Castro Chin |
NeurIPS | 4 |
| 2022 | Programmatic Modeling and Generation of Real-Time Strategic Soccer Environments for Reinforcement LearningabstractThe capability of a reinforcement learning (RL) agent heavily depends on the diversity of the learning scenarios generated by the environment. Generation of diverse realistic scenarios is challenging for real-time strategy (RTS) environments. The RTS environments are characterized by intelligent entities/non-RL agents cooperating and competing with the RL agents with large state and action spaces over a long period of time, resulting in an infinite space of feasible, but not necessarily realistic, scenarios involving complex interaction among different RL and non-RL agents. Yet, most of the existing simulators rely on randomly generating the environments based on predefined settings/layouts and offer limited flexibility and control over the environment dynamics for researchers to generate diverse, realistic scenarios as per their demand. To address this issue, for the first time, we formally introduce the benefits of adopting an existing formal scenario specification language, SCENIC, to assist researchers to model and generate diverse scenarios in an RTS environment in a flexible, systematic, and programmatic manner. To showcase the benefits, we interfaced SCENIC to an existing RTS environment Google Research Football (GRF) simulator and introduced a benchmark consisting of 32 realistic scenarios, encoded in SCENIC, to train RL agents and testing their generalization capabilities. We also show how researchers/RL practitioners can incorporate their domain knowledge to expedite the training process by intuitively modeling stochastic programmatic policies with SCENIC. Abdus Salam Azad, Edward Kim 0005, Qiancheng Wu, Kimin Lee, Ion Stoica, Pieter Abbeel, Alberto L. Sangiovanni-Vincentelli, Sanjit A. Seshia |
AAAI | 4 |
| 2022 | HARP: Autoregressive Latent Video Prediction with High-Fidelity Image GeneratorabstractVideo prediction is an important yet challenging problem; burdened with the tasks of generating future frames and learning environment dynamics. Recently, autoregressive latent video models have proved to be a powerful video prediction tool, by separating the video prediction into two sub-problems: pre-training an image generator model, followed by learning an autoregressive prediction model in the latent space of the image generator. However, successfully generating high-fidelity and high-resolution videos has yet to be seen. In this work, we investigate how to train an autoregressive latent video prediction model capable of predicting high-fidelity future frames with minimal modification to existing models, and produce high-resolution (256x256) videos. Specifically, we scale up prior models by employing a high-fidelity image generator (VQ-GAN) with a causal transformer model, and introduce additional techniques of top-k sampling and data augmentation to further improve video prediction quality. Despite the simplicity, the proposed method achieves competitive performance to state-of-the-art approaches on standard video prediction benchmarks with fewer parameters, and enables high-resolution video prediction on complex and large-scale datasets. Younggyo Seo, Kimin Lee, Fangchen Liu, Stephen James, Pieter Abbeel |
ICIP | 2 |
| 2022 | Reward Uncertainty for Exploration in Preference-based Reinforcement Learning
Xinran Liang, Katherine Shu, Kimin Lee, Pieter Abbeel |
ICLR | 3 |
| 2022 | SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Jongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee, Pieter Abbeel, Kimin Lee |
ICLR | 6 |
| 2022 | Reinforcement Learning with Action-Free Pre-Training from VideosabstractRecent unsupervised pre-training methods have shown to be effective on language and vision domains by learning useful representations for multiple downstream tasks. In this paper, we investigate if such unsupervised pre-training methods can also be effective for vision-based reinforcement learning (RL). To this end, we introduce a framework that learns representations useful for understanding the dynamics via generative pre-training on videos. Our framework consists of two phases: we pre-train an action-free latent video prediction model, and then utilize the pre-trained representations for efficiently learning action-conditional world models on unseen environments. To incorporate additional action inputs during fine-tuning, we introduce a new architecture that stacks an action-conditional latent prediction model on top of the pre-trained action-free prediction model. Moreover, for better exploration, we propose a video-based intrinsic bonus that leverages pre-trained representations. We demonstrate that our framework significantly improves both final performances and sample-efficiency of vision-based RL in a variety of manipulation and locomotion tasks. Code is available at \url{https://github.com/younggyoseo/apv}. Younggyo Seo, Kimin Lee, Stephen James, Pieter Abbeel |
ICML | 2 |
| 2022 | Towards More Generalizable One-shot Visual Imitation LearningabstractA general-purpose robot should be able to master a wide range of tasks and quickly learn a novel one by leveraging past experiences. One-shot imitation learning (OSIL) approaches this goal by training an agent with (pairs of) expert demonstrations, such that at test time, it can directly execute a new task from just one demonstration. However, so far this framework has been limited to training on many variations of one task, and testing on other unseen but similar variations of the same task. In this work, we push for a higher level of generalization ability by investigating a more ambitious multi-task setup. We introduce a diverse suite of vision-based robot manipulation tasks, consisting of 7 tasks, a total of 61 variations, and a continuum of instances within each variation. For consistency and comparison purposes, we first train and evaluate single-task agents (as done in prior few-shot imitation work). We then study the multi-task setting, where multi-task training is followed by (i) one-shot imitation on variations within the training tasks, (ii) one-shot imitation on new tasks, and (iii) fine-tuning on new tasks. Prior state-of-the-art, while performing well within some single tasks, struggles in these harder multi-task settings. To address these limitations, we propose MOSAIC (Multi-task One-Shot Imitation with self-Attention and Contrastive learning), which integrates a self-attention model architecture and a temporal contrastive module to enable better task disambiguation and more robust representation learning. Our experiments show that MOSAIC outperforms prior state of the art in learning efficiency, final performance, and learns a multi-task policy with promising generalization ability via fine-tuning on novel tasks. Mandi Zhao, Fangchen Liu, Kimin Lee, Pieter Abbeel |
ICRA | 3 |
| 2021 | MASKER: Masked Keyword Regularization for Reliable Text ClassificationabstractPre-trained language models have achieved state-of-the-art accuracies on various text classification tasks, e.g., sentiment analysis, natural language inference, and semantic textual similarity. However, the reliability of the fine-tuned text classifiers is an often underlooked performance criterion. For instance, one may desire a model that can detect out-of-distribution (OOD) samples (drawn far from training distribution) or be robust against domain shifts. We claim that one central obstacle to the reliability is the over-reliance of the model on a limited number of keywords, instead of looking at the whole context. In particular, we find that (a) OOD samples often contain in-distribution keywords, while (b) cross-domain samples may not always contain keywords; over-relying on the keywords can be problematic for both cases. In light of this observation, we propose a simple yet effective fine-tuning method, coined masked keyword regularization (MASKER), that facilitates context-based prediction. MASKER regularizes the model to reconstruct the keywords from the rest of the words and make low-confidence predictions without enough context. When applied to various pre-trained language models (e.g., BERT, RoBERTa, and ALBERT), we demonstrate that MASKER improves OOD detection and cross-domain generalization without degrading classification accuracy. Code is available at https://github.com/alinlab/MASKER. Seung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee 0001, Jinwoo Shin |
AAAI | 3 |
| 2021 | Learning to Sample with Local and Global Contexts in Experience Replay Buffer
Kimin Lee, Jinwoo Shin, Eunho Yang, Sung Ju Hwang |
ICLR | 2 |
| 2021 | SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningabstractOff-policy deep reinforcement learning (RL) has been successful in a range of challenging domains. However, standard off-policy RL algorithms can suffer from several issues, such as instability in Q-learning and balancing exploration and exploitation. To mitigate these issues, we present SUNRISE, a simple unified ensemble method, which is compatible with various off-policy RL algorithms. SUNRISE integrates two key ingredients: (a) ensemble-based weighted Bellman backups, which re-weight target Q-values based on uncertainty estimates from a Q-ensemble, and (b) an inference method that selects actions using the highest upper-confidence bounds for efficient exploration. By enforcing the diversity between agents using Bootstrap with random initialization, we show that these different ideas are largely orthogonal and can be fruitfully integrated, together further improving the performance of existing off-policy RL algorithms, such as Soft Actor-Critic and Rainbow DQN, for both continuous and discrete control tasks on both low-dimensional and high-dimensional environments. Kimin Lee, Michael Laskin, Aravind Srinivas, Pieter Abbeel |
ICML | 1 |
| 2021 | PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingabstractConveying complex objectives to reinforcement learning (RL) agents can often be difficult, involving meticulous design of reward functions that are sufficiently informative yet easy enough to provide. Human-in-the-loop RL methods allow practitioners to instead interactively teach agents through tailored feedback; however, such approaches have been challenging to scale since human feedback is very expensive. In this work, we aim to make this process more sample- and feedback-efficient. We present an off-policy, interactive RL algorithm that capitalizes on the strengths of both feedback and off-policy learning. Specifically, we learn a reward model by actively querying a teacher’s preferences between two clips of behavior and use it to train an agent. To enable off-policy learning, we relabel all the agent’s past experience when its reward model changes. We additionally show that pre-training our agents with unsupervised exploration substantially increases the mileage of its queries. We demonstrate that our approach is capable of learning tasks of higher complexity than previously considered by human-in-the-loop methods, including a variety of locomotion and robotic manipulation skills. We also show that our method is able to utilize real-time human feedback to effectively prevent reward exploitation and learn new behaviors that are difficult to specify with standard reward functions. Kimin Lee, Laura M. Smith, Pieter Abbeel |
ICML | 1 |
| 2021 | State Entropy Maximization with Random Encoders for Efficient ExplorationabstractRecent exploration methods have proven to be a recipe for improving sample-efficiency in deep reinforcement learning (RL). However, efficient exploration in high-dimensional observation spaces still remains a challenge. This paper presents Random Encoders for Efficient Exploration (RE3), an exploration method that utilizes state entropy as an intrinsic reward. In order to estimate state entropy in environments with high-dimensional observations, we utilize a k-nearest neighbor entropy estimator in the low-dimensional representation space of a convolutional encoder. In particular, we find that the state entropy can be estimated in a stable and compute-efficient manner by utilizing a randomly initialized encoder, which is fixed throughout training. Our experiments show that RE3 significantly improves the sample-efficiency of both model-free and model-based RL methods on locomotion and navigation tasks from DeepMind Control Suite and MiniGrid benchmarks. We also show that RE3 allows learning diverse behaviors without extrinsic rewards, effectively improving sample-efficiency in downstream tasks. Younggyo Seo, Jinwoo Shin, Honglak Lee, Pieter Abbeel, Kimin Lee |
ICML | 6 |
| 2021 | Decoupling Representation Learning from Reinforcement LearningabstractIn an effort to overcome limitations of reward-driven feature learning in deep reinforcement learning (RL) from images, we propose decoupling representation learning from policy learning. To this end, we introduce a new unsupervised learning (UL) task, called Augmented Temporal Contrast (ATC), which trains a convolutional encoder to associate pairs of observations separated by a short time difference, under image augmentations and using a contrastive loss. In online RL experiments, we show that training the encoder exclusively using ATC matches or outperforms end-to-end RL in most environments. Additionally, we benchmark several leading UL algorithms by pre-training encoders on expert demonstrations and using them, with weights frozen, in RL agents; we find that agents using ATC-trained encoders outperform all others. We also train multi-task encoders on data from multiple environments and show generalization to different downstream RL tasks. Finally, we ablate components of ATC, and introduce a new data augmentation to enable replay of (compressed) latent images from pre-trained encoders when RL requires augmentation. Our experiments span visually diverse RL benchmarks in DeepMind Control, DeepMind Lab, and Atari, and our complete code is available at \url{https://github.com/astooke/rlpyt/tree/master/rlpyt/ul}. Adam Stooke, Kimin Lee, Pieter Abbeel, Michael Laskin |
ICML | 2 |
| 2021 | Reinforcement Learning for Sparse-Reward Object-Interaction Tasks in a First-person Simulated 3D EnvironmentabstractLearning how to execute complex tasks involving multiple objects in a 3D world is challenging when there is no ground-truth information about the objects or any demonstration to learn from. When an agent only receives a signal from task-completion, this makes it challenging to learn the object-representations which support learning the correct object-interactions needed to complete the task. In this work, we formulate learning an attentive object dynamics model as a classification problem, using random object-images to define incorrect labels for our object-dynamics model. We show empirically that this enables object-representation learning that captures an object's category (is it a toaster?), its properties (is it on?), and object-relations (is something inside of it?). With this, our core learner (a relational RL agent) receives the dense training signal it needs to rapidly learn object-interaction tasks. We demonstrate results in the 3D AI2Thor simulated kitchen environment with a range of challenging food preparation tasks. We compare our method's performance to several related approaches and against the performance of an oracle: an agent that is supplied with ground-truth information about objects in the scene. We find that our agent achieves performance closest to the oracle in terms of both learning speed and maximum success rate. Wilka Carvalho, Anthony Liang, Kimin Lee, Sungryull Sohn, Honglak Lee, Richard L. Lewis, Satinder Singh 0001 |
IJCAI | 3 |
| 2021 | Decision Transformer: Reinforcement Learning via Sequence ModelingabstractWe introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer architecture, and associated advances in language modeling such as GPT-x and BERT. In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired return (reward), past states, and actions, our Decision Transformer model can generate future actions that achieve the desired return. Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks. Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, Igor Mordatch |
NeurIPS | 4 |
| 2021 | Improving Computational Efficiency in Visual Reinforcement Learning via Stored EmbeddingsabstractRecent advances in off-policy deep reinforcement learning (RL) have led to impressive success in complex tasks from visual observations. Experience replay improves sample-efficiency by reusing experiences from the past, and convolutional neural networks (CNNs) process high-dimensional inputs effectively. However, such techniques demand high memory and computational bandwidth. In this paper, we present Stored Embeddings for Efficient Reinforcement Learning (SEER), a simple modification of existing off-policy RL methods, to address these computational and memory requirements. To reduce the computational overhead of gradient updates in CNNs, we freeze the lower layers of CNN encoders early in training due to early convergence of their parameters. Additionally, we reduce memory requirements by storing the low-dimensional latent vectors for experience replay instead of high-dimensional images, enabling an adaptive increase in the replay buffer capacity, a useful technique in constrained-memory settings. In our experiments, we show that SEER does not degrade the performance of RL agents while significantly saving computation and memory across a diverse set of DeepMind Control environments and Atari games. Kimin Lee, Aravind Srinivas, Pieter Abbeel |
NeurIPS | 2 |
| 2021 | Improving Transferability of Representations via Augmentation-Aware Self-SupervisionabstractRecent unsupervised representation learning methods have shown to be effective in a range of vision tasks by learning representations invariant to data augmentations such as random cropping and color jittering. However, such invariance could be harmful to downstream tasks if they rely on the characteristics of the data augmentations, e.g., location- or color-sensitive. This is not an issue just for unsupervised learning; we found that this occurs even in supervised learning because it also learns to predict the same label for all augmented samples of an instance. To avoid such failures and obtain more generalizable representations, we suggest to optimize an auxiliary self-supervised loss, coined AugSelf, that learns the difference of augmentation parameters (e.g., cropping positions, color adjustment intensities) between two randomly augmented samples. Our intuition is that AugSelf encourages to preserve augmentation-aware information in learned representations, which could be beneficial for their transferability. Furthermore, AugSelf can easily be incorporated into recent state-of-the-art representation learning methods with a negligible additional training cost. Extensive experiments demonstrate that our simple idea consistently improves the transferability of representations learned by supervised and unsupervised methods in various transfer learning scenarios. The code is available at https://github.com/hankook/AugSelf. Hankook Lee, Kibok Lee 0003, Kimin Lee, Honglak Lee, Jinwoo Shin |
NeurIPS | 3 |
| 2020 | Regularizing Class-Wise Predictions via Self-Knowledge DistillationabstractDeep neural networks with millions of parameters may suffer from poor generalization due to overfitting. To mitigate the issue, we propose a new regularization method that penalizes the predictive distribution between similar samples. In particular, we distill the predictive distribution between different samples of the same label during training. This results in regularizing the dark knowledge (i.e., the knowledge on wrong predictions) of a single network (i.e., a self-knowledge distillation) by forcing it to produce more meaningful and consistent predictions in a class-wise manner. Consequently, it mitigates overconfident predictions and reduces intra-class variations. Our experimental results on various image classification tasks demonstrate that the simple yet powerful method can significantly improve not only the generalization ability but also the calibration performance of modern convolutional neural networks. Sukmin Yun, Jongjin Park, Kimin Lee, Jinwoo Shin |
CVPR | 3 |
| 2020 | Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
Kimin Lee, Kibok Lee 0003, Jinwoo Shin, Honglak Lee |
ICLR | 1 |
| 2020 | Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningabstractModel-based reinforcement learning (RL) enjoys several benefits, such as data-efficiency and planning, by learning a model of the environment’s dynamics. However, learning a global model that can generalize across different dynamics remains a challenge. To tackle this problem, we decompose the task of learning a global dynamics model into two stages: (a) learning a context latent vector that captures the local dynamics, then (b) predicting the next state conditioned on it. In order to encode dynamics-specific information into the context latent vector, we introduce a novel loss function that encourages the context latent vector to be useful for predicting both forward and backward dynamics. The proposed method achieves superior generalization ability across various simulated robotics and control tasks, compared to existing RL schemes. Kimin Lee, Younggyo Seo, Honglak Lee, Jinwoo Shin |
ICML | 1 |
| 2020 | Reinforcement Learning with Augmented DataabstractLearning from visual observations is a fundamental yet challenging problem in Reinforcement Learning (RL). Although algorithmic advances combined with convolutional neural networks have proved to be a recipe for success, current methods are still lacking on two fronts: (a) data-efficiency of learning and (b) generalization to new environments. To this end, we present Reinforcement Learning with Augmented Data (RAD), a simple plug-and-play module that can enhance most RL algorithms. We perform the first extensive study of general data augmentations for RL on both pixel-based and state-based inputs, and introduce two new data augmentations - random translate and random amplitude scale. We show that augmentations such as random translate, crop, color jitter, patch cutout, random convolutions, and amplitude scale can enable simple RL algorithms to outperform complex state-of-the-art methods across common benchmarks. RAD sets a new state-of-the-art in terms of data-efficiency and final performance on the DeepMind Control Suite benchmark for pixel-based control as well as OpenAI Gym benchmark for state-based control. We further demonstrate that RAD significantly improves test-time generalization over existing methods on several OpenAI ProcGen benchmarks. Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, Aravind Srinivas |
NeurIPS | 2 |
| 2020 | Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement LearningabstractModel-based reinforcement learning (RL) has shown great potential in various control tasks in terms of both sample-efficiency and final performance. However, learning a generalizable dynamics model robust to changes in dynamics remains a challenge since the target transition dynamics follow a multi-modal distribution. In this paper, we present a new model-based RL algorithm, coined trajectory-wise multiple choice learning, that learns a multi-headed dynamics model for dynamics generalization. The main idea is updating the most accurate prediction head to specialize each head in certain environments with similar dynamics, i.e., clustering environments. Moreover, we incorporate context learning, which encodes dynamics-specific information from past experiences into the context latent vector, enabling the model to perform online adaptation to unseen environments. Finally, to utilize the specialized prediction heads more effectively, we propose an adaptive planning method, which selects the most accurate prediction head over a recent experience. Our method exhibits superior zero-shot generalization performance across a variety of control tasks, compared to state-of-the-art RL methods. Source code and videos are available at https://sites.google.com/view/trajectory-mcl. Younggyo Seo, Kimin Lee, Ignasi Clavera Gilaberte, Thanard Kurutach, Jinwoo Shin, Pieter Abbeel |
NeurIPS | 2 |
| 2020 | Dynamic Control for On-Demand Interference-Managed WLAN InfrastructuresabstractIn order to handle a high traffic demand, dense wireless local area networks (WLANs) have been deployed rapidly in the past years. However, dense WLANs cause two critical issues: wastage of energy and severe interference. To address these issues, the centralized management of dense WLANs has been emerged as a powerful paradigm for improving energy efficiency as well as avoiding severe interference. In this paper, we study the joint optimization problem of power-operation modes in access points (APs), channel selections and user-AP associations for improving energy efficiency and avoiding interference without sacrificing users' demands. To this end, we first formulate it as a mixed-integer programming using the popular Lyapunov approach, but it turns out to be computationally intractable, i.e., NP-hard. To address the issue, we propose a polynomial-time approximation algorithm and prove that it achieves a constant-factor approximation guarantee under mild assumptions. The main novelty underlying our algorithm design is based on a linear programming relaxation combining with two different greedy rounding schemes, where each achieves a constant-factor approximation in different regimes of parameters. We verify the performance of the proposed algorithm via extensive simulations and also demonstrate its practicability by implementing it at commercial APs using a Software-defined Networking framework. Results from our experiments show that it reduces the wasted energy significantly while maintaining even higher throughput. Seokhyun Kim, Kimin Lee, Yeonkeun Kim, Jinwoo Shin, Seungwon Shin 0001, Song Chong |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | Overcoming Catastrophic Forgetting With Unlabeled Data in the WildabstractLifelong learning with deep neural networks is well-known to suffer from catastrophic forgetting: the performance on previous tasks drastically degrades when learning a new task. To alleviate this effect, we propose to leverage a large stream of unlabeled data easily obtainable in the wild. In particular, we design a novel class-incremental learning scheme with (a) a new distillation loss, termed global distillation, (b) a learning strategy to avoid overfitting to the most recent task, and (c) a confidence-based sampling method to effectively leverage unlabeled external data. Our experimental results on various datasets, including CIFAR and ImageNet, demonstrate the superiority of the proposed methods over prior methods, particularly when a stream of unlabeled data is accessible: our method shows up to 15.8% higher accuracy and 46.5% less forgetting compared to the state-of-the-art method. The code is available at https://github.com/kibok90/iccv2019-inc. Kibok Lee 0003, Kimin Lee, Jinwoo Shin, Honglak Lee |
ICCV | 2 |
| 2019 | Using Pre-Training Can Improve Model Robustness and UncertaintyabstractHe et al. (2018) have called into question the utility of pre-training by showing that training from scratch can often yield similar performance to pre-training. We show that although pre-training may not improve performance on traditional classification metrics, it improves model robustness and uncertainty estimates. Through extensive experiments on label corruption, class imbalance, adversarial examples, out-of-distribution detection, and confidence calibration, we demonstrate large gains from pre-training and complementary effects with task-specific methods. We show approximately a 10% absolute improvement over the previous state-of-the-art in adversarial robustness. In some cases, using pre-training without task-specific methods also surpasses the state-of-the-art, highlighting the need for pre-training when evaluating future methods on robustness and uncertainty tasks. Dan Hendrycks, Kimin Lee, Mantas Mazeika |
ICML | 2 |
| 2019 | Robust Inference via Generative Classifiers for Handling Noisy LabelsabstractLarge-scale datasets may contain significant proportions of noisy (incorrect) class labels, and it is well-known that modern deep neural networks (DNNs) poorly generalize from such noisy training datasets. To mitigate the issue, we propose a novel inference method, termed Robust Generative classifier (RoG), applicable to any discriminative (e.g., softmax) neural classifier pre-trained on noisy datasets. In particular, we induce a generative classifier on top of hidden feature spaces of the pre-trained DNNs, for obtaining a more robust decision boundary. By estimating the parameters of generative classifier using the minimum covariance determinant estimator, we significantly improve the classification accuracy with neither re-training of the deep model nor changing its architectures. With the assumption of Gaussian distribution for features, we prove that RoG generalizes better than baselines under noisy labels. Finally, we propose the ensemble version of RoG to improve its performance by investigating the layer-wise characteristics of DNNs. Our extensive experimental results demonstrate the superiority of RoG given different learning models optimized by several training techniques to handle diverse scenarios of noisy labels. Kimin Lee, Sukmin Yun, Kibok Lee 0003, Honglak Lee, Bo Li 0026, Jinwoo Shin |
ICML | 1 |
| 2018 | Hierarchical Novelty Detection for Visual Object RecognitionabstractDeep neural networks have achieved impressive success in large-scale visual object recognition tasks with a predefined set of classes. However, recognizing objects of novel classes unseen during training still remains challenging. The problem of detecting such novel classes has been addressed in the literature, but most prior works have focused on providing simple binary or regressive decisions, e.g., the output would be "known," "novel," or corresponding confidence intervals. In this paper, we study more informative novelty detection schemes based on a hierarchical classification framework. For an object of a novel class, we aim for finding its closest super class in the hierarchical taxonomy of known classes. To this end, we propose two different approaches termed top-down and flatten methods, and their combination as well. The essential ingredients of our methods are confidence-calibrated classifiers, data relabeling, and the leave-one-out strategy for modeling novel classes under the hierarchical taxonomy. Furthermore, our method can generate a hierarchical embedding that leads to improved generalized zero-shot learning performance in combination with other commonly-used semantic embeddings. Kibok Lee 0003, Kimin Lee, Kyle Min 0001, Yuting Zhang 0001, Jinwoo Shin, Honglak Lee |
CVPR | 2 |
| 2018 | Training Confidence-calibrated Classifiers for Detecting Out-of-Distribution Samples
Kimin Lee, Honglak Lee, Kibok Lee 0003, Jinwoo Shin |
ICLR (Poster) | 1 |
| 2018 | A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial AttacksabstractDetecting test samples drawn sufficiently far away from the training distribution statistically or adversarially is a fundamental requirement for deploying a good classifier in many real-world machine learning applications. However, deep neural networks with the softmax classifier are known to produce highly overconfident posterior distributions even for such abnormal samples. In this paper, we propose a simple yet effective method for detecting any abnormal samples, which is applicable to any pre-trained softmax neural classifier. We obtain the class conditional Gaussian distributions with respect to (low- and upper-level) features of the deep models under Gaussian discriminant analysis, which result in a confidence score based on the Mahalanobis distance. While most prior methods have been evaluated for detecting either out-of-distribution or adversarial samples, but not both, the proposed method achieves the state-of-the-art performances for both cases in our experiments. Moreover, we found that our proposed method is more robust in harsh cases, e.g., when the training dataset has noisy labels or small number of samples. Finally, we show that the proposed method enjoys broader usage by applying it to class-incremental learning: whenever out-of-distribution samples are detected, our classification rule can incorporate new classes well without further training deep models. Kimin Lee, Kibok Lee 0003, Honglak Lee, Jinwoo Shin |
NeurIPS | 1 |
| 2018 | Learning to Specialize with Knowledge Distillation for Visual Question AnsweringabstractVisual Question Answering (VQA) is a notoriously challenging problem because it involves various heterogeneous tasks defined by questions within a unified framework. Learning specialized models for individual types of tasks is intuitively attracting but surprisingly difficult; it is not straightforward to outperform naive independent ensemble approach. We present a principled algorithm to learn specialized models with knowledge distillation under a multiple choice learning (MCL) framework, where training examples are assigned dynamically to a subset of models for updating network parameters. The assigned and non-assigned models are learned to predict ground-truth answers and imitate their own base models before specialization, respectively. Our approach alleviates the limitation of data deficiency in existing MCL frameworks, and allows each model to learn its own specialized expertise without forgetting general knowledge. The proposed framework is model-agnostic and applicable to any tasks other than VQA, e.g., image classification with a large number of labels but few per-class examples, which is known to be difficult under existing MCL schemes. Our experimental results indeed demonstrate that our method outperforms other baselines for VQA and image classification. Jonghwan Mun, Kimin Lee, Jinwoo Shin, Bohyung Han |
NeurIPS | 2 |
| 2017 | Confident Multiple Choice LearningabstractEnsemble methods are arguably the most trustworthy techniques for boosting the performance of machine learning models. Popular independent ensembles (IE) relying on naive averaging/voting scheme have been of typical choice for most applications involving deep neural networks, but they do not consider advanced collaboration among ensemble models. In this paper, we propose new ensemble methods specialized for deep neural networks, called confident multiple choice learning (CMCL): it is a variant of multiple choice learning (MCL) via addressing its overconfidence issue.In particular, the proposed major components of CMCL beyond the original MCL scheme are (i) new loss, i.e., confident oracle loss, (ii) new architecture, i.e., feature sharing and (iii) new training method, i.e., stochastic labeling. We demonstrate the effect of CMCL via experiments on the image classification on CIFAR and SVHN, and the foreground-background segmentation on the iCoseg. In particular, CMCL using 5 residual networks provides 14.05\% and 6.60\% relative reductions in the top-1 error rates from the corresponding IE scheme for the classification task on CIFAR and SVHN, respectively. Kimin Lee, Changho Hwang, KyoungSoo Park, Jinwoo Shin |
ICML | 1 |
| 2016 | Just-in-time WLANs: On-demand interference-managed WLAN infrastructuresabstractIn the past years, the centralized management of dense wireless local area networks has been emerged as a powerful paradigm for improving energy efficiency as well as avoiding severe interference. In this paper, we study the joint optimization on power-operation modes in access points (APs), channel selections and user-AP associations for improving energy efficiency and avoiding interference without sacrificing users' demands. To this end, we first formulate it as a mixed-integer programming using the popular Lyapunov approach, but it turns out to be computationally intractable, i.e., NP-hard. To address the issue, we propose a polynomial-time approximation algorithm and prove that it achieves a constant-factor approximation guarantee under mild assumptions. The main novelty underlying our algorithm design is based on a linear programming relaxation combining with two different greedy rounding schemes, where each achieves a constant-factor approximation in different regimes of parameters. We verify the performance of the proposed algorithm via extensive simulations and also demonstrate its practicability by implementing it at commercial APs using a Software-defined Networking framework, which shows that it reduces the wasted energy significantly while maintaining even higher throughput. Kimin Lee, Yeonkeun Kim, Seokhyun Kim, Jinwoo Shin, Song Chong |
INFOCOM | 1 |
| 2016 | TravelMiner: On the Benefit of Path-Based Mobility PredictionabstractMobility predictions are becoming more valuable in various applications with the rise of mobile devices. Given that existing prediction techniques are composed of two key procedures: 1) profiling past mobility trajectories as sequences of discrete atomic states (e.g., grid locations, semantic locations) and capturing them with an appropriate statistical model, 2) making a prediction on the next state using the statistical model, TravelMiner tackles the former with paths utilized as the atomic states for the first time, where the paths are defined as sub-trajectories with no branches. Comparing to available location-based predictors, TravelMiner makes a fundamental difference in that it is able to predict the sequence of paths rather than locations, which is far more detailed in the perspective of knowing the exact route to follow. TravelMiner enables this benefit by extracting disjoint paths from GPS trajectories via a similarity metric for curves, called Frechet distance and keeping the sequences of such paths in a statistical model, called probabilistic radix tree. Our extensive simulations over the GPS trajectories of 124 users reveal that TravelMiner outperforms other predictors in diverse popular performance metrics including predictability, prediction accuracy and prediction resolution. Jaeseong Jeong, Kyunghan Lee, Beknazar Abdikamalov, Kimin Lee, Song Chong |
SECON | 4 |