EDBT 2026 Demo / reviewers in the wild / expert
Sangdoo Yun
dblp:124/3009
· DBLP profile ↗
68ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0002-0417-8450ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 5 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 6 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language ModelsabstractWe identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse.We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code printing internal variables, among others.Fine-tuned models lose their ability to reason about contextual privacy norms, share information inappropriately with tools, and violate memory boundaries across contexts.Privacy collapse is a "silent failure" because models maintain high performance on standard safety and utility benchmarks whilst exhibiting severe privacy vulnerabilities.Our experiments show evidence of privacy collapse across six models (closed and open weight), five fine-tuning datasets (real-world and controlled data), and two task categories (agentic and memory-based).Our mechanistic analysis reveals that privacy representations are uniquely fragile to fine-tuning, compared to task-relevant features which are preserved.Our results reveal a critical gap in current safety evaluations, in particular for the deployment of specialised agents. 1 Anmol Goel, Cornelius Emde, Seong Joon Oh, Sangdoo Yun, Martin Gubri |
ACL (1) | 4 |
| 2026 | ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale InferenceabstractCapturing professionals’ decision-making in creative workflows (e.g., UI/UX) is essential for reflection, collaboration, and knowledge sharing, yet existing methods often leave rationales incomplete and implicit decisions hidden. To address this, we present the Clear approach, which structures reasoning into cognitive decision steps—linked units of actions, artifacts, and explanations making decisions traceable with generative AI. Building on Clear, we introduce ClearFairy, a think-aloud AI assistant for UI design that detects weak explanations, asks lightweight clarifying questions, and infers missing rationales. In a study with twelve professionals, 85% of ClearFairy’s inferred rationales were accepted (as-is or with revisions). Notably, the system increased “strong explanations”—rationales providing sufficient causal reasoning—from 14% to 83% without adding cognitive demand. Furthermore, exploratory applications demonstrate that captured steps can enhance generative AI agents in Figma, yielding predictions better aligned with professionals and producing coherent outcomes. We release a dataset of 417 decision steps to support future research. Kihoon Son, DaEun Choi, Tae Soo Kim 0002, Young-Ho Kim, Sangdoo Yun, Juho Kim 0001 |
CHI | 5 |
| 2026 | LingoQ: Bridging the Gap between EFL Learning and Work through AI-Generated Work-Related QuizzesabstractNon-native English speakers performing English-related tasks at work struggle to sustain EFL learning, despite their motivation. Often, study materials are disconnected from their work context. Our formative study revealed that reviewing work-related English becomes burdensome with current systems, especially after work. Although workers rely on LLM-based assistants to address their immediate needs, these interactions may not directly contribute to their English skills. We present LingoQ, an AI-mediated system that allows workers to practice English using quizzes generated from their LLM queries during work. LingoQ leverages these on-the-fly queries using AI to generate personalized quizzes that workers can review and practice on their smartphones. We conducted a three-week deployment study with 28 EFL workers to evaluate LingoQ. Participants valued the quality-assured, work-situated quizzes and constantly engaging with the app during the study. This active engagement improved self-efficacy and led to learning gains for beginners and, potentially, for intermediate learners. Drawing on these results, we discuss design implications for leveraging workers’ growing reliance on LLMs to foster proficiency and engagement while respecting work boundaries and ethics. Yeonsun Yang, Sang Won Lee 0002, Jean Y. Song, Sangdoo Yun, Young-Ho Kim |
CHI | 4 |
| 2026 | De-Decay: Defusing Computer Vision Model Degradation through Scalable and Actionable Human-Data AlignmentabstractComputer Vision (CV) models can become outdated after deployment as real-world data evolves, requiring intensive attention from AI engineers to address degraded performance through tasks like data relabeling to update models with new human perceptions. Interactive human-in-the-loop systems have considerable potential to enhance model-steering practices. However, such workflows reveal two challenges: (1) scalability, where labor demands increase with data size, and (2) actionability, where human insights do not readily transform into model revisions. Based on our formative study (S1) on the current challenges faced by CV professionals, we developed De-Decay, an end-to-end Human-Data Alignment system offering scalable label-less assessment and actionable insight transformation . This enables engineers to investigate degradation and auto-retrain models with AI support, such as image clustering and regeneration. Our summative study (S2) showed that De-Decay helped engineers effectively identify and address CV degradation. We discuss how future research can enhance scalability and actionability in AI evaluation systems for aligning AI behaviors with human mental models. Tong Steven Sun, Huining Feng, Jinwei Ye, Sangdoo Yun, Young-Ho Kim, Sungsoo Ray Hong |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2025 | Masking meets Supervision: A Strong Learning AllianceabstractPre-training with random masked inputs has emerged as a novel trend in self-supervised training. However, supervised learning still faces a challenge in adopting masking augmentations, primarily due to unstable training. In this paper, we propose a novel way to involve masking augmentations dubbed Masked Sub-branch (MaskSub). MaskSub consists of the main-branch and sub-branch, the latter being a part of the former. The main-branch undergoes conventional training recipes, while the sub-branch merits intensive masking augmentations, during training. MaskSub tackles the challenge by mitigating adverse effects through a relaxed loss function similar to a self-distillation loss. Our analysis shows that MaskSub improves performance, with the training loss converging faster than in standard training, which suggests our method stabilizes the training process. We further validate MaskSub across diverse training scenarios and models, including DeiT-III training, MAE finetuning, CLIP finetuning, BERT training, and hierarchical architectures (ResNet and Swin Transformer). Our results show that MaskSub consistently achieves impressive performance gains across all the cases. MaskSub provides a practical and effective solution for introducing additional regularization under various training recipes. Code available at https://github.com/naver-ai/augsub Byeongho Heo, Taekyung Kim 0002, Sangdoo Yun, Dongyoon Han |
CVPR | 3 |
| 2025 | Leaky Thoughts: Large Reasoning Models Are Not Private ThinkersabstractWe study privacy leakage in the reasoning traces of large reasoning models used as personal agents.Unlike final outputs, reasoning traces are often assumed to be internal and safe.We challenge this assumption by showing that reasoning traces frequently contain sensitive user data, which can be extracted via prompt injections or accidentally leak into outputs.Through probing and agentic evaluations, we demonstrate that test-time compute approaches, particularly increased reasoning steps, amplify such leakage.While increasing the budget of those test-time compute approaches makes models more cautious in their final answers, it also leads them to reason more verbosely and leak more in their own thinking.This reveals a core tension: reasoning improves utility but enlarges the privacy attack surface.We argue that safety efforts must extend to the model's internal thinking, not just its outputs. 1 * Work done during an internship at Parameter Lab. 1 Code available at github.com/parameterlab/leaky_thoughts.AirGapAgent-R benchmark available at huggingface.co/datasets/parameterlab/leaky_thoughts. Prior work on contextual privacy Tommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun, Seong Joon Oh |
EMNLP | 4 |
| 2025 | A Unified Framework for Motion Reasoning and Generation in Human Interaction
Jeongeun Park 0002, Sangdoo Yun |
ICCV | 3 |
| 2025 | Probabilistic Language-Image Pre-TrainingabstractVision-language models (VLMs) embed aligned image-text pairs into a joint space but often rely on deterministic embeddings, assuming a one-to-one correspondence between images and texts. This oversimplifies real-world relationships, which are inherently many-to-many, with multiple captions describing a single image and vice versa. We introduce Probabilistic Language-Image Pre-training (ProLIP), the first probabilistic VLM pre-trained on a billion-scale image-text dataset using only probabilistic objectives, achieving a strong zero-shot capability (e.g., 74.6% ImageNet zero-shot accuracy with ViT-B/16). ProLIP efficiently estimates uncertainty by an ``uncertainty token'' without extra parameters. We also introduce a novel inclusion loss that enforces distributional inclusion relationships between image-text pairs and between original and masked inputs. Experiments demonstrate that, by leveraging uncertainty estimates, ProLIP benefits downstream tasks and aligns with intuitive notions of uncertainty, e.g., shorter texts being more uncertain and more general inputs including specific ones. Utilizing text uncertainties, we further improve ImageNet accuracy from 74.6% to 75.8% (under a few-shot setting), supporting the practical advantages of our probabilistic approach. The code is available at https://github.com/naver-ai/prolip Sanghyuk Chun, Wonjae Kim, Song Park, Sangdoo Yun |
ICLR | 4 |
| 2025 | DaWin: Training-free Dynamic Weight Interpolation for Robust AdaptationabstractAdapting a pre-trained foundation model on downstream tasks should ensure robustness against distribution shifts without the need to retrain the whole model. Although existing weight interpolation methods are simple yet effective, we argue their static nature limits downstream performance while achieving efficiency. In this work, we propose DaWin, a training-free dynamic weight interpolation method that leverages the entropy of individual models over each unlabeled test sample to assess model expertise, and compute per-sample interpolation coefficients dynamically. Unlike previous works that typically rely on additional training to learn such coefficients, our approach requires no training. Then, we propose a mixture modeling approach that greatly reduces inference overhead raised by dynamic interpolation. We validate DaWin on the large-scale visual recognition benchmarks, spanning 14 tasks across robust fine-tuning -- ImageNet and derived five distribution shift benchmarks -- and multi-task learning with eight classification tasks. Results demonstrate that DaWin achieves significant performance gain in considered settings, with minimal computational overhead. We further discuss DaWin's analytic behavior to explain its empirical success. Changdae Oh, Yixuan Li 0001, Kyungwoo Song, Sangdoo Yun, Dongyoon Han |
ICLR | 4 |
| 2025 | Token Bottleneck: One Token to Remember DynamicsabstractDeriving compact and temporally aware visual representations from dynamic scenes is essential for successful execution of sequential scene understanding tasks such as visual tracking and robotic manipulation. In this paper, we introduce Token Bottleneck (ToBo), a simple yet intuitive self-supervised learning pipeline that squeezes a scene into a bottleneck token and predicts the subsequent scene using minimal patches as hints. The ToBo pipeline facilitates the learning of sequential scene representations by conservatively encoding the reference scene into a compact bottleneck token during the squeeze step. In the expansion step, we guide the model to capture temporal dynamics by predicting the target scene using the bottleneck token along with few target patches as hints. This design encourages the vision backbone to embed temporal dependencies, thereby enabling understanding of dynamic transitions across scenes. Extensive experiments in diverse sequential tasks, including video label propagation and robot manipulation in simulated environments demonstrate the superiority of ToBo over baselines. Moreover, deploying our pre-trained model on physical robots confirms its robustness and effectiveness in real-world environments. We further validate the scalability of ToBo across different model scales. Code is available at https://github.com/naver-ai/tobo. Taekyung Kim 0002, Dongyoon Han, Byeongho Heo, Jeongeun Park 0002, Sangdoo Yun |
NeurIPS | 5 |
| 2025 | KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionabstractTransformer-based large language models (LLMs) cache context as key-value (KV) pairs during inference. As context length grows, KV cache sizes expand, leading to substantial memory overhead and increased attention latency. This paper introduces \textit{KVzip}, a query-agnostic KV cache eviction method enabling effective reuse of compressed KV caches across diverse queries. KVzip quantifies the importance of a KV pair using the underlying LLM to reconstruct original contexts from cached KV pairs, subsequently evicting pairs with lower importance. Extensive empirical evaluations demonstrate that KVzip reduces KV cache size by $3$-$4\times$ and FlashAttention decoding latency by approximately $2\times$, with negligible performance loss in question-answering, retrieval, reasoning, and code comprehension tasks. Evaluations include various models such as LLaMA3.1, Qwen2.5, and Gemma3, with context lengths reaching up to 170K tokens. KVzip significantly outperforms existing query-aware KV eviction methods, which suffer from performance degradation even at a 90\% cache budget ratio under multi-query scenarios. Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song |
NeurIPS | 5 |
| 2025 | C-SEO Bench: Does Conversational SEO Work?abstractLarge Language Models (LLMs) are transforming search engines into Conversational Search Engines (CSE). Consequently, Search Engine Optimization (SEO) is being shifted into Conversational Search Engine Optimization (C-SEO). We are beginning to see dedicated C-SEO methods for modifying web documents to increase their visibility in CSE responses. However, they are often tested only for a limited breadth of application domains; we do not know whether certain C-SEO methods would be effective for a broad range of domains. Moreover, existing evaluations consider only a single-actor scenario where only one web document adopts a C-SEO method; in reality, multiple players are likely to competitively adopt the cutting-edge C-SEO techniques, drawing an analogy from the dynamics we have seen in SEO. We present C-SEO Bench, the first benchmark designed to evaluate C-SEO methods across multiple tasks, domains, and number of actors. We consider two search tasks, question answering and product recommendation, with three domains each. We also formalize a new evaluation protocol with varying adoption rates among involved actors. Our experiments reveal that most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected. Instead, traditional SEO strategies, those aiming to improve the ranking of the source in the LLM context, are significantly more effective. We also observe that as we increase the number of C-SEO adopters, the overall gains decrease, depicting a congested and zero-sum nature of the problem. Haritz Puerto, Martin Gubri, Tommaso Green, Seong Joon Oh, Sangdoo Yun |
NeurIPS | 5 |
| 2024 | Match Me If You Can: Semi-supervised Semantic Correspondence Learning with Unpaired Images
Byeongho Heo, Sangdoo Yun, Seungryong Kim, Dongyoon Han |
ACCV (6) | 3 |
| 2024 | Who Wrote this Code? Watermarking for Code GenerationabstractTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, Gunhee Kim. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Taehyun Lee, Seokhee Hong 0002, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, Gunhee Kim |
ACL (1) | 6 |
| 2024 | Calibrating Large Language Models Using Their Generations OnlyabstractAs large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model's confidence in its prediction becomes even more important.However, finding effective ways to calibrate LLMsespecially when the only interface to the models is their generated text-remains a challenge.We propose APRICOT (Auxiliary prediction of confidence targets): A method to set confidence targets and train an additional model that predicts an LLM's confidence based on its textual input and output alone.This approach has several advantages: It is conceptually simple, does not require access to the target model beyond its output, does not interfere with the language generation, and has a multitude of potential usages, for instance by verbalizing the predicted confidence or adjusting the given answer based on the confidence.We show how our approach performs competitively in terms of calibration error for white-box and blackbox LLMs on closed-book question-answering to detect incorrect LLM answers. Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, Seong Joon Oh |
ACL (1) | 4 |
| 2024 | Language-only Efficient Training of Zero-shot Composed Image RetrievalabstractComposed image retrieval (CIR) task takes a composed query of image and text, aiming to search relative images for both conditions. Conventional CIR approaches need a training dataset composed of triplets of query image, query text, and target image, which is very expensive to collect. Several recent works have worked on the zero-shot (ZS) CIR paradigm to tackle the issue without using pre-collected triplets. However, the existing ZS-CIR methods show limited backbone scalability and generalizability due to the lack of diversity of the input texts during training. We propose a novel CIR framework, only using language for its training. Our LinCIR (Language-only training for CIR) can be trained only with text datasets by a novel self-supervision named self-masking projection (SMP). We project the text latent embedding to the token embedding space and construct a new text by replacing the keyword tokens of the original text. Then, we let the new and original texts have the same latent embedding vector. With this simple strategy, LinCIR is surprisingly efficient and highly effective; LinCIR with CLIP ViT-G backbone is trained in 48 minutes and shows the best ZS-CIR performances on four different CIR benchmarks, CIRCO, GeneCIS, FashionIQ, and CIRR, even outperforming supervised method on FashionIQ. Code is available at github.com/navervision/lincir Geonmo Gu, Sanghyuk Chun, Wonjae Kim, Yoohoon Kang, Sangdoo Yun |
CVPR | 5 |
| 2024 | Rotary Position Embedding for Vision Transformer
Byeongho Heo, Song Park, Dongyoon Han, Sangdoo Yun |
ECCV (10) | 4 |
| 2024 | Model Stock: All We Need Is Just a Few Fine-Tuned Models
Dong-Hwan Jang, Sangdoo Yun, Dongyoon Han |
ECCV (44) | 2 |
| 2024 | HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
Wonjae Kim, Sanghyuk Chun, Taekyung Kim 0002, Dongyoon Han, Sangdoo Yun |
ECCV (40) | 5 |
| 2024 | Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsabstractRecently, GPT-4 has become the de facto evaluator for long-form text generated by large language models (LLMs). However, for practitioners and researchers with large and custom evaluation tasks, GPT-4 is unreliable due to its closed-source nature, uncontrolled versioning, and prohibitive costs. In this work, we propose PROMETHEUS a fully open-source LLM that is on par with GPT-4’s evaluation capabilities when the appropriate reference materials (reference answer, score rubric) are accompanied. For this purpose, we construct a new dataset – FEEDBACK COLLECTION – that consists of 1K fine-grained score rubrics, 20K instructions, and 100K natural language feedback generated by GPT-4. Using the FEEDBACK COLLECTION, we train PROMETHEUS, a 13B evaluation-specific LLM that can assess any given response based on novel and unseen score rubrics and reference materials provided by the user. Our dataset’s versatility and diversity make our model generalize to challenging real-world criteria, such as prioritizing conciseness, child-readability, or varying levels of formality. We show that PROMETHEUS shows a stronger correlation with GPT-4 evaluation compared to ChatGPT on seven evaluation benchmarks (Two Feedback Collection testsets, MT Bench, Vicuna Bench, Flask Eval, MT Bench Human Judgment, and HHH Alignment), showing the efficacy of our model and dataset design. During human evaluation with hand-crafted score rubrics, PROMETHEUS shows a Pearson correlation of 0.897 with human evaluators, which is on par with GPT-4-0613 (0.882), and greatly outperforms ChatGPT (0.392). Remarkably, when assessing the quality of the generated feedback, PROMETHEUS demonstrates a win rate of 58.62% when compared to GPT-4 evaluation and a win rate of 79.57% when compared to ChatGPT evaluation. Our findings suggests that by adding reference materials and training on GPT-4 feedback, we can obtain effective open-source evaluator LMs. Seungone Kim, Jamin Shin, Yejin Choi 0001, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, Minjoon Seo |
ICLR | 7 |
| 2024 | Compressed Context Memory for Online Language Model InteractionabstractThis paper presents a context key/value compression method for Transformer language models in online scenarios, where the context continually expands. As the context lengthens, the attention process demands increasing memory and computations, which in turn reduces the throughput of the language model. To address this challenge, we propose a compressed context memory system that continually compresses the accumulating attention key/value pairs into a compact memory space, facilitating language model inference in a limited memory space of computing environments. Our compression process involves integrating a lightweight conditional LoRA into the language model's forward pass during inference, without the need for fine-tuning the model's entire set of weights. We achieve efficient training by modeling the recursive compression process as a single parallelized forward computation. Through evaluations on conversation, personalization, and multi-task learning, we demonstrate that our approach achieves the performance level of a full context model with $5\times$ smaller context memory size. We further demonstrate the applicability of our approach in a streaming setting with an unlimited context length, outperforming the sliding window approach. Codes are available at https://github.com/snu-mllab/context-memory. Junyoung Yeom, Sangdoo Yun, Hyun Oh Song |
ICLR | 3 |
| 2024 | Toward Interactive Regional Understanding in Vision-Large Language ModelsabstractJungbeom Lee, Sanghyuk Chun, Sangdoo Yun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jungbeom Lee, Sanghyuk Chun, Sangdoo Yun |
NAACL-HLT | 3 |
| 2024 | Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsabstractImproving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation. Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alex Hauptmann 0001, Zhi-Qi Cheng, Kyungwoo Song |
NeurIPS | 5 |
| 2024 | Direct Unlearning Optimization for Robust and Safe Text-to-Image ModelsabstractRecent advancements in text-to-image (T2I) models have greatly benefited from large-scale datasets, but they also pose significant risks due to the potential generation of unsafe content. To mitigate this issue, researchers proposed unlearning techniques that attempt to induce the model to unlearn potentially harmful prompts. However, these methods are easily bypassed by adversarial attacks, making them unreliable for ensuring the safety of generated images. In this paper, we propose Direct Unlearning Optimization (DUO), a novel framework for removing NSFW content from T2I models while preserving their performance on unrelated topics. DUO employs a preference optimization approach using curated paired image data, ensuring that the model learns to remove unsafe visual concepts while retain unrelated features. Furthermore, we introduce an output-preserving regularization term to maintain the model's generative capabilities on safe content. Extensive experiments demonstrate that DUO can robustly defend against various state-of-the-art red teaming methods without significant performance degradation on unrelated topics, as measured by FID and CLIP scores. Our work contributes to the development of safer and more reliable T2I models, paving the way for their responsible deployment in both closed-source and open-source scenarios. Yong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Geonhui Jang, Yonghyun Jeong, Junghyo Jo, Gayoung Lee |
NeurIPS | 2 |
| 2023 | MPCHAT: Towards Multimodal Persona-Grounded ConversationabstractIn order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities.However, to fully describe the multi-faceted nature of persona, image modality can help better reveal the speaker's personal characteristics and experiences in episodic memory (Rubin et al., 2003;Conway, 2009).In this work, we extend persona-based dialogue to the multimodal domain and make two main contributions.First, we present the first multimodal persona-based dialogue dataset named MPCHAT, which extends persona with both text and images to contain episodic memories.Second, we empirically show that incorporating multimodal persona, as measured by three proposed multimodal persona-grounded dialogue tasks (i.e., next response prediction, grounding persona prediction, and speaker identification), leads to statistically significant performance improvements across all tasks.Thus, our work highlights that multimodal persona is crucial for improving multimodal dialogue comprehension, and our MPCHAT serves as a high-quality resource for this research. Jaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee Kim |
ACL (1) | 3 |
| 2023 | Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsabstractGeewook Kim, Hodong Lee, Daehee Kim, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, Seunghyun Park. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Geewook Kim, Hodong Lee, Daehee Kim 0003, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, Seunghyun Park 0001 |
EMNLP | 7 |
| 2023 | Neglected Free Lunch - Learning Image Classifiers Using Annotation ByproductsabstractSupervised learning of image classifiers distills human knowledge into a parametric model fθthrough pairs of images and corresponding labels $\left\{ {\left( {{X_i},{Y_i}} \right)} \right\}_{i = 1}^N$. We argue that this simple and widely used representation of human knowledge neglects rich auxiliary information from the annotation procedure, such as the time-series of mouse traces and clicks left after image selection. Our insight is that such annotation byproducts Z provide approximate human attention that weakly guides the model to focus on the foreground cues, reducing spurious correlations and discouraging shortcut learning. To verify this, we create ImageNet-AB and COCO-AB. They are ImageNet and COCO training sets enriched with sample-wise annotation byproducts, collected by replicating the respective original annotation tasks. We refer to the new paradigm of training models with annotation byproducts as learning using annotation byproducts (LUAB). We show that a simple multitask loss for regressing Z together with Y already improves the generalisability and robustness of the learned models. Compared to the original supervised learning, LUAB does not require extra annotation costs. ImageNet-AB and COCO-AB are at github.com/naverai/NeglectedFreeLunch. Dongyoon Han, Junsuk Choe, Seonghyeok Chun, John Joon Young Chung, Minsuk Chang, Sangdoo Yun, Jean Y. Song, Seong Joon Oh |
ICCV | 6 |
| 2023 | SeiT: Storage-Efficient Vision Training with Tokens Using 1% of Pixel StorageabstractWe need billion-scale images to achieve more generalizable and ground-breaking vision models, as well as massive dataset storage to ship the images (e.g., the LAION-5B dataset needs 240TB storage space). However, it has become challenging to deal with unlimited dataset storage with limited storage infrastructure. A number of storage-efficient training methods have been proposed to tackle the problem, but they are rarely scalable or suffer from severe damage to performance. In this paper, we propose a storage-efficient training strategy for vision classifiers for large-scale datasets (e.g., ImageNet) that only uses 1024 tokens per instance without using the raw level pixels; our token storage only needs <1% of the original JPEG-compressed raw pixels. We also propose token augmentations and a Stem-adaptor module to make our approach able to use the same architecture as pixel-based approaches with only minimal modifications on the stem layer and the carefully tuned optimization settings. Our experimental results on ImageNet-1k show that our method significantly outperforms other storage-efficient training methods with a large gap. We further show the effectiveness of our method in other practical scenarios, storage-efficient pre-training, and continual learning. Code is available at https://github.com/naver-ai/seit. Song Park, Sanghyuk Chun, Byeongho Heo, Wonjae Kim, Sangdoo Yun |
ICCV | 5 |
| 2023 | Exploring Temporally Dynamic Data Augmentation for Video Recognition
Taeoh Kim, Jinhyung Kim, Minho Shim, Sangdoo Yun, Myunggu Kang, Dongyoon Wee, Sangyoun Lee |
ICLR | 4 |
| 2023 | What Do Self-Supervised Vision Transformers Learn?
Namuk Park, Wonjae Kim, Byeongho Heo, Taekyung Kim 0002, Sangdoo Yun |
ICLR | 5 |
| 2023 | ProPILE: Probing Privacy Leakage in Large Language ModelsabstractThe rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of web-collected data, which may inadvertently include sensitive personal data. This paper presents ProPILE, a novel probing tool designed to empower data subjects, or the owners of the PII, with awareness of potential PII leakage in LLM-based services. ProPILE lets data subjects formulate prompts based on their own PII to evaluate the level of privacy intrusion in LLMs. We demonstrate its application on the OPT-1.3B model trained on the publicly available Pile dataset. We show how hypothetical data subjects may assess the likelihood of their PII being included in the Pile dataset being revealed. ProPILE can also be leveraged by LLM service providers to effectively evaluate their own levels of PII leakage with more powerful prompts specifically tuned for their in-house models. This tool represents a pioneering step towards empowering the data subjects for their awareness and control over their own data on the web. Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, Seong Joon Oh |
NeurIPS | 2 |
| 2023 | Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier DataabstractDiagnosing and cleaning data is a crucial step for building robust machine learning systems. However, identifying problems within large-scale datasets with real-world distributions is challenging due to the presence of complex issues such as label errors, under-representation, and outliers. In this paper, we propose a unified approach for identifying the problematic data by utilizing a largely ignored source of information: a relational structure of data in the feature-embedded space. To this end, we present scalable and effective algorithms for detecting label errors and outlier data based on the relational graph structure of data. We further introduce a visualization tool that provides contextual information of a data point in the feature-embedded space, serving as an effective tool for interactively diagnosing data. We evaluate the label error and outlier/out-of-distribution (OOD) detection performances of our approach on the large-scale image, speech, and language domain tasks, including ImageNet, ESC-50, and SST2. Our approach achieves state-of-the-art detection performance on all tasks considered and demonstrates its effectiveness in debugging large-scale real-world datasets across various domains. We release codes at https://github.com/snu-mllab/Neural-Relation-Graph. Sangdoo Yun, Hyun Oh Song |
NeurIPS | 2 |
| 2022 | Weakly Supervised Semantic Segmentation using Out-of-Distribution DataabstractWeakly supervised semantic segmentation (WSSS) methods are often built on pixel-level localization maps obtained from a classifier. However, training on class labels only, classifiers suffer from the spurious correlation between fore-ground and background cues (e.g. train and rail), fundamentally bounding the performance of WSSS. There have been previous endeavors to address this issue with additional supervision. We propose a novel source of information to distinguish foreground from the background: Out-of-Distribution (OoD) data, or images devoid of foreground object classes. In particular, we utilize the hard OoDs that the classifier is likely to make false-positive predictions. These samples typically carry key visual features on the background (e.g. rail) that the classifiers often confuse as foreground (e.g. train), so these cues let classifiers correctly suppress spurious background cues. Acquiring such hard OoDs does not require an extensive amount of annotation efforts; it only incurs a few additional image-level labeling costs on top of the original efforts to collect class labels. We propose a method, W-OoD, for utilizing the hard OoDs. W-OoD achieves state-of-the-art performance on Pascal VOC 2012. The code is available at: https://github.com/naver-ai/w-ood. Jungbeom Lee, Seong Joon Oh, Sangdoo Yun, Junsuk Choe, Eunji Kim 0002, Sungroh Yoon |
CVPR | 3 |
| 2022 | Hypergraph-Induced Semantic Tuplet Loss for Deep Metric LearningabstractIn this paper, we propose Hypergraph-Induced Semantic Tuplet (HIST) loss for deep metric learning that leverages the multilateral semantic relations of multiple samples to multiple classes via hypergraph modeling. We formulate deep metric learning as a hypergraph node classification problem in which each sample in a mini-batch is regarded as a node and each hyperedge models class-specific semantic relations represented by a semantic tuplet. Unlike previous graph-based losses that only use a bundle of pairwise relations, our HIST loss takes advantage of the multilateral semantic relations provided by the semantic tuplets through hypergraph modeling. Notably, by leveraging the rich multilateral semantic relations, HIST loss guides the embedding model to learn class-discriminative visual semantics, contributing to better generalization performance and model robustness against input corruptions. Extensive experiments and ablations provide a strong motivation for the proposed method and show that our HIST loss leads to improved feature learning, achieving state-of-the-art results on three widely used benchmarks. Code is available at https://github.com/ljin0429/HIST. Jongin Lim 0002, Sangdoo Yun, Seulki Park, Jin Young Choi 0002 |
CVPR | 2 |
| 2022 | The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationabstractThe problem of class imbalanced data is that the gener-alization performance of the classifier deteriorates due to the lack of data from minority classes. In this paper, we pro-pose a novel minority over-sampling method to augment di-versified minority samples by leveraging the rich context of the majority classes as background images. To diversify the minority samples, our key idea is to paste an image from a minority class onto rich-context images from a majority class, using them as background images. Our method is simple and can be easily combined with the existing long-tailed recognition methods. We empirically prove the effectiveness of the proposed oversampling method through extensive experiments and ablation studies. Without any architectural changes or complex algorithms, our method achieves state-of-the-art performance on various long-tailed classification benchmarks. Our code is made available at https://github.com/naver-ai/cmo. Seulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun, Jin Young Choi 0002 |
CVPR | 4 |
| 2022 | OCR-Free Document Understanding Transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, Seunghyun Park 0001 |
ECCV (28) | 8 |
| 2022 | Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space Perspective
Luca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli, Sangdoo Yun |
ICLR | 5 |
| 2022 | Dataset Condensation via Efficient Synthetic-Data ParameterizationabstractThe great success of machine learning with massive amounts of data comes at a price of huge computation costs and storage for training and tuning. Recent studies on dataset condensation attempt to reduce the dependence on such massive data by synthesizing a compact training dataset. However, the existing approaches have fundamental limitations in optimization due to the limited representability of synthetic datasets without considering any data regularity characteristics. To this end, we propose a novel condensation framework that generates multiple synthetic data with a limited storage budget via efficient parameterization considering data regularity. We further analyze the shortcomings of the existing gradient matching-based condensation methods and develop an effective optimization technique for improving the condensation of training data information. We propose a unified algorithm that drastically improves the quality of condensed data against the current state-of-the-art on CIFAR-10, ImageNet, and Speech Commands. Seong Joon Oh, Sangdoo Yun, Hwanjun Song, Joonhyun Jeong, Jung-Woo Ha 0001, Hyun Oh Song |
ICML | 4 |
| 2022 | Dataset Condensation with Contrastive SignalsabstractRecent studies have demonstrated that gradient matching-based dataset synthesis, or dataset condensation (DC), methods can achieve state-of-theart performance when applied to data-efficient learning tasks. However, in this study, we prove that the existing DC methods can perform worse than the random selection method when taskirrelevant information forms a significant part of the training dataset. We attribute this to the lack of participation of the contrastive signals between the classes resulting from the class-wise gradient matching strategy. To address this problem, we propose Dataset Condensation with Contrastive signals (DCC) by modifying the loss function to enable the DC methods to effectively capture the differences between classes. In addition, we analyze the new loss function in terms of training dynamics by tracking the kernel velocity. Furthermore, we introduce a bi-level warm-up strategy to stabilize the optimization. Our experimental results indicate that while the existing methods are ineffective for fine-grained image classification tasks, the proposed method can successfully generate informative synthetic datasets for the same tasks. Moreover, we demonstrate that the proposed method outperforms the baselines even on benchmark datasets such as SVHN, CIFAR-10, and CIFAR-100. Finally, we demonstrate the high applicability of the proposed method by applying it to continual learning tasks. Saehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun, Sungroh Yoon |
ICML | 4 |
| 2022 | A Unified Analysis of Mixed Sample Data Augmentation: A Loss Function PerspectiveabstractWe propose the first unified theoretical analysis of mixed sample data augmentation (MSDA), such as Mixup and CutMix. Our theoretical results show that regardless of the choice of the mixing strategy, MSDA behaves as a pixel-level regularization of the underlying training loss and a regularization of the first layer parameters. Similarly, our theoretical results support that the MSDA training strategy can improve adversarial robustness and generalization compared to the vanilla training strategy. Using the theoretical results, we provide a high-level understanding of how different design choices of MSDA work differently. For example, we show that the most popular MSDA methods, Mixup and CutMix, behave differently, e.g., CutMix regularizes the input gradients by pixel distances, while Mixup regularizes the input gradients regardless of pixel distances. Our theoretical results also show that the optimal MSDA strategy depends on tasks, datasets, or model parameters. From these observations, we propose generalized MSDAs, a Hybrid version of Mixup and CutMix (HMix) and Gaussian Mixup (GMix), simple extensions of Mixup and CutMix. Our implementation can leverage the advantages of Mixup and CutMix, while our implementation is very efficient, and the computation cost is almost neglectable as Mixup and CutMix. Our empirical study shows that our HMix and GMix outperform the previous state-of-the-art MSDA methods in CIFAR-100 and ImageNet classification tasks. Chanwoo Park, Sangdoo Yun, Sanghyuk Chun |
NeurIPS | 2 |
| 2021 | Rethinking Channel Dimensions for Efficient Model DesignabstractDesigning an efficient model within the limited computational cost is challenging. We argue the accuracy of a lightweight model has been further limited by the design convention: a stage-wise configuration of the channel dimensions, which looks like a piecewise linear function of the network stage. In this paper, we study an effective channel dimension configuration towards better performance than the convention. To this end, we empirically study how to design a single layer properly by analyzing the rank of the output feature. We then investigate the channel configuration of a model by searching network architectures concerning the channel configuration under the computational cost restriction. Based on the investigation, we propose a simple yet effective channel configuration that can be parameterized by the layer index. As a result, our proposed model following the channel parameterization achieves remarkable performance on ImageNet classification and transfer learning tasks including COCO object detection, COCO instance segmentation, and fine-grained classifications. Code and ImageNet pretrained models are available at https: //github.com/clovaai/rexnet. Dongyoon Han, Sangdoo Yun, Byeongho Heo, Young Joon Yoo |
CVPR | 2 |
| 2021 | Re-Labeling ImageNet: From Single to Multi-Labels, From Global to Localized LabelsabstractImageNet has been the most popular image classification benchmark, but it is also the one with a significant level of label noise. Recent studies have shown that many samples contain multiple classes, despite being assumed to be a single-label benchmark. They have thus proposed to turn ImageNet evaluation into a multi-label task, with exhaustive multi-label annotations per image. However, they have not fixed the training set, presumably because of a formidable annotation cost. We argue that the mismatch between single-label annotations and effectively multi-label images is equally, if not more, problematic in the training setup, where random crops are applied. With the single-label annotations, a random crop of an image may contain an entirely different object from the ground truth, introducing noisy or even incorrect supervision during training. We thus re-label the ImageNet training set with multi-labels. We address the annotation cost barrier by letting a strong image classifier, trained on an extra source of data, generate the multi-labels. We utilize the pixel-wise multi-label predictions before the final pooling layer, in order to exploit the additional location-specific supervision signals. Training on the re-labeled samples results in improved model performances across the board. ResNet-50 attains the top-1 accuracy of 78.9% on ImageNet with our localized multi-labels, which can be further boosted to 80.2% with the CutMix regularization. We show that the models trained with localized multi-labels also outperforms the baselines on transfer learning to object detection and instance segmentation tasks, and various robustness benchmarks. The re-labeled ImageNet training set, pre-trained weights, and the source code are available at https://github.com/naverai/relabel_imagenet. Sangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han, Junsuk Choe, Sanghyuk Chun |
CVPR | 1 |
| 2021 | Rethinking Spatial Dimensions of Vision TransformersabstractVision Transformer (ViT) extends the application range of transformers from language processing to computer vision tasks as being an alternative architecture against the existing convolutional neural networks (CNN). Since the transformer-based architecture has been innovative for computer vision modeling, the design convention towards an effective architecture has been less studied yet. From the successful design principles of CNN, we investigate the role of spatial dimension conversion and its effectiveness on transformer-based architecture. We particularly attend to the dimension reduction principle of CNNs; as the depth increases, a conventional CNN increases channel dimension and decreases spatial dimensions. We empirically show that such a spatial dimension reduction is beneficial to a transformer architecture as well, and propose a novel Pooling-based Vision Transformer (PiT) upon the original ViT model. We show that PiT achieves the improved model capability and generalization performance against ViT. Throughout the extensive experiments, we further show PiT outperforms the baseline on several tasks such as image classification, object detection, and robustness evaluation. Source codes and ImageNet models are available at https://github.com/naver-ai/pit. Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, Seong Joon Oh |
ICCV | 2 |
| 2021 | Normalization Matters in Weakly Supervised Object LocalizationabstractWeakly-supervised object localization (WSOL) enables finding an object using a dataset without any localization information. By simply training a classification model using only image-level annotations, the feature map of the model can be utilized as a score map for localization. In spite of many WSOL methods proposing novel strategies, there has not been any de facto standard about how to normalize the class activation map (CAM). Consequently, many WSOL methods have failed to fully exploit their own capacity because of the misuse of a normalization method. In this paper, we review many existing normalization methods and point out that they should be used according to the property of the given dataset. Additionally, we propose a new normalization method which substantially enhances the performance of any CAM-based WSOL methods. Using the proposed normalization method, we provide a comprehensive evaluation over three datasets (CUB, ImageNet and OpenImages) on three different architectures and observe significant performance gains over the conventional min-max normalization method in all the evaluated cases (See Fig. 1). Jeesoo Kim, Junsuk Choe, Sangdoo Yun, Nojun Kwak |
ICCV | 3 |
| 2021 | AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, Jung-Woo Ha 0001 |
ICLR | 5 |
| 2021 | Progressive Transmission and Inference of Deep Learning ModelsabstractModern image files are usually progressively transmitted and provide a preview before downloading the entire image for improved user experience to cope with a slow network connection. In this paper, with a similar goal, we propose a progressive transmission framework for deep learning models, especially to deal with the scenario where pre-trained deep learning models are transmitted from servers and executed at user devices (e.g., web browser or mobile). Our progressive transmission allows inferring approximate models in the middle of file delivery, and quickly provide an acceptable intermediate outputs. On the server-side, a deep learning model is divided and progressively transmitted to the user devices. Then, the divided pieces are progressively concatenated to construct approximate models on user devices. Experiments show that our method is computationally efficient without increasing the model size and total transmission time while preserving the model accuracy. We further demonstrate that our method can improve the user experience by providing the approximate models especially in a slow connection. Youngsoo Lee, Sangdoo Yun, Yeonghun Kim, Sunghee Choi |
ICMLA | 2 |
| 2021 | Region-based dropout with attention prior for weakly supervised object localization
Junsuk Choe, Dongyoon Han, Sangdoo Yun, Jung-Woo Ha 0001, Seong Joon Oh, Hyunjung Shim |
Pattern Recognit. | 3 |
| 2020 | Learning De-biased Representations with Biased RepresentationsabstractMany machine learning algorithms are trained and evaluated by splitting data from a single source into training and test sets. While such focus on in-distribution learning scenarios has led to interesting advancement, it has not been able to tell if models are relying on dataset biases as shortcuts for successful prediction (e.g., using snow cues for recognising snowmobiles), resulting in biased models that fail to generalise when the bias shifts to a different class. The cross-bias generalisation problem has been addressed by de-biasing training data through augmentation or re-sampling, which are often prohibitive due to the data collection cost (e.g., collecting images of a snowmobile on a desert) and the difficulty of quantifying or expressing biases in the first place. In this work, we propose a novel framework to train a de-biased representation by encouraging it to be different from a set of representations that are biased by design. This tactic is feasible in many scenarios where it is much easier to define a set of biased representations than to define and quantify bias. We demonstrate the efficacy of our method across a variety of synthetic and real-world biases; our experiments show that the method discourages models from taking bias shortcuts, resulting in improved generalisation. Source code is available at https://github.com/clovaai/rebias. Hyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo, Seong Joon Oh |
ICML | 3 |
| 2019 | Knowledge Distillation with Adversarial Samples Supporting Decision BoundaryabstractMany recent works on knowledge distillation have provided ways to transfer the knowledge of a trained network for improving the learning process of a new one, but finding a good technique for knowledge distillation is still an open problem. In this paper, we provide a new perspective based on a decision boundary, which is one of the most important component of a classifier. The generalization performance of a classifier is closely related to the adequacy of its decision boundary, so a good classifier bears a good decision boundary. Therefore, transferring information closely related to the decision boundary can be a good attempt for knowledge distillation. To realize this goal, we utilize an adversarial attack to discover samples supporting a decision boundary. Based on this idea, to transfer more accurate information about the decision boundary, the proposed algorithm trains a student classifier based on the adversarial samples supporting the decision boundary. Experiments show that the proposed method indeed improves knowledge distillation and achieves the state-of-the-arts performance. Byeongho Heo, Minsik Lee 0001, Sangdoo Yun, Jin Young Choi 0002 |
AAAI | 3 |
| 2019 | Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden NeuronsabstractAn activation boundary for a neuron refers to a separating hyperplane that determines whether the neuron is activated or deactivated. It has been long considered in neural networks that the activations of neurons, rather than their exact output values, play the most important role in forming classificationfriendly partitions of the hidden feature space. However, as far as we know, this aspect of neural networks has not been considered in the literature of knowledge transfer. In this paper, we propose a knowledge transfer method via distillation of activation boundaries formed by hidden neurons. For the distillation, we propose an activation transfer loss that has the minimum value when the boundaries generated by the student coincide with those by the teacher. Since the activation transfer loss is not differentiable, we design a piecewise differentiable loss approximating the activation transfer loss. By the proposed method, the student learns a separating boundary between activation region and deactivation region formed by each neuron in the teacher. Through the experiments in various aspects of knowledge transfer, it is verified that the proposed method outperforms the current state-of-the-art. Byeongho Heo, Minsik Lee 0001, Sangdoo Yun, Jin Young Choi 0002 |
AAAI | 3 |
| 2019 | Character Region Awareness for Text DetectionabstractScene text detection methods based on neural networks have emerged recently and have shown promising results. Previous methods trained with rigid word-level bounding boxes exhibit limitations in representing the text region in an arbitrary shape. In this paper, we propose a new scene text detection method to effectively detect text area by exploring each character and affinity between characters. To overcome the lack of individual character level annotations, our proposed framework exploits both the given character-level annotations for synthetic images and the estimated character-level ground-truths for real images acquired by the learned interim model. In order to estimate affinity between characters, the network is trained with the newly proposed representation for affinity. Extensive experiments on six benchmarks, including the TotalText and CTW-1500 datasets which contain highly curved texts in natural images, demonstrate that our character-level text detection significantly outperforms the state-of-the-art detectors. According to the results, our proposed method guarantees high flexibility in detecting complicated scene text images, such as arbitrarily-oriented, curved, or deformed texts. Youngmin Baek, Bado Lee, Dongyoon Han, Sangdoo Yun, Hwalsuk Lee |
CVPR | 4 |
| 2019 | What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisabstractMany new proposals for scene text recognition (STR) models have been introduced in recent years. While each claim to have pushed the boundary of the technology, a holistic and fair comparison has been largely missing in the field due to the inconsistent choices of training and evaluation datasets. This paper addresses this difficulty with three major contributions. First, we examine the inconsistencies of training and evaluation datasets, and the performance gap results from inconsistencies. Second, we introduce a unified four-stage STR framework that most existing STR models fit into. Using this framework allows for the extensive evaluation of previously proposed STR modules and the discovery of previously unexplored module combinations. Third, we analyze the module-wise contributions to performance in terms of accuracy, speed, and memory demand, under one consistent set of training and evaluation datasets. Such analyses clean up the hindrance on the current comparisons to understand the performance gain of the existing modules. Our code is publicly available. Jeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park, Dongyoon Han, Sangdoo Yun, Seong Joon Oh, Hwalsuk Lee |
ICCV | 6 |
| 2019 | A Comprehensive Overhaul of Feature DistillationabstractWe investigate the design aspects of feature distillation methods achieving network compression and propose a novel feature distillation method in which the distillation loss is designed to make a synergy among various aspects: teacher transform, student transform, distillation feature position and distance function. Our proposed distillation loss includes a feature transform with a newly designed margin ReLU, a new distillation feature position, and a partial L2distance function to skip redundant information giving adverse effects to the compression of student. In ImageNet, our proposed method achieves 21.65% of top-1 error with ResNet50, which outperforms the performance of the teacher network, ResNet152. Our proposed method is evaluated on various tasks such as image classification, object detection and semantic segmentation and achieves a significant performance improvement in all tasks. The code is available at bhheo.github.io/overhaul. Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 0001, Nojun Kwak, Jin Young Choi 0002 |
ICCV | 3 |
| 2019 | CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesabstractRegional dropout strategies have been proposed to enhance performance of convolutional neural network classifiers. They have proved to be effective for guiding the model to attend on less discriminative parts of objects (e.g. leg as opposed to head of a person), thereby letting the network generalize better and have better object localization capabilities. On the other hand, current methods for regional dropout removes informative pixels on training images by overlaying a patch of either black pixels or random noise. Such removal is not desirable because it suffers from information loss causing inefficiency in training. We therefore propose the CutMix augmentation strategy: patches are cut and pasted among training images where the ground truth labels are also mixed proportionally to the area of the patches. By making efficient use of training pixels and retaining the regularization effect of regional dropout, CutMix consistently outperforms state-of-the-art augmentation strategies on CIFAR and ImageNet classification tasks, as well as on ImageNet weakly-supervised localization task. Moreover, unlike previous augmentation methods, our CutMix-trained ImageNet classifier, when used as a pretrained model, results in consistent performance gain in Pascal detection and MS-COCO image captioning benchmarks. We also show that CutMix can improve the model robustness against input corruptions and its out-of distribution detection performance. Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Young Joon Yoo, Junsuk Choe |
ICCV | 1 |
| 2018 | Context-Aware Deep Feature Compression for High-Speed Visual TrackingabstractWe propose a new context-aware correlation filter based tracking framework to achieve both high computational speed and state-of-the-art performance among real-time trackers. The major contribution to the high computational speed lies in the proposed deep feature compression that is achieved by a context-aware scheme utilizing multiple expert auto-encoders; a context in our framework refers to the coarse category of the tracking target according to appearance patterns. In the pre-training phase, one expert auto-encoder is trained per category. In the tracking phase, the best expert auto-encoder is selected for a given target, and only this auto-encoder is used. To achieve high tracking performance with the compressed feature map, we introduce extrinsic denoising processes and a new orthogonality loss term for pre-training and fine-tuning of the expert autoencoders. We validate the proposed context-aware framework through a number of experiments, where our method achieves a comparable performance to state-of-the-art trackers which cannot run in real-time, while running at a significantly fast speed of over 100 fps. Jongwon Choi 0002, Hyung Jin Chang, Tobias Fischer 0001, Sangdoo Yun, Kyuewang Lee, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 4 |
| 2018 | Unsupervised Holistic Image Generation from Key Local Patches
Donghoon Lee 0004, Sangdoo Yun, Hwiyeon Yoo, Ming-Hsuan Yang 0001, Songhwai Oh |
ECCV (5) | 2 |
| 2018 | Selective Ensemble Network for Accurate Crowd Density EstimationabstractThis paper proposes a selective ensemble deep network architecture for crowd density estimation and people counting. In contrast to existing deep network-based methods, the proposed method incorporates two sub-networks for local density estimation: one to learn sparse density regions and one to learn dense density regions. Locally estimated density maps from the two sub-networks are selectively combined in ensemble fashion using a gating network to estimate an initial crowd density map. The initial density map is refined as a high resolution map, using another sub-network that draws on contextual information in the image. In training, a novel adaptive loss scheme is applied to resolve an ambiguity in the crowded region. the proposed scheme improves both density map accuracy and counting accuracy by adjusting the weighting value between density loss and counting loss according to the degree of crowdness and training epochs. Experiments using public datasets confirm that the proposed method outperforms state-of-the-art methods. Through self-evaluation, the effectiveness of each part in the network is also verified. Jiyeoup Jeong, Hawook Jeong, Jongin Lim 0002, Jongwon Choi 0002, Sangdoo Yun, Jin Young Choi 0002 |
ICPR | 5 |
| 2018 | Action-Driven Visual Object Tracking With Deep Reinforcement LearningabstractIn this paper, we propose an efficient visual tracker, which directly captures a bounding box containing the target object in a video by means of sequential actions learned using deep neural networks. The proposed deep neural network to control tracking actions is pretrained using various training video sequences and fine-tuned during actual tracking for online adaptation to a change of target and background. The pretraining is done by utilizing deep reinforcement learning (RL) as well as supervised learning. The use of RL enables even partially labeled data to be successfully utilized for semisupervised learning. Through the evaluation of the object tracking benchmark data set, the proposed tracker is validated to achieve a competitive performance at three times the speed of existing deep network-based trackers. The fast version of the proposed method, which operates in real time on graphics processing unit, outperforms the state-of-the-art real-time trackers with an accuracy improvement of more than 8%. Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Attentional Correlation Filter Network for Adaptive Visual TrackingabstractWe propose a new tracking framework with an attentional mechanism that chooses a subset of the associated correlation filters for increased robustness and computational efficiency. The subset of filters is adaptively selected by a deep attentional network according to the dynamic properties of the tracking target. Our contributions are manifold, and are summarised as follows: (i) Introducing the Attentional Correlation Filter Network which allows adaptive tracking of dynamic targets. (ii) Utilising an attentional network which shifts the attention to the best candidate modules, as well as predicting the estimated accuracy of currently inactive modules. (iii) Enlarging the variety of correlation filters which cover target drift, blurriness, occlusion, scale changes, and flexible aspect ratio. (iv) Validating the robustness and efficiency of the attentional mechanism for visual tracking through a number of experiments. Our method achieves similar performance to non real-time trackers, and state-of-the-art performance amongst real-time trackers. Jongwon Choi 0002, Hyung Jin Chang, Sangdoo Yun, Tobias Fischer 0001, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 3 |
| 2017 | Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex ManifoldabstractThis paper proposes a new high dimensional regression method by merging Gaussian process regression into a variational autoencoder framework. In contrast to other regression methods, the proposed method focuses on the case where output responses are on a complex high dimensional manifold, such as images. Our contributions are summarized as follows: (i) A new regression method estimating high dimensional image responses, which is not handled by existing regression algorithms, is proposed. (ii) The proposed regression method introduces a strategy to learn the latent space as well as the encoder and decoder so that the result of the regressed response in the latent space coincide with the corresponding response in the data space. (iii) The proposed regression is embedded into a generative model, and the whole procedure is developed by the variational autoencoder framework. We demonstrate the robustness and effectiveness of our method through a number of experiments on various visual data regression problems. Young Joon Yoo, Sangdoo Yun, Hyung Jin Chang, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 2 |
| 2017 | Action-Decision Networks for Visual Tracking with Deep Reinforcement LearningabstractThis paper proposes a novel tracker which is controlled by sequentially pursuing actions learned by deep reinforcement learning. In contrast to the existing trackers using deep networks, the proposed tracker is designed to achieve a light computation as well as satisfactory tracking accuracy in both location and scale. The deep network to control actions is pre-trained using various training sequences and fine-tuned during tracking for online adaptation to target and background changes. The pre-training is done by utilizing deep reinforcement learning as well as supervised learning. The use of reinforcement learning enables even partially labeled data to be successfully utilized for semi-supervised learning. Through evaluation of the OTB dataset, the proposed tracker is validated to achieve a competitive performance that is three times faster than state-of-the-art, deep network-based trackers. The fast version of the proposed method, which operates in real-time on GPU, outperforms the state-of-the-art real-time trackers. Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002 |
CVPR | 1 |
| 2016 | Visual Path Prediction in Complex Scenes with Crowded Moving ObjectsabstractThis paper proposes a novel path prediction algorithm for progressing one step further than the existing works focusing on single target path prediction. In this paper, we consider moving dynamics of co-occurring objects for path prediction in a scene that includes crowded moving objects. To solve this problem, we first suggest a two-layered probabilistic model to find major movement patterns and their cooccurrence tendency. By utilizing the unsupervised learning results from the model, we present an algorithm to find the future location of any target object. Through extensive qualitative/quantitative experiments, we show that our algorithm can find a plausible future path in complex scenes with a large number of moving objects. Young Joon Yoo, Kimin Yun, Sangdoo Yun, Jonghee Hong, Hawook Jeong, Jin Young Choi 0002 |
CVPR | 3 |
| 2016 | Attention-inspired moving object detection in monocular dashcam videosabstractThis paper proposes a moving object detection algorithm for a monocular dashcam mounted on a vehicle. To deal with dynamic changes of the scene from the dashcam, we propose a new scheme inspired by human-attention inclination for change detection. Humans do not build a detailed visual representation and perceive a change of the scene based on the structure of an interesting region. In this perspective, our method focuses on a sky and road region of the scene and builds an abstracted background model, which is updated with a spatially adaptive learning rate according to the center-focused tendency of the human gaze. To improve the robustness of detection, the final detection map is refined by combining the results from twin processes applied to the original image and the median-filtered image, respectively. In experiments, we have found that our method outperforms state-of-the-art methods qualitatively and quantitatively on a realistic dashcam video. Kimin Yun, Jongin Lim 0002, Sangdoo Yun, Soo Wan Kim, Jin Young Choi 0002 |
ICPR | 3 |
| 2016 | Voting-based 3D object cuboid detection robust to partial occlusion from RGB-D imagesabstractIn this paper, we propose a novel algorithm for 3D object cuboid detection. Contrary to the conventional algorithms based on image segmentation, we propose a part-based voting process to robustly generate cuboids when the object is partially occluded. Our method finds the distinctive parts of RGB-D images and generates the 3D cuboids covering the target objects from the distinctive parts by the proposed probabilistic voting model. To validate the performance of the proposed method, experiments are conducted on the challenging NYU v2 and SUN RGB-D datasets. Experimental results show that our method is computationally efficient and has a competitive performance compared with the state-of-the-art methods. In addition, our method can be combined with the conventional segmentation-based method parallelly, and the combined algorithm is evaluated by experiments to show that it achieves a significant improvement of performance. Sangdoo Yun, Hawook Jeong, Soo Wan Kim, Jin Young Choi 0002 |
WACV | 1 |
| 2015 | Category Attentional Search for Fast Object Detection by Mimicking Human Visual PerceptionabstractIn this paper, we propose a novel selective search method to speed up the object detection via category-based attention scheme. The proposed attentional searching strategy is designed to focus on a small set of selected regions where the object category is expected to exist. The selected regions are estimated by mimicking three properties of the attentional scheme of human visual perception: spotlighting interest regions with low-level saliency (saliency attention), focusing on distinctive features for an object category (feature attention), and estimating potential object position by following human gaze path (gaze attention). Also, the time complexity of each attentional scheme is implemented to be low so that it can hardly affect the computational time. To validate the performance of our method, experiments were conducted on the challenging PASCAL VOC dataset. Experimental results show that our method efficiently generates a small number of candidate boxes for object detection (less than 10ms=image), and the combined object detection system achieves more than 2 times faster performance than the baseline with comparable average precision. Hawook Jeong, Sangdoo Yun, Kwang Moo Yi, Jin Young Choi 0002 |
WACV | 2 |
| 2014 | Multi-task learning with over-sampled time-series representation of a trajectory for traffic motion pattern recognitionabstractThis paper proposes an efficient feature sampling and multi-task learning scheme for traffic scene analysis, where all classifiers are trained simultaneously by exploiting the correlations among different motion patterns. We make feature descriptors by high dimensional embedding of the time series data for traffic pattern representation. They preserve detailed spatio-temporal information of the underlying event. Pattern specific details are extracted from raw trajectories and embedded into feature descriptors, which ensures their great discriminability. Training data scarcity problem is tackled through amplification of the patterns hidden in raw trajectory via strategic oversampling and employment of joint feature selection procedure while training the models. Experimental results on 4 surveillance datasets, show great improvement in the motion pattern recognition performance, importance of joint feature selection and fast incremental learning ability of the proposed framework. Tushar Sandhan, Young Joon Yoo, Hanjoo Yoo, Sangdoo Yun, Moonsub Byeon |
AVSS | 4 |
| 2014 | Visual surveillance briefing system: Event-based video retrieval and summarizationabstractThis paper presents a visual surveillance briefing (VSB) system which provides event-based retrieval and briefing functions. Traditional event-based video retrieval systems usually aim to analyze the appearance of objects rather than the motion information (e.g. trajectory) of objects. The VSB system adopts the video summarization technique which temporally abstracts the retrieved events to understand the motion patterns of objects. We propose various event features including object's appearances and motion patterns for the purpose of event retrieval and design the energy function to abstract the retrieved events in real-time. To avoid the occlusion problem in the briefed events, we propose an animated displaying method that separately presents the global motion and the local motion of moving objects. Effectiveness of the implemented VSB system is evaluated through several surveillance videos. Sangdoo Yun, Kimin Yun, Soo Wan Kim, Young Joon Yoo, Jiyeoup Jeong |
AVSS | 1 |
| 2014 | Self-Organizing Cascaded Structure of Deformable Part Models for Fast Object DetectionabstractIn this paper, we propose a framework which self-organizes the cascaded object detection filters for fast object detection with maintaining high accuracy. The proposed scheme consists of root and part filter modules, which are cascaded in a self-organizing structure. The pruning of non-object regions in low resolution at the root cascade stage is critical for the object detection speed. At root stage, to prune as many non-object regions as possible, we build a root cascade structure using multiple root models. These models are obtained via bagging procedure for non-linear classification of object and non-object parts in an image. Additional speed-up is achieved by determining proper deployment order of part models. We define a discriminability measure for the part models and suggest a self-organizing scheme to generate an efficient order of part models. The proposed method is evaluated through computational experiments with the PASCAL VOC and INRIA datasets, as a result, our method achieves on average more than 2 times faster performance than the original cascade-DPM, with comparable precision scores. Sangdoo Yun, Hawook Jeong, Woo-Sung Kang, Byeongho Heo, Jin Young Choi 0002 |
ICPR | 1 |