VLDB 2026 Research / reviewers in the wild / expert
Dongyoon Han
dblp:151/8876
· DBLP profile ↗
50ranked-venue papers
6as first author
39since 2021 · last 2026
0000-0002-9130-8195ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 6 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 4 first-author · 24 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code GenerationabstractEffective code generation requires both model capability and a problem representation that carefully structures how models reason and plan.Existing approaches augment reasoning steps or inject specific structure into how models think, but leave scattered problem conditions unchanged.Inspired by the way humans organize fragmented information into coherent explanations, we propose STORYCODER, a narrative reformulation framework that transforms code generation questions into coherent natural language narratives, providing richer contextual structure than simple rephrasings.Each narrative consists of three components: a task overview, constraints, and example test cases, guided by the selected algorithm and genre.Experiments across 11 models on HumanEval, LiveCodeBench, and CodeForces demonstrate consistent improvements, with an average gain of 18.7% in zero-shot [email protected] accuracy, our analyses reveal that narrative reformulation guides models toward correct algorithmic strategies, reduces implementation errors, and induces a more modular code structure.The analyses further show that these benefits depend on narrative coherence and genre alignment, suggesting that structured problem representation is important for code generation regardless of model scale or architecture.Our code is available here. Geonhui Jang, Dongyoon Han, Young Joon Yoo |
ACL (1) | 2 |
| 2025 | Masking meets Supervision: A Strong Learning AllianceabstractPre-training with random masked inputs has emerged as a novel trend in self-supervised training. However, supervised learning still faces a challenge in adopting masking augmentations, primarily due to unstable training. In this paper, we propose a novel way to involve masking augmentations dubbed Masked Sub-branch (MaskSub). MaskSub consists of the main-branch and sub-branch, the latter being a part of the former. The main-branch undergoes conventional training recipes, while the sub-branch merits intensive masking augmentations, during training. MaskSub tackles the challenge by mitigating adverse effects through a relaxed loss function similar to a self-distillation loss. Our analysis shows that MaskSub improves performance, with the training loss converging faster than in standard training, which suggests our method stabilizes the training process. We further validate MaskSub across diverse training scenarios and models, including DeiT-III training, MAE finetuning, CLIP finetuning, BERT training, and hierarchical architectures (ResNet and Swin Transformer). Our results show that MaskSub consistently achieves impressive performance gains across all the cases. MaskSub provides a practical and effective solution for introducing additional regularization under various training recipes. Code available at https://github.com/naver-ai/augsub Byeongho Heo, Taekyung Kim 0002, Sangdoo Yun, Dongyoon Han |
CVPR | 4 |
| 2025 | Morphing Tokens Draw Strong Masked Image ModelsabstractMasked image modeling (MIM) has emerged as a promising approach for pre-training Vision Transformers (ViTs). MIMs predict masked tokens token-wise to recover target signals that are tokenized from images or generated by pre-trained models like vision-language models. While using tokenizers or pre-trained models is viable, they often offer spatially inconsistent supervision even for neighboring tokens, hindering models from learning discriminative representations. Our pilot study identifies spatial inconsistency in supervisory signals and suggests that addressing it can improve representation learning. Building upon this insight, we introduce Dynamic Token Morphing (DTM), a novel method that dynamically aggregates tokens while preserving context to generate contextualized targets, thereby likely reducing spatial inconsistency. DTM is compatible with various SSL frameworks; we showcase significantly improved MIM results, barely introducing extra training costs. Our method facilitates MIM training by using more spatially consistent targets, resulting in improved training trends as evidenced by lower losses. Experiments on ImageNet-1K and ADE20K demonstrate DTM's superiority, which surpasses complex state-of-the-art MIM methods. Furthermore, the evaluation of transfer learning on downstream tasks like iNaturalist, along with extensive empirical studies, supports DTM's effectiveness. Taekyung Kim 0002, Byeongho Heo, Dongyoon Han |
ICLR | 3 |
| 2025 | Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language ModelsabstractWith the rapid advancement of test-time compute search strategies to improve the mathematical problem-solving capabilities of large language models (LLMs), the need for building robust verifiers has become increasingly important. However, all these inference strategies rely on existing verifiers originally designed for Best-of-N search, which makes them sub-optimal for tree search techniques at test time. During tree search, existing verifiers can only offer indirect and implicit assessments of partial solutions or under-value prospective intermediate steps, thus resulting in the premature pruning of promising intermediate steps. To overcome these limitations, we propose token-supervised value models (TVMs) -- a new class of verifiers that assign each token a probability that reflects the likelihood of reaching the correct final answer. This new token-level supervision enables TVMs to directly and explicitly evaluate partial solutions, effectively distinguishing between promising and incorrect intermediate steps during tree search at test time. Experimental results demonstrate that combining tree-search-based inference strategies with TVMs significantly improves the accuracy of LLMs in mathematical problem-solving tasks, surpassing the performance of existing verifiers. Jung Hyun Lee, June Yong Yang, Byeongho Heo, Dongyoon Han, Eunho Yang, Kang Min Yoo |
ICLR | 4 |
| 2025 | DaWin: Training-free Dynamic Weight Interpolation for Robust AdaptationabstractAdapting a pre-trained foundation model on downstream tasks should ensure robustness against distribution shifts without the need to retrain the whole model. Although existing weight interpolation methods are simple yet effective, we argue their static nature limits downstream performance while achieving efficiency. In this work, we propose DaWin, a training-free dynamic weight interpolation method that leverages the entropy of individual models over each unlabeled test sample to assess model expertise, and compute per-sample interpolation coefficients dynamically. Unlike previous works that typically rely on additional training to learn such coefficients, our approach requires no training. Then, we propose a mixture modeling approach that greatly reduces inference overhead raised by dynamic interpolation. We validate DaWin on the large-scale visual recognition benchmarks, spanning 14 tasks across robust fine-tuning -- ImageNet and derived five distribution shift benchmarks -- and multi-task learning with eight classification tasks. Results demonstrate that DaWin achieves significant performance gain in considered settings, with minimal computational overhead. We further discuss DaWin's analytic behavior to explain its empirical success. Changdae Oh, Yixuan Li 0001, Kyungwoo Song, Sangdoo Yun, Dongyoon Han |
ICLR | 5 |
| 2025 | Peri-LN: Revisiting Normalization Layer in the Transformer ArchitectureabstractSelecting a layer normalization (LN) strategy that stabilizes training and speeds convergence in Transformers remains difficult, even for today’s large language models (LLM). We present a comprehensive analytical foundation for understanding how different LN strategies influence training dynamics in large-scale Transformers. Until recently, Pre-LN and Post-LN have long dominated practices despite their limitations in large-scale training. However, several open-source models have recently begun silently adopting a third strategy without much explanation. This strategy places normalization layer peripherally around sublayers, a design we term Peri-LN. While Peri-LN has demonstrated promising performance, its precise mechanisms and benefits remain almost unexplored. Our in-depth analysis delineates the distinct behaviors of LN strategies, showing how each placement shapes activation variance and gradient propagation. To validate our theoretical insight, we conduct extensive experiments on Transformers up to $3.2$B parameters, showing that Peri-LN consistently achieves more balanced variance growth, steadier gradient flow, and convergence stability. Our results suggest that Peri-LN warrants broader consideration for large-scale Transformer architectures, providing renewed insights into the optimal placement of LN. Byeongchan Lee 0001, Cheonbok Park, Yeontaek Oh, Taehwan Yoo, Seongjin Shin, Dongyoon Han, Jinwoo Shin, Kang Min Yoo |
ICML | 8 |
| 2025 | NegMerge: Sign-Consensual Weight Merging for Machine UnlearningabstractMachine unlearning aims to selectively remove specific knowledge from a trained model. Existing approaches, such as Task Arithmetic, fine-tune the model on the forget set to create a task vector (i.e., a direction in weight space) for subtraction from the original model's weight. However, their effectiveness is highly sensitive to hyperparameter selection, requiring extensive validation to identify the optimal vector from many fine-tuned candidates. In this paper, we propose a novel method that utilizes all fine-tuned models trained with varying hyperparameters instead of a single selection. Specifically, we aggregate the computed task vectors by retaining only the elements with consistent shared signs. The merged task vector is then negated to induce unlearning on the original model. Evaluations on zero-shot and standard image recognition tasks across twelve datasets and four backbone architectures show that our approach outperforms state-of-the-art methods while requiring similar or fewer computational resources. Code is available at https://github.com/naver-ai/negmerge. Hyoseo Kim, Dongyoon Han, Junsuk Choe |
ICML | 2 |
| 2025 | Token Bottleneck: One Token to Remember DynamicsabstractDeriving compact and temporally aware visual representations from dynamic scenes is essential for successful execution of sequential scene understanding tasks such as visual tracking and robotic manipulation. In this paper, we introduce Token Bottleneck (ToBo), a simple yet intuitive self-supervised learning pipeline that squeezes a scene into a bottleneck token and predicts the subsequent scene using minimal patches as hints. The ToBo pipeline facilitates the learning of sequential scene representations by conservatively encoding the reference scene into a compact bottleneck token during the squeeze step. In the expansion step, we guide the model to capture temporal dynamics by predicting the target scene using the bottleneck token along with few target patches as hints. This design encourages the vision backbone to embed temporal dependencies, thereby enabling understanding of dynamic transitions across scenes. Extensive experiments in diverse sequential tasks, including video label propagation and robot manipulation in simulated environments demonstrate the superiority of ToBo over baselines. Moreover, deploying our pre-trained model on physical robots confirms its robustness and effectiveness in real-world environments. We further validate the scalability of ToBo across different model scales. Code is available at https://github.com/naver-ai/tobo. Taekyung Kim 0002, Dongyoon Han, Byeongho Heo, Jeongeun Park 0002, Sangdoo Yun |
NeurIPS | 2 |
| 2024 | Match Me If You Can: Semi-supervised Semantic Correspondence Learning with Unpaired Images
Byeongho Heo, Sangdoo Yun, Seungryong Kim, Dongyoon Han |
ACCV (6) | 5 |
| 2024 | Rotary Position Embedding for Vision Transformer
Byeongho Heo, Song Park, Dongyoon Han, Sangdoo Yun |
ECCV (10) | 3 |
| 2024 | Similarity of Neural Architectures Using Adversarial Attack Transferability
Jaehui Hwang, Dongyoon Han, Byeongho Heo, Song Park, Sanghyuk Chun, Jong-Seok Lee |
ECCV (68) | 2 |
| 2024 | Model Stock: All We Need Is Just a Few Fine-Tuned Models
Dong-Hwan Jang, Sangdoo Yun, Dongyoon Han |
ECCV (44) | 3 |
| 2024 | Learning with Unmasked Tokens Drives Stronger Vision Learners
Taekyung Kim 0002, Sanghyuk Chun, Byeongho Heo, Dongyoon Han |
ECCV (34) | 4 |
| 2024 | HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
Wonjae Kim, Sanghyuk Chun, Taekyung Kim 0002, Dongyoon Han, Sangdoo Yun |
ECCV (40) | 4 |
| 2024 | DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
Byeongho Heo, Dongyoon Han |
ECCV (3) | 3 |
| 2024 | Leveraging Temporal Contextualization for Video Action Recognition
Minji Kim 0002, Dongyoon Han, Taekyung Kim 0002, Bohyung Han |
ECCV (21) | 2 |
| 2024 | SeiT++: Masked Token Modeling Improves Storage-Efficient Training
Minhyun Lee, Song Park, Byeongho Heo, Dongyoon Han, Hyunjung Shim |
ECCV (28) | 4 |
| 2024 | Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsabstractImproving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation. Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alex Hauptmann 0001, Zhi-Qi Cheng, Kyungwoo Song |
NeurIPS | 4 |
| 2023 | Frequency Selective Augmentation for Video Representation LearningabstractRecent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a mechanism to prevent representation learning from bias towards static information in the video. In this paper, we propose frequency augmentation (FreqAug), a spatio-temporal data augmentation method in the frequency domain for video representation learning. FreqAug stochastically removes specific frequency components from the video so that learned representation captures essential features more from the remaining information for various downstream tasks. Specifically, FreqAug pushes the model to focus more on dynamic features rather than static features in the video via dropping spatial or temporal low-frequency components. To verify the generality of the proposed method, we experiment with FreqAug on multiple self-supervised learning frameworks along with standard augmentations. Transferring the improved representation to five video action recognition and two temporal action localization downstream tasks shows consistent improvements over baselines. Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, Junmo Kim 0002 |
AAAI | 4 |
| 2023 | Can We Find Strong Lottery Tickets in Generative Models?abstractYes. In this paper, we investigate strong lottery tickets in generative models, the subnetworks that achieve good generative performance without any weight update. Neural network pruning is considered the main cornerstone of model compression for reducing the costs of computation and memory. Unfortunately, pruning a generative model has not been extensively explored, and all existing pruning algorithms suffer from excessive weight-training costs, performance degradation, limited generalizability, or complicated training. To address these problems, we propose to find a strong lottery ticket via moment-matching scores. Our experimental results show that the discovered subnetwork can perform similarly or better than the trained dense model even when only 10% of the weights remain. To the best of our knowledge, we are the first to show the existence of strong lottery tickets in generative models and provide an algorithm to find it stably. Our code and supplementary materials are publicly available at https://lait-cvlab.github.io/SLT-in-Generative-Models/. Sangyeop Yeo, Yoojin Jang 0001, Jy-yong Sohn, Dongyoon Han, Jaejun Yoo 0001 |
AAAI | 4 |
| 2023 | The Devil is in the Points: Weakly Semi-Supervised Instance Segmentation via Point-Guided Mask RepresentationabstractIn this paper, we introduce a novel learning scheme named weakly semi-supervised instance segmentation (WS-SIS) with point labels for budget-efficient and high-performance instance segmentation. Namely, we consider a dataset setting consisting of a few fully-labeled images and a lot of point-labeled images. Motivated by the main challenge of semi-supervised approaches mainly derives from the trade-off between false-negative and false-positive instance proposals, we propose a method for WSSIS that can effectively leverage the budget-friendly point labels as a powerful weak supervision source to resolve the challenge. Furthermore, to deal with the hard case where the amount of fully-labeled data is extremely limited, we propose a MaskRefineNet that refines noise in rough masks. We conduct extensive experiments on COCO and BDD 100K datasets, and the proposed method achieves promising results comparable to those of the fully-supervised model, even with 50% of the fully labeled COCO data (38.8% vs. 39.7%). Moreover, when using as little as 5% of fully labeled COCO data, our method shows significantly superior performance over the state-of-the-art semi-supervised learning method (33.7% vs. 24.9%). The code is available at https://github.com/clovaai/PointWSSIS. Beomyoung Kim, Joonhyun Jeong, Dongyoon Han, Sung Ju Hwang |
CVPR | 3 |
| 2023 | Neglected Free Lunch - Learning Image Classifiers Using Annotation ByproductsabstractSupervised learning of image classifiers distills human knowledge into a parametric model fθthrough pairs of images and corresponding labels $\left\{ {\left( {{X_i},{Y_i}} \right)} \right\}_{i = 1}^N$. We argue that this simple and widely used representation of human knowledge neglects rich auxiliary information from the annotation procedure, such as the time-series of mouse traces and clicks left after image selection. Our insight is that such annotation byproducts Z provide approximate human attention that weakly guides the model to focus on the foreground cues, reducing spurious correlations and discouraging shortcut learning. To verify this, we create ImageNet-AB and COCO-AB. They are ImageNet and COCO training sets enriched with sample-wise annotation byproducts, collected by replicating the respective original annotation tasks. We refer to the new paradigm of training models with annotation byproducts as learning using annotation byproducts (LUAB). We show that a simple multitask loss for regressing Z together with Y already improves the generalisability and robustness of the learned models. Compared to the original supervised learning, LUAB does not require extra annotation costs. ImageNet-AB and COCO-AB are at github.com/naverai/NeglectedFreeLunch. Dongyoon Han, Junsuk Choe, Seonghyeok Chun, John Joon Young Chung, Minsuk Chang, Sangdoo Yun, Jean Y. Song, Seong Joon Oh |
ICCV | 1 |
| 2023 | Scratching Visual Transformer's Back with Uniform AttentionabstractThe favorable performance of Vision Transformers (ViTs) is often attributed to the multi-head self-attention (MSA), which enables global interactions at each layer of a ViT model. Previous works acknowledge the property of long-range dependency for the effectiveness in MSA. In this work, we study the role of MSA in terms of the different axis, density. Our preliminary analyses suggest that the spatial interactions of learned attention maps are close to dense interactions rather than sparse ones. This is a curious phenomenon because dense attention maps are harder for the model to learn due to softmax. We interpret this opposite behavior against softmax as a strong preference for the ViT models to include dense interaction. We thus manually insert the dense uniform attention to each layer of the ViT models to supply the much-needed dense interactions. We call this method Context Broadcasting, CB. Our study demonstrates the inclusion of CB takes the role of dense attention and thereby reduces the degree of density in the original attention maps by complying softmax in MSA. We also show that, with negligible costs of CB (1 line in your model code and no additional parameters), both the capacity and generalizability of the ViT models are increased. Nam Hyeon-Woo, Kim Yu-Ji, Byeongho Heo, Dongyoon Han, Seong Joon Oh, Tae-Hyun Oh |
ICCV | 4 |
| 2023 | Generating Instance-level Prompts for Rehearsal-free Continual LearningabstractWe introduce Domain-Adaptive Prompt (DAP), a novel method for continual learning using Vision Transformers (ViT). Prompt-based continual learning has recently gained attention due to its rehearsal-free nature. Currently, the prompt pool, which is suggested by prompt-based continual learning, is key to effectively exploiting the frozen pretrained ViT backbone in a sequence of tasks. However, we observe that the use of a prompt pool creates a domain scalability problem between pre-training and continual learning. This problem arises due to the inherent encoding of group-level instructions within the prompt pool. To address this problem, we propose DAP, a pool-free approach that generates a suitable prompt in an instance-level manner at inference time. We optimize an adaptive prompt generator that creates instance-specific fine-grained instructions required for each input, enabling enhanced model plasticity and reduced forgetting. Our experiments on seven datasets with varying degrees of domain similarity to ImageNet demonstrate the superiority of DAP over state-of-the-art prompt-based methods. Code is publicly available at https://github.com/naver-ai/dap-cl. Dahuin Jung, Dongyoon Han, Jihwan Bang, Hwanjun Song |
ICCV | 2 |
| 2023 | Gramian Attention Heads are Strong yet Efficient Vision LearnersabstractWe introduce a novel architecture design that enhances expressiveness by incorporating multiple head classifiers (i.e., classification heads) instead of relying on channel expansion or additional building blocks. Our approach employs attention-based aggregation, utilizing pairwise feature similarity to enhance multiple lightweight heads with minimal resource overhead. We compute the Gramian matrices to reinforce class tokens in an attention layer for each head. This enables the heads to learn more discriminative representations, enhancing their aggregation capabilities. Furthermore, we propose a learning algorithm that encourages heads to complement each other by reducing correlation for aggregation. Our models eventually surpass state-of-the-art CNNs and ViTs regarding the accuracy-throughput trade-off on ImageNet-1K and deliver remarkable performance across various downstream tasks, such as COCO object instance segmentation, ADE20k semantic segmentation, and fine-grained visual classification datasets. The effectiveness of our framework is substantiated by practical experimental results and further underpinned by generalization error bound. We release the code publicly at: https://github.com/Lab-LVM/imagenet-models. Jong Bin Ryu, Dongyoon Han, Jongwoo Lim |
ICCV | 2 |
| 2023 | GeNAS: Neural Architecture Search with Better GeneralizationabstractNeural Architecture Search (NAS) aims to automatically excavate the optimal network architecture with superior test performance. Recent neural architecture search (NAS) approaches rely on validation loss or accuracy to find the superior network for the target data. In this paper, we investigate a new neural architecture search measure for excavating architectures with better generalization. We demonstrate that the flatness of the loss surface can be a promising proxy for predicting the generalization capability of neural network architectures. We evaluate our proposed method on various search spaces, showing similar or even better performance compared to the state-of-the-art NAS methods. Notably, the resultant architecture found by flatness measure generalizes robustly to various shifts in data distribution (e.g. ImageNet-V2,-A,-O), as well as various tasks such as object detection and semantic segmentation. Joonhyun Jeong, Joonsang Yu, Geondo Park, Dongyoon Han, Young Joon Yoo |
IJCAI | 4 |
| 2023 | Switching Temporary Teachers for Semi-Supervised Semantic SegmentationabstractThe teacher-student framework, prevalent in semi-supervised semantic segmentation, mainly employs the exponential moving average (EMA) to update a single teacher's weights based on the student's. However, EMA updates raise a problem in that the weights of the teacher and student are getting coupled, causing a potential performance bottleneck. Furthermore, this problem may become more severe when training with more complicated labels such as segmentation masks but with few annotated data. This paper introduces Dual Teacher, a simple yet effective approach that employs dual temporary teachers aiming to alleviate the coupling problem for the student. The temporary teachers work in shifts and are progressively improved, so consistently prevent the teacher and student from becoming excessively close. Specifically, the temporary teachers periodically take turns generating pseudo-labels to train a student model and maintain the distinct characteristics of the student model for each epoch. Consequently, Dual Teacher achieves competitive performance on the PASCAL VOC, Cityscapes, and ADE20K benchmarks with remarkably shorter training times than state-of-the-art methods. Moreover, we demonstrate that our approach is model-agnostic and compatible with both CNN- and Transformer-based models. Code is available at https://github.com/naver-ai/dual-teacher. Jaemin Na, Jung-Woo Ha 0001, Hyung Jin Chang, Dongyoon Han, Wonjun Hwang |
NeurIPS | 4 |
| 2023 | TL-ADA: Transferable Loss-based Active Domain Adaptation
Kyeongtak Han, Youngeun Kim, Dongyoon Han, Sungeun Hong |
Neural Networks | 3 |
| 2022 | Demystifying the Neural Tangent Kernel from a Practical Perspective: Can it be trusted for Neural Architecture Search without training?abstractIn Neural Architecture Search (NAS), reducing the cost of architecture evaluation remains one of the most crucial challenges. Among a plethora of efforts to bypass training of each candidate architecture to convergence for evaluation, the Neural Tangent Kernel (NTK) is emerging as a promising theoretical framework that can be utilized to estimate the performance of a neural architecture at initialization. In this work, we revisit several at-initialization metrics that can be derived from the NTK and reveal their key short-comings. Then, through the empirical analysis of the time evolution of NTK, we deduce that modern neural architectures exhibit highly non-linear characteristics, making the NTK-based metrics incapable of reliably estimating the performance of an architecture without some amount of training. To take such non-linear characteristics into account, we introduce Label-Gradient Alignment (LGA), a novel NTK-based metric whose inherent formulation allows it to capture the large amount of non-linear advantage present in modern neural architectures. With minimal amount of training, LGA obtains a meaningful level of rank correlation with the final test accuracy of an architecture. Lastly, we demonstrate that LGA, complemented with few epochs of training, successfully guides existing search algorithms to achieve competitive search performances with significantly less search cost. The code is available at: https://github.com/nute11amok/DemystifyingNTK. Jisoo Mok, Byunggook Na, Dongyoon Han, Sungroh Yoon |
CVPR | 4 |
| 2022 | OCR-Free Document Understanding Transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, Seunghyun Park 0001 |
ECCV (28) | 9 |
| 2022 | Contrastive Vicinal Space for Unsupervised Domain Adaptation
Jaemin Na, Dongyoon Han, Hyung Jin Chang, Wonjun Hwang |
ECCV (34) | 2 |
| 2022 | Learning Features with Parameter-Free Layers
Dongyoon Han, Young Joon Yoo, Beomyoung Kim, Byeongho Heo |
ICLR | 1 |
| 2022 | ViDT: An Efficient and Effective Fully Transformer-based Object Detector
Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, Ming-Hsuan Yang 0001 |
ICLR | 5 |
| 2022 | Time Is MattEr: Temporal Self-supervision for Video TransformersabstractUnderstanding temporal dynamics of video is an essential aspect of learning better video representations. Recently, transformer-based architectural designs have been extensively explored for video tasks due to their capability to capture long-term dependency of input sequences. However, we found that these Video Transformers are still biased to learn spatial dynamics rather than temporal ones, and debiasing the spurious correlation is critical for their performance. Based on the observations, we design simple yet effective self-supervised tasks for video models to learn temporal dynamics better. Specifically, for debiasing the spatial bias, our method learns the temporal order of video frames as extra self-supervision and enforces the randomly shuffled frames to have low-confidence outputs. Also, our method learns the temporal flow direction of video tokens among consecutive frames for enhancing the correlation toward temporal dynamics. Under various video action recognition tasks, we demonstrate the effectiveness of our method and its compatibility with state-of-the-art Video Transformers. Sukmin Yun, Jaehyung Kim 0001, Dongyoon Han, Hwanjun Song, Jung-Woo Ha 0001, Jinwoo Shin |
ICML | 3 |
| 2021 | Rethinking Channel Dimensions for Efficient Model DesignabstractDesigning an efficient model within the limited computational cost is challenging. We argue the accuracy of a lightweight model has been further limited by the design convention: a stage-wise configuration of the channel dimensions, which looks like a piecewise linear function of the network stage. In this paper, we study an effective channel dimension configuration towards better performance than the convention. To this end, we empirically study how to design a single layer properly by analyzing the rank of the output feature. We then investigate the channel configuration of a model by searching network architectures concerning the channel configuration under the computational cost restriction. Based on the investigation, we propose a simple yet effective channel configuration that can be parameterized by the layer index. As a result, our proposed model following the channel parameterization achieves remarkable performance on ImageNet classification and transfer learning tasks including COCO object detection, COCO instance segmentation, and fine-grained classifications. Code and ImageNet pretrained models are available at https: //github.com/clovaai/rexnet. Dongyoon Han, Sangdoo Yun, Byeongho Heo, Young Joon Yoo |
CVPR | 1 |
| 2021 | Re-Labeling ImageNet: From Single to Multi-Labels, From Global to Localized LabelsabstractImageNet has been the most popular image classification benchmark, but it is also the one with a significant level of label noise. Recent studies have shown that many samples contain multiple classes, despite being assumed to be a single-label benchmark. They have thus proposed to turn ImageNet evaluation into a multi-label task, with exhaustive multi-label annotations per image. However, they have not fixed the training set, presumably because of a formidable annotation cost. We argue that the mismatch between single-label annotations and effectively multi-label images is equally, if not more, problematic in the training setup, where random crops are applied. With the single-label annotations, a random crop of an image may contain an entirely different object from the ground truth, introducing noisy or even incorrect supervision during training. We thus re-label the ImageNet training set with multi-labels. We address the annotation cost barrier by letting a strong image classifier, trained on an extra source of data, generate the multi-labels. We utilize the pixel-wise multi-label predictions before the final pooling layer, in order to exploit the additional location-specific supervision signals. Training on the re-labeled samples results in improved model performances across the board. ResNet-50 attains the top-1 accuracy of 78.9% on ImageNet with our localized multi-labels, which can be further boosted to 80.2% with the CutMix regularization. We show that the models trained with localized multi-labels also outperforms the baselines on transfer learning to object detection and instance segmentation tasks, and various robustness benchmarks. The re-labeled ImageNet training set, pre-trained weights, and the source code are available at https://github.com/naverai/relabel_imagenet. Sangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han, Junsuk Choe, Sanghyuk Chun |
CVPR | 4 |
| 2021 | Rethinking Spatial Dimensions of Vision TransformersabstractVision Transformer (ViT) extends the application range of transformers from language processing to computer vision tasks as being an alternative architecture against the existing convolutional neural networks (CNN). Since the transformer-based architecture has been innovative for computer vision modeling, the design convention towards an effective architecture has been less studied yet. From the successful design principles of CNN, we investigate the role of spatial dimension conversion and its effectiveness on transformer-based architecture. We particularly attend to the dimension reduction principle of CNNs; as the depth increases, a conventional CNN increases channel dimension and decreases spatial dimensions. We empirically show that such a spatial dimension reduction is beneficial to a transformer architecture as well, and propose a novel Pooling-based Vision Transformer (PiT) upon the original ViT model. We show that PiT achieves the improved model capability and generalization performance against ViT. Throughout the extensive experiments, we further show PiT outperforms the baseline on several tasks such as image classification, object detection, and robustness evaluation. Source codes and ImageNet models are available at https://github.com/naver-ai/pit. Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, Seong Joon Oh |
ICCV | 3 |
| 2021 | AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, Jung-Woo Ha 0001 |
ICLR | 4 |
| 2021 | Region-based dropout with attention prior for weakly supervised object localization
Junsuk Choe, Dongyoon Han, Sangdoo Yun, Jung-Woo Ha 0001, Seong Joon Oh, Hyunjung Shim |
Pattern Recognit. | 2 |
| 2019 | Character Region Awareness for Text DetectionabstractScene text detection methods based on neural networks have emerged recently and have shown promising results. Previous methods trained with rigid word-level bounding boxes exhibit limitations in representing the text region in an arbitrary shape. In this paper, we propose a new scene text detection method to effectively detect text area by exploring each character and affinity between characters. To overcome the lack of individual character level annotations, our proposed framework exploits both the given character-level annotations for synthetic images and the estimated character-level ground-truths for real images acquired by the learned interim model. In order to estimate affinity between characters, the network is trained with the newly proposed representation for affinity. Extensive experiments on six benchmarks, including the TotalText and CTW-1500 datasets which contain highly curved texts in natural images, demonstrate that our character-level text detection significantly outperforms the state-of-the-art detectors. According to the results, our proposed method guarantees high flexibility in detecting complicated scene text images, such as arbitrarily-oriented, curved, or deformed texts. Youngmin Baek, Bado Lee, Dongyoon Han, Sangdoo Yun, Hwalsuk Lee |
CVPR | 3 |
| 2019 | What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisabstractMany new proposals for scene text recognition (STR) models have been introduced in recent years. While each claim to have pushed the boundary of the technology, a holistic and fair comparison has been largely missing in the field due to the inconsistent choices of training and evaluation datasets. This paper addresses this difficulty with three major contributions. First, we examine the inconsistencies of training and evaluation datasets, and the performance gap results from inconsistencies. Second, we introduce a unified four-stage STR framework that most existing STR models fit into. Using this framework allows for the extensive evaluation of previously proposed STR modules and the discovery of previously unexplored module combinations. Third, we analyze the module-wise contributions to performance in terms of accuracy, speed, and memory demand, under one consistent set of training and evaluation datasets. Such analyses clean up the hindrance on the current comparisons to understand the performance gain of the existing modules. Our code is publicly available. Jeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park, Dongyoon Han, Sangdoo Yun, Seong Joon Oh, Hwalsuk Lee |
ICCV | 5 |
| 2019 | CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesabstractRegional dropout strategies have been proposed to enhance performance of convolutional neural network classifiers. They have proved to be effective for guiding the model to attend on less discriminative parts of objects (e.g. leg as opposed to head of a person), thereby letting the network generalize better and have better object localization capabilities. On the other hand, current methods for regional dropout removes informative pixels on training images by overlaying a patch of either black pixels or random noise. Such removal is not desirable because it suffers from information loss causing inefficiency in training. We therefore propose the CutMix augmentation strategy: patches are cut and pasted among training images where the ground truth labels are also mixed proportionally to the area of the patches. By making efficient use of training pixels and retaining the regularization effect of regional dropout, CutMix consistently outperforms state-of-the-art augmentation strategies on CIFAR and ImageNet classification tasks, as well as on ImageNet weakly-supervised localization task. Moreover, unlike previous augmentation methods, our CutMix-trained ImageNet classifier, when used as a pretrained model, results in consistent performance gain in Pascal detection and MS-COCO image captioning benchmarks. We also show that CutMix can improve the model robustness against input corruptions and its out-of distribution detection performance. Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Young Joon Yoo, Junsuk Choe |
ICCV | 2 |
| 2019 | Learning Receptive Field Size by Learning Filter SizeabstractCovering various receptive fields within a layer is essential to effectively recognize the objects of various sizes and types at a specific layer for a convolutional neural network (CNN). In this work, we propose a novel adaptive learning method which learns the filter size (i.e. the kernel size of a convolutional filter) and distribution to learn the receptive field size. Directly optimizing with respect to the filter size is challenging because the filter size is discrete. To overcome this, we propose a masking technique, which enables the automatic allocation of resources over filters of different sizes and leads to efficient optimization. Through our proposed trainable formulation of the mask, the network self-organizes its structure through the standard backpropagation. The proposed adaptive CNN can be generalized to any single-path structures and multi-path structures as well. The effectiveness of our proposed approach is validated by several benchmark datasets compared with various previous structures on the image classification task for diverse network depths and widths. Furthermore, we demonstrate our adaptive CNN trained on a large-scale dataset can yield improved performance when applying to a transfer learning. Yegang Lee, Heechul Jung, Dongyoon Han, Kyungsu Kim 0003, Junmo Kim 0002 |
WACV | 3 |
| 2018 | Towards Flatter Loss Surface via Nonmonotonic Learning Rate Scheduling
Sihyeon Seong, Yegang Lee, Youngwook Kee, Dongyoon Han, Junmo Kim 0002 |
UAI | 4 |
| 2018 | Unified Simultaneous Clustering and Feature Selection for Unlabeled and Labeled DataabstractThis paper proposes a novel feature selection method, namely, unified simultaneous clustering feature selection (USCFS). A regularized regression with a new type of target matrix is formulated to select the most discriminative features among the original features from labeled or unlabeled data. The regression with -norm regularization allows the projection matrix to represent an effective selection of discriminative features. For unsupervised feature selection, the target matrix discovers label-like information not from the original data points but rather from projected data points, which are of a reduced dimensionality. Without the aid of an affinity graph-based local structure learning method, USCFS allows the target matrix to capture latent cluster centers via orthogonal basis clustering and to simultaneously select discriminative features guided by latent cluster centers. When class labels are available, the target matrix is also able to find latent class labels by regarding the ground-truth class labels as an approximate guide. Hence, supervised feature selection is realized using these latent class labels, which may differ from the ground-truth class labels. Experimental results demonstrate the effectiveness of the proposed method. Specifically, the proposed method outperforms the state-of-the-art methods on diverse real-world data sets for both the supervised and the unsupervised feature selection. Dongyoon Han, Junmo Kim 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Deep Pyramidal Residual NetworksabstractDeep convolutional neural networks (DCNNs) have shown remarkable performance in image classification tasks in recent years. Generally, deep neural network architectures are stacks consisting of a large number of convolutional layers, and they perform downsampling along the spatial dimension via pooling to reduce memory usage. Concurrently, the feature map dimension (i.e., the number of channels) is sharply increased at downsampling locations, which is essential to ensure effective performance because it increases the diversity of high-level attributes. This also applies to residual networks and is very closely related to their performance. In this research, instead of sharply increasing the feature map dimension at units that perform downsampling, we gradually increase the feature map dimension at all units to involve as many locations as possible. This design, which is discussed in depth together with our new insights, has proven to be an effective means of improving generalization ability. Furthermore, we propose a novel residual unit capable of further improving the classification accuracy with our new network architecture. Experiments on benchmark CIFAR-10, CIFAR-100, and ImageNet datasets have shown that our network architecture has superior generalization ability compared to the original residual networks. Dongyoon Han, Jiwhan Kim, Junmo Kim 0002 |
CVPR | 1 |
| 2016 | Salient Region Detection via High-Dimensional Color Transform and Local Spatial SupportabstractIn this paper, we introduce a novel approach to automatically detect salient regions in an image. Our approach consists of global and local features, which complement each other to compute a saliency map. The first key idea of our work is to create a saliency map of an image by using a linear combination of colors in a high-dimensional color space. This is based on an observation that salient regions often have distinctive colors compared with backgrounds in human perception, however, human perception is complicated and highly nonlinear. By mapping the low-dimensional red, green, and blue color to a feature vector in a high-dimensional color space, we show that we can composite an accurate saliency map by finding the optimal linear combination of color coefficients in the high-dimensional color space. To further improve the performance of our saliency estimation, our second key idea is to utilize relative location and color contrast between superpixels as features and to resolve the saliency estimation from a trimap via a learning-based algorithm. The additional local features and learning-based algorithm complement the global estimation from the high-dimensional color transform-based algorithm. The experimental results on three benchmark datasets show that our approach is effective in comparison with the previous state-of-the-art saliency estimation methods. Jiwhan Kim, Dongyoon Han, Yu-Wing Tai, Junmo Kim 0002 |
IEEE Trans. Image Process. | 2 |
| 2015 | Unsupervised Simultaneous Orthogonal basis Clustering Feature SelectionabstractIn this paper, we propose a novel unsupervised feature selection method: Simultaneous Orthogonal basis Clustering Feature Selection (SOCFS). To perform feature selection on unlabeled data effectively, a regularized regression-based formulation with a new type of target matrix is designed. The target matrix captures latent cluster centers of the projected data points by performing orthogonal basis clustering, and then guides the projection matrix to select discriminative features. Unlike the recent unsupervised feature selection methods, SOCFS does not explicitly use the pre-computed local structure information for data points represented as additional terms of their objective functions, but directly computes latent cluster information by the target matrix conducting orthogonal basis clustering in a single unified term of the proposed objective function. It turns out that the proposed objective function can be minimized by a simple optimization algorithm. Experimental results demonstrate the effectiveness of SOCFS achieving the state-of-the-art results with diverse real world datasets. Dongyoon Han, Junmo Kim 0002 |
CVPR | 1 |
| 2015 | Facial age estimation via extended curvature Gabor filterabstractFacial age estimation is a process of identifying the age of a single face in an image or a video. Since age information can be used in many environments such as security, surveillance, and entertainment, age estimation has recently received much attention from researchers. In this paper, we propose an automatic age estimation method via extended curvature Gabor (ECG) features and a learning-based technique. Instead of conventional Gabor Filters, we use ECG filters to extract curvature information from a face image, which is useful for estimating age. We use a feature selection method to reduce the computational complexity and prove the effectiveness of ECG features at the same time. We use a regression algorithm to estimate the age of the test face image. As a result, our work achieves a competitive performance compared with other recent works in terms of age estimation. Jiwhan Kim, Dongyoon Han, Sungryull Sohn, Junmo Kim 0002 |
ICIP | 2 |
| 2014 | Salient Region Detection via High-Dimensional Color TransformabstractIn this paper, we introduce a novel technique to automatically detect salient regions of an image via high-dimensional color transform. Our main idea is to represent a saliency map of an image as a linear combination of high-dimensional color space where salient regions and backgrounds can be distinctively separated. This is based on an observation that salient regions often have distinctive colors compared to the background in human perception, but human perception is often complicated and highly nonlinear. By mapping a low dimensional RGB color to a feature vector in a high-dimensional color space, we show that we can linearly separate the salient regions from the background by finding an optimal linear combination of color coefficients in the high-dimensional color space. Our high dimensional color space incorporates multiple color representations including RGB, CIELab, HSV and with gamma corrections to enrich its representative power. Our experimental results on three benchmark datasets show that our technique is effective, and it is computationally efficient in comparison to previous state-of-the-art techniques. Jiwhan Kim, Dongyoon Han, Yu-Wing Tai, Junmo Kim 0002 |
CVPR | 2 |