VLDB 2026 Research / reviewers in the wild / expert
Byeongho Heo
dblp:142/2705
· DBLP profile ↗
37ranked-venue papers
8as first author
28since 2021 · last 2025
0000-0003-2134-6345ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 7 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Masking meets Supervision: A Strong Learning AllianceabstractPre-training with random masked inputs has emerged as a novel trend in self-supervised training. However, supervised learning still faces a challenge in adopting masking augmentations, primarily due to unstable training. In this paper, we propose a novel way to involve masking augmentations dubbed Masked Sub-branch (MaskSub). MaskSub consists of the main-branch and sub-branch, the latter being a part of the former. The main-branch undergoes conventional training recipes, while the sub-branch merits intensive masking augmentations, during training. MaskSub tackles the challenge by mitigating adverse effects through a relaxed loss function similar to a self-distillation loss. Our analysis shows that MaskSub improves performance, with the training loss converging faster than in standard training, which suggests our method stabilizes the training process. We further validate MaskSub across diverse training scenarios and models, including DeiT-III training, MAE finetuning, CLIP finetuning, BERT training, and hierarchical architectures (ResNet and Swin Transformer). Our results show that MaskSub consistently achieves impressive performance gains across all the cases. MaskSub provides a practical and effective solution for introducing additional regularization under various training recipes. Code available at https://github.com/naver-ai/augsub Byeongho Heo, Taekyung Kim 0002, Sangdoo Yun, Dongyoon Han |
CVPR | 1 |
| 2025 | Morphing Tokens Draw Strong Masked Image ModelsabstractMasked image modeling (MIM) has emerged as a promising approach for pre-training Vision Transformers (ViTs). MIMs predict masked tokens token-wise to recover target signals that are tokenized from images or generated by pre-trained models like vision-language models. While using tokenizers or pre-trained models is viable, they often offer spatially inconsistent supervision even for neighboring tokens, hindering models from learning discriminative representations. Our pilot study identifies spatial inconsistency in supervisory signals and suggests that addressing it can improve representation learning. Building upon this insight, we introduce Dynamic Token Morphing (DTM), a novel method that dynamically aggregates tokens while preserving context to generate contextualized targets, thereby likely reducing spatial inconsistency. DTM is compatible with various SSL frameworks; we showcase significantly improved MIM results, barely introducing extra training costs. Our method facilitates MIM training by using more spatially consistent targets, resulting in improved training trends as evidenced by lower losses. Experiments on ImageNet-1K and ADE20K demonstrate DTM's superiority, which surpasses complex state-of-the-art MIM methods. Furthermore, the evaluation of transfer learning on downstream tasks like iNaturalist, along with extensive empirical studies, supports DTM's effectiveness. Taekyung Kim 0002, Byeongho Heo, Dongyoon Han |
ICLR | 2 |
| 2025 | Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language ModelsabstractWith the rapid advancement of test-time compute search strategies to improve the mathematical problem-solving capabilities of large language models (LLMs), the need for building robust verifiers has become increasingly important. However, all these inference strategies rely on existing verifiers originally designed for Best-of-N search, which makes them sub-optimal for tree search techniques at test time. During tree search, existing verifiers can only offer indirect and implicit assessments of partial solutions or under-value prospective intermediate steps, thus resulting in the premature pruning of promising intermediate steps. To overcome these limitations, we propose token-supervised value models (TVMs) -- a new class of verifiers that assign each token a probability that reflects the likelihood of reaching the correct final answer. This new token-level supervision enables TVMs to directly and explicitly evaluate partial solutions, effectively distinguishing between promising and incorrect intermediate steps during tree search at test time. Experimental results demonstrate that combining tree-search-based inference strategies with TVMs significantly improves the accuracy of LLMs in mathematical problem-solving tasks, surpassing the performance of existing verifiers. Jung Hyun Lee, June Yong Yang, Byeongho Heo, Dongyoon Han, Eunho Yang, Kang Min Yoo |
ICLR | 3 |
| 2025 | Token Bottleneck: One Token to Remember DynamicsabstractDeriving compact and temporally aware visual representations from dynamic scenes is essential for successful execution of sequential scene understanding tasks such as visual tracking and robotic manipulation. In this paper, we introduce Token Bottleneck (ToBo), a simple yet intuitive self-supervised learning pipeline that squeezes a scene into a bottleneck token and predicts the subsequent scene using minimal patches as hints. The ToBo pipeline facilitates the learning of sequential scene representations by conservatively encoding the reference scene into a compact bottleneck token during the squeeze step. In the expansion step, we guide the model to capture temporal dynamics by predicting the target scene using the bottleneck token along with few target patches as hints. This design encourages the vision backbone to embed temporal dependencies, thereby enabling understanding of dynamic transitions across scenes. Extensive experiments in diverse sequential tasks, including video label propagation and robot manipulation in simulated environments demonstrate the superiority of ToBo over baselines. Moreover, deploying our pre-trained model on physical robots confirms its robustness and effectiveness in real-world environments. We further validate the scalability of ToBo across different model scales. Code is available at https://github.com/naver-ai/tobo. Taekyung Kim 0002, Dongyoon Han, Byeongho Heo, Jeongeun Park 0002, Sangdoo Yun |
NeurIPS | 3 |
| 2025 | Improving ViT interpretability with patch-level mask prediction
Junyong Kang, Byeongho Heo, Junsuk Choe |
Pattern Recognit. Lett. | 2 |
| 2024 | Match Me If You Can: Semi-supervised Semantic Correspondence Learning with Unpaired Images
Byeongho Heo, Sangdoo Yun, Seungryong Kim, Dongyoon Han |
ACCV (6) | 2 |
| 2024 | Rotary Position Embedding for Vision Transformer
Byeongho Heo, Song Park, Dongyoon Han, Sangdoo Yun |
ECCV (10) | 1 |
| 2024 | Similarity of Neural Architectures Using Adversarial Attack Transferability
Jaehui Hwang, Dongyoon Han, Byeongho Heo, Song Park, Sanghyuk Chun, Jong-Seok Lee |
ECCV (68) | 3 |
| 2024 | Learning with Unmasked Tokens Drives Stronger Vision Learners
Taekyung Kim 0002, Sanghyuk Chun, Byeongho Heo, Dongyoon Han |
ECCV (34) | 3 |
| 2024 | DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
Byeongho Heo, Dongyoon Han |
ECCV (3) | 2 |
| 2024 | SeiT++: Masked Token Modeling Improves Storage-Efficient Training
Minhyun Lee, Song Park, Byeongho Heo, Dongyoon Han, Hyunjung Shim |
ECCV (28) | 3 |
| 2024 | Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text GuidanceabstractLarge-scale contrastive vision-language pre-trained models provide the zero-shot model achieving competitive performance across a range of image classification tasks without requiring training on downstream data. Recent works have confirmed that while additional fine-tuning of the zero-shot model on the reference data results in enhanced downstream performance, it compromises the model's robustness against distribution shifts. Our investigation begins by examining the conditions required to achieve the goals of robust fine-tuning, employing descriptions based on feature distortion theory and joint energy-based models. Subsequently, we propose a novel robust fine-tuning algorithm, Lipsum-FT, that effectively utilizes the language modeling aspect of the vision-language pre-trained models. Extensive experiments conducted on distribution shift scenarios in DomainNet and ImageNet confirm the superiority of our proposed Lipsum-FT approach over existing robust fine-tuning methods. Giung Nam, Byeongho Heo, Juho Lee 0001 |
ICLR | 2 |
| 2023 | Scratching Visual Transformer's Back with Uniform AttentionabstractThe favorable performance of Vision Transformers (ViTs) is often attributed to the multi-head self-attention (MSA), which enables global interactions at each layer of a ViT model. Previous works acknowledge the property of long-range dependency for the effectiveness in MSA. In this work, we study the role of MSA in terms of the different axis, density. Our preliminary analyses suggest that the spatial interactions of learned attention maps are close to dense interactions rather than sparse ones. This is a curious phenomenon because dense attention maps are harder for the model to learn due to softmax. We interpret this opposite behavior against softmax as a strong preference for the ViT models to include dense interaction. We thus manually insert the dense uniform attention to each layer of the ViT models to supply the much-needed dense interactions. We call this method Context Broadcasting, CB. Our study demonstrates the inclusion of CB takes the role of dense attention and thereby reduces the degree of density in the original attention maps by complying softmax in MSA. We also show that, with negligible costs of CB (1 line in your model code and no additional parameters), both the capacity and generalizability of the ViT models are increased. Nam Hyeon-Woo, Kim Yu-Ji, Byeongho Heo, Dongyoon Han, Seong Joon Oh, Tae-Hyun Oh |
ICCV | 3 |
| 2023 | SeiT: Storage-Efficient Vision Training with Tokens Using 1% of Pixel StorageabstractWe need billion-scale images to achieve more generalizable and ground-breaking vision models, as well as massive dataset storage to ship the images (e.g., the LAION-5B dataset needs 240TB storage space). However, it has become challenging to deal with unlimited dataset storage with limited storage infrastructure. A number of storage-efficient training methods have been proposed to tackle the problem, but they are rarely scalable or suffer from severe damage to performance. In this paper, we propose a storage-efficient training strategy for vision classifiers for large-scale datasets (e.g., ImageNet) that only uses 1024 tokens per instance without using the raw level pixels; our token storage only needs <1% of the original JPEG-compressed raw pixels. We also propose token augmentations and a Stem-adaptor module to make our approach able to use the same architecture as pixel-based approaches with only minimal modifications on the stem layer and the carefully tuned optimization settings. Our experimental results on ImageNet-1k show that our method significantly outperforms other storage-efficient training methods with a large gap. We further show the effectiveness of our method in other practical scenarios, storage-efficient pre-training, and continual learning. Code is available at https://github.com/naver-ai/seit. Song Park, Sanghyuk Chun, Byeongho Heo, Wonjae Kim, Sangdoo Yun |
ICCV | 3 |
| 2023 | What Do Self-Supervised Vision Transformers Learn?
Namuk Park, Wonjae Kim, Byeongho Heo, Taekyung Kim 0002, Sangdoo Yun |
ICLR | 3 |
| 2022 | Joint Global and Local Hierarchical Priors for Learned Image CompressionabstractRecently, learned image compression methods have out-performed traditional hand-crafted ones including BPG. One of the keys to this success is learned entropy models that estimate the probability distribution of the quantized latent representation. Like other vision tasks, most recent learned entropy models are based on convolutional neural networks (CNNs). However, CNNs have a limitation in modeling long-range dependencies due to their nature of local connectivity, which can be a significant bottleneck in image compression where reducing spatial redundancy is a key point. To overcome this issue, we propose a novel entropy model called Information Transformer (Informer) that exploits both global and local information in a content-dependent manner using an attention mechanism. Our experiments show that Informer improves rate-distortion performance over the state-of-the-art methods on the Kodak and Tecnick datasets without the quadratic computational complexity problem. Our source code is available at https://github.com/naver-ai/informer. Jun-Hyuk Kim, Byeongho Heo, Jong-Seok Lee |
CVPR | 2 |
| 2022 | The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationabstractThe problem of class imbalanced data is that the gener-alization performance of the classifier deteriorates due to the lack of data from minority classes. In this paper, we pro-pose a novel minority over-sampling method to augment di-versified minority samples by leveraging the rich context of the majority classes as background images. To diversify the minority samples, our key idea is to paste an image from a minority class onto rich-context images from a majority class, using them as background images. Our method is simple and can be easily combined with the existing long-tailed recognition methods. We empirically prove the effectiveness of the proposed oversampling method through extensive experiments and ablation studies. Without any architectural changes or complex algorithms, our method achieves state-of-the-art performance on various long-tailed classification benchmarks. Our code is made available at https://github.com/naver-ai/cmo. Seulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun, Jin Young Choi 0002 |
CVPR | 3 |
| 2022 | K-centered Patch Sampling for Efficient Video Recognition
Seong Hyeon Park, Jihoon Tack, Byeongho Heo, Jung-Woo Ha 0001, Jinwoo Shin |
ECCV (35) | 3 |
| 2022 | Learning Features with Parameter-Free Layers
Dongyoon Han, Young Joon Yoo, Beomyoung Kim, Byeongho Heo |
ICLR | 4 |
| 2022 | ViDT: An Efficient and Effective Fully Transformer-based Object Detector
Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, Ming-Hsuan Yang 0001 |
ICLR | 6 |
| 2022 | Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationabstractEnsembles of deep neural networks have demonstrated superior performance, but their heavy computational cost hinders applying them for resource-limited environments. It motivates distilling knowledge from the ensemble teacher into a smaller student network, and there are two important design choices for this ensemble distillation: 1) how to construct the student network, and 2) what data should be shown during training. In this paper, we propose a weight averaging technique where a student with multiple subnetworks is trained to absorb the functional diversity of ensemble teachers, but then those subnetworks are properly averaged for inference, giving a single student network with no additional inference cost. We also propose a perturbation strategy that seeks inputs from which the diversities of teachers can be better transferred to the student. Combining these two, our method significantly improves upon previous methods on various image classification tasks. Giung Nam, Hyungi Lee, Byeongho Heo, Juho Lee 0001 |
ICML | 3 |
| 2022 | Rollback Ensemble With Multiple Local Minima in Fine-Tuning Deep Learning NetworksabstractImage retrieval is a challenging problem that requires learning generalized features enough to identify untrained classes, even with very few classwise training samples. In this article, to obtain generalized features further in learning retrieval data sets, we propose a novel fine-tuning method of pretrained deep networks. In the retrieval task, we discovered a phenomenon in which the loss reduction in fine-tuning deep networks is stagnated, even while weights are largely updated. To escape from the stagnated state, we propose a new fine-tuning strategy to roll back some of the weights to the pretrained values. The rollback scheme is observed to drive the learning path to a gentle basin that provides more generalized features than a sharp basin. In addition, we propose a multihead ensemble structure to create synergy among multiple local minima obtained by our rollback scheme. Experimental results show that the proposed learning method significantly improves generalization performance, achieving state-of-the-art performance on the Inshop and SOP data sets. Youngmin Ro, Jongwon Choi 0002, Byeongho Heo, Jin Young Choi 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Show, Attend and Distill: Knowledge Distillation via Attention-based Feature MatchingabstractKnowledge distillation extracts general knowledge from a pretrained teacher network and provides guidance to a target student network. Most studies manually tie intermediate features of the teacher and student, and transfer knowledge through predefined links. However, manual selection often constructs ineffective links that limit the improvement from the distillation. There has been an attempt to address the problem, but it is still challenging to identify effective links under practical scenarios. In this paper, we introduce an effective and efficient feature distillation method utilizing all the feature levels of the teacher without manually selecting the links. Specifically, our method utilizes an attention based meta network that learns relative similarities between features, and applies identified similarities to control distillation intensities of all possible pairs. As a result, our method determines competent links more efficiently than the previous approach and provides better performance on model compression and transfer learning tasks. Further qualitative analyses and ablative studies describe how our method contributes to better distillation. Mingi Ji, Byeongho Heo, Sungrae Park |
AAAI | 2 |
| 2021 | Rethinking Channel Dimensions for Efficient Model DesignabstractDesigning an efficient model within the limited computational cost is challenging. We argue the accuracy of a lightweight model has been further limited by the design convention: a stage-wise configuration of the channel dimensions, which looks like a piecewise linear function of the network stage. In this paper, we study an effective channel dimension configuration towards better performance than the convention. To this end, we empirically study how to design a single layer properly by analyzing the rank of the output feature. We then investigate the channel configuration of a model by searching network architectures concerning the channel configuration under the computational cost restriction. Based on the investigation, we propose a simple yet effective channel configuration that can be parameterized by the layer index. As a result, our proposed model following the channel parameterization achieves remarkable performance on ImageNet classification and transfer learning tasks including COCO object detection, COCO instance segmentation, and fine-grained classifications. Code and ImageNet pretrained models are available at https: //github.com/clovaai/rexnet. Dongyoon Han, Sangdoo Yun, Byeongho Heo, Young Joon Yoo |
CVPR | 3 |
| 2021 | Re-Labeling ImageNet: From Single to Multi-Labels, From Global to Localized LabelsabstractImageNet has been the most popular image classification benchmark, but it is also the one with a significant level of label noise. Recent studies have shown that many samples contain multiple classes, despite being assumed to be a single-label benchmark. They have thus proposed to turn ImageNet evaluation into a multi-label task, with exhaustive multi-label annotations per image. However, they have not fixed the training set, presumably because of a formidable annotation cost. We argue that the mismatch between single-label annotations and effectively multi-label images is equally, if not more, problematic in the training setup, where random crops are applied. With the single-label annotations, a random crop of an image may contain an entirely different object from the ground truth, introducing noisy or even incorrect supervision during training. We thus re-label the ImageNet training set with multi-labels. We address the annotation cost barrier by letting a strong image classifier, trained on an extra source of data, generate the multi-labels. We utilize the pixel-wise multi-label predictions before the final pooling layer, in order to exploit the additional location-specific supervision signals. Training on the re-labeled samples results in improved model performances across the board. ResNet-50 attains the top-1 accuracy of 78.9% on ImageNet with our localized multi-labels, which can be further boosted to 80.2% with the CutMix regularization. We show that the models trained with localized multi-labels also outperforms the baselines on transfer learning to object detection and instance segmentation tasks, and various robustness benchmarks. The re-labeled ImageNet training set, pre-trained weights, and the source code are available at https://github.com/naverai/relabel_imagenet. Sangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han, Junsuk Choe, Sanghyuk Chun |
CVPR | 3 |
| 2021 | Rethinking Spatial Dimensions of Vision TransformersabstractVision Transformer (ViT) extends the application range of transformers from language processing to computer vision tasks as being an alternative architecture against the existing convolutional neural networks (CNN). Since the transformer-based architecture has been innovative for computer vision modeling, the design convention towards an effective architecture has been less studied yet. From the successful design principles of CNN, we investigate the role of spatial dimension conversion and its effectiveness on transformer-based architecture. We particularly attend to the dimension reduction principle of CNNs; as the depth increases, a conventional CNN increases channel dimension and decreases spatial dimensions. We empirically show that such a spatial dimension reduction is beneficial to a transformer architecture as well, and propose a novel Pooling-based Vision Transformer (PiT) upon the original ViT model. We show that PiT achieves the improved model capability and generalization performance against ViT. Throughout the extensive experiments, we further show PiT outperforms the baseline on several tasks such as image classification, object detection, and robustness evaluation. Source codes and ImageNet models are available at https://github.com/naver-ai/pit. Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, Seong Joon Oh |
ICCV | 1 |
| 2021 | AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, Jung-Woo Ha 0001 |
ICLR | 1 |
| 2021 | Motion-aware ensemble of three-mode trackers for unmanned aerial vehicles
Kyuewang Lee, Hyung Jin Chang, Jongwon Choi 0002, Byeongho Heo, Ales Leonardis, Jin Young Choi 0002 |
Mach. Vis. Appl. | 4 |
| 2019 | Knowledge Distillation with Adversarial Samples Supporting Decision BoundaryabstractMany recent works on knowledge distillation have provided ways to transfer the knowledge of a trained network for improving the learning process of a new one, but finding a good technique for knowledge distillation is still an open problem. In this paper, we provide a new perspective based on a decision boundary, which is one of the most important component of a classifier. The generalization performance of a classifier is closely related to the adequacy of its decision boundary, so a good classifier bears a good decision boundary. Therefore, transferring information closely related to the decision boundary can be a good attempt for knowledge distillation. To realize this goal, we utilize an adversarial attack to discover samples supporting a decision boundary. Based on this idea, to transfer more accurate information about the decision boundary, the proposed algorithm trains a student classifier based on the adversarial samples supporting the decision boundary. Experiments show that the proposed method indeed improves knowledge distillation and achieves the state-of-the-arts performance. Byeongho Heo, Minsik Lee 0001, Sangdoo Yun, Jin Young Choi 0002 |
AAAI | 1 |
| 2019 | Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden NeuronsabstractAn activation boundary for a neuron refers to a separating hyperplane that determines whether the neuron is activated or deactivated. It has been long considered in neural networks that the activations of neurons, rather than their exact output values, play the most important role in forming classificationfriendly partitions of the hidden feature space. However, as far as we know, this aspect of neural networks has not been considered in the literature of knowledge transfer. In this paper, we propose a knowledge transfer method via distillation of activation boundaries formed by hidden neurons. For the distillation, we propose an activation transfer loss that has the minimum value when the boundaries generated by the student coincide with those by the teacher. Since the activation transfer loss is not differentiable, we design a piecewise differentiable loss approximating the activation transfer loss. By the proposed method, the student learns a separating boundary between activation region and deactivation region formed by each neuron in the teacher. Through the experiments in various aspects of knowledge transfer, it is verified that the proposed method outperforms the current state-of-the-art. Byeongho Heo, Minsik Lee 0001, Sangdoo Yun, Jin Young Choi 0002 |
AAAI | 1 |
| 2019 | Backbone Cannot Be Trained at Once: Rolling Back to Pre-Trained Network for Person Re-IdentificationabstractIn person re-identification (ReID) task, because of its shortage of trainable dataset, it is common to utilize fine-tuning method using a classification network pre-trained on a large dataset. However, it is relatively difficult to sufficiently finetune the low-level layers of the network due to the gradient vanishing problem. In this work, we propose a novel fine-tuning strategy that allows low-level layers to be sufficiently trained by rolling back the weights of high-level layers to their initial pre-trained weights. Our strategy alleviates the problem of gradient vanishing in low-level layers and robustly trains the low-level layers to fit the ReID dataset, thereby increasing the performance of ReID tasks. The improved performance of the proposed strategy is validated via several experiments. Furthermore, without any addons such as pose estimation or segmentation, our strategy exhibits state-of-the-art performance using only vanilla deep convolutional neural network architecture. Youngmin Ro, Jongwon Choi 0002, Byeongho Heo, Jongin Lim 0002, Jin Young Choi 0002 |
AAAI | 4 |
| 2019 | A Comprehensive Overhaul of Feature DistillationabstractWe investigate the design aspects of feature distillation methods achieving network compression and propose a novel feature distillation method in which the distillation loss is designed to make a synergy among various aspects: teacher transform, student transform, distillation feature position and distance function. Our proposed distillation loss includes a feature transform with a newly designed margin ReLU, a new distillation feature position, and a partial L2distance function to skip redundant information giving adverse effects to the compression of student. In ImageNet, our proposed method achieves 21.65% of top-1 error with ResNet50, which outperforms the performance of the teacher network, ResNet152. Our proposed method is evaluated on various tasks such as image classification, object detection and semantic segmentation and achieves a significant performance improvement in all tasks. The code is available at bhheo.github.io/overhaul. Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 0001, Nojun Kwak, Jin Young Choi 0002 |
ICCV | 1 |
| 2018 | Pose transforming network: Learning to disentangle human posture in variational auto-encoded latent space
Jongin Lim 0002, Young Joon Yoo, Byeongho Heo, Jin Young Choi 0002 |
Pattern Recognit. Lett. | 3 |
| 2017 | Appearance and motion based deep learning architecture for moving object detection in moving cameraabstractBackground subtraction from the given image is a widely used method for moving object detection. However, this method is vulnerable to dynamic background in a moving camera video. In this paper, we propose a novel moving object detection approach using deep learning to achieve a robust performance even in a dynamic background. The proposed approach considers appearance features as well as motion features. To this end, we design a deep learning architecture composed of two networks: an appearance network and a motion network. The two networks are combined to detect moving object robustly to the background motion by utilizing the appearance of the target object in addition to the motion difference. In the experiment, it is shown that the proposed method achieves 50 fps speed in GPU and outperforms state-of-the-art methods for various moving camera videos. Byeongho Heo, Kimin Yun, Jin Young Choi 0002 |
ICIP | 1 |
| 2017 | Deep learning architecture for pedestrian 3-D localization and tracking using multiple camerasabstractIn this paper, we propose a novel deep-learning architecture for accurate 3-D localization and tracking of a pedestrian using multiple cameras. The deep-learning network is composed of two networks: detection network and localization network. The detection network yields the pedestrian detections and the localization network estimates the ground position of a pedestrian within its detection box. In addition, an attentional pass filter is introduced to effectively connect the two networks. Using the detection proposals and their 2-D grounding positions obtained from the two networks, multi-camera multi-target 3-D localization and tracking algorithm is developed through min-cost network flow approach. In the experiments, it is shown that the proposed method improves the performance of 3-D localization and tracking. Kikyung Kim, Byeongho Heo, Moonsub Byeon, Jin Young Choi 0002 |
ICIP | 2 |
| 2014 | Self-Organizing Cascaded Structure of Deformable Part Models for Fast Object DetectionabstractIn this paper, we propose a framework which self-organizes the cascaded object detection filters for fast object detection with maintaining high accuracy. The proposed scheme consists of root and part filter modules, which are cascaded in a self-organizing structure. The pruning of non-object regions in low resolution at the root cascade stage is critical for the object detection speed. At root stage, to prune as many non-object regions as possible, we build a root cascade structure using multiple root models. These models are obtained via bagging procedure for non-linear classification of object and non-object parts in an image. Additional speed-up is achieved by determining proper deployment order of part models. We define a discriminability measure for the part models and suggest a self-organizing scheme to generate an efficient order of part models. The proposed method is evaluated through computational experiments with the PASCAL VOC and INRIA datasets, as a result, our method achieves on average more than 2 times faster performance than the original cascade-DPM, with comparable precision scores. Sangdoo Yun, Hawook Jeong, Woo-Sung Kang, Byeongho Heo, Jin Young Choi 0002 |
ICPR | 4 |
| 2013 | Initialization-Insensitive Visual Tracking through Voting with Salient Local FeaturesabstractIn this paper we propose an object tracking method in case of inaccurate initializations. To track objects accurately in such situation, the proposed method uses "motion saliency" and "descriptor saliency" of local features and performs tracking based on generalized Hough transform (GHT). The proposed motion saliency of a local feature emphasizes features having distinctive motions, compared to the motions which are not from the target object. The descriptor saliency emphasizes features which are likely to be of the object in terms of its feature descriptors. Through these saliencies, the proposed method tries to "learn and find" the target object rather than looking for what was given at initialization, giving robust results even with inaccurate initializations. Also, our tracking result is obtained by combining the results of each local feature of the target and the surroundings with GHT voting, thus is robust against severe occlusions as well. The proposed method is compared against nine other methods, with nine image sequences, and hundred random initializations. The experimental results show that our method outperforms all other compared methods. Kwang Moo Yi, Hawook Jeong, Byeongho Heo, Hyung Jin Chang, Jin Young Choi 0002 |
ICCV | 3 |