EDBT 2026 Demo / reviewers in the wild / expert
Bohyung Han
dblp:73/4880
· DBLP profile ↗
141ranked-venue papers
14as first author
52since 2021 · last 2026
0000-0003-3099-3616ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 120 · 10 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 103 · 9 first-author · 34 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyperPose: Hyper-pose Embeddings for 3D-Aware Generative Models with Self-Supervised Disentangling of Pose and SceneabstractWe propose a novel framework for training 3D-aware Generative Adversarial Networks (GANs) from a collection of 2D images, effectively learning both image distribution and 3D geometric configurations without relying on strong 3D priors such as camera poses, depth information, or target-specific 3D models. To achieve these objectives, we introduce hyper-pose embeddings alongside a novel pose disentanglement technique that effectively separates pose and scene information. This crucial disentanglement helps the generative model overcome the inherent conflict between learning photo-realism and accurate 3D geometry. Furthermore, we propose soft contrastive learning to robustly handle the continuous nature of camera poses, and a non-match loss to further enhance disentanglement and refine embedding training. Experiments on challenging datasets demonstrate the effectiveness of our method in 3D-aware image synthesis, particularly for scenes with complex or diverse objects. Namgi Kim, Bohyung Han |
WACV | 3 |
| 2025 | Human-Centric AI: From Explainability and Trustworthiness to Actionable EthicsabstractTo address the potential risks of AI while supporting innovation and ensuring responsible adoption, there is an urgent need for clear governance frameworks grounded in human-centric values. It is imperative that AI systems operate in ways that are transparent, trustworthy, and ethically sound. Developing truly human-centric AI goes beyond technical innovation. It requires interdisciplinary collaboration and diverse perspectives. This workshop will explore key challenges and emerging solutions in the development of human-centric AI, with a focus on explainability, trustworthiness, fairness, and privacy. We welcome both theoretical contributions and practical case studies that demonstrate how human-centered principles are realized in real-world AI systems. The official workshop webpage is available at https://xai.kaist.ac.kr/Workshop/hcai2025/, which provides comprehensive information about the program. Jaesik Choi, Bohyung Han, Myoung-Wan Koo, Kyungman Bae, Chang Dong Yoo, Simon S. Woo, Wojciech Samek |
CIKM | 2 |
| 2025 | Enhanced Diffusion Sampling via Extrapolation with Multiple ODE SolutionsabstractDiffusion probabilistic models (DPMs), while effective in generating high-quality samples, often suffer from high computational costs due to their iterative sampling process. To address this, we propose an enhanced ODE-based sampling method for DPMs inspired by Richardson extrapolation, which reduces numerical error and improves convergence rates. Our method, RX-DPM, leverages multiple ODE solutions at intermediate time steps to extrapolate the denoised prediction in DPMs. This significantly enhances the accuracy of estimations for the final sample while maintaining the number of function evaluations (NFEs). Unlike standard Richardson extrapolation, which assumes uniform discretization of the time grid, we develop a more general formulation tailored to arbitrary time step scheduling, guided by local truncation error derived from a baseline sampling method. The simplicity of our approach facilitates accurate estimation of numerical solutions without significant computational overhead, and allows for seamless and convenient integration into various DPMs and solvers.
Additionally, RX-DPM provides explicit error estimates, effectively demonstrating the faster convergence as the leading error term's order increases. Through a series of experiments, we show that the proposed method improves the quality of generated samples without requiring additional sampling iterations. Junoh Kang, Bohyung Han |
ICLR | 3 |
| 2025 | Fine-Grained Captioning of Long Videos through Scene Graph ConsolidationabstractRecent advances in vision-language models have led to impressive progress in caption generation for images and short video clips. However, these models remain constrained by their limited temporal receptive fields, making it difficult to produce coherent and comprehensive captions for long videos. While several methods have been proposed to aggregate information across video segments, they often rely on supervised fine-tuning or incur significant computational overhead. To address these challenges, we introduce a novel framework for long video captioning based on graph consolidation. Our approach first generates segment-level captions, corresponding to individual frames or short video intervals, using off-the-shelf visual captioning models. These captions are then parsed into individual scene graphs, which are subsequently consolidated into a unified graph representation that preserves both holistic context and fine-grained details throughout the video. A lightweight graph-to-text decoder then produces the final video-level caption. This framework effectively extends the temporal understanding capabilities of existing models without requiring any additional fine-tuning on long video datasets. Experimental results show that our method significantly outperforms existing LLM-based consolidation approaches, achieving strong zero-shot performance while substantially reducing computational costs. Sanghyeok Chu, Seonguk Seo, Bohyung Han |
ICML | 3 |
| 2025 | FedLPA: Local Prior Alignment for Heterogeneous Federated Generalized Category DiscoveryabstractFederated Generalized Category Discovery (Fed-GCD) aims to train a global model that classifies seen classes while discovering novel ones from data distributed across heterogeneous clients.
Existing GCD methods often rely on unrealistic assumptions, such as prior knowledge of the number of novel classes or balanced class distributions across clients.
We propose Federated Local Prior Alignment (FedLPA), which eliminates these assumptions by grounding learning in client-specific structures and aligning predictions with locally derived priors.
Specifically, each client constructs a similarity graph refined with high-confidence signals from seen classes, and then identifies local concepts and prototypes via Infomap clustering.
Building on these discovered structures, we introduce Local Prior Alignment (LPA), a self-distillation mechanism that aligns batch-level predictions with empirical class prior derived from concept assignments.
Through iterative local structure discovery and adaptive prior refinement, FedLPA achieves robust generalized category discovery under severe data heterogeneity.
Extensive experiments demonstrate that FedLPA significantly outperforms existing federated GCD methods across both fine-grained and standard benchmarks. Geeho Kim, Bohyung Han |
NeurIPS | 3 |
| 2025 | Diffusion-Based Conditional Image Editing Through Optimized Inference with GuidanceabstractWe present a simple but effective training-free approach for text-driven image-to-image translation based on a pre-trained text-to-image diffusion model. Our goal is to gener-ate an image that aligns with the target task while preserving the structure and background of a source image. To this end, we derive the representation guidance with a combination of two objectives: maximizing the similarity to the target prompt based on the CLIP score and minimizing the structural distance to the source latent variable. This guidance improves the fidelity of the generated target image to the given target prompt while maintaining the structure integrity of the source image. To incorporate the representation guidance component, we optimize the target latent variable of diffusion model's reverse process with the guidance. Experimental results demonstrate that our method achieves outstanding image-to-image translation performance on various tasks when combined with the pretrained Stable Diffusion model. Hyunsoo Lee 0005, Minsoo Kang, Bohyung Han |
WACV | 3 |
| 2025 | Re-Evaluating Group Robustness via Adaptive Class-Specific ScalingabstractGroup distributionally robust optimization, which aims to improve robust accuracies-worst-group and unbiased accuracies-is a prominent algorithm used to mitigate spu-rious correlations and address dataset bias. Although ex-isting approaches have reported improvements in robust accuracies, these gains often come at the cost of average accuracy due to inherent trade-offs. To control this trade-off flexibly and efficiently, we propose a simple class-specific scaling strategy, directly applicable to existing debiasing algorithms with no additional training. We further develop an instance-wise adaptive scaling technique to alleviate this trade-off, even leading to improvements in both robust and average accuracies. Our approach reveals that a naive ERM baseline matches or even outperforms the recent debi-asing methods by simply adopting the class-specific scaling technique. Additionally, we introduce a novel unified metric that quantifies the trade-off between the two accuracies as a scalar value, allowing for a comprehensive evaluation of existing algorithms. By tackling the inherent trade-off and offering a performance landscape, our approach provides valuable insights into robust techniques beyond just robust accuracy. We validate the effectiveness of our framework through experiments across datasets in computer vision and natural language processing domains. Seonguk Seo, Bohyung Han |
WACV | 2 |
| 2025 | Revisiting Machine Unlearning with Dimensional AlignmentabstractMachine unlearning, an emerging research topic focusing on data privacy compliance, enables trained models to erase information learned from specific data. While many existing methods indirectly address this issue by intentionally injecting incorrect supervision, they often result in drastic and unpredictable changes to decision boundaries and feature spaces, leading to training instability and undesired side effects. To address this challenge more fundamentally, we first analyze the changes in latent feature spaces between the original and retrained models, and observe that the feature representations of samples not included in training are closely aligned with the feature manifolds of previously seen samples. Building on this insight, we introduce a novel evaluation metric for machine unlearning, coined dimensional alignment, which measures the alignment between the eigenspaces of the forget and retain sets. We incorporate this metric as a regularizer loss to develop a robust and stable unlearning framework, which is further enhanced by a self-distillation loss and an alternating training scheme. Our framework effectively eliminates information from the forget set while preserving knowledge from the retain set. Finally, we identify critical flaws in existing evaluation metrics for machine unlearning and propose new tools that more accurately capture its fundamental objectives. Seonguk Seo, Dongwan Kim, Bohyung Han |
WACV | 3 |
| 2025 | Metric Compatible Training for Online Backfilling in Large-Scale RetrievalabstractBackfilling is the process of re-extracting all gallery embeddings from upgraded models in image retrieval systems. It inevitably spends a prohibitively large amount of computational cost and even entails the downtime of the service. Although backward-compatible learning sidesteps this challenge by tackling query-side representations, this leads to suboptimal solutions in principle because gallery embeddings cannot benefit from model upgrades. We address this dilemma by introducing an online backfilling algorithm, which enables us to achieve a progressive performance improvement during the backfilling process without sacrificing the full performance of the new model after the completion of backfilling. To this end, we first show that a simple distance rank merge is a reasonable option for online backfilling. Then, we incorporate a reverse transformation module for more effective and efficient merging, which is further enhanced by adopting metric-compatible contrastive learning. These two components help to make the distances of old and new models compatible, resulting in desirable merge results during backfilling with no extra computational over-head. Extensive experiments show the benefit of our frame-work on four standard benchmarks in various settings. Seonguk Seo, Mustafa Gökhan Uzunbas, Bohyung Han, Sara Cao, Ser-Nam Lim |
WACV | 3 |
| 2024 | Cross-Class Feature Augmentation for Class Incremental LearningabstractWe propose a novel class incremental learning approach, which incorporates a feature augmentation technique motivated by adversarial attacks. We employ a classifier learned in the past to complement training examples of previous tasks. The proposed approach has an unique perspective to utilize the previous knowledge in class incremental learning since it augments features of arbitrary target classes using examples in other classes via adversarial attacks on a previously learned classifier. By allowing the Cross-Class Feature Augmentations (CCFA), each class in the old tasks conveniently populates samples in the feature space, which alleviates the collapse of the decision boundaries caused by sample deficiency for the previous tasks, especially when the number of stored exemplars is small. This idea can be easily incorporated into existing class incremental learning algorithms without any architecture modification. Extensive experiments on the standard benchmarks show that our method consistently outperforms existing class incremental learning methods by significant margins in various scenarios, especially under an environment with an extremely limited memory budget. Jaeyoo Park, Bohyung Han |
AAAI | 3 |
| 2024 | Observation-Guided Diffusion Probabilistic ModelsabstractWe propose a novel diffusion-based image generation method called the observation-guided diffusion probabilis-tic model (OGDM), which effectively addresses the trade-off between quality control and fast sampling. Our approach reestablishes the training objective by integrating the guidance of the observation process with the Markov chain in a principled way. This is achieved by introducing an additional loss term derived from the observation based on a conditional discriminator on noise level, which employs a Bernoulli distribution indicating whether its in-put lies on the (noisy) real manifold or not. This strat-egy allows us to optimize the more accurate negative log-likelihood induced in the inference stage especially when the number offunction evaluations is limited. The proposed training scheme is also advantageous even when incorpo-rated only into the fine-tuning process, and it is compati-ble with various fast inference strategies since our method yields better denoising networks using the exactly the same inference procedure without incurring extra computational cost. We demonstrate the effectiveness of our training al-gorithm using diverse inference techniques on strong dif-fusion model baselines. Our implementation is available at https://github.com/Junoh-Kang/OGDM_edm. Junoh Kang, Sungik Choi, Bohyung Han |
CVPR | 4 |
| 2024 | Communication-Efficient Federated Learning with Accelerated Client GradientabstractFederated learning often suffers from slow and unstable convergence due to the heterogeneous characteristics of participating client datasets. Such a tendency is aggravated when the client participation ratio is low since the information collected from the clients has large variations. To address this challenge, we propose a simple but effective federated learning framework, which improves the consistency across clients and facilitates the convergence of the server model. This is achieved by making the server broadcast a global model with a lookahead gradient. This strategy enables the proposed approach to convey the projected global update information to participants effectively without additional client memory and extra communication costs. We also regularize local updates by aligning each client with the overshot global model to reduce bias and improve the stability of our algorithm. We provide the theoretical convergence rate of our algorithm and demonstrate remarkable performance gains in terms of accuracy and communication efficiency compared to the state-of-the-art methods, especially with low client participation rates. The source code is available at our project page1.1.https://github.com/geehokim/FedACG: Geeho Kim, Jinkyu Kim 0005, Bohyung Han |
CVPR | 3 |
| 2024 | Robust Image Denoising Through Adversarial Frequency MixupabstractImage denoising approaches based on deep neural net-works often struggle with overfitting to specific noise distributions present in training data. This challenge per-sists in existing real-world denoising networks, which are trained using a limited spectrum of real noise distributions, and thus, show poor robustness to out-of-distribution real noise types. To alleviate this issue, we develop a novel training framework called Adversarial Frequency Mixup (AFM). AFM leverages mixup in the frequency domain to generate noisy images with distinctive and challenging noise characteristics, all the while preserving the properties of authentic real-world noise. Subsequently, incorporating these noisy images into the training pipeline enhances the denoising network's robustness to variations in noise distributions. Extensive experiments and analyses, con-ducted on a wide range of real noise benchmarks demon-strate that denoising networks trained with our proposed framework exhibit significant improvements in robustness to unseen noise distributions. The code is available at https://github.com/dhryougit/AFM. Donghun Ryou, Inju Ha, Hyewon Yoo, Dongwan Kim, Bohyung Han |
CVPR | 5 |
| 2024 | Relaxed Contrastive Learning for Federated LearningabstractWe propose a novel contrastive learning framework to effectively address the challenges of data heterogeneity infederated learning. We first analyze the inconsistency of gradient updates across clients during local training and establish its dependence on the distribution of feature representations, leading to the derivation of the supervised contrastive learning (SCL) objective to mitigate local deviations. In addition, we show that a naïve integration of SCL into federated learning incurs representation collapse, resulting in slow convergence and limited performance gains. To address this issue, we introduce a relaxed contrastive learning loss that imposes a divergence penalty on excessively similar sample pairs within each class. This strategy prevents collapsed representations and enhances feature transferability, facilitating collaborative training and leading to significant performance improvements. Our framework out-performs all existing federated learning approaches by significant margins on the standard benchmarks, as demonstrated by extensive experimental results. The source code is available at our project page11https://github.com/skynbe/FedRCL: Seonguk Seo, Jinkyu Kim 0005, Geeho Kim, Bohyung Han |
CVPR | 4 |
| 2024 | Leveraging Temporal Contextualization for Video Action Recognition
Minji Kim 0002, Dongyoon Han, Taekyung Kim 0002, Bohyung Han |
ECCV (21) | 4 |
| 2024 | Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation
Junsung Lee, Minsoo Kang, Bohyung Han |
ECCV (12) | 3 |
| 2024 | 4D Gaussian Splatting in the Wild with Uncertainty-Aware RegularizationabstractNovel view synthesis of dynamic scenes is becoming important in various applications, including augmented and virtual reality.
We propose a novel 4D Gaussian Splatting (4DGS) algorithm for dynamic scenes from casually recorded monocular videos.
To overcome the overfitting problem of existing work for these real-world videos, we introduce an uncertainty-aware regularization that identifies uncertain regions with few observations and selectively imposes additional priors based on diffusion models and depth smoothness on such regions.
This approach improves both the performance of novel view synthesis and the quality of training image reconstruction.
We also identify the initialization problem of 4DGS in fast-moving dynamic regions, where the Structure from Motion (SfM) algorithm fails to provide reliable 3D landmarks.
To initialize Gaussian primitives in such regions, we present a dynamic region densification method using the estimated depth maps and scene flow.
Our experiments show that the proposed method improves the performance of 4DGS reconstruction from a video captured by a handheld monocular camera and also exhibits promising results in few-shot static scene reconstruction. Mijeong Kim 0002, Jongwoo Lim, Bohyung Han |
NeurIPS | 3 |
| 2024 | FIFO-Diffusion: Generating Infinite Videos from Text without TrainingabstractWe propose a novel inference technique based on a pretrained diffusion model for text-conditional video generation. Our approach, called FIFO-Diffusion, is conceptually capable of generating infinitely long videos without additional training. This is achieved by iteratively performing diagonal denoising, which simultaneously processes a series of consecutive frames with increasing noise levels in a queue; our method dequeues a fully denoised frame at the head while enqueuing a new random noise frame at the tail. However, diagonal denoising is a double-edged sword as the frames near the tail can take advantage of cleaner frames by forward reference but such a strategy induces the discrepancy between training and inference. Hence, we introduce latent partitioning to reduce the training-inference gap and lookahead denoising to leverage the benefit of forward referencing. Practically, FIFO-Diffusion consumes a constant amount of memory regardless of the target video length given a baseline model, while well-suited for parallel inference on multiple GPUs. We have demonstrated the promising results and effectiveness of the proposed methods on existing text-to-video generation baselines. Generated video examples and source codes are available at our project page. Junoh Kang, Bohyung Han |
NeurIPS | 4 |
| 2024 | Hierarchical Visual Feature Aggregation for OCR-Free Document UnderstandingabstractWe present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs).
Our approach employs multi-scale visual features to effectively handle various font sizes within document images.
To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs.
Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes.
Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs.
Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks. Jaeyoo Park, Jeonghyung Park, Bohyung Han |
NeurIPS | 4 |
| 2024 | Randomized Adversarial Style Perturbations for Domain GeneralizationabstractWe propose a novel domain generalization technique, referred to as Randomized Adversarial Style Perturbation (RASP), which is motivated by the observation that the characteristics of each domain are captured by the feature statistics corresponding to its style. The proposed algorithm perturbs the style of a feature in an adversarial direction towards a randomly selected class. By incorporating the perturbed styles into training, we prevent the model from being misled by the unexpected styles observed in unseen target domains. While RASP is effective for handling domain shifts, its naïve integration into the training procedure is prone to degrade the capability of learning knowledge from source domains due to the feature distortions caused by style perturbation. This challenge is alleviated by Normalized Feature Mixup (NFM) during training, which facilitates learning the original features while achieving robustness to perturbed representations. We evaluate the proposed algorithm via extensive experiments on various benchmarks and show that our approach improves domain generalization performance, especially in large-scale benchmarks. Bohyung Han |
WACV | 2 |
| 2023 | Variational Distribution Learning for Unsupervised Text-to-Image GenerationabstractWe propose a text-to-image generation algorithm based on deep neural networks when text captions for images are unavailable during training. In this work, instead of simply generating pseudo-ground-truth sentences of training images using existing image captioning methods, we employ a pretrained CLIP model, which is capable of properly aligning embeddings of images and corresponding texts in a joint space and, consequently, works well on zero-shot recognition tasks. We optimize a text-to-image generation model by maximizing the data log-likelihood conditioned on pairs of image-text CLIP embeddings. To better align data in the two domains, we employ a principled way based on a variational inference, which efficiently estimates an approximate posterior of the hidden text embedding given an image and its CLIP feature. Experimental results validate that the proposed framework outperforms existing approaches by large margins under unsupervised and semi-supervised text-to-image generation settings. Minsoo Kang, Doyup Lee, Jiseob Kim, Saehoon Kim, Bohyung Han |
CVPR | 5 |
| 2023 | On the Stability-Plasticity Dilemma of Class-Incremental LearningabstractA primary goal of class-incremental learning is to strike a balance between stability and plasticity, where models should be both stable enough to retain knowledge learned from previously seen classes, and plastic enough to learn concepts from new classes. While previous works demonstrate strong performance on class-incremental benchmarks, it is not clear whether their success comes from the models being stable, plastic, or a mixture of both. This paper aims to shed light on how effectively recent class-incremental learning algorithms address the stability-plasticity trade-off. We establish analytical tools that measure the stability and plasticity of feature representations, and employ such tools to investigate models trained with various algorithms on large-scale class-incremental benchmarks. Surprisingly, we find that the majority of class-incremental learning algorithms heavily favor stability over plasticity, to the extent that the feature extractor of a model trained on the initial set of classes is no less effective than that of the final incremental model. Our observations not only inspire two simple algorithms that highlight the importance of feature representation analysis, but also suggest that class-incremental learning approaches, in general, should strive for better feature representation learning. Dongwan Kim, Bohyung Han |
CVPR | 2 |
| 2023 | Open-Set Representation Learning through Combinatorial EmbeddingabstractVisual recognition tasks are often limited to dealing with a small subset of classes simply because the labels for the remaining classes are unavailable. We are interested in identifying novel concepts in a dataset through representation learning based on both labeled and unlabeled examples, and extending the horizon of recognition to both known and novel classes. To address this challenging task, we propose a combinatorial learning approach, which naturally clusters the examples in unseen classes using the compositional knowledge given by multiple supervised meta-classifiers on heterogeneous label spaces. The representations given by the combinatorial embedding are made more robust by unsupervised pairwise relation learning. The proposed algorithm discovers novel concepts via a joint optimization for enhancing the discrimitiveness of unseen classes as well as learning the representations of known classes generalizable to novel ones. Our extensive experiments demonstrate remarkable performance gains by the proposed approach on public datasets for image retrieval and image categorization with novel class discovery. Geeho Kim, Junoh Kang, Bohyung Han |
CVPR | 3 |
| 2023 | Multi-Modal Representation Learning with Text-Driven Soft MasksabstractWe propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. First, we generate diverse features for the image-text matching (ITM) task via soft-masking the regions in an image, which are most relevant to a certain word in the cor-responding caption, instead of completely removing them. Since our framework relies only on image-caption pairs with no fine-grained annotations, we identify the relevant regions to each word by computing the word-conditional vi-sual attention using multi-modal encoder. Second, we encourage the model to focus more on hard but diverse examples by proposing a focal loss for the image-text contrastive learning (ITC) objective, which alleviates the inherent limitations of overfitting and bias issues. Last, we perform multi-modal data augmentations for self-supervised learning via mining various examples by masking texts and rendering distortions on images. We show that the combination of these three innovations is effective for learning a pretrained model, leading to outstanding performance on multiple vision-language downstream tasks. Jaeyoo Park, Bohyung Han |
CVPR | 2 |
| 2023 | Conditional Score Guidance for Text-Driven Image-to-Image TranslationabstractWe present a novel algorithm for text-driven image-to-image translation based on a pretrained text-to-image diffusion model.
Our method aims to generate a target image by selectively editing regions of interest in a source image, defined by a modifying text, while preserving the remaining parts.
In contrast to existing techniques that solely rely on a target prompt, we introduce a new score function that additionally considers both the source image and the source text prompt, tailored to address specific translation tasks.
To this end, we derive the conditional score function in a principled way, decomposing it into the standard score and a guiding term for target image generation.
For the gradient computation about the guiding term, we assume a Gaussian distribution for the posterior distribution and estimate its mean and variance to adjust the gradient without additional training.
In addition, to improve the quality of the conditional score guidance, we incorporate a simple yet effective mixup technique, which combines two cross-attention maps derived from the source and target latents.
This strategy is effective for promoting a desirable fusion of the invariant parts in the source image and the edited regions aligned with the target prompt, leading to high-fidelity target image generation.
Through comprehensive experiments, we demonstrate that our approach achieves outstanding image-to-image translation performance on various tasks.
Code is available at https://github.com/Hleephilip/CSG. Hyunsoo Lee 0005, Minsoo Kang, Bohyung Han |
NeurIPS | 3 |
| 2023 | Generative Neural Fields by Mixtures of Neural Implicit FunctionsabstractWe propose a novel approach to learning the generative neural fields represented by linear combinations of implicit basis networks. Our algorithm learns basis networks in the form of implicit neural representations and their coefficients in a latent space by either conducting meta-learning or adopting auto-decoding paradigms. The proposed method easily enlarges the capacity of generative neural fields by increasing the number of basis networks while maintaining the size of a network for inference to be small through their weighted model averaging. Consequently, sampling instances using the model is efficient in terms of latency and memory footprint. Moreover, we customize denoising diffusion probabilistic model for a target task to sample latent mixture coefficients, which allows our final model to generate unseen data effectively. Experiments show that our approach achieves competitive generation performance on diverse benchmarks for images, voxel data, and NeRF scenes without sophisticated designs for specific modalities and domains. Tackgeun You, Mijeong Kim 0002, Jungtaek Kim 0001, Bohyung Han |
NeurIPS | 4 |
| 2023 | Beyond Pretrained Features: Noisy Image Modeling Provides Adversarial DefenseabstractRecent advancements in masked image modeling (MIM) have made it a prevailing framework for self-supervised visual representation learning. The MIM pretrained models, like most deep neural network methods, remain vulnerable to adversarial attacks, limiting their practical application, and this issue has received little research attention. In this paper, we investigate how this powerful self-supervised learning paradigm can provide adversarial robustness to downstream classifiers. During the exploration, we find that noisy image modeling (NIM), a simple variant of MIM that adopts denoising as the pre-text task, reconstructs noisy images surprisingly well despite severe corruption. Motivated by this observation, we propose an adversarial defense method, referred to as De^3, by exploiting the pretrained decoder for denoising. Through De^3, NIM is able to enhance adversarial robustness beyond providing pretrained features. Furthermore, we incorporate a simple modification, sampling the noise scale hyperparameter from random distributions, and enable the defense to achieve a better and tunable trade-off between accuracy and robustness. Experimental results demonstrate that, in terms of adversarial robustness, NIM is superior to MIM thanks to its effective denoising capability. Moreover, the defense provided by NIM achieves performance on par with adversarial training while offering the extra tunability advantage. Source code and models are available at https://github.com/youzunzhi/NIM-AdvDef. Zunzhi You, Daochang Liu, Bohyung Han, Chang Xu 0002 |
NeurIPS | 3 |
| 2023 | Stop or Forward: Dynamic Layer Skipping for Efficient Action RecognitionabstractOne of the challenges for analyzing video contents (e.g., actions) is high computational cost, especially for the tasks that require processing densely sampled frames in a long video. We present a novel efficient action recognition algorithm, which allocates computational resources adaptively to individual frames depending on their relevance and significance. Specifically, our algorithm adopts LSTM-based policy modules and sequentially estimates the usefulness of each frame based on their intermediate representations. If a certain frame is unlikely to be helpful for recognizing actions, our model stops forwarding the features to the rest of the layers and starts to consider the next sampled frame. We further reduce the computational cost of our approach by introducing a simple yet effective early termination strategy during the inference procedure. We evaluate the proposed algorithm on three public benchmarks: ActivityNet-v1.3, Mini-Kinetics, and THUMOS’14. Our experiments show that the proposed approach achieves outstanding trade-off between accuracy and efficiency in action recognition. Jong-Hyeon Seon, Jaedong Hwang, Jonghwan Mun, Bohyung Han |
WACV | 4 |
| 2023 | End-to-end learning for weakly supervised video anomaly detection using Absorbing Markov Chain
Jaeyoo Park, Junha Kim, Bohyung Han |
Comput. Vis. Image Underst. | 3 |
| 2022 | Information-Theoretic Bias Reduction via Causal View of Spurious CorrelationabstractWe propose an information-theoretic bias measurement technique through a causal interpretation of spurious correlation, which is effective to identify the feature-level algorithmic bias by taking advantage of conditional mutual information. Although several bias measurement methods have been proposed and widely investigated to achieve algorithmic fairness in various tasks such as face recognition, their accuracy- or logit-based metrics are susceptible to leading to trivial prediction score adjustment rather than fundamental bias reduction. Hence, we design a novel debiasing framework against the algorithmic bias, which incorporates a bias regularization loss derived by the proposed information-theoretic bias measurement approach. In addition, we present a simple yet effective unsupervised debiasing technique based on stochastic label noise, which does not require the explicit supervision of bias information. The proposed bias measurement and debiasing approaches are validated in diverse realistic scenarios through extensive experiments on multiple standard benchmarks. Seonguk Seo, Joon-Young Lee, Bohyung Han |
AAAI | 3 |
| 2022 | InfoNeRF: Ray Entropy Minimization for Few-Shot Neural Volume RenderingabstractWe present an information-theoretic regularization technique for few-shot novel view synthesis based on neural im-plicit representation. The proposed approach minimizes potential reconstruction inconsistency that happens due to in-sufficient viewpoints by imposing the entropy constraint of the density in each ray. In addition, to alleviate the poten-tial degenerate issue when all training images are acquired from almost redundant viewpoints, we further incorporate the spatial smoothness constraint into the estimated images by restricting information gains from additional rays with slightly different viewpoints. The main idea of our algorithm is to make reconstructed scenes compact along indi-vidual rays and consistent across rays in the neighborhood. The proposed regularizers can be plugged into most of existing neural volume rendering techniques based on NeRF in a straightforward way. Despite its simplicity, we achieve con-sistently improved performance compared to existing neural view synthesis methods by large margins on multiple stan-dard benchmarks. Our codes and models are available in the project website11http://cvlab.snu.ac.kr/research/InfoNeRF. Mijeong Kim 0002, Seonguk Seo, Bohyung Han |
CVPR | 3 |
| 2022 | Pooling Revisited: Your Receptive Field is SuboptimalabstractThe size and shape of the receptive field determine how the network aggregates local features, and affect the overall performance of a model considerably. Many components in a neural network, such as depth, kernel sizes, and strides for convolution and pooling, influence the receptive field. However, they still rely on hyperparameters, and the receptive fields of existing models result in suboptimal shapes and sizes. Hence, we propose a simple yet effective Dynamically Optimized Pooling operation, referred to as DynOPool, which learns the optimized scale factors offeature maps end-to-end. Moreover, DynOPool determines the proper resolution of a feature map by learning the desirable size and shape of its receptive field, which allows an operator in a deeper layer to observe an input image in the optimal scale. Any kind of resizing modules in a deep neural network can be replaced by DynOPool with minimal cost. Also, DynOPool controls the complexity of the model by introducing an additional loss term that constrains computational cost. Our experiments show that the models equipped with the proposed learnable resizing module outperform the baseline algorithms on multiple datasets in image classification and semantic segmentation. Dong-Hwan Jang, Sanghyeok Chu, Joonhyuk Kim, Bohyung Han |
CVPR | 4 |
| 2022 | Class-Incremental Learning by Knowledge Distillation with Adaptive Feature ConsolidationabstractWe present a novel class incremental learning approach based on deep neural networks, which continually learns new tasks with limited memory for storing examples in the previous tasks. Our algorithm is based on knowledge distillation and provides a principled way to maintain the representations of old models while adjusting to new tasks effectively. The proposed method estimates the relationship between the representation changes and the resulting loss increases incurred by model updates. It minimizes the upper bound of the loss increases using the representations, which exploits the estimated importance of each feature map within a backbone model. Based on the importance, the model restricts updates of important features for robustness while allowing changes in less critical features for flexibility. This optimization strategy effectively alleviates the notorious catastrophic forgetting problem despite the limited accessibility of data in the previous tasks. The experimental results show significant accuracy improvement of the proposed algorithm over the existing methods on the standard datasets. Code is available.11https://github.com/kminsoo/AFC Minsoo Kang, Jaeyoo Park, Bohyung Han |
CVPR | 3 |
| 2022 | Self-Supervised Dense Consistency Regularization for Image-to-Image TranslationabstractUnsupervised image-to-image translation has gained considerable attention due to recent impressive advances in generative adversarial networks (GANs). This paper presents a simple but effective regularization technique for improving GAN-based image-to-image translation. To generate images with realistic local semantics and structures, we propose an auxiliary self-supervision loss that enforces point-wise consistency of the overlapping region between a pair of patches cropped from a single real image during training the discriminator of a GAN. Our experiment shows that the proposed dense consistency regularization improves performance substantially on various image-to-image translation scenarios. It also leads to extra performance gains through the combination with instance-level regularization methods. Furthermore, we verify that the proposed model captures domain-specific characteristics more effectively with only a small fraction of training data. Minsu Ko, Eun Ju Cha, Sungjoo Suh, Huijin Lee, Jae-Joon Han, Jinwoo Shin, Bohyung Han |
CVPR | 7 |
| 2022 | Unsupervised Learning of Debiased Representations with Pseudo-AttributesabstractDataset bias is a critical challenge in machine learning since it often leads to a negative impact on a model due to the unintended decision rules captured by spurious correlations. Although existing works often handle this issue based on human supervision, the availability of the proper annotations is impractical and even unrealistic. To better tackle the limitation, we propose a simple but effective unsupervised debiasing technique. Specifically, we first identify pseudo-attributes based on the results from clustering performed in the feature embedding space even without an explicit bias attribute supervision. Then, we employ a novel cluster-wise reweighting scheme to learn debiased representation; the proposed method prevents minority groups from being discounted for minimizing the overall loss, which is desirable for worst-case generalization. The extensive experiments demonstrate the outstanding performance of our approach on multiple standard benchmarks, even achieving the competitive accuracy to the supervised counterpart. The source code is available at our project page11https://github.com/skynbe/pseudo-attributes . Seonguk Seo, Joon-Young Lee, Bohyung Han |
CVPR | 3 |
| 2022 | Towards Sequence-Level Training for Visual Tracking
Minji Kim 0002, Seungkwan Lee, Jungseul Ok, Bohyung Han, Minsu Cho |
ECCV (22) | 4 |
| 2022 | Learning Semantic Segmentation from Multiple Datasets with Label Shifts
Dongwan Kim, Yi-Hsuan Tsai, Yumin Suh, Masoud Faraki, Sparsh Garg, Manmohan Krishna Chandraker, Bohyung Han |
ECCV (28) | 7 |
| 2022 | Multi-Level Branched Regularization for Federated LearningabstractA critical challenge of federated learning is data heterogeneity and imbalance across clients, which leads to inconsistency between local networks and unstable convergence of global models. To alleviate the limitations, we propose a novel architectural regularization technique that constructs multiple auxiliary branches in each local model by grafting local and global subnetworks at several different levels and that learns the representations of the main pathway in the local model congruent to the auxiliary hybrid pathways via online knowledge distillation. The proposed technique is effective to robustify the global model even in the non-iid setting and is applicable to various federated learning frameworks conveniently without incurring extra communication costs. We perform comprehensive empirical studies and demonstrate remarkable performance gains in terms of accuracy and efficiency compared to existing methods. The source code is available at our project page. Jinkyu Kim 0005, Geeho Kim, Bohyung Han |
ICML | 3 |
| 2022 | Online Hybrid Lightweight Representations Learning: Its Application to Visual TrackingabstractThis paper presents a novel hybrid representation learning framework for streaming data, where an image frame in a video is modeled by an ensemble of two distinct deep neural networks; one is a low-bit quantized network and the other is a lightweight full-precision network. The former learns coarse primary information with low cost while the latter conveys residual information for high fidelity to original representations. The proposed parallel architecture is effective to maintain complementary information since fixed-point arithmetic can be utilized in the quantized network and the lightweight model provides precise representations given by a compact channel-pruned network. We incorporate the hybrid representation technique into an online visual tracking task, where deep neural networks need to handle temporal variations of target appearances in real-time. Compared to the state-of-the-art real-time trackers based on conventional deep neural networks, our tracking algorithm demonstrates competitive accuracy on the standard benchmarks with a small fraction of computational cost and memory footprint. Ilchae Jung, Minji Kim 0002, Eunhyeok Park, Bohyung Han |
IJCAI | 4 |
| 2022 | MCL-GAN: Generative Adversarial Networks with Multiple Specialized DiscriminatorsabstractWe propose a framework of generative adversarial networks with multiple discriminators, which collaborate to represent a real dataset more effectively. Our approach facilitates learning a generator consistent with the underlying data distribution based on real images and thus mitigates the chronic mode collapse problem. From the inspiration of multiple choice learning, we guide each discriminator to have expertise in a subset of the entire data and allow the generator to find reasonable correspondences between the latent and real data spaces automatically without extra supervision for training examples. Despite the use of multiple discriminators, the backbone networks are shared across the discriminators and the increase in training cost is marginal. We demonstrate the effectiveness of our algorithm using multiple evaluation metrics in the standard datasets for diverse tasks. Bohyung Han |
NeurIPS | 2 |
| 2022 | Information-Theoretic GAN Compression with Variational Energy-based ModelabstractWe propose an information-theoretic knowledge distillation approach for the compression of generative adversarial networks, which aims to maximize the mutual information between teacher and student networks via a variational optimization based on an energy-based model. Because the direct computation of the mutual information in continuous domains is intractable, our approach alternatively optimizes the student network by maximizing the variational lower bound of the mutual information. To achieve a tight lower bound, we introduce an energy-based model relying on a deep neural network to represent a flexible variational distribution that deals with high-dimensional images and consider spatial dependencies between pixels, effectively. Since the proposed method is a generic optimization algorithm, it can be conveniently incorporated into arbitrary generative adversarial networks and even dense prediction networks, e.g., image enhancement models. We demonstrate that the proposed algorithm achieves outstanding performance in model compression of generative adversarial networks consistently when combined with several existing models. Minsoo Kang, Hyewon Yoo, Eunhee Kang, Sehwan Ki, Hyong-Euk Lee, Bohyung Han |
NeurIPS | 6 |
| 2022 | Locally Hierarchical Auto-Regressive Modeling for Image GenerationabstractWe propose a locally hierarchical auto-regressive model with multiple resolutions of discrete codes. In the first stage of our algorithm, we represent an image with a pyramid of codes using Hierarchically Quantized Variational AutoEncoder (HQ-VAE), which disentangles the information contained in the multi-level codes. For an example of two-level codes, we create two separate pathways to carry high-level coarse structures of input images using top codes while compensating for missing fine details by constructing a residual connection for bottom codes. An appropriate selection of resizing operations for code embedding maps enables top codes to capture maximal information within images and the first stage algorithm achieves better performance on both vector quantization and image generation. The second stage adopts Hierarchically Quantized Transformer (HQ-Transformer) to process a sequence of local pyramids, which consist of a single top code and its corresponding bottom codes. Contrary to other hierarchical models, we sample bottom codes in parallel by exploiting the conditional independence assumption on the bottom codes. This assumption is naturally harvested from our first-stage model, HQ-VAE, where the bottom code learns to describe local details. On class-conditional and text-conditional generation benchmarks, our model shows competitive performance to previous AR models in terms of fidelity of generated images while enjoying lighter computational budgets. Tackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee, Bohyung Han |
NeurIPS | 5 |
| 2022 | Fine-grained neural architecture search for image super-resolution
Seokil Hong, Bohyung Han, Heesoo Myeong, Kyoung Mu Lee |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Beyond homography: nonparametric image alignment via graph convolutional networks
Mijeong Kim 0002, Sanghyeok Chu, Bohyung Han |
Mach. Vis. Appl. | 3 |
| 2021 | Exemplar-Based Open-Set Panoptic Segmentation NetworkabstractWe extend panoptic segmentation to the open-world and introduce an open-set panoptic segmentation (OPS) task. This task requires performing panoptic segmentation for not only known classes but also unknown ones that have not been acknowledged during training. We investigate the practical challenges of the task and construct a benchmark on top of an existing dataset, COCO. In addition, we propose a novel exemplar-based open-set panoptic segmentation network (EOPSN) inspired by exemplar theory. Our approach identifies a new class based on exemplars, which are identified by clustering and employed as pseudoground-truths. The size of each class increases by mining new exemplars based on the similarities to the existing ones associated with the class. We evaluate EOPSN on the proposed benchmark and demonstrate the effectiveness of our proposals. The primary goal of our work is to draw the attention of the community to the recognition in the open- world scenarios. The implementation of our algorithm is available on the project webpage1. Jaedong Hwang, Seoung Wug Oh, Joon-Young Lee, Bohyung Han |
CVPR | 4 |
| 2021 | CoSMo: Content-Style Modulation for Image Retrieval With Text FeedbackabstractWe tackle the task of image retrieval with text feedback, where a reference image and modifier text are combined to identify the desired target image. We focus on designing an image-text compositor, i.e., integrating multi-modal inputs to produce a representation similar to that of the target image. In our algorithm, Content-Style Modulation (CoSMo), we approach this challenge by introducing two modules based on deep neural networks: the content and style modulators. The content modulator performs local updates to the reference image feature after normalizing the style of the image, where a disentangled multi-modal non-local block is employed to achieve the desired content modifications. Then, the style modulator reintroduces global style information to the updated feature. We provide an in-depth view of our algorithm and its design choices, and show that it accomplishes outstanding performance on multiple image-text retrieval benchmarks. Our code can be found at: https://github.com/postBG/CosMo.pytorch Dongwan Kim, Bohyung Han |
CVPR | 3 |
| 2021 | RaScaNet: Learning Tiny Models by Raster-Scanning ImagesabstractDeploying deep convolutional neural networks on ultra-low power systems is challenging due to the extremely limited resources. Especially, the memory becomes a bottleneck as the systems put a hard limit on the size of on-chip memory. Because peak memory explosion in the lower layers is critical even in tiny models, the size of an input image should be reduced with sacrifice in accuracy. To overcome this drawback, we propose a novel Raster-Scanning Network, named RaScaNet, inspired by raster-scanning in image sensors. RaScaNet reads only a few rows of pixels at a time using a convolutional neural network and then sequentially learns the representation of the whole image using a recurrent neural network. The proposed method operates on an ultra-low power system without input size reduction; it requires 15.9–24.3× smaller peak memory and 5.3–12.9× smaller weight memory than the state-of-the-art tiny models. Moreover, RaScaNet fully exploits on-chip SRAM and cache memory of the system as the sum of the peak memory and the weight memory does not exceed 60 KB, improving the power efficiency of the system. In our experiments, we demonstrate the binary classification performance of RaScaNet on Visual Wake Words and Pascal VOC datasets. Jaehyoung Yoo, Changyong Son, Sangil Jung, ByungIn Yoo, Changkyu Choi, Jae-Joon Han, Bohyung Han |
CVPR | 8 |
| 2021 | Class-Incremental Learning for Action Recognition in VideosabstractWe tackle catastrophic forgetting problem in the context of class-incremental learning for video recognition, which has not been explored actively despite the popularity of continual learning. Our framework addresses this challenging task by introducing time-channel importance maps and exploiting the importance maps for learning the representations of incoming examples via knowledge distillation. We also incorporate a regularization scheme in our objective function, which encourages individual features obtained from different time steps in a video to be uncorrelated and eventually improves accuracy by alleviating catastrophic forgetting. We evaluate the proposed approach on brand-new splits of class-incremental action recognition benchmarks constructed upon the UCF101, HMDB51, and Something-Something V2 datasets, and demonstrate the effectiveness of our algorithm in comparison to the existing continual learning methods that are originally designed for image data. Jaeyoo Park, Minsoo Kang, Bohyung Han |
ICCV | 3 |
| 2021 | Variable-Rate Deep Image Compression through Spatially-Adaptive Feature TransformabstractWe propose a versatile deep image compression network based on Spatial Feature Transform (SFT) [45], which takes a source image and a corresponding quality map as inputs and produce a compressed image with variable rates. Our model covers a wide range of compression rates using a single model, which is controlled by arbitrary pixel-wise quality maps. In addition, the proposed framework allows us to perform task-aware image compressions for various tasks, e.g., classification, by efficiently estimating optimized quality maps specific to target tasks for our encoding network. This is even possible with a pretrained network without learning separate models for individual tasks. Our algorithm achieves outstanding rate-distortion trade-off compared to the approaches based on multiple models that are optimized separately for several different target rates. At the same level of compression, the proposed approach successfully improves performance on image classification and text region quality preservation via task-aware quality map estimation without additional model training. The code is available at the project website1. Myungseo Song, Bohyung Han |
ICCV | 3 |
| 2021 | Learning Debiased and Disentangled Representations for Semantic SegmentationabstractDeep neural networks are susceptible to learn biased models with entangled feature representations, which may lead to subpar performances on various downstream tasks. This is particularly true for under-represented classes, where a lack of diversity in the data exacerbates the tendency. This limitation has been addressed mostly in classification tasks, but there is little study on additional challenges that may appear in more complex dense prediction problems including semantic segmentation. To this end, we propose a model-agnostic and stochastic training scheme for semantic segmentation, which facilitates the learning of debiased and disentangled representations. For each class, we first extract class-specific information from the highly entangled feature map. Then, information related to a randomly sampled class is suppressed by a feature selection process in the feature space. By randomly eliminating certain class information in each training iteration, we effectively reduce feature dependencies among classes, and the model is able to learn more debiased and disentangled feature representations. Models trained with our approach demonstrate strong results on multiple semantic segmentation benchmarks, with especially notable performance gains on under-represented classes. Sanghyeok Chu, Dongwan Kim, Bohyung Han |
NeurIPS | 3 |
| 2021 | Learning Student-Friendly Teacher Networks for Knowledge DistillationabstractWe propose a novel knowledge distillation approach to facilitate the transfer of dark knowledge from a teacher to a student. Contrary to most of the existing methods that rely on effective training of student models given pretrained teachers, we aim to learn the teacher models that are friendly to students and, consequently, more appropriate for knowledge transfer. In other words, at the time of optimizing a teacher model, the proposed algorithm learns the student branches jointly to obtain student-friendly representations. Since the main goal of our approach lies in training teacher models and the subsequent knowledge distillation procedure is straightforward, most of the existing knowledge distillation methods can adopt this technique to improve the performance of diverse student models in terms of accuracy and convergence speed. The proposed algorithm demonstrates outstanding accuracy in several well-known knowledge distillation techniques with various combinations of teacher and student models even in the case that their architectures are heterogeneous and there is no prior knowledge about student models at the time of training teacher networks Moon-Hyun Cha, Changwook Jeong, Daesin Kim, Bohyung Han |
NeurIPS | 5 |
| 2021 | Weakly Supervised Instance Segmentation by Deep Community LearningabstractWe present a weakly supervised instance segmentation algorithm based on deep community learning with multiple tasks. This task is formulated as a combination of weakly supervised object detection and semantic segmentation, where individual objects of the same class are identified and segmented separately. We address this problem by designing a unified deep neural network architecture, which has a positive feedback loop of object detection with bounding box regression, instance mask generation, instance segmentation, and feature extraction. Each component of the network makes active interactions with others to improve accuracy, and the end-to-end trainability of our model makes our results more robust and reproducible. The proposed algorithm achieves state-of-the-art performance in the weakly supervised setting without any additional training such as Fast R-CNN and Mask R-CNN on the standard benchmark dataset. The implementation of our algorithm is available on the project webpage: https://cv.snu.ac.kr/research/WSIS_CL. Jaedong Hwang, Jeany Son, Bohyung Han |
WACV | 4 |
| 2020 | Channel Attention Is All You Need for Video Frame InterpolationabstractPrevailing video frame interpolation techniques rely heavily on optical flow estimation and require additional model complexity and computational cost; it is also susceptible to error propagation in challenging scenarios with large motion and heavy occlusion. To alleviate the limitation, we propose a simple but effective deep neural network for video frame interpolation, which is end-to-end trainable and is free from a motion estimation network component. Our algorithm employs a special feature reshaping operation, referred to as PixelShuffle, with a channel attention, which replaces the optical flow computation module. The main idea behind the design is to distribute the information in a feature map into multiple channels and extract motion information by attending the channels for pixel-level frame synthesis. The model given by this principle turns out to be effective in the presence of challenging motion and occlusion. We construct a comprehensive evaluation benchmark and demonstrate that the proposed approach achieves outstanding performance compared to the existing models with a component for optical flow computation. Myungsub Choi, Bohyung Han, Kyoung Mu Lee |
AAAI | 3 |
| 2020 | Real-Time Object Tracking via Meta-Learning: Efficient Model Adaptation and One-Shot Channel PruningabstractWe propose a novel meta-learning framework for real-time object tracking with efficient model adaptation and channel pruning. Given an object tracker, our framework learns to fine-tune its model parameters in only a few gradient-descent iterations during tracking while pruning its network channels using the target ground-truth at the first frame. Such a learning problem is formulated as a meta-learning task, where a meta-tracker is trained by updating its meta-parameters for initial weights, learning rates, and pruning masks through carefully designed tracking simulations. The integrated meta-tracker greatly improves tracking performance by accelerating the convergence of online learning and reducing the cost of feature computation. Experimental evaluation on the standard datasets demonstrates its outstanding accuracy and speed compared to the state-of-the-art methods. Ilchae Jung, Kihyun You, Hyeonwoo Noh, Minsu Cho, Bohyung Han |
AAAI | 5 |
| 2020 | Towards Oracle Knowledge Distillation with Neural Architecture SearchabstractWe present a novel framework of knowledge distillation that is capable of learning powerful and efficient student models from ensemble teacher networks. Our approach addresses the inherent model capacity issue between teacher and student and aims to maximize benefit from teacher models during distillation by reducing their capacity gap. Specifically, we employ a neural architecture search technique to augment useful structures and operations, where the searched network is appropriate for knowledge distillation towards student models and free from sacrificing its performance by fixing the network capacity. We also introduce an oracle knowledge distillation loss to facilitate model search and distillation using an ensemble-based teacher model, where a student network is learned to imitate oracle performance of the teacher. We perform extensive experiments on the image classification datasets—CIFAR-100 and TinyImageNet—using various networks. We also show that searching for a new student model is effective in both accuracy and memory size and that the searched models often outperform their teacher models thanks to neural architecture search with oracle knowledge distillation. Minsoo Kang, Jonghwan Mun, Bohyung Han |
AAAI | 3 |
| 2020 | Context-Aware Zero-Shot RecognitionabstractWe present a novel problem setting in zero-shot learning, zero-shot object recognition and detection in the context. Contrary to the traditional zero-shot learning methods, which simply infers unseen categories by transferring knowledge from the objects belonging to semantically similar seen categories, we aim to understand the identity of the novel objects in an image surrounded by the known objects using the inter-object relation prior. Specifically, we leverage the visual context and the geometric relationships between all pairs of objects in a single image, and capture the information useful to infer unseen categories. We integrate our context-aware zero-shot learning framework into the traditional zero-shot learning techniques seamlessly using a Conditional Random Field (CRF). The proposed algorithm is evaluated on both zero-shot region classification and zero-shot detection tasks. The results on Visual Genome (VG) dataset show that our model significantly boosts performance with the additional visual context compared to traditional methods. Ruotian Luo, Bohyung Han |
AAAI | 3 |
| 2020 | Reinforcing an Image Caption Generator Using Off-Line Human FeedbackabstractHuman ratings are currently the most accurate way to assess the quality of an image captioning model, yet most often the only used outcome of an expensive human rating evaluation is a few overall statistics over the evaluation dataset. In this paper, we show that the signal from instance-level human caption ratings can be leveraged to improve captioning models, even when the amount of caption ratings is several orders of magnitude less than the caption training data. We employ a policy gradient method to maximize the human ratings as rewards in an off-policy reinforcement learning setting, where policy gradients are estimated by samples from a distribution that focuses on the captions in a caption ratings dataset. Our empirical evidence indicates that the proposed method learns to generalize the human raters' judgments to a previously unseen set of images, as judged by a different set of human judges, and additionally on a different, multi-dimensional side-by-side human evaluation procedure. Hongsuck Seo, Piyush Sharma, Tomer Levinboim, Bohyung Han, Radu Soricut |
AAAI | 4 |
| 2020 | Learning to Adapt to Unseen Abnormal Activities Under Weak Supervision
Jaeyoo Park, Junha Kim, Bohyung Han |
ACCV (5) | 3 |
| 2020 | Local-Global Video-Text Interactions for Temporal GroundingabstractThis paper addresses the problem of text-to-video temporal grounding, which aims to identify the time interval in a video semantically relevant to a text query. We tackle this problem using a novel regression-based model that learns to extract a collection of mid-level features for semantic phrases in a text query, which corresponds to important semantic entities described in the query (e.g., actors, objects, and actions), and reflect bi-modal interactions between the linguistic features of the query and the visual features of the video in multiple levels. The proposed method effectively predicts the target time interval by exploiting contextual information from local to global during bi-modal interactions. Through in-depth ablation studies, we find out that incorporating both local and global context in video and text interactions is crucial to the accurate grounding. Our experiment shows that the proposed method outperforms the state of the arts on Charades-STA and ActivityNet Captions datasets by large margins, 7.44% and 4.61% points at Recall@tIoU=0.5 metric, respectively. Jonghwan Mun, Minsu Cho, Bohyung Han |
CVPR | 3 |
| 2020 | Task-Aware Quantization Network for JPEG Image Compression
Jin Young Choi 0002, Bohyung Han |
ECCV (20) | 2 |
| 2020 | URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark
Seonguk Seo, Joon-Young Lee, Bohyung Han |
ECCV (15) | 3 |
| 2020 | Learning to Optimize Domain Specific Normalization for Domain Generalization
Seonguk Seo, Yumin Suh, Dongwan Kim, Geeho Kim, Jongwoo Han, Bohyung Han |
ECCV (22) | 6 |
| 2020 | Traffic Accident Benchmark for Causality Recognition
Tackgeun You, Bohyung Han |
ECCV (7) | 2 |
| 2020 | Efficient Decoupled Neural Architecture Search by Structure And Operation SamplingabstractWe propose a novel neural architecture search algorithm via reinforcement learning by decoupling structure and operation search. Our approach samples candidate models from the multinomial distribution over the policy vectors. The proposed technique improves the efficiency of architecture search significantly compared to the existing methods while achieving competitive classification accuracy and model compactness. Our policy vectors are easily interpretable throughout the training procedure, which allows analyzing the search progress and the identified architectures. Note that, on the contrary, the black-box characteristics of the conventional methods based on RNN controllers hamper understanding training progress in terms of policy parameter updates. Our experiments demonstrate the outstanding performance of our approach compared to the state-of-the-art techniques with a fraction of search cost. Heung-Chang Lee, Do-Guk Kim, Bohyung Han |
ICASSP | 3 |
| 2020 | Operation-Aware Soft Channel Pruning using Differentiable MasksabstractWe propose a simple but effective data-driven channel pruning algorithm, which compresses deep neural networks in a differentiable way by exploiting the characteristics of operations. The proposed approach makes a joint consideration of batch normalization (BN) and rectified linear unit (ReLU) for channel pruning; it estimates how likely the two successive operations deactivate each feature map and prunes the channels with high probabilities. To this end, we learn differentiable masks for individual channels and make soft decisions throughout the optimization procedure, which facilitates to explore larger search space and train more stable networks. The proposed framework enables us to identify compressed models via a joint learning of model parameters and channel pruning without an extra procedure of fine-tuning. We perform extensive experiments and achieve outstanding performance in terms of the accuracy of output networks given the same amount of resources when compared with the state-of-the-art methods. Minsoo Kang, Bohyung Han |
ICML | 2 |
| 2020 | Rotation-Invariant Local-to-Global Representation Learning for 3D Point CloudabstractWe propose a local-to-global representation learning algorithm for 3D point cloud data, which is appropriate to handle various geometric transformations, especially rotation, without explicit data augmentation with respect to the transformations. Our model takes advantage of multi-level abstraction based on graph convolutional neural networks, which constructs a descriptor hierarchy to encode rotation-invariant shape information of an input object in a bottom-up manner. The descriptors in each level are obtained from a neural network based on a graph via stochastic sampling of 3D points, which is effective in making the learned representations robust to the variations of input data. The proposed algorithm presents the state-of-the-art performance on the rotation-augmented 3D object recognition and segmentation benchmarks, and we further analyze its characteristics through comprehensive ablative experiments. Jaeyoo Park, Bohyung Han |
NeurIPS | 3 |
| 2019 | Regularizing Neural Networks via Stochastic Branch LayersabstractWe introduce a novel stochastic regularization technique for deep neural networks, which decomposes a layer into multiple branches with different parameters and merges stochastically sampled combinations of the outputs from the branches during training. Since the factorized branches can collapse into a single branch through a linear operation, inference requires no additional complexity compared to the ordinary layers. The proposed regularization method, referred to as StochasticBranch, is applicable to any linear layers such as fully-connected or convolution layers. The proposed regularizer allows the model to explore diverse regions of the model parameter space via multiple combinations of branches to find better local minima. An extensive set of experiments shows that our method effectively regularizes networks and further improves the generalization performance when used together with other existing regularization techniques. Wonpyo Park, Hongsuck Seo, Bohyung Han, Minsu Cho |
ACML | 3 |
| 2019 | Domain-Specific Batch Normalization for Unsupervised Domain AdaptationabstractWe propose a novel unsupervised domain adaptation framework based on domain-specific batch normalization in deep neural networks. We aim to adapt to both domains by specializing batch normalization layers in convolutional neural networks while allowing them to share all other model parameters, which is realized by a two-stage algorithm. In the first stage, we estimate pseudo-labels for the examples in the target domain using an external unsupervised domain adaptation algorithm-for example, MSTN or CPUA-integrating the proposed domain-specific batch normalization. The second stage learns the final models using a multi-task classification loss for the source and target domains. Note that the two domains have separate batch normalization layers in both stages. Our framework can be easily incorporated into the domain adaptation techniques based on deep neural networks with batch normalization layers. We also present that our approach can be extended to the problem with multiple source domains. The proposed algorithm is evaluated on multiple benchmark datasets and achieves the state-of-the-art accuracy in the standard setting and the multi-source domain adaption scenario. Woong-Gi Chang, Tackgeun You, Seonguk Seo, Suha Kwak, Bohyung Han |
CVPR | 5 |
| 2019 | Streamlined Dense Video CaptioningabstractDense video captioning is an extremely challenging task since accurate and coherent description of events in a video requires holistic understanding of video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first detecting event proposals from a video and then captioning on a subset of the proposals. As a result, the generated sentences are prone to be redundant or inconsistent since they fail to consider temporal dependency between events. To tackle this challenge, we propose a novel dense video captioning framework, which models temporal dependency across events in a video explicitly and leverages visual and linguistic context from prior events for coherent storytelling. This objective is achieved by 1) integrating an event sequence generation network to select a sequence of event proposals adaptively, and 2) feeding the sequence of event proposals to our sequential video captioning network, which is trained by reinforcement learning with two-level rewards - at both event and episode levels - for better context modeling. The proposed technique achieves outstanding performances on ActivityNet Captions dataset in most metrics. Jonghwan Mun, Zhou Ren, Bohyung Han |
CVPR | 5 |
| 2019 | Transfer Learning via Unsupervised Task Discovery for Visual Question AnsweringabstractWe study how to leverage off-the-shelf visual and linguistic data to cope with out-of-vocabulary answers in visual question answering task. Existing large-scale visual datasets with annotations such as image class labels, bounding boxes and region descriptions are good sources for learning rich and diverse visual concepts. However, it is not straightforward how the visual concepts can be captured and transferred to visual question answering models due to missing link between question dependent answering models and visual data without question. We tackle this problem in two steps: 1) learning a task conditional visual classifier, which is capable of solving diverse question-specific visual recognition tasks, based on unsupervised task discovery and 2) transferring the task conditional visual classifier to visual question answering models. Specifically, we employ linguistic knowledge sources such as structured lexical database (e.g. WordNet) and visual descriptions for unsupervised task discovery, and transfer a learned task conditional visual classifier as an answering unit in a visual question answering model. We empirically show that the proposed algorithm generalizes to out-of-vocabulary answers successfully using the knowledge transferred from the visual dataset. Hyeonwoo Noh, Jonghwan Mun, Bohyung Han |
CVPR | 4 |
| 2019 | Learning for Single-Shot Confidence Calibration in Deep Neural Networks Through Stochastic InferencesabstractWe propose a generic framework to calibrate accuracy and confidence of a prediction in deep neural networks through stochastic inferences. We interpret stochastic regularization using a Bayesian model, and analyze the relation between predictive uncertainty of networks and variance of the prediction scores obtained by stochastic inferences for a single example. Our empirical study shows that the accuracy and the score of a prediction are highly correlated with the variance of multiple stochastic inferences given by stochastic depth or dropout. Motivated by this observation, we design a novel variance-weighted confidence-integrated loss function that is composed of two cross-entropy loss terms with respect to ground-truth and uniform distribution, which are balanced by variance of stochastic prediction scores. The proposed loss function enables us to learn deep neural networks that predict confidence calibrated scores using a single inference. Our algorithm presents outstanding confidence calibration performance and improves classification accuracy when combined with two popular stochastic regularization techniques-stochastic depth and dropout-in multiple models and datasets; it alleviates overconfidence issue in deep neural networks significantly by training networks to achieve prediction accuracy proportional to confidence of prediction. Seonguk Seo, Hongsuck Seo, Bohyung Han |
CVPR | 3 |
| 2019 | Stochastic Class-Based Hard Example Mining for Deep Metric LearningabstractPerformance of deep metric learning depends heavily on the capability of mining hard negative examples during training. However, many metric learning algorithms often require intractable computational cost due to frequent feature computations and nearest neighbor searches in a large-scale dataset. As a result, existing approaches often suffer from trade-off between training speed and prediction accuracy. To alleviate this limitation, we propose a stochastic hard negative mining method. Our key idea is to adopt class signatures that keep track of feature embedding online with minor additional cost during training, and identify hard negative example candidates using the signatures. Given an anchor instance, our algorithm first selects a few hard negative classes based on the class-to-sample distances and then performs a refined search in an instance-level only from the selected classes. As most of the classes are discarded at the first step, it is much more efficient than exhaustive search while effectively mining a large number of hard examples. Our experiment shows that the proposed technique improves image retrieval accuracy substantially; it achieves the state-of-the-art performance on the several standard benchmark datasets. Yumin Suh, Bohyung Han, Wonsik Kim, Kyoung Mu Lee |
CVPR | 2 |
| 2019 | Continual Learning by Asymmetric Loss Approximation With Single-Side OverestimationabstractCatastrophic forgetting is a critical challenge in training deep neural networks. Although continual learning has been investigated as a countermeasure to the problem, it often suffers from the requirements of additional network components and the limited scalability to a large number of tasks. We propose a novel approach to continual learning by approximating a true loss function using an asymmetric quadratic function with one of its sides overestimated. Our algorithm is motivated by the empirical observation that the network parameter updates affect the target loss functions asymmetrically. In the proposed continual learning framework, we estimate an asymmetric loss function for the tasks considered in the past through a proper overestimation of its unobserved sides in training new tasks, while deriving the accurate model parameter for the observable sides. In contrast to existing approaches, our method is free from the side effects and achieves the state-of-the-art accuracy that is even close to the upper-bound performance on several challenging benchmark datasets. Dongmin Park, Seokil Hong, Bohyung Han, Kyoung Mu Lee |
ICCV | 3 |
| 2019 | Combinatorial Inference against Label NoiseabstractLabel noise is one of the critical sources that degrade generalization performance of deep neural networks significantly. To handle the label noise issue in a principled way, we propose a unique classification framework of constructing multiple models in heterogeneous coarse-grained meta-class spaces and making joint inference of the trained models for the final predictions in the original (base) class space. Our approach reduces noise level by simply constructing meta-classes and improves accuracy via combinatorial inferences over multiple constituent classifiers. Since the proposed framework has distinct and complementary properties for the given problem, we can even incorporate additional off-the-shelf learning algorithms to improve accuracy further. We also introduce techniques to organize multiple heterogeneous meta-class sets using $k$-means clustering and identify a desirable subset leading to learn compact models. Our extensive experiments demonstrate outstanding performance in terms of accuracy and efficiency compared to the state-of-the-art methods under various synthetic noise configurations and in a real-world noisy dataset. Hongsuck Seo, Geeho Kim, Bohyung Han |
NeurIPS | 3 |
| 2018 | Product Quantized Translation for Fast Nearest Neighbor SearchabstractThis paper proposes a simple nearest neighbor search algorithm, which provides the exact solution in terms of the Euclidean distance efficiently. Especially, we present an interesting approach to improve the speed of nearest neighbor search by proper translations of data and query although the task is inherently invariant to the Euclidean transformations. The proposed algorithm aims to eliminate nearest neighbor candidates effectively using their distance lower bounds in nonlinear embedded spaces, and further improves the lower bounds by transforming data and query through product quantized translations. Although our framework is composed of simple operations only, it achieves the state-of-the-art performance compared to existing nearest neighbor search techniques, which is illustrated quantitatively using various large-scale benchmark datasets in different sizes and dimensions. Yoonho Hwang, Mooyeol Baek, Saehoon Kim, Bohyung Han, Hee-Kap Ahn |
AAAI | 4 |
| 2018 | Forget and Diversify: Regularized Refinement for Weakly Supervised Object Detection
Jeany Son, Solae Lee, Suha Kwak, Minsu Cho, Bohyung Han |
ACCV (4) | 6 |
| 2018 | Progressive Attention Networks for Visual Attribute Prediction
Hongsuck Seo, Zhe Lin 0001, Scott Cohen, Xiaohui Shen, Bohyung Han |
BMVC | 5 |
| 2018 | Weakly Supervised Action Localization by Sparse Temporal Pooling NetworkabstractWe propose a weakly supervised temporal action localization algorithm on untrimmed videos using convolutional neural networks. Our algorithm learns from video-level class labels and predicts temporal intervals of human actions with no requirement of temporal localization annotations. We design our network to identify a sparse subset of key segments associated with target actions in a video using an attention module and fuse the key segments through adaptive temporal pooling. Our loss function is comprised of two terms that minimize the video-level action classification error and enforce the sparsity of the segment selection. At inference time, we extract and score temporal proposals using temporal class activations and class-agnostic attentions to estimate the time intervals that correspond to target actions. The proposed algorithm attains state-of-the-art results on the THUMOS14 dataset and outstanding performance on ActivityNet1.3 even with its weak supervision. Gautam Prasad, Bohyung Han |
CVPR | 4 |
| 2018 | Real-Time MDNet
Ilchae Jung, Jeany Son, Mooyeol Baek, Bohyung Han |
ECCV (4) | 4 |
| 2018 | Attentive Semantic Alignment with Offset-Aware Correlation Kernels
Hongsuck Seo, Jongmin Lee 0005, Deunsol Jung, Bohyung Han, Minsu Cho |
ECCV (4) | 4 |
| 2018 | CPlaNet: Enhancing Image Geolocalization by Combinatorial Partitioning of Maps
Hongsuck Seo, Tobias Weyand, Jack Sim, Bohyung Han |
ECCV (10) | 4 |
| 2018 | Learning to Specialize with Knowledge Distillation for Visual Question AnsweringabstractVisual Question Answering (VQA) is a notoriously challenging problem because it involves various heterogeneous tasks defined by questions within a unified framework. Learning specialized models for individual types of tasks is intuitively attracting but surprisingly difficult; it is not straightforward to outperform naive independent ensemble approach. We present a principled algorithm to learn specialized models with knowledge distillation under a multiple choice learning (MCL) framework, where training examples are assigned dynamically to a subset of models for updating network parameters. The assigned and non-assigned models are learned to predict ground-truth answers and imitate their own base models before specialization, respectively. Our approach alleviates the limitation of data deficiency in existing MCL frameworks, and allows each model to learn its own specialized expertise without forgetting general knowledge. The proposed framework is model-agnostic and applicable to any tasks other than VQA, e.g., image classification with a large number of labels but few per-class examples, which is known to be difficult under existing MCL schemes. Our experimental results indeed demonstrate that our method outperforms other baselines for VQA and image classification. Jonghwan Mun, Kimin Lee, Jinwoo Shin, Bohyung Han |
NeurIPS | 4 |
| 2017 | Weakly Supervised Semantic Segmentation Using Superpixel Pooling NetworkabstractWe propose a weakly supervised semantic segmentation algorithm based on deep neural networks, which relies on image-level class labels only. The proposed algorithm alternates between generating segmentation annotations and learning a semantic segmentation network using the generated annotations. A key determinant of success in this framework is the capability to construct reliable initial annotations given image-level labels only. To this end, we propose Superpixel Pooling Network (SPN), which utilizes superpixel segmentation of input image as a pooling layout to reflect low-level image structure for learning and inferring semantic segmentation. The initial annotations generated by SPN are then used to learn another neural network that estimates pixel-wise semantic labels. The architecture of the segmentation network decouples semantic segmentation task into classification and segmentation so that the network learns class-agnostic shape prior from the noisy annotations. It turns out that both networks are critical to improve semantic segmentation accuracy. The proposed algorithm achieves outstanding performance in weakly supervised semantic segmentation task compared to existing techniques on the challenging PASCAL VOC 2012 segmentation benchmark. Suha Kwak, Seunghoon Hong, Bohyung Han |
AAAI | 3 |
| 2017 | Text-Guided Attention Model for Image CaptioningabstractVisual attention plays an important role to understand images and demonstrates its effectiveness in generating natural language descriptions of images. On the other hand, recent studies show that language associated with an image can steer visual attention in the scene during our cognitive process. Inspired by this, we introduce a text-guided attention model for image captioning, which learns to drive visual attention using associated captions. For this model, we propose an exemplar-based learning approach that retrieves from training data associated captions with each image, and use them to learn attention on visual features. Our attention model enables to describe a detailed state of scenes by distinguishing small or confusable objects effectively. We validate our model on MS-COCO Captioning benchmark and achieve the state-of-the-art performance in standard metrics. Jonghwan Mun, Minsu Cho, Bohyung Han |
AAAI | 3 |
| 2017 | BranchOut: Regularization for Online Ensemble Tracking with Convolutional Neural NetworksabstractWe propose an extremely simple but effective regularization technique of convolutional neural networks (CNNs), referred to as BranchOut, for online ensemble tracking. Our algorithm employs a CNN for target representation, which has a common convolutional layers but has multiple branches of fully connected layers. For better regularization, a subset of branches in the CNN are selected randomly for online learning whenever target appearance models need to be updated. Each branch may have a different number of layers to maintain variable abstraction levels of target appearances. BranchOut with multi-level target representation allows us to learn robust target appearance models with diversity and handle various challenges in visual tracking problem effectively. The proposed algorithm is evaluated in standard tracking benchmarks and shows the state-of-the-art performance even without additional pretraining on external tracking sequences. Bohyung Han, Jack Sim, Hartwig Adam |
CVPR | 1 |
| 2017 | Weakly Supervised Semantic Segmentation Using Web-Crawled VideosabstractWe propose a novel algorithm for weakly supervised semantic segmentation based on image-level class labels only. In weakly supervised setting, it is commonly observed that trained model overly focuses on discriminative parts rather than the entire object area. Our goal is to overcome this limitation with no additional human intervention by retrieving videos relevant to target class labels from web repository, and generating segmentation labels from the retrieved videos to simulate strong supervision for semantic segmentation. During this process, we take advantage of image classification with discriminative localization technique to reject false alarms in retrieved videos and identify relevant spatio-temporal volumes within retrieved videos. Although the entire procedure does not require any additional supervision, the segmentation annotations obtained from videos are sufficiently strong to learn a model for semantic segmentation. The proposed algorithm substantially outperforms existing methods based on the same level of supervision and is even as competitive as the approaches relying on extra annotations. Seunghoon Hong, Donghun Yeo, Suha Kwak, Honglak Lee, Bohyung Han |
CVPR | 5 |
| 2017 | Multi-object Tracking with Quadruplet Convolutional Neural NetworksabstractWe propose Quadruplet Convolutional Neural Networks (Quad-CNN) for multi-object tracking, which learn to associate object detections across frames using quadruplet losses. The proposed networks consider target appearances together with their temporal adjacencies for data association. Unlike conventional ranking losses, the quadruplet loss enforces an additional constraint that makes temporally adjacent detections more closely located than the ones with large temporal gaps. We also employ a multi-task loss to jointly learn object association and bounding box regression for better localization. The whole network is trained end-to-end. For tracking, the target association is performed by minimax label propagation using the metric learned from the proposed network. We evaluate performance of our multi-object tracking algorithm on public MOT Challenge datasets, and achieve outstanding results. Jeany Son, Mooyeol Baek, Minsu Cho, Bohyung Han |
CVPR | 4 |
| 2017 | Superpixel-Based Tracking-by-Segmentation Using Markov ChainsabstractWe propose a simple but effective tracking-by-segmentation algorithm using Absorbing Markov Chain (AMC) on superpixel segmentation, where target state is estimated by a combination of bottom-up and top-down approaches, and target segmentation is propagated to subsequent frames in a recursive manner. Our algorithm constructs a graph for AMC using the superpixels identified in two consecutive frames, where background superpixels in the previous frame correspond to absorbing vertices while all other superpixels create transient ones. The weight of each edge depends on the similarity of scores in the end superpixels, which are learned by support vector regression. Once graph construction is completed, target segmentation is estimated using the absorption time of each superpixel. The proposed tracking algorithm achieves substantially improved performance compared to the state-of-the-art segmentation-based tracking techniques in multiple challenging datasets. Donghun Yeo, Jeany Son, Bohyung Han, Joon Hee Han |
CVPR | 3 |
| 2017 | MarioQA: Answering Questions by Watching Gameplay VideosabstractWe present a framework to analyze various aspects of models for video question answering (VideoQA) using customizable synthetic datasets, which are constructed automatically from gameplay videos. Our work is motivated by the fact that existing models are often tested only on datasets that require excessively high-level reasoning or mostly contain instances accessible through single frame inferences. Hence, it is difficult to measure capacity and flexibility of trained models, and existing techniques often rely on adhoc implementations of deep neural networks without clear insight into datasets and models. We are particularly interested in understanding temporal relationships between video events to solve VideoQA problems; this is because reasoning temporal dependency is one of the most distinct components in videos from images. To address this objective, we automatically generate a customized synthetic VideoQA dataset using Super Mario Bros. gameplay videos so that it contains events with different levels of reasoning complexity. Using the dataset, we show that properly constructed datasets with events in various complexity levels are critical to learn effective models and improve overall performance. Jonghwan Mun, Hongsuck Seo, Ilchae Jung, Bohyung Han |
ICCV | 4 |
| 2017 | Large-Scale Image Retrieval with Attentive Deep Local FeaturesabstractWe propose an attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELE (DEep Local Feature). The new feature is based on convolutional neural networks, which are trained only with image-level annotations on a landmark image dataset. To identify semantically useful local features for image retrieval, we also propose an attention mechanism for key point selection, which shares most network layers with the descriptor. This frame-work can be used for image retrieval as a drop-in replacement for other keypoint detectors and descriptors, enabling more accurate feature matching and geometric verification. Our system produces reliable confidence scores to reject false positives–in particular, it is robust against queries that have no correct match in the database. To evaluate the proposed descriptor, we introduce a new large-scale dataset, referred to as Google-Landmarks dataset, which involves challenges in both database and query such as background clutter, partial occlusion, multiple landmarks, objects in variable scales, etc. We show that DELE outperforms the state-of-the-art global and local descriptors in the large-scale setting by significant margins. Hyeonwoo Noh, André Araújo 0001, Jack Sim, Tobias Weyand, Bohyung Han |
ICCV | 5 |
| 2017 | Regularizing Deep Neural Networks by Noise: Its Interpretation and OptimizationabstractOverfitting is one of the most critical challenges in deep neural networks, and there are various types of regularization methods to improve generalization performance. Injecting noises to hidden units during training, e.g., dropout, is known as a successful regularizer, but it is still not clear enough why such training techniques work well in practice and how we can maximize their benefit in the presence of two conflicting objectives---optimizing to true data distribution and preventing overfitting by regularization. This paper addresses the above issues by 1) interpreting that the conventional training methods with regularization by noise injection optimize the lower bound of the true objective and 2) proposing a technique to achieve a tighter lower bound using multiple noise samples per training example in a stochastic gradient descent iteration. We demonstrate the effectiveness of our idea in several computer vision applications. Hyeonwoo Noh, Tackgeun You, Jonghwan Mun, Bohyung Han |
NIPS | 4 |
| 2017 | Visual Reference Resolution using Attention Memory for Visual DialogabstractVisual dialog is a task of answering a series of inter-dependent questions given an input image, and often requires to resolve visual references among the questions. This problem is different from visual question answering (VQA), which relies on spatial attention ({\em a.k.a. visual grounding}) estimated from an image and question pair. We propose a novel attention mechanism that exploits visual attentions in the past to resolve the current reference in the visual dialog scenario. The proposed model is equipped with an associative attention memory storing a sequence of previous (attention, key) pairs. From this memory, the model retrieves previous attention, taking into account recency, that is most relevant for the current question, in order to resolve potentially ambiguous reference(s). The model then merges the retrieved attention with the tentative one to obtain the final attention for the current question; specifically, we use dynamic parameter prediction to combine the two attentions conditioned on the question. Through extensive experiments on a new synthetic visual dialog dataset, we show that our model significantly outperforms the state-of-the-art (by ~16 % points) in the situation where the visual reference resolution plays an important role. Moreover, the proposed model presents superior performance (~2 % points improvement) in the Visual Dialog dataset, despite having significantly fewer parameters than the baselines. Hongsuck Seo, Andreas M. Lehrmann, Bohyung Han, Leonid Sigal |
NIPS | 3 |
| 2017 | Personalized Image Aesthetic Quality Assessment by Joint Regression and RankingabstractWe propose an image aesthetic quality assessment algorithm, which considers personal taste in addition to generally perceived preference. This problem is formulated by a combination of two different learning frameworks based on support vector machines-Support Vector Regression (SVR) and Ranking SVM (R-SVM), where SVR learns a general model based on public datasets and R-SVM adjusts the model to accommodate personal preference obtained from user interactions. The combined framework, called R-SVR, is represented by a single objective function, which is optimized jointly to learn a model for personalized image aesthetic quality assessment. For the optimization, we use only a small subset of public dataset identified by k-nearest neighbor search instead of using all available training data. This strategy is useful in practice because it reduces training time significantly and alleviates data imbalance problem between regression and ranking. The proposed algorithm is tested through simulation and user study, and we present that our interactive learning algorithm by R-SVR is effective to increase user's satisfaction and improve prediction performance. Kayoung Park, Seunghoon Hong, Mooyeol Baek, Bohyung Han |
WACV | 4 |
| 2017 | Rank-based voting with inclusion relationship for accurate image search
Jaehyeong Cho, Jae-Pil Heo, Bohyung Han, Sung-Eui Yoon |
Vis. Comput. | 4 |
| 2016 | Unsupervised Co-Activity Detection from Multiple Videos Using Absorbing Markov ChainabstractWe propose a simple but effective unsupervised learning algorithm to detect a common activity (co-activity) from a set of videos, which is formulated using absorbing Markov chain in a principled way. In our algorithm, a complete multipartite graph is first constructed, where vertices correspond to subsequences extracted from videos using a temporal sliding window and edges connect between the vertices originated from different videos; the weight of an edge is proportional to the similarity between the features of two end vertices. Then, we extend the graph structure by adding edges between temporally overlapped subsequences in a video to handle variable-length co-activities using temporal locality, and create an absorbing vertex connected from all other nodes. The proposed algorithm identifies a subset of subsequences as co-activity by estimating absorption time in the constructed graph efficiently. The great advantage of our algorithm lies in the properties that it can handle more than two videos naturally and identify multiple instances of a co-activity with variable lengths in a video. Our algorithm is evaluated intensively in a challenging dataset and illustrates outstanding performance quantitatively and qualitatively. Donghun Yeo, Bohyung Han, Joon Hee Han |
AAAI | 2 |
| 2016 | Learning Transferrable Knowledge for Semantic Segmentation with Deep Convolutional Neural NetworkabstractWe propose a novel weakly-supervised semantic segmentation algorithm based on Deep Convolutional Neural Network (DCNN). Contrary to existing weakly-supervised approaches, our algorithm exploits auxiliary segmentation annotations available for different categories to guide segmentations on images with only image-level class labels. To make segmentation knowledge transferrable across categories, we design a decoupled encoder-decoder architecture with attention model. In this architecture, the model generates spatial highlights of each category presented in images using an attention model, and subsequently performs binary segmentation for each highlighted region using decoder. Combining attention model, the decoder trained with segmentation annotations in different categories boosts accuracy of weakly-supervised semantic segmentation. The proposed algorithm demonstrates substantially improved performance compared to the state-of-theart weakly-supervised techniques in PASCAL VOC 2012 dataset when our model is trained with the annotations in 60 exclusive categories in Microsoft COCO dataset. Seunghoon Hong, Junhyuk Oh, Honglak Lee, Bohyung Han |
CVPR | 4 |
| 2016 | Learning to Select Pre-Trained Deep Representations with Bayesian Evidence FrameworkabstractWe propose a Bayesian evidence framework to facilitate transfer learning from pre-trained deep convolutional neural networks (CNNs). Our framework is formulated on top of a least squares SVM (LS-SVM) classifier, which is simple and fast in both training and testing, and achieves competitive performance in practice. The regularization parameters in LS-SVM is estimated automatically without grid search and cross-validation by maximizing evidence, which is a useful measure to select the best performing CNN out of multiple candidates for transfer learning, the evidence is optimized efficiently by employing Aitken's delta-squared process, which accelerates convergence of fixed point update. The proposed Bayesian evidence framework also provides a good solution to identify the best ensemble of heterogeneous CNNs through a greedy algorithm. Our Bayesian evidence framework for transfer learning is tested on 12 visual recognition datasets and illustrates the state-of-the-art performance consistently in terms of prediction accuracy and modeling efficiency. Yong-Deok Kim, Taewoong Jang, Bohyung Han, Seungjin Choi 0001 |
CVPR | 3 |
| 2016 | Learning Multi-domain Convolutional Neural Networks for Visual TrackingabstractWe propose a novel visual tracking algorithm based on the representations from a discriminatively trained Convolutional Neural Network (CNN). Our algorithm pretrains a CNN using a large set of videos with tracking ground-truths to obtain a generic target representation. Our network is composed of shared layers and multiple branches of domain-specific layers, where domains correspond to individual training sequences and each branch is responsible for binary classification to identify target in each domain. We train each domain in the network iteratively to obtain generic target representations in the shared layers. When tracking a target in a new sequence, we construct a new network by combining the shared layers in the pretrained CNN with a new binary classification layer, which is updated online. Online tracking is performed by evaluating the candidate windows randomly sampled around the previous target state. The proposed algorithm illustrates outstanding performance in existing tracking benchmarks. Hyeonseob Nam, Bohyung Han |
CVPR | 2 |
| 2016 | Image Question Answering Using Convolutional Neural Network with Dynamic Parameter PredictionabstractWe tackle image question answering (ImageQA) problem by learning a convolutional neural network (CNN) with a dynamic parameter layer whose weights are determined adaptively based on questions. For the adaptive parameter prediction, we employ a separate parameter prediction network, which consists of gated recurrent unit (GRU) taking a question as its input and a fully-connected layer generating a set of candidate weights as its output. However, it is challenging to construct a parameter prediction network for a large number of parameters in the fully-connected dynamic parameter layer of the CNN. We reduce the complexity of this problem by incorporating a hashing technique, where the candidate weights given by the parameter prediction network are selected using a predefined hash function to determine individual weights in the dynamic parameter layer. The proposed network-joint network with the CNN for ImageQA and the parameter prediction network-is trained end-to-end through back-propagation, where its weights are initialized using a pre-trained CNN and GRU. The proposed algorithm illustrates the state-of-the-art performance on all available public ImageQA benchmarks. Hyeonwoo Noh, Hongsuck Seo, Bohyung Han |
CVPR | 3 |
| 2016 | Interactive motion effects design for a moving object in 4D filmsabstractThis paper presents an algorithm that allows for the rapid design of motion effects for 4D films. Our algorithm is based on a viewer-centered rendering strategy that matches chair motion to the movement of a viewer's visual attention. Object tracking algorithm is used to estimate the movement of visual attention under the assumption that visual attention follows an object of interest. We performed several experiments to find optimal parameters for implementation, such as the required accuracy of object tracking. Our algorithm enables motion effects design to be at least 10 times faster than the current practice of manual authoring. We also assessed the subjective quality of the motion effects generated by our algorithm, and results indicated that our algorithm can provide perceptually plausible motion effects. Jaebong Lee, Bohyung Han, Seungmoon Choi |
VRST | 2 |
| 2016 | Joint Image Clustering and Labeling by Matrix FactorizationabstractWe propose a novel algorithm to cluster and annotate a set of input images jointly, where the images are clustered into several discriminative groups and each group is identified with representative labels automatically. For these purposes, each input image is first represented by a distribution of candidate labels based on its similarity to images in a labeled reference image database. A set of these label-based representations are then refined collectively through a non-negative matrix factorization with sparsity and orthogonality constraints; the refined representations are employed to cluster and annotate the input images jointly. The proposed approach demonstrates performance improvements in image clustering over existing techniques, and illustrates competitive image labeling accuracy in both quantitative and qualitative evaluation. In addition, we extend our joint clustering and labeling framework to solving the weakly-supervised image classification problem and obtain promising results. Seunghoon Hong, Jan Feyereisl, Bohyung Han, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Motion Effects Synthesis for 4D Filmsabstract4D film is an immersive entertainment system that presents various physical effects with a film in order to enhance viewers' experiences. Despite the recent emergence of 4D theaters, production of 4D effects relies on manual authoring. In this paper, we present algorithms that synthesize three classes of motion effects from the audiovisual content of a film. The first class of motion effects is those responding to fast camera motion to enhance the immersiveness of point-of-view shots, delivering fast and dynamic vestibular feedback. The second class moves viewers as closely as possible to the trajectory of slowly moving camera. Such motion provides an illusional effect of observing the scene from a distance while moving slowly within the scene. For these two classes, our algorithms compute the relative camera motion and then map it to a motion command to the 4D chair using appropriate motion mapping algorithms. The last class is for special effects, such as explosions, and our algorithm uses sound for the synthesis of impulses and vibrations. We assessed the subjective quality of our algorithms by user experiments, and results indicated that our algorithms can provide compelling motion effects. Jaebong Lee, Bohyung Han, Seungmoon Choi |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Learning Deconvolution Network for Semantic SegmentationabstractWe propose a novel semantic segmentation algorithm by learning a deep deconvolution network. We learn the network on top of the convolutional layers adopted from VGG 16-layer net. The deconvolution network is composed of deconvolution and unpooling layers, which identify pixelwise class labels and predict segmentation masks. We apply the trained network to each proposal in an input image, and construct the final semantic segmentation map by combining the results from all proposals in a simple manner. The proposed algorithm mitigates the limitations of the existing methods based on fully convolutional networks by integrating deep deconvolution network and proposal-wise prediction, our segmentation method typically identifies detailed structures and handles objects in multiple scales naturally. Our network demonstrates outstanding performance in PASCAL VOC 2012 dataset, and we achieve the best accuracy (72.5%) among the methods trained without using Microsoft COCO dataset through ensemble with the fully convolutional network. Hyeonwoo Noh, Seunghoon Hong, Bohyung Han |
ICCV | 3 |
| 2015 | Tracking-by-Segmentation with Online Gradient Boosting Decision TreeabstractWe propose an online tracking algorithm that adaptively models target appearances based on an online gradient boosting decision tree. Our algorithm is particularly useful for non-rigid and/or articulated objects since it handles various deformations of the target effectively by integrating a classifier operating on individual patches and provides segmentation masks of the target as final results. The posterior of the target state is propagated over time by particle filtering, where the likelihood is computed based mainly on patch-level confidence map associated with a latent target state corresponding to each sample. Once tracking is completed in each frame, our gradient boosting decision tree is updated to adapt new data in a recursive manner. For effective evaluation of segmentation-based tracking algorithms, we construct a new ground-truth that contains pixel-level annotation of segmentation mask. We evaluate the performance of our tracking algorithm based on the measures for segmentation masks, where our algorithm illustrates superior accuracy compared to the state-of-the-art segmentation-based tracking methods. Jeany Son, Ilchae Jung, Kayoung Park, Bohyung Han |
ICCV | 4 |
| 2015 | Online Tracking by Learning Discriminative Saliency Map with Convolutional Neural NetworkabstractWe propose an online visual tracking algorithm by learning discriminative saliency map using Convolutional Neural Network (CNN). Given a CNN pre-trained on a large-scale image repository in offline, our algorithm takes outputs from hidden layers of the network as feature descriptors since they show excellent representation performance in various general visual recognition problems. The features are used to learn discriminative target appearance models using an online Support Vector Machine (SVM). In addition, we construct target-specific saliency map by back-projecting CNN features with guidance of the SVM, and obtain the final tracking result in each frame based on the appearance model generatively constructed with the saliency map. Since the saliency map reveals spatial configuration of target effectively, it improves target localization accuracy and enables us to achieve pixel-level target segmentation. We verify the effectiveness of our tracking algorithm through extensive experiment on a challenging benchmark, where our method illustrates outstanding performance compared to the state-of-the-art tracking algorithms. Seunghoon Hong, Tackgeun You, Suha Kwak, Bohyung Han |
ICML | 4 |
| 2015 | Decoupled Deep Neural Network for Semi-supervised Semantic SegmentationabstractWe propose a novel deep neural network architecture for semi-supervised semantic segmentation using heterogeneous annotations. Contrary to existing approaches posing semantic segmentation as region-based classification, our algorithm decouples classification and segmentation, and learns a separate network for each task. In this architecture, labels associated with an image are identified by classification network, and binary segmentation is subsequently performed for each identified label by segmentation network. The decoupled architecture enables us to learn classification and segmentation networks separately based on the training data with image-level and pixel-wise class labels, respectively. It facilitates to reduce search space for segmentation effectively by exploiting class-specific activation maps obtained from bridging layers. Our algorithm shows outstanding performance compared to other semi-supervised approaches even with much less training images with strong annotations in PASCAL VOC dataset. Seunghoon Hong, Hyeonwoo Noh, Bohyung Han |
NIPS | 3 |
| 2015 | Qualitative Tracking Performance Evaluation without Ground-TruthabstractWe present a qualitative tracking performance evaluation algorithm without ground-truth, where several representative frames are automatically selected and visualized in a principled way. Although tracking algorithms are typically evaluated by quantitative scores based on predefined measures, qualitative evaluation is also useful especially when the ground-truth of a target state is unavailable or unreliable. However, there is no prior study on how to present frames for better qualitative evaluation of tracking algorithms. Motivated by this fact, we propose an unbiased frame selection technique, where salient and unique features in tracking results are captured effectively. Our method identifies a set of representative frames by 1) analyzing the sequence structure using manifold learning, and 2) selecting frames by formulating the task as a facility location problem. By presenting the manifold and the selected frames, one can understand the sequence structure as well as the characteristics of tracking results. The effectiveness of our method is illustrated with single and multiple tracking results for sequences without ground-truth. Bohyung Han, Jihun Hamm |
WACV | 1 |
| 2015 | Occlusion detection using horizontally segmented windows for vehicle tracking
Ahra Jo, Gil-Jin Jang, Bohyung Han |
Multim. Tools Appl. | 3 |
| 2014 | Visual Tracking by Sampling Tree-Structured Graphical Models
Seunghoon Hong, Bohyung Han |
ECCV (1) | 2 |
| 2014 | Generalized Background Subtraction Using Superpixels with Label Integrated Motion Estimation
Jongwoo Lim, Bohyung Han |
ECCV (5) | 2 |
| 2014 | Online Graph-Based Tracking
Hyeonseob Nam, Seunghoon Hong, Bohyung Han |
ECCV (5) | 3 |
| 2014 | Object Localization based on Structural SVM using Privileged Information
Jan Feyereisl, Suha Kwak, Jeany Son, Bohyung Han |
NIPS | 4 |
| 2014 | Macrofeature layout selection for pedestrian localization and its acceleration using GPU
Woonhyun Nam, Bohyung Han, Joon Hee Han |
Comput. Vis. Image Underst. | 2 |
| 2014 | On-Line Video Event Detection by Constraint FlowabstractWe present a novel approach in describing and detecting the composite video events based on scenarios, which constrain the configurations of target events by temporal-logical structures of primitive events. We propose a new scenario description method to represent composite events more fluently and efficiently, and discuss an on-line event detection algorithm based on a combinatorial optimization. For this purpose, constraint flow-a dynamic configuration of scenario constraints-is first generated automatically by our scenario parsing algorithm. Then, composite event detection is formulated by a constrained discrete optimization problem, whose objective is to find the best video interpretation with respect to the constraint flow. Although the search space for the optimization problem is prohibitively large, our on-line event detection algorithm based on constraint flow using dynamic programming reduces the search space dramatically, handles preprocessing errors effectively, and guarantees a globally optimal solution. Experimental results on natural videos demonstrate the effectiveness of our algorithm. Suha Kwak, Bohyung Han, Joon Hee Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Multi-agent Event Detection: Localization and Role AssignmentabstractWe present a joint estimation technique of event localization and role assignment when the target video event is described by a scenario. Specifically, to detect multi-agent events from video, our algorithm identifies agents involved in an event and assigns roles to the participating agents. Instead of iterating through all possible agent-role combinations, we formulate the joint optimization problem as two efficient sub problems-quadratic programming for role assignment followed by linear programming for event localization. Additionally, we reduce the computational complexity significantly by applying role-specific event detectors to each agent independently. We test the performance of our algorithm in natural videos, which contain multiple target events and nonparticipating agents. Suha Kwak, Bohyung Han, Joon Hee Han |
CVPR | 2 |
| 2013 | Orderless Tracking through Model-Averaged Posterior EstimationabstractWe propose a novel offline tracking algorithm based on model-averaged posterior estimation through patch matching across frames. Contrary to existing online and offline tracking methods, our algorithm is not based on temporally-ordered estimates of target state but attempts to select easy-to-track frames first out of the remaining ones without exploiting temporal coherency of target. The posterior of the selected frame is estimated by propagating densities from the already tracked frames in a recursive manner. The density propagation across frames is implemented by an efficient patch matching technique, which is useful for our algorithm since it does not require motion smoothness assumption. Also, we present a hierarchical approach, where a small set of key frames are tracked first and non-key frames are handled by local key frames. Our tracking algorithm is conceptually well-suited for the sequences with abrupt motion, shot changes, and occlusion. We compare our tracking algorithm with existing techniques in real videos with such challenges and illustrate its superior performance qualitatively and quantitatively. Seunghoon Hong, Suha Kwak, Bohyung Han |
ICCV | 3 |
| 2013 | Joint Segmentation and Pose Tracking of Human in Natural VideosabstractWe propose an on-line algorithm to extract a human by foreground/background segmentation and estimate pose of the human from the videos captured by moving cameras. We claim that a virtuous cycle can be created by appropriate interactions between the two modules to solve individual problems. This joint estimation problem is divided into two sub problems, foreground/background segmentation and pose tracking, which alternate iteratively for optimization, segmentation step generates foreground mask for human pose tracking, and human pose tracking step provides fore-ground response map for segmentation. The final solution is obtained when the iterative procedure converges. We evaluate our algorithm quantitatively and qualitatively in real videos involving various challenges, and present its outstanding performance compared to the state-of-the-art techniques for segmentation and pose estimation. Taegyu Lim, Seunghoon Hong, Bohyung Han, Joon Hee Han |
ICCV | 3 |
| 2012 | Online Multi-target Tracking by Large Margin Structured Learning
Suna Kim, Suha Kwak, Jan Feyereisl, Bohyung Han |
ACCV (3) | 4 |
| 2012 | FaceReview: Supporting Interactive Exploration of Linked Heterogeneous Datasets for Unilateral Cleft Lip and Palate
Jinwook Seo, Boeun Kim, Bongshin Lee, Bo Hyoung Kim, Bohyung Han, Nina Anderson, Richard Bruun, Stephen Shusterman |
AMIA | 6 |
| 2012 | A fast nearest neighbor search algorithm by nonlinear embeddingabstractWe propose an efficient algorithm to find the exact nearest neighbor based on the Euclidean distance for large-scale computer vision problems. We embed data points nonlinearly onto a low-dimensional space by simple computations and prove that the distance between two points in the embedded space is bounded by the distance in the original space. Instead of computing the distances in the high-dimensional original space to find the nearest neighbor, a lot of candidates are to be rejected based on the distances in the low-dimensional embedded space; due to this property, our algorithm is well-suited for high-dimensional and large-scale problems. We also show that our algorithm is improved further by partitioning input vectors recursively. Contrary to most of existing fast nearest neighbor search algorithms, our technique reports the exact nearest neighbor - not an approximate one - and requires a very simple preprocessing with no sophisticated data structures. We provide the theoretical analysis of our algorithm and evaluate its performance in synthetic and real data. Yoonho Hwang, Bohyung Han, Hee-Kap Ahn |
CVPR | 2 |
| 2012 | Online Video Segmentation by Bayesian Split-Merge Clustering
Juho Lee 0001, Suha Kwak, Bohyung Han, Seungjin Choi 0001 |
ECCV (4) | 3 |
| 2012 | Seam carving with forward gradient difference mapsabstractWe propose a new energy function for seam carving based on forward gradient differences to preserve regular structures in images. The energy function measures the curvature inconsistency between the pixels that become adjacent after seam removal, and involves the difference of gradient orientation and magnitude of the pixels. Our objective is to minimize the differences induced by the removed seam, and the optimization is performed by dynamic programming based on multiple cumulative energy maps, each of which corresponds to the seam pattern associated with a pixel. The proposed technique preserves straight lines and regular shapes better than the original and improved seam carving, and can be easily combined with other types of energy functions within the seam carving framework. We evaluated the performance of our algorithm by comparing with the original and improved seam carving algorithms using public data. Hyeonwoo Noh, Bohyung Han |
ACM Multimedia | 2 |
| 2012 | Density-Based Multifeature Background Subtraction with Support Vector MachineabstractBackground modeling and subtraction is a natural technique for object detection in videos captured by a static camera, and also a critical preprocessing step in various high-level computer vision applications. However, there have not been many studies concerning useful features and binary segmentation algorithms for this problem. We propose a pixelwise background modeling and subtraction technique using multiple features, where generative and discriminative techniques are combined for classification. In our algorithm, color, gradient, and Haar-like features are integrated to handle spatio-temporal variations for each pixel. A pixelwise generative background model is obtained for each feature efficiently and effectively by Kernel Density Approximation (KDA). Background subtraction is performed in a discriminative manner using a Support Vector Machine (SVM) over background likelihood vectors for a set of features. The proposed algorithm is robust to shadow, illumination changes, spatial variations of background. We compare the performance of the algorithm with other density-based methods using several different feature combinations and modeling techniques, both quantitatively and qualitatively. Bohyung Han, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Modeling and segmentation of floating foreground and background in videos
Taegyu Lim, Bohyung Han, Joon Hee Han |
Pattern Recognit. | 2 |
| 2011 | Dynamic Resource Allocation by Ranking SVM for Particle Filter Tracking
Chang-Kyu Song, Jeany Son, Suha Kwak, Bohyung Han |
BMVC | 4 |
| 2011 | Scenario-based video event recognition by constraint flowabstractWe present a novel approach to representing and recognizing composite video events. A composite event is specified by a scenario, which is based on primitive events and their temporal-logical relations, to constrain the arrangements of the primitive events in the composite event. We propose a new scenario description method to represent composite events fluently and efficiently. A composite event is recognized by a constrained optimization algorithm whose constraints are defined by the scenario. The dynamic configuration of the scenario constraints is represented with constraint flow, which is generated from scenario automatically by our scenario parsing algorithm. The constraint flow reduces the search space dramatically, alleviates the effect of preprocessing errors, and guarantees the globally optimal solution for recognition. We validate our method to describe scenario and construct constraint flow for real videos and illustrate the effectiveness of our composite event recognition algorithm for natural video events. Suha Kwak, Bohyung Han, Joon Hee Han |
CVPR | 2 |
| 2011 | Generalized background subtraction based on hybrid inference by belief propagation and Bayesian filteringabstractWe propose a novel background subtraction algorithm for the videos captured by a moving camera. In our technique, foreground and background appearance models in each frame are constructed and propagated sequentially by Bayesian filtering. We estimate the posterior of appearance, which is computed by the product of the image likelihood in the current frame and the prior appearance propagated from the previous frame. The motion, which transfers the previous appearance models to the current frame, is estimated by nonparametric belief propagation; the initial motion field is obtained by optical flow and noisy and incomplete motions are corrected effectively through the inference procedure. Our framework is represented by a graphical model, where the sequential inference of motion and appearance is performed by the combination of belief propagation and Bayesian filtering. We compare our algorithm with the existing state-of-the-art technique and evaluate its performance quantitatively and qualitatively in several challenging videos. Suha Kwak, Taegyu Lim, Woonhyun Nam, Bohyung Han, Joon Hee Han |
ICCV | 4 |
| 2011 | Learning occlusion with likelihoods for visual trackingabstractWe propose a novel algorithm to detect occlusion for visual tracking through learning with observation likelihoods. In our technique, target is divided into regular grid cells and the state of occlusion is determined for each cell using a classifier. Each cell in the target is associated with many small patches, and the patch likelihoods observed during tracking construct a feature vector, which is used for classification. Since the occlusion is learned with patch likelihoods instead of patches themselves, the classifier is universally applicable to any videos or objects for occlusion reasoning. Our occlusion detection algorithm has decent performance in accuracy, which is sufficient to improve tracking performance significantly. The proposed algorithm can be combined with many generic tracking methods, and we adopt L1 minimization tracker to test the performance of our framework. The advantage of our algorithm is supported by quantitative and qualitative evaluation, and successful tracking and occlusion reasoning results are illustrated in many challenging video sequences. Suha Kwak, Woonhyun Nam, Bohyung Han, Joon Hee Han |
ICCV | 3 |
| 2011 | Personalized video summarization with human in the loopabstractIn automatic video summarization, visual summary is constructed typically based on the analysis of low-level features with little consideration of video semantics. However, the contextual and semantic information of a video is marginally related to low-level features in practice although they are useful to compute visual similarity between frames. Therefore, we propose a novel video summarization technique, where the semantically important information is extracted from a set of keyframes given by human and the summary of a video is constructed based on the automatic temporal segmentation using the analysis of inter-frame similarity to the keyframes. Toward this goal, we model a video sequence with a dissimilarity matrix based on bidirectional similarity measure between every pair of frames, and subsequently characterize the structure of the video by a nonlinear manifold embedding. Then, we formulate video summarization as a variant of the 0-1 knapsack problem, which is solved by dynamic programming efficiently. The effectiveness of our algorithm is illustrated quantitatively and qualitatively using realistic videos collected from YouTube. Bohyung Han, Jihun Hamm, Jack Sim |
WACV | 1 |
| 2011 | Multi-Camera Tracking with Adaptive Resource Allocation
Bohyung Han, Seong-Wook Joo, Larry Davis 0001 |
Int. J. Comput. Vis. | 1 |
| 2010 | Efficient extraction of human motion volumes by trackingabstractWe present an automatic and efficient method to extract spatio-temporal human volumes from video, which combines top-down model-based and bottom-up appearance-based approaches. From the top-down perspective, our algorithm applies shape priors probabilistically to candidate image regions obtained by pedestrian detection, and provides accurate estimates of the human body areas which serve as important constraints for bottom-up processing. Temporal propagation of the identified region is performed with bottom-up cues in an efficient level-set framework, which takes advantage of the sparse top-down information that is available. Our formulation also optimizes the extracted human volume across frames through belief propagation and provides temporally coherent human regions. We demonstrate the ability of our method to extract human body regions efficiently and automatically from a large, challenging dataset collected from YouTube. Juan Carlos Niebles, Bohyung Han, Li Fei-Fei 0001 |
CVPR | 2 |
| 2009 | Probabilistic fusion-based parameter estimation for visual tracking
Bohyung Han, Larry Davis 0001 |
Comput. Vis. Image Underst. | 1 |
| 2009 | Visual Tracking by Continuous Density Propagation in Sequential Bayesian Filtering FrameworkabstractParticle filtering is frequently used for visual tracking problems since it provides a general framework for estimating and propagating probability density functions for nonlinear and non-Gaussian dynamic systems. However, this algorithm is based on a Monte Carlo approach and the cost of sampling and measurement is a problematic issue, especially for high-dimensional problems. We describe an alternative to the classical particle filter in which the underlying density function has an analytic representation for better approximation and effective propagation. The techniques of density interpolation and density approximation are introduced to represent the likelihood and the posterior densities with Gaussian mixtures, where all relevant parameters are automatically determined. The proposed analytic approach is shown to perform more efficiently in sampling in high-dimensional space. We apply the algorithm to real-time tracking problems and demonstrate its performance on real video sequences as well as synthetic examples. Bohyung Han, Ying Zhu 0006, Dorin Comaniciu, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Extracting Moving People from Internet Videos
Juan Carlos Niebles, Bohyung Han, Andras Ferencz, Li Fei-Fei 0001 |
ECCV (4) | 2 |
| 2008 | Sequential Kernel Density Approximation and Its Application to Real-Time Visual TrackingabstractVisual features are commonly modeled with probability density functions in computer vision problems, but current methods such as a mixture of Gaussians and kernel density estimation suffer from either the lack of flexibility, by fixing or limiting the number of Gaussian components in the mixture, or large memory requirement, by maintaining a non-parametric representation of the density. These problems are aggravated in real-time computer vision applications since density functions are required to be updated as new data becomes available. We present a novel kernel density approximation technique based on the mean-shift mode finding algorithm, and describe an efficient method to sequentially propagate the density modes over time. While the proposed density representation is memory efficient, which is typical for mixture densities, it inherits the flexibility of non-parametric methods by allowing the number of components to be variable. The accuracy and compactness of the sequential kernel density approximation technique is illustrated by both simulations and experiments. Sequential kernel density approximation is applied to on-line target appearance modeling for visual tracking, and its performance is demonstrated on a variety of videos. Bohyung Han, Dorin Comaniciu, Ying Zhu 0006, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Probabilistic Fusion Tracking Using Mixture Kernel-Based Bayesian FilteringabstractEven though sensor fusion techniques based on particle filters have been applied to object tracking, their implementations have been limited to combining measurements from multiple sensors by the simple product of individual likelihoods. Therefore, the number of observations is increased as many times as the number of sensors, and the combined observation may become unreliable through blind integration of sensor observations—especially if some sensors are too noisy and non-discriminative. We describe a methodology to model interactions between multiple sensors and to estimate the current state by using a mixture of Bayesian filters—one filter for each sensor, where each filter makes a different level of contribution to estimate the combined posterior in a reliable manner. In this framework, an adaptive particle arrangement system is constructed in which each particle is allocated to only one of the sensors for observation and a different number of samples is assigned to each sensor using prior distribution and partial observations. We apply this technique to visual tracking in logical and physical sensor fusion frameworks, and demonstrate its effectiveness through tracking results. Bohyung Han, Seong-Wook Joo, Larry Davis 0001 |
ICCV | 1 |
| 2005 | Kernel-Based Bayesian Filtering for Object TrackingabstractParticle filtering provides a general framework for propagating probability density functions in nonlinear and non-Gaussian systems. However, the algorithm is based on a Monte Carlo approach and sampling is a problematic issue, especially for high dimensional problems. This paper presents a new kernel-based Bayesian filtering framework, which adopts an analytic approach to better approximate and propagate density functions. In this framework, the techniques of density interpolation and density approximation are introduced to represent the likelihood and the posterior densities by Gaussian mixtures, where all parameters such as the number of mixands, their weight, mean, and covariance are automatically determined. The proposed analytic approach is shown to perform sampling more efficiently in high dimensional space. We apply our algorithm to real-time tracking problems, and demonstrate its performance on real video sequences as well as synthetic examples. Bohyung Han, Ying Zhu 0006, Dorin Comaniciu, Larry Davis 0001 |
CVPR (1) | 1 |
| 2005 | On-Line Density-Based Appearance Modeling for Object TrackingabstractObject tracking is a challenging problem in real-time computer vision due to variations of lighting condition, pose, scale, and view-point over time. However, it is exceptionally difficult to model appearance with respect to all of those variations in advance; instead, on-line update algorithms are employed to adapt to these changes. We present a new on-line appearance modeling technique which is based on sequential density approximation. This technique provides accurate and compact representations using Gaussian mixtures, in which the number of Gaussians is automatically determined. This procedure is performed in linear time at each time step, which we prove by amortized analysis. Features for each pixel and rectangular region are modeled together by the proposed sequential density approximation algorithm, and the target model is updated in scale robustly. We show the performance of our method by simulations and tracking in natural videos Bohyung Han, Larry Davis 0001 |
ICCV | 1 |
| 2005 | Robust observations for object trackingabstractIt is a difficult task to find an observation model that will perform well for long-term visual tracking. In this paper, we propose an adaptive observation enhancement technique based on likelihood images, which are derived from multiple visual features. The most discriminative likelihood image is extracted by principal component analysis (PCA) and incrementally updated frame by frame to reduce temporal tracking error. In the particle filter framework, the feasibility of each sample is computed using this most discriminative likelihood image before the observation process. Integral image is employed for efficient computation of the feasibility of each sample. We illustrate how our enhancement technique contributes to more robust observations through demonstrations. Bohyung Han, Larry Davis 0001 |
ICIP (2) | 1 |
| 2004 | Incremental Density Approximation and Kernel-Based Bayesian Filtering for Object Tracking
Bohyung Han, Dorin Comaniciu, Ying Zhu 0006, Larry Davis 0001 |
CVPR (1) | 1 |
| 2004 | Object tracking by adaptive feature extraction
Bohyung Han, Larry Davis 0001 |
ICIP | 1 |