EDBT 2026 Demo / reviewers in the wild / expert
Bumsub Ham
dblp:03/8108
· DBLP profile ↗
80ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0002-3443-8161ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 53 · 4 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture SearchabstractTransformer architecture search (TAS) aims to automatically discover efficient vision transformers (ViTs), reducing the need for manual design. Existing TAS methods typically train an over-parameterized network (i.e., a supernet) that encompasses all candidate architectures (i.e., subnets). However, subnets partially share weights within the supernet, which leads to interference that degrades the smaller subnets severely. We have found that well-trained small subnets can serve as a good foundation for training larger ones. Motivated by this, we propose a progressive training framework, dubbed GrowTAS, that begins with training small subnets and incorporates larger ones gradually. This enables reducing the interference and stabilizing training. We also introduce GrowTAS+ that fine-tunes a subset of weights only to further enhance the performance of large subnets. Extensive experiments on ImageNet and several transfer learning benchmarks, including CIFAR-10/100, Flowers, CARS, and INAT-19, demonstrate the effectiveness of our approach over current TAS methods. Youngmin Oh 0001, Jeimin Jeon, Donghyeon Baek, Bumsub Ham |
WACV | 5 |
| 2026 | Token-based dynamic bit-width assignment for ViT quantization
Dohyung Kim 0006, Jaehyeon Moon, Junghyup Lee, Jeimin Jeon, Bumsub Ham |
Pattern Recognit. | 6 |
| 2026 | Mixture of zero-cost proxies for training-free network architecture search
Junghyup Lee, Jeimin Jeon, Kwanghoon Sohn, Bumsub Ham |
Pattern Recognit. | 4 |
| 2026 | FVR: Feature variance reduction for post-hoc network calibration
Jongyoun Noh, Hyekang Park, Bumsub Ham |
Pattern Recognit. | 3 |
| 2025 | Maximizing the Position Embedding for Vision Transformers with Global Average PoolingabstractIn vision transformers, position embedding (PE) plays a crucial role in capturing the order of tokens. However, in vision transformer structures, there is a limitation in the expressiveness of PE due to the structure where position embedding is simply added to the token embedding. A layer-wise method that delivers PE to each layer and applies independent LNs for token embedding and PE has been adopted to overcome this limitation. In this paper, we identify the conflicting result that occurs in a layer-wise structure when using the global average pooling (GAP) method instead of the class token. To overcome this problem, we propose MPVG, which maximizes the effectiveness of PE in a layer-wise structure with GAP. Specifically, we identify that PE counterbalances token embedding values at each layer in a layer-wise structure. Furthermore, we recognize that the counterbalancing role of PE is insufficient in the layer-wise structure, and we address this by maximizing the effectiveness of PE through MPVG. Through experiments, we demonstrate that PE performs a counterbalancing role and that maintaining this counterbalancing directionality significantly impacts vision transformers. As a result, the experimental results show that MPVG outperforms existing methods across vision transformers on various tasks. Bumsub Ham, Suhyun Kim 0001 |
AAAI | 2 |
| 2025 | Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear FunctionsabstractNeural architecture search (NAS) enables finding the best-performing architecture from a search space automatically. Most NAS methods exploit an over-parameterized network (i.e., a supernet) containing all possible architectures (i.e., subnets) in the search space. However, the subnets that share the same set of parameters are likely to have different characteristics, interfering with each other during training. To address this, few-shot NAS methods have been proposed that divide the space into a few subspaces and employ a separate supernet for each subspace to limit the extent of weight sharing. They achieve state-of-the-art performance, but the computational cost increases accordingly. We introduce in this paper a novel few-shot NAS method that exploits the number of nonlinear functions to split the search space. To be specific, our method divides the space such that each subspace consists of subnets with the same number of nonlinear functions. Our splitting criterion is efficient, since it does not require comparing gradients of a supernet to split the space. In addition, we have found that dividing the space allows us to reduce the channel dimensions required for each supernet, which enables training multiple supernets in an efficient manner. We also introduce a supernet-balanced sampling (SBS) technique, sampling several subnets at each training step, to train different supernets evenly within a limited number of training steps. Extensive experiments on standard NAS benchmarks demonstrate the effectiveness of our approach. Youngmin Oh 0001, Bumsub Ham |
AAAI | 3 |
| 2025 | Subnet-Aware Dynamic Supernet Training for Neural Architecture SearchabstractN-shot neural architecture search (NAS) exploits a supernet containing all candidate subnets for a given search space. The subnets are typically trained with a static training strategy (e.g., using the same learning rate (LR) scheduler and optimizer for all subnets). This however does not consider that individual subnets have distinct characteristics, leading to two problems: (1) The supernet training is biased towards the low-complexity subnets (unfairness); (2) the momentum update in the supernet is noisy (noisy momentum). We present a dynamic supernet training technique to address these problems by adjusting the training strategy adaptive to the subnets. Specifically, we introduce a complexity-aware LR scheduler (CaLR) that controls the decay ratio of LR adaptive to the complexities of subnets, which alleviates the unfairness problem. We also present a momentum separation technique (MS). It groups the subnets with similar structural characteristics and uses a separate momentum for each group, avoiding the noisy momentum problem. Our approach can be applicable to various N-shot NAS methods with marginal cost, while improving the search performance drastically. We validate the effectiveness of our approach on various search spaces (e.g., NAS-Bench-201, Mobilenet spaces) and datasets (e.g., CIFAR-10/100, ImageNet). Jeimin Jeon, Youngmin Oh 0001, Junghyup Lee, Donghyeon Baek, Dohyung Kim 0006, Chanho Eom, Bumsub Ham |
CVPR | 7 |
| 2025 | Scheduling Weight Transitions for Quantization-Aware Training
Junghyup Lee, Jeimin Jeon, Dohyung Kim 0006, Bumsub Ham |
ICCV | 4 |
| 2025 | ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyabstractCurrent Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primarily rely on automated evaluation methods, but these methods struggle to detect implicit harmful content or produce inaccurate evaluations. Therefore, we found that existing benchmarks have low levels of harmfulness, ambiguous data, and limited diversity in image-text pair combinations. To address these issues, we propose the ELITE benchmark, a high-quality safety evaluation benchmark for VLMs, underpinned by our enhanced evaluation method, the ELITE evaluator. The ELITE evaluator explicitly incorporates a toxicity score to accurately assess harmfulness in multimodal contexts, where VLMs often provide specific, convincing, but unharmful descriptions of images. We filter out ambiguous and low-quality image-text pairs from existing benchmarks using the ELITE evaluator and generate diverse combinations of safe and unsafe image-text pairs. Our experiments demonstrate that the ELITE evaluator achieves superior alignment with human evaluations compared to prior automated methods, and the ELITE benchmark offers enhanced benchmark quality and diversity. By introducing ELITE, we pave the way for safer, more robust VLMs, contributing essential tools for evaluating and mitigating safety risks in real-world applications. Doehyeon Lee, Eugene Choi, Sangyoon Yu, Ashkan Yousefpour, Haon Park, Bumsub Ham, Suhyun Kim 0001 |
ICML | 7 |
| 2025 | AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion ModelsabstractWe present in this paper a novel post-training quantization (PTQ) method, dubbed AccuQuant, for diffusion models. We show analytically and empirically that quantization errors for diffusion models are accumulated over denoising steps in a sampling process. To alleviate the error accumulation problem, AccuQuant minimizes the discrepancies between outputs of a full-precision diffusion model and its quantized version within a couple of denoising steps. That is, it simulates multiple denoising steps of a diffusion sampling process explicitly for quantization, accounting the accumulated errors over multiple denoising steps, which is in contrast to previous approaches to imitating a training process of diffusion models, namely, minimizing the discrepancies independently for each step. We also present an efficient implementation technique for AccuQuant, together with a novel objective, which reduces a memory complexity significantly from $\mathcal{O}(n)$ to $\mathcal{O}(1)$, where $n$ is the number of denoising steps. We demonstrate the efficacy and efficiency of AccuQuant across various tasks and diffusion models on standard benchmarks. Jeongwoo Choi, Byunggwan Son, Jaehyeon Moon, Jeimin Jeon, Bumsub Ham |
NeurIPS | 6 |
| 2025 | Cerberus: Attribute-based person re-identification using semantic IDs
Chanho Eom, Kyunghwan Cho, Hyeonseok Jung, Moonsub Jin, Bumsub Ham |
Expert Syst. Appl. | 6 |
| 2025 | 3DPillars: Pillar-based two-stage 3D object detection
Jongyoun Noh, Junghyup Lee, Hyekang Park, Bumsub Ham |
Expert Syst. Appl. | 4 |
| 2024 | AZ-NAS: Assembling Zero-Cost Proxies for Network Architecture SearchabstractTraining-free network architecture search (NAS) aims to discover high-performing networks with zero-cost proxies, capturing network characteristics related to the final performance. However, network rankings estimated by previous training-free NAS methods have shown weak correlations with the performance. To address this issue, we propose AZ-NAS, a novel approach that leverages the ensemble of various zero-cost proxies to enhance the correlation between a predicted ranking of networks and the ground truth substantially in terms of the performance. To achieve this, we introduce four novel zero-cost proxies that are complementary to each other, analyzing distinct traits of architectures in the views of expressivity, progressivity, trainability, and complexity. The proxy scores can be obtained simultaneously within a single forward and backward pass, making an overall NAS process highly efficient. In order to integrate the rankings predicted by our proxies effectively, we introduce a non-linear ranking aggregation method that highlights the networks highly-ranked consistently across all the proxies. Experimental results conclusively demonstrate the efficacy and efficiency of AZ-NAS, outperforming state-of-the-art methods on standard benchmarks, all while maintaining a reasonable runtime cost. Junghyup Lee, Bumsub Ham |
CVPR | 2 |
| 2024 | Instance-Aware Group Quantization for Vision TransformersabstractPost-training quantization (PTQ) is an efficient model compression technique that quantizes a pretrained full-precision model using only a small calibration set of unla-beled samples without retraining. PTQ methods for convo-lutional neural networks (CNNs) provide quantization re-sults comparable to full-precision counterparts. Directly applying them to vision transformers (ViTs), however, in-curs severe performance degradation, mainly due to the dif-ferences in architectures between CNNs and ViTs. In par-ticular, the distribution of activations for each channel vary drastically according to input instances, making PTQ meth-ods for CNNs inappropriate for ViTs. To address this, we in-troduce instance-aware group quantization for ViTs (IGQ-ViT). To this end, we propose to split the channels of acti-vation maps into multiple groups dynamically for each in-put instance, such that activations within each group share similar statistical properties. We also extend our scheme to quantize softmax attentions across tokens. In addition, the number of groups for each layer is adjusted to minimize the discrepancies between predictions from quantized and full-precision models, under a bit-operation (BOP) constraint. We show extensive experimental results on image classification, object detection, and instance segmentation, with various transformer architectures, demonstrating the effectiveness of our approach. Jaehyeon Moon, Dohyung Kim 0006, Junyong Cheon, Bumsub Ham |
CVPR | 4 |
| 2024 | Toward INT4 Fixed-Point Training via Exploring Quantization Error for Gradients
Dohyung Kim 0006, Junghyup Lee, Jeimin Jeon, Jaehyeon Moon, Bumsub Ham |
ECCV (70) | 5 |
| 2024 | FYI: Flip Your Images for Dataset Distillation
Byunggwan Son, Youngmin Oh 0001, Donghyeon Baek, Bumsub Ham |
ECCV (50) | 4 |
| 2024 | PLoPS: Localization-aware person search with prototypical normalization
Youngmin Oh 0001, Donghyeon Baek, Junghyup Lee, Bumsub Ham |
Pattern Recognit. | 5 |
| 2023 | Camera-Driven Representation Learning for Unsupervised Domain Adaptive Person Re-identificationabstractWe present a novel unsupervised domain adaption method for person re-identification (reID) that generalizes a model trained on a labeled source domain to an unlabeled target domain. We introduce a camera-driven curriculum learning (CaCL) framework that leverages camera labels of person images to transfer knowledge from source to target domains progressively. To this end, we divide target domain dataset into multiple subsets based on the camera labels, and initially train our model with a single subset (i.e., images captured by a single camera). We then gradually exploit more subsets for training, according to a curriculum sequence obtained with a camera-driven scheduling rule. The scheduler considers maximum mean discrepancies (MMD) between each subset and the source domain dataset, such that the subset closer to the source domain is exploited earlier within the curriculum. For each curriculum sequence, we generate pseudo labels of person images in a target domain to train a reID model in a supervised way. We have observed that the pseudo labels are highly biased toward cameras, suggesting that person images obtained from the same camera are likely to have the same pseudo labels, even for different IDs. To address the camera bias problem, we also introduce a camera-diversity (CD) loss encouraging person images of the same pseudo label, but captured across various cameras, to involve more for discriminative feature learning, providing person representations robust to inter-camera variations. Experimental results on standard benchmarks, including real-to-real and synthetic-to-real scenarios, demonstrate the effectiveness of our framework. Dohyung Kim 0006, Younghoon Shin, Yongsang Yoon, Bumsub Ham |
ICCV | 6 |
| 2023 | RankMixup: Ranking-Based Mixup Training for Network CalibrationabstractNetwork calibration aims to accurately estimate the level of confidences, which is particularly important for employing deep neural networks in real-world systems. Recent approaches leverage mixup to calibrate the network’s predictions during training. However, they do not consider the problem that mixtures of labels in mixup may not accurately represent the actual distribution of augmented samples. In this paper, we present RankMixup, a novel mixup-based framework alleviating the problem of the mixture of labels for network calibration. To this end, we propose to use an ordinal ranking relationship between raw and mixup-augmented samples as an alternative supervisory signal to the label mixtures for network calibration. We hypothesize that the network should estimate a higher level of confidence for the raw samples than the augmented ones (Fig. 1). To implement this idea, we introduce a mixup-based ranking loss (MRL) that encourages lower confidences for augmented samples compared to raw ones, maintaining the ranking relationship. We also propose to leverage the ranking relationship among multiple mixup-augmented samples to further improve the calibration capability. Augmented samples with larger mixing coefficients are expected to have higher confidences and vice versa (Fig. 1). That is, the order of confidences should be aligned with that of mixing coefficients. To this end, we introduce a novel loss, M-NDCG, in order to reduce the number of misaligned pairs of the coefficients and confidences. Extensive experimental results on standard benchmarks for network calibration demonstrate the effectiveness of RankMixup. Jongyoun Noh, Hyekang Park, Junghyup Lee, Bumsub Ham |
ICCV | 4 |
| 2023 | ACLS: Adaptive and Conditional Label Smoothing for Network CalibrationabstractWe address the problem of network calibration adjusting miscalibrated confidences of deep neural networks. Many approaches to network calibration adopt a regularization-based method that exploits a regularization term to smooth the miscalibrated confidences. Although these approaches have shown the effectiveness on calibrating the networks, there is still a lack of understanding on the underlying principles of regularization in terms of network calibration. We present in this paper an in-depth analysis of existing regularization-based methods, providing a better understanding on how they affect to network calibration. Specifically, we have observed that 1) the regularization-based methods can be interpreted as variants of label smoothing, and 2) they do not always behave desirably. Based on the analysis, we introduce a novel loss function, dubbed ACLS, that unifies the merits of existing regularization methods, while avoiding the limitations. We show extensive experimental results for image classification and semantic segmentation on standard benchmarks, including CIFAR10, Tiny-ImageNet, ImageNet, and PASCAL VOC, demonstrating the effectiveness of our loss function. Hyekang Park, Jongyoun Noh, Youngmin Oh 0001, Donghyeon Baek, Bumsub Ham |
ICCV | 5 |
| 2022 | Bi-directional Contrastive Learning for Domain Adaptive Semantic Segmentation
Chanho Eom, Wonkyung Lee, Hyekang Park, Bumsub Ham |
ECCV (30) | 5 |
| 2022 | OIMNet++: Prototypical Normalization and Localization-Aware Learning for Person Search
Youngmin Oh 0001, Donghyeon Baek, Junghyup Lee, Bumsub Ham |
ECCV (10) | 5 |
| 2022 | Decomposed Knowledge Distillation for Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation (CISS) labels each pixel of an image with a corresponding object/stuff class continually. To this end, it is crucial to learn novel classes incrementally without forgetting previously learned knowledge. Current CISS methods typically use a knowledge distillation (KD) technique for preserving classifier logits, or freeze a feature extractor, to avoid the forgetting problem. The strong constraints, however, prevent learning discriminative features for novel classes. We introduce a CISS framework that alleviates the forgetting problem and facilitates learning novel classes effectively. We have found that a logit can be decomposed into two terms. They quantify how likely an input belongs to a particular class or not, providing a clue for a reasoning process of a model. The KD technique, in this context, preserves the sum of two terms ($\textit{i.e.}$, a class logit), suggesting that each could be changed and thus the KD does not imitate the reasoning process. To impose constraints on each term explicitly, we propose a new decomposed knowledge distillation (DKD) technique, improving the rigidity of a model and addressing the forgetting problem more effectively. We also introduce a novel initialization method to train new classifiers for novel classes. In CISS, the number of negative training samples for novel classes is not sufficient to discriminate old classes. To mitigate this, we propose to transfer knowledge of negatives to the classifiers successively using an auxiliary classifier, boosting the performance significantly. Experimental results on standard CISS benchmarks demonstrate the effectiveness of our framework. Donghyeon Baek, Youngmin Oh 0001, Junghyup Lee, Bumsub Ham |
NeurIPS | 5 |
| 2022 | ALIFE: Adaptive Logit Regularizer and Feature Replay for Incremental Semantic SegmentationabstractWe address the problem of incremental semantic segmentation (ISS) recognizing novel object/stuff categories continually without forgetting previous ones that have been learned. The catastrophic forgetting problem is particularly severe in ISS, since pixel-level ground-truth labels are available only for the novel categories at training time. To address the problem, regularization-based methods exploit probability calibration techniques to learn semantic information from unlabeled pixels. While such techniques are effective, there is still a lack of theoretical understanding of them. Replay-based methods propose to memorize a small set of images for previous categories. They achieve state-of-the-art performance at the cost of large memory footprint. We propose in this paper a novel ISS method, dubbed ALIFE, that provides a better compromise between accuracy and efficiency. To this end, we first show an in-depth analysis on the calibration techniques to better understand the effects on ISS. Based on this, we then introduce an adaptive logit regularizer (ALI) that enables our model to better learn new categories, while retaining knowledge for previous ones. We also present a feature replay scheme that memorizes features, instead of images directly, in order to reduce memory requirements significantly. Since a feature extractor is changed continually, memorized features should also be updated at every incremental stage. To handle this, we introduce category-specific rotation matrices updating the features for each category separately. We demonstrate the effectiveness of our approach with extensive experiments on standard ISS benchmarks, and show that our method achieves a better trade-off in terms of accuracy and efficiency. Youngmin Oh 0001, Donghyeon Baek, Bumsub Ham |
NeurIPS | 3 |
| 2022 | Disentangled Representations for Short-Term and Long-Term Person Re-IdentificationabstractWe address the problem of person re-identification (reID), that is, retrieving person images from a large dataset, given a query image of the person of interest. A key challenge is to learn person representations robust to intra-class variations, as different persons could have the same attribute, and persons' appearances look different, e.g., with viewpoint changes. Recent reID methods focus on learning person features discriminative only for a particular factor of variations (e.g., human pose), which also requires corresponding supervisory signals (e.g., pose annotations). To tackle this problem, we propose to factorize person images into identity-related and -unrelated features. Identity-related features contain information useful for specifying a particular person (e.g., clothing), while identity-unrelated ones hold other factors (e.g., human pose). To this end, we propose a new generative adversarial network, dubbed identity shuffle GAN (IS-GAN). It disentangles identity-related and -unrelated features from person images through an identity-shuffling technique that exploits identification labels alone without any auxiliary supervisory signals. We restrict the distribution of identity-unrelated features, or encourage the identity-related and -unrelated features to be uncorrelated, facilitating the disentanglement process. Experimental results validate the effectiveness of IS-GAN, showing state-of-the-art performance on standard reID benchmarks, including Market-1501, CUHK03 and DukeMTMC-reID. We further demonstrate the advantages of disentangling person representations on a long-term reID task, setting a new state of the art on a Celeb-reID dataset. Our code and models are available online: https://cvlab-yonsei.github.io/projects/ISGAN/. Chanho Eom, Wonkyung Lee, Bumsub Ham |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Learning Semantic Correspondence Exploiting an Object-Level PriorabstractWe address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal provides an object-level prior for the semantic correspondence task and offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks. Junghyup Lee, Dohyung Kim 0006, Wonkyung Lee, Jean Ponce, Bumsub Ham |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Network Quantization With Element-Wise Gradient ScalingabstractNetwork quantization aims at reducing bit-widths of weights and/or activations, particularly important for implementing deep neural networks with limited hardware resources. Most methods use the straight-through estimator (STE) to train quantized networks, which avoids a zero-gradient problem by replacing a derivative of a discretizer (i.e., a round function) with that of an identity function. Although quantized networks exploiting the STE have shown decent performance, the STE is sub-optimal in that it simply propagates the same gradient without considering discretization errors between inputs and outputs of the discretizer. In this paper, we propose an element-wise gradient scaling (EWGS), a simple yet effective alternative to the STE, training a quantized network better than the STE in terms of stability and accuracy. Given a gradient of the discretizer output, EWGS adaptively scales up or down each gradient element, and uses the scaled gradient as the one for the discretizer input to train quantized networks via backpropagation. The scaling is performed depending on both the sign of each gradient element and an error between the continuous input and discrete output of the discretizer. We adjust a scaling factor adaptively using Hessian information of a network. We show extensive experimental results on the image classification datasets, including CIFAR-10 and ImageNet, with diverse network architectures under a wide range of bit-width settings, demonstrating the effectiveness of our method. Junghyup Lee, Dohyung Kim 0006, Bumsub Ham |
CVPR | 3 |
| 2021 | HVPR: Hybrid Voxel-Point Representation for Single-Stage 3D Object DetectionabstractWe address the problem of 3D object detection, that is, estimating 3D object bounding boxes from point clouds. 3D object detection methods exploit either voxel-based or point-based features to represent 3D objects in a scene. Voxel-based features are efficient to extract, while they fail to preserve fine-grained 3D structures of objects. Point-based features, on the other hand, represent the 3D structures more accurately, but extracting these features is computationally expensive. We introduce in this paper a novel single-stage 3D detection method having the merit of both voxel-based and point-based features. To this end, we propose a new convolutional neural network (CNN) architecture, dubbed HVPR, that integrates both features into a single 3D representation effectively and efficiently. Specifically, we augment the point-based features with a memory module to reduce the computational cost. We then aggregate the features in the memory, semantically similar to each voxel-based one, to obtain a hybrid 3D representation in a form of a pseudo image, allowing to localize 3D objects in a single stage efficiently. We also propose an Attentive Multi-scale Feature Module (AMFM) that extracts scale-aware features considering the sparse and irregular patterns of point clouds. Experimental results on the KITTI dataset demonstrate the effectiveness and efficiency of our approach, achieving a better compromise in terms of speed and accuracy. Jongyoun Noh, Bumsub Ham |
CVPR | 3 |
| 2021 | Background-Aware Pooling and Noise-Aware Loss for Weakly-Supervised Semantic SegmentationabstractWe address the problem of weakly-supervised semantic segmentation (WSSS) using bounding box annotations. Although object bounding boxes are good indicators to segment corresponding objects, they do not specify object boundaries, making it hard to train convolutional neural networks (CNNs) for semantic segmentation. We find that background regions are perceptually consistent in part within an image, and this can be leveraged to discriminate foreground and background regions inside object bounding boxes. To implement this idea, we propose a novel pooling method, dubbed background-aware pooling (BAP), that focuses more on aggregating foreground features inside the bounding boxes using attention maps. This allows to extract high-quality pseudo segmentation labels to train CNNs for semantic segmentation, but the labels still contain noise especially at object boundaries. To address this problem, we also introduce a noise-aware loss (NAL) that makes the networks less susceptible to incorrect labels. Experimental results demonstrate that learning with our pseudo labels already outperforms state-of-the-art weakly- and semi-supervised methods on the PASCAL VOC 2012 dataset, and the NAL further boosts the performance. Youngmin Oh 0001, Bumsub Ham |
CVPR | 3 |
| 2021 | Exploiting a Joint Embedding Space for Generalized Zero-Shot Semantic SegmentationabstractWe address the problem of generalized zero-shot semantic segmentation (GZS3) predicting pixel-wise semantic labels for seen and unseen classes. Most GZS3 methods adopt a generative approach that synthesizes visual features of unseen classes from corresponding semantic ones (e.g., word2vec) to train novel classifiers for both seen and unseen classes. Although generative methods show decent performance, they have two limitations: (1) the visual features are biased towards seen classes; (2) the classifier should be retrained whenever novel unseen classes appear. We propose a discriminative approach to address these limitations in a unified framework. To this end, we leverage visual and semantic encoders to learn a joint embedding space, where the semantic encoder transforms semantic features to semantic prototypes that act as centers for visual features of corresponding classes. Specifically, we introduce boundary-aware regression (BAR) and semantic consistency (SC) losses to learn discriminative features. Our approach to exploiting the joint embedding space, together with BAR and SC terms, alleviates the seen bias problem. At test time, we avoid the retraining process by exploiting semantic prototypes as a nearest-neighbor (NN) classifier. To further alleviate the bias problem, we also propose an inference technique, dubbed Apollonius calibration (AC), that modulates the decision boundary of the NN classifier to the Apollonius circle adaptively. Experimental results demonstrate the effectiveness of our framework, achieving a new state of the art on standard benchmarks. Donghyeon Baek, Youngmin Oh 0001, Bumsub Ham |
ICCV | 3 |
| 2021 | Video-based Person Re-identification with Spatial and Temporal Memory NetworksabstractVideo-based person re-identification (reID) aims to retrieve person videos with the same identity as a query person across multiple cameras. Spatial and temporal distractors in person videos, such as background clutter and partial occlusions over frames, respectively, make this task much more challenging than image-based person reID. We observe that spatial distractors appear consistently in a particular location, and temporal distractors show several patterns, e.g., partial occlusions occur in the first few frames, where such patterns provide informative cues for predicting which frames to focus on (i.e., temporal attentions). Based on this, we introduce a novel Spatial and Temporal Memory Networks (STMN). The spatial memory stores features for spatial distractors that frequently emerge across video frames, while the temporal memory saves attentions which are optimized for typical temporal patterns in person videos. We leverage the spatial and temporal memories to refine frame-level person representations and to aggregate the refined frame-level features into a sequence-level person representation, respectively, effectively handling spatial and temporal distractors in person videos. We also introduce a memory spread loss preventing our model from addressing particular items only in the memories. Experimental results on standard benchmarks, including MARS, DukeMTMC-VideoReID, and LSVID, demonstrate the effectiveness of our method. Chanho Eom, Junghyup Lee, Bumsub Ham |
ICCV | 4 |
| 2021 | Distance-aware QuantizationabstractWe address the problem of network quantization, that is, reducing bit-widths of weights and/or activations to lighten network architectures. Quantization methods use a rounding function to map full-precision values to the nearest quantized ones, but this operation is not differentiable. There are mainly two approaches to training quantized networks with gradient-based optimizers. First, a straight-through estimator (STE) replaces the zero derivative of the rounding with that of an identity function, which causes a gradient mismatch problem. Second, soft quantizers approximate the rounding with continuous functions at training time, and exploit the rounding for quantization at test time. This alleviates the gradient mismatch, but causes a quantizer gap problem. We alleviate both problems in a unified framework. To this end, we introduce a novel quantizer, dubbed a distance-aware quantizer (DAQ), that mainly consists of a distance-aware soft rounding (DASR) and a temperature controller. To alleviate the gradient mismatch problem, DASR approximates the discrete rounding with the kernel soft argmax, which is based on our insight that the quantization can be formulated as a distance-based assignment problem between full-precision values and quantized ones. The controller adjusts the temperature parameter in DASR adaptively according to the input, addressing the quantizer gap problem. Experimental results on standard benchmarks show that DAQ outperforms the state of the art significantly for various bit-widths without bells and whistles. Dohyung Kim 0006, Junghyup Lee, Bumsub Ham |
ICCV | 3 |
| 2021 | Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal CorrespondencesabstractWe address the problem of visible-infrared person re-identification (VI-reID), that is, retrieving a set of person images, captured by visible or infrared cameras, in a cross-modal setting. Two main challenges in VI-reID are intraclass variations across person images, and cross-modal discrepancies between visible and infrared images. Assuming that the person images are roughly aligned, previous approaches attempt to learn coarse image- or rigid part-level person representations that are discriminative and generalizable across different modalities. However, the person images, typically cropped by off-the-shelf object detectors, are not necessarily well-aligned, which distract discriminative person representation learning. In this paper, we introduce a novel feature learning framework that addresses these problems in a unified way. To this end, we propose to exploit dense correspondences between cross-modal person images. This allows to address the cross-modal discrepancies in a pixel-level, suppressing modality-related features from person representations more effectively. This also encourages pixel-wise associations between cross-modal local features, further facilitating discriminative feature learning for VI-reID. Extensive experiments and analyses on standard VI-reID benchmarks demonstrate the effectiveness of our approach, which significantly outperforms the state of the art. Hyunjong Park, Junghyup Lee, Bumsub Ham |
ICCV | 4 |
| 2021 | Deformable Kernel Networks for Joint Image Filtering
Jean Ponce, Bumsub Ham |
Int. J. Comput. Vis. | 3 |
| 2020 | Relation Network for Person Re-IdentificationabstractPerson re-identification (reID) aims at retrieving an image of the person of interest from a set of images typically captured by multiple cameras. Recent reID methods have shown that exploiting local features describing body parts, together with a global feature of a person image itself, gives robust feature representations, even in the case of missing body parts. However, using the individual part-level features directly, without considering relations between body parts, confuses differentiating identities of different persons having similar attributes in corresponding parts. To address this issue, we propose a new relation network for person reID that considers relations between individual body parts and the rest of them. Our model makes a single part-level feature incorporate partial information of other body parts as well, supporting it to be more discriminative. We also introduce a global contrastive pooling (GCP) method to obtain a global feature of a person image. We propose to use contrastive features for GCP to complement conventional max and averaging pooling techniques. We show that our model outperforms the state of the art on the Market1501, DukeMTMC-reID and CUHK03 datasets, demonstrating the effectiveness of our approach on discriminative person representations. Hyunjong Park, Bumsub Ham |
AAAI | 2 |
| 2020 | Learning Memory-Guided Normality for Anomaly DetectionabstractWe address the problem of anomaly detection, that is, detecting anomalous events in a video sequence. Anomaly detection methods based on convolutional neural networks (CNNs) typically leverage proxy tasks, such as reconstructing input video frames, to learn models describing normality without seeing anomalous samples at training time, and quantify the extent of abnormalities using the reconstruction error at test time. The main drawbacks of these approaches are that they do not consider the diversity of normal patterns explicitly, and the powerful representation capacity of CNNs allows to reconstruct abnormal video frames. To address this problem, we present an unsupervised learning approach to anomaly detection that considers the diversity of normal patterns explicitly, while lessening the representation capacity of CNNs. To this end, we propose to use a memory module with a new update scheme where items in the memory record prototypical patterns of normal data. We also present novel feature compactness and separateness losses to train the memory, boosting the discriminative power of both memory items and deeply learned features from normal data. Experimental results on standard benchmarks demonstrate the effectiveness and efficiency of our approach, which outperforms the state of the art. Hyunjong Park, Jongyoun Noh, Bumsub Ham |
CVPR | 3 |
| 2020 | Learning with Privileged Information for Efficient Image Super-Resolution
Wonkyung Lee, Junghyup Lee, Dohyung Kim 0006, Bumsub Ham |
ECCV (24) | 4 |
| 2020 | Temporally Consistent Depth Prediction With Flow-Guided Memory UnitsabstractPredicting depth from a monocular video sequence is an important task for autonomous driving. Although it has advanced considerably in the past few years, recent methods based on convolutional neural networks (CNNs) discard temporal coherence in the video sequence and estimate depth independently for each frame, which often leads to undesired inconsistent results over time. To address this problem, we propose to memorize temporal consistency in the video sequence, and leverage it for the task of depth prediction. To this end, we introduce a two-stream CNN with a flow-guided memory module, where each stream encodes visual and temporal features, respectively. The memory module, implemented using convolutional gated recurrent units (ConvGRUs), inputs visual and temporal features sequentially together with optical flow tailored to our task. It memorizes trajectories of individual features selectively and propagates spatial information over time, enforcing a long-term temporal consistency to prediction results. We evaluate our method on the KITTI benchmark dataset in terms of depth prediction accuracy, temporal consistency and runtime, and achieve a new state of the art. We also provide an extensive experimental analysis, clearly demonstrating the effectiveness of our approach to memorizing temporal consistency for depth prediction. Chanho Eom, Hyunjong Park, Bumsub Ham |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | SFNet: Learning Object-Aware Semantic CorrespondenceabstractWe address the problem of semantic correspondence, that is, establishing a dense flow field between images depicting different instances of the same object or scene category. We propose to use images annotated with binary foreground masks and subjected to synthetic geometric deformations to train a convolutional neural network (CNN) for this task. Using these masks as part of the supervisory signal offers a good compromise between semantic flow methods, where the amount of training data is limited by the cost of manually selecting point correspondences, and semantic alignment ones, where the regression of a single global geometric transformation between images may be sensitive to image-specific details such as background clutter. We propose a new CNN architecture, dubbed SFNet, which implements this idea. It leverages a new and differentiable version of the argmax function for end-to-end training, with a loss that combines mask and flow consistency with smoothness terms. Experimental results demonstrate the effectiveness of our approach, which significantly outperforms the state of the art on standard benchmarks. Junghyup Lee, Dohyung Kim 0006, Jean Ponce, Bumsub Ham |
CVPR | 4 |
| 2019 | Learning Disentangled Representation for Robust Person Re-identificationabstractWe address the problem of person re-identification (reID), that is, retrieving person images from a large dataset, given a query image of the person of interest. The key challenge is to learn person representations robust to intra-class variations, as different persons can have the same attribute and the same person's appearance looks different with viewpoint changes. Recent reID methods focus on learning discriminative features but robust to only a particular factor of variations (e.g., human pose) and this requires corresponding supervisory signals (e.g., pose annotations). To tackle this problem, we propose to disentangle identity-related and -unrelated features from person images. Identity-related features contain information useful for specifying a particular person (e.g.,clothing), while identity-unrelated ones hold other factors (e.g., human pose, scale changes). To this end, we introduce a new generative adversarial network, dubbed identity shuffle GAN (IS-GAN), that factorizes these features using identification labels without any auxiliary information. We also propose an identity shuffling technique to regularize the disentangled features. Experimental results demonstrate the effectiveness of IS-GAN, largely outperforming the state of the art on standard reID benchmarks including the Market-1501, CUHK03 and DukeMTMC-reID. Our code and models will be available online at the time of the publication. Chanho Eom, Bumsub Ham |
NeurIPS | 2 |
| 2019 | OCEAN: Object-centric arranging network for self-supervised visual representations learning
Changjae Oh, Bumsub Ham, Hansung Kim 0001, Adrian Hilton 0001, Kwanghoon Sohn |
Expert Syst. Appl. | 2 |
| 2019 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. Unlike traditional dense correspondence approaches for estimating depth or optical flow, semantic correspondence estimation poses additional challenges due to intra-class appearance and shape variations among different instances within the same object or scene category. To robustly match points across semantically similar images, we formulate FCSS using local self-similarity (LSS), which is inherently insensitive to intra-class appearance variations. LSS is incorporated through a proposed convolutional self-similarity (CSS) layer, where the sampling patterns and the self-similarity measure are jointly learned in an end-to-end and multi-scale manner. Furthermore, to address shape variations among different object instances, we propose a convolutional affine transformer (CAT) layer that estimates explicit affine transformation fields at each pixel to transform the sampling patterns and corresponding receptive fields. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in most existing datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS significantly outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Structure-Texture Image Decomposition Using Deep Variational PriorsabstractMost variational formulations for structure-texture image decomposition force structure images to have small norm in some functional spaces, and share a common notion of edges, i.e., large-gradients or -intensity differences. However, such definition makes it difficult to distinguish structure edges from oscillations that have fine spatial scale but high contrast. In this paper, we introduce a new model by learning deep variational prior for structure images without explicit training data. An alternating direction method of multiplier (ADMM) algorithm and its modular structure are adopted to plug deep variational priors into an iterative smoothing process. The central observations are that convolution neural networks (CNNs) can replace the total variation prior, and are indeed powerful to capture the natures of structure and texture. We show that our learned priors using CNNs successfully differentiate highamplitude details from structure edges, and avoid halo artifacts. Different from previous data-driven smoothing schemes, our formulation provides another degree of freedom to produce continuous smoothing effects. Experimental results demonstrate the effectiveness of our approach on various computational photography and image processing applications, including texture removal, detail manipulation, HDR tone-mapping, and nonphotorealistic abstraction. Youngjung Kim, Bumsub Ham, Minh N. Do, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2018 | Robust Guided Image Filtering Using Nonconvex PotentialsabstractFiltering images using a guidance signal, a process called guided or joint image filtering, has been used in various tasks in computer vision and computational photography, particularly for noise reduction and joint upsampling. This uses an additional guidance signal as a structure prior, and transfers the structure of the guidance signal to an input image, restoring noisy or altered image structure. The main drawbacks of such a data-dependent framework are that it does not consider structural differences between guidance and input images, and that it is not robust to outliers. We propose a novel SD (for static/dynamic) filter to address these problems in a unified framework, and jointly leverage structural information from guidance and input images. Guided image filtering is formulated as a nonconvex optimization problem, which is solved by the majorize-minimization algorithm. The proposed algorithm converges quickly while guaranteeing a local minimum. The SD filter effectively controls the underlying image structure at different scales, and can handle a variety of types of data from different sensors. It is robust to outliers and other artifacts such as gradient reversal and global intensity shift, and has good edge-preserving smoothing properties. We demonstrate the flexibility and effectiveness of the proposed SD filter in a variety of applications, including depth upsampling, scale-space filtering, texture removal, flash/non-flash denoising, and RGB/NIR denoising. Bumsub Ham, Minsu Cho, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Proposal Flow: Semantic Correspondences from Object ProposalsabstractFinding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic flow, dubbed proposal flow, that establishes reliable correspondences using object proposals. Unlike prevailing semantic flow approaches that operate on pixels or regularly sampled local regions, proposal flow benefits from the characteristics of modern object proposals, that exhibit high repeatability at multiple scales, and can take advantage of both local and geometric consistency constraints among proposals. We also show that the corresponding sparse proposal flow can effectively be transformed into a conventional dense flow field. We introduce two new challenging datasets that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow. We use these benchmarks to compare different matching algorithms, object proposals, and region features within proposal flow, to the state of the art in semantic flow. This comparison, along with experiments on standard datasets, demonstrates that proposal flow significantly outperforms existing semantic flow methods in various settings. Bumsub Ham, Minsu Cho, Cordelia Schmid, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to existing CNN-based descriptors, FCSS is inherently insensitive to intra-class appearance variations because of its LSS-based structure, while maintaining the precise localization ability of deep neural networks. The sampling patterns of local structure and the self-similarity measure are jointly learned within the proposed network in an end-to-end and multi-scale manner. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in existing image datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Sangryul Jeon, Stephen Lin 0001, Kwanghoon Sohn |
CVPR | 3 |
| 2017 | SCNet: Learning Semantic CorrespondenceabstractThis paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearance only. We propose instead a convolutional neural network architecture, called SCNet, for learning a geometrically plausible model for semantic correspondence. SCNet uses region proposals as matching primitives, and explicitly incorporates geometric consistency in its loss function. It is trained on image pairs obtained from the PASCAL VOC 2007 keypoint dataset, and a comparative evaluation on several standard benchmarks demonstrates that the proposed approach substantially outperforms both recent deep learning architectures and previous methods based on hand-crafted features. Kai Han 0001, Rafael S. Rezende, Bumsub Ham, Kwan-Yee Kenneth Wong, Minsu Cho, Cordelia Schmid, Jean Ponce |
ICCV | 3 |
| 2017 | Convolutional cost aggregation for robust stereo matchingabstractAlthough convolutional neural network (CNN)-based stereo matching methods have become increasingly popular thanks to their robustness, they primarily have been focused on the matching cost computation. By leveraging CNNs, we present a novel method for matching cost aggregation to boost the stereo matching performance. Our insight is to learn the convolution kernel within CNN architecture for cost aggregation in a fully convolutional manner. Tailored to cost aggregation problem, our method differs from handcrafted methods in terms of its convolutional aggregation through optimally learned CNNs. First, the matching cost is aggregated with cost volume unary network, and then optimized with explicit disparity boundary, estimated through disparity boundary pairwise network, within a global energy minimization. Experiments demonstrate that our method outperforms conventional hand-crafted aggregation methods. Somi Jeong, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn |
ICIP | 3 |
| 2017 | Unsupervised stereo matching using correspondence consistencyabstractDeep convolutional neural networks (CNNs) have shown revolutionary performance improvements for matching cost computation in stereo matching. However, conventional CNN-based approaches to learn the network in a supervised manner require a large number of ground-truth disparity maps, which limits their applicability. To overcome this limitation, we present a novel framework to learn a CNNs architecture for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. Exploiting the correspondence consistency between stereo images as supervision, our method selects the training samples in each iteration during network training and uses them to learn the network. To boost the performance, we also propose a multi-scale cost computation scheme. Experimental results show that our method outperforms the state-of-the-art methods including even supervised learning based methods on various benchmarks. Sunghun Joung, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn |
ICIP | 3 |
| 2017 | Deep stereo confidence prediction for depth estimationabstractWe present a novel method that predicts a confidence to improve the accuracy of an estimated depth map in stereo matching. In contrast to existing learning based approaches relying on hand-crafted confidence features, we cast this problem into a convolutional neural network, learned using both a matching cost volume and its associated disparity map. As the size of the matching cost volume varies depending on a search range of stereo image pairs, we propose to use a top-K matching probability volume layer so that an input size for convolutional layers remains unchanged. Experimental results demonstrate that the proposed method outperforms the state-of-the-art confidence estimation approaches on various benchmarks. Sunok Kim, Dongbo Min, Bumsub Ham, Seungryong Kim, Kwanghoon Sohn |
ICIP | 3 |
| 2017 | Robust interactive image segmentation using structure-aware labeling
Changjae Oh, Bumsub Ham, Kwanghoon Sohn |
Expert Syst. Appl. | 2 |
| 2017 | DASC: Robust Dense Descriptor for Multi-Modal and Multi-Spectral Correspondence EstimationabstractEstablishing dense correspondences between multiple images is a fundamental task in many applications. However, finding a reliable correspondence between multi-modal or multi-spectral images still remains unsolved due to their challenging photometric and geometric variations. In this paper, we propose a novel dense descriptor, called dense adaptive self-correlation (DASC), to estimate dense multi-modal and multi-spectral correspondences. Based on an observation that self-similarity existing within images is robust to imaging modality variations, we define the descriptor with a series of an adaptive self-correlation similarity measure between patches sampled by a randomized receptive field pooling, in which a sampling pattern is obtained using a discriminative learning. The computational redundancy of dense descriptors is dramatically reduced by applying fast edge-aware filtering. Furthermore, in order to address geometric variations including scale and rotation, we propose a geometry-invariant DASC (GI-DASC) descriptor that effectively leverages the DASC through a superpixel-based representation. For a quantitative evaluation of the GI-DASC, we build a novel multi-modal benchmark as varying photometric and geometric conditions. Experimental results demonstrate the outstanding performance of the DASC and GI-DASC in many cases of dense multi-modal and multi-spectral correspondences. Seungryong Kim, Dongbo Min, Bumsub Ham, Minh N. Do, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Fast Domain Decomposition for Global Image SmoothingabstractEdge-preserving smoothing (EPS) can be formulated as minimizing an objective function that consists of data and regularization terms. At the price of high-computational cost, this global EPS approach is more robust and versatile than a local one that typically has a form of weighted averaging. In this paper, we introduce an efficient decomposition-based method for global EPS that minimizes the objective function of L2 data and (possibly non-smooth and non-convex) regularization terms in linear time. Different from previous decompositionbased methods, which require solving a large linear system, our approach solves an equivalent constrained optimization problem, resulting in a sequence of 1-D sub-problems. This enables applying fast linear time solver for weighted-least squares and -L1 smoothing problems. An alternating direction method of multipliers algorithm is adopted to guarantee fast convergence. Our method is fully parallelizable, and its runtime is even comparable to the state-of-the-art local EPS approaches. We also propose a family of fast majorization-minimization algorithms that minimize an objective with non-convex regularization terms. Experimental results demonstrate the effectiveness and flexibility of our approach in a range of image processing and computational photography applications. Youngjung Kim, Dongbo Min, Bumsub Ham, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2016 | Point-Cut: Interactive Image Segmentation Using Point Supervision
Changjae Oh, Bumsub Ham, Kwanghoon Sohn |
ACCV (1) | 2 |
| 2016 | Proposal FlowabstractFinding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic flow, dubbed proposal flow, that establishes reliable correspondences using object proposals. Unlike prevailing semantic flow approaches that operate on pixels or regularly sampled local regions, proposal flow benefits from the characteristics of modern object proposals, that exhibit high repeatability at multiple scales, and can take advantage of both local and geometric consistency constraints among proposals. We also show that proposal flow can effectively be transformed into a conventional dense flow field. We introduce a new dataset that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow. We use this benchmark to compare different matching algorithms, object proposals, and region features within proposal flow, to the state of the art in semantic flow. This comparison, along with experiments on standard datasets, demonstrates that proposal flow significantly outperforms existing semantic flow methods in various settings. Bumsub Ham, Minsu Cho, Cordelia Schmid, Jean Ponce |
CVPR | 1 |
| 2016 | Real-time rear obstacle detection using reliable disparity for driver assistance
Hunjae Yoo, Jongin Son, Bumsub Ham, Kwanghoon Sohn |
Expert Syst. Appl. | 3 |
| 2016 | Structure Selective Depth Superresolution for RGB-D CamerasabstractThis paper describes a method for high-quality depth superresolution. The standard formulations of image-guided depth upsampling, using simple joint filtering or quadratic optimization, lead to texture copying and depth bleeding artifacts. These artifacts are caused by inherent discrepancy of structures in data from different sensors. Although there exists some correlation between depth and intensity discontinuities, they are different in distribution and formation. To tackle this problem, we formulate an optimization model using a nonconvex regularizer. A nonlocal affinity established in a high-dimensional feature space is used to offer precisely localized depth boundaries. We show that the proposed method iteratively handles differences in structure between depth and intensity images. This property enables reducing texture copying and depth bleeding artifacts significantly on a variety of range data sets. We also propose a fast alternating direction method of multipliers algorithm to solve our optimization problem. Our solver shows a noticeable speed up compared with the conventional majorize-minimize algorithm. Extensive experiments with synthetic and real-world data sets demonstrate that the proposed method is superior to the existing methods. Youngjung Kim, Bumsub Ham, Changjae Oh, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2015 | Robust image filtering using joint static and dynamic guidanceabstractRegularizing images under a guidance signal has been used in various tasks in computer vision and computational photography, particularly for noise reduction and joint upsampling. The aim is to transfer fine structures of guidance signals to input images, restoring noisy or altered structures. One of main drawbacks in such a data-dependent framework is that it does not handle differences in structure between guidance and input images. We address this problem by jointly leveraging structural information of guidance and input images. Image filtering is formulated as a nonconvex optimization problem, which is solved by the majorization-minimization algorithm. The proposed algorithm converges quickly while guaranteeing a local minimum. It effectively controls image structures at different scales and can handle a variety of types of data from different sensors. We demonstrate the flexibility and effectiveness of our model in several applications including depth super-resolution, scale-space filtering, texture removal, flash/non-flash denoising, and RGB/NIR denoising. Bumsub Ham, Minsu Cho, Jean Ponce |
CVPR | 1 |
| 2015 | DASC: Dense adaptive self-correlation descriptor for multi-modal and multi-spectral correspondenceabstractEstablishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramatically advanced in recent studies. However, finding reliable visual correspondence in multi-modal or multi-spectral images still remains unsolved. In this paper, we propose a novel dense matching descriptor, called dense adaptive self-correlation (DASC), to effectively address this kind of matching scenarios. Based on the observation that a self-similarity existing within images is less sensitive to modality variations, we define the descriptor with a series of an adaptive self-correlation similarity for patches within a local support window. To further improve the matching quality and runtime efficiency, we propose a randomized receptive field pooling, in which a sampling pattern is optimized with a discriminative learning. Moreover, the computational redundancy that arises when computing densely sampled descriptor over an entire image is dramatically reduced by applying fast edge-aware filtering. Experiments demonstrate the outstanding performance of the DASC descriptor in many cases of multi-modal and multi-spectral correspondence. Seungryong Kim, Dongbo Min, Bumsub Ham, Seungchul Ryu, Minh N. Do, Kwanghoon Sohn |
CVPR | 3 |
| 2015 | Depth Analogy: Data-Driven Approach for Single Image Depth Estimation Using Gradient SamplesabstractInferring scene depth from a single monocular image is a highly ill-posed problem in computer vision. This paper presents a new gradient-domain approach, called depth analogy, that makes use of analogy as a means for synthesizing a target depth field, when a collection of RGB-D image pairs is given as training data. Specifically, the proposed method employs a non-parametric learning process that creates an analogous depth field by sampling reliable depth gradients using visual correspondence established on training image pairs. Unlike existing data-driven approaches that directly select depth values from training data, our framework transfers depth gradients as reconstruction cues, which are then integrated by the Poisson reconstruction. The performance of most conventional approaches relies heavily on the training RGB-D data used in the process, and such a dependency severely degenerates the quality of reconstructed depth maps when the desired depth distribution of an input image is quite different from that of the training data, e.g., outdoor versus indoor scenes. Our key observation is that using depth gradients in the reconstruction is less sensitive to scene characteristics, providing better cues for depth recovery. Thus, our gradient-domain approach can support a great variety of training range datasets that involve substantial appearance and geometric variations. The experimental results demonstrate that our (depth) gradient-domain approach outperforms existing data-driven approaches directly working on depth domain, even when only uncorrelated training datasets are available. Sunghwan Choi, Dongbo Min, Bumsub Ham, Youngjung Kim, Changjae Oh, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2015 | Unsupervised Texture Flow Estimation Using Appearance-Space Clustering and CorrespondenceabstractThis paper presents a texture flow estimation method that uses an appearance-space clustering and a correspondence search in the space of deformed exemplars. To estimate the underlying texture flow, such as scale, orientation, and texture label, most existing approaches require a certain amount of user interactions. Strict assumptions on a geometric model further limit the flow estimation to such a near-regular texture as a gradient-like pattern. We address these problems by extracting distinct texture exemplars in an unsupervised way and using an efficient search strategy on a deformation parameter space. This enables estimating a coherent flow in a fully automatic manner, even when an input image contains multiple textures of different categories. A set of texture exemplars that describes the input texture image is first extracted via a medoid-based clustering in appearance space. The texture exemplars are then matched with the input image to infer deformation parameters. In particular, we define a distance function for measuring a similarity between the texture exemplar and a deformed target patch centered at each pixel from the input image, and then propose to use a randomized search strategy to estimate these parameters efficiently. The deformation flow field is further refined by adaptively smoothing the flow field under guidance of a matching confidence score. We show that a local visual similarity, directly measured from appearance space, explains local behaviors of the flow very well, and the flow field can be estimated very efficiently when the matching criterion meets the randomized search strategy. Experimental results on synthetic and natural images show that the proposed method outperforms existing methods. Sunghwan Choi, Dongbo Min, Bumsub Ham, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2015 | Depth Superresolution by TransductionabstractThis paper presents a depth superresolution (SR) method that uses both of a low-resolution (LR) depth image and a high-resolution (HR) intensity image. We formulate depth SR as a graph-based transduction problem. In particular, the HR intensity image is represented as an undirected graph, in which pixels are characterized as vertices, and their relations are encoded as an affinity function. When the vertices initially labeled with certain depth hypotheses (from the LR depth image) are regarded as input queries, all the vertices are scored with respect to the relevances to these queries by a classifying function. Each vertex is then labeled with the depth hypothesis that receives the highest relevance score. We design the classifying function by considering the local and global structures of the HR intensity image. This approach enables us to address a depth bleeding problem that typically appears in current depth SR methods. Furthermore, input queries are assigned in a probabilistic manner, making depth SR robust to noisy depth measurements. We also analyze existing depth SR methods in the context of transduction, and discuss their theoretic relations. Intensive experiments demonstrate the superiority of the proposed method over state-of-the-art methods both qualitatively and quantitatively. Bumsub Ham, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2014 | Robust Stereo Matching Using Probabilistic Laplacian Surface Propagation
Seungryong Kim, Bumsub Ham, Seungchul Ryu, Seon Joo Kim, Kwanghoon Sohn |
ACCV (1) | 2 |
| 2014 | Local self-similarity frequency descriptor for multispectral feature matchingabstractThis paper describes a robust feature descriptor called the local self-similarity frequency (LSSF) for the multispectral RGB-NIR feature matching, which uses the frequency response of the local internal layout of self-similarities. A nonlinear relationship between multi-spectral image pairs makes conventional descriptors be sensitive to spectral deformation. To alleviate this problem, the LSSF employs a weighted correlation surface reducing the discrepancy between mul-tispectral images. Furthermore, the LSSF provides a rotation invariance exploiting the frequency response of maximal values on logpolar bins based on the fact that a cyclic shift on the log-polar representation leads only a phase shift in a frequency domain. Experimental results show that LSSF outperforms state-of-the-art descriptors in terms of a recognition rate for multispectral RGB-NIR image pairs. Seungryong Kim, Seungchul Ryu, Bumsub Ham, Junhyung Kim, Kwanghoon Sohn |
ICIP | 3 |
| 2014 | Mahalanobis Distance Cross-Correlation for Illumination-Invariant Stereo MatchingabstractA robust similarity measure called the Mahalanobis distance cross-correlation (MDCC) is proposed for illumination-invariant stereo matching, which uses a local color distribution within support windows. It is shown that the Mahalanobis distance between the color itself and the average color is preserved under affine transformation. The MDCC converts pixels within each support window into the Mahalanobis distance transform (MDT) space. The similarity between MDT pairs is then computed using the cross-correlation with an asymmetric weight function based on the Mahalanobis distance. The MDCC considers correlation on cross-color channels, thus providing robustness to affine illumination variation. Experimental results show that the MDCC outperforms state-of-the-art similarity measures in terms of stereo matching for image pairs taken under different illumination conditions. Seungryong Kim, Bumsub Ham, Bongjoe Kim, Kwanghoon Sohn |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Probability-Based Rendering for View SynthesisabstractIn this paper, a probability-based rendering (PBR) method is described for reconstructing an intermediate view with a steady-state matching probability (SSMP) density function. Conventionally, given multiple reference images, the intermediate view is synthesized via the depth image-based rendering technique in which geometric information (e.g., depth) is explicitly leveraged, thus leading to serious rendering artifacts on the synthesized view even with small depth errors. We address this problem by formulating the rendering process as an image fusion in which the textures of all probable matching points are adaptively blended with the SSMP representing the likelihood that points among the input reference images are matched. The PBR hence becomes more robust against depth estimation errors than existing view synthesis approaches. The MP in the steady-state, SSMP, is inferred for each pixel via the random walk with restart (RWR). The RWR always guarantees visually consistent MP, as opposed to conventional optimization schemes (e.g., diffusion or filtering-based approaches), the accuracy of which heavily depends on parameters used. Experimental results demonstrate the superiority of the PBR over the existing view synthesis approaches both qualitatively and quantitatively. Especially, the PBR is effective in suppressing flicker artifacts of virtual video rendering although no temporal aspect is considered. Moreover, it is shown that the depth map itself calculated from our RWR-based method (by simply choosing the most probable matching point) is also comparable with that of the state-of-the-art local stereo matching methods. Bumsub Ham, Dongbo Min, Changjae Oh, Minh N. Do, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2014 | Fast Global Image Smoothing Based on Weighted Least SquaresabstractThis paper presents an efficient technique for performing a spatially inhomogeneous edge-preserving image smoothing, called fast global smoother. Focusing on sparse Laplacian matrices consisting of a data term and a prior term (typically defined using four or eight neighbors for 2D image), our approach efficiently solves such global objective functions. In particular, we approximate the solution of the memory-and computation-intensive large linear system, defined over a d-dimensional spatial domain, by solving a sequence of 1D subsystems. Our separable implementation enables applying a linear-time tridiagonal matrix algorithm to solve d three-point Laplacian matrices iteratively. Our approach combines the best of two paradigms, i.e., efficient edge-preserving filters and optimization-based smoothing. Our method has a comparable runtime to the fast edge-preserving filters, but its global optimization formulation overcomes many limitations of the local filtering approaches. Our method also achieves high-quality results as the state-of-the-art optimization-based techniques, but runs ∼10-30 times faster. Besides, considering the flexibility in defining an objective function, we further propose generalized fast algorithms that perform Lγ norm smoothing (0 < γ < 2) and support an aggregated (robust) data term for handling imprecise data constraints. We demonstrate the effectiveness and efficiency of our techniques in a range of image processing and computer graphics applications. Dongbo Min, Sunghwan Choi, Jiangbo Lu, Bumsub Ham, Kwanghoon Sohn, Minh N. Do |
IEEE Trans. Image Process. | 4 |
| 2013 | Fast image retargeting via axis-aligned importance scalingabstractIn this paper, we propose an image retargeting method that resizes an image by the axis-aligned importance scaling. The proposed method operates on the axis-aligned deformation space, where the mesh structure is parameterized in the 1D vector, i.e., quads on the same column (or row) share the single parameter. The unknown variables are thus dependent on the grid resolution only, allowing a fast and simple implementation on a moderate CPU. The optimal parameters are inferred in an iterative manner: these parameters are updated by scaling initial ones according to the deformation error. It is measured at each iteration by aggregating the transition cost for the deformed quad. Experimental results show that the proposed method preserves visually salient features without foldover artifacts better than competing methods. In addition, the optimal parameters can be calculated within 0.02 ms on a single-core CPU. Sunghwan Choi, Bumsub Ham, Kwanghoon Sohn |
ICIP | 2 |
| 2013 | ABFT: Anisotropic binary feature transform based on structure tensor spaceabstractLocal feature matching is a fundamental step for many computer vision applications. Recently, binary feature transforms have been popularly proposed to improve the computational efficiency while preserving high matching performance. However, it is sensitive to noise and geometrical distortion such as affine transformation. In this paper, we propose ABFT framework, composed of a noise robust feature detection and affine invariant binary feature description based on a structure tensor space. Experimental results show that ABFT outperforms other state-of-the-art feature transforms in terms of the repeatability, recognition rate, and computational time. Seungryong Kim, Hunjae Yoo, Seungchul Ryu, Bumsub Ham, Kwanghoon Sohn |
ICIP | 4 |
| 2013 | Contextual information based visual saliency modelabstractAutomatic detection of visual saliency has been considered a very important task because of a wide range of applications such as object detection, image quality assessment, image segmentation, and more. Thanks to active researches in this field, many effective saliency models have been developed. Nevertheless, several challenging problems are still remain unsolved, such as detecting saliency in complex scene and providing high resolution and accurate saliency maps. In order to address such challenging problems, we propose a visual saliency model based on the concept of contextual information. First, we introduce a general framework for detecting saliency of an image using contextual information. Then, the proposed saliency model based on color and shape features is proposed. Quantitative and qualitative comparisons with seven state-of-the-art models on the public database show that the proposed model achieves excellent performance. Especially, the proposed model can provide good performance on challenging images including images with cluttered background and repeating distractors compared to the other models. Seungchul Ryu, Bumsub Ham, Kwanghoon Sohn |
ICIP | 2 |
| 2013 | Space-Time Hole Filling With Random Walks in View Extrapolation for 3D VideoabstractIn this paper, a space-time hole filling approach is presented to deal with a disocclusion when a view is synthesized for the 3D video. The problem becomes even more complicated when the view is extrapolated from a single view, since the hole is large and has no stereo depth cues. Although many techniques have been developed to address this problem, most of them focus only on view interpolation. We propose a space-time joint filling method for color and depth videos in view extrapolation. For proper texture and depth to be sampled in the following hole filling process, the background of a scene is automatically segmented by the random walker segmentation in conjunction with the hole formation process. Then, the patch candidate selection process is formulated as a labeling problem, which can be solved with random walks. The patch candidates that best describe the hole region are dynamically selected in the space-time domain, and the hole is filled with the optimal patch for ensuring both spatial and temporal coherence. The experimental results show that the proposed method is superior to state-of-the-art methods and provides both spatially and temporally consistent results with significantly reduced flicker artifacts. Sunghwan Choi, Bumsub Ham, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2013 | Revisiting the Relationship Between Adaptive Smoothing and Anisotropic Diffusion With Modified FiltersabstractAnisotropic diffusion has been known to be closely related to adaptive smoothing and discretized in a similar manner. This paper revisits a fundamental relationship between two approaches. It is shown that adaptive smoothing and anisotropic diffusion have different theoretical backgrounds by exploring their characteristics with the perspective of normalization, evolution step size, and energy flow. Based on this principle, adaptive smoothing is derived from a second order partial differential equation (PDE), not a conventional anisotropic diffusion, via the coupling of Fick's law with a generalized continuity equation where a "source" or "sink" exists, which has not been extensively exploited. We show that the source or sink is closely related to the asymmetry of energy flow as well as the normalization term of adaptive smoothing. It enables us to analyze behaviors of adaptive smoothing, such as the maximum principle and stability with a perspective of a PDE. Ultimately, this relationship provides new insights into application-specific filtering algorithm design. By modeling the source or sink in the PDE, we introduce two specific diffusion filters, the robust anisotropic diffusion and the robust coherence enhancing diffusion, as novel instantiations which are more robust against the outliers than the conventional filters. Bumsub Ham, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2013 | A Generalized Random Walk With Restart and its Application in Depth Up-Sampling and Interactive SegmentationabstractIn this paper, the origin of random walk with restart (RWR) and its generalization are described. It is well known that the random walk (RW) and the anisotropic diffusion models share the same energy functional, i.e., the former provides a steady-state solution and the latter gives a flow solution. In contrast, the theoretical background of the RWR scheme is different from that of the diffusion-reaction equation, although the restarting term of the RWR plays a role similar to the reaction term of the diffusion-reaction equation. The behaviors of the two approaches with respect to outliers reveal that they possess different attributes in terms of data propagation. This observation leads to the derivation of a new energy functional, where both volumetric heat capacity and thermal conductivity are considered together, and provides a common framework that unifies both the RW and the RWR approaches, in addition to other regularization methods. The proposed framework allows the RWR to be generalized (GRWR) in semilocal and nonlocal forms. The experimental results demonstrate the superiority of GRWR over existing regularization approaches in terms of depth map up-sampling and interactive image segmentation. Bumsub Ham, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2012 | Probabilistic Correspondence Matching using Random Walk with RestartabstractThis paper presents a probabilistic method for correspondence matching with a framework of the random walk with restart (RWR). The matching cost is reformulated as a corresponding probability, which enables the RWR to be utilized for matching the correspondences. There are mainly two advantages in our method. First, the proposed method guarantees the non-trivial steady-state solution of a given initial matching probability due to the restarting term in the RWR. It means the number of iteration, a crucial parameter which influences the performance of algorithm, is not needed in contrast to the conventional methods. This gives the consistent results regardless of the evolution time. Second, only an adjacent neighborhood is considered when the matching probabilities are inferred, which lowers the computational complexity while not sacrificing performance. Experimental results show that the performance of the proposed method is competitive to that of state-of-the-art methods both qualitatively and quantitatively. Changjae Oh, Bumsub Ham, Kwanghoon Sohn |
BMVC | 2 |
| 2012 | Robust Scale-Space Filter Using Second-Order Partial Differential EquationsabstractThis paper describes a robust scale-space filter that adaptively changes the amount of flux according to the local topology of the neighborhood. In a manner similar to modeling heat or temperature flow in physics, the robust scale-space filter is derived by coupling Fick's law with a generalized continuity equation in which the source or sink is modeled via a specific heat capacity. The filter plays an essential part in two aspects. First, an evolution step size is adaptively scaled according to the local structure, enabling the proposed filter to be numerically stable. Second, the influence of outliers is reduced by adaptively compensating for the incoming flux. We show that classical diffusion methods represent special cases of the proposed filter. By analyzing the stability condition of the proposed filter, we also verify that its evolution step size in an explicit scheme is larger than that of the diffusion methods. The proposed filter also satisfies the maximum principle in the same manner as the diffusion. Our experimental results show that the proposed filter is less sensitive to the evolution step size, as well as more robust to various outliers, such as Gaussian noise, impulsive noise, or a combination of the two. Bumsub Ham, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2011 | Hole filling with random walks using occlusion constraints in view synthesisabstractIn this paper, we propose a hole filling technique which coherently reconstructs the hole region during the view synthesis. The holes can be filled successfully in case that the virtual camera locates between real cameras by using interpolation. However, they cannot be handled in case that the virtual camera locates beyond the field of view of the real camera. We address this problem by jointly using image completion technique and random walks. First, occlusion constraint is imposed in order to guide the filling order. It is observed that the holes occur in a similar pattern because of the geometric characteristic of the camera configuration. This observation named vertical prior in this paper is also used to label each pixel on the fill front with foreground or background. Second, the probabilities estimated by random walks are utilized to find the patch candidates and to select the optimal patch. The experimental results show that the proposed method gives visually pleasing results over both interpolation and conventional image completion method. Sunghwan Choi, Bumsub Ham, Kwanghoon Sohn |
ICIP | 2 |
| 2011 | Cost aggregation with anisotropic diffusion in feature space for hybrid stereo matchingabstractIn this paper, we present a cost aggregation using anisotropic diffusion on a feature space for hybrid stereo matching. Stereo matching can be classified into two categories: feature-based and area-based approaches. Feature-based approaches generate accurate but sparse disparity maps. On the other hand, area-based approaches generate dense but unreliable disparity maps, especially at depth discontinuities and homogeneous regions. We hence propose a stereo matching algorithm having advantages of both approaches. We study how to design a correspondence algorithm without modeling any depth cues except disparity. A procedure of depth perception is modeled via anisotropic diffusion on the feature space in terms of coherence. Based on the assumption that similar local feature space has similar disparity, we define the feature space and its similarity and then introduce feature confidences into the proposed model. Experimental results show that the performance of the proposed method is comparable to that of the state-of-the-art methods. Bumsub Ham, Dongbo Min, Kwanghoon Sohn |
ICIP | 1 |
| 2010 | Visual fatigue evaluation and enhancement for 2D-plus-depth videoabstractA 3D video is expected to be a representative technique of realistic system but still has some problems such as visual fatigue and headache. In this paper, we propose a visual fatigue evaluation algorithm to predict the degree of visual fatigue from a 2D-plus-depth video. Spatial and temporal characteristics of the depth video are main factors of visual fatigue for autostereoscopic displays. Using depth image directly, we estimate spatial and temporal complexities, depth position and scene movement of the 3D video. Then, the overall visual fatigue of the 3D video is evaluated to have higher correlation with subjective fatigue evaluation by a linear regression. Moreover we control the pixel value of depth image from the 3D video which may induce severe fatigue to make more comfortable 3D video. The results of proposed algorithm show a considerable correlation with subjective visual fatigue. Jaeseob Choi, Donghyun Kim 0010, Bumsub Ham, Sunghwan Choi, Kwanghoon Sohn |
ICIP | 3 |
| 2009 | Spatial and temporal up-conversion technique for depth videoabstractThis paper proposes a novel framework for up-conversion of depth video resolution both in spatial and in time domain. Time-of-flight (TOF) sensors are widely used in computer vision fields. Although TOF sensors provide depth video in real time, there are some problems in a sense that it provides a low resolution and a low frame-rate depth video. We propose a cheaper solution that enhances depth video obtained by TOF sensor by combining it with CCD camera. The proposed method provides high quality video as a cheaper solution for low resolution, and low frame-rate depth video. It is useful when depth video is used in various applications such as 3DTV, free-view TV, teleconference system. High-quality depth video can be obtained by motion compensated frame interpolation (MCFI) and extended joint bilateral upsampling (JBU). Experimental results show that depth video obtained by the proposed method has satisfactory quality. Jinwook Choi, Dongbo Min, Bumsub Ham, Kwanghoon Sohn |
ICIP | 3 |
| 2009 | Virtual view rendering using super-resolution with multiview imagesabstractThis paper presents a new approach to solve the problem of quality degradation of a synthesized view, when a virtual camera moves forward. Interpolation techniques using only two neighboring views are generally applied when a virtual view is synthesized. Because the size of an object increases when the virtual camera moves forward, conventional methods have usually addressed this problem by interpolation techniques in order to synthesize a virtual view. However, as it generates a degraded view such as blurred images, we prevent a synthesized view from being blurred by using more images in multiview camera configuration. That is, this problem is solved by applying super-resolution concept which reconstructs a high resolution image from several low resolution images. Data fusion is performed by geometric warping with disparity maps of the multiple images followed by deblurring. Experimental results show that the image quality can further be improved by reducing blurring and halo effects in comparison with the interpolation method. Bumsub Ham, Dongbo Min, Jinwook Choi, Kwanghoon Sohn |
ICIP | 1 |