Dongbo Min

dblp:44/1149 · also Dong Bo Min · DBLP profile ↗
← Back
96ranked-venue papers
15as first author
30since 2021 · last 2026
0000-0003-4825-5240ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 78 · 13 first-author · 20 since 2021Artificial intelligence and machine learning · 45 · 3 first-author · 22 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Difficulty-Aware Label-Guided Denoising for Monocular 3D Object Detection
abstract
Monocular 3D object detection is a cost-effective solution for applications like autonomous driving and robotics, but remains fundamentally ill-posed due to inherently ambiguous depth cues. Recent DETR-based methods attempt to mitigate this through global attention and auxiliary depth prediction, yet they still struggle with inaccurate depth estimates. Moreover, these methods often overlook instance-level detection difficulty, such as occlusion, distance, and truncation, leading to suboptimal detection performance. We propose MonoDLGD, a novel Difficulty-Aware Label-Guided Denoising framework that adaptively perturbs and reconstructs ground-truth labels based on detection uncertainty. Specifically, MonoDLGD applies stronger perturbations to easier instances and weaker ones into harder cases, and then reconstructs them to effectively provide explicit geometric supervision. By jointly optimizing label reconstruction and 3D object detection, MonoDLGD encourages geometry-aware representation learning and improves robustness to varying levels of object complexity. Extensive experiments on the KITTI benchmark demonstrate that MonoDLGD achieves state-of-the-art performance across all difficulty levels.
Soyul Lee, Seungmin Baek, Dongbo Min
AAAI3
2026 AnyBald: Toward Realistic Diffusion-Based Hair Removal In-The-Wild
abstract
We present AnyBald, a novel framework for realistic hair removal from portrait images captured under diverse in-the-wild conditions. One of the key challenges in this task is the lack of high-quality paired data, as existing datasets are often low-quality, with limited viewpoint variation and diversity, making it difficult to handle real-world cases. To address this, we construct a scalable data augmentation pipeline that synthesizes high-quality hair and non-hair image pairs across diverse real-world scenarios, enabling effective generalization and scalable supervision. With this enriched dataset, we present a new hair removal framework that reformulates pretrained latent diffusion in-painting using learnable text prompts, removing the need for explicit masking at inference. In doing so, our model achieves natural hair removal with semantic preservation via implicit localization. To further enhance spatial precision, we introduce a regularization loss that guides the model to attend specifically to hair regions. Extensive experiments demonstrate that AnyBald outperforms in removing hair while preserving identity and semantics across various in-the-wild domains. Our project page is here: https://vision3d-lab.github.io/anybald/.
Yongjun Choi, Seungoh Han, Soomin Kim 0006, Sumin Son, Mohsen Rohani, Edgar Maucourant, Dongbo Min, Kyungdon Joo
WACV7
2025 TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task Learning
abstract
Transfer learning paradigm has driven substantial advancements in various vision tasks. However, as state-of-the-art models continue to grow, classical full fine-tuning often becomes computationally impractical, particularly in multi-task learning (MTL) setup where training complexity increases proportional to the number of tasks. Consequently, recent studies have explored Parameter-Efficient Fine-Tuning (PEFT) for MTL architectures. Despite some progress, these approaches still exhibit limitations in capturing fine-grained, task-specific features that are crucial to MTL. In this paper, we introduce Task-Adaptive Dynamic transFormer, termed TADFormer, a novel PEFT framework that performs task-aware feature adaptation in the fine-grained manner by dynamically considering task-specific input contexts. TADFormer proposes the parameter-efficient prompting for task adaptation and the Dynamic Task Filter (DTF) to capture task information conditioned on input contexts. Experiments on the PASCAL-Context benchmark demonstrate that the proposed method achieves higher accuracy in dense scene understanding tasks, while reducing the number of trainable parameters by up to 8.4 times when compared to full fine-tuning of MTL models. TADFormer also demonstrates superior parameter efficiency and accuracy compared to recent PEFT methods.
Seungmin Baek, Soyul Lee, Hayeon Jo, Hyesong Choi, Dongbo Min
CVPR5
2025 Hybrid-Tta: Continual Test-Time Adaptation Via Dynamic Domain Shift Detection
abstract
Continual Test Time Adaptation (CTTA) has emerged as a critical approach for bridging the domain gap between the controlled training environments and the real-world scenarios, enhancing model adaptability and robustness. Existing CTTA methods, typically categorized into Full-Tuning (FT) and Efficient-Tuning (ET), struggle with effectively addressing domain shifts. To overcome these challenges, we propose Hybrid-TTA, a holistic approach that dynamically selects instance-wise tuning method for optimal adaptation. Our approach introduces the Dynamic Domain Shift Detection (DDSD) strategy, which identifies domain shifts by leveraging temporal correlations in input sequences and dynamically switches between FT and ET to adapt to varying domain shifts effectively. Additionally, the Masked Image Modeling based Adaptation (MIMA) framework is integrated to ensure domain-agnostic robustness with minimal computational overhead. Our Hybrid-TTA achieves a notable 1.6%p improvement in mIoU on the Cityscapes-to-ACDC benchmark dataset, surpassing previous state-of-the-art methods and offering a robust solution for real-world continual adaptation challenges.
Hyewon Park, Jueun Ko, Dongbo Min
ICCV4
2025 RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo
abstract
Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense ground-truth labels. While recent Test-Time Adaptation (TTA) methods offer promising solutions, most rely on static target domain assumptions and input-invariant adaptation strategies, limiting their effectiveness under continual shifts. In this paper, we propose RobIA, a novel Robust, Instance-Aware framework for Continual Test-Time Adaptation (CTTA) in stereo depth estimation. RobIA integrates two key components: (1) Attend-and-Excite Mixture-of-Experts (AttEx-MoE), a parameter-efficient module that dynamically routes input to frozen experts via lightweight self-attention mechanism tailored to epipolar geometry, and (2) Robust AdaptBN Teacher, a PEFT-based teacher model that provides dense pseudo-supervision by complementing sparse handcrafted labels. This strategy enables input-specific flexibility, broad supervision coverage, improving generalization under domain shift. Extensive experiments demonstrate that RobIA achieves superior adaptation performance across dynamic target domains while maintaining computational efficiency.
Jueun Ko, Hyewon Park, Hyesong Choi, Dongbo Min
NeurIPS4
2024 Emerging Property of Masked Token for Effective Pre-training
Hyesong Choi, Hunsang Lee, Seyoung Joung, Hyejin Park 0003, Jiyeong Kim, Dongbo Min
ECCV (76)6
2024 Salience-Based Adaptive Masking: Revisiting Token Dynamics for Enhanced Pre-training
Hyesong Choi, Hyejin Park 0003, Kwang Moo Yi, Sungmin Cha, Dongbo Min
ECCV (78)5
2024 Dynamic Guidance Adversarial Distillation with Enhanced Teacher Knowledge
Hyejin Park 0003, Dongbo Min
ECCV (72)2
2024 Improving Self-Supervised Vision Transformers for Visual Control
abstract
Despite the tremendous success of vision transformer (ViT) architectures in a broad range of computer vision tasks, the potential of ViT for vision-based deep reinforcement learning (RL) has not been fully explored yet. To improve the performance of the ViT model in visual RL, we propose a simple yet effective approach for self-supervised learning by utilizing the structural capability of a single ViT model, which can learn multiple, distinct representations through extra learnable token embeddings. To this end, in addition to an RL token used for RL input, which corresponds to the classification token in computer vision, we introduce additional extra tokens that are tailored to two auxiliary self-supervised tasks specialized to learn visual and environmental dynamics representations. By interacting with embeddings of these extra tokens through self-attention, our approach provides additional learning signals to the ViT encoder, enabling it to learn more comprehensive representations that are beneficial to RL tasks. In experiments on benchmarks including the DeepMind Control Suite (DMControl) and Atari games, we demonstrate that the proposed approach outperforms the baselines that utilize ViT encoders, particularly achieving state-of-the-art performance in 4 out of 5 tasks in DMControl.
Wonil Song, Kwanghoon Sohn, Dongbo Min
ICIP3
2024 A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations
abstract
In the rapidly evolving domain of vision-based deep reinforcement learning (RL), a pivotal challenge is to achieve generalization capability to dynamic environmental changes reflected in visual observations. Our work delves into the intricacies of this problem, identifying two key issues that appear in previous approaches for visual RL generalization: (i) imbalanced saliency and (ii) observational overfitting. Imbalanced saliency is a phenomenon where an RL agent disproportionately identifies salient features across consecutive frames in a frame stack. Observational overfitting occurs when the agent focuses on certain background regions rather than task-relevant objects. To address these challenges, we present a simple yet effective framework for generalization in visual RL (SimGRL) under dynamic scene perturbations. First, to mitigate the imbalanced saliency problem, we introduce an architectural modification to the image encoder to stack frames at the feature level rather than the image level. Simultaneously, to alleviate the observational overfitting problem, we propose a novel technique called shifted random overlay augmentation, which is specifically designed to learn robust representations capable of effectively handling dynamic visual scenes. Extensive experiments demonstrate the superior generalization capability of SimGRL, achieving state-of-the-art performance in benchmarks including the DeepMind Control Suite.
Wonil Song, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
NeurIPS4
2024 Fine-Grained Background Representation for Weakly Supervised Semantic Segmentation
abstract
Generating reliable pseudo masks from image-level labels is challenging in the weakly supervised semantic segmentation (WSSS) task due to the lack of spatial information. Prevalent class activation map (CAM)-based solutions are challenged to discriminate the foreground (FG) objects from the suspicious background (BG) pixels (a.k.a. co-occurring) and learn the integral object regions. This paper proposes a simple fine-grained background representation (FBR) method to discover and represent diverse BG semantics and address the co-occurring problems. We abandon using the class prototype or pixel-level features for BG representation. Instead, we develop a novel primitive, negative region of interest (NROI), to capture the fine-grained BG semantic information and conduct the pixel-to-NROI contrast to distinguish the confusing BG pixels. We also present an active sampling strategy to mine the FG negatives on-the-fly, enabling efficient pixel-to-pixel intra-foreground contrastive learning to activate the entire object region. Thanks to the simplicity of design and convenience in use, our proposed method can be seamlessly plugged into various models, yielding new state-of-the-art results under various WSSS settings across benchmarks. Leveraging solely image-level (I) labels as supervision, our method achieves 73.2 mIoU and 45.6 mIoU segmentation results on Pascal Voc and MS COCO test sets, respectively. Furthermore, by incorporating saliency maps as an additional supervision signal (I+S), we attain 74.9 mIoU on Pascal Voc test set. Concurrently, our FBR approach demonstrates meaningful performance gains in weakly-supervised instance segmentation (WSIS) tasks, showcasing its robustness and strong generalization capabilities across diverse domains.
Xu Yin, Woobin Im, Dongbo Min, Yuchi Huo, Sung-Eui Yoon
IEEE Trans. Circuits Syst. Video Technol.3
2024 Revisiting Domain-Adaptive Semantic Segmentation via Knowledge Distillation
abstract
Numerous methods for unsupervised domain adaptation (UDA) have been proposed in semantic segmentation, achieving remarkable improvements. These methods are categorized into an adversarial learning-based approach that utilizes an additional discriminator and image translation model, and a self-supervised approach that uses a teacher model to generate pseudo labels. Among them, the self-supervised UDA approaches based on a self-training show excellent adaptability in semantic segmentation. However, erroneous estimates of the pseudo ground truths (PGTs) used in the self-training may often lead to inaccurate updates in the teacher model. Although several attempts have been made to address this issue, the teacher model updated through exponential moving average (EMA) still has a risk of propagating inaccuracies from the PGTs. Inspired by the fact that UDA shares similar principles with knowledge distillation (KD), we revisit the self-training based UDA approach from the perspective of KD and propose a novel UDA approach that employs two different teacher models. Specifically, we utilize both an EMA-updated teacher model to generate PGTs and a frozen teacher model pretrained with source data to transfer knowledge on a feature space. Since the frozen teacher model has no constraint on the model architecture unlike the EMA updated teacher model, we can effectively leverage a better representation power from the larger frozen teacher. Extensive experiments on various backbones (DeepLab-V2 [40] and DAFormer [73]) and scenarios (GTA5 → Cityscapes and SYNTHIA → Cityscapes) show that the proposed method improves segmentation performance in the target domain with its scalability. In particular, our method achieves comparable or better performance than state-of-the-arts even with a lightweight backbone.
Seongwon Jeong, Jiyeong Kim, Sungheui Kim, Dongbo Min
IEEE Trans. Image Process.4
2023 Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning
abstract
Recent vision-based reinforcement learning (RL) methods have found extracting high-level features from raw pixels with self-supervised learning to be effective in learning policies. However, these methods focus on learning global representations of images, and disregard local spatial structures present in the consecutively stacked frames. In this paper, we propose a novel approach, termed self-supervised Paired Similarity Representation Learning (PSRL) for effectively encoding spatial structures in an unsupervised manner. Given the input frames, the latent volumes are first generated individually using an encoder, and they are used to capture the variance in terms of local spatial structures, i.e., correspondence maps among multiple frames. This enables for providing plenty of fine-grained samples for training the encoder of deep RL. We further attempt to learn the global semantic representations in the action aware transform module that predicts future state representations using action vectors as a medium. The proposed method imposes similarity constraints on the three latent volumes; transformed query representations by estimated pixel-wise correspondence, predicted query representations from the action aware transform model, and target representations of future state, guiding action aware transform with locality-inherent volume. Experimental results on complex tasks in Atari Games and DeepMind Control Suite demonstrate that the RL methods are significantly boosted by the proposed self-supervised learning of paired similarity representations.
Hyesong Choi, Hunsang Lee, Wonil Song, Sangryul Jeon, Kwanghoon Sohn, Dongbo Min
CVPR6
2023 Environment Agnostic Representation for Visual Reinforcement learning
abstract
Generalization capability of vision-based deep reinforcement learning (RL) is indispensable to deal with dynamic environment changes that exist in visual observations. The high-dimensional space of the visual input, however, imposes challenges in adapting an agent to unseen environments. In this work, we propose Environment Agnostic Reinforcement learning (EAR), which is a compact framework for domain generalization of the visual deep RL. Environmentagnostic features (EAFs) are extracted by leveraging three novel objectives based on feature factorization, reconstruction, and episode-aware state shifting, so that policy learning is accomplished only with vital features. EAR is a simple single-stage method with a low model complexity and a fast inference time, ensuring a high reproducibility, while attaining state-of-the-art performance in the DeepMind Control Suite and DrawerWorld benchmarks. Code is available at: https://github.com/doihye/EAR.
Hyesong Choi, Hunsang Lee, Seongwon Jeong, Dongbo Min
ICCV4
2023 Learning disentangled skills for hierarchical reinforcement learning through trajectory autoencoder with weak labels
Wonil Song, Sangryul Jeon, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
Expert Syst. Appl.5
2023 Stereo Confidence Estimation via Locally Adaptive Fusion and Knowledge Distillation
abstract
Stereo confidence estimation aims to estimate the reliability of the estimated disparity by stereo matching. Different from the previous methods that exploit the limited input modality, we present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. Moreover, we propose a knowledge distillation framework to learn more compact confidence estimation networks as student networks. By transferring the knowledge from LAF-Net as teacher networks, the student networks that solely take as input a disparity can achieve comparable performance. To transfer more informative knowledge, we also propose a module to learn the locally-varying temperature in a softmax function. We further extend this framework to a multiview scenario. Experimental results show that LAF-Net and its variations outperform the state-of-the-art stereo confidence methods on various benchmarks.
Sunok Kim, Seungryong Kim, Dongbo Min, Pascal Frossard, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Contour-Aware Equipotential Learning for Semantic Segmentation
abstract
With increasing demands for high-quality semantic segmentation in the industry, hard-distinguishing semantic boundaries have posed a significant threat to existing solutions. Inspired by real-life experience, i.e., combining varied observations contributes to higher visual recognition confidence, we present the equipotential learning (EPL) method. This novel module transfers the predicted/ground-truth semantic labels to a self-defined potential domain to learn and infer decision boundaries along customized directions. The conversion to the potential domain is implemented via a lightweight differentiable anisotropic convolution without incurring any parameter overhead. Besides, the designed two loss functions, the point loss and the equipotential line loss implement anisotropic field regression and category-level contour learning, respectively, enhancing prediction consistencies in the inter/intra-class boundary areas. More importantly, EPL is agnostic to network architectures, and thus it can be plugged into most existing segmentation models. This paper is the first attempt to address the boundary segmentation problem with field regression and contour learning. Meaningful performance improvements on Pascal Voc 2012 and Cityscapes demonstrate that the proposed EPL module can benefit the off-the-shelf fully convolutional network models when recognizing semantic boundary areas. Besides, intensive comparisons and analysis show the favorable merits of EPL for distinguishing semantically-similar and irregular-shaped categories.
Xu Yin, Dongbo Min, Yuchi Huo, Sung-Eui Yoon
IEEE Trans. Multim.2
2022 Pin the Memory: Learning to Generalize Semantic Segmentation
abstract
The rise of deep neural networks has led to several break-throughs for semantic segmentation. In spite of this, a model trained on source domain often fails to work properly in new challenging domains, that is directly concerned with the generalization capability of the model. In this paper, we present a novel memory-guided domain generalization method for semantic segmentation based on meta-learning framework. Especially, our method abstracts the conceptual knowledge of semantic classes into categorical memory which is constant beyond the domains. Upon the meta-learning concept, we repeatedly train memory-guided networks and simulate virtual test to 1) learn how to memorize a domain-agnostic and distinct information of classes and 2) offer an externally settled memory as a class-guidance to reduce the ambiguity of representation in the test data of arbitrary unseen domain. To this end, we also propose memory divergence and feature cohesion losses, which encourage to learn memory reading and update processes for category-aware domain generalization. Extensive experiments for semantic segmentation demonstrate the superior generalization capability of our method over state-of-the-art works on various benchmarks.11https://github.com/Genie-Kim/PintheMemory
Jin Kim 0005, Jiyoung Lee 0005, Jungin Park, Dongbo Min, Kwanghoon Sohn
CVPR4
2022 KNN Local Attention for Image Restoration
abstract
Recent works attempt to integrate the non-local operation with CNNs or Transformer, achieving remarkable performance in image restoration tasks. The global similarity, however, has the problems of the lack of locality and the high computational complexity that is quadratic to an input resolution. The local attention mechanism alleviates these issues by introducing the inductive bias of the locality with convolution-like operators. However, by focusing only on adjacent positions, the local attention suffers from an insufficient receptive field for image restoration. In this paper, we propose a new attention mechanism for image restoration, called k-NN Image Transformer (KiT), that rectifies the above mentioned limitations. Specifically, the KiT groups k-nearest neighbor patches with locality sensitive hashing (LSH), and the grouped patches are aggregated into each query patch by performing a pair-wise local attention. In this way, the pair-wise operation establishes nonlocal connectivity while maintaining the desired properties of the local attention, i.e., inductive bias of locality and linear complexity to input resolution. The proposed method outperforms state-of-the-art restoration approaches on image denoising, deblurring and deraining benchmarks. The code will be available soon.
Hunsang Lee, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
CVPR4
2022 PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation
Kwonyoung Kim, Jungin Park, Jiyoung Lee 0005, Dongbo Min, Kwanghoon Sohn
ECCV (38)4
2022 Sequential Cross Attention Based Multi-Task Learning
abstract
In multi-task learning (MTL) for visual scene understanding, it is crucial to transfer useful information between multiple tasks with minimal interferences. In this paper, we propose a novel architecture that effectively transfers informative features by applying the attention mechanism to the multi-scale features of the tasks. Since applying the attention module directly to all possible features in terms of scale and task requires a high complexity, we propose to apply the attention module sequentially for the task and scale. The cross-task attention module (CTAM) is first applied to facilitate the exchange of relevant information between the multiple task features of the same scale. The cross-scale attention module (CSAM) then aggregates useful information from feature maps at different resolutions in the same task. Also, we attempt to capture long range dependencies through the self-attention module in the feature extraction network. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the NYUD-v2 and PASCAL-Context dataset. Our code is available at https://github.com/kimsunkyung/SCA-MTL
Sunkyung Kim, Hyesong Choi, Dongbo Min
ICIP3
2022 Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence
abstract
Existing pipelines of semantic correspondence commonly include extracting high-level semantic features for the invariance against intra-class variations and background clutters. This architecture, however, inevitably results in a low-resolution matching field that additionally requires an ad-hoc interpolation process as a post-processing for converting it into a high-resolution one, certainly limiting the overall performance of matching results. To overcome this, inspired by recent success of implicit neural representation, we present a novel method for semantic correspondence, called Neural Matching Field (NeMF). However, complicacy and high-dimensionality of a 4D matching field are the major hindrances, which we propose a cost embedding network to process a coarse cost volume to use as a guidance for establishing high-precision matching field through the following fully-connected network. Nevertheless, learning a high-dimensional matching field remains challenging mainly due to computational complexity, since a na\"ive exhaustive inference would require querying from all pixels in the 4D space to infer pixel-wise correspondences. To overcome this, we propose adequate training and inference procedures, which in the training phase, we randomly sample matching candidates and in the inference phase, we iteratively performs PatchMatch-based inference and coordinate optimization at test time. With these combined, competitive results are attained on several standard benchmarks for semantic correspondence. Code and pre-trained weights are available at~\url{https://ku-cvlab.github.io/NeMF/}.
Sunghwan Hong, Jisu Nam, Seokju Cho, Susung Hong, Sangryul Jeon, Dongbo Min, Seungryong Kim
NeurIPS6
2022 Pyramidal Semantic Correspondence Networks
abstract
This paper presents a deep architecture, called pyramidal semantic correspondence networks (PSCNet), that estimates locally-varying affine transformation fields across semantically similar images. To deal with large appearance and shape variations that commonly exist among different instances within the same object category, we leverage a pyramidal model where the affine transformation fields are progressively estimated in a coarse-to-fine manner so that the smoothness constraint is naturally imposed. Different from the previous methods which directly estimate global or local deformations, our method first starts to estimate the transformation from an entire image and then progressively increases the degree of freedom of the transformation by dividing coarse cell into finer ones. To this end, we propose two spatial pyramid models by dividing an image in a form of quad-tree rectangles or into multiple semantic elements of an object. Additionally, to overcome the limitation of insufficient training data, a novel weakly-supervised training scheme is introduced that generates progressively evolving supervisions through the spatial pyramid models by leveraging a correspondence consistency across image pairs. Extensive experimental results on various benchmarks including TSS, Proposal Flow-WILLOW, Proposal Flow-PASCAL, Caltech-101, and SPair-71k demonstrate that the proposed method outperforms the lastest methods for dense semantic correspondence.
Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 On the Confidence of Stereo Matching in a Deep-Learning Era: A Quantitative Evaluation
abstract
Stereo matching is one of the most popular techniques to estimate dense depth maps by finding the disparity between matching pixels on two, synchronized and rectified images. Alongside with the development of more accurate algorithms, the research community focused on finding good strategies to estimate the reliability, i.e., the confidence, of estimated disparity maps. This information proves to be a powerful cue to naively find wrong matches as well as to improve the overall effectiveness of a variety of stereo algorithms according to different strategies. In this paper, we review more than ten years of developments in the field of confidence estimation for stereo matching. We extensively discuss and evaluate existing confidence measures and their variants, from hand-crafted ones to the most recent, state-of-the-art learning based methods. We study the different behaviors of each measure when applied to a pool of different stereo algorithms and, for the first time in literature, when paired with a state-of-the-art deep stereo network. Our experiments, carried out on five different standard datasets, provide a comprehensive overview of the field, highlighting in particular both strengths and limitations of learning-based strategies.
Matteo Poggi, Seungryong Kim, Fabio Tosi, Sunok Kim, Filippo Aleotti, Dongbo Min, Kwanghoon Sohn, Stefano Mattoccia
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Mining Better Samples for Contrastive Learning of Temporal Correspondence
abstract
We present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appropriately adjusting their hardness during training. This allows us to suppress the adverse impact of ambiguous matches and prevent a trivial solution from being yielded by too hard or too easy negative samples. To accomplish this, we incorporate three different criteria that ranges from a pixel-level matching confidence to a video-level one into a bottom-up pipeline, and plan a curriculum that is aware of current representation power for the adaptive hardness of negative samples during training. With the proposed method, state-of-the-art performance is attained over the latest approaches on several video label propagation tasks.
Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
CVPR2
2021 Adaptive confidence thresholding for monocular depth estimation
abstract
Self-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to leverage pseudo ground truth depth maps of stereo images generated from self-supervised stereo matching methods. The confidence map of the pseudo ground truth depth map is estimated to mitigate performance degeneration by inaccurate pseudo depth maps. To cope with the prediction error of the confidence map itself, we also leverage the threshold network that learns the threshold dynamically conditioned on the pseudo depth maps. The pseudo depth labels filtered out by the thresholded confidence map are used to supervise the monocular depth network. Furthermore, we propose the probabilistic framework that refines the monocular depth map with the help of its uncertainty map through the pixel-adaptive convolution (PAC) layer. Experimental results demonstrate superior performance to state-of-the-art monocular depth estimation methods. Lastly, we exhibit that the proposed threshold learning can also be used to improve the performance of existing confidence estimation approaches.
Hyesong Choi, Hunsang Lee, Sunkyung Kim, Sunok Kim, Seungryong Kim, Kwanghoon Sohn, Dongbo Min
ICCV7
2021 Self-Balanced Learning for Domain Generalization
abstract
Domain generalization aims to learn a prediction model on multi-domain source data such that the model can generalize to a target domain with unknown statistics. Most existing approaches have been developed under the assumption that the source data is well-balanced in terms of both domain and class. However, real-world training data collected with different composition biases often exhibits severe distribution gaps for domain and class, leading to substantial performance degradation. In this paper, we propose a self-balanced domain generalization framework that adaptively learns the weights of losses to alleviate the bias caused by different distributions of the multi-domain source data. The self-balanced scheme is based on an auxiliary reweighting network that iteratively updates the weight of loss conditioned on the domain and class information by leveraging balanced meta data. Experimental results demonstrate the effectiveness of our method overwhelming state-of-the-art works for domain generalization.
Jin Kim 0005, Jiyoung Lee 0005, Jungin Park, Dongbo Min, Kwanghoon Sohn
ICIP4
2021 Deep monocular depth estimation leveraging a large-scale outdoor stereo dataset
Jaehoon Cho, Dongbo Min, Youngjung Kim, Kwanghoon Sohn
Expert Syst. Appl.2
2021 Dense Cross-Modal Correspondence Estimation With the Deep Self-Correlation Descriptor
abstract
We present the deep self-correlation (DSC) descriptor for establishing dense correspondences between images taken under different imaging modalities, such as different spectral ranges or lighting conditions. We encode local self-similar structure in a pyramidal manner that yields both more precise localization ability and greater robustness to non-rigid image deformations. Specifically, DSC first computes multiple self-correlation surfaces with randomly sampled patches over a local support window, and then builds pyramidal self-correlation surfaces through average pooling on the surfaces. The feature responses on the self-correlation surfaces are then encoded through spatial pyramid pooling in a log-polar configuration. To better handle geometric variations such as scale and rotation, we additionally propose the geometry-invariant DSC (GI-DSC) that leverages multi-scale self-correlation computation and canonical orientation estimation. In contrast to descriptors based on deep convolutional neural networks (CNNs), DSC and GI-DSC are training-free (i.e., handcrafted descriptors), are robust to cross-modality, and generalize well to various modality variations. Extensive experiments demonstrate the state-of-the-art performance of DSC and GI-DSC on challenging cases of cross-modal image pairs having photometric and/or geometric variations.
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Adversarial Confidence Estimation Networks for Robust Stereo Matching
abstract
Stereo matching aiming to perceive the 3-D geometry of a scene facilitates numerous computer vision tasks used in advanced driver assistance systems (ADAS). Although numerous methods have been proposed for this task by leveraging deep convolutional neural networks (CNNs), stereo matching still remains an unsolved problem due to its inherent matching ambiguities. To overcome these limitations, we present a method for jointly estimating disparity and confidence from stereo image pairs through deep networks. We accomplish this through a minmax optimization to learn the generative cost aggregation networks and discriminative confidence estimation networks in an adversarial manner. Concretely, the generative cost aggregation networks are trained to accurately generate disparities at both confident and unconfident pixels from an input matching cost that are indistinguishable by the discriminative confidence estimation networks, while the discriminative confidence estimation networks are trained to distinguish the confident and unconfident disparities. In addition, to fully exploit complementary information of matching cost, disparity, and color image in confidence estimation, we present a dynamic fusion module. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks including real driving scenes.
Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Intell. Transp. Syst.2
2020 Guided Semantic Flow
Sangryul Jeon, Dongbo Min, Seungryong Kim, Jihwan Choe, Kwanghoon Sohn
ECCV (28)2
2020 Discrete-Continuous Transformation Matching for Dense Semantic Correspondence
abstract
Techniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Furthermore, leveraging correspondence consistency and confidence-guided filtering in each iteration facilitates the convergence of our method. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks and applications.
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Unsupervised Low-Light Image Enhancement Using Bright Channel Prior
abstract
Recent approaches for low-light image enhancement achieve excellent performance through supervised learning based on convolutional neural networks. However, it is still challenging to collect a large amount of low-/normal-light image pairs in real environments for training the networks. In this letter, we propose an unsupervised learning approach for single low-light image enhancement using the bright channel prior (BCP) that the brightest pixel in a small patch is likely to be close to 1. An unsupervised loss function is defined with the pseudo ground-truth generated using the BCP. An enhancement network, consisting of a simple encoder-decoder, is then trained using the unsupervised loss function. To the best of our knowledge, this is the first attempt that enhances a low-light image through unsupervised learning. Furthermore, we introduce saturation loss and self-attention map for preserving image details and naturalness in the enhanced result. The performance of the proposed method is validated on various public datasets. Experimental results demonstrate that the proposed unsupervised approach achieves competitive performance over state-of-the-art methods based on supervised learning.
Hunsang Lee, Kwanghoon Sohn, Dongbo Min
IEEE Signal Process. Lett.3
2020 Single Image Deraining Using Time-Lapse Data
abstract
Leveraging on recent advances in deep convolutional neural networks (CNNs), single image deraining has been studied as a learning task, achieving an outstanding performance over traditional hand-designed approaches. Current CNNs based deraining approaches adopt the supervised learning framework that uses a massive training data generated with synthetic rain streaks, having a limited generalization ability on real rainy images. To address this problem, we propose a novel learning framework for single image deraining that leverages time-lapse sequences instead of the synthetic image pairs. The deraining networks are trained using the time-lapse sequences in which both camera and scenes are static except for time-varying rain streaks. Specifically, we formulate a background consistency loss such that the deraining networks consistently generate the same derained images from the time-lapse sequences. We additionally introduce two loss functions, the structure similarity loss that encourages the derained image to be similar with an input rainy image and the directional gradient loss using the assumption that the estimated rain streaks are likely to be sparse and have dominant directions. To consider various rain conditions, we leverage a dynamic fusion module that effectively fuses multi-scale features. We also build a novel large-scale time-lapse dataset providing real world rainy images containing various rain conditions. Experiments demonstrate that the proposed method outperforms state-of-the-art techniques on synthetic and real rainy images both qualitatively and quantitatively. On the high-level vision tasks under severe rainy conditions, it has been shown that the proposed method can be utilized as a pre-preprocessing step for subsequent tasks.
Jaehoon Cho, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.3
2020 Learning Deeply Aggregated Alternating Minimization for General Inverse Problems
abstract
Regularization-based image restoration is one of the most powerful tools in image processing and computer vision thanks to its flexibility for handling various inverse problems. However, designing an optimal regularization function still remains unsolved since natural images and related scene types have a complex structure. In this paper, we present a general and principled framework, called deeply aggregated alternating minimization (DeepAM). We design a convolutional neural network (CNN) to implicitly parameterize the regularizer of the alternating minimization (AM) algorithm. Contrary to the conventional AM algorithm based on a point-wise proximal mapping, the DeepAM projects intermediate estimate into a set of natural images via deep aggregation. Since the CNN is fully integrated into the AM procedure, all parameters can be jointly optimized through end-to-end training. These properties enable the DeepAM to converge with a small number of iterations, while maintaining an algorithmic simplicity. We show that the DeepAM outperforms state-of-the-art methods, including nonlocal-based methods, Plug-and-Play regularization, and recent data-driven approaches. The effectiveness of our framework is demonstrated in a variety of image restoration tasks: Guassian denoising, deraining, deblurring, super-resolution, color-guided depth upsampling, and RGB/NIR restoration.
Hyungjoo Jung, Youngjung Kim, Dongbo Min, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
IEEE Trans. Image Process.3
2019 LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence Estimation
abstract
We present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. The attention inference networks encode the importance of tri-modal confidence features and then concatenate them using the attention maps in an adaptive and dynamic fashion. This enables us to make an optimal fusion of the heterogeneous features, compared to a simple concatenation technique that is commonly used in conventional approaches. In addition, to encode the confidence features with locally-varying receptive fields, the scale inference networks learn the scale map and warp the fused confidence features through convolutional spatial transformer networks. Finally, the confidence map is progressively estimated in the recursive refinement networks to enforce a spatial context and local consistency. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks.
Sunok Kim, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
CVPR3
2019 Semantic Attribute Matching Networks
abstract
We present semantic attribute matching networks (SAM-Net) for jointly establishing correspondences and transferring attributes across semantically similar images, which intelligently weaves the advantages of the two tasks while overcoming their limitations. SAM-Net accomplishes this through an iterative process of establishing reliable correspondences by reducing the attribute discrepancy between the images and synthesizing attribute transferred images using the learned correspondences. To learn the networks using weak supervisions in the form of image pairs, we present a semantic attribute matching loss based on the matching similarity between an attribute transferred source feature and a warped target feature. With SAM-Net, the state-of-the-art performance is attained on several benchmarks for semantic matching and attribute transfer.
Seungryong Kim, Dongbo Min, Somi Jeong, Sunok Kim, Sangryul Jeon, Kwanghoon Sohn
CVPR2
2019 Joint Learning of Semantic Alignment and Object Landmark Detection
abstract
Convolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In this paper, we present a joint learning approach for obtaining dense correspondences and discovering object landmarks from semantically similar images. Based on the key insight that the two tasks can mutually provide supervisions to each other, our networks accomplish this through a joint loss function that alternatively imposes a consistency constraint between the two tasks, thereby boosting the performance and addressing the lack of training data in a principled manner. To the best of our knowledge, this is the first attempt to address the lack of training data for the two tasks through the joint learning. To further improve the robustness of our framework, we introduce a probabilistic learning formulation that allows only reliable matches to be used in the joint learning process. With the proposed method, state-of-the-art performance is attained on several benchmarks for semantic matching and landmark detection.
Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
ICCV2
2019 FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence
abstract
We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. Unlike traditional dense correspondence approaches for estimating depth or optical flow, semantic correspondence estimation poses additional challenges due to intra-class appearance and shape variations among different instances within the same object or scene category. To robustly match points across semantically similar images, we formulate FCSS using local self-similarity (LSS), which is inherently insensitive to intra-class appearance variations. LSS is incorporated through a proposed convolutional self-similarity (CSS) layer, where the sampling patterns and the self-similarity measure are jointly learned in an end-to-end and multi-scale manner. Furthermore, to address shape variations among different object instances, we propose a convolutional affine transformer (CAT) layer that estimates explicit affine transformation fields at each pixel to transform the sampling patterns and corresponding receptive fields. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in most existing datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS significantly outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks.
Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin 0001, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Unified Confidence Estimation Networks for Robust Stereo Matching
abstract
We present a deep architecture that estimates a stereo confidence, which is essential for improving the accuracy of stereo matching algorithms. In contrast to existing methods based on deep convolutional neural networks (CNNs) that rely on only one of the matching cost volume or estimated disparity map, our network estimates the stereo confidence by using the two heterogeneous inputs simultaneously. Specifically, the matching probability volume is first computed from the matching cost volume with residual networks and a pooling module in a manner that yields greater robustness. The confidence is then estimated through a unified deep network that combines confidence features extracted both from the matching probability volume and its corresponding disparity. In addition, our method extracts the confidence features of the disparity map by applying multiple convolutional filters with varying sizes to an input disparity map. To learn our networks in a semi-supervised manner, we propose a novel loss function that use confident points to compute the image reconstruction loss. To validate the effectiveness of our method in a disparity post-processing step, we employ three post-processing approaches; cost modulation, ground control points-based propagation, and aggregated ground control points-based propagation. Experimental results demonstrate that our method outperforms state-of-the-art confidence estimation methods on various benchmarks.
Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Image Process.2
2018 PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
ECCV (6)3
2018 Recurrent Transformer Networks for Semantic Correspondence
abstract
We present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convolutional activations. By directly estimating the transformations between an image pair, rather than employing spatial transformer networks to independently normalize each individual image, we show that greater accuracy can be achieved. This process is conducted in a recursive manner to refine both the transformation estimates and the feature representations. In addition, a technique is presented for weakly-supervised training of RTNs that is based on a proposed classification loss. With RTNs, state-of-the-art performance is attained on several benchmarks for semantic correspondence.
Seungryong Kim, Stephen Lin 0001, Sangryul Jeon, Dongbo Min, Kwanghoon Sohn
NeurIPS4
2018 Deep Monocular Depth Estimation via Integration of Global and Local Predictions
abstract
Recent works on machine learning have greatly advanced the accuracy of single image depth estimation. However, the resulting depth images are still over-smoothed and perceptually unsatisfying. This paper casts depth prediction from single image as a parametric learning problem. Specifically, we propose a deep variational model that effectively integrates heterogeneous predictions from two convolutional neural networks (CNNs), named global and local networks. They have contrasting network architecture and are designed to capture depth information with complementary attributes. These intermediate outputs are then combined in the integration network based on the variational framework. By unrolling the optimization steps of Split Bregman (SB) iterations in the integration network, our model can be trained in an end-to-end manner. This enables one to simultaneously learn an efficient parameterization of the CNNs and hyper-parameter in the variational method. Finally, we offer a new dataset of 0.22 million RGB-D images captured by Microsoft Kinect v2. Our model generates realistic and discontinuity-preserving depth prediction without involving any low-level segmentation or superpixels. Intensive experiments demonstrate the superiority of the proposed method in a range of RGB-D benchmarks including both indoor and outdoor scenarios.
Youngjung Kim, Hyungjoo Jung, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.3
2018 Fast 2D Complex Gabor Filter With Kernel Decomposition
abstract
2D complex Gabor filtering has found numerous applications in the fields of computer vision and image processing. Especially, in some applications, it is often needed to compute 2D complex Gabor filter bank consisting of filtering outputs at multiple orientations and frequencies. Although several approaches for fast Gabor filtering have been proposed, they focus primarily on reducing the runtime for performing filtering once at specific orientation and frequency. To obtain the Gabor filter bank, the existing methods are repeatedly applied with respect to multiple orientations and frequencies. In this paper, we propose a novel approach that efficiently computes the 2D complex Gabor filter bank by reducing the computational redundancy that arises when performing filtering at multiple orientations and frequencies. The proposed method first decomposes the Gabor kernel to allow a fast convolution with the Gaussian kernel in a separable manner. This enables reducing the runtime of the Gabor filter bank by reusing intermediate results computed at a specific orientation. By extending this idea, we also propose a fast approach for 2D localized sliding discrete Fourier transform that uses the Gaussian kernel in order to lend spatial localization ability as in the Gabor filter. Experimental results demonstrate that the proposed method runs faster than the state-of-the-art methods, while maintaining similar filtering quality.
Suhyuk Um, Dongbo Min
IEEE Trans. Image Process.3
2017 Deeply Aggregated Alternating Minimization for Image Restoration
abstract
Regularization-based image restoration has remained an active research topic in image processing and computer vision. It often leverages a guidance signal captured in different-fields as an additional cue. In this work, we present a general framework for image restoration, called deeply aggregated alternating minimization (DeepAM). We propose to train deep neural network to advance two of the steps in the conventional AM algorithm: proximal mapping and β-continuation. Both steps are learned from a large dataset in an end-to-end manner. The proposed framework enables the convolutional neural networks (CNNs) to operate as a regularizer in the AM algorithm. We show that our learned regularizer via deep aggregation outperforms the recent data-driven approaches as well as the nonlocal-based methods. The flexibility and effectiveness of our framework are demonstrated in several restoration tasks, including single image denoising, RGB-NIR restoration, and depth superresolution.
Youngjung Kim, Hyungjoo Jung, Dongbo Min, Kwanghoon Sohn
CVPR3
2017 FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence
abstract
We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to existing CNN-based descriptors, FCSS is inherently insensitive to intra-class appearance variations because of its LSS-based structure, while maintaining the precise localization ability of deep neural networks. The sampling patterns of local structure and the self-similarity measure are jointly learned within the proposed network in an end-to-end and multi-scale manner. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in existing image datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks.
Seungryong Kim, Dongbo Min, Bumsub Ham, Sangryul Jeon, Stephen Lin 0001, Kwanghoon Sohn
CVPR2
2017 DCTM: Discrete-Continuous Transformation Matching for Semantic Flow
abstract
Techniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks.
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
ICCV2
2017 Depth prediction from a single image with conditional adversarial networks
abstract
Recent works on machine learning have greatly advanced the accuracy of depth estimation from a single image. However, resulting depth images are still visually unsatisfactory, often producing poor boundary localization and spurious regions. In this paper, we formulate this problem from single images as a deep adversarial learning framework. A two-stage convolutional network is designed as a generator to sequentially predict global and local structures of the depth image. At the heart of our approach is a training criterion based on adversarial discriminator which attempts to distinguish between real and generated depth images as accurately as possible. Our model enables more realistic and structure-preserving depth prediction from a single image, compared to state-of-the-arts approaches. An experimental comparison demonstrates the effectiveness of our approach on large RGB-D dataset.
Hyungjoo Jung, Youngjung Kim, Dongbo Min, Changjae Oh, Kwanghoon Sohn
ICIP3
2017 Deep stereo confidence prediction for depth estimation
abstract
We present a novel method that predicts a confidence to improve the accuracy of an estimated depth map in stereo matching. In contrast to existing learning based approaches relying on hand-crafted confidence features, we cast this problem into a convolutional neural network, learned using both a matching cost volume and its associated disparity map. As the size of the matching cost volume varies depending on a search range of stereo image pairs, we propose to use a top-K matching probability volume layer so that an input size for convolutional layers remains unchanged. Experimental results demonstrate that the proposed method outperforms the state-of-the-art confidence estimation approaches on various benchmarks.
Sunok Kim, Dongbo Min, Bumsub Ham, Seungryong Kim, Kwanghoon Sohn
ICIP2
2017 DASC: Robust Dense Descriptor for Multi-Modal and Multi-Spectral Correspondence Estimation
abstract
Establishing dense correspondences between multiple images is a fundamental task in many applications. However, finding a reliable correspondence between multi-modal or multi-spectral images still remains unsolved due to their challenging photometric and geometric variations. In this paper, we propose a novel dense descriptor, called dense adaptive self-correlation (DASC), to estimate dense multi-modal and multi-spectral correspondences. Based on an observation that self-similarity existing within images is robust to imaging modality variations, we define the descriptor with a series of an adaptive self-correlation similarity measure between patches sampled by a randomized receptive field pooling, in which a sampling pattern is obtained using a discriminative learning. The computational redundancy of dense descriptors is dramatically reduced by applying fast edge-aware filtering. Furthermore, in order to address geometric variations including scale and rotation, we propose a geometry-invariant DASC (GI-DASC) descriptor that effectively leverages the DASC through a superpixel-based representation. For a quantitative evaluation of the GI-DASC, we build a novel multi-modal benchmark as varying photometric and geometric conditions. Experimental results demonstrate the outstanding performance of the DASC and GI-DASC in many cases of dense multi-modal and multi-spectral correspondences.
Seungryong Kim, Dongbo Min, Bumsub Ham, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 PatchMatch Filter: Edge-Aware Filtering Meets Randomized Search for Visual Correspondence
abstract
Though many tasks in computer vision can be formulated elegantly as pixel-labeling problems, a typical challenge discouraging such a discrete formulation is often due to computational efficiency. Recent studies on fast cost volume filtering based on efficient edge-aware filters provide a fast alternative to solve discrete labeling problems, with the complexity independent of the support window size. However, these methods still have to step through the entire cost volume exhaustively, which makes the solution speed scale linearly with the label space size. When the label space is huge or even infinite, which is often the case for (subpixel-accurate) stereo and optical flow estimation, their computational complexity becomes quickly unacceptable. Developed to search approximate nearest neighbors rapidly, the PatchMatch method can significantly reduce the complexity dependency on the search space size. But, its pixel-wise randomized search and fragmented data access within the 3D cost volume seriously hinder the application of efficient cost slice filtering. This paper presents a generic and fast computational framework for general multi-labeling problems called PatchMatch Filter (PMF). We explore effective and efficient strategies to weave together these two fundamental techniques developed in isolation, i.e., PatchMatch-based randomized search and efficient edge-aware image filtering. By decompositing an image into compact superpixels, we also propose superpixel-based novel search strategies that generalize and improve the original PatchMatch method. Further motivated to improve the regularization strength, we propose a simple yet effective cross-scale consistency constraint, which handles labeling estimation for large low-textured regions more reliably than a single-scale PMF algorithm. Focusing on dense correspondence field estimation in this paper, we demonstrate PMF's applications in stereo and optical flow. Our PMF methods achieve top-tier correspondence accuracy but run much faster than other related competing methods, often giving over 10-100 times speedup.
Jiangbo Lu, Yu Li 0003, Hongsheng Yang, Dongbo Min, Wei Yong Eng, Minh N. Do
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Cross-Scale Cost Aggregation for Stereo Matching
abstract
This paper proposes a generic framework that enables a multiscale interaction in the cost aggregation step of stereo matching algorithms. Inspired by the formulation of image filters, we first reformulate cost aggregation from a weighted least-squares (WLS) optimization perspective and show that different cost aggregation methods essentially differ in the choices of similarity kernels. Our key motivation is that while the human stereo vision system processes information at both coarse and fine scales interactively for the correspondence search, state-of-the-art approaches aggregate costs at the finest scale of the input stereo images only, ignoring inter-consistency across multiple scales. This motivation leads us to introduce an inter-scale regularizer into the WLS optimization objective to enforce the consistency of the cost volume among the neighboring scales. The new optimization objective with the inter-scale regularization is convex, and thus, it is easily and analytically solved. Minimizing this new objective leads to the proposed framework. Since the regularization term is independent of the similarity kernel, various cost aggregation approaches, including discrete and continuous parameterization methods, can be easily integrated into the proposed framework. We show that the cross-scale framework is important as it effectively and efficiently expands state-of-the-art cost aggregation methods and leads to significant improvements, when evaluated on Middlebury, Middlebury Third, KITTI, and New Tsukuba data sets.
Kang Zhang 0004, Yuqiang Fang, Dongbo Min, Lifeng Sun, Shiqiang Yang, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.3
2017 Fast Domain Decomposition for Global Image Smoothing
abstract
Edge-preserving smoothing (EPS) can be formulated as minimizing an objective function that consists of data and regularization terms. At the price of high-computational cost, this global EPS approach is more robust and versatile than a local one that typically has a form of weighted averaging. In this paper, we introduce an efficient decomposition-based method for global EPS that minimizes the objective function of L2 data and (possibly non-smooth and non-convex) regularization terms in linear time. Different from previous decompositionbased methods, which require solving a large linear system, our approach solves an equivalent constrained optimization problem, resulting in a sequence of 1-D sub-problems. This enables applying fast linear time solver for weighted-least squares and -L1 smoothing problems. An alternating direction method of multipliers algorithm is adopted to guarantee fast convergence. Our method is fully parallelizable, and its runtime is even comparable to the state-of-the-art local EPS approaches. We also propose a family of fast majorization-minimization algorithms that minimize an objective with non-convex regularization terms. Experimental results demonstrate the effectiveness and flexibility of our approach in a range of image processing and computational photography applications.
Youngjung Kim, Dongbo Min, Bumsub Ham, Kwanghoon Sohn
IEEE Trans. Image Process.2
2017 Feature Augmentation for Learning Confidence Measure in Stereo Matching
abstract
Confidence estimation is essential for refining stereo matching results through a post-processing step. This problem has recently been studied using a learning-based approach, which demonstrates a substantial improvement on conventional simple non-learning based methods. However, the formulation of learning-based methods that individually estimates the confidence of each pixel disregards spatial coherency that might exist in the confidence map, thus providing a limited performance under challenging conditions. Our key observation is that the confidence features and resulting confidence maps are smoothly varying in the spatial domain, and highly correlated within the local regions of an image. We present a new approach that imposes spatial consistency on the confidence estimation. Specifically, a set of robust confidence features is extracted from each superpixel decomposed using the Gaussian mixture model, and then these features are concatenated with pixel-level confidence features. The features are then enhanced through adaptive filtering in the feature domain. In addition, the resulting confidence map, estimated using the confidence features with a random regression forest, is further improved through K-nearest neighbor based aggregation scheme on both pixel- and superpixel-level. To validate the proposed confidence estimation scheme, we employ cost modulation or ground control points based optimization in stereo matching. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches on various benchmarks including challenging outdoor scenes.
Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Image Process.2
2016 Deep Self-correlation Descriptor for Dense Cross-Modal Correspondence
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
ECCV (8)2
2016 Fast Guided Global Interpolation for Depth and Motion
Yu Li 0003, Dongbo Min, Minh N. Do, Jiangbo Lu
ECCV (3)2
2016 ANCC flow: Adaptive normalized cross-correlation with evolving guidance aggregation for dense correspondence estimation
abstract
Adaptive normalized cross-correlation (ANCC) cost function works well between images under photometric distortions, but its heavy computational burden often limits its applications. To overcome this limitation, this paper proposes a robust and efficient computational framework, called ANCC flow, designed for establishing dense correspondences between images under severe photometric variations. We first simplify the weight of ANCC in an asymmetric manner by considering a source image weight only. It is then efficiently computed by applying constant-time edge-aware filters without loss of its matching accuracy. Additionally, to deal with a large discrete label space effectively, which is a challenging issue in a flow field estimation, we propose a randomized label space sampling strategy similar to PatchMatch filer (PMF) optimization. The robustness of the asymmetric ANCC and the cost filter is further enhanced through an evolving weight computation, where a flow field computed in a previous iteration is utilized to build current edge-aware weights. Experimental results demonstrate the outstanding performance of ANCC flow in many cases of dense correspondence estimations under severe photometric and geometric variations.
Seungryong Kim, Dongbo Min, Kwanghoon Sohn
ICIP2
2016 EMCCD color correction based on spectral sensitivity analysis
Jongin Son, Minsung Kang, Dongbo Min, Kwanghoon Sohn
Multim. Tools Appl.3
2015 DASC: Dense adaptive self-correlation descriptor for multi-modal and multi-spectral correspondence
abstract
Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramatically advanced in recent studies. However, finding reliable visual correspondence in multi-modal or multi-spectral images still remains unsolved. In this paper, we propose a novel dense matching descriptor, called dense adaptive self-correlation (DASC), to effectively address this kind of matching scenarios. Based on the observation that a self-similarity existing within images is less sensitive to modality variations, we define the descriptor with a series of an adaptive self-correlation similarity for patches within a local support window. To further improve the matching quality and runtime efficiency, we propose a randomized receptive field pooling, in which a sampling pattern is optimized with a discriminative learning. Moreover, the computational redundancy that arises when computing densely sampled descriptor over an entire image is dramatically reduced by applying fast edge-aware filtering. Experiments demonstrate the outstanding performance of the DASC descriptor in many cases of multi-modal and multi-spectral correspondence.
Seungryong Kim, Dongbo Min, Bumsub Ham, Seungchul Ryu, Minh N. Do, Kwanghoon Sohn
CVPR2
2015 SPM-BP: Sped-Up PatchMatch Belief Propagation for Continuous MRFs
abstract
Markov random fields are widely used to model many computer vision problems that can be cast in an energy minimization framework composed of unary and pairwise potentials. While computationally tractable discrete optimizers such as Graph Cuts and belief propagation (BP) exist for multi-label discrete problems, they still face prohibitively high computational challenges when the labels reside in a huge or very densely sampled space. Integrating key ideas from PatchMatch of effective particle propagation and resampling, PatchMatch belief propagation (PMBP) has been demonstrated to have good performance in addressing continuous labeling problems and runs orders of magnitude faster than Particle BP (PBP). However, the quality of the PMBP solution is tightly coupled with the local window size, over which the raw data cost is aggregated to mitigate ambiguity in the data constraint. This dependency heavily influences the overall complexity, increasing linearly with the window size. This paper proposes a novel algorithm called sped-up PMBP (SPM-BP) to tackle this critical computational bottleneck and speeds up PMBP by 50-100 times. The crux of SPM-BP is on unifying efficient filter-based cost aggregation and message passing with PatchMatch-based particle generation in a highly effective way. Though simple in its formulation, SPM-BP achieves superior performance for sub-pixel accurate stereo and optical-flow on benchmark datasets when compared with more complex and task-specific approaches.
Yu Li 0003, Dongbo Min, Michael S. Brown, Minh N. Do, Jiangbo Lu
ICCV2
2015 Depth Analogy: Data-Driven Approach for Single Image Depth Estimation Using Gradient Samples
abstract
Inferring scene depth from a single monocular image is a highly ill-posed problem in computer vision. This paper presents a new gradient-domain approach, called depth analogy, that makes use of analogy as a means for synthesizing a target depth field, when a collection of RGB-D image pairs is given as training data. Specifically, the proposed method employs a non-parametric learning process that creates an analogous depth field by sampling reliable depth gradients using visual correspondence established on training image pairs. Unlike existing data-driven approaches that directly select depth values from training data, our framework transfers depth gradients as reconstruction cues, which are then integrated by the Poisson reconstruction. The performance of most conventional approaches relies heavily on the training RGB-D data used in the process, and such a dependency severely degenerates the quality of reconstructed depth maps when the desired depth distribution of an input image is quite different from that of the training data, e.g., outdoor versus indoor scenes. Our key observation is that using depth gradients in the reconstruction is less sensitive to scene characteristics, providing better cues for depth recovery. Thus, our gradient-domain approach can support a great variety of training range datasets that involve substantial appearance and geometric variations. The experimental results demonstrate that our (depth) gradient-domain approach outperforms existing data-driven approaches directly working on depth domain, even when only uncorrelated training datasets are available.
Sunghwan Choi, Dongbo Min, Bumsub Ham, Youngjung Kim, Changjae Oh, Kwanghoon Sohn
IEEE Trans. Image Process.2
2015 Unsupervised Texture Flow Estimation Using Appearance-Space Clustering and Correspondence
abstract
This paper presents a texture flow estimation method that uses an appearance-space clustering and a correspondence search in the space of deformed exemplars. To estimate the underlying texture flow, such as scale, orientation, and texture label, most existing approaches require a certain amount of user interactions. Strict assumptions on a geometric model further limit the flow estimation to such a near-regular texture as a gradient-like pattern. We address these problems by extracting distinct texture exemplars in an unsupervised way and using an efficient search strategy on a deformation parameter space. This enables estimating a coherent flow in a fully automatic manner, even when an input image contains multiple textures of different categories. A set of texture exemplars that describes the input texture image is first extracted via a medoid-based clustering in appearance space. The texture exemplars are then matched with the input image to infer deformation parameters. In particular, we define a distance function for measuring a similarity between the texture exemplar and a deformed target patch centered at each pixel from the input image, and then propose to use a randomized search strategy to estimate these parameters efficiently. The deformation flow field is further refined by adaptively smoothing the flow field under guidance of a matching confidence score. We show that a local visual similarity, directly measured from appearance space, explains local behaviors of the flow very well, and the flow field can be estimated very efficiently when the matching criterion meets the randomized search strategy. Experimental results on synthetic and natural images show that the proposed method outperforms existing methods.
Sunghwan Choi, Dongbo Min, Bumsub Ham, Kwanghoon Sohn
IEEE Trans. Image Process.2
2015 Depth Superresolution by Transduction
abstract
This paper presents a depth superresolution (SR) method that uses both of a low-resolution (LR) depth image and a high-resolution (HR) intensity image. We formulate depth SR as a graph-based transduction problem. In particular, the HR intensity image is represented as an undirected graph, in which pixels are characterized as vertices, and their relations are encoded as an affinity function. When the vertices initially labeled with certain depth hypotheses (from the LR depth image) are regarded as input queries, all the vertices are scored with respect to the relevances to these queries by a classifying function. Each vertex is then labeled with the depth hypothesis that receives the highest relevance score. We design the classifying function by considering the local and global structures of the HR intensity image. This approach enables us to address a depth bleeding problem that typically appears in current depth SR methods. Furthermore, input queries are assigned in a probabilistic manner, making depth SR robust to noisy depth measurements. We also analyze existing depth SR methods in the context of transduction, and discuss their theoretic relations. Intensive experiments demonstrate the superiority of the proposed method over state-of-the-art methods both qualitatively and quantitatively.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.2
2014 Cross-Scale Cost Aggregation for Stereo Matching
abstract
Human beings process stereoscopic correspondence across multiple scales. However, this bio-inspiration is ignored by state-of-the-art cost aggregation methods for dense stereo correspondence. In this paper, a generic cross-scale cost aggregation framework is proposed to allow multi-scale interaction in cost aggregation. We firstly reformulate cost aggregation from a unified optimization perspective and show that different cost aggregation methods essentially differ in the choices of similarity kernels. Then, an inter-scale regularizer is introduced into optimization and solving this new optimization problem leads to the proposed framework. Since the regularization term is independent of the similarity kernel, various cost aggregation methods can be integrated into the proposed general framework. We show that the cross-scale framework is important as it effectively and efficiently expands state-of-the-art cost aggregation methods and leads to significant improvements, when evaluated on Middlebury, KITTI and New Tsukuba datasets.
Kang Zhang 0004, Yuqiang Fang, Dongbo Min, Lifeng Sun, Shiqiang Yang, Shuicheng Yan, Qi Tian 0001
CVPR3
2014 Randomized texture flow estimation using visual similarity
abstract
Exploring underlying texture flows defined with orientation and scale is of a great interest on a variety of vision-related tasks. However, existing methods often fail to capture accurate flows due to over-parameterization of texture deformation or employ a costly global optimization which makes the algorithm computationally demanding. In this paper, we address this inverse problem by casting it as a randomized correspondence search along with a locally-adaptive vector field smoothing. When a small example patch is given as a reference, a randomized deformable matching is performed on the very densely quantized label space, enabling an efficient estimation of texture deformation without quality degeneration, e.g., due to quantization artifacts which often appear in the optimization-driven discrete approaches. The visual similarity with respect to the deformation parameters is directly measured with an input texture image on an appearance space. The locally-adaptive smoothing is then applied to the intermediate flow field, resulting in a good continuation of the resultant texture flow. Experimental results on both synthetic and natural images show that the proposed method improves the performance in terms of both runtime efficiency and/or visual quality, compared to the existing methods.
Sunghwan Choi, Dongbo Min, Kwanghoon Sohn
ICIP2
2014 Reliability-Based Multiview Depth Enhancement Considering Interview Coherence
abstract
Color-plus-depth video format has been increasingly popular in 3-D video applications, such as auto-stereoscopic 3-D TV and freeview TV. The performance of these applications is heavily dependent on the quality of depth maps since intermediate views are synthesized using the corresponding depth maps. This paper presents a novel framework for obtaining high-quality multiview color-plus-depth video using a hybrid sensor, which consists of multiple color cameras and depth sensors. Given multiple high-resolution color images and low quality depth maps obtained from the color cameras and depth sensors, we improve the quality of the depth map corresponding to each color view by increasing its spatial resolution and enforcing interview coherence. Specifically, a new up-sampling method considering the interview coherence is proposed to enhance multiview depth maps. This approach can improve the performance of the existing up-sampling algorithms, such as joint bilateral up-sampling and weighted mode filtering, which have been developed to enhance a single-view depth map only. In addition, an adaptive approach of fusing multiple input low-resolution depth maps is proposed based on the reliability that considers camera geometry and depth validity. The proposed framework can be extended into the temporal domain for temporally consistent depth maps. Experimental results demonstrate that the proposed method provides better multiview depth quality than the conventional single-view-based methods. We also show that it provides comparable results, yet much more efficiently, to other fusion approaches that employ both depth sensors and stereo matching algorithm together. Moreover, it is shown that the proposed method significantly reduces bit rates required to compress the multiview color-plus-depth video.
Jinwook Choi, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.2
2014 Probability-Based Rendering for View Synthesis
abstract
In this paper, a probability-based rendering (PBR) method is described for reconstructing an intermediate view with a steady-state matching probability (SSMP) density function. Conventionally, given multiple reference images, the intermediate view is synthesized via the depth image-based rendering technique in which geometric information (e.g., depth) is explicitly leveraged, thus leading to serious rendering artifacts on the synthesized view even with small depth errors. We address this problem by formulating the rendering process as an image fusion in which the textures of all probable matching points are adaptively blended with the SSMP representing the likelihood that points among the input reference images are matched. The PBR hence becomes more robust against depth estimation errors than existing view synthesis approaches. The MP in the steady-state, SSMP, is inferred for each pixel via the random walk with restart (RWR). The RWR always guarantees visually consistent MP, as opposed to conventional optimization schemes (e.g., diffusion or filtering-based approaches), the accuracy of which heavily depends on parameters used. Experimental results demonstrate the superiority of the PBR over the existing view synthesis approaches both qualitatively and quantitatively. Especially, the PBR is effective in suppressing flicker artifacts of virtual video rendering although no temporal aspect is considered. Moreover, it is shown that the depth map itself calculated from our RWR-based method (by simply choosing the most probable matching point) is also comparable with that of the state-of-the-art local stereo matching methods.
Bumsub Ham, Dongbo Min, Changjae Oh, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Image Process.2
2014 Fast Global Image Smoothing Based on Weighted Least Squares
abstract
This paper presents an efficient technique for performing a spatially inhomogeneous edge-preserving image smoothing, called fast global smoother. Focusing on sparse Laplacian matrices consisting of a data term and a prior term (typically defined using four or eight neighbors for 2D image), our approach efficiently solves such global objective functions. In particular, we approximate the solution of the memory-and computation-intensive large linear system, defined over a d-dimensional spatial domain, by solving a sequence of 1D subsystems. Our separable implementation enables applying a linear-time tridiagonal matrix algorithm to solve d three-point Laplacian matrices iteratively. Our approach combines the best of two paradigms, i.e., efficient edge-preserving filters and optimization-based smoothing. Our method has a comparable runtime to the fast edge-preserving filters, but its global optimization formulation overcomes many limitations of the local filtering approaches. Our method also achieves high-quality results as the state-of-the-art optimization-based techniques, but runs ∼10-30 times faster. Besides, considering the flexibility in defining an objective function, we further propose generalized fast algorithms that perform Lγ norm smoothing (0 < γ < 2) and support an aggregated (robust) data term for handling imprecise data constraints. We demonstrate the effectiveness and efficiency of our techniques in a range of image processing and computer graphics applications.
Dongbo Min, Sunghwan Choi, Jiangbo Lu, Bumsub Ham, Kwanghoon Sohn, Minh N. Do
IEEE Trans. Image Process.1
2013 Patch Match Filter: Efficient Edge-Aware Filtering Meets Randomized Search for Fast Correspondence Field Estimation
abstract
Though many tasks in computer vision can be formulated elegantly as pixel-labeling problems, a typical challenge discouraging such a discrete formulation is often due to computational efficiency. Recent studies on fast cost volume filtering based on efficient edge-aware filters have provided a fast alternative to solve discrete labeling problems, with the complexity independent of the support window size. However, these methods still have to step through the entire cost volume exhaustively, which makes the solution speed scale linearly with the label space size. When the label space is huge, which is often the case for (sub pixel-accurate) stereo and optical flow estimation, their computational complexity becomes quickly unacceptable. Developed to search approximate nearest neighbors rapidly, the Patch Match method can significantly reduce the complexity dependency on the search space size. But, its pixel-wise randomized search and fragmented data access within the 3D cost volume seriously hinder the application of efficient cost slice filtering. This paper presents a generic and fast computational framework for general multi-labeling problems called Patch Match Filter (PMF). For the very first time, we explore effective and efficient strategies to weave together these two fundamental techniques developed in isolation, i.e., Based-based randomized search and efficient edge-aware image filtering. By decompositing an image into compact super pixels, we also propose super pixel-based novel search strategies that generalize and improve the original Patch Match method. Focusing on dense correspondence field estimation in this paper, we demonstrate PMF's applications in stereo and optical flow. Our PMF methods achieve state-of-the-art correspondence accuracy but run much faster than other competing methods, often giving over 10-times speedup for large label space cases.
Jiangbo Lu, Hongsheng Yang, Dongbo Min, Minh N. Do
CVPR3
2013 Joint Histogram-Based Cost Aggregation for Stereo Matching
abstract
This paper presents a novel method for performing efficient cost aggregation in stereo matching. The cost aggregation problem is reformulated from the perspective of a histogram, giving us the potential to reduce the complexity of the cost aggregation in stereo matching significantly. Differently from previous methods which have tried to reduce the complexity in terms of the size of an image and a matching window, our approach focuses on reducing the computational redundancy that exists among the search range, caused by a repeated filtering for all the hypotheses. Moreover, we also reduce the complexity of the window-based filtering through an efficient sampling scheme inside the matching window. The tradeoff between accuracy and complexity is extensively investigated by varying the parameters used in the proposed method. Experimental results show that the proposed method provides high-quality disparity maps with low complexity and outperforms existing local methods. This paper also provides new insights into complexity-constrained stereo-matching algorithm design.
Dongbo Min, Jiangbo Lu, Minh N. Do
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Stereo/multiview picture quality: Overview and recent advances
Stefan Winkler 0001, Dongbo Min
Signal Process. Image Commun.2
2013 Efficient Techniques for Depth Video Compression Using Weighted Mode Filtering
abstract
This paper proposes efficient techniques to compress a depth video by taking into account coding artifacts, spatial resolution, and dynamic range of the depth data. Due to abrupt signal changes on object boundaries, a depth video compressed by conventional video coding standards often introduces serious coding artifacts over object boundaries, which severely affect the quality of a synthesized view. We suppress the coding artifacts by proposing an efficient postprocessing method based on a weighted mode filtering and utilizing it as an in-loop filter. In addition, the proposed filter is also tailored to efficiently reconstruct the depth video from the reduced spatial resolution and the low dynamic range. The down/upsampling coding approaches for the spatial resolution and the dynamic range are used together with the proposed filter in order to further reduce the bit rate. We verify the proposed techniques by applying them to an efficient compression of multiview-plus-depth data, which has emerged as an efficient data representation for 3-D video. Experimental results show that the proposed techniques significantly reduce the bit rate while achieving a better quality of the synthesized view in terms of both objective and subjective measures.
Dongbo Min, Minh N. Do
IEEE Trans. Circuits Syst. Video Technol.2
2013 Revisiting the Relationship Between Adaptive Smoothing and Anisotropic Diffusion With Modified Filters
abstract
Anisotropic diffusion has been known to be closely related to adaptive smoothing and discretized in a similar manner. This paper revisits a fundamental relationship between two approaches. It is shown that adaptive smoothing and anisotropic diffusion have different theoretical backgrounds by exploring their characteristics with the perspective of normalization, evolution step size, and energy flow. Based on this principle, adaptive smoothing is derived from a second order partial differential equation (PDE), not a conventional anisotropic diffusion, via the coupling of Fick's law with a generalized continuity equation where a "source" or "sink" exists, which has not been extensively exploited. We show that the source or sink is closely related to the asymmetry of energy flow as well as the normalization term of adaptive smoothing. It enables us to analyze behaviors of adaptive smoothing, such as the maximum principle and stability with a perspective of a PDE. Ultimately, this relationship provides new insights into application-specific filtering algorithm design. By modeling the source or sink in the PDE, we introduce two specific diffusion filters, the robust anisotropic diffusion and the robust coherence enhancing diffusion, as novel instantiations which are more robust against the outliers than the conventional filters.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.2
2013 A Generalized Random Walk With Restart and its Application in Depth Up-Sampling and Interactive Segmentation
abstract
In this paper, the origin of random walk with restart (RWR) and its generalization are described. It is well known that the random walk (RW) and the anisotropic diffusion models share the same energy functional, i.e., the former provides a steady-state solution and the latter gives a flow solution. In contrast, the theoretical background of the RWR scheme is different from that of the diffusion-reaction equation, although the restarting term of the RWR plays a role similar to the reaction term of the diffusion-reaction equation. The behaviors of the two approaches with respect to outliers reveal that they possess different attributes in terms of data propagation. This observation leads to the derivation of a new energy functional, where both volumetric heat capacity and thermal conductivity are considered together, and provides a common framework that unifies both the RW and the RWR approaches, in addition to other regularization methods. The proposed framework allows the RWR to be generalized (GRWR) in semilocal and nonlocal forms. The experimental results demonstrate the superiority of GRWR over existing regularization approaches in terms of depth map up-sampling and interactive image segmentation.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.2
2012 Cross-based local multipoint filtering
abstract
This paper presents a cross-based framework of performing local multipoint filtering efficiently. We formulate the filtering process as a local multipoint regression problem, consisting of two main steps: 1) multipoint estimation, calculating the estimates for a set of points within a shape-adaptive local support, and 2) aggregation, fusing a number of multipoint estimates available for each point. Compared with the guided filter that applies the linear regression to all pixels covered by a fixed-sized square window non-adaptively, the proposed filtering framework is a more generalized form. Two specific filtering methods are instantiated from this framework, based on piecewise constant and piecewise linear modeling, respectively. Leveraging a cross-based local support representation and integration technique, the proposed filtering methods achieve theoretically strong results in an efficient manner, with the two main steps' complexity independent of the filtering kernel size. We demonstrate the strength of the proposed filters in various applications including stereo matching, depth map enhancement, edge-preserving smoothing, color image denoising, detail enhancement, and flash/no-flash denoising.
Jiangbo Lu, Keyang Shi, Dongbo Min, Liang Lin 0004, Minh N. Do
CVPR3
2012 Weighted mode filtering and its applications to depth video enhancement and coding
abstract
This paper presents a novel approach for improving the quality of depth video. Given a high-quality color image and its corresponding low-quality depth image, we handle various artifacts which may exist on the depth video by applying a weighted mode filtering method based on a joint histogram. When the histogram is generated, the weight based on color similarity between reference and neighboring pixels on the color image is computed and then used for counting each bin on the joint histogram of the depth map. A final solution is determined by seeking a global mode on the histogram. Experimental results show that the proposed method has outstanding performance and is very efficient in various applications such as depth video enhancement and compression.
Dongbo Min, Jiangbo Lu, Minh N. Do
ICASSP1
2012 Efficient edge-preserving interpolation and in-loop filters for depth map compression
abstract
Due to abrupt signal changes on object boundaries, a depth video compressed by conventional video coding standards often introduces serious coding artifacts over the boundaries, which severely affect the quality of a synthesized view. In this paper, we propose an edge-preserving depth interpolation filter based on weighted mode filtering to provide more accurate fractional-pixel samples in the motion-compensated interpolation for an effective inter-coding of the depth video. In addition, an efficient post-processing method is also proposed to further suppress the coding artifacts on the depth video and utilized as an in-loop filter. Experimental results show the proposed methods can significantly improve the synthesized view quality in terms of both objective and subjective measures.
Dongbo Min, Minh N. Do
ICIP2
2012 Robust Scale-Space Filter Using Second-Order Partial Differential Equations
abstract
This paper describes a robust scale-space filter that adaptively changes the amount of flux according to the local topology of the neighborhood. In a manner similar to modeling heat or temperature flow in physics, the robust scale-space filter is derived by coupling Fick's law with a generalized continuity equation in which the source or sink is modeled via a specific heat capacity. The filter plays an essential part in two aspects. First, an evolution step size is adaptively scaled according to the local structure, enabling the proposed filter to be numerically stable. Second, the influence of outliers is reduced by adaptively compensating for the incoming flux. We show that classical diffusion methods represent special cases of the proposed filter. By analyzing the stability condition of the proposed filter, we also verify that its evolution step size in an explicit scheme is larger than that of the diffusion methods. The proposed filter also satisfies the maximum principle in the same manner as the diffusion. Our experimental results show that the proposed filter is less sensitive to the evolution step size, as well as more robust to various outliers, such as Gaussian noise, impulsive noise, or a combination of the two.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.2
2012 Depth Video Enhancement Based on Weighted Mode Filtering
abstract
This paper presents a novel approach for depth video enhancement. Given a high-resolution color video and its corresponding low-quality depth video, we improve the quality of the depth video by increasing its resolution and suppressing noise. For that, a weighted mode filtering method is proposed based on a joint histogram. When the histogram is generated, the weight based on color similarity between reference and neighboring pixels on the color image is computed and then used for counting each bin on the joint histogram of the depth map. A final solution is determined by seeking a global mode on the histogram. We show that the proposed method provides the optimal solution with respect to L(1) norm minimization. For temporally consistent estimate on depth video, we extend this method into temporally neighboring frames. Simple optical flow estimation and patch similarity measure are used for obtaining the high-quality depth video in an efficient manner. Experimental results show that the proposed method has outstanding performance and is very efficient, compared with existing methods. We also show that the temporally consistent enhancement of depth video addresses a flickering problem and improves the accuracy of depth video.
Dongbo Min, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.1
2011 High level synthesis of stereo matching: Productivity, performance, and software constraints
abstract
FPGAs are an attractive platform for applications with high computation demand and low energy consumption requirements. However, design effort for FPGA implementations remains high - often an order of magnitude larger than design effort using high level languages. Instead of this time-consuming process, high level synthesis (HLS) tools generate hardware implementations from high level languages (HLL) such as C/C++/SystemC. Such tools reduce design effort: high level descriptions are more compact and less error prone. HLS tools promise hardware development abstracted from software designer knowledge of the implementation platform. In this paper, we examine several implementations of stereo matching, an active area of computer vision research that uses techniques also common for image de-noising, image retrieval, feature matching and face recognition. We present an unbiased evaluation of the suitability of using HLS for typical stereo matching software, usability and productivity of AutoPilot (a state of the art HLS tool), and the performance of designs produced by AutoPilot. Based on our study, we provide guidelines for software design, limitations of mapping general purpose software to hardware using HLS, and future directions for HLS tool development. For the stereo matching algorithms, we demonstrate between 3.5X and 67.9X speedup over software (but less than achievable by manual RTL design) with a five-fold reduction in design effort vs. manual hardware design.
Kyle Rupnow, Yun Liang 0001, Dongbo Min, Minh N. Do, Deming Chen
FPT4
2011 A revisit to MRF-based depth map super-resolution and enhancement
abstract
This paper presents a Markov Random Field (MRF)-based approach for depth map super-resolution and enhancement. Given a low-resolution or moderate quality depth map, we study the problem of enhancing its resolution or quality with a registered high-resolution color image. Different from the previous methods, this MRF-based approach is based on a novel data term formulation that fits well to the unique characteristics of depth maps. We also discuss a few important design choices that boost the performance of general MRF-based methods. Experimental results show that our proposed approach achieves high resolution depth maps at more desirable quality, both qualitatively and quantitatively. It can also be applied to enhance the depth maps derived with state-of-the-art stereo methods, resulting in the raised ranking based on the Middlebury benchmark.
Jiangbo Lu, Dongbo Min, Ramanpreet Singh Pahwa, Minh N. Do
ICASSP2
2011 A revisit to cost aggregation in stereo matching: How far can we reduce its computational redundancy?
abstract
This paper presents a novel method for performing an efficient cost aggregation in stereo matching. The cost aggregation problem is re-formulated with a perspective of a histogram, and it gives us a potential to reduce the complexity of the cost aggregation significantly. Different from the previous methods which have tried to reduce the complexity in terms of the size of an image and a matching window, our approach focuses on reducing the computational redundancy which exists among the search range, caused by a repeated filtering for all disparity hypotheses. Moreover, we also reduce the complexity of the window-based filtering through an efficient sampling scheme inside the matching window. The trade-off between accuracy and complexity is extensively investigated into parameters used in the proposed method. Experimental results show that the proposed method provides high-quality disparity maps with low complexity. This work provides new insights into complexity-constrained stereo matching algorithm design.
Dongbo Min, Jiangbo Lu, Minh N. Do
ICCV1
2011 Cost aggregation with anisotropic diffusion in feature space for hybrid stereo matching
abstract
In this paper, we present a cost aggregation using anisotropic diffusion on a feature space for hybrid stereo matching. Stereo matching can be classified into two categories: feature-based and area-based approaches. Feature-based approaches generate accurate but sparse disparity maps. On the other hand, area-based approaches generate dense but unreliable disparity maps, especially at depth discontinuities and homogeneous regions. We hence propose a stereo matching algorithm having advantages of both approaches. We study how to design a correspondence algorithm without modeling any depth cues except disparity. A procedure of depth perception is modeled via anisotropic diffusion on the feature space in terms of coherence. Based on the assumption that similar local feature space has similar disparity, we define the feature space and its similarity and then introduce feature confidences into the proposed model. Experimental results show that the performance of the proposed method is comparable to that of the state-of-the-art methods.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
ICIP2
2010 Disparity search range estimation: Enforcing temporal consistency
abstract
This paper presents a new approach for estimating the disparity search range in stereo video that enforces temporal consistency. Reliable search range estimation is very important since an incorrect estimate causes most stereo matching methods to get trapped in local minima or produce unstable results over time. In this work, the search range is estimated based on a disparity histogram that is generated with sparse feature matching algorithms such as SURF. To achieve more stable results over time, we further propose to enforce temporal consistency by calculating a weighted sum of temporally-neighboring histograms, where the weights are determined by the similarity of depth distribution between frames. Experimental results show that this proposed method yields accurate disparity search ranges for several challenging stereo videos and is robust to various forms of noise, scene complexity and camera configurations.
Dongbo Min, Sehoon Yea, Zafer Arican, Anthony Vetro
ICASSP1
2010 3D JBU based depth video filtering for temporal fluctuation reduction
abstract
In this paper, we propose a three-dimensional Joint Bilateral Up-sampling (3D JBU) for the depth video which can be applied to a 3DTV system based on 2D-plus-depth video. Recently, a number of researches have been done to improve a resolution and frame-rate of the depth video. We proposed a novel method that enhances depth video obtained by Time-of-Flight (TOF) sensor by combining it with Charge-coupled Device (CCD) camera [1]. However, this method does not consider the temporal coherence of the depth video. It may cause an eye fatigue on the 3D display and increase bit rates on video coding, since it is possible to generate a temporal fluctuation problem. Therefore, in order to solve these problems, we propose a 3D JBU model which is extended conventional JBU into the temporal domain of depth video. Experimental results show that depth video obtained by the proposed method provides satisfactory quality.
Jinwook Choi, Dongbo Min, Donghyun Kim 0010, Kwanghoon Sohn
ICIP2
2010 Occlusion handling based on support and decision
abstract
This paper proposes a novel method for handling occluded pixels in stereo images based on a probabilistic voting framework that utilizes a novel support-and-decision process. Occlusion handling aims to assign a reasonable disparity value to occluded pixels in the disparity maps. In an initial step, disparities and their corresponding supports at the occluded pixels are calculated using a probabilistic voting method using the disparities at visible pixels. In this way, the visible pixel information is propagated when the disparities and supports at the occluded pixels are computed. The final disparities for occluded pixels are then computed through an iterative support-and-decision process to propagate the information inside the occluded pixel region. An acceleration technique is also proposed to improve the performance of the iterative support-and-decision process. Experimental results show that the proposed occlusion handling method works well for several challenging stereo images.
Dongbo Min, Sehoon Yea, Anthony Vetro
ICIP1
2010 An asymmetric post-processing for correspondence problem
Dongbo Min, Kwanghoon Sohn
Signal Process. Image Commun.1
2009 Spatial and temporal up-conversion technique for depth video
abstract
This paper proposes a novel framework for up-conversion of depth video resolution both in spatial and in time domain. Time-of-flight (TOF) sensors are widely used in computer vision fields. Although TOF sensors provide depth video in real time, there are some problems in a sense that it provides a low resolution and a low frame-rate depth video. We propose a cheaper solution that enhances depth video obtained by TOF sensor by combining it with CCD camera. The proposed method provides high quality video as a cheaper solution for low resolution, and low frame-rate depth video. It is useful when depth video is used in various applications such as 3DTV, free-view TV, teleconference system. High-quality depth video can be obtained by motion compensated frame interpolation (MCFI) and extended joint bilateral upsampling (JBU). Experimental results show that depth video obtained by the proposed method has satisfactory quality.
Jinwook Choi, Dongbo Min, Bumsub Ham, Kwanghoon Sohn
ICIP2
2009 Virtual view rendering using super-resolution with multiview images
abstract
This paper presents a new approach to solve the problem of quality degradation of a synthesized view, when a virtual camera moves forward. Interpolation techniques using only two neighboring views are generally applied when a virtual view is synthesized. Because the size of an object increases when the virtual camera moves forward, conventional methods have usually addressed this problem by interpolation techniques in order to synthesize a virtual view. However, as it generates a degraded view such as blurred images, we prevent a synthesized view from being blurred by using more images in multiview camera configuration. That is, this problem is solved by applying super-resolution concept which reconstructs a high resolution image from several low resolution images. Data fusion is performed by geometric warping with disparity maps of the multiple images followed by deblurring. Experimental results show that the image quality can further be improved by reducing blurring and halo effects in comparison with the interpolation method.
Bumsub Ham, Dongbo Min, Jinwook Choi, Kwanghoon Sohn
ICIP2
2009 2D/3D freeview video generation for 3DTV system
Dongbo Min, Donghyun Kim 0010, SangUn Yun, Kwanghoon Sohn
Signal Process. Image Commun.1
2008 Stereo matching with asymmetric occlusion handling in weighted least square framework
abstract
This paper presents a novel method for stereo matching with occlusion handling. In order to estimate optimal cost, we define an energy function and solve the iterative equation with the numerical method. We improve performance and convergence rate by using several acceleration techniques. The proposed method is computationally efficient since it does not use color segmentation or any global optimization techniques. For occlusion handling, which has not been performed effectively by any conventional cost aggregation approaches, we combine the occlusion problem with the proposed minimization scheme. Asymmetric information is used so that few additional computational loads are necessary. Experimental results show that performance is comparable to that of many state-of-the-art methods.
Dongbo Min, Kwanghoon Sohn
ICASSP1
2008 2D/3D freeview video generation for 3DTV system
abstract
In this paper, we propose a new approach of synthesizing novel views in multiview camera configurations. We introduce the semi iV-view & iV-depth framework in order to estimate disparity maps efficiently and correctly. This framework reduces redundancy on disparity estimation by using information from neighboring views. The occlusion problem is handled by using cost functions computed with multiview images. The proposed method provides a 2D/3D freeview video. User can select 2D/3D modes of freeview video and control 3D depth perception by adjusting several parameters in 3D freeview video. Experimental results show that the proposed method yields the accurate disparity maps and provides seamless freeview videos.
Dongbo Min, Donghyun Kim 0010, Kwanghoon Sohn
ICIP1
2008 Asymmetric post-processing for stereo correspondence
abstract
This paper presents a novel approach that performs post-processing for stereo correspondence. We improve the performance of stereo correspondence by performing consistency check and adaptive filtering in an iterative filtering scheme. The consistency check is done with asymmetric information only so that very few additional computational loads are necessary. The proposed post-filtering method can be used in various methods for stereo correspondence without any modification. We demonstrate the validity of the proposed method by applying it to hierarchical belief propagation and semi-global matching.
Dongbo Min, Juhyun Oh, Kwanghoon Sohn
ICPR1
2008 Freeview rendering with trinocular camera
abstract
The paper presents a method for synthesizing novel view from the virtual camera with trinocular camera configuration. We propose a cost aggregation method with weighted least square for stereo matching, and address the occlusion problem by using cost functions computed with multiview images. We avoid the redundancy of estimating disparity maps for all the images by using simple geometry transfer method. The novel view is synthesized by view-dependent geometries on the 3D translation of virtual camera. Experimental results show that the proposed method yields the accurate disparity maps and the synthesized novel view is satisfactory enough to provide a viewer freeview videos.
Dongbo Min, Donghyun Kim 0010, SangUn Yuri, Kwanghoon Sohn
ISCAS1
2008 Cost Aggregation and Occlusion Handling With WLS in Stereo Matching
abstract
This paper presents a novel method for cost aggregation and occlusion handling for stereo matching. In order to estimate optimal cost, given a per-pixel difference image as observed data, we define an energy function and solve the minimization problem by solving the iterative equation with the numerical method. We improve performance and increase the convergence rate by using several acceleration techniques such as the Gauss-Seidel method, the multiscale approach, and adaptive interpolation. The proposed method is computationally efficient since it does not use color segmentation or any global optimization techniques. For occlusion handling, which has not been performed effectively by any conventional cost aggregation approaches, we combine the occlusion problem with the proposed minimization scheme. Asymmetric information is used so that few additional computational loads are necessary. Experimental results show that performance is comparable to that of many state-of-the-art methods. The proposed method is in fact the most successful among all cost aggregation methods based on standard stereo test beds.
Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.1
2006 Edge-preserving joint motion-disparity estimation in stereo image sequences
Dongbo Min, Hansung Kim 0001, Kwanghoon Sohn
Signal Process. Image Commun.1