VLDB 2026 Research / reviewers in the wild / expert
Sunok Kim
dblp:172/9755
· DBLP profile ↗
26ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-9665-4214ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion model-based data augmentation for land cover segmentation in Pol-SAR imagery
Keunhoon Choi, Sunok Kim, Kwanghoon Sohn |
Pattern Recognit. | 2 |
| 2025 | Deep Spiking Neural Network for Energy-Efficient SAR Ship DetectionabstractIn this letter, we introduce the first spiking-based network optimized for synthetic aperture radar (SAR) ship detection and compare its performance with conventional neural networks (CNNs). Spiking neural networks (SNNs) offer significant advantages over traditional artificial neural networks (ANNs) by resulting in highly efficient computation. Unlike ANNs, SNNs only perform calculations when spikes occur, leading to lower power consumption and reduced computational costs, making them ideal for energy-constrained and onboard applications. Furthermore, we conduct experiments to analyze the power differences between the SNN and traditional ANN-based detection models. The results demonstrate the potential advantages of SNNs in terms of power efficiency and computational load in satellite-based target detection. Minjung Yoo, Juhyeon Han, Sunok Kim |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Learning confidence measure with transformer in stereo matching
Jini Yang, Minjung Yoo, Jaehoon Cho, Sunok Kim |
Pattern Recognit. | 4 |
| 2024 | A Prototype Unit for Image De-raining using Time-Lapse Data
Jaehoon Cho, Minjung Yoo, Jini Yang, Sunok Kim |
BMVC | 4 |
| 2024 | Noise Robust SAR Image Classification Using Siamese Spiking Neural NetworksabstractIn advancements in artificial neural network (ANN) learning, significant enhancements have been made to the accuracy and robustness of synthetic aperture radar (SAR) image classification. However, using ANNs involves high computational costs. On the other hand, spiking neural networks (SNNs), recognized as the third generation of neural networks, have been introduced due to their energy efficiency. Recent research has explored the use of SNNs for SAR classification, but deeper architectures and noise-robust settings have not been fully explored, resulting in a gap in accuracy compared to ANN models. In this letter, we propose a deep Siamese SNN with speckle noise augmentation that addresses the limitations of shallow SNNs in previous studies, enhances robustness against noise, and leverages information maximization through Poisson encoding and soft resetting. We have validated the effectiveness of our model against various ANN and SNN models on different variations of the MSTAR dataset, including those with limited samples and noisy images. Our results demonstrate the potential of SNNs to achieve comparable performance to ANNs in SAR image classification. Jini Yang, Sunok Kim |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Discriminative action tubelet detector for weakly-supervised action detection
Jiyoung Lee 0005, Seungryong Kim, Sunok Kim, Kwanghoon Sohn |
Pattern Recognit. | 3 |
| 2023 | Unsupervised Deep Asymmetric Stereo Matching with Spatially-Adaptive Self-SimilarityabstractUnsupervised stereo matching has received a lot of attention since it enables the learning of disparity estimation without ground-truth data. However, most of the unsupervised stereo matching algorithms assume that the left and right images have consistent visual properties, i.e., symmetric, and easily fail when the stereo images are asymmetric. In this paper, we present a novel spatially-adaptive self-similarity (SASS) for unsupervised asymmetric stereo matching. It extends the concept of self-similarity and generates deep features that are robust to the asymmetries. The sampling patterns to calculate self-similarities are adaptively generated throughout the image regions to effectively encode diverse patterns. In order to learn the effective sampling patterns, we design a contrastive similarity loss with positive and negative weights. Consequently, SASS is further encouraged to encode asymmetry-agnostic features, while maintaining the distinctiveness for stereo correspondence. We present extensive experimental results including ablation studies and comparisons with different methods, demonstrating effectiveness of the proposed method under resolution and noise asymmetries. Taeyong Song, Sunok Kim, Kwanghoon Sohn |
CVPR | 2 |
| 2023 | Stereo Confidence Estimation via Locally Adaptive Fusion and Knowledge DistillationabstractStereo confidence estimation aims to estimate the reliability of the estimated disparity by stereo matching. Different from the previous methods that exploit the limited input modality, we present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. Moreover, we propose a knowledge distillation framework to learn more compact confidence estimation networks as student networks. By transferring the knowledge from LAF-Net as teacher networks, the student networks that solely take as input a disparity can achieve comparable performance. To transfer more informative knowledge, we also propose a module to learn the locally-varying temperature in a softmax function. We further extend this framework to a multiview scenario. Experimental results show that LAF-Net and its variations outperform the state-of-the-art stereo confidence methods on various benchmarks. Sunok Kim, Seungryong Kim, Dongbo Min, Pascal Frossard, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Mask-Guided Attention and Episode Adaptive Weights for Few-Shot SegmentationabstractFew-shot segmentation aims to segment objects with novel classes in a query image, given a support set which consists of few annotated support images. A key factor in few-shot segmentation is to effectively exploit information for the target classes from the support set. In addition, we argue that the overall quality of information available in each training episode varies depending on the given support samples. In this paper, we propose Mask-Guided Attention module to extract more beneficial features for few-shot segmentation from the support images. Taking advantage of the support masks, the area correlated to the foreground object is highlighted and enables the support encoder to extract comprehensive support features with contextual information. Furthermore, we propose Episode Adaptive Weight to balance the training between different episodes. It adaptively adjusts loss weight according to the difficulty of each episode determined by self-supervised segmentation loss of support images and encourages the model to pay more attention to more difficult episodes. Extensive experimental results including comparisons with the state-of-the-art methods and ablation studies demonstrate the effectiveness of the proposed method. Hyeongjun Kwon, Taeyong Song, Sunok Kim, Kwanghoon Sohn |
ICIP | 3 |
| 2022 | Meta-confidence estimation for stereo matchingabstractWe propose a novel framework to estimate the confidence of a disparity map taking into account, for the first time, the uncertainty affecting the confidence estimation process itself. Conversely to other tasks such as disparity estimation, the uncertainty of confidence directly hints that the confidence should be increased if initially low, but with high uncertainty, decreased otherwise. By modelling such a cue in the form of a second-level confidence, or meta-confidence, our solution allows for finding incorrect predictions inferred by confidence estimator and for learning a correction for them. Our strategy is suited for any state-of-the-art method known in literature, either implemented using random forest classifiers or deep neural networks. Especially, for deep neural networks-based models, we present a multi-headed confidence estimator followed by an uncertainty network, so as to predict mean confidence and meta-confidence within a single network without the cost of lower accuracy, a known limitation in literature for uncertainty estimation. Experimental results on a variety of stereo algorithms and confidence estimation models prove that the modeled meta-confidence is meaningful of the reliability of the estimated confidence and allows for refining it. Seungryong Kim, Matteo Poggi, Sunok Kim, Kwanghoon Sohn, Stefano Mattoccia |
ICRA | 3 |
| 2022 | Deep Cascade Network for Noise-Robust SAR Ship Detection With Label AugmentationabstractDeep learning has recently made an impressive advance in ship detection in Synthetic Aperture Radar (SAR) images. In spite of this advancement, conventional deep detection networks often suffer from speckle noise that inherently occurs in SAR images. However, despeckling researches have focused only on improving the visual quality of the SAR images. Despeckling without considering subsequent task may cause loss of semantic information and result in performance degradation. In this letter, we propose a deep cascade framework for noise-robust SAR ship detection that sequentially performs despeckle and detection. We effectively train our cascade network using pseudo SAR images with SAR-like structures and additional detection annotations. We also propose semantic conservative loss that allows these two tasks to cooperate with each other. Experimental results including comparisons to previous methods and extensive ablation studies show the effectiveness of our proposed method. Keunhoon Choi, Taeyong Song, Sunok Kim, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Shape-Robust SAR Ship Detection via Context-Preserving Augmentation and Deep Contrastive RoI LearningabstractWith recent advances in deep-learning techniques and the advantages of synthetic aperture radar (SAR) images, deep-learning-based SAR ship detection has attracted a lot of attention. In this letter, we aim to build an SAR ship detection framework that is robust to target shape variations. Focusing on shape variation caused by radar shadow, we propose an instance-level data augmentation (DA) method. We leverage ground-truth annotations for bounding box and instance segmentation mask to design a sophisticated pipeline to simulate target information loss while preserving contextual information. In order to enhance the capacity to model shape variation, we design network architecture using deformable convolutional networks (DCNs). Furthermore, we introduce contrastive region of interest (RoI) loss to encourage similarity between original and augmented target RoI features, while encouraging background features to be distinguished from the target RoI features. We present extensive experiments to demonstrate the effectiveness of the proposed method. Taeyong Song, Sunok Kim, Kwanghoon Sohn |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | On the Confidence of Stereo Matching in a Deep-Learning Era: A Quantitative EvaluationabstractStereo matching is one of the most popular techniques to estimate dense depth maps by finding the disparity between matching pixels on two, synchronized and rectified images. Alongside with the development of more accurate algorithms, the research community focused on finding good strategies to estimate the reliability, i.e., the confidence, of estimated disparity maps. This information proves to be a powerful cue to naively find wrong matches as well as to improve the overall effectiveness of a variety of stereo algorithms according to different strategies. In this paper, we review more than ten years of developments in the field of confidence estimation for stereo matching. We extensively discuss and evaluate existing confidence measures and their variants, from hand-crafted ones to the most recent, state-of-the-art learning based methods. We study the different behaviors of each measure when applied to a pool of different stereo algorithms and, for the first time in literature, when paired with a state-of-the-art deep stereo network. Our experiments, carried out on five different standard datasets, provide a comprehensive overview of the field, highlighting in particular both strengths and limitations of learning-based strategies. Matteo Poggi, Seungryong Kim, Fabio Tosi, Sunok Kim, Filippo Aleotti, Dongbo Min, Kwanghoon Sohn, Stefano Mattoccia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech SeparationabstractIn this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depends on the accuracy of audio-visual synchronization and the effectiveness of their representations. To overcome the frame discontinuity problem between two modalities due to transmission delay mismatch or jitter, we propose a cross-modal affinity network (CaffNet) that learns global correspondence as well as locally-varying affinities between audio and visual streams. Given that the global term provides stability over a temporal sequence at the utterance-level, this resolves the label permutation problem characterized by inconsistent assignments. By extending the proposed cross-modal affinity on the complex network, we further improve the separation performance in the complex spectral domain. Experimental results verify that the proposed methods outperform conventional ones on various datasets, demonstrating their advantages in real-world scenarios. Jiyoung Lee 0005, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang, Kwanghoon Sohn |
CVPR | 3 |
| 2021 | Adaptive confidence thresholding for monocular depth estimationabstractSelf-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to leverage pseudo ground truth depth maps of stereo images generated from self-supervised stereo matching methods. The confidence map of the pseudo ground truth depth map is estimated to mitigate performance degeneration by inaccurate pseudo depth maps. To cope with the prediction error of the confidence map itself, we also leverage the threshold network that learns the threshold dynamically conditioned on the pseudo depth maps. The pseudo depth labels filtered out by the thresholded confidence map are used to supervise the monocular depth network. Furthermore, we propose the probabilistic framework that refines the monocular depth map with the help of its uncertainty map through the pixel-adaptive convolution (PAC) layer. Experimental results demonstrate superior performance to state-of-the-art monocular depth estimation methods. Lastly, we exhibit that the proposed threshold learning can also be used to improve the performance of existing confidence estimation approaches. Hyesong Choi, Hunsang Lee, Sunkyung Kim, Sunok Kim, Seungryong Kim, Kwanghoon Sohn, Dongbo Min |
ICCV | 4 |
| 2021 | Adversarial Confidence Estimation Networks for Robust Stereo MatchingabstractStereo matching aiming to perceive the 3-D geometry of a scene facilitates numerous computer vision tasks used in advanced driver assistance systems (ADAS). Although numerous methods have been proposed for this task by leveraging deep convolutional neural networks (CNNs), stereo matching still remains an unsolved problem due to its inherent matching ambiguities. To overcome these limitations, we present a method for jointly estimating disparity and confidence from stereo image pairs through deep networks. We accomplish this through a minmax optimization to learn the generative cost aggregation networks and discriminative confidence estimation networks in an adversarial manner. Concretely, the generative cost aggregation networks are trained to accurately generate disparities at both confident and unconfident pixels from an input matching cost that are indistinguishable by the discriminative confidence estimation networks, while the discriminative confidence estimation networks are trained to distinguish the confident and unconfident disparities. In addition, to fully exploit complementary information of matching cost, disparity, and color image in confidence estimation, we present a dynamic fusion module. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks including real driving scenes. Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Multi-Modal Recurrent Attention Networks for Facial Expression RecognitionabstractRecent deep neural networks based methods have achieved state-of-the-art performance on various facial expression recognition tasks. Despite such progress, previous researches for facial expression recognition have mainly focused on analyzing color recording videos only. However, the complex emotions expressed by people with different skin colors under different lighting conditions through dynamic facial expressions can be fully understandable by integrating information from multi-modal videos. We present a novel method to estimate dimensional emotion states, where color, depth, and thermal recording videos are used as a multi-modal input. Our networks, called multi-modal recurrent attention networks (MRAN), learn spatiotemporal attention volumes to robustly recognize the facial expression based on attention-boosted feature volumes. We leverage the depth and thermal sequences as guidance priors for color sequence to selectively focus on emotional discriminative regions. We also introduce a novel benchmark for multi-modal facial expression recognition, termed as multi-modal arousal-valence facial expression recognition (MAVFER), which consists of color, depth, and thermal recording videos with corresponding continuous arousal-valence scores. The experimental results show that our method can achieve the state-of-the-art results in dimensional facial expression recognition on color recording datasets including RECOLA, SEWA and AFEW, and a multi-modal recording dataset including MAVFER. Jiyoung Lee 0005, Sunok Kim, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2019 | LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence EstimationabstractWe present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. The attention inference networks encode the importance of tri-modal confidence features and then concatenate them using the attention maps in an adaptive and dynamic fashion. This enables us to make an optimal fusion of the heterogeneous features, compared to a simple concatenation technique that is commonly used in conventional approaches. In addition, to encode the confidence features with locally-varying receptive fields, the scale inference networks learn the scale map and warp the fused confidence features through convolutional spatial transformer networks. Finally, the confidence map is progressively estimated in the recursive refinement networks to enforce a spatial context and local consistency. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks. Sunok Kim, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
CVPR | 1 |
| 2019 | Semantic Attribute Matching NetworksabstractWe present semantic attribute matching networks (SAM-Net) for jointly establishing correspondences and transferring attributes across semantically similar images, which intelligently weaves the advantages of the two tasks while overcoming their limitations. SAM-Net accomplishes this through an iterative process of establishing reliable correspondences by reducing the attribute discrepancy between the images and synthesizing attribute transferred images using the learned correspondences. To learn the networks using weak supervisions in the form of image pairs, we present a semantic attribute matching loss based on the matching similarity between an attribute transferred source feature and a warped target feature. With SAM-Net, the state-of-the-art performance is attained on several benchmarks for semantic matching and attribute transfer. Seungryong Kim, Dongbo Min, Somi Jeong, Sunok Kim, Sangryul Jeon, Kwanghoon Sohn |
CVPR | 4 |
| 2019 | Context-Aware Emotion Recognition NetworksabstractTraditional techniques for emotion recognition have focused on the facial expression analysis only, thus providing limited ability to encode context that comprehensively represents the emotional responses. We present deep networks for context-aware emotion recognition, called CAER-Net, that exploit not only human facial expression but also context information in a joint and boosting manner. The key idea is to hide human faces in a visual scene and seek other contexts based on an attention mechanism. Our networks consist of two sub-networks, including two-stream encoding networks to separately extract the features of face and context regions, and adaptive fusion networks to fuse such features in an adaptive fashion. We also introduce a novel benchmark for context-aware emotion recognition, called CAER, that is appropriate than existing benchmarks both qualitatively and quantitatively. On several benchmarks, CAER-Net proves the effect of context for emotion recognition. Our dataset is available at http://caer-dataset.github.io. Jiyoung Lee 0005, Seungryong Kim, Sunok Kim, Jungin Park, Kwanghoon Sohn |
ICCV | 3 |
| 2019 | Unified Confidence Estimation Networks for Robust Stereo MatchingabstractWe present a deep architecture that estimates a stereo confidence, which is essential for improving the accuracy of stereo matching algorithms. In contrast to existing methods based on deep convolutional neural networks (CNNs) that rely on only one of the matching cost volume or estimated disparity map, our network estimates the stereo confidence by using the two heterogeneous inputs simultaneously. Specifically, the matching probability volume is first computed from the matching cost volume with residual networks and a pooling module in a manner that yields greater robustness. The confidence is then estimated through a unified deep network that combines confidence features extracted both from the matching probability volume and its corresponding disparity. In addition, our method extracts the confidence features of the disparity map by applying multiple convolutional filters with varying sizes to an input disparity map. To learn our networks in a semi-supervised manner, we propose a novel loss function that use confident points to compute the image reconstruction loss. To validate the effectiveness of our method in a disparity post-processing step, we employ three post-processing approaches; cost modulation, ground control points-based propagation, and aggregated ground control points-based propagation. Experimental results demonstrate that our method outperforms state-of-the-art confidence estimation methods on various benchmarks. Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2018 | Spatiotemporal Attention Based Deep Neural Networks for Emotion RecognitionabstractWe propose a spatiotemporal attention based deep neural networks for dimensional emotion recognition in facial videos. To learn the spatiotemporal attention that selectively focuses on emotional sailient parts within facial videos, we formulate the spatiotemporal encoder-decoder network using Convolutional LSTM (ConvLSTM) modules, which can be learned implicitly without any pixel-level annotations. By leveraging the spatiotemporal attention, we also formulate the 3D convolutional neural networks (3D-CNNs) to robustly recognize the dimensional emotion in facial videos. The experimental results show that our method can achieve the state-of-the-art results in dimensional emotion recognition with the highest concordance correlation coefficient (CCC) on RECOLA and AV+EC 2017 dataset. Jiyoung Lee 0005, Sunok Kim, Seungryong Kim, Kwanghoon Sohn |
ICASSP | 2 |
| 2017 | Deep stereo confidence prediction for depth estimationabstractWe present a novel method that predicts a confidence to improve the accuracy of an estimated depth map in stereo matching. In contrast to existing learning based approaches relying on hand-crafted confidence features, we cast this problem into a convolutional neural network, learned using both a matching cost volume and its associated disparity map. As the size of the matching cost volume varies depending on a search range of stereo image pairs, we propose to use a top-K matching probability volume layer so that an input size for convolutional layers remains unchanged. Experimental results demonstrate that the proposed method outperforms the state-of-the-art confidence estimation approaches on various benchmarks. Sunok Kim, Dongbo Min, Bumsub Ham, Seungryong Kim, Kwanghoon Sohn |
ICIP | 1 |
| 2017 | Modality-Invariant Image Classification Based on Modality Uniqueness and Dictionary LearningabstractWe present a unified framework for the image classification of image sets taken under varying modality conditions. Our method is motivated by a key observation that the image feature distribution is simultaneously influenced by the semantic-class and the modality category label, which limits the performance of conventional methods for that task. With this insight, we introduce modality uniqueness as a discriminative weight that divides each modality cluster from all other clusters. By leveraging the modality uniqueness, our framework is formulated as unsupervised modality clustering and classifier learning based on modality-invariant similarity kernel. Specifically, in the assignment step, each training image is first assigned to the most similar cluster according to its modality. In the update step, based on the current cluster hypothesis, the modality uniqueness and the sparse dictionary are updated. These two steps are formulated in an iterative manner. Based on the final clusters, a modality-invariant marginalized kernel is then computed, where the similarities between the reconstructed features of each modality are aggregated across all clusters. Our framework enables the reliable inference of semantic-class category for an image, even across large photometric variations. Experimental results show that our method outperforms conventional methods on various benchmarks, such as landmark identification under severely varying weather conditions, domain-adapting image classification, and RGB and near-infrared image classification. Seungryong Kim, Rui Cai 0002, Kihong Park, Sunok Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 4 |
| 2017 | Feature Augmentation for Learning Confidence Measure in Stereo MatchingabstractConfidence estimation is essential for refining stereo matching results through a post-processing step. This problem has recently been studied using a learning-based approach, which demonstrates a substantial improvement on conventional simple non-learning based methods. However, the formulation of learning-based methods that individually estimates the confidence of each pixel disregards spatial coherency that might exist in the confidence map, thus providing a limited performance under challenging conditions. Our key observation is that the confidence features and resulting confidence maps are smoothly varying in the spatial domain, and highly correlated within the local regions of an image. We present a new approach that imposes spatial consistency on the confidence estimation. Specifically, a set of robust confidence features is extracted from each superpixel decomposed using the Gaussian mixture model, and then these features are concatenated with pixel-level confidence features. The features are then enhanced through adaptive filtering in the feature domain. In addition, the resulting confidence map, estimated using the confidence features with a random regression forest, is further improved through K-nearest neighbor based aggregation scheme on both pixel- and superpixel-level. To validate the proposed confidence estimation scheme, we employ cost modulation or ground control points based optimization in stereo matching. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches on various benchmarks including challenging outdoor scenes. Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2015 | Learning depth from a single image using visual-depth wordsabstractEstimating depth from a single monocular image is a fundamental problem in computer vision. Traditional methods for such estimation usually require complicated and sometimes labor-intensive processing. In this paper, we propose a new perspective for this problem and suggest a new gradient-domain learning framework which is much simpler and more efficient. Inspired by the observation that there is substantial co-occurrence of image edges and depth discontinuities in natural scenes, we learn the relationship between local appearance features and corresponding depth gradients by making use of the K-means clustering algorithm within the image feature space. We then encode each cluster centroid with its associated depth gradients, which defines visual-depth words that model the image-depth relationship very well. This enables one to estimate the scene depth for an arbitrary image by simply selecting proper depth gradients from a compact dictionary of visual-depth words, followed by a Poisson surface reconstruction. Experimental results demonstrate that the proposed gradient-domain approach outperforms state-of-the-art methods both qualitatively and quantitatively and is generic over (unseen) scene categories which are not used for training. Sunok Kim, Sunghwan Choi, Kwanghoon Sohn |
ICIP | 1 |