Kwanghoon Sohn

dblp:21/2373 · DBLP profile ↗
← Back
217ranked-venue papers
7as first author
67since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 161 · 4 first-author · 41 since 2021Artificial intelligence and machine learning · 90 · 1 first-author · 46 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 9 since 2021Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 3D semantic image synthesis with geometric and semantic consistency
Jihyun Kim 0009, Changjae Oh, Hoseok Do, Sunghwan Choi, Kwanghoon Sohn
Expert Syst. Appl.5
2026 Geospatial Domain Adaptation With Truncated Parameter-Efficient Fine-Tuning
abstract
Parameter-efficient fine-tuning (PEFT) adapts large pre-trained foundation models to downstream tasks, such as remote sensing scene classification, by learning a small set of additional parameters while keeping the pre-trained parameters frozen. While PEFT offers substantial training efficiency over full fine-tuning, it still incurs high inference costs due to reliance on both pre-trained and task-specific parameters. To address this limitation, we propose a novel PEFT approach with model truncation, termed TruncPEFT, enabling efficiency gains to persist during inference. Observing that predictions from final and intermediate layers often exhibit high agreement, we truncate a set of final layers and replace them with a lightweight attention module. Additionally, we introduce a token dropping strategy to mitigate interclass interference, reducing the model’s sensitivity to visual similarities between different classes in remote sensing data. Extensive experiments on seven remote sensing scene classification datasets demonstrate the effectiveness of the proposed method, significantly improving training, inference, and GPU memory efficiencies while achieving comparable or even better performance than prior PEFT methods and full fine-tuning.
Kwonyoung Kim, Jungin Park, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.3
2026 Diffusion model-based data augmentation for land cover segmentation in Pol-SAR imagery
Keunhoon Choi, Sunok Kim, Kwanghoon Sohn
Pattern Recognit.3
2026 Mixture of zero-cost proxies for training-free network architecture search
Junghyup Lee, Jeimin Jeon, Kwanghoon Sohn, Bumsub Ham
Pattern Recognit.3
2025 Mixture of Submodules for Domain Adaptive Person Search
abstract
Existing technique on domain adaptive person search commonly utilizes the unified framework for jointly localizing and identifying the person across domains. This framework, however, inevitably results in the gradient conflict problem, particularly in cross-domain scenarios with contradictory objectives, as the unified framework employs shared parameters to simultaneously address person detection and re-identification tasks across the domains. To over-come this, we present a novel mixture of submodules framework, dubbed MoS, that dynamically modulates the combination of submodules depending on the specific task to perform person detection and re-identification, separately. We further design the mixtures of submodules that vary depending on the domain, enabling domain-specific knowledge transfer. Especially, we decompose the main model into several submodules and employ diverse mixtures of submodules that vary depending on the tasks and domains through the conditional routing policy. In addition, we also present counterpart domain sample generation that synthesizes the augmented sample and uses them to learn domain invariant representation for person re-identification through the contrastive domain alignment. We conduct experiments to demonstrate the effectiveness of our MoS over the existing domain adaptive person search method and provide ablation studies.
Seungryong Kim, Kwanghoon Sohn
CVPR3
2025 Faster Parameter-Efficient Tuning with Token Redundancy Reduction
abstract
Parameter-efficient tuning (PET) aims to transfer pre-trained foundation models to downstream tasks by learning a small number of parameters. Compared to traditional fine-tuning, which updates the entire model, PET significantly reduces storage and transfer costs for each task regardless of exponentially increasing pre-trained model capacity. However, most PET methods inherit the inference latency of their large backbone models and often introduce additional computational overhead due to additional modules (e.g. adapters), limiting their practicality for compute-intensive applications. In this paper, we propose Faster Parameter-Efficient Tuning (FPET), a novel approach that enhances inference speed and training efficiency while maintaining high storage efficiency. Specifically, we introduce a plug-and-play token redundancy reduction module delicately designed for PET. This module refines tokens from the self-attention layer using an adapter to learn the accurate similarity between tokens and cuts off the tokens through a fully-differentiable token merging strategy, which uses a straight-through estimator for optimal token reduction. Experimental results prove that our FPET achieves faster inference and higher memory efficiency than the pre-trained backbone while keeping competitive performance on par with state-of-the-art PET methods. The code is available at https://github.com/kyk120/fpet.
Kwonyoung Kim, Jungin Park, Jin Kim 0005, Hyeongjun Kwon, Kwanghoon Sohn
CVPR5
2025 Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
abstract
View-invariant representation learning from egocentric (first-person, ego) and exocentric (third-person, exo) videos is a promising approach toward generalizing video understanding systems across multiple viewpoints. However, this area has been underexplored due to the substantial differences in perspective, motion patterns, and context between ego and exo views. In this paper, we propose a novel masked ego-exo modeling that promotes both causal temporal dynamics and cross-view alignment, called Bootstrap Your Own Views (BYOV), for fine-grained view-invariant video representation learning from unpaired ego-exo videos. We highlight the importance of capturing the compositional nature of human actions as a basis for robust cross-view understanding. Specifically, self-view masking and cross-view masking predictions are designed to learn view-invariant and powerful representations concurrently. Experimental results demonstrate that our BYOV significantly surpasses existing approaches with notable gains across all metrics in four downstream ego-exo video tasks. The code is available at https://github.com/park-jungin/byov.
Jungin Park, Jiyoung Lee 0005, Kwanghoon Sohn
CVPR3
2025 Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
abstract
Abstract Video summarization aims to select keyframes that are visually diverse and can represent the whole story of a given video. Previous approaches have focused on global interlinkability between frames in a video by temporal modeling. However, fine-grained visual entities, such as objects, are also highly related to the main content of the video. Moreover, language-guided video summarization, which has recently been studied, requires a comprehensive linguistic understanding of complex real-world videos. To consider how all the objects are semantically related to each other, this paper regards video summarization as a language-guided spatiotemporal graph modeling problem. We present recursive spatiotemporal graph networks, called VideoGraph , which formulate the objects and frames as nodes of the spatial and temporal graphs, respectively. The nodes in each graph are connected and aggregated with graph edges, representing the semantic relationships between the nodes. To prevent the edges from being configured with visual similarity, we incorporate language queries derived from the video into the graph node representations, enabling them to contain semantic knowledge. In addition, we adopt a recursive strategy to refine initial graphs and correctly classify each frame node as a keyframe. In our experiments, VideoGraph achieves state-of-the-art performance on several benchmarks for generic and query-focused video summarization in both supervised and unsupervised manners. The code is available at https://github.com/park-jungin/videograph .
Jungin Park, Kwanghoon Sohn
Int. J. Comput. Vis.3
2025 Source-Free Domain Adaptation for Remote Sensing Object Detection Using Low-Confidence Pseudolabels
Jin Kim 0005, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.5
2025 Prototype-Guided Attention Distillation for Discriminative Person Search
abstract
Person search aims to localize a person of interest in a large image gallery captured by multiple, non-overlapping cameras. Prevalent unified methods have suffered from (1) noisy proposals with mis-detection and occlusion, and (2) large appearance variation within a class, which deteriorates the prototype-based metric learning. To address these problems, we introduce a Prototype-guided Attention Distillation, shortly PAD, which exploits a prototype (a typical representation of an identity) as a guidance to the attention module to consistently highlight identity-inherent regions across different poses. To utilize the knowledge encoded in prototypes for matching unseen IDs, PAD conducts attention distillation to guide student Re-ID queries by deeply mimicking attention maps from the prototype query. Additionally, to address large intra-class variation induced by pose or camera views, we extend PAD with multiple part prototypes representing consistent local regions across different instances. Furthermore, we exploit an adaptive momentum strategy for robust attention distillation in PAD to update more distinct prototypes. Extensive experiments conducted on CUHK-SYSU and PRW demonstrate the effectiveness of PAD, showcasing state-of-the-art performance. Moreover, our distilled attention surprisingly highlights distinguished multiple regions for person search.
Hanjae Kim, Jiyoung Lee 0005, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Diffusion-Driven GAN Inversion for Multi-Modal Face Image Generation
abstract
We present a new multimodal face image generation method that converts a text prompt and a visual input, such as a semantic mask or scribble map, into a photorealistic face image. To do this, we combine the strengths of Generative Adversarial networks (GANs) and diffusion models (DMs) by employing the multimodal features in the DM into the latent space of the pretrained GANs. We present a simple mapping and a style modulation network to link two models and convert meaningful representations in feature maps and attention maps into latent codes. With GAN inversion, the estimated latent codes can be used to generate 2D or 3D-aware facial images. We further present a multi-step training strategy that reflects textual and structural representations into the generated image. Our proposed network produces realistic 2D, multi-view, and stylized face images, which align well with inputs. We validate our method by using pretrained 2D and 3D GANs, and our results outperform existing methods. Our project page is available at https://github.com/1211sh/Diffusion-driven_GAN-Inversion/.
Jihyun Kim 0009, Changjae Oh, Hoseok Do, Kwanghoon Sohn
CVPR5
2024 Improving Visual Recognition with Hyperbolical Visual Hierarchy Mapping
abstract
Visual scenes are naturally organized in a hierarchy, where a coarse semantic is recursively comprised of several fine details. Exploring such a visual hierarchy is crucial to recognize the complex relations of visual elements, leading to a comprehensive scene understanding. In this paper, we propose a Visual Hierarchy Mapper (Hi-Mapper), a novel approach for enhancing the structured understanding of the pre-trained Deep Neural Networks (DNNs). Hi-Mapper investigates the hierarchical organization of the visual scene by 1) pre-defining a hierarchy tree through the encapsulation of probability densities; and 2) learning the hierarchical relations in hyperbolic space with a novel hierarchical contrastive loss. The pre-defined hierarchy tree recursively interacts with the visual features of the pre-trained DNNs through hierarchy decomposition and encoding procedures, thereby effectively identifying the visual hierarchy and enhancing the recognition of an entire scene. Extensive experiments demonstrate that Hi-Mapper significantly enhances the representation capability of DNNs, leading to an improved performance on various tasks, including image classification and dense prediction tasks. The code is available at https://github.com/kwonjunn0l/Hi-Mapper.
Hyeongjun Kwon, Jinhyun Jang, Jin Kim 0005, Kwonyoung Kim, Kwanghoon Sohn
CVPR5
2024 EBDM: Exemplar-Guided Image Translation with Brownian-Bridge Diffusion Models
Eungbean Lee, Somi Jeong, Kwanghoon Sohn
ECCV (13)3
2024 Enhancing Source-Free Domain Adaptive Object Detection with Low-Confidence Pseudo Label Distillation
Ilhoon Yoon, Hyeongjun Kwon, Jin Kim 0005, Hyunsung Jang, Kwanghoon Sohn
ECCV (84)6
2024 Improving Self-Supervised Vision Transformers for Visual Control
abstract
Despite the tremendous success of vision transformer (ViT) architectures in a broad range of computer vision tasks, the potential of ViT for vision-based deep reinforcement learning (RL) has not been fully explored yet. To improve the performance of the ViT model in visual RL, we propose a simple yet effective approach for self-supervised learning by utilizing the structural capability of a single ViT model, which can learn multiple, distinct representations through extra learnable token embeddings. To this end, in addition to an RL token used for RL input, which corresponds to the classification token in computer vision, we introduce additional extra tokens that are tailored to two auxiliary self-supervised tasks specialized to learn visual and environmental dynamics representations. By interacting with embeddings of these extra tokens through self-attention, our approach provides additional learning signals to the ViT encoder, enabling it to learn more comprehensive representations that are beneficial to RL tasks. In experiments on benchmarks including the DeepMind Control Suite (DMControl) and Atari games, we demonstrate that the proposed approach outperforms the baselines that utilize ViT encoders, particularly achieving state-of-the-art performance in 4 out of 5 tasks in DMControl.
Wonil Song, Kwanghoon Sohn, Dongbo Min
ICIP2
2024 Bridging Vision and Language Spaces with Assignment Prediction
abstract
This paper introduces VLAP, a novel approach that bridges pretrained vision models and large language models (LLMs) to make frozen LLMs understand the visual world. VLAP transforms the embedding space of pretrained vision models into the LLMs' word embedding space using a single linear layer for efficient and general-purpose visual and language understanding. Specifically, we harness well-established word embeddings to bridge two modality embedding spaces. The visual and text representations are simultaneously assigned to a set of word embeddings within pretrained LLMs by formulating the assigning procedure as an optimal transport problem. We predict the assignment of one modality from the representation of another modality data, enforcing consistent assignments for paired multimodal data. This allows vision and language representations to contain the same information, grounding the frozen LLMs' word embedding space in visual data. Moreover, a robust semantic taxonomy of LLMs can be preserved with visual data since the LLMs interpret and reason linguistic information from correlations between word embeddings. Experimental results show that VLAP achieves substantial improvements over the previous linear transformation-based approaches across a range of vision-language tasks, including image captioning, visual question answering, and cross-modal retrieval. We also demonstrate the learned visual representations hold a semantic taxonomy of LLMs, making visual semantic arithmetic possible.
Jungin Park, Jiyoung Lee 0005, Kwanghoon Sohn
ICLR3
2024 A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations
abstract
In the rapidly evolving domain of vision-based deep reinforcement learning (RL), a pivotal challenge is to achieve generalization capability to dynamic environmental changes reflected in visual observations. Our work delves into the intricacies of this problem, identifying two key issues that appear in previous approaches for visual RL generalization: (i) imbalanced saliency and (ii) observational overfitting. Imbalanced saliency is a phenomenon where an RL agent disproportionately identifies salient features across consecutive frames in a frame stack. Observational overfitting occurs when the agent focuses on certain background regions rather than task-relevant objects. To address these challenges, we present a simple yet effective framework for generalization in visual RL (SimGRL) under dynamic scene perturbations. First, to mitigate the imbalanced saliency problem, we introduce an architectural modification to the image encoder to stack frames at the feature level rather than the image level. Simultaneously, to alleviate the observational overfitting problem, we propose a novel technique called shifted random overlay augmentation, which is specifically designed to learn robust representations capable of effectively handling dynamic visual scenes. Extensive experiments demonstrate the superior generalization capability of SimGRL, achieving state-of-the-art performance in benchmarks including the DeepMind Control Suite.
Wonil Song, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
NeurIPS3
2024 Layer-wise Auto-Weighting for Non-Stationary Test-Time Adaptation
abstract
Given the inevitability of domain shifts during inference in real-world applications, test-time adaptation (TTA) is essential for model adaptation after deployment. However, the real-world scenario of continuously changing target distributions presents challenges including catastrophic forgetting and error accumulation. Existing TTA methods for non-stationary domain shifts, while effective, incur excessive computational load, making them impractical for on-device settings. In this paper, we introduce a layer-wise auto-weighting algorithm for continual and gradual TTA that autonomously identifies layers for preservation or concentrated adaptation. By leveraging the Fisher Information Matrix (FIM), we first design the learning weight to selectively focus on layers associated with log-likelihood changes while preserving unrelated ones. Then, we further propose an exponential min-max scaler to make certain layers nearly frozen while mitigating outliers. This minimizes forgetting and error accumulation, leading to efficient adaptation to non-stationary target distribution. Experiments on CIFAR-10C, CIFAR-100C, and ImageNet-C show our method outperforms conventional continual and gradual TTA approaches while significantly reducing computational load, highlighting the importance of FIM-based learning weight in adapting to continuously or gradually shifting target domains.1
Jin Kim 0005, Hyeongjun Kwon, Ilhoon Yoon, Kwanghoon Sohn
WACV5
2024 Discriminative action tubelet detector for weakly-supervised action detection
Jiyoung Lee 0005, Seungryong Kim, Sunok Kim, Kwanghoon Sohn
Pattern Recognit.4
2023 Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning
abstract
Recent vision-based reinforcement learning (RL) methods have found extracting high-level features from raw pixels with self-supervised learning to be effective in learning policies. However, these methods focus on learning global representations of images, and disregard local spatial structures present in the consecutively stacked frames. In this paper, we propose a novel approach, termed self-supervised Paired Similarity Representation Learning (PSRL) for effectively encoding spatial structures in an unsupervised manner. Given the input frames, the latent volumes are first generated individually using an encoder, and they are used to capture the variance in terms of local spatial structures, i.e., correspondence maps among multiple frames. This enables for providing plenty of fine-grained samples for training the encoder of deep RL. We further attempt to learn the global semantic representations in the action aware transform module that predicts future state representations using action vectors as a medium. The proposed method imposes similarity constraints on the three latent volumes; transformed query representations by estimated pixel-wise correspondence, predicted query representations from the action aware transform model, and target representations of future state, guiding action aware transform with locality-inherent volume. Experimental results on complex tasks in Atari Games and DeepMind Control Suite demonstrate that the RL methods are significantly boosted by the proposed self-supervised learning of paired similarity representations.
Hyesong Choi, Hunsang Lee, Wonil Song, Sangryul Jeon, Kwanghoon Sohn, Dongbo Min
CVPR5
2023 PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-Identification
abstract
Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains unexplored. In this paper, we present a novel data augmentation technique, dubbed PartMix, that synthesizes the augmented samples by mixing the part descriptors across the modalities to improve the performance of part-based VI-ReID models. Especially, we synthesize the positive and negative samples within the same and across different identities and regularize the backbone model through contrastive learning. In addition, we also present an entropy-based mining strategy to weaken the adverse impact of unreliable positive and negative samples. When incorporated into existing part-based VI-ReID model, PartMix consistently boosts the performance. We conduct experiments to demonstrate the effectiveness of our PartMix over the existing VI-ReID methods and provide ablation studies.
Seungryong Kim, Jungin Park, Seongheon Park, Kwanghoon Sohn
CVPR5
2023 Probabilistic Prompt Learning for Dense Prediction
abstract
Recent progress in deterministic prompt learning has become a promising alternative to various downstream vision tasks, enabling models to learn powerful visual representations with the help of pre-trained vision-language models. However, this approach results in limited performance for dense prediction tasks that require handling more complex and diverse objects, since a single and deterministic description cannot sufficiently represent the entire image. In this paper, we present a novel probabilistic prompt learning to fully exploit the vision-language knowledge in dense prediction tasks. First, we introduce learnable class-agnostic attribute prompts to describe universal attributes across the object class. The attributes are combined with class information and visual-context knowledge to define the class-specific textual distribution. Text representations are sampled and used to guide the dense prediction task using the probabilistic pixel-text matching loss, enhancing the stability and generalization capability of the proposed method. Extensive experiments on different dense prediction tasks and ablation studies demonstrate the effectiveness of our proposed method.
Hyeongjun Kwon, Taeyong Song, Somi Jeong, Jin Kim 0005, Jinhyun Jang, Kwanghoon Sohn
CVPR6
2023 Dual-Path Adaptation from Image to Video Transformers
abstract
In this paper, we efficiently transfer the surpassing representation power of the vision foundation models, such as ViT and Swin, for video understanding with only a few trainable parameters. Previous adaptation methods have simultaneously considered spatial and temporal modeling with a unified learnable module but still suffered from fully leveraging the representative capabilities of image transformers. We argue that the popular dual-path (two-stream) architecture in video models can mitigate this problem. We propose a novel DualPath adaptation separated into spatial and temporal adaptation paths, where a lightweight bottleneck adapter is employed in each transformer block. Especially for temporal dynamic modeling, we incorporate consecutive frames into a grid-like frameset to precisely imitate vision transformers' capability that extrapolates relationships between tokens. In addition, we extensively investigate the multiple baselines from a unified perspective in video understanding and compare them with DualPath. Experimental results on four action recognition benchmarks prove that pretrained image transformers with DualPath can be effectively generalized beyond the data domain.
Jungin Park, Jiyoung Lee 0005, Kwanghoon Sohn
CVPR3
2023 Unsupervised Deep Asymmetric Stereo Matching with Spatially-Adaptive Self-Similarity
abstract
Unsupervised stereo matching has received a lot of attention since it enables the learning of disparity estimation without ground-truth data. However, most of the unsupervised stereo matching algorithms assume that the left and right images have consistent visual properties, i.e., symmetric, and easily fail when the stereo images are asymmetric. In this paper, we present a novel spatially-adaptive self-similarity (SASS) for unsupervised asymmetric stereo matching. It extends the concept of self-similarity and generates deep features that are robust to the asymmetries. The sampling patterns to calculate self-similarities are adaptively generated throughout the image regions to effectively encode diverse patterns. In order to learn the effective sampling patterns, we design a contrastive similarity loss with positive and negative weights. Consequently, SASS is further encouraged to encode asymmetry-agnostic features, while maintaining the distinctiveness for stereo correspondence. We present extensive experimental results including ablation studies and comparisons with different methods, demonstrating effectiveness of the proposed method under resolution and noise asymmetries.
Taeyong Song, Sunok Kim, Kwanghoon Sohn
CVPR3
2023 Knowing Where to Focus: Event-aware Transformer for Video Grounding
abstract
Recent DETR-based video grounding models have made the model directly predict moment timestamps without any hand-crafted components, such as a pre-defined proposal or non-maximum suppression, by learning moment queries. However, their input-agnostic moment queries inevitably overlook an intrinsic temporal structure of a video, providing limited positional information. In this paper, we formulate an event-aware dynamic moment query to enable the model to take the input-specific content and positional information of the video into account. To this end, we present two levels of reasoning: 1) Event reasoning that captures distinctive event units constituting a given video using a slot attention mechanism; and 2) moment reasoning that fuses the moment queries with a given sentence through a gated fusion transformer layer and learns interactions between the moment queries and video-sentence representations to predict moment timestamps. Extensive experiments demonstrate the effectiveness and efficiency of the event-aware dynamic moment queries, outperforming state-of-the-art approaches on several video grounding benchmarks. The code is publicly available at https://github.com/jinhyunj/EaTR.
Jinhyun Jang, Jungin Park, Jin Kim 0005, Hyeongjun Kwon, Kwanghoon Sohn
ICCV5
2023 Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning
abstract
Compositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer from grasping the contextuality between attribute and object, as well as the discriminability of visual features, and the long-tailed distribution of real-world compositional data. We propose a simple and scalable framework called Composition Transformer (CoT) to address these issues. CoT employs object and attribute experts in distinctive manners to generate representative embeddings, using the visual network hierarchically. The object expert extracts representative object embeddings from the final layer in a bottom-up manner, while the attribute expert makes attribute embeddings in a top-down manner with a proposed object-guided attention module that models contextuality explicitly. To remedy biased prediction caused by imbalanced data distribution, we develop a simple minority attribute augmentation (MAA) that synthesizes virtual samples by mixing two images and oversampling minority attribute classes. Our method achieves SoTA performance on several benchmarks, including MIT-States, C-GQA, and VAW-CZSL. We also demonstrate the effectiveness of CoT in improving visual discrimination and addressing the model bias from the imbalanced data distribution. The code is available at https://github.com/HanjaeKim98/CoT.
Hanjae Kim, Jiyoung Lee 0005, Seongheon Park, Kwanghoon Sohn
ICCV4
2023 Face Photo-Sketch Synthesis Via Domain-Invariant Feature Embedding
abstract
Face photo-sketch synthesis involves transforming photos into sketches and vice versa. A well-transformed image should preserve its original identity characteristics and naturalness. However, identity preservation remains a challenge because of the large discrepancy between the photo and sketch domains. To this end, we propose a novel face photo-sketch synthesis framework that uses domain-invariant feature embedding (DIFE). The DIFE framework generates images assuming the domain-invariant feature of an image pair for the same person to be the identity information. A joint feature embedding module considers latent features from two different domains as input and transfers them into the domain-invariant latent space. Subsequently, a semantic-aware decoder completes the desired image guided by multiscale facial parsing masks. Experimental results demonstrate that the DIFE method outperforms state-of-the-art approaches visually and perceptually.
Yeji Choi, Kwanghoon Sohn, Ig-Jae Kim
ICIP2
2023 Language-free Training for Zero-shot Video Grounding
abstract
Given an untrimmed video and a language query depicting a specific temporal moment in the video, video grounding aims to localize the time interval by understanding the text and video simultaneously. One of the most challenging issues is an extremely time- and cost-consuming annotation collection, including video captions in a natural language form and their corresponding temporal regions. In this paper, we present a simple yet novel training framework for video grounding in the zero-shot setting, which learns a network with only video data without any annotation. Inspired by the recent language-free paradigm, i.e. training without language data, we train the network without compelling the generation of fake (pseudo) text queries into a natural language form. Specifically, we propose a method for learning a video grounding model by selecting a temporal interval as a hypothetical correct answer and considering the visual feature selected by our method in the interval as a language feature, with the help of the well-aligned visual-language space of CLIP. Extensive experiments demonstrate the prominence of our language-free training framework, outperforming the existing zero-shot video grounding method and even several weakly-supervised approaches with large margins on two standard datasets.
Dahye Kim 0004, Jungin Park, Jiyoung Lee 0005, Seongheon Park, Kwanghoon Sohn
WACV5
2023 Normality Guided Multiple Instance Learning for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised Video Anomaly Detection (wVAD) aims to distinguish anomalies from normal events based on video-level supervision. Most existing works utilize Multiple Instance Learning (MIL) with ranking loss to tackle this task. These methods, however, rely on noisy predictions from a MIL-based classifier for target instance selection in ranking loss, degrading model performance. To overcome this problem, we propose Normality Guided Multiple Instance Learning (NG-MIL) framework, which encodes diverse normal patterns from noise-free normal videos into prototypes for constructing a similarity-based classifier. By ensembling predictions of two classifiers, our method could refine the anomaly scores, reducing training instability from weak labels. Moreover, we introduce normality clustering and normality guided triplet loss constraining inner bag instances to boost the effect of NG-MIL and increase the discriminability of classifiers. Extensive experiments on three public datasets (ShanghaiTech, UCF-Crime, XD-Violence) demonstrate that our method is comparable to or better than existing weakly supervised methods, achieving state-of-the-art results.
Seongheon Park, Hanjae Kim, Dahye Kim 0004, Kwanghoon Sohn
WACV5
2023 Learning disentangled skills for hierarchical reinforcement learning through trajectory autoencoder with weak labels
Wonil Song, Sangryul Jeon, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
Expert Syst. Appl.4
2023 Learning Semantic Keypoints for Object Detection in Aerial Images
abstract
Object detection in aerial images has achieved remarkable progress with the advent of deep convolutional neural networks (CNNs). It is, however, still a challenging task since the objects in aerial images are arbitrarily oriented and often densely packed. In this letter, we propose a novel method for oriented object detection in aerial images that represents objects as rotation equivariant semantic keypoints. Unlike conventional methods that represent object rotation according to angles from each axis in the Cartesian coordinate system, we represent object using a canonical orientation to ensure rotation equivariance. We accomplish this by representing an object as semantic keypoints, where each keypoint of the object consistently corresponds to the semantic part, regardless of rotation variation. To this end, we define the “head” point of the object as the canonical orientation and the remaining bounding box vectors as semantic keypoints in clockwise order. To discriminate visual attributes between different categories, we further use category-specific semantic keypoints, so that object classification and localization can be jointly solved in a cooperative manner. Our experiments demonstrate the effectiveness of rotation equivariant semantic keypoints on oriented object detection.
Sunghun Joung, Taeyong Song, Hanjae Kim, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.5
2023 Stereo Confidence Estimation via Locally Adaptive Fusion and Knowledge Distillation
abstract
Stereo confidence estimation aims to estimate the reliability of the estimated disparity by stereo matching. Different from the previous methods that exploit the limited input modality, we present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. Moreover, we propose a knowledge distillation framework to learn more compact confidence estimation networks as student networks. By transferring the knowledge from LAF-Net as teacher networks, the student networks that solely take as input a disparity can achieve comparable performance. To transfer more informative knowledge, we also propose a module to learn the locally-varying temperature in a softmax function. We further extend this framework to a multiview scenario. Experimental results show that LAF-Net and its variations outperform the state-of-the-art stereo confidence methods on various benchmarks.
Sunok Kim, Seungryong Kim, Dongbo Min, Pascal Frossard, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 COAT: Correspondence-driven Object Appearance Transfer
Sangryul Jeon, Zhe Lin 0001, Scott Cohen, Zhihong Ding, Kwanghoon Sohn
BMVC6
2022 Pin the Memory: Learning to Generalize Semantic Segmentation
abstract
The rise of deep neural networks has led to several break-throughs for semantic segmentation. In spite of this, a model trained on source domain often fails to work properly in new challenging domains, that is directly concerned with the generalization capability of the model. In this paper, we present a novel memory-guided domain generalization method for semantic segmentation based on meta-learning framework. Especially, our method abstracts the conceptual knowledge of semantic classes into categorical memory which is constant beyond the domains. Upon the meta-learning concept, we repeatedly train memory-guided networks and simulate virtual test to 1) learn how to memorize a domain-agnostic and distinct information of classes and 2) offer an externally settled memory as a class-guidance to reduce the ambiguity of representation in the test data of arbitrary unseen domain. To this end, we also propose memory divergence and feature cohesion losses, which encourage to learn memory reading and update processes for category-aware domain generalization. Extensive experiments for semantic segmentation demonstrate the superior generalization capability of our method over state-of-the-art works on various benchmarks.11https://github.com/Genie-Kim/PintheMemory
Jin Kim 0005, Jiyoung Lee 0005, Jungin Park, Dongbo Min, Kwanghoon Sohn
CVPR5
2022 KNN Local Attention for Image Restoration
abstract
Recent works attempt to integrate the non-local operation with CNNs or Transformer, achieving remarkable performance in image restoration tasks. The global similarity, however, has the problems of the lack of locality and the high computational complexity that is quadratic to an input resolution. The local attention mechanism alleviates these issues by introducing the inductive bias of the locality with convolution-like operators. However, by focusing only on adjacent positions, the local attention suffers from an insufficient receptive field for image restoration. In this paper, we propose a new attention mechanism for image restoration, called k-NN Image Transformer (KiT), that rectifies the above mentioned limitations. Specifically, the KiT groups k-nearest neighbor patches with locality sensitive hashing (LSH), and the grouped patches are aggregated into each query patch by performing a pair-wise local attention. In this way, the pair-wise operation establishes nonlocal connectivity while maintaining the desired properties of the local attention, i.e., inductive bias of locality and linear complexity to input resolution. The proposed method outperforms state-of-the-art restoration approaches on image denoising, deblurring and deraining benchmarks. The code will be available soon.
Hunsang Lee, Hyesong Choi, Kwanghoon Sohn, Dongbo Min
CVPR3
2022 Probabilistic Representations for Video Contrastive Learning
abstract
This paper presents Probabilistic Video Contrastive Learning, a self-supervised representation learning method that bridges contrastive learning with probabilistic representation. We hypothesize that the clips composing the video have different distributions in short-term duration, but can represent the complicated and sophisticated video distribution through combination in a common embedding space. Thus, the proposed method represents video clips as normal distributions and combines them into a Mixture of Gaussians to model the whole video distribution. By sampling embeddings from the whole video distribution, we can circumvent the careful sampling strategy or transformations to generate augmented views of the clips, unlike previous deterministic methods that have mainly focused on such sample generation strategies for contrastive learning. We further propose a stochastic contrastive loss to learn proper video distributions and handle the inherent uncertainty from the nature of the raw video. Experimental results verify that our probabilistic embedding stands as a state-of-the-art video representation learning for action recognition and video retrieval on the most popular benchmarks, including UCF101 and HMDB51.
Jungin Park, Jiyoung Lee 0005, Ig-Jae Kim, Kwanghoon Sohn
CVPR4
2022 PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation
Kwonyoung Kim, Jungin Park, Jiyoung Lee 0005, Dongbo Min, Kwanghoon Sohn
ECCV (38)5
2022 Multi-Domain Unsupervised Image-to-Image Translation with Appearance Adaptive Convolution
abstract
Over the past few years, image-to-image (I2I) translation methods have been proposed to translate a given image into diverse outputs. Despite the impressive results, they mainly focus on the I2I translation between two domains, so the multi-domain I2I translation still remains a challenge. To address this problem, we propose a novel multi-domain unsupervised image-to-image translation (MDUIT) framework that leverages the decomposed content feature and appearance adaptive convolution to translate an image into a target appearance while preserving the given geometric content. We also exploit a contrast learning objective, which improves the disentanglement ability and effectively utilizes multi-domain image data in the training process by pairing the semantically similar images. This allows our method to learn the diverse mappings between multiple visual domains with only a single framework. We show that the proposed method produces visually diverse and plausible results in multiple domains compared to the state-of-the-art methods.
Somi Jeong, Jiyoung Lee 0005, Kwanghoon Sohn
ICASSP3
2022 PASTS: Toward Effective Distilling Transformer for Panoramic Semantic Segmentation
abstract
Recently, panoramic imaging system has been attracting a lot of attention in various real-world applications due to its all-around sensing abilities. Despite the success of semantic segmentation, the performance of panoramic segmentation is still poor because the number of annotated panoramic datasets is insufficient and existing methods cannot handle the structural distortions in panoramic images caused by wide FoV. In this paper, we present a novel PAnoramic Segmentation Transformers (PASTs) trained by a knowledge distillation strategy with teacher-student branches. We first train the teacher using labeled pinhole images. The knowledge learned from the teacher is transferred to the student via feature distillation. To this end, we exploit the distorted pinhole images to force the attention and the prediction from the teacher consistent with those from the student. In addition, we adopt the entropy loss to train the student with unlabeled panoramic images. Experimental results demonstrate the effectiveness of our method, both qualitatively and quantitatively.
Jihyun Kim 0009, Somi Jeong, Kwanghoon Sohn
ICIP3
2022 Mask-Guided Attention and Episode Adaptive Weights for Few-Shot Segmentation
abstract
Few-shot segmentation aims to segment objects with novel classes in a query image, given a support set which consists of few annotated support images. A key factor in few-shot segmentation is to effectively exploit information for the target classes from the support set. In addition, we argue that the overall quality of information available in each training episode varies depending on the given support samples. In this paper, we propose Mask-Guided Attention module to extract more beneficial features for few-shot segmentation from the support images. Taking advantage of the support masks, the area correlated to the foreground object is highlighted and enables the support encoder to extract comprehensive support features with contextual information. Furthermore, we propose Episode Adaptive Weight to balance the training between different episodes. It adaptively adjusts loss weight according to the difficulty of each episode determined by self-supervised segmentation loss of support images and encourages the model to pay more attention to more difficult episodes. Extensive experimental results including comparisons with the state-of-the-art methods and ablation studies demonstrate the effectiveness of the proposed method.
Hyeongjun Kwon, Taeyong Song, Sunok Kim, Kwanghoon Sohn
ICIP4
2022 Meta-confidence estimation for stereo matching
abstract
We propose a novel framework to estimate the confidence of a disparity map taking into account, for the first time, the uncertainty affecting the confidence estimation process itself. Conversely to other tasks such as disparity estimation, the uncertainty of confidence directly hints that the confidence should be increased if initially low, but with high uncertainty, decreased otherwise. By modelling such a cue in the form of a second-level confidence, or meta-confidence, our solution allows for finding incorrect predictions inferred by confidence estimator and for learning a correction for them. Our strategy is suited for any state-of-the-art method known in literature, either implemented using random forest classifiers or deep neural networks. Especially, for deep neural networks-based models, we present a multi-headed confidence estimator followed by an uncertainty network, so as to predict mean confidence and meta-confidence within a single network without the cost of lower accuracy, a known limitation in literature for uncertainty estimation. Experimental results on a variety of stereo algorithms and confidence estimation models prove that the modeled meta-confidence is meaningful of the reliability of the estimated confidence and allows for refining it.
Seungryong Kim, Matteo Poggi, Sunok Kim, Kwanghoon Sohn, Stefano Mattoccia
ICRA4
2022 Deep Cascade Network for Noise-Robust SAR Ship Detection With Label Augmentation
abstract
Deep learning has recently made an impressive advance in ship detection in Synthetic Aperture Radar (SAR) images. In spite of this advancement, conventional deep detection networks often suffer from speckle noise that inherently occurs in SAR images. However, despeckling researches have focused only on improving the visual quality of the SAR images. Despeckling without considering subsequent task may cause loss of semantic information and result in performance degradation. In this letter, we propose a deep cascade framework for noise-robust SAR ship detection that sequentially performs despeckle and detection. We effectively train our cascade network using pseudo SAR images with SAR-like structures and additional detection annotations. We also propose semantic conservative loss that allows these two tasks to cooperate with each other. Experimental results including comparisons to previous methods and extensive ablation studies show the effectiveness of our proposed method.
Keunhoon Choi, Taeyong Song, Sunok Kim, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.6
2022 Enriching SAR Ship Detection via Multistage Domain Alignment
abstract
The advent of deep learning has made a significant advance in ship detection in synthetic aperture radar (SAR) images. However, it is still challenging since the amount of labeled SAR samples for training is not sufficient. Moreover, SAR images are corrupted by speckle noise, making them complex and difficult to interpret even by human experts. In this letter, we propose a novel SAR ship detection framework that leverages label-rich electro-optical (EO) images for more plentiful feature representations, and delicately addresses the speckle noise in SAR images. To this end, we first introduce a multistage domain alignment module that reduces the distribution discrepancies between EO and SAR feature maps at local, global, and instance levels. This allows enriching SAR representations by gradually instilling cross-domain knowledge from a large-scale EO image dataset. We further design a blind-spot layer for feature extraction to suppress the influence of speckles. Experimental results on the high-resolution SAR images dataset (HRSID) show that our detection performance achieves average precision (AP) 5.5% better than the current state-of-the-arts that exploits SAR images only. Our method significantly improves the detection performance with higher speckle noises, demonstrating stronger robustness than the conventional methods.
Somi Jeong, Youngjung Kim, Sung-Ho Kim 0004, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.4
2022 Semantic Equalization Learning for Semi-Supervised SAR Building Segmentation
abstract
Synthetic aperture radar (SAR) building segmentation, which is one of the fundamental tasks in the remote sensing community, has been achieved remarkable performance using convolutional neural networks (CNNs). Since most methods do not consider distinctive characteristics of SAR images, they tend to be biased towards simple and large buildings while ignoring small and complex-shaped ones. To build a general and powerful SAR building segmentation model, in this letter, we introduce a semi-supervised learning (SSL) framework with semantic equalization learning (SEL). Concretely, we leverage labeled SAR and EO image pairs and unlabeled SAR images for SSL to extract representative SAR features with the help of context-rich EO features. Moreover, SEL aims to balance the training of well and poor-performing samples via our purposed data augmentation technique and the objective functions. It consists of a semantic proportional CutMix (SP-CutMix) module to increase the sampling probability of under-performed samples during the training phase, and an equalized segmentation loss (ESL) to adjust the loss contribution depending on difficulties. By doing so, our method prevents the model from being biased to easy samples and increases the performance of difficult building samples. Experimental results on the SpaceNet-6 benchmark demonstrate the effectiveness of our framework, especially by significantly improving the most challenging scenarios, that is less labeled data available.
Eungbean Lee, Somi Jeong, Jun Hee Kim, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.4
2022 Shape-Robust SAR Ship Detection via Context-Preserving Augmentation and Deep Contrastive RoI Learning
abstract
With recent advances in deep-learning techniques and the advantages of synthetic aperture radar (SAR) images, deep-learning-based SAR ship detection has attracted a lot of attention. In this letter, we aim to build an SAR ship detection framework that is robust to target shape variations. Focusing on shape variation caused by radar shadow, we propose an instance-level data augmentation (DA) method. We leverage ground-truth annotations for bounding box and instance segmentation mask to design a sophisticated pipeline to simulate target information loss while preserving contextual information. In order to enhance the capacity to model shape variation, we design network architecture using deformable convolutional networks (DCNs). Furthermore, we introduce contrastive region of interest (RoI) loss to encourage similarity between original and augmented target RoI features, while encouraging background features to be distinguished from the target RoI features. We present extensive experiments to demonstrate the effectiveness of the proposed method.
Taeyong Song, Sunok Kim, Kwanghoon Sohn
IEEE Geosci. Remote. Sens. Lett.3
2022 Pyramidal Semantic Correspondence Networks
abstract
This paper presents a deep architecture, called pyramidal semantic correspondence networks (PSCNet), that estimates locally-varying affine transformation fields across semantically similar images. To deal with large appearance and shape variations that commonly exist among different instances within the same object category, we leverage a pyramidal model where the affine transformation fields are progressively estimated in a coarse-to-fine manner so that the smoothness constraint is naturally imposed. Different from the previous methods which directly estimate global or local deformations, our method first starts to estimate the transformation from an entire image and then progressively increases the degree of freedom of the transformation by dividing coarse cell into finer ones. To this end, we propose two spatial pyramid models by dividing an image in a form of quad-tree rectangles or into multiple semantic elements of an object. Additionally, to overcome the limitation of insufficient training data, a novel weakly-supervised training scheme is introduced that generates progressively evolving supervisions through the spatial pyramid models by leveraging a correspondence consistency across image pairs. Extensive experimental results on various benchmarks including TSS, Proposal Flow-WILLOW, Proposal Flow-PASCAL, Caltech-101, and SPair-71k demonstrate that the proposed method outperforms the lastest methods for dense semantic correspondence.
Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 On the Confidence of Stereo Matching in a Deep-Learning Era: A Quantitative Evaluation
abstract
Stereo matching is one of the most popular techniques to estimate dense depth maps by finding the disparity between matching pixels on two, synchronized and rectified images. Alongside with the development of more accurate algorithms, the research community focused on finding good strategies to estimate the reliability, i.e., the confidence, of estimated disparity maps. This information proves to be a powerful cue to naively find wrong matches as well as to improve the overall effectiveness of a variety of stereo algorithms according to different strategies. In this paper, we review more than ten years of developments in the field of confidence estimation for stereo matching. We extensively discuss and evaluate existing confidence measures and their variants, from hand-crafted ones to the most recent, state-of-the-art learning based methods. We study the different behaviors of each measure when applied to a pool of different stereo algorithms and, for the first time in literature, when paired with a state-of-the-art deep stereo network. Our experiments, carried out on five different standard datasets, provide a comprehensive overview of the field, highlighting in particular both strengths and limitations of learning-based strategies.
Matteo Poggi, Seungryong Kim, Fabio Tosi, Sunok Kim, Filippo Aleotti, Dongbo Min, Kwanghoon Sohn, Stefano Mattoccia
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 Memory-Guided Image De-Raining Using Time-Lapse Data
abstract
This paper addresses the problem of single image de-raining, that is, the task of recovering clean and rain-free background scenes from a single image obscured by a rainy artifact. Although recent advances adopt real-world time-lapse data to overcome the need for paired rain-clean images, they are limited to fully exploit the time-lapse data. The main cause is that, in terms of network architectures, they could not capture long-term rain streak information in the time-lapse data during training owing to the lack of memory components. To address this problem, we propose a novel network architecture combining the time-lapse data and, the memory network that explicitly helps to capture long-term rain streak information. Our network comprises the encoder-decoder networks and a memory network. The features extracted from the encoder are read and updated in the memory network that contains several memory items to store rain streak-aware feature representations. With the read/update operation, the memory network retrieves relevant memory items in terms of the queries, enabling the memory items to represent the various rain streaks included in the time-lapse data. To boost the discriminative power of memory features, we also present a novel background selective whitening (BSW) loss for capturing only rain streak information in the memory network by erasing the background information. Experimental results on standard benchmarks demonstrate the effectiveness and superiority of our approach.
Jaehoon Cho, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Image Process.3
2021 Deep Low-Contrast Image Enhancement using Structure Tensor Representation
Hyungjoo Jung, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
AAAI4
2021 Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation
abstract
Existing techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain itself or estimated category, providing the limitation to encode the domains having a multi-modal data distribution. To overcome this limitation, we introduce a learnable clustering module, and a novel domain adaptation framework, called cross-domain grouping and alignment. To cluster the samples across domains with an aim to maximize the domain alignment without forgetting precise segmentation ability on the source domain, we present two loss functions, in particular, for encouraging semantic consistency and orthogonality among the clusters. We also present a loss so as to solve a class imbalance problem, which is the other limitation of the previous methods. Our experiments show that our method consistently boosts the adaptation performance in semantic segmentation, outperforming the state-of-the-arts on various domain adaptation settings.
Sunghun Joung, Seungryong Kim, Jungin Park, Ig-Jae Kim, Kwanghoon Sohn
AAAI6
2021 Wide and Narrow: Video Prediction from Context and Motion
Jaehoon Cho, Jiyoung Lee 0005, Changjae Oh, Wonil Song, Kwanghoon Sohn
BMVC5
2021 Mining Better Samples for Contrastive Learning of Temporal Correspondence
abstract
We present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appropriately adjusting their hardness during training. This allows us to suppress the adverse impact of ambiguous matches and prevent a trivial solution from being yielded by too hard or too easy negative samples. To accomplish this, we incorporate three different criteria that ranges from a pixel-level matching confidence to a video-level one into a bottom-up pipeline, and plan a curriculum that is aware of current representation power for the adaptive hardness of negative samples during training. With the proposed method, state-of-the-art performance is attained over the latest approaches on several video label propagation tasks.
Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
CVPR4
2021 Memory-Guided Unsupervised Image-to-Image Translation
abstract
We present a novel unsupervised framework for instance-level image-to-image translation. Although recent advances have been made by incorporating additional object annotations, existing methods often fail to handle images with multiple disparate objects. The main cause is that, during inference, they apply a global style to the whole image and do not consider the large style discrepancy between instance and background, or within instances. To address this problem, we propose a class-aware memory network that explicitly reasons about local style variations. A key-values memory structure, with a set of read/update operations, is introduced to record class-wise style variations and access them without requiring an object detector at the test time. The key stores a domain-agnostic content representation for allocating memory items, while the values encode domain-specific style representations. We also present a feature contrastive loss to boost the discriminative power of memory items. We show that by incorporating our memory, we can transfer class-aware and accurate style representations across domains. Experimental results demonstrate that our model outperforms recent instance-level methods and achieves state-of-the-art performance.
Somi Jeong, Youngjung Kim, Eungbean Lee, Kwanghoon Sohn
CVPR4
2021 Prototype-Guided Saliency Feature Learning for Person Search
abstract
Existing person search methods integrate person detection and re-identification (re-ID) module into a unified system. Though promising results have been achieved, the misalignment problem, which commonly occurs in person search, limits the discriminative feature representation for re-ID. To overcome this limitation, we introduce a novel framework to learn the discriminative representation by utilizing prototype in OIM loss. Unlike conventional methods using prototype as a representation of person identity, we utilize it as guidance to allow the attention network to consistently highlight multiple instances across different poses. Moreover, we propose a new prototype update scheme with adaptive momentum to increase the discriminative ability across different instances. Extensive ablation experiments demonstrate that our method can significantly enhance the feature discriminative power, outperforming the state-of-the-art results on two person search benchmarks including CUHK-SYSU and PRW.
Hanjae Kim, Sunghun Joung, Ig-Jae Kim, Kwanghoon Sohn
CVPR4
2021 Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech Separation
abstract
In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depends on the accuracy of audio-visual synchronization and the effectiveness of their representations. To overcome the frame discontinuity problem between two modalities due to transmission delay mismatch or jitter, we propose a cross-modal affinity network (CaffNet) that learns global correspondence as well as locally-varying affinities between audio and visual streams. Given that the global term provides stability over a temporal sequence at the utterance-level, this resolves the label permutation problem characterized by inconsistent assignments. By extending the proposed cross-modal affinity on the complex network, we further improve the separation performance in the complex spectral domain. Experimental results verify that the proposed methods outperform conventional ones on various datasets, demonstrating their advantages in real-world scenarios.
Jiyoung Lee 0005, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang, Kwanghoon Sohn
CVPR5
2021 Bridge To Answer: Structure-Aware Graph Interaction Network for Video Question Answering
abstract
This paper presents a novel method, termed Bridge to Answer, to infer correct answers for questions about a given video by leveraging adequate graph interactions of heterogeneous crossmodal graphs. To realize this, we learn question conditioned visual graphs by exploiting the relation between video and question to enable each visual node using question-to-visual interactions to encompass both visual and linguistic cues. In addition, we propose bridged visual-to-visual interactions to incorporate two complementary visual information on appearance and motion by placing the question graph as an intermediate bridge. This bridged architecture allows reliable message passing through compositional semantics of the question to generate an appropriate answer. As a result, our method can learn the question conditioned visual representations attributed to appearance and motion that show powerful capability for video question answering. Extensive experiments prove that the proposed method provides effective and superior performance than state-of-the-art methods on several benchmarks.
Jungin Park, Jiyoung Lee 0005, Kwanghoon Sohn
CVPR3
2021 Adaptive confidence thresholding for monocular depth estimation
abstract
Self-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to leverage pseudo ground truth depth maps of stereo images generated from self-supervised stereo matching methods. The confidence map of the pseudo ground truth depth map is estimated to mitigate performance degeneration by inaccurate pseudo depth maps. To cope with the prediction error of the confidence map itself, we also leverage the threshold network that learns the threshold dynamically conditioned on the pseudo depth maps. The pseudo depth labels filtered out by the thresholded confidence map are used to supervise the monocular depth network. Furthermore, we propose the probabilistic framework that refines the monocular depth map with the help of its uncertainty map through the pixel-adaptive convolution (PAC) layer. Experimental results demonstrate superior performance to state-of-the-art monocular depth estimation methods. Lastly, we exhibit that the proposed threshold learning can also be used to improve the performance of existing confidence estimation approaches.
Hyesong Choi, Hunsang Lee, Sunkyung Kim, Sunok Kim, Seungryong Kim, Kwanghoon Sohn, Dongbo Min
ICCV6
2021 Learning Canonical 3D Object Representation for Fine-Grained Recognition
abstract
We propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its appearance, while eliminating the effect of camera viewpoint, in a canonical configuration. Unlike conventional methods modeling spatial variation in 2D images only, our method is capable of reconfiguring the appearance feature in a canonical 3D space, thus enabling the subsequent object classifier to be invariant under 3D geometric variation. Our representation also allows us to go beyond existing methods, by incorporating 3D shape variation as an additional cue for object recognition. To learn the model without ground-truth 3D annotation, we deploy a differentiable renderer in an analysis-by-synthesis frame- work. By incorporating 3D shape and appearance jointly in a deep representation, our method learns the discriminative representation of the object and achieves competitive performance on fine-grained image recognition and vehicle re-identification. We also demonstrate that the performance of 3D shape reconstruction is improved by learning fine-grained shape deformation in a boosting manner.
Sunghun Joung, Seungryong Kim, Ig-Jae Kim, Kwanghoon Sohn
ICCV5
2021 Semantic-Aware Network for Aerial-To-Ground Image Synthesis
abstract
Aerial-to-ground image synthesis is an emerging and challenging problem that aims to synthesize a ground image from an aerial image. Due to the highly different layout and object representation between the aerial and ground images, existing approaches usually fail to transfer the components of the aerial scene into the ground scene. In this paper, we propose a novel framework to explore the challenges by imposing enhanced structural alignment and semantic awareness. We introduce a novel semantic-attentive feature transformation module that allows to reconstruct the complex geographic structures by aligning the aerial feature to the ground layout. Furthermore, we propose semantic-aware loss functions by leveraging a pre-trained segmentation network. The network is enforced to synthesize realistic objects across various classes by separately calculating losses for different classes and balancing them. Extensive experiments including comparisons with previous methods and ablation studies show the effectiveness of the proposed framework both qualitatively and quantitatively.
Jinhyun Jang, Taeyong Song, Kwanghoon Sohn
ICIP3
2021 Self-Balanced Learning for Domain Generalization
abstract
Domain generalization aims to learn a prediction model on multi-domain source data such that the model can generalize to a target domain with unknown statistics. Most existing approaches have been developed under the assumption that the source data is well-balanced in terms of both domain and class. However, real-world training data collected with different composition biases often exhibits severe distribution gaps for domain and class, leading to substantial performance degradation. In this paper, we propose a self-balanced domain generalization framework that adaptively learns the weights of losses to alleviate the bias caused by different distributions of the multi-domain source data. The self-balanced scheme is based on an auxiliary reweighting network that iteratively updates the weight of loss conditioned on the domain and class information by leveraging balanced meta data. Experimental results demonstrate the effectiveness of our method overwhelming state-of-the-art works for domain generalization.
Jin Kim 0005, Jiyoung Lee 0005, Jungin Park, Dongbo Min, Kwanghoon Sohn
ICIP5
2021 Stereo-augmented Depth Completion from a Single RGB-LiDAR image
abstract
Depth completion is an important task in computer vision and robotics applications, which aims at predicting accurate dense depth from a single RGB-LiDAR image. Convolutional neural networks (CNNs) have been widely used for depth completion to learn a mapping function from sparse to dense depth. However, recent methods do not exploit any 3D geometric cues during the inference stage and mainly rely on sophisticated CNN architectures. In this paper, we present a cascade and geometrically inspired learning framework for depth completion, consisting of three stages: view extrapolation, stereo matching, and depth refinement. The first stage extrapolates a virtual (right) view using a single RGB (left) and its LiDAR data. We then mimic the binocular stereo-matching, and as a result, explicitly encode geometric constraints during depth completion. This stage augments the final refinement process by providing additional geometric reasoning. We also introduce a distillation framework based on teacher-student strategy to effectively train our network. Knowledge from a teacher model privileged with real stereo pairs is transferred to the student through feature distillation. Experimental results on KITTI depth completion benchmark demonstrate that the proposed method is superior to state-of-the-art methods.
Keunhoon Choi, Somi Jeong, Youngjung Kim, Kwanghoon Sohn
ICRA4
2021 Privileged Knowledge Distillation for SAR Building Extraction
abstract
Automatic building footprint extraction from SAR imagery is one of the critical tasks in the remote sensing community. CNN has been recently explored in building extraction tasks and achieved improved performance. However, due to the scarcity of training data, it suffers from overfitting problem. This paper presents a novel knowledge distillation based framework consisting of teacher and student networks. Regarding EO image as privileged information, the teacher network learns to extract the rich EO and SAR image pair features. The student network then learns to estimate the building footprints from only SAR images based on the privileged knowledge from the teacher network. Experimental results on SpaceNet-6 benchmark demonstrate the effectiveness of our framework, which explicitly improves the performance of SAR segmentation network.
Eungbean Lee, Somi Jeong, Kwanghoon Sohn
IGARSS3
2021 CATs: Cost Aggregation Transformers for Visual Correspondence
abstract
We propose a novel cost aggregation network, called Cost Aggregation Transformers (CATs), to find dense correspondences between semantically similar images with additional challenges posed by large intra-class appearance and geometric variations. Cost aggregation is a highly important process in matching tasks, which the matching accuracy depends on the quality of its output. Compared to hand-crafted or CNN-based methods addressing the cost aggregation, in that either lacks robustness to severe deformations or inherit the limitation of CNNs that fail to discriminate incorrect matches due to limited receptive fields, CATs explore global consensus among initial correlation map with the help of some architectural designs that allow us to fully leverage self-attention mechanism. Specifically, we include appearance affinity modeling to aid the cost aggregation process in order to disambiguate the noisy initial correlation maps and propose multi-level aggregation to efficiently capture different semantics from hierarchical feature representations. We then combine with swapping self-attention technique and residual connections not only to enforce consistent matching, but also to ease the learning process, which we find that these result in an apparent performance boost. We conduct experiments to demonstrate the effectiveness of the proposed model over the latest methods and provide extensive ablation studies. Code and trained models are available at https://sunghwanhong.github.io/CATs/.
Seokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee, Kwanghoon Sohn, Seungryong Kim
NeurIPS5
2021 Deep monocular depth estimation leveraging a large-scale outdoor stereo dataset
Jaehoon Cho, Dongbo Min, Youngjung Kim, Kwanghoon Sohn
Expert Syst. Appl.4
2021 Dense Cross-Modal Correspondence Estimation With the Deep Self-Correlation Descriptor
abstract
We present the deep self-correlation (DSC) descriptor for establishing dense correspondences between images taken under different imaging modalities, such as different spectral ranges or lighting conditions. We encode local self-similar structure in a pyramidal manner that yields both more precise localization ability and greater robustness to non-rigid image deformations. Specifically, DSC first computes multiple self-correlation surfaces with randomly sampled patches over a local support window, and then builds pyramidal self-correlation surfaces through average pooling on the surfaces. The feature responses on the self-correlation surfaces are then encoded through spatial pyramid pooling in a log-polar configuration. To better handle geometric variations such as scale and rotation, we additionally propose the geometry-invariant DSC (GI-DSC) that leverages multi-scale self-correlation computation and canonical orientation estimation. In contrast to descriptors based on deep convolutional neural networks (CNNs), DSC and GI-DSC are training-free (i.e., handcrafted descriptors), are robust to cross-modality, and generalize well to various modality variations. Extensive experiments demonstrate the state-of-the-art performance of DSC and GI-DSC on challenging cases of cross-modal image pairs having photometric and/or geometric variations.
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Multi-Task Learning Framework for Motion Estimation and Dynamic Scene Deblurring
abstract
Motion blur, which disturbs human and machine perceptions of a scene, has been considered an unnecessary artifact that should be removed. However, the blur can be a useful clue to understanding the dynamic scene, since various sources of motion generate different types of artifacts. Motivated by the relationship between motion and blur, we propose a motion-aware feature learning framework for dynamic scene deblurring through multi-task learning. Our multi-task framework simultaneously estimates a deblurred image and a motion field from a blurred image. We design the encoder-decoder architectures for two tasks, and the encoder part is shared between them. Our motion estimation network could effectively distinguish between different types of blur, which facilitates image deblurring. Understanding implicit motion information through image deblurring could improve the performance of motion estimation. In addition to sharing the network between two tasks, we propose a reblurring loss function to optimize the overall parameters in our multi-task architecture. We provide an intensive analysis of complementary tasks to show the effectiveness of our multi-task framework. Furthermore, the experimental results demonstrate that the proposed method outperforms the state-of-the-art deblurring methods with respect to both qualitative and quantitative evaluations.
Hyungjoo Jung, Youngjung Kim, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
IEEE Trans. Image Process.5
2021 Adversarial Confidence Estimation Networks for Robust Stereo Matching
abstract
Stereo matching aiming to perceive the 3-D geometry of a scene facilitates numerous computer vision tasks used in advanced driver assistance systems (ADAS). Although numerous methods have been proposed for this task by leveraging deep convolutional neural networks (CNNs), stereo matching still remains an unsolved problem due to its inherent matching ambiguities. To overcome these limitations, we present a method for jointly estimating disparity and confidence from stereo image pairs through deep networks. We accomplish this through a minmax optimization to learn the generative cost aggregation networks and discriminative confidence estimation networks in an adversarial manner. Concretely, the generative cost aggregation networks are trained to accurately generate disparities at both confident and unconfident pixels from an input matching cost that are indistinguishable by the discriminative confidence estimation networks, while the discriminative confidence estimation networks are trained to distinguish the confident and unconfident disparities. In addition, to fully exploit complementary information of matching cost, disparity, and color image in confidence estimation, we present a dynamic fusion module. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks including real driving scenes.
Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Intell. Transp. Syst.4
2020 Stereoscopic Image Super-Resolution with Stereo Consistent Feature
abstract
We present a first attempt for stereoscopic image super-resolution (SR) for recovering high-resolution details while preserving stereo-consistency between stereoscopic image pair. The most challenging issue in the stereoscopic SR is that the texture details should be consistent for corresponding pixels in stereoscopic SR image pair. However, existing stereo SR methods cannot maintain the stereo-consistency, thus causing 3D fatigue to the viewers. To address this issue, in this paper, we propose a self and parallax attention mechanism (SPAM) to aggregate the information from its own image and the counterpart stereo image simultaneously, thus reconstructing high-quality stereoscopic SR image pairs. Moreover, we design an efficient network architecture and effective loss functions to enforce stereo-consistency constraint. Finally, experimental results demonstrate the superiority of our method over state-of-the-art SR methods in terms of both quantitative metrics and qualitative visual quality while maintaining stereo-consistency between stereoscopic image pair.
Wonil Song, Sungil Choi, Somi Jeong, Kwanghoon Sohn
AAAI4
2020 Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation
abstract
Existing techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcome this limitation, we introduce a learnable module, cylindrical convolutional networks (CCNs), that exploit cylindrical representation of a convolutional kernel defined in the 3D space. CCNs extract a view-specific feature through a view-specific convolutional kernel to predict object category scores at each viewpoint. With the view-specific feature, we simultaneously determine objective category and viewpoints using the proposed sinusoidal soft-argmax module. Our experiments demonstrate the effectiveness of the cylindrical convolutional networks on joint object detection and viewpoint estimation.
Sunghun Joung, Seungryong Kim, Hanjae Kim, Ig-Jae Kim, Junghyun Cho, Kwanghoon Sohn
CVPR7
2020 Guided Semantic Flow
Sangryul Jeon, Dongbo Min, Seungryong Kim, Jihwan Choe, Kwanghoon Sohn
ECCV (28)5
2020 SumGraph: Video Summarization via Recursive Graph Modeling
Jungin Park, Jiyoung Lee 0005, Ig-Jae Kim, Kwanghoon Sohn
ECCV (25)4
2020 Shape-Adaptive Kernel Network for Dense Object Detection
abstract
Dense object detectors that are applied over a regular, dense grid have advanced and drawn their attention in recent days. Their fully convolutional nature greatly advances the computational efficiency of object detectors compared to the two-stage detectors. However, the lack of the ability to adjust shape variation on a regular grid is still limited. In this paper we introduce a new framework, shape-adaptive kernel network, to handle spatial manipulation of input data in convolutional kernel space. At the heart of out approach is to align the original kernel space recovering shape variation of each input feature on regular grid. To this end, we propose a shape-adaptive kernel sampler to adjust dynamic convolutional kernel conditioned on input. To increase the flexibility of geometric transformation, a cascade refinement module is designed, which first estimates the global transformation grid and then estimates local offset in convolutional kernel space. Our experiments demonstrate the effectiveness of the shape-adaptive kernel network for dense object detection on various benchmarks.
Hanjae Kim, Sunghun Joung, Ig-Jae Kim, Kwanghoon Sohn
ICIP4
2020 Simultaneous Deep Stereo Matching and Dehazing with Feature Attention
Taeyong Song, Youngjung Kim, Changjae Oh, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
Int. J. Comput. Vis.6
2020 Discrete-Continuous Transformation Matching for Dense Semantic Correspondence
abstract
Techniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Furthermore, leveraging correspondence consistency and confidence-guided filtering in each iteration facilitates the convergence of our method. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks and applications.
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Unsupervised Low-Light Image Enhancement Using Bright Channel Prior
abstract
Recent approaches for low-light image enhancement achieve excellent performance through supervised learning based on convolutional neural networks. However, it is still challenging to collect a large amount of low-/normal-light image pairs in real environments for training the networks. In this letter, we propose an unsupervised learning approach for single low-light image enhancement using the bright channel prior (BCP) that the brightest pixel in a small patch is likely to be close to 1. An unsupervised loss function is defined with the pseudo ground-truth generated using the BCP. An enhancement network, consisting of a simple encoder-decoder, is then trained using the unsupervised loss function. To the best of our knowledge, this is the first attempt that enhances a low-light image through unsupervised learning. Furthermore, we introduce saturation loss and self-attention map for preserving image details and naturalness in the enhanced result. The performance of the proposed method is validated on various public datasets. Experimental results demonstrate that the proposed unsupervised approach achieves competitive performance over state-of-the-art methods based on supervised learning.
Hunsang Lee, Kwanghoon Sohn, Dongbo Min
IEEE Signal Process. Lett.2
2020 Single Image Deraining Using Time-Lapse Data
abstract
Leveraging on recent advances in deep convolutional neural networks (CNNs), single image deraining has been studied as a learning task, achieving an outstanding performance over traditional hand-designed approaches. Current CNNs based deraining approaches adopt the supervised learning framework that uses a massive training data generated with synthetic rain streaks, having a limited generalization ability on real rainy images. To address this problem, we propose a novel learning framework for single image deraining that leverages time-lapse sequences instead of the synthetic image pairs. The deraining networks are trained using the time-lapse sequences in which both camera and scenes are static except for time-varying rain streaks. Specifically, we formulate a background consistency loss such that the deraining networks consistently generate the same derained images from the time-lapse sequences. We additionally introduce two loss functions, the structure similarity loss that encourages the derained image to be similar with an input rainy image and the directional gradient loss using the assumption that the estimated rain streaks are likely to be sparse and have dominant directions. To consider various rain conditions, we leverage a dynamic fusion module that effectively fuses multi-scale features. We also build a novel large-scale time-lapse dataset providing real world rainy images containing various rain conditions. Experiments demonstrate that the proposed method outperforms state-of-the-art techniques on synthetic and real rainy images both qualitatively and quantitatively. On the high-level vision tasks under severe rainy conditions, it has been shown that the proposed method can be utilized as a pre-preprocessing step for subsequent tasks.
Jaehoon Cho, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.4
2020 Unsupervised Deep Image Fusion With Structure Tensor Representations
abstract
Convolutional neural networks (CNNs) have facilitated substantial progress on various problems in computer vision and image processing. However, applying them to image fusion has remained challenging due to the lack of the labelled data for supervised learning. This paper introduces a deep image fusion network (DIF-Net), an unsupervised deep learning framework for image fusion. The DIF-Net parameterizes the entire processes of image fusion, comprising of feature extraction, feature fusion, and image reconstruction, using a CNN. The purpose of DIF-Net is to generate an output image which has an identical contrast to high-dimensional input images. To realize this, we propose an unsupervised loss function using the structure tensor representation of the multi-channel image contrasts. Different from traditional fusion methods that involve time-consuming optimization or iterative procedures to obtain the results, our loss function is minimized by a stochastic deep learning solver with large-scale examples. Consequently, the proposed method can produce fused images that preserve source image details through a single forward network trained without reference ground-truth labels. The proposed method has broad applicability to various image fusion problems, including multi-spectral, multi-focus, and multi-exposure image fusions. Quantitative and qualitative evaluations show that the proposed technique outperforms existing state-of-the-art approaches for various applications.
Hyungjoo Jung, Youngjung Kim, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
IEEE Trans. Image Process.5
2020 Learning Deeply Aggregated Alternating Minimization for General Inverse Problems
abstract
Regularization-based image restoration is one of the most powerful tools in image processing and computer vision thanks to its flexibility for handling various inverse problems. However, designing an optimal regularization function still remains unsolved since natural images and related scene types have a complex structure. In this paper, we present a general and principled framework, called deeply aggregated alternating minimization (DeepAM). We design a convolutional neural network (CNN) to implicitly parameterize the regularizer of the alternating minimization (AM) algorithm. Contrary to the conventional AM algorithm based on a point-wise proximal mapping, the DeepAM projects intermediate estimate into a set of natural images via deep aggregation. Since the CNN is fully integrated into the AM procedure, all parameters can be jointly optimized through end-to-end training. These properties enable the DeepAM to converge with a small number of iterations, while maintaining an algorithmic simplicity. We show that the DeepAM outperforms state-of-the-art methods, including nonlocal-based methods, Plug-and-Play regularization, and recent data-driven approaches. The effectiveness of our framework is demonstrated in a variety of image restoration tasks: Guassian denoising, deraining, deblurring, super-resolution, color-guided depth upsampling, and RGB/NIR restoration.
Hyungjoo Jung, Youngjung Kim, Dongbo Min, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn
IEEE Trans. Image Process.6
2020 Multi-Modal Recurrent Attention Networks for Facial Expression Recognition
abstract
Recent deep neural networks based methods have achieved state-of-the-art performance on various facial expression recognition tasks. Despite such progress, previous researches for facial expression recognition have mainly focused on analyzing color recording videos only. However, the complex emotions expressed by people with different skin colors under different lighting conditions through dynamic facial expressions can be fully understandable by integrating information from multi-modal videos. We present a novel method to estimate dimensional emotion states, where color, depth, and thermal recording videos are used as a multi-modal input. Our networks, called multi-modal recurrent attention networks (MRAN), learn spatiotemporal attention volumes to robustly recognize the facial expression based on attention-boosted feature volumes. We leverage the depth and thermal sequences as guidance priors for color sequence to selectively focus on emotional discriminative regions. We also introduce a novel benchmark for multi-modal facial expression recognition, termed as multi-modal arousal-valence facial expression recognition (MAVFER), which consists of color, depth, and thermal recording videos with corresponding continuous arousal-valence scores. The experimental results show that our method can achieve the state-of-the-art results in dimensional facial expression recognition on color recording datasets including RECOLA, SEWA and AFEW, and a multi-modal recording dataset including MAVFER.
Jiyoung Lee 0005, Sunok Kim, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Image Process.4
2020 Unsupervised Stereo Matching Using Confidential Correspondence Consistency
abstract
Stereo matching aims to perceive the 3D geometric configuration of scenes and facilitates a variety of computer vision in advanced driver assistance systems (ADAS) applications. Recently, deep convolutional neural networks (CNNs) have shown dramatic performance improvements for computing the matching cost in the stereo matching. However, the performance of CNN-based approaches relies heavily on datasets, requiring a large number of ground truth data which needs tremendous works. To overcome this limitation, we present a novel framework to learn CNNs for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. By exploiting the correspondence consistency between stereo images, our method selects putative positive samples in each training iteration and utilizes them to train the networks. We further propose a positive sample propagation scheme to leverage additional training samples. Our unsupervised learning method is evaluated with two kinds of network architectures, simple and precise CNNs, and shows comparable performance to that of the state-of-the-art methods including both supervised and unsupervised learning approaches on KITTI, Middlebury, HCI, and Yonsei datasets. This extensive evaluation demonstrates that the proposed learning framework can be applied to deal with various real driving conditions.
Sunghun Joung, Seungryong Kim, Kihong Park, Kwanghoon Sohn
IEEE Trans. Intell. Transp. Syst.4
2020 High-Precision Depth Estimation Using Uncalibrated LiDAR and Stereo Fusion
abstract
We address the problem of 3D reconstruction from uncalibrated LiDAR point cloud and stereo images. Since the usage of each sensor alone for 3D reconstruction has weaknesses in terms of density and accuracy, we propose a deep sensor fusion framework for high-precision depth estimation. The proposed architecture consists of calibration network and depth fusion network, where both networks are designed considering the trade-off between accuracy and efficiency for mobile devices. The calibration network first corrects an initial extrinsic parameter to align the input sensor coordinate systems. The accuracy of calibration is markedly improved by formulating the calibration in the depth domain. In the depth fusion network, complementary characteristics of sparse LiDAR and dense stereo depth are then encoded in a boosting manner. Since training data for the LiDAR and stereo depth fusion are rather limited, we introduce a simple but effective approach to generate pseudo ground truth labels from the raw KITTI dataset. The experimental evaluation verifies that the proposed method outperforms current state-of-the-art methods on the KITTI benchmark. We also collect data using our proprietary multi-sensor acquisition platform and verify that the proposed method generalizes across different sensor settings and scenes.
Kihong Park, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Intell. Transp. Syst.3
2019 LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence Estimation
abstract
We present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. The attention inference networks encode the importance of tri-modal confidence features and then concatenate them using the attention maps in an adaptive and dynamic fashion. This enables us to make an optimal fusion of the heterogeneous features, compared to a simple concatenation technique that is commonly used in conventional approaches. In addition, to encode the confidence features with locally-varying receptive fields, the scale inference networks learn the scale map and warp the fused confidence features through convolutional spatial transformer networks. Finally, the confidence map is progressively estimated in the recursive refinement networks to enforce a spatial context and local consistency. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks.
Sunok Kim, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
CVPR4
2019 Semantic Attribute Matching Networks
abstract
We present semantic attribute matching networks (SAM-Net) for jointly establishing correspondences and transferring attributes across semantically similar images, which intelligently weaves the advantages of the two tasks while overcoming their limitations. SAM-Net accomplishes this through an iterative process of establishing reliable correspondences by reducing the attribute discrepancy between the images and synthesizing attribute transferred images using the learned correspondences. To learn the networks using weak supervisions in the form of image pairs, we present a semantic attribute matching loss based on the matching similarity between an attribute transferred source feature and a warped target feature. With SAM-Net, the state-of-the-art performance is attained on several benchmarks for semantic matching and attribute transfer.
Seungryong Kim, Dongbo Min, Somi Jeong, Sunok Kim, Sangryul Jeon, Kwanghoon Sohn
CVPR6
2019 Joint Learning of Semantic Alignment and Object Landmark Detection
abstract
Convolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In this paper, we present a joint learning approach for obtaining dense correspondences and discovering object landmarks from semantically similar images. Based on the key insight that the two tasks can mutually provide supervisions to each other, our networks accomplish this through a joint loss function that alternatively imposes a consistency constraint between the two tasks, thereby boosting the performance and addressing the lack of training data in a principled manner. To the best of our knowledge, this is the first attempt to address the lack of training data for the two tasks through the joint learning. To further improve the robustness of our framework, we introduce a probabilistic learning formulation that allows only reliable matches to be used in the joint learning process. With the proposed method, state-of-the-art performance is attained on several benchmarks for semantic matching and landmark detection.
Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
ICCV4
2019 Context-Aware Emotion Recognition Networks
abstract
Traditional techniques for emotion recognition have focused on the facial expression analysis only, thus providing limited ability to encode context that comprehensively represents the emotional responses. We present deep networks for context-aware emotion recognition, called CAER-Net, that exploit not only human facial expression but also context information in a joint and boosting manner. The key idea is to hide human faces in a visual scene and seek other contexts based on an attention mechanism. Our networks consist of two sub-networks, including two-stream encoding networks to separately extract the features of face and context regions, and adaptive fusion networks to fuse such features in an adaptive fashion. We also introduce a novel benchmark for context-aware emotion recognition, called CAER, that is appropriate than existing benchmarks both qualitatively and quantitatively. On several benchmarks, CAER-Net proves the effect of context for emotion recognition. Our dataset is available at http://caer-dataset.github.io.
Jiyoung Lee 0005, Seungryong Kim, Sunok Kim, Jungin Park, Kwanghoon Sohn
ICCV5
2019 Unpaired Cross-Spectral Pedestrian Detection Via Adversarial Feature Learning
abstract
Even though there exist significant advances in recent studies, existing methods for pedestrian detection still have shown limited performances under challenging illumination conditions especially at nighttime. To address this, cross-spectral pedestrian detection methods have been presented using color and thermal, and shown substantial performance gains on the challenging circumstances. However, their paired cross-spectral settings have limited applicability in real-world scenarios. To overcome this, we propose a novel learning framework for cross-spectral pedestrian detection in an unpaired setting. Based on an assumption that features from color and thermal images share their characteristics in a common feature space to benefit their complement information, we design the separate feature embedding networks for color and thermal images followed by sharing detection networks. To further improve the cross-spectral feature representation, we apply an adversarial learning scheme to intermediate features of the color and thermal images. Experiments demonstrate the outstanding performance of the proposed method on the KAIST multi-spectral benchmark in comparison to the state-of-the-art methods.
Sunghun Joung, Kihong Park, Seungryong Kim, Kwanghoon Sohn
ICIP5
2019 Graph Regularization Network with Semantic Affinity for Weakly-Supervised Temporal Action Localization
abstract
This paper presents a novel deep architecture for weakly-supervised temporal action localization that not only generates segment-level action responses but also propagates segment-level responses to the neighborhood in a form of graph Laplacian regularization. Specifically, our approach consists of two sub-modules; a class activation module to estimate the action score map over time through the action classifiers, and a graph regularization module to refine the estimated action score map by solving a quadratic programming problem with the predicted segment-level semantic affinities. Since these two modules are integrated with fully differentiable layers, the proposed networks can be jointly trained in an end-to-end manner. Experimental results on Thumos14 and ActivityNet1.2 demonstrate that the proposed method provides outstanding performances in weakly-supervised temporal action localization.
Jungin Park, Jiyoung Lee 0005, Sangryul Jeon, Seungryong Kim, Kwanghoon Sohn
ICIP5
2019 OCEAN: Object-centric arranging network for self-supervised visual representations learning
Changjae Oh, Bumsub Ham, Hansung Kim 0001, Adrian Hilton 0001, Kwanghoon Sohn
Expert Syst. Appl.5
2019 FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence
abstract
We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. Unlike traditional dense correspondence approaches for estimating depth or optical flow, semantic correspondence estimation poses additional challenges due to intra-class appearance and shape variations among different instances within the same object or scene category. To robustly match points across semantically similar images, we formulate FCSS using local self-similarity (LSS), which is inherently insensitive to intra-class appearance variations. LSS is incorporated through a proposed convolutional self-similarity (CSS) layer, where the sampling patterns and the self-similarity measure are jointly learned in an end-to-end and multi-scale manner. Furthermore, to address shape variations among different object instances, we propose a convolutional affine transformer (CAT) layer that estimates explicit affine transformation fields at each pixel to transform the sampling patterns and corresponding receptive fields. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in most existing datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS significantly outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks.
Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin 0001, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Learning to Find Unpaired Cross-Spectral Correspondences
abstract
We present a deep architecture and learning framework for establishing correspondences across cross-spectral visible and infrared images in an unpaired setting. To overcome the unpaired cross-spectral data problem, we design the unified image translation and feature extraction modules to be learned in a joint and boosting manner. Concretely, the image translation module is learned only with the unpaired cross-spectral data, and the feature extraction module is learned with an input image and its translated image. By learning two modules simultaneously, the image translation module generates the translated image that preserves not only the domain-specific attributes with separate latent spaces but also the domain-agnostic contents with feature consistency constraint. In an inference phase, the cross-spectral feature similarity is augmented by intra-spectral similarities between the features extracted from the translated images. Experimental results show that this model outperforms the state-of-the-art unpaired image translation methods and cross-spectral feature descriptors on various visible and infrared benchmarks.
Somi Jeong, Seungryong Kim, Kihong Park, Kwanghoon Sohn
IEEE Trans. Image Process.4
2019 Structure-Texture Image Decomposition Using Deep Variational Priors
abstract
Most variational formulations for structure-texture image decomposition force structure images to have small norm in some functional spaces, and share a common notion of edges, i.e., large-gradients or -intensity differences. However, such definition makes it difficult to distinguish structure edges from oscillations that have fine spatial scale but high contrast. In this paper, we introduce a new model by learning deep variational prior for structure images without explicit training data. An alternating direction method of multiplier (ADMM) algorithm and its modular structure are adopted to plug deep variational priors into an iterative smoothing process. The central observations are that convolution neural networks (CNNs) can replace the total variation prior, and are indeed powerful to capture the natures of structure and texture. We show that our learned priors using CNNs successfully differentiate highamplitude details from structure edges, and avoid halo artifacts. Different from previous data-driven smoothing schemes, our formulation provides another degree of freedom to produce continuous smoothing effects. Experimental results demonstrate the effectiveness of our approach on various computational photography and image processing applications, including texture removal, detail manipulation, HDR tone-mapping, and nonphotorealistic abstraction.
Youngjung Kim, Bumsub Ham, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Image Process.4
2019 Unified Confidence Estimation Networks for Robust Stereo Matching
abstract
We present a deep architecture that estimates a stereo confidence, which is essential for improving the accuracy of stereo matching algorithms. In contrast to existing methods based on deep convolutional neural networks (CNNs) that rely on only one of the matching cost volume or estimated disparity map, our network estimates the stereo confidence by using the two heterogeneous inputs simultaneously. Specifically, the matching probability volume is first computed from the matching cost volume with residual networks and a pooling module in a manner that yields greater robustness. The confidence is then estimated through a unified deep network that combines confidence features extracted both from the matching probability volume and its corresponding disparity. In addition, our method extracts the confidence features of the disparity map by applying multiple convolutional filters with varying sizes to an input disparity map. To learn our networks in a semi-supervised manner, we propose a novel loss function that use confident points to compute the image reconstruction loss. To validate the effectiveness of our method in a disparity post-processing step, we employ three post-processing approaches; cost modulation, ground control points-based propagation, and aggregated ground control points-based propagation. Experimental results demonstrate that our method outperforms state-of-the-art confidence estimation methods on various benchmarks.
Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Image Process.4
2018 Deep Network for Simultaneous Stereo Matching and Dehazing
Taeyong Song, Youngjung Kim, Changjae Oh, Kwanghoon Sohn
BMVC4
2018 PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn
ECCV (6)4
2018 Spatiotemporal Attention Based Deep Neural Networks for Emotion Recognition
abstract
We propose a spatiotemporal attention based deep neural networks for dimensional emotion recognition in facial videos. To learn the spatiotemporal attention that selectively focuses on emotional sailient parts within facial videos, we formulate the spatiotemporal encoder-decoder network using Convolutional LSTM (ConvLSTM) modules, which can be learned implicitly without any pixel-level annotations. By leveraging the spatiotemporal attention, we also formulate the 3D convolutional neural networks (3D-CNNs) to robustly recognize the dimensional emotion in facial videos. The experimental results show that our method can achieve the state-of-the-art results in dimensional emotion recognition with the highest concordance correlation coefficient (CCC) on RECOLA and AV+EC 2017 dataset.
Jiyoung Lee 0005, Sunok Kim, Seungryong Kim, Kwanghoon Sohn
ICASSP4
2018 Multi-Task Self-Supervised Visual Representation Learning for Monocular Road Segmentation
abstract
Training deep networks commonly follows the supervised learning paradigm, which requires large-scale semantically-labeled data. The construction of such dataset is one of the major challenges when approaching to Advanced Driver Assistance Systems (ADAS) due to the expense of human annotation. In this paper, we explore whether unsupervised stereo-based cues can be used to learn high-level semantics for monocular road detection. Specifically, we estimate drivable space and surface normals from stereo images, which are used for pseudo ground-truth to train a convolutional neural network (CNN) as a multi-task learning scheme. Combining these multiple self-supervision tasks enables CNN to jointly encode the knowledge of obstacle and ground-plane into a single frame. We demonstrate that the feature representation learned by our multi-task approach synergistically provides a rich knowledge about geometrical characteristics. Experiments on the KITTI road dataset show that our representation outperforms state-of-the-art road detection approaches.
Laehoon Cho, Youngjung Kim, Hyungjoo Jung, Changjae Oh, Jaesung Youn, Kwanghoon Sohn
ICME6
2018 High-Precision Depth Estimation with the 3D LiDAR and Stereo Fusion
abstract
We present a deep convolutional neural network (CNN) architecture for high-precision depth estimation by jointly utilizing sparse 3D LiDAR and dense stereo depth information. In this network, the complementary characteristics of sparse 3D LiDAR and dense stereo depth are simultaneously encoded in a boosting manner. Tailored to the LiDAR and stereo fusion problem, the proposed network differs from previous CNNs in the incorporation of a compact convolution module, which can be deployed with the constraints of mobile devices. As training data for the LiDAR and stereo fusion is rather limited, we introduce a simple yet effective approach for reproducing the raw KITTI dataset. The raw LiDAR scans are augmented by adapting an off-the-shelf stereo algorithm and a confidence measure. We evaluate the proposed network on the KITTI benchmark and data collected by our multi-sensor acquisition system. Experiments demonstrate that the proposed network generalizes across datasets and is significantly more accurate than various baseline approaches.
Kihong Park, Seungryong Kim, Kwanghoon Sohn
ICRA3
2018 AVSU: Workshop on Audio-Visual Scene Understanding for Immersive Multimedia
abstract
This workshop aims to provide a forum to exchange ideas in scene understanding techniques researched in audio and visual communities, and to ultimately unlock the creative potential of joint audio-visual signal processing to deliver a step change in various multimedia applications. Papers and talks presented in this workshop will contribute to the emerging technology for audio and visual information that can improve traditional approaches for multimedia content production and reproduction. The goals of this workshop are to (1) present and discuss the latest trends in audio and computer vision fields for the common research goals, (2) understand state-of-the-art techniques and bottlenecks in the other's discipline for the common topics, (3) investigate research opportunities of joint audio-visual scene understandings in multimedia content production. This workshop will be a good opportunity to bring together leading experts in audio processing and computer vision, and will bridge the gap between two research fields in multimedia content production and reproduction.
Adrian Hilton 0001, Hong-Goo Kang, Hansung Kim 0001, Kwanghoon Sohn
ACM Multimedia4
2018 CoVieW'18: The 1st Workshop and Challenge on Comprehensive Video Understanding in the Wild
abstract
The 1st Workshop and Challenge on Comprehensive Video Understanding in the Wild, dubbed CoVieW'18, is held in Seoul, Korea on October 22, 2018, in conjuction with ACM Multimedia 2018. The workshop aims to solve the joint and comprehensive understanding problem in untrimmed videos with a particular emphasis on joint action and scene recognition. The workshop encourages researchers to participate in joint action and scene recognition challenge in untrimmed videos and to report their results. The workshop program includes 1 keynote speech, 2 invited speakers, 6 regular and challenge papers. The developments made in the workshop will deliver a step change in a variety of video applications.
Kwanghoon Sohn, Ming-Hsuan Yang 0001, Hyeran Byun, Jongwoo Lim, Gee-Sern Hsu, Stephen Lin 0001, Euntai Kim, Seungryong Kim
ACM Multimedia1
2018 Session details: Demo + Video + Makers' Program
Kwanghoon Sohn, Yong Man Ro
ACM Multimedia1
2018 Recurrent Transformer Networks for Semantic Correspondence
abstract
We present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convolutional activations. By directly estimating the transformations between an image pair, rather than employing spatial transformer networks to independently normalize each individual image, we show that greater accuracy can be achieved. This process is conducted in a recursive manner to refine both the transformation estimates and the feature representations. In addition, a technique is presented for weakly-supervised training of RTNs that is based on a proposed classification loss. With RTNs, state-of-the-art performance is attained on several benchmarks for semantic correspondence.
Seungryong Kim, Stephen Lin 0001, Sangryul Jeon, Dongbo Min, Kwanghoon Sohn
NeurIPS5
2018 Unified multi-spectral pedestrian detection based on probabilistic fusion networks
Kihong Park, Seungryong Kim, Kwanghoon Sohn
Pattern Recognit.3
2018 Deep Monocular Depth Estimation via Integration of Global and Local Predictions
abstract
Recent works on machine learning have greatly advanced the accuracy of single image depth estimation. However, the resulting depth images are still over-smoothed and perceptually unsatisfying. This paper casts depth prediction from single image as a parametric learning problem. Specifically, we propose a deep variational model that effectively integrates heterogeneous predictions from two convolutional neural networks (CNNs), named global and local networks. They have contrasting network architecture and are designed to capture depth information with complementary attributes. These intermediate outputs are then combined in the integration network based on the variational framework. By unrolling the optimization steps of Split Bregman (SB) iterations in the integration network, our model can be trained in an end-to-end manner. This enables one to simultaneously learn an efficient parameterization of the CNNs and hyper-parameter in the variational method. Finally, we offer a new dataset of 0.22 million RGB-D images captured by Microsoft Kinect v2. Our model generates realistic and discontinuity-preserving depth prediction without involving any low-level segmentation or superpixels. Intensive experiments demonstrate the superiority of the proposed method in a range of RGB-D benchmarks including both indoor and outdoor scenarios.
Youngjung Kim, Hyungjoo Jung, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.4
2017 Deeply Aggregated Alternating Minimization for Image Restoration
abstract
Regularization-based image restoration has remained an active research topic in image processing and computer vision. It often leverages a guidance signal captured in different-fields as an additional cue. In this work, we present a general framework for image restoration, called deeply aggregated alternating minimization (DeepAM). We propose to train deep neural network to advance two of the steps in the conventional AM algorithm: proximal mapping and β-continuation. Both steps are learned from a large dataset in an end-to-end manner. The proposed framework enables the convolutional neural networks (CNNs) to operate as a regularizer in the AM algorithm. We show that our learned regularizer via deep aggregation outperforms the recent data-driven approaches as well as the nonlocal-based methods. The flexibility and effectiveness of our framework are demonstrated in several restoration tasks, including single image denoising, RGB-NIR restoration, and depth superresolution.
Youngjung Kim, Hyungjoo Jung, Dongbo Min, Kwanghoon Sohn
CVPR4
2017 FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence
abstract
We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to existing CNN-based descriptors, FCSS is inherently insensitive to intra-class appearance variations because of its LSS-based structure, while maintaining the precise localization ability of deep neural networks. The sampling patterns of local structure and the self-similarity measure are jointly learned within the proposed network in an end-to-end and multi-scale manner. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in existing image datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks.
Seungryong Kim, Dongbo Min, Bumsub Ham, Sangryul Jeon, Stephen Lin 0001, Kwanghoon Sohn
CVPR6
2017 DCTM: Discrete-Continuous Transformation Matching for Semantic Flow
abstract
Techniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks.
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
ICCV4
2017 Multispectral human co-segmentation via joint convolutional neural networks
abstract
We present a novel human body co-segmentation method for unregistered multispectral, color and thermal, images by leveraging CNNs. The main challenges for that tasks are no-alignment between color and thermal images and an absent of ground truth human segmentation labels. To solve these limitations, our key-insight is to formulate the segmentation network for each modality that solve two sub-tasks, correspondence and classification, in a joint and iterative manner. We formulate the learning framework between multispectral images in a way that training labels for one modality are used to learn the network for the other modality. We estimate dense correspondences between multispectral image pairs using intermediate convolutional activations of CNNs and perform human segmentation for each modality through the conditional random fields (CRF) optimization using unary and pairwise fusion. These two steps are formulated as an iterative framework, enables the network to converge on an optimal solution. Experimental results show that our proposed method outperforms conventional state-of-the-art methods on the VAP benchmark consisting of unregistered multispectral color and thermal images.
Sungil Choi, Seungryong Kim, Kihong Park, Kwanghoon Sohn
ICIP4
2017 Convolutional feature pyramid fusion via attention network
abstract
We present a novel fusion scheme between multiple intermediate convolutional features within convolutional neurual network (CNN) for dense correspondence estimation. In contrast to existing CNN-based descriptors that utilize a single convolutional activation, our approach jointly uses multiple intermediate features of CNN through the attention weight that balances the contribution of each features. We formulate the overall network as two sub-networks, correspondence network and attention network. The correspondence network is designed to provide multiple intermediate matching costs while the attention network is to learn the optimal weight between them. These two networks are learned in a joint manner to boost the correspondence estimation performance. Experiments demonstrate that our proposed method outperforms the state-of-the-art methods on various correspondence estimation tasks including depth estimation, optical flow, and semantic correspondence.
Sangryul Jeon, Seungryong Kim, Kwanghoon Sohn
ICIP3
2017 Convolutional cost aggregation for robust stereo matching
abstract
Although convolutional neural network (CNN)-based stereo matching methods have become increasingly popular thanks to their robustness, they primarily have been focused on the matching cost computation. By leveraging CNNs, we present a novel method for matching cost aggregation to boost the stereo matching performance. Our insight is to learn the convolution kernel within CNN architecture for cost aggregation in a fully convolutional manner. Tailored to cost aggregation problem, our method differs from handcrafted methods in terms of its convolutional aggregation through optimally learned CNNs. First, the matching cost is aggregated with cost volume unary network, and then optimized with explicit disparity boundary, estimated through disparity boundary pairwise network, within a global energy minimization. Experiments demonstrate that our method outperforms conventional hand-crafted aggregation methods.
Somi Jeong, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn
ICIP4
2017 Unsupervised stereo matching using correspondence consistency
abstract
Deep convolutional neural networks (CNNs) have shown revolutionary performance improvements for matching cost computation in stereo matching. However, conventional CNN-based approaches to learn the network in a supervised manner require a large number of ground-truth disparity maps, which limits their applicability. To overcome this limitation, we present a novel framework to learn a CNNs architecture for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. Exploiting the correspondence consistency between stereo images as supervision, our method selects the training samples in each iteration during network training and uses them to learn the network. To boost the performance, we also propose a multi-scale cost computation scheme. Experimental results show that our method outperforms the state-of-the-art methods including even supervised learning based methods on various benchmarks.
Sunghun Joung, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn
ICIP4
2017 Depth prediction from a single image with conditional adversarial networks
abstract
Recent works on machine learning have greatly advanced the accuracy of depth estimation from a single image. However, resulting depth images are still visually unsatisfactory, often producing poor boundary localization and spurious regions. In this paper, we formulate this problem from single images as a deep adversarial learning framework. A two-stage convolutional network is designed as a generator to sequentially predict global and local structures of the depth image. At the heart of our approach is a training criterion based on adversarial discriminator which attempts to distinguish between real and generated depth images as accurately as possible. Our model enables more realistic and structure-preserving depth prediction from a single image, compared to state-of-the-arts approaches. An experimental comparison demonstrates the effectiveness of our approach on large RGB-D dataset.
Hyungjoo Jung, Youngjung Kim, Dongbo Min, Changjae Oh, Kwanghoon Sohn
ICIP5
2017 Deep stereo confidence prediction for depth estimation
abstract
We present a novel method that predicts a confidence to improve the accuracy of an estimated depth map in stereo matching. In contrast to existing learning based approaches relying on hand-crafted confidence features, we cast this problem into a convolutional neural network, learned using both a matching cost volume and its associated disparity map. As the size of the matching cost volume varies depending on a search range of stereo image pairs, we propose to use a top-K matching probability volume layer so that an input size for convolutional layers remains unchanged. Experimental results demonstrate that the proposed method outperforms the state-of-the-art confidence estimation approaches on various benchmarks.
Sunok Kim, Dongbo Min, Bumsub Ham, Seungryong Kim, Kwanghoon Sohn
ICIP5
2017 Automatic 2D-to-3D conversion using multi-scale deep neural network
abstract
We present a multi-scale deep convolutional neural network (CNN) for the task of automatic 2D-to-3D conversion. Traditional methods, which make a virtual view from a reference view, consist of separate stages i.e., depth (or disparity) estimation for the reference image and depth image-based rendering (DIBR) with estimated depth. In contrast, we reformulate the view synthesis task as an image reconstruction problem with a spatial transformer module and directly make stereo image pairs with a unified CNN framework without ground-truth depth as a supervision. We further propose a multi-scale deep architecture to capture the large displacements between images from coarse-level and enhance the detail from fine-level. Experimental results demonstrate the effectiveness of the proposed method over state-of-the-art approaches both qualitatively and quantitatively on the KITTI driving dataset.
Jiyoung Lee 0005, Hyungjoo Jung, Youngjung Kim, Kwanghoon Sohn
ICIP4
2017 Pedestrian proposal generation using depth-aware scale estimation
abstract
In this work, we propose an efficient method that generates pedestrian proposals suitable for the autonomous vehicle. Our main intuition is that depth information provides an important cue to assign the scale of pedestrian proposals. Based on the observation that in a 3-D world coordinate the scales of pedestrians are almost similar, we formulate the scales of pedestrian patches by projecting 3-D models to an image plane with its corresponding depth. We also introduce a scale-aware binary description using both color and depth images. By using this descriptor, the regression models are trained to rank the pedestrian proposal candidates and adjust the proposal bounding boxes for an accurate localization. Our algorithm achieves significant performance gains compared to conventional proposal generation methods on the challenging KITTI dataset.
Kihong Park, Seungryong Kim, Kwanghoon Sohn
ICIP3
2017 Personness estimation for real-time human detection on mobile devices
abstract
One aim of detection proposal methods is to reduce the computational overhead of object detection. However, most of the existing methods have significant computational overhead for real-time detection on mobile devices. A fast and accurate proposal method of human detection called personness estimation is proposed, which facilitates real-time human detection on mobile devices and can be effectively integrated into part-based detection, achieving high detection performance at a low computational cost. Our work is based on two observations: (i) normed gradients, which are designed for generic objectness estimation, effectively generate high-quality detection proposals for the person category; (ii) fusing the normed gradients with color attributes improves the performance of proposal generation for human detection. Thus, the candidate windows generated by the personness estimation will very likely contain human subjects. The human detection is then guided by the candidate windows, offering high detection performance even when the detection task terminates prior to completion. This interruptible detection scheme, called anytime detection, enables real-time human detection on mobile devices. Furthermore, we introduce a new evaluation methodology called time-recall curves to practically evaluate our approach. The applicability of our proposed method is demonstrated in extensive experiments on a publicly available dataset and a real mobile device, facilitating acquisition and enhancement of portrait photographs (e.g. selfie) on widespread mobile platforms.
Kyuwon Kim, Changjae Oh, Kwanghoon Sohn
Expert Syst. Appl.3
2017 Robust interactive image segmentation using structure-aware labeling
Changjae Oh, Bumsub Ham, Kwanghoon Sohn
Expert Syst. Appl.3
2017 DASC: Robust Dense Descriptor for Multi-Modal and Multi-Spectral Correspondence Estimation
abstract
Establishing dense correspondences between multiple images is a fundamental task in many applications. However, finding a reliable correspondence between multi-modal or multi-spectral images still remains unsolved due to their challenging photometric and geometric variations. In this paper, we propose a novel dense descriptor, called dense adaptive self-correlation (DASC), to estimate dense multi-modal and multi-spectral correspondences. Based on an observation that self-similarity existing within images is robust to imaging modality variations, we define the descriptor with a series of an adaptive self-correlation similarity measure between patches sampled by a randomized receptive field pooling, in which a sampling pattern is obtained using a discriminative learning. The computational redundancy of dense descriptors is dramatically reduced by applying fast edge-aware filtering. Furthermore, in order to address geometric variations including scale and rotation, we propose a geometry-invariant DASC (GI-DASC) descriptor that effectively leverages the DASC through a superpixel-based representation. For a quantitative evaluation of the GI-DASC, we build a novel multi-modal benchmark as varying photometric and geometric conditions. Experimental results demonstrate the outstanding performance of the DASC and GI-DASC in many cases of dense multi-modal and multi-spectral correspondences.
Seungryong Kim, Dongbo Min, Bumsub Ham, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 LAT: Local area transform for cross modal correspondence matching
Seungchul Ryu, Seungryong Kim, Kwanghoon Sohn
Pattern Recognit.3
2017 Modality-Invariant Image Classification Based on Modality Uniqueness and Dictionary Learning
abstract
We present a unified framework for the image classification of image sets taken under varying modality conditions. Our method is motivated by a key observation that the image feature distribution is simultaneously influenced by the semantic-class and the modality category label, which limits the performance of conventional methods for that task. With this insight, we introduce modality uniqueness as a discriminative weight that divides each modality cluster from all other clusters. By leveraging the modality uniqueness, our framework is formulated as unsupervised modality clustering and classifier learning based on modality-invariant similarity kernel. Specifically, in the assignment step, each training image is first assigned to the most similar cluster according to its modality. In the update step, based on the current cluster hypothesis, the modality uniqueness and the sparse dictionary are updated. These two steps are formulated in an iterative manner. Based on the final clusters, a modality-invariant marginalized kernel is then computed, where the similarities between the reconstructed features of each modality are aggregated across all clusters. Our framework enables the reliable inference of semantic-class category for an image, even across large photometric variations. Experimental results show that our method outperforms conventional methods on various benchmarks, such as landmark identification under severely varying weather conditions, domain-adapting image classification, and RGB and near-infrared image classification.
Seungryong Kim, Rui Cai 0002, Kihong Park, Sunok Kim, Kwanghoon Sohn
IEEE Trans. Image Process.5
2017 Fast Domain Decomposition for Global Image Smoothing
abstract
Edge-preserving smoothing (EPS) can be formulated as minimizing an objective function that consists of data and regularization terms. At the price of high-computational cost, this global EPS approach is more robust and versatile than a local one that typically has a form of weighted averaging. In this paper, we introduce an efficient decomposition-based method for global EPS that minimizes the objective function of L2 data and (possibly non-smooth and non-convex) regularization terms in linear time. Different from previous decompositionbased methods, which require solving a large linear system, our approach solves an equivalent constrained optimization problem, resulting in a sequence of 1-D sub-problems. This enables applying fast linear time solver for weighted-least squares and -L1 smoothing problems. An alternating direction method of multipliers algorithm is adopted to guarantee fast convergence. Our method is fully parallelizable, and its runtime is even comparable to the state-of-the-art local EPS approaches. We also propose a family of fast majorization-minimization algorithms that minimize an objective with non-convex regularization terms. Experimental results demonstrate the effectiveness and flexibility of our approach in a range of image processing and computational photography applications.
Youngjung Kim, Dongbo Min, Bumsub Ham, Kwanghoon Sohn
IEEE Trans. Image Process.4
2017 Feature Augmentation for Learning Confidence Measure in Stereo Matching
abstract
Confidence estimation is essential for refining stereo matching results through a post-processing step. This problem has recently been studied using a learning-based approach, which demonstrates a substantial improvement on conventional simple non-learning based methods. However, the formulation of learning-based methods that individually estimates the confidence of each pixel disregards spatial coherency that might exist in the confidence map, thus providing a limited performance under challenging conditions. Our key observation is that the confidence features and resulting confidence maps are smoothly varying in the spatial domain, and highly correlated within the local regions of an image. We present a new approach that imposes spatial consistency on the confidence estimation. Specifically, a set of robust confidence features is extracted from each superpixel decomposed using the Gaussian mixture model, and then these features are concatenated with pixel-level confidence features. The features are then enhanced through adaptive filtering in the feature domain. In addition, the resulting confidence map, estimated using the confidence features with a random regression forest, is further improved through K-nearest neighbor based aggregation scheme on both pixel- and superpixel-level. To validate the proposed confidence estimation scheme, we employ cost modulation or ground control points based optimization in stereo matching. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches on various benchmarks including challenging outdoor scenes.
Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn
IEEE Trans. Image Process.4
2016 Point-Cut: Interactive Image Segmentation Using Point Supervision
Changjae Oh, Bumsub Ham, Kwanghoon Sohn
ACCV (1)3
2016 Deep Self-correlation Descriptor for Dense Cross-Modal Correspondence
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn
ECCV (8)4
2016 Unified Depth Prediction and Intrinsic Image Decomposition from a Single Image via Joint Convolutional Neural Fields
Seungryong Kim, Kihong Park, Kwanghoon Sohn, Stephen Lin 0001
ECCV (8)3
2016 ANCC flow: Adaptive normalized cross-correlation with evolving guidance aggregation for dense correspondence estimation
abstract
Adaptive normalized cross-correlation (ANCC) cost function works well between images under photometric distortions, but its heavy computational burden often limits its applications. To overcome this limitation, this paper proposes a robust and efficient computational framework, called ANCC flow, designed for establishing dense correspondences between images under severe photometric variations. We first simplify the weight of ANCC in an asymmetric manner by considering a source image weight only. It is then efficiently computed by applying constant-time edge-aware filters without loss of its matching accuracy. Additionally, to deal with a large discrete label space effectively, which is a challenging issue in a flow field estimation, we propose a randomized label space sampling strategy similar to PatchMatch filer (PMF) optimization. The robustness of the asymmetric ANCC and the cost filter is further enhanced through an evolving weight computation, where a flow field computed in a previous iteration is utilized to build current edge-aware weights. Experimental results demonstrate the outstanding performance of ANCC flow in many cases of dense correspondence estimations under severe photometric and geometric variations.
Seungryong Kim, Dongbo Min, Kwanghoon Sohn
ICIP3
2016 Edge-aware image smoothing using commute time distances
abstract
Most edge-aware smoothing methods are based on the Euclidean distance to measure the similarity between adjacent pixels. This paper exploits the properties of the commute time to extend the notion of “similarity” in this context. The intuition is that since the commute time reflects the effect of all possible weighted paths between nodes (pixels), it can account for the global distribution of image features. The commute time is characterized by eigenvectors of a large Laplacian matrix, which is very costly even with sophisticated eigen-solver. To this end, we further employ a multiscale algorithm for approximating the eigenvector computation efficiently. It is analogous to the classical Nystrom's method for low rank matrix approximation. However, we do not depend on long-range connections between nodes, allowing one to include spatial coordinates in defining feature space. Extensive experimental validation demonstrates the benefits of using the commute time in a range of image processing applications, such as edge-aware image smoothing, texture filtering, and local edit propagation.
Youngjung Kim, Changjae Oh, Kwanghoon Sohn
ICIP3
2016 Multi-spectral pedestrian detection based on accumulated object proposal with fully convolutional networks
abstract
This paper presents a method for detecting a pedestrian by leveraging multi-spectral image pairs. Our approach is based on the observation that a multi-spectral image, especially far-infrared (FIR) image, enables us to overcome inherent limitations for pedestrian detection under challenging circumstances, such as even dark environments. For that task, multi-spectral color-FIR image pairs are used in a synergistic manner for pedestrian detection through deep convolutional neural networks (CNNs) learning and support vector regression (SVR). For inferring the confidence of a pedestrian, we first learn CNNs between color images (or FIR images) and bounding box annotations of pedestrians, respectively. Furthermore, for each object proposal, we extract intermediate activation features from network, and learn the probability of pedestrian using SVR. To improve the detection performance, the learned probability of pedestrian for each proposal is accumulated on the image domain. Based on the pedestrian confidence estimated from each network and accumulated pedestrian probabilities, the most probable pedestrian is finally localized among object proposal candidates. Thanks to its high robustness of multi-spectral imaging in dark environments and its high discriminative power of deep CNNs, our framework is shown to surpass state-of-the-art pedestrian detection methods on multi-spectral pedestrian benchmark.
Seungryong Kim, Kihong Park, Kwanghoon Sohn
ICPR4
2016 Fast illumination-robust foreground detection using hierarchical distribution map for real-time video surveillance system
Jongin Son, Seungryong Kim, Kwanghoon Sohn
Expert Syst. Appl.3
2016 Real-time rear obstacle detection using reliable disparity for driver assistance
Hunjae Yoo, Jongin Son, Bumsub Ham, Kwanghoon Sohn
Expert Syst. Appl.4
2016 EMCCD color correction based on spectral sensitivity analysis
Jongin Son, Minsung Kang, Dongbo Min, Kwanghoon Sohn
Multim. Tools Appl.4
2016 Structure Selective Depth Superresolution for RGB-D Cameras
abstract
This paper describes a method for high-quality depth superresolution. The standard formulations of image-guided depth upsampling, using simple joint filtering or quadratic optimization, lead to texture copying and depth bleeding artifacts. These artifacts are caused by inherent discrepancy of structures in data from different sensors. Although there exists some correlation between depth and intensity discontinuities, they are different in distribution and formation. To tackle this problem, we formulate an optimization model using a nonconvex regularizer. A nonlocal affinity established in a high-dimensional feature space is used to offer precisely localized depth boundaries. We show that the proposed method iteratively handles differences in structure between depth and intensity images. This property enables reducing texture copying and depth bleeding artifacts significantly on a variety of range data sets. We also propose a fast alternating direction method of multipliers algorithm to solve our optimization problem. Our solver shows a noticeable speed up compared with the conventional majorize-minimize algorithm. Extensive experiments with synthetic and real-world data sets demonstrate that the proposed method is superior to the existing methods.
Youngjung Kim, Bumsub Ham, Changjae Oh, Kwanghoon Sohn
IEEE Trans. Image Process.4
2015 Real-time Human Detection based on Personness Estimation
abstract
In this work, we study a real-time human detection method for mobile devices using window proposals. We find that the normed gradients, designed for generic objectness estimation, are also able to rapidly generate high quality object windows for a singlecategory object. We also notice that fusing the normed gradients with additional color feature improves the performance of objectness estimation for the single-category object. Based on these observations, we propose an efficient method, which we call personness estimation, to produce candidate windows that are highly likely to contain a person. The produced candidate windows are used to search over feature maps of an image so that a human detection method can achieve high detection performance within a short period of time. We further present how personness estimation can be efficiently combined into part-based human detection. Our experiments indicate that the proposed method is directly applicable to mobile devices, and allows real-time human detection.
Kyuwon Kim, Kwanghoon Sohn
BMVC2
2015 Randomized Global Transformation Approach for Dense Correspondence
Kihong Park, Seungryong Kim, Seungchul Ryu, Kwanghoon Sohn
BMVC4
2015 DASC: Dense adaptive self-correlation descriptor for multi-modal and multi-spectral correspondence
abstract
Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramatically advanced in recent studies. However, finding reliable visual correspondence in multi-modal or multi-spectral images still remains unsolved. In this paper, we propose a novel dense matching descriptor, called dense adaptive self-correlation (DASC), to effectively address this kind of matching scenarios. Based on the observation that a self-similarity existing within images is less sensitive to modality variations, we define the descriptor with a series of an adaptive self-correlation similarity for patches within a local support window. To further improve the matching quality and runtime efficiency, we propose a randomized receptive field pooling, in which a sampling pattern is optimized with a discriminative learning. Moreover, the computational redundancy that arises when computing densely sampled descriptor over an entire image is dramatically reduced by applying fast edge-aware filtering. Experiments demonstrate the outstanding performance of the DASC descriptor in many cases of multi-modal and multi-spectral correspondence.
Seungryong Kim, Dongbo Min, Bumsub Ham, Seungchul Ryu, Minh N. Do, Kwanghoon Sohn
CVPR6
2015 A majorize-minimize approach for high-quality depth upsampling
abstract
This paper describes a non-convex model that is carefully designed for high quality depth upsampling. Modern depth sensors such as time-of-flight cameras provide a promising depth measurement with video rate, but suffer from noise and low resolution. To tackle these limitations, we formulate an optimization problem using a robust potential function. In this formulation, a nonlocal principle established in the high-dimensional feature space is used to disambiguate the up-sampling problem. We also derive a numerical algorithm based on the majorization-minimization approach for efficient optimization. The proposed model iteratively creates a new affinity space that determines the influence of neighboring pixels by jointly considering spatial distance, appearance, and current estimates. This behavior enables one to significantly reduce annoying artifacts on a variety of range dataset, including a challenging real measurement. Extensive experiments demonstrate that the proposed model achieves competitive performance with state-of-the-art methods.
Youngjung Kim, Sunghwan Choi, Changjae Oh, Kwanghoon Sohn
ICIP4
2015 Learning depth from a single image using visual-depth words
abstract
Estimating depth from a single monocular image is a fundamental problem in computer vision. Traditional methods for such estimation usually require complicated and sometimes labor-intensive processing. In this paper, we propose a new perspective for this problem and suggest a new gradient-domain learning framework which is much simpler and more efficient. Inspired by the observation that there is substantial co-occurrence of image edges and depth discontinuities in natural scenes, we learn the relationship between local appearance features and corresponding depth gradients by making use of the K-means clustering algorithm within the image feature space. We then encode each cluster centroid with its associated depth gradients, which defines visual-depth words that model the image-depth relationship very well. This enables one to estimate the scene depth for an arbitrary image by simply selecting proper depth gradients from a compact dictionary of visual-depth words, followed by a Poisson surface reconstruction. Experimental results demonstrate that the proposed gradient-domain approach outperforms state-of-the-art methods both qualitatively and quantitatively and is generic over (unseen) scene categories which are not used for training.
Sunok Kim, Sunghwan Choi, Kwanghoon Sohn
ICIP3
2015 Sparse edit propagation for high resolution image using support vector machines
abstract
In this paper, we formulate image edit propagation as a task of machine learning to handle a high resolution image efficiently. Conventional graph-based methods solve the edit propagation by minimizing an energy function which considers the relationship between a reference pixel and its spatially neighboring ones. It is becoming a time-consuming and memory-requiring task due to the increase of the image size. Inspired by the observation that similar features get analogous edits, the edit propagation is casted as a classification problem using support vector machines in the feature space. A classifier is trained with initial sparse edits given by user interaction, and then the rest of the features are classified and manipulated. In experiments, the proposed method is applied to an image recoloring to verify the performance. Experimental results show that the proposed method gives competitive editing results comparing to other state-of-the-art methods.
Changjae Oh, Seungchul Ryu, Youngjung Kim, Jihyun Kim 0009, Taewoong Park, Kwanghoon Sohn
ICIP6
2015 Fast affine-invariant image matching based on global Bhattacharyya measure with adaptive tree
abstract
Establishing visual correspondence is one of the most fundamental tasks in many applications of computer vision fields. In this paper we propose a robust image matching to address the affine variation problems between two images taken under different viewpoints. Unlike the conventional approach finding the correspondence with local feature matching on fully affine transformed-images, which provides many outliers with a time consuming scheme, our approach is to find only one global correspondence and then utilizes the local feature matching to estimate the most reliable inliers between two images. In order to estimate a global image correspondence very fast as varying affine transformation in affine space of reference and query images, we employ a Bhattacharyya similarity measure between two images. Furthermore, an adaptive tree with affine transformation model is employed to dramatically reduce the computational complexity. Our approach represents the satisfactory results for severe affine transformed-images while providing a very low computational time. Experimental results show that the proposed affine-invariant image matching is twice faster than the state-of-the-art methods at least, and provides better correspondence performance under viewpoint change conditions.
Jongin Son, Seungryong Kim, Kwanghoon Sohn
ICIP3
2015 Rear obstacle detection system with fisheye stereo camera using HCT
Deukhyeon Kim, Jinwook Choi, Hunjae Yoo, Ukil Yang, Kwanghoon Sohn
Expert Syst. Appl.5
2015 A multi-vision sensor-based fast localization system with image matching for challenging outdoor environments
Jongin Son, Seungryong Kim, Kwanghoon Sohn
Expert Syst. Appl.3
2015 Real-time illumination invariant lane detection for lane departure warning system
Jongin Son, Hunjae Yoo, Kwanghoon Sohn
Expert Syst. Appl.4
2015 Depth Analogy: Data-Driven Approach for Single Image Depth Estimation Using Gradient Samples
abstract
Inferring scene depth from a single monocular image is a highly ill-posed problem in computer vision. This paper presents a new gradient-domain approach, called depth analogy, that makes use of analogy as a means for synthesizing a target depth field, when a collection of RGB-D image pairs is given as training data. Specifically, the proposed method employs a non-parametric learning process that creates an analogous depth field by sampling reliable depth gradients using visual correspondence established on training image pairs. Unlike existing data-driven approaches that directly select depth values from training data, our framework transfers depth gradients as reconstruction cues, which are then integrated by the Poisson reconstruction. The performance of most conventional approaches relies heavily on the training RGB-D data used in the process, and such a dependency severely degenerates the quality of reconstructed depth maps when the desired depth distribution of an input image is quite different from that of the training data, e.g., outdoor versus indoor scenes. Our key observation is that using depth gradients in the reconstruction is less sensitive to scene characteristics, providing better cues for depth recovery. Thus, our gradient-domain approach can support a great variety of training range datasets that involve substantial appearance and geometric variations. The experimental results demonstrate that our (depth) gradient-domain approach outperforms existing data-driven approaches directly working on depth domain, even when only uncorrelated training datasets are available.
Sunghwan Choi, Dongbo Min, Bumsub Ham, Youngjung Kim, Changjae Oh, Kwanghoon Sohn
IEEE Trans. Image Process.6
2015 Unsupervised Texture Flow Estimation Using Appearance-Space Clustering and Correspondence
abstract
This paper presents a texture flow estimation method that uses an appearance-space clustering and a correspondence search in the space of deformed exemplars. To estimate the underlying texture flow, such as scale, orientation, and texture label, most existing approaches require a certain amount of user interactions. Strict assumptions on a geometric model further limit the flow estimation to such a near-regular texture as a gradient-like pattern. We address these problems by extracting distinct texture exemplars in an unsupervised way and using an efficient search strategy on a deformation parameter space. This enables estimating a coherent flow in a fully automatic manner, even when an input image contains multiple textures of different categories. A set of texture exemplars that describes the input texture image is first extracted via a medoid-based clustering in appearance space. The texture exemplars are then matched with the input image to infer deformation parameters. In particular, we define a distance function for measuring a similarity between the texture exemplar and a deformed target patch centered at each pixel from the input image, and then propose to use a randomized search strategy to estimate these parameters efficiently. The deformation flow field is further refined by adaptively smoothing the flow field under guidance of a matching confidence score. We show that a local visual similarity, directly measured from appearance space, explains local behaviors of the flow very well, and the flow field can be estimated very efficiently when the matching criterion meets the randomized search strategy. Experimental results on synthetic and natural images show that the proposed method outperforms existing methods.
Sunghwan Choi, Dongbo Min, Bumsub Ham, Kwanghoon Sohn
IEEE Trans. Image Process.4
2015 Depth Superresolution by Transduction
abstract
This paper presents a depth superresolution (SR) method that uses both of a low-resolution (LR) depth image and a high-resolution (HR) intensity image. We formulate depth SR as a graph-based transduction problem. In particular, the HR intensity image is represented as an undirected graph, in which pixels are characterized as vertices, and their relations are encoded as an affinity function. When the vertices initially labeled with certain depth hypotheses (from the LR depth image) are regarded as input queries, all the vertices are scored with respect to the relevances to these queries by a classifying function. Each vertex is then labeled with the depth hypothesis that receives the highest relevance score. We design the classifying function by considering the local and global structures of the HR intensity image. This approach enables us to address a depth bleeding problem that typically appears in current depth SR methods. Furthermore, input queries are assigned in a probabilistic manner, making depth SR robust to noisy depth measurements. We also analyze existing depth SR methods in the context of transduction, and discuss their theoretic relations. Intensive experiments demonstrate the superiority of the proposed method over state-of-the-art methods both qualitatively and quantitatively.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.3
2014 Robust Stereo Matching Using Probabilistic Laplacian Surface Propagation
Seungryong Kim, Bumsub Ham, Seungchul Ryu, Seon Joo Kim, Kwanghoon Sohn
ACCV (1)5
2014 Randomized texture flow estimation using visual similarity
abstract
Exploring underlying texture flows defined with orientation and scale is of a great interest on a variety of vision-related tasks. However, existing methods often fail to capture accurate flows due to over-parameterization of texture deformation or employ a costly global optimization which makes the algorithm computationally demanding. In this paper, we address this inverse problem by casting it as a randomized correspondence search along with a locally-adaptive vector field smoothing. When a small example patch is given as a reference, a randomized deformable matching is performed on the very densely quantized label space, enabling an efficient estimation of texture deformation without quality degeneration, e.g., due to quantization artifacts which often appear in the optimization-driven discrete approaches. The visual similarity with respect to the deformation parameters is directly measured with an input texture image on an appearance space. The locally-adaptive smoothing is then applied to the intermediate flow field, resulting in a good continuation of the resultant texture flow. Experimental results on both synthetic and natural images show that the proposed method improves the performance in terms of both runtime efficiency and/or visual quality, compared to the existing methods.
Sunghwan Choi, Dongbo Min, Kwanghoon Sohn
ICIP3
2014 Data-driven single image depth estimation using weighted median statistics
abstract
In this paper, a data-driven approach is proposed for automatically estimating a plausible depth map from a single monocular image based on the weighted median statistics (WMS). Instead of using complicated parametric models for learning frameworks that are typically employed in existing methods, we cast the estimation as a simple yet effective statistical approach. It assigns perceptually proper depth values to an input image in accordance with a data-driven depth prior. Based on the assumption that similar scenes are likely to have similar depth structure, the depth prior is computed from the WMS of k-nearest neighbor 3D pairs in a large 3D image repository. We show that the WMS captures the underlying depth structure of the input image very well, even though the visual appearance of nearest neighbor images are not tightly aligned. The depth map is then inferred according to the depth prior by making use of the edge-aware image filtering technique, resulting in a discontinuity-preserving smooth depth map. Experimental results demonstrate that our method outperforms state-of-the-art methods in terms of both accuracy and efficiency.
Youngjung Kim, Sunghwan Choi, Kwanghoon Sohn
ICIP3
2014 Local self-similarity frequency descriptor for multispectral feature matching
abstract
This paper describes a robust feature descriptor called the local self-similarity frequency (LSSF) for the multispectral RGB-NIR feature matching, which uses the frequency response of the local internal layout of self-similarities. A nonlinear relationship between multi-spectral image pairs makes conventional descriptors be sensitive to spectral deformation. To alleviate this problem, the LSSF employs a weighted correlation surface reducing the discrepancy between mul-tispectral images. Furthermore, the LSSF provides a rotation invariance exploiting the frequency response of maximal values on logpolar bins based on the fact that a cyclic shift on the log-polar representation leads only a phase shift in a frequency domain. Experimental results show that LSSF outperforms state-of-the-art descriptors in terms of a recognition rate for multispectral RGB-NIR image pairs.
Seungryong Kim, Seungchul Ryu, Bumsub Ham, Junhyung Kim, Kwanghoon Sohn
ICIP5
2014 Synthesis quality prediction model based on distortion intolerance
abstract
Free-viewpoint video system will provide viewers with freedom to navigate through the scene at different viewpoints. In the system, arbitrary viewpoints of videos are synthesized by the depth image-based rendering with multi-view plus depth videos. Despite the widespread of technologies for free-viewpoint video system, the field of quality assessment for the free-viewpoint video, especially the quality prediction of a synthesized image, has not yet been thoroughly investigated. This paper analyzes how distortions in color and depth images influence on the quality of a synthesized image. Then, an objective quality prediction model for a synthesized image is proposed based on the concept of intolerance of synthesis distortion. Experimental results show that the proposed model provides outstanding performance in predicting the quality of a synthesized image compared to other models.
Seungchul Ryu, Seungryong Kim, Kwanghoon Sohn
ICIP3
2014 No-reference perceptual blur model based on inherent sharpness
abstract
An objective blurriness model is useful in various image processing applications. Especially, a no-reference model is expected as a highly desirable approach due to its applicability to wide range of applications. Blurriness of an image is known as to be commonly induced by the attenuation of spatial high-frequency and thus most conventional researches focused on a model estimating the amount of spatial high-frequency. However, the human-perceived blurriness might be varied across image contents. Very few researches have been investigated the human visual system model for the blurriness perception. To address the lack of an efficient model, this paper presents the blurriness perception model designed as a spatially varying function based on the inherent sharpness. Pixel-wise perceptual blurriness is computed employing the blurriness perception model and then integrated into an overall blurriness index using saliency information. The experimental comparisons with state-of-the-arts blurriness models for extensive public databases show that the proposed model is well-correlated with the subjective scores across different content of images, and outperforms the compared models.
Seungchul Ryu, Kwanghoon Sohn
ICIP2
2014 Reliability-Based Multiview Depth Enhancement Considering Interview Coherence
abstract
Color-plus-depth video format has been increasingly popular in 3-D video applications, such as auto-stereoscopic 3-D TV and freeview TV. The performance of these applications is heavily dependent on the quality of depth maps since intermediate views are synthesized using the corresponding depth maps. This paper presents a novel framework for obtaining high-quality multiview color-plus-depth video using a hybrid sensor, which consists of multiple color cameras and depth sensors. Given multiple high-resolution color images and low quality depth maps obtained from the color cameras and depth sensors, we improve the quality of the depth map corresponding to each color view by increasing its spatial resolution and enforcing interview coherence. Specifically, a new up-sampling method considering the interview coherence is proposed to enhance multiview depth maps. This approach can improve the performance of the existing up-sampling algorithms, such as joint bilateral up-sampling and weighted mode filtering, which have been developed to enhance a single-view depth map only. In addition, an adaptive approach of fusing multiple input low-resolution depth maps is proposed based on the reliability that considers camera geometry and depth validity. The proposed framework can be extended into the temporal domain for temporally consistent depth maps. Experimental results demonstrate that the proposed method provides better multiview depth quality than the conventional single-view-based methods. We also show that it provides comparable results, yet much more efficiently, to other fusion approaches that employ both depth sensors and stereo matching algorithm together. Moreover, it is shown that the proposed method significantly reduces bit rates required to compress the multiview color-plus-depth video.
Jinwook Choi, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.3
2014 Mahalanobis Distance Cross-Correlation for Illumination-Invariant Stereo Matching
abstract
A robust similarity measure called the Mahalanobis distance cross-correlation (MDCC) is proposed for illumination-invariant stereo matching, which uses a local color distribution within support windows. It is shown that the Mahalanobis distance between the color itself and the average color is preserved under affine transformation. The MDCC converts pixels within each support window into the Mahalanobis distance transform (MDT) space. The similarity between MDT pairs is then computed using the cross-correlation with an asymmetric weight function based on the Mahalanobis distance. The MDCC considers correlation on cross-color channels, thus providing robustness to affine illumination variation. Experimental results show that the MDCC outperforms state-of-the-art similarity measures in terms of stereo matching for image pairs taken under different illumination conditions.
Seungryong Kim, Bumsub Ham, Bongjoe Kim, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.4
2014 No-Reference Quality Assessment for Stereoscopic Images Based on Binocular Quality Perception
abstract
Quality perception of 3-D images is one of the most important parameters for accelerating advances in 3-D imaging fields. Despite active research in recent years for understanding the quality perception of 3-D images, binocular quality perception of asymmetric distortions in stereoscopic images is not thoroughly comprehended. In this paper, we explore the relationship between the perceptual quality of stereoscopic images and visual information, and introduce a model for binocular quality perception. Based on this binocular quality perception model, a no-reference quality metric for stereoscopic images is proposed. The proposed metric is a top-down method modeling the binocular quality perception of the human visual system in the context of blurriness and blockiness. Perceptual blurriness and blockiness scores of left and right images were computed using local blurriness, blockiness, and visual saliency information and then combined into an overall quality index using the binocular quality perception model. Experiments for image and video databases show that the proposed metric provides consistent correlations with subjective quality scores. The results also show that the proposed metric provides higher performance than existing full-reference methods even though the proposed method is a no-reference approach.
Seungchul Ryu, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.2
2014 Probability-Based Rendering for View Synthesis
abstract
In this paper, a probability-based rendering (PBR) method is described for reconstructing an intermediate view with a steady-state matching probability (SSMP) density function. Conventionally, given multiple reference images, the intermediate view is synthesized via the depth image-based rendering technique in which geometric information (e.g., depth) is explicitly leveraged, thus leading to serious rendering artifacts on the synthesized view even with small depth errors. We address this problem by formulating the rendering process as an image fusion in which the textures of all probable matching points are adaptively blended with the SSMP representing the likelihood that points among the input reference images are matched. The PBR hence becomes more robust against depth estimation errors than existing view synthesis approaches. The MP in the steady-state, SSMP, is inferred for each pixel via the random walk with restart (RWR). The RWR always guarantees visually consistent MP, as opposed to conventional optimization schemes (e.g., diffusion or filtering-based approaches), the accuracy of which heavily depends on parameters used. Experimental results demonstrate the superiority of the PBR over the existing view synthesis approaches both qualitatively and quantitatively. Especially, the PBR is effective in suppressing flicker artifacts of virtual video rendering although no temporal aspect is considered. Moreover, it is shown that the depth map itself calculated from our RWR-based method (by simply choosing the most probable matching point) is also comparable with that of the state-of-the-art local stereo matching methods.
Bumsub Ham, Dongbo Min, Changjae Oh, Minh N. Do, Kwanghoon Sohn
IEEE Trans. Image Process.5
2014 Fast Global Image Smoothing Based on Weighted Least Squares
abstract
This paper presents an efficient technique for performing a spatially inhomogeneous edge-preserving image smoothing, called fast global smoother. Focusing on sparse Laplacian matrices consisting of a data term and a prior term (typically defined using four or eight neighbors for 2D image), our approach efficiently solves such global objective functions. In particular, we approximate the solution of the memory-and computation-intensive large linear system, defined over a d-dimensional spatial domain, by solving a sequence of 1D subsystems. Our separable implementation enables applying a linear-time tridiagonal matrix algorithm to solve d three-point Laplacian matrices iteratively. Our approach combines the best of two paradigms, i.e., efficient edge-preserving filters and optimization-based smoothing. Our method has a comparable runtime to the fast edge-preserving filters, but its global optimization formulation overcomes many limitations of the local filtering approaches. Our method also achieves high-quality results as the state-of-the-art optimization-based techniques, but runs ∼10-30 times faster. Besides, considering the flexibility in defining an objective function, we further propose generalized fast algorithms that perform Lγ norm smoothing (0 < γ < 2) and support an aggregated (robust) data term for handling imprecise data constraints. We demonstrate the effectiveness and efficiency of our techniques in a range of image processing and computer graphics applications.
Dongbo Min, Sunghwan Choi, Jiangbo Lu, Bumsub Ham, Kwanghoon Sohn, Minh N. Do
IEEE Trans. Image Process.5
2013 Blind blockiness measure based on marginal distribution of wavelet coefficient and saliency
abstract
The objective measurement of blockiness plays an important role in many applications, such as the quality assessment of an image, and the design of image and video coding system. However, most of the existing no-reference blockiness metrics do not consider important influences of grid distortion of an image on the performance of the metric. In this paper, we propose a new blockiness metric, which is robust to grid distortion, based on the marginal distribution of local wavelet coefficients and saliency information. Experiments for several public image databases showed that the proposed metric provides consistent correlations with subjective blockiness scores and outperforms other existing no-reference blockiness metrics.
Seungchul Ryu, Kwanghoon Sohn
ICASSP2
2013 Fast image retargeting via axis-aligned importance scaling
abstract
In this paper, we propose an image retargeting method that resizes an image by the axis-aligned importance scaling. The proposed method operates on the axis-aligned deformation space, where the mesh structure is parameterized in the 1D vector, i.e., quads on the same column (or row) share the single parameter. The unknown variables are thus dependent on the grid resolution only, allowing a fast and simple implementation on a moderate CPU. The optimal parameters are inferred in an iterative manner: these parameters are updated by scaling initial ones according to the deformation error. It is measured at each iteration by aggregating the transition cost for the deformed quad. Experimental results show that the proposed method preserves visually salient features without foldover artifacts better than competing methods. In addition, the optimal parameters can be calculated within 0.02 ms on a single-core CPU.
Sunghwan Choi, Bumsub Ham, Kwanghoon Sohn
ICIP3
2013 ABFT: Anisotropic binary feature transform based on structure tensor space
abstract
Local feature matching is a fundamental step for many computer vision applications. Recently, binary feature transforms have been popularly proposed to improve the computational efficiency while preserving high matching performance. However, it is sensitive to noise and geometrical distortion such as affine transformation. In this paper, we propose ABFT framework, composed of a noise robust feature detection and affine invariant binary feature description based on a structure tensor space. Experimental results show that ABFT outperforms other state-of-the-art feature transforms in terms of the repeatability, recognition rate, and computational time.
Seungryong Kim, Hunjae Yoo, Seungchul Ryu, Bumsub Ham, Kwanghoon Sohn
ICIP5
2013 Contextual information based visual saliency model
abstract
Automatic detection of visual saliency has been considered a very important task because of a wide range of applications such as object detection, image quality assessment, image segmentation, and more. Thanks to active researches in this field, many effective saliency models have been developed. Nevertheless, several challenging problems are still remain unsolved, such as detecting saliency in complex scene and providing high resolution and accurate saliency maps. In order to address such challenging problems, we propose a visual saliency model based on the concept of contextual information. First, we introduce a general framework for detecting saliency of an image using contextual information. Then, the proposed saliency model based on color and shape features is proposed. Quantitative and qualitative comparisons with seven state-of-the-art models on the public database show that the proposed model achieves excellent performance. Especially, the proposed model can provide good performance on challenging images including images with cluttered background and repeating distractors compared to the other models.
Seungchul Ryu, Bumsub Ham, Kwanghoon Sohn
ICIP3
2013 Advanced motion vector coding framework for multiview video sequences
Seungchul Ryu, Jungdong Seo, Kwanghoon Sohn
Multim. Tools Appl.4
2013 Exact order based feature descriptor for illumination robust image matching
Bongjoe Kim, Hunjae Yoo, Kwanghoon Sohn
Pattern Recognit.3
2013 Visual Comfort Enhancement for Stereoscopic Video Based on Binocular Fusion Characteristics
abstract
A well-known problem in stereoscopic videos is visual fatigue. However, conventional depth adjustment methods provide little guidance in deciding the amount of depth control or determining the condition for depth control. We propose a depth adjustment method based on binocular fusion characteristics, where the fusion time is used as the parameter for adjustment. Visual comfort enhancement is implemented with the horizontal image shift approach for depth adjustment. The speeded-up robust feature is used to estimate maximum disparity and disparity distribution, while a face detection algorithm is used to estimate the viewing distance with a single web camera. Binocular fusion characteristics are acquired in advance by measuring the time required for fusion under various conditions, including foreground disparity, background disparity, and focal distance for random-dot stereograms. Finally, a subjective evaluation is conducted in fixed and free-to-move viewing conditions, and the results show that comfortable videos were generated based on the proposed depth adjustment method.
Donghyun Kim 0010, Sunghwan Choi, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.3
2013 Space-Time Hole Filling With Random Walks in View Extrapolation for 3D Video
abstract
In this paper, a space-time hole filling approach is presented to deal with a disocclusion when a view is synthesized for the 3D video. The problem becomes even more complicated when the view is extrapolated from a single view, since the hole is large and has no stereo depth cues. Although many techniques have been developed to address this problem, most of them focus only on view interpolation. We propose a space-time joint filling method for color and depth videos in view extrapolation. For proper texture and depth to be sampled in the following hole filling process, the background of a scene is automatically segmented by the random walker segmentation in conjunction with the hole formation process. Then, the patch candidate selection process is formulated as a labeling problem, which can be solved with random walks. The patch candidates that best describe the hole region are dynamically selected in the space-time domain, and the hole is filled with the optimal patch for ensuring both spatial and temporal coherence. The experimental results show that the proposed method is superior to state-of-the-art methods and provides both spatially and temporally consistent results with significantly reduced flicker artifacts.
Sunghwan Choi, Bumsub Ham, Kwanghoon Sohn
IEEE Trans. Image Process.3
2013 Revisiting the Relationship Between Adaptive Smoothing and Anisotropic Diffusion With Modified Filters
abstract
Anisotropic diffusion has been known to be closely related to adaptive smoothing and discretized in a similar manner. This paper revisits a fundamental relationship between two approaches. It is shown that adaptive smoothing and anisotropic diffusion have different theoretical backgrounds by exploring their characteristics with the perspective of normalization, evolution step size, and energy flow. Based on this principle, adaptive smoothing is derived from a second order partial differential equation (PDE), not a conventional anisotropic diffusion, via the coupling of Fick's law with a generalized continuity equation where a "source" or "sink" exists, which has not been extensively exploited. We show that the source or sink is closely related to the asymmetry of energy flow as well as the normalization term of adaptive smoothing. It enables us to analyze behaviors of adaptive smoothing, such as the maximum principle and stability with a perspective of a PDE. Ultimately, this relationship provides new insights into application-specific filtering algorithm design. By modeling the source or sink in the PDE, we introduce two specific diffusion filters, the robust anisotropic diffusion and the robust coherence enhancing diffusion, as novel instantiations which are more robust against the outliers than the conventional filters.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.3
2013 A Generalized Random Walk With Restart and its Application in Depth Up-Sampling and Interactive Segmentation
abstract
In this paper, the origin of random walk with restart (RWR) and its generalization are described. It is well known that the random walk (RW) and the anisotropic diffusion models share the same energy functional, i.e., the former provides a steady-state solution and the latter gives a flow solution. In contrast, the theoretical background of the RWR scheme is different from that of the diffusion-reaction equation, although the restarting term of the RWR plays a role similar to the reaction term of the diffusion-reaction equation. The behaviors of the two approaches with respect to outliers reveal that they possess different attributes in terms of data propagation. This observation leads to the derivation of a new energy functional, where both volumetric heat capacity and thermal conductivity are considered together, and provides a common framework that unifies both the RW and the RWR approaches, in addition to other regularization methods. The proposed framework allows the RWR to be generalized (GRWR) in semilocal and nonlocal forms. The experimental results demonstrate the superiority of GRWR over existing regularization approaches in terms of depth map up-sampling and interactive image segmentation.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.3
2013 Gradient-Enhancing Conversion for Illumination-Robust Lane Detection
abstract
Lane detection is important in many advanced driver-assistance systems (ADAS). Vision-based lane detection algorithms are widely used and generally use gradient information as a lane feature. However, gradient values between lanes and roads vary with illumination change, which degrades the performance of lane detection systems. In this paper, we propose a gradient-enhancing conversion method for illumination-robust lane detection. Our proposed gradient-enhancing conversion method produces a new gray-level image from an RGB color image based on linear discriminant analysis. The converted images have large gradients at lane boundaries. To deal with illumination changes, the gray-level conversion vector is dynamically updated. In addition, we propose a novel lane detection algorithm, which uses the proposed conversion method, adaptive Canny edge detector, Hough transform, and curve model fitting method. We performed several experiments in various illumination environments and confirmed that the gradient is maximized at lane boundaries on the road. The detection rate of the proposed lane detection algorithm averages 96% and is greater than 93% in very poor environments.
Hunjae Yoo, Ukil Yang, Kwanghoon Sohn
IEEE Trans. Intell. Transp. Syst.3
2012 Probabilistic Correspondence Matching using Random Walk with Restart
abstract
This paper presents a probabilistic method for correspondence matching with a framework of the random walk with restart (RWR). The matching cost is reformulated as a corresponding probability, which enables the RWR to be utilized for matching the correspondences. There are mainly two advantages in our method. First, the proposed method guarantees the non-trivial steady-state solution of a given initial matching probability due to the restarting term in the RWR. It means the number of iteration, a crucial parameter which influences the performance of algorithm, is not needed in contrast to the conventional methods. This gives the consistent results regardless of the evolution time. Second, only an adjacent neighborhood is considered when the matching probabilities are inferred, which lowers the computational complexity while not sacrificing performance. Experimental results show that the performance of the proposed method is competitive to that of state-of-the-art methods both qualitatively and quantitatively.
Changjae Oh, Bumsub Ham, Kwanghoon Sohn
BMVC3
2012 Real-time depth range estimation and its application to mobile stereo camera
abstract
This paper proposes a real-time depth range estimation algorithm that produces a specialized depth map for mobile stereo camera and its applications. The proposed algorithm effectively controls the unwanted depth noise by using a simple depth fusion technique which is based-on the properties of different matching window size. Search range estimation used for the proposed algorithm improves both the algorithm efficiency and the accuracy of the proposed depth range map. In order to show the performance of the proposed algorithm, shooting guide functions such as auto-disparity control and comfortable zone indicator functions are implemented by using the proposed algorithm. Additionally, the out of focus images are represented as an image enhancement function. The proposed algorithm and the experimental results support that the proposed algorithm effectively controls the background artifacts and provide accurate dynamic range of depth with low computational complexity.
Hyoungchul Shin, Kwanghoon Sohn
CCNC2
2012 Stereoscopic image quality metric based on binocular perception model
abstract
Measuring a perceptual quality of an image is one of the important tasks in various applications such as image coding, processing, enhancement, and monitoring system. Although active researches have been made for objective quality assessment of 2D images for some decades, still very few efforts have been concentrated on 3D image quality assessment. In this paper, we propose a new quality metric for stereoscopic images based on the binocular perception model considering asymmetric property of a stereoscopic image pair. Experiments for publicly available databases show that the proposed metric provides consistent correlations with subjective quality scores. The results also show that the proposed metric outperforms state-of-the-arts metrics.
Seungchul Ryu, Donghyun Kim 0010, Kwanghoon Sohn
ICIP3
2012 High-speed inter-view frame mode decision procedure for multi-view video coding
Xingang Liu, Laurence T. Yang, Kwanghoon Sohn
Future Gener. Comput. Syst.3
2012 Low-Cost H.264/AVC Inter Frame Mode Decision Algorithm for Mobile Communication Systems
Xingang Liu, Kwanghoon Sohn, Meikang Qiu, Minho Jo 0001, Hoh Peter In
Mob. Networks Appl.2
2012 Depth-based direct mode for multiview video coding
Seungchul Ryu, Kwanghoon Sohn
Signal Process. Image Commun.2
2012 Effect of Vergence-Accommodation Conflict and Parallax Difference on Binocular Fusion for Random Dot Stereogram
abstract
Recently, various studies of human factors have been conducted to reveal stereoscopic characteristics of the human visual system and visual fatigue. In this paper, we investigate the effect of vergence-accommodation conflict and parallax difference on binocular fusion for random dot stereograms. The aim of this paper is to provide a study on visual fatigue induced by the conflict. We measured the time required for fusion under various conditions that include foreground parallax, background parallax, focal distance, aperture size, and corrugation frequency. The results show that foreground parallax and parallax difference between foreground parallax and background parallax have significant influences on fusion time. In addition, we verify the relationship between fusion time and visual fatigue by conducting a subjective evaluation of stereoscopic images.
Donghyun Kim 0010, Sunghwan Choi, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.3
2012 Robust Scale-Space Filter Using Second-Order Partial Differential Equations
abstract
This paper describes a robust scale-space filter that adaptively changes the amount of flux according to the local topology of the neighborhood. In a manner similar to modeling heat or temperature flow in physics, the robust scale-space filter is derived by coupling Fick's law with a generalized continuity equation in which the source or sink is modeled via a specific heat capacity. The filter plays an essential part in two aspects. First, an evolution step size is adaptively scaled according to the local structure, enabling the proposed filter to be numerically stable. Second, the influence of outliers is reduced by adaptively compensating for the incoming flux. We show that classical diffusion methods represent special cases of the proposed filter. By analyzing the stability condition of the proposed filter, we also verify that its evolution step size in an explicit scheme is larger than that of the diffusion methods. The proposed filter also satisfies the maximum principle in the same manner as the diffusion. Our experimental results show that the proposed filter is less sensitive to the evolution step size, as well as more robust to various outliers, such as Gaussian noise, impulsive noise, or a combination of the two.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.3
2011 Complexity Control Scheme for H.264/AVC Inter Frame Encoding
abstract
In this paper, a complexity control scheme (CCS) for Inter frame mode decision (MD) is proposed for H.264/AVC encoder to speed-up the original encoding process. The information extracted from macroblock (MB), which can be used to pre-estimate the optimal mode of the MB is investigated and utilized to eliminate the redundant mode candidates. The simulation results show that the proposed algorithm can reduce over 80% Inter frame encoding time with little quality loss. It can be widely implemented in the mobile communication systems with H.264/AVC standard to realize the real-time video signal coding.
Xingang Liu, Laurence T. Yang, Kwanghoon Sohn
HPCC3
2011 Hole filling with random walks using occlusion constraints in view synthesis
abstract
In this paper, we propose a hole filling technique which coherently reconstructs the hole region during the view synthesis. The holes can be filled successfully in case that the virtual camera locates between real cameras by using interpolation. However, they cannot be handled in case that the virtual camera locates beyond the field of view of the real camera. We address this problem by jointly using image completion technique and random walks. First, occlusion constraint is imposed in order to guide the filling order. It is observed that the holes occur in a similar pattern because of the geometric characteristic of the camera configuration. This observation named vertical prior in this paper is also used to label each pixel on the fill front with foreground or background. Second, the probabilities estimated by random walks are utilized to find the patch candidates and to select the optimal patch. The experimental results show that the proposed method gives visually pleasing results over both interpolation and conventional image completion method.
Sunghwan Choi, Bumsub Ham, Kwanghoon Sohn
ICIP3
2011 Cost aggregation with anisotropic diffusion in feature space for hybrid stereo matching
abstract
In this paper, we present a cost aggregation using anisotropic diffusion on a feature space for hybrid stereo matching. Stereo matching can be classified into two categories: feature-based and area-based approaches. Feature-based approaches generate accurate but sparse disparity maps. On the other hand, area-based approaches generate dense but unreliable disparity maps, especially at depth discontinuities and homogeneous regions. We hence propose a stereo matching algorithm having advantages of both approaches. We study how to design a correspondence algorithm without modeling any depth cues except disparity. A procedure of depth perception is modeled via anisotropic diffusion on the feature space in terms of coherence. Based on the assumption that similar local feature space has similar disparity, we define the feature space and its similarity and then introduce feature confidences into the proposed model. Experimental results show that the performance of the proposed method is comparable to that of the state-of-the-art methods.
Bumsub Ham, Dongbo Min, Kwanghoon Sohn
ICIP3
2011 Embedded compression algorithm using error-aware quantization and hybrid DPCM/BTC coding
abstract
An embedded compression (EC) is a technique that can compress on-the-fly the frame data when stored in video memory; it has been used to effectively reduce the system memory requirements. In this paper, we propose a lossy EC algorithm that uses hybrid coding scheme of feedforward DPCM and modified four-level BTC, which aims at the target specification of the 4 × 2 block-wise random access at 50 % compression ratio. The algorithmic characteristics of the proposed algorithm are twofold: the first is the error-aware quantization scheme to definitely reduce the total error in a block and the last is the hybrid DPCM/BTC coding scheme to effectively handle various texture patterns that can occur in a block. The architectural advantage stems from the feedforward DPCM that breaks off the inherent dependency between the prediction and quantization present in the conventional DPCM, which helps reduce the implementation burden as well as speed up the overall processing. Comparative evaluation shows that the proposed algorithm outperforms the existing ones by 0.77–6.39 dB.
Kiwon Yoo, Changsu Han, Useok Kang, Kwanghoon Sohn
ICME4
2011 Visual Fatigue Prediction for Stereoscopic Image
abstract
In this letter, we propose a visual fatigue prediction metric which can replace subjective evaluation for stereoscopic images. It detects stereoscopic impairments caused by inappropriate shooting parameters or camera misalignment which induces excessive horizontal and vertical disparities. Pearson's correlation was measured between the proposed metrics and the subjective results by usingk-fold cross-validation, acquiring ranges of 78-87% with sparse feature and 74-85% with dense feature.
Donghyun Kim 0010, Kwanghoon Sohn
IEEE Trans. Circuits Syst. Video Technol.2
2010 Visual fatigue evaluation and enhancement for 2D-plus-depth video
abstract
A 3D video is expected to be a representative technique of realistic system but still has some problems such as visual fatigue and headache. In this paper, we propose a visual fatigue evaluation algorithm to predict the degree of visual fatigue from a 2D-plus-depth video. Spatial and temporal characteristics of the depth video are main factors of visual fatigue for autostereoscopic displays. Using depth image directly, we estimate spatial and temporal complexities, depth position and scene movement of the 3D video. Then, the overall visual fatigue of the 3D video is evaluated to have higher correlation with subjective fatigue evaluation by a linear regression. Moreover we control the pixel value of depth image from the 3D video which may induce severe fatigue to make more comfortable 3D video. The results of proposed algorithm show a considerable correlation with subjective visual fatigue.
Jaeseob Choi, Donghyun Kim 0010, Bumsub Ham, Sunghwan Choi, Kwanghoon Sohn
ICIP5
2010 3D JBU based depth video filtering for temporal fluctuation reduction
abstract
In this paper, we propose a three-dimensional Joint Bilateral Up-sampling (3D JBU) for the depth video which can be applied to a 3DTV system based on 2D-plus-depth video. Recently, a number of researches have been done to improve a resolution and frame-rate of the depth video. We proposed a novel method that enhances depth video obtained by Time-of-Flight (TOF) sensor by combining it with Charge-coupled Device (CCD) camera [1]. However, this method does not consider the temporal coherence of the depth video. It may cause an eye fatigue on the 3D display and increase bit rates on video coding, since it is possible to generate a temporal fluctuation problem. Therefore, in order to solve these problems, we propose a 3D JBU model which is extended conventional JBU into the temporal domain of depth video. Experimental results show that depth video obtained by the proposed method provides satisfactory quality.
Jinwook Choi, Dongbo Min, Donghyun Kim 0010, Kwanghoon Sohn
ICIP4
2010 Depth adjustment for stereoscopic image using visual fatigue prediction and depth-based view synthesis
abstract
A well-known problem in stereoscopic images is visual fatigue. In order to reduce visual fatigue, we propose depth adjustment method which controls the amounts of parallax of stereoscopic images by using visual fatigue prediction and depth-based view synthesis. We predict the visual fatigue level by examining the horizontal and vertical disparity characteristics of 3D images, and if the contents are judged to contain visual fatigue, depth adjustment is applied by using depth-based view synthesis. We present a method for extracting disparity characteristics from sparse corresponding features. Depth-based view synthesis algorithm is proposed to handle the hole regions in rendering process. We measured the correlations between visual fatigue prediction metrics and the subjective results, acquiring the ranges of 79% to 85%. Then, we performed depth-based view synthesis to the contents which are predicted to cause visual fatigue. A subjective evaluation showed that the proposed depth adjustment method generated comfortable stereoscopic images.
Donghyun Kim 0010, Kwanghoon Sohn
ICME2
2010 An asymmetric post-processing for correspondence problem
Dongbo Min, Kwanghoon Sohn
Signal Process. Image Commun.2
2010 Hardware design of motion data decoding process for H.264/AVC
Kiwon Yoo, Kwanghoon Sohn
Signal Process. Image Commun.2
2009 Spatial and temporal up-conversion technique for depth video
abstract
This paper proposes a novel framework for up-conversion of depth video resolution both in spatial and in time domain. Time-of-flight (TOF) sensors are widely used in computer vision fields. Although TOF sensors provide depth video in real time, there are some problems in a sense that it provides a low resolution and a low frame-rate depth video. We propose a cheaper solution that enhances depth video obtained by TOF sensor by combining it with CCD camera. The proposed method provides high quality video as a cheaper solution for low resolution, and low frame-rate depth video. It is useful when depth video is used in various applications such as 3DTV, free-view TV, teleconference system. High-quality depth video can be obtained by motion compensated frame interpolation (MCFI) and extended joint bilateral upsampling (JBU). Experimental results show that depth video obtained by the proposed method has satisfactory quality.
Jinwook Choi, Dongbo Min, Bumsub Ham, Kwanghoon Sohn
ICIP4
2009 Virtual view rendering using super-resolution with multiview images
abstract
This paper presents a new approach to solve the problem of quality degradation of a synthesized view, when a virtual camera moves forward. Interpolation techniques using only two neighboring views are generally applied when a virtual view is synthesized. Because the size of an object increases when the virtual camera moves forward, conventional methods have usually addressed this problem by interpolation techniques in order to synthesize a virtual view. However, as it generates a degraded view such as blurred images, we prevent a synthesized view from being blurred by using more images in multiview camera configuration. That is, this problem is solved by applying super-resolution concept which reconstructs a high resolution image from several low resolution images. Data fusion is performed by geometric warping with disparity maps of the multiple images followed by deblurring. Experimental results show that the image quality can further be improved by reducing blurring and halo effects in comparison with the interpolation method.
Bumsub Ham, Dongbo Min, Jinwook Choi, Kwanghoon Sohn
ICIP4
2009 HVS-aware ROI-based illumination and color restoration
abstract
We present the framework for Human visual system (HVS)-aware Region of interest (ROI)-based Illumination and Color Restoration called ‘HRICR’ for restoring distorted images taken under the arbitrary illumination environment. The proposed method is effective for appropriate illumination and vivid color restoration as well as suppression of artifacts such as HALO and noise amplification. We introduce the ROI-based parameter estimation dealing with small shadow area against spacious well-exposed background in an image for the touch-screen camera. The perceptual difference threshold is obtained using just-noticeable-difference (JND) based on HVS in response to user interaction.
Heechul Han, Kwanghoon Sohn
ICIP2
2009 Practical Pan-Tilt-Zoom-Focus Camera Calibration for Augmented Reality
Juhyun Oh, Seungjin Nam, Kwanghoon Sohn
ICVS3
2009 Fast mode-decision for H.264/AVC based on inter-frame correlations
Hyunsuk Ko, Kiwon Yoo, Kwanghoon Sohn
Signal Process. Image Commun.3
2009 2D/3D freeview video generation for 3DTV system
Dongbo Min, Donghyun Kim 0010, SangUn Yun, Kwanghoon Sohn
Signal Process. Image Commun.4
2008 Stereo matching with asymmetric occlusion handling in weighted least square framework
abstract
This paper presents a novel method for stereo matching with occlusion handling. In order to estimate optimal cost, we define an energy function and solve the iterative equation with the numerical method. We improve performance and convergence rate by using several acceleration techniques. The proposed method is computationally efficient since it does not use color segmentation or any global optimization techniques. For occlusion handling, which has not been performed effectively by any conventional cost aggregation approaches, we combine the occlusion problem with the proposed minimization scheme. Asymmetric information is used so that few additional computational loads are necessary. Experimental results show that performance is comparable to that of many state-of-the-art methods.
Dongbo Min, Kwanghoon Sohn
ICASSP2
2008 2D/3D freeview video generation for 3DTV system
abstract
In this paper, we propose a new approach of synthesizing novel views in multiview camera configurations. We introduce the semi iV-view & iV-depth framework in order to estimate disparity maps efficiently and correctly. This framework reduces redundancy on disparity estimation by using information from neighboring views. The occlusion problem is handled by using cost functions computed with multiview images. The proposed method provides a 2D/3D freeview video. User can select 2D/3D modes of freeview video and control 3D depth perception by adjusting several parameters in 3D freeview video. Experimental results show that the proposed method yields the accurate disparity maps and provides seamless freeview videos.
Dongbo Min, Donghyun Kim 0010, Kwanghoon Sohn
ICIP3
2008 VLSI architecture design of motion vector processor for H.264/AVC
abstract
H.264/AVC has considerably complex derivation process of motion data in comparison with that of previous video standards. It mainly results from advanced motion vector prediction process to cope with various macroblock partitions and spatial/temporal direct modes. This paper addresses the efficient hardware design of the motion vector processor of full-compliant H.264/AVC High Profile (HP) decoder and its FPGA implementation. It has the processing capability of HD1080 (1920 × 1088) at 60 frames per second (fps) that is asymptotic to Level 4.2 of the standard. To do this, several design considerations are investigated and the solutions for them are presented. The proposed design was realized with 41 K logic gates and 4,608 bits SRAM at the operating frequency of 266 MHz and was completely conformed by means of Allegro compliance bitstreams on an FPGA platform.
Kiwon Yoo, Jae Hun Lee, Kwanghoon Sohn
ICIP3
2008 Static text region detection in video sequences using color and orientation consistencies
abstract
Motion compensated error of the static text in the frame rate conversion (FRC) is most annoying artifact because people cannot read it. In this paper, we present a novel static text region detection algorithm for preventing the motion compensation error in FRC. We use some consistent properties of the static text that the color of the text is spatio-temporally consistent and the orientation of the text boundary is preserved in consecutive frames. We observe whether each pixelpsilas consistency is preserved for several frames, and then we decide it as a static text. Our algorithm can not only perfectly extract the static text region but also easily be implemented in hardware because of its ease.
Kwanghoon Sohn
ICPR2
2008 Asymmetric post-processing for stereo correspondence
abstract
This paper presents a novel approach that performs post-processing for stereo correspondence. We improve the performance of stereo correspondence by performing consistency check and adaptive filtering in an iterative filtering scheme. The consistency check is done with asymmetric information only so that very few additional computational loads are necessary. The proposed post-filtering method can be used in various methods for stereo correspondence without any modification. We demonstrate the validity of the proposed method by applying it to hierarchical belief propagation and semi-global matching.
Dongbo Min, Juhyun Oh, Kwanghoon Sohn
ICPR3
2008 Freeview rendering with trinocular camera
abstract
The paper presents a method for synthesizing novel view from the virtual camera with trinocular camera configuration. We propose a cost aggregation method with weighted least square for stereo matching, and address the occlusion problem by using cost functions computed with multiview images. We avoid the redundancy of estimating disparity maps for all the images by using simple geometry transfer method. The novel view is synthesized by view-dependent geometries on the 3D translation of virtual camera. Experimental results show that the proposed method yields the accurate disparity maps and the synthesized novel view is satisfactory enough to provide a viewer freeview videos.
Dongbo Min, Donghyun Kim 0010, SangUn Yuri, Kwanghoon Sohn
ISCAS4
2008 Surrounding adaptive color image enhancement based on CIECAM02
abstract
In this paper, we propose a CIECAM02-based color image enhancement method which is particularly robust to scenes with bright surrounding. The proposed method detects color edges using a distance metric based on the characteristics of a human visual system (HVS). The CIECAM02 is inherently strong considering both HVS and surrounding conditions. A major problem of HVS is that the dark region appears darker under a bright surrounding condition, leading to masking of details within the dark region. This phenomenon causes deterioration of edges which are among the most important and sensitive components for the HVS. To overcome this problem, we propose to weight the deteriorated edges at bright scenes. Adaptively, we estimate the surrounding image by the CIECAM02, and then use a vector gradient edge detector with a newly proposed distance metric to perform the weighting. The proposed method is seen to enhance edges without introducing unwanted color artifacts. We subjectively confirm the performance with clearly enhanced images.
Minsung Kang, Bongjoe Kim, Kar-Ann Toh, Kwanghoon Sohn
SMC4
2008 Stereoscopic video coding and disparity estimation for low bitrate applications based on MPEG-4 multiple auxiliary components
Kwanghoon Sohn
Signal Process. Image Commun.3
2008 Cost Aggregation and Occlusion Handling With WLS in Stereo Matching
abstract
This paper presents a novel method for cost aggregation and occlusion handling for stereo matching. In order to estimate optimal cost, given a per-pixel difference image as observed data, we define an energy function and solve the minimization problem by solving the iterative equation with the numerical method. We improve performance and increase the convergence rate by using several acceleration techniques such as the Gauss-Seidel method, the multiscale approach, and adaptive interpolation. The proposed method is computationally efficient since it does not use color segmentation or any global optimization techniques. For occlusion handling, which has not been performed effectively by any conventional cost aggregation approaches, we combine the occlusion problem with the proposed minimization scheme. Asymmetric information is used so that few additional computational loads are necessary. Experimental results show that performance is comparable to that of many state-of-the-art methods. The proposed method is in fact the most successful among all cost aggregation methods based on standard stereo test beds.
Dongbo Min, Kwanghoon Sohn
IEEE Trans. Image Process.2
2007 A Pose Robust Multi-View Face Recognition System using Plane of Pose Tolerance
abstract
In this paper, we propose a pose robust multi-view face recognition system using plane of pose tolerance (PPT). The proposed system defines the PPT which represents the tendency of pose variation well presented in multiple images. Compared with the traditional multi-view face recognition system, the proposed system is more robust and accurate especially to the error caused by pose estimation since it uses the property of multiple images. We obtain 91% face recognition rate for the proposed system and accomplished 15% improvement when compared with the traditional one.
Bongjoe Kim, Hyoungchul Shin, Kwanghoon Sohn
ICASSP (1)3
2006 3D Face Recognition with Geometrically Localized Surface Shape Indexes
abstract
This paper describes a pose invariant three-dimensional (3D) face recognition method using distinctive facial features. A face has its structural components like the eyes, nose and mouth. The positions and the shapes of the facial components are very important characteristics of a face. We extract invariant facial feature points on those components using the facial geometry from a normalized face data and calculate relative features using these feature points. We also calculate a shape index on each area of facial feature point to represent curvature characteristics of facial components. We perform recognition by using weighted distance matching, support vector machine (SVM) and independent component analysis (ICA)
Hyoungchul Shin, Kwanghoon Sohn
ICARCV2
2006 Edge-preserving joint motion-disparity estimation in stereo image sequences
Dongbo Min, Hansung Kim 0001, Kwanghoon Sohn
Signal Process. Image Commun.3
2005 3D reconstruction from stereo images for interactions between real and virtual objects
Hansung Kim 0001, Kwanghoon Sohn
Signal Process. Image Commun.2
2005 Joint Motion and Disparity Fields Estimation for Stereoscopic Video Sequences
King Ngi Ngan, JeongEun Lim, Kwanghoon Sohn
Signal Process. Image Commun.4
2004 3D head pose estimation using range images for face recognition
abstract
This paper describes a robust three-dimensional (3D) head pose estimation method based on face geometry and linear regression model. Given an unknown range image, we extract six invariant facial features based on face curvature characteristics. For estimating the head pose, we estimate the initial head pose using the SVD method, and perform a refinement procedure to compensate for remaining errors for the X and Y axis. To compensate for the Z axis, we orthogonally project feature points onto the Z plane and estimate the face center line based on the linear regression model. Experimental results show that less than a 0.5 degree error on average for each axis has been achieved.
Hwanjong Song, Kwanghoon Sohn
ICARCV2
2004 Accurate bit rate and quality control for 3D multiview sequences
abstract
This paper presents an accurate bit rate and quality control algorithm for 3D multiview sequences. We remodel the quadratic rate-distortion model for 3D multiview sequences based on their picture types. The proposed method consists of two levels for more accurate bits rate control. In a frame level, it sets quantization parameters based on HVS (Human Visual System) and reduces error bits by remodeling rate-quantization. In a MB level, rate control is activated to provide stricter buffer regulations and higher bit rate encoding based on the target bits and the quantization parameters calculated at the frame level. The proposed algorithm shows improvements in rate control as well as in PSNR compared to the conventional methods. It also provides more accurate rate control and higher image quality than TM5 and TMN8 at various bits rates.
JeongEun Lim, Kwanghoon Sohn
VCIP2
2004 Efficient disparity vector coding for multiview sequences
Seungchul Choi, Sukhee Cho, Kwanghoon Sohn
Signal Process. Image Commun.4
2004 A multiview sequence CODEC with view scalability
JeongEun Lim, King Ngi Ngan, Kwanghoon Sohn
Signal Process. Image Commun.4
2003 Hierarchical disparity estimation with energy-based regularization
abstract
In this paper, we propose a hierarchical disparity estimation algorithm with energy-based regularization. Initial disparity vectors are obtained from downsampled stereo images using a feature-based region-dividing disparity estimation technique. Dense disparities are estimated from these initial vectors with shape-adaptive windows in full resolution images. Finally, the vector fields are regularized with the minimization of the energy functional which considers both fidelity and smoothness of the fields. The first two steps provide highly reliable disparity vectors, so that local minimum problem can be avoided in regularization step. The proposed algorithm generates accurate disparity map which is smooth inside objects while preserving its discontinuities in boundaries. Experimental results are presented to illustrate the capabilities of the proposed disparity estimation technique.
Hansung Kim 0001, Kwanghoon Sohn
ICIP (1)2
2003 3D Reconstruction of Stereo Images for Interaction between Real and Virtual Worlds
abstract
Mixed reality is different from the virtual reality in that users can feel immersed in a space which is composed of not only virtual but also real objects. Thus, it is essential to realize seamless integration and interaction of the virtual and real worlds. We need depth information of the real scene to synthesize the real and virtual objects. We propose a two-stage algorithm to find smooth and precise disparity vector fields with sharp object boundaries in a stereo image pair for depth estimation. Hierarchical region-dividing disparity estimation increases the efficiency and the reliability of the estimation process, and a shape-adaptive window provides high reliability of the fields around the object boundary region. At the second stage, the vector fields are regularized with a energy model which produces smooth fields while preserving their discontinuities resulting from the object boundaries. The vector fields are used to reconstruct 3D surface of the real scene. Simulation results show that the proposed algorithm provides accurate and spatially correlated disparity vector fields in various kinds of images, and synthesized 3D models produce natural space where the virtual objects interact with the real world as if they are in the same world.
Hansung Kim 0001, Seung-Jun Yang, Kwanghoon Sohn
ISMAR3
2002 Grouped zerotree wavelet image coding for very low bit rate
abstract
The zerotree wavelet coding is a very efficient image and video compression method using wavelet transform. In this paper, we introduce a grouped zerotree wavelet image coding technique with high performance that is independent on a scale of the wavelet transform. A key method that we employ is to merge the adjacent zerotrees and to transmit the merged zerotrees by one symbol. The basic idea of our algorithm can be easily applied to zerotree wavelet coders for video. Simulation results show that our algorithm outperforms previously reported zerotree wavelet coding at low bit rates and in all scales of the wavelet transform. It also preserves the advantages of the zerotree wavelet coding technique for multimedia applications.
Woo-Young Jang, Byung-Hoan Chon, Seh-Woong Jeong, Kwanghoon Sohn
ICIP (3)4
2000 Selective coding of human faces using wavelets
abstract
Proposes an automated selective coding algorithm for human faces using neural networks and wavelets. In the proposed coding algorithm, we try to preserve information about human faces as much as possible without compromising the overall compression efficiency. In particular, we want to eliminate the artifacts near the boundary between the face and the background. We first extract the facial area using the eye location information and the skin color information. When we allocate bits at each level in the wavelet transform, we allocate more bits to the area that corresponds to the facial area. Experiments show that we can obtain a crisper facial area at the expense of the background.
Jaeyoung Seol, Kwanghoon Sohn
SMC2
2000 Error-resilient video coding technique based on wavelet transform
Kwanghoon Sohn, Woo-Young Jang
VCIP1
1998 A mean field annealing approach to robust corner detection
abstract
This paper is an extension of our previous paper to improve the capability of detecting corners. We proposed a method of boundary smoothing for curvature estimation using a constrained regularization technique in the previous paper. We propose another approach to boundary smoothing for curvature estimation in this paper to improve the capability of detecting corners. The method is based on a minimization strategy known as mean field annealing which is a deterministic approximation to simulated annealing. It removes the noise while preserving corners very well. Thus, we can detect corners easier and better in this approach than in the constrained regularization approach. Finally, some matching results based on the corners detected by corner sharpness in the mean field annealing approach are presented as a demonstration of the power of the proposed algorithm.
Kwanghoon Sohn, Winser E. Alexander
IEEE Trans. Syst. Man Cybern. Part B1
1996 Recognition of partially occluded target objects
abstract
This paper presents a new method of consistent object representation which can be used for partially occluded target object recognition. We proposed a boundary smoothing method for curvature estimation using a constrained regularization technique. Even though the method is effective in detecting corners due to the use of corner sharpness to increase the robustness of the proposed algorithm, it does not preserve corners well. We propose another approach to boundary smoothing for curvature estimation using a mean field annealing technique to improve the capability of detecting corners. It removes the noise while preserving corners very well. In addition, we show some matching results in an occlusion environment based on the corners detected by corner sharpness with the mean field annealing approach using a hybrid Hopfield (1985) neural network.
Kwanghoon Sohn
ICIP (3)1
1994 A Constrained Regularization Approach to Robust Corner Detection
abstract
This paper presents a method of optimal boundary smoothing for curvature estimation and a method of corner detection for consistent representation of objects for computer vision applications. The existing methods for curvature estimation have a common problem in determining a unique smoothing factor. We propose a constrained regularization (CR) approach to overcome that problem. The curvature function computed on the preprocessed boundary, which is obtained by the CR approach, gives consistent corner detection results. Ideal corners rarely exist for a real boundary. They are often rounded due to the smoothing effects of the preprocessing. In addition, a human recognizes both sharp corners and slightly rounded segments as corners. Hence, we establish a criterion, called "corner sharpness", which is qualitatively similar to a human's capability to detect corners.>
Kwanghoon Sohn, Winser E. Alexander, Wesley E. Snyder
IEEE Trans. Syst. Man Cybern. Syst.1
1992 Curvature estimation and unique corner point detection for boundary representation
abstract
Computing a curvature function on a digitized boundary is an ill-posed problem due to the discrete nature of the boundary. The authors use a constrained regularization technique to obtain the optimal smooth boundary before computing the curvature function. A corner sharpness is defined for robust corner point detection. Matching results in the presence of occlusion using a 2-D Hopfield neural network are also shown to produce excellent results using this boundary representation. The human cognition system recognizes both ideal corner points and slightly rounded segments as corner points. A criterion to mimic a human's capability of detecting corner points and to compensate for the smoothing effect of the preprocessing in detecting corner points in the curvature function space is established.>
Kwanghoon Sohn, Winser E. Alexander, Yonghoon Kim, Wesley E. Snyder
ICRA1