VLDB 2026 Research / reviewers in the wild / expert
Seungryong Kim
dblp:141/9955
· DBLP profile ↗
132ranked-venue papers
21as first author
80since 2021 · last 2026
0000-0003-2927-6273ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 14 first-author · 73 since 2021Graphics, computer vision, multimedia, augmented reality and games · 84 · 13 first-author · 45 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHIMERA: Controllable High-quality Image-Mask Extraction for Reliable Diffusion-based Anomaly SynthesisabstractWe present CHIMERA, a novel framework for generating realistic, generalizable, and prompt-driven industrial anomalies from natural language instructions. Our method addresses two key challenges in text-guided anomaly synthesis: (1) the scarcity of scalable, high-quality paired anomaly data and (2) the difficulty of efficiently adapting large diffusion models to domain-specific tasks without overfitting. To tackle these challenges, we first introduce a Vision-Language Model (VLM)-guided data curation pipeline that automatically generates semantically rich and spatially grounded captions from normal images, enabling effective dataset augmentation without manual annotations. Building upon this, we propose a parameter-efficient fine-tuning strategy that adapts a pre-trained Diffusion Transformer (Stable Diffusion 3) using lightweight LoRA adapters. By aligning structured prompts with the model's pre-trained language-vision prior and introducing auxiliary attention-based mask supervision, our method prevents overfitting, enhances spatial consistency, and ensures efficient training even with limited data. Extensive experiments show that CHIMERA is the first unified framework to achieve controllable, scalable, and generalizable industrial anomaly generation by integrating VLM-guided data curation with efficient diffusion-based training, significantly improving anomaly detection in low-data and unseen scenarios. JoungBin Lee, Hyunkoo Lee, Jini Yang, Chaehyun Kim, Jung Yi, Seok Hwangbo, Hyeoncheol Lee, Minho Chun, Eunjo Jeong, Seungryong Kim |
AAAI | 10 |
| 2026 | Video Camera Trajectory Editing with Generative Rendering from Estimated GeometryabstractWe introduce a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view video data for training. Traditional reconstruction methods struggle with extreme trajectory changes, and existing generative models for dynamic novel view synthesis cannot handle in-the-wild videos. Our approach consists of two steps: estimating temporally consistent geometry, and generative rendering guided by this geometry. By integrating geometric priors, the generative model focuses on synthesizing realistic details where the estimated geometry is uncertain. We eliminate the need for extensive 4D training data through a factorized fine-tuning framework that separately trains spatial and temporal components using multi-view image and video data. Our method outperforms baselines in producing plausible videos from novel camera trajectories, especially in extreme extrapolation scenarios on real-world footage. Junyoung Seo, Jisang Han, Jaewoo Jung, Siyoon Jin, Joungbin Lee, Takuya Narihira, Kazumi Fukuda, Takashi Shibuya 0001, Donghoon Ahn, Shoukang Hu, Seungryong Kim, Yuki Mitsufuji |
AAAI | 11 |
| 2025 | MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head GenerationabstractConventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models have attempted to address these limitations and improve fidelity. However, they still face challenges, such as intensive sampling times and difficulties in maintaining temporal consistency due to the high stochasticity of diffusion models. To overcome these challenges, we propose a novel motion-disentangled diffusion model for high-quality talking head generation, called MoDiTalker. We introduce two modules: the Audio-To-Motion (AToM) module, designed to generate synchronized lip movements from audio, and the Motion-To-Video (MToV) module, designed to produce high-quality talking head videos based on the generated motions. AToM excels in capturing subtle lip movements by leveraging an audio attention mechanism. Additionally, MToV enhances temporal consistency by utilizing an efficient tri-plane representation. Our experiments on standard benchmarks demonstrate that our model outperforms existing GAN-based and diffusion-based models. We also provide comprehensive ablation studies and user study results. Siyoon Jin, Jisu Nam, Seungryong Kim |
AAAI | 7 |
| 2025 | Multi-Granularity Video Object SegmentationabstractCurrent benchmarks for video segmentation are limited to annotating only salient objects (i.e., foreground instances). Despite their impressive architectural designs, previous works trained on these benchmarks have struggled to adapt to realworld scenarios. Thus, developing a new video segmentation dataset aimed at tracking multi-granularity segmentation target in the video scene is necessary. In this work, we aim to generate multi-granularity video segmentation dataset that is annotated for both salient and non-salient masks. To achieve this, we propose a large-scale, densely annotated multi-granularity video object segmentation (MUG-VOS) dataset that includes various types and granularities of mask annotations. We automatically collected a training set that assists in tracking both salient and non-salient objects, and we also curated a human-annotated test set for reliable evaluation. In addition, we present memory-based mask propagation model (MMPM), trained and evaluated on MUG-VOS dataset, which leads to the best performance among the existing video object segmentation methods and Segment SAM-based video segmentation methods. Sangbeom Lim, Seongchan Kim, Seungjun An, Seokju Cho, Hongsuck Seo, Seungryong Kim |
AAAI | 6 |
| 2025 | Cross-View Completion Models are Zero-shot Correspondence EstimatorsabstractIn this work, we explore new perspectives on crossview completion learning by drawing an analogy to self-supervised correspondence learning. Through our analysis, we demonstrate that the cross-attention map within crossview completion models captures correspondence more effectively than other correlations derived from encoder or decoder features. We verify the effectiveness of the cross-attention map by evaluating on both zero-shot matching and learning-based geometric matching and multi-frame depth estimation. Honggyu An, Jin Hyeon Kim, Seonghoon Park 0002, Jaewoo Jung, Jisang Han, Sunghwan Hong, Seungryong Kim |
CVPR | 7 |
| 2025 | Seurat: From Moving Points to DepthabstractAccurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However, humans often perceive relative depth intuitively by observing variations in the size and spacing of objects as they move. Inspired by this, we propose a novel method that infers relative depth by examining the spatial relationships and temporal evolution of a set of tracked 2D trajectories. Specifically, we use off-the-shelf point tracking models to capture 2D trajectories. Then, our approach employs spatial and temporal transformers to process these trajectories and directly infer depth changes over time. Evaluated on the TAPVid-3D benchmark, our method demonstrates robust zero-shot performance, generalizing effectively from synthetic to real-world datasets. Results indicate that our approach achieves temporally smooth, high-accuracy depth predictions across diverse domains. Seokju Cho, Seungryong Kim, Joon-Young Lee |
CVPR | 3 |
| 2025 | ControlFace: Harnessing Facial Parametric Control for Face RiggingabstractManipulation of facial images to meet specific controls such as pose, expression, and lighting, also known as face rigging, is a complex task in computer vision. Existing methods are limited by their reliance on image datasets, which necessitates individual-specific fine-tuning and limits their ability to retain fine-grained identity and semantic details, reducing practical usability. To overcome these limitations, we introduce ControlFace, a novel face rigging method conditioned on 3DMM renderings that enables flexible, high-fidelity control. We employ a dual-branch U-Nets: one, referred to as FaceNet, captures identity and fine details, while the other focuses on generation. To enhance control precision, the control mixer module encodes the correlated features between the target-aligned control and reference-aligned control, and a novel guidance method, reference control guidance, steers the generation process for better control adherence. By training on a facial video dataset, we fully utilize FaceNet’s rich representations while ensuring control adherence. Extensive experiments demonstrate ControlFace’s superior performance in identity preservation and control precision, highlighting its practicality. Please see the project website: https://cvlab-kaist.github.io/ControlFace/. Woo-seok Jang, Youngjun Hong, Geonho Cha, Seungryong Kim |
CVPR | 4 |
| 2025 | Exploring Temporally-Aware Features for Point TrackingabstractPoint tracking in videos is a fundamental task with applications in robotics, video editing, and more. While many vision tasks benefit from pre-trained feature backbones to improve generalizability, point tracking has primarily relied on simpler backbones trained from scratch on synthetic data, which may limit robustness in real-world scenarios. Additionally, point tracking requires temporal awareness to ensure coherence across frames, but using temporally-aware features is still underexplored. Most current methods often employ a two-stage process: an initial coarse prediction followed by a refinement stage to inject temporal information and correct errors from the coarse stage. These approach, however, is computationally expensive and potentially redundant if the feature backbone itself captures sufficient temporal information.In this work, we introduce Chrono, a feature backbone specifically designed for point tracking with built-in temporal awareness. Leveraging pre-trained representations from self-supervised learner DINOv2 and enhanced with a temporal adapter, Chrono effectively captures long-term temporal context, enabling precise prediction even without the refinement stage. Experimental results demonstrate that Chrono achieves state-of-the-art performance in a refiner-free setting on the TAP-Vid-DAVIS and TAP-Vid-Kinetics datasets, among common feature backbones used in point tracking as well as DINOv2, with exceptional efficiency. Project page: https://cvlab-kaist.github.io/Chrono/ Inès Hyeonsu Kim, Seokju Cho, Jung Yi, Joon-Young Lee, Seungryong Kim |
CVPR | 6 |
| 2025 | Identity-preserving Distillation Sampling by Fixed-Point IteratorabstractScore distillation sampling (SDS) demonstrates a powerful capability for text-conditioned 2D image and 3D object generation by distilling the knowledge from learned score functions. However, SDS often suffers from blurriness caused by noisy gradients. When SDS meets the image editing, such degradations can be reduced by adjusting bias shifts using reference pairs, but the de-biasing techniques are still corrupted by erroneous gradients. To this end, we introduce Identity-preserving Distillation Sampling (IDS), which compensates for the gradient leading to undesired changes in the results. Based on the analysis that these errors come from the text-conditioned scores, a new regularization technique, called fixed-point iterative regularization (FPR), is proposed to modify the score itself, driving the preservation of the identity even including poses and structures. Thanks to a self-correction by FPR, the proposed method provides clear and unambiguous representations corresponding to the given prompts in image-to-image editing and editable neural radiance field (NeRF). The structural consistency between the source and the edited data is obviously maintained compared to other state-of-the-art methods. Our code is https://github.com/shhh0620/IDS Seonhwa Kim, Soobin Park, Donghoon Ahn, Seungryong Kim, Kyong Hwan Jin, Eun Ju Cha |
CVPR | 6 |
| 2025 | Mixture of Submodules for Domain Adaptive Person SearchabstractExisting technique on domain adaptive person search commonly utilizes the unified framework for jointly localizing and identifying the person across domains. This framework, however, inevitably results in the gradient conflict problem, particularly in cross-domain scenarios with contradictory objectives, as the unified framework employs shared parameters to simultaneously address person detection and re-identification tasks across the domains. To over-come this, we present a novel mixture of submodules framework, dubbed MoS, that dynamically modulates the combination of submodules depending on the specific task to perform person detection and re-identification, separately. We further design the mixtures of submodules that vary depending on the domain, enabling domain-specific knowledge transfer. Especially, we decompose the main model into several submodules and employ diverse mixtures of submodules that vary depending on the tasks and domains through the conditional routing policy. In addition, we also present counterpart domain sample generation that synthesizes the augmented sample and uses them to learn domain invariant representation for person re-identification through the contrastive domain alignment. We conduct experiments to demonstrate the effectiveness of our MoS over the existing domain adaptive person search method and provide ablation studies. Seungryong Kim, Kwanghoon Sohn |
CVPR | 2 |
| 2025 | Visual Persona: Foundation Model for Full-Body Human CustomizationabstractWe introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach captures detailed full-body appearance, aligning with text descriptions for body structure and scene variations. Training this model requires large-scale paired human data, consisting of multiple images per individual with consistent full-body identities, which is notoriously difficult to obtain. To address this, we propose a data curation pipeline leveraging vision-language models to evaluate full-body appearance consistency, resulting in Visual Persona-500K—a dataset of 580k paired human images across 100k unique identities. For precise appearance transfer, we introduce a transformer encoder-decoder architecture adapted to a pre-trained text-to-image diffusion model, which augments the input image into distinct body regions, encodes these regions as local appearance features, and projects them into dense identity embeddings independently to condition the diffusion model for synthesizing customized images. Visual Persona consistently surpasses existing approaches, generating high-quality, customized images from in-the-wild inputs. Extensive ablation studies validate design choices, and we demonstrate the versatility of Visual Persona across various downstream tasks. Jisu Nam, Soowon Son, Jing Shi 0005, Difan Liu, Feng Liu 0015, Seungryong Kim, Yang Zhou 0009 |
CVPR | 7 |
| 2025 | AM-Adapter: Appearance Matching Adapter for Exemplar-Based Semantic Image Synthesis In-the-Wild
Siyoon Jin, Jisu Nam, Dahyun Chung, Yeong-Seok Kim, Joonhyung Park, Heonjeong Chu, Seungryong Kim |
ICCV | 8 |
| 2025 | S4M: Boosting Semi-Supervised Instance Segmentation with SAM
Heeji Yoon, Heeseong Shin, Eunbeen Hong, Hyunwook Choi, Hansang Cho, Daun Jeong, Seungryong Kim |
ICCV | 7 |
| 2025 | PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting for Novel View SynthesisabstractWe consider the problem of novel view synthesis from unposed images in a single feed-forward. Our framework capitalizes on fast speed, scalability, and high-quality 3D reconstruction and view synthesis capabilities of 3DGS, where we further extend it to offer a practical solution that relaxes common assumptions such as dense image views, accurate camera poses, and substantial image overlaps. We achieve this through identifying and addressing unique challenges arising from the use of pixel-aligned 3DGS: misaligned 3D Gaussians across different views induce noisy or sparse gradients that destabilize training and hinder convergence, especially when above assumptions are not met. To mitigate this, we employ pre-trained monocular depth estimation and visual correspondence models to achieve coarse alignments of 3D Gaussians. We then introduce lightweight, learnable modules to refine depth and pose estimates from the coarse alignments, improving the quality of 3D reconstruction and novel view synthesis. Furthermore, the refined estimates are leveraged to estimate geometry confidence scores, which assess the reliability of 3D Gaussian centers and condition the prediction of Gaussian parameters accordingly. Extensive evaluations on large-scale real-world datasets demonstrate that PF3plat sets a new state-of-the-art across all benchmarks, supported by comprehensive ablation studies validating our design choices. We will make the code and weights publicly available. Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jisang Han, Jiaolong Yang, Chong Luo 0001, Seungryong Kim |
ICML | 7 |
| 2025 | Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow ModelsabstractRecent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose ``HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head’s attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection. We validate our method on modern large-scale DiT-based text-to-image models including Stable Diffusion 3 and FLUX.1, demonstrating superior performance in both general quality enhancement and style-specific guidance. Our work provides the first head-level analysis of attention perturbation in diffusion models, uncovering interpretable specialization within attention layers and enabling practical design of effective perturbation strategies. Donghoon Ahn, Woo-seok Jang, Jaewon Min, Sangwu Lee, Sayak Paul, Seungryong Kim |
NeurIPS | 9 |
| 2025 | Enhancing 3D Reconstruction for Dynamic ScenesabstractIn this work, we address the task of 3D reconstruction in dynamic scenes, where object motions frequently degrade the quality of previous 3D pointmap regression methods, such as DUSt3R, that are originally designed for static 3D scene reconstruction. Although these methods provide an elegant and powerful solution in static settings, they struggle in the presence of dynamic motions that disrupt alignment based solely on camera poses. To overcome this, we propose D$^2$USt3R that directly regresses Static-Dynamic Aligned Pointmaps (SDAP) that simultaneiously capture both static and dynamic 3D scene geometry. By explicitly incorporating both spatial and temporal aspects, our approach successfully encapsulates 3D dense correspondence to the proposed pointmaps, enhancing downstream tasks. Extensive experimental evaluations demonstrate that our proposed approach consistently achieves superior 3D reconstruction performance across various datasets featuring complex motions. Jisang Han, Honggyu An, Jaewoo Jung, Takuya Narihira, Junyoung Seo, Kazumi Fukuda, Chaehyun Kim, Sunghwan Hong, Yuki Mitsufuji, Seungryong Kim |
NeurIPS | 10 |
| 2025 | First Attentions Last: Better Exploiting First Attentions for Efficient Parallel TrainingabstractAs training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However, existing transformer designs suffer from significant communication overhead, especially in Tensor Parallelism (TP), where each block’s MHA–MLP connection requires an all-reduce communication. Through our investigation, we show that the MHA-MLP connections can be bypassed for efficiency, while the attention output of the first layer can serve as an alternative signal for the bypassed connection. Motivated by the observations, we propose FAL (First Attentions Last), an efficient transformer architecture that redirects the first MHA output to the MLP inputs of the following layers, eliminating the per-block MHA-MLP connections. This removes the all-reduce communication and enables parallel execution of MHA and MLP on a single GPU. We also introduce FAL+, which adds the normalized first attention output to the MHA outputs of the following layers to augment the MLP input for the model quality. Our evaluation shows that FAL reduces multi-GPU training time by up to 44%, improves single-GPU throughput by up to 1.18×, and achieves better perplexity compared to the baseline GPT. FAL+ achieves even lower perplexity without increasing the training time than the baseline. Codes are available at: https://casl-ku.github.io/FAL/ Gyudong Kim, Hyukju Na, Jin Kyu Kim, Hyunsung Jang, Jaegi Hwang, Namkoo Ha, Seungryong Kim, Younggeun Kim 0001 |
NeurIPS | 8 |
| 2025 | Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersabstractText-to-image diffusion models excel at translating language prompts into photorealistic images by implicitly grounding textual concepts through their cross-modal attention mechanisms. Recent multi-modal diffusion transformers extend this by introducing joint self-attention over concatenated image and text tokens, enabling richer and more scalable cross-modal alignment. However, a detailed understanding of how and where these attention maps contribute to image generation remains limited. In this paper, we introduce Seg4Diff (Segmentation for Diffusion), a systematic framework for analyzing the attention structures of MM-DiT, with a focus on how specific layers propagate semantic information from text to image. Through comprehensive analysis, we identify a semantic grounding expert layer, a specific MM-DiT block that consistently aligns text tokens with spatially coherent image regions, naturally producing high-quality semantic segmentation masks. We further demonstrate that applying a lightweight fine-tuning scheme with mask-annotated image data enhances the semantic grouping capabilities of these layers and thereby improves both segmentation performance and generated image fidelity. Our findings demonstrate that semantic grouping is an emergent property of diffusion transformers and can be selectively amplified to advance both segmentation and generation performance, paving the way for unified models that bridge visual perception and generation. Chaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon, Anurag Arnab, Hongsuck Seo, Sunghwan Hong, Seungryong Kim |
NeurIPS | 8 |
| 2025 | Active Test-time Vision-Language NavigationabstractVision-Language Navigation (VLN) policies trained on offline datasets often exhibit degraded task performance when deployed in unfamiliar navigation environments at test time, where agents are typically evaluated without access to external interaction or feedback. Entropy minimization has emerged as a practical solution for reducing prediction uncertainty at test time; however, it can suffer from accumulated errors, as agents may become overconfident in incorrect actions without sufficient contextual grounding. To tackle these challenges, we introduce ATENA (Active TEst-time Navigation Agent), a test-time active learning framework that enables a practical human-robot interaction via episodic feedback on uncertain navigation outcomes. In particular, ATENA learns to increase certainty in successful episodes and decrease it in failed ones, improving uncertainty calibration. Here, we propose mixture entropy optimization, where entropy is obtained from a combination of the action and pseudo-expert distributions—a hypothetical action distribution assuming the agent's selected action to be optimal—controlling both prediction confidence and action preference. In addition, we propose a self-active learning strategy that enables an agent to evaluate its navigation outcomes based on confident predictions. As a result, the agent stays actively engaged throughout all iterations, leading to well-grounded and adaptive decision-making. Extensive evaluations on challenging VLN benchmarks—REVERIE, R2R, and R2R-CE—demonstrate that ATENA successfully overcomes distributional shifts at test time, outperforming the compared baseline methods across various settings. Heeju Ko, Sung June Kim, Gyeongrok Oh, Jeongyoon Yoon, Honglak Lee, Sujin Jang, Seungryong Kim, Sangpil Kim |
NeurIPS | 7 |
| 2025 | Emergent Temporal Correspondences from Video Diffusion TransformersabstractRecent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish and represent temporal correspondences across frames? We introduce DiffTrack, the first quantitative analysis framework designed to answer this question. DiffTrack constructs a dataset of prompt-generated video with pseudo ground-truth tracking annotations and proposes novel evaluation metrics to systematically analyze how each component within the full 3D attention mechanism of DiTs (e.g., representations, layers, and timesteps) contributes to establishing temporal correspondences. Our analysis reveals that query-key similarities in specific (but not all) layers play a critical role in temporal matching, and that this matching becomes increasingly prominent throughout denoising. We demonstrate practical applications of DiffTrack in zero-shot point tracking, where it achieves state-of-the-art performance compared to existing vision foundation and self-supervised video models. Further, we extend our findings to motion-enhanced video generation with a novel guidance method that improves temporal consistency of generated videos without additional training. We believe our work offers crucial insights into the inner workings of video DiTs and establishes a foundation for further research and applications leveraging their temporal understanding. Jisu Nam, Soowon Son, Dahyun Chung, Siyoon Jin, Junhwa Hur, Seungryong Kim |
NeurIPS | 7 |
| 2025 | Domain Generalization using Large Pretrained Models with Mixture-of-AdaptersabstractLearning robust vision models that perform well in out-of-distribution (OOD) situations is an important task for model deployment in real-world settings. Despite extensive research in this field, many proposed methods have only shown minor performance improvements compared to the simplest empirical risk minimization (ERM) approach, which was evaluated on a benchmark with a limited hy-perparameter search space. Our focus in this study is on leveraging the knowledge of large pretrained models to improve handling of OOD scenarios and tackle domain generalization problems. However, prior research has revealed that naively fine-tuning a large pretrained model can impair OOD robustness. Thus, we employ parameter-efficient fine-tuning (PEFT) techniques to effectively preserve OOD robustness while working with large models. Our extensive experiments and analysis confirm that the most effective approaches involve ensembling diverse models and increasing the scale of pretraining. As a result, we achieve state-of-the-art performance in domain generalization tasks. Our code and project page are available at: https://cvlab-kaist.github.io/MoA Gyuseong Lee, Woo-seok Jang, Jinhyeon Kim, Jaewoo Jung, Seungryong Kim |
WACV | 5 |
| 2025 | Semantically complex audio to video generation with audio source separation
Jaehwan Jeong, Sumin In, Seungryong Kim, Saerom Kim, Wooyeol Baek, Sang Ho Yoon, Eugenio Culurciello, Sangpil Kim |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | SplitNet: Learnable Clean-Noisy Label Splitting for Learning with Noisy LabelsabstractAbstract Annotating the dataset with high-quality labels is crucial for deep networks’ performance, but in real-world scenarios, the labels are often contaminated by noise. To address this, some methods were recently proposed to automatically split clean and noisy labels among training data, and learn a semi-supervised learner in a Learning with Noisy Labels (LNL) framework. However, they leverage a handcrafted module for clean-noisy label splitting, which induces a confirmation bias in the semi-supervised learning phase and limits the performance. In this paper, for the first time, we present a learnable module for clean-noisy label splitting, dubbed SplitNet, and a novel LNL framework which complementarily trains the SplitNet and main network for the LNL task. We also propose to use a dynamic threshold based on split confidence by SplitNet to optimize the semi-supervised learner better. To enhance SplitNet training, we further present a risk hedging method. Our proposed method performs at a state-of-the-art level, especially in high noise ratio settings on various LNL benchmarks. Kwangrok Ryoo, Hansang Cho, Seungryong Kim |
Int. J. Comput. Vis. | 4 |
| 2025 | Correction: SplitNet: Learnable Clean-Noisy Label Splitting for Learning with Noisy Labels
Kwangrok Ryoo, Hansang Cho, Seungryong Kim |
Int. J. Comput. Vis. | 4 |
| 2025 | DiffFace: Diffusion-based face swapping with facial guidance
Seokju Cho, Junyoung Seo, Jisu Nam, Kyuchul Lee, Seungryong Kim, Kwanghee Lee |
Pattern Recognit. | 7 |
| 2024 | Context Enhanced Transformer for Single Image Object Detection in Video DataabstractWith the increasing importance of video data in real-world applications, there is a rising need for efficient object detection methods that utilize temporal information. While existing video object detection (VOD) techniques employ various strategies to address this challenge, they typically depend on locally adjacent frames or randomly sampled images within a clip. Although recent Transformer-based VOD methods have shown promising results, their reliance on multiple inputs and additional network complexity to incorporate temporal information limits their practical applicability. In this paper, we propose a novel approach to single image object detection, called Context Enhanced TRansformer (CETR), by incorporating temporal context into DETR using a newly designed memory module. To efficiently store temporal information, we construct a class-wise memory that collects contextual information across data. Additionally, we present a classification-based sampling technique to selectively utilize the relevant memory for the current image. In the testing, We introduce a test-time memory adaptation method that updates individual memory functions by considering the test distribution. Experiments with CityCam and ImageNet VID datasets exhibit the efficiency of the framework on various video systems. The project page and code will be made available at: https://ku-cvlab.github.io/CETR. Seungjun An, Seonghoon Park 0002, Gyeongnyeon Kim, Jeongyeol Baek, Byeongwon Lee, Seungryong Kim |
AAAI | 6 |
| 2024 | Few-Shot Neural Radiance Fields under Unconstrained IlluminationabstractIn this paper, we introduce a new challenge for synthesizing novel view images in practical environments with limited input multi-view images and varying lighting conditions. Neural radiance fields (NeRF), one of the pioneering works for this task, demand an extensive set of multi-view images taken under constrained illumination, which is often unattainable in real-world settings. While some previous works have managed to synthesize novel views given images with different illumination, their performance still relies on a substantial number of input multi-view images. To address this problem, we suggest ExtremeNeRF, which utilizes multi-view albedo consistency, supported by geometric alignment. Specifically, we extract intrinsic image components that should be illumination-invariant across different views, enabling direct appearance comparison between the input and novel view under unconstrained illumination. We offer thorough experimental results for task evaluation, employing the newly created NeRF Extreme benchmark—the first in-the-wild benchmark for novel view synthesis under multiple viewing directions and varying illuminations. SeokYeong Lee, Junyong Choi, Seungryong Kim, Ig-Jae Kim, Junghyun Cho |
AAAI | 3 |
| 2024 | Match Me If You Can: Semi-supervised Semantic Correspondence Learning with Unpaired Images
Byeongho Heo, Sangdoo Yun, Seungryong Kim, Dongyoon Han |
ACCV (6) | 4 |
| 2024 | FlowTrack: Revisiting Optical Flow for Long-Range Dense TrackingabstractIn the domain of video tracking, existing methods often grapple with a trade-off between spatial density and temporal range. Current approaches in dense optical flow estimators excel in providing spatially dense tracking but are limited to short temporal spans. Conversely, recent advancements in long-range trackers offer extended temporal coverage but at the cost of spatial sparsity. This paper introduces FlowTrack, a novel framework designed to bridge this gap. FlowTrack combines the strengths of both paradigms by 1) chaining confident flow predictions to maximize efficiency and 2) automatically switching to an error compensation module in instances of flow prediction inaccuracies. This dual strategy not only offers efficient dense tracking over extended temporal spans but also ensures robustness against error accumulations and occlusions, common pitfalls of naive flow chaining. Furthermore, we demonstrate that chained flow itself can serve as an effective guide for an error compensation module, even for occluded points. Our framework achieves state-of-the-art accuracy for longrange tracking on the DAVIS dataset, and renders 50% speed-up when performing dense tracking. Seokju Cho, Seungryong Kim, Joon-Young Lee |
CVPR | 3 |
| 2024 | CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions. In this work, we introduce a novel cost-based approach to adapt vision-language foundation models, notably CLIP, for the intricate task of semantic segmentation. Through aggregating the cosine similarity score, i. e., the cost volume between image and text embeddings, our method potently adapts CLIP for segmenting seen and unseen classes by fine-tuning its encoders, addressing the challenges faced by existing methods in handling unseen classes. Building upon this, we explore methods to effectively aggregate the cost volume considering its multi-modal nature of being established between image and text embeddings. Furthermore, we examine various methods for efficiently fine-tuning CLIP. Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Hongsuck Seo, Seungryong Kim |
CVPR | 6 |
| 2024 | Unifying Correspondence, Pose and NeRF for Generalized Pose-Free Novel View SynthesisabstractThis work delves into the task of pose-free novel view synthesis from stereo pairs, a challenging and pioneering task in 3D vision. Our innovative framework, unlike any before, seamlessly integrates 2D correspondence matching, camera pose estimation, and NeRF rendering, fostering a synergistic enhancement of these tasks. We achieve this through designing an architecture that utilizes a shared representation, which serves as a foundation for enhanced 3D geometry understanding. Capitalizing on the inherent in-terplay between the tasks, our unified framework is trained end-to-end with the proposed training strategy to improve overall model accuracy. Through extensive evaluations across diverse indoor and outdoor scenes from two real-world datasets, we demonstrate that our approach achieves substantial improvement over previous methodologies, es-pecially in scenarios characterized by extreme viewpoint changes and the absence of accurate camera poses. The project page and code will be made available at: https://ku-cvlab.github.io/CoPoNeRF/. Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jiaolong Yang, Seungryong Kim, Chong Luo 0001 |
CVPR | 5 |
| 2024 | DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationabstractThe objective of text-to-image (T2I) personalization is to customize a diffusion model to a user-provided reference concept, generating diverse images of the concept aligned with the target prompts. Conventional methods representing the reference concepts using unique text embeddings often fail to accurately mimic the appearance of the reference. To address this, one solution may be explicitly conditioning the reference images into the target denoising process, known as key-value replacement. However, prior works are constrained to local editing since they disrupt the structure path of the pretrained T2I model. To overcome this, we propose a novel plug-in method, called DreamMatcher, which reformulates T2I personalization as semantic matching. Specifically, DreamMatcher replaces the target values with reference values aligned by semantic matching, while leaving the structure path unchanged to preserve the versatile capability of pretrained T2I models for generating diverse structures. We also introduce a semantic-consistent masking strategy to isolate the personalized concept from irrelevant regions introduced by the target prompts. Compatible with existing T2I models, DreamMatcher shows significant improvements in complex scenarios. Intensive analyses demonstrate the effectiveness of our approach. Jisu Nam, Heesu Kim, Siyoon Jin, Seungryong Kim, Seunggyu Chang |
CVPR | 5 |
| 2024 | Self-rectifying Diffusion Sampling with Perturbed-Attention Guidance
Donghoon Ahn, Hyoungwon Cho, Jaewon Min, Woo-seok Jang, Seonhwa Kim, Hyun Hee Park, Kyong Hwan Jin, Seungryong Kim |
ECCV (73) | 9 |
| 2024 | Local All-Pair Correspondence for Point Tracking
Seokju Cho, Jisu Nam, Honggyu An, Seungryong Kim, Joon-Young Lee |
ECCV (10) | 5 |
| 2024 | Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation
Haneol Lee, Kwanghee Lee, Seungryong Kim, Jaejun Yoo 0001 |
ECCV (52) | 6 |
| 2024 | Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual CorrespondenceabstractThis paper introduces a Transformer-based integrative feature and cost aggregation network designed for dense matching tasks. In the context of dense matching, many works benefit from one of two forms of aggregation: feature aggregation, which pertains to the alignment of similar features, or cost aggregation, a procedure aimed at instilling coherence in the flow estimates across neighboring pixels. In this work, we first show that feature aggregation and cost aggregation exhibit distinct characteristics and reveal the potential for substantial benefits stemming from the judicious use of both aggregation processes. We then introduce a simple yet effective architecture that harnesses self- and cross-attention mechanisms to show that our approach unifies feature aggregation and cost aggregation and effectively harnesses the strengths of both techniques. Within the proposed attention layers, the features and cost volume both complement each other, and the attention layers are interleaved through a coarse-to-fine design to further promote accurate correspondence estimation. Finally at inference, our network produces multi-scale predictions, computes their confidence scores, and selects the most confident flow for final prediction. Our framework is evaluated on standard benchmarks for semantic matching, and also applied to geometric matching, where we show that our approach achieves significant improvements compared to existing methods. Sunghwan Hong, Seokju Cho, Seungryong Kim, Stephen Lin 0001 |
ICLR | 3 |
| 2024 | Diffusion Model for Dense MatchingabstractThe objective for establishing dense correspondence between paired images con- sists of two terms: a data term and a prior term. While conventional techniques focused on defining hand-designed prior terms, which are difficult to formulate, re- cent approaches have focused on learning the data term with deep neural networks without explicitly modeling the prior, assuming that the model itself has the capacity to learn an optimal prior from a large-scale dataset. The performance improvement was obvious, however, they often fail to address inherent ambiguities of matching, such as textureless regions, repetitive patterns, large displacements, or noises. To address this, we propose DiffMatch, a novel conditional diffusion-based framework designed to explicitly model both the data and prior terms for dense matching. This is accomplished by leveraging a conditional denoising diffusion model that explic- itly takes matching cost and injects the prior within generative process. However, limited input resolution of the diffusion model is a major hindrance. We address this with a cascaded pipeline, starting with a low-resolution model, followed by a super-resolution model that successively upsamples and incorporates finer details to the matching field. Our experimental results demonstrate significant performance improvements of our method over existing approaches, and the ablation studies validate our design choices along with the effectiveness of each component. Code and pretrained weights are available at https://ku-cvlab.github.io/DiffMatch. Jisu Nam, Gyuseong Lee, Hyoungwon Cho, Seungryong Kim |
ICLR | 7 |
| 2024 | Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D GenerationabstractText-to-3D generation has shown rapid progress in recent days with the advent of score distillation sampling (SDS), a methodology of using pretrained text-to-2D diffusion models to optimize a neural radiance field (NeRF) in a zero-shot setting. However, the lack of 3D awareness in the 2D diffusion model often destabilizes previous methods from generating a plausible 3D scene. To address this issue, we propose 3DFuse, a novel framework that incorporates 3D awareness into the pretrained 2D diffusion model, enhancing the robustness and 3D consistency of score distillation-based methods. Specifically, we introduce a consistency injection module which constructs a 3D point cloud from the text prompt and utilizes its projected depth map at given view as a condition for the diffusion model. The 2D diffusion model, through its generative capability, robustly infers dense structure from the sparse point cloud depth map and generates a geometrically consistent and coherent 3D scene. We also introduce a new technique called semantic coding that reduces semantic ambiguity of the text prompt for improved results. Our method can be easily adapted to various text-to-3D baselines, and we experimentally demonstrate how our method notably enhances the 3D consistency of generated scenes in comparison to previous baselines, achieving state-of-the-art performance in geometric robustness and fidelity. Junyoung Seo, Woo-seok Jang, Minseop Kwak, Inès Hyeonsu Kim, Jaehoon Ko, Jin-Hwa Kim, Jiyoung Lee 0005, Seungryong Kim |
ICLR | 9 |
| 2024 | Retrieval-Augmented Score Distillation for Text-to-3D GenerationabstractText-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have been released, fine-tuning the diffusion model on the multi-view datasets becomes a mainstream to solve the 3D inconsistency problem. However, it has confronted with fundamental difficulties regarding the limited quality and diversity of 3D data, compared with 2D data. To sidestep these trade-offs, we explore a retrieval-augmented approach tailored for score distillation, dubbed ReDream. We postulate that both expressiveness of 2D diffusion models and geometric consistency of 3D assets can be fully leveraged by employing the semantically relevant assets directly within the optimization process. To this end, we introduce novel framework for retrieval-based quality enhancement in text-to-3D generation. We leverage the retrieved asset to incorporate its geometric prior in the variational objective and adapt the diffusion model's 2D prior toward view consistency, achieving drastic improvements in both geometry and fidelity of generated scenes. We conduct extensive experiments to demonstrate that ReDream exhibits superior quality with increased geometric consistency. Project page is available at https://ku-cvlab.github.io/ReDream/. Junyoung Seo, Susung Hong, Woo-seok Jang, Inès Hyeonsu Kim, Minseop Kwak, Doyup Lee, Seungryong Kim |
ICML | 7 |
| 2024 | MaskingDepth: Masked Consistency Regularization for Semi-Supervised Monocular Depth EstimationabstractWe propose MaskingDepth, a semi-supervised learning framework for monocular depth estimation. MaskingDepth is designed to enforce consistency between the depths obtained from strongly-augmented images and the pseudo-depths derived from weakly-augmented images, which enables mitigating the reliance on large ground-truth depth quantities. In this framework, we leverage uncertainty estimation to only retain high-confident depth predictions from the weakly-augmented branch as pseudo-depths. We also present a novel data augmentation, dubbed K-way disjoint masking, that takes advantage of a naïve token masking strategy as an augmentation, while avoiding its scale ambiguity problem between depths from weakly-and strongly-augmented branches and risk of missing small-scale objects. Experiments on KITTI and NYU-Depth-v2 datasets demonstrate the effectiveness of each component, its robustness to the use of fewer depth-annotated images, and superior performance compared to other state-of-the-art semi-supervised learning methods for monocular depth estimation. Jong-Beom Baek, Gyeongnyeon Kim, Seonghoon Park 0002, Honggyu An, Matteo Poggi, Seungryong Kim |
IROS | 6 |
| 2024 | GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian Splatting
Kyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong, Jaehoon Ko, Sangjun Ahn, Seungryong Kim |
ACM Multimedia | 7 |
| 2024 | GenWarp: Single Image to Novel Views with Semantic-Preserving Generative WarpingabstractGenerating novel views from a single image remains a challenging task due to the complexity of 3D scenes and the limited diversity in the existing multi-view datasets to train a model on. Recent research combining large-scale text-to-image (T2I) models with monocular depth estimation (MDE) has shown promise in handling in-the-wild images. In these methods, an input view is geometrically warped to novel views with estimated depth maps, then the warped image is inpainted by T2I models. However, they struggle with noisy depth maps and loss of semantic details when warping an input view to novel viewpoints. In this paper, we propose a novel approach for single-shot novel view synthesis, a semantic-preserving generative warping framework that enables T2I generative models to learn where to warp and where to generate, through augmenting cross-view attention with self-attention. Our approach addresses the limitations of existing methods by conditioning the generative model on source view images and incorporating geometric warping signals. Qualitative and quantitative evaluations demonstrate that our model outperforms existing methods in both in-domain and out-of-domain scenarios. Project page is available at https://GenWarp-NVS.github.io. Junyoung Seo, Kazumi Fukuda, Takashi Shibuya 0001, Takuya Narihira, Naoki Murata, Shoukang Hu, Chieh-Hsin Lai, Seungryong Kim, Yuki Mitsufuji |
NeurIPS | 8 |
| 2024 | Towards Open-Vocabulary Semantic Segmentation Without Semantic LabelsabstractLarge-scale vision-language models like CLIP have demonstrated impressive open-vocabulary capabilities for image-level tasks, excelling in recognizing what objects are present. However, they struggle with pixel-level recognition tasks like semantic segmentation, which require understanding where the objects are located. In this work, we propose a novel method, PixelCLIP, to adapt the CLIP image encoder for pixel-level understanding by guiding the model on where, which is achieved using unlabeled images and masks generated from vision foundation models such as SAM and DINO. To address the challenges of leveraging masks without semantic labels, we devise an online clustering algorithm using learnable class names to acquire general semantic concepts. PixelCLIP shows significant performance improvements over CLIP and competitive results compared to caption-supervised methods in open-vocabulary semantic segmentation. Heeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho, Anurag Arnab, Hongsuck Seo, Seungryong Kim |
NeurIPS | 7 |
| 2024 | InstaFormer++: Multi-Domain Instance-Aware Image-to-Image Translation with Transformer
Jong-Beom Baek, Eunjae Ha, Homin Jung, Seungryong Kim |
Int. J. Comput. Vis. | 7 |
| 2024 | Depth-aware guidance with self-estimated depth representations of diffusion models
Gyeongnyeon Kim, Woo-seok Jang, Gyuseong Lee, Susung Hong, Junyoung Seo, Seungryong Kim |
Pattern Recognit. | 6 |
| 2024 | Controllable Style Transfer via Test-time Training of Implicit Neural Representation
Youngjo Min, Younghun Jung, Seungryong Kim |
Pattern Recognit. | 4 |
| 2024 | Discriminative action tubelet detector for weakly-supervised action detection
Jiyoung Lee 0005, Seungryong Kim, Sunok Kim, Kwanghoon Sohn |
Pattern Recognit. | 2 |
| 2023 | MIDMs: Matching Interleaved Diffusion Models for Exemplar-Based Image TranslationabstractWe present a novel method for exemplar-based image translation, called matching interleaved diffusion models (MIDMs). Most existing methods for this task were formulated as GAN-based matching-then-generation framework. However, in this framework, matching errors induced by the difficulty of semantic matching across cross-domain, e.g., sketch and photo, can be easily propagated to the generation step, which in turn leads to the degenerated results. Motivated by the recent success of diffusion models, overcoming the shortcomings of GANs, we incorporate the diffusion models to overcome these limitations. Specifically, we formulate a diffusion-based matching-and-generation framework that interleaves cross-domain matching and diffusion steps in the latent space by iteratively feeding the intermediate warp into the noising process and denoising it to generate a translated image. In addition, to improve the reliability of diffusion process, we design confidence-aware process using cycle-consistency to consider only confident regions during translation. Experimental results show that our MIDMs generate more plausible images than state-of-the-art methods. Junyoung Seo, Gyuseong Lee, Seokju Cho, Jiyoung Lee 0005, Seungryong Kim |
AAAI | 5 |
| 2023 | PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-IdentificationabstractModern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains unexplored. In this paper, we present a novel data augmentation technique, dubbed PartMix, that synthesizes the augmented samples by mixing the part descriptors across the modalities to improve the performance of part-based VI-ReID models. Especially, we synthesize the positive and negative samples within the same and across different identities and regularize the backbone model through contrastive learning. In addition, we also present an entropy-based mining strategy to weaken the adverse impact of unreliable positive and negative samples. When incorporated into existing part-based VI-ReID model, PartMix consistently boosts the performance. We conduct experiments to demonstrate the effectiveness of our PartMix over the existing VI-ReID methods and provide ablation studies. Seungryong Kim, Jungin Park, Seongheon Park, Kwanghoon Sohn |
CVPR | 2 |
| 2023 | LANIT: Language-Driven Image-to-Image Translation for Unlabeled DataabstractExisting techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent truly-unsupervised methods adopt clustering approaches to easily provide per-sample one-hot domain labels. However, they cannot account for the real-world setting: one sample may have multiple attributes. In addition, the semantics of the clusters are not easily coupled to human understanding. To overcome these, we present LANguage-driven Image-to-image Translation model, dubbed LANIT. We leverage easy-to-obtain candidate attributes given in texts for a dataset: the similarity between images and attributes indicates per-sample domain labels. This formulation naturally enables multi-hot labels so that users can specify the target domain with a set of attributes in language. To account for the case that the initial prompts are inaccurate, we also present prompt learning. We further present domain regularization loss that enforces translated images to be mapped to the corresponding domain. Experiments on several standard benchmarks demonstrate that LANIT achieves comparable or superior performance to existing models. The code is available at github.com/KU-CVLAB/LANIT. Seokju Cho, Jaejun Yoo 0001, Youngjung Uh, Seungryong Kim |
CVPR | 7 |
| 2023 | Improving Sample Quality of Diffusion Models Using Self-Attention GuidanceabstractDenoising diffusion models (DDMs) have attracted attention for their exceptional generation quality and diversity. This success is largely attributed to the use of class- or text-conditional diffusion guidance methods, such as classifier and classifier-free guidance. In this paper, we present a more comprehensive perspective that goes beyond the traditional guidance methods. From this generalized perspective, we introduce novel condition- and training-free strategies to enhance the quality of generated images. As a simple solution, blur guidance improves the suitability of intermediate samples for their fine-scale information and structures, enabling diffusion models to generate higher quality samples with a moderate guidance scale. Improving upon this, Self-Attention Guidance (SAG) uses the intermediate self-attention maps of diffusion models to enhance their stability and efficacy. Specifically, SAG adversarially blurs only the regions that diffusion models attend to at each iteration and guides them accordingly. Our experimental re sults show that our SAG improves the performance of various diffusion models, including ADM, IDDPM, Stable Diffusion, and DiT. Moreover, combining SAG with conventional guidance methods leads to further improvement. Susung Hong, Gyuseong Lee, Woo-seok Jang, Seungryong Kim |
ICCV | 4 |
| 2023 | GeCoNeRF: Few-shot Neural Radiance Fields via Geometric ConsistencyabstractWe present a novel framework to regularize Neural Radiance Field (NeRF) in a few-shot setting with a geometry-aware consistency regularization. The proposed approach leverages a rendered depth map at unobserved viewpoint to warp sparse input images to the unobserved viewpoint and impose them as pseudo ground truths to facilitate learning of NeRF. By encouraging such geometry-aware consistency at a feature-level instead of using pixel-level reconstruction loss, we regularize the NeRF at semantic and structural levels while allowing for modeling view dependent radiance to account for color variations across viewpoints. We also propose an effective method to filter out erroneous warped solutions, along with training strategies to stabilize training during optimization. We show that our model achieves competitive results compared to state-of-the-art few-shot NeRF models. Minseop Kwak, Jiuhn Song, Seungryong Kim |
ICML | 3 |
| 2023 | Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D GenerationabstractExisting score-distilling text-to-3D generation techniques, despite their considerable promise, often encounter the view inconsistency problem. One of the most notable issues is the Janus problem, where the most canonical view of an object (\textit{e.g}., face or head) appears in other views. In this work, we explore existing frameworks for score-distilling text-to-3D generation and identify the main causes of the view inconsistency problem---the embedded bias of 2D diffusion models. Based on these findings, we propose two approaches to debias the score-distillation frameworks for view-consistent text-to-3D generation. Our first approach, called score debiasing, involves cutting off the score estimated by 2D diffusion models and gradually increasing the truncation value throughout the optimization process. Our second approach, called prompt debiasing, identifies conflicting words between user prompts and view prompts using a language model, and adjusts the discrepancy between view prompts and the viewing direction of an object. Our experimental results show that our methods improve the realism of the generated 3D objects by significantly reducing artifacts and achieve a good trade-off between faithfulness to the 2D diffusion models and 3D consistency with little overhead. Our project page is available at~\url{https://susunghong.github.io/Debiased-Score-Distillation-Sampling/}. Susung Hong, Donghoon Ahn, Seungryong Kim |
NeurIPS | 3 |
| 2023 | DäRF: Boosting Radiance Fields from Sparse Input Views with Monocular Depth AdaptationabstractNeural radiance field (NeRF) shows powerful performance in novel view synthesis and 3D geometry reconstruction, but it suffers from critical performance degradation when the number of known viewpoints is drastically reduced. Existing works attempt to overcome this problem by employing external priors, but their success is limited to certain types of scenes or datasets. Employing monocular depth estimation (MDE) networks, pretrained on large-scale RGB-D datasets, with powerful generalization capability may be a key to solving this problem: however, using MDE in conjunction with NeRF comes with a new set of challenges due to various ambiguity problems exhibited by monocular depths. In this light, we propose a novel framework, dubbed DäRF, that achieves robust NeRF reconstruction with a handful of real-world images by combining the strengths of NeRF and monocular depth estimation through online complementary training. Our framework imposes the MDE network's powerful geometry prior to NeRF representation at both seen and unseen viewpoints to enhance its robustness and coherence. In addition, we overcome the ambiguity problems of monocular depths through patch-wise scale-shift fitting and geometry distillation, which adapts the MDE network to produce depths aligned accurately with NeRF geometry. Experiments show our framework achieves state-of-the-art results both quantitatively and qualitatively, demonstrating consistent and reliable performance in both indoor and outdoor real-world datasets. Jiuhn Song, Seonghoon Park 0002, Honggyu An, Seokju Cho, Minseop Kwak, Sungjin Cho, Seungryong Kim |
NeurIPS | 7 |
| 2023 | 3D GAN Inversion with Pose OptimizationabstractWith the recent advances in NeRF-based 3D aware GANs quality, projecting an image into the latent space of these 3D-aware GANs has a natural advantage over 2D GAN inversion: not only does it allow multi-view consistent editing of the projected image, but it also enables 3D reconstruction and novel view synthesis when given only a single image. However, the explicit viewpoint control acts as a main hindrance in the 3D GAN inversion process, as both camera pose and latent code have to be optimized simultaneously to reconstruct the given image. Most works that explore the latent space of the 3D-aware GANs rely on ground-truth camera viewpoint or deformable 3D model, thus limiting their applicability. In this work, we introduce a generalizable 3D GAN inversion method that infers camera viewpoint and latent code simultaneously to enable multi-view consistent semantic image editing. The key to our approach is to leverage pre-trained estimators for better initialization and utilize the pixel-wise depth calculated from NeRF parameters to better reconstruct the given image. We conduct extensive experiments on image reconstruction and editing both quantitatively and qualitatively, and further compare our results with 2D GAN-based editing to demonstrate the advantages of utilizing the latent space of 3D GANs. Additional results and visualizations are available at https://3dgan-inversion.github.io/. Jaehoon Ko, Kyusun Cho, Daewon Choi, Kwangrok Ryoo, Seungryong Kim |
WACV | 5 |
| 2023 | CATs++: Boosting Cost Aggregation With Convolutions and TransformersabstractCost aggregation is a process in image matching tasks that aims to disambiguate the noisy matching scores. Existing methods generally tackle this by hand-crafted or CNN-based methods, which either lack robustness to severe deformations or inherit the limitation of CNNs that fail to discriminate incorrect matches due to limited receptive fields and inadaptability. In this paper, we introduce Cost Aggregation with Transformers (CATs) to tackle this by exploring global consensus among initial correlation map with the help of some architectural designs that allow us to benefit from global receptive fields of self-attention mechanism. To this end, we include appearance affinity modeling, which helps to disambiguate the noisy initial correlation maps. Furthermore, we introduce some techniques, including multi-level aggregation to exploit rich semantics prevalent at different feature levels and swapping self-attention to obtain reciprocal matching scores to act as a regularization. Although CATs can attain competitive performance, it may face some limitations, i.e., high computational costs, which may restrict its applicability only at limited resolution and hurt performance. To overcome this, we propose CATs++, an extension of CATs. Concretely, we introduce early convolutions prior to cost aggregation with a transformer to control the number of tokens and inject some convolutional inductive bias, then propose a novel transformer architecture for both efficient and effective cost aggregation, which results in apparent performance boost and cost reduction. With the reduced costs, we are able to compose our network with a hierarchical structure to process higher-resolution inputs. We show that the proposed method with these integrated outperforms the previous state-of-the-art methods by large margins. Codes and pretrained weights are available at: https://ku-cvlab.github.io/CATs-PlusPlus-Project-Page/. Seokju Cho, Sunghwan Hong, Seungryong Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Stereo Confidence Estimation via Locally Adaptive Fusion and Knowledge DistillationabstractStereo confidence estimation aims to estimate the reliability of the estimated disparity by stereo matching. Different from the previous methods that exploit the limited input modality, we present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. Moreover, we propose a knowledge distillation framework to learn more compact confidence estimation networks as student networks. By transferring the knowledge from LAF-Net as teacher networks, the student networks that solely take as input a disparity can achieve comparable performance. To transfer more informative knowledge, we also propose a module to learn the locally-varying temperature in a softmax function. We further extend this framework to a multiview scenario. Experimental results show that LAF-Net and its variations outperform the state-of-the-art stereo confidence methods on various benchmarks. Sunok Kim, Seungryong Kim, Dongbo Min, Pascal Frossard, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Call for Customized Conversation: Customized Conversation Grounding Persona and KnowledgeabstractHumans usually have conversations by making use of prior knowledge about a topic and background information of the people whom they are talking to. However, existing conversational agents and datasets do not consider such comprehensive information, and thus they have a limitation in generating the utterances where the knowledge and persona are fused properly. To address this issue, we introduce a call For Customized conversation (FoCus) dataset where the customized answers are built with the user's persona and Wikipedia knowledge. To evaluate the abilities to make informative and customized utterances of pre-trained language models, we utilize BART and GPT-2 as well as transformer-based models. We assess their generation abilities with automatic scores and conduct human evaluations for qualitative results. We examine whether the model reflects adequate persona and knowledge with our proposed two sub-tasks, persona grounding (PG) and knowledge grounding (KG). Moreover, we show that the utterances of our data are constructed with the proper knowledge and persona through grounding quality assessment. Yoonna Jang, Jungwoo Lim, Yuna Hur, Dongsuk Oh, Suhyune Son, Yeonsoo Lee, Dong-Hoon Shin, Seungryong Kim, Heuiseok Lim |
AAAI | 8 |
| 2022 | Deep Translation Prior: Test-Time Training for Photorealistic Style TransferabstractRecent techniques to solve photorealistic style transfer within deep convolutional neural networks (CNNs) generally require intensive training from large-scale datasets, thus having limited applicability and poor generalization ability to unseen images or styles. To overcome this, we propose a novel framework, dubbed Deep Translation Prior (DTP), to accomplish photorealistic style transfer through test-time training on given input image pair with untrained networks, which learns an image pair-specific translation prior and thus yields better performance and generalization. Tailored for such test-time training for style transfer, we present novel network architectures, with two sub-modules of correspondence and generation modules, and loss functions consisting of contrastive content, style, and cycle consistency losses. Our framework does not require offline training phase for style transfer, which has been one of the main challenges in existing methods, but the networks are to be solely learned during test time. Experimental results prove that our framework has a better generalization ability to unseen image pairs and even outperforms the state-of-the-art methods. Seungryong Kim |
AAAI | 3 |
| 2022 | InstaFormer: Instance-Aware Image-to-Image Translation with TransformerabstractWe present a novel Transformer-based network architecture for instance-aware image-to-image translation, dubbed InstaFormer, to effectively integrate global- and instance-level information. By considering extracted content featuresfrom an image as tokens, our networks discover global consensus of content features by considering context information through a self-attention module in Transformers. By augmenting such tokens with an instance-level feature extracted from the content feature with respect to bounding box information, our framework is capable of learning an interaction between object instances and the global image, thus boosting the instance-awareness. We replace layer normalization (LayerNorm) in standard Transformers with adaptive instance normalization (AdaIN) to enable a multi-modal translation with style codes. In addition, to improve the instance-awareness and translation quality at object regions, we present an instance-level content contrastive loss defined between input and translated image. We conduct experiments to demonstrate the effectiveness of our InstaFormer over the latest methods and provide extensive ablation studies. Jong-Beom Baek, Gyeongnyeon Kim, Seungryong Kim |
CVPR | 5 |
| 2022 | Semi-Supervised Learning of Semantic Correspondence with Pseudo-LabelsabstractEstablishing dense correspondences across semantically similar images remains a challenging task due to the significant intra-class variations and background clutters. Traditionally, a supervised learning was used for training the models, which required tremendous manually-labeled data, while some methods suggested a self-supervised or weakly-supervised learning to mitigate the reliance on the labeled data, but with limited performance. In this paper, we present a simple, but effective solution for semantic correspondence that learns the networks in a semi-supervised manner by supplementing few ground-truth correspondences via utilization of a large amount of confident correspondences as pseudo-labels, called SemiMatch. Specifically, our framework generates the pseudo-labels using the model's prediction itself between source and weakly-augmented target, and uses pseudo-labels to learn the model again between source and strongly-augmented target, which improves the robustness of the model. We also present a novel confidence measure for pseudo-labels and data augmentation tailored for semantic correspondence. In experiments, SemiMatch achieves state-of-the-art performance on various benchmarks. Kwangrok Ryoo, Junyoung Seo, Gyuseong Lee, Hansang Cho, Seungryong Kim |
CVPR | 7 |
| 2022 | Cost Aggregation with 4D Convolutional Swin Transformer for Few-Shot Segmentation
Sunghwan Hong, Seokju Cho, Jisu Nam, Stephen Lin 0001, Seungryong Kim |
ECCV (29) | 5 |
| 2022 | ConMatch: Semi-supervised Learning with Confidence-Guided Consistency Regularization
Youngjo Min, Gyuseong Lee, Junyoung Seo, Kwangrok Ryoo, Seungryong Kim |
ECCV (30) | 7 |
| 2022 | Joint Learning of Feature Extraction and Cost Aggregation for Semantic CorrespondenceabstractEstablishing dense correspondences across semantically similar images is one of the challenging tasks due to the significant intra-class variations and background clutters. To solve these problems, numerous methods have been proposed, focused on learning feature extractor or cost aggregation independently, which yields sub-optimal performance. In this paper, we propose a novel framework for jointly learning feature extraction and cost aggregation for semantic correspondence. By exploiting the pseudo labels from each module, the networks consisting of feature extraction and cost aggregation modules are simultaneously learned in a boosting fashion. Moreover, to ignore unreliable pseudo labels, we present a confidence-aware contrastive loss function for learning the networks in a weakly-supervised manner. We demonstrate our competitive results on standard benchmarks for semantic correspondence. Youngjo Min, Mira Kim, Seungryong Kim |
ICASSP | 4 |
| 2022 | Meta-Learned Initialization For 3D Human RecoveryabstractWe propose a novel framework for 3D human recovery from a single image that integrates meta-learning with test-time optimization. Compared to previous optimization-based or learning-based methods that showed limited generalization ability, the test-time optimization framework enables us to estimate an optimal network parameter, which yields better performance and generalization ability, but it is highly sensitive to an initial network parameter. To alleviate this, we present a meta-learning framework to learn better initial parameters for test-time optimization. Experimental results on standard benchmarks show that the proposed framework boost the test-time optimization performance compared to state-of-the-arts. Mira Kim, Youngjo Min, Seungryong Kim |
ICIP | 4 |
| 2022 | Semi-Supervised Learning with Mutual Distillation for Monocular Depth EstimationabstractWe propose a semi-supervised learning framework for monocular depth estimation. Compared to existing semi-supervised learning methods, which inherit limitations of both sparse supervised and unsupervised loss functions, we achieve the complementary advantages of both loss functions, by building two separate network branches for each loss and distilling each other through the mutual distillation loss function. We also present to apply different data augmentation to each branch, which improves the robustness. We conduct experiments to demonstrate the effectiveness of our framework over the latest methods and provide extensive ablation studies. Jong-Beom Baek, Gyeongnyeon Kim, Seungryong Kim |
ICRA | 3 |
| 2022 | Meta-confidence estimation for stereo matchingabstractWe propose a novel framework to estimate the confidence of a disparity map taking into account, for the first time, the uncertainty affecting the confidence estimation process itself. Conversely to other tasks such as disparity estimation, the uncertainty of confidence directly hints that the confidence should be increased if initially low, but with high uncertainty, decreased otherwise. By modelling such a cue in the form of a second-level confidence, or meta-confidence, our solution allows for finding incorrect predictions inferred by confidence estimator and for learning a correction for them. Our strategy is suited for any state-of-the-art method known in literature, either implemented using random forest classifiers or deep neural networks. Especially, for deep neural networks-based models, we present a multi-headed confidence estimator followed by an uncertainty network, so as to predict mean confidence and meta-confidence within a single network without the cost of lower accuracy, a known limitation in literature for uncertainty estimation. Experimental results on a variety of stereo algorithms and confidence estimation models prove that the modeled meta-confidence is meaningful of the reliability of the estimated confidence and allows for refining it. Seungryong Kim, Matteo Poggi, Sunok Kim, Kwanghoon Sohn, Stefano Mattoccia |
ICRA | 1 |
| 2022 | Neural Matching Fields: Implicit Representation of Matching Fields for Visual CorrespondenceabstractExisting pipelines of semantic correspondence commonly include extracting high-level semantic features for the invariance against intra-class variations and background clutters. This architecture, however, inevitably results in a low-resolution matching field that additionally requires an ad-hoc interpolation process as a post-processing for converting it into a high-resolution one, certainly limiting the overall performance of matching results. To overcome this, inspired by recent success of implicit neural representation, we present a novel method for semantic correspondence, called Neural Matching Field (NeMF). However, complicacy and high-dimensionality of a 4D matching field are the major hindrances, which we propose a cost embedding network to process a coarse cost volume to use as a guidance for establishing high-precision matching field through the following fully-connected network. Nevertheless, learning a high-dimensional matching field remains challenging mainly due to computational complexity, since a na\"ive exhaustive inference would require querying from all pixels in the 4D space to infer pixel-wise correspondences. To overcome this, we propose adequate training and inference procedures, which in the training phase, we randomly sample matching candidates and in the inference phase, we iteratively performs PatchMatch-based inference and coordinate optimization at test time. With these combined, competitive results are attained on several standard benchmarks for semantic correspondence. Code and pre-trained weights are available at~\url{https://ku-cvlab.github.io/NeMF/}. Sunghwan Hong, Jisu Nam, Seokju Cho, Susung Hong, Sangryul Jeon, Dongbo Min, Seungryong Kim |
NeurIPS | 7 |
| 2022 | Pyramidal Semantic Correspondence NetworksabstractThis paper presents a deep architecture, called pyramidal semantic correspondence networks (PSCNet), that estimates locally-varying affine transformation fields across semantically similar images. To deal with large appearance and shape variations that commonly exist among different instances within the same object category, we leverage a pyramidal model where the affine transformation fields are progressively estimated in a coarse-to-fine manner so that the smoothness constraint is naturally imposed. Different from the previous methods which directly estimate global or local deformations, our method first starts to estimate the transformation from an entire image and then progressively increases the degree of freedom of the transformation by dividing coarse cell into finer ones. To this end, we propose two spatial pyramid models by dividing an image in a form of quad-tree rectangles or into multiple semantic elements of an object. Additionally, to overcome the limitation of insufficient training data, a novel weakly-supervised training scheme is introduced that generates progressively evolving supervisions through the spatial pyramid models by leveraging a correspondence consistency across image pairs. Extensive experimental results on various benchmarks including TSS, Proposal Flow-WILLOW, Proposal Flow-PASCAL, Caltech-101, and SPair-71k demonstrate that the proposed method outperforms the lastest methods for dense semantic correspondence. Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | On the Confidence of Stereo Matching in a Deep-Learning Era: A Quantitative EvaluationabstractStereo matching is one of the most popular techniques to estimate dense depth maps by finding the disparity between matching pixels on two, synchronized and rectified images. Alongside with the development of more accurate algorithms, the research community focused on finding good strategies to estimate the reliability, i.e., the confidence, of estimated disparity maps. This information proves to be a powerful cue to naively find wrong matches as well as to improve the overall effectiveness of a variety of stereo algorithms according to different strategies. In this paper, we review more than ten years of developments in the field of confidence estimation for stereo matching. We extensively discuss and evaluate existing confidence measures and their variants, from hand-crafted ones to the most recent, state-of-the-art learning based methods. We study the different behaviors of each measure when applied to a pool of different stereo algorithms and, for the first time in literature, when paired with a state-of-the-art deep stereo network. Our experiments, carried out on five different standard datasets, provide a comprehensive overview of the field, highlighting in particular both strengths and limitations of learning-based strategies. Matteo Poggi, Seungryong Kim, Fabio Tosi, Sunok Kim, Filippo Aleotti, Dongbo Min, Kwanghoon Sohn, Stefano Mattoccia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Memory-Guided Image De-Raining Using Time-Lapse DataabstractThis paper addresses the problem of single image de-raining, that is, the task of recovering clean and rain-free background scenes from a single image obscured by a rainy artifact. Although recent advances adopt real-world time-lapse data to overcome the need for paired rain-clean images, they are limited to fully exploit the time-lapse data. The main cause is that, in terms of network architectures, they could not capture long-term rain streak information in the time-lapse data during training owing to the lack of memory components. To address this problem, we propose a novel network architecture combining the time-lapse data and, the memory network that explicitly helps to capture long-term rain streak information. Our network comprises the encoder-decoder networks and a memory network. The features extracted from the encoder are read and updated in the memory network that contains several memory items to store rain streak-aware feature representations. With the read/update operation, the memory network retrieves relevant memory items in terms of the queries, enabling the memory items to represent the various rain streaks included in the time-lapse data. To boost the discriminative power of memory features, we also present a novel background selective whitening (BSW) loss for capturing only rain streak information in the memory network by erasing the background information. Experimental results on standard benchmarks demonstrate the effectiveness and superiority of our approach. Jaehoon Cho, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2021 | Cross-Domain Grouping and Alignment for Domain Adaptive Semantic SegmentationabstractExisting techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain itself or estimated category, providing the limitation to encode the domains having a multi-modal data distribution. To overcome this limitation, we introduce a learnable clustering module, and a novel domain adaptation framework, called cross-domain grouping and alignment. To cluster the samples across domains with an aim to maximize the domain alignment without forgetting precise segmentation ability on the source domain, we present two loss functions, in particular, for encouraging semantic consistency and orthogonality among the clusters. We also present a loss so as to solve a class imbalance problem, which is the other limitation of the previous methods. Our experiments show that our method consistently boosts the adaptation performance in semantic segmentation, outperforming the state-of-the-arts on various domain adaptation settings. Sunghun Joung, Seungryong Kim, Jungin Park, Ig-Jae Kim, Kwanghoon Sohn |
AAAI | 3 |
| 2021 | RobustNet: Improving Domain Generalization in Urban-Scene Segmentation via Instance Selective WhiteningabstractEnhancing the generalization capability of deep neural networks to unseen domains is crucial for safety-critical applications in the real world such as autonomous driving. To address this issue, this paper proposes a novel instance selective whitening loss to improve the robustness of the segmentation networks for unseen domains. Our approach disentangles the domain-specific style and domain-invariant content encoded in higher-order statistics (i.e., feature covariance) of the feature representations and selectively removes only the style information causing domain shift. As shown in Fig. 1, our method provides reasonable predictions for (a) low-illuminated, (b) rainy, and (c) unseen structures. These types of images are not included in the training dataset, where the baseline shows a significant performance drop, contrary to ours. Being simple yet effective, our approach improves the robustness of various backbone networks without additional computational cost. We conduct extensive experiments in urban-scene segmentation and show the superiority of our approach to existing work. Our code is available at this link1. Sungha Choi, Sanghun Jung, Huiwon Yun, Joanne Taery Kim, Seungryong Kim, Jaegul Choo |
CVPR | 5 |
| 2021 | Mining Better Samples for Contrastive Learning of Temporal CorrespondenceabstractWe present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appropriately adjusting their hardness during training. This allows us to suppress the adverse impact of ambiguous matches and prevent a trivial solution from being yielded by too hard or too easy negative samples. To accomplish this, we incorporate three different criteria that ranges from a pixel-level matching confidence to a video-level one into a bottom-up pipeline, and plan a curriculum that is aware of current representation power for the adaptive hardness of negative samples during training. With the proposed method, state-of-the-art performance is attained over the latest approaches on several video label propagation tasks. Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
CVPR | 3 |
| 2021 | Adaptive confidence thresholding for monocular depth estimationabstractSelf-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to leverage pseudo ground truth depth maps of stereo images generated from self-supervised stereo matching methods. The confidence map of the pseudo ground truth depth map is estimated to mitigate performance degeneration by inaccurate pseudo depth maps. To cope with the prediction error of the confidence map itself, we also leverage the threshold network that learns the threshold dynamically conditioned on the pseudo depth maps. The pseudo depth labels filtered out by the thresholded confidence map are used to supervise the monocular depth network. Furthermore, we propose the probabilistic framework that refines the monocular depth map with the help of its uncertainty map through the pixel-adaptive convolution (PAC) layer. Experimental results demonstrate superior performance to state-of-the-art monocular depth estimation methods. Lastly, we exhibit that the proposed threshold learning can also be used to improve the performance of existing confidence estimation approaches. Hyesong Choi, Hunsang Lee, Sunkyung Kim, Sunok Kim, Seungryong Kim, Kwanghoon Sohn, Dongbo Min |
ICCV | 5 |
| 2021 | Deep Matching Prior: Test-Time Optimization for Dense CorrespondenceabstractConventional techniques to establish dense correspondences across visually or semantically similar images focused on designing a task-specific matching prior, which is difficult to model in general. To overcome this, recent learning-based methods have attempted to learn a good matching prior within a model itself on large training data. The performance improvement was apparent, but the need for sufficient training data and intensive learning hinders their applicability. Moreover, using the fixed model at test time does not account for the fact that a pair of images may require their own prior, thus providing limited performance and poor generalization to unseen images.In this paper, we show that an image pair-specific prior can be captured by solely optimizing the untrained matching networks on an input pair of images. Tailored for such test-time optimization for dense correspondence, we present a residual matching network and a confidence-aware contrastive loss to guarantee a meaningful convergence. Experiments demonstrate that our framework, dubbed Deep Matching Prior (DMP), is competitive, or even outperforms, against the latest learning-based methods on several benchmarks for geometric matching and semantic matching, even though it requires neither large training data nor intensive learning. With the networks pre-trained, DMP attains state-of-the-art performance on all benchmarks. Sunghwan Hong, Seungryong Kim |
ICCV | 2 |
| 2021 | Learning Canonical 3D Object Representation for Fine-Grained RecognitionabstractWe propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its appearance, while eliminating the effect of camera viewpoint, in a canonical configuration. Unlike conventional methods modeling spatial variation in 2D images only, our method is capable of reconfiguring the appearance feature in a canonical 3D space, thus enabling the subsequent object classifier to be invariant under 3D geometric variation. Our representation also allows us to go beyond existing methods, by incorporating 3D shape variation as an additional cue for object recognition. To learn the model without ground-truth 3D annotation, we deploy a differentiable renderer in an analysis-by-synthesis frame- work. By incorporating 3D shape and appearance jointly in a deep representation, our method learns the discriminative representation of the object and achieves competitive performance on fine-grained image recognition and vehicle re-identification. We also demonstrate that the performance of 3D shape reconstruction is improved by learning fine-grained shape deformation in a boosting manner. Sunghun Joung, Seungryong Kim, Ig-Jae Kim, Kwanghoon Sohn |
ICCV | 2 |
| 2021 | CATs: Cost Aggregation Transformers for Visual CorrespondenceabstractWe propose a novel cost aggregation network, called Cost Aggregation Transformers (CATs), to find dense correspondences between semantically similar images with additional challenges posed by large intra-class appearance and geometric variations. Cost aggregation is a highly important process in matching tasks, which the matching accuracy depends on the quality of its output. Compared to hand-crafted or CNN-based methods addressing the cost aggregation, in that either lacks robustness to severe deformations or inherit the limitation of CNNs that fail to discriminate incorrect matches due to limited receptive fields, CATs explore global consensus among initial correlation map with the help of some architectural designs that allow us to fully leverage self-attention mechanism. Specifically, we include appearance affinity modeling to aid the cost aggregation process in order to disambiguate the noisy initial correlation maps and propose multi-level aggregation to efficiently capture different semantics from hierarchical feature representations. We then combine with swapping self-attention technique and residual connections not only to enforce consistent matching, but also to ease the learning process, which we find that these result in an apparent performance boost. We conduct experiments to demonstrate the effectiveness of the proposed model over the latest methods and provide extensive ablation studies. Code and trained models are available at https://sunghwanhong.github.io/CATs/. Seokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee, Kwanghoon Sohn, Seungryong Kim |
NeurIPS | 6 |
| 2021 | Dense Cross-Modal Correspondence Estimation With the Deep Self-Correlation DescriptorabstractWe present the deep self-correlation (DSC) descriptor for establishing dense correspondences between images taken under different imaging modalities, such as different spectral ranges or lighting conditions. We encode local self-similar structure in a pyramidal manner that yields both more precise localization ability and greater robustness to non-rigid image deformations. Specifically, DSC first computes multiple self-correlation surfaces with randomly sampled patches over a local support window, and then builds pyramidal self-correlation surfaces through average pooling on the surfaces. The feature responses on the self-correlation surfaces are then encoded through spatial pyramid pooling in a log-polar configuration. To better handle geometric variations such as scale and rotation, we additionally propose the geometry-invariant DSC (GI-DSC) that leverages multi-scale self-correlation computation and canonical orientation estimation. In contrast to descriptors based on deep convolutional neural networks (CNNs), DSC and GI-DSC are training-free (i.e., handcrafted descriptors), are robust to cross-modality, and generalize well to various modality variations. Extensive experiments demonstrate the state-of-the-art performance of DSC and GI-DSC on challenging cases of cross-modal image pairs having photometric and/or geometric variations. Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Adversarial Confidence Estimation Networks for Robust Stereo MatchingabstractStereo matching aiming to perceive the 3-D geometry of a scene facilitates numerous computer vision tasks used in advanced driver assistance systems (ADAS). Although numerous methods have been proposed for this task by leveraging deep convolutional neural networks (CNNs), stereo matching still remains an unsolved problem due to its inherent matching ambiguities. To overcome these limitations, we present a method for jointly estimating disparity and confidence from stereo image pairs through deep networks. We accomplish this through a minmax optimization to learn the generative cost aggregation networks and discriminative confidence estimation networks in an adversarial manner. Concretely, the generative cost aggregation networks are trained to accurately generate disparities at both confident and unconfident pixels from an input matching cost that are indistinguishable by the discriminative confidence estimation networks, while the discriminative confidence estimation networks are trained to distinguish the confident and unconfident disparities. In addition, to fully exploit complementary information of matching cost, disparity, and color image in confidence estimation, we present a dynamic fusion module. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks including real driving scenes. Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | DUNIT: Detection-Based Unsupervised Image-to-Image TranslationabstractImage-to-image translation has made great strides in recent years, with current techniques being able to handle unpaired training images and to account for the multi-modality of the translation problem. Despite this, most methods treat the image as a whole, which makes the results they produce for content-rich scenes less realistic. In this paper, we introduce a Detection-based Unsupervised Image-to-image Translation (DUNIT) approach that explicitly accounts for the object instances in the translation process. To this end, we extract separate representations for the global image and for the instances, which we then fuse into a common representation from which we generate the translated image. This allows us to preserve the detailed content of object instances, while still modeling the fact that we aim to produce an image of a single consistent scene. We introduce an instance consistency loss to maintain the coherence between the detections. Furthermore, by incorporating a detector into our architecture, we can still exploit object instances at test time. As evidenced by our experiments, this allows us to outperform the state-of-the-art unsupervised image-to-image translation methods. Furthermore, our approach can also be used as an unsupervised domain adaptation strategy for object detection, and it also achieves state-of-the-art performance on this task. Deblina Bhattacharjee, Seungryong Kim, Guillaume Vizier, Mathieu Salzmann |
CVPR | 2 |
| 2020 | Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint EstimationabstractExisting techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcome this limitation, we introduce a learnable module, cylindrical convolutional networks (CCNs), that exploit cylindrical representation of a convolutional kernel defined in the 3D space. CCNs extract a view-specific feature through a view-specific convolutional kernel to predict object category scores at each viewpoint. With the view-specific feature, we simultaneously determine objective category and viewpoints using the proposed sinusoidal soft-argmax module. Our experiments demonstrate the effectiveness of the cylindrical convolutional networks on joint object detection and viewpoint estimation. Sunghun Joung, Seungryong Kim, Hanjae Kim, Ig-Jae Kim, Junghyun Cho, Kwanghoon Sohn |
CVPR | 2 |
| 2020 | Guided Semantic Flow
Sangryul Jeon, Dongbo Min, Seungryong Kim, Jihwan Choe, Kwanghoon Sohn |
ECCV (28) | 3 |
| 2020 | Volumetric Transformer Networks
Seungryong Kim, Sabine Süsstrunk, Mathieu Salzmann |
ECCV (28) | 1 |
| 2020 | Discrete-Continuous Transformation Matching for Dense Semantic CorrespondenceabstractTechniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Furthermore, leveraging correspondence consistency and confidence-guided filtering in each iteration facilitates the convergence of our method. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks and applications. Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Single Image Deraining Using Time-Lapse DataabstractLeveraging on recent advances in deep convolutional neural networks (CNNs), single image deraining has been studied as a learning task, achieving an outstanding performance over traditional hand-designed approaches. Current CNNs based deraining approaches adopt the supervised learning framework that uses a massive training data generated with synthetic rain streaks, having a limited generalization ability on real rainy images. To address this problem, we propose a novel learning framework for single image deraining that leverages time-lapse sequences instead of the synthetic image pairs. The deraining networks are trained using the time-lapse sequences in which both camera and scenes are static except for time-varying rain streaks. Specifically, we formulate a background consistency loss such that the deraining networks consistently generate the same derained images from the time-lapse sequences. We additionally introduce two loss functions, the structure similarity loss that encourages the derained image to be similar with an input rainy image and the directional gradient loss using the assumption that the estimated rain streaks are likely to be sparse and have dominant directions. To consider various rain conditions, we leverage a dynamic fusion module that effectively fuses multi-scale features. We also build a novel large-scale time-lapse dataset providing real world rainy images containing various rain conditions. Experiments demonstrate that the proposed method outperforms state-of-the-art techniques on synthetic and real rainy images both qualitatively and quantitatively. On the high-level vision tasks under severe rainy conditions, it has been shown that the proposed method can be utilized as a pre-preprocessing step for subsequent tasks. Jaehoon Cho, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2020 | Multi-Modal Recurrent Attention Networks for Facial Expression RecognitionabstractRecent deep neural networks based methods have achieved state-of-the-art performance on various facial expression recognition tasks. Despite such progress, previous researches for facial expression recognition have mainly focused on analyzing color recording videos only. However, the complex emotions expressed by people with different skin colors under different lighting conditions through dynamic facial expressions can be fully understandable by integrating information from multi-modal videos. We present a novel method to estimate dimensional emotion states, where color, depth, and thermal recording videos are used as a multi-modal input. Our networks, called multi-modal recurrent attention networks (MRAN), learn spatiotemporal attention volumes to robustly recognize the facial expression based on attention-boosted feature volumes. We leverage the depth and thermal sequences as guidance priors for color sequence to selectively focus on emotional discriminative regions. We also introduce a novel benchmark for multi-modal facial expression recognition, termed as multi-modal arousal-valence facial expression recognition (MAVFER), which consists of color, depth, and thermal recording videos with corresponding continuous arousal-valence scores. The experimental results show that our method can achieve the state-of-the-art results in dimensional facial expression recognition on color recording datasets including RECOLA, SEWA and AFEW, and a multi-modal recording dataset including MAVFER. Jiyoung Lee 0005, Sunok Kim, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2020 | Unsupervised Stereo Matching Using Confidential Correspondence ConsistencyabstractStereo matching aims to perceive the 3D geometric configuration of scenes and facilitates a variety of computer vision in advanced driver assistance systems (ADAS) applications. Recently, deep convolutional neural networks (CNNs) have shown dramatic performance improvements for computing the matching cost in the stereo matching. However, the performance of CNN-based approaches relies heavily on datasets, requiring a large number of ground truth data which needs tremendous works. To overcome this limitation, we present a novel framework to learn CNNs for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. By exploiting the correspondence consistency between stereo images, our method selects putative positive samples in each training iteration and utilizes them to train the networks. We further propose a positive sample propagation scheme to leverage additional training samples. Our unsupervised learning method is evaluated with two kinds of network architectures, simple and precise CNNs, and shows comparable performance to that of the state-of-the-art methods including both supervised and unsupervised learning approaches on KITTI, Middlebury, HCI, and Yonsei datasets. This extensive evaluation demonstrates that the proposed learning framework can be applied to deal with various real driving conditions. Sunghun Joung, Seungryong Kim, Kihong Park, Kwanghoon Sohn |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | High-Precision Depth Estimation Using Uncalibrated LiDAR and Stereo FusionabstractWe address the problem of 3D reconstruction from uncalibrated LiDAR point cloud and stereo images. Since the usage of each sensor alone for 3D reconstruction has weaknesses in terms of density and accuracy, we propose a deep sensor fusion framework for high-precision depth estimation. The proposed architecture consists of calibration network and depth fusion network, where both networks are designed considering the trade-off between accuracy and efficiency for mobile devices. The calibration network first corrects an initial extrinsic parameter to align the input sensor coordinate systems. The accuracy of calibration is markedly improved by formulating the calibration in the depth domain. In the depth fusion network, complementary characteristics of sparse LiDAR and dense stereo depth are then encoded in a boosting manner. Since training data for the LiDAR and stereo depth fusion are rather limited, we introduce a simple but effective approach to generate pseudo ground truth labels from the raw KITTI dataset. The experimental evaluation verifies that the proposed method outperforms current state-of-the-art methods on the KITTI benchmark. We also collect data using our proprietary multi-sensor acquisition platform and verify that the proposed method generalizes across different sensor settings and scenes. Kihong Park, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence EstimationabstractWe present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attention and scale maps to fuse the tri-modal confidence features. The attention inference networks encode the importance of tri-modal confidence features and then concatenate them using the attention maps in an adaptive and dynamic fashion. This enables us to make an optimal fusion of the heterogeneous features, compared to a simple concatenation technique that is commonly used in conventional approaches. In addition, to encode the confidence features with locally-varying receptive fields, the scale inference networks learn the scale map and warp the fused confidence features through convolutional spatial transformer networks. Finally, the confidence map is progressively estimated in the recursive refinement networks to enforce a spatial context and local consistency. Experimental results show that this model outperforms the state-of-the-art methods on various benchmarks. Sunok Kim, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
CVPR | 2 |
| 2019 | Semantic Attribute Matching NetworksabstractWe present semantic attribute matching networks (SAM-Net) for jointly establishing correspondences and transferring attributes across semantically similar images, which intelligently weaves the advantages of the two tasks while overcoming their limitations. SAM-Net accomplishes this through an iterative process of establishing reliable correspondences by reducing the attribute discrepancy between the images and synthesizing attribute transferred images using the learned correspondences. To learn the networks using weak supervisions in the form of image pairs, we present a semantic attribute matching loss based on the matching similarity between an attribute transferred source feature and a warped target feature. With SAM-Net, the state-of-the-art performance is attained on several benchmarks for semantic matching and attribute transfer. Seungryong Kim, Dongbo Min, Somi Jeong, Sunok Kim, Sangryul Jeon, Kwanghoon Sohn |
CVPR | 1 |
| 2019 | Joint Learning of Semantic Alignment and Object Landmark DetectionabstractConvolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In this paper, we present a joint learning approach for obtaining dense correspondences and discovering object landmarks from semantically similar images. Based on the key insight that the two tasks can mutually provide supervisions to each other, our networks accomplish this through a joint loss function that alternatively imposes a consistency constraint between the two tasks, thereby boosting the performance and addressing the lack of training data in a principled manner. To the best of our knowledge, this is the first attempt to address the lack of training data for the two tasks through the joint learning. To further improve the robustness of our framework, we introduce a probabilistic learning formulation that allows only reliable matches to be used in the joint learning process. With the proposed method, state-of-the-art performance is attained on several benchmarks for semantic matching and landmark detection. Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
ICCV | 3 |
| 2019 | Context-Aware Emotion Recognition NetworksabstractTraditional techniques for emotion recognition have focused on the facial expression analysis only, thus providing limited ability to encode context that comprehensively represents the emotional responses. We present deep networks for context-aware emotion recognition, called CAER-Net, that exploit not only human facial expression but also context information in a joint and boosting manner. The key idea is to hide human faces in a visual scene and seek other contexts based on an attention mechanism. Our networks consist of two sub-networks, including two-stream encoding networks to separately extract the features of face and context regions, and adaptive fusion networks to fuse such features in an adaptive fashion. We also introduce a novel benchmark for context-aware emotion recognition, called CAER, that is appropriate than existing benchmarks both qualitatively and quantitatively. On several benchmarks, CAER-Net proves the effect of context for emotion recognition. Our dataset is available at http://caer-dataset.github.io. Jiyoung Lee 0005, Seungryong Kim, Sunok Kim, Jungin Park, Kwanghoon Sohn |
ICCV | 2 |
| 2019 | Unpaired Cross-Spectral Pedestrian Detection Via Adversarial Feature LearningabstractEven though there exist significant advances in recent studies, existing methods for pedestrian detection still have shown limited performances under challenging illumination conditions especially at nighttime. To address this, cross-spectral pedestrian detection methods have been presented using color and thermal, and shown substantial performance gains on the challenging circumstances. However, their paired cross-spectral settings have limited applicability in real-world scenarios. To overcome this, we propose a novel learning framework for cross-spectral pedestrian detection in an unpaired setting. Based on an assumption that features from color and thermal images share their characteristics in a common feature space to benefit their complement information, we design the separate feature embedding networks for color and thermal images followed by sharing detection networks. To further improve the cross-spectral feature representation, we apply an adversarial learning scheme to intermediate features of the color and thermal images. Experiments demonstrate the outstanding performance of the proposed method on the KAIST multi-spectral benchmark in comparison to the state-of-the-art methods. Sunghun Joung, Kihong Park, Seungryong Kim, Kwanghoon Sohn |
ICIP | 4 |
| 2019 | Graph Regularization Network with Semantic Affinity for Weakly-Supervised Temporal Action LocalizationabstractThis paper presents a novel deep architecture for weakly-supervised temporal action localization that not only generates segment-level action responses but also propagates segment-level responses to the neighborhood in a form of graph Laplacian regularization. Specifically, our approach consists of two sub-modules; a class activation module to estimate the action score map over time through the action classifiers, and a graph regularization module to refine the estimated action score map by solving a quadratic programming problem with the predicted segment-level semantic affinities. Since these two modules are integrated with fully differentiable layers, the proposed networks can be jointly trained in an end-to-end manner. Experimental results on Thumos14 and ActivityNet1.2 demonstrate that the proposed method provides outstanding performances in weakly-supervised temporal action localization. Jungin Park, Jiyoung Lee 0005, Sangryul Jeon, Seungryong Kim, Kwanghoon Sohn |
ICIP | 4 |
| 2019 | Satellite Image-Based Ship Classification Method with Sentinel-1 Iw Mode DataabstractClassification of a ship based on satellite imagery usually results differently depending on the type of polarization used in image generation. Also, the different ship's orientation in each image degrades the performance of image-based classification. Given these points, Sentinel-1 data also needs some methods to classify the type of ship. For the classification, we have produced a ship dataset, KIOST-OpenSARShip, which was modified from the OpenSARShip dataset. We compared the brightness of each pixel of ship images generated by different polarizations. Based on this, we created a new image dataset. Then, we increased the similarity between ship images of the same type by aligning the heading direction in the ship images. As a result, our new datasets improve classification performances in some cases compared to using the OpenSARShip. The results of composite images from the VV- and VH-polarized images show up to 19.34% higher accuracy than those using only the one polarized images. In the future, we will improve the performance of the ship classification method considering various characteristics of the ship. Seungryong Kim, Jeongju Bae, Chan-Su Yang |
IGARSS | 1 |
| 2019 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. Unlike traditional dense correspondence approaches for estimating depth or optical flow, semantic correspondence estimation poses additional challenges due to intra-class appearance and shape variations among different instances within the same object or scene category. To robustly match points across semantically similar images, we formulate FCSS using local self-similarity (LSS), which is inherently insensitive to intra-class appearance variations. LSS is incorporated through a proposed convolutional self-similarity (CSS) layer, where the sampling patterns and the self-similarity measure are jointly learned in an end-to-end and multi-scale manner. Furthermore, to address shape variations among different object instances, we propose a convolutional affine transformer (CAT) layer that estimates explicit affine transformation fields at each pixel to transform the sampling patterns and corresponding receptive fields. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in most existing datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS significantly outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin 0001, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Learning to Find Unpaired Cross-Spectral CorrespondencesabstractWe present a deep architecture and learning framework for establishing correspondences across cross-spectral visible and infrared images in an unpaired setting. To overcome the unpaired cross-spectral data problem, we design the unified image translation and feature extraction modules to be learned in a joint and boosting manner. Concretely, the image translation module is learned only with the unpaired cross-spectral data, and the feature extraction module is learned with an input image and its translated image. By learning two modules simultaneously, the image translation module generates the translated image that preserves not only the domain-specific attributes with separate latent spaces but also the domain-agnostic contents with feature consistency constraint. In an inference phase, the cross-spectral feature similarity is augmented by intra-spectral similarities between the features extracted from the translated images. Experimental results show that this model outperforms the state-of-the-art unpaired image translation methods and cross-spectral feature descriptors on various visible and infrared benchmarks. Somi Jeong, Seungryong Kim, Kihong Park, Kwanghoon Sohn |
IEEE Trans. Image Process. | 2 |
| 2019 | Unified Confidence Estimation Networks for Robust Stereo MatchingabstractWe present a deep architecture that estimates a stereo confidence, which is essential for improving the accuracy of stereo matching algorithms. In contrast to existing methods based on deep convolutional neural networks (CNNs) that rely on only one of the matching cost volume or estimated disparity map, our network estimates the stereo confidence by using the two heterogeneous inputs simultaneously. Specifically, the matching probability volume is first computed from the matching cost volume with residual networks and a pooling module in a manner that yields greater robustness. The confidence is then estimated through a unified deep network that combines confidence features extracted both from the matching probability volume and its corresponding disparity. In addition, our method extracts the confidence features of the disparity map by applying multiple convolutional filters with varying sizes to an input disparity map. To learn our networks in a semi-supervised manner, we propose a novel loss function that use confident points to compute the image reconstruction loss. To validate the effectiveness of our method in a disparity post-processing step, we employ three post-processing approaches; cost modulation, ground control points-based propagation, and aggregated ground control points-based propagation. Experimental results demonstrate that our method outperforms state-of-the-art confidence estimation methods on various benchmarks. Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2018 | PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
ECCV (6) | 2 |
| 2018 | Spatiotemporal Attention Based Deep Neural Networks for Emotion RecognitionabstractWe propose a spatiotemporal attention based deep neural networks for dimensional emotion recognition in facial videos. To learn the spatiotemporal attention that selectively focuses on emotional sailient parts within facial videos, we formulate the spatiotemporal encoder-decoder network using Convolutional LSTM (ConvLSTM) modules, which can be learned implicitly without any pixel-level annotations. By leveraging the spatiotemporal attention, we also formulate the 3D convolutional neural networks (3D-CNNs) to robustly recognize the dimensional emotion in facial videos. The experimental results show that our method can achieve the state-of-the-art results in dimensional emotion recognition with the highest concordance correlation coefficient (CCC) on RECOLA and AV+EC 2017 dataset. Jiyoung Lee 0005, Sunok Kim, Seungryong Kim, Kwanghoon Sohn |
ICASSP | 3 |
| 2018 | High-Precision Depth Estimation with the 3D LiDAR and Stereo FusionabstractWe present a deep convolutional neural network (CNN) architecture for high-precision depth estimation by jointly utilizing sparse 3D LiDAR and dense stereo depth information. In this network, the complementary characteristics of sparse 3D LiDAR and dense stereo depth are simultaneously encoded in a boosting manner. Tailored to the LiDAR and stereo fusion problem, the proposed network differs from previous CNNs in the incorporation of a compact convolution module, which can be deployed with the constraints of mobile devices. As training data for the LiDAR and stereo fusion is rather limited, we introduce a simple yet effective approach for reproducing the raw KITTI dataset. The raw LiDAR scans are augmented by adapting an off-the-shelf stereo algorithm and a confidence measure. We evaluate the proposed network on the KITTI benchmark and data collected by our multi-sensor acquisition system. Experiments demonstrate that the proposed network generalizes across datasets and is significantly more accurate than various baseline approaches. Kihong Park, Seungryong Kim, Kwanghoon Sohn |
ICRA | 2 |
| 2018 | CoVieW'18: The 1st Workshop and Challenge on Comprehensive Video Understanding in the WildabstractThe 1st Workshop and Challenge on Comprehensive Video Understanding in the Wild, dubbed CoVieW'18, is held in Seoul, Korea on October 22, 2018, in conjuction with ACM Multimedia 2018. The workshop aims to solve the joint and comprehensive understanding problem in untrimmed videos with a particular emphasis on joint action and scene recognition. The workshop encourages researchers to participate in joint action and scene recognition challenge in untrimmed videos and to report their results. The workshop program includes 1 keynote speech, 2 invited speakers, 6 regular and challenge papers. The developments made in the workshop will deliver a step change in a variety of video applications. Kwanghoon Sohn, Ming-Hsuan Yang 0001, Hyeran Byun, Jongwoo Lim, Gee-Sern Hsu, Stephen Lin 0001, Euntai Kim, Seungryong Kim |
ACM Multimedia | 8 |
| 2018 | Recurrent Transformer Networks for Semantic CorrespondenceabstractWe present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convolutional activations. By directly estimating the transformations between an image pair, rather than employing spatial transformer networks to independently normalize each individual image, we show that greater accuracy can be achieved. This process is conducted in a recursive manner to refine both the transformation estimates and the feature representations. In addition, a technique is presented for weakly-supervised training of RTNs that is based on a proposed classification loss. With RTNs, state-of-the-art performance is attained on several benchmarks for semantic correspondence. Seungryong Kim, Stephen Lin 0001, Sangryul Jeon, Dongbo Min, Kwanghoon Sohn |
NeurIPS | 1 |
| 2018 | Unified multi-spectral pedestrian detection based on probabilistic fusion networks
Kihong Park, Seungryong Kim, Kwanghoon Sohn |
Pattern Recognit. | 2 |
| 2017 | Reliable smart energy IoT-cloud service operation with container orchestrationabstractWe discuss the prototype implementation of IoT-Cloud services for efficient cooling management for small-size server room. We leverage open-source-based miniaturized playground for IoT-Cloud services and realize smart energy IoT-Cloud service based on the container-leveraged service orchestration with Docker containers and Docker Swarm. ln addition, we attempt to facilitate the automated and reliable operation of IoT-Cloud services with container-based orchestration. Seungryong Kim, Chorwon Kim, Jongwon Kim 0001 |
APNOMS | 1 |
| 2017 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to existing CNN-based descriptors, FCSS is inherently insensitive to intra-class appearance variations because of its LSS-based structure, while maintaining the precise localization ability of deep neural networks. The sampling patterns of local structure and the self-similarity measure are jointly learned within the proposed network in an end-to-end and multi-scale manner. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in existing image datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Sangryul Jeon, Stephen Lin 0001, Kwanghoon Sohn |
CVPR | 1 |
| 2017 | DCTM: Discrete-Continuous Transformation Matching for Semantic FlowabstractTechniques for dense semantic correspondence have provided limited ability to deal with the geometric variations that commonly exist between semantically similar images. While variations due to scale and rotation have been examined, there is a lack of practical solutions for more complex deformations such as affine transformations because of the tremendous size of the associated solution space. To address this problem, we present a discrete-continuous transformation matching (DCTM) framework where dense affine transformation fields are inferred through a discrete label optimization in which the labels are iteratively updated via continuous regularization. In this way, our approach draws solutions from the continuous space of affine transformations in a manner that can be computed efficiently through constant-time edge-aware filtering and a proposed affine-varying CNN-based descriptor. Experimental results show that this model outperforms the state-of-the-art methods for dense semantic correspondence on various benchmarks. Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
ICCV | 1 |
| 2017 | Multispectral human co-segmentation via joint convolutional neural networksabstractWe present a novel human body co-segmentation method for unregistered multispectral, color and thermal, images by leveraging CNNs. The main challenges for that tasks are no-alignment between color and thermal images and an absent of ground truth human segmentation labels. To solve these limitations, our key-insight is to formulate the segmentation network for each modality that solve two sub-tasks, correspondence and classification, in a joint and iterative manner. We formulate the learning framework between multispectral images in a way that training labels for one modality are used to learn the network for the other modality. We estimate dense correspondences between multispectral image pairs using intermediate convolutional activations of CNNs and perform human segmentation for each modality through the conditional random fields (CRF) optimization using unary and pairwise fusion. These two steps are formulated as an iterative framework, enables the network to converge on an optimal solution. Experimental results show that our proposed method outperforms conventional state-of-the-art methods on the VAP benchmark consisting of unregistered multispectral color and thermal images. Sungil Choi, Seungryong Kim, Kihong Park, Kwanghoon Sohn |
ICIP | 2 |
| 2017 | Convolutional feature pyramid fusion via attention networkabstractWe present a novel fusion scheme between multiple intermediate convolutional features within convolutional neurual network (CNN) for dense correspondence estimation. In contrast to existing CNN-based descriptors that utilize a single convolutional activation, our approach jointly uses multiple intermediate features of CNN through the attention weight that balances the contribution of each features. We formulate the overall network as two sub-networks, correspondence network and attention network. The correspondence network is designed to provide multiple intermediate matching costs while the attention network is to learn the optimal weight between them. These two networks are learned in a joint manner to boost the correspondence estimation performance. Experiments demonstrate that our proposed method outperforms the state-of-the-art methods on various correspondence estimation tasks including depth estimation, optical flow, and semantic correspondence. Sangryul Jeon, Seungryong Kim, Kwanghoon Sohn |
ICIP | 2 |
| 2017 | Convolutional cost aggregation for robust stereo matchingabstractAlthough convolutional neural network (CNN)-based stereo matching methods have become increasingly popular thanks to their robustness, they primarily have been focused on the matching cost computation. By leveraging CNNs, we present a novel method for matching cost aggregation to boost the stereo matching performance. Our insight is to learn the convolution kernel within CNN architecture for cost aggregation in a fully convolutional manner. Tailored to cost aggregation problem, our method differs from handcrafted methods in terms of its convolutional aggregation through optimally learned CNNs. First, the matching cost is aggregated with cost volume unary network, and then optimized with explicit disparity boundary, estimated through disparity boundary pairwise network, within a global energy minimization. Experiments demonstrate that our method outperforms conventional hand-crafted aggregation methods. Somi Jeong, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn |
ICIP | 2 |
| 2017 | Unsupervised stereo matching using correspondence consistencyabstractDeep convolutional neural networks (CNNs) have shown revolutionary performance improvements for matching cost computation in stereo matching. However, conventional CNN-based approaches to learn the network in a supervised manner require a large number of ground-truth disparity maps, which limits their applicability. To overcome this limitation, we present a novel framework to learn a CNNs architecture for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. Exploiting the correspondence consistency between stereo images as supervision, our method selects the training samples in each iteration during network training and uses them to learn the network. To boost the performance, we also propose a multi-scale cost computation scheme. Experimental results show that our method outperforms the state-of-the-art methods including even supervised learning based methods on various benchmarks. Sunghun Joung, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn |
ICIP | 2 |
| 2017 | Deep stereo confidence prediction for depth estimationabstractWe present a novel method that predicts a confidence to improve the accuracy of an estimated depth map in stereo matching. In contrast to existing learning based approaches relying on hand-crafted confidence features, we cast this problem into a convolutional neural network, learned using both a matching cost volume and its associated disparity map. As the size of the matching cost volume varies depending on a search range of stereo image pairs, we propose to use a top-K matching probability volume layer so that an input size for convolutional layers remains unchanged. Experimental results demonstrate that the proposed method outperforms the state-of-the-art confidence estimation approaches on various benchmarks. Sunok Kim, Dongbo Min, Bumsub Ham, Seungryong Kim, Kwanghoon Sohn |
ICIP | 4 |
| 2017 | Pedestrian proposal generation using depth-aware scale estimationabstractIn this work, we propose an efficient method that generates pedestrian proposals suitable for the autonomous vehicle. Our main intuition is that depth information provides an important cue to assign the scale of pedestrian proposals. Based on the observation that in a 3-D world coordinate the scales of pedestrians are almost similar, we formulate the scales of pedestrian patches by projecting 3-D models to an image plane with its corresponding depth. We also introduce a scale-aware binary description using both color and depth images. By using this descriptor, the regression models are trained to rank the pedestrian proposal candidates and adjust the proposal bounding boxes for an accurate localization. Our algorithm achieves significant performance gains compared to conventional proposal generation methods on the challenging KITTI dataset. Kihong Park, Seungryong Kim, Kwanghoon Sohn |
ICIP | 2 |
| 2017 | DASC: Robust Dense Descriptor for Multi-Modal and Multi-Spectral Correspondence EstimationabstractEstablishing dense correspondences between multiple images is a fundamental task in many applications. However, finding a reliable correspondence between multi-modal or multi-spectral images still remains unsolved due to their challenging photometric and geometric variations. In this paper, we propose a novel dense descriptor, called dense adaptive self-correlation (DASC), to estimate dense multi-modal and multi-spectral correspondences. Based on an observation that self-similarity existing within images is robust to imaging modality variations, we define the descriptor with a series of an adaptive self-correlation similarity measure between patches sampled by a randomized receptive field pooling, in which a sampling pattern is obtained using a discriminative learning. The computational redundancy of dense descriptors is dramatically reduced by applying fast edge-aware filtering. Furthermore, in order to address geometric variations including scale and rotation, we propose a geometry-invariant DASC (GI-DASC) descriptor that effectively leverages the DASC through a superpixel-based representation. For a quantitative evaluation of the GI-DASC, we build a novel multi-modal benchmark as varying photometric and geometric conditions. Experimental results demonstrate the outstanding performance of the DASC and GI-DASC in many cases of dense multi-modal and multi-spectral correspondences. Seungryong Kim, Dongbo Min, Bumsub Ham, Minh N. Do, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | LAT: Local area transform for cross modal correspondence matching
Seungchul Ryu, Seungryong Kim, Kwanghoon Sohn |
Pattern Recognit. | 2 |
| 2017 | Modality-Invariant Image Classification Based on Modality Uniqueness and Dictionary LearningabstractWe present a unified framework for the image classification of image sets taken under varying modality conditions. Our method is motivated by a key observation that the image feature distribution is simultaneously influenced by the semantic-class and the modality category label, which limits the performance of conventional methods for that task. With this insight, we introduce modality uniqueness as a discriminative weight that divides each modality cluster from all other clusters. By leveraging the modality uniqueness, our framework is formulated as unsupervised modality clustering and classifier learning based on modality-invariant similarity kernel. Specifically, in the assignment step, each training image is first assigned to the most similar cluster according to its modality. In the update step, based on the current cluster hypothesis, the modality uniqueness and the sparse dictionary are updated. These two steps are formulated in an iterative manner. Based on the final clusters, a modality-invariant marginalized kernel is then computed, where the similarities between the reconstructed features of each modality are aggregated across all clusters. Our framework enables the reliable inference of semantic-class category for an image, even across large photometric variations. Experimental results show that our method outperforms conventional methods on various benchmarks, such as landmark identification under severely varying weather conditions, domain-adapting image classification, and RGB and near-infrared image classification. Seungryong Kim, Rui Cai 0002, Kihong Park, Sunok Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2017 | Feature Augmentation for Learning Confidence Measure in Stereo MatchingabstractConfidence estimation is essential for refining stereo matching results through a post-processing step. This problem has recently been studied using a learning-based approach, which demonstrates a substantial improvement on conventional simple non-learning based methods. However, the formulation of learning-based methods that individually estimates the confidence of each pixel disregards spatial coherency that might exist in the confidence map, thus providing a limited performance under challenging conditions. Our key observation is that the confidence features and resulting confidence maps are smoothly varying in the spatial domain, and highly correlated within the local regions of an image. We present a new approach that imposes spatial consistency on the confidence estimation. Specifically, a set of robust confidence features is extracted from each superpixel decomposed using the Gaussian mixture model, and then these features are concatenated with pixel-level confidence features. The features are then enhanced through adaptive filtering in the feature domain. In addition, the resulting confidence map, estimated using the confidence features with a random regression forest, is further improved through K-nearest neighbor based aggregation scheme on both pixel- and superpixel-level. To validate the proposed confidence estimation scheme, we employ cost modulation or ground control points based optimization in stereo matching. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches on various benchmarks including challenging outdoor scenes. Sunok Kim, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2016 | Deep Self-correlation Descriptor for Dense Cross-Modal Correspondence
Seungryong Kim, Dongbo Min, Stephen Lin 0001, Kwanghoon Sohn |
ECCV (8) | 1 |
| 2016 | Unified Depth Prediction and Intrinsic Image Decomposition from a Single Image via Joint Convolutional Neural Fields
Seungryong Kim, Kihong Park, Kwanghoon Sohn, Stephen Lin 0001 |
ECCV (8) | 1 |
| 2016 | ANCC flow: Adaptive normalized cross-correlation with evolving guidance aggregation for dense correspondence estimationabstractAdaptive normalized cross-correlation (ANCC) cost function works well between images under photometric distortions, but its heavy computational burden often limits its applications. To overcome this limitation, this paper proposes a robust and efficient computational framework, called ANCC flow, designed for establishing dense correspondences between images under severe photometric variations. We first simplify the weight of ANCC in an asymmetric manner by considering a source image weight only. It is then efficiently computed by applying constant-time edge-aware filters without loss of its matching accuracy. Additionally, to deal with a large discrete label space effectively, which is a challenging issue in a flow field estimation, we propose a randomized label space sampling strategy similar to PatchMatch filer (PMF) optimization. The robustness of the asymmetric ANCC and the cost filter is further enhanced through an evolving weight computation, where a flow field computed in a previous iteration is utilized to build current edge-aware weights. Experimental results demonstrate the outstanding performance of ANCC flow in many cases of dense correspondence estimations under severe photometric and geometric variations. Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
ICIP | 1 |
| 2016 | Multi-spectral pedestrian detection based on accumulated object proposal with fully convolutional networksabstractThis paper presents a method for detecting a pedestrian by leveraging multi-spectral image pairs. Our approach is based on the observation that a multi-spectral image, especially far-infrared (FIR) image, enables us to overcome inherent limitations for pedestrian detection under challenging circumstances, such as even dark environments. For that task, multi-spectral color-FIR image pairs are used in a synergistic manner for pedestrian detection through deep convolutional neural networks (CNNs) learning and support vector regression (SVR). For inferring the confidence of a pedestrian, we first learn CNNs between color images (or FIR images) and bounding box annotations of pedestrians, respectively. Furthermore, for each object proposal, we extract intermediate activation features from network, and learn the probability of pedestrian using SVR. To improve the detection performance, the learned probability of pedestrian for each proposal is accumulated on the image domain. Based on the pedestrian confidence estimated from each network and accumulated pedestrian probabilities, the most probable pedestrian is finally localized among object proposal candidates. Thanks to its high robustness of multi-spectral imaging in dark environments and its high discriminative power of deep CNNs, our framework is shown to surpass state-of-the-art pedestrian detection methods on multi-spectral pedestrian benchmark. Seungryong Kim, Kihong Park, Kwanghoon Sohn |
ICPR | 2 |
| 2016 | Fast illumination-robust foreground detection using hierarchical distribution map for real-time video surveillance system
Jongin Son, Seungryong Kim, Kwanghoon Sohn |
Expert Syst. Appl. | 2 |
| 2015 | Randomized Global Transformation Approach for Dense Correspondence
Kihong Park, Seungryong Kim, Seungchul Ryu, Kwanghoon Sohn |
BMVC | 2 |
| 2015 | DASC: Dense adaptive self-correlation descriptor for multi-modal and multi-spectral correspondenceabstractEstablishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramatically advanced in recent studies. However, finding reliable visual correspondence in multi-modal or multi-spectral images still remains unsolved. In this paper, we propose a novel dense matching descriptor, called dense adaptive self-correlation (DASC), to effectively address this kind of matching scenarios. Based on the observation that a self-similarity existing within images is less sensitive to modality variations, we define the descriptor with a series of an adaptive self-correlation similarity for patches within a local support window. To further improve the matching quality and runtime efficiency, we propose a randomized receptive field pooling, in which a sampling pattern is optimized with a discriminative learning. Moreover, the computational redundancy that arises when computing densely sampled descriptor over an entire image is dramatically reduced by applying fast edge-aware filtering. Experiments demonstrate the outstanding performance of the DASC descriptor in many cases of multi-modal and multi-spectral correspondence. Seungryong Kim, Dongbo Min, Bumsub Ham, Seungchul Ryu, Minh N. Do, Kwanghoon Sohn |
CVPR | 1 |
| 2015 | Fast affine-invariant image matching based on global Bhattacharyya measure with adaptive treeabstractEstablishing visual correspondence is one of the most fundamental tasks in many applications of computer vision fields. In this paper we propose a robust image matching to address the affine variation problems between two images taken under different viewpoints. Unlike the conventional approach finding the correspondence with local feature matching on fully affine transformed-images, which provides many outliers with a time consuming scheme, our approach is to find only one global correspondence and then utilizes the local feature matching to estimate the most reliable inliers between two images. In order to estimate a global image correspondence very fast as varying affine transformation in affine space of reference and query images, we employ a Bhattacharyya similarity measure between two images. Furthermore, an adaptive tree with affine transformation model is employed to dramatically reduce the computational complexity. Our approach represents the satisfactory results for severe affine transformed-images while providing a very low computational time. Experimental results show that the proposed affine-invariant image matching is twice faster than the state-of-the-art methods at least, and provides better correspondence performance under viewpoint change conditions. Jongin Son, Seungryong Kim, Kwanghoon Sohn |
ICIP | 2 |
| 2015 | A multi-vision sensor-based fast localization system with image matching for challenging outdoor environments
Jongin Son, Seungryong Kim, Kwanghoon Sohn |
Expert Syst. Appl. | 2 |
| 2014 | Robust Stereo Matching Using Probabilistic Laplacian Surface Propagation
Seungryong Kim, Bumsub Ham, Seungchul Ryu, Seon Joo Kim, Kwanghoon Sohn |
ACCV (1) | 1 |
| 2014 | Local self-similarity frequency descriptor for multispectral feature matchingabstractThis paper describes a robust feature descriptor called the local self-similarity frequency (LSSF) for the multispectral RGB-NIR feature matching, which uses the frequency response of the local internal layout of self-similarities. A nonlinear relationship between multi-spectral image pairs makes conventional descriptors be sensitive to spectral deformation. To alleviate this problem, the LSSF employs a weighted correlation surface reducing the discrepancy between mul-tispectral images. Furthermore, the LSSF provides a rotation invariance exploiting the frequency response of maximal values on logpolar bins based on the fact that a cyclic shift on the log-polar representation leads only a phase shift in a frequency domain. Experimental results show that LSSF outperforms state-of-the-art descriptors in terms of a recognition rate for multispectral RGB-NIR image pairs. Seungryong Kim, Seungchul Ryu, Bumsub Ham, Junhyung Kim, Kwanghoon Sohn |
ICIP | 1 |
| 2014 | Synthesis quality prediction model based on distortion intoleranceabstractFree-viewpoint video system will provide viewers with freedom to navigate through the scene at different viewpoints. In the system, arbitrary viewpoints of videos are synthesized by the depth image-based rendering with multi-view plus depth videos. Despite the widespread of technologies for free-viewpoint video system, the field of quality assessment for the free-viewpoint video, especially the quality prediction of a synthesized image, has not yet been thoroughly investigated. This paper analyzes how distortions in color and depth images influence on the quality of a synthesized image. Then, an objective quality prediction model for a synthesized image is proposed based on the concept of intolerance of synthesis distortion. Experimental results show that the proposed model provides outstanding performance in predicting the quality of a synthesized image compared to other models. Seungchul Ryu, Seungryong Kim, Kwanghoon Sohn |
ICIP | 2 |
| 2014 | Mahalanobis Distance Cross-Correlation for Illumination-Invariant Stereo MatchingabstractA robust similarity measure called the Mahalanobis distance cross-correlation (MDCC) is proposed for illumination-invariant stereo matching, which uses a local color distribution within support windows. It is shown that the Mahalanobis distance between the color itself and the average color is preserved under affine transformation. The MDCC converts pixels within each support window into the Mahalanobis distance transform (MDT) space. The similarity between MDT pairs is then computed using the cross-correlation with an asymmetric weight function based on the Mahalanobis distance. The MDCC considers correlation on cross-color channels, thus providing robustness to affine illumination variation. Experimental results show that the MDCC outperforms state-of-the-art similarity measures in terms of stereo matching for image pairs taken under different illumination conditions. Seungryong Kim, Bumsub Ham, Bongjoe Kim, Kwanghoon Sohn |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | ABFT: Anisotropic binary feature transform based on structure tensor spaceabstractLocal feature matching is a fundamental step for many computer vision applications. Recently, binary feature transforms have been popularly proposed to improve the computational efficiency while preserving high matching performance. However, it is sensitive to noise and geometrical distortion such as affine transformation. In this paper, we propose ABFT framework, composed of a noise robust feature detection and affine invariant binary feature description based on a structure tensor space. Experimental results show that ABFT outperforms other state-of-the-art feature transforms in terms of the repeatability, recognition rate, and computational time. Seungryong Kim, Hunjae Yoo, Seungchul Ryu, Bumsub Ham, Kwanghoon Sohn |
ICIP | 1 |