Sunghoon Im 0001

dblp:174/1228-1 · DBLP profile ↗
← Back
49ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0001-9776-8101ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 6 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 23 since 2021Systems, architecture and hardware · 6 · 2 since 2021
YearPublicationVenuePosition
2026 Infinite-Story: A Training-Free Consistent Text-to-Image Generation
abstract
We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6x faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling.
Jihun Park 0001, Kyoungmin Lee 0001, Jongmin Gim 0001, Hyeonseo Jo, Minseok Oh, Wonhyeok Choi, Kyumin Hwang, Jae-Yeul Kim, Minwoo Choi, Sunghoon Im 0001
AAAI10
2026 CascadeOcc: Rethinking 3D Occupancy World Models With Cascaded VQ Representations
abstract
This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for autonomous driving. Occupancy world models—forecasting the future driving environment and planning the driving trajectory—effectively bridge perception and planning, but current approaches often heavily rely on external modalities or large language models, failing to fully exploit the inherent structural potential of occupancy representations themselves. To enhance representational capacity for complex 3D scenes, we integrate a cascaded Vector Quantized (VQ) mechanism into an autoregressive framework. Following a coarse-to-fine principle, CascadeOcc progressively refines fine-grained details from global structures through a multi-scale architecture. Additionally, we incorporate a TimeMixer to capture multi-scale temporal dependencies, establishing a dual-hierarchy mechanism in both space and time. Experimental results on 4D occupancy forecasting and motion planning benchmarks demonstrate that CascadeOcc achieves superior performance among vision-centric approaches, validating that optimizing inherent representations is a powerful alternative to relying on external foundation models.
Kyumin Hwang, Wonhyeok Choi, Jae-Yeul Kim, Jihun Park 0001, Daehee Park 0001, Sunghoon Im 0001
IEEE Signal Process. Lett.6
2025 Intrinsic Image Decomposition for Robust Self-supervised Monocular Depth Estimation on Reflective Surfaces
abstract
Self-supervised monocular depth estimation (SSMDE) has gained attention in the field of deep learning as it estimates depth without requiring ground truth depth maps. This approach typically uses a photometric consistency loss between a synthesized image, generated from the estimated depth, and the original image, thereby reducing the need for extensive dataset acquisition. However, the conventional photometric consistency loss relies on the Lambertian assumption, which often leads to significant errors when dealing with reflective surfaces that deviate from this model. To address this limitation, we propose a novel framework that incorporates intrinsic image decomposition into SSMDE. Our method synergistically trains for both monocular depth estimation and intrinsic image decomposition. The accurate depth estimation facilitates multi-image consistency for intrinsic image decomposition by aligning different view coordinate systems, while the decomposition process identifies reflective areas and excludes corrupted gradients from the depth training process. Furthermore, our framework introduces a pseudo-depth generation and knowledge distillation technique to further enhance the performance of the student model across both reflective and non-reflective surfaces. Comprehensive evaluations on multiple datasets show that our approach significantly outperforms existing SSMDE baselines in depth prediction, especially on reflective surfaces.
Wonhyeok Choi, Kyumin Hwang, Minwoo Choi, Kiljoon Han, Wonjoon Choi, Mingyu Shin, Sunghoon Im 0001
AAAI7
2025 Towards Lossless Implicit Neural Representation via Bit Plane Decomposition
abstract
We quantify the upper bound on the size of the implicit neural representation (INR) model from a digital perspective. The upper bound of the model size increases exponentially as the required bit-precision increases. To this end, we present a bit-plane decomposition method that makes INR predict bit-planes, producing the same effect as reducing the upper bound of the model size. We validate our hypothesis that reducing the upper bound leads to faster convergence with constant model size. Our method achieves lossless representation in 2D image and audio fitting, even for high bit-depth signals, such as 16-bit, which was previously unachievable. We pioneered the presence of bit bias, which INR prioritizes as the most significant bit (MSB). We expand the application of the INR task to bit depth expansion, lossless image compression, and extreme network quantization. Our source code is available at https://github.com/WooKyoungHan/LosslessINR.
Woo Kyoung Han, Byeonghun Lee, Hyunmin Cho, Sunghoon Im 0001, Kyong Hwan Jin
CVPR4
2025 Style-Editor: Text-driven Object-centric Style Editing
abstract
We present Text-driven object-centric style editing model named Style-Editor, a novel method that guides style editing at an object-centric level using textual inputs. The core of Style-Editor is our Patch-wise Co-Directional (PCD) loss, meticulously designed for precise object-centric editing that are closely aligned with the input text. This loss combines a patch directional loss for text-guided style direction and a patch distribution consistency loss for even CLIP embedding distribution across object regions. It ensures a seamless and harmonious style editing across object regions. Key to our method are the Text-Matched Patch Selection (TMPS) and Pre-fixed Region Selection (PRS) modules for identifying object locations via text, eliminating the need for segmentation masks. Lastly, we introduce an Adaptive Background Preservation (ABP) loss to maintain the original style and structural essence of the image’s background. This loss is applied to dynamically identified background areas. Extensive experiments underline the effectiveness of our approach in creating visually coherent and textually aligned style editing.
Jihun Park 0001, Jongmin Gim 0001, Kyoungmin Lee 0001, Seunghun Lee 0002, Sunghoon Im 0001
CVPR5
2025 JPEG Processing Neural Operator for Backward-Compatible Coding
abstract
Despite significant advances in learning-based lossy compression algorithms, standardizing codecs remains a critical challenge. In this paper, we present the JPEG Processing Neural Operator (JPNeO), a next-generation JPEG algorithm that maintains full backward compatibility with the current JPEG format. Our JPNeO improves chroma component preservation and enhances reconstruction fidelity compared to existing artifact removal methods by incorporating neural operators in both the encoding and decoding stages. JPNeO achieves practical benefits in terms of reduced memory usage and parameter count. We further validate our hypothesis about the existence of a space with high mutual information through empirical evidence. In summary, the JPNeO functions as a high-performance out-of-the-box image compression pipeline without changing source coding's protocol. Our source code is available at https://github.com/WooKyoungHan/JPNeO.
Woo Kyoung Han, Yongjun Lee, Byeonghun Lee, Sanghyun Park 0004, Sunghoon Im 0001, Kyong Hwan Jin
ICCV5
2025 LOMM: Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
abstract
In this paper, we introduce Latest Object Memory (LOM), a system for robustly tracking and continuously updating the latest states of objects by explicitly modeling their presence across video frames. LOM enables consistent tracking and accurate identity management across frames, enhancing both performance and reliability through the video segmentation process. Building upon LOM, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation, significantly improving long-term instance tracking. This enables consistent tracking and accurate identity management across frames, enhancing both performance and reliability through the video segmentation process. Moreover, we introduce Decoupled Object Association (DOA), a strategy that separately handles newly appearing and already existing objects. By leveraging our memory system, DOA accurately assigns object indices, improving matching accuracy and ensuring stable identity consistency, even in dynamic scenes where objects frequently appear and disappear. Extensive experiments and ablation studies demonstrate the superiority of our method over traditional approaches, setting a new state-of-the-art in video instance segmentation. Notably, our LOMM achieves an AP score of 54.0 on YouTube-VIS 2022, a dataset known for its challenging long videos. Project page: this https URL.
Seunghun Lee 0002, Jiwan Seo, Minwoo Choi, Kiljoon Han, Zane Durante, Ehsan Adeli-Mosabbeb, Sunghoon Im 0001
ICCV9
2025 CAVIS: Context-Aware Video Instance Segmentation
Seunghun Lee 0002, Jiwan Seo, Kiljoon Han, Minwoo Choi, Sunghoon Im 0001
ICCV5
2025 Self-supervised Monocular Depth Estimation Robust to Reflective Surface Leveraged by Triplet Mining
abstract
Self-supervised monocular depth estimation (SSMDE) aims to predict the dense depth map of a monocular image, by learning depth from RGB image sequences, eliminating the need for ground-truth depth labels. Although this approach simplifies data acquisition compared to supervised methods, it struggles with reflective surfaces, as they violate the assumptions of Lambertian reflectance, leading to inaccurate training on such surfaces. To tackle this problem, we propose a novel training strategy for an SSMDE by leveraging triplet mining to pinpoint reflective regions at the pixel level, guided by the camera geometry between different viewpoints. The proposed reflection-aware triplet mining loss specifically penalizes the inappropriate photometric error minimization on the localized reflective regions while preserving depth accuracy on non-reflective areas. We also incorporate a reflection-aware knowledge distillation method that enables a student model to selectively learn the pixel-level knowledge from reflective and non-reflective regions. This results in robust depth estimation across areas. Evaluation results on multiple datasets demonstrate that our method effectively enhances depth quality on reflective surfaces and outperforms state-of-the-art SSMDE baselines.
Wonhyeok Choi, Kyumin Hwang, Minwoo Choi, Sunghoon Im 0001
ICLR5
2024 Content-Adaptive Style Transfer: A Training-Free Approach with VQ Autoencoders
Jongmin Gim 0001, Jihun Park 0001, Kyoungmin Lee 0001, Sunghoon Im 0001
ACCV (5)4
2024 JDEC: JPEG Decoding via Enhanced Continuous Cosine Coefficients
abstract
We propose a practical approach to JPEG image de-coding, utilizing a local implicit neural representation with continuous cosine formulation. The JPEG algorithm sig-nificantly quantizes discrete cosine transform (DCT) spec-tra to achieve a high compression rate, inevitably resulting in quality degradation while encoding an image. We have designed a continuous cosine spectrum estimator to address the quality degradation issue that restores the distorted spectrum. By leveraging local DCT formulations, our network has the privilege to exploit dequantization and upsampling simultaneously. Our proposed model enables decoding compressed images directly across different quality factors using a single pre-trained model without relying on a conventional JPEG decoder. As a result, our proposed network achieves state-of-the-art performance in flexible color image JPEG artifact removal tasks. Our source code is available at https://github.com/WooKyoungHan/Jdec.
Woo Kyoung Han, Sunghoon Im 0001, Jaedeok Kim, Kyong Hwan Jin
CVPR2
2024 BurstM: Deep Burst Multi-scale SR Using Fourier Space with Optical Flow
EungGu Kang, Byeonghun Lee, Sunghoon Im 0001, Kyong Hwan Jin
ECCV (42)3
2024 Rethinking LiDAR Domain Generalization: Single Source as Multiple Density Domains
Jae-Yeul Kim, Jungwan Woo, Sunghoon Im 0001
ECCV (20)4
2024 Multi-task Learning for Real-time Autonomous Driving Leveraging Task-adaptive Attention Generator
abstract
Real-time processing is crucial in autonomous driving systems due to the imperative of instantaneous decision-making and rapid response. In real-world scenarios, autonomous vehicles are continuously tasked with interpreting their surroundings, analyzing intricate sensor data, and making decisions within split seconds to ensure safety through numerous computer vision tasks. In this paper, we present a new real-time multi-task network adept at three vital autonomous driving tasks: monocular 3D object detection, semantic segmentation, and dense depth estimation. To counter the challenge of negative transfer — the prevalent issue in multi-task learning — we introduce a task-adaptive attention generator. This generator is designed to automatically discern interrelations across the three tasks and arrange the task-sharing pattern, all while leveraging the efficiency of the hard-parameter sharing approach. To the best of our knowledge, the proposed model is pioneering in its capability to concurrently handle multiple tasks, notably 3D object detection, while maintaining real-time processing speeds. Our rigorously optimized network, when tested on the Cityscapes-3D datasets, consistently outperforms various base-line models. Moreover, an in-depth ablation study substantiates the efficacy of the methodologies integrated into our framework.
Wonhyeok Choi, Mingyu Shin, Hyukzae Lee, Jaehoon Cho, Jaehyeon Park, Sunghoon Im 0001
ICRA6
2024 Density-aware Domain Generalization for LiDAR Semantic Segmentation
abstract
3D LiDAR-based perception has made remarkable advancements, leading to the widespread adoption of LiDAR in autonomous driving systems. Despite these technological strides, variations in LiDAR sensors and environmental conditions can significantly deteriorate the performance of perception models, primarily due to changes in the density of point clouds. Recent studies in domain generalization have aimed to mitigate this challenge; however, they often rely on the availability of sequential data and ego-motion, which limits their applicability. To address these limitations, we propose two novel methods that enable network operation in a density-aware fashion without any constraints, thereby ensuring consistent performance despite fluctuations in point cloud density. First, we design the network to be density-aware by utilizing the kernel occupancy information from the 3D sparse convolution as geometric features. Subsequently, we further enhance density awareness by incorporating voxel-wise density prediction as an auxiliary task in a self-supervised manner. Our method demonstrates superior performance over current state-of-the-art approaches, achieving this without the need for specific data prerequisites. Our approach is compatible with a variety of 3D backbone architectures, enhancing domain generalization performance by 18.4% while adding a minimal computational overhead of only 7ms.
Jae-Yeul Kim, Jungwan Woo, Ukcheol Shin, Jean Oh, Sunghoon Im 0001
IROS5
2024 Offline-to-Online Knowledge Distillation for Video Instance Segmentation
abstract
In this paper, we present offline-to-online knowledge distillation (OOKD) for video instance segmentation (VIS), which transfers a wealth of video knowledge from an offline model to an online model for consistent prediction. Unlike previous methods that have adopted either an online or offline model, our single online model takes advantage of both models by distilling offline knowledge. To transfer knowledge correctly, we propose query filtering and association (QFA), which filters irrelevant queries to exact instances. Our KD with QFA increases the robustness of feature matching by encoding object-centric features from a single frame supplemented by long-range global information. We also propose a simple data augmentation scheme for knowledge distillation in the VIS task that fairly transfers the knowledge of all classes into the online model. Extensive experiments show that our method significantly improves the performance in video instance segmentation, especially for challenging datasets, including long, dynamic sequences. Our method also achieves state-of-the-art performance on YTVIS-21, YTVIS-22, and OVIS datasets, with mAP scores of 46.1%, 43.6%, and 31.1%, respectively.
Seunghun Lee 0002, Hyeon Kang, Sunghoon Im 0001
WACV4
2024 Implicit Neural Image Stitching With Enhanced and Blended Feature Reconstruction
abstract
Existing frameworks for image stitching often provide visually reasonable stitchings. However, they suffer from blurry artifacts and disparities in illumination, depth level, etc. Although the recent learning-based stitchings relax such disparities, the required methods impose sacrifice of image qualities failing to capture high-frequency details for stitched images. To address the problem, we propose a novel approach, implicit Neural Image Stitching (NIS) that extends arbitrary-scale super-resolution. Our method estimates Fourier coefficients of images for quality-enhancing warps. Then, the suggested model blends color mismatches and misalignment in the latent space and decodes the features into RGB values of stitched images. Our experiments show that our approach achieves improvement in resolving the low-definition imaging of the previous deep image stitching with favorable accelerated image-enhancing methods. Our source code is available at https://github.com/minshu-kim/NIS.
Byeonghun Lee, Sunghoon Im 0001, Kyong Hwan Jin
WACV4
2024 A Study on the Generality of Neural Network Structures for Monocular Depth Estimation
abstract
Monocular depth estimation has been widely studied, and significant improvements in performance have been recently reported. However, most previous works are evaluated on a few benchmark datasets, such as KITTI datasets, and none of the works provide an in-depth analysis of the generalization performance of monocular depth estimation. In this paper, we deeply investigate the various backbone networks (e.g.CNN and Transformer models) toward the generalization of monocular depth estimation. First, we evaluate state-of-the-art models on both in-distribution and out-of-distribution datasets, which have never been seen during network training. Then, we investigate the internal properties of the representations from the intermediate layers of CNN-/Transformer-based models using synthetic texture-shifted datasets. Through extensive experiments, we observe that the Transformers exhibit a strong shape-bias rather than CNNs, which have a strong texture-bias. We also discover that texture-biased models exhibit worse generalization performance for monocular depth estimation than shape-biased models. We demonstrate that similar aspects are observed in real-world driving datasets captured under diverse environments. Lastly, we conduct a dense ablation study with various backbone networks which are utilized in modern strategies. The experiments demonstrate that the intrinsic locality of the CNNs and the self-attention of the Transformers induce texture-bias and shape-bias, respectively.
Jinwoo Bae, Kyumin Hwang, Sunghoon Im 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Deep Digging into the Generalization of Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has been widely studied recently. Most of the work has focused on improving performance on benchmark datasets, such as KITTI, but has offered a few experiments on generalization performance. In this paper, we investigate the backbone networks (e.g., CNNs, Transformers, and CNN-Transformer hybrid models) toward the generalization of monocular depth estimation. We first evaluate state-of-the-art models on diverse public datasets, which have never been seen during the network training. Next, we investigate the effects of texture-biased and shape-biased representations using the various texture-shifted datasets that we generated. We observe that Transformers exhibit a strong shape bias and CNNs do a strong texture-bias. We also find that shape-biased models show better generalization performance for monocular depth estimation compared to texture-biased models. Based on these observations, we newly design a CNN-Transformer hybrid network with a multi-level adaptive feature fusion module, called MonoFormer. The design intuition behind MonoFormer is to increase shape bias by employing Transformers while compensating for the weak locality bias of Transformers by adaptively fusing multi-level representations. Extensive experiments show that the proposed method achieves state-of-the-art performance with various public datasets. Our method also shows the best generalization ability among the competitive methods.
Jinwoo Bae, Sungho Moon, Sunghoon Im 0001
AAAI3
2023 Multi-Target Domain Adaptation with Class-Wise Attribute Transfer in Semantic Segmentation
Changjae Kim, Seunghun Lee 0002, Sunghoon Im 0001
BMVC3
2023 Dynamic Neural Network for Multi-Task Learning Searching across Diverse Network Topologies
abstract
In this paper, we present a new MTL framework that searches for structures optimized for multiple tasks with diverse graph topologies and shares features among tasks. We design a restricted DAG-based central network with read-in/read-out layers to build topologically diverse task-adaptive structures while limiting search space and time. We search for a single optimized network that serves as multiple task adaptive sub-networks using our three-stage training process. To make the network compact and discretized, we propose a flow-based reduction algorithm and a squeeze loss used in the training process. We evaluate our optimized network on various public MTL datasets and show ours achieves state-of-the-art performance. An extensive ablation study experimentally validates the effectiveness of the sub-module and schemes in our framework.
Wonhyeok Choi, Sunghoon Im 0001
CVPR2
2023 Depth-discriminative Metric Learning for Monocular 3D Object Detection
abstract
Monocular 3D object detection poses a significant challenge due to the lack of depth information in RGB images. Many existing methods strive to enhance the object depth estimation performance by allocating additional parameters for object depth estimation, utilizing extra modules or data. In contrast, we introduce a novel metric learning scheme that encourages the model to extract depth-discriminative features regardless of the visual attributes without increasing inference time and model size. Our method employs the distance-preserving function to organize the feature space manifold in relation to ground-truth object depth. The proposed $(K,B,\epsilon)$-quasi-isometric loss leverages predetermined pairwise distance restriction as guidance for adjusting the distance among object descriptors without disrupting the non-linearity of the natural feature manifold. Moreover, we introduce an auxiliary head for object-wise depth estimation, which enhances depth quality while maintaining the inference time. The broad applicability of our method is demonstrated through experiments that show improvements in overall performance when integrated into various baselines. The results show that our method consistently improves the performance of various baselines by 23.51\% and 5.78\% on average across KITTI and Waymo, respectively.
Wonhyeok Choi, Mingyu Shin, Sunghoon Im 0001
NeurIPS3
2023 A Large-Scale Virtual Dataset and Egocentric Localization for Disaster Responses
abstract
With the increasing social demands of disaster response, methods of visual observation for rescue and safety have become increasingly important. However, because of the shortage of datasets for disaster scenarios, there has been little progress in computer vision and robotics in this field. With this in mind, we present the first large-scale synthetic dataset of egocentric viewpoints for disaster scenarios. We simulate pre- and post-disaster cases with drastic changes in appearance, such as buildings on fire and earthquakes. The dataset consists of more than 300K high-resolution stereo image pairs, all annotated with ground-truth data for the semantic label, depth in metric scale, optical flow with sub-pixel precision, and surface normal as well as their corresponding camera poses. To create realistic disaster scenes, we manually augment the effects with 3D models using physically-based graphics tools. We train various state-of-the-art methods to perform computer vision tasks using our dataset, evaluate how well these methods recognize the disaster situations, and produce reliable results of virtual scenes as well as real-world images. We also present a convolutional neural network-based egocentric localization method that is robust to drastic appearance changes, such as the texture changes in a fire, and layout changes from a collapse. To address these key challenges, we propose a new model that learns a shape-based representation by training on stylized images, and incorporate the dominant planes of query images as approximate scene coordinates. We evaluate the proposed method using various scenes including a simulated disaster dataset to demonstrate the effectiveness of our method when confronted with significant changes in scene layout. Experimental results show that our method provides reliable camera pose predictions despite vastly changed conditions.
Hae-Gon Jeon, Sunghoon Im 0001, Byeong-Uk Lee, François Rameau, Dong-Geol Choi, Jean Oh, In-So Kweon, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 ADAS: A Direct Adaptation Strategy for Multi-Target Domain Adaptive Semantic Segmentation
abstract
In this paper, we present a direct adaptation strategy (ADAS), which aims to directly adapt a single model to multiple target domains in a semantic segmentation task without pretrained domain-specific models. To do so, we design a multi-target domain transfer network (MTDT-Net) that aligns visual attributes across domains by transferring the domain distinctive features through a new target adaptive denormalization (TAD) module. Moreover, we propose a bi-directional adaptive region selection (BARS) that reduces the attribute ambiguity among the class labels by adaptively selecting the regions with consistent feature statistics. We show that our single MTDT-Net can synthesize visually pleasing domain transferred images with complex driving datasets, and BARS effectively filters out the unnecessary region of training images for each target domain. With the collaboration of MTDT-Net and BARS, our ADAS achieves state-of-the-art performance for multi-target domain adaptation (MTDA). To the best of our knowledge, our method is the first MTDA method that directly adapts to multiple domains in semantic segmentation.
Seunghun Lee 0002, Wonhyeok Choi, Changjae Kim, Minwoo Choi, Sunghoon Im 0001
CVPR5
2022 Facial Depth and Normal Estimation Using Single Dual-Pixel Camera
Minjun Kang, Jaesung Choe, Hyowon Ha, Hae-Gon Jeon, Sunghoon Im 0001, In-So Kweon, Kuk-Jin Yoon
ECCV (8)5
2022 CMSNet: Deep Color and Monochrome Stereo
Hae-Gon Jeon, Sunghoon Im 0001, Jaesung Choe, Minjun Kang, Joon-Young Lee, Martial Hebert
Int. J. Comput. Vis.2
2022 Self-Supervised Monocular Depth and Motion Learning in Dynamic Scenes: Semantic Prior to Rescue
Seokju Lee, François Rameau, Sunghoon Im 0001, In-So Kweon
Int. J. Comput. Vis.3
2022 ProFeat: Unsupervised image clustering via progressive feature refinement
abstract
Unsupervised image clustering is a chicken-and-egg problem that involves representation learning and clustering. To resolve the inter-dependency between them, many approaches that iteratively perform the two tasks have been proposed, but their accuracy is limited due to inaccurate intermediate representations and clusters. To overcome this, this paper proposes ProFeat, a novel iterative approach to unsupervised image clustering based on progressive feature refinement. To learn discriminative features for clustering while avoiding adversarial influence from inaccurate intermediate clusters, ProFeat rigorously divides representation learning and clustering by modeling a neural network for clustering as a composition of an embedding and a clustering function and introducing an auxiliary embedding function. ProFeat progressively refines representations using confident samples from intermediate clusters using an extended contrastive loss. This paper also proposes ensemble-based feature refinement for more robust clustering. Our experiments demonstrate that ProFeat achieves superior results compared to previous methods.
Sunghoon Im 0001, Sunghyun Cho
Pattern Recognit. Lett.2
2021 Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection Consistency
abstract
We present an end-to-end joint training framework that explicitly models 6-DoF motion of multiple dynamic objects, ego-motion, and depth in a monocular camera setup without supervision. Our technical contributions are three-fold. First, we highlight the fundamental difference between inverse and forward projection while modeling the individual motion of each rigid object, and propose a geometrically correct projection pipeline using a neural forward projection module. Second, we design a unified instance-aware photometric and geometric consistency loss that holistically imposes self-supervisory signals for every background and object region. Lastly, we introduce a general-purpose auto-annotation scheme using any off-the-shelf instance segmentation and optical flow models to produce video instance segmentation maps that will be utilized as input to our training pipeline. These proposed elements are validated in a detailed ablation study. Through extensive experiments conducted on the KITTI and Cityscapes dataset, our framework is shown to outperform the state-of-the-art depth and motion estimation methods. Our code, dataset, and models are publicly available.
Seokju Lee, Sunghoon Im 0001, Stephen Lin 0001, In-So Kweon
AAAI2
2021 ZeBRA: Precisely Destroying Neural Networks with Zero-Data Based Repeated Bit Flip Attack
Dahoon Park, Kon-Woo Kwon, Sunghoon Im 0001
BMVC3
2021 DRANet: Disentangling Representation and Adaptation Networks for Unsupervised Cross-Domain Adaptation
abstract
In this paper, we present DRANet, a network architecture that disentangles image representations and transfers the visual attributes in a latent space for unsupervised cross-domain adaptation. Unlike the existing domain adaptation methods that learn associated features sharing a domain, DRANet preserves the distinctiveness of each domain’s characteristics. Our model encodes individual representations of content (scene structure) and style (artistic appearance) from both source and target images. Then, it adapts the domain by incorporating the transferred style factor into the content factor along with learnable weights specified for each domain. This learning framework allows bi/multi-directional domain adaptation with a single encoder-decoder network and aligns their domain shift. Additionally, we propose a content-adaptive domain transfer module that helps retain scene structure while transferring style. Extensive experiments show our model successfully separates content-style factors and synthesizes visually pleasing domain-transferred images. The proposed method demonstrates state-of-the-art performance on standard digit classification tasks as well as semantic segmentation tasks.
Seunghun Lee 0002, Sunghyun Cho, Sunghoon Im 0001
CVPR3
2021 VolumeFusion: Deep Depth Fusion for 3D Scene Reconstruction
abstract
To reconstruct a 3D scene from a set of calibrated views, traditional multi-view stereo techniques rely on two distinct stages: local depth maps computation and global depth maps fusion. Recent studies concentrate on deep neural architectures for depth estimation by using conventional depth fusion method or direct 3D reconstruction network by regressing Truncated Signed Distance Function (TSDF). In this paper, we advocate that replicating the traditional two stages framework with deep neural networks improves both the interpretability and the accuracy of the results. As mentioned, our network operates in two steps: 1) the local computation of the local depth maps with a deep MVS technique, and, 2) the depth maps and images’ features fusion to build a single TSDF volume. In order to improve the matching performance between images acquired from very different viewpoints (e.g., large-baseline and rotations), we introduce a rotation-invariant 3D convolution kernel called PosedConv. The effectiveness of the proposed architecture is underlined via a large series of experiments conducted on the ScanNet dataset where our approach compares favorably against both traditional and deep learning techniques.
Jaesung Choe, Sunghoon Im 0001, François Rameau, Minjun Kang, In-So Kweon
ICCV2
2021 Deep Depth from Uncalibrated Small Motion Clip
abstract
We propose a novel approach to infer a high-quality depth map from a set of images with small viewpoint variations. In general, techniques for depth estimation from small motion consist of camera pose estimation and dense reconstruction. In contrast to prior approaches that recover scene geometry and camera motions using pre-calibrated cameras, we introduce in this paper a self-calibrating bundle adjustment method tailored for small motion which enables computation of camera poses without the need for camera calibration. For dense depth reconstruction, we present a convolutional neural network called DPSNet (Deep Plane Sweep Network) whose design is inspired by best practices of traditional geometry-based approaches. Rather than directly estimating depth or optical flow correspondence from image pairs as done in many previous deep learning methods, DPSNet takes a plane sweep approach that involves building a cost volume from deep features using the plane sweep algorithm, regularizing the cost volume, and regressing the depth map from the cost volume. The cost volume is constructed using a differentiable warping process that allows for end-to-end training of the network. Through the effective incorporation of conventional multiview stereo concepts within a deep learning framework, the proposed method achieves state-of-the-art results on a variety of challenging datasets.
Sunghoon Im 0001, Hyowon Ha, Hae-Gon Jeon, Stephen Lin 0001, In-So Kweon
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Learning Shape-based Representation for Visual Localization in Extremely Changing Conditions
abstract
Visual localization is an important task for applications such as navigation and augmented reality, but is a challenging problem when there are changes in scene appearances through day, seasons, or environments. In this paper, we present a convolutional neural network (CNN)-based approach for visual localization across normal to drastic appearance variations such as pre- and post-disaster cases. Our approach aims to address two key challenges: (1) to reduce the biases based on scene textures as in traditional CNNs, our model learns a shape-based representation by training on stylized images; (2) to make the model robust against layout changes, our approach uses the estimated dominant planes of query images as approximate scene coordinates. Our method is evaluated on various scenes including a simulated disaster dataset to demonstrate the effectiveness of our method in significant changes of scene layout. Experimental results show that our method provides reliable camera pose predictions in various changing conditions.
Hae-Gon Jeon, Sunghoon Im 0001, Jean Oh, Martial Hebert
ICRA2
2020 Ring Difference Filter for Fast and Noise Robust Depth From Focus
abstract
Depth from focus (DfF) is a method of estimating the depth of a scene by using information acquired through changes in the focus of a camera. Within the DfF framework of, the focus measure (FM) forms the foundation which determines the accuracy of the output. With the results from the FM, the role of a DfF pipeline is to determine and recalculate unreliable measurements while enhancing those that are reliable. In this paper, we propose a new FM, which we call the "ring difference filter" (RDF), that can more accurately and robustly measure focus. FMs can usually be categorized as confident local methods or noise robust non-local methods. The RDF's unique ring-and-disk structure allows it to have the advantages of both local and non-local FMs. We then describe an efficient pipeline that utilizes the RDF's properties. Part of this pipeline is our proposed RDF-based cost aggregation method, which is able to robustly refine the initial results in the presence of image noise. Our method is able to reproduce results that are on par with or even better than those of state-of-the-art methods, while spending less time in computation.
Hae-Gon Jeon, Jaeheung Surh, Sunghoon Im 0001, In-So Kweon
IEEE Trans. Image Process.3
2019 DPSNet: End-to-end Deep Plane Sweep Stereo
Sunghoon Im 0001, Hae-Gon Jeon, Stephen Lin 0001, In-So Kweon
ICLR (Poster)1
2019 Depth Completion with Deep Geometry and Context Guidance
abstract
In this paper, we present an end-to-end convolutional neural network (CNN) for depth completion. Our network consists of a geometry network and a context network. The geometry network, a single encoder-decoder network, learns to optimize a multi-task loss to generate an initial propagated depth map and a surface normal. The complementary outputs allow it to correctly propagate initial sparse depth points in slanted surfaces. The context network extracts a local and a global feature of an image to compute a bilateral weight, which enables it to preserve edges and fine details in the depth maps. At the end, a final output is produced by multiplying the initially propagated depth map with the bilateral weight. In order to validate the effectiveness and the robustness of our network, we performed extensive ablation studies and compared the results against state-of-the-art CNN-based depth completions, where we showed promising results on various scenes.
Byeong-Uk Lee, Hae-Gon Jeon, Sunghoon Im 0001, In-So Kweon
ICRA3
2019 DISC: A Large-scale Virtual Dataset for Simulating Disaster Scenarios
abstract
In this paper, we present the first large-scale synthetic dataset for visual perception in disaster scenarios, and analyze state-of-the-art methods for multiple computer vision tasks with reference baselines. We simulated before and after disaster scenarios such as fire and building collapse for fifteen different locations in realistic virtual worlds. The dataset consists of more than 300K high-resolution stereo image pairs, all annotated with ground-truth data for semantic segmentation, depth, optical flow, surface normal estimation and camera pose estimation. To create realistic disaster scenes, we manually augmented the effects with 3D models using physical-based graphics tools. We use our dataset to train state-of-the-art methods and evaluate how well these methods can recognize the disaster situations and produce reliable results on virtual scenes as well as real-world images. The results obtained from each task are then used as inputs to the proposed visual odometry network for generating 3D maps of buildings on fire. Finally, we discuss challenges for future research.
Hae-Gon Jeon, Sunghoon Im 0001, Byeong-Uk Lee, Dong-Geol Choi, Martial Hebert, In-So Kweon
IROS2
2019 Learning Residual Flow as Dynamic Motion from Stereo Videos
abstract
We present a method for decomposing the 3D scene flow observed from a moving stereo rig into stationary scene elements and dynamic object motion. Our unsupervised learning framework jointly reasons about the camera motion, optical flow, and 3D motion of moving objects. Three cooperating networks predict stereo matching, camera motion, and residual flow, which represents the flow component due to object motion and not from camera motion. Based on rigid projective geometry, the estimated stereo depth is used to guide the camera motion estimation, and the depth and camera motion are used to guide the residual flow estimation. We also explicitly estimate the 3D scene flow of dynamic objects based on the residual flow and scene depth. Experiments on the KITTI dataset demonstrate the effectiveness of our approach and show that our method outperforms other state-of-the-art algorithms on the optical flow and visual odometry tasks.
Seokju Lee, Sunghoon Im 0001, Stephen Lin 0001, In-So Kweon
IROS2
2019 Accurate 3D Reconstruction from Small Motion Clip for Rolling Shutter Cameras
abstract
Structure from small motion has become an important topic in 3D computer vision as a method for estimating depth, since capturing the input is so user-friendly. However, major limitations exist with respect to the form of depth uncertainty, due to the narrow baseline and the rolling shutter effect. In this paper, we present a dense 3D reconstruction method from small motion clips using commercial hand-held cameras, which typically cause the undesired rolling shutter artifact. To address these problems, we introduce a novel small motion bundle adjustment that effectively compensates for the rolling shutter effect. Moreover, we propose a pipeline for a fine-scale dense 3D reconstruction that models the rolling shutter effect by utilizing both sparse 3D points and the camera trajectory from narrow-baseline images. In this reconstruction, the sparse 3D points are propagated to obtain an initial depth hypothesis using a geometry guidance term. Then, the depth information on each pixel is obtained by sweeping the plane around each depth search space near the hypothesis. The proposed framework shows accurate dense reconstruction results suitable for various sought-after applications. Both qualitative and quantitative evaluations show that our method consistently generates better depth maps compared to state-of-the-art methods.
Sunghoon Im 0001, Hyowon Ha, Gyeongmin Choe, Hae-Gon Jeon, Kyungdon Joo, In-So Kweon
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Robust Depth Estimation Using Auto-Exposure Bracketing
abstract
As the computing power of hand-held devices grows, there has been increasing interest in the capture of depth information, to enable a variety of photographic applications. However, under low-light conditions, most devices still suffer from low imaging quality and inaccurate depth acquisition. To address the problem, we present a robust depth estimation method from a short burst shot with varied intensity (i.e., Auto-exposure bracketing) and/or strong noise (i.e., High ISO). Our key idea synergistically combines deep convolutional neural networks with geometric understanding of the scene. We introduce a geometric transformation between optical flow and depth tailored for burst images, enabling our learning-based multi-view stereo matching to be performed effectively. We then describe our depth estimation pipeline that incorporates this geometric transformation into our residual-flow network. It allows our framework to produce an accurate depth map even with a bracketed image sequence. We demonstrate that our method outperforms state-of-the-art methods for various datasets captured by a smartphone and a DSLR camera. Moreover, we show that the estimated depth is applicable for image quality enhancement and photographic editing.
Sunghoon Im 0001, Hae-Gon Jeon, In-So Kweon
IEEE Trans. Image Process.1
2018 Robust Depth Estimation From Auto Bracketed Images
abstract
As demand for advanced photographic applications on hand-held devices grows, these electronics require the capture of high quality depth. However, under low-light conditions, most devices still suffer from low imaging quality and inaccurate depth acquisition. To address the problem, we present a robust depth estimation method from a short burst shot with varied intensity (i.e., Auto Bracketing) or strong noise (i.e., High ISO). We introduce a geometric transformation between flow and depth tailored for burst images, enabling our learning-based multi-view stereo matching to be performed effectively. We then describe our depth estimation pipeline that incorporates the geometric transformation into our residual-flow network. It allows our framework to produce an accurate depth map even with a bracketed image sequence. We demonstrate that our method outperforms state-of-the-art methods for various datasets captured by a smartphone and a DSLR camera. Moreover, we show that the estimated depth is applicable for image quality enhancement and photographic editing.
Sunghoon Im 0001, Hae-Gon Jeon, In-So Kweon
CVPR1
2017 Noise Robust Depth from Focus Using a Ring Difference Filter
abstract
Depth from focus (DfF) is a method of estimating depth of a scene by using the information acquired through the change of the focus of a camera. Within the framework of DfF, the focus measure (FM) forms the foundation on which the accuracy of the output is determined. With the result from the FM, the role of a DfF pipeline is to determine and recalculate unreliable measurements while enhancing those that are reliable. In this paper, we propose a new FM that more accurately and robustly measures focus, which we call the ring difference filter (RDF). FMs can usually be categorized as confident local methods or noise robust non-local methods. RDFs unique ring-and-disk structure allows it to have the advantageous sides of both local and non-local FMs. We then describe an efficient pipeline that utilizes the properties that the RDF brings. Our method is able to reproduce results that are on par with or even better than those of the state-of-the-art, while spending less time in computation.
Jaeheung Surh, Hae-Gon Jeon, Yunwon Park, Sunghoon Im 0001, Hyowon Ha, In-So Kweon
CVPR4
2017 Geometry Guided Three-Dimensional Propagation for Depth From Small Motion
abstract
In this letter, we present an accurate Depth from Small Motion approach, which reconstructs three-dimensional (3-D) depth from image sequences with extremely narrow baselines. We start with estimating sparse 3-D points and camera poses via the structure from motion method. For dense depth reconstruction, we propose a novel depth propagation using a geometric guidance term that considers not only the geometric constraint from the surface normal, but also color consistency. In addition, we propose an accurate surface normal estimation method with a multiple range search so that the normal vector can guide the direction of the depth propagation precisely. The major benefit of our depth propagation method is that it obtains detailed structures of a scene without fronto-parallel bias. We validate our method using various indoor and outdoor datasets, and both qualitative and quantitative experimental results show that our new algorithm consistently generates better 3-D depth information than the results of existing state-of-the-art methods.
Seunghak Shin, Sunghoon Im 0001, Inwook Shim, Hae-Gon Jeon, In-So Kweon
IEEE Signal Process. Lett.2
2016 High-Quality Depth from Uncalibrated Small Motion Clip
abstract
We propose a novel approach that generates a highquality depth map from a set of images captured with a small viewpoint variation, namely small motion clip. As opposed to prior methods that recover scene geometry and camera motions using pre-calibrated cameras, we introduce a self-calibrating bundle adjustment tailored for small motion. This allows our dense stereo algorithm to produce a high-quality depth map for the user without the need for camera calibration. In the dense matching, the distributions of intensity profiles are analyzed to leverage the benefit of having negligible intensity changes within the scene due to the minuscule variation in viewpoint. The depth maps obtained by the proposed framework show accurate and extremely fine structures that are unmatched by previous literature under the same small motion configuration.
Hyowon Ha, Sunghoon Im 0001, Jaesik Park, Hae-Gon Jeon, In-So Kweon
CVPR2
2016 Stereo Matching with Color and Monochrome Cameras in Low-Light Conditions
abstract
Consumer devices with stereo cameras have become popular because of their low-cost depth sensing capability. However, those systems usually suffer from low imaging quality and inaccurate depth acquisition under low-light conditions. To address the problem, we present a new stereo matching method with a color and monochrome camera pair. We focus on the fundamental trade-off that monochrome cameras have much better light-efficiency than color-filtered cameras. Our key ideas involve compensating for the radiometric difference between two cross-spectral images and taking full advantage of complementary data. Consequently, our method produces both an accurate depth map and high-quality images, which are applicable for various depth-aware image processing. Our method is evaluated using various datasets and the performance of our depth estimation consistently outperforms state-of-the-art methods.
Hae-Gon Jeon, Joon-Young Lee, Sunghoon Im 0001, Hyowon Ha, In-So Kweon
CVPR3
2016 All-Around Depth from Small Motion with a Spherical Panoramic Camera
Sunghoon Im 0001, Hyowon Ha, François Rameau, Hae-Gon Jeon, Gyeongmin Choe, In-So Kweon
ECCV (3)1
2015 High Quality Structure from Small Motion for Rolling Shutter Cameras
abstract
We present a practical 3D reconstruction method to obtain a high-quality dense depth map from narrow-baseline image sequences captured by commercial digital cameras, such as DSLRs or mobile phones. Depth estimation from small motion has gained interest as a means of various photographic editing, but important limitations present themselves in the form of depth uncertainty due to a narrow baseline and rolling shutter. To address these problems, we introduce a novel 3D reconstruction method from narrow-baseline image sequences that effectively handles the effects of a rolling shutter that occur from most of commercial digital cameras. Additionally, we present a depth propagation method to fill in the holes associated with the unknown pixels based on our novel geometric guidance model. Both qualitative and quantitative experimental results show that our new algorithm consistently generates better 3D depth maps than those by the state-of-the-art method.
Sunghoon Im 0001, Hyowon Ha, Gyeongmin Choe, Hae-Gon Jeon, Kyungdon Joo, In-So Kweon
ICCV1
2015 Depth from accidental motion using geometry prior
abstract
We present a method to reconstruct dense 3D points from small camera motion. We begin with estimating sparse 3D points and camera poses by Structure from Motion (SfM) method with homography decomposition. Although the estimated points are optimized via bundle adjustment and gives reliable accuracy, the reconstructed points are sparse because it heavily depends on the extracted features of a scene. To handle this, we propose a depth propagation method using both a color prior from the images and a geometry prior from the initial points. The major benefit of our method is that we can easily handle the regions with similar colors but different depths by using the surface normal estimated from the initial points. We design our depth propagation framework into the cost minimization process. The cost function is linearly designed, which makes our optimization tractable. We demonstrate the effectiveness of our approach by comparing with a conventional method using various real-world examples.
Sunghoon Im 0001, Gyeongmin Choe, Hae-Gon Jeon, In-So Kweon
ICIP1