EDBT 2026 Demo / reviewers in the wild / expert
Donghyeon Cho
dblp:142/2590
· DBLP profile ↗
46ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-2184-921XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 32 · 7 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GPGS: Consistent 3D Object Removal via Geometry-Aware 3D Inpainting and Projected Image Refinement in 3D Gaussian SplattingabstractObject removal in 3D space is a key technology for immersive applications such as virtual reality (VR), augmented reality (AR), and the metaverse. While recent approaches have attempted to address this task using 2D inpainting models, they often suffer from two major limitations: (1) inaccurate geometric restoration in the removed regions, and (2) visual inconsistency across multiple viewpoints. To address these challenges, we propose GPGS, a novel pipeline built upon the 3D Gaussian Splatting (3DGS) framework. First, we perform geometry-aware 3D inpainting by leveraging a pre-trained point cloud completion model and a coarse-to-fine inference strategy, enabling accurate restoration of unseen 3D structures. Next, we introduce a projected image refinement method that improves the appearance of novel-view projections by addressing view-dependent artifacts such as brightness shifts and texture misalignments. GPGS further enhances overall scene consistency through fine-tuning of the original 3DGS scene using the refined multi-view images. Experimental results show that our GPGS makes geometrically accurate and visually coherent outputs, even in challenging 360° panoramic scenes, significantly outperforming existing methods. Yongjoon Lee, Donghyeon Cho |
AAAI | 2 |
| 2026 | CURE: Controllable Unified Image Restoration for Complex Degradations
Boseong Kim, Donghyeon Cho |
ICPR (1) | 2 |
| 2026 | Consistent Scene Understanding in 3D Gaussian Splatting via Multi-cue Mask Refinement
Hyunjoon Park, Donghyeon Cho |
ICPR (9) | 2 |
| 2026 | Preserving instance-level characteristics for multi-instance generation
Jaehak Ryu, Sungwon Moon, Donghyeon Cho |
Image Vis. Comput. | 3 |
| 2025 | Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
Joowon Kim, Ziseok Lee, Donghyeon Cho, Sanghyun Jo, Yeonsung Jung, Eunho Yang |
ICCV | 3 |
| 2025 | Progressive Artwork Outpainting Via Latent Diffusion Models
Dae-Young Song, Jung-Jae Yu, Donghyeon Cho |
ICCV | 3 |
| 2025 | MELON: Learning Multi-Aspect Modality Preferences for Accurate Multimedia RecommendationabstractExisting multimedia recommender systems have made the best efforts to predict user preferences for items by utilizing behavioral similarities between users and the modality features of items a user has interacted with. However, we identify two key limitations in existing methods regarding preferences for modality features: (L1) although preferences for modality features is an important aspect of users' preferences, existing methods only leverage neighbors with similar interactions and do not consider the neighbors who may have similar preferences for modality features while having different interactions; (L2) although modality features of a user and an item may have a complex geometric relationship in the latent space, existing methods overlook and face challenges in precisely capturing this relationship. To address these two limitations, we propose a novel multimedia recommendation framework, named MELON, which is based on two core ideas: (Idea 1) Modality-cEntered embedding extraction; (Idea 2) reLatiOnship-ceNtered embedding extraction. We validate the effectiveness and validity of MELON through extensive experiments with four real-world datasets, showing 10.51% higher accuracy compared to the best competitor in terms of recall@10. The code and dataset of MELON is available at https://github.com/Bigdasgit/MELON. Dongho Jeong, Taeri Kim 0001, Donghyeon Cho, Sang-Wook Kim |
SIGIR | 3 |
| 2024 | CORE-MPI: Consistency Object Removal with Embedding MultiPlane ImageabstractNovel view synthesis is attractive for social media, but it often contains unwanted details such as personal information that needs to be edited out for a better experience. Multiplane image (MPI) is desirable for social media because of its generality but it is complex and computationally expensive, making object removal challenging. To address these challenges, we propose CORE-MPI, which employs embedding images to improve the consistency and accessibility of MPI object removal. CORE-MPI allows for real-time transmission and interaction with embedding images on social media, facilitating object removal with a single mask. However, recovering the geometric information hidden in the embedding images is a significant challenge. Therefore, we propose a dual-network approach, where one network focuses on color restoration and the other on inpainting the embedding image including geometric information. For the training of CORE-MPI, we introduce a pseudo-reference loss aimed at proficient color recovery, even in complex scenes or with large masks. Furthermore, we present a disparity consistency loss to preserve the geometric consistency of the inpainted region. We demonstrate the effectiveness of CORE-MPI on RealEstate10K and UCSD datasets. Donggeun Yoon, Donghyeon Cho |
CVPR | 2 |
| 2024 | Probabilistic Weather Forecasting with Deterministic Guidance-Based Diffusion Model
Donggeun Yoon, Doyi Kim, Yeji Choi, Donghyeon Cho |
ECCV (30) | 5 |
| 2024 | CSSR: Cross-and Self-feature Transformer with High-Frequency Feature Alignment for Reference-Based Super-Resolution
Seonggwan Ko, Donghyeon Cho |
ICPR (21) | 2 |
| 2024 | Consistent Object Removal from Masked Neural Radiance Fields by Estimating Never-Seen Regions in All-Views
Yongjoon Lee, Jaehak Ryu, Donggeun Yoon, Donghyeon Cho |
ICPR (22) | 4 |
| 2024 | Reference-based Burst Super-resolution
Seonggwan Ko, Yeong Jun Koh, Donghyeon Cho |
ACM Multimedia | 3 |
| 2023 | DIFu: Depth-Guided Implicit Function for Clothed Human ReconstructionabstractRecently, implicit function (IF)-based methods for clothed human reconstruction using a single image have received a lot of attention. Most existing methods rely on a 3D embedding branch using volume such as the skinned multi-person linear (SMPL) model, to compensate for the lack of information in a single image. Beyond the SMPL, which provides skinned parametric human 3D information, in this paper, we propose a new IF-based method, DIFu, that utilizes a projected depth prior containing textured and non-parametric human 3D information. In particular, DIFu consists of a generator, an occupancy prediction network, and a texture prediction network. The generator takes an RGB image of the human front-side as input, and hallucinates the human back-side image. After that, depth maps for front/back images are estimated and projected into 3D volume space. Finally, the occupancy prediction network extracts a pixel-aligned feature and a voxel-aligned feature through a 2D encoder and a 3D encoder, respectively, and estimates occupancy using these features. Note that voxel-aligned features are obtained from the projected depth maps, thus it can contain detailed 3D information such as hair and cloths. Also, colors of each query point are also estimated with the texture inference branch. The effectiveness of DIFu is demonstrated by comparing to recent IF-based models quantitatively and qualitatively. Dae-Young Song, HeeKyung Lee, Jeongil Seo, Donghyeon Cho |
CVPR | 4 |
| 2023 | IFQA: Interpretable Face Quality AssessmentabstractExisting face restoration models have relied on general assessment metrics that do not consider the characteristics of facial regions. Recent works have therefore assessed their methods using human studies, which is not scalable and involves significant effort. This paper proposes a novel face-centric metric based on an adversarial framework where a generator simulates face restoration and a discriminator assesses image quality. Specifically, our per-pixel discriminator enables interpretable evaluation that cannot be provided by traditional metrics. Moreover, our metric emphasizes facial primary regions considering that even minor changes to the eyes, nose, and mouth significantly affect human cognition. Our face-oriented metric consistently surpasses existing general or facial image quality assessment metrics by impressive margins. We demonstrate the generalizability of the proposed strategy in various architectural designs and challenging scenarios. Interestingly, we find that our IFQA can lead to performance improvement as an objective function. The code and models are available at https://github.com/VCLLab/IFQA. Byungho Jo, Donghyeon Cho, In Kyu Park, Sungeun Hong |
WACV | 2 |
| 2023 | Cloth-Changing Person Re-Identification With Noisy Patch FilteringabstractCloth-changing person re-identification (ReID) aims to recognize the identities of persons even when changing cloth. Even with the same person, changes in clothing can cause significant visual variations, making it challenging to identify them. Therefore, there is a need for a technique that can extract their inherent characteristics while being invariant to changes in clothing. Recent studies have utilized additional information such as gait, body parsing maps, and 3D shape to address cloth variance. However, these methods require additional processing and their effectiveness is dependent on the quality of the information provided. In this letter, we propose a two-stream model for ReID based on a single RGB image, consisting of a baseline ReID network and cloth-unrelated ReID networks. The baseline network takes the full RGB image as input, while the cloth-unrelated networks only receive cropped patches of the RGB image that include only cloth-independent identity cues such as the face and legs. We also propose a noisy patch filtering module (NPFM) to remove interference from noise patches during training. Finally, we combine features from baseline and cloth-unrelated ReID networks to perform cloth-changing person ReID. Our approach achieves a significant improvement of$+7.4\%$for Rank-1 and$+1.7\%$for mAP in cloth-changing environments compared to the recent state-of-the-art method. The effectiveness of our model for the cloth-changing person ReID is verified on the benchmark dataset with various ablation studies. In summary, our method can extract the unique features of an individual person in a cloth-changing environment and is robustly trained even with occluded or poorly captured data. Therefore, it can be more versatile in practical environments than existing ReID methods. Hyeok-Joon Kweon, Donghyeon Cho |
IEEE Signal Process. Lett. | 2 |
| 2022 | Lightweight Alpha Matting Network Using Distillation-Based Channel Pruning
Donggeun Yoon, Jinsun Park, Donghyeon Cho |
ACCV (3) | 3 |
| 2022 | Weakly-Supervised Stitching Network for Real-World Panoramic Image Generation
Dae-Young Song, Geonsoo Lee, Heekyung Lee, Gi-Mun Um, Donghyeon Cho |
ECCV (16) | 5 |
| 2022 | Channel Sampler in Hyperspectral Images for Vehicle DetectionabstractSince hyperspectral images (HSIs) contain visual information of multiple wavelengths, invisible signals to human eyes can also be detected. Therefore, it can be widely used for target object detection in bad weather and disaster environments. However, the channel dimension of the HSI is very large, and thus it is very inefficient to apply the existing object detector naively. In this letter, we present a lightweight convolutional neural network (CNN)-based channel sampler to estimate the importance score of each channel in the HSI. Based on the importance score of each channel, we can generate single-channel images that achieve the best object detection performance, as well as analyze the impact of the wavelength in the HSI on object detection performance. The proposed sampler is trained by a self-supervised adversarial learning method that recovers the original input HSI from the generated single-channel image. Therefore, our channel sampler can be seamlessly combined with any existing detectors. For experiments, we build a hyperspectral dataset for vehicle detection and then show the effectiveness of our method through various ablation studies. Geonsoo Lee, Jaekyu Lee, Jeonghyun Baek, Hoseong Kim, Donghyeon Cho |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Learning Lightweight Low-Light Enhancement Network Using Pseudo Well-Exposed ImagesabstractRecently, there has been growing attention on deep learning-based low-light image enhancement algorithms. With this interest, various synthetic low-light image datasets have been released publicly. However, real-world low-light and well-exposed image pair datasets are still lacking. In this paper, we propose a real-world low-light image dataset and a practical lightweight low-light image enhancement network. In order to construct a large-scale real-world low-light dataset, we have not only captured under-exposed images by ourselves but also collected under-exposed images from the Internet. Then, we produce pseudo well-exposed images for each low-light image. Using pairs of a real-world low-light image and a pseudo well-exposed image, we present a lightweight deep CNN model through knowledge distillation. Experimental results demonstrate the effectiveness and practicality of the proposed method on various datasets. Seonggwan Ko, Jinsun Park, Byungjoo Chae, Donghyeon Cho |
IEEE Signal Process. Lett. | 4 |
| 2021 | Restore From Restored: Video Restoration With Pseudo Clean VideoabstractIn this study, we propose a self-supervised video denoising method called "restore-from-restored." This method fine-tunes a pre-trained network by using a pseudo clean video during the test phase. The pseudo clean video is obtained by applying a noisy video to the baseline network. By adopting a fully convolutional neural network (FCN) as the baseline, we can improve video denoising performance without accurate optical flow estimation and registration steps, in contrast to many conventional video restoration methods, due to the translation equivariant property of the FCN. Specifically, the proposed method can take advantage of plentiful similar patches existing across multiple consecutive frames (i.e., patch-recurrence); these patches can boost the performance of the baseline network by a large margin. We analyze the restoration performance of the fine-tuned video denoising networks with the proposed self-supervision-based learning algorithm, and demonstrate that the FCN can utilize recurring patches without requiring accurate registration among adjacent frames. In our experiments, we apply the proposed method to state-of-the-art denoisers and show that our fine-tuned networks achieve a considerable improvement in denoising performance. Donghyeon Cho, Tae Hyun Kim 0006 |
CVPR | 2 |
| 2021 | OCR-based Inventory Management Algorithms Robust to Damaged ImagesabstractAccurate and fast inventory management algorithms are essential in the modern distribution industry. However, the configuration process of inventory management algorithms is very expensive, and the direct comprehensive management of inventory procedures is labor intensive and inaccurate. Therefore, in this paper, we propose an optical character recognition (OCR)-based inventory management algorithm to resolve these practical issues. The main purpose of our inventory management algorithm is to automatically inspect whether a list of items and the actual items match. To this end, our method consists of three steps, namely, text detection, text recognition, and text matching. In addition, to expand our algorithm to real-world applications, we propose adversarial training to ensure robustness against various damaged images, including corruption, blur, and inappropriate viewpoints. To train the network, we construct a new inventory management dataset (IMD) consisting of 10,000 sheets in real retail store environments. We verify our algorithm on the public dataset and our new IMD. As a result, we experimentally demonstrate that our method is not only robust against various damaged images but is also easily applicable in both large and small scale distributions stores at a low cost. Our code and dataset is available at https://blog.airlab.re.kr/Deform-and-Recover/. Daehan Kim, Hyeyoon Kang, Donghyeon Cho, Dong-Geol Choi |
ICRA | 4 |
| 2021 | Self-Supervised Feature Enhancement Networks for Small Object Detection in Noisy ImagesabstractRecent CNN-based approaches have shown impressive improvements in object detection, but detecting small objects in images is still a challenging task. Small object detection becomes more difficult if the image contains a lot of noise, which is frequent in real environments. The main reason is that the ratio of visual signal to noise on small objects is very low, making it difficult to extract rich features for detection. To address this issue, we propose a feature enhancement network (FEN) that is trained in a self-supervised manner. Specifically, FEN takes features from input images whose values randomly were erased, then predicts the erased values by aggregating neighboring values. This scheme enables FEN to improve features using surrounding values, which have great effects on enriching features from small-object regions during the test phase. To verify the robustness of our method against small object detection from noisy images, we adopt vehicle detection in aerial images as the main target task. The proposed method consistently outperformed the baseline methods in our experiments. We further present a variety of empirical studies, quantitatively and qualitatively, for in-depth analysis. Geonsoo Lee, Sungeun Hong, Donghyeon Cho |
IEEE Signal Process. Lett. | 3 |
| 2021 | End-to-End Image Stitching Network via Multi-Homography EstimationabstractIn this letter, we propose an end-to-end stitching network, which takes two images with a narrow field of view (FOV) as inputs, and produces a single image with a wide FOV. Our method estimates multiple homographies to cover the depth differences in the scene and is therefore robust against parallax distortion. In particular, global warping maps are generated using estimated multiple homographies and adjusted by local displacement maps. The final result is made by warping input images multiple times using the warping maps and then merging warped images with the weight maps. Multiple homographies, local displacement maps, and weight maps are generated simultaneously by our stitching network. To train the stitching network, we construct a dataset using the CARLA simulator. Then, using this dataset, our network is trained by end-to-end supervised learning based on appearance matching loss and depth layer loss. In experiments, we show that our method is superior to existing methods both qualitatively and quantitatively. Also, we provide various empirical studies for in-depth analysis as well as the result of the expansion to 360°panoramas. Dae-Young Song, Gi-Mun Um, Heekyung Lee, Donghyeon Cho |
IEEE Signal Process. Lett. | 4 |
| 2020 | Global-and-Local Relative Position Embedding for Unsupervised Video Summarization
Yunjae Jung, Donghyeon Cho, Sanghyun Woo, In-So Kweon |
ECCV (25) | 2 |
| 2020 | Fast Adaptation to Super-Resolution Networks via Meta-learning
Seobin Park, Jinsu Yoo, Donghyeon Cho, Tae Hyun Kim 0006 |
ECCV (27) | 3 |
| 2020 | Rank order coding based spiking convolutional neural network architecture with energy-efficient membrane voltage updates
Hoyoung Tang, Donghyeon Cho, Dongwoo Lew, Jongsun Park 0001 |
Neurocomputing | 2 |
| 2020 | Semi-Calibrated Photometric StereoabstractWhile conventional calibrated photometric stereo methods assume that light intensities and sensor exposures are known or unknown but identical across observed images, this assumption easily breaks down in practical settings due to individual light bulb's characteristics and limited control over sensors. This paper studies the effect of unknown and possibly non-uniform light intensities and sensor exposures among observed images on the shape recovery based on photometric stereo. This leads to the development of a "semi-calibrated" photometric stereo method, where the light directions are known but light intensities (and sensor exposures) are unknown. We show that the semi-calibrated photometric stereo becomes a bilinear problem, whose general form is difficult to solve, but in the photometric stereo context, there exists a unique solution for the surface normal and light intensities (or sensor exposures). We further show that there exists a linear solution method for the problem, and develop efficient and stable solution methods. The semi-calibrated photometric stereo is advantageous over conventional calibrated photometric stereo in accurate determination of surface normal, because it relaxes the assumption of known light intensity ratios/sensor exposures. The experimental results show superior accuracy of the semi-calibrated photometric stereo in comparison to conventional methods in practical settings. Donghyeon Cho, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Towards Privacy-Preserving Domain AdaptationabstractThis study suggests a new domain adaptation paradigm that can address potential data-privacy issues. Despite promising results of existing domain adaptation methods, they have a strong constraint where the source and target samples are accessible during a training phase. However, direct usage of source samples possibly causes data-privacy issues especially when each label of source domain acts as an individual's identifier such as biometric information. To address data-privacy problems in conventional domain adaptation, we propose privacy-preserving domain adaptation (PPDA). Our main hypothesis is that if we train our target model initialized from a pre-trained source model in a self-learning manner, we can successfully transfer knowledge from a labeled source domain to an unlabeled target domain. In our preliminary study, we observe that target samples with low self-entropy measured from the pre-trained source model achieves sufficiently high accuracy. From this key observation, we first select the reliable samples based on self-entropy and define them as class prototypes. We then assign pseudo labels to the target samples through the similarity between target samples and class prototypes. To further reduce the uncertainty of the pseudo labeling process, we also introduce a sample-level reweighting scheme. Surprisingly, our PPDA model outperforms conventional domain adaptation methods on public datasets even though we do not directly access any source data. Youngeun Kim, Donghyeon Cho, Sungeun Hong |
IEEE Signal Process. Lett. | 2 |
| 2020 | Lightweight Deep CNN for Natural Image Matting via Similarity-Preserving Knowledge DistillationabstractRecently, alpha matting has witnessed remarkable growth by wide and deep convolutional neural networks. However, previous deep learning-based alpha matting methods require a high computational cost to be used in real environments including mobile devices. In this letter, a lightweight natural image matting network with a similarity-preserving knowledge distillation is developed. The similarity-preserving knowledge distillation makes pairwise similarities from a compact student network similar to those from a teacher network. The pairwise similarity measured on spatial, channel, and batch units enables to transfer knowledge of the teacher to the student. Based on the similarity-preserving knowledge distillation, we not only design a lighter and smaller student network than the teacher one but also achieve superior performance compared to that of the student network without the knowledge distillation. In addition, the proposed algorithm can be seamlessly applied to various deep image matting algorithms. Therefore, our algorithm is effective for mobile applications (e.g., human portrait matting), which are in growing demand. The effectiveness of the proposed algorithm is verified on two public benchmark datasets. Donggeun Yoon, Jinsun Park, Donghyeon Cho |
IEEE Signal Process. Lett. | 3 |
| 2019 | Discriminative Feature Learning for Unsupervised Video SummarizationabstractIn this paper, we address the problem of unsupervised video summarization that automatically extracts key-shots from an input video. Specifically, we tackle two critical issues based on our empirical observations: (i) Ineffective feature learning due to flat distributions of output importance scores for each frame, and (ii) training difficulty when dealing with longlength video inputs. To alleviate the first problem, we propose a simple yet effective regularization loss term called variance loss. The proposed variance loss allows a network to predict output scores for each frame with high discrepancy which enables effective feature learning and significantly improves model performance. For the second problem, we design a novel two-stream network named Chunk and Stride Network (CSNet) that utilizes local (chunk) and global (stride) temporal view on the video features. Our CSNet gives better summarization results for long-length videos compared to the existing methods. In addition, we introduce an attention mechanism to handle the dynamic information in videos. We demonstrate the effectiveness of the proposed methods by conducting extensive ablation studies and show that our final model achieves new state-of-the-art results on two benchmark datasets. Yunjae Jung, Donghyeon Cho, Dahun Kim, Sanghyun Woo, In-So Kweon |
AAAI | 2 |
| 2019 | Self-Supervised Video Representation Learning with Space-Time Cubic PuzzlesabstractSelf-supervised tasks such as colorization, inpainting and zigsaw puzzle have been utilized for visual representation learning for still images, when the number of labeled images is limited or absent at all. Recently, this worthwhile stream of study extends to video domain where the cost of human labeling is even more expensive. However, the most of existing methods are still based on 2D CNN architectures that can not directly capture spatio-temporal information for video applications. In this paper, we introduce a new self-supervised task called as Space-Time Cubic Puzzles to train 3D CNNs using large scale video dataset. This task requires a network to arrange permuted 3D spatio-temporal crops. By completing Space-Time Cubic Puzzles, the network learns both spatial appearance and temporal relation of video frames, which is our final goal. In experiments, we demonstrate that our learned 3D representation is well transferred to action recognition tasks, and outperforms state-of-the-art 2D CNN-based competitors on UCF101 and HMDB51 datasets. Dahun Kim, Donghyeon Cho, In-So Kweon |
AAAI | 2 |
| 2019 | Video Retargeting: Trade-off between Content Preservation and Spatio-temporal ConsistencyabstractAs new display technologies (i.e. foldable phone and modular display) with variable aspect ratios emerge, content-aware video retargeting has attracted much attention from both academia and industry. The content-aware video retargeting aims to adjust the aspect ratio of a video sequence while preserving both, its content and its spatio-temporal consistency. This is a particularly challenging task since these two properties may drastically differ and contradict depending on the video characteristics. In this paper, we explore this conflict in the context of video retargeting, then we propose an appropriate solution to alleviate this issue using a deep recurrent convolutional neural network architecture. First of all, we present a method to generate multiple ground-truth labels under various aspect ratios. Using this dataset, our network is trained to predict various retargeted video candidates from a single input sequence. The resulting candidates present different properties, some of them with more emphasis on the content preservation while the others focus on the spatio-temporal consistency. Among the generated candidates, the final result which satisfy the best compromise is selected. A large set of qualitative and quantitative experiments shows the ability of our method for the content-aware video retargeting. Donghyeon Cho, Yunjae Jung, François Rameau, Dahun Kim, Sanghyun Woo, In-So Kweon |
ACM Multimedia | 1 |
| 2019 | Preserving Semantic and Temporal Consistency for Unpaired Video-to-Video TranslationabstractIn this paper, we investigate the problem of unpaired video-to-video translation. Given a video in the source domain, we aim to learn the conditional distribution of the corresponding video in the target domain, without seeing any pairs of corresponding videos. While significant progress has been made in the unpaired translation of images, directly applying these methods to an input video leads to low visual quality due to the additional time dimension. In particular, previous methods suffer from semantic inconsistency (i.e., semantic label flipping) and temporal flickering artifacts. To alleviate these issues, we propose a new framework that is composed of carefully-designed generators and discriminators, coupled with two core objective functions: 1) content preserving loss and 2) temporal consistency loss. Extensive qualitative and quantitative evaluations demonstrate the superior performance of the proposed method against previous approaches. We further apply our framework to a domain adaptation task and achieve favorable results. Kwanyong Park, Sanghyun Woo, Dahun Kim, Donghyeon Cho, In-So Kweon |
ACM Multimedia | 4 |
| 2019 | Deep Convolutional Neural Network for Natural Image Matting Using Initial Alpha MattesabstractWe propose a deep convolutional neural network (CNN) method for natural image matting. Our method takes multiple initial alpha mattes of the previous methods and normalized RGB color images as inputs, and directly learns an end-to-end mapping between the inputs and reconstructed alpha mattes. Among the various existing methods, we focus on using two simple methods as initial alpha mattes: the closed-form matting and KNN matting. They are complementary to each other in terms of local and nonlocal principles. A major benefit of our method is that it can "recognize" different local image structures and then combine the results of local (closed-form matting) and nonlocal (KNN matting) mattings effectively to achieve higher quality alpha mattes than both of the inputs. Furthermore, we verify extendability of the proposed network to different combinations of initial alpha mattes from more advanced techniques such as KL divergence matting and information-flow matting. On the top of deep CNN matting, we build an RGB guided JPEG artifacts removal network to handle JPEG block artifacts in alpha matting. Extensive experiments demonstrate that our proposed deep CNN matting produces visually and quantitatively high-quality alpha mattes. We perform deeper experiments including studies to evaluate the importance of balancing training data and to measure the effects of initial alpha mattes and also consider results from variant versions of the proposed network to analyze our proposed DCNN matting. In addition, our method achieved high ranking in the public alpha matting evaluation dataset in terms of the sum of absolute differences, mean squared errors, and gradient errors. Also, our RGB guided JPEG artifacts removal network restores the damaged alpha mattes from compressed images in JPEG format. Donghyeon Cho, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Image Process. | 1 |
| 2018 | Double JPEG Detection in Mixed JPEG Quality Factors Using Deep Convolutional Neural Network
Jin-Seok Park, Donghyeon Cho, Wonhyuk Ahn, Heung-Kyu Lee |
ECCV (5) | 2 |
| 2018 | Spike Counts Based Low Complexity Learning with Binary SynapseabstractTwo main difficulties encountered when implementing the neuromorphic system that supports pre- and post-synaptic textbf spike timing based real-time learning, are 1) huge memory size to store synaptic weights and 2) large amount of computations to update the synaptic weights. In this paper, we present a textbf spike counts based learning method that can significantly relieve the hardware burden of the real time unsupervised learning process. The novel learning approach uses both pre- and post-synaptic textbf spike counts as decision metrics in the following two ways: First, the synaptic weights are updated following the proposed simplified mean-based weight update rules, where binary feature images are produced based on pre-synaptic spike counts. Using the feature image, 1 bit synaptic weights are updated only once for each input image. When updating the weights, the post-synaptic spiking counts are also used to select the most active excitatory neuron. The actual weight updates are performed only for the weights connected to the dominant excitatory neuron, leading to further reduction of the weight updates without sacrificing accuracy. The simulation results show that using only 1 bit synaptic weights with 400 output neurons, the proposed learning approach achieves 82 % recognition accuracy with MNIST test set. In addition, the number of weight updates is reduced by 15.3 times compared to the state-of-the-art learning method. Hoyoung Tang, Heetak Kim, Donghyeon Cho, Jongsun Park 0001 |
IJCNN | 3 |
| 2018 | LinkNet: Relational Embedding for Scene GraphabstractObjects and their relationships are critical contents for image understanding. A scene graph provides a structured description that captures these properties of an image. However, reasoning about the relationships between objects is very challenging and only a few recent works have attempted to solve the problem of generating a scene graph from an image. In this paper, we present a novel method that improves scene graph generation by explicitly modeling inter-dependency among the entire object instances. We design a simple and effective relational embedding module that enables our model to jointly represent connections among all related objects, rather than focus on an object in isolation. Our novel method significantly benefits two main parts of the scene graph generation task: object classification and relationship classification. Using it on top of a basic Faster R-CNN, our model achieves state-of-the-art results on the Visual Genome benchmark. We further push the performance by introducing global context encoding module and geometrical layout encoding module. We validate our final model, LinkNet, through extensive ablation studies, demonstrating its efficacy in scene graph generation. Sanghyun Woo, Dahun Kim, Donghyeon Cho, In-So Kweon |
NeurIPS | 3 |
| 2018 | Learning Image Representations by Completing Damaged Jigsaw PuzzlesabstractIn this paper, we explore methods of complicating selfsupervised tasks for representation learning. That is, we do severe damage to data and encourage a network to recover them. First, we complicate each of three powerful self-supervised task candidates: jigsaw puzzle, inpainting, and colorization. In addition, we introduce a novel complicated self-supervised task called "Completing damaged jigsaw puzzles" which is puzzles with one piece missing and the other pieces without color. We train a convolutional neural network not only to solve the puzzles, but also generate the missing content and colorize the puzzles. The recovery of the aforementioned damage pushes the network to obtain robust and general-purpose representations. We demonstrate that complicating the self-supervised tasks improves their original versions and that our final task learns more robust and transferable representations compared to the previous methods, as well as the simple combination of our candidate tasks. Our approach achieves state-of-the-art performance in transfer learning on PASCAL classification and semantic segmentation. Dahun Kim, Donghyeon Cho, Donggeun Yoo, In-So Kweon |
WACV | 2 |
| 2017 | A Unified Approach of Multi-scale Deep and Hand-Crafted Features for Defocus EstimationabstractIn this paper, we introduce robust and synergetic hand-crafted features and a simple but efficient deep feature from a convolutional neural network (CNN) architecture for defocus estimation. This paper systematically analyzes the effectiveness of different features, and shows how each feature can compensate for the weaknesses of other features when they are concatenated. For a full defocus map estimation, we extract image patches on strong edges sparsely, after which we use them for deep and hand-crafted feature extraction. In order to reduce the degree of patch-scale dependency, we also propose a multi-scale patch extraction strategy. A sparse defocus map is generated using a neural network classifier followed by a probability-joint bilateral filter. The final defocus map is obtained from the sparse defocus map with guidance from an edge-preserving filtered input image. Experimental results show that our algorithm is superior to state-of-the-art algorithms in terms of defocus estimation. Our work can be used for applications such as segmentation, blur magnification, all-in-focus image generation, and 3-D estimation. Jinsun Park, Yu-Wing Tai, Donghyeon Cho, In-So Kweon |
CVPR | 3 |
| 2017 | Weakly- and Self-Supervised Learning for Content-Aware Deep Image RetargetingabstractThis paper proposes a weakly- and self-supervised deep convolutional neural network (WSSDCNN) for content-aware image retargeting. Our network takes a source image and a target aspect ratio, and then directly outputs a retargeted image. Retargeting is performed through a shift reap, which is a pixel-wise mapping from the source to the target grid. Our method implicitly learns an attention map, which leads to r content-aware shift map for image retargeting. As a result, discriminative parts in an image are preserved, while background regions are adjusted seamlessly. In the training phase, pairs of an image and its image-level annotation are used to compute content and structure tosses. We demonstrate the effectiveness of our proposed method for a retargeting application with insightful analyses. Donghyeon Cho, Jinsun Park, Tae-Hyun Oh, Yu-Wing Tai, In-So Kweon |
ICCV | 1 |
| 2017 | Two-Phase Learning for Weakly Supervised Object LocalizationabstractWeakly supervised semantic segmentation and localization have a problem of focusing only on the most important parts of an image since they use only image-level annotations. In this paper, we solve this problem fundamentally via two-phase learning. Our networks are trained in two steps. In the first step, a conventional fully convolutional network (FCN) is trained to find the most discriminative parts of an image. In the second step, the activations on the most salient parts are suppressed by inference conditional feedback, and then the second learning is performed to find the area of the next most important parts. By combining the activations of both phases, the entire portion of the target object can be captured. Our proposed training scheme is novel and can be utilized in well-designed techniques for weakly supervised semantic segmentation, salient region detection, and object location prediction. Detailed experiments demonstrate the effectiveness of our two-phase learning in each task. Dahun Kim, Donghyeon Cho, Donggeun Yoo |
ICCV | 2 |
| 2017 | Automatic Trimap Generation and Consistent Matting for Light-Field ImagesabstractIn this paper, we introduce an automatic approach to generate trimaps and consistent alpha mattes of foreground objects in a light-field image. Our method first performs binary segmentation to roughly segment a light-field image into foreground and background based on depth and color. Next, we estimate accurate trimaps through analyzing color distribution along the boundary of the segmentation using guided image filter and KL-divergence. In order to estimate consistent alpha mattes across sub-images, we utilize the epipolar plane image (EPI) where colors and alphas along the same epipolar line must be consistent. Since EPI of foreground and background are mixed in the matting area, we propagate the EPI from definite foreground/background regions to unknown regions by assuming depth variations within unknown regions are spatially smooth. Using the EPI constraint, we derive two solutions to estimate alpha when color samples along epipolar line are known, and unknown. To further enhance consistency, we refine the estimated alpha mattes by using the multi-image matting Laplacian with an additional EPI smoothness constraint. In experimental evaluations, we have created a dataset where the ground truth alpha mattes of light-field images were obtained by using the blue screen technique. A variety of experiments show that our proposed algorithm produces both visually and quantitatively high-quality alpha mattes for light-field images. Donghyeon Cho, Sunyeong Kim, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Photometric Stereo Under Non-uniform Light Intensities and Exposures
Donghyeon Cho, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
ECCV (2) | 1 |
| 2016 | Natural Image Matting Using Deep Convolutional Neural Networks
Donghyeon Cho, Yu-Wing Tai, In-So Kweon |
ECCV (2) | 1 |
| 2014 | Consistent Matting for Light Field Images
Donghyeon Cho, Sunyeong Kim, Yu-Wing Tai |
ECCV (4) | 1 |
| 2013 | Modeling the Calibration Pipeline of the Lytro Camera for High Quality Light-Field Image ReconstructionabstractLight-field imaging systems have got much attention recently as the next generation camera model. A light-field imaging system consists of three parts: data acquisition, manipulation, and application. Given an acquisition system, it is important to understand how a light-field camera converts from its raw image to its resulting refocused image. In this paper, using the Lytro camera as an example, we describe step-by-step procedures to calibrate a raw light-field image. In particular, we are interested in knowing the spatial and angular coordinates of the micro lens array and the resampling process for image reconstruction. Since Lytro uses a hexagonal arrangement of a micro lens image, additional treatments in calibration are required. After calibration, we analyze and compare the performances of several resampling methods for image reconstruction with and without calibration. Finally, a learning based interpolation method is proposed which demonstrates a higher quality image reconstruction than previous interpolation methods including a method used in Lytro software. Donghyeon Cho, Minhaeng Lee, Sunyeong Kim, Yu-Wing Tai |
ICCV | 1 |