EDBT 2026 Demo / reviewers in the wild / expert
Kyoung Mu Lee
dblp:17/4029
· DBLP profile ↗
224ranked-venue papers
7as first author
61since 2021 · last 2026
0000-0001-7210-1036ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 180 · 5 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 169 · 4 first-author · 49 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Particle Diffusion Matching: Random Walk Correspondence Search for the Alignment of Standard and Ultra-Widefield Fundus ImagesabstractWe propose a robust alignment technique for Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which are challenging to align due to differences in scale, appearance, and the scarcity of distinctive features. Our method, termed Particle Diffusion Matching (PDM), performs alignment through an iterative Random Walk Correspondence Search (RWCS) guided by a diffusion model. At each iteration, the model estimates displacement vectors for particle points by considering local appearance, the structural distribution of particles, and an estimated global transformation, enabling progressive refinement of correspondences even under difficult conditions. PDM achieves state-of-the-art performance across multiple retinal image alignment benchmarks, showing substantial improvement on a primary dataset of SFI-UWFI pairs and demonstrating its effectiveness in real-world clinical scenarios. By providing accurate and scalable correspondence estimation, PDM overcomes the limitations of existing methods and facilitates the integration of complementary retinal image modalities. This diffusion-guided search strategy offers a new direction for improving downstream supervised learning, disease diagnosis, and multi-modal image analysis in ophthalmology. Kanggeon Lee, Soochahn Lee, Kyoung Mu Lee |
IEEE Trans. Image Process. | 3 |
| 2025 | OmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable CapabilitiesabstractFeed-forward 3D Gaussian splatting (3DGS) models have gained significant popularity due to their ability to generate scenes immediately without needing per-scene optimization. Although omnidirectional images are becoming more popular since they reduce the computation required for image stitching to composite a holistic scene, existing feed-forward models are only designed for perspective images. The unique optical properties of omnidirectional images make it difficult for feature encoders to correctly understand the context of the image and make the Gaussian non-uniform in space, which hinders the image quality synthesized from novel views. We propose OmniSplat, a training-free fast feed-forward 3DGS generation framework for omnidirectional images. We adopt a Yin-Yang grid and decompose images based on it to reduce the domain gap between omnidirectional and perspective images. The Yin-Yang grid can use the existing CNN structure as it is, but its quasi-uniform characteristic allows the decomposed image to be similar to a perspective image, so it can exploit the strong prior knowledge of the learned feed-forward network. OmniSplat demonstrates higher reconstruction accuracy than existing feed-forward networks trained on perspective images. The code is available on: https://github.com/esw0116/OmniSplat. Suyoung Lee, Jaeyoung Chung 0002, Kihoon Kim, Jaeyoo Huh, Minsoo Lee, Kyoung Mu Lee |
CVPR | 7 |
| 2025 | SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion ModelsabstractWe introduce SemanticDraw, a new paradigm of interactive content creation where high-quality images are generated in near real-time from given multiple hand-drawn regions, each encoding prescribed semantic meaning. In order to maximize the productivity of content creators and to fully realize their artistic imagination, it requires both quick interactive interfaces and fine-grained regional controls in their tools. Despite astonishing generation quality from recent diffusion models, we find that existing approaches for regional controllability are very slow (52 seconds for 512 × 512 image) while not compatible with acceleration methods such as LCM, blocking their huge potential in interactive content creation. From this observation, we build our solution for interactive content creation in two steps: (1) we establish compatibility between region-based controls and acceleration techniques for diffusion models, maintaining high fidelity of multi-prompt image generation with ×10 reduced number of inference steps, (2) we increase the generation throughput with our new multi-prompt stream batch pipeline, enabling low-latency generation from multiple, region-based text prompts on a single RTX 2080 Ti GPU. Our proposed framework is generalizable to any existing diffusion models and acceleration schedulers, allowing sub-second (0.64 seconds) image content creation application upon well-established image diffusion models. The code is https://github.com/ironjr/semantic-draw. Jaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu Lee |
CVPR | 4 |
| 2025 | DeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single ImageabstractMost existing methods of 3D clothed human reconstruction from a single image treat the clothed human as a single object without distinguishing between cloth and human body. In this regard, we present DeClotH, which separately reconstructs 3D cloth and human body from a single image. This task remains largely unexplored due to the extreme occlusion between cloth and the human body, making it challenging to infer accurate geometries and textures. Moreover, while recent 3D human reconstruction methods have achieved impressive results using text-to-image diffusion models, directly applying such an approach to this problem often leads to incorrect guidance, particularly in reconstructing 3D cloth. To address these challenges, we propose two core designs in our framework. First, to alleviate the occlusion issue, we leverage 3D template models of cloth and human body as regularizations, which provide strong geometric priors to prevent erroneous reconstruction by the occlusion. Second, we introduce a cloth diffusion model specifically designed to provide contextual information about cloth appearance, thereby enhancing the reconstruction of 3D cloth. Qualitative and quantitative experiments demonstrate that our proposed approach is highly effective in reconstructing both 3D cloth and the human body. Hyeongjin Nam, Jeongtaek Oh, Kyoung Mu Lee |
CVPR | 4 |
| 2025 | Exploiting Diffusion Prior for Task-Driven Image Restoration
Jaeha Kim, Junghun Oh, Kyoung Mu Lee |
ICCV | 3 |
| 2025 | Auto-Regressive Transformation for Image AlignmentabstractExisting methods for image alignment struggle in cases involving feature-sparse regions, extreme scale and field-of-view differences, and large deformations, often resulting in suboptimal accuracy. Robustness to these challenges can be improved through iterative refinement of the transform field while focusing on critical regions in multi-scale image representations. We thus propose Auto-Regressive Transformation (ART), a novel method that iteratively estimates the coarse-to-fine transformations through an auto-regressive pipeline. Leveraging hierarchical multi-scale features, our network refines the transform field parameters using randomly sampled points at each scale. By incorporating guidance from the cross-attention layer, the model focuses on critical regions, ensuring accurate alignment even in challenging, feature-limited conditions. Extensive experiments demonstrate that ART significantly outperforms state-of-the-art methods on planar images and achieves comparable performance on 3D scene images, establishing it as a powerful and versatile solution for precise image alignment. Kanggeon Lee, Soochahn Lee, Kyoung Mu Lee |
ICCV | 3 |
| 2025 | PARTE: Part-Guided Texturing for 3D Human Reconstruction from a Single ImageabstractThe misaligned human texture across different human parts is one of the main limitations of existing 3D human reconstruction methods. Each human part, such as a jacket or pants, should maintain a distinct texture without blending into others. The structural coherence of human parts serves as a crucial cue to infer human textures in the invisible regions of a single image. However, most existing 3D human reconstruction methods do not explicitly exploit such part segmentation priors, leading to misaligned textures in their reconstructions. In this regard, we present PARTE, which utilizes 3D human part information as a key guide to reconstruct 3D human textures. Our framework comprises two core components. First, to infer 3D human part information from a single image, we propose a 3D part segmentation module (PartSegmenter) that initially reconstructs a textureless human surface and predicts human part labels based on the textureless surface. Second, to incorporate part information into texture reconstruction, we introduce a part-guided texturing module (PartTexturer), which acquires prior knowledge from a pre-trained image generation network on texture alignment of human parts. Extensive experiments demonstrate that our framework achieves state-of-the-art quality in 3D human reconstruction. The project page is available at https://hygenie1228.github.io/PARTE/. Hyeongjin Nam, Gyeongsik Moon, Kyoung Mu Lee |
ICCV | 4 |
| 2025 | Find A Winning Sign: Sign Is All We Need to Win the LotteryabstractThe Lottery Ticket Hypothesis (LTH) posits the existence of a sparse subnetwork (a.k.a. winning ticket) that can generalize comparably to its over-parameterized counterpart when trained from scratch.
The common approach to finding a winning ticket is to preserve the original strong generalization through Iterative Pruning (IP) and transfer information useful for achieving the learned generalization by applying the resulting sparse mask to an untrained network.
However, existing IP methods still struggle to generalize their observations beyond ad-hoc initialization and small-scale architectures or datasets, or they bypass these challenges by applying their mask to trained weights instead of initialized ones.
In this paper, we demonstrate that the parameter sign configuration plays a crucial role in conveying useful information for generalization to any randomly initialized network.
Through linear mode connectivity analysis, we observe that a sparse network trained by an existing IP method can retain its basin of attraction if its parameter signs and normalization layer parameters are preserved.
To take a step closer to finding a winning ticket, we alleviate the reliance on normalization layer parameters by preventing high error barriers along the linear path between the sparse network trained by our method and its counterpart with initialized normalization layer parameters.
Interestingly, across various architectures and datasets, we observe that any randomly initialized network can be optimized to exhibit low error barriers along the linear path to the sparse network trained by our method by inheriting its sparsity and parameter sign information, potentially achieving performance comparable to the original.
The code is available at https://github.com/JungHunOh/AWS_ICLR2025.git. Junghun Oh, Sungyong Baik, Kyoung Mu Lee |
ICLR | 3 |
| 2025 | Learning Dense Hand Contact Estimation from Imbalanced DataabstractHands are essential to human interaction, and exploring contact between hands and the world can promote comprehensive understanding of their function. Recently, there have been growing number of hand interaction datasets that cover interaction with object, other hand, scene, and body. Despite the significance of the task and increasing high-quality data, how to effectively learn dense hand contact estimation remains largely underexplored. There are two major challenges for learning dense hand contact estimation. First, there exists class imbalance issue from hand contact datasets where majority of regions are not in contact. Second, hand contact datasets contain spatial imbalance issue with most of hand contact exhibited in finger tips, resulting in challenges for generalization towards contacts in other hand regions. To tackle these issues, we present a framework that learns dense HAnd COntact estimation (HACO) from imbalanced data. To resolve the class imbalance issue, we introduce balanced contact sampling, which builds and samples from multiple sampling groups that fairly represent diverse contact statistics for both contact and non-contact vertices. Moreover, to address the spatial imbalance issue, we propose vertex-level class-balanced (VCB) loss, which incorporates spatially varying contact distribution by separately reweighting loss contribution of each vertex based on its contact frequency across dataset. As a result, we effectively learn to predict dense hand contact estimation with large-scale hand contact data without suffering from class and spatial imbalance issue. The codes are available at https://github.com/dqj5182/HACO_RELEASE. Daniel Sungho Jung, Kyoung Mu Lee |
NeurIPS | 2 |
| 2025 | Difficulty, Diversity, and Plausibility: Dynamic Data-Free QuantizationabstractWithout access to the original training data, data-free quantization (DFQ) aims to recover the performance loss induced by quantization. Most previous works have focused on using an original network to extract the train data information, which is instilled into surrogate synthesized images. However, existing DFQ methods do not take into account important aspects of quantization: the extent of a computational-cost-and-accuracy trade-off varies for each image, depending on its task difficulty. To handle such varying trade-offs, several efforts have been made to dynamically allocate bit-widths for each image. Such dynamic quantization, however, remains challenging and unexplored in the data-free domain, because synthesized images of previous works fail to possess properties in natural test images that are crucial for learning the appropriate dynamic allocation policy: difficulty, its diversity, and its plausibility. By contrast, we propose a data-free quantization framework that is dynamic-friendly, by modeling varying extents of task difficulties with plausibility. We generate plausibly difficult images with soft labels, whose probabilities are allocated to a group of similar classes. Images with diverse and plausible difficulties enable us to train the framework to dynamically handle the varying trade-offs. Consequently, our framework achieves better accuracy-complexity Pareto front than existing data-free quantization approaches. Cheeun Hong, Sungyong Baik, Junghun Oh, Kyoung Mu Lee |
WACV | 4 |
| 2025 | LucidDreamer: Domain-Free Generation of 3D Gaussian Splatting ScenesabstractGenerating high-quality 3D scenes is a critical challenge in computer vision, driven by advances in 3D graphics and the growing demand for immersive environments. While object-centric 3D generation has achieved significant progress, scene generation remains difficult due to the scarcity of large-scale 3D scene datasets and scalability constraints of conventional 3D representations, which hinder efficient large-scale expansion. To address these challenges, we propose LucidDreamer, a novel pipeline that synthesizes diverse, high-quality, and expandable 3D scenes using a unified 3D Gaussian splatting representation. Our approach employs an iterative Navigation-Dreaming-Alignment process, leveraging 2D image generation and depth estimation to construct photorealistic, scalable 3D environments. By iteratively generating images and navigating through the scene, LucidDreamer fully utilizes the power of image generation models, enabling the creation of highly detailed and expandable 3D scenes. LucidDreamer supports various input modalities, including text, RGB, and RGBD, and enables dynamic modifications during generation. Experimental results demonstrate that LucidDreamer outperforms existing methods in generating high-quality, diverse, structurally consistent, and navigable 3D scenes. Jaeyoung Chung 0002, Suyoung Lee, Hyeongjin Nam, Jaerin Lee, Kyoung Mu Lee |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | AdaBM: On-the-Fly Adaptive Bit Mapping for Image Super-ResolutionabstractAlthough image super-resolution (SR) problem has ex-perienced unprecedented restoration accuracy with deep neural networks, it has yet limited versatile applications due to the substantial computational costs. Since differ-ent input images for SR face different restoration difficul-ties, adapting computational costs based on the input image, referred to as adaptive inference, has emerged as a promising solution to compress SR networks. Specifically, adapting the quantization bit-widths has successfully re-duced the inference and memory cost without sacrificing the accuracy. However, despite the benefits of the resul-tant adaptive network, existing works rely on time-intensive quantization-aware training with full access to the origi-nal training pairs to learn the appropriate bit allocation policies, which limits its ubiquitous usage. To this end, we introduce the first on-the-fly adaptive quantization frame-work that accelerates the processing time from hours to sec-onds. We formulate the bit allocation problem with only two bit mapping modules: one to map the input image to the image-wise bit adaptation factor and one to obtain the layer-wise adaptation factors. These bit mappings are cali-brated and fine-tuned using only a small number of calibration images. We achieve competitive performance with the previous adaptive quantization methods, while the processing time is accelerated by × 2000. Codes are available at https://github.com/Cheeun/AdaBM. Cheeun Hong, Kyoung Mu Lee |
CVPR | 2 |
| 2024 | Beyond Image Super-Resolution for Image Recognition with Task-Driven Perceptual LossabstractIn real-world scenarios, image recognition tasks, such as semantic segmentation and object detection, often pose greater challenges due to the lack of information available within low-resolution (LR) content. Image super-resolution (SR) is one of the promising solutions for addressing the challenges. However, due to the ill-posed property of SR, it is challenging for typical SR methods to restore task-relevant high-frequency contents, which may dilute the advantage of utilizing the SR method. Therefore, in this paper, we propose Super-Resolution for Image Recognition (SR4IR) that effectively guides the generation of SR images beneficial to achieving satisfactory image recognition performance when processing LR images. The critical component of our SR4IR is the task-driven perceptual (TDP) loss that enables the SR network to acquire task-specific knowledge from a network tailored for a specific task. Moreover, we propose a cross-quality patch mix and an alternate training framework that significantly enhances the efficacy of the TDP loss by addressing potential problems when employing the TDP loss. Through extensive experiments, we demonstrate that our SR4IR achieves outstanding task performance by generating SR images useful for a specific image recognition task, including semantic segmentation, object detection, and image classification. The implementation code is available at https://github.com/JaehaKim97ISR4IR. Jaeha Kim, Junghun Oh, Kyoung Mu Lee |
CVPR | 3 |
| 2024 | Joint Reconstruction of 3D Human and Object via Contact-Based Refinement TransformerabstractHuman-object contact serves as a strong cue to understand how humans physically interact with objects. Nev-ertheless, it is not widely explored to utilize human-object contact information for the joint reconstruction of 3D human and object from a single image. In this work, we present a novel joint 3D human-object reconstruction method (CONTHO) that effectively exploits contact information between humans and objects. There are two core designs in our system: 1) 3D-guided contact estimation and 2) contact-based 3D human and object refinement. First, for accurate human-object contact estimation, CONTHO initially reconstructs 3D humans and objects and utilizes them as explicit 3D guidance for contact estimation. Second, to refine the initial reconstructions of 3D human and object, we propose a novel contact-based refinement Transformer that effectively aggregates human features and object features based on the estimated human-object contact. The proposed contact-based refinement prevents the learning of erroneous correlation between human and object, which enables accurate 3D reconstruction. As a result, our CON-THO achieves state-of-the-art performance in both human-object contact estimation and joint reconstruction of 3D human and object. The code is publicly available11https://github.com/dqj5182/CONTHO_RELEASE. Hyeongjin Nam, Daniel Sungho Jung, Gyeongsik Moon, Kyoung Mu Lee |
CVPR | 4 |
| 2024 | CNC-Net: Self-Supervised Learning for CNC Machining OperationsabstractCNC manufacturing is a process that employs computer numerical control (CNC) machines to govern the move-ments of various industrial tools and machinery, encom-passing equipment ranging from grinders and lathes to mills and CNC routers. However, the reliance on man-ual CNC programming has become a bottleneck, and the requirement for expert knowledge can result in significant costs. Therefore, we introduce a pioneering approach named CNC-Net, representing the use of deep neural net-works (DNNs) to simulate CNC machines and grasp intri-cate operations when supplied with raw materials. CNC-Net constitutes a self-supervised framework that exclu-sively takes an input 3D model and subsequently gener-ates the essential operation parameters required by the CNC machine to construct the object. Our method has the potential to transformative automation in manufac-turing by offering a cost-effective alternative to the high costs of manual CNC programming while maintaining ex-ceptional precision in 3D object production. Our ex-periments underscore the effectiveness of our CNC-Net in constructing the desired 3D objects through the uti-lization of CNC operations. Notably, it excels in pre-serving finer local details, exhibiting a marked enhance-ment in precision compared to the state-of-the-art 3D CAD reconstruction approaches. The codes are available at https://github.com/myavartanoo/CNC-Net_PyTorch. Mohsen Yavartanoo, Sangmin Hong, Reyhaneh Neshatavar, Kyoung Mu Lee |
CVPR | 4 |
| 2024 | Overcoming Distribution Mismatch in Quantizing Image Super-Resolution Networks
Cheeun Hong, Kyoung Mu Lee |
ECCV (13) | 2 |
| 2024 | CLOSER: Towards Better Representation Learning for Few-Shot Class-Incremental Learning
Junghun Oh, Sungyong Baik, Kyoung Mu Lee |
ECCV (49) | 3 |
| 2024 | 3D Hand Sequence Recovery from Real Blurry Images and Event Stream
Joonkyu Park, Gyeongsik Moon, Weipeng Xu, Evan Kaseman, Takaaki Shiratori, Kyoung Mu Lee |
ECCV (59) | 6 |
| 2024 | ODGS: 3D Scene Reconstruction from Omnidirectional Images with 3D Gaussian SplattingsabstractOmnidirectional (or 360-degree) images are increasingly being used for 3D applications since they allow the rendering of an entire scene with a single image. Existing works based on neural radiance fields demonstrate successful 3D reconstruction quality on egocentric videos, yet they suffer from long training and rendering times. Recently, 3D Gaussian splatting has gained attention for its fast optimization and real-time rendering. However, directly using a perspective rasterizer to omnidirectional images results in severe distortion due to the different optical properties between the two image domains. In this work, we present ODGS, a novel rasterization pipeline for omnidirectional images with geometric interpretation. For each Gaussian, we define a tangent plane that touches the unit sphere and is perpendicular to the ray headed toward the Gaussian center. We then leverage a perspective camera rasterizer to project the Gaussian onto the corresponding tangent plane. The projected Gaussians are transformed and combined into the omnidirectional image, finalizing the omnidirectional rasterization process. This interpretation reveals the implicit assumptions within the proposed pipeline, which we verify through mathematical proofs. The entire rasterization process is parallelized using CUDA, achieving optimization and rendering speeds 100 times faster than NeRF-based methods. Our comprehensive experiments highlight the superiority of ODGS by delivering the best reconstruction and perceptual quality across various datasets. Additionally, results on roaming datasets demonstrate that ODGS effectively restores fine details, even when reconstructing large 3D scenes. The source code is available on our project page (https://github.com/esw0116/ODGS). Suyoung Lee, Jaeyoung Chung 0002, Jaeyoo Huh, Kyoung Mu Lee |
NeurIPS | 4 |
| 2024 | GS-Blur: A 3D Scene-Based Dataset for Realistic Image DeblurringabstractTo train a deblurring network, an appropriate dataset with paired blurry and sharp images is essential.Existing datasets collect blurry images either synthetically by aggregating consecutive sharp frames or using sophisticated camera systems to capture real blur.However, these methods offer limited diversity in blur types (blur trajectories) or require extensive human effort to reconstruct large-scale datasets, failing to fully reflect real-world blur scenarios.To address this, we propose GS-Blur, a dataset of synthesized realistic blurry images created using a novel approach.To this end, we first reconstruct 3D scenes from multi-view images using 3D Gaussian Splatting~(3DGS), then render blurry images by moving the camera view along the randomly generated motion trajectories.By adopting various camera trajectories in reconstructing our GS-Blur, our dataset contains realistic and diverse types of blur, offering a large-scale dataset that generalizes well to real-world blur.Using GS-Blur with various deblurring methods, we demonstrate its ability to generalize effectively compared to previous synthetic or real blur datasets, showing significant improvements in deblurring performance.We will publicly release our dataset. Joonkyu Park, Kyoung Mu Lee |
NeurIPS | 3 |
| 2024 | ICF-SRSR: Invertible scale-Conditional Function for Self-Supervised Real-world Single Image Super-ResolutionabstractSingle image super-resolution (SISR) is a challenging ill-posed problem that aims to up-sample a given low-resolution (LR) image to a high-resolution (HR) counterpart. Due to the difficulty in obtaining real LR-HR training pairs, recent approaches are trained on simulated LR images degraded by simplified down-sampling operators, e.g., bicubic. Such an approach can be problematic in practice due to the large gap between the synthesized and real-world LR images. To alleviate the issue, we propose a novel Invertible scale-Conditional Function (ICF), which can scale an input image and then restore the original input with different scale conditions. Using the proposed ICF, we construct a novel self-supervised SISR framework (ICF-SRSR) to handle the real-world SR task without using any paired/unpaired training data. Furthermore, our ICF-SRSR can generate realistic and feasible LR-HR pairs, which can make existing supervised SISR networks more robust. Extensive experiments demonstrate the effectiveness of our method in handling SISR in a fully self-supervised manner. Our ICF-SRSR demonstrates superior performance compared to the existing methods trained on synthetic paired images in real-world scenarios and exhibits comparable performance compared to state-of-the-art supervised/unsupervised methods on public benchmark datasets. The code is available from this link. Reyhaneh Neshatavar, Mohsen Yavartanoo, Sanghyun Son 0002, Kyoung Mu Lee |
WACV | 4 |
| 2024 | Learning to Learn Task-Adaptive Hyperparameters for Few-Shot LearningabstractThe objective of few-shot learning is to design a system that can adapt to a given task with only few examples while achieving generalization. Model-agnostic meta-learning (MAML), which has recently gained the popularity for its simplicity and flexibility, learns a good initialization for fast adaptation to a task under few-data regime. However, its performance has been relatively limited especially when novel tasks are different from tasks previously seen during training. In this work, instead of searching for a better initialization, we focus on designing a better fast adaptation process. Consequently, we propose a new task-adaptive weight update rule that greatly enhances the fast adaptation process. Specifically, we introduce a small meta-network that can generate per-step hyperparameters for each given task: learning rate and weight decay coefficients. The experimental results validate that learning a good weight update rule for fast adaptation is the equally important component that has drawn relatively less attention in the recent few-shot learning approaches. Surprisingly, fast adaptation from random initialization with ALFA can already outperform MAML. Furthermore, the proposed weight-update rule is shown to consistently improve the task-adaptation capability of MAML across diverse problem domains: few-shot classification, cross-domain few-shot classification, regression, visual tracking, and video frame interpolation. Sungyong Baik, Myungsub Choi, Janghoon Choi, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | CoLaNet: Adaptive Context and Latent Information Blending for Face Image InpaintingabstractFace inpainting, the task of filling up missing regions in a face image plausibly, has witnessed great advances with deep learning-based approaches. To fill in the missing region, existing methods either use information from the surrounding visible region of the input image itself (i.e., context) or use prior knowledge obtained from the training data (i.e., latent). However, we find that exclusive usage of the two types of information is sub-optimal; whether the context-based approach is effective or the latent-based approach is effective is different for each missing region. To this end, we propose CoLaNet, a novel framework that adaptively blends context and latent information to inpaint face images. Specifically, the two types of information are balanced based on the attention between the missing region and the rest of the image. The regions strongly correlated to the visible region leverage context information more. Consequently, the adaptive utilization of context and latent information leads to better inpainting performance in various face images. Joonkyu Park, Cheeun Hong, Sungyong Baik, Kyoung Mu Lee |
IEEE Signal Process. Lett. | 4 |
| 2024 | Learning Controllable ISP for Image EnhancementabstractWe present a plug-and-play Image Signal Processor (ISP) for image enhancement to better produce diverse image styles than the previous works. Our proposed method, ContRollable Image Signal Processor (CRISP), explicitly controls the parameters of the ISP that determine output image styles. ISP parameters for high-quality (HQ) image styles are encoded into low-dimensional latent codes, allowing fast and easy style adjustments. We empirically show that CRISP covers a wide range of image styles with high efficiency. On the MIT-Adobe FiveK dataset, CRISP can very closely estimate the reference styles produced by human experts and achieves better MOS with diverse image styles. Compared with the state-of-the-art method, our ISP comprises only 19 parameters, allowing CRISP to have 2× smaller parameters and 100× reduced FLOPs for an image output. CRISP outperforms previous works in PSNR and FLOPs with several scenarios for style adjustments. Kyoung Mu Lee |
IEEE Trans. Image Process. | 2 |
| 2023 | MultiAct: Long-Term 3D Human Motion Generation from Multiple Action LabelsabstractWe tackle the problem of generating long-term 3D human motion from multiple action labels. Two main previous approaches, such as action- and motion-conditioned methods, have limitations to solve this problem. The action-conditioned methods generate a sequence of motion from a single action. Hence, it cannot generate long-term motions composed of multiple actions and transitions between actions. Meanwhile, the motion-conditioned methods generate future motions from initial motion. The generated future motions only depend on the past, so they are not controllable by the user's desired actions. We present MultiAct, the first framework to generate long-term 3D human motion from multiple action labels. MultiAct takes account of both action and motion conditions with a unified recurrent generation system. It repetitively takes the previous motion and action label; then, it generates a smooth transition and the motion of the given action. As a result, MultiAct produces realistic long-term motion controlled by the given sequence of multiple action labels. The code is publicly available in https://github.com/TaeryungLee/MultiAct RELEASE. Taeryung Lee, Gyeongsik Moon, Kyoung Mu Lee |
AAAI | 3 |
| 2023 | ACL-SPC: Adaptive Closed-Loop System for Self-Supervised Point Cloud CompletionabstractPoint cloud completion addresses filling in the missing parts of a partial point cloud obtained from depth sensors and generating a complete point cloud. Although there has been steep progress in the supervised methods on the synthetic point cloud completion task, it is hardly applicable in real-world scenarios due to the domain gap between the synthetic and real-world datasets or the requirement of prior information. To overcome these limitations, we propose a novel self-supervised framework ACL-SPC for point cloud completion to train and test on the same data. ACL-SPC takes a single partial input and attempts to output the complete point cloud using an adaptive closed-loop (ACL) system that enforces the output same for the variation of an input. We evaluate our ACL-SPC on various datasets to prove that it can successfully learn to complete a partial point cloud as the first self-supervised scheme. Results show that our method is comparable with unsupervised methods and achieves superior performance on the real-world dataset compared to the supervised methods trained on the synthetic dataset. Extensive experiments justify the necessity of self-supervised learning and the effectiveness of our proposed method for the real-world point cloud completion task. The code is publicly available from this link. Sangmin Hong, Mohsen Yavartanoo, Reyhaneh Neshatavar, Kyoung Mu Lee |
CVPR | 4 |
| 2023 | Recovering 3D Hand Mesh Sequence from a Single Blurry Image: A New Dataset and Temporal UnfoldingabstractHands, one of the most dynamic parts of our body, suffer from blur due to their active movements. However, previous 3D hand mesh recovery methods have mainly focused on sharp hand images rather than considering blur due to the absence of datasets providing blurry hand images. We first present a novel dataset BlurHand, which contains blurry hand images with 3D groundtruths. The BlurHand is constructed by synthesizing motion blur from sequential sharp hand images, imitating realistic and natural motion blurs. In addition to the new dataset, we propose BlurHandNet, a baseline network for accurate 3D hand mesh recovery from a blurry hand image. Our BlurHandNet unfolds a blurry input image to a 3D hand mesh sequence to utilize temporal information in the blurry input image, while previous works output a static single hand mesh. We demonstrate the usefulness of BlurHand for the 3D hand mesh recovery from blurry images in our experiments. The proposed BlurHandNet produces much more robust results on blurry images while generalizing well to in-the-wild images. The training codes and BlurHand dataset are available at https://github.com/laehaKim97IBlurHand_RELEASE. Yeonguk Oh, Joonkyu Park, Jaeha Kim, Gyeongsik Moon, Kyoung Mu Lee |
CVPR | 5 |
| 2023 | Human Part-wise 3D Motion Context Learning for Sign Language RecognitionabstractIn this paper, we propose P3D, the human part-wise motion context learning framework for sign language recognition. Our main contributions lie in two dimensions: learning the part-wise motion context and employing the pose ensemble to utilize 2D and 3D pose jointly. First, our empirical observation implies that part-wise context encoding benefits the performance of sign language recognition. While previous methods of sign language recognition learned motion context from the sequence of the entire pose, we argue that such methods cannot exploit part-specific motion context. In order to utilize part-wise motion context, we propose the alternating combination of a part-wise encoding Transformer (PET) and a whole-body encoding Transformer (WET). PET encodes the motion contexts from a part sequence, while WET merges them into a unified context. By learning part-wise motion context, our P3D achieves superior performance on WLASL compared to previous state-of-the-art methods. Second, our framework is the first to ensemble 2D and 3D poses for sign language recognition. Since the 3D pose holds rich motion context and depth information to distinguish the words, our P3D outperformed the previous state-of-the-art methods employing a pose ensemble. Taeryung Lee, Yeonguk Oh, Kyoung Mu Lee |
ICCV | 3 |
| 2023 | ExBluRF: Efficient Radiance Fields for Extreme Motion Blurred ImagesabstractWe present ExBluRF, a novel view synthesis method for extreme motion blurred images based on efficient radiance fields optimization. Our approach consists of two main components: 6-DOF camera trajectory-based motion blur formulation and voxel-based radiance fields. From extremely blurred images, we optimize the sharp radiance fields by jointly estimating the camera trajectories that generate the blurry images. In training, multiple rays along the camera trajectory are accumulated to reconstruct single blurry color, which is equivalent to the physical motion blur operation. We minimize the photo-consistency loss on blurred image space and obtain the sharp radiance fields with camera trajectories that explain the blur of all images. The joint optimization on the blurred image space demands painfully increasing computation and resources proportional to the blur size. Our method solves this problem by replacing the MLP-based framework to low-dimensional 6-DOF camera poses and voxel-based radiance fields. Compared with the existing works, our approach restores much sharper 3D scenes from challenging motion blurred views with the order of 10× less training time and GPU memory consumption. Jeongtaek Oh, Jaesung Rim, Sunghyun Cho, Kyoung Mu Lee |
ICCV | 5 |
| 2023 | Cyclic Test-Time Adaptation on Monocular Video for 3D Human Mesh ReconstructionabstractDespite recent advances in 3D human mesh reconstruction, domain gap between training and test data is still a major challenge. Several prior works tackle the domain gap problem via test-time adaptation that fine-tunes a network relying on 2D evidence (e.g., 2D human keypoints) from test images. However, the high reliance on 2D evidence during adaptation causes two major issues. First, 2D evidence induces depth ambiguity, preventing the learning of accurate 3D human geometry. Second, 2D evidence is noisy or partially non-existent during test time, and such imperfect 2D evidence leads to erroneous adaptation. To overcome the above issues, we introduce CycleAdapt, which cyclically adapts two networks: a human mesh reconstruction network (HMRNet) and a human motion denoising network (MDNet), given a test video. In our framework, to alleviate high reliance on 2D evidence, we fully supervise HMRNet with generated 3D supervision targets by MDNet. Our cyclic adaptation scheme progressively elaborates the 3D supervision targets, which compensate for imperfect 2D evidence. As a result, our CycleAdapt achieves state-of-the-art performance compared to previous test-time adaptation methods. The codes are available in here. Hyeongjin Nam, Daniel Sungho Jung, Yeonguk Oh, Kyoung Mu Lee |
ICCV | 4 |
| 2023 | Content-Aware Local GAN for Photo-Realistic Super-ResolutionabstractRecently, GAN has successfully contributed to making single-image super-resolution (SISR) methods produce more realistic images. However, natural images have complex distribution in the real world, and a single classifier in the discriminator may not have enough capacity to classify real and fake samples, making the preceding SR network generate unpleasing noise and artifacts. To solve the problem, we propose a novel content-aware local GAN framework, CAL-GAN, which processes a large and complicated distribution of real-world images by dividing them into smaller subsets based on similar contents. Our mixture of classifiers (MoC) design allocates different super-resolved patches to corresponding expert classifiers. Additionally, we introduce novel routing and orthogonality loss terms so that different classifiers can handle various contents and learn separable features. By feeding similar distributions into the corresponding specialized classifiers, CAL-GAN enhances the representation power of existing super-resolution models, achieving state-of-the-art perceptual performance on standard benchmarks and real-world images without modifying the generator-side architecture. The codes are available at https://github.com/jkpark0825/CAL_GAN. Joonkyu Park, Sanghyun Son 0002, Kyoung Mu Lee |
ICCV | 3 |
| 2023 | Rethinking Self-Supervised Visual Representation Learning in Pre-training for 3D Human Pose and Shape Estimation
Hongsuk Choi, Hyeongjin Nam, Taeryung Lee, Gyeongsik Moon, Kyoung Mu Lee |
ICLR | 5 |
| 2023 | NERDS: A General Framework to Train Camera Denoisers from Raw-RGB Noisy Image Pairs
Kyoung Mu Lee |
ICLR | 2 |
| 2023 | Dense Depth-Guided Generalizable NeRFabstractNeural rendering approaches enable photo-realistic rendering on novel view synthesis tasks while their per-scene optimization remains an issue for scalability. Recent methods introduce novel neural radiance field (NeRF) frameworks that generalize to unseen scenes on-the-fly by combining multi-view stereo with differentiable volume rendering. These generalizable NeRF methods synthesize the colors of 3D ray points by learning the consistency of image features projected from given nearby views. Since the consistency is computed on the 2D projected image space, it is vulnerable to occlusion and local shape variation by viewing direction. To solve this problem, we present dense depth-guided generalizable NeRF that leverages the depth as the signed distance between the ray point and the object surface of the scene. We first generate the dense depth maps from sparse 3D points of structure from motion (SfM) which is an inevitable step to obtain camera poses. Next, the dense depth maps are exploited as complementary features invariant to the sparsity of nearby views and mask for occlusion handling. Experiments demonstrate that our approach outperforms existing generalizable NeRF methods for widely used real and synthetic datasets. Kyoung Mu Lee |
IEEE Signal Process. Lett. | 2 |
| 2022 | MonoNHR: Monocular Neural Human RendererabstractExisting neural human rendering methods struggle with a single image input due to the lack of information in in-visible areas and the depth ambiguity of pixels in visible areas. In this regard, we propose Monocular Neural Human Renderer (MonoNHR), a novel approach that renders robust free-viewpoint images of an arbitrary human given only a single image. MonoNHR is the first method that (i) renders human subjects never seen during training in a monocular setup, and (ii) is trained in a weakly-supervised manner without geometry supervision. First, we propose to disentangle 3D geometry and texture features and to condition the texture inference on the 3D geometry features. Second, we introduce a Mesh Inpainter module that inpaints the occluded parts exploiting human structural priors such as symmetry. Experiments on ZJU-MoCap, AIST and HUMBI datasets show that our approach significantly outperforms the recent methods adapted to the monocular case. Hongsuk Choi, Gyeongsik Moon, Matthieu Armando, Vincent Leroy 0003, Kyoung Mu Lee, Grégory Rogez |
3DV | 5 |
| 2022 | Learning to Estimate Robust 3D Human Mesh from In-the-Wild Crowded ScenesabstractWe consider the problem of recovering a single person's 3D human mesh from in-the-wild crowded scenes. While much progress has been in 3D human mesh estimation, existing methods struggle when test input has crowded scenes. The first reason for the failure is a domain gap between training and testing data. A motion capture dataset, which provides accurate 3D labels for training, lacks crowd data and impedes a network from learning crowded scene-robust image features of a target person. The second reason is a feature processing that spatially averages the feature map of a localized bounding box containing multiple people. Averaging the whole feature map makes a target person's feature indistinguishable from others. We present 3DCrowdNet that firstly explicitly targets in-the-wild crowded scenes and estimates a robust 3D human mesh by addressing the above issues. First, we leverage 2D human pose estimation that does not require a motion capture dataset with 3D labels for training and does not suffer from the domain gap. Second, we propose a joint-based regressor that distinguishes a target person's feature from others. Our joint-based regressor preserves the spatial activation of a target by sampling features from the target's joint locations and regresses human model parameters. As a result, 3DCrowdNet learns target-focused features and effectively excludes the irrelevant features of nearby persons. We conduct experiments on various benchmarks and prove the robustness of 3D CrowdNet to the in-the-wild crowded scenes both quantitatively and qualitatively. Codes are available here11https://github.com/hongsukchoi/3DCrowdNet_RELEASE. Hongsuk Choi, Gyeongsik Moon, Joonkyu Park, Kyoung Mu Lee |
CVPR | 4 |
| 2022 | AP-BSN: Self-Supervised Denoising for Real-World Images via Asymmetric PD and Blind-Spot NetworkabstractBlind-spot network (BSN) and its variants have made significant advances in self-supervised denoising. Never-theless, they are still bound to synthetic noisy inputs due to less practical assumptions like pixel-wise independent noise. Hence, it is challenging to deal with spatially corre-lated real-world noise using self-supervised BSN. Recently, pixel-shuffle downsampling (PD) has been proposed to re-move the spatial correlation of real-world noise. However, it is not trivial to integrate PD and BSN directly, which prevents the fully self-supervised denoising model on real-world images. We propose an Asymmetric PD (AP) to ad-dress this issue, which introduces different P D stride factors for training and inference. We systematically demonstrate that the proposed AP can resolve inherent trade-offs caused by specific PD stride factors and make BSN applicable to practical scenarios. To this end, we develop AP-BSN, a state-of-the-art self-supervised denoising method for real-world sRGB images. We further propose random-replacing refinement, which significantly improves the performance of our AP-BSN without any additional parameters. Extensive studies demonstrate that our method outperforms the other self-supervised and even unpaired denoising methods by a large margin, without using any additional knowledge, e.g., noise level, regarding the underlying unknown noise. Wooseok Lee, Sanghyun Son 0002, Kyoung Mu Lee |
CVPR | 3 |
| 2022 | CVF-SID: Cyclic multi-Variate Function for Self-Supervised Image Denoising by Disentangling Noise from ImageabstractRecently, significant progress has been made on image denoising with strong supervision from large-scale datasets. However, obtaining well-aligned noisy-clean training image pairs for each specific scenario is complicated and costly in practice. Consequently, applying a conventional supervised denoising network on in-the-wild noisy inputs is not straightforward. Although several studies have challenged this problem without strong supervision, they rely on less practical assumptions and cannot be applied to practical situations directly. To address the aforementioned challenges, we propose a novel and powerful self-supervised denoising method called CVF-SID based on a Cyclic multi-Variate Function (CVF) module and a self-supervised image disentangling (SID) framework. The CVF module can output multiple decomposed variables of the input and take a combination of the outputs back as an input in a cyclic manner. Our CVF-SID can disentangle a clean image and noise maps from the input by leveraging various self-supervised loss terms. Unlike several methods that only consider the signal-independent noise models, we also deal with signal-dependent noise components for real-world applications. Furthermore, we do not rely on any prior assumptions about the underlying noise distribution, making CVF-SID more generalizable toward realistic noise. Extensive experiments on real-world datasets show that CVF-SID achieves state-of-the-art self-supervised image denoising performance and is comparable to other existing approaches. The code is publicly available from this link. Reyhaneh Neshatavar, Mohsen Yavartanoo, Sanghyun Son 0002, Kyoung Mu Lee |
CVPR | 4 |
| 2022 | Attentive Fine-Grained Structured Sparsity for Image RestorationabstractImage restoration tasks have witnessed great performance improvement in recent years by developing large deep models. Despite the outstanding performance, the heavy computation demanded by the deep models has restricted the application of image restoration. To lift the restriction, it is required to reduce the size of the networks while maintaining accuracy. Recently, N:M structured pruning has appeared as one of the effective and practical pruning approaches for making the model efficient with the accuracy constraint. However, it fails to account for different computational complexities and performance requirements for different layers of an image restoration network. To further optimize the trade-off between the efficiency and the restoration accuracy, we propose a novel pruning method that determines the pruning ratio for N:M structured sparsity at each layer. Extensive experimental results on super-resolution and deblurring tasks demonstrate the efficacy of our method which outperforms previous pruning methods significantly. PyTorch implementation for the proposed methods will be publicly available at https://github.com/JungHunOh/SLS_CVPR2022 Junghun Oh, Seungjun Nah, Cheeun Hong, Kyoung Mu Lee |
CVPR | 6 |
| 2022 | HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkabstractHands are often severely occluded by objects, which makes 3D hand mesh estimation challenging. Previous works often have disregarded information at occluded regions. However, we argue that occluded regions have strong correlations with hands so that they can provide highly beneficial information for complete 3D hand mesh estimation. Thus, in this work, we propose a novel 3D hand mesh estimation network HandOccNet, that can fully exploits the information at occluded regions as a secondary means to enhance image features and make it much richer. To this end, we design two successive Transformer-based modules, called feature injecting transformer (FIT) and self-enhancing transformer (SET). FIT injects hand information into occluded region by considering their correlation. SET refines the output of FIT by using a self-attention mechanism. By injecting the hand information to the occluded region, our HandOccNet reaches the state-of-the-art performance on 3D hand mesh benchmarks that contain challenging hand-object occlusions. The codes are available in: https://github.com/namepllet/HandOccNet. Joonkyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi, Kyoung Mu Lee |
CVPR | 5 |
| 2022 | CADyQ: Content-Aware Dynamic Quantization for Image Super-Resolution
Cheeun Hong, Sungyong Baik, Seungjun Nah, Kyoung Mu Lee |
ECCV (7) | 5 |
| 2022 | 3D Clothed Human Reconstruction in the Wild
Gyeongsik Moon, Hyeongjin Nam, Takaaki Shiratori, Kyoung Mu Lee |
ECCV (2) | 4 |
| 2022 | Clean Images are Hard to Reblur: Exploiting the Ill-Posed Inverse Task for Dynamic Scene Deblurring
Seungjun Nah, Sanghyun Son 0002, Jaerin Lee, Kyoung Mu Lee |
ICLR | 4 |
| 2022 | DAQ: Channel-Wise Distribution-Aware Quantization for Deep Image Super-Resolution NetworksabstractSince the resurgence of deep neural networks (DNNs), image super-resolution (SR) has recently seen a huge progress in improving the quality of low resolution images, however at the great cost of computations and resources. Recently, there has been several efforts to make DNNs more efficient via quantization. However, SR demands pixel-level accuracy in the system, it is more difficult to perform quantization without significantly sacrificing SR performance. To this end, we introduce a new ultra-low precision yet effective quantization approach specifically designed for SR. In particular, we observe that in recent SR networks, each channel has different distribution characteristics. Thus we propose a channel-wise distribution-aware quantization scheme. Experimental results demonstrate that our proposed quantization, dubbed Distribution-Aware Quantization (DAQ), manages to greatly reduce the computational and resource costs without the significant sacrifice in SR performance, compared to other quantization methods. Cheeun Hong, Sungyong Baik, Junghun Oh, Kyoung Mu Lee |
WACV | 5 |
| 2022 | Batch Normalization Tells You Which Filter is ImportantabstractThe goal of filter pruning is to search for unimportant filters to remove in order to make convolutional neural networks (CNNs) efficient without sacrificing the performance in the process. The challenge lies in finding information that can help determine how important or relevant each filter is with respect to the final output of neural networks. In this work, we share our observation that the batch normalization (BN) parameters of pre-trained CNNs can be used to estimate the feature distribution of activation outputs, without processing of training data. Upon observation, we propose a simple yet effective filter pruning method by evaluating the importance of each filter based on the BN parameters of pre-trained CNNs. The experimental results on CIFAR-10 and ImageNet demonstrate that the proposed method can achieve outstanding performance with and without fine-tuning in terms of the trade-off between the accuracy drop and the reduction in computational complexity and number of parameters of pruned networks. Junghun Oh, Sungyong Baik, Cheeun Hong, Kyoung Mu Lee |
WACV | 5 |
| 2022 | Fine-grained neural architecture search for image super-resolution
Seokil Hong, Bohyung Han, Heesoo Myeong, Kyoung Mu Lee |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | Learning to Forget for Meta-Learning via Task-and-Layer-Wise AttenuationabstractFew-shot learning is an emerging yet challenging problem in which the goal is to achieve generalization from only few examples. Meta-learning tackles few-shot learning via the learning of prior knowledge shared across tasks and using it to learn new tasks. One of the most representative meta-learning algorithms is the model-agnostic meta-learning (MAML), which formulates prior knowledge as a common initialization, a shared starting point from where a learner can quickly adapt to unseen tasks. However, forcibly sharing an initialization can lead to conflicts among tasks and the compromised (undesired by tasks) location on optimization landscape, thereby hindering task adaptation. Furthermore, the degree of conflict is observed to vary not only among the tasks but also among the layers of a neural network. Thus, we propose task-and-layer-wise attenuation on the compromised initialization to reduce its adverse influence on task adaptation. As attenuation dynamically controls (or selectively forgets) the influence of the compromised prior knowledge for a given task and each layer, we name our method Learn to Forget (L2F). Experimental results demonstrate that the proposed method greatly improves the performance of the state-of-the-art MAML-based frameworks across diverse domains: few-shot classification, cross-domain few-shot classification, regression, reinforcement learning, and visual tracking. Sungyong Baik, Junghoon Oh, Seokil Hong, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Test-Time Adaptation for Video Frame Interpolation via Meta-LearningabstractVideo frame interpolation is a challenging problem that involves various scenarios depending on the variety of foreground and background motions, frame rate, and occlusion. Therefore, generalizing across different scenes is difficult for a single network with fixed parameters. Ideally, one could have a different network for each scenario, but this will be computationally infeasible for practical applications. In this work, we propose MetaVFI, an adaptive video frame interpolation algorithm that uses additional information readily available at test time but has not been exploited in previous works. We initially show the benefits of test-time adaptation through simple fine-tuning of a network and then greatly improve its efficiency by incorporating meta-learning. Thus, we obtain significant performance gains with only a single gradient update without introducing any additional parameters. Moreover, the proposed MetaVFI algorithm is model-agnostic which can be easily combined with any video frame interpolation network. We show that our adaptive framework greatly improves the performance of baseline video frame interpolation networks on multiple benchmark datasets. Myungsub Choi, Janghoon Choi, Sungyong Baik, Tae Hyun Kim 0006, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Editorial
Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Toward Real-World Super-Resolution via Adaptive Downsampling ModelsabstractMost image super-resolution (SR) methods are developed on synthetic low-resolution (LR) and high-resolution (HR) image pairs that are constructed by a predetermined operation, e.g., bicubic downsampling. As existing methods typically learn an inverse mapping of the specific function, they produce blurry results when applied to real-world images whose exact formulation is different and unknown. Therefore, several methods attempt to synthesize much more diverse LR samples or learn a realistic downsampling model. However, due to restrictive assumptions on the downsampling process, they are still biased and less generalizable. This study proposes a novel method to simulate an unknown downsampling process without imposing restrictive prior knowledge. We propose a generalizable low-frequency loss (LFL) in the adversarial training framework to imitate the distribution of target LR images without using any paired examples. Furthermore, we design an adaptive data loss (ADL) for the downsampler, which can be adaptively learned and updated from the data during the training loops. Extensive experiments validate that our downsampling model can facilitate existing SR methods to perform more accurate reconstructions on various synthetic and real-world examples than the conventional approaches. Sanghyun Son 0002, Jaeha Kim, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | PolyNet: Polynomial Neural Network for 3D Shape Recognition with PolyShape Representationabstract3D shape representation and its processing have substantial effects on 3D shape recognition. The polygon mesh as a 3D shape representation has many advantages in computer graphics and geometry processing. However, there are still some challenges for the existing deep neural network (DNN)-based methods on polygon mesh representation, such as handling the variations in the degree and permutations of the vertices and their pairwise distances. To overcome these challenges, we propose a DNN-based method (PolyNet) and a specific polygon mesh representation (PolyShape) with a multi-resolution structure. PolyNet contains two operations; (1) a polynomial convolution (PolyConv) operation with learnable coefficients, which learns continuous distributions as the convolutional filters to share the weights across different vertices, and (2) a polygonal pooling (PolyPool) procedure by utilizing the multi-resolution structure of PolyShape to aggregate the features in a much lower dimension. Our experiments demonstrate the strength and the advantages of PolyNet on both 3D shape classification and retrieval tasks compared to existing polygon mesh-based methods and its superiority in classifying graph representations of images. The code is publicly available from this link. Mohsen Yavartanoo, Shih-Hsuan Hung, Reyhaneh Neshatavar, Yue Zhang 0009, Kyoung Mu Lee |
3DV | 5 |
| 2021 | Recurrence-in-Recurrence Networks for Video Deblurring
Joonkyu Park, Seungjun Nah, Kyoung Mu Lee |
BMVC | 3 |
| 2021 | Beyond Static Features for Temporally Consistent 3D Human Pose and Shape From a VideoabstractDespite the recent success of single image-based 3D human pose and shape estimation methods, recovering temporally consistent and smooth 3D human motion from a video is still challenging. Several video-based methods have been proposed; however, they fail to resolve the single image-based methods’ temporal inconsistency issue due to a strong dependency on a static feature of the current frame. In this regard, we present a temporally consistent mesh recovery system (TCMR). It effectively focuses on the past and future frames’ temporal information without being dominated by the current static feature. Our TCMR significantly outperforms previous video-based methods in temporal consistency with better per-frame 3D pose and shape accuracy. We also release the codes. Hongsuk Choi, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 4 |
| 2021 | SRWarp: Generalized Image Super-Resolution under Arbitrary TransformationabstractDeep CNNs have achieved significant successes in image processing and its applications, including single image super-resolution (SR). However, conventional methods still resort to some predetermined integer scaling factors, e.g., ×2 or ×4. Thus, they are difficult to be applied when arbitrary target resolutions are required. Recent approaches ex-tend the scope to real-valued upsampling factors, even with varying aspect ratios to handle the limitation. In this pa-per, we propose the SRWarp framework to further generalize the SR tasks toward an arbitrary image transformation. We interpret the traditional image warping task, specifically when the input is enlarged, as a spatially-varying SR problem. We also propose several novel formulations, including the adaptive warping layer and multiscale blending, to reconstruct visually favorable results in the transformation process. Compared with previous methods, we do not con-strain the SR model on a regular grid but allow numerous possible deformations for flexible and diverse image editing. Extensive experiments and ablation studies justify the necessity and demonstrate the advantage of the proposed SRWarp method under various transformations. Sanghyun Son 0002, Kyoung Mu Lee |
CVPR | 2 |
| 2021 | Meta-Learning with Task-Adaptive Loss Function for Few-Shot LearningabstractIn few-shot learning scenarios, the challenge is to generalize and perform well on new unseen examples when only very few labeled examples are available for each task. Model-agnostic meta-learning (MAML) has gained the popularity as one of the representative few-shot learning methods for its flexibility and applicability to diverse problems. However, MAML and its variants often resort to a simple loss function without any auxiliary loss function or regularization terms that can help achieve better generalization. The problem lies in that each application and task may require different auxiliary loss function, especially when tasks are diverse and distinct. Instead of attempting to hand-design an auxiliary loss function for each application and task, we introduce a new meta-learning framework with a loss function that adapts to each task. Our proposed framework, named Meta-Learning with Task-Adaptive Loss Function (MeTAL), demonstrates the effectiveness and the flexibility across various domains, such as few-shot classification and few-shot regression. Sungyong Baik, Janghoon Choi, Dohee Cho, Jaesik Min, Kyoung Mu Lee |
ICCV | 6 |
| 2021 | Motion-Aware Dynamic Architecture for Efficient Frame InterpolationabstractVideo frame interpolation aims to synthesize accurate intermediate frames given a low-frame-rate video. While the quality of the generated frames is increasingly getting better, state-of-the-art models have become more and more computationally expensive. However, local regions with small or no motion can be easily interpolated with simple models and do not require such heavy compute, whereas some regions may not be correct even after inference through a large model. Thus, we propose an effective framework that assigns varying amounts of computation for different regions. Our dynamic architecture first calculates the approximate motion magnitude to use as a proxy for the difficulty levels for each region, and decides the depth of the model and the scale of the input. Experimental results show that static regions pass through a smaller number of layers, while the regions with larger motion are downscaled for better motion reasoning. In doing so, we demonstrate that the proposed framework can significantly reduce the computation cost (FLOPs) while maintaining the performance, often up to 50% when interpolating a 2K resolution video. Myungsub Choi, Suyoung Lee, Kyoung Mu Lee |
ICCV | 4 |
| 2021 | C2N: Practical Generative Noise Modeling for Real-World DenoisingabstractLearning-based image denoising methods have been bounded to situations where well-aligned noisy and clean images are given, or samples are synthesized from predetermined noise models, e.g., Gaussian. While recent generative noise modeling methods aim to simulate the unknown distribution of real-world noise, several limitations still exist. In a practical scenario, a noise generator should learn to simulate the general and complex noise distribution without using paired noisy and clean images. However, since existing methods are constructed on the unrealistic assumption of real-world noise, they tend to generate implausible patterns and cannot express complicated noise maps. Therefore, we introduce a Clean-to-Noisy image generation framework, namely C2N, to imitate complex real-world noise without using any paired examples. We construct the noise generator in C2N accordingly with each component of real-world noise characteristics to express a wide range of noise accurately. Combined with our C2N, conventional denoising CNNs can be trained to outperform existing unsupervised methods on challenging real-world benchmarks by a large margin. Geonwoon Jang, Wooseok Lee, Sanghyun Son 0002, Kyoung Mu Lee |
ICCV | 4 |
| 2021 | Searching for Controllable Image Restoration NetworksabstractWe present a novel framework for controllable image restoration that can effectively restore multiple types and levels of degradation of a corrupted image. The proposed model, named TASNet, is automatically determined by our neural architecture search algorithm, which optimizes the efficiency-accuracy trade-off of the candidate model architectures. Specifically, we allow TASNet to share the early layers across different restoration tasks and adaptively adjust the remaining layers with respect to each task. The shared task-agnostic layers greatly improve the efficiency while the task-specific layers are optimized for restoration quality, and our search algorithm seeks for the best balance between the two. We also propose a new data sampling strategy to further improve the overall restoration performance. As a result, TASNet achieves significantly faster GPU latency and lower FLOPs compared to the existing state-of-the-art models, while also showing visually more pleasing outputs. The source code and pre-trained models are available at https://github.com/ghimhw/TASNet. Sungyong Baik, Myungsub Choi, Janghoon Choi, Kyoung Mu Lee |
ICCV | 5 |
| 2021 | 3DIAS: 3D Shape Reconstruction with Implicit Algebraic Surfacesabstract3D Shape representation has substantial effects on 3D shape reconstruction. Primitive-based representations approximate a 3D shape mainly by a set of simple implicit primitives, but the low geometrical complexity of the primitives limits the shape resolution. Moreover, setting a sufficient number of primitives for an arbitrary shape is challenging. To overcome these issues, we propose a constrained implicit algebraic surface as the primitive with few learnable coefficients and higher geometrical complexities and a deep neural network to produce these primitives. Our experiments demonstrate the superiorities of our method in terms of representation power compared to the state-of-the-art methods in single RGB image 3D shape reconstruction. Furthermore, we show that our method can semantically learn segments of 3D shapes in an unsupervised manner. The code is publicly available from this link. Mohsen Yavartanoo, Jaeyoung Chung 0002, Reyhaneh Neshatavar, Kyoung Mu Lee |
ICCV | 4 |
| 2021 | Structure-Resonant Discriminator for Image Super-ResolutionabstractConvolutional neural networks are data models. Their design should embrace the structural properties of the data being modeled, e.g., the natural images. We argue that this also holds for discriminators of adversarial training frameworks for photo-realistic image restoration. We develop this idea to highlight three essential structural features of natural images: translation equivariance, rotation invariance, and hierarchy of scale. The analysis leads to a new discriminator, Structure-Resonant Discriminator (SRD), which can capture image structures in need. The proposed SRD is demonstrated in the perceptual single image super-resolution task. By replacing only the discriminator, our method restores more visually pleasing high-resolution images than the previous state-of-the-art techniques, while exhibits the least distortions. Jaerin Lee, Kyoung Mu Lee |
ICME | 2 |
| 2021 | DynaVSR: Dynamic Adaptive Blind Video Super-ResolutionabstractMost conventional supervised super-resolution (SR) algorithms assume that low-resolution (LR) data is obtained by downscaling high-resolution (HR) data with a fixed known kernel, but such an assumption often does not hold in real scenarios. Some recent blind SR algorithms have been proposed to estimate different downscaling kernels for each input LR image. However, they suffer from heavy computational overhead, making them infeasible for direct application to videos. In this work, we present DynaVSR, a novel meta-learning-based framework for real-world video SR that enables efficient downscaling model estimation and adaptation to the current input. Specifically, we train a multi-frame downscaling module with various types of synthetic blur kernels, which is seamlessly combined with a video SR network for input-aware adaptation. Experimental results show that DynaVSR consistently improves the performance of the state-of-the-art video SR models by a large margin, with an order of magnitude faster inference time compared to the existing blind SR approaches. Suyoung Lee, Myungsub Choi, Kyoung Mu Lee |
WACV | 3 |
| 2020 | Channel Attention Is All You Need for Video Frame InterpolationabstractPrevailing video frame interpolation techniques rely heavily on optical flow estimation and require additional model complexity and computational cost; it is also susceptible to error propagation in challenging scenarios with large motion and heavy occlusion. To alleviate the limitation, we propose a simple but effective deep neural network for video frame interpolation, which is end-to-end trainable and is free from a motion estimation network component. Our algorithm employs a special feature reshaping operation, referred to as PixelShuffle, with a channel attention, which replaces the optical flow computation module. The main idea behind the design is to distribute the information in a feature map into multiple channels and extract motion information by attending the channels for pixel-level frame synthesis. The model given by this principle turns out to be effective in the presence of challenging motion and occlusion. We construct a comprehensive evaluation benchmark and demonstrate that the proposed approach achieves outstanding performance compared to the existing models with a component for optical flow computation. Myungsub Choi, Bohyung Han, Kyoung Mu Lee |
AAAI | 5 |
| 2020 | Visual Tracking by TridentAlign and Context Embedding
Janghoon Choi, Junseok Kwon, Kyoung Mu Lee |
ACCV (2) | 3 |
| 2020 | Domain Adaptation of Learned Featuresfor Visual Localization
Sungyong Baik, Hyo Jin Kim 0004, Tianwei Shen, Eddy Ilg, Kyoung Mu Lee, Chris Sweeney |
BMVC | 5 |
| 2020 | Learning to Forget for Meta-LearningabstractFew-shot learning is a challenging problem where the goal is to achieve generalization from only few examples. Model-agnostic meta-learning (MAML) tackles the problem by formulating prior knowledge as a common initialization across tasks, which is then used to quickly adapt to unseen tasks. However, forcibly sharing an initialization can lead to conflicts among tasks and the compromised (undesired by tasks) location on optimization landscape, thereby hindering the task adaptation. Further, we observe that the degree of conflict differs among not only tasks but also layers of a neural network. Thus, we propose task-and-layer-wise attenuation on the compromised initialization to reduce its influence. As the attenuation dynamically controls (or selectively forgets) the influence of prior knowledge for a given task and each layer, we name our method as L2F (Learn to Forget). The experimental results demonstrate that the proposed method provides faster adaptation and greatly improves the performance. Furthermore, L2F can be easily applied and improve other state-of-the-art MAML-based frameworks, illustrating its simplicity and generalizability. Sungyong Baik, Seokil Hong, Kyoung Mu Lee |
CVPR | 3 |
| 2020 | Scene-Adaptive Video Frame Interpolation via Meta-LearningabstractVideo frame interpolation is a challenging problem because there are different scenarios for each video depending on the variety of foreground and background motion, frame rate, and occlusion. It is therefore difficult for a single network with fixed parameters to generalize across different videos. Ideally, one could have a different network for each scenario, but this is computationally infeasible for practical applications. In this work, we propose to adapt the model to each video by making use of additional information that is readily available at test time and yet has not been exploited in previous works. We first show the benefits of 'test-time adaptation' through simple fine-tuning of a network, then we greatly improve its efficiency by incorporating meta-learning. We obtain significant performance gains with only a single gradient update without any additional parameters. Finally, we show that our meta-learning framework can be easily employed to any video frame interpolation network and can consistently improve its performance on multiple benchmark datasets. Myungsub Choi, Janghoon Choi, Sungyong Baik, Tae Hyun Kim 0006, Kyoung Mu Lee |
CVPR | 5 |
| 2020 | Pose2Mesh: Graph Convolutional Network for 3D Human Pose and Mesh Recovery from a 2D Human Pose
Hongsuk Choi, Gyeongsik Moon, Kyoung Mu Lee |
ECCV (7) | 3 |
| 2020 | I2L-MeshNet: Image-to-Lixel Prediction Network for Accurate 3D Human Pose and Mesh Estimation from a Single RGB Image
Gyeongsik Moon, Kyoung Mu Lee |
ECCV (7) | 2 |
| 2020 | DeepHandMesh: A Weakly-Supervised Deep Encoder-Decoder Framework for High-Fidelity Hand Mesh Modeling
Gyeongsik Moon, Takaaki Shiratori, Kyoung Mu Lee |
ECCV (2) | 3 |
| 2020 | InterHand2.6M: A Dataset and Baseline for 3D Interacting Hand Pose Estimation from a Single RGB Image
Gyeongsik Moon, Shoou-I Yu, He Wen 0001, Takaaki Shiratori, Kyoung Mu Lee |
ECCV (20) | 5 |
| 2020 | Multi Image Depth from Defocus Network with Boundary Cue for Dual Aperture CameraabstractIn this paper, we estimate depth information using two defocused images from dual aperture camera. Recent advances in deep learning techniques have increased the accuracy of depth estimation. Besides, methods of using a defocused image in which an object is blurred according to a distance from a camera have been widely studied. We further improve the accuracy of the depth estimation by training the network using two images with different degrees of depth-of-field. Using images taken with different apertures for the same scene, we can determine the degree of blur in an image more accurately. In this work, we propose a novel deep convolutional network that estimates depth map using dual aperture images based on boundary cue. Our proposed method achieves state-of-the-art performance on a synthetically modified NYU-v2 dataset. In addition, we built a new camera using fast variable apertures to build a test environment in the real world. In particular, we collected a new dataset which consists of real world vehicle driving scenes. Our proposed work shows excellent performance in the new dataset. Gwangmo Song, Yumee Kim, Kukjin Chun, Kyoung Mu Lee |
ICASSP | 4 |
| 2020 | Meta-Learning with Adaptive HyperparametersabstractDespite its popularity, several recent works question the effectiveness of MAML when test tasks are different from training tasks, thus suggesting various task-conditioned methodology to improve the initialization. Instead of searching for better task-aware initialization, we focus on a complementary factor in MAML framework, inner-loop optimization (or fast adaptation). Consequently, we propose a new weight update rule that greatly enhances the fast adaptation process. Specifically, we introduce a small meta-network that can adaptively generate per-step hyperparameters: learning rate and weight decay coefficients. The experimental results validate that the Adaptive Learning of hyperparameters for Fast Adaptation (ALFA) is the equally important ingredient that was often neglected in the recent few-shot learning approaches. Surprisingly, fast adaptation from random initialization with ALFA can already outperform MAML. Sungyong Baik, Myungsub Choi, Janghoon Choi, Kyoung Mu Lee |
NeurIPS | 5 |
| 2020 | Triplanar convolution with shared 2D kernels for 3D classification and shape retrieval
Euyoung Kim, Seung Yeon Shin, Soochahn Lee, Kyong Joon Lee, Kyoung Ho Lee, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 6 |
| 2020 | Bi-Directional Seed Attention Network for Interactive Image SegmentationabstractIn interactive segmentation, the role of seed information provided by the user is significant. A seed is a clue to ease the ambiguity of the problem by making the object segmentation task interactive. However, in most deep network-based works, seed information has been used as an additional channel for input images. In this paper, we propose a novel bi-directional attention module for more actively using seed information. The proposed bi-directional seed attention module (BSA) operates based on the feature map of the segmentation network and the input seed map. Through our attention module, the network feature map is affected by the seed map, while the feature also updates the seed information. As a result, our system concentrates on the seed information and more accurately derives the segmentation result required by the user. We have conducted validation experiments on the four standard benchmark datasets, including SBD, GrabCut, Berkeley, and DAVIS, and achieved the state-of-the-art results. Gwangmo Song, Kyoung Mu Lee |
IEEE Signal Process. Lett. | 2 |
| 2019 | PoseFix: Model-Agnostic General Human Pose Refinement NetworkabstractMulti-person pose estimation from a 2D image is an essential technique for human behavior understanding. In this paper, we propose a human pose refinement network that estimates a refined pose from a tuple of an input image and input pose. The pose refinement was performed mainly through an end-to-end trainable multi-stage architecture in previous methods. However, they are highly dependent on pose estimation models and require careful model design. By contrast, we propose a model-agnostic pose refinement method. According to a recent study, state-of-the-art 2D human pose estimation methods have similar error distributions. We use this error statistics as prior information to generate synthetic poses and use the synthesized poses to train our model. In the testing stage, pose estimation results of any other methods can be input to the proposed method. Moreover, the proposed model does not require code or knowledge about other methods, which allows it to be easily used in the post-processing step. We show that the proposed approach achieves better performance than the conventional multi-stage refinement models and consistently improves the performance of various state-of-the-art pose estimation methods on the commonly used benchmark. The code is available in. Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 3 |
| 2019 | Recurrent Neural Networks With Intra-Frame Iterations for Video DeblurringabstractRecurrent neural networks (RNNs) are widely used for sequential data processing. Recent state-of-the-art video deblurring methods bank on convolutional recurrent neural network architectures to exploit the temporal relationship between neighboring frames. In this work, we aim to improve the accuracy of recurrent models by adapting the hidden states transferred from past frames to the frame being processed so that the relations between video frames could be better used. We iteratively update the hidden state via re-using RNN cell parameters before predicting an output deblurred frame. Since we use existing parameters to update the hidden state, our method improves accuracy without additional modules. As the architecture remains the same regardless of iteration number, fewer iteration models can be considered as a partial computational path of the models with more iterations. To take advantage of this property, we employ a stochastic method to optimize our iterative models better. At training time, we randomly choose the iteration number on the fly and apply a regularization loss that favors less computation unless there are considerable reconstruction gains. We show that our method exhibits state-of-the-art video deblurring performance while operating in real-time speed. Seungjun Nah, Sanghyun Son 0002, Kyoung Mu Lee |
CVPR | 3 |
| 2019 | Stochastic Class-Based Hard Example Mining for Deep Metric LearningabstractPerformance of deep metric learning depends heavily on the capability of mining hard negative examples during training. However, many metric learning algorithms often require intractable computational cost due to frequent feature computations and nearest neighbor searches in a large-scale dataset. As a result, existing approaches often suffer from trade-off between training speed and prediction accuracy. To alleviate this limitation, we propose a stochastic hard negative mining method. Our key idea is to adopt class signatures that keep track of feature embedding online with minor additional cost during training, and identify hard negative example candidates using the signatures. Given an anchor instance, our algorithm first selects a few hard negative classes based on the class-to-sample distances and then performs a refined search in an instance-level only from the selected classes. As most of the classes are discarded at the first step, it is much more efficient than exhaustive search while effectively mining a large number of hard examples. Our experiment shows that the proposed technique improves image retrieval accuracy substantially; it achieves the state-of-the-art performance on the several standard benchmark datasets. Yumin Suh, Bohyung Han, Wonsik Kim, Kyoung Mu Lee |
CVPR | 4 |
| 2019 | Deep Meta Learning for Real-Time Target-Aware Visual TrackingabstractIn this paper, we propose a novel on-line visual tracking framework based on the Siamese matching network and meta-learner network, which run at real-time speeds. Conventional deep convolutional feature-based discriminative visual tracking algorithms require continuous re-training of classifiers or correlation filters, which involve solving complex optimization tasks to adapt to the new appearance of a target object. To alleviate this complex process, our proposed algorithm incorporates and utilizes a meta-learner network to provide the matching network with new appearance information of the target objects by adding target-aware feature space. The parameters for the target-specific feature space are provided instantly from a single forward-pass of the meta-learner network. By eliminating the necessity of continuously solving complex optimization tasks in the course of tracking, experimental results demonstrate that our algorithm performs at a real-time speed while maintaining competitive performance among other state-of-the-art tracking algorithms. Janghoon Choi, Junseok Kwon, Kyoung Mu Lee |
ICCV | 3 |
| 2019 | Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageabstractAlthough significant improvement has been achieved recently in 3D human pose estimation, most of the previous methods only treat a single-person case. In this work, we firstly propose a fully learning-based, camera distance-aware top-down approach for 3D multi-person pose estimation from a single RGB image. The pipeline of the proposed system consists of human detection, absolute 3D human root localization, and root-relative 3D single-person pose estimation modules. Our system achieves comparable results with the state-of-the-art 3D single-person pose estimation models without any ground truth information and significantly outperforms previous 3D multi-person pose estimation methods on publicly available datasets. The code is available in1,2. Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
ICCV | 3 |
| 2019 | Continual Learning by Asymmetric Loss Approximation With Single-Side OverestimationabstractCatastrophic forgetting is a critical challenge in training deep neural networks. Although continual learning has been investigated as a countermeasure to the problem, it often suffers from the requirements of additional network components and the limited scalability to a large number of tasks. We propose a novel approach to continual learning by approximating a true loss function using an asymmetric quadratic function with one of its sides overestimated. Our algorithm is motivated by the empirical observation that the network parameter updates affect the target loss functions asymmetrically. In the proposed continual learning framework, we estimate an asymmetric loss function for the tasks considered in the past through a proper overestimation of its unobserved sides in training new tasks, while deriving the accurate model parameter for the observable sides. In contrast to existing approaches, our method is free from the side effects and achieves the state-of-the-art accuracy that is even close to the upper-bound performance on several challenging benchmark datasets. Dongmin Park, Seokil Hong, Bohyung Han, Kyoung Mu Lee |
ICCV | 4 |
| 2019 | Learning to Remember Past to Predict Future for Visual TrackingabstractFast and reliable adaptability to appearance variations of any target object has been the holy grail of visual tracking. Recently, Siamese-based trackers have demonstrated outstanding speed, however at the cost of adaptability and accuracy. We propose to model a temporal evolution of appearance features, allowing for adaptability without online training. Specifically, we introduce a memory-augmented convolutional recurrent neural network (RNN), named Past-to-Future (P2FNet), that takes appearance features as an input at each frame and predicts the next-frame features. RNN allows for fast adaptability to dynamically varying appearance, while the memory provides the generalization capability over longer sequences via template management. For reliability, we propose a new augmentation to train RNN to disregard corrupted features. A novel visualization method illustrates the reliable template management of the memory. The experimental results on benchmarks demonstrate the tracker shows competitive performance among real-time state-of-the-art trackers. Sungyong Baik, Junseok Kwon, Kyoung Mu Lee |
ICIP | 3 |
| 2019 | Deep vessel segmentation by learning graphical connectivity
Seung Yeon Shin, Soochahn Lee, Il Dong Yun, Kyoung Mu Lee |
Medical Image Anal. | 4 |
| 2019 | Joint Weakly and Semi-Supervised Deep Learning for Localization and Classification of Masses in Breast Ultrasound ImagesabstractWe propose a framework for localization and classification of masses in breast ultrasound images. We have experimentally found that training convolutional neural network-based mass detectors with large, weakly annotated datasets presents a non-trivial problem, while overfitting may occur with those trained with small, strongly annotated datasets. To overcome these problems, we use a weakly annotated dataset together with a smaller strongly annotated dataset in a hybrid manner. We propose a systematic weakly and semi-supervised training scenario with appropriate training loss selection. Experimental results show that the proposed method can successfully localize and classify masses with less annotation effort. The results trained with only 10 strongly annotated images along with weakly annotated images were comparable to results trained from 800 strongly annotated images, with the 95% confidence interval (CI) of difference -3%-5%, in terms of the correct localization (CorLoc) measure, which is the ratio of images with intersection over union with ground truth higher than 0.5. With the same number of strongly annotated images, additional weakly annotated images can be incorporated to give a 4.5% point increase in CorLoc, from 80% to 84.50% (with 95% CIs 76%-83.75% and 81%-88%). The effects of different algorithmic details and varied amount of data are presented through ablative analysis. Seung Yeon Shin, Soochahn Lee, Il Dong Yun, Sun Mi Kim, Kyoung Mu Lee |
IEEE Trans. Medical Imaging | 5 |
| 2018 | SPNet: Deep 3D Object Classification and Retrieval Using Stereographic Projection
Mohsen Yavartanoo, Euyoung Kim, Kyoung Mu Lee |
ACCV (5) | 3 |
| 2018 | V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation From a Single Depth MapabstractMost of the existing deep learning-based methods for 3D hand and human pose estimation from a single depth map are based on a common framework that takes a 2D depth map and directly regresses the 3D coordinates of keypoints, such as hand or human body joints, via 2D convolutional neural networks (CNNs). The first weakness of this approach is the presence of perspective distortion in the 2D depth map. While the depth map is intrinsically 3D data, many previous methods treat depth maps as 2D images that can distort the shape of the actual object through projection from 3D to 2D space. This compels the network to perform perspective distortion-invariant estimation. The second weakness of the conventional approach is that directly regressing 3D coordinates from a 2D image is a highly nonlinear mapping, which causes difficulty in the learning procedure. To overcome these weaknesses, we firstly cast the 3D hand and human pose estimation problem from a single depth map into a voxel-to-voxel prediction that uses a 3D voxelized grid and estimates the per-voxel likelihood for each keypoint. We design our model as a 3D CNN that provides accurate estimates while running in real-time. Our system outperforms previous methods in almost all publicly available 3D hand and human pose estimation datasets and placed first in the HANDS 2017 frame-based 3D hand pose estimation challenge. The code is available in1. Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 3 |
| 2018 | SeedNet: Automatic Seed Generation With Deep Reinforcement Learning for Robust Interactive SegmentationabstractIn this paper, we propose an automatic seed generation technique with deep reinforcement learning to solve the interactive segmentation problem. One of the main issues of the interactive segmentation problem is robust and consistent object extraction with less human effort. Most of the existing algorithms highly depend on the distribution of inputs, which differs from one user to another and hence need sequential user interactions to achieve adequate performance. In our system, when a user first specifies a point on the desired object and a point in the background, a sequence of artificial user input is automatically generated for precisely segmenting the desired object. The proposed system allows the user to reduce the number of input significantly. This problem is difficult to cast as a supervised learning problem because it is not possible to define globally optimal user input at some stage of the interactive segmentation task. Hence, we formulate automatic seed generation problem as Markov Decision Process (MDP) and then optimize it by reinforcement learning with Deep Q-Network (DQN). We train our network on the MSRA10K dataset and show that the network achieves notable performance improvement from inaccurate initial segmentation on both seen and unseen datasets. Gwangmo Song, Heesoo Myeong, Kyoung Mu Lee |
CVPR | 3 |
| 2018 | Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future GoalsabstractIn this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints. Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001 |
CVPR | 6 |
| 2018 | Task-Aware Image Downscaling
Myungsub Choi, Bee Lim, Kyoung Mu Lee |
ECCV (4) | 4 |
| 2018 | Joint Blind Motion Deblurring and Depth Estimation of Light Field
Haesol Park, In Kyu Park, Kyoung Mu Lee |
ECCV (16) | 4 |
| 2018 | Clustering Convolutional Kernels to Compress Deep Neural Networks
Sanghyun Son 0002, Seungjun Nah, Kyoung Mu Lee |
ECCV (8) | 3 |
| 2018 | Part-Aligned Bilinear Representations for Person Re-identification
Yumin Suh, Jingdong Wang 0001, Siyu Tang 0001, Tao Mei 0001, Kyoung Mu Lee |
ECCV (14) | 5 |
| 2018 | Depth Estimation Network for Dual Defocused Images with Different Depth-of-FieldabstractIn this work, we propose an algorithm to estimate the depth map of a scene using defocused images. In particular, the depth map is estimated using two defocused images with different depth-of-field for the same scene. Similar to the approach of the general depth from defocus (DFD), the proposed algorithm obtains the depth information from the blurredness of the object. Moreover, our proposed algorithm dramatically improves the accuracy by using both the shallow and deep depth-of-field images, simultaneously. Especially, we propose a novel depth estimation network for dual defocused images using convolutional neural network (CNN). We evaluate our proposed network on the NYU-v2 dataset and show superior performance compared to the existing techniques. Gwangmo Song, Kyoung Mu Lee |
ICIP | 2 |
| 2018 | Session details: Keynote 4
Kyoung Mu Lee |
ACM Multimedia | 1 |
| 2018 | 2D-3D pose consistency-based conditional random fields for 3D human pose estimation
Ju Yong Chang, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 2 |
| 2018 | Real-time visual tracking by deep reinforced decision making
Janghoon Choi, Junseok Kwon, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 3 |
| 2018 | Dynamic Video Deblurring Using a Locally Adaptive Blur ModelabstractState-of-the-art video deblurring methods cannot handle blurry videos recorded in dynamic scenes since they are built under a strong assumption that the captured scenes are static. Contrary to the existing methods, we propose a new video deblurring algorithm that can deal with general blurs inherent in dynamic scenes. To handle general and locally varying blurs caused by various sources, such as moving objects, camera shake, depth variation, and defocus, we estimate pixel-wise varying non-uniform blur kernels. We infer bidirectional optical flows to handle motion blurs, and also estimate Gaussian blur maps to remove optical blur from defocus. Therefore, we propose a single energy model that jointly estimates optical flows, defocus blur maps and latent frames. We also provide a framework and efficient solvers to minimize the proposed energy model. By optimizing the energy model, we achieve significant improvements in removing general blurs, estimating optical flows, and extending depth-of-field in blurry frames. Moreover, in this work, to evaluate the performance of non-uniform deblurring methods objectively, we have constructed a new realistic dataset with ground truths. In addition, extensive experimental results on publicly available challenging videos demonstrate that the proposed method produces qualitatively superior performance than the state-of-the-art methods which often fail in either deblurring or optical flow estimation. Tae Hyun Kim 0006, Seungjun Nah, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Robust Light Field Depth Estimation Using Occlusion-Noise Aware Data CostsabstractDepth estimation is essential in many light field applications. Numerous algorithms have been developed using a range of light field properties. However, conventional data costs fail when handling noisy scenes in which occlusion is present. To address this problem, we introduce a light field depth estimation method that is more robust against occlusion and less sensitive to noise. Two novel data costs are proposed, which are measured using the angular patch and refocus image, respectively. The constrained angular entropy cost (CAE) reduces the effects of the dominant occluder and noise in the angular patch, resulting in a low cost. The constrained adaptive defocus cost (CAD) provides a low cost in the occlusion region, while also maintaining robustness against noise. Integrating the two data costs is shown to significantly improve the occlusion and noise invariant capability. Cost volume filtering and graph cut optimization are applied to improve the accuracy of the depth map. Our experimental results confirm the robustness of the proposed method and demonstrate its ability to produce high-quality depth maps from a range of scenes. The proposed method outperforms other state-of-the-art light field depth estimation methods in both qualitative and quantitative evaluations. Williem 0001, In Kyu Park, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Deep Multi-scale Convolutional Neural Network for Dynamic Scene DeblurringabstractNon-uniform blind deblurring for general dynamic scenes is a challenging computer vision problem as blurs arise not only from multiple object motions but also from camera shake, scene depth variation. To remove these complicated motion blurs, conventional energy optimization based methods rely on simple assumptions such that blur kernel is partially uniform or locally linear. Moreover, recent machine learning based methods also depend on synthetic blur datasets generated under these assumptions. This makes conventional deblurring methods fail to remove blurs where blur kernel is difficult to approximate or parameterize (e.g. object motion boundaries). In this work, we propose a multi-scale convolutional neural network that restores sharp images in an end-to-end manner where blur is caused by various sources. Together, we present multi-scale loss function that mimics conventional coarse-to-fine approaches. Furthermore, we propose a new large-scale dataset that provides pairs of realistic blurry image and the corresponding ground truth sharp image that are obtained by a high-speed camera. With the proposed model trained on this dataset, we demonstrate empirically that our method achieves the state-of-the-art performance in dynamic scene deblurring not only qualitatively, but also quantitatively. Seungjun Nah, Tae Hyun Kim 0006, Kyoung Mu Lee |
CVPR | 3 |
| 2017 | Online Video Deblurring via Dynamic Temporal Blending Network
Tae Hyun Kim 0006, Kyoung Mu Lee, Bernhard Schölkopf, Michael Hirsch 0001 |
ICCV | 2 |
| 2017 | Joint Estimation of Camera Pose, Depth, Deblurring, and Super-Resolution from a Blurred Image SequenceabstractThe conventional methods for estimating camera poses and scene structures from severely blurry or low resolution images often result in failure. The off-the-shelf deblurring or super-resolution methods may show visually pleasing results. However, applying each technique independently before matching is generally unprofitable because this naive series of procedures ignores the consistency between images. In this paper, we propose a pioneering unified framework that solves four problems simultaneously, namely, dense depth reconstruction, camera pose estimation, super-resolution, and deblurring. By reflecting a physical imaging process, we formulate a cost minimization problem and solve it using an alternating optimization technique. The experimental results on both synthetic and real videos show high-quality depth maps derived from severely degraded images that contrast the failures of naive multi-view stereo methods. Our proposed method also produces outstanding deblurred and super-resolved images unlike the independent application or combination of conventional video deblurring, super-resolution methods. Haesol Park, Kyoung Mu Lee |
ICCV | 2 |
| 2017 | Multi-modal/multi-scale convolutional neural network based in-loop filter design for next generation video codecabstractIn this paper, we propose a novel in-loop filter design for video compression. Our approach aims to replace existing deblocking filter and SAO (Sample Adaptive Offset) of HEVC standard with multi-modal/multi-scale convolutional neural network (MMS-net). The proposed CNN architecture consists of two sub-networks of different scales. An input image is down-sampled first and restored through the lower scale network, then the output image from it is fed into higher scale network concatenated with the original input image. Moreover, to boost the restoration performance, the proposed architecture utilizes information resides in the coded sequence. Specifically, the compression parameters from coding tree units (CTU) are exploited as input to CNN, which helps to alleviate blocking artifacts on the reconstructed images. In the experiments, our method reduces the average BD-rate by 4.55% and 8.5%, respectively, compared with the conventional neural network based approach [1] and HEVC reference software HM16.7 [2] in `All Intra - Main' configuration. Jihong Kang, Sungjei Kim, Kyoung Mu Lee |
ICIP | 3 |
| 2017 | Adaptive Visual Tracking with Minimum Uncertainty Gap EstimationabstractA novel tracking algorithm is proposed, which robustly tracks a target by finding the state that minimizes the likelihood uncertainty. Likelihood uncertainty is estimated by determining the gap between the lower and upper bounds of likelihood. By minimizing the gap between the two bounds, the proposed method identifies the confident and reliable state of the target. In this study, the state that provides the Minimum Uncertainty Gap (MUG) between likelihood bounds is shown to be more reliable than the state that provides the maximum likelihood only, especially when severe illumination changes, occlusions, and pose variations occur. A rigorous derivation of the lower and upper bounds of the likelihood for the visual tracking problem is provided to address this issue. Additionally, an efficient inference algorithm that uses Interacting Markov Chain Monte Carlo (IMCMC) approach is presented to find the best state that maximizes the average of the lower and upper bounds of likelihood while minimizing the gap between the two bounds. We extend our method to update the target model adaptively. To update the model, the current observation is combined with a previous target model with the adaptive weight, which is calculated according to the goodness of the current observation. The goodness of the observation is measured using the proposed uncertainty gap estimation of likelihood. Experimental results demonstrate that the proposed method robustly tracks the target in realistic videos and outperforms conventional tracking methods. Junseok Kwon, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Look Wider to Match Image Patches With Convolutional Neural NetworksabstractWhen a human matches two images, the viewer has a natural tendency to view the wide area around the target pixel to obtain clues of right correspondence. However, designing a matching cost function that works on a large window in the same way is difficult. The cost function is typically not intelligent enough to discard the information irrelevant to the target pixel, resulting in undesirable artifacts. In this letter, we propose a novel convolutional neural network (CNN) module to learn a stereo matching cost with a large-sized window. Unlike conventional pooling layers with strides, the proposed per-pixel pyramid-pooling layer can cover a large area without a loss of resolution and detail. Therefore, the learned matching cost function can successfully utilize the information from a large area without introducing the fattening effect. The proposed method is robust despite the presence of weak textures, depth discontinuity, illumination, and exposure difference. The proposed method achieves near-peak performance on the Middlebury benchmark. Haesol Park, Kyoung Mu Lee |
IEEE Signal Process. Lett. | 2 |
| 2016 | Deeply-Recursive Convolutional Network for Image Super-ResolutionabstractWe propose an image super-resolution method (SR) using a deeply-recursive convolutional network (DRCN). Our network has a very deep recursive layer (up to 16 recursions). Increasing recursion depth can improve performance without introducing new parameters for additional convolutions. Albeit advantages, learning a DRCN is very hard with a standard gradient descent method due to exploding/ vanishing gradients. To ease the difficulty of training, we propose two extensions: recursive-supervision and skip-connection. Our method outperforms previous methods by a large margin. Jung Kwon Lee, Kyoung Mu Lee |
CVPR | 3 |
| 2016 | Accurate Image Super-Resolution Using Very Deep Convolutional NetworksabstractWe present a highly accurate single-image superresolution (SR) method. Our method uses a very deep convolutional network inspired by VGG-net used for ImageNet classification [19]. We find increasing our network depth shows a significant improvement in accuracy. Our final model uses 20 weight layers. By cascading small filters many times in a deep network structure, contextual information over large image regions is exploited in an efficient way. With very deep networks, however, convergence speed becomes a critical issue during training. We propose a simple yet effective training procedure. We learn residuals only and use extremely high learning rates (104 times higher than SRCNN [6]) enabled by adjustable gradient clipping. Our proposed method performs better than existing methods in accuracy and visual improvements in our results are easily noticeable. Jung Kwon Lee, Kyoung Mu Lee |
CVPR | 3 |
| 2016 | A Sequential Approach to 3D Human Pose Estimation: Separation of Localization and Identification of Body Joints
Ho Yub Jung, Yumin Suh, Gyeongsik Moon, Kyoung Mu Lee |
ECCV (5) | 4 |
| 2016 | Extraction of Coronary Vessels in Fluoroscopic X-Ray Sequences Using Vessel Correspondence OptimizationabstractWe present a method to extract coronary vessels from fluoroscopic x-ray sequences. Given the vessel structure for the source frame, vessel correspondence candidates in the subsequent frame are generated by a novel hierarchical search scheme to overcome the aperture problem. Optimal correspondences are determined within a Markov random field optimization framework. Post-processing is performed to extract vessel branches newly visible due to the inflow of contrast agent. Quantitative and qualitative evaluation conducted on a dataset of 18 sequences demonstrate the effectiveness of the proposed method. Seung Yeon Shin, Soochahn Lee, Kyoung Jin Noh, Il Dong Yun, Kyoung Mu Lee |
MICCAI (3) | 5 |
| 2016 | Special Issue on Visual Tracking
Xue Mei, Tianzhu Zhang 0001, Huchuan Lu, Ming-Hsuan Yang 0001, Kyoung Mu Lee, Horst Bischof |
Comput. Vis. Image Underst. | 5 |
| 2015 | MRF optimization by graph approximationabstractGraph cuts-based algorithms have achieved great success in energy minimization for many computer vision applications. These algorithms provide approximated solutions for multi-label energy functions via move-making approach. This approach fuses the current solution with a proposal to generate a lower-energy solution. Thus, generating the appropriate proposals is necessary for the success of the move-making approach. However, not much research efforts has been done on the generation of “good” proposals, especially for non-metric energy functions. In this paper, we propose an application-independent and energy-based approach to generate “good” proposals. With these proposals, we present a graph cuts-based move-making algorithm called GA-fusion (fusion with graph approximation-based proposals). Extensive experiments support that our proposal generation is effective across different classes of energy functions. The proposed algorithm outperforms others both on real and synthetic problems. Wonsik Kim, Kyoung Mu Lee |
CVPR | 2 |
| 2015 | Generalized video deblurring for dynamic scenesabstractSeveral state-of-the-art video deblurring methods are based on a strong assumption that the captured scenes are static. These methods fail to deblur blurry videos in dynamic scenes. We propose a video deblurring method to deal with general blurs inherent in dynamic scenes, contrary to other methods. To handle locally varying and general blurs caused by various sources, such as camera shake, moving objects, and depth variation in a scene, we approximate pixel-wise kernel with bidirectional optical flows. Therefore, we propose a single energy model that simultaneously estimates optical flows and latent frames to solve our deblurring problem. We also provide a framework and efficient solvers to optimize the energy model. By minimizing the proposed energy function, we achieve significant improvements in removing blurs and estimating accurate optical flows in blurry frames. Extensive experimental results demonstrate the superiority of the proposed method in real and challenging videos that state-of-the-art methods fail in either deblurring or optical flow estimation. Tae Hyun Kim 0006, Kyoung Mu Lee |
CVPR | 2 |
| 2015 | Subgraph matching using compactness prior for robust feature correspondenceabstractFeature correspondence plays a central role in various computer vision applications. It is widely formulated as a graph matching problem due to its robust performance under challenging conditions, such as background clutter, object deformation and repetitive patterns. A variety of fast and accurate algorithms have been proposed for graph matching. However, most of them focus on improving the recall of the solution while rarely considering its precision, thus inducing a solution with numerous outliers. To address both precision and recall feature correspondence should rather be formulated as a subgraph matching problem. This paper proposes a new subgraph matching formulation which uses a compactness prior, an additional constraint that prefers sparser solutions and effectively eliminates outliers. To solve the new optimization problem, we propose a meta-algorithm based on Markov chain Monte Carlo. By constructing Markov chain on the restricted search space instead of the original solution space, our method approximates the solution effectively. The experiments indicate that our proposed formulation and algorithm significantly improve the baseline performance under challenging conditions when both outliers and deformation noise are present. Yumin Suh, Kamil Adamczewski, Kyoung Mu Lee |
CVPR | 3 |
| 2015 | Discrete Tabu Search for Graph MatchingabstractGraph matching is a fundamental problem in computer vision. In this paper, we propose a novel graph matching algorithm based on tabu search [13]. The proposed method solves graph matching problem by casting it into an equivalent weighted maximum clique problem of the corresponding association graph, which we further penalize through introducing negative weights. Subsequent tabu search optimization allows for overcoming the convention of using positive weights. The method's distinct feature is that it utilizes the history of search to make more strategic decisions while looking for the optimal solution, thus effectively escaping local optima and in practice achieving superior results. The proposed method, unlike the existing algorithms, enables direct optimization in the original discrete space while encouraging rather than artificially enforcing hard one-to-one constraint, thus resulting in better solution. The experiments demonstrate the robustness of the algorithm in a variety of settings, presenting the state-of-the-art results. The code is available at http://cv.snu.ac.kr/research/~DTSGM/. Kamil Adamczewski, Yumin Suh, Kyoung Mu Lee |
ICCV | 3 |
| 2015 | Large margin learning of hierarchical semantic similarity for image classification
Ju Yong Chang, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 2 |
| 2015 | A Unified Framework for Event Summarization and Rare Event Detection from Multiple ViewsabstractA novel approach for event summarization and rare event detection is proposed. Unlike conventional methods that deal with event summarization and rare event detection independently, our method solves them in a single framework by transforming them into a graph editing problem. In our approach, a video is represented by a graph, each node of which indicates an event obtained by segmenting the video spatially and temporally. The edges between nodes describe the relationship between events. Based on the degree of relations, edges have different weights. After learning the graph structure, our method finds subgraphs that represent event summarization and rare events in the video by editing the graph, that is, merging its subgraphs or pruning its edges. The graph is edited to minimize a predefined energy model with the Markov Chain Monte Carlo (MCMC) method. The energy model consists of several parameters that represent the causality, frequency, and significance of events. We design a specific energy model that uses these parameters to satisfy each objective of event summarization and rare event detection. The proposed method is extended to obtain event summarization and rare event detection results across multiple videos captured from multiple views. For this purpose, the proposed method independently learns and edits each graph of individual videos for event summarization or rare event detection. Then, the method matches the extracted multiple graphs to each other, and constructs a single composite graph that represents event summarization or rare events from multiple views. Experimental results show that the proposed approach accurately summarizes multiple videos in a fully unsupervised manner. Moreover, the experiments demonstrate that the approach is advantageous in detecting rare transition of events. Junseok Kwon, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Scanline Sampler without Detailed Balance: An Efficient MCMC for MRF OptimizationabstractMarkov chain Monte Carlo (MCMC) is an elegant tool, widely used in variety of areas. In computer vision, it has been used for the inference on the Markov random field model (MRF). However, MCMC less concerned than other deterministic approaches although it converges to global optimal solution in theory. The major obstacle is its slow convergence. To come up with faster sampling method, we investigate two ideas: breaking detailed balance and updating multiple nodes at a time. Although detailed balance is considered to be essential element of MCMC, it actually is not the necessary condition for the convergence. In addition, exploiting the structure of MRF, we introduce a new kernel which updates multiple nodes in a scanline rather than a single node. Those two ideas are integrated in a novel way to develop an efficient method called scanline sampler without detailed balance. In experimental section, we apply our method to the OpenGM2 benchmark of MRF optimization and show the proposed method achieves faster convergence than the conventional approaches. Wonsik Kim, Kyoung Mu Lee |
CVPR | 2 |
| 2014 | Segmentation-Free Dynamic Scene DeblurringabstractMost state-of-the-art dynamic scene deblurring methods based on accurate motion segmentation assume that motion blur is small or that the specific type of motion causing the blur is known. In this paper, we study a motion segmentation-free dynamic scene deblurring method, which is unlike other conventional methods. When the motion can be approximated to linear motion that is locally (pixel-wise) varying, we can handle various types of blur caused by camera shake, including out-of-plane motion, depth variation, radial distortion, and so on. Thus, we propose a new energy model simultaneously estimating motion flow and the latent image based on robust total variation (TV)-L1 model. This approach is necessary to handle abrupt changes in motion without segmentation. Furthermore, we address the problem of the traditional coarse-to-fine deblurring framework, which gives rise to artifacts when restoring small structures with distinct motion. We thus propose a novel kernel re-initialization method which reduces the error of motion flow propagated from a coarser level. Moreover, a highly effective convex optimization-based solution mitigating the computational difficulties of the TV-L1 model is established. Comparative experimental results on challenging real blurry images demonstrate the efficiency of the proposed method. Tae Hyun Kim 0006, Kyoung Mu Lee |
CVPR | 2 |
| 2014 | Interval Tracker: Tracking by Interval AnalysisabstractThis paper proposes a robust tracking method that uses interval analysis. Any single posterior model necessarily includes a modeling uncertainty (error), and thus, the posterior should be represented as an interval of probability. Then, the objective of visual tracking becomes to find the best state that maximizes the posterior and minimizes its interval simultaneously. By minimizing the interval of the posterior, our method can reduce the modeling uncertainty in the posterior. In this paper, the aforementioned objective is achieved by using the M4 estimation, which combines the Maximum a Posterior (MAP) estimation with Minimum Mean-Square Error (MMSE), Maximum Likelihood (ML), and Minimum Interval Length (MIL) estimations. In the M4 estimation, our method maximizes the posterior over the state obtained by the MMSE estimation. The method also minimizes interval of the posterior by reducing the gap between the lower and upper bounds of the posterior. The gap is reduced when the likelihood is maximized by the ML estimation and the interval length of the state is minimized by the MIL estimation. The experimental results demonstrate that M4 estimation can be easily integrated into conventional tracking methods and can greatly enhance their tracking accuracy. In several challenging datasets, our method outperforms state-of-the-art tracking methods. Junseok Kwon, Kyoung Mu Lee |
CVPR | 2 |
| 2014 | Robust Visual Tracking with Double Bounding Box Model
Junseok Kwon, Junha Roh, Kyoung Mu Lee, Luc Van Gool |
ECCV (1) | 3 |
| 2014 | Alpha Matting of Motion-Blurred Objects in Bracket Sequence Images
Heesoo Myeong, Stephen Lin 0001, Kyoung Mu Lee |
ECCV (3) | 3 |
| 2014 | Stereo reconstruction using high-order likelihoods
Ho Yub Jung, Haesol Park, In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 4 |
| 2014 | Tracking by Sampling and IntegratingMultiple TrackersabstractWe propose the visual tracker sampler, a novel tracking algorithm that can work robustly in challenging scenarios, where several kinds of appearance and motion changes of an object can occur simultaneously. The proposed tracking algorithm accurately tracks a target by searching for appropriate trackers in each frame. Since the real-world tracking environment varies severely over time, the trackers should be adapted or newly constructed depending on the current situation, so that each specific tracker takes charge of a certain change in the object. To do this, our method obtains several samples of not only the states of the target but also the trackers themselves during the sampling process. The trackers are efficiently sampled using the Markov Chain Monte Carlo (MCMC) method from the predefined tracker space by proposing new appearance models, motion models, state representation types, and observation types, which are the important ingredients of visual trackers. All trackers are then integrated into one compound tracker through an Interacting MCMC (IMCMC) method, in which the trackers interactively communicate with one another while running in parallel. By exchanging information with others, each tracker further improves its performance, thus increasing overall tracking performance. Experimental results show that our method tracks the object accurately and reliably in realistic videos, where appearance and motion drastically change over time, and outperforms even state-of-the-art tracking methods. Junseok Kwon, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | A Geometric Particle Filter for Template-Based Visual TrackingabstractExisting approaches to template-based visual tracking, in which the objective is to continuously estimate the spatial transformation parameters of an object template over video frames, have primarily been based on deterministic optimization, which as is well-known can result in convergence to local optima. To overcome this limitation of the deterministic optimization approach, in this paper we present a novel particle filtering approach to template-based visual tracking. We formulate the problem as a particle filtering problem on matrix Lie groups, specifically the three-dimensional Special Linear group SL(3) and the two-dimensional affine group Aff(2). Computational performance and robustness are enhanced through a number of features: (i) Gaussian importance functions on the groups are iteratively constructed via local linearization; (ii) the inverse formulation of the Jacobian calculation is used; (iii) template resizing is performed; and (iv) parent-child particles are developed and used. Extensive experimental results using challenging video sequences demonstrate the enhanced performance and robustness of our particle filtering-based approach to template-based visual tracking. We also show that our approach outperforms several state-of-the-art template-based visual tracking methods via experiments using the publicly available benchmark data set. Junghyun Kwon, Hee Seok Lee, Frank C. Park 0001, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Minimum Uncertainty Gap for Robust Visual TrackingabstractWe propose a novel tracking algorithm that robustly tracks the target by finding the state which minimizes uncertainty of the likelihood at current state. The uncertainty of the likelihood is estimated by obtaining the gap between the lower and upper bounds of the likelihood. By minimizing the gap between the two bounds, our method finds the confident and reliable state of the target. In the paper, the state that gives the Minimum Uncertainty Gap (MUG) between likelihood bounds is shown to be more reliable than the state which gives the maximum likelihood only, especially when there are severe illumination changes, occlusions, and pose variations. A rigorous derivation of the lower and upper bounds of the likelihood for the visual tracking problem is provided to address this issue. Additionally, an efficient inference algorithm using Interacting Markov Chain Monte Carlo is presented to find the best state that maximizes the average of the lower and upper bounds of the likelihood and minimizes the gap between two bounds simultaneously. Experimental results demonstrate that our method successfully tracks the target in realistic videos and outperforms conventional tracking methods. Junseok Kwon, Kyoung Mu Lee |
CVPR | 2 |
| 2013 | Dense 3D Reconstruction from Severely Blurred Images Using a Single Moving CameraabstractMotion blur frequently occurs in dense 3D reconstruction using a single moving camera, and it degrades the quality of the 3D reconstruction. To handle motion blur caused by rapid camera shakes, we propose a blur-aware depth reconstruction method, which utilizes a pixel correspondence that is obtained by considering the effect of motion blur. Motion blur is dependent on 3D geometry, thus parameter zing blurred appearance of images with scene depth given camera motion is possible and a depth map can be accurately estimated from the blur-considered pixel correspondence. The estimated depth is then converted into pixel-wise blur kernels, and non-uniform motion blur is easily removed with low computational cost. The obtained blur kernel is depth-dependent, thus it effectively addresses scene-depth variation, which is a challenging problem in conventional non-uniform deblurring methods. Hee Seok Lee, Kyoung Mu Lee |
CVPR | 2 |
| 2013 | Simultaneous Super-Resolution of Depth and Images Using a Single CameraabstractIn this paper, we propose a convex optimization framework for simultaneous estimation of super-resolved depth map and images from a single moving camera. The pixel measurement error in 3D reconstruction is directly related to the resolution of the images at hand. In turn, even a small measurement error can cause significant errors in reconstructing 3D scene structure or camera pose. Therefore, enhancing image resolution can be an effective solution for securing the accuracy as well as the resolution of 3D reconstruction. In the proposed method, depth map estimation and image super-resolution are formulated in a single energy minimization framework with a convex function and solved efficiently by a first-order primal-dual algorithm. Explicit inter-frame pixel correspondences are not required for our super-resolution procedure, thus we can avoid a huge computation time and obtain improved depth map in the accuracy and resolution as well as high-resolution images with reasonable time. The superiority of our algorithm is demonstrated by presenting the improved depth map accuracy, image super-resolution results, and camera pose estimation. Hee Seok Lee, Kyoung Mu Lee |
CVPR | 2 |
| 2013 | Tensor-Based High-Order Semantic Relation Transfer for Semantic Scene SegmentationabstractWe propose a novel nonparametric approach for semantic segmentation using high-order semantic relations. Conventional context models mainly focus on learning pairwise relationships between objects. Pairwise relations, however, are not enough to represent high-level contextual knowledge within images. In this paper, we propose semantic relation transfer, a method to transfer high-order semantic relations of objects from annotated images to unlabeled images analogous to label transfer techniques where label information are transferred. We first define semantic tensors representing high-order relations of objects. Semantic relation transfer problem is then formulated as semi-supervised learning using a quadratic objective function of the semantic tensors. By exploiting low-rank property of the semantic tensors and employing Kronecker sum similarity, an efficient approximation algorithm is developed. Based on the predicted high-order semantic relations, we reason semantic segmentation and evaluate the performance on several challenging datasets. Heesoo Myeong, Kyoung Mu Lee |
CVPR | 2 |
| 2013 | Dynamic Scene DeblurringabstractMost conventional single image deblurring methods assume that the underlying scene is static and the blur is caused by only camera shake. In this paper, in contrast to this restrictive assumption, we address the deblurring problem of general dynamic scenes which contain multiple moving objects as well as camera shake. In case of dynamic scenes, moving objects and background have different blur motions, so the segmentation of the motion blur is required for deblurring each distinct blur motion accurately. Thus, we propose a novel energy model designed with the weighted sum of multiple blur data models, which estimates different motion blurs and their associated pixel-wise weights, and resulting sharp image. In this framework, the local weights are determined adaptively and get high values when the corresponding data models have high data fidelity. And, the weight information is used for the segmentation of the motion blur. Non-local regularization of weights are also incorporated to produce more reliable segmentation results. A convex optimization-based method is used for the solution of the proposed energy model. Experimental results demonstrate that our method outperforms conventional approaches in deblurring both dynamic scenes and static scenes. Tae Hyun Kim 0006, Byeongjoo Ahn, Kyoung Mu Lee |
ICCV | 3 |
| 2013 | Optical Flow via Locally Adaptive Fusion of Complementary Data CostsabstractMany state-of-the-art optical flow estimation algorithms optimize the data and regularization terms to solve ill-posed problems. In this paper, in contrast to the conventional optical flow framework that uses a single or fixed data model, we study a novel framework that employs locally varying data term that adaptively combines different multiple types of data models. The locally adaptive data term greatly reduces the matching ambiguity due to the complementary nature of the multiple data models. The optimal number of complementary data models is learnt by minimizing the redundancy among them under the minimum description length constraint (MDL). From these chosen data models, a new optical flow estimation energy model is designed with the weighted sum of the multiple data models, and a convex optimization-based highly effective and practical solution that finds the optical flow, as well as the weights is proposed. Comparative experimental results on the Middlebury optical flow benchmark show that the proposed method using the complementary data models outperforms the state-of-the art methods. Tae Hyun Kim 0006, Hee Seok Lee, Kyoung Mu Lee |
ICCV | 3 |
| 2013 | Window annealing for pixel-labeling problems
Ho Yub Jung, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 2 |
| 2013 | Multi-object reconstruction from dynamic scenes: An object-centered approach
Young Min Shin, Minsu Cho, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 3 |
| 2013 | Geometric particle swarm optimization for robust visual ego-motion estimation via particle filtering
Young Ki Baik, Junghyun Kwon, Hee Seok Lee, Kyoung Mu Lee |
Image Vis. Comput. | 4 |
| 2013 | Joint Depth Map and Color Consistency Estimation for Stereo Images with Different Illuminations and CamerasabstractAbstract—In this paper, we propose a method that infers both accurate depth maps and color-consistent stereo images for radiometrically varying stereo images. In general, stereo matching and performing color consistency between stereo images are a chicken-and-egg problem since it is not a trivial task to simultaneously achieve both goals. Hence, we have developed an iterative framework in which these two processes can boost each other. First, we transform the input color images to log-chromaticity color space, from which a linear relationship can be established during constructing a joint pdf of transformed left and right color images. From this joint pdf, we can estimate a linear function that relates the corresponding pixels in stereo images. Based on this linear property, we present a new stereo matching cost by combining Mutual Information (MI), SIFT descriptor, and segment-based plane-fitting to robustly find correspondence for stereo image pairs which undergo radiometric variations. Meanwhile, we devise a Stereo Color Histogram Equalization (SCHE) method to produce color-consistent stereo image pairs, which conversely boost the disparity map estimation. Experimental results show that our method produces both accurate depth maps and color-consistent stereo images, even for stereo images with severe radiometric differences. Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Learning Full Pairwise Affinities for Spectral SegmentationabstractSegmenting a single image into multiple coherent groups remains a challenging task in the field of computer vision. Particularly, spectral segmentation which uses the global information embedded in the spectrum of a given image's affinity matrix is a major trend in image segmentation. This paper focuses on the problem of efficiently learning a full range of pairwise affinities gained by integrating local grouping cues for spectral segmentation. We first construct a sparse multilayer graph whose nodes are both the pixels and the oversegmented regions obtained by an unsupervised segmentation algorithm. By applying the semi-supervised learning strategy to this graph, the intra and interlayer affinities between all pairs of nodes can be estimated without iteration. These pairwise affinities are then applied into the spectral segmentation algorithms. In this paper, two types of spectral segmentation algorithms are introduced: $(K)$-way segmentation and hierarchical segmentation. Our algorithms provide high-quality segmentations which preserve object details by directly incorporating the full-range connections. Moreover, since our full affinity matrix is defined by the inverse of a sparse matrix, its eigendecomposition can be efficiently computed. The experimental results on the BSDS and MSRC image databases demonstrate the superiority of our segmentation algorithms in terms of relevance and accuracy compared with existing popular methods. Kyoung Mu Lee, Sang Uk Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Wang-Landau Monte Carlo-Based Tracking Methods for Abrupt MotionsabstractWe propose a novel tracking algorithm based on the Wang-Landau Monte Carlo (WLMC) sampling method for dealing with abrupt motions efficiently. Abrupt motions cause conventional tracking methods to fail because they violate the motion smoothness constraint. To address this problem, we introduce the Wang-Landau sampling method and integrate it into a Markov Chain Monte Carlo (MCMC)-based tracking framework. By employing the novel density-of-states term estimated by the Wang-Landau sampling method into the acceptance ratio of MCMC, our WLMC-based tracking method alleviates the motion smoothness constraint and robustly tracks the abrupt motions. Meanwhile, the marginal likelihood term of the acceptance ratio preserves the accuracy in tracking smooth motions. The method is then extended to obtain good performance in terms of scalability, even on a high-dimensional state space. Hence, it covers drastic changes in not only position but also scale of a target. To achieve this, we modify our method by combining it with the N-fold way algorithm and present the N-Fold Wang-Landau (NFWL)-based tracking method. The N-fold way algorithm helps estimate the density-of-states with a smaller number of samples. Experimental results demonstrate that our approach efficiently samples the states of the target, even in a whole state space, without loss of time, and tracks the target accurately and robustly when position and scale are changing severely. Junseok Kwon, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Highly Nonrigid Object Tracking via Patch-Based Dynamic Appearance ModelingabstractA novel tracking algorithm is proposed for targets with drastically changing geometric appearances over time. To track such objects, we develop a local patch-based appearance model and provide an efficient online updating scheme that adaptively changes the topology between patches. In the online update process, the robustness of each patch is determined by analyzing the likelihood landscape of the patch. Based on this robustness measure, the proposed method selects the best feature for each patch and modifies the patch by moving, deleting, or newly adding it over time. Moreover, a rough object segmentation result is integrated into the proposed appearance model to further enhance it. The proposed framework easily obtains segmentation results because the local patches in the model serve as good seeds for the semi-supervised segmentation task. To solve the complexity problem attributable to the large number of patches, the Basin Hopping (BH) sampling method is introduced into the tracking framework. The BH sampling method significantly reduces computational complexity with the help of a deterministic local optimizer. Thus, the proposed appearance model could utilize a sufficient number of patches. The experimental results show that the present approach could track objects with drastically changing geometric appearance accurately and robustly. Junseok Kwon, Kyoung Mu Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Progressive graph matching: Making a move of graphs via probabilistic votingabstractGraph matching is widely used in a variety of scientific fields, including computer vision, due to its powerful performance, robustness, and generality. Its computational complexity, however, limits the permissible size of input graphs in practice. Therefore, in real-world applications, the initial construction of graphs to match becomes a critical factor for the matching performance, and often leads to unsatisfactory results. In this paper, to resolve the issue, we propose a novel progressive framework which combines probabilistic progression of graphs with matching of graphs. The algorithm efficiently re-estimates in a Bayesian manner the most plausible target graphs based on the current matching result, and guarantees to boost the matching objective at the subsequent graph matching. Experimental evaluation demonstrates that our approach effectively handles the limits of conventional graph matching and achieves significant improvement in challenging image matching problems. Minsu Cho, Kyoung Mu Lee |
CVPR | 2 |
| 2012 | Mode-seeking on graphs via random walksabstractMode-seeking has been widely used as a powerful data analysis technique for clustering and filtering in a metric feature space. We introduce a versatile and efficient mode-seeking method for “graph” representation where general embedding of relational data is possible beyond metric spaces. Exploiting the global structure of the graph by random walks, our method intrinsically combines mode-seeking with ranking on the graph, and performs robust analysis by seeking high-ranked authoritative data and suppressing low-ranked noise and outliers. This enables mode-seeking to be applied to a large class of challenging real-world problems involving graph representation which frequently arises in computer vision. We demonstrate our method on various synthetic experiments and real applications dealing with noisy and complex data such as scene summarization and object-based image matching. Minsu Cho, Kyoung Mu Lee |
CVPR | 2 |
| 2012 | A unified framework for event summarization and rare event detectionabstractIn this paper, we have proposed an unified framework for event summarization and rare event detection and presented the graph-structure learning and editing method to solve these problems efficiently. The experimental results demonstrated that the proposed method outperformed conventional algorithms in complex and crowded public scenes by exploiting and utilizing causality, frequency, and significance of relations of events. Junseok Kwon, Kyoung Mu Lee |
CVPR | 2 |
| 2012 | Learning object relationships via graph-based context modelabstractIn this paper, we propose a novel framework for modeling image-dependent contextual relationships using graph-based context model. This approach enables us to selectively utilize the contextual relationships suitable for an input query image. We introduce a context link view of contextual knowledge, where the relationship between a pair of annotated regions is represented as a context link on a similarity graph of regions. Link analysis techniques are used to estimate the pairwise context scores of all pairs of unlabeled regions in the input image. Our system integrates the learned context scores into a Markov Random Field (MRF) framework in the form of pairwise cost and infers the semantic segmentation result by MRF optimization. Experimental results on object class segmentation show that the proposed graph-based context model outperforms the current state-of-the-art methods. Heesoo Myeong, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 3 |
| 2012 | Robust visual tracking using autoregressive hidden Markov ModelabstractRecent studies on visual tracking have shown significant improvement in accuracy by handling the appearance variations of the target object. Whereas most studies present schemes to extract the time-invariant characteristics of the target and adaptively update the appearance model, the present paper concentrates on modeling the probabilistic dependency between sequential target appearances (Fig. 1-(a)). To actualize this interest, a new Bayesian tracking framework is formulated under the autoregressive Hidden Markov Model (AR-HMM), where the probabilistic dependency between sequential target appearances is implied. During the learning phase at each time step, the proposed tracker separates formerly seen target samples into several clusters based on their visual similarity, and learns cluster-specific classifiers as multiple appearance models, each of which represents a certain type of the target appearance. Then the dependency between these appearance models is learned. During the searching phase, the target state is estimated by inferring the most probable appearance model under the consideration of its dependency on formerly utilized appearance models. The proposed method is tested on 12 challenging video sequences containing targets with abrupt appearance variations, and demonstrates that it outperforms current state-of-the-art methods in accuracy. Dong Woo Park, Junseok Kwon, Kyoung Mu Lee |
CVPR | 3 |
| 2012 | Abnormal Object Detection by Canonical Scene-Based Contextual Model
Sangdon Park 0001, Wonsik Kim, Kyoung Mu Lee |
ECCV (3) | 3 |
| 2012 | Graph Matching via Sequential Monte Carlo
Yumin Suh, Minsu Cho, Kyoung Mu Lee |
ECCV (3) | 3 |
| 2011 | Hyper-graph matching via reweighted random walksabstractEstablishing correspondences between two feature sets is a fundamental issue in computer vision, pattern recognition, and machine learning. This problem can be well formulated as graph matching in which nodes represent feature points while edges describe pairwise relations between feature points. Recently, several researches have tried to embed higher-order relations of feature points by hyper-graph matching formulations. In this paper, we generalize the previous hyper-graph matching formulations to cover relations of features in arbitrary orders, and propose a novel state-of-the-art algorithm by reinterpreting the random walk concept on the hyper-graph in a probabilistic manner. Adopting personalized jumps with a reweighting scheme, the algorithm effectively reflects the one-to-one matching constraints during the random walk process. Comparative experiments on synthetic data and real images show that the proposed method clearly outperforms existing algorithms especially in the presence of noise and outliers. Minsu Cho, Kyoung Mu Lee |
CVPR | 3 |
| 2011 | Stereo reconstruction using high order likelihoodabstractUnder the popular Bayesian approach, a stereo problem can be formulated by defining likelihood and prior. Likelihoods are often associated with unary terms and priors are defined by pair-wise or higher order cliques in Markov random field (MRF). In this paper, we propose to use high order likelihood model in stereo. Numerous conventional patch based matching methods such as normalized cross correlation, Laplacian of Gaussian, or census filters are designed under the naive assumption that all the pixels of a patch have the same disparities. However, patch-wise cost can be formulated as higher order cliques for MRF so that the matching cost is a function of image patch's disparities. A patch obtained from the projected image by a disparity map should provide a better match without the blurring effect around disparity discontinuities. Among patch-wise high order matching costs, the census filter approach can be easily reduced to pair-wise cliques. The experimental results on census filter-based high order likelihood demonstrate the advantages of high order likelihood over independent identically distributed unary model. Ho Yub Jung, Kyoung Mu Lee, Sang Uk Lee |
ICCV | 2 |
| 2011 | Tracking by Sampling TrackersabstractWe propose a novel tracking framework called visual tracker sampler that tracks a target robustly by searching for the appropriate trackers in each frame. Since the real-world tracking environment varies severely over time, the trackers should be adapted or newly constructed depending on the current situation. To do this, our method obtains several samples of not only the states of the target but also the trackers themselves during the sampling process. The trackers are efficiently sampled using the Markov Chain Monte Carlo method from the predefined tracker space by proposing new appearance models, motion models, state representation types, and observation types, which are the basic important components of visual trackers. Then, the sampled trackers run in parallel and interact with each other while covering various target variations efficiently. The experiment demonstrates that our method tracks targets accurately and robustly in the real-world tracking environments and outperforms the state-of-the-art tracking methods. Junseok Kwon, Kyoung Mu Lee |
ICCV | 2 |
| 2011 | Simultaneous localization, mapping and deblurringabstractHandling motion blur is one of important issues in visual SLAM. For a fast-moving camera, motion blur is an unavoidable effect and it can degrade the results of localization and reconstruction severely. In this paper, we present a unified algorithm to handle motion blur for visual SLAM, including the blur-robust data association method and the fast deblurring method. In our framework, camera motion and 3-D point structures are reconstructed by SLAM, and the information from SLAM makes the estimation of motion blur quite easy and effective. Reversely, estimating motion blur enables robust data association and drift-free localization of SLAM with blurred images. The blurred images are recovered by fast deconvolution using SLAM data, and more features are extracted and registered to the map so that the SLAM procedure can be continued even with the blurred images. In this way, visual SLAM and deblurring are solved simultaneously, and improve each other's results significantly. Hee Seok Lee, Junghyun Kwon, Kyoung Mu Lee |
ICCV | 3 |
| 2011 | GPU-friendly multi-view stereo reconstruction using surfel representation and graph cuts
Ju Yong Chang, Haesol Park, In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 4 |
| 2011 | A hybrid approach for MRF optimization problems: Combination of stochastic sampling and deterministic algorithms
Wonsik Kim, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 2 |
| 2011 | Robust Stereo Matching Using Adaptive Normalized Cross-CorrelationabstractA majority of the existing stereo matching algorithms assume that the corresponding color values are similar to each other. However, it is not so in practice as image color values are often affected by various radiometric factors such as illumination direction, illuminant color, and imaging device changes. For this reason, the raw color recorded by a camera should not be relied on completely, and the assumption of color consistency does not hold good between stereo images in real scenes. Therefore, the performance of most conventional stereo matching algorithms can be severely degraded under the radiometric variations. In this paper, we present a new stereo matching measure that is insensitive to radiometric variations between left and right images. Unlike most stereo matching measures, we use the color formation model explicitly in our framework and propose a new measure, called the Adaptive Normalized Cross-Correlation (ANCC), for a robust and accurate correspondence measure. The advantage of our method is that it is robust to lighting geometry, illuminant color, and camera parameter changes between left and right images, and does not suffer from the fattening effect unlike conventional Normalized Cross-Correlation (NCC). Experimental results show that our method outperforms other state-of-the-art stereo methods under severely different radiometric conditions between stereo images. Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Ghost-Free High Dynamic Range Imaging
Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee, Youngsu Moon, Joonhyuk Cha |
ACCV (4) | 2 |
| 2010 | Authority-shift clustering: Hierarchical clustering by authority seeking on graphsabstractIn this paper, a novel hierarchical clustering method using link analysis techniques is introduced. The algorithm is formulated as an authority seeking procedure on graphs, which computes the shifts toward nodes with high authority scores. For the authority shift, we adopted the personalized PageRank score of the graph. Based on the concept of authority seeking, we achieve hierarchical clustering by iteratively propagating the authority scores to other nodes and shifting authority nodes. This scheme solves the chicken-egg difficulty in hierarchical clustering by a semiglobal bottom-up approach exploiting the global structure of the graph. The experimental evaluation demonstrates that our algorithm is more powerful compared with existing graph-based approaches in clustering and image segmentation tasks. Minsu Cho, Kyoung Mu Lee |
CVPR | 2 |
| 2010 | Unsupervised detection and segmentation of identical objectsabstractWe address an unsupervised object detection and segmentation problem that goes beyond the conventional assumptions of one-to-one object correspondences or modeltest settings between images. Our method can detect and segment identical objects directly from a single image or a handful of images without any supervision. To detect and segment all the object-level correspondences from the given images, a novel multi-layer match-growing method is proposed that starts from initial local feature matches and explores the images by intra-layer expansion and inter-layer merge. It estimates geometric relations between object entities and establishes ‘object correspondence networks’ that connect matching objects. Experiments demonstrate robust performance of our method on challenging datasets. Minsu Cho, Young Min Shin, Kyoung Mu Lee |
CVPR | 3 |
| 2010 | Learning full pairwise affinities for spectral segmentationabstractThis paper studies the problem of learning a full range of pairwise affinities gained by integrating local grouping cues for spectral segmentation. The overall quality of the spectral segmentation depends mainly on the pairwise pixel affinities. By employing a semi-supervised learning technique, optimal affinities are learnt from the test image without iteration. We first construct a multi-layer graph with pixels and regions, generated by the mean shift algorithm, as nodes. By applying the semi-supervised learning strategy to this graph, we can estimate the intra- and inter-layer affinities between all pairs of nodes together. These pair-wise affinities are then used to simultaneously cluster all pixel and region nodes into visually coherent groups across all layers in a single multi-layer framework of Normalized Cuts. Our algorithm provides high-quality segmentations with object details by directly incorporating the full range connections in the spectral framework. Since the full affinity matrix is defined by the inverse of a sparse matrix, its eigen-decomposition is efficiently computed. The experimental results on Berkeley and MSRC image databases demonstrate the relevance and accuracy of our algorithm as compared to existing popular methods. Kyoung Mu Lee, Sang Uk Lee |
CVPR | 2 |
| 2010 | Nonparametric higher-order learning for interactive segmentationabstractIn this paper, we deal with a generative model for multilabel, interactive segmentation. To estimate the pixel likelihoods for each label, we propose a new higher-order formulation additionally imposing the soft label consistency constraint whereby the pixels in the regions, generated by unsupervised image segmentation algorithms, tend to have the same label. In contrast with previous works which focus on the parametric model of the higher-order cliques for adding this soft constraint, we address a nonparametric learning technique to recursively estimate the region likelihoods as higher-order cues from the resulting likelihoods of pixels included in the regions. Therefore the main idea of our algorithm is to design two quadratic cost functions of pixel and region likelihoods, that are supplementary to each other, in a proposed multi-layer graph and to estimate them simultaneously by a simple optimization technique. In this manner, we consider long-range connections between the regions that facilitate propagation of local grouping cues across larger image areas. The experiments on challenging data sets show that integration of higher-order cues quantitatively and qualitatively improves the segmentation results with detailed boundaries and reduces sensitivity with respect to seed quantity and placement. Kyoung Mu Lee, Sang Uk Lee |
CVPR | 2 |
| 2010 | Visual tracking decompositionabstractWe propose a novel tracking algorithm that can work robustly in a challenging scenario such that several kinds of appearance and motion changes of an object occur at the same time. Our algorithm is based on a visual tracking decomposition scheme for the efficient design of observation and motion models as well as trackers. In our scheme, the observation model is decomposed into multiple basic observation models that are constructed by sparse principal component analysis (SPCA) of a set of feature templates. Each basic observation model covers a specific appearance of the object. The motion model is also represented by the combination of multiple basic motion models, each of which covers a different type of motion. Then the multiple basic trackers are designed by associating the basic observation models and the basic motion models, so that each specific tracker takes charge of a certain change in the object. All basic trackers are then integrated into one compound tracker through an interactive Markov Chain Monte Carlo (IMCMC) framework in which the basic trackers communicate with one another interactively while run in parallel. By exchanging information with others, each tracker further improves its performance, which results in increasing the whole performance of tracking. Experimental results show that our method tracks the object accurately and reliably in realistic videos where the appearance and motion are drastically changing over time. Junseok Kwon, Kyoung Mu Lee |
CVPR | 2 |
| 2010 | Monocular SLAM with locally planar landmarks via geometric rao-blackwellized particle filtering on Lie groupsabstractWe propose a novel geometric Rao-Blackwellized particle filtering framework for monocular SLAM with locally planar landmarks. We represent the states for the camera pose and the landmark plane normal as SE(3) and SO(3), respectively, which are both Lie groups. The measurement error is also represented as another Lie group SL(3) corresponding to the space of homography matrices. We then formulate the unscented transformation on Lie groups for optimal importance sampling and landmark estimation via unscented Kalman filter. The feasibility of our framework is demonstrated via various experiments. Junghyun Kwon, Kyoung Mu Lee |
CVPR | 2 |
| 2010 | Reweighted Random Walks for Graph Matching
Minsu Cho, Kyoung Mu Lee |
ECCV (5) | 3 |
| 2010 | Continuous Markov Random Field Optimization Using Fusion Move Driven Markov Chain Monte Carlo TechniqueabstractMany vision applications have been formulated as Markov Random Field (MRF) problems. Although many of them are discrete labeling problems, continuous formulation often achieves great improvement on the qualities of the solutions in some applications such as stereo matching and optical flow. In continuous formulation, however, it is much more difficult to optimize the target functions. In this paper, we propose a new method called fusion move driven Markov Chain Monte Carlo method (MCMC-F) that combines the Markov Chain Monte Carlo method and the fusion move to solve continuous MRF problems effectively. This algorithm exploits powerful fusion move while it fully explore the whole solution space. We evaluate it using the stereo matching problem. We empirically demonstrate that the proposed algorithm is more stable and always finds lower energy states than the state-of-the art optimization techniques. Wonsik Kim, Kyoung Mu Lee |
ICPR | 2 |
| 2010 | A Unified Probabilistic Approach to Feature Matching and Object SegmentationabstractThis paper deals with feature matching and segmentation of common objects in a pair of images, simultaneously. For the feature matching problem, the matching likelihoods of all feature correspondences are obtained by combining their discriminative power with the spatial coherence constraint that favors their spatial aggregation via object segmentation. At the same time, for the object segmentation problem, our algorithm estimates the object likelihood that each subregion is a commonly existing part in two images by the affinity propagation of the resulted matching likelihoods. Since these two problems are related to each other, our main idea to solve them is to integrate all the priors about them into a unified framework, that consists of several correlated quadratic cost functions. Eventually, all matching and object likelihoods are estimated simultaneously as a solution of linear system of equations. Based on these likelihoods, we finally recover the optimal feature matches and the common object parts by imposing simple sequential mapping and thresholding techniques, respectively. The experiments demonstrate the superiority of our algorithm compared with the conventional methods. Kyoung Mu Lee, Sang Uk Lee |
ICPR | 2 |
| 2010 | A Graph Matching Algorithm Using Data-Driven Markov Chain Monte Carlo SamplingabstractWe propose a novel stochastic graph matching algorithm based on data-driven Markov Chain Monte Carlo (DDMCMC) sampling technique. The algorithm explores the solution space efficiently and avoid local minima by taking advantage of spectral properties of the given graphs in data-driven proposals. Thus, it enables the graph matching to be robust to deformation and outliers arising from the practical correspondence problems. Our comparative experiments using synthetic and real data demonstrate that the algorithm outperforms the state-of-the-art graph matching algorithms. Minsu Cho, Kyoung Mu Lee |
ICPR | 3 |
| 2010 | Co-recognition of Actions in Video PairsabstractIn this paper, we present a method that recognizes single or multiple common actions between a pair of video sequences. We establish an energy function that evaluates geometric and photometric consistency, and solve the action recognition problem by optimizing the energy function. The proposed stochastic inference algorithm based on the Monte Carlo method explores the video pair from the local spatio-temporal interest point matches to find the common actions. Our algorithm works in unsupervised way without prior knowledge about the type and the number of common actions. Experiments show that our algorithm produces promising results on single and multiple action recognition. Young Min Shin, Minsu Cho, Kyoung Mu Lee |
ICPR | 3 |
| 2010 | FPGA Design and Implementation of a Real-Time Stereo Vision SystemabstractStereo vision is a well-known ranging method because it resembles the basic mechanism of the human eye. However, the computational complexity and large amount of data access make real-time processing of stereo vision challenging because of the inherent instruction cycle delay within conventional computers. In order to solve this problem, the past 20 years of research have focused on the use of dedicated hardware architecture for stereo vision. This paper proposes a fully pipelined stereo vision system providing a dense disparity image with additional sub-pixel accuracy in real-time. The entire stereo vision process, such as rectification, stereo matching, and post-processing, is realized using a single field programmable gate array (FPGA) without the necessity of any external devices. The hardware implementation is more than 230 times faster when compared to a software program operating on a conventional computer, and shows stronger performance over previous hardware-related studies. Seunghun Jin, Jung Uk Cho, Xuan Dai Pham, Kyoung Mu Lee, Sung-Kee Park, Jaewook Jeon |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | A Probabilistic Model for Correspondence Problems Using Random Walks with Restart
Kyoung Mu Lee, Sang Uk Lee |
ACCV (3) | 2 |
| 2009 | Bilateral Symmetry Detection via Symmetry-GrowingabstractWe present a novel and robust method for localizing and segmenting bilaterally symmetric patterns from real-world images. On the basis of symmetrically matched pairs of local features, our method expands and merges condent local symmetric region matches by exploiting both photometric similarity and geometric consistency via our new symmetry-growing framework. It overcomes the limitations of the previous local-feature based approaches by efciently exploring the image space to grow symmetry beyond the detected symmetric features. The experimental evaluation demonstrates that our method successfully detects and segments multiple symmetric patterns from real-world images, and clearly outperforms the state-of-the-art methods in accuracy and robustness. Minsu Cho, Kyoung Mu Lee |
BMVC | 2 |
| 2009 | Incorporating Higher-Order Cues in Image ColorizationabstractColorization problem is to find the colors of all pixels X = {xn}n=1,...,|X |, given a grayscale image I with scribbles S with the desired colors. We work in the YUV color space where Y = {yn}n=1,...,|X | is the monochromatic luminance channel, which we will refer to simply as intensity, while U = {un}n=1,...,|X | and V = {vn}n=1,...,|X | are the chrominance channels, encoding the color. Our goal is to complete both the U and V channels, given Y = I. We deal with the only U channel in this paper, since the V channel can be treated in the same manner. In this paper, we propose a new multi-layer graph model and an energy formulation that can incorporate higher-order cues for reliable colorization of natural images. In contrast to most existing energy functions [3] with unary and pairwise constraints, we address the problem of imposing a high-order constraint whereby pixels constituting each region tend to have similar colors to the representative color of the region they belong to. The representative colors of the regions that are generated by unsupervised image segmentation algorithms, act as higher-order cues. Unlike previous parametric models [2], they are automatically obtained by a nonparametric learning technique that estimates them from the resulting pixel colors in a recursive fashion. We formulate this problem in terms of two quadratic energy functions of pixel and region colors, that are supplementary to each other, in our proposed multi-layer graph model and estimate them by a simple optimization technique that minimizes both functions simultaneously. Our proposed algorithm works as follows. We first design an undirected graph G = (Q,E) where the nodes Q = {X ,R} consist of two types: pixels X and regions R, generated by an unsupervised segmentation algorithm such as Mean Shift [1], and the edges E are the links between two nodes as shown in Fig. 1(a). Each pixel xn ∈ X initially has an intensity yn ∈ Y . For each region rk ∈ R, we can generate its properties ȳk as the mean intensity of the inner pixels xn ∈ rk: ȳk = 1 |rk| ∑xn∈rk yn. We then formulate both quadratic energy functions JX and JR for estimating the pixel colors U = {un}n=1,...,|X | and the region colors Ū = {ūk}k=1,...,|R|, respectively, as follows. Kyoung Mu Lee, Sang Uk Lee |
BMVC | 2 |
| 2009 | Mutual information-based stereo matching combined with SIFT descriptor in log-chromaticity color spaceabstractRadiometric variations between input images can seriously degrade the performance of stereo matching algorithms. In this situation, mutual information is a very popular and powerful measure which can find any global relationship of intensities between two input images taken from unknown sources. The mutual information-based method, however, is still ambiguous or erroneous as regards local radiometric variations, since it only accounts for global variation between images, and does not contain spatial information properly. In this paper, we present a new method based on mutual information combined with SIFT descriptor to find correspondence for images which undergo local as well as global radiometric variations. We transform the input color images to log-chromaticity color space from which a linear relationship can be established. To incorporate spatial information in mutual information, we utilize the SIFT descriptor which includes near pixel gradient histogram to construct a joint probability in log-chromaticity color space. By combining the mutual information as an appearance measure and the SIFT descriptor as a geometric measure, we devise a robust and accurate stereo system. Experimental results show that our method is superior to the state-of-the art algorithms including conventional mutual information-based methods and window correlation methods under various radiometric changes. Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee |
CVPR | 2 |
| 2009 | Markov Chain Monte Carlo combined with deterministic methods for Markov random field optimizationabstractMany vision problems have been formulated as energy minimization problems and there have been significant advances in energy minimization algorithms. The most widely-used energy minimization algorithms include graph cuts, belief propagation and tree-reweighted message passing. Although they have obtained good results, they are still unsatisfactory when it comes to more difficult MRF problems such as non-submodular energy functions, highly connected MRFs, and high-order clique potentials. There have also been other approaches, known as stochastic sampling-based algorithms, which include simulated annealing, Markov chain Monte Carlo and population based Markov chain Monte Carlo. They are applicable to any general energy models but they are usually slower than deterministic methods. In this paper, we propose new algorithms which elegantly combine stochastic and deterministic methods. Sampling-based methods are boosted by deterministic methods so that they can rapidly move to lower energy states and easily jump over energy barriers. In different point of view, the sampling-based method prevents deterministic methods from getting stuck at local minima. Consequently, a combination of both approaches substantially increases the quality of the solutions. We present a thorough analysis of the proposed methods in synthetic MRF problems by controlling the hardness of the problems. We also demonstrate experimental results for the photomontage problem which is the most difficult one among the standard MRF benchmark problems. Wonsik Kim, Kyoung Mu Lee |
CVPR | 2 |
| 2009 | Tracking of a non-rigid object via patch-based dynamic appearance modeling and adaptive Basin Hopping Monte Carlo samplingabstractWe propose a novel tracking algorithm for the target of which geometric appearance changes drastically over time. To track it, we present a local patch-based appearance model and provide an efficient scheme to evolve the topology between local patches by on-line update. In the process of on-line update, the robustness of each patch in the model is estimated by a new method of measurement which analyzes the landscape of local mode of the patch. This patch can be moved, deleted or newly added, which gives more flexibility to the model. Additionally, we introduce the Basin Hopping Monte Carlo (BHMC) sampling method to our tracking problem to reduce the computational complexity and deal with the problem of getting trapped in local minima. The BHMC method makes it possible for our appearance model to consist of enough numbers of patches. Since BHMC uses the same local optimizer that is used in the appearance modeling, it can be efficiently integrated into our tracking framework. Experimental results show that our approach tracks the object whose geometric appearance is drastically changing, accurately and robustly. Junseok Kwon, Kyoung Mu Lee |
CVPR | 2 |
| 2009 | Visual tracking via geometric particle filtering on the affine group with optimal importance functionsabstractWe propose a geometric method for visual tracking, in which the 2-D affine motion of a given object template is estimated in a video sequence by means of coordinate-invariant particle filtering on the 2-D affine group Aff(2). Tracking performance is further enhanced through a geometrically defined optimal importance function, obtained explicitly via Taylor expansion of a principal component analysis based measurement function on Aff(2). The efficiency of our approach to tracking is demonstrated via comparative experiments. Junghyun Kwon, Kyoung Mu Lee, Frank C. Park 0001 |
CVPR | 2 |
| 2009 | A genetic algorithm with local map for path planning in dynamic environmentsabstractIn this paper, a new genetic algorithm (GA) for solving the path planning in dynamic environments is proposed. The new genetic algorithm uses local maps, therefore, does not require the knowledge of exact or estimated position of the destination point as other approaches in the literature. Consequently, the new GA could be used to solve the problem of dynamic path planning under an assumption that makes the problem more restrictive but more close to reality (in searching tasks). Ivan Koryakovskiy, Nguyen Xuan Hoai, Kyoung Mu Lee |
GECCO | 3 |
| 2009 | Feature correspondence and deformable object matching via agglomerative correspondence clusteringabstractWe present an efficient method for feature correspondence and object-based image matching, which exploits both photometric similarity and pairwise geometric consistency from local invariant features. We formulate object-based image matching as an unsupervised multi-class clustering problem on a set of candidate feature matches, and propose a novel pairwise dissimilarity measure and a robust linkage model in the framework of hierarchical agglomerative clustering. The algorithm handles significant amount of outliers and deformation as well as multiple clusters, thus enabling simultaneous feature matching and clustering from real-world image pairs with significant clutter and multiple deformable objects. The experimental evaluation on feature correspondence, object recognition, and object-based image matching demonstrates that our method is robust to both outliers and deformation, and applicable to a wide range of image matching problems. Minsu Cho, Kyoung Mu Lee |
ICCV | 3 |
| 2009 | Simultaneous color consistency and depth map estimation for radiometrically varying stereo imagesabstractIn this paper, we propose a new method that infers accurate depth maps and color-consistent images between radiometrically varying stereo images, simultaneously. In general, stereo matching and performing color consistency between stereo images are a chicken-and-egg problem. Color consistency enhances the performance of stereo matching, while accurate correspondences from stereo disparities improve color consistency between stereo images. We devise a new iterative framework in which these two processes can boost each other. For robust stereo matching, we utilize the mutual information-based method combined with the SIFT descriptor from which we can estimate the joint pdf in log-chromaticity color space. From this joint pdf, we can estimate a linear relationship between the corresponding pixels in stereo images. Using this linear relationship and the estimated depth maps, we devise a stereo color histogram equalization method to make color-consistent stereo images which conversely boost the disparity map estimation. Experimental results show that our method produces both accurate depth maps and color-consistent stereo images even for stereo images with severe radiometric differences. Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee |
ICCV | 2 |
| 2009 | Edge-preserving colorization using data-driven Random Walks with RestartabstractIn this paper, we consider the colorization problem of grayscale images in which some color scribbles are initially given. Our proposed method is based on the weighted color blending of the scribbles. Unlike previous works which utilize the shortest distance as the blending weights, we employ a new intrinsic distance measure based on the random walks with restart (RWR), known as a very successful technique for defining the relevance between two nodes in a graph. In our work, we devise new modified data-driven RWR framework that can incorporate locally adaptive and data-driven restarting probabilities. In this new framework, the restarting probability of each pixel becomes dependent on its edgeness, generated by the Canny detector. Since this data-driven RWR enforces color consistency in the areas bounded by the edges, it produces more reliable edge-preserving colorization results that are less sensitive to the size and position of each scribble. Moreover, if the additional information about the scribbles which indicate the foreground object is available, our method can be readily applied to the object segmentation and matting. Experiments on several synthetic, cartoon and natural images demonstrate that our method achieves much high quality colorization results compared with the state-of-the-art methods. Kyoung Mu Lee, Sang Uk Lee |
ICIP | 2 |
| 2009 | Multi-robot SLAM using ceiling visionabstractIn this paper we present a new vision-based SLAM approach for multi-robot formulation. For a cooperative map reconstruction, the robots have to know each other's relative poses, but estimating these at the start of operation puts a limit on real applications. In our study, the robots start the single SLAM with their own global coordinate, and merge their maps during the operation by detecting the overlapped region of their maps. The robots automatically recognize the occurrence of map overlapping by matching their current frame with the maps built by other robots. With the robust data association technique from the ceiling-vision based SLAM, the proposed algorithm robustly detects the overlapping regions and estimates the accurate transformations for map alignment. In our experiment, we have verified that our algorithm successfully enables the multi-robot SLAM without any initial correspondence or encounter of robots. Hee Seok Lee, Kyoung Mu Lee |
IROS | 2 |
| 2009 | Multiswarm Particle Filter for vision based SLAMabstractParticle filters have been widely used as a powerful optimization tool for nonlinear, non-Gaussian dynamic models such as simultaneous localization and mapping (SLAM) and visual tracking. Particle filters, however, often suffer from particle impoverishment, which is caused by a mismatch between proposal distribution and target distribution. To solve this problem, we propose a new method to improve the efficiency of particle filters by employing the particle swarm optimization (PSO), which is a kind of swarm intelligence algorithm. The PSO, especially its variant for dynamic models, is combined with the generic particle filter to get samples that are well matched with target distribution. The resulting filter is applied to a vision based SLAM system and its performance is tested. We present experimental results that demonstrate improved accuracy in localization and mapping at the same or less computational cost than the conventional particle filters. Hee Seok Lee, Kyoung Mu Lee |
IROS | 2 |
| 2009 | Stereo Matching Using Population-Based MCMC
Wonsik Kim, Joonyoung Park, Kyoung Mu Lee |
Int. J. Comput. Vis. | 3 |
| 2008 | Illumination and camera invariant stereo matchingabstractColor information can be used as a basic and crucial cue for finding correspondence in a stereo matching algorithm. In a real scene, however, image colors are affected by various geometric and radiometric factors. For this reason, the raw color recorded by a camera is not a reliable cue, and the color consistency assumption is no longer valid between stereo images in real scenes. Hence the performance of most conventional stereo matching algorithms can be severely degraded under the radiometric variations. In this paper, we present a new stereo matching algorithm that is invariant to various radiometric variations between left and right images. Unlike most stereo algorithms, we explicitly employ the color formation model in our framework and propose a new measure called Adaptive Normalized Cross Correlation (ANCC) for a robust and accurate correspondence measure. ANCC is invariant to lighting geometry, illuminant color and camera parameter changes between left and right images, and does not suffer from fattening effects unlike conventional Normalized Cross Correlation (NCC). Experimental results show that our algorithm outperforms other stereo algorithms under severely different radiometric conditions between stereo images. Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee |
CVPR | 2 |
| 2008 | Co-recognition of Image Pairs by Data-Driven Monte Carlo Image Exploration
Minsu Cho, Young Min Shin, Kyoung Mu Lee |
ECCV (4) | 3 |
| 2008 | Window Annealing over Square Lattice Markov Random Field
Ho Yub Jung, Kyoung Mu Lee, Sang Uk Lee |
ECCV (2) | 2 |
| 2008 | Toward Global Minimum through Combined Local Minima
Ho Yub Jung, Kyoung Mu Lee, Sang Uk Lee |
ECCV (4) | 2 |
| 2008 | Generative Image Segmentation Using Random Walks with Restart
Kyoung Mu Lee, Sang Uk Lee |
ECCV (3) | 2 |
| 2008 | Tracking of Abrupt Motion Using Wang-Landau Monte Carlo Estimation
Junseok Kwon, Kyoung Mu Lee |
ECCV (1) | 2 |
| 2008 | MAP-MRF approach for binarization of degraded document imageabstractWe propose an algorithm for the binarization of document images degraded by uneven light distribution, based on the Markov Random Field modeling with Maximum A Posteriori probability (MAP-MRF) estimation. While the conventional algorithms use the decision based on the thresholding, the proposed algorithm makes a soft decision based on the probabilistic model. To work with the MAP-MRF framework we formulate an energy function by a likelihood model and a generalized Potts prior model. Then we construct a graph for the energy, and obtain the optimized result by using the well-known graph cut algorithm. Experimental results show that our approach is more robust to various types of images than the previous hard decision approaches. Jung Gap Kuk, Nam Ik Cho, Kyoung Mu Lee |
ICIP | 3 |
| 2008 | Occlusion invariant face recognition using selective local non-negative matrix factorization basis images
Hyun Jun Oh, Kyoung Mu Lee, Sang Uk Lee |
Image Vis. Comput. | 2 |
| 2008 | Shape from shading using graph cuts
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
Pattern Recognit. | 2 |
| 2007 | Stereo Matching Using Population-Based MCMC
Joonyoung Park, Wonsik Kim, Kyoung Mu Lee |
ACCV (2) | 3 |
| 2007 | Multiview normal field integration using level set methodsabstractIn this paper, we propose a new method to integrate multiview normal fields using level sets. In contrast with conventional normal integration algorithms used in shape from shading and photometric stereo that reconstruct a 2.5D surface using a single-view normal field, our algorithm can combine multiview normal fields simultaneously and recover the full 3D shape of a target object. We formulate this multiview normal integration problem by an energy minimization framework and find an optimal solution in a least square sense using a variational technique. A level set method is applied to solve the resultant geometric PDE that minimizes the proposed error functional. It is shown that the resultant flow is composed of the well known mean curvature and flux maximizing flows. In particular, we apply the proposed algorithm to the problem of 3D shape modelling in a multiview photometric stereo setting. Experimental results for various synthetic data show the validity of our approach. Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
CVPR | 2 |
| 2007 | Partially Occluded Object-Specific Segmentation in View-Based RecognitionabstractWe present a novel object-specific segmentation method which can be used in view-based object recognition systems. Previous object segmentation approaches generate inexact results especially in partially occluded and cluttered environment because their top-down strategies fail to explain the details of various specific objects. On the contrary, our segmentation method efficiently exploits the information of the matched model views in view-based recognition because the aligned model view to the input image can serve as the best top-down cue for object segmentation. In this paper, we cast the problem of partially occluded object segmentation as that of labelling displacement and foreground status simultaneously for each pixel between the aligned model view and an input image. The problem is formulated by a maximum a posteriori Markov random field (MAP-MRF) model which minimizes a particular energy function. Our method overcomes complex occlusion and clutter and provides accurate segmentation boundaries by combining a bottom-up segmentation cue together. We demonstrate the efficiency and robustness of it by experimental results on various objects under occluded and cluttered environments. Minsu Cho, Kyoung Mu Lee |
CVPR | 2 |
| 2007 | Simultaneous Depth Reconstruction and Restoration of Noisy Stereo Images using Non-local Pixel DistributionabstractIn this paper, we propose a new algorithm that solves both the stereo matching and the image denoising problem simultaneously for a pair of noisy stereo images. Most stereo algorithms employ L1 or L2 intensity error-based data costs in the MAP-MRF framework by assuming the naive intensity-constancy. These data costs make typical stereo algorithms suffer from the effect of noise severely. In this study, a new robust stereo algorithm to noise is presented that performs the stereo matching and the image denoising simultaneously. In our approach, we redefine the data cost by two terms. The first term is the restored intensity difference, instead of the observed intensity difference. The second term is the non-local pixel distribution dissimilarity around the matched pixels. We adopted the NL-means (Non Local-means) algorithm for restoring the intensity value as a function of disparity. And a pixel distribution dissimilarity is calculated by using PMHD (Perceptually Modified Hausdorff Distance). The restored intensity values in each image are determined by inferring optimal disparity map at the same time. Experimental results show that the proposed algorithm is more robust and accurate than other conventional algorithms in both stereo matching and denoising. Yong Seok Heo, Kyoung Mu Lee, Sang Uk Lee |
CVPR | 2 |
| 2007 | Stereo matching using iterative reliable disparity map expansion in the color-spatial-disparity space
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
Pattern Recognit. | 2 |
| 2006 | Stereo Matching Using Iterated Graph Cuts and Mean Shift Filtering
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
ACCV (1) | 2 |
| 2006 | Occlusion Invariant Face Recognition Using Selective LNMF Basis Images
Hyun Jun Oh, Kyoung Mu Lee, Sang Uk Lee, Chung-Hyuk Yim |
ACCV (1) | 2 |
| 2006 | A New Stereo Matching Model Using Visibility Constraint Based on Disparity Consistency
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
ACIVS | 2 |
| 2006 | Stereo Matching Using Scanline Disparity Discontinuity Optimization
Ho Yub Jung, Kyoung Mu Lee, Sang Uk Lee |
ACIVS | 2 |
| 2006 | A Novel Stochastic Attributed Relational Graph Matching Based on Relation Vector Space Analysis
Bo Gun Park, Kyoung Mu Lee, Sang Uk Lee |
ACIVS | 2 |
| 2006 | A New Similarity Measure for Random Signatures: Perceptually Modified Hausdorff Distance
Bo Gun Park, Kyoung Mu Lee, Sang Uk Lee |
ACIVS | 2 |
| 2006 | A New 3-D Model Retrieval System Based on Aspect-Transition Descriptor
Soochahn Lee, Sehyuk Yoon, Il Dong Yun, Duck Hoon Kim, Kyoung Mu Lee, Sang Uk Lee |
ECCV (4) | 5 |
| 2006 | Visual SLAM with Line and Corner FeaturesabstractWe propose a new vision-based SLAM (simultaneous localization and mapping) technique using both line and corner features as landmarks in the scene. The proposed SLAM algorithm uses an extended Kalman filter based framework to localize and reconstruct 3D line and corner landmarks at the same time and in real time. It provides more accurate localization and map building results than conventional corner feature only-based techniques. Moreover, the reconstructed 3D line landmarks enhance the performance of the robot relocation when robot's pose remains uncertain with corner information only. Experimental results show that the hybrid landmark based SLAM, using lines and corners, produces better performance than corner only one's Woo Yeon Jeong, Kyoung Mu Lee |
IROS | 2 |
| 2005 | A Dense Stereo Matching Using Two-Pass Dynamic Programming with Generalized Ground Control PointsabstractA method for solving dense stereo matching problem is presented in this paper. First, a new generalized ground control points (GGCPs) scheme is introduced, where one or more disparity candidates for the true disparity of each pixel are assigned by local matching using the oriented spatial filters. By allowing "all" pixels to have multiple candidates for their true disparities, GGCPs not only guarantee to provide a sufficient number of starting pixels needed for guiding the subsequent matching process, but also remarkably reduce the risk of false match, improving the previous GCP-based approaches where the number of the selected control points tends to be inversely proportional to the reliability. Second, by employing a two-pass dynamic programming technique that performs optimization both along and across the scanlines, we solve the typical inter-scanline inconsistency problem. Moreover, combined with the GGCPs, the stability and efficiency of the optimization are improved significantly. Experimental results for the standard data sets show that the proposed algorithm achieves comparable results to the state-of-the-arts with much less computational cost. Jae-Chul Kim, Kyoung Mu Lee, Byoung-Tae Choi, Sang Uk Lee |
CVPR (2) | 2 |
| 2005 | Asymmetric multi-phase deformable model for colon segmentationabstractIn virtual colonography, precise segmentation is essential for accurate diagnosis. For the segmentation of colon wall, we propose a novel multi-phase deformable model using a level set method. By defining an asymmetric energy functional, the proposed model can simultaneously segment regions with different characteristics. Compared with the conventional multi-phase models, it shows better convergence without ambiguity. Experimental results with real CT images demonstrate that the proposed algorithm outperforms other methods. Yongseok Yoo, Kyoung Mu Lee, Il Dong Yun, Sang Uk Lee |
ICIP (2) | 2 |
| 2005 | CV-SLAM: a new ceiling vision-based SLAM techniqueabstractWe propose a fast and robust CV-SLAM (ceiling vision-based simultaneous localization and mapping) technique using a single ceiling vision sensor. The proposed algorithm is suitable for system that demands very high localization accuracy such as an intelligent robot vacuum cleaner. A single camera looking upward direction (called ceiling vision system) is mounted on the robot, and salient image features are detected and tracked through the image sequence. Compared with the conventional frontal view systems, the ceiling vision has advantage in tracking, since it involves only rotation and affine transform without scale change. And, in this paper, we solve the rotation and affine transform problems using 3D gradient orientation estimation method and multi-view description of landmarks. By applying these methods to the solution for data association, we can reconstruct the 3D landmark map in real-time through the extend Kalman filter based SLAM framework. Furthermore, relocation problem is solved efficiently by using a wide base line matching between the reconstructed 3D map and a 2D ceiling image. Experimental results demonstrate the accuracy and robustness of the proposed algorithm in real environments. Woo Yeon Jeong, Kyoung Mu Lee |
IROS | 2 |
| 2005 | Face Recognition Using Face-ARG MatchingabstractIn this paper, we propose a novel line feature-based face recognition algorithm. A face is represented by the Face-ARG model, where all the geometric quantities and the structural information are encoded in an Attributed Relational Graph (ARG) structure, then the partial ARG matching is done for matching Face-ARG's. Experimental results demonstrate that the proposed algorithm is quite robust to various facial expression changes, varying illumination conditions and occlusion, even when a single sample per person is given. Bo Gun Park, Kyoung Mu Lee, Sang Uk Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Progressive encoding of binary voxel models using pyramidal decomposition
Musik Kwon, Chang-Su Kim 0001, Kyoung Mu Lee, Sang Uk Lee |
J. Vis. Commun. Image Represent. | 3 |
| 2004 | Perceptual grouping of line features in 3-D space: a model-based framework
In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
Pattern Recognit. | 2 |
| 2003 | Shape from shading using graph cutsabstractThis paper describes a new semiglobal method for SFS (shape-from-shading) using graph cuts. The new algorithm combines the local method proposed by Lee and Rosenfeld (1985) and the global method using energy minimization technique. By employing a new global energy minimization formulation, the convex/concave ambiguity problem of the Lee and Rosenfeld method can be resolved efficiently. A new combinatorial optimization technique, graph cuts method is used for the minimization of the proposed energy functional. Experimental results on a variety of synthetic and real-world images show that the proposed algorithm reconstructs the 3-D shape of objects very efficiently. Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
ICIP (1) | 2 |
| 2003 | A statistical error analysis for voxel coloringabstractThis paper presents an error analysis for voxel coloring [S.M. Seitz, et al. (1997), K.N. Kutulakos, et al. (2000)], which is one of the well known methods to reconstruct 3D shape from 2D calibrated multiple-view images. In order to analyze the errors arising in the reconstruction process of voxel coloring algorithms, we first model several noise sources in the analytic or statistical way, and then examine the effects of each noise component on the reconstructed 3D model. Specifically, in order to analyze the statistical errors, we focus on the distribution of the image variance, which is employed as photo consistency measurement. And also, we show that how specular components induce errors in reconstructing 3D model. The results of this analysis are very useful for evaluating the statistical confidence of the reconstructed 3D model as well as finding the optimal threshold for the occupancy decision. Musik Kwon, Kyoung Mu Lee, Sang Uk Lee |
ICIP (1) | 2 |
| 2003 | Recognition of partially occluded objects using probabilistic ARG (attributed relational graph)-based matching
Bo Gun Park, Kyoung Mu Lee, Sang Uk Lee, Jin Hak Lee |
Comput. Vis. Image Underst. | 2 |
| 2003 | 3D target recognition based on projective invariant relationships
Bong Seop Song, Kyoung Mu Lee, Sang Uk Lee, Il Dong Yun |
J. Vis. Commun. Image Represent. | 2 |
| 2003 | Models and algorithms for efficient multiresolution topology estimation of measured 3-D range dataabstractIn this paper, we propose a new efficient topology estimation algorithm to construct a multiresolution polygonal mesh from measured three-dimensional (3-D) range data. The topology estimation problem is defined under the constraints of cognition, compactness, and regularity, and the algorithm is designed to be applied to either a cloud of points or a dense mesh. The proposed algorithm initially segments the range data into a finite number of Voronoi patches using the K-means clustering algorithm. Each patch is then approximated by an appropriate polygonal and eventually a triangular mesh model. In order to improve the equiangularity of the mesh, we employ a dynamic mesh model, in which the mesh finds its equilibrium state adaptively, according to the equiangularity constraint. Experimental results demonstrate that satisfactory equiangular triangular mesh models can be constructed rapidly at various resolutions, while yielding tolerable modeling error. In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2002 | Face detection using the 1st-order RCE classifierabstractWe present a new face detection algorithm based on the 1st-order reduced Coulomb energy (RCE) classifier. The algorithm locates frontal views of human faces at any degree of rotation and scale in complex scenes. The face candidates and their orientations are first determined by computing the Hausdorff distance between a simple face abstraction model and binarized test windows in an image pyramid. Then, after normalizing the energy, each face candidate is verified by two subsequent classifiers; a binary image classifier and the 1st-order RCE classifier While the binary image classifier is employed as a pre-classifier to discard nonfaces with minimum computational complexity, the 1st-order RCE classifier is used as the main face classifier for final verification. An optimal training method to construct the representative face model database is also presented. Experimental results show that the proposed algorithm yields a high detection ratio, while yielding no false alarm. Byeong Hwan Jeon, Sang Uk Lee, Kyoung Mu Lee |
ICIP (2) | 3 |
| 2002 | Adaptive rate control algorithms for low bit rate video under networks supporting bandwidth renegotiation
Hwangjun Song, Kyoung Mu Lee |
Signal Process. Image Commun. | 2 |
| 2001 | A new robust 3D motion estimation under perspective projectionabstractWe present a new 3D camera motion estimation technique using the optical flow from a pair of images taken under a perspective projection. The problem formulation leads to the solution of an overdetermined nonlinear system of equations with respect to the motion parameters. By employing an efficient initial guess algorithm which uses a weak perspective projection and an image coordinate normalization technique, the nonlinear solution can be obtained robustly and accurately. The proposed method has been tested on both several synthetic and real image sequences. The results show that the performance of the proposed algorithm is quite superior to the conventional ones even under more general and noisy situations. Hye Ri Cho, Kyoung Mu Lee, Sang Uk Lee |
ICIP (3) | 2 |
| 2001 | Model-Based Object Recognition Using Geometric Invariants of Points and Lines
Bong Seop Song, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 2 |
| 2001 | Multi-image matching for a general motion stereo camera model
Ja Seong Ku, Kyoung Mu Lee, Sang Uk Lee |
Pattern Recognit. | 2 |
| 2000 | Recognition and Reconstruction of 3-D Objects Using Model-Based Perceptual GroupingabstractWe address a new algorithm for recognition and reconstruction of 3D polyhedral objects, based on perceptual grouping and graph search technique. Perceptual grouping is performed in a model-based framework, in which decision tree classifier is employed for learning and retrieving geometric information of the 3D model object. On the other hand, in order to extract the polygonal patch structure, initial grouping result is represented by a Gestalt graph. Polygonal patch hypotheses are then generated by graph search and verified by the consistency test with the model. In the experiments, it is shown that the model-based grouping reduces the number of the generated hypotheses efficiently, and furthermore, robust recognition and reconstruction are achieved by means of the graph search technique. In Kyu Park, Sang Uk Lee, Kyoung Mu Lee |
ICPR | 3 |
| 2000 | A Line Feature Matching Technique Based on an Eigenvector Approach
Sang Ho Park, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 2 |
| 1999 | Perceptual Grouping of 3-D Features in Aerial Image Using Decision Tree ClassifierabstractWe address a new perceptual grouping algorithm for aerial images, which employs a decision tree classifier and hierarchical multilevel grouping strategy in a bottom-up fashion. In our approach, grouping is performed perceptually on 3D features extracted from 2D images, in which the gestalt principles including collinearity, parallelism and L-typed convergence are encoded by the decision tree learning technique. The decision tree is constructed using training samples obtained from the given 3D reference model. Then, each pair of the extracted 3D line features of an input image is classified into one of the learned gestalt primitives. On the other hand, in multilevel grouping procedure, grouping of collated features are performed from lower to higher level, yielding the structured target model. In order to evaluate the proposed algorithm, experiments are carried out on RADIUS model board images. The results show that grouping is performed effectively to extract man-made structures in aerial images. In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
ICIP (1) | 2 |
| 1998 | Direct Shape from Texture Using a Parametric Surface Model and an Adaptive Filtering TechniqueabstractIn this research, we propose a new iterative shape from texture (SFT) algorithm which extracts accurate surface depth information of a curved object covered with fairly homogeneous texture directly. The shape information can be inferred from the rate of texture distortion depicted in an image, and therefore the modeling of the projection and surface geometry as well as the estimation of local texture variation are crucial in obtaining accurate surface shape of an object. By introducing semi-perspective projection camera model and a parametric surface model, we establish a new SFT problem formulation called the textural irradiance equation which relates the local texture density called textural intensity to finite surface parameters. Moreover, by adopting an adaptive multiscale filtering scheme for local texture density estimation, in which the scale or frequency band of a local edge filter is chosen adaptively according to the local shape information, we greatly enhance the accuracy of the estimation of the projected local texture densities, and the final reconstructed shape. We demonstrate the performance of the proposed algorithm by the test with several synthetic and real texture images. Kyoung Mu Lee, C.-C. Jay Kuo |
CVPR | 1 |
| 1998 | Multi-Image Matching for a General Motion Stereo Camera ModelabstractThe aim of motion stereo is to extract the 3-D information of an object from images of a moving camera using the geometric relationships between corresponding points. This paper presents an accurate and robust motion stereo algorithm employing multiple images, taken under a general motion. The object functions for individual stereo pairs are represented, with respect to the distance, then these object functions are integrated considering the position of cameras and the shape of the object functions. By integrating the general motion stereo images, we not only reduce the ambiguities in correspondence, but also improve the precision of the reconstruction. Also by introducing an adaptive window technique, we can alleviate the effect of projective distortion in matching features and improve the accuracy greatly. Experimental results on a synthetic and real data set are presented to demonstrate the performance of the proposed algorithm. Ja Seong Ku, Kyoung Mu Lee, Sang Uk Lee |
ICIP (2) | 2 |
| 1998 | Recognition of 2D Object Contours Using Starting-Point-Independent Wavelet Coefficient Matching
Hee Soo Yang, Sang Uk Lee, Kyoung Mu Lee |
J. Vis. Commun. Image Represent. | 3 |
| 1997 | Shape from Shading with a Generalized Reflectance Map Model
Kyoung Mu Lee, C.-C. Jay Kuo |
Comput. Vis. Image Underst. | 1 |
| 1996 | Shape from Photometric Ratio and Stereo
Kyoung Mu Lee, C.-C. Jay Kuo |
J. Vis. Commun. Image Represent. | 1 |
| 1993 | Shape from Shading with a Linear Triangular Element Surface ModelabstractThe authors propose to combine a triangular element surface model with a linearized reflectance map to formulate the shape-from-shading problem. The main idea is to approximate a smooth surface by the union of triangular surface patches called triangular elements and express the approximating surface as a linear combination of a set of nodal basis functions. Since the surface normal of a triangular element is uniquely determined by the heights of its three vertices (or nodes), image brightness can be directly related to nodal heights using the linearized reflectance map. The surface height can then be determined by minimizing a quadratic cost functional corresponding to the squares of brightness errors and solved effectively with the multigrid computational technique. The proposed method does not require any integrability constraint or artificial assumptions on boundary conditions. Simulation results for synthetic and real images are presented to illustrate the performance and efficiency of the method.> Kyoung Mu Lee, C.-C. Jay Kuo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Shape reconstruction from photometric stereoabstractTwo iterative algorithms for shape reconstruction based on multiple images taken under different lighting conditions, known as photometric stereo, are proposed. It is shown that single-image shape-from-shading (SFS) algorithms have an inherent problem, i.e., the accuracy of the reconstructed surface height is related to the slope of the reflectance map function defined on the gradient space. This observation motivates the authors to generalize the single-image SFS algorithm to two photometric stereo SFS algorithms aiming at more accurate surface reconstruction. The two algorithms directly determine the surface height by minimizing a quadratic cost functional, which is defined to be the square of the brightness error obtained from each individual image in a parallel or cascade manner. The optimal illumination condition that leads to best shape reconstruction is examined.> Kyoung Mu Lee, C.-C. Jay Kuo |
CVPR | 1 |