Nam Ik Cho

dblp:70/424 · DBLP profile ↗
← Back
196ranked-venue papers
10as first author
49since 2021 · last 2026
0000-0001-5297-4649ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 177 · 8 first-author · 46 since 2021Artificial intelligence and machine learning · 55 · 22 since 2021Systems, architecture and hardware · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Mamba-VOS: Efficient Video Object Segmentation with Selective State Space Models
Cheolhun Jang, Wontae Kim 0002, Daehyun Ji, Nam Ik Cho
ICPR (5)4
2026 STPose: Unseen Object Pose Estimation with a Single Template via Query-Aware 3D Reconstruction
Jaeguk Kim, Nam Ik Cho
ICPR (2)2
2026 DINOLight: Robust Ambient Light Normalization with Self-supervised Visual Prior Integration
Youngjin Oh, Junhyeong Kwon, Nam Ik Cho
ICPR (1)3
2026 M2Pose: Robust 6D Object Pose Estimation via Multi-frequency Surface Encoding and Multi-reference Voting
Jaewoo Park 0005, Jaeguk Kim, Nam Ik Cho
ICPR (1)3
2026 Faithful and Realistic Image Compression at Extreme-Low Bitrate via Pretrained Discrete Priors
abstract
Image compression is fundamentally governed by the distortion-perception trade-off, where achieving high perceptual realism often necessitates a sacrifice in pixel-wise fidelity. This challenge is significantly magnified at extreme low bitrates due to severe information constraints, creating a critical bottleneck for data representation. To address this representation bottleneck, we propose a framework that leverages pretrained discrete priors derived from large-scale vision encoders. Rather than training from scratch, we employ clustering-based initialization to map high-capacity discrete priors into a bitrate-constrained space. This ensures that reconstructed images achieve enhanced content fidelity while preserving superior perceptual realism. Furthermore, to overcome the practical bitrate overhead where arithmetic coding occasionally exceeds the theoretical uniform bound in discrete latent spaces, we introduce a dynamic entropy coding scheme. By deterministically switching between arithmetic and uniform coding on a per-image basis, our method ensures near-optimal compression efficiency. Extensive experiments confirm that our method improves content preservation while maintaining high-fidelity perceptual quality.
Nam Ik Cho
IEEE Signal Process. Lett.2
2025 APR-RD: Complemental Two Steps for Self-Supervised Real Image Denoising
abstract
Recent advancements in self-supervised denoising have made it possible to train models without needing a large amount of noisy-clean image pairs. A significant development in this area is the use of blind-spot networks (BSNs), which use single noisy images as training pairs by masking some input information to prevent noise transmission to the network output. Researchers have shown that BSNs are capable of reconstructing clean pixels from various types of independent pixel-wise degradations, such as synthetic additive white Gaussian noise (AWGN). However, unlike synthetic noise, real noise often contains highly correlated components which can induce noise transmission and reduce the performance of BSNs. To address the spatial correlation of real noise, we propose the Adjacent Pixel Replacer (APR), which decorrelates noise without a downsampling process that is widely adopted in previous research. The dissimilarity in our APR-generated pairs serves as relatively different noise components during training. Hence, it enables the BSN to block noise transmission while utilizing clean information effectively. As a result, BSN can utilize denser information to reconstruct the corresponding center pixel. We also propose Recharged Distillation (RD) to enhance high-frequency textures without additional network modifications. This method selectively refines clean information from recharged noisy pixels during distillation. Extensive experimental results demonstrate that our proposed method outperforms the existing state-of-the-art self-supervised denoising methods in real sRGB space.
Nam Ik Cho
AAAI2
2025 Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-Resolution
abstract
Transformer-based Super-Resolution (SR) methods have demonstrated superior performance compared to convolutional neural network (CNN)-based SR approaches due to their capability to capture long-range dependencies. However, their high computational complexity necessitates the development of lightweight approaches for practical use. To address this challenge, we propose the Attention-Sharing Information Distillation (ASID) network, a lightweight SR network that integrates attention-sharing and an information distillation structure specifically designed for Transformer-based SR methods. We modify the information distillation scheme, originally designed for efficient CNN operations, to reduce the computational load of stacked self-attention layers, effectively addressing the efficiency bottleneck. Additionally, we introduce attention-sharing across blocks to further minimize the computational cost of self-attention operations. By combining these strategies, ASID achieves competitive performance with existing SR methods while requiring only around 300K parameters – significantly fewer than existing CNN-based and Transformer-based SR models. Furthermore, ASID outperforms state-of-the-art SR methods when the number of parameters is matched, demonstrating its efficiency and effectiveness.
Karam Park, Jae Woong Soh, Nam Ik Cho
AAAI3
2025 RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objects
abstract
Estimating the 6D pose of unseen objects from monocular RGB images remains a challenging problem, especially due to the lack of prior object-specific knowledge. To tackle this issue, we propose RefPose, an innovative approach to object pose estimation that leverages a reference image and geometric correspondence as guidance. RefPose first predicts an initial pose by using object templates to render the reference image and establish the geometric correspondence needed for the refinement stage. During the refinement stage, RefPose estimates the geometric correspondence of the query based on the generated references and iteratively refines the pose through a render-and-compare approach. To enhance this estimation, we introduce a correlation volume-guided attention mechanism that effectively captures correlations between the query and reference images. Unlike traditional methods that depend on pre-defined object models, RefPose dynamically adapts to new object shapes by leveraging a reference image and geometric correspondence. This results in robust performance across previously unseen objects. Extensive evaluation on the BOP benchmark datasets shows that RefPose achieves state-of-the-art results while maintaining a competitive runtime.
Jaeguk Kim, Jaewoo Park 0005, Keuntek Lee, Nam Ik Cho
CVPR4
2025 Lightweight and Fast Real-Time Image Enhancement via Decomposition of the Spatial-Aware Lookup Tables
Wontae Kim 0002, Keuntek Lee, Nam Ik Cho
ICCV3
2025 Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity
abstract
Semantic watermarking techniques for latent diffusion models (LDMs) are robust against regeneration attacks, but often suffer from detection performance degradation due to the loss of frequency integrity. To tackle this problem, we propose a novel embedding method called Hermitian Symmetric Fourier Watermarking (SFW), which maintains frequency integrity by enforcing Hermitian symmetry. Additionally, we introduce a center-aware embedding strategy that reduces the vulnerability of semantic watermarking due to cropping attacks by ensuring robust information retention. To validate our approach, we apply these techniques to existing semantic watermarking schemes, enhancing their frequency-domain structures for better robustness and retrieval accuracy. Extensive experiments demonstrate that our methods achieve state-of-the-art verification and identification performance, surpassing previous approaches across various attack scenarios. Ablation studies confirm the impact of SFW on detection capabilities, the effectiveness of the center-aware embedding against cropping, and how message capacity influences identification accuracy. Notably, our method achieves the highest detection accuracy while maintaining superior image fidelity, as evidenced by FID and CLIP scores. Conclusively, our proposed SFW is shown to be an effective framework for balancing robustness and image fidelity, addressing the inherent trade-offs in semantic watermarking. Code available at https://github.com/thomas11809/SFWMark
Sung Ju Lee 0004, Nam Ik Cho
ICCV2
2025 Adapting Foundation Features via Cross-View Contrastive Learning for Unseen Object Pose Estimation
abstract
Recent advances in visual foundation models, such as DI-NOv2, have shown impressive generalization capabilities, making them a strong baseline for complex tasks like unseen object pose estimation. However, a significant domain gap between real images and rendered templates, combined with limited 3D awareness, hinders their effectiveness in precise correspondence matching and pose estimation. In this paper, we propose a new adaptation method that reduces this domain gap and enhances multi-view consistency through cross-view contrastive learning. Our approach aligns features from real and rendered images by using sampled poses and 2D-2D correspondences while also leveraging advanced augmentations and PCA-based loss computation to improve robustness and efficiency. When applied to the FoundPose model, our method achieves state-of-the-art performance on the BOP benchmark, surpassing previous methods with minimal fine-tuning.
Jaeguk Kim, Nam Ik Cho
ICIP2
2025 Towards Controllable Real Image Denoising With Camera Parameters
abstract
Recent deep learning-based image denoising methods have shown impressive performance; however, many lack the flexibility to adjust the denoising strength based on the noise levels, camera settings, and user preferences. In this paper, we introduce a new controllable denoising framework that adaptively removes noise from images by utilizing information from camera parameters. Specifically, we focus on ISO, shutter speed, and F-number, which are closely related to noise levels. We convert these selected parameters into a vector to control and enhance the performance of the denoising network. Experimental results show that our method seamlessly adds controllability to standard denoising neural networks and improves their performance. Code is available at https://github.com/OBAKSA/CPADNet.
Youngjin Oh, Junhyeong Kwon, Keuntek Lee, Nam Ik Cho
ICIP4
2025 SFS-NeRF: Enhancing Geometry Consistency in Few-Shot Novel View Synthesis Through Surface-Aware Neural Rendering
abstract
We propose a new pipeline for few-shot novel view synthesis called Surface-aware Few-Shot Neural Radiance Fields (SFS-NeRF). This approach enhances surface rigidity, resulting in more realistic rendering outcomes. Our model uti-lizes geometric constraints to better maintain the rigid surface of objects, effectively addressing the limitations encountered in previous methods that rely on sparse inputs. To achieve accurate color reconstruction of rendered images based on viewing direction, we employ Spherical Harmonics (SH) parameters. Additionally, we use the Laplace distribution to model the transparency function, which offers sharper max-ima compared to the logistic sigmoid distribution used in earlier works. During the training phase, we implement occlusion regularization and frequency regularization to further improve both the training and rendering processes. Experiments show that our method outperforms the baseline model on both the Blender and DTU datasets.
Seonji Park, Nam Ik Cho
ICIP2
2025 Efficient Implicit Neural Representations for Videos with Feature Modulation
abstract
Implicit neural representations for videos (NeRV) have gained attraction due to their ability to embed videos into neural networks effectively. A major challenge in implicit neural representations for videos is achieving high performance with small-sized networks. Therefore, it is crucial to minimize the intrinsic redundancy of the videos embedded within the network. In this paper, we propose a new method called ModNeRV (NeRV with feature modulation), which aims to reduce redundancy using base grids and feature modulation. The base grids, structured with multiple spatial resolutions, capture the general content of the video that is shared across frames. Frame-Specific features are then generated through frame-wise feature modulation applied to these base grids. Additionally, feature modulation is further utilized within ModNeRV blocks to generate details specific to each frame. As a result, ModNeRV demonstrates superior performance compared to previous NeRV-based methods in video reconstruction. Specifically, on the "Big Buck Bunny" and UVG datasets, ModNeRV achieves an increase in PSNR of 9.5 dB and 2.5 dB over NeRV with model sizes of 0.75M and 3M, respectively. Furthermore, in terms of video compression, ModNeRV outperforms both previous NeRV-based methods and traditional compression techniques such as H.264 (AVC) and H.265 (HEVC).
Haeyoon Yang, Nam Ik Cho
ICIP2
2025 Diffusion on Demand: Selective Caching and Modulation for Efficient Generation
abstract
Diffusion transformers demonstrate significant potential for various generation tasks but are challenged by high computational cost. Recently, feature caching methods have been introduced to improve inference efficiency by storing features at certain timesteps and reusing them at subsequent timesteps. However, their effectiveness is limited as they rely only on choosing between cached features and performing model inference. Motivated by high cosine similarity between features across consecutive timesteps, we propose a cache-based framework that reuses features and selectively adapts them through linear modulation. In our framework, the selection is performed via a modulation gate, and both the gate and modulation parameters are learned. Extensive experiments show that our method achieves similar generation performance to the original sampler while requiring significantly less computation. For example, FLOPs and inference latency are reduced by $2.93\times$ and $2.15\times$ for DiT-XL/2 and by $2.83\times$ and $1.50\times$ for PixArt-$\alpha$, respectively. We find that modulation is effective when applied to as little as 2\% of layers, resulting in negligible computation overhead.
Hee Min Choi, Hyoa Kang, Dokwan Oh, Nam Ik Cho
NeurIPS4
2025 Counting Guidance for High Fidelity Text-to-Image Synthesis
abstract
Recently, there have been significant improvements in the quality and performance of text-to-image generation, largely due to the impressive results attained by diffusion models. However, text-to-image diffusion models sometimes struggle to create high-fidelity content for the given input prompt. One specific issue is their difficulty in generating the precise number of objects specified in the text prompt. For example, when provided with the prompt “five apples and ten lemons on a table,” images generated by diffusion models often contain an incorrect number of objects. In this paper, we present a method to improve diffusion models so that they accurately produce the correct object count based on the input prompt. We adopt a counting network that performs reference-less class-agnostic counting for any given image. We calculate the gradients of the counting network and refine the predicted noise for each step. To address the presence of multiple types of objects in the prompt, we utilize novel attention map guidance to obtain high-quality masks for each object. Finally, we guide the denoising process using the calculated gradients for each object. Through extensive experiments and evaluation, we demonstrate that the proposed method significantly enhances the fidelity of diffusion models with respect to object count.
Wonjun Kang, Kevin Galim, Hyung Il Koo, Nam Ik Cho
WACV4
2025 Self-Supervised Learning with Probabilistic Density Labeling for Rainfall Probability Estimation
abstract
Numerical weather prediction (NWP) models are fundamental in meteorology for simulating and forecasting the behavior of various atmospheric variables. The accuracy of precipitation forecasts and the acquisition of sufficient lead time are crucial for preventing hazardous weather events. However, the performance of NWP models is limited by the nonlinear and unpredictable patterns of extreme weather phenomena driven by temporal dynamics. In this regard, we propose a Self-Supervised Learning with Probabilistic Density Labeling (SSLPDL) for estimating rainfall probability by post-processing NWP forecasts. Our post-processing method uses self-supervised learning (SSL) with masked modeling for reconstructing atmospheric physics variables, enabling the model to learn the dependency between variables. The pre-trained encoder is then utilized in transfer learning to a precipitation segmentation task. Furthermore, we introduce a straightforward labeling approach based on probability density to address the class imbalance in extreme weather phenomena like heavy rain events. Experimental results show that SSLPDL surpasses other precipitation forecasting models in regional precipitation post-processing and demonstrates competitive performance in extending forecast lead times. Our code is available at https://github.com/joonha425/SSLPDL.
Junha Lee, Sojung An, Su Jeong You, Nam Ik Cho
WACV4
2025 Partial Filter-Sharing: Improved Parameter-sharing Method for Single Image Super-Resolution Networks
abstract
Numerous deep learning techniques have been developed for Single Image Super-Resolution (SISR), leading to significant performance improvements. However, these techniques have also resulted in a substantial increase in parameter size. As a result, there is a growing interest in reducing network complexity for more practical usage while still maintaining high SR quality. One such method is parameter-sharing, which includes recursive, recurrent, and multi-scale learning approaches. However, sharing identical kernels across layers or up-scaling tasks can reduce the network's representational capacity. To address this, we propose Partial filter-Sharing (PS), a new parameter-sharing method that preserves the network's representational power more effectively than previous approaches. Instead of sharing a single filter, PS shares segments of filters, called partial filters, across layers. This approach enables parameter-sharing layers to use diverse filters for each layer or task, striking a balance between parameter efficiency and the network's representational ability without imposing excessive computational or parameter overhead. Furthermore, the PS framework provides precise control over the network's performance and complexity by adjusting the quantity of partial filters. Extensive experiments demonstrate that our PS framework outperforms traditional parameter-sharing super-resolution (SR) methods without incurring excessive additional parameters or computational cost. Our code is available here: https://github.com/saturnian77/Partial_filter-Sharing.
Karam Park, Nam Ik Cho
WACV2
2025 Explicit Guidance for Robust Video Frame Interpolation Against Discontinuous Motions
abstract
Nowadays, many videos contain graphic elements such as logos, subtitles, and user interfaces. These overlayed elements exhibit discontinuous motions, characterized by static or instantaneous motions that are neither spatially nor temporally coherent. As existing Video Frame Inter-polation (VFI) methods rely on motion-compensation techniques, they work best for videos with continuous motions but face limitations against videos with discontinuous motion. In this paper, we propose a simple framework to enhance the robustness of existing VFI models against discontinuous motion. We first identify key properties that distinguish discontinuous from continuous motion. They are then leveraged by the Discontinuity map (D-map) estimator to explicitly guide the classification of continuous and discontinuous areas through a coherence mask and an additional supervisory signal. Our framework separately interpolates the predicted continuous and discontinuous regions to achieve state-of-the-art performance against synthetic discontinuous motions while also generalizing well to real-world discontinuous motions. Moreover, our framework's ‘plug-and-play’ design enables easy application to existing VFI models without the need for retraining and maintains strong performance on continuous motion.
Nam Ik Cho
WACV2
2025 UPVIS: upsampled video query for offline video instance segmentation
Junho Jo, Haesoo Chung, Joon Seok Lee, Dongyoon Wee, Nam Ik Cho
Multim. Tools Appl.5
2025 GeoFlow: geometry-guided optical flow refinement for 6D object pose estimation
Jaeguk Kim, Jaewoo Park 0005, Nam Ik Cho
Multim. Tools Appl.3
2025 Single-stage table structure recognition approach via efficient sequence modeling
Yeonsik Jo, Soonyoung Lee, Nam Ik Cho
Pattern Recognit. Lett.5
2025 Blind Image Super-Resolution With Efficient Network Design Using Frequency Domain Information
abstract
Blind Image Super-Resolution (BSR) tackles the challenge of enhancing image resolution that has been degraded by unknown kernels. Although existing BSR models have achieved remarkable results, they often demand significant computational resources to manage various degradation kernels. Recent studies have utilized large models with parameter counts ranging from 4M to 20M, resulting in computational costs exceeding 200G Multi-Adds. In this paper, we introduce a blind super-resolution method, which is the first lightweight non-iterative approach for BSR. By leveraging the connection between degradation kernel shapes and the frequency-domain characteristics of low-resolution images, we simplify the kernel estimation process, thereby reducing the overall model complexity. Additionally, we employ a constrained least squares approach to refine the low-resolution image using the estimated kernel, which serves as the input to the hierarchical Transformer blocks. Our approach delivers competitive performance while requiring only 1.7M parameters and 2G Multi-Adds. Experimental results demonstrate that our method achieves comparable results with significantly smaller network size
Sunwoo Cho, Nam Ik Cho
IEEE Signal Process. Lett.2
2025 Dataset Distillation for Super-Resolution Without Class Labels and Pre-Trained Models
abstract
Training deep neural networks has become increasingly demanding, requiring large datasets and significant computational resources, especially as model complexity advances. Data distillation methods, which aim to improve data efficiency, have emerged as promising solutions to this challenge. In the field of single image super-resolution (SISR), the reliance on large training datasets highlights the importance of these techniques. Recently, a generative adversarial network (GAN) inversion-based data distillation framework for SR was proposed, showing potential for better data utilization. However, the current method depends heavily on pre-trained SR networks and class-specific information, limiting its generalizability and applicability. To address these issues, we introduce a new data distillation approach for image SR that does not need class labels or pre-trained SR models. In particular, we first extract high- gradient patches and categorize images based on CLIP features, then fine-tune a diffusion model on the selected patches to learn their distribution and synthesize distilled training images. Experimental results show that our method achieves state-of-the-art performance while using significantly less training data and requiring less computational time. Specifically, when we train a baseline Transformer model for SR with only 0.68% of the original dataset, the performance drop is just 0.3 dB. In this case, diffusion model fine-tuning takes 4 hours, and SR model training completes within 1 hour, much shorter than the 11-hour training time with the full dataset.
Sunwoo Cho, Yejin Jung, Nam Ik Cho, Jae Woong Soh
IEEE Signal Process. Lett.3
2025 Efficiently Trained Real Image Dehazing Network With Dual Discrete Priors for Enhanced Naturalness
abstract
In this paper, we efficiently train a dehazing network with enhanced performance by introducing new network architectures and objective functions. Our dehazing network uses high-quality discrete priors from a vector quantization network pretrained on clean images. To mitigate the prolonged pretraining time of existing methods, we analyze the metrics related to discrete priors and propose criteria for early stopping, significantly reducing training time. Furthermore, we introduce dual branches, namely the texture and structure branches, into the dehazing network. The branches act as priors, consisting of pretrained components. To enhance naturalness, we apply our new Structure Alignment Loss with the structure branch which is active only during training, and adopt losses in the frequency domain. Moreover, our analysis of the quantization gap between real and synthetic data shows that additional domain adaptation is unnecessary. Experiments demonstrate that our method outperforms strong baselines on real-world datasets.
Nam Ik Cho
IEEE Signal Process. Lett.2
2024 Image-Adaptive 3D Lookup Tables for Real-Time Image Enhancement with Bilateral Grids
Wontae Kim 0002, Nam Ik Cho
ECCV (49)2
2024 Panel-Specific Degradation Representation for Raw Under-Display Camera Image Restoration
Youngjin Oh, Keuntek Lee, Dae-Hyun Lee, Nam Ik Cho
ECCV (48)5
2024 RFG-HDR: Representative Feature-Guided Transformer For Multi-Exposure High Dynamic Range Imaging
abstract
Multi-exposure fusion is a high dynamic range (HDR) imaging technique that combines multiple low dynamic range (LDR) images of a scene with varying exposure times to produce a single high-quality HDR image. Since each LDR frame is captured with a different exposure time (bias), it is crucial to extract meaningful features from each differently-exposed LDR frame for producing high-quality HDR images. This paper introduces a new contrastive learning method that provides a versatile way of extracting characteristic features from LDR frames by considering the relationship between LDR frames. Additionally, we introduce Representative Feature-Guided Transformer (RFGHDR), a new architecture that utilizes contrastive-learned representations to improve frame alignment and merging. Based on extensive experiments on various datasets, we have found that the RFG-HDR performs better than existing multi-exposure HDR imaging methods in terms of various evaluation metrics. Our work will be released on https://github.com/KeuntekLee/RFG-HDR.
Keuntek Lee, Gu Yong Park, Nam Ik Cho
ICIP4
2024 Multi-Reference Flow-Guided Cross-Domain Reconstruction For General Object 6D Pose Estimation
abstract
Estimating the pose of an unseen object is challenging since the pose in object space and the 3D shape are not trainable. Previous methods relied on either template matching with numerous references or Transformer-based local feature matching. However, both approaches focused solely on the appearance of the rendered image and did not consider the object’s shape properties. Therefore, we propose an optical flow-based cross-domain reconstruction method to leverage the object’s geometric features during pose estimation. Additionally, we introduce a cycle-consistent loss to utilize reconstruction self-supervision and a novel deformable aggregator to effectively integrate misaligned features from each reference. Our experiments demonstrate that the proposed method successfully estimates the unseen object’s geometric features and shows competitive performance in general object pose estimation while maintaining fast inference time.
Jaewoo Park 0005, Jaeguk Kim, Nam Ik Cho
ICIP3
2024 SANERV: Scene-Adaptive Neural Representation for Videos
abstract
Neural Representations for Videos (NeRV) is a neural network model that learns video representation through image-wise implicit neural representation (INR). Although early NeRV-based research has improved representation learning, it overlooked video properties. In this paper, we propose a new method called Scene-Adaptive Neural Representation for Videos (SANeRV) that considers videos’ general properties and adapts to individual video characteristics. To achieve this, we incorporate temporal redundancy, which is a typical feature of videos, by splitting the learning process into two parts: global representation and residual representation. Moreover, we adjust the network architecture based on the presence of scene changes and the dynamism of the video to cater to individual video characteristics. Experiments show that SANeRV achieves state-of-the-art performance in video regression and compression tasks for benchmark datasets.
Hochang Rhee, Haesoo Chung, Junho Jo, Nam Ik Cho
ICIP5
2024 Enhancing Multi-exposure High Dynamic Range Imaging with Overlapped Codebook for Improved Representation Learning
Keuntek Lee, Nam Ik Cho
ICPR (32)3
2023 Perception-Oriented Single Image Super-Resolution using Optimal Objective Estimation
abstract
Single-image super-resolution (SISR) networks trained with perceptual and adversarial losses provide high-contrast outputs compared to those of networks trained with distortion-oriented losses, such as$L1$or$L2$. However, it has been shown that using a single perceptual loss is insufficient for accurately restoring locally varying diverse shapes in images, often generating undesirable artifacts or unnatural details. For this reason, combinations of various losses, such as perceptual, adversarial, and distortion losses, have been attempted, yet it remains challenging to find optimal combinations. Hence, in this paper, we propose a new SISR framework that applies optimal objectives for each region to generate plausible results in overall areas of high-resolution outputs. Specifically, the framework comprises two models: a predictive model that infers an optimal objective map for a given low-resolution (LR) input and a generative model that applies a target objective map to produce the corresponding SR output. The generative model is trained over our proposed objective trajectory representing a set of essential objectives, which enables the single network to learn various SR results corresponding to combined losses on the trajectory. The predictive model is trained using pairs of LR images and corresponding optimal objective maps searched from the objective trajectory. Experimental results on five benchmarks show that the proposed method outperforms state-of-the-art perception-driven SR methods in LPIPS, DISTS, PSNR, and SSIM metrics. The visual results also demonstrate the superiority of our method in perception-oriented reconstruction. The code is available at https://github.com/seungho-snu/SROOE.
Seung Ho Park, Young-Su Moon, Nam Ik Cho
CVPR3
2023 LAN-HDR: Luminance-based Alignment Network for High Dynamic Range Video Reconstruction
abstract
As demands for high-quality videos continue to rise, high-resolution and high-dynamic range (HDR) imaging techniques are drawing attention. To generate an HDR video from low dynamic range (LDR) images, one of the critical steps is the motion compensation between LDR frames, for which most existing works employed the optical flow algorithm. However, these methods suffer from flow estimation errors when saturation or complicated motions exist. In this paper, we propose an end-to-end HDR video composition framework, which aligns LDR frames in the feature space and then merges aligned features into an HDR frame, without relying on pixel-domain optical flow. Specifically, we propose a luminance-based alignment network for HDR (LAN-HDR) consisting of an alignment module and a hallucination module. The alignment module aligns a frame to the adjacent reference by evaluating luminance-based attention, excluding color information. The hallucination module generates sharp details, especially for washed-out areas due to saturation. The aligned and hallucinated features are then blended adaptively to complement each other. Finally, we merge the features to generate a final HDR frame. In training, we adopt a temporal loss, in addition to frame reconstruction losses, to enhance temporal consistency and thus reduce flickering. Extensive experiments demonstrate that our method performs better or comparable to state-of-the-art methods on several benchmarks. Codes are available at https://github.com/haesoochung/LAN-HDR.
Haesoo Chung, Nam Ik Cho
ICCV2
2023 Self-supervised Image Denoising with Downsampled Invariance Loss and Conditional Blind-Spot Network
abstract
There have been many image denoisers using deep neural networks, which outperform conventional model-based methods by large margins. Recently, self-supervised methods have attracted attention because constructing a large real noise dataset for supervised training is an enormous burden. The most representative self-supervised denoisers are based on blind-spot networks, which exclude the receptive field’s center pixel. However, excluding any input pixel is abandoning some information, especially when the input pixel at the corresponding output position is excluded. In addition, a standard blind-spot network fails to reduce real camera noise due to the pixel-wise correlation of noise, though it successfully removes independently distributed synthetic noise. Hence, to realize a more practical denoiser, we propose a novel self-supervised training framework that can remove real noise. For this, we derive the theoretic upper bound of a supervised loss where the network is guided by the downsampled blinded output. Also, we design a conditional blind-spot network (C-BSN), which selectively controls the blindness of the network to use the center pixel information. Furthermore, we exploit a random subsampler to decorrelate noise spatially, making the C-BSN free of visual artifacts that were often seen in downsample-based methods. Extensive experiments show that the proposed C-BSN achieves state-of-the-art performance on real-world datasets as a self-supervised denoiser and shows qualitatively pleasing results without any post-processing or refinement.
Yeong Il Jang, Keuntek Lee, Gu Yong Park, Seyun Kim, Nam Ik Cho
ICCV5
2023 OCVOS: Object-Centric Representation for Video Object Segmentation
abstract
Semi-supervised video object segmentation (VOS) methods aim to segment target objects with the help of pixel-level annotations in the first frame. Many methods employ Transformer-based attention modules to propagate the given annotations in the first frame to the most similar patch or pixel in the following frames. Although they have shown impressive results, they can still be prone to errors in challenging scenes with multiple overlapping objects. To tackle this problem, we propose an object-centric VOS (OCVOS) method that exploits query-based Transformer decoder blocks. After aggregating target object information with typical matching-based approaches, the Transformer networks extract object-wise information by interacting with object queries. In this way, the proposed method considers not only global and contextual information but also object-centric representations. We validate its effectiveness in inducing object-wise information compared to existing methods on the DAVIS and YouTube-VOS benchmarks.
Junho Jo, Dongyoon Wee, Nam Ik Cho
ICIP3
2023 Lossless Compression of Raw Images by Learning the Prediction and Frequency Decomposition
abstract
This paper presents a lossless color filter array (CFA) image compression method that attempts to maximize the use of image correlations from two perspectives. First, we adopt a hierarchical approach for decomposing an input into subimages and proceed with the subimage encoding with a new encoding order. The new encoding scheme is shown to provide better prediction performances compared to conventional raster scan orders. Secondly, for each subimage, we compress the low-frequency components first and use them as additional input for encoding the remaining high-frequency components. In this scenario, the high-frequency components show improved compression efficiency since they use the low-frequency ones as strong prior. Experiments show that the proposed method achieves state-of-the-art performance for 8∼16 bits raw images.
Hochang Rhee, Nam Ik Cho
VCIP2
2023 WHFL: Wavelet-Domain High Frequency Loss for Sketch-to-Image Translation
abstract
Even a rough sketch can effectively convey the descriptions of objects, as humans can imagine the original shape from the sketch. The sketch-to-photo translation is a computer vision task that enables a machine to do this imagination, taking a binary sketch image and generating plausible RGB images corresponding to the sketch. Hence, deep neural networks for this task should learn to generate a wide range of frequencies because most parts of the input (binary sketch image) are composed of DC signals. In this paper, we propose a new loss function named Wavelet-domain High-Frequency Loss (WHFL) to overcome the limitations of previous methods that tend to have a bias toward low frequencies. The proposed method emphasizes the loss on the high frequencies by designing a new weight matrix imposing larger weights on the high bands. Unlike existing handcraft methods that control frequency weights using binary masks, we use the matrix with finely controlled elements according to frequency scales. The WHFL is designed in a multi-scale form, which lets the loss function focus more on the high frequency according to decomposition levels. We use the WHFL as a complementary loss in addition to conventional ones defined in the spatial domain. Experiments show we can improve the qualitative and quantitative results in both spatial and frequency domains. Additionally, we attempt to verify the WHFL’s high-frequency generation capability by defining a new evaluation metric named Unsigned Euclidean Distance Field Error (UEDFE).
Nam Ik Cho
WACV2
2023 A Dynamic Residual Self-Attention Network for Lightweight Single Image Super-Resolution
abstract
Deep learning methods have shown outstanding performance in many applications, including single-image super-resolution (SISR). With residual connection architecture, deeply stacked convolutional neural networks provide a substantial performance boost for SISR, but their huge parameters and computational loads are impractical for real-world applications. Thus, designing lightweight models with acceptable performance is one of the major tasks in current SISR research. The objective of lightweight network design is to balance a computational load and reconstruction performance. Most of the previous methods have manually designed complex and predefined fixed structures, which generally required a large number of experiments and lacked flexibility in the diversity of input image statistics. In this paper, we propose a dynamic residual self-attention network (DRSAN) for lightweight SISR, while focusing on the automated design of residual connections between building blocks. The proposed DRSAN has dynamic residual connections based on dynamic residual attention (DRA), which adaptively changes its structure according to input statistics. Specifically, we propose a dynamic residual module that explicitly models the DRA by finding the interrelation between residual paths and input image statistics, as well as assigning proper weights to each residual path. We also propose a residual self-attention (RSA) module to further boost the performance, which produces 3-dimensional attention maps without additional parameters by cooperating with residual structures. The proposed dynamic scheme, exploiting the combination of DRA and RSA, shows an efficient trade-off between computational complexity and network performance. Experimental results show that the DRSAN performs better than or comparable to existing state-of-the-art lightweight models for SISR.
Karam Park, Jae Woong Soh, Nam Ik Cho
IEEE Trans. Multim.3
2022 LC-FDNet: Learned Lossless Image Compression with Frequency Decomposition Network
abstract
Recent learning-based lossless image compression methods encode an image in the unit of subimages and achieve comparable performances to conventional non-learning algorithms. However, these methods do not consider the performance drop in the high-frequency region, giving equal consideration to the low and high-frequency areas. In this paper, we propose a new lossless image compression method that proceeds the encoding in a coarse-to-fine manner to separate and process low and high-frequency regions differently. We initially compress the low-frequency components and then use them as additional input for encoding the remaining high-frequency region. The low-frequency components act as a strong prior in this case, which leads to improved estimation in the high-frequency area. In addition, we design the frequency decomposition process to be adptive to color channel, spatial location, and image characteristics. As a result, our method derives an image-specific optimal ratio of low/high-frequency components. Experiments show that the proposed method achieves state-of-the-art performance for benchmark high-resolution datasets.
Hochang Rhee, Yeong Il Jang, Seyun Kim, Nam Ik Cho
CVPR4
2022 Deep Hash Distillation for Image Retrieval
Young Kyun Jang, Geonmo Gu, Byungsoo Ko, Isaac Kang, Nam Ik Cho
ECCV (14)5
2022 DProST: Dynamic Projective Spatial Transformer Network for 6D Pose Estimation
Jaewoo Park 0005, Nam Ik Cho
ECCV (6)2
2022 Disentangled Feature-Guided Multi-Exposure High Dynamic Range Imaging
abstract
Multi-exposure high dynamic range (HDR) imaging aims to generate an HDR image from multiple differently exposed low dynamic range (LDR) images. It is a challenging task due to two major problems: (1) there are usually misalignments among the input LDR images, and (2) LDR images often have incomplete information due to under-/over-exposure. In this paper, we propose a disentangled feature-guided HDR network (DFGNet) to alleviate the above-stated problems. Specifically, we first extract and disentangle exposure features and spatial features of input LDR images. Then, we process these features through the proposed DFG modules, which produce a high-quality HDR image. Experiments show that the proposed DFGNet achieves outstanding performance on a benchmark dataset. Our code and more results are available at https://github.com/KeuntekLee/DFGNet.
Keuntek Lee, Yeong Il Jang, Nam Ik Cho
ICASSP3
2022 Episode Difficulty Based Sampling Method for Few-Shot Classification
Hochang Rhee, Nam Ik Cho
ICIP2
2022 Self-Supervised Pretraining for Deep Hash-Based Image Retrieval
abstract
Deep hashing aims to produce discriminative binary hash codes for fast image retrieval through a deep baseline network and additional trainable hash function. In a supervised deep hashing network, the baseline network is generally initialized with classification-based pretrained models, and the overall hashing network is trained in a supervised fashion. However, since classification and retrieval are two different tasks, it is necessary to reconsider the initial model for the baseline network. In this paper, we propose to use a self-supervised pretrained model as the baseline for the first time. We investigate the impact of pretrained model types by comparing deep hashing networks that use the baseline network with 1) randomly initialized weights, 2) conventional supervised pretrained weights, and 3) proposed self-supervised pretrained weights. As a result, we confirm that the performance of deep hashing differs depending on the initial baseline setting, and the proposed self-supervised baseline model shows comparable or better performance over the supervised one. Our code is released at https://github.com/HaeyoonYang/SSPH.
Haeyoon Yang, Young Kyun Jang, Isaac Kang, Nam Ik Cho
ICIP4
2022 High Dynamic Range Imaging of Dynamic Scenes with Saturation Compensation but without Explicit Motion Compensation
abstract
High dynamic range (HDR) imaging is a highly challenging task since a large amount of information is lost due to the limitations of camera sensors. For HDR imaging, some methods capture multiple low dynamic range (LDR) images with altering exposures to aggregate more information. However, these approaches introduce ghosting artifacts when significant inter-frame motions are present. Moreover, although multi-exposure images are given, we have little information in severely over-exposed areas. Most existing methods focus on motion compensation, i.e., alignment of multiple LDR shots to reduce the ghosting artifacts, but they still produce unsatisfying results. These methods also rather overlook the need to restore the saturated areas. In this paper, we generate well-aligned multi-exposure features by reformulating a motion alignment problem into a simple brightness adjustment problem. In addition, we propose a coarse-to-fine merging strategy with explicit saturation compensation. The saturated areas are reconstructed with similar well-exposed content using adaptive contextual attention. We demonstrate that our method outperforms the state-of-the-art methods regarding qualitative and quantitative evaluations.
Haesoo Chung, Nam Ik Cho
WACV2
2022 Deep-learning and graph-based approach to table structure recognition
Jaewoo Park 0005, Hyung Il Koo, Nam Ik Cho
Multim. Tools Appl.4
2022 Variational Deep Image Restoration
abstract
This paper presents a new variational inference framework for image restoration and a convolutional neural network (CNN) structure that can solve the restoration problems described by the proposed framework. Earlier CNN-based image restoration methods primarily focused on network architecture design or training strategy with non-blind scenarios where the degradation models are known or assumed. For a step closer to real-world applications, CNNs are also blindly trained with the whole dataset, including diverse degradations. However, the conditional distribution of a high-quality image given a diversely degraded one is too complicated to be learned by a single CNN. Therefore, there have also been some methods that provide additional prior information to train a CNN. Unlike previous approaches, we focus more on the objective of restoration based on the Bayesian perspective and how to reformulate the objective. Specifically, our method relaxes the original posterior inference problem to better manageable sub-problems and thus behaves like a divide-and-conquer scheme. As a result, the proposed framework boosts the performance of several restoration problems compared to the previous ones. Specifically, our method delivers state-of-the-art performance on Gaussian denoising, real-world noise reduction, blind image super-resolution, and JPEG compression artifacts reduction. Our code and more details are available on our project page, https://github.com/JWSoh/VDIR.
Jae Woong Soh, Nam Ik Cho
IEEE Trans. Image Process.2
2022 Inverse-Based Approach to Explaining and Visualizing Convolutional Neural Networks
abstract
This article presents a new method for understanding and visualizing convolutional neural networks (CNNs). Most existing approaches to this problem focus on a global score and evaluate the pixelwise contribution of inputs to the score. The analysis of CNNs for multilabeled outputs or regression has not yet been considered in the literature, despite their success on image classification tasks with well-defined global scores. To address this problem, we propose a new inverse-based approach that computes the inverse of a feedforward pass to identify activations of interest in lower layers. We developed a layerwise inverse procedure based on two observations: 1) inverse results should have consistent internal activations to the original forward pass and 2) a small amount of activation in inverse results is desirable for human interpretability. Experimental results show that the proposed method allows us to analyze CNNs for classification and regression in the same framework. We demonstrated that our method successfully finds attributions in the inputs for image classification with comparable performance to state-of-the-art methods. To visualize the tradeoff between various methods, we developed a novel plot that shows the tradeoff between the amount of activations and the rate of class reidentification. In the case of regression, our method showed that conventional CNNs for single image super-resolution overlook a portion of frequency bands that may result in performance degradation.
Hyuk Jin Kwon, Hyung Il Koo, Jae Woong Soh, Nam Ik Cho
IEEE Trans. Neural Networks Learn. Syst.4
2021 Self-supervised Product Quantization for Deep Unsupervised Image Retrieval
abstract
Supervised deep learning-based hash and vector quantization are enabling fast and large-scale image retrieval systems. By fully exploiting label annotations, they are achieving outstanding retrieval performances compared to the conventional methods. However, it is painstaking to assign labels precisely for a vast amount of training data, and also, the annotation process is error-prone. To tackle these issues, we propose the first deep unsupervised image retrieval method dubbed Self-supervised Product Quantization (SPQ) network, which is label-free and trained in a self-supervised manner. We design a Cross Quantized Contrastive learning strategy that jointly learns codewords and deep visual descriptors by comparing individually transformed images (views). Our method analyzes the image contents to extract descriptive features, allowing us to understand image representations for accurate retrieval. By conducting extensive experiments on benchmarks, we demonstrate that the proposed method yields state-of-the-art results even without supervised pretraining.
Young Kyun Jang, Nam Ik Cho
ICCV2
2020 Generalized Product Quantization Network for Semi-Supervised Image Retrieval
abstract
Image retrieval methods that employ hashing or vector quantization have achieved great success by taking advantage of deep learning. However, these approaches do not meet expectations unless expensive label information is sufficient. To resolve this issue, we propose the first quantization-based semi-supervised image retrieval scheme: Generalized Product Quantization (GPQ) network. We design a novel metric learning strategy that preserves semantic similarity between labeled data, and employ entropy regularization term to fully exploit inherent potentials of unlabeled data. Our solution increases the generalization capacity of the quantization network, which allows overcoming previous limitations in the retrieval community. Extensive experimental results demonstrate that GPQ yields state-of-the-art performance on large-scale real image benchmark datasets.
Young Kyun Jang, Nam Ik Cho
CVPR2
2020 Transfer Learning From Synthetic to Real-Noise Denoising With Adaptive Instance Normalization
abstract
Real-noise denoising is a challenging task because the statistics of real-noise do not follow the normal distribution, and they are also spatially and temporally changing. In order to cope with various and complex real-noise, we propose a well-generalized denoising architecture and a transfer learning scheme. Specifically, we adopt an adaptive instance normalization to build a denoiser, which can regularize the feature map and prevent the network from overfitting to the training set. We also introduce a transfer learning scheme that transfers knowledge learned from synthetic-noise data to the real-noise denoiser. From the proposed transfer learning, the synthetic-noise denoiser can learn general features from various synthetic-noise data, and the real-noise denoiser can learn the real-noise characteristics from real data. From the experiments, we find that the proposed denoising method has great generalization ability, such that our network trained with synthetic-noise achieves the best performance for Darmstadt Noise Dataset (DND) among the methods from published papers. We can also see that the proposed transfer learning scheme robustly works for real-noise images through the learning with a very small number of labeled data.
Yoonsik Kim, Jae Woong Soh, Gu Yong Park, Nam Ik Cho
CVPR4
2020 Meta-Transfer Learning for Zero-Shot Super-Resolution
abstract
Convolutional neural networks (CNNs) have shown dramatic improvements in single image super-resolution (SISR) by using large-scale external samples. Despite their remarkable performance based on the external dataset, they cannot exploit internal information within a specific image. Another problem is that they are applicable only to the specific condition of data that they are supervised. For instance, the low-resolution (LR) image should be a "bicubic" downsampled noise-free image from a high-resolution (HR) one. To address both issues, zero-shot super-resolution (ZSSR) has been proposed for flexible internal learning. However, they require thousands of gradient updates, i.e., long inference time. In this paper, we present Meta-Transfer Learning for Zero-Shot Super-Resolution (MZSR), which leverages ZSSR. Precisely, it is based on finding a generic initial parameter that is suitable for internal learning. Thus, we can exploit both external and internal information, where one single gradient update can yield quite considerable results. With our method, the network can quickly adapt to a given image condition. In this respect, our method can be applied to a large spectrum of image conditions within a fast adaptation process.
Jae Woong Soh, Sunwoo Cho, Nam Ik Cho
CVPR3
2020 Channel-Wise Progressive Learning For Lossless Image Compression
abstract
This paper presents a channel-wise progressive coding system for lossless compression of color images. We follow the classical lossless compression scheme of LOCO-I and CALIC, where pixel values and coding contexts are predicted and forwarded to the entropy coder for compression. The contribution is that we jointly estimate the pixel values and coding contexts from neighboring pixels by training a simple multilayer perceptron in a residual and channel-wise progressive manner. Specifically, we obtain accurate pixel prediction along with coding contexts that reflect the magnitude of local activity very well. These results are sent to an adaptive arithmetic coder that appropriately encodes the prediction error according to the corresponding coding context. Experimental results demonstrate the effectiveness of the proposed method in high-resolution datasets.
Hochang Rhee, Yeong Il Jang, Seyun Kim, Nam Ik Cho
ICIP4
2020 Neural Architecture Search for Image Super-Resolution Using Densely Constructed Search Space: DeCoNAS
abstract
The recent progress of deep convolutional neural networks has enabled great success in single image super-resolution (SISR) and many other vision tasks. Their performances are also being increased by deepening the networks and developing more sophisticated network structures. However, finding an optimal structure for the given problem is a difficult task, even for human experts. For this reason, neural architecture search (NAS) methods have been introduced, which automate the procedure of constructing the structures. In this paper, we expand the NAS to the super-resolution domain and find a lightweight densely connected network named DeCoNASNet. We use a hierarchical search strategy to find the best connection with local and global features. In this process, we define a complexity-based penalty for solving image super-resolution, which can be considered a multi-objective problem. Experiments show that our DeCoNASNet outperforms the state-of-the-art lightweight super-resolution networks designed by handcraft methods and existing NAS-based design.
Joonyoung Ahn, Nam Ik Cho
ICPR2
2020 User-Independent Gaze Estimation by Extracting Pupil Parameter and Its Mapping to the Gaze Angle
abstract
Since gaze estimation plays a crucial role in recognizing human intentions, it has been researched for a long time, and its accuracy is ever increasing. However, due to the wide variation in eye shapes and focusing abilities between the individuals, accuracies of most algorithms vary depending on each person in the test group, especially when the initial calibration is not well performed. To alleviate the user-dependency, we attempt to derive features that are general for most people and use them as the input to a deep network instead of using the images as the input. Specifically, we use the pupil shape as the core feature because it is directly related to the 3D eyeball rotation, and thus the gaze direction. While existing deep learning methods learn the gaze point by extracting various features from the image, we focus on the mapping function from the eyeball rotation to the gaze point by using the pupil shape as the input. It is shown that the accuracy of gaze point estimation also becomes robust for the uncalibrated points by following the characteristics of the mapping function. Also, our gaze network learns the gaze difference to facilitate the re-calibration process to fix the calibration-drift problem that typically occurs with glass-type or head-mount devices.
Sang Yoon Han, Nam Ik Cho
ICPR2
2020 Improving Explainability of Integrated Gradients with Guided Non-Linearity
abstract
Along with the performance improvements of neural network models, developing methods that enable the explanation of their behavior is a significant research topic. For convolutional neural networks, the explainability is usually achieved with attribution (heatmap) that visualizes pixel-level importance or contribution of input to its corresponding result. This attribution should reflect the relation (dependency) between inputs and outputs, which has been studied with a variety of methods, e.g., derivative of an output with respect to an input pixel value, a weighted sum of gradients, amount of output changes to input perturbations, and so on. In this paper, we present a new method that improves the measure of attribution, and incorporates it into the integrated gradients method. To be precise, rather than using the conventional chain-rule, we propose a method called guided non-linearity that propagates gradients more effectively through non-linear units (e.g., ReLU and max-pool) so that only positive gradients backpropagate through nonlinear units. Our method is inspired by the mechanism of action potential generation in postsynaptic neurons, where the firing of action potentials depends on the sum of excitatory (EPSP) and inhibitory postsynaptic potentials (IPSP). We believe that paths consisting of EPSP-giving-neurons faithfully reflect the contribution of inputs to the output, and we make gradients flow only along those paths (i.e., paths of positive chain reactions). Experiments with 5 deep neural networks have shown that the proposed method outperforms others in terms of the deletion metrics, and yields fine-grained and more human-interpretable attribution.
Hyuk Jin Kwon, Hyung Il Koo, Nam Ik Cho
ICPR3
2020 Single Image Super-Resolution with Dynamic Residual Connection
abstract
Deep convolutional neural networks have shown significant improvement in the single image super-resolution (SISR) field. Recently, there have been attempts to solve the SISR problem using lightweight networks, considering limited computational resources for real-world applications. Especially for lightweight networks, balancing between parameter demand and performance is very difficult to adjust, and most lightweight SISR networks are manually designed based on a huge number of brute-force experiments. Besides, a critical key to the network performance relies on the skip connection of building blocks that are repeatedly in the architecture. Notably, in previous works, these connections are pre-defined and manually determined by human researchers. Hence, they are less flexible to the input image statistics, and there can be a better solution for the given number of parameters. Therefore, we focus on the automated design of networks regarding the connection of basic building blocks (residual networks), and as a result, propose a dynamic residual attention network (DRAN). The proposed method allows the network to dynamically select residual paths depending on the input image, based on the idea of attention mechanism. For this, we design a dynamic residual module that determines the residual paths between the basic building blocks for the given input image. By finding optimal residual paths between the blocks, the network can selectively bypass informative features needed to reconstruct the target high-resolution (HR) image. Experimental results show that our proposed DRAN outperforms most of the existing state-of-the-arts lightweight models in SISR.
Karam Park, Jae Woong Soh, Nam Ik Cho
ICPR3
2020 Deep Universal Blind Image Denoising
abstract
Image denoising is an essential part of many image processing and computer vision tasks due to inevitable noise corruption during image acquisition. Traditionally, many researchers have investigated image priors for the denoising, within the Bayesian perspective based on image properties and statistics. Recently, deep convolutional neural networks (CNNs) have shown great success in image denoising by incorporating large-scale synthetic datasets. However, they both have pros and cons. While the deep CNNs are powerful for removing the noise with known statistics, they tend to lack flexibility and practicality for the blind and real-world noise. Moreover, they cannot easily employ explicit priors. On the other hand, traditional non-learning methods can involve explicit image priors, but they require considerable computation time and cannot exploit large-scale external datasets. In this paper, we present a CNN-based method that leverages the advantages of both methods based on the Bayesian perspective. Concretely, we divide the blind image denoising problem into sub-problems and conquer each inference problem separately. As the CNN is a powerful tool for inference, our method is rooted in CNNs and propose a novel design of network for efficient inference. With our proposed method, we can successfully remove blind and real-world noise, with a moderate number of parameters of universal CNN.
Jae Woong Soh, Nam Ik Cho
ICPR2
2020 Handwritten Text Segmentation via End-to-End Learning of Convolutional Neural Networks
Junho Jo, Hyung Il Koo, Jae Woong Soh, Nam Ik Cho
Multim. Tools Appl.4
2020 Dual Path Denoising Network for Real Photographic Noise
abstract
This letter presents a convolutional neural network (CNN) for image denoising, especially for the reduction of real noises. As a network topology, we adopt the dual path network (DPN) that combines the advantages of residual and densely connected networks. Using the DPN as a basic building block, we design a network that connects the DPN in dual path again with an attention mechanism. For efficient denoising of real noise images, we build a training set where noisy images are obtained from a heteroscedastic Gaussian noise model and in-camera pipeline. In addition, we augment the synthetic training set with a relatively small number of real noise data. In the experiments, the proposed method is shown to provide state-of-the-art performance in reducing both synthetic and real noises.
Yeong Il Jang, Yoonsik Kim, Nam Ik Cho
IEEE Signal Process. Lett.3
2020 A Pseudo-Blind Convolutional Neural Network for the Reduction of Compression Artifacts
abstract
This paper presents methods based on convolutional neural networks (CNNs) for removing compression artifacts. We modify the Inception module for the image restoration problem and use it as a building block for constructing blind and non-blind artifact removal networks. It is known that a CNN trained in a non-blind scenario (known compression quality factor) performs better than the one trained in a blind scenario (unknown factor), and our network is not an exception. However, the blind system is more practical because the compression quality factor is not always available or does not reflect the actual quality when the image is a transcoded or requantized image. Hence, in this paper, we also propose a pseudo-blind system that estimates the quality factor for a given compressed image and then applies a network that is trained with a similar quality factor. For this purpose, we propose a CNN that estimates the compression quality factor and prepare several non-blind artifact removal networks that are trained for some specific compression quality factors. We train the networks and conduct experiments on widely used compression standards, such as JPEG, MPEG-2, H.264, and HEVC. In addition, we conduct experiments for dynamically changing and transcoded videos to demonstrate the effectiveness of the quality estimation method. The experimental results show that the proposed pseudo-blind network performs better than the blind one for the various cases stated above and requires fewer computations.
Yoonsik Kim, Jae Woong Soh, Jaewoo Park 0005, Byeongyong Ahn, Hyun-Seung Lee 0001, Young-Su Moon, Nam Ik Cho
IEEE Trans. Circuits Syst. Video Technol.7
2019 Deep Face Image Retrieval for Cancelable Biometric Authentication
abstract
This paper presents a cancelable biometric system for face authentication by exploiting the convolutional neural network (CNN)-based face image retrieval system. For the cancelable biometrics we must build a template that achieves good performance while maintaining some essential conditions. First the same template should not be used in different applications. Second if the compromise event occurs original biometric data should not be retrieved from the template. Last the template should be easily discarded and recreated. Hence we propose a Deep Table-based Hashing (DTH) framework that encodes CNN-based features into a binary code by utilizing the index of the hashing table. We employ noise embedding and intra-normalization that distorts biometric data which enhances the non-invertibility. For training we propose a new segment-clustering loss and pairwise Hamming loss with two classification losses. The final authentication results are obtained by voting on the outcome of the retrieval system. Experiments conducted on two large scale face image datasets demonstrate that the proposed method works as a proper cancelable biometric system.
Young Kyun Jang, Nam Ik Cho
AVSS2
2019 Natural and Realistic Single Image Super-Resolution With Explicit Natural Manifold Discrimination
abstract
Recently, many convolutional neural networks for single image super-resolution (SISR) have been proposed, which focus on reconstructing the high-resolution images in terms of objective distortion measures. However, the networks trained with objective loss functions generally fail to reconstruct the realistic fine textures and details that are essential for better perceptual quality. Recovering the realistic details remains a challenging problem, and only a few works have been proposed which aim at increasing the perceptual quality by generating enhanced textures. However, the generated fake details often make undesirable artifacts and the overall image looks somewhat unnatural. Therefore, in this paper, we present a new approach to reconstructing realistic super-resolved images with high perceptual quality, while maintaining the naturalness of the result. In particular, we focus on the domain prior properties of SISR problem. Specifically, we define the naturalness prior in the low-level domain and constrain the output image in the natural manifold, which eventually generates more natural and realistic images. Our results show better naturalness compared to the recent super-resolution algorithms including perception-oriented ones.
Jae Woong Soh, Gu Yong Park, Junho Jo, Nam Ik Cho
CVPR4
2019 GF-CapsNet: Using Gabor Jet and Capsule Networks for Facial Age, Gender, and Expression Recognition
abstract
The convolutional neural network (CNN) works very well in many computer vision tasks including the face-related problems. However, in the case of age estimation and facial expression recognition (FER), the accuracy provided by the CNN is still not good enough to be used for the real-world problems. It seems that the CNN does not well find the subtle differences in thickness and amount of wrinkles on the face, which are the essential features for the age estimation and FER. Also, the face images in the real world have many variations due to the face rotation and illumination, where the CNN is not robust in finding the rotated objects when not every possible variation is in the training data. To alleviate these problems, we first propose to use the Gabor filter responses of faces as the input to the CNN, along with the original face image. This method enhances the wrinkles on the face so that the face-related features are found in the earlier stage of convolutional layers, and hence the overall performance is increased. We also adopt the idea of capsule network, which is shown to be robust to the rotation of objects and be able to capture the relationship of facial landmarks. We show that the performance of age estimation and FER are improved by using the capsule network than using the plain CNNs. Moreover, by using the Gabor responses as the input to the capsule network, the overall performances of face-related problems are increased compared to the recent CNN-based methods.
Sepidehsadat Hosseini, Nam Ik Cho
FG2
2019 Adaptively Tuning a Convolutional Neural Network by Gating Process for Image Denoising
abstract
This paper presents a new framework that controls feature maps of a convolutional neural network (CNN) according to the noise level such that the network can have different properties to different levels. Unlike the conventional non-blind approach which reloads all the parameters of CNN or switches to other CNNs for different noise levels, we adjust the CNN activation feature maps without changing the parameters at the test phase. For this, we additionally construct a noise level indicator network, which gives appropriate weighting values to the feature maps for the given situation. The noise level indicator network is so simple that it can be implemented as a low-dimensional look-up table at the test phase and thus does not increase the overall complexity. From the experiments on noise reduction, we can observe that the proposed method achieves better performance compared to the baseline network.
Yoonsik Kim, Jae Woong Soh, Nam Ik Cho
ICIP3
2019 Age Estimation Using Trainable Gabor Wavelet Layers In A Convolutional Neural Network
abstract
In this paper, we propose a trainable Gabor wavelet (TGW) layer and cascade it with a convolutional neural network (CNN) for the age estimation. Unlike an existing method that uses fixed (hand-tuned) Gabor filters at the head of a CNN, we use Gabor wavelets that can be adapted for the given input as well as for the targeting task. This is enabled by (a) estimating hyperparameters of Gabor wavelets from the input and (b) using a 1 × 1 convolution layer for the selection of orientation parameter. The proposed TGW layers are trained with the standard gradient-descent method and can be easily incorporated with conventional CNNs in an end-to-end training manner. We conduct experiments on the Adience dataset and show that the proposed network outperforms the baseline CNN without TGW layers and efficiently used trainable parameters than ordinary CNN based methods.
Hyuk Jin Kwon, Hyung Il Koo, Jae Woong Soh, Nam Ik Cho
ICIP4
2019 Deep Hierarchical Single Image Super-Resolution by Exploiting Controlled Diverse Context Features
abstract
This paper presents a hierarchical convolutional neural network (CNN) for single image super-resolution (SISR), which exploits the controlled multi-context features. We focus on the method to extract more enriched features than the case of using fixed size kernels. For this, we attempt to bring out the best of the given parameter capacity through the design of some sophisticated networks in a hierarchical manner. First, we exploit the multi-kernel dilated convolution for extracting multi-size contexts from the image and combine them with the proposed trainable parameters. The multi-kernel network with some new pre-and post-processing blocks forms our basic building block. Then the basic building blocks are densely connected with a new feature fusion schemes, which makes the upper level building block. Then, by connecting the upper level blocks, we can use various features which can enrich the representation of the images. In the experiments, it is shown that the proposed method achieves significant PSNR gain compared to recent lightweight models with comparable numbers of parameters.
Jae Woong Soh, Gu Yong Park, Nam Ik Cho
ISM3
2019 Pupil Center Detection Based on the UNet for the User Interaction in VR and AR Environments
abstract
Finding the location of a pupil center is important for the human-computer interaction especially for the user interface in AR/VR devices. In this paper, we propose an indirect use of the convolutional neural network (CNN) for the task, which first segments the pupil region by a CNN, and then finds the center of mass of the region. For this, we create a dataset by labeling the pupil area on 111,581 images from 29 IR video sequences. We also label the pupil region of widely used datasets to test and validate our method on a variety of inputs. Experiments show that the proposed method provides better accuracies than the conventional ones, showing robustness to the noise.
Sang Yoon Han, Yoonsik Kim, Sang Hwa Lee, Nam Ik Cho
VR4
2019 Generation of high dynamic range illumination from a single image for the enhancement of undesirably illuminated images
Jae Woong Soh, Nam Ik Cho
Multim. Tools Appl.3
2018 Learning Background Subtraction by Video Synthesis and Multi-scale Recurrent Networks
Sung-Kwon Choo, Wonkyo Seo, Dong-ju Jeong, Nam Ik Cho
ACCV (6)4
2018 Deep Clustering and Block Hashing Network for Face Image Retrieval
Young Kyun Jang, Dong-ju Jeong, Seok Hee Lee, Nam Ik Cho
ACCV (6)4
2018 Classification-Based Supervised Hashing with Complementary Networks for Image Search
Dong-ju Jeong, Sung-Kwon Choo, Wonkyo Seo, Nam Ik Cho
BMVC4
2018 A Multi-Exposure Image Fusion Based on the Adaptive Weights Reflecting the Relative Pixel Intensity and Global Gradient
abstract
This paper presents a new multi-exposure fusion algorithm. The conventional approach is to define a weight map for each of the multi-exposure images, and then obtain the fusion image as their weighted sum. Most of existing methods focused on finding weight functions that assign larger weights to the pixels in better-exposed regions. While the conventional methods apply the same function to each of the multi-exposure images, we propose a function that considers all the multi-exposure images simultaneously to reflect the relative intensity between the images and global gradients. Specifically, we define two kinds of weight functions for this. The first is to measure the importance of a pixel value relative to the overall brightness and neighboring exposure images. The second is to reflect the importance of a pixel value when it is in a range with relatively large global gradient compared to other exposures. The proposed method needs modest computational complexity owing to the simple weight functions, and yet it achieves visually pleasing results and gets high scores according to an image quality measure.
Nam Ik Cho
ICIP3
2018 Multi-scale Recurrent Encoder-Decoder Network for Dense Temporal Classification
abstract
The temporal events in video sequences often have long-term dependencies which are difficult to be handled by a convolutional neural network (CNN). Especially, the dense pixel-wise prediction of video frames is a difficult problem for the CNN because huge memories and a large number of parameters are needed to learn the temporal correlation. To overcome these difficulties, we propose a recurrent encoder-decoder network which compresses the spatiotemporal features at the encoder and restores them to the original sized results at the decoder. We adopt a convolutional long short-term memory (LSTM) into the encoder-decoder architecture, which successfully learns the spatiotemporal relation with relatively a small number of parameters. The proposed network is applied to one of the dense pixel-prediction problems, specifically, the background subtraction in video sequences. The proposed network is trained with limited duration of video frames, and yet it shows good generalization performance for different videos and time duration. Also, by additional video specific learning, it shows the best performance on a benchmark dataset (CDnet 2014).
Sung-Kwon Choo, Wonkyo Seo, Dong-ju Jeong, Nam Ik Cho
ICPR4
2018 Scene text rectification using glyph and character alignment properties
abstract
Scene text images usually suffer from perspective distortions, and hence their rectification has been an essential pre-processing step for many applications. Existing methods for scene text rectification mainly exploited the glyph property, which means that the characters in many languages have horizontal/vertical strokes and also have some symmetries in their shapes. In this paper, we propose to use an additional property that the characters need to be well aligned when rectified. For this, character alignment, as well as glyph properties, are encoded in the proposed cost function, and its minimization generates the transformation parameters. For encoding the alignment constraints, we perform the character segmentation using a projection profile method before optimizing the cost function. Since better segmentation needs better rectification and vice versa, the overall algorithm is designed to perform character segmentation and rectification iteratively. We evaluate our method on real and synthetic scene text images, and the experimental results show that our method achieves higher optical character recognition (OCR) rate than the previous approaches and also yields visually pleasing results.
Tae Ho Kil, Hyung Il Koo, Nam Ik Cho
ICPR3
2018 Cancelable Biometrics Using Noise Embedding
abstract
This paper presents a cancelable biometric (CB) scheme for iris recognition system. The CB approaches are roughly classified into two categories depending on whether the method stresses more on non-invertibility or on discriminability. The former is to use non-invertibly transformed data for the recognition instead of the original, so that the impostors cannot retrieve the original biometric information from the stolen data. The latter is to use a salting method that mixes random signals generated by user-specific keys so that the imposters cannot retrieve the original data without the key. The proposed CB can be considered a combination of these methods, which applies a non-invertible transform to the salted data for binary biocode input. We use the reduced random permutation and binary salting (RRP-BS) method as the biometric salting, and use the Hadamard product for enhancing the non-invertibility of salted data. Moreover, we generate several templates for an input, and define non-coherent and coherent matching regions among these templates. We show that salting the non-coherent matching regions is less influential on the overall performance. Specifically, embedding the noise in this region does not affect the performance, while making the data difficult to be inverted to the original.
Dae-Hyun Lee, Sang Hwa Lee, Nam Ik Cho
ICPR3
2018 Co-Salient Object Detection Based on Deep Saliency Networks and Seed Propagation Over an Integrated Graph
abstract
This paper presents a co-salient object detection method to find common salient regions in a set of images. We utilize deep saliency networks to transfer co-saliency prior knowledge and better capture high-level semantic information. The resulting initial co-saliency maps are enhanced by seed propagation steps over an integrated graph. The deep saliency networks are trained in a supervised manner to avoid weakly supervised online learning and exploit them not only to extract high-level features but also to produce both intra- and inter-image saliency maps. Through a refinement step, the initial co-saliency maps can uniformly highlight co-salient regions and locate accurate object boundaries. To handle input image groups inconsistent in size, we propose to pool multi-regional descriptors including both within-segment and within-group information. In addition, the integrated multilayer graph is constructed to find the regions that the previous steps may not detect by seed propagation with low-level descriptors. In this paper, we utilize the useful complementary components of high- and low-level information and several learning-based steps. Our experiments have demonstrated that the proposed approach outperforms comparable co-saliency detection methods on widely used public databases and can also be directly applied to co-segmentation tasks.
Dong-ju Jeong, Insung Hwang, Nam Ik Cho
IEEE Trans. Image Process.3
2018 Aging Management Using a Reconfigurable Switch Network for Arrays of Nonideal Power Cells
Donghwa Shin, Nam Ik Cho, Byunghee Kang, Naehyuck Chang
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Skin detection based on multi-seed propagation in a multi-layer graph for regional and color consistency
abstract
We propose a new skin detection method based on multi-seeds propagation in a multi-layer graph representation of an image. Initially, some of nodes in the graph are set to be foreground or background seeds based on a simple Bayesian skin detector, and they are propagated through the graph to find the skin probability in the manner of semi-supervised learning. The graph is designed to consider not only local and global coherence but also to consider the color consistency by constructing a multilayer graph of image and cluster layers. Extensive experiments on several datasets are conducted, which demonstrate that our method outperforms the existing methods in terms of various quantitative measures, such as accuracy, precision, recall and F-measure.
Insung Hwang, Yoonsik Kim, Nam Ik Cho
ICASSP3
2017 Regional deep feature aggregation for image retrieval
abstract
This paper presents a method to aggregate deep features for an object-based image retrieval system. Several recent works have demonstrated that it is quite important to selectively aggregate features with a weighting scheme and extract features from the limited regions likely to contain specific objects. Hence, the proposed method is to find possible candidate regions in an image, extract region descriptors from each region, and match images in a region-by-region manner. To adhere to using a pre-trained network without retraining or spatial verification, several candidate regions are found in an image and a more sophisticated pooling scheme is used for better performance. Specifically, salient points with active responses are detected in the image and clustered to form the candidate regions. In each region, we aggregate activations of a convolutional layer with the emphasis on more active spatial positions, and generate region descriptors effective for the object-based image retrieval. Our experiments show that the proposed method performs well on several public datasets, especially for the images showing the varied shapes or positions of an object.
Dong-ju Jeong, Sung-Kwon Choo, Wonkyo Seo, Nam Ik Cho
ICASSP4
2017 Robust Document Image Dewarping Method Using Text-Lines and Line Segments
abstract
Conventional text-line based document dewarping methods have problems when handling complex layout and/or very few text-lines. When there are few aligned text-lines in the image, this usually means that photos, graphics and/or tables take large portion of the input instead. Hence, for the robust document dewarping, we propose to use line segments in the image in addition to the aligned text-lines. Based on the assumption and observation that many of the line segments in the image are horizontally or vertically aligned in the well-rectified images, we encode this property into the cost function in addition to the text-line alignment cost. By minimizing the function, we can obtain transformation parameters for camera pose, page curve, etc., which are used for document rectification. Considering that there are many outliers in line segment directions and missed text-lines in some cases, the overall algorithm is designed in an iterative manner. At each step, we remove text components and line segments that are not well aligned, and then minimize the cost function with the updated information. Experimental results show that the proposed method is robust to the variety of page layouts.
Tae Ho Kil, Wonkyo Seo, Hyung Il Koo, Nam Ik Cho
ICDAR4
2017 Co-saliency detection via seed propagation over the integrated graph with a cluster layer
abstract
This paper presents a method to detect common salient regions in a set of images. Since saliency and co-saliency detection are usually used as a pre-processing step for image processing and vision tasks, it is important to consider the complexity as well as the performance of an algorithm. Thus, we adhere to using low-level features and propose to detect co-salient objects with a new multilayer graph model in a bottom-up manner. Input images are represented by intra- and inter-image graphs composed of superpixels and their clusters, and each of these nodes obtains its initial co-saliency value from several cues. To generate resultant co-saliency values, foreground and background seeds are defined at parts of the unified multilayer graph, over which the seeds are propagated. Our experiments show that the proposed algorithm outperforms comparable methods on widely used public datasets, especially for the images that have various features.
Insung Hwang, Dong-ju Jeong, Nam Ik Cho
ICIP4
2017 Convolutional neural networks and training strategies for skin detection
abstract
This paper presents two convolutional neural networks (CNN) and their training strategies for skin detection. The first CNN, consisting of 20 convolution layers with 3 × 3 filters, is a kind of VGG network. The second is composed of 20 networkin-network (NiN) layers which can be considered a modification of Inception structure. When training these networks for human skin detection, we consider patch-based and whole-image-based training. The first method focuses on local features such as skin color and texture, and the second on the human-related shape features as well as color and texture. Experiments show that the proposed CNNs yield better performance than the conventional methods and also than the existing deep-learning based method. Also, it is found that the NiN structure generally shows higher accuracy than the VGG-based structure. The experiments also show that the whole-image-based training that learns the shape features yields better accuracy than the patch-based learning that focuses on local color and texture only.
Yoonsik Kim, Insung Hwang, Nam Ik Cho
ICIP3
2017 Gaze estimation using 3-D eyeball model under HMD circumstance
abstract
This paper presents a gaze estimation algorithm using 3-D eyeball model and 2-D pupil center - inner eye corner(PC-IEC) vector. The conventional methods using feature points in the eye images need lots of calibration markers and long calibration time. However, since the pupil and gaze movements are closely related to the 3-D rotation of eyeball, the long and complicated calibrations are not necessary. This paper derives the relationship between the 3-D eyeball model and 2-D PC- IEC vector with a single reference calibration point which is located at the center of the screen. Also, the proposed algorithm compensates for the eyeball movements using the eyelid height against the inner eye corner. According to the experiment, the proposed method estimates the gaze within 2 degree error.
Sang Yoon Han, Sang Hwa Lee, Nam Ik Cho
MMSP3
2017 Saliency detection based on seed propagation in a multilayer graph
Insung Hwang, Sang Hwa Lee, Nam Ik Cho
Multim. Tools Appl.4
2017 Rectification of planar targets using line segments
Jaehyun An, Hyung Il Koo, Nam Ik Cho
Mach. Vis. Appl.3
2017 Open-Contour Tracking Using a New State-Space Model and Nonrigid Motion Training
abstract
Object tracking in a video sequence is usually achieved by tracking the bounding box over the object or the object’s boundary, each of which has somewhat different applications. In this paper, we present a new open-contour tracking algorithm based on a Bayesian framework in which the contour is a part of the object’s boundary. We first propose a new state-space model for the representation of contours, which can handle the rigid and nonrigid motions of contours independently. This model enables us to focus on the nonrigid motions during the training, and the model works for challenging rigid motion scenarios. In addition, for the robust tracking of contours, we propose a measurement function that considers the contrast on object boundaries, target appearance, and temporal coherence. We applied the proposed method to two kinds of open-contours targets, and the experimental results show that the proposed method achieves superior performance to the conventional contour tracking methods. The proposed method is also compared with recent bounding box tracking methods for the object tracking purposes, and the comparison shows that the proposed method works robustly to fast motions and yields a more accurate estimate of an object’s location than the conventional bounding box tracking methods.
Seon Heo, Hyung Il Koo, Nam Ik Cho
IEEE Trans. Circuits Syst. Video Technol.3
2016 Image co-saliency detection based on clustering and diffusion process
abstract
This paper presents a co-saliency detection algorithm based on clustering and diffusion process. For each image in a set, intra saliency maps are constructed from the measure of boundary and contrast priors. Then, segmented regions of all images are clustered based on features of colors, intra saliency and coherence of saliency. Co-saliency of each cluster is computed from combination of foreground probability and coherence of the cluster. The co-saliency of cluster is propagated over the segmented regions according to affinity between the cluster and segments. In addition, we adopt an intra image diffusion process from a graph with learned fully affinity in order to improve spatial consistency of co-saliency maps. Experimental results show that our algorithm yields better results compared to the state-of-the-art methods in terms of precision-recall curve, visual plausibility and computational cost.
Insung Hwang, Gu Yong Park, Nam Ik Cho
VCIP4
2016 Fast and simple text replacement algorithm for text-based augmented reality
abstract
In this paper, we present a novel text-based augmented reality system that performs optical character recognition on natural images and replaces the recognized texts with other informative contents. For the goal, we implement text detection and recognition functions, and develop an image augmentation algorithm for the realistic contents replacement. To be precise, we reconstruct background with a linear interpolation method and insert new contents to the reconstructed backgrounds. Finally, we get natural results by applying proper geometric distortions to them. In order to reduce visual artifacts caused by noisy boundaries, we also develop an optimal path selection method based on dynamic programming. Experimental results show that our method provides very natural results and runs in real-time even in mobile devices.
Hyung Il Koo, Beom Su Kim, Young Ki Baik, Nam Ik Cho
VCIP4
2015 LPI adaptive descreening method with Hadamard domain analysis
abstract
This paper presents a new line-per-inch (LPI) adaptive descreening algorithm for restoring a high-quality image when scanning a printed matter. The LPI of scanned image is estimated and image region is classified into flat or edge region by analyzing the distribution in Hadamard transform coefficients. Then, an LPI adaptive edge preserving filtering is applied to reduce the halftone patterns not only in flat region but also in edge region. Proposed method has hardware-friendly structure without iteration and cascades of filters, and also it needs only integer adder and shift operations. Experimental result shows that the proposed method performs adaptively for a variety of LPI contents.
Hyun-Seung Lee 0001, Ji Young Yi, Nam Ik Cho
ICIP3
2015 Junction-based table detection in camera-captured document images
Wonkyo Seo, Hyung Il Koo, Nam Ik Cho
Int. J. Document Anal. Recognit.3
2015 Document dewarping via text-line based optimization
Beom Su Kim, Hyung Il Koo, Nam Ik Cho
Pattern Recognit.3
2015 Efficient Unwrap Representation of Faces for Video Editing
abstract
Unwrap mosaic is a method for decomposing a video into a 2D texture and a dense mapping that enable the reconstruction of the video from the texture. This representation is useful in some frameworks because we can edit videos by simply retouching 2D textures. However, the complexity of conventional approaches is too high to be adopted in time-critical applications (it takes up to several hours). In this letter, we focus on face-related applications such as face editing and replacement, and propose a face unwrap approach for these applications. To be precise, we adopt the view-based active appearance model (AAM) trackers and estimate dense mappings from the AAM results. The AAM also provides pose information which is also exploited in building the texture map. Experimental results show that our method is very efficient compared with the conventional unwrap mosaic approach. Moreover, based on the proposed system, we develop face-related applications.
Byeongyong Ahn, Hyung Il Koo, Hong Il Kim, Jichull Jeong, Nam Ik Cho
IEEE Signal Process. Lett.5
2015 Word Segmentation Method for Handwritten Documents based on Structured Learning
abstract
Segmentation of handwritten document images into text-lines and words is an essential task for optical character recognition. However, since the features of handwritten document are irregular and diverse depending on the person, it is considered a challenging problem. In order to address the problem, we formulate the word segmentation problem as a binary quadratic assignment problem that considers pairwise correlations between the gaps as well as the likelihoods of individual gaps. Even though many parameters are involved in our formulation, we estimate all parameters based on the Structured SVM framework so that the proposed method works well regardless of writing styles and written languages without user-defined parameters. Experimental results on ICDAR 2009/2013 handwriting segmentation databases show that proposed method achieves the state-of-the-art performance on Latin-based and Indian languages.
Jewoong Ryu, Hyung Il Koo, Nam Ik Cho
IEEE Signal Process. Lett.3
2014 Content-based image retrieval using color features of salient regions
abstract
This paper presents a content-based color image retrieval system based on color features from the salient regions and their spatial relationship. The proposed method first extracts the salient regions by a color contrast method, and finds several dominant colors for each region. Then, the spatial distribution of each dominant color is described as a binary map. Specifically, the salient region is partitioned into small sub-blocks, and each sub-block is assigned as 1 or 0 according to the number of pixels corresponding to the dominant color. The set of binary maps define the spatial distribution of dominant colors within and across the salient regions, which approximately reflect the objects' shapes and the spatial relationship of the objects. A simple matching method for this description is also proposed, which needs very few computations for each image matching. According to the experiments with several widely used color image databases, the proposed method shows better retrieval performance than the state-of-the-art and previous color-based methods. The proposed algorithm is suitable for color image retrieval on the web and mobile systems, because it needs very few computations which are mostly binary logical operations.
Jaehyun An, Sang Hwa Lee, Nam Ik Cho
ICIP3
2014 Language-Independent Text-Line Extraction Algorithm for Handwritten Documents
abstract
Text-line extraction in handwritten documents is an important step for document image understanding, and a number of algorithms have been proposed to address this problem. However, most of them exploit features of specific languages and work only for a given language. In order to overcome this limitation, we develop a language-independent text-line extraction algorithm. Our method is based on connected components (CCs), however, unlike conventional methods, we analyze strokes and partition under-segmented CCs into normalized ones. Due to this normalization, the proposed method is able to estimate the states of CCs for a range of different languages and writing styles. From the estimated states, we build a cost function whose minimization yields text-lines. Experimental results show that the proposed method yields the state-of-the-art performance on Latin-based and Chinese script databases. Further, we submitted the proposed algorithm to the ICDAR 2013 handwriting segmentation competition and our method showed the best text-line extraction performance among 10 participant methods.
Jewoong Ryu, Hyung Il Koo, Nam Ik Cho
IEEE Signal Process. Lett.3
2014 Lossless Compression of Color Filter Array Images by Hierarchical Prediction and Context Modeling
abstract
This paper presents an encoder for the lossless compression of color filter array (CFA) data, which consists of a hierarchical predictor and context-adaptive arithmetic encoder. In hierarchical prediction, the subsampled images are encoded in order; each of the subimages contains only one color component (red, green, or blue) in the case of a Bayer CFA image. By subsampling, the green pixels are separated into two sets, one of which is encoded by a conventional grayscale encoder, and then is used to predict the green pixels in the other set. Both the sets of greens are then used to predict the reds, and the green and red pixels are used to predict the blues. Throughout this process, the predictors are designed considering the direction of the edges in the neighborhood. By gathering some information from the prediction process, such as edge activity and neighboring errors, the magnitude of prediction error is also estimated. From this, the probability distribution function of prediction error conditioned on neighboring pixels, i.e., the context is estimated, and context-adaptive arithmetic encoding is applied to reduce the resulting bits further. The experimental results on real and simulated CFA images show that the proposed method produces less bits per pixel than the conventional lossless image compression methods and recently developed lossless CFA compression algorithms.
Seyun Kim, Nam Ik Cho
IEEE Trans. Circuits Syst. Video Technol.2
2014 Hierarchical Prediction and Context Adaptive Coding for Lossless Color Image Compression
abstract
This paper presents a new lossless color image compression algorithm, based on the hierarchical prediction and context-adaptive arithmetic coding. For the lossless compression of an RGB image, it is first decorrelated by a reversible color transform and then Y component is encoded by a conventional lossless grayscale image compression method. For encoding the chrominance images, we develop a hierarchical scheme that enables the use of upper, left, and lower pixels for the pixel prediction, whereas the conventional raster scan prediction methods use upper and left pixels. An appropriate context model for the prediction error is also defined and the arithmetic coding is applied to the error signal corresponding to each context. For several sets of images, it is shown that the proposed method further reduces the bit rates compared with JPEG2000 and JPEG-XR.
Seyun Kim, Nam Ik Cho
IEEE Trans. Image Process.2
2013 A dehazing algorithm using dark channel prior and contrast enhancement
abstract
This paper proposes a dehazing algorithm based on dark channel prior and contrast enhancement approaches. The conventional dark channel prior method removes haze and thus restores colors of objects in the scene, but it does not consider the enhancement of image contrast. On the contrary, the image contrast method improves the local contrast of objects, but the colors are often distorted due to the over-stretching of contrast. The proposed algorithm combines the advantages of these two conventional approaches for keeping the color while dehazing. For this, an optimization function is proposed to balance between the contrast and colors distortion, where the contrast measure follows the conventional image statistics and the hue component is used to constrain the color changes. According to the experimental results, the proposed approach compensates for the disadvantages of conventional methods, and enhances contrast with less color distortion.
Tae Ho Kil, Sang Hwa Lee, Nam Ik Cho
ICASSP3
2013 Luminance adapted skin color modeling for the robust detection of skin areas
abstract
Statistical models of color channels have been used for the detection of skin areas. However, since the distribution of colors also change as the luminance varies, color distribution models without considering the luminance variation do not work well for the images taken under various illumination conditions. Hence we propose a new skin detection algorithm that considers the luminance value in modeling the color distribution. For implementing this idea, we need a sample of skin color in the image, which can be obtained from the face. It is noted that the faces can be detected without color feature, but by using only structural features such as eyes, mouth, etc. In our algorithm, the eye detector and elliptical boundary are used to detect the face. From the face region, joint Gaussian distributions of color components with respect to the luminance values are obtained. Then, each pixel in the image is classified into skin or non-skin using the statistics obtained from the face region. Experimental results show that the proposed skin color model outperforms the conventional methods in terms of the detection accuracy and F-score.
Insung Hwang, Sang Hwa Lee, Byungseok Min, Nam Ik Cho
ICIP4
2013 Single image dehazing based on reliability map of dark channel prior
abstract
The dark channel prior is generally a powerful prior for single image dehazing, but it is invalid when there are objects which have similar colors to the atmospheric light. It is difficult to find the accurate transmission factor in this case, and thus the results show some color distortion. This paper proposes a new algorithm for the estimation of correct transmission map regardless of object color. The proposed method defines a reliability map that depicts how much the objects or areas meet the dark channel prior assumption, and then estimate the transmission map using the reliable pixels only. The transmission factors at the unreliable pixels are estimated by linear fitting using the reliable neighboring pixels. Experiments show that the dehazed images by the proposed method have less color distortion than the results by the conventional algorithm.
Tae Ho Kil, Sang Hwa Lee, Nam Ik Cho
ICIP3
2013 A mobile spherical mosaic system
abstract
This paper proposes a new mobile mosaic system using the spherical surface for the panoramic image synthesis. The proposed system consists of image acquisition interface, exposure compensation, local image alignment, spherical projection, and blending. The image acquisition interface guides the users to capture images by displaying a wire frame on the spherical surface. The camera pose and direction are estimated by Gyro sensor and accelerometer. We correct the different brightness levels of images using local means and variances in the overlapped regions. Then, the images are locally realigned since the sensor information has much noise. We implement the successive template matching on the spherical surface, which reduces the sensor errors and misalignment in the 3-D rotational motion. After getting the 3-D rotation information of images, the images are rotated in the virtual 3-D space and projected on the spherical surface. We derive the spherical projection using the radius of virtual mosaic sphere and focal length. Finally, the overlapped multiple images are blended only at the boundaries on the spherical surface. We have implemented the system in the usual tablet PC and mobile phone. According to the various experiments, the proposed spherical mosaic system composes the full environment around the user in real-time.
Sang Hwa Lee, Ji In Jeon, Sung-Kwon Choo, Nam Ik Cho, Jong-Il Park
ICIP4
2012 Reduction of ghost effect in exposure fusion by detecting the ghost pixels in saturated and non-saturated regions
abstract
This paper proposes a multiple exposure image fusion algorithm with reduced ghost. The basic idea is to adjust the weight map in the conventional fusion method in such a way that the ghost pixels are excluded. For this, pixels that cause ghost effect are detected in both of the saturated and non-saturated regions. In order to detect ghost pixels in the non-saturated region, we use the photometric relation and Gaussian mixture modeling (GMM) of a zero mean normalized cross correlation (ZNCC) map between a given exposure image and the reference. From this, we can also obtain static region where we construct an intensity mapping function (IMF) to detect the ghost pixels in saturated regions. Experimental results show that the proposed method generates high quality image without noticeable ghost effect, and yields less artifacts than the conventional methods.
Jaehyun An, Seong Jong Ha, Nam Ik Cho
ICASSP3
2012 Fast text line extraction in document images
abstract
This paper proposes an algorithm for fast text line extraction in document image. Instead of binarization or multi-oriented Gaussian blurring of an image as in the conventional methods, we use integral image and design filters that are proper to detect text regions on the integral image. After the filtering, the center points in the regions are discovered by cascade text region verification followed by non-maximum suppression. Finally, text lines are extracted by grouping the points on the same line. The proposed method is tested with document images taken in various environments, and it is shown to be faster than the conventional ones while its performance is comparable.
Seong Jong Ha, Bora Jin, Nam Ik Cho
ICIP3
2012 A gradient guided deinterlacing algorithm
abstract
This paper proposes new intra and inter-deinterlacing algorithm based on the gradient domain image/video processing approach. From the interlaced (field) images, gradient field images are generated and then the gradients of missing lines are estimated to generate gradient images which correspond to progressive frames. The proposed intra-deinterlacing is basically an edge-oriented interpolation, which interpolates the gradients of missing pixels along the optimal spatial orientation. Finding the optimal orientation among all possible ones is formulated as a labeling problem with Markov random field (MRF) framework. For obtaining better results for fast moving video sequences, this method is extended to inter-deinterlacing, which considers the temporal orientations as well as the spatial ones. With the synthesized gradient frame images and the original pixel values of the field images, we then formulate a linear equation that generates the final progressive frame images. Like other gradient domain image processing applications, the integrity of edges is the main advantage of the proposed method.
Bora Jin, Jung Gap Kuk, Nam Ik Cho
ICIP3
2012 Image registration by using a descriptor for repetitive patterns
abstract
This paper proposes a new feature-based image registration method based on the description of feature clusters. This method can find larger number of correspondences than the conventional methods using singleton feature descriptors, which often fail in repetitive patterns. The reason for the failure of conventional methods in a repeating pattern is due to the existence of too many similar features, which in turn gives geometrically inconsistent matching or do not survive ratio test. Hence the proposed method follows the strategy that first separate the similar features from the repetitive patterns from the others. Then the similar features in a pattern are grouped into a set that is described by a support vector descriptor in terms of the cluster's center and radius. Once the same pattern in different images are matched, the geometric cue is added to find many geometrically consistent correspondences of the features. In the experiments, it has been demonstrated that the larger number of geometrically consistent correspondences from the repetitive pattern give more accurate registration, and thus more pleasing results in image stitching and panoramic image generation.
Seong Jong Ha, Seyun Kim, Nam Ik Cho
VCIP3
2012 A lossless color image compression method based on a new reversible color transform
abstract
In many conventional lossless color image compression methods, the pixels or lines from each color component are interleaved, and then they are predicted and coded. Also, it has been reported that the reversible color transform (RCT) followed by a grayscale encoder gives higher coding gain than the independent compression of each channel does. In this paper, we propose a lossless color image compression method that concentrates on the efficient coding of chrominance channels with a new color transform and hierarchical coding of chrominance channel pixels. Specifically, we first transform an input image with R, G, and B color space into Y CuCvcolor space using the proposed RCT, which shows better decorrelation performance than the existing RCT. After the color transformation, the luminance channel Y is compressed by a conventional lossless image coder, such as JPEG-LS, CALIC, or JPEG2000 lossless. Unlike the luminance channel, the chrominance channels Cuand Cvare relatively smooth and have different statistical characteristic. Therefore, the chrominance channels are differently encoded based on a hierarchical decomposition and directional prediction. Finally, effective context modeling for prediction residuals is adopted. Experimental results show that the proposed method improves the compression performance by 40% over the conventional channel independent compression methods and 5% over the existing methods that exploit the channel correlation.
Seyun Kim, Nam Ik Cho
VCIP2
2012 Discrimination and description of repetitive patterns for enhancing the performance of feature-based recognition
Seong Jong Ha, Sang Hwa Lee, Nam Ik Cho
Image Vis. Comput.3
2012 Text-Line Extraction in Handwritten Chinese Documents Based on an Energy Minimization Framework
abstract
Text-line extraction in unconstrained handwritten documents remains a challenging problem due to nonuniform character scale, spatially varying text orientation, and the interference between text lines. In order to address these problems, we propose a new cost function that considers the interactions between text lines and the curvilinearity of each text line. Precisely, we achieve this goal by introducing normalized measures for them, which are based on an estimated line spacing. We also present an optimization method that exploits the properties of our cost function. Experimental results on a database consisting of 853 handwritten Chinese document images have shown that our method achieves a detection rate of 99.52% and an error rate of 0.32%, which outperforms conventional methods.
Hyung Il Koo, Nam Ik Cho
IEEE Trans. Image Process.2
2011 A multi-exposure image fusion algorithm without ghost effect
abstract
This paper proposes a new multi-exposure fusion algorithm for HDR imaging. The advantage of the fusion based method is that it does not actually generate HDR images so that the compression to LDR is also not needed. However, the conventional exposure fusion suffers from ghost effects when there are moving object or hand trembling because the out put is just a weighted sum of multiple exposure input images. In order to alleviate this problem, we propose two kinds of weights in addition to the conventional weights that reflect the contrast, saturation and well-exposedness. The first is the weight that removes the pixels that can cause ghost effects, based on the observation that pixels near the boundary of moving objects often violate the rule that more highly exposed pixel have higher intensity level. The second is to reduce the weight when the correlation between an exposed image and the reference image is low. Experimental results show that the proposed method removes the ghost effectively, and yield less artifacts than the conventional methods.
Jaehyun An, Sang Heon Lee 0004, Jung Gap Kuk, Nam Ik Cho
ICASSP4
2011 High dynamic range (HDR) imaging by gradient domain fusion
abstract
This paper proposes a new HDR imaging method in the gradient domain based on the fusion of two images with different exposure. We first formulate an energy function for the binary labeling of each pixel based on the brightness and contrast measure of each image, which gives a binary map that determines whether the under-exposed image is better than the overly exposed one or not. Then a target gradient field is generated by combining the gradient fields of two images according to the binary map. Since the gradient field so combined is generally not integrable, we modify it via Poisson solver. This gradient field is considered that of HDR image and directly compressed to the display scale (in the gradient domain). Experiments show that the gradient domain fusion provides better results than the image domain methods, when fusing two differently exposed images.
Jung Gap Kuk, Nam Ik Cho, Sang Uk Lee
ICASSP2
2011 A new image denoising method based on thewavelet domain nonlocal means filtering
abstract
We present a new image denoising method based on the non local means filtering in the wavelet domain. A noisy image is first decomposed into subbands by wavelet transform and the nonlocal means filter is applied to each subband. It is also noted that the performance of the nonlocal means filter depends on the kernel bandwidth (size of the filter) and the image properties. Hence we propose a method to adjust the kernel bandwidth for each of the subband images, based on the estimation of noise statistics. This filtering method preserves the wavelet coefficients corresponding to the structures, while effectively suppressing noisy ones. Experimental results show that the proposed method provides comparable or sometimes higher peak signal-to-noise ratio (PSNR) than the state-of-the-art wavelet denoising methods and the spatial nonlocal means filter. Subjective comparison also shows that the proposed method provides better contrast than the spatial nonlocal means filter, and less ringing artifacts that commonly arise in the conventional wavelet denoising.
Su Jeong You, Nam Ik Cho
ICASSP2
2011 Discrimination and description of repetitive patterns for enhancing object recognition performance
abstract
Objects with repetitive patterns are not well recognized by SIFT/SURF based matching because the features from those patterns are too similar and thus it is often difficult to find the homography of matched pairs. In this paper, we propose a new feature matching strategy to alleviate this problem by differentiating repetitive patterns from the other salient ones and also by developing a way of utilizing the patterns for robust feature matching. Specifically, we develop a classifier that tells whether the features are from repetitive patterns or salient features, based on mean shift clustering followed by support vector data description. Then the homography is found over the salient features by excluding the repetitive features at first, which is then validated and refined by the patterns. The proposed method is tested with the examples of matching the buildings with repeating patterns, and it is shown to be robuster and more reliable than the conventional ones.
Seong Jong Ha, Sang Hwa Lee, Nam Ik Cho
ICIP3
2011 Improved H.264/AVC lossless intra compression using multiple partition prediction for 4×4 intra block
abstract
The sample by sample DPCM (SbS DPCM) is an important prediction technique for the H.264/AVC lossless intra compression. In this paper, we propose a new prediction method that is more efficient than the conventional SbS DPCM, thereby improving the overall compression performance. The proposed method prepares 5 partition patterns for each 4×4 block such as 4×4 (no partition), 4×2, 2×4, 2×2 and 1×1. The pixels in each partition is intra predicted by SbS DPCM and the best partition which produces minimum bit is selected as the partition pattern for the 4×4 block. Also, the number of available intra prediction directions is determined according to the partition pattern to avoid too much side-information transmission. The experimental results show that the proposed method gives 3.62 % point bit rate saving on average and 4.74 % point bit rate saving at maximum compared to the conventional SbS DPCM.
Sang Heon Lee 0004, Jewoong Ryu, Nam Ik Cho
ICIP3
2011 Color filter array demosaicking using optimized edge direction map
abstract
This paper proposes a new color filter array de-mosaicking method with emphasis on the edge estimation. In many existing approaches, the demosaicking is considered a directional interpolation problem, and thus finding the correct edge direction is a very important factor. However, these methods sometimes fail to determine an accurate interpolating direction because they use local information from neighboring pixels. For the estimation of edge direction using global information, we employ an MRF framework where the energy function is formulated by defining new notions of interpolation risk and pixel connectivity. Minimizing this function gives the edge directions, and the green channel is interpolated along the edges. Then we iterate the luminance update and color correction using the high frequencies from green channel. The algorithm is tested with the commonly used images, and it is shown to yield higher CPSNR than the state-of-the-art methods in many images, up to 2.7dB at maximum and 0.4dB on average. Subjective comparison also shows that the proposed method produces less artifacts on complex structures.
Seyun Kim, Nam Ik Cho
MMSP2
2011 Design of Interchannel MRF Model for Probabilistic Multichannel Image Processing
abstract
In this paper, we present a novel framework that exploits an informative reference channel in the processing of another channel. We formulate the problem as a maximum a posteriori estimation problem considering a reference channel and develop a probabilistic model encoding the interchannel correlations based on Markov random fields. Interestingly, the proposed formulation results in an image-specific and region-specific linear filter for each site. The strength of filter response can also be controlled in order to transfer the structural information of a channel to the others. Experimental results on satellite image fusion and chrominance image interpolation with denoising show that our method provides improved subjective and objective performance compared with conventional approaches.
Hyung Il Koo, Nam Ik Cho
IEEE Trans. Image Process.2
2010 A Color to Grayscale Conversion Considering Local and Global Contrast
Jung Gap Kuk, Jae Hyun Ahn, Nam Ik Cho
ACCV (4)3
2010 Rectification of figures and photos in document images using bounding box interface
abstract
This paper proposes an algorithm for the segmentation and rectification of figures and photos in document images. The algorithm requires just a rough user-provided bounding box for the objects in a single-view image. On receiving the user's bounding box, it takes about 1-2 seconds to segment and rectify mega-pixel sized figures. The main feature of the algorithm is a novel segmentation method that exploits the properties of printed figures. Specifically, a set of boundary candidates is generated using the properties, and the optimal boundary in the set is found by using an alternating optimization scheme. This segmentation result is further refined so that it is well localized to the true boundary. In addition to our segmentation method, we also propose a new boundary interpolation method for the rectification of segmented figures. The method improves the quality of output by largely removing perspective distortions compared to conventional boundary interpolation methods. Experimental results on a variety of images show that the method is efficient, robust, and easy to use.
Hyung Il Koo, Nam Ik Cho
CVPR2
2010 State Estimation in a Document Image and Its Application in Text Block Identification and Text Line Extraction
Hyung Il Koo, Nam Ik Cho
ECCV (2)2
2010 H.264/AVC based color filter array compression with inter-channel prediction model
abstract
In the case of images captured by color filter array (CFA), the compression-first approach (compression of CFA) is shown to be better than the demosaick-first approach (general image compression). In this paper, we propose a new H.264/AVC based compression-first approach that exploits the correlation between the color channels. For this, we propose an inter-channel prediction model described by a linear equation in the RGB space. Specfically, the R and B pixels are predicted by a neighboring G pixel or demosaicked G pixel at the same location. The proposed prediction scheme further reduces the variance of prediction error, and as a result the proposed method provides about 1.45dB gain on average and 2.2dB gain at maximum for the green channel over the conventional compression-frist scheme. For the non-green channel, average gain is about 1.0dB and maximum gain is about 1.8dB.
Sang Heon Lee 0004, Nam Ik Cho
ICIP2
2010 A video object segmentation algorithm based on the feature learning and shape tracking
abstract
This paper proposes a video object segmentation algorithm based on the conditional random field (CRF) framework. A foreground object in the first frame is segmented by training the CRF on user interaction, i.e., by using user scribbles corresponding to foreground and background respectively for CRF training. The data term of the energy function in this CRF framework is designed as a function of the score of texture-color classifier trained by AdaBoost. From the second frame, a weighted data term that encodes the shape of the object is added to this energy function. The boundary pixels of the current frame are predicted by the optical flow, and a smaller cost is given to a pixel closer to the boundary and vice versa. Also, a confidence of optical flow is defined, and a larger weight is given to the data term when the confident is high. As a result, the data term related with the shape becomes important when the motion estimation is reliable, and conversely the color-texture term becomes important otherwise. Experimental results show that the proposed data term keeps the boundary correctly in most cases and provides comparable result when compared to a state-of-the-art method.
Sang-Hak Lee, Hyung Il Koo, Nam Ik Cho
ICIP3
2010 DSP integration of sound source localization and multi-channel Wiener filter
abstract
This paper describes a DSP integration of sound source localization (SSL) and multi-channel Wiener filter (MWF). To develop a robot audition system, we integrated SSL module and MWF module into a DSP system. SSL is a module to perceive the direction of a human user's call. It measures time delay of arrival among microphones and estimates the direction of sound source. Also, it post-processes the resulted estimations of direction by histogram to perceive the direction robustly under noisy environment. MWF is a module to reduce background noises from raw voice signal to enhance the performance of robot's speech recognition. It gathers information of background noises during noise-period and then reduces noises during voice-period. This SSL-MWF combination system will be a cheap, high-performing and convenient solution for robot audition.
Byoung-gi Lee, Hyun-dong Kim, Jongsuk Choi, Seyun Kim, Nam Ik Cho
ICRA5
2010 A new image projection method for panoramic image stitching
abstract
We propose a new image projection method in an attempt to reduce the perceptual distortion in panoramic image mosaics. Specifically, we reduce the stretching distortion of some image patches and bending of straight lines. Since the stretching distortion usually occurs when projecting a viewing sphere to the cylindrical image surface in an oblique direction, we propose to use an adjustable cylindrical surface to match the viewing direction with the equator of the cylindrical surface. Also, in order to find the trade-off between the stretching distortion and bending of straight lines, we also adjust the curvature of cylindrical surface according to the object of interest in the image. The warping function from the viewing sphere to the adjustable image surface is derived and the amount of distortion caused by this warping function is also defined. From the measure of distortion, the optimal pose of the cylindrical image plane and its curvature are determined, and the image on the viewing sphere is projected on the optimal plane. The experimental results show that the proposed method produces the panoramic image with less distortion than the existing methods.
Beom Su Kim, Hyung Il Koo, Nam Ik Cho
MMSP3
2010 A motion vector prediction method for multi-view video coding
Sang Heon Lee 0004, Sang Hwa Lee, Jeong Hyu Yang, Nam Ik Cho
J. Vis. Commun. Image Represent.4
2010 Image segmentation algorithms based on the machine learning of features
Sang-Hak Lee, Hyung Il Koo, Nam Ik Cho
Pattern Recognit. Lett.3
2010 Improved linear soft-input soft-output detection via soft feedback successive interference cancellation
abstract
We propose an improved minimum mean square error (MMSE) vertical Bell Labs layered space-time (V-BLAST) detection technique, called a soft input, soft output, and soft feedback (SIOF) V-BLAST detector, for turbo multi-input multioutput (turbo-MIMO) systems. We derive a symbol estimator by minimizing the power of the interference plus noise, given a priori probabilities of undetected layer symbols and a posteriori probabilities for past detected layer symbols. For a low-complexity implementation, an approximate SIOF algorithm is presented, which allows for a time-invariant realization of the symbol ordering and an MMSE filtering process. Another implementation, referred to as the iterative SIOF algorithm is introduced, which decides on symbol detection order based on a posteriori symbol probabilities to improve the detection performance. Simulations performed on a space-time bit-interleaved coded modulation (STBICM) architecture over quasi-static MIMO fading channels demonstrate that the SIOF V-BLAST detector provides performance gains over previous turbo-BLAST detectors, most notably when more transmit antennas are used.
Andrew C. Singer, Jungwoo Lee 0001, Nam Ik Cho
IEEE Trans. Commun.4
2009 A new method to find an optimal warping function in image stitching
abstract
In image stitching applications, it is very important to find a suitable warping function for the visual quality of a composite (stitching result). In this paper, we present a new mathematical criterion to select an optimal warping function among a set of possible candidates (e.g., parametric family). The proposed criterion can be considered as a direct view condition for image stitching , i.e., it is desirable that each part of the composite image looks like its corresponding input image. More specifically, we do not use an explicit modeling of a compositing surface, but, we focus on the differential properties of a warping function. That is we design a cost function so that the Jacobian matrix of a warping function is close to a shape preserving matrix such as rotation and reflection matrices. The proposed cost function can be effectively minimized by using Levenberg-Marquardt algorithm. The experimental results show that the proposed method results in visually pleasing stitched results because the original shape of each image is preserved in the composite.
Hyung Il Koo, Beom Su Kim, Nam Ik Cho
ICASSP3
2009 An overlap save algorithm for block convolution with reduced complexity
abstract
We propose a block convolution algorithm that requires shorter length FFT than the conventional overlap save algorithm (OSA). It is shown that the OSA can be split into two separate processes related to the previous and current data blocks. Hence, only current data block needs to be transformed in the proposed OSA, whereas the concatenated block of previous and current data is transformed in the conventional method. As a result, the number of arithmetic operations for the block convolution is reduced. Also, the reduced transform size gives additional advantage in data manipulation when implemented on DSP and PC.
Jung Gap Kuk, Se Yoon Kim, Nam Ik Cho
ICASSP3
2009 Graph cuts using a Riemannian metric induced by tensor voting
abstract
In this paper, we present a new algorithm that combines the advantages of tensor voting into graph cuts. Tensor voting has been a popular tool for a number of early vision problems since it can use principles of perceptual grouping, which are not well considered in graph cuts. We attempt to encode the power of tensor voting into an energy minimization framework. For this, we assume that the tensor map obtained by tensor voting induces a Riemannian metric in image domain, and the metric is constructed according to the conventional ways of tensor interpretation. Finally, by embedding the induced Riemannian metric into the graph via edge weights, the graph cuts algorithm can have priors considering principles of perceptual grouping. The proposed method can be used in the labeling of occluded regions, object segmentation using only edge information, and boundary regularization.
Hyung Il Koo, Nam Ik Cho
ICCV2
2009 Feature Based Binarization of Document Images Degraded by Uneven Light Condition
abstract
This paper proposes a document image binarization method, which is especially robust to the images degraded by uneven light condition, such as the camera captured document images. A descriptor that captures the regional properties around a given pixel is first defined for this purpose. For each pixel, the descriptor is defined as a vector composed of filter responses with varying length. This descriptor is shown to give highly discriminating pattern with respect to the background region, text region, and near text region. Of course there are misclassified pixels, which are then relabeled using an energy optimization method, specifically by using the graph cut method. For this, we devise an appropriate energy function that leads to clear and correct binarization. The proposed descriptor is also used for the skew detection, and thus correcting the skewed documents.
Jung Gap Kuk, Nam Ik Cho
ICDAR2
2009 Eliminating structure misalignments using robust matching and image editing based on seam carving
abstract
In this paper, we propose an algorithm that generates a natural composite from the misaligned images. The image stitching for this purpose is usually performed in two steps: correspondence matching of salient features followed by appropriated warping or editing. The proposed correspondence matching problem is formulated as the one-dimensional registration along the stitching boundary, where an appropriate energy function is proposed. The designed energy function consists of three complementary terms that encode appearance, smoothness and ordering of points, whereas the existing method considers the edge strength and ordering. Then we develop an algorithm that makes several salient points move to desired positions by using seam carving/inserting, which produces visually pleasing results compared to the conventional warping methods with less computations. The experimental results show that the proposed method efficiently and robustly generates natural composite images.
Hyung Il Koo, Jung Gap Kuk, Nam Ik Cho
ICIP3
2009 Intra prediction method based on the linear relationship between the channels for YUV 4: 2: 0 intra coding
abstract
In general, the transformation from RGB to YUV color space reduces the correlation between the channels. But some sources in the YUV space still have strong inter-channel correlation, which can be modeled as a linear function. Based on this linear model, we propose a new intra chrominance prediction method for color image compression in YUV 4∶2∶0 color space. A block of chrominance pixels to be encoded is predicted from the reconstructed luminance signal using the proposed prediction scheme and the residual is encoded based on the H.264/AVC. Also, an implicit prediction method that can obviate the transmission of side information is proposed. The experimental results show that there are about 0.35∼1.0dB gains for the luminance (Y) channel, and 0.5∼3.0dB gains for the chrominance (Cb or Cr) channels at the medium to high bit-rates. The proposed method also shows some gains at the low bit-rates.
Sang Heon Lee 0004, Nam Ik Cho
ICIP2
2009 An unsupervised image segmentation algorithm based on the machine learning of appropriate features
abstract
This paper proposes a new approach to the feature based unsupervised image segmentation. The difficulty with the conventional unsupervised segmentation lies in finding appropriate features that discriminate a meaningful region from the others. In this paper, the appropriate features are automatically learnt by machine learning with boosting scheme. At the initial step, the image is split into many small regions (blocks at first) and strong classifiers for every region, which discriminate the region from the others, are found by AdaBoosting. Each strong classifier so obtained is the weighted sum of several popular weak classifiers (features), which best describes the coherence of the region and thus well discriminates the region from the others. The output of this classifier is used in designing the energy function for the labeling, in the form of conditional random fields (CRFs). Minimization of the energy function produces the labeling result which reflects the property learnt by the classifier. For the labeling result, the machine learning is again performed and the process iterates until some conditions are met. Experimental results show that the proposed method provides competitive result compared to the conventional feature based methods.
Sang-Hak Lee, Hyung Il Koo, Nam Ik Cho
ICIP3
2009 Stereo matching using hierarchical belief propagation along ambiguity gradient
abstract
This paper proposes a stereo matching algorithm based on hierarchical belief propagation and occlusion handling. We define a new order for message passing in belief propagation instead of the scanline approach. The primary assumption is that a pixel with a well-defined minimum in its likelihood field is more likely to contain a correct disparity, when compared to a pixel having an ill-defined minimum with several local minima. The order for message passing is determined by the variance of likelihood field at each pixel. The variances evaluate the ambiguity of likelihood fields, and the messages are hierarchically updated along the gradient of ambiguity. The experimental results show that the proposed method estimates the disparities correctly in the hard regions such as large occlusions and textureless regions. The proposed algorithm is currently tied with the best performing algorithm on the Middlebury stereo site.
Sumit Srivastava, Seong Jong Ha, Sang Hwa Lee, Nam Ik Cho, Sang Uk Lee
ICIP4
2009 A new intra prediction method using channel correlations for the H.264/AVC intra coding
abstract
This paper introduces a new intra prediction algorithm that exploits the correlations between the color channels, for the H.264/AVC intra coding. The main idea is to use the weighted sum of neighboring pixels as a new chroma prediction tool, where the weighting coefficients are determined considering the inter-channel correlations. Specifically the weighting coefficients are determined under the assumption that the similarity of a chrominance channel intensity increases as the distance of the luminance channel intensity decreases. The proposed method is added to the conventional chroma prediction modes with a little modification in the chroma prediction mode binarization. The experiments show that the proposed method provides up to 1.0 dB gain for the color-abundant images.
Sang Heon Lee 0004, Jae Won Moon, Jae Woan Byun, Nam Ik Cho
PCS4
2009 Automatic Defect Classification Using Frequency and Spatial Features in a Boosting Scheme
abstract
An automatic defect classification algorithm is proposed in a boosting manner. The proposed method exploits the histogram of spatial orientation and frequency features. Specifically, the spatial gradient orientations of defect image are accumulated to be a histogram, and they are trained by SVM to construct a classifier. The frequency features are the projection of 2-D Haar patterns on the frequency responses. The classifiers using these spatial and frequency features are combined in a boosting manner to improve the classification performance. According to the experiments with 100 training and testing sets, the proposed boosting method improves the classification performance compared with the previous works using optical features such as colors, shapes, and sizes of defects.
Hong Il Kim, Sang Hwa Lee, Nam Ik Cho
IEEE Signal Process. Lett.3
2009 DCT-Based Embedded Image Compression With a New Coefficient Sorting Method
abstract
This letter presents a new sorting algorithm for the DCT-based embedded image compression. The magnitudes of DCT coefficients are estimated from the ones in the neighboring blocks, and the coefficients with larger estimates are given higher priority in encoding. This priority information is also used as contexts for arithmetic coding. The proposed method finds more significant coefficients earlier than the conventional methods, thereby providing higher coding gain.
Han Sae Song, Nam Ik Cho
IEEE Signal Process. Lett.2
2009 Composition of a Dewarped and Enhanced Document Image From Two View Images
abstract
In this paper, we propose an algorithm to compose a geometrically dewarped and visually enhanced image from two document images taken by a digital camera at different angles. Unlike the conventional works that require special equipment or assumptions on the contents of books or complicated image acquisition steps, we estimate the unfolded book or document surface from the corresponding points between two images. For this purpose, the surface and camera matrices are estimated using structure reconstruction, 3-D projection analysis, and random sample consensus-based curve fitting with the cylindrical surface model. Because we do not need any assumption on the contents of books, the proposed method can be applied not only to optical character recognition (OCR), but also to the high-quality digitization of pictures in documents. In addition to the dewarping for a structurally better image, image mosaic is also performed for further improving the visual quality. By finding better parts of images (with less out of focus blur and/or without specular reflections) from either of views, we compose a better image by stitching and blending them. These processes are formulated as energy minimization problems that can be solved using a graph cut method. Experiments on many kinds of book or document images show that the proposed algorithm robustly works and yields visually pleasing results. Also, the OCR rate of the resulting image is comparable to that of document images from a flatbed scanner.
Hyung Il Koo, Nam Ik Cho
IEEE Trans. Image Process.3
2008 An improved soft feedback V-Blast detection technique for TURBO-MIMO systems
abstract
In this paper, an improved minimum mean square error (MMSE) soft feedback detector, called the soft input, soft output, and soft feedback (SIOF) symbol detector, is proposed for turbo multi-input multi-output (TURBO-MIMO) systems. The SIOF symbol detector is derived by minimizing the power of interference plus noise, given a priori probabilities of yet undetected layers and a posteriori probabilities of detected layers. As a result, soft feedback interference cancellation based on a posteriori information is derived, yielding symbol detection robust to error propagation effects. Furthermore, a low complexity implementation using approximate detection ordering and linear filtering is introduced. Simulations performed for block fading channels show that the SIOF symbol detector exhibits performance gains over the existing TURBO-BLAST algorithm [3].
Andrew C. Singer, Jungwoo Lee 0001, Nam Ik Cho
ICASSP4
2008 Principal subspace modification for multi-channel Wiener filter in multi-microphone noise reduction
abstract
In multi-microphone noise reduction for single desired speech signal, the principal subspace based multi-channel Wiener filter provides better performance compared with the conventional multi-channel Wiener filter. The principal subspace vector estimates the acoustic transfer function vector up to a scaling factor. However, as input SNR becomes lower, the error increases in the acoustic transfer function vector estimation. In this paper, we propose the principal subspace modification which is controlled by the angle between the principal subspace vector and the steering vector of the desired speech signal. In the simulation, the proposed method is evaluated with multi-channel speech data which are degraded by interfering noise coming from other direction. The simulation results show that the modification of principal subspace vector allows better performance compared to the conventional principal subspace based multi-channel Wiener filter.
Gibak Kim, Nam Ik Cho
ICASSP2
2008 Image denoising based on a statistical model for wavelet coefficients
abstract
In this paper, we propose a new statistical model for the relationship of wavelet coefficients and its application to image denoising. The magnitude of a wavelet coefficient usually shows high correlations with the nearby ones. This property has been exploited in many wavelet-based image processing techniques. However, conventional works consider only the local neighborhood of a coefficient when inferring its hidden state. Consequently, the image context is not faithfully reflected and thus there are sometimes visually annoying artifacts. We attempt to alleviate this problem by developing a new statistical model for the random field that is consisted of hidden variables of the overall band and thus includes global relationship of wavelet coefficients. In this model, the image context is encoded by the relations of hidden states, and the state plane is efficiently inferred by the sum-product algorithm. In the experiment, the proposed model is incorporated with the state-of-the-art denoising algorithm, namely BLS GSM (Bayes Least Square - Gaussian Scale Mixture). The results show that the proposed algorithm suppresses many annoying artifacts that exist in the conventional denoising methods, and thus improves the subjective quality.
Hyung Il Koo, Nam Ik Cho
ICASSP2
2008 Pnoramic mosaic system for mobile devices
abstract
This paper deals with a panoramic mosaic system for mobile devices. The proposed system is optimized for mobile devices by integer-programmable algorithms and auto-shot user interface. The proposed system consists of auto-shot interface, transform onto cylindrical surface, color compensation, local alignment, image stitching, and blending. The auto-shot interface senses user's camera motion and takes pictures when the camera motion is matched to the predefined motion model. The captured images are projected onto a cylindrical mosaic surface by modified projection equation. Exposure difference is removed by color compensation, where the luminance differences between images are adjusted and color saturation is eliminated. The projected images are locally aligned by fast hierarchical hexagonal search technique. Then, the optimal boundary of overlapped images are determined using dynamic programming and synthesized seamlessly. According to the experiments using mobile devices, the proposed system shows good performance compared with other PC-based mosaic algorithms.
Seong Jong Ha, Sang Hwa Lee, Yu Ri Ahn, Nam Ik Cho
ICIP4
2008 Camera-based document digitization using multiple images
abstract
Recently, there have been some attempts to use a potable digital camera for the document digitization. But unlike the conventional flatbed scanners, the document images taken by digital cameras suffer from the perspective distortion and the geometric distortion. These problems deteriorate not only the subjective quality but also the character recognition rate. In this paper, a new document dewarping algorithm that utilizes two input images is presented. The proposed algorithm requires neither auxiliary hardwares such that measure the 3D shape nor the impractical assumptions on the pose of books and cameras. Instead, the 3D shape of the book (document) surface is reconstructed from the geometric correspondence of two images and some post-processings. Contrast to the conventional works, the proposed algorithm do not need to find textlines. So it can be applied to not only text regions but also pictures, tables and mathematical equations. That is, we can rectify distorted images irrespective of languages or contents without auxiliary hardwares.
Hyung Il Koo, Nam Ik Cho
ICIP3
2008 Boosting image segmentation
abstract
This paper presents a new approach to image segmentation, based on the conditional random fields (CRF) modeling and AdaBoost. In the proposed segmentation algorithm, the discriminating characteristics are first learned online using a training machine, and then the learnt characteristics are used to improve the region segmentation. The proposed algorithm is devised to include any kind of features even if they have different semantics, and to learn the difference of regions by selecting and combining only a few discriminating features among them. These novel properties are accomplished by a new Gibbs energy derived from CRF, AdaBoost, and probabilistic interpretation of its strong classifier. Experimental results on various images show the effectiveness of the proposed method.
Hyung Il Koo, Nam Ik Cho
ICIP2
2008 MAP-MRF approach for binarization of degraded document image
abstract
We propose an algorithm for the binarization of document images degraded by uneven light distribution, based on the Markov Random Field modeling with Maximum A Posteriori probability (MAP-MRF) estimation. While the conventional algorithms use the decision based on the thresholding, the proposed algorithm makes a soft decision based on the probabilistic model. To work with the MAP-MRF framework we formulate an energy function by a likelihood model and a generalized Potts prior model. Then we construct a graph for the energy, and obtain the optimized result by using the well-known graph cut algorithm. Experimental results show that our approach is more robust to various types of images than the previous hard decision approaches.
Jung Gap Kuk, Nam Ik Cho, Kyoung Mu Lee
ICIP2
2008 Non-rigid image registration based on the globally optimized correspondences
abstract
In this paper, we propose a new approach to the non-rigid image registration. This problem can be easily attacked if we can find regularly distributed correspondence points over the whole image or over the objects of interest. Dense and stable image registration can be achieved by using some natural mapping (e.g., thin plate spline) of these correspondences. However, the problems with conventional correspondence matching methods are that the features can rarely be found at the textureless regions and the matching accuracy is degraded at the parts with non-rigid motions. In order to find the regularly spaced correspondences and their accurate matching even under the non-rigid motion, we place mesh nodes over the image and develop a new cost function that considers three complementary terms: similarity, smoothness and some topological constraint that prevents unlikely mappings. Experimental results demonstrate that the proposed method can find correct correspondences in the presence of non-rigid motions, multi-layers (motion discontinuity) and even in the textureless regions. Experimental results also show that the proposed method can be applied to old film restoration as well as image registration.
Hyung Il Koo, Jung Gap Kuk, Nam Ik Cho
ICPR3
2008 Extending the lifetime of media recorders constrained by battery and flash memory size
abstract
The lifetime of a stand-alone media recorder is a function of both the battery size and flash memory size. In this paper, we present a power management framework for media recorders that significantly enhances their lifetime while minimizing the flash memory usage and maintaining the same level of recording quality. This is achieved by implementing a mixture of encoding algorithms of different complexities that generate data with different compression ratios, and in turn balancing the energy consumption and the flash memory usage.
Younghyun Kim 0001, Youngjin Cho, Naehyuck Chang, Chaitali Chakrabarti, Nam Ik Cho
ISLPED5
2008 Frequency domain multi-channel noise reduction based on the spatial subspace decomposition and noise eigenvalue modification
Gibak Kim, Nam Ik Cho
Speech Commun.2
2007 Noise Eigenvalue Modification Methods for Spatial Subspace Based Multi-Channel Speech Enhancement
abstract
In this paper, frequency domain multi-channel filtering schemes are proposed for speech enhancement, based on the subspace decomposition of spatial spectral matrices. For better estimation of noise statistics, which is important for most speech enhancers, we propose noise eigenvalue modification methods for the correction of noise spatial spectral matrix. These methods are based on the rank-1 property of the speech spatial spectral matrix for single desired speech source. Simulation results show that the proposed methods yield better performance compared to the conventional multi-channel Wiener filtering.
Gibak Kim, Nam Ik Cho
ICASSP (4)2
2007 Prior Model for the MRF Modeling of Multi-Channel Images
abstract
In multi-channel images (e.g. color images with R, G, B channel, and multi-spectral images), there exist higher-order correlations among the channels. We develop a new MRP-MAP (Markov random field - maximum a posteriori) framework that can be used for various multi-channel image processing. Main features of the proposed framework is that the higher-order correlation between the channels is considered, whereas it is not well addressed in the conventional works. Given a channel image, the prior probability of another channel is computed based on the MRF modeling that the channel correlation is described as piecewise linear relationship. An optimization algorithm for the MAP estimation is also developed. The effectiveness of the proposed priors is demonstrated with a simple application, i.e., image denoising.
Hyung Il Koo, Nam Ik Cho
ICASSP (1)2
2007 Hybrid Resolution Switching Method for Low Bit Rate Video Coding
abstract
This paper proposes a video coding method using hybrid resolution switching for low bit rate environments. The proposed method encodes I pictures in high resolution and B and P pictures in low resolution. The decimated inter-frame pictures are encoded in the usual H.264/AVC framework, and they are reconstructed to high resolution ones using motion information, interpolation filter, and some residual signals in high resolution. The proposed video coding scheme shows better compression performances in low bit rate environments than the traditional algorithms based on H.264/AVC. The resolution switching method increases the coding efficiency in low bit rates, mainly because the side information of motion estimation is reduced. It is expected that the proposed method can also improve higher bit rate cases, if some parameters optimization and deblocking filter are appropriately applied.
Sang Heon Lee 0004, Sang Hwa Lee, Nam Ik Cho
ICIP (6)3
2007 Stereo Matching using Multi-Directional Dynamic Programming and Edge Orientations
abstract
This paper proposes a stereo matching algorithm which employs an adaptive multi-directional dynamic programming (DP) scheme using edge orientations. A new energy function is defined in order to consider the discontinuity of disparity and occlusions, which is minimized by the multi-directional DP scheme. Chain codes are introduced to find the accurate edge orientations which provide the DP scheme with optimal multidirectional paths. The proposed algorithm eliminates the streaking problem of conventional DP based algorithms, and estimates more accurate disparity information in boundary areas. The experimental results using the Middlebury stereo images [7] demonstrate that our algorithm shows better performance than previous DP based approaches.
Min-Chul Sung, Sang Hwa Lee, Nam Ik Cho
ICIP (1)3
2007 Voice activity detection using the phase vector in microphone array
Gibak Kim, Nam Ik Cho
INTERSPEECH2
2007 Bayesian object extraction from uncalibrated image pairs
Hyung Il Koo, Sang Hwa Lee, Nam Ik Cho
Signal Process. Image Commun.3
2006 Low-Power Adaptive FIR Equalizer Via Soft Error Cancellation
abstract
In this paper, we present an adaptive FIR equalizer which reduces power dissipation by employing a new algorithmic error correction technique. Building on the voltage over-scaling (VOS) technique, we formulate the statistical estimation of timing errors that may be caused by VOS, called soft errors to detect and cancel them at a system level. We derive a minimum variance unbiased estimator, and develop an adaptive and power-optimized algorithm for an adaptive equalizer. Up to 30% power savings are demonstrated with negligible performance loss for an example, 16-tap minimum mean square error (MMSE) FIR equalizer.
Andrew C. Singer, Nam Ik Cho
ICASSP (4)3
2006 Interpolation of Multi-Spectral Images Inwavelet Domain for Satellite Image Fusion
abstract
This paper presents an image interpolation technique for satellite image fusion in the wavelet domain. For the fusion of satellite images, we need to interpolate and match the low resolution multi-spectral images (MSIs) to the high resolution panchromatic images. But since the typical spline-based interpolation methods employed in the conventional image fusion entails blurriness of edges, we propose a wavelet-domain image interpolation method that creates high frequency details based on the estimation of non-existent higher band coefficients from the relationship of available coefficients in lower scales. We model the relationship of coefficients in vertical as well as horizontal direction by the Markov stochastic model, and also find the coefficients of higher scale in this respect. The estimated coefficients are further refined by adding the maximum a posteriori (MAP) estimation process. The proposed interpolation technique is employed into the most popular image fusion algorithms, namely wavelet, principle component analysis (PCA), and intensity-hue-saturation (IHS) transformation based algorithms, instead of the conventional bilinear or bicubic interpolation methods. The experimental results show that the fused image based on the proposed interpolation instead of the conventional bilinear and bicubic interpolation algorithms employed in the conventional wavelet, principle component analysis (PCA), and intensity-hue-saturation (IHS) transformation based fusion algorithms, we incorporate the proposed wavelet based interpolation method.
Hak Chang Kim, Sang Hwa Lee, Nam Ik Cho
ICIP4
2006 Bayesian Image Interpolation Based on the Learning and Estimation of Higher Bandwavelet Coefficients
abstract
This paper presents an image interpolation algorithm based on the estimation of higher band wavelet coefficients. We presume that a given image is the LL band of the wavelet coefficients of a high resolution image that does not actually exist and is target of the interpolation. The proposed method estimates the higher band coefficients by learning the correlation of coefficients across the scale. According to the wavelet theory, a sequence of wavelet coefficients has extreme at the point that corresponds to the singularity of signal, and the extremes across the wavelet scale have some relationship. The main point of the wavelet domain interpolation is to exploit these properties of wavelet coefficients for estimating the extreme points in the higher frequency bands. In this paper, the relationship between the wavelet coefficients across the scale is described by Markov stochastic model, and each wavelet coefficient is modeled by Gaussian mixture that has multiple means and variances. For the enhanced subjective quality of interpolated image through the above modeling, we added refinement process using maximum a posteriori (MAP) technique. Comparison with the existing wavelet-domain and edge-preserving interpolation algorithms shows that the proposed method provides improved objective and subjective quality.
Sang Hwa Lee, Nam Ik Cho
ICIP3
2006 MAP-Based Object Extraction from Uncalibrated Image Pair
abstract
In this paper, we propose a new MAP algorithm for foreground objects extraction from uncalibrated image pairs. The segmentation is performed in the MAP framework with MRF (Markov random field) modeling of images. The proposed algorithm estimates several spatial transformations between two images by corresponding SIFT (scale invariant feature transform) points and sequential RANSAC (random sample consensus) algorithm. The area-ratio criterion is applied to the each transformation so that we select the transformation of foreground object. Using these transformations, we compute the likelihood of color-segmented subregions. We model the prior information which is based on smoothness condition. Finally, the object extraction is performed by the Bayesian belief propagation. Experiments on various image pairs and video sequences show promising results in extracting the foreground objects from the backgrounds.
Hyung Il Koo, Sang Hwa Lee, Nam Ik Cho, Seong Keun Kim, Dong Hahk Lee, Sunghoon Lee
ICIP3
2006 Stochastic Approach to Separate Diffuse and Specular Reflections
abstract
This paper presents separation of specular and diffuse reflection components from an image pair. The proposed approach is based on the dichromatic reflectance model and Markov random field models. The proposed method estimates specular and diffuse components by minimizing observation color noise and prior potential of specular reflectance. The specular reflection component is modelled as an MRF and estimated in maximum a posteriori framework. This paper proposes likelihood term of specular component and the prior model of specular reflectance based on Phong's shading. Some experiments show that the proposed approach separates the specular and diffuse reflection components effectively. The separated specular components can be utilized in image-based lighting which renders a scene with virtual lighting sources.
Sang Hwa Lee, Hyung Il Koo, Nam Ik Cho, Jong-Il Park
ICIP3
2006 Video Transcoding for Packet Loss Resilience Based on the Multiple Descriptions
abstract
This paper proposes error resilient video transcoding structures based on the multiple description (MD) scheme. Two structures are proposed for different use, namely low complexity simple MD transcoding structure (simple MD) and adaptive MD transcoding structure (adaptive MD). Simple MD structure always converts and splits the given bit stream into two descriptions regardless of channel condition. On the other hand, the adaptive MD structure considers the channel condition, and it switches between the MD and single description (SD) mode depending on the packet loss rate (PLR). In the switching process, the best mode which produces less end-to-end cost is chosen between SD and MD under given channel condition. The expected end-to-end cost is optimally estimated based on the rate-distortion theory at the time of encoding. The simulation results show that the simple MD structure provides efficient performance than the conventional error resilient transcoding structures based on the spatial and temporal error localization. And the adaptive MD structure shows higher PSNR than the conventional methods under the comparable complexity
Il Koo Kim, Nam Ik Cho
ICME2
2006 Two-microphone voice activity detection in the presence of coherent interference
Gibak Kim, Nam Ik Cho
INTERSPEECH2
2006 Hybrid multiple description video coding using optimal DCT coefficient splitting and SD/MD switching
Il Koo Kim, Nam Ik Cho
Signal Process. Image Commun.2
2006 A requantization algorithm for the transcoding of JPEG images
Jae Won Moon, Jong Seok Lee, Nam Ik Cho
Signal Process. Image Commun.3
2005 Energy-efficient digital filtering using ML-based error correction (ML-EC) technique
abstract
We present a maximum likelihood-based error correction (ML-EC) technique which achieves significant power savings in digital filtering. Although voltage over-scaling (VOS) can achieve high energy efficiency, it can introduce "soft errors" which severely degrade the performance of the filter. The proposed scheme detects, estimates and corrects these soft errors via an ML-based algorithm that achieves up to 47% power savings without any SNR loss and up to 60% power savings with a 1.5 dB SNR loss for an example case study of a frequency-selective low-pass filter.
Byonghyo Shim, Andrew C. Singer, Nam Ik Cho
ICASSP (4)4
2005 Design of perfect reconstruction QMF lattice with signed powers-of-two coefficients using CORDIC algorithm
abstract
The lattice structure has several advantages over the tapped delay line form, especially for the hardware implementation of general digital filters. It is also efficient for the implementation of quadrature mirror filters (QMF), because the perfect reconstruction (PR) is conserved even under the severe coefficient quantization. Moreover, if lattice coefficients are implemented by signed powers-of-two (SPT), the hardware complexity can also be reduced. But the discrete space represented by the SPT is sparse when the number of non-zero bits is small. This paper proposes an orthogonal QMF lattice with SPT coefficients that can provide much denser discrete coefficient space than the conventional structure. For this purpose, we employ the CORDIC algorithm that is structurally related to the PR lattice filter with SPT coefficients. The paraunitariness of CORDIC subrotation also continues to hold the PR condition to our wishes. Since the proposed architecture provides denser coefficient space, it shows less coefficient quantization error than the conventional QMF lattice.
Nam Ik Cho
ICASSP (4)2
2005 A fast algorithm for the conversion of DCT coefficients to H.264 transform coefficients
abstract
This paper proposes a fast algorithm that converts DCT coefficients into integer transform coefficients, for the transform domain transcoding from MPEG-x to H.264. For the transcoding in the same resolution, the 8 /spl times/ 8 DCT coefficients are converted to four 4/spl times/ 4 integer transform coefficients by decomposing the conversion matrix into sparse ones. For the reduction of resolution by half, we also propose an algorithm that converts DCT coefficients in the lower band into 4 /spl times/ 4 integer transform coefficients. The sparse matrices derived in this paper require fewer computations than the direct and conventional conversion matrices, and thus the overall transcoding using the proposed algorithm requires less computational complexity.
Chan Yul Park, Nam Ik Cho
ICIP (3)2
2005 Blocking artifact reduction method based on non-iterative POCS in the DCT domain
abstract
The post-processing method based on projections onto convex sets (POCS) has demonstrated good performance for blocking artifact reduction. However the iterative procedure in POCS requires a lot of computations, and would be infeasible for practical post-processing applications with real-time constraints. In this paper, we propose non-iterative post-processing method based on POCS in DCT domain for complexity reduction. In DCT domain POCS (D-POCS), the low-pass filtering (LPF) is performed in the DCT domain to remove the inverse DCT and forward DCT modules from POCS. Through the investigation of LPF in iterative POCS, we define kth order LPF which is equivalent to the LPF with k iterations. By combining DCT domain filtering and the kth order LPF concept, we define kth order DCT domain LPF. Simulation results show that the proposed D-POCS without iteration gives very close PSNR and subjective quality performance compared to the conventional POCS with iterations, while it requires much less computational complexity. If we take into account typical sparseness in DCT coefficients, the D-POCS method gives tremendous complexity reduction. Hence the proposed D-POCS would be an attractive method for practical real-time post-processing applications.
Changboon Yim, Nam Ik Cho
ICIP (2)2
2005 An efficient multicategory classifier based on AdaBoosting
abstract
In this paper, we propose an efficient multicategory classifier based on AdaBoosting scheme. The multicategory problems can be solved by multiple use of two-category classifiers or by use of a single classifier with multiple discriminant functions. In the case of boosting algorithms, since the use of simple classifier is one of the most important ingredients, they have focused on two-category classifier for each weak classifier. But for applying the two-category booster to m-category problems, we need O(m/sup 2/) boosters instead of O(m) ones arrangement scheme of the boosters as like detector-pyramid (S.Z. Li and Z. Zhang, 2004). We propose a multicategory boosting algorithm named M-Booster, where each weak classifier is the multicategory classifier. We focused on efficient method to extract the features and update the weights of data. The label for the each category is represented by m-dimensional vector, and the weights for the feature and other parameters are also modified accordingly. We have performed simulation for the artificial data and the face data with different rotation angles. It is shown that the use of single M-Booster can solve the multicategory problems more efficiently than the method based on 2-category classifiers and previous method (Adaboost.MH) (Y. Freund and R.E. Schapire, 1997).
Hong Il Kim, Sang Hwa Lee, Nam Ik Cho
ICMLA3
2005 Automatic defect classification using boosting
abstract
This paper deals with automatic defect classification (ADC) in semiconductor fabrication. The defects such as particle and scratch are automatically classified using a boosting approach. The boosting scheme is based on the Kullback-Leibler distance and linear projection along feature vectors. The paper generates the linear features which discriminate the defects maximally. The features are the linear combinations of Haar-like patterns in the frequency domain. By learning in a boosting manner, the particle and scratch are recognized out of the other defects. And, we propose another feature in the spatial domain which is based on the orientation histogram in the local region. The spatial feature is combined with frequency domain features in a boosting manner. According to the experiments with various defect samples, the accuracy of defect classification is larger than 92% on the average, and scratch is especially recognized with 98% purity. More improvement is expected by the new spatial features such as defect colors, shapes, textures, and so on.
Sang Hwa Lee, Hong Il Kim, Nam Ik Cho, Yu Han Jeong, Ki Suk Chung, Chung Sam Jun
ICMLA3
2004 Disparity estimation using color coherence and stochastic diffusion
abstract
This paper deals with disparity estimation based on Markov random field (MRF) models and color coherence. The disparity and line fields are explicitly modeled as MRFs, and are estimated by the stochastic diffusion. The potential functions are defined from the novel stochastic models between disparity and line fields. The color information is also utilized to model the textureless regions where the disparity estimation generally fails. And, the derived potential functions are minimized by the novel energy minimization method called stochastic diffusion. The stochastic diffusion diffuses the potential space using the probability distribution of neighboring fields, and searches for the optimal fields in the converged potential space. Some experiments show good performances of disparity estimation, which are compared with the other methods in a Webpage.
Sang Hwa Lee, Nam Ik Cho, Jong-Il Park
ICIP2
2004 Segmentation based disparity estimation using color and depth information
abstract
The well-known cooperative stereo uses two dimensional rectangular window for a local block matching, and three dimensional box-shaped volume for a global optimization procedure. In many cases, appropriate selections of these matching regions can provide satisfactory matching results. This paper presents a new method for iteratively modifying sizes and shapes of matching regions based on color and depth information. This algorithm computes the aggregated matching costs with two ideas. The first idea is to select matching regions based on object boundaries to avoid projective distortion. This provides the reliable matching scores as well as the prevention of the foreground fattening phenomenon. The second idea is to iteratively modify the segmentation map by merging the regions where the disparities are likely to be the same. Experimental results show that the proposed algorithm provides more accurate disparity map than other algorithms. Especially, the computed disparity map shows the advantage of our algorithm in disparity discontinuity regions.
Sang Hwa Lee, Nam Ik Cho
ICIP3
2003 Fixed point error analysis of CORDIC processor based on the variance propagation
abstract
The effects of angle approximation and rounding in the CORDIC processor have been intensively studied for the determination of design parameters. However, the conventional analyses provide only the error bound which results in large discrepancy between the analysis and the actual implementation. Moreover, some of the signal processing architectures require the specification in terms of the mean squared error (MSE) as in the design specification of FFT processor for OFDM. This paper proposes a fixed point MSE analysis based on the variance propagation for more accurate error expression of the CORDIC processor. It is shown that the proposed analysis can also be applied to the modified CORDIC algorithms. As an example of application, an FFT processor for OFDM using the CORDIC processor is presented. The results show close match between the analysis and simulation.
Nam Ik Cho
ICASSP (2)2
2003 Error resilient video coding using optimal multiple description of DCT coefficients
abstract
This paper proposes an algorithm for the robust transmission of video in error prone environment using multiple description (MD) scheme and optimal splitting of DCT coefficients. We use the redundancy rate-distortion (RRD) criteria to split a one-layer compressed video stream into two correlated streams or descriptions. For optimal splitting, recursive structure is exploited and Lagrange optimization and dynamic programming are used. Compared to the existing RRD based methods, the proposed algorithm can find more optimal points on the RRD curve. For compliance with the standards, the proposed MD video coder is implemented based on the H.263. Hence, each description can be decoded independently by the H.263 decoder. Also, several descriptions can be decoded into a single stream by additional simple merge stage and the H.263 decoder. Simulation results show that the proposed MD video coder yields better performance than the conventional MD splitting algorithms based on the RRD criteria at all redundancy rates.
Il Koo Kim, Nam Ik Cho
ICIP (2)2
2003 A relevance feedback algorithm based on the clustering and Parzen window
abstract
A relevance feedback algorithm based on the nonparametric approach is proposed. In the feature space, the algorithm generates multiple hyper-spheres around the regions where the images relevant with the query are densely populated, whereas the conventional algorithm searches the images in a single hyper-ellipsoid region. Then the Parzen window approach is applied to estimate the probability of relevance of each image in these multiple clusters (hyper-spheres). As a result, the relevance region in the feature space expands rapidly and covers arbitrarily shaped spaces with a small number of parameters. Also, since the user needs to determine only the positive images not the ambiguous negative ones, it is more convenient to use compared to some of the existing algorithms requiring negative feedback.
Hyung Il Koo, Nam Ik Cho
ICIP (2)2
2003 Simultaneous object extraction and disparity estimation using stochastic diffusion
abstract
This paper proposes a Bayesian maximum a posteriori (MAP) method to estimate disparity field and to extract objects in the stereoscopic images. The disparity and segmentation fields are modelled as the Markov random fields (MRFs), and are estimated by a stochastic approach called stochastic diffusion. The stochastic diffusion is a new energy minimization method to search for the solution fields in the MAP estimation. A clustering method utilizes the disparity field and color information to classify the regions into foreground objects or background. The line field is also included to improve the detection of the object boundaries. According to the some experiments, the proposed approach shows good performances in the foreground objects extraction even in the complex background.
Sang Hwa Lee, Nam Ik Cho, Yasuaki Kanatsugu, Jong-Il Park
ICIP (3)2
2003 Occlusion detection and stereo matching in a stochastic method
abstract
This paper proposes a stochastic approach to estimate the occlusion and disparity fields of stereoscopic images. The fields are estimated by the Bayesian maximum a posteriori (MAP) framework and Markov random field (MRF) models. The occlusion field model is based on the stochastic observation that the probability distribution in MAP estimator is relatively unstable and uniform at occluded regions. The occlusion is explicitly modelled as MRF, and is estimated in an energy optimization method called stochastic diffusion. The detected occluded region is compensated for the re-estimated disparity field in the same stochastic diffusion and the dynamic programming approach. Experimental results show good occlusion detection and disparity estimation. These results show the novel stochastic approach is suitable for occlusion detection and disparity estimation.
Sang Hwa Lee, Nam Ik Cho, Yasuaki Kanatsugu, Jong-Il Park
ICIP (1)3
2003 Fixed point error analysis of CORDIC processor based on the variance propagation
abstract
The effects of angle approximation and founding in the CORDIC processor have been intensively studied for the determination of design parameters. However, the conventional analyses provide only the error bound which results in large discrepancy between the analysis and the actual implementation. Moreover, some of the signal processing architectures require the specification in terms of the mean squared error (MSE) as in the design specification of FFT processor for OFDM. This paper proposes a fixed point MSE analysis based on the variance propagation for more accurate error expression of CORDIC processor. It is shown that the proposed analysis can also be applied to the modified CORDIC algorithms. As an example of application, an FFT processor for OFDM using the CORDIC processor is presented. The results show close match between the analysis and simulation.
Nam Ik Cho
ICME2
2003 Rate-distortion optimization of the image compression algorithm based on the warped discrete cosine transform
Il Koo Kim, Nam Ik Cho, Sanjit K. Mitra
Signal Process.2
2003 Relevance feedback in content-based image retrieval system by selective region growing in the feature space
Jung Won Kwak, Nam Ik Cho
Signal Process. Image Commun.2
2002 Facial feature tracking by robust face segmentation and scalable rotational BMA
Jung S. Kim, Nam Ik Cho, Seok-Cheol Kee, Sang Uk Lee
VCIP2
2002 Scene change detection by feature extraction from strong edge blocks
Han Sae Song, Il Koo Kim, Nam Ik Cho
VCIP3
2002 Suppression of narrow-band interference in DS-spread spectrum systems using adaptive IIR notch filter
Nam Ik Cho
Signal Process.2
2001 Narrow-band interference suppression in direct sequence spread spectrum systems using a lattice IIR notch filter
abstract
This paper proposes an algorithm for the suppression of narrow-band interference in direct sequence spread spectrum (DSSS) systems, based on the open loop adaptive IIR notch filtering. The center frequency of the interference is monitored on-line by the adaptive lattice IIR notch filter in Cho and Lee (1993) or by time-frequency analysis in Amin (1997). The power of the interference signal is also estimated from the adaptive filters. Another lattice IIR notch filter is placed in front of the receiver, the notch of which is controlled by the frequency estimate to remove the interference. However, the IIR notch filter with the zeros on the unit circle also removes the information signal at the notch frequency while removing the interference and causes data distortion. Hence, the depth of the notch should also be adjusted for the trade-off between data distortion and effective interference reduction. The objective function for adjusting the depth of the notch is defined as the overall signal to noise ratio (SNR). The SNR is expressed as a function of filter parameters and the notch depth that maximizes the SNR is found. Simulation results show that the proposed algorithm yields better performance than the existing FIR notch filter and the conventional FIR LMS algorithm with very long taps.
Nam Ik Cho
ICASSP2
2001 Relevance feedback in image retrieval system by region growing in the feature space
abstract
We propose a relevance feedback algorithm based on region growing in the feature space, where most conventional algorithms are based on weight updating. By adaptively expanding match regions based on the user's feedback, the region can have arbitrary shape in the feature space, whereas the weight update methods have typical hyper-ellipsoidal shape. As a result, the proposed algorithm shows matching results with higher precision.
Jung Won Kwak, Jae Jun Lee, Nam Ik Cho
MMSP3
2001 Reduction of blocking artifacts by cepstral filtering
Nam Ik Cho
Signal Process.1
2000 Reduction of blocking artifacts by a modeled lowpass filter output
abstract
This paper proposes an algorithm for the reduction of blocking artifacts in the transform coded images and videos. The algorithm is based on the filtering of block boundaries. But, the proposed technique is not the actual filtering in the sense that the output is not obtained by multiplying the pixel values with the filter coefficients. Instead of filtering the pixel values through the filter coefficients, the output of the filter is modeled as a typical step response function of the lowpass filter or a functions with similar shapes. Then the blocky signal is just replaced by the modeled output signal which is generated by the library function such as exp or sin. The algorithm can be easily implemented by a lookup-table and very few multiplications, while the objective and subjective performance is comparable to those of other algorithms based on the conventional filtering and POCS (projection onto convex sets).
Nam Ik Cho, Bong Gyun Roh, Sang Uk Lee
ISCAS1
2000 Warped discrete cosine transform and its application in image compression
abstract
This paper introduces the concept of warped discrete cosine transform (WDCT) and an image compression algorithm based on the WDCT. The proposed WDCT is a cascade connection of a conventional DCT and all-pass filters whose parameters can be adjusted to provide frequency warping. Because only the first-order all-pass filters are considered, the WDCT can be implemented by a Laguerre network connected with the DCT. For the more efficient software implementation, we propose truncated and approximated FIR filter banks which can be used instead of the Laguerre network. As a result, the input-output relationship of the WDCT can be represented by a single matrix-vector multiplication, like the DCT. In the proposed image-compression scheme, the frequency response of the all-pass filter is controlled by a fixed set of parameters from which a specified warping parameter is used for a specified frequency range. Also, for each parameter, the corresponding WDCT matrices are computed a priori. For each image block, the best parameter is chosen from the set and the index is sent to the decoder as side information along with the result of corresponding WDCT matrix computation. At the decoder, an inverse WDCT is performed to reconstruct the image. The WDCT based compression outperforms the DCT based compression, for high bit rate applications and for images with high-frequency components. It results in 1.1-3.1-dB PSNR gain over conventional DCT at 1.5 bpp for natural images, and provides more gain for compound images with texts.
Nam Ik Cho, Sanjit K. Mitra
IEEE Trans. Circuits Syst. Video Technol.1
1999 An Image Compression Algorithm Using Warped Discrete Cosine Transform
abstract
This paper introduces the concept of warped discrete cosine transform (WDCT) and an image compression algorithm based on the WDCT. The proposed WDCT is a cascade connection of a conventional DCT and all-pass filters whose parameters can be adjusted to provide frequency warping. In the proposed image compression scheme, the frequency response of the all-pass filter is controlled by a set of parameters with each parameter for a specified frequency range. For each image block, the best parameter is chosen from the set and is sent to the decoder as a side information along with the result of corresponding WDCT coefficients. It is shown that the WDCT can be implemented by a single matrix computation like the DCT. The proposed algorithm outperforms the DCT based compression, for high bit rate applications and for images with high frequency components.
Nam Ik Cho, Sanjit K. Mitra
ICIP (2)1
1999 An adaptive quantization algorithm for video coding
abstract
This paper proposes an adaptive quantization algorithm for video coding using the information obtained from the previously encoded image. Before quantizing the discrete cosine transform coefficients, the properties of reconstruction error of each macro block (MB) are estimated from the previous frame. For the estimation of the error of current MB, a block with the size of MB in the previous frame is chosen. Since the original and reconstructed images of the previous frame are available in the encoder, we can evaluate the tendency of reconstruction error of this block in advance. Then, this error is considered as the expected error of the current MB if it is quantized with the same step size and bit rate. Comparing the error of the MB with the average of overall MBs, if it is larger than the average, a small step size is given for this MB, and vice versa. As a result, the error distribution of the MB is more concentrated to the average, yielding low variance and improved image quality. Especially for low bit applications, the proposed algorithm yields much smaller error variance and higher peak signal-to-noise ratio compared to the conventional TM5. We also propose a modified algorithm for efficient hardware implementation.
Nam Ik Cho, Heesub Lee, Sang Uk Lee
IEEE Trans. Circuits Syst. Video Technol.1
1998 Design and Implementation of Format Conversion Filters for MPEG-4
Nam Ik Cho, Kichul Kim, Sang Uk Lee
J. Vis. Commun. Image Represent.1
1995 On the performance analysis and applications of the subband adaptive digital filter
Yoon Gi Yang, Nam Ik Cho, Sang Uk Lee
Signal Process.2
1993 On the tracking properties of the ALNF (adaptive lattice notch filter)
Nam Ik Cho, Sang Uk Lee
ICASSP (3)1
1993 Design of FIR filter over a discrete coefficient space with applications to HDTV signal processing
Nam Ik Cho, Sang Uk Lee, Kiho Kim
ISCAS1
1993 On the performance analysis of the subband adaptive digital filter
Yoon Gi Yang, Nam Ik Cho, Sang Uk Lee
ISCAS2
1991 A fast algorithm for 2-D DCT
abstract
Recently, a novel fast algorithm for 2-D N*N DCT (discrete cosine transform), where N=2/sup m/, was proposed. Only half the number of multiplications required for the conventional row-column approach are needed. However, the relationship between the input-output indices for the postaddition stage in the algorithm is seemingly very irregular. In the present work, the authors derive general and systematic expressions for the relation of the postaddition stage in the 2-D DCT algorithm by representing it in matrix form and developing a method for partitioning the matrices. The results show that the signal flow graph from input to output has a recursive structure where the structure for smaller N appears recursively for larger N. Hence, one can obtain an organized and regular structure for the input-output relation in the postaddition stage.>
Nam Ik Cho, Il Dong Yun, Sang Uk Lee
ICASSP1
1990 On the performance analysis of ALE using an IIR lattice notch filter
abstract
An attempt is made at a quantitative understanding of the ALE (adaptive line enhancer) using an adaptive IIR (infinite impulse response) lattice notch filter. The stability of the algorithm is discussed, and an expression for the asymptotic bias of the frequency estimate is provided. The transient behavior of the algorithm is investigated. It is verified by simulation that the analysis substantiates the fast convergent properties shown by N.I. Cho et al. (1989). The cascade structure for retrieving multiple sinusoids is also discussed. The simulation results indicate that the cascade structure based on the lattice ALE algorithm provides results comparable to those described in the recent literature with much less computational complexity (i.e. O(N) versus O(N/sup 2/)), where N is the number of input sinusoids.>
Nam Ik Cho, Sang Uk Lee
ICASSP1