VLDB 2026 Research / reviewers in the wild / expert
Munchurl Kim
dblp:84/4965
· DBLP profile ↗
113ranked-venue papers
2as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 1 first-author · 30 since 2021Artificial intelligence and machine learning · 39 · 1 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoBGS: Motion Deblurring Dynamic 3D Gaussian Splatting for Blurry Monocular VideoabstractWe present MoBGS, a novel motion deblurring 3D Gaussian Splatting (3DGS) framework capable of reconstructing sharp and high-quality novel spatio-temporal views from blurry monocular videos in an end-to-end manner. Existing dynamic novel view synthesis (NVS) methods are highly sensitive to motion blur in casually captured videos, resulting in significant degradation of rendering quality. While recent approaches address motion-blurred inputs for NVS, they primarily focus on static scene reconstruction and lack dedicated motion modeling for dynamic objects. To overcome these limitations, our MoBGS introduces a novel Blur-adaptive Latent Camera Estimation (BLCE) method using a proposed Blur-adaptive Neural Ordinary Differential Equation (ODE) solver for effective latent camera trajectory estimation, improving global camera motion deblurring. In addition, we propose a Latent Camera-induced Exposure Estimation (LCEE) method to ensure consistent deblurring of both a global camera and local object motions. Extensive experiments on the Stereo Blur dataset and real-world blurry videos show that our MoBGS significantly outperforms the very recent methods, achieving state-of-the-art performance for dynamic NVS under motion blur. Minh-Quan Viet Bui, Jongmin Park 0001, Juan Luis Gonzalez 0001, Jaeho Moon, Jihyong Oh, Munchurl Kim |
AAAI | 6 |
| 2026 | DeepHQ: Learned Hierarchical Quantizer for Progressive Deep Image CodingabstractResearch on entropy model-based Learned Image Compression (LIC) has been actively progressing, leading to rapid advancements in coding efficiency. Beyond improvements in coding efficiency, LIC methods have also been explored for practical codec development. Despite these advancements, research on learned Progressive Image Coding (PIC) remains in its early stages. PIC aims to encode multiple quality levels into a single bitstream, improving bitstream versatility and achieving higher compression efficiency than simulcast compression. Existing learned PIC methods hierarchically quantize transformed latent representations with varying quantization step sizes. More specifically, these approaches progressively compress the additional information needed for quality improvement, considering that a wider quantization interval for lower-quality compression includes multiple narrower subintervals for higher-quality compression. However, they rely on handcrafted quantization hierarchies, leading to suboptimal compression efficiency. In this article, we propose a learned PIC method that first exploits learned quantization step sizes for each quantization layer. We also incorporate selective compression, ensuring that only essential representation components are retained in each quantization layer. Our experimental results demonstrate that the proposed method significantly enhances coding efficiency compared to the existing approaches while also reducing decoding time and model size. The source code is publicly available at https://github.com/JooyoungLeeETRI/DeepHQ . Jooyoung Lee 0004, Se Yoon Jeong, Munchurl Kim |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | U-Know-DiffPAN: An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-SharpeningabstractConventional methods for PAN-sharpening often struggle to restore fine details due to limitations in leveraging high-frequency information. Moreover, diffusion-based approaches lack sufficient conditioning to fully utilize Panchromatic (PAN) images and low-resolution multi-spectral (LRMS) inputs effectively. To address these challenges, we propose an uncertainty-aware knowledge distillation diffusion framework with details enhancement for PAN-sharpening, called U-Know-DiffPAN. The U-Know-DiffPAN incorporates uncertainty-aware knowledge distillation for effective transfer of feature details from our teacher model to a student one. The teacher model in our U-Know-DiffPAN captures frequency details through freqeuncy selective attention, facilitating accurate reverse process learning. By conditioning the encoder on compact vector representations of PAN and LRMS and the decoder on Wavelet transforms, we enable rich frequency utilization. So, the high-capacity teacher model distills frequency-rich features into a lightweight student model aided by an un certainty map. From this, the teacher model can guide the student model to focus on difficult image regions for PAN-sharpening via the usage of the uncertainty map. Extensive experiments on diverse datasets demonstrate the robustness and superior performance of our U-Know-DiffPAN over very recent state-of-the-art PAN-sharpening methods. The project page is available at https://kaist-viclab.github.io/U-Know-DiffPAN-site/. Sungpyo Kim, Jeonghyeok Do, Jaehyup Lee, Munchurl Kim |
CVPR | 4 |
| 2025 | MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has made significant strides in scene representation and neural rendering, with intense efforts focused on adapting it for dynamic scenes. Despite delivering remarkable rendering quality and speed, existing methods struggle with storage demands and the representation of complex real-world motions. To address these challenges, we propose MoDec-GS, a memory-efficient Gaussian splatting framework designed to reconstruct novel views in challenging scenarios with complex motions. We introduce Global-to-Local Motion Decomposition (GLMD) to effectively capture dynamic motions in a coarse-to-fine manner. This approach leverages Global Canonical Scaffolds (Global CS) and Local Canonical Scaffolds (Local CS), which extend static Scaffold representation to dynamic video reconstruction. For Global CS, we propose Global Anchor Deformation (GAD) to efficiently represent global dynamics along complex motions, by directly deforming the implicit Scaffold attributes which are anchor position, offset, and local context features. Next, we finely adjust local motions via the Local Gaussian Deformation (LGD) of Local CS explicitly. Additionally, we introduce Temporal Interval Adjustment (TIA) to automatically control the temporal coverage of each Local CS during training, enabling MoDec-GS to find optimal interval assignments based on the specified number of temporal segments. Extensive evaluations demonstrate that MoDec-GS achieves an average 70% reduction in model size over state-of-the-art methods for dynamic 3D Gaussians from real-world dynamic videos while maintaining or even improving rendering quality. Sangwoon Kwak, Joonsoo Kim, Jun Young Jeong, Won-Sik Cheong, Jihyong Oh, Munchurl Kim |
CVPR | 6 |
| 2025 | ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image ObjectsabstractWeakly supervised Oriented Object Detection (WS-OOD) has gained attention as a cost-effective alternative to fully supervised methods, providing efficiency and high accuracy. Among weakly supervised approaches, horizontal bounding box (HBox) supervised OOD stands out for its ability to directly leverage existing HBox annotations while achieving the highest accuracy under weak supervision settings. This paper introduces adaptive bounding box scaling and symmetry-prior-based orientation prediction, called ABBSPO that is a framework for WS-OOD. Our ABBSPO addresses the limitations of previous HBox-supervised OOD methods, which compare ground truth (GT) HBoxes directly with predicted RBoxes’ minimum circumscribed rectangles, often leading to inaccuracies. To overcome this, we propose: (i) Adaptive Bounding Box Scaling (ABBS) that appropriately scales the GT HBoxes to optimize for the size of each predicted RBox, ensuring more accurate prediction for RBoxes’ scales; and (ii) a Symmetric Prior Angle (SPA) loss that uses the inherent symmetry of aerial objects for self-supervised learning, addressing the issue in previous methods where learning fails if they consistently make incorrect predictions for all three augmented views (original, rotated, and flipped). Extensive experimental results demonstrate that our ABBSPO achieves state-of-the-art results, outperforming existing methods. Hyugjae Chang, Jaeho Moon, Jaehyup Lee, Munchurl Kim |
CVPR | 5 |
| 2025 | SplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular VideoabstractSynthesizing novel views from in-the-wild monocular videos is challenging due to scene dynamics and the lack of multi-view cues. To address this, we propose SplineGS, a COLMAP-free dynamic 3D Gaussian Splatting (3DGS) framework for high-quality reconstruction and fast rendering from monocular videos. At its core is a novel Motion-Adaptive Spline (MAS) method, which represents continuous dynamic 3D Gaussian trajectories using cubic Her-mite splines with a small number of control points. For MAS, we introduce a Motion-Adaptive Control points Pruning (MACP) method to model the deformation of each dynamic 3D Gaussian across varying motions, progressively pruning control points while maintaining dynamic modeling integrity. Additionally, we present a joint optimization strategy for camera parameter estimation and 3D Gaussian attributes, leveraging photometric and geometric consistency. This eliminates the need for Structure-from-Motion preprocessing and enhances SplineGS’s robustness in real-world conditions. Experiments show that SplineGS significantly outperforms state-of-the-art methods in novel view synthesis quality for dynamic scenes from monocular videos, achieving thousands times faster rendering speed. Jongmin Park 0001, Minh-Quan Viet Bui, Juan Luis Gonzalez 0001, Jaeho Moon, Jihyong Oh, Munchurl Kim |
CVPR | 6 |
| 2025 | BiM-VFI: Bidirectional Motion Field-Guided Frame Interpolation for Video with Non-uniform MotionsabstractExisting Video Frame interpolation (VFI) models tend to suffer from time-to-location ambiguity when trained with video of non-uniform motions, such as accelerating, decelerating, and changing directions, which often yield blurred interpolated frames. In this paper, we propose (i) a novel motion description map, Bidirectional Motion field (BiM), to effectively describe non-uniform motions; (ii) a BiM-guided Flow Net (BiMFN) with Content-Aware Upsampling Network (CAUN) for precise optical flow estimation; and (iii) Knowledge Distillation for VFI-centric Flow supervision (KDVCF) to supervise the motion estimation of VFI model with VFI-centric teacher flows. The proposed VFI is called a Bidirectional Motion field-guided VFI (BiM-VFI) model. Extensive experiments show that our BiM-VFI model significantly surpasses the recent state-of-the-art VFI methods by 26% and 45% improvements in LPIPS and STLPIPS respectively, yielding interpolated frames with much fewer blurs at arbitrary time instances. Wonyong Seo, Jihyong Oh, Munchurl Kim |
CVPR | 3 |
| 2025 | Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-Shot Skeleton-Based Action Recognition
Jeonghyeok Do, Munchurl Kim |
ICCV | 2 |
| 2025 | PAN-Crafter: Learning Modality-Consistent Alignment for Pan-SharpeningabstractPAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11$\times$ faster inference time and 0.63$\times$ the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions. Jeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee, Munchurl Kim |
ICCV | 5 |
| 2025 | One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesabstractZero-shot depth estimation (DE) models exhibit strong generalization performance as they are trained on large-scale datasets. However, existing models struggle with high-resolution images due to the discrepancy in image resolutions of training (with smaller resolutions) and inference (for high resolutions). Processing them at full resolution leads to decreased estimation accuracy on depth with tremendous memory consumption, while downsampling to the training resolution results in blurred edges in the estimated depth images. Prevailing high-resolution depth estimation methods adopt a patch-based approach, which introduces depth discontinuity issues when reassembling the estimated depth patches, resulting in test-time inefficiency. Additionally, to obtain fine-grained depth details, these methods rely on synthetic datasets due to the real-world sparse ground truth depth, leading to poor generalizability. To tackle these limitations, we propose Patch Refine Once (PRO), an efficient and generalizable tile-based framework. Our PRO consists of two key components: (i) Grouped Patch Consistency Training that enhances test-time efficiency while mitigating the depth discontinuity problem by jointly processing four overlapping patches and enforcing a consistency loss on their overlapping regions within a single backpropagation step, and (ii) Bias Free Masking that prevents the DE models from overfitting to dataset-specific biases, enabling better generalization to real-world datasets even after training on synthetic data. Zero-shot evaluations on Booster, ETH3D, Middlebury 2014, and NuScenes demonstrate that our PRO can be seamlessly integrated into existing depth estimation models. Byeongjun Kwon, Munchurl Kim |
ICCV | 2 |
| 2025 | A Dual-Decoder-VAE-Based Latent Diffusion Model for PAN-SharpeningabstractHigh-resolution (HR) electro-optical (EO) satellites generally obtain multispectral (MS) images of a lower spatial resolution than their corresponding panchromatic (PAN) images owing to physical constraints. Despite these challenges, HR MS images remain critical in fields such as defense, industrial monitoring, and disaster response. Therefore, PAN-sharpening techniques have been widely studied. Recent progress in PAN-sharpening has been driven by deep learning, with diffusion models (DMs) emerging as a promising direction. Recent diffusion-based PAN-sharpening methods in the pixel domain generate high-quality PAN-sharpened (PS) images, but they generally require 25–50 denoising steps, imposing high computational complexity. To mitigate these limitations, we attempted to utilize a latent DM (LDM) for the PAN-sharpening task. However, the conventional variational autoencoder (CVAE) in LDM cannot accurately reconstruct satellite EO images of high bit depths, relatively smaller dynamic ranges within the bit depths, and multichannel characteristics. In this study, we first identify the limitations of CVAE and propose a dual-decoder-VAE (DDV), which is more suitable for satellite EO images. Furthermore, we introduce a DDV-based latent diffusion PAN-sharpening model (DDV-LDP). DDV-LDP achieves 0.06 dB and 0.98 dB higher PSNR values than a state-of-the-art diffusion-based method (UKnowDif-T) on the KOMPSAT-3A and WorldView-III datasets, respectively, even with a 99.5% reduction in testing time. Hyun-Ho Kim, Munchurl Kim |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | MoBluRF: Motion Deblurring Neural Radiance Fields for Blurry Monocular VideoabstractNeural Radiance Fields (NeRF), initially developed for static scenes, have inspired many video novel view synthesis techniques. However, the challenge for video view synthesis arises from motion blur, a consequence of object or camera movements during exposure, which hinders the precise synthesis of sharp spatio-temporal views. In response, we propose a novel motion deblurring NeRF framework for blurry monocular video, called MoBluRF, consisting of a Base Ray Initialization (BRI) stage and a Motion Decomposition-based Deblurring (MDD) stage. In the BRI stage, we coarsely reconstruct dynamic 3D scenes and jointly initialize the base rays which are further used to predict latent sharp rays, using the inaccurate camera pose information from the given blurry frames. In the MDD stage, we introduce a novel Incremental Latent Sharp-rays Prediction (ILSP) approach for the blurry monocular video frames by decomposing the latent sharp rays into global camera motion and local object motion components. We further propose two loss functions for effective geometry regularization and decomposition of static and dynamic scene components without any mask supervision. Experiments show that MoBluRF outperforms qualitatively and quantitatively the recent state-of-the-art methods with large margins. Minh-Quan Viet Bui, Jongmin Park 0001, Jihyong Oh, Munchurl Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | FBS-PS: Fully Band-Separable PAN-Sharpening Considering the Physical Characteristics of Electro-Optical SensorsabstractElectro-optical (EO) satellites are primarily used for reconnaissance, national defense, and cartography. However, most high-resolution (HR) EO satellites obtain images at a lower resolution (LR) in the multispectral (MS) band compared to the panchromatic (PAN) band due to technical limitations. Deep learning-based PAN-sharpening methods have been continuously developed to address the growing demand for MS images with the same ground sample distance (GSD) as PAN images. The improvements in deep learning-based PAN-sharpening methods have focused on enhancing the network structure and often overlooked the physical characteristics of satellite EO sensors, leading to artifacts such as noise and color distortions in the PAN-sharpened (PS) images. Thus, we propose a fully band-separable PAN-sharpening (FBS-PS) method and its more elaborate quality-centric model, called FBS-PS+, that process MS images separately, effectively considering the physical properties of the corresponding EO sensors in the acquisition when generating PS images. This helps prevent unrelated information from being mixed among MS images and enables accurate feature extractions. Therefore, the generated PS images have less noise and reduced color distortions than previous PAN-sharpening methods that typically fuse MS images from front-end layers. In addition, we design a novel training method and loss function to handle the problem of misregistered MS and PAN images. Our FBS-PS+ outperforms all other PAN-sharpening methods in most reference-based quality metrics, while our FBS-PS, lightweight and faster, achieves comparable quality performance to local-global transformer enhanced unfolding network (LGTEUN) that has$7.95\times $more parameters and requires$4.06\times $more floating-point operations per second (FLOPs). Hyun-Ho Kim, Munchurl Kim |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Novel View Synthesis with View-Dependent Effects from a Single ImageabstractIn this paper, we address single image-based novel view synthesis (NVS) by firstly integrating view-dependent effects (VDE) into the process. Our approach leverages camera motion priors to model VDE, treating negative disparity as the representation of these effects in the scene. By identifying that specularities align with camera motion, we infuse VDEs into input images by aggregating pixel colors along the negative depth region of epipolar lines. Additionally, we introduce a ‘relaxed volumetric rendering’ approximation, enhancing efficiency by computing densities in a single pass for NVS from single images. Notably, our method learns single-image NVS from image sequences alone, making it a fully self-supervised learning approach that requires no depth or camera pose annotations. We present extensive experimental results and show that our proposed method can learn NVS with VDEs, outperforming the SOTA single-view NVS methods on the RealEstate10k and MannequinChallenge datasets. Visit our project site11https://kaist-viclab.github.io/monovde-site. Juan Luis Gonzalez 0001, Munchurl Kim |
CVPR | 2 |
| 2024 | From-Ground-To-Objects: Coarse-to-Fine Self-supervised Monocular Depth Estimation of Dynamic Objects with Ground Contact PriorabstractSelf-supervised monocular depth estimation (DE) is an approach to learning depth without costly depth ground truths. However, it often struggles with moving objects that violate the static scene assumption during training. To address this issue, we introduce a coarse-to-fine training strategy leveraging the ground contacting prior based on the observation that most moving objects in outdoor scenes contact the ground. In the coarse training stage, we exclude the objects in dynamic classes from the reprojection loss calculation to avoid inaccurate depth learning. To provide precise supervision on the depth of the objects, we present a novel Ground-contacting-prior Disparity Smoothness Loss (GDS-Loss) that encourages a DE network to align the depth of the objects with their ground-contacting points. Subsequently, in the fine training stage, we refine the DE network to learn the detailed depth of the objects from the reprojection loss, while ensuring accurate DE on the moving object regions by employing our regularization loss with a cost-volume-based weighting factor. Our overall coarse-to-fine training strategy can easily be integrated with existing DE methods without any modifications, significantly enhancing DE performance on challenging Cityscapes and KITTI datasets, especially in the moving object regions. Jaeho Moon, Juan Luis Gonzalez 0001, Byeongjun Kwon, Munchurl Kim |
CVPR | 4 |
| 2024 | FMA-Net: Flow-Guided Dynamic Filtering and Iterative Feature Refinement with Multi-Attention for Joint Video Super-Resolution and DeblurringabstractWe present a joint learning scheme of video super-resolution and deblurring, called VSRDB, to restore clean high-resolution (HR) videos from blurry low-resolution (LR) ones. This joint restoration problem has drawn much less attention compared to single restoration problems. In this paper, we propose a novel flow-guided dynamic filtering (FGDF) and iterative feature refinement with multi-attention (FRMA), which constitutes our VSRDB framework, denoted as FMA-Net. Specifically, our proposed FGDF enables precise estimation of both spatiotemporally-variant degradation and restoration kernels that are aware of motion trajectories through sophisticated motion representation learning. Compared to conventional dynamic filtering, the FGDF enables the FMA-Net to effectively handle large motions into the VSRDB. Additionally, the stacked FRMA blocks trained with our novel temporal anchor (TA) loss, which temporally anchors and sharpens features, refine features in a coarse-to-fine manner through iterative updates. Extensive experiments demonstrate the superiority of the proposed FMA-Net over state-of-the-art methods in terms of both quantitative and qualitative quality. Codes and pretrained models are available at: https://kaist-viclab.github.io/fmanetsite. Geunhyuk Youk, Jihyong Oh, Munchurl Kim |
CVPR | 3 |
| 2024 | SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
Jeonghyeok Do, Munchurl Kim |
ECCV (41) | 2 |
| 2024 | MacDC: Masking-augmented Collaborative Domain Congregation for Multi-target Domain Adaptation in Semantic SegmentationabstractThis paper addresses the challenges in multi-target domain adaptive (MTDA) for semantic segmentation, aiming to learn a single model capable of adapting to multi-target domains. Existing methods solely focus on visual appearance (style) discrepancies, overlooking contextual variations across multi-target domains, resulting in limited performance. We propose a novel approach termed Masking-augmented Collaborative Domain Congregation (MacDC) to handle both style gap and contextual gap among multi-target domains. MacDC achieves this goal by generating image-level and region-level intermediate domains among multi-target domains. To further strengthen contextual alignment, MacDC applies multi-context masking that enforces the model’s understanding of diverse contexts. Notably, MacDC directly learns a single model for multi-target domain adaptation, significantly reducing training times and model parameters. Despite its simplicity, MacDC demonstrates superior performance compared to state-of-the-art MTDA segmentation methods on the syn-to-real and real-to-real benchmarks. Xu Yin, Chenshuang Zhang, Munchurl Kim |
IV | 5 |
| 2024 | Segmentation-Guided Context Learning Using EO Object Labels for Stable SAR-to-EO TranslationabstractRecently, the analysis and use of synthetic aperture radar (SAR) imagery have become crucial for surveillance, military operations, and environmental monitoring. A common challenge with SAR images is the presence of speckle noise, which can hinder their interpretability. To enhance the clarity of SAR images, this letter introduces a novel SAR-to-electro-optical (EO) image translation (SET) network, called SGCL-SET, which first incorporates EO object label information for stable translation. We use a pretrained segmentation network to provide the segmentation regions with their labels into learning the SET. Our SGCL-SET can be trained to effectively learn the translation for the regions of confusing contexts using the segmentation and label information. Through comprehensive experiments on our KOMPSAT dataset, our SGCL-SET significantly outperforms all the previous methods with large margins across nine image quality evaluation metrics. Jaehyup Lee, Hyun-Ho Kim, Munchurl Kim |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Deep Spectral Blending Network for Color Bleeding Reduction in PAN-Sharpening ImagesabstractHigh-resolution (HR) satellites generally transmit multispectral (MS) images at a lower resolution than that of panchromatic (PAN) images. However, satellite image users often prefer MS images to have the same resolution as the corresponding PAN images. Therefore, PAN-sharpening (PS), a technique for obtaining HRMS images by utilizing low-resolution (LR) MS images and HRPAN images, has been a subject of study for several decades. Nevertheless, in most PS methods, various considerations are often ignored, including disparities in physical sensor locations, sensor distortions, geometric variations among acquired images, and registration errors. Owing to these missed factors, increasing the resolution by generating PS images from MS images results in increased registration errors, leading to color bleeding. Furthermore, when obtaining PS images from LRMS images, interpolation of spectral information can lead to image blurring. To address these issues, we propose a novel spectral blending network (SBN) that incorporates spectral alignment blocks (SABs) and a half-instance and half-attention block (HHB) to alleviate both color bleeding and registration errors, producing high-quality PS images with low complexity, respectively. Our SBN achieves superior performance with 1.40~4.55 dB higher peak signal-to-noise ratio (PSNR) for KOMPSAT-3A data and 1.11~3.27 dB higher PSNR for WorldView-III data, as well as with significantly lower computational complexity 42.5~99.3% lower floating point operations per second (FLOPs) than other state-of-the-art methods. Hyun-Ho Kim, Munchurl Kim |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Self-Supervised Monocular Depth Estimation With Positional Shift Depth Variance and Adaptive Disparity QuantizationabstractRecently, attempts to learn the underlying 3D structures of a scene from monocular videos in a fully self-supervised fashion have drawn much attention. One of the most challenging aspects of this task is to handle independently moving objects as they break the rigid-scene assumption. In this paper, we show for the first time that pixel positional information can be exploited to learn SVDE (Single View Depth Estimation) from videos. The proposed moving object (MO) masks, which are induced by the depth variance to shifted positional information (SPI) and are referred to as 'SPIMO' masks, are highly robust and consistently remove independently moving objects from the scenes, allowing for robust and consistent learning of SVDE from videos. Additionally, we introduce a new adaptive quantization scheme that assigns the best per-pixel quantization curve for depth discretization, improving the fine granularity and accuracy of the final aggregated depth maps. Finally, we employ existing boosting techniques in a new way that self-supervises moving object depths further. With these features, our pipeline is robust against moving objects and generalizes well to high-resolution images, even when trained with small patches, yielding state-of-the-art (SOTA) results with four- to eight-fold fewer parameters than the previous SOTA techniques that learn from videos. We present extensive experiments on KITTI and CityScapes that show the effectiveness of our method. Juan Luis Gonzalez 0001, Jaeho Moon, Munchurl Kim |
IEEE Trans. Image Process. | 3 |
| 2024 | A VVC Intra Rate Control With Small Bit Fluctuations Using a Lagrange Multiplier AdjustmentabstractSince the emergence of high-quality multimedia processing applications such as video streaming, digital editing, archiving, etc. these days, an intra coding rate control (RC) is becoming an indispensable and important technology. In this paper, a frame-level intra RC scheme for Versatile Video Coding (VVC) using a Lagrange multiplier adjustment (LMA) is proposed. The VVC test model (VTM) uses an R-λ model-based rate control. However, the estimation performance of target bits based on an R-λ-QP relation is decreased because the distortion dependencies among consecutive frames are not considered especially for intra RC. Thus, in a rate-distortion optimization (RDO) based encoding, the λ values determined for given quantization parameter (QP) values should be elaborately controlled to increase the target bits estimation performance. In our work, we focus on the intra RC scheme by taking advantage of particle-filtering-based prediction (PFP) for distortion estimates, and precise per-frame λ values can be derived for an appropriate RDO process that can lead to small bit-fluctuations. Our extensive experimental results demonstrate that our RC scheme using the per-frame LMA is superior to the default RC (VTM-16.0rc1) method and the state-of-the-art RC methods withsignificantmargins of average 15.57%, 15.31% and 31.13% improvements in terms of the normalized root mean square error (NRMSE) for All Intra (AI) configuration of VVC, respectively. Myung Han Hyun, Bumshik Lee, Munchurl Kim |
IEEE Trans. Multim. | 3 |
| 2023 | Modernizing Old Photos Using Multiple References via Photorealistic Style TransferabstractThis paper firstly presents old photo modernization using multiple references by performing stylization and enhancement in a unified manner. In order to modernize old photos, we propose a novel multi-reference-based old photo modernization (MROPM) framework consisting of a network MROPM-Net and a novel synthetic data generation scheme. MROPM-Net stylizes old photos using multiple references via photorealistic style transfer (PST) and further enhances the results to produce modern-looking images. Meanwhile, the synthetic data generation scheme trains the network to effectively utilize multiple references to perform modernization. To evaluate the performance, we propose a new old photos benchmark dataset (CHD) consisting of diverse natural indoor and outdoor scenes. Extensive experiments show that the proposed method outperforms other baselines in performing modernization on real old photos, even though no old photos were used during training. Moreover, our method can appropriately select styles from multiple references for each semantic region in the old photo to further improve the modernization performance. Agus Gunawan, Soo Ye Kim, Hyeonjun Sim, Munchurl Kim |
CVPR | 5 |
| 2023 | COMPASS: High-Efficiency Deep Image Compression with Arbitrary-scale Spatial ScalabilityabstractRecently, neural network (NN)-based image compression studies have actively been made and has shown impressive performance in comparison to traditional methods. However, most of the works have focused on non-scalable image compression (single-layer coding) while spatially scalable image compression has drawn less attention although it has many applications. In this paper, we propose a novel NN-based spatially scalable image compression method, called COMPASS, which supports arbitrary-scale spatial scalability. Our proposed COMPASS has a very flexible structure where the number of layers and their respective scale factors can be arbitrarily determined during inference. To reduce the spatial redundancy between adjacent layers for arbitrary scale factors, our COMPASS adopts an inter-layer arbitrary scale prediction method, called LIFF, based on implicit neural representation. We propose a combined RD loss function to effectively train multiple layers. Experimental results show that our COMPASS achieves BD-rate gain of -58.33% and -47.17% at maximum compared to SHVC and the state-of-the-art NN-based spatially scalable image compression method, respectively, for various combinations of scale factors. Our COMPASS also shows comparable or even better coding efficiency than the single-layer coding for various scale factors. Jongmin Park 0001, Jooyoung Lee 0004, Munchurl Kim |
ICCV | 3 |
| 2023 | SGSR: A Saliency-Guided Image Super-Resolution NetworkabstractThe human visual system finds salient regions in images and allows the cognitive ability to focus on them. Hence, such salient regions play substantial roles in determining the quality of images. However, the existing super-resolution (SR) methods restore all regions of low-resolution images in the same manner. In this paper, we first propose a saliency-guided image super-resolution (SGSR) network where its restoration ability concentrates on the salient regions in the natural images. For this, we propose a saliency learning scheme using newly computed saliency scores for object regions. Then, by providing the saliency features to the saliency-guided attention (SGA) module and using a novel saliency-weighted loss function, the SGSR maximizes the image quality of salient regions and suppresses the excessive generation of unnecessary structures in backgrounds. To the best of our knowledge, this SGSR is the first attempt to induce discriminatory results guided by saliency in the field of natural image SR. Dayeon Kim, Munchurl Kim |
ICIP | 2 |
| 2023 | Optical-Flow-Aided Self-Supervised Single View Depth Estimation from Monocular VideosabstractRecent advances in deep-learning-based accurate optical flow (OF) estimation have sparked interest in building hardware dedicated to it, making OF a real-time commodity for downstream computer vision tasks. Relative camera pose, and depth estimations (DE) are closely related to OF, but obtaining one from the others is not a trivial task. While previous works have attempted simultaneous learning of OF and DE, or utilized OF-nets only for camera pose prediction and sparse depth map supervision, we propose OF-aided self-supervised DE, that is, pre-computed OF is used as a network input. One of the main challenges in OF-aided DE is preventing the DE network from learning an incorrect OF magnitude prior, which is valid for rigid regions but breaks for scenes with dynamic objects or static camera sequences. We achieve OF-aided DE by instead transforming the input OF into a 3D surface normals space, which provides a well-informed geometrical input while being invariant to the OF magnitude. In combination with a new training strategy, we show that our models with OF-aided DE achieve state-of-the-art (SOTA) results on the KITTI dataset. Juan Luis Gonzalez 0001, Munchurl Kim |
VCIP | 2 |
| 2023 | MorphVAD: Efficient Video Anomaly Detection Using Morphological TransformationabstractVideo anomaly detection aims to identify unusual events within video sequences. Approaches that rely on reconstruction and prediction have shown impressive success in this area and continue to be actively explored. However, recent techniques often rely on pre-trained external components such as optical flow or segmentation maps to handle the challenge of removing background that constitutes a substantial portion of video content. These additional modules, due to their computational complexity, introduce a noteworthy trade-off, particularly in real-time video anomaly detection, where timely monitoring is paramount. So, this paper introduces an efficient video anomaly detection method (MorphVAD) that is based on a novel morphology-based masking module (Morph-Mask). The Morph-Mask harnesses the straightforward concept of morphological transformations to create masks that highlight foreground elements. So, our MorphVAD strategically incorporates these masks during the training phase, focusing on retaining only foreground-related information in memory. Through illustrative experiments and various evaluations, we demonstrate the efficiency and effectiveness of these masks detection, showing the significant enhancements in video anomaly detection performance. Jeonghyeok Do, Munchurl Kim |
VCIP | 2 |
| 2023 | CFCA-SET: Coarse-to-Fine Context-Aware SAR-to-EO Translation With Auxiliary Learning of SAR-to-NIR TranslationabstractSatellite Synthetic Aperture Radar (SAR) images are immensely valuable because they can be obtained regardless of weather and time conditions. However, SAR images have fatal noise and less contextual information, thus making it harder and less interpretable. So, translation of SAR to Electro-Optical (EO) images is highly required for easier interpretation. In this paper, we propose a novel coarse-to-fine context-aware SAR-to-EO image translation (CFCA-SET) framework and a misalignment-resistant loss for the misaligned pairs of SAR-EO images. With our auxiliary learning of SAR-to-Near-Infrared translation, CFCA-SET consists of a two-stage training: (i) the low-resolution SAR-to-EO translation is learned in the coarse stage via a local self-attention module that helps diminish the SAR noise, and (ii) the resulting output is used as guidance in the fine stage to generate the SAR colorization of high resolution. Our proposed auxiliary learning of SAR-to-NIR translation can successfully lead CFCA-SET to learn distinguishable characteristics of various SAR objects with less confusion in a context-aware manner. To handle the inevitable misalignment problem between SAR and EO images, we newly design a misalignment-resistant loss function. Extensive experimental results show that our CFCA-SET can generate more recognizable and understandable EO-like images compared to other methods in terms of nine image quality metrics. Our CFCA-SET surpasses the state-of-the-art methods for two (QXS and CASET) datasets with the improvements: PSNR (3.6%, 29%), ERGAS (7.4%, 30%), SSIM (15%, 15%), SAM (21%, 38%), Ds (16%, 13%), QNR (1.5%, 3.1%), CHD (18%, 12%), LPIPS (4.2%, 8%), and FID (9.0%, 33%). Jaehyup Lee, Hyebin Cho, Hyun-Ho Kim, Munchurl Kim |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Transformer-Based Synthetic-to-Measured SAR Image Translation via Learning of Representational FeaturesabstractDeep-learning-based target recognition in synthetic aperture radar (SAR) images has been actively studied in recent years. However, it is very costly to collect large numbers of labeled SAR images, especially measured SAR target images of various classes, to train high-performance classification networks. To solve the problem of insufficient SAR data, electromagnetic computational tools have often been developed and used to synthesize the measured SAR target images from data modeling. However, despite the use of sophisticated SAR image modeling, there is a large domain gap between synthetic SAR images and measured images such that networks trained with synthetic SAR images tend to show poor classification performance when tested on measured SAR target images. In this paper, we propose a novel transformer-based synthetic-to-measured SAR target image translation network, referred to as SAR-SMT Net, to bridge the gap between synthetic and measured SAR target images. SAR-SMT Net takes synthetic SAR target images as input and estimates the latent representational features of their corresponding measured SAR images to faithfully adjust the global context and scattering characteristics of the input synthetic SAR target images to the corresponding measured SAR values. In addition, we propose five challenging experimental scenarios that can validate SAR image translation performance outcomes. Experimentally, SAR-SMT Net as proposed here outperforms previous state-of-the-art methods in the experiment scenarios, demonstrating feasible generalization ability when used to translate synthetic SAR target images into their corresponding measured SAR target images with a high level of fidelity, even for unseen target classes at unseen azimuth angles. Geunhyuk Youk, Munchurl Kim |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Layered Depth Refinement with Mask GuidanceabstractDepth maps are used in a wide range of applications from 3D rendering to 2D image effects such as Bokeh. However, those predicted by single image depth estimation (SIDE) models often fail to capture isolated holes in objects and/or have inaccurate boundary regions. Meanwhile, high-quality masks are much easier to obtain, using commercial auto-masking tools or off-the-shelf methods of segmentation and matting or even by manual editing. Hence, in this paper, we formulate a novel problem of mask-guided depth refinement that utilizes a generic mask to refine the depth prediction of SIDE models. Our framework performs layered refinement and inpainting/outpainting, decomposing the depth map into two separate layers signified by the mask and the inverse mask. As datasets with both depth and mask annotations are scarce, we propose a self-supervised learning scheme that uses arbitrary masks and RGB-D datasets. We empirically show that our method is robust to different types of masks and initial depth predictions, accurately refining depth values in inner and outer mask boundary regions. We further analyze our model with an ablation study and demonstrate results on real applications. More information can be found on our project page.11https://sooyekim.github.io/MaskDepth/ Soo Ye Kim, Jianming Zhang 0001, Simon Niklaus, Simon Chen, Zhe Lin 0001, Munchurl Kim |
CVPR | 7 |
| 2022 | DeMFI: Deep Joint Deblurring and Multi-frame Interpolation with Flow-Guided Attentive Correlation and Recursive Boosting
Jihyong Oh, Munchurl Kim |
ECCV (7) | 2 |
| 2022 | Selective compression learning of latent representations for variable-rate image compressionabstractRecently, many neural network-based image compression methods have shown promising results superior to the existing tool-based conventional codecs. However, most of them are often trained as separate models for different target bit rates, thus increasing the model complexity. Therefore, several studies have been conducted for learned compression that supports variable rates with single models, but they require additional network modules, layers, or inputs that often lead to complexity overhead, or do not provide sufficient coding efficiency. In this paper, we firstly propose a selective compression method that partially encodes the latent representations in a fully generalized manner for deep learning-based variable-rate image compression. The proposed method adaptively determines essential representation elements for compression of different target quality levels. For this, we first generate a 3D importance map as the nature of input content to represent the underlying importance of the representation elements. The 3D importance map is then adjusted for different target quality levels using importance adjustment curves. The adjusted 3D importance map is finally converted into a 3D binary mask to determine the essential representation elements for compression. The proposed method can be easily integrated with the existing compression models with a negligible amount of overhead increase. Our method can also enable continuously variable-rate compression via simple interpolation of the importance adjustment curves among different quality levels. The extensive experimental results show that the proposed method can achieve comparable compression efficiency as those of the separately trained reference compression models and can reduce decoding time owing to the selective compression. Jooyoung Lee 0004, Seyoon Jeong, Munchurl Kim |
NeurIPS | 3 |
| 2022 | Self-Supervised Deep Monocular Depth Estimation With Ambiguity BoostingabstractWe propose a novel two-stage training strategy with ambiguity boosting for the self-supervised learning of single view depths from stereo images. Our proposed two-stage learning strategy first aims to obtain a coarse depth prior by training an auto-encoder network for a stereoscopic view synthesis task. This prior knowledge is then boosted and used to self-supervise the model in the second stage of training in our novel ambiguity boosting loss. Our ambiguity boosting loss is a confidence-guided type of data augmentation loss that improves the accuracy and consistency of generated depth maps under several transformations of the single-image input. To show the benefits of the proposed two-stage training strategy with boosting, our two previous depth estimation (DE) networks, one with t-shaped adaptive kernels and the other with exponential disparity volumes, are extended with our new learning strategy, referred to as DBoosterNet-t and DBoosterNet-e, respectively. Our self-supervised DBoosterNets are competitive, and in some cases even better, compared to the most recent supervised SOTA methods, and are remarkably superior to the previous self-supervised methods for monocular DE on the challenging KITTI dataset. We present intensive experimental results, showing the efficacy of our method for the self-supervised monocular DE task. Juan Luis Gonzalez 0001, Munchurl Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single-View Depth Estimation With Neural Positional Encoding and Distilled Matting LossabstractIn this paper, we propose a self-supervised singleview pixel-level accurate depth estimation network, called PLADE-Net. The PLADE-Net is the first work that shows remarkable accuracy levels, exceeding 95% in terms of the δ1metric on the challenging KITTI dataset. Our PLADENet is based on a new network architecture with neural positional encoding and a novel loss function that borrows from the closed-form solution of the matting Laplacian to learn pixel-level accurate depth estimation from stereo images. Neural positional encoding allows our PLADENet to obtain more consistent depth estimates by letting the network reason about location-specific image properties such as projection (and potentially lens) distortions. Our novel distilled matting Laplacian loss allows our network to predict sharp depths at object boundaries and more consistent depths in highly homogeneous regions. Our proposed method outperforms all previous self-supervised single-view depth estimation methods by a large margin on the challenging KITTI dataset, with unparalleled levels of accuracy. Furthermore, our PLADE-Net, naively extended for stereo inputs, outperforms the most recent self-supervised stereo methods, even without any advanced blocks like 1D correlations, 3D convolutions, or spatial pyramid pooling. We present extensive ablation studies and experiments that support our method’s effectiveness on the KITTI, CityScapes, and Make3D datasets. Juan Luis Gonzalez 0001, Munchurl Kim |
CVPR | 2 |
| 2021 | KOALAnet: Blind Super-Resolution Using Kernel-Oriented Adaptive Local AdjustmentabstractBlind super-resolution (SR) methods aim to generate a high quality high resolution image from a low resolution image containing unknown degradations. However, natural images contain various types and amounts of blur: some may be due to the inherent degradation characteristics of the camera, but some may even be intentional, for aesthetic purposes (e.g. Bokeh effect). In the case of the latter, it becomes highly difficult for SR methods to disentangle the blur to remove, and that to leave as is. In this paper, we propose a novel blind SR framework based on kernel-oriented adaptive local adjustment (KOALA) of SR features, called KOALAnet, which jointly learns spatially-variant degradation and restoration kernels in order to adapt to the spatially-variant blur characteristics in real images. Our KOALAnet outperforms recent blind SR methods for synthesized LR images obtained with randomized degradations, and we further show that the proposed KOALAnet produces the most natural results for artistic photographs with intentional blur, which are not over-sharpened, by effectively handling images mixed with in-focus and out-of-focus areas. Soo Ye Kim, Hyeonjun Sim, Munchurl Kim |
CVPR | 3 |
| 2021 | SIPSA-Net: Shift-Invariant Pan Sharpening With Moving Object Alignment for Satellite ImageryabstractPan-sharpening is a process of merging a high-resolution (HR) panchromatic (PAN) image and its corresponding low-resolution (LR) multi-spectral (MS) image to create an HR-MS and pan-sharpened image. However, due to the different sensors’ locations, characteristics and acquisition time, PAN and MS image pairs often tend to have various amounts of misalignment. Conventional deep-learning-based methods that were trained with such misaligned PAN-MS image pairs suffer from diverse artifacts such as double-edge and blur artifacts in the resultant PAN-sharpened images. In this paper, we propose a novel framework called shift-invariant pan-sharpening with moving object alignment (SIPSA-Net) which is the first method to take into account such large misalignment of moving object regions for PAN sharpening. The SISPA-Net has a feature alignment module (FAM) that can adjust one feature to be aligned to another feature, even between the two different PAN and MS domains. For better alignment in pan-sharpened images, a shift-invariant spectral loss is newly designed, which ignores the inherent misalignment in the original MS input, thereby having the same effect as optimizing the spectral loss with a well-aligned MS image. Extensive experimental results show that our SIPSA-Net can generate pan-sharpened images with remarkable improvements in terms of visual quality and alignment, compared to the state-of-the-art methods. Jaehyup Lee, Soomin Seo, Munchurl Kim |
CVPR | 3 |
| 2021 | XVFI: eXtreme Video Frame InterpolationabstractIn this paper, we firstly present a dataset (X4K1000FPS) of 4K videos of 1000 fps with the extreme motion to the research community for video frame interpolation (VFI), and propose an extreme VFI network, called XVFI-Net, that first handles the VFI for 4K videos with large motion. The XVFI-Net is based on a recursive multi-scale shared structure that consists of two cascaded modules for bidirectional optical flow learning between two input frames (BiOF-I) and for bidirectional optical flow learning from target to input frames (BiOF-T). The optical flows are stably approximated by a complementary flow reversal (CFR) proposed in BiOF-T module. During inference, the BiOF-I module can start at any scale of input while the BiOF-T module only operates at the original input scale so that the inference can be accelerated while maintaining highly accurate VFI performance. Extensive experimental results show that our XVFI-Net can successfully capture the essential information of objects with extremely large motions and complex textures while the state-of-the-art methods exhibit poor performance. Furthermore, our XVFI-Net framework also performs comparably on the previous lower resolution benchmark dataset, which shows a robustness of our algorithm as well. All source codes, pre-trained models, and proposed X4K1000FPS datasets are publicly available at https://github.com/JihyongOh/XVFI. Hyeonjun Sim, Jihyong Oh, Munchurl Kim |
ICCV | 3 |
| 2021 | A Novel Rate and Distortion Estimation Method Using Particle Filtering Based Prediction for Intra-Predictive Coding of Deep Block Partitioning StructuresabstractIn this paper, we propose a new R/D estimation method for intra-predictive coding with deep block partitioning structures. In our proposed R/D prediction, we adopt a particle filtering based prediction (PFP) to precisely predict intermediate R/D estimates for the next frame in a stochastic manner, which helps increasing the prediction accuracy of fast changing R/D values. Then, based on the intermediate R/D estimates by PFP, we infer an optimal model parameter of the TC's probability density function (pdf) via convex optimization. We found that the proposed method brings about more stable R/D estimation performance thanks to both the improved prediction accuracy using the PFP for abrupt changes in true R/D values and the precise estimation of the optimal model parameter. Experimental results show that our method significantly reduces the normalized root mean square error from average 3.17 to 0.79 (74.90% reduction) for rate and from average 2.32 to 0.82 (64.61% reduction) for distortion, compared to the state-of-the art method. Myung Han Hyun, Bumshik Lee, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | SPAM-Net: A CNN-Based SAR Target Recognition Network With Pose Angle Marginalization LearningabstractRecently, deep convolutional neural networks (CNNs) have started to be applied for automatic target recognition (ATR) problems of synthetic aperture radar (SAR) images. In conventional SAR-ATR algorithms, the pose angle information of the target has been importantly used. However, recent deep learning-based SAR-ATR algorithms often only utilize the intensity information. In this paper, based on the prior works that the pose angle is an important latent variable for boosting target recognition performance, we propose a CNN-based SAR target recognition network with pose angle marginalization learning, called SPAM-Net that marginalizes the conditional probabilities of SAR targets over their pose angles to precisely estimate the true class probabilities. The proposed SPAM-Net consists of two sub-nets: (i) a sub-net for class-conditional probability estimation, called CP sub-net, and (ii) a sub-net for pose angle probability estimation, called PP sub-net. The two sub-nets are jointly learned via an end-to-end manner in a Bayesian framework so that the SPAM-Net incorporates the pose angle information into target recognition task effectively. The SPAM-Net outperforms our baseline network that does not utilize the pose angle information. In the experiments, we intensively analyze the effectiveness of pose angle information for SAR-ATR, revealing that more accurate pose angle information helps the SPAM-Net precisely estimate target classes for the misclassified target groups that are obtained by the baseline network. Furthermore, our method also outperforms the other state-of-the-art SAR-ATR algorithms, yielding the correct target recognition rate with average 99.61%. Jihyong Oh, Gwang-Young Youm, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | A Novel Just-Noticeable-Difference-Based Saliency-Channel Attention Residual Network for Full-Reference Image Quality PredictionsabstractRecently, due to the strength of deep convolutional neural networks (CNN), many CNN-based image quality assessment (IQA) models have been studied. However, previous CNN-based IQA models likely have yet to utilize the characteristics of the human visual system (HVS) fully for IQA problems when they simply entrust everything to the CNN, expecting it to learn from a training dataset. Therefore, the performance capabilities of such deep-learning-based methods are somewhat saturated. However, in this article, we propose a novel saliency-channel attention residual network based on the just-noticeable-difference (JND) concept for full-reference image quality assessments (FR-IQA). It is referred to as JND-SalCAR and shows significant improvements in large IQA datasets with various types of distortion. The proposed JND-SalCAR effectively learns how to incorporate human psychophysical characteristics, such as visual saliency and JND, into image quality predictions. In the proposed network, a SalCAR block is devised so that perceptually important features can be extracted with the help of saliency-based spatial attention and channel attention schemes. In addition, a saliency map serves as a guideline for predicting a patch weight map in order to afford stable training of end-to-end optimization for the JND-SalCAR. To the best of our knowledge, our work presents the first HVS-inspired trainable FR-IQA network that considers both visual saliency and the JND characteristics of the HVS. When the visual saliency map and the JND probability map are explicitly given as priors, they can be usefully combined to predict IQA scores rated by humans more precisely, eventually leading to performance improvements and faster convergence. The experimental results show that the proposed JND-SalCAR significantly outperforms all recent state-of-the-art FR-IQA methods on large IQA datasets in terms of the Spearman rank order coefficient (SRCC) and the Pearson linear correlation coefficient (PLCC). Soomin Seo, Sehwan Ki, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | FISR: Deep Joint Frame Interpolation and Super-Resolution with a Multi-Scale Temporal LossabstractSuper-resolution (SR) has been widely used to convert low-resolution legacy videos to high-resolution (HR) ones, to suit the increasing resolution of displays (e.g. UHD TVs). However, it becomes easier for humans to notice motion artifacts (e.g. motion judder) in HR videos being rendered on larger-sized display devices. Thus, broadcasting standards support higher frame rates for UHD (Ultra High Definition) videos (4K@60 fps, 8K@120 fps), meaning that applying SR only is insufficient to produce genuine high quality videos. Hence, to up-convert legacy videos for realistic applications, not only SR but also video frame interpolation (VFI) is necessitated. In this paper, we first propose a joint VFI-SR framework for up-scaling the spatio-temporal resolution of videos from 2K 30 fps to 4K 60 fps. For this, we propose a novel training scheme with a multi-scale temporal loss that imposes temporal regularization on the input video sequence, which can be applied to any general video-related task. The proposed structure is analyzed in depth with extensive experiments. Soo Ye Kim, Jihyong Oh, Munchurl Kim |
AAAI | 3 |
| 2020 | JSI-GAN: GAN-Based Joint Super-Resolution and Inverse Tone-Mapping with Pixel-Wise Task-Specific Filters for UHD HDR VideoabstractJoint learning of super-resolution (SR) and inverse tone-mapping (ITM) has been explored recently, to convert legacy low resolution (LR) standard dynamic range (SDR) videos to high resolution (HR) high dynamic range (HDR) videos for the growing need of UHD HDR TV/broadcasting applications. However, previous CNN-based methods directly reconstruct the HR HDR frames from LR SDR frames, and are only trained with a simple L2 loss. In this paper, we take a divide-and-conquer approach in designing a novel GAN-based joint SR-ITM network, called JSI-GAN, which is composed of three task-specific subnets: an image reconstruction subnet, a detail restoration (DR) subnet and a local contrast enhancement (LCE) subnet. We delicately design these subnets so that they are appropriately trained for the intended purpose, learning a pair of pixel-wise 1D separable filters via the DR subnet for detail restoration and a pixel-wise 2D local filter by the LCE subnet for contrast enhancement. Moreover, to train the JSI-GAN effectively, we propose a novel detail GAN loss alongside the conventional GAN loss, which helps enhancing both local details and contrasts to reconstruct high quality HR HDR results. When all subnets are jointly trained well, the predicted HR HDR results of higher quality are obtained with at least 0.41 dB gain in PSNR over those generated by the previous methods. The official Tensorflow code is available at https://github.com/JihyongOh/JSI-GAN. Soo Ye Kim, Jihyong Oh, Munchurl Kim |
AAAI | 3 |
| 2020 | Why Are Deep Representations Good Perceptual Quality Features?
Taimoor Tariq, Okan Tarhan Tursun, Munchurl Kim, Piotr Didyk |
ECCV (22) | 3 |
| 2020 | Pan-Sharpening With Color-Aware Perceptual Loss And Guided Re-ColorizationabstractIn remote sensing, “pan-sharpening” is the task of enhancing the spatial resolution of a multi-spectral (MS) image by exploiting the high-frequency information in a panchromatic (PAN) reference image. We present a novel color-aware perceptual (CAP) loss for learning the task of pan-sharpening. Our CAP loss is designed to focus on the deep features of a pre-trained VGG network that are more sensitive to spatial details and ignore color information to allow the network to extract the structural information from the PAN image while keeping the color from the lower resolution MS image. Additionally, we propose “guided re-colorization”, which generates a pan-sharpened image with real colors from the MS input by “picking” the closest MS pixel color for each pan-sharpened pixel, as a human operator would do in manual colorization. Such a re-colorized (RC) image is completely aligned with the pan-sharpened (PS) network output and can be used as a self-supervision signal during training, or to enhance the colors in the PS image during test. We present several experiments where our network trained with our CAP loss generates naturally looking pan-sharpened images with fewer artifacts and outperforms the state-of-the-arts on the WorldView3 dataset in terms of ERGAS, SCC, and QNR metrics. Juan Luis Gonzalez 0001, Soomin Seo, Munchurl Kim |
ICIP | 3 |
| 2020 | Deep 3D Pan via local adaptive "t-shaped" convolutions with global and local adaptive dilations
Juan Luis Gonzalez 0001, Munchurl Kim |
ICLR | 2 |
| 2020 | A CNN-Based Multi-scale Super-Resolution Architecture on FPGA for 4K/8K UHD Applications
Jae-Seok Choi, Jaehyup Lee, Munchurl Kim |
MMM (2) | 4 |
| 2020 | Forget About the LiDAR: Self-Supervised Depth Estimators with MED Probability VolumesabstractSelf-supervised depth estimators have recently shown results comparable to the supervised methods on the challenging single image depth estimation (SIDE) task, by exploiting the geometrical relations between target and reference views in the training data. However, previous methods usually learn forward or backward image synthesis, but not depth estimation, as they cannot effectively neglect occlusions between the target and the reference images. Previous works rely on rigid photometric assumptions or on the SIDE network to infer depth and occlusions, resulting in limited performance. On the other hand, we propose a method to "Forget About the LiDAR" (FAL), with Mirrored Exponential Disparity (MED) probability volumes for the training of monocular depth estimators from stereo images. Our MED representation allows us to obtain geometrically inspired occlusion maps with our novel Mirrored Occlusion Module (MOM), which does not impose a learning burden on our FAL-net. Contrary to the previous methods that learn SIDE from stereo pairs by regressing disparity in the linear space, our FAL-net regresses disparity by binning it into the exponential space, which allows for better detection of distant and nearby objects. We define a two-step training strategy for our FAL-net: It is first trained for view synthesis and then fine-tuned for depth estimation with our MOM. Our FAL-net is remarkably light-weight and outperforms the previous state-of-the-art methods with 8$\times$ fewer parameters and 3$\times$ faster inference speeds on the challenging KITTI dataset. We present extensive experimental results on the KITTI, CityScapes, and Make3D datasets to verify our method's effectiveness. To the authors' best knowledge, the presented method performs the best among all the previous self-supervised methods until now. Juan Luis Gonzalez 0001, Munchurl Kim |
NeurIPS | 2 |
| 2020 | S3: A Spectral-Spatial Structure Loss for Pan-Sharpening NetworksabstractRecently, many deep-learning-based pan-sharpening methods have been proposed for generating high-quality pan-sharpened (PS) satellite images. These methods focused on various types of convolutional neural network (CNN) structures, which were trained by simply minimizing a spectral loss between network outputs and the corresponding high-resolution (HR) multi-spectral (MS) target images. However, owing to different sensor characteristics and acquisition times, HR panchromatic (PAN) and low-resolution MS image pairs tend to have large pixel misalignments, especially for moving objects in the images. Conventional CNNs trained with only the spectral loss with these satellite image data sets often produce PS images of low visual quality including double-edge artifacts along strong edges and ghosting artifacts on moving objects. In this letter, we propose a novel loss function, called a spectral-spatial structure (S3) loss, based on the correlation maps between MS targets and PAN inputs. Our proposed S3 loss can be very effectively used for pan-sharpening with various types of CNN structures, resulting in significant visual improvements on PS images with suppressed artifacts. Jae-Seok Choi, Munchurl Kim |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Learning-Based Low-Complexity Reverse Tone Mapping With Linear MappingabstractAlthough high dynamic range (HDR) display has become popular recently, the legacy content such as standard dynamic range (SDR) video is still in service and needs to be properly converted on HDR display devices. Therefore, it is desirable for HDR TV sets to have the capability of automatically converting input SDR video into HDR video, which is called reverse tone mapping (RTM). In this paper, we propose a novel learning-based low-complexity RTM scheme that not only expands the suppressed dynamic ranges (DR) of the SDR videos (or images), but also effectively restores lost detail in the SDR videos. Most existing conventional RTM schemes have focused on how to expand the DR of global contrast, resulting in limitations in recovering lost detail of SDR videos. On the other hand, the recent convolutional neural network-based approaches show promising results, but they are too complex to be applied on the users' devices in practice. In this paper, our learning-based RTM scheme is computationally simple but effective in recovering lost detail. To learn the SDR-to-HDR relation, training “SDR-HDR” images are first separated into their base layer components and detail layer components by applying a guided filter. The detail layer components of the “SDR-HDR” pairs are used to train the SDR-to-HDR mapping. The mapping matrices are computed based on kernel ridge regression. In the meantime, the global contrast of the base layers is expanded by a nonlinear function that suppresses darker regions and amplifies brighter regions to fit the full DR of a target HDR display. To verify the effectiveness of our learning-based RTM scheme, we performed subjective quality assessment for images and videos. The experimental results show that our RTM scheme outperforms the existing RTM scheme with the successful restoration of lost detail in SDR images. Dae-Eun Kim, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep SR-ITM: Joint Learning of Super-Resolution and Inverse Tone-Mapping for 4K UHD HDR ApplicationsabstractRecent modern displays are now able to render high dynamic range (HDR), high resolution (HR) videos of up to 8K UHD (Ultra High Definition). Consequently, UHD HDR broadcasting and streaming have emerged as high quality premium services. However, due to the lack of original UHD HDR video content, appropriate conversion technologies are urgently needed to transform the legacy low resolution (LR) standard dynamic range (SDR) videos into UHD HDR versions. In this paper, we propose a joint super-resolution (SR) and inverse tone-mapping (ITM) framework, called Deep SR-ITM, which learns the direct mapping from LR SDR video to their HR HDR version. Joint SR and ITM is an intricate task, where high frequency details must be restored for SR, jointly with the local contrast, for ITM. Our network is able to restore fine details by decomposing the input image and focusing on the separate base (low frequency) and detail (high frequency) layers. Moreover, the proposed modulation blocks apply location-variant operations to enhance local contrast. The Deep SR-ITM shows good subjective quality with increased contrast and details, outperforming the previous joint SR-ITM method. Soo Ye Kim, Jihyong Oh, Munchurl Kim |
ICCV | 3 |
| 2019 | A Novel Monocular Disparity Estimation Network with Domain Transformation and Ambiguity LearningabstractConvolutional neural networks (CNN) have shown state-of-the-art results for low-level computer vision problems such as stereo and monocular disparity estimations, but still, have much room to further improve their performance in terms of accuracy, numbers of parameters, etc. Recent works have uncovered the advantages of using an unsupervised scheme to train CNN's to estimate monocular disparity, where only the relatively-easy-to-obtain stereo images are needed for training. We propose a novel encoder-decoder architecture that outperforms previous unsupervised monocular depth estimation networks by (i) taking into account ambiguities, (ii) efficient fusion between encoder and decoder features with rectangular convolutions and (iii) domain transformations between encoder and decoder. Our architecture outperforms the Monodepth baseline in all metrics, even with a considerable reduction of parameters. Furthermore, our architecture is capable of estimating a full disparity map in a single forward pass, whereas the baseline requires two passes. We perform extensive experiments to verify the effectiveness of our method on the KITTI dataset. Juan Luis Gonzalez 0001, Munchurl Kim |
ICIP | 2 |
| 2019 | Video Super-Resolution Based on 3D-CNNS with Consideration of Scene ChangeabstractIn video super-resolution, the spatio-temporal coherence between, and among the frames must be exploited appropriately for the accurate prediction of the high resolution frames. Although 2D-CNNs are powerful in modelling images, 3D-CNNs are more suitable for spatio-temporal feature extraction as they can preserve the temporal information. To this end, we propose an effective 3D-CNN for video super-resolution that does not require motion alignment as preprocessing. The proposed 3DSRnet maintains the temporal depth of spatio-temporal feature maps to maximally capture the temporally nonlinear characteristics between low and high resolution frames, and adopts residual learning in conjunction with the sub-pixel outputs. It outperforms the state-of-the-art method with average 0.45 dB and 0.36 dB higher in PSNR, for scale 3 and 4, in the Vidset4 benchmark. Our 3DSRnet first deals with the performance drop due to scene change, which is important in practice but has not been previously considered. Soo Ye Kim, Jeongyeon Lim, Taeyoung Na, Munchurl Kim |
ICIP | 4 |
| 2019 | A Real-Time Convolutional Neural Network for Super-Resolution on FPGA With Applications to 4K UHD 60 fps Video ServicesabstractIn this paper, we present a novel hardware-friendly super-resolution (SR) method based on a convolutional neural network (CNN) and its dedicated hardware (HW) on field programmable gate array (FPGA). Although CNN-based SR methods have shown very promising results for SR, their computational complexities are prohibitive for hardware implementation. To the best of our knowledge, we are the first to implement a real-time CNN-based SR HW that upscales 2K full high-definition video to 4K ultra high-definition (UHD) video at 60 frames per second (fps). In our dedicated CNN-based SR HW, low-resolution input frames are processed line-by-line, and the number of convolutional filter parameters is reduced significantly by incorporating depth-wise separable convolutions with a residual connection. Our CNN-based SR HW incorporates a cascade of 1D convolutions having large receptive fields along horizontal lines while keeping vertical receptive fields minimal, which allows us to save required line memory space in achieving comparable SR performance against full 2D convolution operations. For efficient HW implementation, we use a simple and effective quantization method with little peak signal-to-noise ratio (PSNR) degradation. Also, we propose a compression method to efficiently store intermediate feature map data to reduce the number of line memories used in HW. Our HW implementation on the FPGA generates 4K UHD frames of higher PSNR values at 60 fps and shows better visual quality, compared with conventional CNN-based SR methods that are trained and tested in software. Jae-Seok Choi, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Fast Computation of Integer DCT-V, DCT-VIII, and DST-VII for Video CodingabstractJoint exploration model (JEM) reference codecs of ISO/IEC and ITU-T utilize multiple types of integer transforms based on DCT and DST of various transform sizes for intra- and inter-predictive coding, which has brought a significant improvement in coding efficiency. JEM adopts three types of integer DCTs (DCT-II, DCT-V, and DCT-VIII), and two types of integer DSTs (DST-I and DST-VII). The fast computations of Integer DCT-II and DST-I are well known, but few studies have been performed for the other types such as DCT-V, DCT-VIII, and DST-VII for all transform sizes. In this paper, we present fast computation methods of N-point DCT-V and DCT-VIII. For this, we first decompose the DCT-VIII into a preprocessing matrix, the DST-VII and a post-processing matrix to quickly compute it by using the linear relation between DCT-VIII and DST-VII. Then, we approximate integer kernels of N = 4, 8, 16, and 32 for DCT-V, DCT-VIII, and DST-VII with norm scaling and bit-shift to be compatible with quantization in each stage of multiplications between decomposed matrices for video coding. In various experiments, the proposed fast computation methods have shown to effectively reduce the total complexity of the matrix operations with little loss in BDBR performance. In particular, our methods reduce the number of addition and multiplication operations by 38% and 80.3%, respectively, in average, compared to the original JEM. Woon-Sung Park, Bumshik Lee, Munchurl Kim |
IEEE Trans. Image Process. | 3 |
| 2018 | Single Image Super-Resolution Using Lightweight CNN with Maxout Units
Jae-Seok Choi, Munchurl Kim |
ACCV (6) | 2 |
| 2018 | A Multi-purpose Convolutional Neural Network for Simultaneous Super-Resolution and High Dynamic Range Image Reconstruction
Soo Ye Kim, Munchurl Kim |
ACCV (3) | 2 |
| 2018 | ITM-CNN: Learning the Inverse Tone Mapping from Low Dynamic Range Video to High Dynamic Range Displays Using Convolutional Neural Networks
Soo Ye Kim, Dae-Eun Kim, Munchurl Kim |
ACCV (3) | 3 |
| 2018 | Learning-Based Just-Noticeable-Quantization- Distortion Modeling for Perceptual Video CodingabstractConventional predictive video coding-based approaches are reaching the limit of their potential coding efficiency improvements, because of severely increasing computation complexity. As an alternative approach, perceptual video coding (PVC) has attempted to achieve high coding efficiency by eliminating perceptual redundancy, using just-noticeable-distortion (JND) directed PVC. The previous JNDs were modeled by adding white Gaussian noise or specific signal patterns into the original images, which were not appropriate in finding JND thresholds due to distortion with energy reduction. In this paper, we present a novel discrete cosine transform-based energy-reduced JND model, called ERJND, that is more suitable for JND-based PVC schemes. Then, the proposed ERJND model is extended to two learning-based just-noticeable-quantization-distortion (JNQD) models as preprocessing that can be applied for perceptual video coding. The two JNQD models can automatically adjust JND levels based on given quantization step sizes. One of the two JNQD models, called LR-JNQD, is based on linear regression and determines the model parameter for JNQD based on extracted handcraft features. The other JNQD model is based on a convolution neural network (CNN), called CNN-JNQD. To our best knowledge, our paper is the first approach to automatically adjust JND levels according to quantization step sizes for preprocessing the input to video encoders. In experiments, both the LR-JNQD and CNN-JNQD models were applied to high efficiency video coding (HEVC) and yielded maximum (average) bitrate reductions of 38.51% (10.38%) and 67.88% (24.91%), respectively, with little subjective video quality degradation, compared with the input without preprocessing applied. Sehwan Ki, Sung-Ho Bae, Munchurl Kim, Hyunsuk Ko |
IEEE Trans. Image Process. | 3 |
| 2018 | Hierarchical Extended Bilateral Motion Estimation-Based Frame Rate Upconversion Using Learning-Based Linear MappingabstractWe present a novel and effective learning-based frame rate upconversion (FRUC) scheme, using linear mapping. The proposed learning-based FRUC scheme consists of: 1) a new hierarchical extended bilateral motion estimation (HEBME) method; 2) a light-weight motion deblur (LWMD) method; and 3) a synthesis-based motion-compensated frame interpolation (S-MCFI) method. First, the HEBME method considerably enhances the accuracy of the motion estimation (ME), which can lead to a significant improvement of the FRUC performance. The proposed HEBME method consists of two ME pyramids with a three-layered hierarchy, where the motion vectors (MVs) are searched in a coarse-to-fine manner via each pyramid. The found MVs are further refined in an enhanced resolution of four times by jointly combining the MVs from the two pyramids. The HEBME method employs a new elaborate matching criterion for precise ME which effectively combines a bilateral absolute difference, an edge variance, pixel variances, and an MV difference among two consecutive blocks and its neighboring blocks. Second, the LWMD method uses the MVs found by the HEBME method and removes the small motion blurs in original frames via transformations by linear mapping. Third, the S-MCFI method finally generates interpolated frames by applying linear mapping kernels for the deblurred original frames. In consequence, our FRUC scheme is capable of precisely generating interpolated frames based on the HEBME for accurate ME, the S-MCFI for elaborate frame interpolation, and the LWMD for contrast enhancement. The experimental results show that our FRUC significantly outperforms the state-of-the-art non-deep learning-based schemes with an average of 1.42 dB higher in the peak signal-to-noise-ratio and shows comparable performance with the state-of-the-art deep learning-based scheme. Sung-Jun Yoon, Hyun-Ho Kim, Munchurl Kim |
IEEE Trans. Image Process. | 3 |
| 2017 | Just-noticeable-quantization-distortion based preprocessing for perceptual video codingabstractConventional predictive video coding may no longer become capable of effectively accommodating the demand of high quality video services with continuously increasing spatiotemporal resolutions as before since it is reaching the limit of its coding efficiency improvement. As an alternative, perceptual video coding (PVC) is being exploited by effectively removing perceptual redundancy for coding efficiency improvement, one of which is just-noticeable-distortion (JND) directed PVC. Unfortunately, the previous JND modeling is not often suitable for JND-directed PVC approaches because quantization effects are not considered. Thus, we presents a new DCT-domain JND model that considers the quantization operation in video compression into JND modeling for PVC, which is calledjust noticeable quantization distortion (JNQD) model. Our proposed JNQD model can be applied as preprocessing prior to any video compression scheme by adding a parameter to adapt the model to quantization step sizes. For experiments, our JNQD models have been applied to High Efficiency Video Coding (HEVC) and yielded the maximum and average bitrate reductions of 37.35% and 13.05%, respectively with little subjective video quality degradation, compared to the input without preprocessing applied. Moreover, it can be applicable for any encoder as preprocessing, which can have a large flexibility compared to previous encoder-dependent schemes. Sehwan Ki, Munchurl Kim, Hyunsuk Ko |
VCIP | 2 |
| 2017 | A DCT-Based Total JND Profile for Spatiotemporal and Foveated Masking EffectsabstractIn image and video processing fields, Discrete Cosine Transform (DCT)-based just-noticeable difference (JND) profiles have effectively been utilized to remove perceptual redundancies in pictures for compression. In this paper, we solve two problems that are often intrinsic to the conventional DCT-based JND profiles: 1) no foveated masking (FM) JND model has been incorporated in modeling the DCT-based JND profiles and 2) the conventional temporal masking (TM) JND models assume that all moving objects in frames can be well tracked by the eyes and that they are projected on the fovea regions of the eyes, which is not a realistic assumption and may result in poor estimation of JND values for untracked moving objects (or image regions). To solve these two problems, we first propose a generalized JND model for joint effects between TM and FM effects. With this model, called the temporal-foveated masking (TFM) JND model, JND thresholds for any tracked/untracked and moving/still image regions can be elaborately estimated. Finally, the TFM-JND model is incorporated into a total DCT-based JND profile with a spatial contrast sensitivity function, luminance masking, and contrast masking JND models. In addition, we propose a JND adjustment method for our total JND profile to avoid overestimation of JND values for image blocks of fixed sizes with various image characteristics. To validate the effectiveness of the total JND profile, an experiment involving a subjective distortion-visibility assessment has been conducted. The experiment results show that the proposed total DCT-based JND profile yields significant performance improvement with much higher capability of distortion concealment (average 5.6-dB lower PSNR) compared with state-of-the-art JND profiles. The MATLAB source code of the proposed total DCT-based JND profile is publicly available online at https://sites.google.com/site/sunghobaecv/jnd. Sung-Ho Bae, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Single Image Super-Resolution Using Global Regression Based on Multiple Local Linear MappingsabstractSuper-resolution (SR) has become more vital, because of its capability to generate high-quality ultra-high definition (UHD) high-resolution (HR) images from low-resolution (LR) input images. Conventional SR methods entail high computational complexity, which makes them difficult to be implemented for up-scaling of full-high-definition input images into UHD-resolution images. Nevertheless, our previous super-interpolation (SI) method showed a good compromise between Peak-Signal-to-Noise Ratio (PSNR) performances and computational complexity. However, since SI only utilizes simple linear mappings, it may fail to precisely reconstruct HR patches with complex texture. In this paper, we present a novel SR method, which inherits the large-to-small patch conversion scheme from SI but uses global regression based on local linear mappings (GLM). Thus, our new SR method is called GLM-SI. In GLM-SI, each LR input patch is divided into 25 overlapped subpatches. Next, based on the local properties of these subpatches, 25 different local linear mappings are applied to the current LR input patch to generate 25 HR patch candidates, which are then regressed into one final HR patch using a global regressor. The local linear mappings are learned cluster-wise in our off-line training phase. The main contribution of this paper is as follows: Previously, linear-mapping-based conventional SR methods, including SI only used one simple yet coarse linear mapping to each patch to reconstruct its HR version. On the contrary, for each LR input patch, our GLM-SI is the first to apply a combination of multiple local linear mappings, where each local linear mapping is found according to local properties of the current LR patch. Therefore, it can better approximate nonlinear LR-to-HR mappings for HR patches with complex texture. Experiment results show that the proposed GLM-SI method outperforms most of the state-of-the-art methods, and shows comparable PSNR performance with much lower computational complexity when compared with a super-resolution method based on convolutional neural nets (SRCNN15). Compared with the previous SI method that is limited with a scale factor of 2, GLM-SI shows superior performance with average 0.79 dB higher in PSNR, and can be used for scale factors of 3 or higher. Jae-Seok Choi, Munchurl Kim |
IEEE Trans. Image Process. | 2 |
| 2016 | Topic-tracking-based dynamic user modeling with TV recommendation applications
Eunhui Kim, Munchurl Kim |
Appl. Intell. | 2 |
| 2016 | A Novel Image Quality Assessment With Globally and Locally Consilient Visual Quality PerceptionabstractComputational models for image quality assessment (IQA) have been developed by exploring effective features that are consistent with the characteristics of a human visual system (HVS) for visual quality perception. In this paper, we first reveal that many existing features used in computational IQA methods can hardly characterize visual quality perception for local image characteristics and various distortion types. To solve this problem, we propose a new IQA method, called the structural contrast-quality index (SC-QI), by adopting a structural contrast index (SCI), which can well characterize local and global visual quality perceptions for various image characteristics with structural-distortion types. In addition to SCI, we devise some other perceptually important features for our SC-QI that can effectively reflect the characteristics of HVS for contrast sensitivity and chrominance component variation. Furthermore, we develop a modified SC-QI, called structural contrast distortion metric (SC-DM), which inherits desirable mathematical properties of valid distance metricability and quasi-convexity. So, it can effectively be used as a distance metric for image quality optimization problems. Extensive experimental results show that both SC-QI and SC-DM can very well characterize the HVS's properties of visual quality perception for local image characteristics and various distortion types, which is a distinctive merit of our methods compared with other IQA methods. As a result, both SC-QI and SC-DM have better performances with a strong consilience of global and local visual quality perception as well as with much lower computation complexity, compared with the state-of-the-art IQA methods. The MATLAB source codes of the proposed SC-QI and SC-DM are publicly available online at https://sites.google.com/site/sunghobaecv/iqa. Sung-Ho Bae, Munchurl Kim |
IEEE Trans. Image Process. | 2 |
| 2016 | DCT-QM: A DCT-Based Quality Degradation Metric for Image Quality Optimization ProblemsabstractRecent development of computational image quality assessment methods has shown to give very promising results in measuring perceptual visual quality for distorted images. However, most of them are difficult to be applied for optimization problems due to the lack of desirable mathematical properties, such as differentiability, convexity, and valid distance metricability. This paper proposes a novel Discrete Cosine Transform (DCT)-based quality degradation metric, called DCT-QM, which is based on the probability summation theory with a psychometric function for neural responses in the receptive fields of visual cortex in psychophysics. Consequently, the DCT-QM is formulated as a weighted mean L2norm in the DCT domain, which is very easy to implement and inherits the three desirable mathematical properties, that is, differentiability, convexity, and valid distance metricability, for image quality optimization problems. The extensive experimental results show that the proposed DCT-QM has promising results for many practical distortion types in image processing problems by showing high consistency with perceived visual quality. Sung-Ho Bae, Munchurl Kim |
IEEE Trans. Image Process. | 2 |
| 2016 | HEVC-Based Perceptually Adaptive Video Coding Using a DCT-Based Local Distortion Detection Probability ModelabstractDiscrete Cosine Transform (DCT)-based just noticeable difference (JND) profiles have widely been applied into human perception-based video coding in order to reduce perceptual redundancy, which is one of the main goals of perceptual video coding (PVC). However, there are two problems for this approach: 1) the JND value of each transform coefficient is estimated for a fixed-sized DCT kernel (e.g., 8 × 8), but flexible coding structures with variable-sized transform units have been utilized in standard video coding frameworks, such high efficiency video coding (HEVC) and 2) the DCT transform coefficients are suppressed by the amounts of JND values for the removal of perceptual redundancy, but the DCT transform coefficients of residues are not sufficiently suppressed due to many small transform coefficient values in mid- and high-frequency regions below the JND values. In order to solve these problems, we propose a more generalized visibility model in the DCT domain, called the DCT-based local distortion detection probability (LDDP) model that can estimate a degree of distortion visibility for any distribution of the transform coefficients of any sized DCT kernel for residues. Furthermore, we propose an HEVC-compliant LDDP-based PVC scheme where transform coefficients are sufficiently suppressed based on the LDDP model. The proposed PVC scheme is implemented in the HEVC Test Model (HM 11.0) reference software to show the effectiveness of the LDDP-based PVC scheme. Objective and subjective tests for encoded test sequences are performed. The experimental results show that the LDDP-based PVC scheme achieves a significant performance improvement of bitrate reduction at the similar visual quality levels compared with the original HM 11.0. Sung-Ho Bae, Jaeil Kim, Munchurl Kim |
IEEE Trans. Image Process. | 3 |
| 2016 | Super-Interpolation With Edge-Orientation-Based Mapping Kernels for Low Complex 2× UpscalingabstractWith the advent of ultrahigh-definition (UHD) video services, super-resolution (SR) techniques are often required to generate high-resolution (HR) images from low-resolution (LR) images, such as HD images. To generate such HR images and a video of UHD resolutions in limited computing devices with hardware and software, low complex but excellent SR methods are particularly required. In this paper, we present a novel and fast SR method, called super-interpolation (SI), by unifying an interpolation step and a quality-enhancement step. The proposed SI method utilizes edge-orientation (EO)-based pre-learned kernels, which inherits the simplicity of interpolation and the quality enhancement of SR. It performs SR directly from the initial resolution of an input image to the target resolution of an up-scaled output image without requiring any intermediate interpolated image. The proposed SI method involves offline training and online up-scaling phases. In the offline training phase, training LR image patches are clustered based on their edge orientations into different EO classes for which class-dependent linear mapping functions are learned between training LR and HR image patches. In up-scaling phase, an HR output image patch for each LR input image patch is generated by applying an appropriate linear mapping function selected based on the EO of LR input image patch. Our proposed SI method is intensively compared with the ten state-of-the-art SR methods for common image sets and many HD/UHD images. The experimental results show that the SI method yields the smallest running time and requires relatively small hardware resources. It outperforms the six state-of-the-art methods in average (peak signal-to-noise ratio) PSNR/(structural similarity) SSIM, and exhibits competitive or somewhat lower PSNR/SSIM performance compared with the others. Jae-Seok Choi, Munchurl Kim |
IEEE Trans. Image Process. | 2 |
| 2016 | A CU-Level Rate and Distortion Estimation Scheme for RDO of Hardware-Friendly HEVC Encoders Using Low-Complexity Integer DCTsabstractIn this paper, a low complexity coding unit (CU)-level rate and distortion estimation scheme is proposed for High Efficiency Video Coding (HEVC) hardware-friendly implementation where a Walsh-Hadamard transform (WHT)-based low-complexity integer discrete cosine transform (DCT) is employed for distortion estimation. Since HEVC adopts quadtree structures of coding blocks with hierarchical coding depths, it becomes more difficult to estimate accurate rate and distortion values without actually performing transform, quantization, inverse transform, de-quantization, and entropy coding. Furthermore, DCT for rate-distortion optimization (RDO) is computationally high, because it requires a number of multiplication and addition operations for various transform block sizes of 4-, 8-, 16-, and 32-orders and requires recursive computations to decide the optimal depths of CU or transform unit. Therefore, full RDO-based encoding is highly complex, especially for low-power implementation of HEVC encoders. In this paper, a rate and distortion estimation scheme is proposed in CU levels based on a low-complexity integer DCT that can be computed in terms of WHT whose coefficients are produced in prediction stages. For rate and distortion estimation in CU levels, two orthogonal matrices of 4×4 and 8×8 , which are applied to WHT that are newly designed in a butterfly structure only with addition and shift operations. By applying the integer DCT based on the WHT and newly designed transforms in each CU block, the texture rate can precisely be estimated after quantization using the number of non-zero quantized coefficients and the distortion can also be precisely estimated in transform domain without de-quantization and inverse transform required. In addition, a non-texture rate estimation is proposed by using a pseudoentropy code to obtain accurate total rate estimates. The proposed rate and the distortion estimation scheme can effectively be used for HW-friendly implementation of HEVC encoders with 9.8% loss over HEVC full RDO, which much less than 20.3% and 30.2% loss of a conventional approach and Hadamard-only scheme, respectively. Bumshik Lee, Munchurl Kim |
IEEE Trans. Image Process. | 2 |
| 2016 | An All-Zero Block Detection Scheme for Low-Complexity HEVC EncodersabstractIn this paper, an all-zero block detection scheme is proposed prior to DCT to reduce the encoding complexity for high efficiency video coding (HEVC). Since many coding blocks tend to have all zero coefficients after DCT and quantization, it is worthwhile to detect all-zero-quantized blocks for input residual blocks before DCT so that subsequent transform and quantization can be skipped. Unlike previous coding standards, HEVC adopts large transform sizes such as 16 × 16 and 8 × 8. It becomes more difficult to accurately detect all-zero blocks in HEVC because the large transform blocks contains more variety of content characteristics than smaller ones, thus making it ineffective the existing all-zero block (AZB) detection schemes for large transform blocks in HEVC. In this paper, a novel AZB detection scheme is proposed for the case that Hadamard transform is used as a distortion metric for RDO in HEVC. Statistical upper bounds to be all-zero blocks are derived using the relationship between Walsh Hadamard and DCT transform kernels. Then, a small number of quantized coefficients in a upper left corner of a transform block, which are obtained using the relations between Hadamard transform and DCT, are examined for AZB detection. For 32 × 32 blocks, DC coefficients of 8 × 8 sub-blocks are further examined for AZB detection. The experimental results demonstrate that the proposed scheme detects 87.79% of actual AZBs with 2.87% false alarm rate in average, outperforming the state-of-the-art method. Computational complexity to detect AZB is almost negligible compared to the conventional method. Bumshik Lee, Jaehong Jung, Munchurl Kim |
IEEE Trans. Multim. | 3 |
| 2015 | Single image super-resolution based on self-examples using context-dependent subpatchesabstractSelf-example-based super-resolution (SR) methods utilize internal dictionaries to reconstruct a high-resolution (HR) image from a single low-resolution (LR) input image. In general, a square-sized patch is used to find the LR-HR correspondences in the dictionaries. However, this may be a difficult issue because the LR input image and the dictionaries are of different scales. Inspired by this observation, we propose a novel self-example-based SR method, using context-dependent multi-shaped subpatches. Each LR input patch is segmented into multiple subpatches according to the context of the patch, enabling us to extract the better LR-HR correspondences. Our experimental results show that the proposed subpatch-based SR generates competitive high-quality HR images compared to state-of-the-art methods, with visually sharper edges that result in better visual quality. Jae-Seok Choi, Sung-Ho Bae, Munchurl Kim |
ICIP | 3 |
| 2015 | A spatial class LDA model for classification of sports scene imagesabstractRecently, the bag-of-visual words (BoW) models have widely been studied in computer vision area. Owing to the limit of the BoW models that only consider the distributions of visual words in images, the Latent Dirichlet Allocation (LDA) model has drawn an attention to discover the structure of the visual word distributions over latent topics which can represent semantic objects in images. In order to reflect the spatial information of images, the LDA model has been extended to so-called a spatial LDA model for image segmentation, which is not applicable for image classification. Therefore, in this paper, we propose a spatial class LDA (scLDA) model for image classification where the topic distributions over visual words are found per image segments and a class-specific-simplex LDA (cssLDA) model is applied for image classification. From our experimental results, the proposed scLDA model outperforms the previous LDA models in terms of correct classification rates. Jin Jeon, Munchurl Kim |
ICIP | 2 |
| 2015 | A novel image quality assessment based on an adaptive feature for image characteristics and distortion typesabstractIn this paper, we reveal that many conventional features used in computational image quality assessment (IQA) methods can hardly characterize perceived distortions on various image characteristics and distortion types, thus resulting in relatively low prediction performance of visual quality scores. To solve this problem, we propose a new IQA method, called Structural Contrast-Quality Index (SC-QI) which is based on structural contrast index (SCI) as a very effective feature. SCI can adaptively quantify perceived distortions depending on various image characteristics and distortions types. In addition to SCI, some other perceptually important features that reflect effects of contrast sensitivity function and chrominance component variation are also combined into the proposed SC-QI. Our comprehensive experiments on three large IQA datasets verify that the proposed SC-QI outperforms the state-of-the-art ones while accompanying lower computational complexity. Sung-Ho Bae, Munchurl Kim |
VCIP | 2 |
| 2015 | A novel SSIM index for image quality assessment using a new luminance adaptation effect model in pixel intensity domainabstractThe Structural SIMimarity (SSIM) is one of the most prominent image quality assessment (IQA) methods due to its high prediction performance and wide applicability for image quality optimization problems. To reflect the luminance adaptation (LA) characteristic of human visual system (HVS), SSIM is modelled to have high consistency with Weber's law. However, it inevitably has some intrinsic faults that wrongly incorporate the LA effect into SSIM. In this paper, we firstly analyze that Weber's law in the conventional SSIM index cannot precisely reflect the LA effect due to two reasons: (i) it is reported that Weber's law model is not able to precisely be fitted in the measured experimental results for the LA effect; (ii) SSIM is calculated with pixel intensity values, but Weber's law is applied for luminance values which have non-linear relations with the pixel intensity values. To solve this problem, we first theoretically derive a new LA effect model in pixel intensity domain using a Gamma correction function and a power-law model. We then devise a weight function for the LA effect, called LA-based local weight function (LALF) which allows the proposed LA effect model to be precisely incorporated into SSIM index. To verify the effectiveness of the proposed LALF-based SSIM, we perform comprehensive experiments on four large IQA databases. Experimental results show that the proposed LALF helps performance improvement of the SSIM index. Sung-Ho Bae, Munchurl Kim |
VCIP | 2 |
| 2015 | A Novel Fast CU Encoding Scheme Based on Spatiotemporal Encoding Parameters for HEVC Inter CodingabstractRecently, a new video coding standard, High Efficiency Video Coding (HEVC), has shown greatly improved coding efficiency by adopting hierarchical structures of coding unit (CU), prediction unit (PU), and transform unit (TU). To best achieve the coding efficiency, the best combinations of CU, PU, and TU must be found in the sense of the minimum rate-distortion (R-D) costs. Owing to this, a large computational complexity occurs. Among these CU, PU, and TU, the determination of CU sizes most significantly affects the R-D performance of HEVC encoders, which causes large computational costs in operation with PU and TU size determinations. In spite of recent works in the complexity reduction of HEVC encoders, most of the research has focused on the complexity reduction with fast CU split in intra slice coding and with early TU split in both intra and inter slice. In this paper, we propose a fast and an efficient CU encoding scheme based on the spatiotemporal encoding parameters of HEVC encoders, which consists of an improved early CU SKIP detection method and a fast CU split decision method. For the current CU block under encoding, the proposed scheme utilizes sample-adaptive-offset parameters as the spatial encoding parameter to estimate the texture complexity that affects the CU partition. In addition, the motion vectors, TU size, and coded block flag information are used as the temporal encoding parameters to estimate the temporal complexity that also affects the CU partition. The proposed scheme effectively utilizes the spatiotemporal encoding parameters that are the byproducts during the encoding process of HEVC without additionally required computation. The proposed novel fast CU encoding scheme significantly reduces the total encoding time with negligible RD-performance loss. The experimental results show that the proposed scheme achieves the total encoding time savings of average 49.6% and 42.7% only with average 1.4% and 1.0% bit-rate losses for various test sequences under random access and low delay B conditions, respectively. The proposed scheme has an advantage on the implementation for parallel processing in pipeline structures of HEVC encoders due to its independency with neighboring CU blocks. Sangsoo Ahn, Bumshik Lee, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Binocular Suppression-Based Stereoscopic Video Coding by Joint Rate Control With KKT Conditions for a Hybrid Video Codec SystemabstractAsymmetric video coding based on binocular suppression provides a prospect for improving coding efficiency since one view in stereoscopic video can be encoded at lower bitrates than the other view without a loss of perceptual video quality. In this paper, we propose a joint rate control scheme for asymmetric stereoscopic video coding in a hybrid video codec system, which supports binocular suppression-based asymmetric stereoscopic video coding in a theoretical basis within an optimization framework. To obtain an optimal solution to quantization steps for joint rate control, we apply the Karush-Kuhn-Tucker (KKT) conditions in minimizing the perceptual distortion of decoded video under the conditions that: 1) the sum of bits generated from two video encoders is constrained within a given bit budget; 2) the quality of a primary view is superior to that of an auxiliary view with an allowable difference between two view qualities; and 3) the distortion of an auxiliary view is less than a given threshold value for high bitrate encoding. As a result, the proposed rate control scheme effectively enhances the perceptual quality based on an optimization framework by taking into account the binocular suppression and also simultaneously controlling the rates of both the left and right views of stereoscopic video under the constraints with a given target bit budget and an allowed quality difference. Experimental results demonstrate that, compared with the independent rate control scheme, the proposed joint rate control scheme with KKT conditions for asymmetric stereoscopic video coding not only demonstrates comparable 3-D perceived visual quality with an overall gain of 0.33 dB, but also attains 2-D visual quality gain by an average of 2.49 dB while accurately satisfying the given constraints in the proposed optimization framework Yongjun Chang, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | An HEVC-Compliant Perceptual Video Coding Scheme Based on JND Models for Variable Block-Sized Transform KernelsabstractIn this paper, a High Efficiency Video Coding (HEVC)-compliant perceptual video coding (PVC) scheme is introduced based on just-noticeable difference (JND) models in both transform and pixel domains. We adopt an existing pixel-domain JND model for the transform skip mode of HEVC and propose a transform-domain JND model for the transform nonskip modes of HEVC. The proposed transform-domain JND model is designed by considering the spatial JND characteristics such as contrast sensitivity, luminance adaptation, and contrast masking effects as well as by considering the summation effects of variable block-sized transforms in HEVC. A temporal JND model is additionally incorporated into the proposed transform-domain JND model to further reduce perceptual redundancy. To incorporate the transform- and pixel-domain JND models into the encoding process in an HEVC-compliant manner, the transform coefficients and residues are suppressed in harmonization with the transform/quantization process and the quantization-only process of HEVC, respectively. To make the JND-based suppression effective, a distortion compensation factor is also proposed to reflect the perceptual distortion in the rate-distortion optimization-based encoding process. Based on subjective quality assessments of the encoded bit streams of test sequences, the proposed HEVC-compliant PVC scheme yields remarkable bitrate reductions of a maximum 49.10% and an average 16.10% with negligible subjective quality loss, compared with an HEVC reference software HEVC test model (HM 11.0). In addition, the proposed HEVC-compliant PVC scheme increases the encoding complexity of HM 11.0 only by an average of 11.25%. Jaeil Kim, Sung-Ho Bae, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | LDA-Based Unified Topic Modeling for Similar TV User Grouping and TV Program RecommendationabstractSocial TV is a social media service via TV and social networks through which TV users exchange their experiences about TV programs that they are viewing. For social TV service, two technical aspects are envisioned: grouping of similar TV users to create social TV communities and recommending TV programs based on group and personal interests for personalizing TV. In this paper, we propose a unified topic model based on grouping of similar TV users and recommending TV programs as a social TV service. The proposed unified topic model employs two latent Dirichlet allocation (LDA) models. One is a topic model of TV users, and the other is a topic model of the description words for viewed TV programs. The two LDA models are then integrated via a topic proportion parameter for TV programs, which enforces the grouping of similar TV users and associated description words for watched TV programs at the same time in a unified topic modeling framework. The unified model identifies the semantic relation between TV user groups and TV program description word groups so that more meaningful TV program recommendations can be made. The unified topic model also overcomes an item ramp-up problem such that new TV programs can be reliably recommended to TV users. Furthermore, from the topic model of TV users, TV users with similar tastes can be grouped as topics, which can then be recommended as social TV communities. To verify our proposed method of unified topic-modeling-based TV user grouping and TV program recommendation for social TV services, in our experiments, we used real TV viewing history data and electronic program guide data from a seven-month period collected by a TV poll agency. The experimental results show that the proposed unified topic model yields an average 81.4% precision for 50 topics in TV program recommendation and its performance is an average of 6.5% higher than that of the topic model of TV users only. For TV user prediction with new TV programs, the average prediction precision was 79.6%. Also, we showed the superiority of our proposed model in terms of both topic modeling performance and recommendation performance compared to two related topic models such as polylingual topic model and bilingual topic model. Shinjee Pyo, Eunhui Kim, Munchurl Kim |
IEEE Trans. Cybern. | 3 |
| 2015 | Efficient In-Loop Filtering Across Tile Boundaries for Multi-Core HEVC Hardware Decoders With 4 K/8 K-UHD Video ApplicationsabstractHEVC is a next generation video coding standard designed with modern coding techniques to be especially efficient for coding high-resolution video such as 4 K/8 K-ultra high- definition (UHD) video. Among the advanced coding tools of HEVC, tiles and wavefront parallel processing (WPP) have been newly adopted for parallel processing of such high-resolution (4 K/8 K-UHD) video. To realize UHD video services over portable devices with limited battery power, it is essential to implement multi-core-based and dedicated HEVC hardware decoders that support the tile- and wavefront-based parallel processing. By doing so, each frame is divided into a multiple number of picture partitions which can then be processed by multiple hardware decoder cores in parallel. However, in-loop filtering (ILF) at tile boundaries cannot be easily parallelized by a multi-core HEVC hardware decoder because of the data dependency between samples in different tiles. In this paper, an efficient control method for ILF across tile boundaries is proposed for multi-core HEVC hardware decoders. The proposed method does not require additional in-loop filters for ILF across the tile boundaries and it allows a decoder core to continue to process the next coding tree unit (CTU) without waiting for other decoders until they finish their ILF processing for the neighboring CTUs in other tiles. From experiments, we show the effectiveness of our ILF control method via a quad-core HEVC decoder for 4 K-UHD video implemented on a prototyping FPGA board. Seunghyun Cho, Hyunmi Kim, Hui Yong Kim, Munchurl Kim |
IEEE Trans. Multim. | 4 |
| 2014 | Object tracking based on online partial instance learning with multiple local strong classifiersabstractIn this paper, we propose a new appearance model based on Partial Instance Learning (PIL) with multiple local strong classifiers. The key idea of PIL is that image examples are divided into several partial image examples (or local-images), each of which is then independently trained with a local strong classifier. Finally, a tracker is updated for the optimal solution in the sense that the joint probability of partial image examples for each input image example becomes the largest. The proposed PIL method can be considered a risk diversification strategy for unpredictable partial occlusions or appearance changes of an object. Also, it can be regarded as a divide-and-conquer method of Online Boosting (OB), so that PIL only requires approximately 20% of computations compared with other OB methods in terms of iterations taken for learning process. Experiment results show that the proposed PIL-based object tracking method achieves better performance in tracking accuracy and much faster processing speed than other compared real-time based ones. Sung-Ho Bae, Munchurl Kim |
ICIP | 2 |
| 2014 | Performance analysis of hierarchical transform coding with a large kernel for video codecsabstractIn this study, the performance of hierarchical transform coding is analysed with design of an order‐16 integer transform kernel. The proposed hierarchical transform‐coding structure is constructed with a set of 4 × 4, 8 × 8 and 16 × 16 integer transforms of variable transform block sizes, which takes the advantages of both lower and higher transform kernels by flexibly adapting to varying image characteristics of video sequences with homogeneous and complex regions. The proposed hierarchical transform‐coding structure is implemented as an extension to H.264/advanced video coding joint model. The authors show the effectiveness of the hierarchical variable‐sized block transform scheme by analysing the quantisation effects and the correlation among neighbouring pixels in video sequences of different spatial resolutions. The experimental results show that: (i) the variable‐sized block transform scheme with the hierarchical structure is advantageous to the texture regions with strong local edges and (ii) the higher‐order‐16 integer transform kernel itself is more effective for the homogeneous texture regions, which are often encountered in higher resolution sequences. Therefore these two features can complementarily work in an rate‐distortion (RD) optimised manner for various characteristics of the input signals. Bumshik Lee, Munchurl Kim, Hui Yong Kim, Jin Soo Choi |
IET Image Process. | 2 |
| 2014 | A Frame-Level Rate Control Scheme Based on Texture and Nontexture Rate Models for High Efficiency Video CodingabstractIn this paper, a frame-level rate control scheme is proposed based on texture and nontexture rate models for High Efficiency Video Coding (HEVC). Due to more complicated coding structures and the adoption of new coding tools, the statistical characteristics of transform residues are significantly different depending on the depth levels of coding units (CUs) from which the residues are obtained. A new texture rate model is constructed for the transform residues, which are categorized into three types of CUs: low-, medium- and high-textured CUs. One single Laplacian probability PDF model is used for each residue category to derive a rate-quantization model. Based on the Laplacian PDF, a simplified rate model for texture bits is derived using entropy. In addition, an analytic rate model for nontexture bits is proposed, which also takes into account the different characteristics of nontexture bits occurring in various depths of CUs in HEVC. The nontexture bitrates are modeled based on the linear relation between the total nontexture data and the dominant nontexture data in each CU category. Based on the proposed rate models for the texture and nontexture bits, accurate rate control can be achieved owing to more precise rate estimation. The experimental results show that the proposed rate control scheme achieves the average PSNR with 0.44 dB higher and the average PSNR standard deviation of 0.32 point lower with the buffer status levels maintained very close to target buffer levels, compared to the conventional methods. Finally, the proposed rate control scheme remarkably outperforms the conventional schemes especially for the sequences of complex texture and large motion. Bumshik Lee, Munchurl Kim, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A Novel No-Reference PSNR Estimation Method With Regard to Deblocking Filtering Effect in H.264/AVC BitstreamsabstractPeak signal-to-noise ratio (PSNR) monitoring is an important application for video quality assessment of video systems at the receiver sides where no-reference PSNR estimation is essential. Most of the PSNR estimation methods for H.264/AVC bitstreams ignore or do not consider the effect of deblocking filtering. Instead, they only focus on estimating the mean squared error (MSE) due to quantization. However, the PSNR estimation affected by deblocking filtering cannot be negligible for sequences of large picture resolutions. In this paper, we first present an MSE estimation method on H.264/AVC bitstreams by considering the deblocking filtering effect so that more accurate PSNR estimation can be made. For this, the total MSE between the original and reconstructed frames is separated into two terms for PSNR estimation: one due to quantization error and the other due to the deblocking filtering effect in H.264/AVC. In the proposed PSNR estimation, the contribution of deblocking filtering to the total MSE is quantified by a compensation factor of each encoded picture type between the original and the deblocked frames. Experimental results show that the proposed method effectively reflects the contribution of deblocking filtering to PSNR estimation, thus yielding more accurate PSNR estimates. Taeyoung Na, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A Novel Generalized DCT-Based JND Profile Based on an Elaborate CM-JND Model for Variable Block-Sized Transforms in Monochrome ImagesabstractIn this paper, we propose a new DCT-based just noticeable difference (JND) profile incorporating the spatial contrast sensitivity function, the luminance adaptation effect, and the contrast masking (CM) effect. The proposed JND profile overcomes two limitations of conventional JND profiles: 1) the CM JND models in the conventional JND profiles employed simple texture complexity metrics, which are not often highly correlated with perceived complexity, especially for unstructured patterns. So, we proposed a new texture complexity metric that considers not only contrast intensity, but also structureness of image patterns, called the structural contrast index. We also newly found out that, as the structural contrast index of a background texture pattern increases, the modulation factors for CM-JND show a bandpass property in frequency. Based on this observation, a new CM-JND is modeled as a function of DCT frequency and the proposed structural contrast index, showing significantly high correlations with measured CM-JND values and 2) while the conventional DCT-based JND profiles are only applicable for specific transform block sizes, our proposed DCT-based JND profile is first designed to be applicable to any size of transform by deriving a new summation effect function, which can also be appropriately applied for quad-tree transform of high efficiency video coding. For the overall performance, the proposed DCT-based JND profile shows more tolerance for distortions with better perceptual quality than other JND profiles under comparison. Sung-Ho Bae, Munchurl Kim |
IEEE Trans. Image Process. | 2 |
| 2013 | Fast decision of CU partitioning based on SAO parameter, motion and PU/TU split information for HEVCabstractHigh Efficiency Video Coding (HEVC) has recently been standardized with a significant improvement of coding efficiency compared to its preceding video coding standards. To achieve high coding efficiency improvements, HEVC adopts deeper hierarchical block coding structures of coding unit (CU), prediction unit (PU) and transform unit (TU), to better adapt the complex texture and motion natures of various video sequences. However, this causes a dramatically increased computational and structural complexity of HEVC encoders and decoders. In this paper, we propose a fast decision method of CU partitioning, which significantly reduces total encoding time with negligible RD-performance loss for HEVC. For this, the proposed method effectively utilizes the available side information such as sample adaptive offset (SAO) parameter values, PU sizes, MV sizes and coded block flag (cbf) data, so that the required computational complexity for fast CU split decision is minimized. Our experiment results show that the proposed method reduces the total encoding time to average 43.2% only with average 1.58% bit-rate increase for ten test sequences of HEVC. Sangsoo Ahn, Munchurl Kim, Seongmo Park |
PCS | 2 |
| 2013 | Automatic and personalized recommendation of TV program contents using sequential pattern mining for smart TV user interaction
Shinjee Pyo, Eunhui Kim, Munchurl Kim |
Multim. Syst. | 3 |
| 2013 | Low complexity object detection and tracking with inter-layer graph mapping and intra-layer graph refinement in H.264/SVC bitstreams
M. S. Houari Sabirin, Munchurl Kim |
Pattern Recognit. Lett. | 2 |
| 2013 | A Novel DCT-Based JND Model for Luminance Adaptation Effect in DCT FrequencyabstractMany conventional DCT based Just Noticeable Distortion (JND) models incorporate luminance adaptation (LA) effect of the human visual system (HVS). The conventional LA-JND models exploit only background luminance to estimate JND values. In this letter, we reveal that the LA effect of HVS depends not only on background luminance but also on frequency in DCT domain. In addition, we first propose a novel DCT-based LA-JND model that takes into account its frequency characteristics. From our psychophysical experiment results, we found that the LA-JND threshold exhibits quasi-parabolic shapes with lower and higher curvatures in lower and higher frequency ranges, respectively. Our subjective evaluation shows that the proposed LA-JND model yields almost invisible distortions for test images with average PSNR of 29.11 dB, which is 2.63 dB lower than the other models for comparison, showing higher consistency with real visual perception for the LA in DCT domain. Sung-Ho Bae, Munchurl Kim |
IEEE Signal Process. Lett. | 2 |
| 2013 | Fast CU Splitting and Pruning for Suboptimal CU Partitioning in HEVC Intra CodingabstractHigh Efficiency Video Coding (HEVC), a new video coding standard currently being established, adopts a quadtree-based Coding Unit (CU) block partitioning structure that is flexible in adapting various texture characteristics of images. However, this causes a dramatic increase in computational complexity compared to previous video coding standards due to the necessity of finding the best CU partitions. In this paper, a fast CU splitting and pruning method is presented for HEVC intra coding, which allows for significant reduction in computational complexity with small degradations in rate-distortion (RD) performance. The proposed fast splitting and pruning method is performed in two complementary steps: 1) early CU split decision and 2) early CU pruning decision. For CU blocks, the early CU splitting and pruning tests are performed at each CU depth level according to a Bayes decision rule method based on low-complexity RD costs and full RD costs, respectively. The statistical parameters for the early CU split and pruning tests are periodically updated on the fly for each CU depth level to cope with varying signal characteristics. Experimental results show that our proposed fast CU splitting and pruning method reduces the computational complexity of the current HM to about 50% in encoding time with only 0.6% increases in BD rate. Seunghyun Cho, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | A Hybrid Stereoscopic Video Coding Scheme Based on MPEG-2 and HEVC for 3DTV ServicesabstractRecently, 3DTV has drawn much attention as a new broadcasting service. In spite of technical advances in 3DTV broadcasting services, there are two barriers that hinder the new services from being launched for terrestrial broadcasting: the lack of available bandwidth for transmission of additional-view video via terrestrial channel and backward compatibility with the legacy 2-D HDTV services. As an alternative, a hybrid stereoscopic video coding scheme is proposed for a stereoscopic 3DTV service, where one-view video is transmitted through the legacy broadcasting systems and the other one is delivered via broadband networks to which TV terminals are connected. In this paper, the proposed hybrid stereoscopic video coding scheme incorporates an MPEG-2 codec for backward compatibility with a legacy 2-D HDTV service and an HEVC codec with inter-view prediction coding as an extension to encode additional-view video sequences with high coding efficiency. The proposed inter-view prediction coding scheme in the extended HEVC incorporates an advanced motion and disparity vector prediction (AMDVP) method for enhanced motion- and disparity-compensated coding. The experimental results demonstrate that the proposed hybrid stereoscopic video coding scheme achieves an average coding efficiency of 38.22% (32.86%) in BD-rate or 1.394 dB (1.235 dB) in BD-PSNR for the out-band (in-band) scenario. Taeyoung Na, Sangsoo Ahn, M. S. Houari Sabirin, Munchurl Kim, Byungsun Kim, Sangjin Hahm, Keunsik Lee |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Moving Object Detection and Tracking Using a Spatio-Temporal Graph in H.264/AVC Bitstreams for Video SurveillanceabstractThis paper presents a spatio-temporal graph-based method of detecting and tracking moving objects by treating the encoded blocks with non-zero motion vectors and/or non-zero residues as potential parts of objects in H.264/AVC bitstreams. A spatio-temporal graph is constructed by first clustering the encoded blocks of potential object parts into block groups, each of which is defined as an attributed subgraph where the attributes of the vertices represent the positions, motion vectors and residues of the blocks. In order to remove false-positive blocks and to track the real objects, temporal connections between subgraphs in two consecutive frames are constructed and the similarities between subgraphs are computed, which constitutes a spatio-temporal graph. We show the experimental results that the proposed spatio-temporal graph-based representation of potential object blocks enables effective detection for the small-sized objects and the objects with small motion vectors and residues, and allows for reliable tracking of the detected objects even under occlusion. The identification of the detected moving objects is determined as rectangular regions of interest (ROIs) for which the ROI sizes and positions are adaptively adjusted to give the best approximation of the real shapes and positions of the objects. M. S. Houari Sabirin, Munchurl Kim |
IEEE Trans. Multim. | 2 |
| 2011 | A perceptual quality assessment metric using temporal complexity and disparity information for stereoscopic videoabstractVisual discomfort/fatigue problems in 3D video have become an important issue. In this paper, we examine the factors that affect human perception of depth and visual comfort from stereoscopic video. For this, we conduct subjective quality assessments from which we extract four factors - temporal variance, disparity variation in intra-frames, disparity variation in inter-frames and disparity distribution of frame boundary areas. Based on these four factors, we design a no-reference stereoscopic video quality perception model (SV-QPM) as an objective quality assessment metric for stereoscopic video. The proposed SV-QPM does not require depth map but utilize the disparity information by simple estimation, and the model parameters are estimated based on linear regression. The experimental results show that our proposed model exhibits high consistency with subjective quality assessment results in terms of a Pearson correlation coefficient value of 0.808, and the prediction performance exhibits good consistency with zero outlier ratio value. Kwangsung Ha, Munchurl Kim |
ICIP | 2 |
| 2011 | Hybrid video codec based joint rate control of stereoscopic video for terrestrial broadcasting servicesabstractRecently, terrestrial stereoscopic 3DTV broadcasting services are under preparation. In order to support backward compatibility for the current 2DTV services with MPEG-2 video, one way for the new terrestrial stereoscopic 3DTV broadcasting services is to adopt a hybrid video codec system in which the current MPEG-2 standard is used for the left or right channel while a more efficient video codec is used for the other channel of a stereoscopic video. In this case, it becomes important to maintain the similar visual quality for output bitstreams given target bitrates. In this paper, a hybrid-codec based joint bit rate control scheme is proposed for stereoscopic video encoding, which is based on an optimization framework. The experimental results show that our proposed joint rate control scheme can effectively maintain similar visual qualities of the left-channel bitstreams by MPEG-2 and the right-channel bitstreams by H.264/AVC under target bitrate constraints. Yongjun Chang, Munchurl Kim |
ICME | 2 |
| 2011 | Graph-based object detection and tracking in H.264/AVC bitstreams for surveillance videoabstractIn this paper we present a novel method to detect and track moving objects in H.264/AVC bitstreams by processing motion vector and residue information. The encoded blocks with nonzero motion vectors and residues are first detected as moving object candidates. A spatio-temporal graph in video sequences is then constructed to represent groups of blocks in each frame and their associations to the other groups of blocks in subsequent frames. Identification and refinement of ROIs for moving objects being tracked are done by graph matching and adaptive ROI-size adjustment. The experimental results show that the proposed method can correctly identify real moving objects from frame to frame and can effectively detect small-sized objects and objects with small motion vectors and residues, as well as by recognizing moving objects even under occlusion. M. S. Houari Sabirin, Jaeil Kim, Munchurl Kim |
ICME | 3 |
| 2011 | Hybrid Codec-Based Intra-Frame Joint Rate Control for Stereoscopic VideoabstractAn intra-frame joint rate control scheme is first proposed for a hybrid coder to encode stereoscopic video, which is based on an optimization framework with a gradient-based quadratic rate-quantization model and a gradient-based linear distortion-quantization model. The proposed rate control scheme jointly works on the left and right encoders for stereoscopic video input by controlling the output bit rates of both encoders in the sense that the sum of the two decoded video qualities is maximized and the quality difference is maintained around a desired level for a given target bit budget at the same time. In experiments, the proposed intra-frame joint rate control scheme for the hybrid coder produces the average 0.62 dB gain in PSNR, 64.93% reduction in the mean PSNR differences and 72.04% reduction in the MSE of PSNR difference, compared with the independent rate control schemes of the MPEG-2 TM5 and H.264 JM 16. Yongjun Chang, Munchurl Kim |
IEEE Signal Process. Lett. | 2 |
| 2011 | Modeling Rates and Distortions Based on a Mixture of Laplacian Distributions for Inter-Predicted Residues in Quadtree Coding of HEVCabstractInprobability model based rate control of video coding, modeling of residual distribution is important in predicting precise distortions so as to determine appropriate quantization parameter values. For this, single probability model approaches have been popularly taken which may fail to model the underlying statistical characteristics of different residues from variable block-sized coding. In this letter, new rate and distortion models based on a mixture of multiple Laplacian distributions are presented for the transform coefficients of inter-predicted residues in quadtree coding. The proposed mixture model of multiple Laplacian distributions is tested for the High Efficiency Video Coding (HEVC) Test Model (HM) with quadtree-structured Coding Unit and Transform Unit. The experimental results show that the proposed model achieves more accurate results of rate and distortion estimation than the single probability models. Bumshik Lee, Munchurl Kim |
IEEE Signal Process. Lett. | 2 |
| 2011 | A Selective Protection Scheme for Scalable Video CodingabstractSelective protection can exploit dependency coding properties to effectively perform partial protection on scalable video coding (SVC) bitstreams since protecting frames in lower scalability layers affects visual quality of the reconstructed frames in higher scalability layers. In this paper, we propose a selective protection scheme that maximizes protection effects by the minimum number of encoded frames to be protected in the SVC bitstream domain. We first model the SVC dependency coding structure as a directed acyclic graph which is characterized with an estimated visual quality value as the attribute at each node. In addition, a visual quality estimation model is proposed based on the proportions of intra-predicted and inter-predicted MBs, amounts of residues, and estimated visual quality of reference frames. The proposed selective protection scheme traverses the dependency graph to find optimal protection paths that can give the maximum visual quality degradation. Experimental results show that, compared to the existing protection schemes, the proposed selective protection scheme reduces computation complexity in the number of protected frames, the amount of protected data, and the protection time saving. The SVC file format specification supports the carriage of selectively protected bitstreams based on the concept of our selective protection in dependency coding structure of SVC. Hendry, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | A Low Complexity Mode Decision Method for Spatial Scalability CodingabstractIn this paper, a fast mode decision method for the spatial higher layers (SHLs) in scalable video coding (SVC) is proposed based on coding dependency between two adjacent spatial lower and higher layers. The proposed fast mode decision method detects zero motion and zero transform coefficient blocks in the current spatial layer using the already encoded information of the corresponding blocks from the lower layer. The information for zero motion vectors and zero transform coefficients is used to induce a reduced set of candidate modes in SHLs of SVC, which reduces the computation complexity up to about 75% of the total encoding time while maintaining the coding performance with negligible amounts of degradation in peak signal-to-noise ratio values and bitrates. Bumshik Lee, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | A hierarchical variable-sized block transform coding scheme for coding efficiency improvement on H.264/AVCabstractIn this paper, a rate-distortion optimized variable block transform coding scheme is proposed based on a hierarchical structured transform for macroblock (MB) coding with a set of the order-4 and −8 integer cosine transform (ICT) kernels of H.264/AVC as well as a new order-16 ICT kernel. The set of order-4, −8 and −16 ICT kernels are applied for inter-predictive coding in square (4×4, 8×8 or 16×16) or non-square (16×8 or 8×16) transform for each MB in a hierarchical structured manner. The proposed hierarchical variable-sized block transform scheme using the order-16 ICT kernel achieves significant bitrate reduction up to 15%, compared to the High profile of H.264/AVC. Even if the number of candidates for the transform types increases, the encoding time can be reduced to average 4–6% over the H.264/AVC Bumshik Lee, Jaeil Kim, Sangsoo Ahn, Munchurl Kim, Hui Yong Kim, Jong-Ho Kim, Jin Soo Choi |
PCS | 4 |
| 2009 | An ROI/xROI Based Rate Control Algorithm in H.264|AVC for Video Telephony Applications
Changhee Kim, Taeyoung Na, Jeongyeon Lim, Youngho Joo, Kimun Kim, Jaewoan Byun, Munchurl Kim |
PSIVT | 7 |
| 2008 | Implementation on a Real-Time SVC Encoder for Mobile BroadcastingabstractScalable Video Coding (SVC) can be applicable in mobile broadcasting environment due to the flexibility of spatial, temporal and quality scalability. Recently, SVC technology becomes mature rapidly but its reference SW encoder isn't optimized yet. Therefore, we have developed a real-time SW SVC encoder for broadcasting. In this paper, we show our SVC encoder that can provide two spatial layers: QVGA(320 x 240) and VGA(640 x 480). The base layer can be fully compatible with H.264/AVC. Our encoder is performing real-time operation on a normal PC by optimizing SVC algorithm. Sangjin Hahm, Changseob Park, Keunsoo Park, Munchurl Kim |
CCNC | 4 |
| 2008 | An efficient intra mode selection scheme for inter frame coding in H.264|AVC video codingabstractAn efficient intra mode decision scheme is proposed to reduce the computational complexity of inter frame coding in for the H.264/AVC. To decrease this computational burden, we propose an adaptive thresholding method based on distribution characteristics of the sum of the absolute differences (SAD) for the best inter mode when the intra mode is the final coding mode. We also include a simple refinement process by using spatial correlation between neighbouring marcoblocks (MBs) and the current MB. The performance of the proposed scheme is presented in terms of reduction in encoding times and PSNR values, and the increments in bit amounts. Jong-Ho Kim, Byung-Gyu Kim, Munchurl Kim |
ICME | 3 |
| 2008 | A target advertisement system based on TV viewer's profile reasoning
Jeongyeon Lim, Munjo Kim, Bumshik Lee, Munchurl Kim, Heekyung Lee, Hankyu Lee |
Multim. Tools Appl. | 4 |
| 2008 | An Optimal Adaptation Framework for Streaming Multiple Video ObjectsabstractWith multi-angle video contents available at a user terminal, the user can select his/her preferred alternate views among the given multiple video streams captured at different view angles for a same event. This enhanced experience often entails streaming problems in real-time over the network. In order to cope with this problem, multi-angle video contents are encoded at different bit rates and the appropriate video streams are then selected or transcoded for delivery to meet such bandwidth constraints. Therefore, the selection and transcoding operations become essential processing steps to adapt to the time-varying bandwidth of the network. In this paper, we propose an optimal adaptation framework for streaming multiple video streams by jointly formulating the selection and transcoding problems into a unified optimization problem. Jeongyeon Lim, Munchurl Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Automatic user preference learning for personalized electronic program guide applicationsabstractAbstract In this article, we introduce a user preference model contained in the User Interaction Tools Clause of the MPEG‐7 Multimedia Description Schemes, which is described by a UserPreferences description scheme (DS) and a UsageHistory description scheme (DS). Then we propose a user preference learning algorithm by using a Bayesian network to which weighted usage history data on multimedia consumption is taken as input. Our user preference learning algorithm adopts a dynamic learning method for learning real‐time changes in a user's preferences from content consumption history data by weighting these choices in time. Finally, we address a user preference–based television program recommendation system on the basis of the user preference learning algorithm and show experimental results for a large set of realistic usage‐history data of watched television programs. The experimental results suggest that our automatic user reference learning method is well suited for a personalized electronic program guide (EPG) application. Jeongyeon Lim, Sanggil Kang, Munchurl Kim |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2006 | An Automatic Personal TV Scheduler Based on HMM for Intelligent Broadcasting Services
Agus Syawal Yudhistira, Munchurl Kim, Hieyong Kim, Hankyu Lee |
PSIVT | 2 |
| 2006 | Intelligent broadcasting system and services for personalized semantic contents consumption
Sung Ho Jin, Tae Meon Bae, Yong Man Ro, Hoirin Kim, Munchurl Kim |
Expert Syst. Appl. | 5 |
| 2005 | Visual content adaptation according to user perception characteristicsabstractAdapting multimedia content to users' preferences and perceptual characteristics is a key direction for enabling personalized multimedia services. In this paper, we address the problem of tailoring visual content within the MPEG-21 Digital Item Adaptation (DIA) framework to meet users' visual perception characteristics. In particular, we present methods for adapting visual content to accommodate color vision deficiency and low-vision capabilities. In addition, we present methods for adapting visual content according to user preferences for color temperature. Finally, we report on experiments that adapt visual content within the MPEG-21 DIA framework. Jeho Nam, Yong Man Ro, Youngsik Huh, Munchurl Kim |
IEEE Trans. Multim. | 4 |
| 2004 | Modeling the user preference on broadcasting contents using Bayesian networksabstractIn this paper, we introduce a new supervised learning method of a Bayesian network for user preference models. Unlike other preference models, our method traces the trend of a user preference as time passes. It allows us to do online learning so we do not need the exhaustive data collection. The tracing of the trend can be done by modifying the frequency of attributes in order to force the old preference to be correlated with the current preference under the assumption that the current preference is correlated with the near future preference. The objective of our learning method is to force the mutual information to be reinforced by modifying the frequency of the attributes in the old preference by providing weights to the attributes. With developing mathematical derivation of our learning method, experimental results on the learning and reasoning performance on TV genre preference using a real set of TV program watching history data. Sanggil Kang, Jeongyeon Lim, Munchurl Kim |
VCIP | 3 |
| 2003 | Semantic transcoding of video based on regions of interest
Jeongyeon Lim, Munchurl Kim, Jong-Nam Kim, Kyeongsoo Kim |
VCIP | 2 |
| 2001 | Moving object segmentation in video sequences by user interaction and automatic object tracking
Munchurl Kim, Jun Geun Jeon, Jinsuk Kwak, Myoung Ho Lee, Chieteuk Ahn |
Image Vis. Comput. | 1 |
| 2000 | Summary description schemes for efficient video navigation and browsing
Jae-Gon Kim, Hyun Sung Chang, Munchurl Kim, Jinwoong Kim, Hyung-Myung Kim |
VCIP | 3 |
| 1999 | A VOP generation tool: automatic segmentation of moving objects in image sequences based on spatio-temporal informationabstractThe new MPEG-4 video coding standard enables content-based functionalities. In order to support the philosophy of the MPEG-4 visual standard, each frame of video sequences should be represented in terms of video object planes (VOPs). In other words, video objects to be encoded in still pictures or video sequences should be prepared before the encoding process starts. Therefore, it requires a prior decomposition of sequences into VOPs so that each VOP represents a moving object. This paper addresses an image segmentation method for separating moving objects from the background in image sequences. The proposed method utilizes the following spatio-temporal information. (1) For localization of moving objects in the image sequence, two consecutive image frames in the temporal direction are examined and a hypothesis testing is performed by comparing two variance estimates from two consecutive difference images, which results in an F-test. (2) Spatial segmentation is performed to divide each image into semantic regions and to find precise object boundaries of the moving objects. The temporal segmentation yields a change detection mask that indicates moving areas (foreground) and nonmoving areas (background), and spatial segmentation produces spatial segmentation masks. A combination of the spatial and temporal segmentation masks produces VOPs faithfully. This paper presents various experimental results. Munchurl Kim, Jae-Gark Choi, Daehee Kim 0005, Hyung Lee, Myoung Ho Lee, Chieteuk Ahn, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Target discrimination in synthetic aperture radar using artificial neural networksabstractThis paper addresses target discrimination in synthetic aperture radar (SAR) imagery using linear and nonlinear adaptive networks. Neural networks are extensively used for pattern classification but here the goal is discrimination. We show that the two applications require different cost functions. We start by analyzing with a pattern recognition perspective the two-parameter constant false alarm rate (CFAR) detector which is widely utilized as a target detector in SAR. Then we generalize its principle to construct the quadratic gamma discriminator (QGD), a nonparametrically trained classifier based on local image intensity. The linear processing element of the QCD is further extended with nonlinearities yielding a multilayer perceptron (MLP) which we call the NL-QGD (nonlinear QGD). MLPs are normally trained based on the L(2) norm. We experimentally show that the L(2) norm is not recommended to train MLPs for discriminating targets in SAR. Inspired by the Neyman-Pearson criterion, we create a cost function based on a mixed norm to weight the false alarms and the missed detections differently. Mixed norms can easily be incorporated into the backpropagation algorithm, and lead to better performance. Several other norms (L(8), cross-entropy) are applied to train the NL-QGD and all outperformed the L(2) norm when validated by receiver operating characteristics (ROC) curves. The data sets are constructed from TABILS 24 ISAR targets embedded in 7 km(2) of SAR imagery (MIT/LL mission 90). José C. Príncipe, Munchurl Kim, John W. Fisher III |
IEEE Trans. Image Process. | 2 |