VLDB 2026 Research / reviewers in the wild / expert
Seung-Won Jung
dblp:80/6256
· DBLP profile ↗
54ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0002-0319-4467ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 12 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Systems, architecture and hardware · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal Compressed Image Restoration via Codec-Aware Conditioning with Reinforcement LearningabstractWe address the task of universal compressed image restoration, which involves recovering high-quality images degraded by a wide range of codecs and compression levels. While prior methods have made significant progress, they typically target specific degradation types and struggle to generalize across both traditional and learning-based codecs. To overcome this limitation, we propose a unified framework that leverages codec-aware conditioning and reinforcement learning-based fine-tuning. Specifically, we introduce a conditioning module that encodes both codec type and compression level, enabling the restoration network to adapt its behavior to diverse degradation settings. To further improve generalization, we incorporate reward-based objectives during fine-tuning, providing complementary signals that enhance training across both conventional and learned compression schemes. Experimental results demonstrate the effectiveness of our method in restoring images across a wide range of compression artifacts and scenarios. Changwoo Han 0001, Hongil Kim 0001, Donghyun Kim 0017, Sung-Chang Lim, Seung-Won Jung |
AAAI | 5 |
| 2026 | An Anisotropic Cross-View Texture Transfer With Multi-Reference Non-Local Attention for CT Slice InterpolationabstractComputed tomography (CT) is one of the most widely used non-invasive imaging modalities for medical diagnosis. In clinical practice, CT images are usually acquired with large slice thicknesses due to the high cost of memory storage and operation time, resulting in an anisotropic CT volume with much lower inter-slice resolution than in-plane resolution. Since such inconsistent resolution may lead to difficulties in disease diagnosis, deep learning-based volumetric super-resolution methods have been developed to improve inter-slice resolution. Most existing methods conduct single-image super-resolution on the through-plane or synthesize intermediate slices from adjacent slices; however, the anisotropic characteristic of 3D CT volume has not been well explored. In this paper, we propose a novel cross-view texture transfer approach for CT slice interpolation by fully utilizing the anisotropic nature of 3D CT volume. Specifically, we design a unique framework that takes high-resolution in-plane texture details as a reference and transfers them to low-resolution through-plane images. To this end, we introduce a multi-reference non-local attention module that extracts meaningful features for reconstructing through-plane high-frequency details from multiple in-plane images. Through extensive experiments, we demonstrate that our method performs significantly better in CT slice interpolation than existing competing methods on public CT datasets including a real-paired benchmark, verifying the effectiveness of the proposed framework. The source code of this work is available at https://github.com/khuhm/ACVTT. Kwang-Hyun Uhm, Hyunjun Cho, Sung-Hoo Hong, Seung-Won Jung |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor ScenesabstractWe propose a diffusion-based inverse rendering framework that decomposes a single RGB image into geometry, material, and lighting. Inverse rendering is inherently ill-posed, making it difficult to predict a single accurate solution. To address this challenge, recent generative model-based methods aim to present a range of possible solutions. However, finding a single accurate solution and generating diverse solutions can be conflicting. In this paper, we propose a channel-wise noise scheduling approach that allows a single diffusion model architecture to achieve two conflicting objectives. The resulting two diffusion models, trained with different channel-wise noise schedules, can predict a single highly accurate solution and present multiple possible solutions. The experimental results demonstrate the superiority of our two models in terms of both diversity and accuracy, which translates to enhanced performance in downstream applications such as object insertion and material editing. Junyong Choi, Min-Cheol Sagong, SeokYeong Lee, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
CVPR | 4 |
| 2025 | Text Embedding Knows How to Quantize Text-Guided Diffusion Models
Hongjae Lee, Myungjun Son, Dongjea Kang, Seung-Won Jung |
ICCV | 4 |
| 2025 | IM-LUT: Interpolation Mixing Look-Up Tables for Image Super-Resolution
Sejin Park 0002, Kyong Hwan Jin, Seung-Won Jung |
ICCV | 4 |
| 2025 | Quantization-friendly super-resolution: Unveiling the benefits of activation normalization
Dongjea Kang, Myungjun Son, Hongjae Lee, Seung-Won Jung |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | MAIR++: Improving Multi-View Attention Inverse Rendering With Implicit Lighting RepresentationabstractIn this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level inverse rendering, scene-level inverse rendering has primarily been studied using single-view images due to the lack of a dataset containing high dynamic range multi-view images with ground-truth geometry, material, and spatially-varying lighting. To improve the quality of scene-level inverse rendering, a novel framework called Multi-view Attention Inverse Rendering (MAIR) was recently introduced. MAIR performs scene-level multi-view inverse rendering by expanding the OpenRooms dataset, designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Although MAIR showed impressive results, its lighting representation is fixed to spherical Gaussians, which limits its ability to render images realistically. Consequently, MAIR cannot be directly used in applications such as material editing. Moreover, its multi-view aggregation networks have difficulties extracting rich features because they only focus on the mean and variance between multi-view features. In this paper, we propose its extended version, called MAIR++. MAIR++ addresses the aforementioned limitations by introducing an implicit lighting representation that accurately captures the lighting conditions of an image while facilitating realistic rendering. Furthermore, we design a directional attention-based multi-view aggregation network to infer more intricate relationships between views. Experimental results show that MAIR++ not only outperforms MAIR and single-view-based methods but also demonstrates robust performance on unseen real-world scenes. Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Intrinsic Decomposition-Based Curriculum Learning for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) aims to restore the visual quality of images captured under poor illumination conditions, a task that remains challenging due to complex degradations such as overexposure, noise, and low contrast. In this paper, we propose a novel curriculum learning framework that facilitates effective LLIE model training by modulating sample selection according to estimated difficulty. Our key insight is that residual signals obtained via intrinsic decomposition capture image characteristics such as color spill and indirect lighting, which are strongly correlated with reconstruction difficulty. We use the magnitude of these residuals as a proxy for difficulty, enabling a curriculum strategy that begins with easier samples and gradually incorporates more difficult ones. Extensive experiments demonstrate that the proposed method consistently improves performance across various LLIE baseline models and datasets. Being model-agnostic and plug-and-play, our method offers meaningful gains through curriculum learning without requiring additional annotations or architectural modifications. Jinah Kim, Geon-Ho Park, Juhyun Bae, Seung-Won Jung |
IEEE Signal Process. Lett. | 4 |
| 2025 | A Nuclei-Focused Strategy for Automated Histopathology Grading of Renal Cell CarcinomaabstractThe rising incidence of kidney cancer underscores the need for precise and reproducible diagnostic methods. In particular, renal cell carcinoma (RCC), the most prevalent type of kidney cancer, requires accurate nuclear grading for better prognostic prediction. Recent advances in deep learning have facilitated end-to-end diagnostic methods using contextual features in histopathological images. However, most existing methods focus only on image-level features or lack an effective process for aggregating nuclei prediction results, limiting their diagnostic accuracy. In this paper, we introduce a novel framework, Nuclei feature Assisted Patch-level RCC grading (NuAP-RCC), that leverages nuclei-level features for enhanced patch-level RCC grading. Our approach employs a nuclei-level RCC grading network to extract grade-aware features, which serve as node features in a graph. These node features are aggregated using graph neural networks to capture the morphological characteristics and distributions of the nuclei. The aggregated features are then combined with global image-level features extracted by convolutional neural networks, resulting in a final feature for accurate RCC grading. In addition, we present a new dataset for patch-level RCC grading. Experimental results demonstrate the superior accuracy and generalizability of NuAP-RCC across datasets from different medical institutions, achieving a 6.15% improvement in accuracy over the second-best model on the USM-RCC dataset. Hyunjun Cho, Dongjin Shin, Kwang-Hyun Uhm, Sung-Jea Ko, Yosep Chong, Seung-Won Jung |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Spectrum Translation for Refinement of Image Generation (STIG) Based on Contrastive Learning and Spectral Filter ProfileabstractCurrently, image generation and synthesis have remarkably progressed with generative models. Despite photo-realistic results, intrinsic discrepancies are still observed in the frequency domain. The spectral discrepancy appeared not only in generative adversarial networks but in diffusion models. In this study, we propose a framework to effectively mitigate the disparity in frequency domain of the generated images to improve generative performance of both GAN and diffusion models. This is realized by spectrum translation for the refinement of image generation (STIG) based on contrastive learning. We adopt theoretical logic of frequency components in various generative networks. The key idea, here, is to refine the spectrum of the generated image via the concept of image-to-image translation and contrastive learning in terms of digital signal processing. We evaluate our framework across eight fake image datasets and various cutting-edge models to demonstrate the effectiveness of STIG. Our framework outperforms other cutting-edges showing significant decreases in FID and log frequency distance of spectrum. We further emphasize that STIG improves image quality by decreasing the spectral anomaly. Additionally, validation results present that the frequency-based deepfake detector confuses more in the case where fake spectrums are manipulated by STIG. Seokjun Lee, Seung-Won Jung, Hyunseok Seo |
AAAI | 2 |
| 2024 | PD-CR: Patch-Based Diffusion Using Constrained Refinement for Image RestorationabstractDiffusion models, which are state-of-the-art generative models, have been widely applied to image restoration tasks. However, most image restoration methods based on diffusion models require a large amount of computational memory, making it difficult to use them with high-resolution images. Although patch-based diffusion models have emerged to address this problem, these models are limited in effectively mitigating boundary artifacts and producing results close to the ground truth. In this paper, we propose Patch-based Diffusion using Constrained Refinement (PD-CR) that refines the noise estimated by patch-based diffusion models to produce a restored image while keeping the luminance of the input degraded image. Leveraging patch-based diffusion models, the proposed method can handle a high-resolution image as input with minimal memory requirements. Our experiments on various image restoration tasks, such as image denoising and raindrop removal, demonstrate that the proposed method is better than or on par with the state-of-the-art methods. Hyunjun Cho, Hong-Kyu Shin, Yurim Jang, Sung-Jea Ko, Seung-Won Jung |
IEEE Signal Process. Lett. | 5 |
| 2024 | Revisiting PID Control for Power-Constrained Video DisplayabstractDue to the significant power consumption of emissive display panels and the limited battery capacity of electronic devices, it is necessary to implement power-constrained display operations for video content. One of the effective power-constrained display techniques is power-constrained contrast enhancement (PCCE), which aims to reduce the power demands of the display while maintaining the quality of the content. However, PCCE has been mainly studied in the image domain. In this letter, we explore how PCCE can be applied to the video domain by revisiting the proportional-integral-derivative (PID) control. The proposed method consists of three key components: 1) The proportional term determines a power-saving ratio for each video frame based on its luminance, resulting in aggressive power-saving in bright video frames; 2) the integral term ensures that the target power-saving ratio is achieved over the entire video; 3) the differential term prevents abrupt temporal fluctuations in the power-saving ratio of consecutive video frames. The experimental results demonstrate that the proposed PID-based PCCE method achieves superior performance on two public video datasets. In particular, the proposed method achieves 5.8% improvements in VMAF compared to the previous method on the TVSum dataset. Yurim Jang, Hyunjun Cho, Geon-Ho Park, Min-Jae Yoo, Seung-Won Jung |
IEEE Signal Process. Lett. | 5 |
| 2024 | CAWM: Class-Aware Weight Map for Improved Semi-Supervised Nuclei SegmentationabstractDue to the rich histopathological information of nuclei in whole slide images, nuclei segmentation becomes essential for medical analysis. Since collecting sufficient pixel-wise annotations for supervised training of nuclei segmentation networks is challenging, semi-supervised nuclei segmentation methods have been extensively studied. In particular, many of them use pseudo-labels generated from unlabeled images for training the segmentation model. In this Letter, we propose a new pseudo-label handling method for semi-supervised nuclei segmentation. Specifically, based on our observation that nuclear features within the same image share high similarities, we define confidence maps for pseudo-labels and use them to adapt consistency regularization and contrastive loss measures. From extensive experiments on three public datasets, we demonstrate the effectiveness of the proposed method compared with other semi-supervised training methods. Seohoon Lim, Zhixin Xu, Yosep Chong, Seung-Won Jung |
IEEE Signal Process. Lett. | 4 |
| 2024 | RefQSR: Reference-Based Quantization for Image Super-Resolution NetworksabstractSingle image super-resolution (SISR) aims to reconstruct a high-resolution image from its low-resolution observation. Recent deep learning-based SISR models show high performance at the expense of increased computational costs, limiting their use in resource-constrained environments. As a promising solution for computationally efficient network design, network quantization has been extensively studied. However, existing quantization methods developed for SISR have yet to effectively exploit image self-similarity, which is a new direction for exploration in this study. We introduce a novel method called reference-based quantization for image super-resolution (RefQSR) that applies high-bit quantization to several representative patches and uses them as references for low-bit quantization of the rest of the patches in an image. To this end, we design dedicated patch clustering and reference-based quantization modules and integrate them into existing SISR network quantization methods. The experimental results demonstrate the effectiveness of RefQSR on various SISR networks and quantization methods. Hongjae Lee, Jun-Sang Yoo 0002, Seung-Won Jung |
IEEE Trans. Image Process. | 3 |
| 2024 | Pose and Shape Estimation of Humans in VehiclesabstractAs autonomous driving technologies advance, occupants are expected to be free from driving, diversifying interaction scenarios with vehicles. Despite the growing importance of in-vehicle occupant monitoring systems, most existing systems focus on the face or head tracking of occupants, and only a few studies have attempted to detect their poses. In this paper, we present the first in-vehicle environment-specialized framework for the joint estimation of 3D human pose and shape from a single image. To this end, we introduce a new dataset called Human In VEhicles (HIVE), which contains a large collection of synthesized humans with different shapes and poses in vehicle images. HIVE provides RGB and NIR in-vehicle image pairs with ground-truth 2D and 3D pose and shape annotations, respectively. In addition, to exploit the different characteristics of humans in vehicles and unconstrained environments, we present a new pose prior penalizing poses that deviate from in-vehicle poses. The pose prior is derived using a variational autoencoder trained with in-vehicle human pose data. By using the proposed HIVE dataset and pose prior along with an elaborately designed two-stage training procedure, our method exhibits significantly improved pose and shape estimation performance compared with state-of-the-art methods for real-world test images captured in vehicles under different conditions. Kwang-Lim Ko, Jun-Sang Yoo 0002, Changwoo Han 0001, Jungyeop Kim, Seung-Won Jung |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Cross-Domain Denoising for Low-Dose Multi-Frame Spiral Computed TomographyabstractComputed tomography (CT) has been used worldwide as a non-invasive test to assist in diagnosis. However, the ionizing nature of X-ray exposure raises concerns about potential health risks such as cancer. The desire for lower radiation doses has driven researchers to improve reconstruction quality. Although previous studies on low-dose computed tomography (LDCT) denoising have demonstrated the effectiveness of learning-based methods, most were developed on the simulated data. However, the real-world scenario differs significantly from the simulation domain, especially when using the multi-slice spiral scanner geometry. This paper proposes a two-stage method for the commercially available multi-slice spiral CT scanners that better exploits the complete reconstruction pipeline for LDCT denoising across different domains. Our approach makes good use of the high redundancy of multi-slice projections and the volumetric reconstructions while leveraging the over-smoothing issue in conventional cascaded frameworks caused by aggressive denoising. The dedicated design also provides a more explicit interpretation of the data flow. Extensive experiments on various datasets showed that the proposed method could remove up to 70% of noise without compromised spatial resolution, while subjective evaluations by two experienced radiologists further supported its superior performance against state-of-the-art methods in clinical practice. Code is available at https://github.com/YCL92/TMD-LDCT. Yucheng Lu 0001, Zhixin Xu, Moon Hyung Choi, Seung-Won Jung |
IEEE Trans. Medical Imaging | 5 |
| 2023 | MAIR: Multi-View Attention Inverse Rendering with 3D Spatially-Varying Lighting EstimationabstractWe propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene, multi-view images in object-level inverse rendering have been taken for granted. However, owing to the absence of multi-view HDR synthetic dataset, scene-level inverse rendering has mainly been studied using single-view image. We were able to successfully perform scene-level inverse rendering using multi-view images by expanding OpenRooms dataset and designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Our experiments show that the proposed method not only achieves better performance than single-view-based methods, but also achieves robust performance on unseen real-world scene. Also, our sophisticated 3D spatially-varying lighting volume allows for photorealistic object insertion in any 3D location. Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
CVPR | 4 |
| 2023 | Video Object Segmentation-aware Video Frame InterpolationabstractVideo frame interpolation (VFI) is a very active re-search topic due to its broad applicability to many applications, including video enhancement, video encoding, and slow-motion effects. VFI methods have been advanced by improving the overall image quality for challenging sequences containing occlusions, large motion, and dynamic texture. This mainstream research direction neglects that foreground and background regions have different importance in perceptual image quality. Moreover, accurate synthesis of moving objects can be of utmost importance in computer vision applications. In this paper, we propose a video object segmentation (VOS)-aware training framework called VOS-VFI that allows VFI models to interpolate frames with more precise object boundaries. Specifically, we exploit VOS as an auxiliary task to help train VFI models by providing additional loss functions, including segmentation loss and bi-directional consistency loss. From extensive experiments, we demonstrate that VOS-VFI can boost the performance of existing VFI models by rendering clear object boundaries. Moreover, VOS-VFI displays its effectiveness on multiple benchmarks for different applications, including video object segmentation, object pose estimation, and visual tracking. The code is available at https://github.com/junsang7777/VOS-VFI Jun-Sang Yoo 0002, Hongjae Lee, Seung-Won Jung |
ICCV | 3 |
| 2023 | Semantic and Instance-Aware Pixel-Adaptive Convolution for Panoptic SegmentationabstractAlthough the weight-sharing property of convolution is one of the major reasons for the success of convolution neural networks, the content-agnostic operation is insufficient for several tasks requiring content-adaptive processing, including panoptic segmentation. Inspired by several recent works on content-adaptive convolutions, we introduce the GuidedPAKA, the first content-adaptive convolution method specialized for panoptic segmentation. Specifically, GuidedPAKA learns the pixel-adaptive kernel attention consisting of the channel and spatial kernel attentions. Instead of commonly used self-attention operation, we guide the channel and spatial kernel attentions using their respective supervision signals, i.e., semantic segmentation maps and local instance affinities. Consequently, these kernel attentions extract features helpful for panoptic segmentation. Experimental results show that the proposed GuidedPAKA improves the performance of panoptic segmentation when integrated into the baseline model. Sumin Song, Min-Cheol Sagong, Seung-Won Jung, Sung-Jea Ko |
ICIP | 3 |
| 2023 | Multispectral-to-RGB Knowledge Distillation for Remote Sensing Image Scene ClassificationabstractScene classification is a fundamental task in the remote sensing (RS) field, assigning semantic labels to RS images. Multispectral (MS) images play an essential role in scene classification as they contain richer spectral information than red, green, blue (RGB) images. However, MS images are not always available due to the higher cost and complexity of MS sensors compared to RGB sensors. To improve scene classification performance using only RGB images, in this letter, we propose a novel MS-to-RGB knowledge distillation (MS2RGB-KD) framework that transfers MS knowledge from a teacher model to a student model. Specifically, our MS2RGB-KD drives a student model that requires only an RGB image as input to mimic the feature representations of different modalities extracted by the teacher model. Moreover, we introduce novel loss functions that encourage the student model to preserve intramodal and intermodal relationships of the feature representations in the teacher model. Experiments on the EuroSAT dataset demonstrate the effectiveness of MS2RGB-KD compared with other KD baselines. Hong-Kyu Shin, Kwang-Hyun Uhm, Seung-Won Jung, Sung-Jea Ko |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | RZSR: Reference-Based Zero-Shot Super-Resolution With Depth Guided Self-ExemplarsabstractRecent methods for single image super-resolution (SISR) have demonstrated outstanding performance in generating high-resolution (HR) images from low-resolution (LR) images. However, most of these methods show their superiority using synthetically generated LR images, and their generalizability to real-world images is often not satisfactory. In this paper, we pay attention to two well-known strategies developed for robust super-resolution (SR), i.e., reference-based SR (RefSR) and zero-shot SR (ZSSR), and propose an integrated solution, called reference-based zero-shot SR (RZSR). Following the principle of ZSSR, we train an image-specific SR network at test time using training samples extracted only from the input image itself. To advance ZSSR, we obtain reference image patches with rich textures and high-frequency details which are also extracted only from the input image using cross-scale matching. To this end, we construct an internal reference dataset and retrieve reference image patches from the dataset using depth information. Using LR patches and their corresponding HR reference patches, we train a RefSR network that is embodied with a non-local attention module. Experimental results demonstrate the superiority of the proposed RZSR compared to the previous ZSSR methods and robustness to unseen images compared to other fully supervised SISR methods. Jun-Sang Yoo 0002, Yucheng Lu 0001, Seung-Won Jung |
IEEE Trans. Multim. | 4 |
| 2022 | RORD: A Real-world Object Removal Dataset
Min-Cheol Sagong, Yoon-Jae Yeo, Seung-Won Jung, Sung-Jea Ko |
BMVC | 3 |
| 2022 | XYDeblur: Divide and Conquer for Single Image DeblurringabstractMany convolutional neural networks (CNNs) for single image deblurring employ a U-Net structure to estimate latent sharp images. Having long been proven to be effective in image restoration tasks, a single lane of encoder-decoder architecture overlooks the characteristic of deblurring, where a blurry image is generated from complicated blur kernels caused by tangled motions. Toward an effective network architecture for single image deblurring, we present complemental sub-solution learning with a one-encoder-two-decoder architecture. Observing that multiple decoders successfully learn to decompose encoded feature information into directional components, we further improve both the network efficiency and the deblurring performance by rotating and sharing kernels exploited in the decoders, which prevents the decoders from separating unnecessary components such as color shift. As a result, our proposed network shows superior results compared to U-Net while preserving the network parameters, and using the proposed network as the base network can improve the performance of existing state-of-the-art deblurring networks. Seowon Ji, Jeongmin Lee 0005, Seung-Wook Kim 0002, Jun-Pyo Hong, Seung-Jin Baek, Seung-Won Jung, Sung-Jea Ko |
CVPR | 6 |
| 2022 | RAWtoBit: A Fully End-to-end Camera ISP Network
Wooseok Jeong, Seung-Won Jung |
ECCV (19) | 2 |
| 2022 | Progressive Joint Low-Light Enhancement and Noise Removal for Raw ImagesabstractLow-light imaging on mobile devices is typically challenging due to insufficient incident light coming through the relatively small aperture, resulting in low image quality. Most of the previous works on low-light imaging focus either only on a single task such as illumination adjustment, color enhancement, or noise removal; or on a joint illumination adjustment and denoising task that heavily relies on short-long exposure image pairs from specific camera models. These approaches are less practical and generalizable in real-world settings where camera-specific joint enhancement and restoration is required. In this paper, we propose a low-light imaging framework that performs joint illumination adjustment, color enhancement, and denoising to tackle this problem. Considering the difficulty in model-specific data collection and the ultra-high definition of the captured images, we design two branches: a coefficient estimation branch and a joint operation branch. The coefficient estimation branch works in a low-resolution space and predicts the coefficients for enhancement via bilateral learning, whereas the joint operation branch works in a full-resolution space and progressively performs joint enhancement and denoising. In contrast to existing methods, our framework does not need to recollect massive data when adapted to another camera model, which significantly reduces the efforts required to fine-tune our approach for practical usage. Through extensive experiments, we demonstrate its great potential in real-world low-light imaging applications. Yucheng Lu 0001, Seung-Won Jung |
IEEE Trans. Image Process. | 2 |
| 2022 | A Unified Multi-Phase CT Synthesis and Classification Framework for Kidney Cancer Diagnosis With Incomplete DataabstractMulti-phase computed tomography (CT) is widely adopted for the diagnosis of kidney cancer due to the complementary information among phases. However, the complete set of multi-phase CT is often not available in practical clinical applications. In recent years, there have been some studies to generate the missing modality image from the available data. Nevertheless, the generated images are not guaranteed to be effective for the diagnosis task. In this paper, we propose a unified framework for kidney cancer diagnosis with incomplete multi-phase CT, which simultaneously recovers missing CT images and classifies cancer subtypes using the completed set of images. The advantage of our framework is that it encourages a synthesis model to explicitly learn to generate missing CT phases that are helpful for classifying cancer subtypes. We further incorporate lesion segmentation network into our framework to exploit lesion-level features for effective cancer classification in the whole CT volumes. The proposed framework is based on fully 3D convolutional neural networks to jointly optimize both synthesis and classification of 3D CT volumes. Extensive experiments on both in-house and external datasets demonstrate the effectiveness of our framework for the diagnosis with incomplete data compared with state-of-the-art baselines. In particular, cancer subtype classification using the completed CT data by our method achieves higher performance than the classification using the given incomplete data. Kwang-Hyun Uhm, Seung-Won Jung, Moon Hyung Choi, Sung-Hoo Hong, Sung-Jea Ko |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Rethinking Coarse-to-Fine Approach in Single Image DeblurringabstractCoarse-to-fine strategies have been extensively used for the architecture design of single image deblurring networks. Conventional methods typically stack sub-networks with multi-scale input images and gradually improve sharpness of images from the bottom sub-network to the top sub-network, yielding inevitably high computational costs. Toward a fast and accurate deblurring network design, we revisit the coarse-to-fine strategy and present a multi-input multi-output U-net (MIMO-UNet). The MIMO-UNet has three distinct features. First, the single encoder of the MIMO-UNet takes multi-scale input images to ease the difficulty of training. Second, the single decoder of the MIMO-UNet outputs multiple deblurred images with different scales to mimic multi-cascaded U-nets using a single U-shaped network. Last, asymmetric feature fusion is introduced to merge multi-scale features in an efficient manner. Extensive experiments on the GoPro and RealBlur datasets demonstrate that the proposed network outperforms the state-of-the-art methods in terms of both accuracy and computational complexity. Source code is available for research purposes at https://github.com/chosj95/MIMO-UNet. Sung-Jin Cho 0002, Seowon Ji, Jun-Pyo Hong, Seung-Won Jung, Sung-Jea Ko |
ICCV | 4 |
| 2021 | Parametric Shape Estimation of Human Body Under Wide ClothingabstractThe shape of the human body plays an important role in many applications, such as those involving personal healthcare and virtual clothing try-ons. However, accurate body shape measurements typically require the user to be wearing a minimal amount of clothing, which is not practical in many situations. To resolve this issue using deep learning techniques, we need a paired dataset of ground-truth naked human body shapes and their corresponding color images with clothes. As it is practically impossible to collect enough of this kind of data from real-world environments to train a deep neural network, in this paper, we present the Synthetic dataset of Human Avatars under wiDE gaRment (SHADER). The SHADER dataset consists of 300,000 paired ground-truth naked and dressed images of 1,500 synthetic humans with different body shapes, poses, garments, skin tones, and backgrounds. To take full advantage of SHADER, we propose a novel silhouette confidence measure and show that our silhouette confidence prediction network can help improve the performance of state-of-the-art shape estimation networks for human bodies under clothing. The experimental results demonstrate the effectiveness of the proposed approach. The code and dataset are available athttps://github.com/YCL92/SHADER. Yucheng Lu 0001, Jin-Hyuck Cha, Sekyoung Youm, Seung-Won Jung |
IEEE Trans. Multim. | 4 |
| 2020 | Simple Yet Effective Way for Improving the Performance of Depth Map Super-ResolutionabstractIn depth map super-resolution (SR), a high-resolution color image plays an important role as guidance for preventing blurry depth boundaries. However, excessive/deficient use of the color image features often causes performance degradation such as texture-copying/edge-smoothing in flat/boundary areas. To alleviate these problems, this letter presents a simple yet effective method for enhancing the performance of the SR without requiring significant modifications to the original SR network. To this end, we present a self-selective concatenation (SSC), which is a substitute for the conventional feature concatenation. In the upsampling layers of the SR network, the SSC extracts spatial and channel attention from both color and depth features such that color features can be selectively used for depth SR. Specifically, the SSC learns to use sufficient color features for rendering sharp depth boundaries, whereas their effects are reduced in smooth regions to prevent texture-copying. The proposed SSC can be included in any existing SR networks that have the encoder-decoder structure. The experimental results show that the proposed method can further improve the performances of existing SR networks in terms of the root mean squared error and peak signal-to-noise ratio. Yoon-Jae Yeo, Min-Cheol Sagong, Yong-Goo Shin, Seung-Won Jung, Sung-Jea Ko |
IEEE Signal Process. Lett. | 4 |
| 2019 | Foreground extraction via dual-side cameras on a mobile device using long short-term trajectory analysis
Yucheng Lu 0001, Se-Song Kim, Seung-Won Jung |
Image Vis. Comput. | 4 |
| 2018 | Interactive Image Segmentation Using Semi-transparent Wearable GlassesabstractSince it is difficult to automatically and precisely extract an object of interest, interactive image segmentation techniques exploit user-provided segmentation seeds. In previous interactive segmentation applications, the segmentation seeds are typically provided by mouse clicks or finger touches. In this paper, the segmentation of an object is studied from the scene that the user sees through semi-transparent wearable glasses. In this application scenario, a front-view camera is used to obtain the segmentation seeds from the user's fingertip position. In particular, two segmentation methodologies called transparent segmentation and semi-transparent segmentation are considered to determine an effective segmentation scheme for the wearable glasses. Extensive user studies are performed to evaluate the user preferences and the segmentation accuracies of the two methodologies. Kyumok Kim, Seung-Won Jung |
IEEE Trans. Multim. | 2 |
| 2017 | A new depth image quality metric using a pair of color and depth images
Seung-Won Jung, Chee Sun Won |
Multim. Tools Appl. | 2 |
| 2017 | Depth completion for kinect v2 sensor
Wanbin Song, Anh Vu Le, Seok Min Yun, Seung-Won Jung, Chee Sun Won |
Multim. Tools Appl. | 4 |
| 2017 | Near-reversible efficient image resizing for devices supporting different spatial resolutions
Chee Sun Won, Seung-Won Jung |
J. Supercomput. | 2 |
| 2016 | Text-aware image dehazing using stroke width transformabstractHaze removal, which is also referred to as image dehazing, has been extensively used to improve the visibility in images captured under inclement weather. In particular, the dark channel prior (DCP)-based single image dehazing has received the greatest amount of interest due to its superior performance. However, since the DCP is based on the characteristics of natural outdoor images, its reliability tends to decrease especially when an image contains man-made textures. In this paper, we present a DCP-based single image dehazing method that is robust when text or text-like patterns are present in the image. The proposed method first estimates the text likelihood from a hazy image using the stroke width transform (SWT) and uses the estimated likelihood to correct the DCP. The experimental results show that the proposed algorithm outperforms the conventional DCP-based dehazing methods. Jinwon Park, Kyumok Kim, Chee Sun Won, Seung-Won Jung |
ICIP | 5 |
| 2016 | All-in-focus and multi-focus color image reconstruction from a database of color and depth image pairs
Seung-Won Jung, Jong Hyuk Park 0001, Young-Sik Jeong |
Multim. Tools Appl. | 1 |
| 2016 | Rotated top-bottom dual-kinect for improved field of view
Wanbin Song, Seok Min Yun, Seung-Won Jung, Chee Sun Won |
Multim. Tools Appl. | 3 |
| 2016 | Lossless embedding of depth hints in JPEG compressed color images
Seung-Won Jung |
Signal Process. | 1 |
| 2016 | Order-Preserving Condensation of Moving Objects in Surveillance VideosabstractVision-based detection of illegal or accidental activities in urban traffic has attracted great interest. Since state-of-the-art online automated detection algorithms are far from perfect, much research effort on offline video surveillance has been made to prevent police or security staff from observing all recorded video frames unnecessarily. To solve the problem, this study focuses on video condensation, which provides fast monitoring of moving objects in a long duration of surveillance videos. Considering the computational complexity and the condensation ratio as the two main criteria for efficient video condensation, we propose a video condensation algorithm, which consists of the following: 1) initial condensation by discarding frames of nonmoving objects; 2) intra-GoFM (group of frames with moving objects) condensation; and 3) inter-GoFM condensation. In the intra-GoFM and inter-GoFM condensation, spatiotemporal static pixels within each GoFM and temporal static pixels between two consecutive GoFMs are dropped to shorten the temporal distances between consecutive moving objects. Experimental results show that our video condensation saves a significant amount of computational loads compared with the previous methods without sacrificing the condensation ratio and visual quality. Hai Thanh Nguyen 0002, Seung-Won Jung, Chee Sun Won |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Adaptive post-filtering of JPEG compressed images considering compressed domain lossless data hiding
Seung-Won Jung |
Inf. Sci. | 1 |
| 2014 | Image Contrast Enhancement Using Color and Depth HistogramsabstractIn this letter, we propose a new global contrast enhancement algorithm using the histograms of color and depth images. On the basis of the histogram-modification framework, the color and depth image histograms are first partitioned into sub-intervals using the Gaussian mixture model. The positions partitioning the color histogram are then adjusted such that spatially neighboring pixels with the similar intensity and depth values can be grouped into the same sub-interval. By estimating the mapping curve of the contrast enhancement for each sub-interval, the global image contrast can be improved without over-enhancing the local image contrast. Experimental results demonstrate the effectiveness of the proposed algorithm. Seung-Won Jung |
IEEE Signal Process. Lett. | 1 |
| 2014 | Learning-Based Filter Selection Scheme for Depth Image Super ResolutionabstractDepth images that have the same spatial resolution as color images are required in many applications, such as multiview rendering and 3-D texture modeling. Since a depth sensor usually has poorer spatial resolution compared with a color image sensor, many depth image super-resolution methods have been investigated in the literature. With an assumption that no one super-resolution method can universally outperform the other methods, in this paper we introduce a learning-based selection scheme for different super-resolution methods. In our case study, three distinctive mean-type, max-type, and median-type filtering methods are selected as candidate methods. In addition, a new frequency-domain feature vector is designed to enhance the discriminability of the methods. Given the candidate methods and feature vectors, a classifier is trained such that the best method can be selected for each depth pixel. The effectiveness of the proposed scheme is first demonstrated using the synthetic data set. The noise-free and noisy low-resolution depth images are constructed, and the quantitative performance evaluation is performed by measuring the difference between the ground-truth high-resolution depth images and the resultant depth images. The proposed algorithm is then applied to real color and time-of-flight depth cameras. The experimental results demonstrate that the proposed algorithm outperforms the conventional algorithms both quantitatively and qualitatively. Seung-Won Jung, Ouk Choi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | A Consensus-Driven Approach for Structure and Texture Aware Depth Map UpsamplingabstractThis paper presents a method for increasing spatial resolution of a depth map using its corresponding high-resolution (HR) color image as a guide. Most of the previous methods rely on the assumption that depth discontinuities are highly correlated with color boundaries, leading to artifacts in the regions where the assumption is broken. To prevent scene texture from being erroneously transferred to reconstructed scene surfaces, we propose a framework for dividing the color image into different regions and applying different methods tailored to each region type. For the region classification, we first segment the low-resolution (LR) depth map into regions of smooth surfaces, and then use them to guide the segmentation of the color image. Using the consensus of multiple image segmentations obtained by different super-pixel generation methods, the color image is divided into continuous and discontinuous regions: in the continuous regions, their HR depth values are interpolated from LR depth samples without exploiting the color information. In the discontinuous regions, their HR depth values are estimated by sequentially applying more complicated depth-histogram-based methods. Through experiments, we show that each step of our method improves depth map upsampling both quantitatively and qualitatively. We also show that our method can be extended to handle real data with occluded regions caused by the displacement between color and depth sensors. Ouk Choi, Seung-Won Jung |
IEEE Trans. Image Process. | 2 |
| 2013 | Enhancement of Image and Depth Map Using Adaptive Joint Trilateral FilterabstractIn this paper, we present an adaptive joint trilateral filter (AJTF), which consists of domain, range, and depth filters. The AJTF is used for the joint enhancement of images and depth maps, which is achieved by suppressing the noise and sharpening the edges simultaneously. For improving the sharpness of the image and depth map, the AJTF parameters, the offsets, and the standard deviations of the range and depth filters are determined in such a way that image edges that match well with depth edges are emphasized. To this end, pattern matching between local patches in the image and depth map is performed and the matching result is utilized to adjust the AJTF parameters. Experimental results show that the AJTF produces sharpness-enhanced images and depth maps without overshoot and undershoot artifacts, while successfully reducing noise as well. A comparison of the performance of the AJTF with those of conventional image and depth enhancement algorithms shows that the proposed algorithm is effective. Seung-Won Jung |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | A Modified Model of the Just Noticeable Depth Difference and Its Application to Depth Sensation EnhancementabstractThe just noticeable depth difference (JNDD) describes the threshold of human perception of the difference in the depth. In flat-panel-based three-dimensional (3-D) displays, the JNDD is typically measured by changing the depth difference between displayed image objects until the difference is perceivable. However, not only the depth, but also the perceived size changes when the depth difference increases. In this paper, we present a modified JNDD measurement method that adjusts the physical size of the object such that the perceived size of the object is maintained. We then apply the proposed JNDD measurement method to depth sensation enhancement. When the depth value difference between the objects is increased to enable the viewer to perceive the depth difference, the size of the objects is adjusted to maintain the perceived size of the objects. In addition, since the size change of the objects can produce a whole region, a depth-adaptive hole-inpainting technique is proposed to compensate for the hole region with high accuracy. The experimental results demonstrate the effectiveness of the proposed method. Seung-Won Jung |
IEEE Trans. Image Process. | 1 |
| 2012 | Depth Map Based Image Enhancement Using Color StereopsisabstractColor stereopsis is a phenomenon in the human visual system (HVS) in which a long wavelength color is perceived as being located closer than a short wavelength color. Although many psychophysical studies on color stereopsis have been carried out, its practical application is not fully investigated. In this letter, we propose a new image enhancement algorithm using color stereopsis between red and blue colors. First, the relationship of red and blue colors to the depth perception is analyzed. Then, a simple but practical image enhancement algorithm is presented based on the analysis result. The experimental results demonstrate the effectiveness of the proposed algorithm. Seung-Won Jung, Sung-Jea Ko |
IEEE Signal Process. Lett. | 1 |
| 2012 | Sharpness Enhancement of Stereo Images Using Binocular Just-Noticeable DifferenceabstractIn this paper, we propose a new sharpness enhancement algorithm for stereo images. Although the stereo image and its applications are becoming increasingly prevalent, there has been very limited research on specialized image enhancement solutions for stereo images. Recently, a binocular just-noticeable-difference (BJND) model that describes the sensitivity of the human visual system to luminance changes in stereo images has been presented. We introduce a novel application of the BJND model for the sharpness enhancement of stereo images. To this end, an overenhancement problem in the sharpness enhancement of stereo images is newly addressed, and an efficient solution for reducing the overenhancement is proposed. The solution is found within an optimization framework with additional constraint terms to suppress the unnecessary increase in luminance values. In addition, the reliability of the BJND model is taken into account by estimating the accuracy of stereo matching. Experimental results demonstrate that the proposed algorithm can provide sharpness-enhanced stereo images without producing excessive distortion. Seung-Won Jung, Jae-Yun Jeong, Sung-Jea Ko |
IEEE Trans. Image Process. | 1 |
| 2012 | Depth Sensation Enhancement Using the Just Noticeable Depth DifferenceabstractIn this paper, we present a novel depth sensation enhancement algorithm considering the behavior of human visual system (HVS) toward stereoscopic image displays. On the basis of the recent studies on the just noticeable depth difference (JNDD), which represents a threshold that a human can perceive the depth difference between objects, we modify the depth image such that neighboring objects in the depth image can have a depth value difference of at least the JNDD. This modification is modeled via an energy minimization framework using three energy terms defined as depth data preservation, depth-order preservation, and depth difference expansion. The depth data term enforces minimal changes in the depth image with an additional weighting function that controls the direction of depth changes. The depth-order term restricts the inversion of the local and global depth orders among objects, and the JNDD term leads to an increase in the depth differences between segments. Throughout subjective quality evaluation on a stereoscopic image display, it is demonstrated that the human depth sensation is effectively improved by the proposed algorithm. Seung-Won Jung, Sung-Jea Ko |
IEEE Trans. Image Process. | 1 |
| 2011 | A New Histogram Modification Based Reversible Data Hiding Algorithm Considering the Human Visual SystemabstractIn this letter, we propose an improved histogram modification based reversible data hiding technique. In the proposed algorithm, unlike the conventional reversible techniques, a data embedding level is adaptively adjusted for each pixel with a consideration of the human visual system (HVS) characteristics. To this end, an edge and the just noticeable difference (JND) values are estimated for every pixel, and the estimated values are used to determine the embedding level. This pixel level adjustment can effectively reduce the distortion caused by data embedding. The experimental results and performance comparison with other reversible data hiding algorithms are presented to demonstrate the validity of the proposed algorithm. Seung-Won Jung, Sung-Jea Ko |
IEEE Signal Process. Lett. | 1 |
| 2010 | Fast Mode Decision Using All-Zero Block Detection for Fidelity and Spatial Scalable Video CodingabstractIn scalable video coding (SVC) as an extension of H.264/advanced video coding (AVC), a computationally expensive exhaustive mode decision is employed to select the best coding mode for each macroblock (MB). In order to reduce computational complexity, we propose a fast mode decision algorithm for SVC which uses an all-zero block (AZB) detection. Based on the empirical analysis of the inter-layer correlation of the AZB, the MB at the enhancement layer (EL) is predicted whether to be the AZB. Then, only predicted MBs are examined by the AZB detection algorithm. Since the mode decision can be terminated by detecting the AZB, we determine the processing order of the candidate modes at the EL in order to terminate the mode decision in the early stage. The proposed algorithm can be combined with other conventional fast mode decision methods in order to further reduce the computational complexity of those methods. Experimental results show that the proposed algorithm can significantly speed up the encoding process, especially for low bit-rate video sequences, without deteriorating the coding efficiency of SVC. Seung-Won Jung, Seung-Jin Baek, Chun-Su Park, Sung-Jea Ko |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | A Novel Multiple Image Deblurring Technique Using Fuzzy Projection onto Convex SetsabstractIn this letter, we present a novel image restoration algorithm to deblur the image without estimating the image blur. The proposed algorithm performs deblurring by merging differently blurred multiple images in the spectrum domain using the fuzzy projection onto convex sets (POCS). Experimental results demonstrate that the merged single image contains much less blur than the multiple blurred images. Seung-Won Jung, Sung-Jea Ko |
IEEE Signal Process. Lett. | 1 |
| 2009 | Estimation-Based Interlayer Intra Prediction for Scalable Video CodingabstractThe scalable video coding (SVC) standard adopts a simple interlayer intra prediction (ILIP) method for encoding scalable video sequences. In the conventional ILIP, a prediction signal for the macroblock (MB) of the enhancement layer (EL) is obtained by simply upsampling the colocated block of the base layer (BL). We propose an improved ILIP method by generalizing the original one adopted in the SVC standard. In the proposed ILIP method, the MB of the EL is predicted using all MBs of the BL. Experimental results show that the proposed algorithm can reduce the bit rate by 1.91% to 6.44%, as compared with the conventional ILIP, while the average PSNR is not decreased. Chun-Su Park, Seung-Jin Baek, Seung-Won Jung, Hye-Soo Kim, Sung-Jea Ko |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Overlapped Block Motion Compensation Based on Irregular GridabstractIn this work, we propose a hybrid motion compensation integrating the advantages of control grid interpolation (CGI) and overlapped block motion compensation (OBMC). We consider control points of CGI and overlapped window of OBMC as sampling points and spread function of motion vector, respectively. Then, a whole motion field in a frame is composed of motion vectors on sampling points and their spreadings. In this view-point, the conventional OBMC is considered as regular grid based OBMC (RG-OBMC) while the proposed OBMC is called irregular grid based OBMC (IG-OBMC). Experimental results demonstrate that the proposed IG-OBMC achieves improved visual quality as well as better PSNR performance than conventional motion compensation methods. Byeong-Doo Choi, Jong-Woo Han, Seung-Won Jung, Ju-Hun Nam, Sung-Jea Ko |
ICIP | 3 |
| 2006 | Improved Differential Energy Watermarking for Embedding Watermark
Goo-Rak Kwon, Seung-Won Jung, Sang-Jae Nam, Sung-Jea Ko |
IWDW | 2 |