Hyeongmin Lee

dblp:217/0406 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0001-5923-5419ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Probing intrinsic bias: Internal attention feature analysis for social bias evaluation in diffusion models
Hyeongmin Lee, Kyungjune Baek
Neurocomputing1
2025 Temporal Smoothness-Aware Rate-Distortion Optimized 4D Gaussian Splatting
abstract
Dynamic 4D Gaussian Splatting (4DGS) effectively extends the high-speed rendering capabilities of 3D Gaussian Splatting (3DGS) to represent volumetric videos. However, the large number of Gaussians, substantial temporal redundancies, and especially the absence of an entropy-aware compression framework result in large storage requirements. Consequently, this poses significant challenges for practical deployment, efficient edge-device processing, and data transmission. In this paper, we introduce a novel end-to-end RD-optimized compression framework tailored for 4DGS, aiming to enable flexible, high-fidelity rendering across varied computational platforms. Leveraging Fully Explicit Dynamic Gaussian Splatting (Ex4DGS), one of the state-of-the-art 4DGS methods, as our baseline, we start from the existing 3DGS compression methods for compatibility while effectively addressing additional challenges introduced by the temporal axis. In particular, instead of storing motion trajectories independently per point, we employ a wavelet transform to reflect the real-world smoothness prior, significantly enhancing storage efficiency. This approach yields significantly improved compression ratios and provides a user-controlled balance between compression efficiency and rendering quality. Extensive experiments demonstrate the effectiveness of our method, achieving up to 91$\times$ compression compared to the original Ex4DGS model while maintaining high visual fidelity. These results highlight the applicability of our framework for real-time dynamic scene rendering in diverse scenarios, from resource-constrained edge devices to high-performance environments. The source code is available at https://github.com/HyeongminLEE/RD4DGS.
Hyeongmin Lee, Kyungjune Baek
NeurIPS1
2024 CLIPtone: Unsupervised Learning for Text-Based Image Tone Adjustment
abstract
Recent image tone adjustment (or enhancement) approaches have predominantly adopted supervised learning for learning human-centric perceptual assessment. However, these approaches are constrained by intrinsic challenges of supervised learning. Primarily, the requirement for expertly-curated or retouched images escalates the data acquisition expenses. Moreover, their coverage of target styles is confined to stylistic variants inferred from the training data. To surmount the above challenges, we propose an unsupervised learning-based approach for text-based image tone adjustment, CLIPtone, that extends an existing image enhancement method to accommodate natural language descriptions. Specifically, we design a hyper-network to adaptively modulate the pretrained parameters of a back-bone model based on a text description. To assess whether an adjusted image aligns with its text description without a ground-truth image, we utilize CLIP, which is trained on a vast set of language-image pairs and thus encompasses the knowledge of human perception. The major advantages of our approach are threefold: (i) minimal data collection expenses, (ii) support for a range of adjustments, and (iii) the ability to handle novel text descriptions unseen in training. The efficacy of the proposed method is demonstrated through comprehensive experiments including a user study.
Hyeongmin Lee, Kyoungkook Kang, Jungseul Ok, Sunghyun Cho
CVPR1
2024 UGPNet: Universal Generative Prior for Image Restoration
abstract
Recent image restoration methods can be broadly categorized into two classes: (1) regression methods that recover the rough structure of the original image without synthesizing high-frequency details and (2) generative methods that synthesize perceptually-realistic high-frequency details even though the resulting image deviates from the original structure of the input. While both directions have been extensively studied in isolation, merging their benefits with a single framework has been rarely studied. In this paper, we propose UGPNet, a universal image restoration framework that can effectively achieve the benefits of both approaches by simply adopting a pair of an existing regression model and a generative model. UGPNet first restores the image structure of a degraded input using a regression model and synthesizes a perceptually-realistic image with a generative model on top of the regressed output. UGPNet then combines the regressed output and the synthesized output, resulting in a final result that faithfully reconstructs the structure of the original image in addition to perceptually-realistic textures. Our extensive experiments on deblurring, denoising, and super-resolution demonstrate that UGPNet can successfully exploit both regression and generative methods for high-fidelity image restoration.
Hwayoon Lee, Kyoungkook Kang, Hyeongmin Lee, Seung-Hwan Baek, Sunghyun Cho
WACV3
2024 A Nonlinear, Regularized, and Data-independent Modulation for Continuously Interactive Image Processing Network
Hyeongmin Lee, Taeoh Kim, Hanbin Son, Sangwook Baek, Minsu Cheon, Sangyoun Lee
Int. J. Comput. Vis.1
2023 Exploring Discontinuity for Video Frame Interpolation
abstract
Video frame interpolation (VFI) is the task that synthesizes the intermediate frame given two consecutive frames. Most of the previous studies have focused on appropriate frame warping operations and refinement modules for the warped frames. These studies have been conducted on natural videos containing only continuous motions. However, many practical videos contain various unnatural objects with discontinuous motions such as logos, user interfaces and subtitles. We propose three techniques that can make the existing deep learning-based VFI architectures robust to these elements. First is a novel data augmentation strategy called figure-text mixing (FTM) which can make the models learn discontinuous motions during training stage without any extra dataset. Second, we propose a simple but effective module that predicts a map called discontinuity map (D-map), which densely distinguishes between areas of continuous and discontinuous motions. Lastly, we propose loss functions to give supervisions of the discontinuous motion areas which can be applied along with FTM and D-map. We additionally collect a special test benchmark called Graphical Discontinuous Motion (GDM) dataset consisting of some mobile games and chatting videos. Applied to the various state-of-the-art VFI networks, our method significantly improves the interpolation qualities on the videos from not only GDM dataset, but also the existing benchmarks containing only continuous motions such as Vimeo90K, UCF101, and DAVIS.
Hyeongmin Lee, Chajin Shin, Hanbin Son, Sangyoun Lee
CVPR2
2022 Expanded Adaptive Scaling Normalization for End to End Image Compression
Chajin Shin, Hyeongmin Lee, Hanbin Son, Dogyoon Lee, Sangyoun Lee
ECCV (17)2
2022 Enhanced Standard Compatible Image Compression Framework Based on Auxiliary Codec Networks
abstract
Recent deep neural network-based research to enhance image compression performance can be divided into three categories: learnable codecs, postprocessing networks, and compact representation networks. The learnable codec has been designed for end-to-end learning beyond the conventional compression modules. The postprocessing network increases the quality of decoded images using example-based learning. The compact representation network is learned to reduce the capacity of an input image, reducing the bit rate while maintaining the quality of the decoded image. However, these approaches are not compatible with existing codecs or are not optimal for increasing coding efficiency. Specifically, it is difficult to achieve optimal learning in previous studies using a compact representation network due to the inaccurate consideration of the codecs. In this paper, we propose a novel standard compatible image compression framework based on auxiliary codec networks (ACNs). In addition, ACNs are designed to imitate image degradation operations of the existing codec, which delivers more accurate gradients to the compact representation network. Therefore, compact representation and postprocessing networks can be learned effectively and optimally. We demonstrate that the proposed framework based on the JPEG and High Efficiency Video Coding standard substantially outperforms existing image compression algorithms in a standard compatible manner.
Hanbin Son, Taeoh Kim, Hyeongmin Lee, Sangyoun Lee
IEEE Trans. Image Process.3
2021 Regularization Strategy for Point Cloud via Rigidly Mixed Sample
abstract
Data augmentation is an effective regularization strategy to alleviate the overfitting, which is an inherent drawback of the deep neural networks. However, data augmentation is rarely considered for point cloud processing despite many studies proposing various augmentation methods for image data. Actually, regularization is essential for point clouds since lack of generality is more likely to occur in point cloud due to small datasets. This paper proposes a Rigid Subset Mix (RSMix)1, a novel data augmentation method for point clouds that generates a virtual mixed sample by replacing part of the sample with shape-preserved subsets from another sample. RSMix preserves structural information of the point cloud sample by extracting subsets from each sample without deformation using a neighboring function. The neighboring function was carefully designed considering unique properties of point cloud, unordered structure and non-grid. Experiments verified that RSMix successfully regularized the deep neural networks with remarkable improvement for shape classification. We also analyzed various combinations of data augmentations including RSMix with single and multi-view evaluations, based on abundant ablation studies.
Dogyoon Lee, Jaeha Lee, Junhyeop Lee, Hyeongmin Lee, Minhyeok Lee, Sungmin Woo, Sangyoun Lee
CVPR4
2020 AdaCoF: Adaptive Collaboration of Flows for Video Frame Interpolation
abstract
Video frame interpolation is one of the most challenging tasks in video processing research. Recently, many studies based on deep learning have been suggested. Most of these methods focus on finding locations with useful information to estimate each output pixel using their own frame warping operations. However, many of them have Degrees of Freedom (DoF) limitations and fail to deal with the complex motions found in real world videos. To solve this problem, we propose a new warping module named Adaptive Collaboration of Flows (AdaCoF). Our method estimates both kernel weights and offset vectors for each target pixel to synthesize the output frame. AdaCoF is one of the most generalized warping modules compared to other approaches, and covers most of them as special cases of it. Therefore, it can deal with a significantly wide domain of complex motions. To further improve our framework and synthesize more realistic outputs, we introduce dual-frame adversarial loss which is applicable only to video frame interpolation tasks. The experimental results show that our method outperforms the state-of-the-art methods for both fixed training set environments and the Middlebury benchmark. Our source code is available at https://github.com/HyeongminLEE/AdaCoF-pytorch
Hyeongmin Lee, Taeoh Kim, Tae-Young Chung, Daehyun Pak, Yuseok Ban, Sangyoun Lee
CVPR1
2020 Extrapolative-Interpolative Cycle-Consistency Learning For Video Frame Extrapolation
abstract
Video frame extrapolation is a task to predict future frames when the past frames are given. Unlike previous studies that usually have been focused on the design of modules or construction of networks, we propose a novel ExtrapolativeInterpolative Cycle (EIC) loss using pre-trained frame interpolation module to improve extrapolation performance. Cycle-consistency loss has been used for stable prediction between two function spaces in many visual tasks. We formulate this cycle-consistency using two mapping functions; frame extrapolation and interpolation. Since it is easier to predict intermediate frames than to predict future frames in terms of the object occlusion and motion uncertainty, interpolation module can give guidance signal effectively for training the extrapolation function. EIC loss can be applied to any existing extrapolation algorithms and guarantee consistent prediction in the short future as well as long future frames. Experimental results show that simply adding EIC loss to the existing baseline increases extrapolation performance on both UCF101 [1] and KITTI [2] datasets.
Hyeongmin Lee, Taeoh Kim, Sangyoun Lee
ICIP2
2019 N-RPN: Hard Example Learning For Region Proposal Networks
abstract
The region proposal task is to generate a set of candidate regions that contain an object. In this task, it is most important to propose as many candidates of ground-truth as possible in a fixed number of proposals. In a typical image, however, there are too few hard negative examples compared to the vast number of easy negatives, so region proposal networks struggle to train on hard negatives. Because of this problem, networks tend to propose hard negatives as candidates, while failing to propose ground-truth candidates, which leads to poor performance. In this paper, we propose a Negative Region Proposal Network(nRPN) to improve Region Proposal Network(RPN). The nRPN learns from the RPN's false positives and provide hard negative examples to the RPN. Our proposed nRPN leads to a reduction in false positives and better RPN performance. An RPN trained with an nRPN achieves performance improvements on the PASCAL VOC 2007 dataset.
MyeongAh Cho, Tae-Young Chung, Hyeongmin Lee, Sangyoun Lee
ICIP3
2019 SF-CNN: A Fast Compression Artifacts Removal via Spatial-To-Frequency Convolutional Neural Networks
abstract
In this paper, we propose SF-CNN, a fast convolutional neural network structure for JPEG image compression artifacts removal. Recently, Convolutional Neural Network (CNN)-based image restoration has shown great performance improvement. However, its heavy computational cost makes it difficult to apply to other uses such as high-level vision tasks. Since heavy computation arises from maintaining the spatial resolution of an input image, some works make a structure that is composed of spatial downsampling and upsampling operations. SF-CNN takes Spatial input and predicts residual Frequency using downsampling operations only. Since every 8×8 pixel is grouped and spatially invariant in the JPEG DCT domain, it is possible to down sample the input by a factor of 8 to reduce the computational cost. We show this simple structure is effective for compression artifacts removal. Our scalable baseline networks achieve results comparable to to the reference networks in reduced computations.
Taeoh Kim, Hyeongmin Lee, Hanbin Son, Sangyoun Lee
ICIP2
2018 Collabonet: Collaboration of Generative Models by Unsupervised Classification
abstract
Designing models for learning dataset with complex distributions is one of the main challenges that still remains in machine learning areas. We propose CollaboNet, which can divide a large dataset into sub-datasets, train two generative models separately, and let two models work together to achieve better performance. The proposed algorithm divides a large dataset without label since the capability difference between two generative models in performing tasks on each data is the main criterion for dividing a large dataset. In other words, the classification model can be trained by unsupervised manner. Autoencoder experiments for pure MNIST and the datasets combined artificially from two image sets shows that CollaboNet successfully splits large datasets without labels, improving the performance of generative models.
Hyeongmin Lee, Taeoh Kim, Eungyeol Song, Sangyoun Lee
ICIP1