EDBT 2026 Demo / reviewers in the wild / expert
Jooyoung Lee 0004
dblp:10/1064-4
· DBLP profile ↗
13ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-0753-0699ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gradient-Guided Diffusion-Based Restoration of Extremely Compressed Backgrounds for Video Coding for MachinesabstractVideo coding for machines (VCM) is an emerging approach in video compression designed to optimize content for machine analysis tasks. Although VCM was initially developed for machine vision, scalable coding frameworks have been developed to support both machine-driven analysis and human viewing as required. In this work, we focus on scenarios where high-quality encoding of regions of interest (ROIs) for machine vision and low-bitrate encoding of the background (BG) for human vision. At the decoder, severely degraded BG quality in reconstructed frames makes them unsuitable for viewing; therefore, restoring the degraded BGs by leveraging high-quality ROIs is essential. To this end, we propose the Gradient-Guided Diffusion Restoration (GGDR) algorithm, which integrates a pretrained generative diffusion model with content-aware supervision and adaptive refinement mechanisms to restore severely degraded regions robustly while maintaining visual consistency across the entire frame. The GGDR algorithm consists of two key components: (i) a content-aware supervision mechanism that preserves salient features and structural information in the input image, ensuring superior performance even with challenging high-variance inputs and (ii) a refinement block that guides the generation process of the pretrained diffusion model based on a degradation model and structural guidance. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art algorithms both qualitatively and quantitatively. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Naeun Yang, Chul Lee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | TSDF Volume Compression Using Sampling-Enhancing Residual Block and Selective Latent Code EncodingabstractThis article introduces a novel deep learning method for efficiently compressing truncated signed distance function (TSDF) volumes. Previous works divide TSDF volumes into blocks and encode each block into the same number of latent codes, regardless of the geometric complexity stored in each block. This results in higher bitrates and increased complexity in arithmetic coding, as both complex and simple geometric blocks require the same number of latent codes. To address these inefficiencies, we propose a Hyperprior-based Latent Code Selection (HyperLCS) that dynamically adjusts the number of latent codes based on the geometric complexity of TSDF blocks. Through geometry-complexity-adaptive selective coding, HyperLCS reduces unnecessary bit allocation and arithmetic coding complexity, leading to improved coding efficiency, lower bitrates, and faster compression times. Furthermore, we introduce the Sampling-Enhancing Residual Block (SERB), a modified residual block designed to compensate for feature loss during spatial sampling by calculating residuals at the input resolution and adjusting the sampled output. Experimentally, SERB demonstrated improved TSDF volume compression performance compared to conventional residual blocks with the same capacity in terms of weight parameters. By combining HyperLCS and SERB, our method achieves superior TSDF volume compression performance, maintaining high data fidelity even at high compression rates. Experimental results demonstrate substantial improvements over existing techniques. Soowoong Kim, Jooyoung Lee 0004, Gun Bang, Seungjoon Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | DeepHQ: Learned Hierarchical Quantizer for Progressive Deep Image CodingabstractResearch on entropy model-based Learned Image Compression (LIC) has been actively progressing, leading to rapid advancements in coding efficiency. Beyond improvements in coding efficiency, LIC methods have also been explored for practical codec development. Despite these advancements, research on learned Progressive Image Coding (PIC) remains in its early stages. PIC aims to encode multiple quality levels into a single bitstream, improving bitstream versatility and achieving higher compression efficiency than simulcast compression. Existing learned PIC methods hierarchically quantize transformed latent representations with varying quantization step sizes. More specifically, these approaches progressively compress the additional information needed for quality improvement, considering that a wider quantization interval for lower-quality compression includes multiple narrower subintervals for higher-quality compression. However, they rely on handcrafted quantization hierarchies, leading to suboptimal compression efficiency. In this article, we propose a learned PIC method that first exploits learned quantization step sizes for each quantization layer. We also incorporate selective compression, ensuring that only essential representation components are retained in each quantization layer. Our experimental results demonstrate that the proposed method significantly enhances coding efficiency compared to the existing approaches while also reducing decoding time and model size. The source code is publicly available at https://github.com/JooyoungLeeETRI/DeepHQ . Jooyoung Lee 0004, Se Yoon Jeong, Munchurl Kim |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Content-Aware Supervision For Diffusion-Based Restoration of Extremely Compressed Background For VCMabstractWe propose content-aware supervision (CAS) techniques for diffusion-based restoration of an extremely compressed background for video coding for machines (VCM). First, we develop a CAS block to exploit prior information in an input image to reconstruct the noisy image, which is used as the input for the pretrained diffusion model. Then, we construct a refinement block to guide the pretrained diffusion model at each diffusion step by incorporating a degradation model and correction gradient estimation. Experimental results demonstrate the proposed algorithm outperforms state-of-the-art algorithms. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Naeun Yang, Chul Lee |
ICIP | 3 |
| 2024 | End-to-End Learnable Multi-Scale Feature Compression for VCMabstractThe proliferation of deep learning-based machine vision applications has given rise to a new type of compression, so called video coding for machine (VCM). VCM differs from traditional video coding in that it is optimized for machine vision performance instead of human visual quality. In the feature compression track of MPEG-VCM, multi-scale features extracted from images are subject to compression. Recent feature compression works have demonstrated that the versatile video coding (VVC) standard-based approach can achieve a BD-rate reduction of up to 96% against MPEG-VCM feature anchor. However, it is still sub-optimal as VVC was not designed for extracted features but for natural images. Moreover, the high encoding complexity of VVC makes it difficult to design a lightweight encoder without sacrificing performance. To address these challenges, we propose a novel multi-scale feature compression method that enables both the end-to-end optimization on the extracted features and the design of lightweight encoders. The proposed model combines a learnable compressor with a multi-scale feature fusion network so that the redundancy in the multi-scale features is effectively removed. Instead of simply cascading the fusion network and the compression network, we integrate the fusion and encoding processes in an interleaved way. Our model first encodes a larger-scale feature to obtain a latent representation and then fuses the latent with a smaller-scale feature. This process is successively performed until the smallest-scale feature is fused and then the encoded latent at the final stage is entropy-coded for transmission. The results show that our model outperforms previous approaches by at least 52% BD-rate reduction and has$\times 5$to$\times 27$times less encoding time for object detection. It is noteworthy that our model can attain near-lossless task performance with only 0.002-0.003% of the uncompressed feature data size. Yeongwoong Kim, Hyewon Jeong, Janghyun Yu, Younhee Kim, Jooyoung Lee 0004, Seyoon Jeong, Hui Yong Kim |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | COMPASS: High-Efficiency Deep Image Compression with Arbitrary-scale Spatial ScalabilityabstractRecently, neural network (NN)-based image compression studies have actively been made and has shown impressive performance in comparison to traditional methods. However, most of the works have focused on non-scalable image compression (single-layer coding) while spatially scalable image compression has drawn less attention although it has many applications. In this paper, we propose a novel NN-based spatially scalable image compression method, called COMPASS, which supports arbitrary-scale spatial scalability. Our proposed COMPASS has a very flexible structure where the number of layers and their respective scale factors can be arbitrarily determined during inference. To reduce the spatial redundancy between adjacent layers for arbitrary scale factors, our COMPASS adopts an inter-layer arbitrary scale prediction method, called LIFF, based on implicit neural representation. We propose a combined RD loss function to effectively train multiple layers. Experimental results show that our COMPASS achieves BD-rate gain of -58.33% and -47.17% at maximum compared to SHVC and the state-of-the-art NN-based spatially scalable image compression method, respectively, for various combinations of scale factors. Our COMPASS also shows comparable or even better coding efficiency than the single-layer coding for various scale factors. Jongmin Park 0001, Jooyoung Lee 0004, Munchurl Kim |
ICCV | 2 |
| 2023 | Restoration of Extremely Compressed Background for VCM Using Guided Generative PriorsabstractWe propose a learning-based image restoration algorithm for a single decoded image with a high-quality foreground and an extremely degraded background for video coding for machines (VCM). First, we develop an encoder that extracts multiscale features and learns latent vectors. Then, a background generator with style and feature fusion blocks generates guided features that contain the prior background information in the input image. Finally, the decoder restores the degraded background region by merging the image features from the encoder and prior background information from the generator. Experimental results show that the proposed algorithm achieves better performance than state-of-the-art algorithms. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Chul Lee |
ICIP | 3 |
| 2023 | Pixel-Unshuffled Multi-level Feature Map Compression for FCVCMabstractThe feature compression process for machine tasks involves several steps: feeding a video into the task network, extracting intermediate feature maps, compressing these maps into a bitstream on the client side, transmitting the bitstream to a resource-rich server, decoding it, and ultimately completing the specific machine task. In this paper, we present a multi-level feature compression method designed for machine tasks. We introduce an efficient and effective feature reshaping and merging module within the PCA-based feature coding scheme. This module utilizes pixel-unshuffled operations to reshape the multi-level features, merges them into a single map, and then performs a transformation. Our proposed method achieves a BD-rate gain of 49.69% and 66.3% in comparison to the previous computational low cost PCA-based feature coding method for object detection and instance segmentation tasks, respectively. Younhee Kim, Seyoon Jeong, Jooyoung Lee 0004, Jongseok Lee, Minsub Kim |
VCIP | 3 |
| 2023 | An Advanced Multi-Scale Feature Compression using Selective Learning Strategy for Video Coding for MachinesabstractMachine vision-based applications have witnessed widespread adoption in diverse fields. Efficiently processing and compressing the vast amounts of video data collected by machines is crucial for these applications. To address this need, the Moving Picture Experts Group (MPEG) is developing a new' video coding standard known as Video Coding for Machines (VCM), specifically optimized for video consumed by machines in vision applications. This paper proposes an advanced multi-scale feature compression (advMSFC) method with a selective learning strategy (SLS) for the Feature Compression for VCM (FCVCM). By applying the SLS, the proposed method converts multi-scale features into a single-scale feature arranged based on channel-wise importance, enabling adaptive feature channel truncation based on the QP. The truncated feature is efficiently compressed using the latest video codec, Versatile Video Coding (VVC). The proposed method outperforms the VCM feature anchor in instance segmentation, object detection, and object tracking tasks, achieving significant Bjontegaard delta-rate (BD-rate) gains. The adaptability of our method using a single trained model for various QPs shows promise for efficient feature compression in machine vision applications. Yong-Uk Yoon, Gyu-Woong Han, Jooyoung Lee 0004, Seyoon Jeong, Jae-Gon Kim |
VCIP | 3 |
| 2023 | MEDO: Minimizing Effective Distortions Only for Machine-Oriented Visual Feature CompressionabstractIn search for efficient feature compression technologies for machine consumption, MPEG recently issued a call for proposal (CfP) on feature compression for video coding for machine (FCVCM). One issue in feature compression is that the input feature maps generally have high redundancy in them. Various researches to reduce such redundancy have been made. For example, a recent study called L-MSFC (learnable multi-scale feature compression), which effectively combines multi-scale feature fusion and compression in an end-to-end learnable framework, showed up to 98% BD rate gain over the anchor model defined in the FCVCM CfP. Despite these advances in FCVCM, relation between distortions in feature maps and performance of vision tasks has stayed relatively unexplored. In this paper, we propose a novel loss function called MEDO (minimizing effective distortions only) based on our hypothesis that distortions below some threshold do not improve task performance. Experimental results on instance segmentation task show that our MEDO loss on top of L-MSFC improves the overall rate-mAP performance without compromising complexity. Being more practical for real-world uses, we also present an extension to L-MSFC for variable-rate support with a single model. Curie Yoon, Dalhong Lim, Yeongwoong Kim, Hyewon Jeong, Hui Yong Kim, Jooyoung Lee 0004, Younhee Kim, Seyoon Jeong |
VCIP | 6 |
| 2022 | Selective compression learning of latent representations for variable-rate image compressionabstractRecently, many neural network-based image compression methods have shown promising results superior to the existing tool-based conventional codecs. However, most of them are often trained as separate models for different target bit rates, thus increasing the model complexity. Therefore, several studies have been conducted for learned compression that supports variable rates with single models, but they require additional network modules, layers, or inputs that often lead to complexity overhead, or do not provide sufficient coding efficiency. In this paper, we firstly propose a selective compression method that partially encodes the latent representations in a fully generalized manner for deep learning-based variable-rate image compression. The proposed method adaptively determines essential representation elements for compression of different target quality levels. For this, we first generate a 3D importance map as the nature of input content to represent the underlying importance of the representation elements. The 3D importance map is then adjusted for different target quality levels using importance adjustment curves. The adjusted 3D importance map is finally converted into a 3D binary mask to determine the essential representation elements for compression. The proposed method can be easily integrated with the existing compression models with a negligible amount of overhead increase. Our method can also enable continuously variable-rate compression via simple interpolation of the importance adjustment curves among different quality levels. The extensive experimental results show that the proposed method can achieve comparable compression efficiency as those of the separately trained reference compression models and can reduce decoding time owing to the selective compression. Jooyoung Lee 0004, Seyoon Jeong, Munchurl Kim |
NeurIPS | 1 |
| 2019 | Context-adaptive Entropy Model for End-to-end Optimized Image Compression
Jooyoung Lee 0004, Seunghyun Cho, Seung-Kwon Beack |
ICLR (Poster) | 1 |
| 2018 | GPU-based real-time super-resolution system for high-quality UHD video up-conversion
Dae Yeol Lee, Jooyoung Lee 0004, Ji-Hoon Choi, Jong-Ok Kim, Hui Yong Kim, Jin Soo Choi |
J. Supercomput. | 2 |