EDBT 2026 Demo / reviewers in the wild / expert
Seyoon Jeong
dblp:04/4661 · also Se-Yoon Jeong
· DBLP profile ↗
18ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-1675-4814ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gradient-Guided Diffusion-Based Restoration of Extremely Compressed Backgrounds for Video Coding for MachinesabstractVideo coding for machines (VCM) is an emerging approach in video compression designed to optimize content for machine analysis tasks. Although VCM was initially developed for machine vision, scalable coding frameworks have been developed to support both machine-driven analysis and human viewing as required. In this work, we focus on scenarios where high-quality encoding of regions of interest (ROIs) for machine vision and low-bitrate encoding of the background (BG) for human vision. At the decoder, severely degraded BG quality in reconstructed frames makes them unsuitable for viewing; therefore, restoring the degraded BGs by leveraging high-quality ROIs is essential. To this end, we propose the Gradient-Guided Diffusion Restoration (GGDR) algorithm, which integrates a pretrained generative diffusion model with content-aware supervision and adaptive refinement mechanisms to restore severely degraded regions robustly while maintaining visual consistency across the entire frame. The GGDR algorithm consists of two key components: (i) a content-aware supervision mechanism that preserves salient features and structural information in the input image, ensuring superior performance even with challenging high-variance inputs and (ii) a refinement block that guides the generation process of the pretrained diffusion model based on a degradation model and structural guidance. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art algorithms both qualitatively and quantitatively. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Naeun Yang, Chul Lee |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Content-Aware Supervision For Diffusion-Based Restoration of Extremely Compressed Background For VCMabstractWe propose content-aware supervision (CAS) techniques for diffusion-based restoration of an extremely compressed background for video coding for machines (VCM). First, we develop a CAS block to exploit prior information in an input image to reconstruct the noisy image, which is used as the input for the pretrained diffusion model. Then, we construct a refinement block to guide the pretrained diffusion model at each diffusion step by incorporating a degradation model and correction gradient estimation. Experimental results demonstrate the proposed algorithm outperforms state-of-the-art algorithms. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Naeun Yang, Chul Lee |
ICIP | 4 |
| 2024 | End-to-End Learnable Multi-Scale Feature Compression for VCMabstractThe proliferation of deep learning-based machine vision applications has given rise to a new type of compression, so called video coding for machine (VCM). VCM differs from traditional video coding in that it is optimized for machine vision performance instead of human visual quality. In the feature compression track of MPEG-VCM, multi-scale features extracted from images are subject to compression. Recent feature compression works have demonstrated that the versatile video coding (VVC) standard-based approach can achieve a BD-rate reduction of up to 96% against MPEG-VCM feature anchor. However, it is still sub-optimal as VVC was not designed for extracted features but for natural images. Moreover, the high encoding complexity of VVC makes it difficult to design a lightweight encoder without sacrificing performance. To address these challenges, we propose a novel multi-scale feature compression method that enables both the end-to-end optimization on the extracted features and the design of lightweight encoders. The proposed model combines a learnable compressor with a multi-scale feature fusion network so that the redundancy in the multi-scale features is effectively removed. Instead of simply cascading the fusion network and the compression network, we integrate the fusion and encoding processes in an interleaved way. Our model first encodes a larger-scale feature to obtain a latent representation and then fuses the latent with a smaller-scale feature. This process is successively performed until the smallest-scale feature is fused and then the encoded latent at the final stage is entropy-coded for transmission. The results show that our model outperforms previous approaches by at least 52% BD-rate reduction and has$\times 5$to$\times 27$times less encoding time for object detection. It is noteworthy that our model can attain near-lossless task performance with only 0.002-0.003% of the uncompressed feature data size. Yeongwoong Kim, Hyewon Jeong, Janghyun Yu, Younhee Kim, Jooyoung Lee 0004, Seyoon Jeong, Hui Yong Kim |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Restoration of Extremely Compressed Background for VCM Using Guided Generative PriorsabstractWe propose a learning-based image restoration algorithm for a single decoded image with a high-quality foreground and an extremely degraded background for video coding for machines (VCM). First, we develop an encoder that extracts multiscale features and learns latent vectors. Then, a background generator with style and feature fusion blocks generates guided features that contain the prior background information in the input image. Finally, the decoder restores the degraded background region by merging the image features from the encoder and prior background information from the generator. Experimental results show that the proposed algorithm achieves better performance than state-of-the-art algorithms. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Chul Lee |
ICIP | 4 |
| 2023 | Pixel-Unshuffled Multi-level Feature Map Compression for FCVCMabstractThe feature compression process for machine tasks involves several steps: feeding a video into the task network, extracting intermediate feature maps, compressing these maps into a bitstream on the client side, transmitting the bitstream to a resource-rich server, decoding it, and ultimately completing the specific machine task. In this paper, we present a multi-level feature compression method designed for machine tasks. We introduce an efficient and effective feature reshaping and merging module within the PCA-based feature coding scheme. This module utilizes pixel-unshuffled operations to reshape the multi-level features, merges them into a single map, and then performs a transformation. Our proposed method achieves a BD-rate gain of 49.69% and 66.3% in comparison to the previous computational low cost PCA-based feature coding method for object detection and instance segmentation tasks, respectively. Younhee Kim, Seyoon Jeong, Jooyoung Lee 0004, Jongseok Lee, Minsub Kim |
VCIP | 2 |
| 2023 | An Advanced Multi-Scale Feature Compression using Selective Learning Strategy for Video Coding for MachinesabstractMachine vision-based applications have witnessed widespread adoption in diverse fields. Efficiently processing and compressing the vast amounts of video data collected by machines is crucial for these applications. To address this need, the Moving Picture Experts Group (MPEG) is developing a new' video coding standard known as Video Coding for Machines (VCM), specifically optimized for video consumed by machines in vision applications. This paper proposes an advanced multi-scale feature compression (advMSFC) method with a selective learning strategy (SLS) for the Feature Compression for VCM (FCVCM). By applying the SLS, the proposed method converts multi-scale features into a single-scale feature arranged based on channel-wise importance, enabling adaptive feature channel truncation based on the QP. The truncated feature is efficiently compressed using the latest video codec, Versatile Video Coding (VVC). The proposed method outperforms the VCM feature anchor in instance segmentation, object detection, and object tracking tasks, achieving significant Bjontegaard delta-rate (BD-rate) gains. The adaptability of our method using a single trained model for various QPs shows promise for efficient feature compression in machine vision applications. Yong-Uk Yoon, Gyu-Woong Han, Jooyoung Lee 0004, Seyoon Jeong, Jae-Gon Kim |
VCIP | 4 |
| 2023 | MEDO: Minimizing Effective Distortions Only for Machine-Oriented Visual Feature CompressionabstractIn search for efficient feature compression technologies for machine consumption, MPEG recently issued a call for proposal (CfP) on feature compression for video coding for machine (FCVCM). One issue in feature compression is that the input feature maps generally have high redundancy in them. Various researches to reduce such redundancy have been made. For example, a recent study called L-MSFC (learnable multi-scale feature compression), which effectively combines multi-scale feature fusion and compression in an end-to-end learnable framework, showed up to 98% BD rate gain over the anchor model defined in the FCVCM CfP. Despite these advances in FCVCM, relation between distortions in feature maps and performance of vision tasks has stayed relatively unexplored. In this paper, we propose a novel loss function called MEDO (minimizing effective distortions only) based on our hypothesis that distortions below some threshold do not improve task performance. Experimental results on instance segmentation task show that our MEDO loss on top of L-MSFC improves the overall rate-mAP performance without compromising complexity. Being more practical for real-world uses, we also present an extension to L-MSFC for variable-rate support with a single model. Curie Yoon, Dalhong Lim, Yeongwoong Kim, Hyewon Jeong, Hui Yong Kim, Jooyoung Lee 0004, Younhee Kim, Seyoon Jeong |
VCIP | 8 |
| 2022 | Selective compression learning of latent representations for variable-rate image compressionabstractRecently, many neural network-based image compression methods have shown promising results superior to the existing tool-based conventional codecs. However, most of them are often trained as separate models for different target bit rates, thus increasing the model complexity. Therefore, several studies have been conducted for learned compression that supports variable rates with single models, but they require additional network modules, layers, or inputs that often lead to complexity overhead, or do not provide sufficient coding efficiency. In this paper, we firstly propose a selective compression method that partially encodes the latent representations in a fully generalized manner for deep learning-based variable-rate image compression. The proposed method adaptively determines essential representation elements for compression of different target quality levels. For this, we first generate a 3D importance map as the nature of input content to represent the underlying importance of the representation elements. The 3D importance map is then adjusted for different target quality levels using importance adjustment curves. The adjusted 3D importance map is finally converted into a 3D binary mask to determine the essential representation elements for compression. The proposed method can be easily integrated with the existing compression models with a negligible amount of overhead increase. Our method can also enable continuously variable-rate compression via simple interpolation of the importance adjustment curves among different quality levels. The extensive experimental results show that the proposed method can achieve comparable compression efficiency as those of the separately trained reference compression models and can reduce decoding time owing to the selective compression. Jooyoung Lee 0004, Seyoon Jeong, Munchurl Kim |
NeurIPS | 2 |
| 2022 | Modelling Surround-aware Contrast Sensitivity for HDR DisplaysabstractAbstract Despite advances in display technology, many existing applications rely on psychophysical datasets of human perception gathered using older, sometimes outdated displays. As a result, there exists the underlying assumption that such measurements can be carried over to the new viewing conditions of more modern technology. We have conducted a series of psychophysical experiments to explore contrast sensitivity using a state‐of‐the‐art HDR display, taking into account not only the spatial frequency and luminance of the stimuli but also their surrounding luminance levels. From our data, we have derived a novel surround‐aware contrast sensitivity function (CSF), which predicts human contrast sensitivity more accurately. We additionally provide a practical version that retains the benefits of our full model, while enabling easy backward compatibility and consistently producing good results across many existing applications that make use of CSF models. We show examples of effective HDR video compression using a transfer function derived from our CSF, tone‐mapping and improved accuracy in visual difference prediction. Shinyoung Yi 0001, Daniel S. Jeon, Ana Serrano, Seyoon Jeong, Hui Yong Kim, Diego Gutierrez, Min H. Kim 0001 |
Comput. Graph. Forum | 4 |
| 2022 | A Subjective and Objective Study of Space-Time Subsampled Video QualityabstractVideo dimensions are continuously increasing to provide more realistic and immersive experiences to global streaming and social media viewers. However, increments in video parameters such as spatial resolution and frame rate are inevitably associated with larger data volumes. Transmitting increasingly voluminous videos through limited bandwidth networks in a perceptually optimal way is a current challenge affecting billions of viewers. One recent practice adopted by video service providers is space-time resolution adaptation in conjunction with video compression. Consequently, it is important to understand how different levels of space-time subsampling and compression affect the perceptual quality of videos. Towards making progress in this direction, we constructed a large new resource, called the ETRI-LIVE Space-Time Subsampled Video Quality (ETRI-LIVE STSVQ) database, containing 437 videos generated by applying various levels of combined space-time subsampling and video compression on 15 diverse video contents. We also conducted a large-scale human study on the new dataset, collecting about 15,000 subjective judgments of video quality. We provide a rate-distortion analysis of the collected subjective scores, enabling us to investigate the perceptual impact of space-time subsampling at different bit rates. We also evaluated and compare the performance of leading video quality models on the new database. The new ETRI-LIVE STSVQ database is being made freely available at (https://live.ece.utexas.edu/research/ETRI-LIVE_STSVQ/index.html). Dae Yeol Lee, Somdyuti Paul, Christos G. Bampis, Hyunsuk Ko, Seyoon Jeong, Blake Homan, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2020 | Perceptual Video Coding using Deep Neural Network Based JND ModelabstractWe propose a perceptual video coding (PVC) method that uses the deep neural network (DNN) based just noticeable difference (JND) suppression model. The proposed JND suppression model's goal is to reduce the perceptual redundancy of the input video prior to the encoding process through a DNN, and further improve the compression efficiency while minimally affecting the perceptual quality. Dae Yeol Lee, Seyoon Jeong, Seunghyun Cho |
DCC | 3 |
| 2019 | A New No-Reference Method for Judder Artifact AssessmentabstractThis paper proposes a new metric to measure judder artifacts of video sequences. The judder artifacts appear as non-smooth motions in hold-type displays when the frame rate is low and object motion is fast. To analyze the judder artifacts in video sequences, the proposed judder metric is defined by analyzing the effects and cross-relations of judder features in the video sequences. The judder features include spatial features of image gradients and temporal features of motion vectors and the frame rate. In addition, sensitivity of the human visual system (HVS) is considered to determine the perceptual judder artifacts because it has special characteristics of sensitivity to image brightness and contrast. Therefore, a sensitivity map of the HVS is used to mask the judder artifacts. Then, the degree of perceptual judder artifacts is estimated using the judder features and a regression model. The experimental results demonstrate that the proposed judder metric is highly correlated with the subjective assessment results. Se Ri Oh, Seyoon Jeong, Pyeong Gang Heo, Hui Yong Kim, Hyun Wook Park |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Understanding and Removal of False Contour in HEVC Compressed ImagesabstractA contour-like artifact called false contour is often observed in large smooth areas of decoded images and video. Without loss of generality, we focus on detection and removal of false contours resulting from the state-of-the-art High Efficiency Video Coding codec. First, we identify the cause of false contours by explaining the human perceptual experiences on them with specific experiments. Next, we propose a precise pixel-based false contour detection method based on the evolution of a false contour candidate (FCC) map. The number of points in the FCC map becomes fewer by imposing more constraints step by step. Special attention is paid to separating false contours from real contours such as edges and textures in the video source. Then, a decontour method is designed to remove false contours in the exact contour position while preserving edge/texture details. Extensive experimental results are provided to demonstrate the superior performance of the proposed false contour detection and removal method in both compressed images and videos. Qin Huang 0006, Hui Yong Kim, Wen-Jiin Tsai, Seyoon Jeong, Jin Soo Choi, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Measure and Prediction of HEVC Perceptually Lossy/Lossless Boundary QP ValuesabstractEvaluation of coding efficiency is traditionally modeled as a continuous rate-distortion (R-D) function, where the peak signal-to-noise ratio (PSNR) is adopted as the quality measure. Although the PSNR-versus-bitrate curve offers some useful tradeoff information between video quality and coding bit-rates, it does not take human perceptual experience into account. In this work, by following the recent image/video quality assessment framework based on the just-noticeable-difference (JND) notion, we conduct a subjective test for HEVC (High Efficiency Video Codec) video to measure the QP value that lies in the boundary of perceptually lossless and lossy coded bit streams for each human subject. This is also known as the first JND point. It is observed that the statistics of the first JND points of 30 subjects follows the normal distribution for a great majority of test sequences. Finally, a machine-learning approach is proposed to predict the mean of the group-based JND distribution based on extracted video features. It is shown by experimental results that the mean JND point can be predicted accurately. Qin Huang 0006, Haiqiang Wang, Sung-Chang Lim, Hui Yong Kim, Seyoon Jeong, C.-C. Jay Kuo |
DCC | 5 |
| 2014 | Offset Compensation Method for Skip Mode in Hybrid Video CodingabstractThe skip mode has been adopted to reduce the bitrate in hybrid video coding such as H.264/MPEG-4 Advanced Video Coding (AVC). To improve the coding performance by reducing the dc distortion in the skip mode, we propose a skip with offset method, which uses offset compensation with the estimated offset only in the skip mode. In the proposed method, no flag bit is used for offset because incorrectly estimated offsets are mostly filtered out through the rate distortion mode decision. Only a semantic change is needed at the skip mode in the proposed method. The existing skip mode can therefore be easily replaced with the proposed method. The proposed method shows better coding performance than previous methods, including the weighted prediction method, in H.264/AVC. Furthermore, the proposed method shows a significant coding gain of -7.60% and -5.81% in its averaged bitrate differences compared, respectively, with baseline and high profiles of H.264/AVC at 720p high definition sequences. Seyoon Jeong, Hyun Wook Park |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Fast block mode decision scheme for B-picture coding in H.264/AVCabstractThe recent H.264/AVC video coding standard provides a higher coding efficiency than previous standards. H.264/AVC achieves a bit rate saving of more than 50 % with many new technologies, but it is computationally complex. Most of fast mode decision algorithms have focused on Baseline profile of H.264/AVC which does not consider B-picture coding. In this paper, a fast block mode decision scheme for B-pictures in High profile and Main profile is proposed to reduce the computational complexity for H.264/AVC. To reduce the block mode decision complexity in B-pictures of High profile, we use the SAD value after 16×16 block motion estimation. This SAD value is used for the classification feature to divide all block modes into some proper candidate search block modes. A differential mode allocation method is also used for the list (list 0, list 1) of B-slices based on the SAD value of the 16 × 16 block mode. The proposed algorithm shows the average speed-up factors of 41.9 ∼ 58.57% for IBBPBB sequences with a negligible bit increment and a minimal loss of image quality. Jong-Ho Kim, Hyo-Sung Kim, Byung-Gyu Kim, Hui Yong Kim, Seyoon Jeong, Jin Soo Choi |
ICIP | 5 |
| 2006 | Interactive Multi-View Visual Contents Authoring SystemabstractThis paper describes issues and consideration on authoring of interactive multi-view visual content based on MPEG-4. The issues include types of multi-view visual content; functionalities for user-interaction; scene composition for rendering; and multi-view visual content file format. The MPEG-4 standard, which aims to provide an object based audiovisual coding tool, has been developed to address the emerging needs from communications, interactive broadcasting as well as from mixed service models resulting from technological convergence. Due to the feature of object based coding, the use of MPEG-4 can resolve the format diversity problem of multi-view visual contents while providing high interactivity to users. Throughout this paper, we will present which issues need to be determined and how currently available tools can be effectively utilized for interactive multi-view visual content creation Injae Lee, Myungseok Ki, Seyoon Jeong, Kyuheon Kim |
ICME | 3 |
| 2004 | Image Navigation: A Massively Interactive Model for Similarity Retrieval of Images
Seyoon Jeong, Kyuseo Han, Byungtae Chun, Younglae J. Bae |
Int. J. Comput. Vis. | 2 |