Soonbin Lee

dblp:243/3649 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-8951-0335ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Computer networks · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CodecGS-IR: Implicit Representation for Decoder-friendly Video-based Gaussian Splat Compression
abstract
Coding of Gaussian splats has drawn the attention of academia and standardization bodies lately. A commonly used approach involves projecting explicit 3D attributes onto 2D planes of a video and compressing it with existing video coding standards. However, such an approach produces excessively high sample rates, resulting in a significant bottleneck that might make it unsuitable for existing hardware decoders if the number of splats is high. This paper presents a framework that utilizes an implicit representation, where Gaussian splat attributes are represented by compact feature planes. By reducing the dependency of video resolution on the number of splats, our approach significantly reduces the required sample rate, making it suitable for deployed devices. In addition, the proposed framework achieves a superior rate-distortion trade-off, providing high-fidelity reconstruction at low bitrates without the excessive sample rates associated with the conventional video-based anchor.
Soonbin Lee, Simon Sasse, Yago Sánchez de la Fuente, Robert Skupin, Tomas M. Borges, Cornelius Hellge, Thomas Schierl
DCC1
2025 Compression of 3D Gaussian Splatting with Optimized Feature Planes and Standard Video Codecs
abstract
3D Gaussian Splatting is a recognized method for 3D scene representation, known for its high rendering quality and speed. However, its substantial data requirements present challenges for practical applications. In this paper, we introduce an efficient compression technique that significantly reduces storage overhead by using compact representation. We propose a unified architecture that combines point cloud data and feature planes through a progressive tri-plane structure. Our method utilizes 2D feature planes, enabling continuous spatial representation. To further optimize these representations, we incorporate entropy modeling in the frequency domain, specifically designed for standard video codecs. We also propose channel-wise bit allocation to achieve a better trade-off between bitrate consumption and feature plane representation. Consequently, our model effectively leverages spatial correlations within the feature planes to enhance rate-distortion performance using standard, non-differentiable video codecs. Experimental results demonstrate that our method outperforms existing methods in data compactness while maintaining high rendering quality. Our project page is available at https://fraunhoferhhi.github.io/CodecGS
Soonbin Lee, Fangwen Shu, Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge
ICCV1
2024 ECRF: Entropy-Constrained Neural Radiance Fields Compression with Frequency Domain Optimization
abstract
Explicit feature-grid based NeRF models have shown promising results in terms of rendering quality and significant speed-up in training. However, these methods often require a significant amount of data to represent a single scene or object. In this work, we present a compression model that aims to minimize the entropy in the frequency domain in order to effectively reduce the data size. First, we propose using the discrete cosine transform (DCT) on the tensorial radiance fields to compress the feature-grid. This feature-grid is transformed into coefficients, which are then quantized and entropy encoded, following a similar approach to the traditional video coding pipeline. Furthermore, to achieve a higher level of sparsity, we propose using an entropy parameterization technique for the frequency domain, specifically for DCT coefficients of the feature-grid. Since the transformed coefficients are optimized during the training phase, the proposed model does not require any fine-tuning or additional information. Our model only requires a lightweight compression pipeline for encoding and decoding, making it easier to apply volumetric radiance field methods for real-world applications. Experimental results demonstrate that our proposed frequency domain entropy model can achieve superior compression performance across various datasets.
Soonbin Lee, Fangwen Shu, Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge
MMSP1
2024 DATRA-MIV: Decoder-Adaptive Tiling and Rate Allocation for MPEG Immersive Video
abstract
The emerging immersive video coding standard moving picture experts group (MPEG) immersive video (MIV), which is ongoing standardization by MPEG-Immersive (MPEG-I) group, enables six degrees of freedom in a virtual reality environment that represents both natural and computer-generated scenes using multi-view video compression. The MIV eliminates the redundancy between multi-view videos and merges the residuals into multiple pictures, called an atlas. Thus, bitstreams with encoded atlases are generated and corresponding number of decoders are needed, which is challenging for the lightweight device with a single decoder. This article proposes a decoder-adaptive tiling and rate allocation method for MIV to overcome the challenge. First, the proposed method divides atlases into subpictures considering two aspects: (i) subpicture bitstream extracting and merging into one bitstream to use a single decoder and (ii) separation of each source view from the atlases for rate allocation. Second, the atlases are encoded by versatile video coding (VVC), using an extractable subpicture to divide the atlases into subpictures. Third, each subpicture bitstream is extracted, and asymmetric quality allocation for each subpictures is conducted by considering the residuals in the subpicture. Fourth, mixed-quality subpictures were merged by using the proposed bitstream merger. Fifth, the merged bitstream is decoded by using a single decoder. Finally, the viewing area of the user is synthesized by using the reconstructed atlases. Experimental results with the VVC test model (VTM) show that the proposed method achieves a 21.37% Bjøntegaard delta rate saving for immersive video peak signal-to-noise ratio and a 26.76% decoding runtime saving compared to the VTM anchor configuration. Moreover, it supports bitstreams for multiple decoders and single decoder without re-encoding, transcoding, or a substantial increase of the server-side storage.
JongBeom Jeong, Soonbin Lee, Eun-Seok Ryu
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Implementing Partial Atlas Selector for Viewport-dependent MPEG Immersive Video Streaming
abstract
The ISO/IEC 23090-12 MPEG Immersive Video (MIV) standard technology, which provides immersive volumetric scenes with six degrees of freedom (6DoF), has recently been the subject of research and development efforts. The key concept of MIV technology is to generate an atlas that is a minimal representation of the multiple source view, with a low pixel rate to limit the number of existing video decoder instantiations. However, this atlas generation process produces dependencies between views in the reconstruction. This inability of conventional MIV to independently transmit and decode portions of the source view is a major challenge for 6DoF viewport-dependent streaming. This paper proposes a framework that can independently select and transmit only the atlas of the required area when rendering immersive content. This paper also presents a visibility calculation method to determine the importance of each atlas for viewport rendering. Experiments with a limited pixel rate under experimental conditions have shown that a highly efficient 6DoF viewport-dependent streaming system is achievable. The proposed method has been implemented with high-level syntax conformance in the MIV test model software, so this framework can be deployed with various adaptive streaming systems along with MIV bitstream in the future.
Soonbin Lee, JongBeom Jeong, Eun-Seok Ryu
NOSSDAV1
2023 Fine-grained Single-layer Tiling for Viewport-Adaptive 360-degree Video Streaming
abstract
Tile-based streaming is widely adopted for viewport-adaptive 360-degree video streaming due to its potential for bitrate reduction. However, for legacy devices equipped with a single decoder, managing multiple tile bitstreams poses a challenge, particularly with regards to synchronizing multiple decoders. This paper introduces Butterfly360, a fine-grained single-layer tiling method that advances two main ideas: (i) an optimal tiling scheme decision method for single-layer merging of tiles using various tile sizes, and (ii) a simple rate adaptation method for viewport tiles. Experimental results demonstrate that the proposed method offers advantages in terms of bitrate savings and reduction in decoding runtime. Furthermore, the proposed method facilitates the use of either multiple tile bitstreams or a single-layer merged bitstream.
JongBeom Jeong, Jun-Hyeong Park, Soonbin Lee, Eun-Seok Ryu
VCIP3
2023 Entropy-Constrained Implicit Neural Representations for Deep Image Compression
abstract
Implicit neural representations (INRs) for various data types have gained popularity in the field of deep learning owing to their effectiveness. However, previous studies on INRs have only focused on recovering original representations. This paper investigated an image compression model based on INRs using a model compression technique for entropy-constrained neural networks. Specifically, the proposed model trains a multilayer perceptron (MLP) to overfit a single image and then uses its weights to optimize its compressed representation using additive uniform noise. Accordingly, the proposed model efficiently minimizes the size of the model weight in an end-to-end manner. This training optimization process is fairly desirable for adjusting the rate of distortion for image compression. In contrast to other model compression techniques, the proposed model is implemented without additional training process or memory cost. By introducing entropy loss, this paper demonstrated that the proposed model can be used to preserve high image quality while maintaining smaller model size. The experimental results demonstrated that the proposed model achieved comparable performance to conventional image compression models without incurring high storage costs.
Soonbin Lee, JongBeom Jeong, Eun-Seok Ryu
IEEE Signal Process. Lett.1
2022 Atlas level rate distortion optimization for 6DoF immersive video compression
abstract
The Moving Picture Experts Group (MPEG) has started an immersive media standard project to enable multi-view video and depth representation in three-dimensional (3D) scenes. The MPEG immersive video (MIV) standard explores the six degree of freedom (6DoF) technologies of immersive content to support motion parallax. Despite the standard being designed to compress multi-view immersive media, MIV coding has not been investigated from the perspective of bit allocation. This paper presents an efficient bit allocation scheme for atlas level compression. The proposed model establishes a model of view synthesis distortion and analyzes the impact of distortion on complete views and patches. This paper also introduces packing alignment to separate two types of patches and characterize the distortion for each MIV atlas. By considering these characteristics, the proposed model derives a bitrate ratio between texture and geometry for model-based view-rendering optimization. Experimental results showed that the proposed method achieved a more accurate reconstruction of sequences under common test conditions (CTCs).
Soonbin Lee, JongBeom Jeong, Eun-Seok Ryu
NOSSDAV1
2021 DWS-BEAM: Decoder-Wise Subpicture Bitstream Extracting and Merging for MPEG Immersive Video
abstract
With the new immersive video coding standard MPEG immersive video (MIV) and versatile video coding (VVC), six degrees of freedom (6DoF) virtual reality (VR) streaming technology is emerging for both computer-generated and natural content videos. This paper addresses the decoder-wise subpicture bitstream extracting and merging (DWS-BEAM) method for MIV and proposes two main ideas: (i) a selective streaming-aware subpicture allocation method using a motion-constrained tile set (MCTS), (ii) a decoder-wise subpicture extracting and merging method for single-pass decoding. In the experiments using the VVC test model (VTM), the proposed method shows 1.23% BD-rate saving for immersive video PSNR (IV-PSNR) and 15.78% decoding runtime saving compared to the VTM anchor. Moreover, while the MIV test model requires four decoders, the proposed method only requires one decoder.
JongBeom Jeong, Soonbin Lee, Eun-Seok Ryu
VCIP2
2020 Towards Viewport-dependent 6DoF 360 Video Tiled Streaming for Virtual Reality Systems
abstract
Previous studies of 360-degree video streaming with regard to virtual reality allowed users to move their head freely, while their position is fixed according to the camera's location in virtual reality. One of the approaches to overcome the problem is transmitting multiview video to provide six degrees of freedom (6DoF). However, 6DoF streaming system implementation is challenging because multiple high-quality video streaming requires several decoders and a high bandwidth. Therefore, this paper proposes a viewport-dependent high-efficiency video coding (HEVC)-compliant tiled streaming system on test model for immersive video (TMIV), MPEG-Immersive multiview compression reference software. This paper proposes a 6DoF viewport tile selector (VTS) for multiple 360-degree video tiled streaming. Furthermore, this paper introduces a viewport-dependent multiple-tile extractor. The proposed system detects the user's head movement, selects the tile sets that correspond to the user's viewport, extracts tile bitstreams, and generates single bitstream. The extracted bitstream is transmitted and decoded to render the user's viewport The proposed viewport-dependent streaming method can reduce the decoding time as well as the bandwidth. Experimental results demonstrated 12.04% bjontegaard delta rate (BD-rate) saving for the luma peak signal-to-noise ratio (PSNR) compared to those obtained via the TMIV anchor without tiled encoding and a 55.51% decoding time saving compared to those obtained via the TMIV anchor with the existing tiled streaming method.
JongBeom Jeong, Soonbin Lee, Il-Woong Ryu, Tuan Thanh Le, Eun-Seok Ryu
ACM Multimedia2
2019 Motion-constrained tile set based 360-degree video streaming using saliency map prediction
abstract
In 360-degree video streaming, Most solutions are based on tile-based streaming that divides videos into tiles and streams the high-quality tiles corresponding to the user's viewport areas. However, these methods cannot transmit different combinations of tile coding efficiently. In this paper, we experimented with streaming 360-degree videos using a motion-constrained tile set (MCTS) technique that allows encoding with constraining motion vectors such that each tile can be decoded and transmitted independently. Moreover, we have used a tile-based approach using a saliency map that integrates the information of human visual attention with the contents to deliver high-quality tiles to the region of interest (ROI). We encoded the 360-degree videos at various quality representations with MCTS techniques and assigned a tile quality representation using a saliency map predicted by the existing convolutional neural network (CNN) model. We proposed a novel heuristic algorithm to assign appropriate quality to the tiles on the centerline. Consequently, mixed quality videos based on the saliency map enable efficient streaming in 360-degree videos. Using the Salient360! dataset, the proposed method shows an improvement in terms of bandwidth with little loss of viewport image quality.
Soonbin Lee, Dongmin Jang, JongBeom Jeong, Eun-Seok Ryu
NOSSDAV1