EDBT 2026 Demo / reviewers in the wild / expert
Gun Bang
dblp:202/8718
· DBLP profile ↗
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-4355-599XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MPEG Explorations Toward 3D Gaussian Splat Coding and Standardizationabstract3D Gaussian splats (3DGS) have rapidly gained traction as a 3D scene representation technique that enables efficient real-time rendering and highfidelity novel view synthesis. This paper reports on ongoing MPEG GSC efforts conducted jointly by the MPEG Video Coding group (WG 4) and the Coding of 3D Graphics and Haptics group (WG 7) to define a practical and interoperable compression framework for 3DGS. MPEG GSC is planning GSC standardization with short-term and long-term timelines to address market requirements. The short-term objective is to standardize coding tools that build on proven MPEG ecosystems while introducing only the minimal set of extensions, syntax, and processing required for INRIA-3DGS format (referred to in MPEG as I-3DGS). In the long-term, MPEG is also investigating broader alternatives for 3DGS representation and compression, including approaches that integrate training during compression, collectively referred to as Alternative-3DGS (A-3DGS). More specifically, this paper focuses on the I-3DGS and explores both geometrybased and video-based coding frameworks within MPEG GSC. Gun Bang, Yiyi Liao, Alexandre Zaghetto, Marius Preda, Lu Yu 0003 |
DCC | 1 |
| 2026 | TSDF Volume Compression Using Sampling-Enhancing Residual Block and Selective Latent Code EncodingabstractThis article introduces a novel deep learning method for efficiently compressing truncated signed distance function (TSDF) volumes. Previous works divide TSDF volumes into blocks and encode each block into the same number of latent codes, regardless of the geometric complexity stored in each block. This results in higher bitrates and increased complexity in arithmetic coding, as both complex and simple geometric blocks require the same number of latent codes. To address these inefficiencies, we propose a Hyperprior-based Latent Code Selection (HyperLCS) that dynamically adjusts the number of latent codes based on the geometric complexity of TSDF blocks. Through geometry-complexity-adaptive selective coding, HyperLCS reduces unnecessary bit allocation and arithmetic coding complexity, leading to improved coding efficiency, lower bitrates, and faster compression times. Furthermore, we introduce the Sampling-Enhancing Residual Block (SERB), a modified residual block designed to compensate for feature loss during spatial sampling by calculating residuals at the input resolution and adjusting the sampled output. Experimentally, SERB demonstrated improved TSDF volume compression performance compared to conventional residual blocks with the same capacity in terms of weight parameters. By combining HyperLCS and SERB, our method achieves superior TSDF volume compression performance, maintaining high data fidelity even at high compression rates. Experimental results demonstrate substantial improvements over existing techniques. Soowoong Kim, Jooyoung Lee 0004, Gun Bang, Seungjoon Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | TSDF-Based Efficient Motion-Compensated Temporal Interpolation for 3D Dynamic SequencesabstractThis paper introduces a method for efficiently interpolating 3D dynamic sequences using truncated signed distance function (TSDF) volumes. The method calculates bi-directional motions between TSDF volumes of two frames and refines them to reconstruct intermediate frames. Unlike point cloud-based methods, which can suffer from varying and irregular point densities, the uniform and dense grid structure of TSDF offers a consistent framework for estimating the true motion of objects within a scene. In our experiments, the TSDF-based method offers more precise and reliable smooth motion prediction compared to the often error-prone surface depiction in point clouds. Experimental results demonstrate improved accuracy and reduced computational complexity, making it suitable for real-time applications. Soowoong Kim, Minseong Kwon, Gun Bang, Seungjoon Yang |
AAAI | 4 |
| 2025 | DA4NeRF: Depth-aware Augmentation technique for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) demonstrate impressive capabilities in rendering novel views of specific scenes by learning an implicit volumetric representation from posed RGB images without any depth information. View synthesis is the computational process of synthesizing novel images of a scene from different viewpoints, based on a set of existing images. One big problem is the need for a large number of images in the training datasets for neural network-based view synthesis frameworks. The challenge of data augmentation for view synthesis applications has not been addressed yet. NeRF models require comprehensive scene coverage in multiple views to accurately estimate radiance and density at any point. In cases without sufficient coverage of scenes with different viewing directions, cannot effectively interpolate or extrapolate unseen scene parts. In this paper, we introduce a new pipeline to tackle this data augmentation problem using depth data. We use MPEG's Depth Estimation Reference Software and Reference View Synthesizer to add novel non-existent views to the training sets needed for the NeRF framework. Experimental results show that our approach improves the quality of the rendered images using NeRF's model. The average quality increased by 6.4 dB in terms of Peak Signal-to-Noise Ratio (PSNR), with the highest increase being 11 dB. Our approach not only adds the ability to handle the sparsely captured multiview content to be used in the NeRF framework, but also makes NeRF more accurate and useful for creating high-quality virtual views. Hamed Razavi Khosroshahi, Jaime Sancho, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Neural Volumetric Video Coding With Hierarchical Coded Representation of Dynamic VolumeabstractThis article proposes a novel multi-view (MV) video coding technique that leverages a four-dimensional (4D) voxel-grid representation to enhance coding efficiency, particularly in novel view synthesis. Although the voxel grid approximation provides a continuous representation for dynamic scenes, its volumetric nature requires substantial storage. The compression of MV videos can be interpreted as the compression of dense features. However, the substantial size of these features poses a significant problem relative to the generation of dynamic scenes at arbitrary viewpoints. To address this challenge, this study introduces a hierarchical coded representation of dynamic volumes based on low-rank tensor decomposition of volumetric features and develops effective coding techniques based on this representation. The proposed method employs a two-level coding strategy to capture the temporal characteristics of the decomposed features. At a higher level, spatial features are encoded, representing 3D structural information, with time-invariant components over short intervals of an MV video sequence. At a lower level, temporal features are encoded to capture the dynamics of current scenes. The spatial features are shared in a group, and temporal features are encoded at each time step. The experimental results demonstrate that the proposed technique outperforms existing MV video coding standards and current state-of-the-art methods, providing superior rate-distortion performance in the novel view synthesis of MV video compression. Ju-Yeon Shin, Jung-Kyung Lee, Gun Bang, Junsik Kim 0002, Je-Won Kang |
IEEE Trans. Multim. | 3 |
| 2024 | Sync-NeRF: Generalizing Dynamic NeRFs to Unsynchronized VideosabstractRecent advancements in 4D scene reconstruction using neural radiance fields (NeRF) have demonstrated the ability to represent dynamic scenes from multi-view videos. However, they fail to reconstruct the dynamic scenes and struggle to fit even the training views in unsynchronized settings. It happens because they employ a single latent embedding for a frame while the multi-view images at the same frame were actually captured at different moments. To address this limitation, we introduce time offsets for individual unsynchronized videos and jointly optimize the offsets with NeRF. By design, our method is applicable for various baselines and improves them with large margins. Furthermore, finding the offsets always works as synchronizing the videos without manual effort. Experiments are conducted on the common Plenoptic Video Dataset and a newly built Unsynchronized Dynamic Blender Dataset to verify the performance of our method. Project page: https://seoha-kim.github.io/sync-nerf Seoha Kim, Jeongmin Bae 0001, Youngsik Yun, Hahyun Lee, Gun Bang, Youngjung Uh |
AAAI | 5 |
| 2024 | Per-Gaussian Embedding-Based Deformation for Deformable 3D Gaussian Splatting
Jeongmin Bae 0001, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, Youngjung Uh |
ECCV (15) | 5 |
| 2024 | A Practical Approach to Depth-Aware Augmentation for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have demonstrated exceptional performance in generating novel views of scenes by learning implicit volumetric representations from calibrated RGB images, without depth information. A major limitation is the need for large training datasets in neural network-based view synthesis frameworks. The challenge of effective data augmentation for view synthesis remains unresolved. NeRF models require extensive scene coverage from multiple views to accurately estimate radiance and density. Insufficient coverage reduces the model’s ability to interpolate or extrapolate unseen parts of the scene effectively. In this paper, we propose a novel pipeline to address this data augmentation issue using depth map information. We use depth image-based rendering (DIBR) to overcome the lack of enough views for training NeRF. Experimental results indicate that our approach enhances the quality of rendered images using the NeRF framework, achieving an average peak signal-to-noise ratio (PSNR) increase of 7.2 dB, with a maximum improvement of 12 dB. Hamed Razavi Khosroshahi, Jaime Sancho, Daniele Bonatto, Sarah Fachada, Gun Bang, Gauthier Lafruit, Eduardo Juárez Martínez, Mehrdad Teratani |
VCIP | 5 |
| 2024 | Dynamic Volumetric Video Coding with Tensor DecompositionabstractRecently, volumetric video coding based on neural radiance fields has gained significant attention for storing and transmitting three-dimensional (3D) scenes captured from multi-view video. Because the neural networks are trained to produce novel view synthesis of surrounding 3D scenes, compressing the model and then rendering the colors and geometry through the decompressed model can be utilized as a 3D video coding system. However, although this approach provides superior performance compared to conventional 3D video coding standards using depth video, challenges remain in reducing overall model sizes to improve coding efficiency. In this paper, we propose a novel dynamic volumetric video coding technique that employs a Group of Volume (GoV) to divide multi-view video sequences into smaller chunks, addressing complex temporal dynamics. Our method uses volumetric video features represented with 3D spatial and temporal tensor matrices and vectors and encodes them with the GoVs. The tensors are compressed by existing 2D video codec, allowing for fast rendering and easing deployment. Experimental results validate that our method not only reduces memory footprint but also maintains high-quality rendering as compared to state-of-the-art studies. Ju-Yeon Shin, Yeoneui Kim, Je-Won Kang, Gun Bang |
VCIP | 4 |
| 2024 | Efficient immersive video coding using specular detection for high rendering quality
Yongho Choi, The Van Le, Gun Bang |
Multim. Tools Appl. | 3 |
| 2006 | Data Broadcasting and Interactive TelevisionabstractThis paper provides an overview of the digital television (DTV) data broadcast service and interactive service technologies that have been deployed over the last ten years. We show how these trials have led to the development of data protocol and software middleware specifications, worldwide. Particular attention is given to the series of standards established by the Advanced Television System Committee. Experimental deployments to both Personal Computer(PC) and Set-Top-Box (STB)/spl I.bar/receivers are considered, with an emphasis on the services that have introduced new business models for DTV operators. Regis J. Crinon, Dinkar Bhat, David Catapano, Gomer Thomas, James van Loo, Gun Bang |
Proc. IEEE | 6 |