VLDB 2026 Research / reviewers in the wild / expert
Ronggang Wang
dblp:64/6287
· DBLP profile ↗
17ranked-venue papers in the field
0as first author
13since 2021 · last 2025
0000-0003-0873-0465ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 15Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Feature Prediction for 3D Gaussian Splatting CompressionabstractRecently, 3D Gaussian Splatting (3DGS) has emerged as a promising scene representation technique for novel view synthesis. However, the large number of Gaussians used to represent the 3D scene poses challenges for storage and transmission. Existing advanced compression methods focus on designing various context models for entropy modeling, but lack the exploration of prediction techniques. In this paper, we dig into feature correlations in the anchor-based Gaussian representation and propose two types of feature prediction techniques to further reduce the scene redundancy. Firstly, we observe that context features contain rich scene priors, which are also helpful for reconstruction but often ignored by previous methods. To this end, we design a Context-based Weighted Prediction module to adaptively aggregate the anchor feature and the context feature for rendering, which can reduce the storage costs in anchors. Secondly, a high degree of similarity is discovered between different feature channels. To utilize the cross-channel correlations, we propose a Cross-channel Residual Prediction module, which further reduces the bit cost for coding anchor features. Extensive experiments show that our method can further enhance compression performance while maintaining rendering quality compared to existing 3DGS compression methods. Our code is available at https://github.com/Pomelomm/FP-GS. Luyang Tang, Yongqi Zhai, Chunhui Yang, Ronggang Wang |
DCC | 5 |
| 2025 | MVCNet: An End to End Network for Multi-View Video CodingabstractThe rapid advancement of immersive visual applications has drawn significant attention to multi-view video compression. However, no end-to-end learning compression model is proposed for multi-view video sequences with six degrees of freedom. In this paper, we first propose an end-to-end model MVCNet to enhance multi-view video compression performance as shown in Fig. 1. MVCNet eliminates spatial and temporal redundancy in multi-view data effectively. In our methods, an efficient encoding structure is designed for compression, which utilizes spatial and temporal information among frames and views to improve compression performance. Furthermore, we propose a hybrid prediction module, which combines different prediction methods to provide satisfactory images and reduce the bit rate. Besides, we demonstrate a strategy of the fusion network to perform adaptive reconstruction. Chunhui Yang, Luyang Tang, Yongqi Zhai, Ronggang Wang |
DCC | 4 |
| 2025 | Compressed 3D Gaussian Splatting Model with Residual RenderingabstractRecently, 3D Gaussian Splatting (3D-GS) techniques have effectively driven the development of novel view synthesis due to the fast rendering speed and high-quality rendering. In this paper, we propose a compressed 3D Gaussian splatting model with residual rendering to enhance rendering quality and reduce storage. In our model, we design a residual rendering strategy to enrich the scene details, which leverages residual Gaussian points to optimize basic Gaussian points and supplement missing information. Additionally, an octree-based geometric compression model is introduced to compress geometric location. Chunhui Yang, Luyang Tang, Yongqi Zhai, Ronggang Wang |
DCC | 4 |
| 2025 | DeepFGS: Fine-Grained Scalable Coding for Learned Image CompressionabstractScalable coding, which can adapt to channel bandwidth variation, performs well in today's complex network environment. However, most existing scalable compression methods face two challenges: reduced compression performance and insufficient scalability. To overcome the above problems, this paper proposes a learned fine-grained scalable image compression framework, namely DeepFGS. Specifically, we introduce a feature separation backbone to divide the image information into basic and scalable features, then redistribute the features channel by channel through an information rearrangement strategy. In this way, we can generate a continuously scalable bitstream via one-pass encoding. For entropy coding, we design a mutual entropy model to fully explore the correlation between the basic and scalable features. In addition, we reuse the decoder to reduce the parameters and computational complexity. Experiments demonstrate that our proposed DeepFGS outperforms previous learning-based scalable image compression models and traditional scalable image codecs in both PSNR and MS-SSIM metrics. Yongqi Zhai, Luyang Tang, Wei Jiang 0031, Ronggang Wang |
DCC | 5 |
| 2025 | L-LBVC: Long-Term Motion Estimation and Prediction for Learned Bi-Directional Video CompressionabstractRecently, learned video compression (LVC) has shown superior performance under lowdelay configuration. However, the performance of learned bi-directional video compression (LBVC) still lags behind traditional bi-directional coding. The performance gap mainly arises from inaccurate long-term motion estimation and prediction of distant frames, especially in large motion scenes. To solve these two critical problems, this paper proposes a novel LBVC framework, namely L-LBVC. Firstly, we propose an adaptive motion estimation module that can handle both short-term and long-term motions. Specifically, we directly estimate the optical flows for adjacent frames and non-adjacent frames with small motions. For non-adjacent frames with large motions, we recursively accumulate local flows between adjacent frames to estimate long-term flows. Secondly, we propose an adaptive motion prediction module that can largely reduce the bit cost for motion coding. To improve the accuracy of long-term motion prediction, we adaptively downsample reference frames during testing to match the motion ranges observed during training. Experiments show that our L-LBVC significantly outperforms previous state-of-the-art LVC methods and even surpasses VVC (VTM) on some test datasets under random access configuration. Yongqi Zhai, Luyang Tang, Wei Jiang 0031, Ronggang Wang |
DCC | 5 |
| 2025 | MLIIC: Meta-Learned Implicit Image Codec with 15× Faster Encoding Speed and Higher PerformanceabstractImplicit Neural Representation (INR) has introduced a novel paradigm for image compression, achieving competitive Rate-Distortion (RD) performance with low decoding complexity. Existing INR-based codecs typically comprise three core components: (1) Multilayer Perceptron (MLP) networks, (2) an entropy coding module, and (3) a set of latent grids. Encoding a specific image involves overfitting these components to the image. However, current approaches often initiate overfitting from scratch, utilizing random or zero-initialized parameters. This approach necessitates tens of minutes to several hours for full overfitting, rendering it highly inefficient and impractical. To address this limitation, we propose MLIIC: a Meta-Learned Implicit Image Codec built upon the state-of-the-art INR-based image codec. Our enhanced meta-learning methodology provides a generalizable initialization that reduces baseline encoding time by an order of magnitude. Empirical results demonstrate that MLIIC not only achieves more than 15 × faster encoding speed but also exhibits superior RD performance compared to baseline initialized with random or zero parameters. Wei Jiang 0031, Yongqi Zhai, Ronggang Wang |
DCC | 5 |
| 2024 | An imperceptible adversarial attack against reconstruction for learned image compressionabstractLearned image compression has achieved better performance than traditional coding methods in terms of rate-distortion performance. However, the robustness of compression models themselves is rarely paid attention to by coding community. In this work, we explore the potential threats of image compression model, and design an imperceptible adversarial perturbation generation method based on gradient optimization. The image with our generated adversarial perturbation will lead to serious distortion on decoder side when the image is reconstructed. Specifically, we use a similar method based on Fast Gradient Sign Method (FGSM) to optimize a noise and generate an adversarial perturbation against image reconstruction. Furthermore, in order to improve the imperceptibility of our attack, we restrict the optimized noise to the high frequency region of the chrominance components of a YUV image, inspired by the characteristics of human vision system (HVS). See Figure 1 for more details. Experiments on four types of popular image compression models show that our adversarial attack can cause serious distortion on decoder side of the model while keeping the perturbation undetectable to human eyes. We hope that our work could arouse the concern of coding community to the robustness and security of AI intelligent coding technology. Jingui Ma, Ronggang Wang |
DCC | 2 |
| 2024 | SSNVC: Single Stream Neural Video Compression with Implicit Temporal InformationabstractNeural Video Compression (NVC) techniques have achieved remarkable performance, even surpassing the best traditional lossy video codec. However, most existing NVC methods [1] heavily rely on transmitting Motion Vector (MV) to generate accurate contextual features, which has following drawbacks. (1) Compressing and transmitting MV requires specialized MV encoder and decoder, which makes modules redundant. (2) Due to the existence of MV Encoder-Decoder, the training strategy is complex. In this paper, we propose Single Stream Neural Video Compression, SS-NVC. It implicitly utilizes temporal information to eliminate temporal redundancy in video sequence. Without MV encoder-decoder [2] , it only needs to transmit single bit-stream in channel and use single-stage training strategy, which can greatly simplify training and compression process of NVC. Besides, we reimplement window-based attention intra-frame image compression with channel-wise and checkerboard auto-regression entropy model, enhance contextual encoder with mixing global and local context module, and redesign Dense-UNet frame generator with stronger generation capability to improve SSNVC’s compression performance. Experiment results show that SSNVC can achieve competitive performance on multiple benchmarks. Haihang Ruan, Zhihuang Xie, Ronggang Wang, Xiangyu Yue 0001 |
DCC | 4 |
| 2024 | UCVC: A Unified Contextual Video Compression Framework with Joint P-frame and B-frame CodingabstractThis paper presents a learned video compression method in response to video compression track of the 6th Challenge on Learned Image Compression (CLIC), at DCC 2024. Specifically, we propose a unified contextual video compression framework (UCVC) for joint Pframe and B-frame coding. Each non-intra frame refers to two neighboring decoded frames, which can be either both from the past for P-frame compression, or one from the past and one from the future for B-frame compression. In training stage, the model parameters are jointly optimized with both P-frames and B-frames. Benefiting from the designs, the framework can support both P-frame and B-frame coding and achieve comparable compression efficiency with that specifically designed for P-frame or B-frame. As for challenge submission, we report the optimal compression efficiency by selecting appropriate frame types for each test sequence. Our team name is PKUSZ-LVC. Wei Jiang 0031, Yongqi Zhai, Chunhui Yang, Ronggang Wang |
DCC | 5 |
| 2024 | Hybrid Local-Global Context Learning for Neural Video CompressionabstractIn neural video codecs, current state-of-the-art methods typically adopt multi-scale motion compensation to handle diverse motions. These methods estimate and compress either optical flow or deformable offsets to reduce inter-frame redundancy. However, flow-based methods often suffer from inaccurate motion estimation in complicated scenes. Deformable convolution-based methods are more robust but have a higher bit cost for motion coding. In this paper, we propose a hybrid context generation module, which combines the advantages of the above methods in an optimal way and achieves accurate compensation at a low bit cost. Specifically, considering the characteristics of features at different scales, we adopt flow-guided deformable compensation at largest-scale to produce accurate alignment in de-tailed regions. For smaller-scale features, we perform flow-based warping to save the bit cost for motion coding. Furthermore, we design a local-global context enhancement module to fully explore the local-global information of previous reconstructed signals. Experimental results demonstrate that our proposed Hybrid Local-Global Context learning (HLGC) method can significantly enhance the state-of-the-art methods on standard test datasets. Yongqi Zhai, Wei Jiang 0031, Chunhui Yang, Luyang Tang, Ronggang Wang |
DCC | 6 |
| 2023 | Butterfly: Multiple Reference Frames Feature Propagation Mechanism for Neural Video CompressionabstractUsing more reference frames can significantly improve the compression efficiency in neural video compression. However, in low-latency scenarios, most existing neural video compression frameworks usually use the previous one frame as reference. Or a few frameworks which use the previous multiple frames as reference only adopt a simple multi-reference frames propagation mechanism. In this paper, we present a more reasonable multi-reference frames propagation mechanism for neural video compression, called butterfly multi-reference frame propagation mechanism (Butterfly), which allows a more effective feature fusion of multireference frames. By this, we can generate more accurate temporal context conditional prior for Contextual Coding Module. Besides, when the number of decoded frames does not meet the required number of reference frames, we duplicate the nearest reference frame to achieve the requirement, which is better than duplicating the furthest one. Experiment results show that our method can significantly outperform the previous state-of-the-art (SOTA), and our neural codec can achieve -7.6% bitrate save on HEVC Class D dataset when compares with our base single-reference frame model with the same compression configuration. Haihang Ruan, Litian Li, Ronggang Wang |
DCC | 6 |
| 2022 | Evaluating the Throughput of Video Transcoding in Cloud ServicesabstractIn this paper, we propose a method to evaluate the throughput of video transcoding in cloud services. This method can quickly estimate the maximum number of transcoding video concurrence on the current transcoding unit. Yangang Cai, Zhenyu Wang 0002, Ronggang Wang |
DCC | 4 |
| 2022 | Jointly Training of Binary 3D CNN Features for Action RecognitionabstractThis paper presents a novel method to train the quantized feature with the action recognition task jointly. A quantization and inverse-quantization layers are introduced to the 3D CNN. The quantization and the action recognition loss functions are minimized jointly. That is, the method aims to learn the feature not only to improve action recognition accuracy but also reduce the information loss of the quantization. The framework is shown in Fig. (1). Yangang Cai, Peiyin Xing, Zhenyu Wang 0002, Ronggang Wang |
DCC | 4 |
| 2019 | Separable KLT for Intra Coding in Versatile Video Coding (VVC)abstractAfter the works on the state-of-the-art High Efficiency Video Coding (HEVC) standard, the standard organizations continued to study the potential video coding technologies for the next generation of video coding standard, named Versatile Video Coding (VVC). Transform is a key technique for compression efficiency, and core experiment 6 (CE6) is carried out to explore the transform related coding tools. In this paper, we propose a novel separable transform based on Karhunen-Loève Transform (KLT) to eliminate the horizontal and vertical correlations in the residual samples of intra coding. In the proposed method, the weaknesses of the traditional KLT are addressed. The separable KLT is developed as an alternative transform type in addition to DCT-II, and the transform matrices from 4×4 to 64×64 are trained from intra residual samples. Experimental results show the proposed method can achieve 2.7% bitrate saving averagely on top of the reference software of VVC (VTM-1.1), and the consistent performance improvement on test set also validates the strong generalization capacity of the proposed separable KLT. Kui Fan, Ronggang Wang, Weisi Lin, Jong-Uk Hou, Ling-Yu Duan, Ge Li 0002, Wen Gao 0001 |
DCC | 2 |
| 2018 | Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001 |
Inf. Sci. | 2 |
| 2016 | Regional Subspace Projection Coding for Image RetrievalabstractFor image retrieval task, hamming embedding, being proved to be one of the state-of-the-art methods, has been prevalently utilised. The basic idea is to project local features into orthogonal space randomly, in which the binary signature is generated based on a single partition of feature space. However, the binary signature generation process is coarse and heuristic. On the one hand, the same projection is carried out for all visual word space without consideration of difference among subspaces. On the other hand, the projection matrix is generated randomly regardless of the distribution of feature data. Therefore, the performance of hamming embedding is limited and far from the optimal. In this paper, we firstly analyse the limitation of hamming em- bedding and compare different orthogonal projection methods. Then we propose a regional subspace projection coding method that is based on the distribution of local features assigned to each visual word. Finally, our experiments on two benchmark datasets demonstrate that our proposed method outperforms current state-of-the-art methods. Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang |
ICMR | 3 |
| 2012 | A Compact Stereoscopic Video Representation for 3D Video Generation and CodingabstractWe propose a novel compact representation for stereoscopic videos - a 2D video and its depth cues. Depth cues are derived from an interactive labeling process during 2D-to-3D video conversion, they are contour points of foreground objects and a background geometric model. By using such cues and image features of 2D video frames, depth maps of the frames can be recovered. Compared with traditional 3D video representation, the proposed one is more compact. We also design algorithms to encode and decode the depth cues. The representation benefits both 3D video generation and coding. Experimental results demonstrate that the bit rate can be saved about 10%-50% in coding 3D videos compared with multi-view video coding and 2D+depth methods. A system coupling 2D-to-3D video conversion and coding (CVCC) is proposed to verify advantages of the representation. Zhebin Zhang, Ronggang Wang, Yizhou Wang 0001, Wen Gao 0001 |
DCC | 2 |