Chuanbo Tang

dblp:352/2771 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Neural Video Compression with Reference Hierarchy
abstract
Efficient reference structures are essential in video compression, enabling the exploitation of temporal dependencies across frames to reduce redundancy. In this paper, we delve into the inter-frame reference management mechanism in neural video codecs (NVCs). Previous schemes have inherited the reference propagation mechanism with the guidance of predefined reference structure, but the reference modeling across diverse reference sources remains underexplored. Moreover, the mismatch between the reference structure used for motion estimation and motion compensation limits the effectiveness of inter-frame prediction. To address the above limitations, we propose the unified reference hierarchy that integrates a learned hierarchical reference structure into the existing inherent reference propagation mechanism. Specifically, we first propose the hierarchical reference structure (HRS) to manage the multiple temporal contexts in the propagated reference feature, where a hierarchy-aware reference modulation module is integrated to select the most relevant reference features across different quality levels under the guidance of the reference balance loss. In addition, we propose the HRS-guided feature-wise inter-frame prediction that learns the low-rank approximation of the selected reference feature for ensuring the consistency and improving the inter-frame prediction performance. We conduct experiments on a state-of-the-art NVC, DCVC-DC. Experimental results show that our codec achieves an average 26% bitrate saving over H.266/VVC, and a 28.2% bitrate reduction compared to DCVC-DC without increasing the decoding complexity.
Chuanbo Tang, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001
AAAI1
2026 NVC-1B: Scaling up Neural Video Coding Models
abstract
Emerging large models have achieved notable progress in the fields of natural language processing and computer vision. However, large models for neural video coding are still unexplored. In this paper, we try to explore how to build a large neural video coding model. Based on a small baseline model, we gradually scale up the model sizes of its different coding parts, including the motion encoder-decoder, motion entropy model, contextual encoder-decoder, contextual entropy model, and temporal context mining module, and analyze the influence of model sizes on video compression performance. Then, we explore using different architectures, including CNN, mixed CNN-Transformer, and Transformer architectures, to implement the neural video coding model and analyze the influence of model architectures on video compression performance. Based on our exploration results, we design the first neural video coding model having more than 1 billion parameters - NVC-1B. Experimental results show that our large model achieves a significant video compression performance improvement over recent state-of-the-art neural video compression models. With the continuous advancement in hardware and the successful on-device deployment of large models, we anticipate that our proposed large neural video coding model can bring video coding technologies to the next level.
Chuanbo Tang, Xihua Sheng, Li Li 0040, Dong Liu 0002, Feng Wu 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020s
abstract
Image/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD.
Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005
IEEE Trans. Multim.3
2025 Augmented Deep Contexts for Spatially Embedded Video Coding
abstract
Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-of-the-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https://github.com/EsakaK/SEVC.
Yifan Bian, Chuanbo Tang, Li Li 0040, Dong Liu 0002
CVPR2
2025 Neural Video Compression with Context Modulation
abstract
Efficient video coding is highly dependent on exploiting the temporal redundancy, which is usually achieved by extracting and leveraging the temporal context in the emerging conditional coding-based neural video codec (NVC). Although the latest NVC has achieved remarkable progress in improving the compression performance, the inherent temporal context propagation mechanism lacks the ability to sufficiently leverage the reference information, limiting further improvement. In this paper, we address the limitation by modulating the temporal context with the reference frame in two steps. Specifically, we first propose the flow orientation to mine the inter-correlation between the reference frame and prediction frame for generating the additional oriented temporal context. Moreover, we introduce the context compensation to leverage the oriented context to modulate the propagated temporal context generated from the propagated reference feature. Through the synergy mechanism and decoupling loss supervision, the irrelevant propagated information can be effectively eliminated to ensure better context modeling. Experimental results demonstrate that our codec achieves on average 22.7% bitrate reduction over the advanced traditional video codec H.266/VVC, and offers an average 10.1% bitrate saving over the previous state-of-the-art NVC DCVC-FM. The code is available at https://github.com/Austin4USTC/DCMVC.
Chuanbo Tang, Zhuoyuan Li 0001, Yifan Bian, Li Li 0040, Dong Liu 0002
CVPR1
2025 Enhanced Inter-frame Dependency Modeling with State Space Models for Neural Video Compression
abstract
In recent years, neural video compression (NVC) has achieved remarkable rate-distortion performance, surpassing traditional codecs by leveraging learnable architectures and endto-end optimization. However, under large intra-period settings, most NVC methods still lag behind conventional codecs. To address this issue, we first analyze the importance of the frame generator in inter-frame dependency modeling, where the propagated features and the quality of the reconstructed frames play a critical role in prediction accuracy and compression efficiency. Then, we enhance the spatiotemporal context modeling by incorporating state space models (SSMs) into the frame generator, aiming to improve the quality of propagated features. Specifically, we introduce a 2D-Selective-Scan (SS2D) mechanism that aggregates global features across frames to construct more informative and expressive reference representations. This design enables effective long-range temporal modeling while maintaining computational efficiency. Experimental results show that our method achieves an average bitrate saving of 13.6% over VTM in the RGB color space, evaluated on 300 frames with intra-period set to −1.
Chuanbo Tang, Yifan Bian, Li Li 0040, Dong Liu 0002
VCIP2
2025 Long-Term Motion Modeling for Neural Video Compression
abstract
Recent advances in end-to-end neural video compression (NVC) have attracted increasing attention due to their remarkable improvement in compression efficiency and adaptability. The accuracy of motion modeling in inter-frame prediction plays a critical role in further boosting compression performance. However, most existing NVC methods still rely on a single reference frame in the motion estimation (ME) stage, while long-term information is mainly exploited only in the motion compensation stage. This limitation hinders the ability to capture complex motion patterns, resulting in sub-optimal prediction accuracy. We propose a Long-Term Motion Modeling framework for Deep Video Compression (LMVC). Specifically, we introduce a ConvLSTM-based Long-Term Motion Modeling module into the ME stage, which leverages historical reconstructed frames to continuously refine the temporal content of reference frames. This design improves the consistency and robustness of motion estimation. Experimental results demonstrate that the proposed LMVC achieves an average BD-rate reduction of 18.7% compared to the VTM software across multiple public datasets and outperforms state-of-the-art neural video compression approaches, verifying the effectiveness and superiority of our motion modeling strategy.
Weihao Shi, Xiongzhuang Liang, Shujuan Zhang, Chuanbo Tang
VCIP6
2024 Offline and Online Optical Flow Enhancement for Deep Video Compression
abstract
Video compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these networks often adopt pre-trained optical flow estimation networks for motion estimation. The optical flows, however, may be less suitable for video compression due to the following two factors. First, the optical flow estimation networks were trained to perform inter-frame prediction as accurately as possible, but the optical flows themselves may cost too many bits to encode. Second, the optical flow estimation networks were trained on synthetic data, and may not generalize well enough to real-world videos. We address the twofold limitations by enhancing the optical flows in two stages: offline and online. In the offline stage, we fine-tune a trained optical flow estimation network with the motion information provided by a traditional (non-deep) video compression scheme, e.g. H.266/VVC, as we believe the motion information of H.266/VVC achieves a better rate-distortion trade-off. In the online stage, we further optimize the latent features of the optical flows with a gradient descent-based algorithm for the video to be compressed, so as to enhance the adaptivity of the optical flows. We conduct experiments on two state-of-the-art deep video compression schemes, DCVC and DCVC-DC. Experimental results demonstrate that the proposed offline and online enhancement together achieves on average 13.4% bitrate saving for DCVC and 4.1% bitrate saving for DCVC-DC on the tested videos, without increasing the model or computational complexity of the decoder side.
Chuanbo Tang, Xihua Sheng, Zhuoyuan Li 0001, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002
AAAI1
2024 Uniformly Accelerated Motion Model for Inter Prediction
abstract
Inter prediction is a key technology in video coding to reduce the temporal redundancy. In natural videos, there are usually moving objects with changing velocity, resulting in complex motion fields that are difficult to represent compactly. In Versatile Video Coding (VVC), existing inter prediction methods usually assume uniform speed motion between consecutive frames, which may not well handle the complex motion fields in the real world. To address these issues, we introduce a uniformly accelerated motion model (UAMM) to exploit velocity and acceleration of moving objects between the video frames, and further combine them to assist in the inter prediction methods to handle the motion change in the temporal domain. First, we review the theory of UAMM. Second, we propose UAMM-based parameter derivation and extrapolation schemes in the coding process. Third, we integrate the UAMM into existing inter prediction modes (Merge, MMVD, CIIP) to achieve higher prediction accuracy. The proposed method is implemented into the VVC reference software, VTM version 12.0. Experimental results show that the proposed method achieves up to 0.38% BD-rate reduction compared to the VTM anchor, under the Low-delay P configuration, with a slight increase of time complexity on the encoding/decoding side.
Zhuoyuan Li 0001, Yao Li 0016, Chuanbo Tang, Li Li 0040, Dong Liu 0002, Feng Wu 0001
VCIP3