EDBT 2026 Demo / reviewers in the wild / expert
Xihua Sheng
dblp:307/5102
· DBLP profile ↗
17ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-7350-4929ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CartoonCodec: Generative talking face video coding with cartoon-style customizationabstractThe rapid growth of video-based social applications has intensified the demand for efficient transmission and personalized cartoon-style customization. Existing solutions typically apply talking face video and text codecs, followed by cartoon-style control algorithms, which often result in unsatisfactory compression performance and high inference latency. In this paper, we propose CartoonCodec, an efficient generative framework that unifies face video coding and cartoon-style control into an end-to-end process. Our framework encodes facial video sequences into compact motion feature representations, which are transmitted together with compressed text prompts. To integrate control into the coding process while ensuring efficient compression, a text-guided adaptive layer selection mechanism relying on the compact representations is introduced to dynamically select and optimize the most influential layers in the generators. Furthermore, to facilitate stylization over the decoupled spaces, we propose a self-supervised domain stylization training strategy that constructs both multimodal and unimodal data pairs, enabling the use of diverse loss functions. Extensive experiments demonstrate that CartoonCodec outperforms baselines in compression efficiency for video reconstruction and cartoon-style control tasks while maintaining competitive inference efficiency. CartoonCodec provides key insights for advancing face video communication with cartoon-style customization. The project page can be found at https://github.com/xiaonae/CartoonCodec/tree/main . Xihua Sheng, Meng Wang 0017, Long Xu 0001, Shiqi Wang 0001, Sam Kwong |
Neurocomputing | 2 |
| 2026 | NVC-1B: Scaling up Neural Video Coding ModelsabstractEmerging large models have achieved notable progress in the fields of natural language processing and computer vision. However, large models for neural video coding are still unexplored. In this paper, we try to explore how to build a large neural video coding model. Based on a small baseline model, we gradually scale up the model sizes of its different coding parts, including the motion encoder-decoder, motion entropy model, contextual encoder-decoder, contextual entropy model, and temporal context mining module, and analyze the influence of model sizes on video compression performance. Then, we explore using different architectures, including CNN, mixed CNN-Transformer, and Transformer architectures, to implement the neural video coding model and analyze the influence of model architectures on video compression performance. Based on our exploration results, we design the first neural video coding model having more than 1 billion parameters - NVC-1B. Experimental results show that our large model achieves a significant video compression performance improvement over recent state-of-the-art neural video compression models. With the continuous advancement in hardware and the successful on-device deployment of large models, we anticipate that our proposed large neural video coding model can bring video coding technologies to the next level. Chuanbo Tang, Xihua Sheng, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | DRFC: An End-to-End Deep Dynamic RF Signal Compression FrameworkabstractRadio frequency (RF) signals have gained widespread adoption in intelligent perception systems due to their unique advantages, including non-line-of-sight propagation capability, robustness in low-light environments, and inherent privacy preservation. However, their substantial data volumes, generated by the dual-polarization direction characteristic, result in significant challenges to data storage and transmission. To address this, we propose the first end-to-end deep dynamic RF signal compression (DRFC) framework, which primarily focuses on exploiting cross-directional correlation in dynamic RF signals. The proposed framework incorporates four key innovations: (1) a mask-guided RF motion estimation module that leverages Doppler shifts and electromagnetic noise characteristics to identify regions of significant motion using a threshold-based mask, significantly improving motion estimation accuracy; (2) a cross-directional RF motion entropy model that utilizes cross-directional RF motion latent priors to refine the probability distribution for motion entropy coding; (3) a cross-directional RF context mining module that predicts RF contexts from temporal and cross-directional reference signals, adaptively fusing these contexts with confidence maps to maximize complementary information utilization; and (4) a cross-directional RF contextual entropy model that incorporates cross-directional RF contextual latent priors to optimize contextual entropy modeling. Experimental results demonstrate the superiority of our framework over existing codecs. Our DRFC framework achieves significant bitrate savings on benchmark datasets, establishing a strong baseline for future research in this field. Xihua Sheng, Peilin Chen 0001, Shiqi Wang 0001, Dapeng Oliver Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | CoFaCo: Controllable Generative Talking Face Video CodingabstractEfficient talking face video coding and control are crucial in modern video communication, reshaping how individuals connect, collaborate, and interact. Coding seeks to reduce transmission costs, while control enables the realization of user-customizable facial expressions and head poses in the transmitted videos. However, the compression efficiency of the common par-adigm of applying control algorithms before video coding is not satisfactory. In this paper, we propose an efficient, Controllable Generative Talking Face Video Coding (CoFaCo) framework, wh-ich seamlessly integrates control into the coding process. Specific-ally, CoFaCo projects talking face videos into ultra-compact and semantic feature representations that can be customized by users before compression. To enable independent controls of pose and expression, we design a set of sophisticated losses to accurately de-couple the pose and expression direction codes. Once the decoupled direction codes and the semantic face representations are obtained, the pose and expression control modules can be effectively learned to generate decoupled, controlled pose and expression direction codes. The controlled direction codes are subsequently smoothed to enhance temporal consistency in the controlled video output by the generators. Experimental results demonstrate that CoFaCo achieves competitive compression efficiency in ultra-low bit rate video reconstruction and control tasks, providing valuab-le insights for advancing face video communication with diverse control capabilities. Xihua Sheng, Meng Wang 0017, Fu-Zhao Ou, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2026 | USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020sabstractImage/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD. Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 7 |
| 2026 | Distributed Deep Point Cloud Feature Compression for Vehicle-to-Vehicle Cooperative PerceptionabstractWith the rapid development of autonomous driving, intermediate fusion cooperative perception has been proposed to broaden the sensing range of a single vehicle by fusing the point cloud features transmitted from surrounding vehicles. However, transmission of these features demands substantial bandwidth, necessitating the creation of efficient methods for compressing point cloud features. Our basic compression idea is to remove two types of semantic redundancies in Vehicle-to-Vehicle (V2V) cooperative perception scenarios: inter-redundancy between the transmitted features and the ego feature, and intra-redundancy among the symbols of transmitted features. To this end, we propose a Distributed deep Point cloud Feature Compression (DPFC) scheme with rate-perception optimization. Different from previous schemes that merely reduce the dimension of transmitted features using an autoencoder, we design a distributed feature compression network conditioned on the ego feature to remove inter-redundancy. To further remove intra-redundancy, we introduce rate-perception optimization to exploit in-depth spatial redundancy and symbol redundancy. Experimental results demonstrate that our proposed DPFC can significantly reduce the bandwidth cost while achieving higher perceptive performance compared with the autoencoder. Xihua Sheng, Li Li 0040, Dong Liu 0002, Houqiang Li |
IEEE Trans. Multim. | 2 |
| 2025 | An Information-Theoretic Regularizer for Lossy Neural Image CompressionabstractLossy image compression networks aim to minimize the latent entropy of images while adhering to specific distortion constraints. However, optimizing the neural network can be challenging due to its nature of learning quantized latent representations. In this paper, our key finding is that minimizing the latent entropy is, to some extent, equivalent to maximizing the conditional source entropy, an insight that is deeply rooted in information-theoretic equalities. Building on this insight, we propose a novel structural regularization method for the neural image compression task by incorporating the negative conditional source entropy into the training objective, such that both the optimization efficacy and the model's generalization ability can be promoted. The proposed information-theoretic regularizer is interpretable, plug-and-play, and imposes no inference overheads. Extensive experiments demonstrate its superiority in regularizing the models and further squeezing bits from the latent representation across various compression structures and unseen domains. Yingwen Zhang, Meng Wang 0017, Xihua Sheng, Peilin Chen 0001, Li Zhang 0006, Shiqi Wang 0001 |
ICCV | 3 |
| 2025 | Prediction and Reference Quality Adaptation for Learned Video CompressionabstractTemporal prediction is one of the most important technologies for video compression. Various prediction coding modes are designed in traditional video codecs. Traditional video codecs will adaptively to decide the optimal coding mode according to the prediction quality and reference quality. Recently, learned video codecs have made great progress. However, they did not effectively address the problem of prediction and reference quality adaptation, which limits the effective utilization of temporal prediction and reduction of reconstruction error propagation. Therefore, in this paper, we first propose a confidence-based prediction quality adaptation (PQA) module to provide explicit discrimination for the spatial and channel-wise prediction quality difference. With this module, the prediction with low quality will be suppressed and that with high quality will be enhanced. The codec can adaptively decide which spatial or channel location of predictions to use. Then, we further propose a reference quality adaptation (RQA) module and an associated repeat-long training strategy to provide dynamic spatially variant filters for diverse reference qualities. With these filters, our codec can adapt to different reference qualities, making it easier to achieve the target reconstruction quality and reduce the reconstruction error propagation. Experimental results verify that our proposed modules can effectively help our codec achieve a higher compression performance. Xihua Sheng, Li Li 0040, Dong Liu 0002, Houqiang Li |
IEEE Trans. Image Process. | 1 |
| 2025 | Bi-Directional Deep Contextual Video CompressionabstractDeep video compression has made impressive process in recent years, with the majority of advancements concentrated on P-frame coding. Although efforts to enhance B-frame coding are ongoing, their compression performance is still far behind that of traditional bi-directional video codecs. In this article, we introduce a bi-directional deep contextual video compression scheme tailored for B-frames, termed DCVC-B, to improve the compression performance of deep B-frame coding. Our scheme mainly has three key innovations. First, we develop a bi-directional motion difference context propagation method for effective motion difference coding, which significantly reduces the bit cost of bi-directional motions. Second, we propose a bi-directional contextual compression model and a corresponding bi-directional temporal entropy model, to make better use of the multi-scale temporal contexts. Third, we propose a hierarchical quality structure-based training strategy, leading to an effective bit allocation across large groups of pictures (GOP). Experimental results show that our DCVC-B achieves an average reduction of 26.6% in BD-Rate compared to the reference software for H.265/HEVC under random access conditions. Remarkably, it surpasses the performance of the H.266/VVC reference software on certain test datasets under the same configuration. We anticipate our work can provide valuable insights and bring up deep B-frame coding to the next level. Xihua Sheng, Li Li 0040, Dong Liu 0002, Shiqi Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Offline and Online Optical Flow Enhancement for Deep Video CompressionabstractVideo compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these networks often adopt pre-trained optical flow estimation networks for motion estimation. The optical flows, however, may be less suitable for video compression due to the following two factors. First, the optical flow estimation networks were trained to perform inter-frame prediction as accurately as possible, but the optical flows themselves may cost too many bits to encode. Second, the optical flow estimation networks were trained on synthetic data, and may not generalize well enough to real-world videos. We address the twofold limitations by enhancing the optical flows in two stages: offline and online. In the offline stage, we fine-tune a trained optical flow estimation network with the motion information provided by a traditional (non-deep) video compression scheme, e.g. H.266/VVC, as we believe the motion information of H.266/VVC achieves a better rate-distortion trade-off. In the online stage, we further optimize the latent features of the optical flows with a gradient descent-based algorithm for the video to be compressed, so as to enhance the adaptivity of the optical flows. We conduct experiments on two state-of-the-art deep video compression schemes, DCVC and DCVC-DC. Experimental results demonstrate that the proposed offline and online enhancement together achieves on average 13.4% bitrate saving for DCVC and 4.1% bitrate saving for DCVC-DC on the tested videos, without increasing the model or computational complexity of the decoder side. Chuanbo Tang, Xihua Sheng, Zhuoyuan Li 0001, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
AAAI | 2 |
| 2024 | VNVC: A Versatile Neural Video Coding Framework for Efficient Human-Machine VisionabstractAlmost all digital videos are coded into compact representations before being transmitted. Such compact representations need to be decoded back to pixels before being displayed to humans and - as usual - before being enhanced/analyzed by machine vision algorithms. Intuitively, it is more efficient to enhance/analyze the coded representations directly without decoding them into pixels. Therefore, we propose a versatile neural video coding (VNVC) framework, which targets learning compact representations to support both reconstruction and direct enhancement/analysis, thereby being versatile for both human and machine vision. Our VNVC framework has a feature-based compression loop. In the loop, one frame is encoded into compact representations and decoded to an intermediate feature that is obtained before performing reconstruction. The intermediate feature can be used as reference in motion compensation and motion estimation through feature-based temporal context mining and cross-domain motion encoder-decoder to compress the following frames. The intermediate feature is directly fed into video reconstruction, video enhancement, and video analysis networks to evaluate its effectiveness. The evaluation shows that our framework with the intermediate feature achieves high compression efficiency for video reconstruction and satisfactory task performances with lower complexities. Xihua Sheng, Li Li 0040, Dong Liu 0002, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Spatial Decomposition and Temporal Fusion Based Inter Prediction for Learned Video CompressionabstractVideo compression performance is closely related to the accuracy of inter prediction. It tends to be difficult to obtain accurate inter prediction for the local video regions with inconsistent motion and occlusion. Traditional video coding standards propose various technologies to handle motion inconsistency and occlusion, such as recursive partitions, geometric partitions, and long-term references. However, existing learned video compression schemes focus on obtaining an overall minimized prediction error averaged over all regions while ignoring the motion inconsistency and occlusion in local regions. In this paper, we propose a spatial decomposition and temporal fusion based inter prediction for learned video compression. To handle motion inconsistency, we propose to decompose the video into structure and detail (SDD) components first. Then we perform SDD-based motion estimation and SDD-based temporal context mining for the structure and detail components to generate short-term temporal contexts. To handle occlusion, we propose to propagate long-term temporal contexts by recurrently accumulating the temporal information of each historical reference feature and fuse them with short-term temporal contexts. With the SDD-based motion model and long short-term temporal contexts fusion, our proposed learned video codec can obtain more accurate inter prediction. Comprehensive experimental results demonstrate that our codec outperforms the reference software of H.266/VVC on all common test datasets for both PSNR and MS-SSIM. Xihua Sheng, Li Li 0040, Dong Liu 0002, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | LSSVC: A Learned Spatially Scalable Video Coding SchemeabstractTraditional block-based spatially scalable video coding has been studied for over twenty years. While significant advancements have been made, the scope for further improvement in compression performance is limited. Inspired by the success of learned video coding, we propose an end-to-end learned spatially scalable video coding scheme, LSSVC, which provides a new solution for scalable video coding. In LSSVC, we propose to use the motion, texture, and latent information of the base layer (BL) as interlayer information for compressing the enhancement layer (EL). To reduce interlayer redundancy, we design three modules to leverage the upsampled interlayer information. Firstly, we design a contextual motion vector (MV) encoder-decoder, which utilizes the upsampled BL motion information to help compress high-resolution MV. Secondly, we design a hybrid temporal-layer context mining module to learn more accurate contexts from the EL temporal features and the upsampled BL texture information. Thirdly, we use the upsampled BL latent information as an interlayer prior for the entropy model to estimate more accurate probability distribution parameters for the high-resolution latents. Experimental results show that our scheme surpasses H.265/SHVC reference software by a large margin. Our code is available at https://github.com/EsakaK/LSSVC. Yifan Bian, Xihua Sheng, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Image Process. | 2 |
| 2023 | Joint Optimized Point Cloud Compression for 3d Object DetectionabstractAs a large amount of point clouds are fed into machines for 3D object detection in autonomous driving, efficient point cloud compression methods are urgently needed. Existing point cloud compression methods are usually optimized for high signal fidelity rather than high detection accuracy. However, higher signal fidelity may be unnecessary for object detection. In this work, we propose a learning-based point cloud compression framework for 3D object detection by jointly optimizing the point cloud compression and 3D object detection network. Since point clouds commonly need to be pre-processed before performing detection, simply connecting the codec and detector will disable the gradient back-propagation due to some non-differentiable operations. Therefore, we design a gradient bridge function to enable the gradient back-propagation from the detector to the codec. In addition, joint optimization of two networks from scratch is easy to make the training process unstable. We propose a progressive training strategy to stabilize the training process. Experimental results demonstrate that our framework can achieve significant compression performance improvement under the same detection accuracy. Xihua Sheng, Li Li 0040, Dong Liu 0002 |
ICIP | 3 |
| 2023 | Temporal Context Mining for Learned Video CompressionabstractApplying deep learning to video compression has attracted increasing attention in recent few years. In this work, we address end-to-end learned video compression with a special focus on better learning and utilizing temporal contexts. We propose to propagate not only the last reconstructed frame but also the feature before obtaining the reconstructed frame for temporal context mining. From the propagated feature, we learn multi-scale temporal contexts and re-fill the learned temporal contexts into the modules of our compression scheme, including the contextual encoder-decoder, the frame generator, and the temporal context encoder. We discard the parallelization-unfriendly auto-regressive entropy model to pursue a more practical encoding and decoding time. Experimental results show that our proposed scheme achieves a higher compression ratio than the existing learned video codecs. Our scheme also outperforms x264 and x265 (representing industrial software for H.264 and H.265, respectively) as well as the official reference software for H.264, H.265, and H.266 (JM, HM, and VTM, respectively). Specifically, when intra period is 32 and oriented to PSNR, our scheme outperforms H.265–HM by 14.4% bit rate saving; when oriented to MS-SSIM, our scheme outperforms H.266–VTM by 21.1% bit rate saving. Xihua Sheng, Jiahao Li 0001, Bin Li 0012, Li Li 0040, Dong Liu 0002, Yan Lu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Attribute Artifacts Removal for Geometry-Based Point Cloud CompressionabstractGeometry-based point cloud compression (G-PCC) can achieve remarkable compression efficiency for point clouds. However, it still leads to serious attribute compression artifacts, especially under low bitrate scenarios. In this paper, we propose a Multi-Scale Graph Attention Network (MS-GAT) to remove the artifacts of point cloud attributes compressed by G-PCC. We first construct a graph based on point cloud geometry coordinates and then use the Chebyshev graph convolutions to extract features of point cloud attributes. Considering that one point may be correlated with points both near and far away from it, we propose a multi-scale scheme to capture the short- and long-range correlations between the current point and its neighboring and distant points. To address the problem that various points may have different degrees of artifacts caused by adaptive quantization, we introduce the quantization step per point as an extra input to the proposed network. We also incorporate a weighted graph attentional layer into the network to pay special attention to the points with more attribute artifacts. To the best of our knowledge, this is the first attribute artifacts removal method for G-PCC. We validate the effectiveness of our method over various point clouds. Objective comparison results show that our proposed method achieves an average of 9.74% BD-rate reduction compared with Predlift and 10.13% BD-rate reduction compared with RAHT. Subjective comparison results present that visual artifacts such as color shifting, blurring, and quantization noise are reduced. Xihua Sheng, Li Li 0040, Dong Liu 0002, Zhiwei Xiong |
IEEE Trans. Image Process. | 1 |
| 2022 | Deep-PCAC: An End-to-End Deep Lossy Compression Framework for Point Cloud AttributesabstractThe large data volume of point clouds poses severe challenges for efficient storage and transmission in recent years. In this paper, we propose the first--to our best knowledge--end-to-end deep framework for compressing point cloud attributes. Specifically, we propose a point cloud lossy attribute autoencoder, which directly encodes and decodes point cloud attributes with the help of geometry, instead of voxelizing or projecting the points. In the autoencoder, we propose a second-order point convolution that utilizes the spatial correlations between more points and the nonlinear relationship between attribute features. We introduce a dense point-inception block, which derives from a combination of an inception-style block and a dense block, to improve feature propagation. In addition, we devise a multiscale loss to guide the autoencoder in focusing attention on the coarse-grained points with better coverage of the entire point cloud, which makes it easier for the autoencoder to obtain better optimization of the qualities of all points. Experimental results show that our proposed framework still has a performance gap compared with the state-of-the-art algorithms in the MPEG G-PCC reference software TMC13. However, it does outperform the RAHT-RLGR, which is one of the core transforms used in TMC13 without many well-designed techniques that make TMC13 what it is today. It outperforms RAHT-RLGR by 2.63 dB, 1.77 dB, and 3.40 dB on average in terms of the BD-PSNR for the Y, U, and V components. A subjective quality comparison demonstrates that our framework can preserve more textures and reduce blocking and color noise artifacts. Xihua Sheng, Li Li 0040, Dong Liu 0002, Zhiwei Xiong, Zhu Li 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |