Siwei Ma 0001

dblp:40/5402 · DBLP profile ↗
← Back
68ranked-venue papers in the field
1as first author
43since 2021 · last 2026
0000-0002-2731-5403ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 67 (1 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 An Information-Guided Learned Framework for Free-View Image Coding
abstract
In this paper, we propose an information-guided learned image compression framework, which achieves efficient free-view image compression. Specifically, we model the inter-view correlations as inter-view prior based on the proposed feature transform module, which is used to guide the encoding and decoding process of different views. Specifically, we design the feature transform module to generate the inter-view prior. Different from prior information modeled in pixel domain, the inter-view prior in feature domain has higher dimensions and can provide richer and more correlated condition information, which can guide the encoding and decoding process of different view images effectively. Furthermore, we design a multi-prior fusion module, which can fuse the information from different reference views to achieve more accurate guidance. We evaluate the proposed framework by comparing the rate-distortion performance averaged over the BlendedMVS dataset. Extensive experiments demonstrate that the proposed framework can achieve the superior performance in compressing free-view images.
Wenhong Duan, Hui Yuan 0001, Siwei Ma 0001
DCC5
2026 A Search Method for Approximate Optimal Rate Control Solution via Reinforcement Learning
abstract
Recent studies on video rate control (RC) have introduced accurate and high performance methods but have not explored the optimal RC solution. The optimal solution is crucial for improving RC methods and providing labels for supervised learning. To find approximate optimal RC solutions within a limited time, we are the first to propose a reinforcement learning based search method for finding approximate optimal RC solutions within a limited time. Specifically, the RC problem for a video is first modeled as a Markov decision process (MDP). Then, with the MDP model, we develop a search method based on the deep Q-network method, which consists of exploration and exploitation steps. During exploration, an agent is created, consisting of two multilayer perceptrons and a replay memory, and trained within the MDP to estimate the value function, while superior RC solutions are recorded throughout the training process. After training, RC solution is estimated by the trained agent using the value function and a greedy strategy during the exploitation step. Finally, the approximate optimal RC solution is determined based on the RC solutions from both two steps. In addition, the time complexity of proposed method is controllable, specifically,$O(m n)$where$m$denotes the number of training epoch. Experimental results show that the bit-rate error and compression quality of the solutions found by proposed method approach the optimal solutions, with differences of only less than 0.005% and 0.399%, respectively, and are achieved in a significantly shorter time compared to the brute force search.
Longtao Feng, Qian Yin 0002, Jiaqi Zhang 0007, Yuwen He, Siwei Ma 0001
DCC6
2026 Virtual Reference Frame Synthesis for Video Coding via Local-Global Spatiotemporal Context Modeling
abstract
Inter prediction is a fundamental component of modern video coding, where the quality of reference frames critically affects motion compensation accuracy and overall compression efficiency. However, relying solely on reconstructed low-temporal-layer frames imposes significant limitations, as these frames often suffer from compression artifacts that degrade prediction quality. To overcome this limitation, we propose a Local-Global spatiotemporal context modeling-based virtual reference frame generation network (LGCM-Net) that synthesizes high-quality reference frames based on reconstructed frames, as shown in Fig. 1. The proposed network integrates hierarchical feature extraction with long-range dependency modeling, where QP-conditioned modulation is applied to shallow features to adapt them to quantization-induced quality variations, enabling temporally and structurally consistent reference generation closely aligned with the to-be-coded frame. Moreover, a coarse-to-fine multi-stage optical flow refinement mechanism is employed to progressively enhance motion accuracy, and a residual refiner further compensates remaining motion estimation errors and reconstruction artifacts to deliver a more accurate final prediction. The proposed method achieves$5.37 \%, 9.96 \%$, and 9.91% BD-rate reduction under the Random Access configuration in VVC reference Software (VTM-11.0_nnvc-10.0 w/o NN Coding tools) for the$\mathrm{Y}, \mathrm{U}$, and V components, respectively.
Yanchen Zhao, Xuewei Meng, Jiaqi Zhang 0007, Kai Zhang 0007, Siwei Ma 0001
DCC6
2026 Prompt-Optimization with Contextual Mining for Cross-Modal Image Compression
abstract
Recent advances in cross-modal compression(CMC) have opened new horizons for perceptual image coding at ultra-low bitrates (below 0.1 bpp) within a generative compression paradigm, but reconstruction fidelity is often compromised, yielding visually plausible yet semantically inconsistent reconstructions. While prompt engineering with contextual optimization has been extensively explored in generative models, its potential for controlling perception-fidelity trade-offs in image compression remains largely under-explored. To address these challenges, we propose PO-CMC, a novel diffusion-based cross-modal image compression approach that introduces contextual prompt optimization to achieve efficient and perceptually faithful reconstruction. The proposed method comprises three synergistic components: an optimized image codec that produces a compact structural prior, a contextual prompt module that adaptively encodes semantic cues into compact textual embeddings, and a diffusion-based decoder that fuses the structural and semantic priors to reconstruct high-fidelity images. Extensive experiments show that PO-CMC achieves superior perceptual quality while maintaining comparable reconstruction fidelity, yielding an average BD-rate saving of 72.5 % and 79.8 % over VVC at equivalent LPIPS and DISTS levels, respectively.
Shenpeng Song, Zhimeng Huang, Junlong Gao, Chuanmin Jia, Siwei Ma 0001
DCC5
2026 Efficient Feedforward Human-Centric Video Compression via 3D Gaussian Generation
abstract
In this paper, we propose a feed-forward framework (Fig. 1) for human video compression based on 3D generative reconstruction. Recent surveys [1], [2] highlight the need for more efficient and semantically aligned solutions. Our approach disentangles video content into complementary structural and motion layers: the structural layer encodes regularized texture from a single frame, while the motion layer leverages the SMPL-X prior to represent complex dynamics with a compact set of pose and shape parameters. A hierarchical coding scheme exploits the heterogeneity of these representations for improved efficiency. After decoding, a feed-forward 3D reconstruction pipeline with facial feature extraction is employed, in which a multimodal transformer and a Gaussian head synthesize parametric cues that are fused with motion signals for accurate animation and high-fidelity rendering. Experiments show over$1000 \times$compression while preserving structural and semantic fidelity. The method consistently outperforms strong baselines (especially at$0.04-0.1 \text{kbpp})$, with significant gains in rate-distortion, FVD, and perceptual quality, as well as robust generalization across identities and scenes.
Haocheng Tang, Ruoke Yan, Jiaqi Zhang 0007, Siwei Ma 0001
DCC7
2026 End-to-End RGB-IR Joint Image Compression with Channel-Wise Cross-Modality Entropy Model
abstract
RGB-IR(RGB-Infrared) image pairs are frequently applied simultaneously in various applications like intelligent surveillance. However, as the number of modalities increases, the required data storage and transmission costs also double. Therefore, efficient RGB-IR data compression is essential. This work proposes a joint compression framework for RGB-IR image pair. Specifically, to fully utilize cross-modality prior information for accurate context probability modeling within and between modalities, we propose a Channel-wise Cross-modality Entropy Model (CCEM). Among CCEM, a Low-frequency Context Extraction Block (LCEB) and a Low-frequency Context Fusion Block (LCFB) are designed for extracting and aggregating the global low-frequency information from both modalities, which assist the model in predicting entropy parameters more accurately. Experimental results demonstrate that our approach outperforms existing RGB-IR image pair and single-modality compression methods on LLVIP and KAIST datasets. For instance, the proposed framework achieves a 23.1% bit rate saving on LLVIP dataset compared to the state-of-the-art RGB-IR image codec presented at CVPR 2022.
Fangtao Zhou, Qizhang, Tiange Zhang, Xiaofeng Huang, Zhao Wang 0004, Siwei Ma 0001
DCC8
2026 Rethink Feature Coding for Machine Under Ultra-Low Bitrate: Framework and Optimization
abstract
This paper presents FCM-ULB, a feature coding framework designed for machine vision in extremely low-bitrate scenarios. It introduces a spatial-channel attention (SC-Attn) block to better capture global feature dependencies, enabling aggressive downsampling while preserving task-relevant information. In addition, a three-step training strategy—warm-up of vision backbone, rate-distortion optimization of feature compression, and joint fine-tuning—further improves compression and maintains task accuracy. Experiments on COCO and OpenImages show that FCM-ULB achieves much higher accuracy than existing standards like FCTMv4 at compression ratios up to$10,000 \times$, demonstrating its effectiveness for bandwidth-limited vision applications.
Rongao Yuan, Zhimeng Huang, Siwei Ma 0001
DCC5
2026 L-STEC: Learned Video Compression with Long-Term Spatio-Temporal Enhanced Context
abstract
Neural Video Compression has emerged in recent years, with condition-based frameworks outperforming traditional codecs. However, most existing methods rely solely on the previous frame's features to predict temporal context, leading to two critical issues. First, the short reference window misses long-term dependencies and fine texture details. Second, propagating only feature-level information accumulates errors over frames, causing prediction inaccuracies and loss of subtle textures. To address these, we propose the Long-term Spatio-Temporal Enhanced Context (L-STEC) method. We first extend the reference chain with LSTM to capture long-term dependencies. We then incorporate warped spatial context from the pixel domain, fusing spatio-temporal information through a multi-receptive field network to better preserve reference details. Experimental results show that L-STEC significantly improves compression by enriching contextual information, achieving 37.01% bitrate savings in PSNR and 31.65% in MS-SSIM compared to DCVC-TCM, outperforming both VTM-17.0 and DCVC-FM and establishing new state-of-the-art performance.
Tiange Zhang, Zhimeng Huang, Xiandong Meng, Kai Zhang 0007, Zhipin Deng, Siwei Ma 0001
DCC6
2026 Content Adaptive Based Motion Alignment Framework for Learned Video Compression
abstract
Recent advances in end-to-end video compression have shown promising results owing to their unified end-to-end learning optimization. However, such generalized frameworks often lack content-specific adaptation, leading to suboptimal compression performance. To address this, this paper proposes a content adaptive based motion alignment framework that improves performance by adapting encoding strategies to diverse content characteristics. Specifically, we first introduce a two-stage flow-guided deformable warping mechanism that refines motion compensation with coarse-to-fine offset prediction and mask modulation, enabling precise feature alignment. Second, we propose a multi-reference quality aware strategy that adjusts distortion weights based on reference quality, and applies it to hierarchical training to reduce error propagation. Third, we integrate a training-free module that downsamples frames by motion magnitude and resolution to obtain smooth motion estimation. Experimental results on standard test datasets demonstrate that our framework CAMA achieves significant improvements over state-of-the-art Neural Video Compression models, achieving a 24.95% BD-rate (PSNR) savings over our baseline model DCVC-TCM, while also outperforming reproduced DCVC-DC and traditional codec HM-16.25.
Tiange Zhang, Xiandong Meng, Siwei Ma 0001
DCC3
2026 Lightweight CNN-Based In-Loop Filtering for Video Coding with Hardware-Aware Optimizations
abstract
Neural network-based in-loop filtering significantly enhances video compression efficiency. However, high computational complexity hinders their deployment in real-time and ultra-high-definition scenarios. To address this, we propose a lightweight CNN-based in-loop filter for the luma component. In terms of model design, we utilize a U-Net-like architecture to learn the residual signal, incorporating depthwise separable 3 × 3 convolutions and 1 × 1 convolutions to reduce computational complexity, which results in a low complexity of only$37.707 \text{kMACs} /$pixel. For deployment optimization, we implement memory linearization to improve cache efficiency and combine blocked matrix multiplication with SIMD to maximize parallelism, ensuring cross-platform compatibility without third-party libraries. Experimental results on AVS4 EVM-0.9 (All-Intra) on a CPU platform show that the proposed method achieves BD-rate reductions of$1.36 \%, 0.35 \%$, and 0.34% for$\mathrm{Y}, \mathrm{U}$, and V components, respectively. Furthermore, the optimizations lead to a 91.5% reduction in decoding time, resulting in a decoding complexity of 7757% compared to the anchor.
Yanchen Zhao, Xuewei Meng, Jiaqi Zhang 0007, Haocheng Tang, Lin Li 0062, Siwei Ma 0001
DCC8
2026 Beyond CNN Filters: Diffusion-Based Post-Processing for VVC Intra Coding
abstract
Video post-processing can significantly enhance compressed video quality. Traditional methods are mostly based on handcrafted designs, such as deblocking and deringing algorithms, which rely on fixed heuristic rules and exhibit limited flexibility and adaptability. In recent years, with the rapid development of deep learning, Neural Network-based (NN-based) video post-processing methods have demonstrated remarkable coding performance. Among these, Transformer-based or multi-frame joint enhancement methods have improved reconstruction quality but still heavily depend on the prediction of known pixels, making it difficult to effectively restore high-frequency texture information lost due to the lossy compression. In contrast, diffusion models leverage their powerful generative priors and progressive denoising mechanisms to synthesize more natural and realistic high-frequency details, providing a promising solution for video post-processing. Inspired by recent advancements in conditional generative modeling, we propose a diffusion-based post-processing filter. The design incorporates a quantization parameter adaptive module and a block-based inference strategy with overlapping blocks to balance visual quality and efficiency. Experimental results demonstrate that the proposed method achieves significant improvements in both subjective and objective quality on the VTM-11.0.
Yanchen Zhao, Zhimeng Huang, Jiaqi Zhang 0007, Lin Li 0062, Siwei Ma 0001
DCC7
2025 STACO: Spatio-Temporal Adaptive Context Optimization for Neural Video Compression
abstract
This paper introduces the Spatio-Temporal Adaptive Context Optimization (STACO) method, which enhances the quality of contextual prediction across various resolutions, essential for subsequent compression. The STACO takes predicted contexts$C_t^{\{1,2,3\}}$as input and improves their quality by aligning them better with decoded features$f_{t}$, thus boosting coding efficiency. The STACO comprises Quality Perception Units and Consistency Synergy Modules, arranged in a hierarchical stacked architecture. This multi-scale design enables simultaneous processing of contexts at different spatial resolutions and facilitates information exchange through upsampling and downsampling. Enhanced contexts$\tilde{C}_{t}^{\{1,2,3\}}$are output after passing through residual connections, ensuring better alignment with reconstructed features. Using VTM-11.0 as anchor, the STACO significantly improves compression efficiency on common test condition (CTC) in HEVC, achieving an average BD-rate reduction of 17.99% for PSNR and 43.84% for MS-SSIM. By incorporating spatial quality mapping and temporal propagation, STACO offers a significant advancement in video compression.
Kexiang Feng, Shuhong Liao, Zhimeng Huang, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC7
2025 A Fast Bit Allocation Refinement for Video Rate Control
abstract
Since the introduction of hierarchical picture prediction structure in the advanced video coding (AVC), the hierarchical coding structure (HCS) has been widely adopted and continuously improved in video coding standards. Correspondingly, the HCS-based bit allocation methods in rate control have also emerged endlessly. Considering that pictures in higher temporal levels (TLs) of HCS usually refer to pictures in lower TLs, most methods tend to allocate more bits to pictures in lower TLs. However, these methods do not fully consider the correlation of picture quality in different TLs, which leads to the bit allocation waste and the coding performance degradation. To address this issue, we propose a fast bit allocation refinement method that can adapt to different video rate control approaches. Fig. 1 shows the overall framework of the proposed method. In general, our method is to appropriately adjust the bit allocation of pictures in lower TLs according to the relationship between the quality of pictures in different TLs. Specifically, based on the hyperbolic rate-distortion (RD) model and initial allocated bits, the quality of picture in higher TLs is first predicted and then used to estimate the quality of picture in lower TLs. Subsequently, the bits of picture in lower TLs are derived using estimated quality and its RD model. Finally, the final allocated bits of picture in lower TLs are adjusted by comparing the estimated and initial allocated bits. Experimental results show that our method can improve the coding performance of different rate control methods without introducing latency and encoding complexity.
Longtao Feng, Qian Yin 0002, Jiaqi Zhang 0007, Lin Li 0062, Siwei Ma 0001
DCC6
2025 Recurrent Intra Prediction Mode for Future Video Coding
abstract
Intra prediction is a crucial component of hybrid video coding framework due to its remarkable ability to reduce spatial redundancy in video signals. Unlike the single-mode based intra prediction in HEVC and VVC, intra fusion prediction methods, that combine the results of multiple angular prediction modes, were newly adopted by Enhanced Compression Model (ECM). However, intra fusion prediction over-relies on local spatial correlations and neglects potential texture similarities in non-adjacent regions. To overcome these limitations and elevate the accuracy of luma intra prediction, a Recurrent Intra Prediction Mode (RIPM) is proposed in this paper, which is composed of two sub-modules, i.e., Recurrent Intra Merge Mode (RIMM) and Recurrent Block Vector Substitution Module (RBVSM). RIMM utilizes the recurrent spatial texture information of the adjacent and non-adjacent spaces for adaptive mode derivation and prediction within the intra fusion prediction framework. RBVSM is a sophisticated mechanism for adaptive prediction mode selection and weight assignment during intra fusion prediction, resulting in enhanced coding performance with minimal impact on computational complexity. The proposed method, implemented on top of ECM-12.0, demonstrates a 0.095% BD-rate gain for the luma component under All Intra configuration, with negligible complexity increase. Currently, RIPM is under study in Exploration Experiments (EE) for ECM in JVET.
Jiaye Fu, Xuewei Meng, Siwei Ma 0001, Jiaqi Zhang 0007, Yao-Jen Chang, Vadim Seregin, Marta Karczewicz
DCC3
2025 Rethinking Bjøntegaard Delta for Compression Efficiency Evaluation: Are we Calculating it Precisely and Reliably?
abstract
For decades, the Bjøntegaard Delta (BD) has been the metric for evaluating codec Rate-Distortion (R-D) performance. Yet, in most studies, BD is determined using just 4–5 R-D data points, could this be sufficient? As codecs and quality metrics advance, does the conventional BD estimation still hold up? Crucially, are the performance improvements of new codecs and tools genuine, or merely artifacts of estimation flaws? We address these concerns by reevaluating BD estimation. We have established a large-scale, high-precision R-D dataset to verify the accuracy of existing BD estimation algorithms. Moreover, we propose a robust method for high-precision BD estimation across diverse compression scenarios, enhanced by a reliability assessment to determine the probability distribution of BD values from R-D sample points. This approach both assesses the reliability of BD calculations and serves as a precise BD estimator. Our method's validity is confirmed through extensive testing on a dataset we constructed. Our findings advocate for the adoption of rigorous R-D sampling and reliability metrics in future compression research to ensure the validity and reliability of results. Our code and additional experimental details are publicly accessible at https://github.com/fgvfgfg564/BDCI.
Xinyu Hang, Shenpeng Song, Zhimeng Huang, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC5
2025 Image Coding for Machine with Visual-Language Mimic Feature Learning
abstract
This paper propose a Image Coding for Machine (ICM) framework with Visual-Language Mimic Feature Learning (VLM-ICM). VLM-ICM decouples the position and semantic information into language modality and extracts universal features from the input image. Language, inherently more semantically compact, helps reduce the bitrate. Meanwhile, the universal features in VLM-ICM, guided by the language at the decoder side, allow for flexible domain adaptation, thereby enhancing versatility and practicality.
Zhimeng Huang, Junlong Gao, Jiaqi Zhang 0007, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia
DCC5
2025 Point Cloud-Assisted Neural Image Compression
abstract
High-efficient image compression is a critical requirement. In several scenarios where multiple modalities of data are captured by different sensors, the auxiliary information from other modalities are not fully leveraged by existing image-only codecs, leading to suboptimal compression efficiency. In this paper, we increase image compression performance with the assistance of point cloud, which is widely adopted in the area of autonomous driving. As depicted in Figure 1 (a), we have unified the digital representation of image and point cloud, and propose the point cloud-assisted neural image codec (PCA-NIC) to enhance the preservation of image texture and structure by utilizing the high-dimensional point cloud information. As depicted in Figure 1 (b), we further introduce a multi-modal feature fusion transform module (MMFFT) to capture more representative image features, remove redundant information between channels and modalities that are not relevant to the image content.
Ziqun Li, Qi Zhang 0042, Xiaofeng Huang, Zhao Wang 0004, Siwei Ma 0001
DCC5
2025 Dynamic Temporal Reference Aggregation for Neural Video Compression
abstract
Neural Video Compression (NVC) has advanced significantly in recent years, with improvements in inter prediction techniques. In inter prediction, most NVC approaches utilize pixel information or temporal features from neighboring frames as reference information, while using optical flow to represent motion information. In this paper, we introduce an innovative and efficient method for Dynamic Temporal Reference Aggregation (DTRA). The proposed DTRA consists of two components: Temporal Information Compensation (TIC) and Feature Level Motion Information Enhancement (MIE). The TIC module generates compensation information by leveraging long-term temporal information from the decoding buffer, enriching the semantic content of the reference features and enhancing their texture details. The MIE module refines the motion features at the encoder side and divides the motion information into multiple groups for diverse motion alignment at the decoder side, thereby improving the motion compensation. Extensive experiments demonstrate the effectiveness of the proposed method, achieving an average bitrate savings of 9.67% compared to state-of-the-art (SOTA) approaches.
Shuhong Liao, Kexiang Feng, Zhimeng Huang, Siwei Ma 0001, Chuanmin Jia
DCC4
2025 Enhanced Decoder-Side Secondary Transform Derivation for Video Coding Beyond AVS3
abstract
The Decoder-side Secondary Transform Derivation (DSTD) method has been adopted by the exploration software of the fourth generation Audio Video coding Standard (AVS4), Exploration Video Model (EVM). Although DSTD can significantly improve the coding performance, it increases the encoding complexity simultaneously. To reduce the encoding complexity of DSTD, three enhanced DSTD methods are proposed in this paper, which consists of Deleting Mode Optimization (DMO), Intra Mode Dependent Optimization (IMDO) and Interleaved Intra Mode Dependent Optimization (IIMDO). The process of DMO is similar to that of the original DSTD method, but all the contents related to diagonal flipping have been removed. IMDO divides intra prediction modes into three areas based on the angle of the intra prediction mode, as shown in Fig. 1. Each area will correspondingly reduce a secondary transform type. IIMDO divides intra prediction modes into four areas, as shown in Fig. 2. The experimental results demonstrate that the proposed methods can effectively reduce the encoding complexity with negligible coding performance loss. The first enhanced method, Deleting Mode Optimization, has been adopted by EVM.
Yuhuai Zhang, Jiaqi Zhang 0007, Weijia Jiang, Siwei Ma 0001
DCC6
2025 FAPC: Frequency-Based Adaptive Pixel Correction for Compressed Screen Content
abstract
Screen content is an important category of video. The statistical distribution of pixels in screen content exhibits substantial differences compared to camera-captured video, leading to different compression needs and challenges. Hitherto, most Screen Content Coding (SCC) tools are designed for block-based hybrid coding frameworks, which may not be suitable for emerging wavelet-based and learning-based coding frameworks. Consequently, a plug-and-play SCC tool independent of coding frameworks is lacking in the current video coding landscape. In this paper, an out-loop coding method, Frequency-based Adaptive Pixel Correction (FAPC), is proposed to improve the SCC performance for arbitrary codecs. First, a True Color Value (TCV) table is established based on the most frequently occurring pixel values. Then, the reconstructed pixels are corrected according to the TCV table. To realize precise pixel correction, an adaptive threshold derivation method is meticulously designed to control the pixel correction process. Furthermore, an inheritance coding strategy is proposed to reduce the overhead of parameter transmission. The proposed method has been integrated into three different coding frameworks. Simulation results demonstrate that the proposed method can achieve 10.66%, 3.55% and 10.67% luma component BD-BR gains on the three coding frameworks, respectively. These results prove the superiority and universality of the proposed method.
Zetian Song, Jiaqi Zhang 0007, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC6
2025 Compressed Domain Prior-Guided Video Super-Resolution for Cloud Gaming Content
abstract
Cloud gaming is an advanced form of Internet service that necessitates local terminals to decode within limited resources and time latency. Super-Resolution (SR) techniques are often employed on these terminals as an efficient way to reduce the required bit-rate bandwidth for cloud gaming. However, insufficient attention has been paid to SR of compressed game video content. Most SR networks amplify block artifacts and ringing effects in decoded frames while ignoring edge details of game content, leading to unsatisfactory reconstruction results. In this paper, we propose a novel lightweight network called Coding Prior-Guided Super-Resolution (CPGSR) to address the SR challenges in compressed game video content. First, we design a Compressed Domain Guided Block (CDGB) to extract features of different depths from coding priors, which are subsequently integrated with features from the U-net backbone. Then, a series of re-parameterization blocks are utilized for reconstruction. Ultimately, inspired by the quantization in video coding, we propose a partitioned focal frequency loss to effectively guide the model's focus on preserving high-frequency information. Extensive experiments demonstrate the advancement of our approach.
Qizhe Wang, Qian Yin 0002, Zhimeng Huang, Weijia Jiang, Siwei Ma 0001, Jiaqi Zhang 0007
DCC6
2025 A Hardware-Friendly AVS3 Entropy Decoder Architecture for 8K Ultra-High-Definition Video
abstract
Logarithmic binary arithmetic coding (LBAC) is a entropy coding method first introduced in AVS2 and now used in the third generation of audio video coding standard (AVS3). While LBAC provides high coding efficiency, the data dependencies in AVS3 make it challenging to parallelize parsing syntax element. To enhance the throughput of AVS3 LBAC entropy decoding engine, a sophisticated hardware implementation architecture is meticulously designed in the paper. The proposed method decouples the bitstream decoding process into four distinct modules. To ensure the decoding efficiency and performance, an advanced parallel pipeline is carefully designed. Furthermore, we develop sub-branch state prediction mechanism and grouped decoding method to further improve the decoding efficiency. The proposed AVS3 entropy decoder has been implemented on the S10 FPGA System Board, and experimental results show that the proposed method could achieve real-time and seamless decoding efficiency for 8K AVS3 bitstreams, even at bitrates as high as 207.1Mbps.
Jiaqi Zhang 0007, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001
DCC5
2025 MoRLACS: A Monocular RGBD-based Locomotion Approach for CAVE Systems
abstract
Navigation within Cave Automatic Virtual Environment (CAVE) systems often faces challenges due to limited physical space and the necessity for seamless user interaction. Traditional solutions typically rely on multi-view tracking systems or constrained locomotion techniques, which can interrupt immersion and hinder usability. In this paper, we introduce MoRLACS, a novel locomotion approach for CAVE systems that leverages a single RGBD camera. This hybrid framework integrates small-scale physical walking with controller-based large-scale exploration through a tailored guidance method. By accurately tracking the user's head position in the real world and synchronizing it with the virtual camera, MoRLACS enables natural walking within confined CAVE spaces and supports extended interaction in larger virtual environments. Preliminary user experiments demonstrate the approach's effectiveness, revealing improvements in usability and a heightened sense of presence. These findings underscore the potential of MoRLACS to enrich user experiences in immersive CAVE settings and offer valuable design insights for integrating 3D sensor data into multimedia interaction frameworks.
Haopeng Lu, Qian Yin 0002, Li Song 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICMR7
2024 Extreme Image Compression Using Fine-tuned VQGANs
abstract
Recent advances in generative compression methods have demonstrated remarkable progress in enhancing the perceptual quality of compressed data, especially in scenarios with low bitrates. However, their efficacy and applicability to achieve extreme compression ratios (< 0.05 bpp) remain constrained. In this work, we propose a simple yet effective coding framework by introducing vector quantization (VQ)–based generative models into the image compression domain. The main insight is that the codebook learned by the VQGAN model yields a strong expressive capacity, facilitating efficient compression of continuous information in the latent space while maintaining reconstruction quality. Specifically, an image can be represented as VQ-indices by finding the nearest codeword, which can be encoded using lossless compression methods into bitstreams. We propose clustering a pre-trained large-scale codebook into smaller codebooks through the K-means algorithm, yielding variable bitrates and different levels of reconstruction quality within the coding framework. Furthermore, we introduce a transformer to predict lost indices and restore images in unstable environments. Extensive qualitative and quantitative experiments on various benchmark datasets demonstrate that the proposed framework outperforms state-of-the-art codecs in terms of perceptual quality-oriented metrics and human perception at extremely low bitrates (≤ 0.04 bpp). Remarkably, even with the loss of up to 20% of indices, the images can be effectively restored with minimal perceptual loss.
Qi Mao 0002, Tinghan Yang, Meng Wang 0017, Shiqi Wang 0001, Libiao Jin, Siwei Ma 0001
DCC8
2024 Decoder-side Secondary Transform Derivation for Video Coding beyond AVS3
abstract
Secondary transform was adopted into the third generation Audio Video coding Standard (AVS3) to improve the intra-coded residual coding by applying a 4×4 secondary transform kernel. However, the adaptability of the single 4×4 transform kernel is limited for various residual data. In order to achieve higher residual coding gains, we propose a Decoder-side Secondary Transform Derivation (DSTD) method. Specifically, DSTD expands the maximum range of secondary transform from 4×4 to 8×8, where an 8×8 size transform kernel is introduced to further enhance the capability of compacting residuals. In particularly, three flipped secondary transform types are employed to extend transform candidates, including horizontal, vertical and diagonal flipping types. The boundary continuity is utilized to derive the transform type. Experimental results show that the proposed method can achieve 0.51% and 0.18% BD-rate savings on average under All Intra (AI) and Random Access (RA) configurations, respectively. DSTD has been adopted into the Exploration Video Model (EVM) for AVS4.
Yuhuai Zhang, Huiwen Ren, Shiqi Wang 0001, Siwei Ma 0001
DCC6
2024 A Neural-network Enhanced Video Coding Framework beyond ECM
abstract
In this paper, a hybrid video compression framework is proposed that serves as a demonstrative showcase of deep learning-based approaches extending beyond the confines of traditional coding methodologies. The proposed hybrid framework is founded upon the Enhanced Compression Model (ECM), which is a further enhancement of the Versatile Video Coding (VVC) standard. We have augmented the latest ECM reference software with well-designed coding techniques, including block partitioning, deep learning-based loop filter, and the activation of block importance mapping (BIM) which was integrated but previously inactive within ECM, further enhancing coding performance. We evaluate the coding performance of the proposed framework with extensive experiments on the JVET dataset compared with ECM10.0 and VTM-11.0. Due to the testing environment and the coding complexity of the ECM, we did not conduct testing on Class A. The QPs are set as 22, 27, 32, 37, and 42. Compared with ECM-10.0, our method achieves 6.26%, 13.33%, and 12.33% BD-rate savings for the Y, U, and V components under random access (RA) configuration. The traditional hybrid coding framework combined with the three coding tools can further improve compression efficiency and has great potential for performance improvement.
Yanchen Zhao, Chuanmin Jia, Qizhe Wang, Yue Li 0015, Chaoyi Lin, Kai Zhang 0007, Li Zhang 0006, Siwei Ma 0001
DCC10
2024 A Dynamic Point Cloud Dataset for MPEG Point Cloud Compression and Performance Analysis
abstract
Recent years witnessed the development in MPEG point cloud compression (PCC). However, the exploration of inter-frame coding may be impeded due to the lack of dynamic point clouds (point cloud sequences). To promote the development of PCC technology, we propose Dynamic3D , a dynamic 3D point cloud dataset with high-quality real-captured 3D persons and objects. There are several appealing properties: 1) Dynamic scenes: It contains five sequences and each sequence comprises 600 frames with temporal variation; 2) Complex content: instead of a single person or object in the existing dataset from MPEG, our established dataset contains multiple persons or both person and objects; 3) Realistic capture: the color industrial cameras and infrared cameras are used for data acquisition. This dataset provides the vast exploration space for PCC, especially the elimination of temporal redundancy. Extensive simulations are conducted on this dataset by using the reference software of MPEG G-PCC and V-PCC, i.e., (GeS-TM and TMC2), delivering observations, analysis and opportunities for the future research of PCC.
Lili Zhao 0001, Qian Yin 0002, Lancao Ren, Lei Yang 0063, Chuanmin Jia, Siwei Ma 0001
DCC6
2023 Rate-Distortion Optimization for Cross Modal Compression
abstract
Recently, cross modal compression (CMC) is proposed to compress highly redundant visual data into a compact, common, human-comprehensible domain (such as text) to preserve semantic fidelity for semantic-related applications. However, CMC only achieves a certain level of semantic fidelity at a constant rate, and the model aims to optimize the probability of the ground truth text but not directly semantic fidelity. To tackle the problems, we propose a novel scheme named rate-distortion optimized CMC (RDO-CMC). Specifically, we model the text generation process as a Markov decision process and propose rate-distortion reward which is used in reinforcement learning to optimize text generation. In rate-distortion reward, the distortion measures both the semantic fidelity and naturalness of the encoded text. The rate for the text is estimated by the sum of the amount of information of all the tokens in the text since the amount of information of each token is a lower bound of coding bits. Experimentally, RDO-CMC effectively controls the rate in the CMC framework and achieves competitive performance on MSCOCO dataset.
Junlong Gao, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC4
2023 Learning to Compress Unmanned Aerial Vehicle (UAV) Captured Video: Benchmark and Analysis
abstract
In this paper, we propose to build a novel benchmark and neural video coding task named learning based Unmanned Aerial Vehicle (UAV) video coding. We collect the UAV videos with different content variations, including in-door and out-door scenes, object-scale variations and viewpoint distance, different climate condition etc. Then we encode those properly-selected videos using popular end-to-end optimized video codecs and conventional hybrid codecs, to form a comprehensive benchmark for learned drone video compression. We also provide a detailed analysis and envision the challenge of such task for future research. The main contributions of this paper are three folds. First, we construct a comprehensive benchmark for the task of drone video compression which consists of the rate-distortion (R-D) behavior of both hybrid and learned video codecs. To our knowledge, it is the first attempt in end-to-end optimized solution to compress drone videos. Second, we provide the review and analysis of the learned drone video compression schemes and further discuss the challenges of encoding UAV videos. Third, this benchmark and related research is accomplished as a milestone MPAI End-to-end Video (EEV) coding project. The proposed benchmark has constructed a solid baseline for compressing UAV videos and facilitates the future research works for related task.
Chuanmin Jia, Huifang Sun, Siwei Ma 0001, Wen Gao 0001
DCC4
2023 An Efficient Rate Control Scheme for Video Compression in Low-latency Interoperable Interfaces
abstract
Lightweight video compression has effectively alleviated the tension between growing transmission demands and expensive integration upgrades. Effective rate control algorithms are believed to be the crucial bottleneck for quality improvement during those ultra-high throughput coding processes. This paper proposes a novel rate control (RC) scheme that constructs a contextual adaptive bit estimation model through clustering historical compression information into block-gradient complexity categories. A buffer-aware tuning method and a flexible quantization parameter (QP) mapping algorithm are designed to determine the Luma/Chroma QP distribution where a simplified Lagrangian multiplier is further defined to preserve the stability of the overall compression process. As a result, the constant-bitrate compression towards low-latency interoperable ASICs is implemented with a promising RC performance.
Huiwen Ren, Zetian Song, Yan Wang 0011, Shanshe Wang, Fangdong Chen, Shiliang Pu, Siwei Ma 0001, Wen Gao 0001
DCC9
2023 An Adaptive Intra-frame Quantization Parameter Derivation Model Jointing with Inter-frame Analysis
abstract
This paper proposes a novel quantization parameter (QP) derivation module that constructs several spatiotemporal characteristics into key-frame QP determination through an efficient pre-analysis progress. A series of simplified prediction modes and a histogram statistic are employed to model the reference quality that key-frames provide to subsequent frames. An adaptive delta-QP value is generated to address the conflict between the low compression efficiency of intra-only frames and the critical predictive basis of temporal-underlying frames. The experimental result shows that the proposed method reduces 41.91% of peak-to-valley bitrate difference while leading to a 0.02% BDBR performance change, indicates that high-quality key-frames may not be indispensable in nowadays video compression frameworks. The proposed method has shown that the pre-analysis based QP optimization for intra-only frames is promising for the enhancement of transmission bandwidth utilization, which may hopefully provide new inspiration for bit allocation and rate control designs.
Huiwen Ren, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC3
2022 A Smart Reference Picture Resampling Approach for VVC
abstract
Resampling-based coding, i.e. down-sampling before encoding and up-sampling after decoding, has been recognized to be an effective tool for compressing high-resolution videos at low bitrates. The newest video coding standard, Versatile Video Coding (VVC), supports resampling-based coding via a mechanism named Reference Picture Resampling (RPR), where the spatial resolution can be changed without inserting an intra frame. Intuitively, it is not wise to utilize a single resolution throughout the whole video, because frames with different contents may prefer different coding resolutions. In this paper, we propose a smart reference picture resampling approach, namely smart-RPR, where the coding-resolution of a frame is determined based on the property of the frame without multiple-pass encoding. Specifically, we first down- and up-sample a frame without considering compression and compare the up-sampled frame with the original frame to obtain the resampling distortion, which is then compared with a threshold to decide whether to code the frame in a resampling way. Then, we build up an exponential model to approximate the optimal threshold. In addition, we also study how to derive the coding parameters of the down-sampled frame to achieve better performance. Simulation results on the VTM-12.0 show that the proposed method could achieve 2.72%, 5.29%, and 10.82% BD-rate reductions for Y, Cb, and Cr components, respectively, with lower encoding and decoding complexity.
Tianliang Fu, Kai Zhang 0007, Yue Li 0015, Li Zhang 0006, Shanshe Wang, Siwei Ma 0001
DCC6
2022 Rate Distortion Characteristic Modeling for Neural Image Compression
abstract
End-to-end optimized neural image compression (NIC) has obtained superior lossy compression performance recently. In this paper, we consider the problem of rate-distortion (R-D) characteristic analysis and modeling for NIC. We make efforts to formulate the essential mathematical functions to describe the R-D behavior of NIC using deep networks. Thus arbitrary bit-rate points could be elegantly realized by leveraging such model via a single trained network. We propose a plugin-in module to learn the relationship between the target bit-rate and the binary representation for the latent variable of auto-encoder. The proposed scheme resolves the problem of training distinct models to reach different points in the R-D space. Furthermore, we model the rate and distortion characteristic of NIC as a function of the coding parameter$\lambda$respectively. Our experiments show our proposed method is easy to adopt and realizes state-of-the-art continuous bit-rate coding performance, which implies that our approach would benefit the practical deployment of NIC.
Chuanmin Jia, Ziqing Ge, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC4
2022 Coarse-to-fine Prediction With Local and Nonlocal Correlations for Intra Coding
abstract
Recently many efforts have been devoted to learning non-linear predictions from neighboring samples with deep neural networks. However, existing methods mainly generate predictions with local reference samples, regardless of nonlocal self-similarity. In this paper, we aim to incorporate local and nonlocal correlations for intra prediction and propose a two-stage coarse-to-fine network (CTFN), which is integrated into VVC codec as an optional intra prediction mode. The prediction process of CTFN is decomposed into two stages. In the first stage, we train a set of networks to generate a coarse result with local reference samples. In the second stage, we extract sufficient features from nonlocal region using the coarse result as priors and transform the features into a fine prediction result. In particular, a patch-wise attention layer (PAL) is designed in the second stage that can fully explore nonlocal correlations in feature domain and assign weights to each nonlocal feature adaptively, as shown in Fig. 1. As such, the proposed CTFN can not only learn a non-linear mapping from local context, but also explicitly borrow similar features from nonlocal region in a weighted form. Different from image inpainting tasks, the patch synthesis problem is converted to patch matching problem with the CTFN, yielding more reliable predictions. More-over, we construct a classified dataset based on Pearson Correlation Coefficient for network training to better handle contents that are highly correlated. Experiments on VTM-11.0 show that the proposed network achieves 1.77% ED-rate reductions under all intra configuration, which outperforms the state-of-the-art methods.
Meng Lei, Xuewei Meng, Chuanmin Jia, Shanshe Wang, Zhipeng Cheng, Siwei Ma 0001
DCC6
2022 Parametric Non-local In-loop Filter for Future Video Coding
abstract
In-loop filter has been comprehensively explored during the development of video coding standards to suppress compression artifacts. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take advantage of the image local similarity. Although some non-local based in-loop filters can make up for this short-coming, the unsupervised parameter selection scheme, which is widely used by non-local filters, limits the content adaptability. Given this, we propose a parametric non-local in-loop filter (PNLF) that fully considers the non-local characteristics and trains the filter coefficients based on the video content. In the filtering process, the reference samples based on the non-local similarity are first derived for each to-be-filtered sample. Then to-be-filtered samples are grouped into specific classes based on multiple features. For each class, filter coefficients are online trained in the encoder and transmitted to the decoder. Finally, the filtering process is conducted using the online-selected coefficients. Simulation results reveal that the proposed approach achieves 0.70%, 1.43%, and 2.09% bit-rate savings on average compared to VTM-11.0 under All Intra (AI), Random Access (RA), and Low-Delay B (LDB) configurations, respectively. The sequences used in the experiment include Class AI, A2, B, C, D, E, F, and SCC. Compared to the non-local structure-based filter (NLSF) [1], our proposed PNLF with fast block matching scheme [2] applied on B-frames and P-frames can achieve better performance gain with lower software and hardware complexity under RA and LDB configurations.
Xuewei Meng, Chuanmin Jia, Xinfeng Zhang 0001, Meng Lei, Shanshe Wang, Lin Li 0062, Siwei Ma 0001
DCC7
2022 Fast Partition Mode Decision via a Plug-in Fully Connected Network for Video Coding
abstract
Flexible coding unit partitioning such as quad-tree nested binary-tree and ternary-tree adopted by the emerging enhanced compression model (ECM) brings promising coding performance improvement. Meanwhile, the computational complexity increases dramatically, which may block the exploration and validation of new coding tools. This paper investigates a partition mode early pruning scheme via a fully connected network to reduce the encoding complexity for the ECM. In particular, we carefully select features and devise the fully connected network, which could seamlessly cooperate with the encoder, revealing promising learning and inference capability. Experimental results demonstrate that the proposed method achieves 15%~50% encoding time savings with moderate bit-rate increasing on the ECM, and the extra complexity regarding the fully connected network and feature extraction is negligible.
Jiaqi Zhang 0007, Meng Wang 0017, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC6
2022 Analysis on Compressed Domain: A Multi-Task Learning Approach
abstract
Image compression approaches based on deep learning have achieved remarkable success. Existing studies mainly focus on human vision and machine analysis tasks taking reconstructed images as input. However, those methods need images to be decoded before performing downstream visual tasks, which motivates us to explore how to directly conduct visual analysis using the compressed data without decoding. The overview of our proposed model is shown as Fig. 1(a). Specifically, a task-agnostic learning-based compression model is proposed, which effectively supports various compressed domain-based analytical tasks meanwhile reserves outstanding re-constructed perceptual quality compared with traditional and learning-based codecs. To obtain the extremely compacted data representation with essential semantic infor-mation, we take the help of the generative model on decoder part. Then, we propose a multi-task learning model which can directly obtain semantic information from the compressed visual data. The pipeline of the proposed model is detailedly illus-trated in Fig. 1(b). In addition, joint optimization strategy is adopted to achieve the best balance point among compression efficiency, reconstructed image quality, and the downstream visual tasks' performance. Experimental results verify that our proposed compressed domain-based multi-task analysis model outperforms the reconstructed image-based method on transmission efficiency, saving more than ten times of bit-rate consumption while preserving comparable visual analysis precision (i.e., classification and segmentation tasks) when compared with RGB image input models, which is evaluated on the CelebA-HO dataset.
Yuefeng Zhang, Chuanmin Jia, Jianhui Chang, Siwei Ma 0001
DCC4
2022 Interpretable Learned Image Compression: A Frequency Transform Decomposition Perspective
abstract
Image compression is a key problem in this age of information explosion. With the help of machine learning, recent studies have shown that learning-based image compression methods tend to surpass traditional codecs. Image compression can be split into three steps: transform, quantization, and entropy estimation. However, the transform step in traditional codecs lacks flexibility because of the strict mathematical premise while the transform in most learning-based codecs neglects its intrinsic interpretation. After observing compression degradation degree varies on different frequency bands as illustrated as Fig. 1(a), we propose an end-to-end compression model from the frequency perspective with a frequency-pyramid transform and a frequency-aware fusion module. The right of the Fig. 1(a) displays each frequency layer's component of the proposed model from low to high-frequency splits. Intuitively, we can infer that the low-frequency part contains the global structure while the high-frequency part gets finer details, satisfying the feature of human visual system (HVS). The proposed model are detailedly shown in Fig. 1(b) that independent probability estimation models are set for each frequency split. Extensive experiments are conducted to demonstrate that our model outperforms all traditional codecs (e.g., JPEG, JPEG2000, HEVC, and VVC) on MS-SSIM metric on both Kodak and CLIC2020 professional test datasets. Taking BPG-4:4:4 as the anchor, our proposed model achieves 11.6% BD-rate reduction under PSNR measurement, which is evaluated on the Kodak dataset.
Yuefeng Zhang, Chuanmin Jia, Siwei Ma 0001
DCC4
2021 Intra Block Partition Structure Prediction via Convolutional Neural Network
abstract
In video coding, block partition segments images into non-overlap blocks for individual coding, the structure of which is becoming more and more flexible along with the development of video coding standards. Multiple types of tree structures have been proposed recently, which extensively improved the complexity of the encoding process due to the recursive rate-distortion search for the optimal partition. In this paper, a two-stage Convolutional Neural Network (CNN) based partition structure prediction method is proposed to bypass the decision process of the block size in intra frame coding. Specifically, the Coding Unit (CU) partition is first represented in sub-block granularity and predicted by the end-to-end trained CNNs. Then, the final partition structure compatible with the coding standard is derived from the prediction results directly. Experimental results show that the proposed CNN based partition method achieves about 56 times speedup (97% time-saving) with 9% BD-rate degradation against the reference software of the latest AVS3 coding standard (IEEE Standard 1857.10).
Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC4
2021 Quad-Treea Based Sample Refinement Filter for Video Coding
abstract
In-loop filter is a crucial module in video coding, which can improve both subjective and object quality of reconstructed videos. In this paper, a new sample-based classification method is first proposed using features extracted from different stages of the existing in-loop filter process. Based on this method, an adaptive three-layer Quad-tree Based Sample Refinement Filter (QSRF) algorithm is designed to further improve the coding efficiency. Experimental results show that the proposed QSRF algorithm achieves 0.39%, 0.77% and 0.70% BD-rate savings for random access, lowdelay B and lowdelay P configurations compared to AVS3 reference software, respectively. Moreover, the proposed method can also improve visual quality of reconstructed videos significantly.
Yunrui Jian, Jiaqi Zhang 0007, Chuanmin Jia, Suhong Wang, Shanshe Wang, Siwei Ma 0001
DCC6
2021 Optimized Adaptive Loop Filter in Versatile Video Coding
abstract
In the Versatile Video Coding (VVC) standard, adaptive loop filter (ALF), including Geometry transformation-based Adaptive Loop Filter (GALF) and Cross Component Adaptive Loop Filter (CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is unfriendly to the software and hardware design. Therefore, we propose an optimized ALF framework, including the parallel design of GALF and CCALF, the adaptive parameter decision of GALF, and one-pass CCALF scheme by effectively estimating the CCALF filtering distortion without conducting filter operation. Compared to VTM-8.0, the proposed method can reduce the picture buffer access from 152 to 1 and achieve roughly 25% time-savings of the ALF module with negligible coding performance change under RA configuration. Some of the proposed methods have been adopted in the VVC reference software.
Xuewei Meng, Jiaqi Zhang 0007, Chuanmin Jia, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC6
2021 Flow-Grounded Dynamic Texture Synthesis for Video Compression
abstract
The basic ingredients of modern video coding standards are block-based prediction and transforms. However, when dealing with video contents containing dynamic textures (DT), the existing prediction schemes usually failed due to temporal variability and randomness of DT, which results in more bit cost on residual coding compared with other contents. In view of this point, a novel video compression scheme for DT is proposed in this work. In particular, wavelet-based analysis on motion characteristics of DT is firstly presented and based on the analysis, we introduce a flow-grounded texture synthesis method for video compression. Instead of conventional inter prediction, synthesized DT contents are used for reconstruction at the decoder. The proposed scheme has been fully integrated into the test model of Versatile Video Coding standard, VTM-10.0, for validation and a subjective test has also been carried out. Experimental results show that bitrate savings can be achieved by 40% on average at comparable visual quality.
Suhong Wang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC4
2021 Point AE-DCGAN: A deep learning model for 3D point cloud lossy geometry compression
abstract
3D point cloud has been widely applied in virtual reality and augmented reality. A complex 3D scene always needs a large number of the point cloud to represent and demands a lot of space to store. Thus, point cloud compression becomes a crucial issue to research. In this paper, we propose a novel lossy geometric compression method of autoencoder based on DCGAN optimization. This method can reconstruct a high-quality point cloud and solves a large area of missing points in the process of compression and decompression. To improve the point cloud codec performance, we propose a multi-scale 3D deconvolution hopping connection structure to obtain a better-quality reconstructed point cloud under low bit rates. Our approach is the first GAN-based point cloud compression algorithm to our knowledge. Compared with state-of-the-art methods on the MVUB dataset, our approach achieves a better rate-distortion performance and visual quality.
Zhijun Fang 0001, Yongbin Gao, Siwei Ma 0001, Yaochu Jin, Anjie Wang
DCC4
2020 Gradient-Based Early Termination of CU Partition in VVC Intra Coding
abstract
The quaternary tree with nested binary and ternary tree structure is an efficient coding unit (CU) partitioning method adopted in Versatile Video Coding (VVC). Compared with quaternary tree only in HEVC, its flexible block sizes improve the coding performance significantly at the cost of computation load increase due to the recursive and nested searching for the best CU structure. In this paper, an early termination algorithm is proposed to skip unnecessary searches in CU size decision. Based on directional gradients, pre-determine the likelihood of binary partition or ternary partition in horizontal or vertical direction for current block, thus the impertinent partition pattern is skipped. The experimental results show that the proposed method is able to save the encoding time up to 51% with about 1.2% BD-rate degradation compared with VVC software reference VTM5.0.
Tao Zhang 0013, Chenchen Gu, Xinfeng Zhang 0001, Siwei Ma 0001
DCC5
2020 Sub-Sampled Cross-Component Prediction for Chroma Component Coding
abstract
Cross-component prediction, which takes advantage of inter-channel correlations, predicts the chroma block with the luma reconstructed block according to associated linear model. Instead of involving all available reference samples in building the linear model, in this paper, we propose a sub-sampled approach that utilizes at most four neighboring chroma samples and their corresponding down-sampled luma samples, leading to significantly reduced operations in the derivation of model parameters at both encoder and decoder. The proposed scheme is hardware friendly in terms of the overheads of memory access and clock cycles, and greatly benefits the practical implementations of the emerging video coding standard in real applications. Extensive experiments reveal that the proposed sub-sampled method provides simple operations and robust coding performance, leading to the adoption by Versatile Video Coding (VVC) Standard and the third generation Audio Video Coding Standard (AVS3).
Meng Wang 0017, Li Zhang 0006, Kai Zhang 0007, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC7
2019 Perceptual Video Coding Based on Visual Saliency Modulated Just Noticeable Distortion
abstract
To reduce the perceptual redundancy in the video coding process, human visual system (HVS)-based visual attention and visual sensitivity can be utilized due to their intrinsic natures. Just Noticeable Distortion (JND) is one of widely used models to simulate human visual sensitivity, while visual saliency map has been popular for years in image processing to describe the visual attention feature, which has been proved by the ability to enhance the visual sensitivity effect. In this paper, we proposed a perceptual video coding (PVC) scheme with visual saliency modulated JND model to suppress the DCT coefficient without resulting in noteworthy subjective quality degradation. The experimental results show that the PVC scheme with the proposed VS-JND model can save bit rates up to 35.58% in high bit rates case with the similar subjective quality compared with that of HEVC software reference code HM 16.12.
Ruiqin Xiong, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC5
2019 Bi-Intra Prediction for Versatile Video Coding
abstract
This paper presents a novel Bi-Intra Prediction (BIP) algorithm to improve intra coding performance for the next generation of video coding. In the proposed algorithm, a new predictor is generated by combining two existing intra prediction modes as an additional mode, being able to provide more accurate prediction. To remove the number of unnecessary combinations, restrictions on the search candidates and block sizes are carried out based on the statistical analyses. In addition, an efficient mode coding method of syntax elements in BIP is introduced to improve the coding performance. Moreover, a rough mode decision scheme is adopted to avoid high computation complexity in encoder side. Experimental results show that compared with the Versatile Video Coding reference software VTM-1.0, the proposed algorithm achieves 0.71%, 0.46% and 0.43% BD-Rate gains on average for Y, U and V components under all intra configuration, respectively, and 0.32%, 0.39%, 0.49% BD-Rate gains under random access configuration.
Congrui Li, Zhenghui Zhao, Xiang Zhang 0004, Siwei Ma 0001
DCC5
2019 Adaptive Wavelet Domain Filter for Versatile Video Coding (VVC)
abstract
Owing to the ability of removing compression artifacts, extensive in-loop filters have been proposed for video coding standards. They are performed after the reconstruction of all coding units (CUs), however, none of them has been taken into account in the mode decision when coding each CU. To address this issue and make the rate-distortion optimization (RDO) more precise for each CU, we introduce a low-pass filter when checking the rate-distortion cost after the reconstruction of each CU. Specifically, based on Haar wavelet, the reconstructed block is transformed to the frequency domain, and then an adaptive wavelet domain filter (AWF) is proposed to suppress the quantization noises in coded blocks. To be adaptive, the filter strength varies from CU to CU according to the texture complexity and quantization parameters (QPs). Experimental results show that the proposed method can reduce the compression artifacts and improve both the objective and subjective quality.
Suhong Wang, Xiang Zhang 0004, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC4
2019 Extended Quad-Tree Partitioning for Future Video Coding
abstract
The quad-tree plus binary-tree (QTBT) coding unit (CU) partitioning structure, which has been adopted to the next generation video coding standard, shows promising coding performance when compared with the conventional quad-tree structure in HEVC. In this paper, we propose the Extended Quad-tree (EQT) partitioning, which further extends the QTBT scheme and increases the partitioning exibility. More specifcally, EQT splits a parent CU into four sub-CUs of dierent sizes, which can adequately model the local image content that cannot be elaborately characterized with QTBT. Meanwhile, EQT partitioning allows the interleaving with BT partitioning for enhanced adaptability. Experimental results on the JEM7-QTBT-Only platform show that EQT brings better coding performance with 3.17%, 3.20% and 3.06% BD-Rate gains under random access, low-delay P and low-delay B configurations, respectively.
Meng Wang 0017, Li Zhang 0006, Kai Zhang 0007, Hongbin Liu 0004, Shiqi Wang 0001, Sam Kwong, Siwei Ma 0001
DCC8
2018 Locally Refined Motion Compensation for Future Video Coding
abstract
Motion compensation plays a key role in high efficiency video coding. The popular video compression standards, such as H.264/AVC and HEVC, adopt block based motion compensation technique due to its high compression efficiency and relatively low computational complexity. However, block based motion compensation may not be in accordance with the actual object boundary, potentially leading to low prediction accuracy especially in the high-texture areas. In this paper, we propose a locally refined motion compensation method to address this issue. In particular, the image segmentation is applied on the prediction block indicated by a motion vector rather than the original block to avoid explicit signaling. Furthermore, the local content is analyzed to select one segmented region and subsequently the prediction of this region is generated based on the local motion filed. Experimental results show that the proposed algorithm can achieve 0.8%, 1.1% and 1.7% bitrate savings for Random Access, Lowdelay-B and Lowdelay-P configurations respectively without introducing noticeable computational complexity.
Zhao Wang 0004, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC5
2018 A Group Variational Transformation Neural Network for Fractional Interpolation of Video Coding
abstract
Motion compensation is an important technology in video coding to remove the temporal redundancy between coded video frames. In motion compensation, fractional interpolation is used to obtain more reference blocks at sub-pixel level. Existing video coding standards commonly use fixed interpolation filters for fractional interpolation, which are not efficient enough to handle diverse video signals well. In this paper, we design a group variational transformation convolutional neural network (GVTCNN) to improve the fractional interpolation performance of the luma component in motion compensation. GVTCNN infers samples at different sub-pixel positions from the input integer-position sample. It first extracts a shared feature map from the integer-position sample to infer various sub-pixel position samples. Then a group variational transformation technique is used to transform a group of copied shared feature maps to samples at different sub-pixel positions. Experimental results have identified the interpolation efficiency of our GVTCNN. Compared with the interpolation method of High Efficiency Video Coding, our method achieves 1.9% bit saving on average and up to 5.6% bit saving under low-delay P configuration.
Sifeng Xia, Wenhan Yang, Yueyu Hu, Siwei Ma 0001, Jiaying Liu 0001
DCC4
2017 Wireless Image SoftCast Using Compressive Gradient
abstract
Summary form only given: Based on observations that the visual quality has strong correlation with image gradients, gradient based image SoftCast (G-Cast) [1] advocates to convey visual information by delivering image gradients. In G-Cast, both horizontal gradients and vertical gradients needs to be transmitted, even if the channel bandwidth is insufficient. This paper propose to send out the random projection measurements of the gradients instead of delivering gradients directly, so that data size can be reduced to an arbitrary ratio and channel bandwidth occupation can be lowered. We name this scheme as compressive gradient based SoftCast (CG-Cast). At CG-Cast sender, after generated by gradient transform, the gradients are down-sampled by random projection, sample rate of which is set according to the channel bandwidth condition. Then the produced measurements are sent out for raw OFDM transmission. A few lowfrequency components are also transmitted to tell the global luminance. At CG-Cast receiver, the received noisy measurements are used for compressive gradient based reconstruction procedure, which utilizes sparsity in gradient domain and non-local similarity in spatial domain [2]. The proposed method is compared with SoftCast [3] and compressive sensing (CS) in bandwidth limited and power constrained scenarios. To make fair comparison, these three schemes are tested under the same channel signal-to-noise ratio (CSNR) conditions to transmit equal amount of data for reconstruction, using equivalent power and bandwidth. CG-Cast outperforms SoftCast and CS in terms of SSIM and gradient signal-to-noise ratio (GSNR) at different bandwidth ratios. Comparing with SoftCast in different channel conditions, the average SSIM gain of all the tested images varies from 0.04 to 0.13, and the average GSNR gain ranges from 1.5dB to 2.9dB. CS is rather unstable in noisy conditions. SoftCast performs better than CS because of its power allocation.
Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001, Wen Gao 0001
DCC4
2017 Band-Wise Adaptive Sparsity Regularization for Quantized Compressed Sensing Exploiting Nonlocal Similarity
abstract
The theory of compressive sensing (CS) has attracted considerable research interests from signal and image processing communities. And in practice, because of the considerations of data storage and transmission, scalar quantization is necessary to be implemented on the CS measurements. In this paper, we propose an adaptive bandwise sparsity regularization to handle the recovery problem of quantized compressive sensing. The sparsity regularization constraints every patch by using bandwise distribution model in transform domain. In addition, we bring in the quantization cost function to quantify the influence of measurement quantization. Experimental results demonstrate that our CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes with both unquantized measurements and quantized measurements.
Ruiqin Xiong, Xinfeng Zhang 0001, Siwei Ma 0001
DCC4
2017 Effective Quadtree Plus Binary Tree Block Partition Decision for Future Video Coding
abstract
Block partition structure has been recognized as a crucial module in video coding scheme. Recently, a quadtree plus binary tree (QTBT) block partition structure has been proposed in the Joint Video Exploration Team (JVET) development. Compared to the quadtree structure in HEVC, QTBT can achieve better coding performance with hugely increased encoding complexity. Here, we propose an effective QTBT partition decision algorithm to achieve a good trade-off between computational complexity and coding performance. In particular, at the Coding Tree Unit level, the partition parameters of QTBT are dynamically derived to adapt to the local characteristics without transmitting any overhead. Subsequently, at the Coding Unit level, a joint-classifier decision tree structure is designed to eliminate unnecessary iterations and meanwhile control the risk of false prediction. Experimental results show that the proposed algorithm can achieve 64% encoding time reduction on average with only 1.26% increase in terms of bit rate. This greatly benefits the practical implementations of QTBT in real application scenarios.
Zhao Wang 0004, Shiqi Wang 0001, Jian Zhang 0018, Shanshe Wang, Siwei Ma 0001
DCC5
2017 Globally Variance-Constrained Sparse Representation for Rate-Distortion Optimized Image Representation
abstract
Sparse representation is efficient to approximately recover signals by a linear composition of a few bases from an over-complete dictionary. However, in the scenario of data compression, its efficiency and popularity are hindered due to the extra overhead for encoding the sparse coefficients. Therefore, how to establish an accurate rate model in sparse coding and dictionary learning becomes meaningful, which has been not fully exploited in the context of sparse representation. According to the Shannon entropy inequality, the variance of data source can bound its entropy, thus can reflect the actual coding bits. Therefore, a Globally Variance-Constrained Sparse Representation (GVCSR) model is proposed, where a variance-constrained rate term is introduced to the conventional sparse representation. To solve the non-convex optimization problem, we employ the Alternating Direction Method of Multipliers (ADMM) for sparse coding and dictionary learning, both of which have shown state-of-the-art rate-distortion performance in image representation.
Xiang Zhang 0004, Siwei Ma 0001, Zhouchen Lin, Jian Zhang 0018, Shiqi Wang 0001, Wen Gao 0001
DCC2
2016 Adaptive Motion Vector Resolution Scheme for Enhanced Video Coding
abstract
In the state-of-the-art H.265/HEVC video coding standard, the motion vector is always fixed to be 1/4-pixel resolution for the entire video sequence regardless of the different video contents, which is not efficient for prediction coding. In this paper, we propose a frame level adaptive motion vector resolution selection scheme based on a rate-distortion model in terms of motion vector resolution. In the proposed rate-distortion model, the relationship between the distortion and the motion vector resolution is approximated with a linear model. And a rate model of motion vector is built, which reflects the relationship between the coding bits of motion vector and its value. With the proposed rate-distortion model, an optimal motion vector resolution minimizing the total rate-distortion cost will be selected for each frame. Experimental results show that the proposed scheme can achieve 1.5%, 1.3% and 2.5% BD-rate gain on average for Random Access, Lowdelay-B and Lowdelay-P configurations without complexity increment.
Zhao Wang 0004, Jian Zhang 0018, Nan Zhang 0015, Siwei Ma 0001
DCC4
2016 Structure-driven Adaptive Non-local Filter for High Efficiency Video Coding (HEVC)
abstract
Deblocking filter (DF) Is High Efficiency Video Coding (HEVC) is Only Applied to all Samples Adjacent to prediction units (PU), or transform units (TU), which actually exists two issues. The first one is that DF in HEVC does not fully exploit nonlocal similarity structure information in video. The second one is that DF is HEVC does not consider the inside pixels, which often suffer from quantization distrotion. To alleviate these issues, in this paper, a structure-driven adaptive non-local filter (SANF) Is Proposed By Simultaneously Enforcing The Intrinsic Local Sparsity And The Non-Local Self-Similarity Of Each Frame. Not only SANF deals with the boundary pixels, but also the inside area, which is able to effectively reduce block artifacts while enhancing the quality of the deblocked frames. Applying SANF to luma and chroma components after DF, simulation results demonstrate that the proposed SANF can save BD-rate reduction up to 10.3% with ALF off. For luma component, SANF achieves 4.1%. 3.3%, 4.4% BD-rate saving for all intra, low delay B and random access configurations, respectively with ALF off. furthermore, the performance with ALF on is also discussed.
Jian Zhang 0018, Chuanmin Jia, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001
DCC4
2016 From Visual Search to Video Compression: A Compact Representation Framework for Video Feature Descriptors
abstract
Visual feature descriptors have been successfully deployed in a wide range of applications, e.g. visual retrieval and analysis. To transmit these descriptors over bandwidth-limited networks, a high efficiency feature coding technique is highly desired to maximize compression capability and achieve compact feature representations. In this paper, a hybrid visual feature descriptor compression framework is presented and implemented in the encoding and decoding loops of texture videos. In particular, the multiple-hypothesis prediction is employed to effectively remove redundancies originated not only from spatial and temporal similarities, but also from reconstructed video frames. As the ultimate purpose of the transmitted descriptors is retrieval, the rate-accuracy optimization (RAO) technique is proposed to obtain the best tradeoff between the rate and retrieval performance. Such paradigm enables the conventional video stream to achieve high efficient retrieval/analysis with very low bitrate consumption. Moreover, we also demonstrate that texture video compression can also benefit from the additional information provided by the transmitted descriptors, leading to significantly improvement of coding efficiency on top of the high efficiency video coding (HEVC) standard. Extensive simulations have shown that the proposed method can offer significant bitrate reduction in representing both the descriptors and texture video frames, and meanwhile providing desirable retrieval performance.
Xiang Zhang 0004, Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Xinfeng Zhang 0001, Wen Gao 0001
DCC2
2016 Nonconvex Lp Nuclear Norm based ADMM Framework for Compressed Sensing
abstract
Compressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression methodology. Recent studies further show that image prior models play an important role in image CS recovery. By exploiting the non-local self-similarity of natural images and clustering similar patches, low-rank prior model is adopted in this paper. Different from traditional nuclear norm, we extend thelp(0plpnuclear norm prior model for image CS recovery, which is able to more accurately enforce image structural sparsity and self-similarity at the same time. The proposed optimization problem is efficiently solved within the alternative direction multiplier method (ADMM) framework. Experimental results demonstrate that the proposedlpnuclear norm based ADMM framework for image CS recovery framework exhibits good convergence and achieves significant performance improvements over the current state-of-the-art methods.
Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001
DCC3
2016 Compressive-Sensed Image Coding via Stripe-based DPCM
abstract
These years have seen the advances of compressive sensing (CS), but efficient coding of sensed measurements is still an issue. In this paper, we propose an image coding system based on the compressive sensing paradigm via stripe-based differential pulse-code modulation (DPCM). In the system, we sample and encode an image in a unit of multiple rows, which we call a stripe. Through extensive experiments, we observe that the correlation between measurements of adjacent stripes are much higher than that of the neighboring blocks. Based on this, we combine the stripe-based CS acquisition with the DPCM framework and design a mechanism that predicts a stripe of measurements from its preceding stripe of measurements. The produced measurement residuals are then quantized and entropy-encoded into binary coding bits, which are tremendously reduced compared to the traditional block-based framework. Furthermore, we provide an image CS reconstruction algorithm corresponding to the stripe-based acquisition. Experiments verify that the reconstruction quality is no worse or even better than the block-based case when much lower bitrate is consumed. In a rate-distortion point of view, the proposed system also outperforms the methods using block-based sampling and achieves the state-of-the-art performance for compressive-sensed image coding.
Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001
DCC3
2014 G-CAST: Gradient Based Image SoftCast for Perception-Friendly Wireless Visual Communication
abstract
Conventional image and video communication systems are usually designed with the objective being to maximize the fidelity of reconstructed images measured by mean square errors (MSE). It is well known that the fidelity metric MSE may not reflect the visual quality perceived by human eyes. Recent advancements in image quality assessment tell us that the structural similarity (SSIM), especially the gradient similarity, reveals the perceptual fidelity of images more reliably. Inspired by this observation, this paper proposes a new image communication approach, which conveys the visual information in an image by transmitting the image gradients and recovers the image from the received gradient data at decoder side using statistical image prior knowledge. In particular, we designed a gradient-based image SoftCast scheme for wireless scenarios. Experimental results show that the proposed scheme can produce reconstruction images with much better perceptual quality. The advantage in perceptual quality is verified by the quality improvement measured by the metrics SSIM and gradient signal-to-noise ratio (GSNR).
Ruiqin Xiong, Hangfan Liu, Siwei Ma 0001, Xiaopeng Fan 0001, Feng Wu 0001, Wen Gao 0001
DCC3
2013 Low Complexity Rate Distortion Optimization for HEVC
abstract
The emerging High Efficiency Video Coding (HEVC) standard has improved the coding efficiency drastically, and can provide equivalent subjective quality with more than 50% bit rate reduction compared to its predecessor H.264/AVC. As expected, the improvement on coding efficiency is obtained at the expense of more intensive computation complexity. In this paper, based on an overall analysis of computation complexity in HEVC encoder, a low complexity rate distortion optimization (RDO) coding scheme is proposed by reducing the number of available candidates for evaluation in terms of the intra prediction mode decision, reference frame selection and CU splitting. With the proposed scheme, the RDO technique of HEVC can be implemented in a low-complexity way for complexity-constrained encoders. Experimental results demonstrate that, compared with the original HEVC reference encoder implementation, the proposed algorithms can achieve about 30% reduced encoding time on average with ignorable coding performance degradation (0.8%).
Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Liang Zhao 0007, Qin Yu 0003, Wen Gao 0001
DCC1
2012 Compressed Sensing Recovery via Collaborative Sparsity
abstract
Compressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression approach. Its theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. So one of the most significant challenges in CS is to seek a domain where a signal can exhibit a high degree of sparsity and hence be recovered faithfully. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for compressed sensing recovery via collaborative sparsity (RCoS), which enforces local two-dimensional sparsity and nonlocal three-dimensional sparsity simultaneously in an adaptive hybrid space-transform domain, thus substantially utilizing intrinsic sparsities of natural images and greatly confining the CS solution space. In addition, an efficient augmented Lagrangian based technique is developed to solve the above optimization problem. Experimental results on a wide range of natural images are presented to demonstrate the efficacy of the new CS recovery strategy.
Jian Zhang 0018, Debin Zhao, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001
DCC5
2011 Transductive Regression with Local and Global Consistency for Image Super-Resolution
abstract
In this paper, we propose a novel image super-resolution algorithm, referred to as interpolation based on transductive regression with local and global consistency (TRLGC). Our algorithm first constructs a set of local interpolation models which can predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Furthermore, a graph-Laplacian based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the accumulated loss of the locally linear regression, square error of prediction bias on the available LR samples and the manifold regularization term, which could be solved with a closed-form solution as a convex optimization problem. In this way, a transductive regression algorithm with local and global consistency is developed. Experimental results on benchmark test images demonstrate that the proposed image super-resolution method achieves very competitive performance with the state-of-the-art algorithms.
Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun
DCC4
2010 Error Resilient Dual Frame Motion Compensation with Uneven Quality Protection
abstract
Summary form only given. In this paper, an error resilient JU-DFMC is proposed for video transmission over error-prone channels. In the proposed error resilient JU-DFMC, a new error resilient prediction structure of DFMC is firstly presented. The LQF can adaptively select reference frame according to different packet loss rate. Then the MB information is divided into two partition header information (A) and texture coefficients (B). Based on the partition, an end-to-end distortion model is applied for macroblock (MB) level mode decision. Finally a frame level rate distortion cost scheme is proposed to determine how many times the header information will be transmitted in a high quality frame (HQF). The HQF (LTR) is given more protection. The experimental results show that the proposed method can achieve better performance than the previous DFMC schemes. In the future, how to determine LQF header transmission times will be further exploited.
Debin Zhao, Siwei Ma 0001
DCC3
2010 Auto Regressive Model and Weighted Least Squares Based Packet Video Error Concealment
abstract
In this paper, auto regressive (AR) model is applied to error concealment for block-based packet video encoding. Each pixel within the corrupted block is restored as the weighted summation of corresponding pixels within the previous frame in a linear regression manner. Two novel algorithms using weighted least squares method are proposed to derive the AR coefficients. First, we present a coefficient derivation algorithm under the spatial continuity constraint, in which the summation of the weighted square errors within the available neighboring blocks is minimized. The confident weight of each sample is inversely proportional to the distance between the sample and the corrupted block. Second, we provide a coefficient derivation algorithm under the temporal continuity constraint, where the summation of the weighted square errors around the target pixel within the previous frame is minimized. The confident weight of each sample is proportional to the similarity of geometric proximity as well as the intensity gray level. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to increase the peak signal-to-noise ratio (PSNR) compared to other methods.
Yongbing Zhang 0002, Xinguang Xiang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001
DCC3
2009 Compression-Induced Rendering Distortion Analysis for Texture/Depth Rate Allocation in 3D Video Compression
abstract
In 3D video applications, the virtual view is generally rendered by the compressed texture and depth. The texture and depth compression with different bit-rate overheads can lead to different virtual view rendering qualities. In this paper, we analyze the compression-induced rendering distortion for the virtual view. Based on the 3D warping principle, we first address how the texture and depth compression affects the virtual view quality, and then derive an upper bound for the compression-induced rendering distortion. The derived distortion bound depends on the compression-induced depth error and texture intensity error. Simulation results demonstrate that the theoretical upper bound is an approximate indication of the rendering quality and can be used to guide sequence-level texture/depth rate allocation for 3D video compression.
Yanwei Liu 0001, Siwei Ma 0001, Qingming Huang, Debin Zhao, Wen Gao 0001, Nan Zhang 0015
DCC2
2007 An Enhanced Robust Entropy Coder for Video Codecs Based on Context-Adaptive Reversible VLC
abstract
This paper proposes an enhanced RVLC coder, context-adaptive reversible variable length coder (CRVLC), for DCT coefficients by using the techniques of data sub-partitioning and context modeling. The data sub-partitioning means that the data part of DCT coefficients is split into several small sub-partitions. As each sub-partition can be reversibly decoded by RVLC, more data as well as higher error resilience can be obtained. The context modeling exploits the correlation of DCT coefficients for further compression. This modeling defines the contexts by hierarchical-dependent information. The information is also available in the backward decoding, so that it supports the reversible decoding. And with it the data outputted by CRVLC can be naturally placed into multiple sub-partitions.
Qiang Wang 0011, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
DCC3