Chuanmin Jia

dblp:177/8156 · DBLP profile ↗
← Back
24ranked-venue papers in the field
2as first author
23since 2021 · last 2026
0000-0002-7418-6245ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 24 (2 first)
YearPublicationVenuePosition
2026 LSE-Codec: An Arbitrary Frame-Rate Compliant Lossless Compression Model for Neuromorphic Spike Camera
abstract
Spike cameras represent a novel class of neuromorphic imaging devices that capture visual scenes as binary spike streams through temporal integration of light intensity. While offering exceptional dynamic range and energy efficiency, they generate massive binary data volumes that require efficient compression. Existing codecs fail to exploit the unique structure of spike data and lack support for arbitrary temporal resolution. We propose LSE-Codec, a neural lossless compression framework specifically designed for spike cameras that operates independently of frame-rate, as shown in Fig. 1. Our approach introduces Local Spike Embedding (LSE) to reorganize sparse binary spike patterns into compact 8 -bit symbols while preserving spatial structure, followed by a hierarchical autoencoder with autoregressive entropy modeling to predict spike distributions. Evaluated on different datasets under aligned test conditions, our method achieves state-of-the-art performance, reducing bit rates by over 11% compared to JPEG-XL (best-effort mode). This work bridges a critical gap in spike vision systems with frame-rate agnostic lossless coding, enabling efficient storage and transmission of spike data without sacrificing fidelity.
Fanke Dong, Yiyang Zhou, Chuanmin Jia
DCC3
2026 Prompt-Optimization with Contextual Mining for Cross-Modal Image Compression
abstract
Recent advances in cross-modal compression(CMC) have opened new horizons for perceptual image coding at ultra-low bitrates (below 0.1 bpp) within a generative compression paradigm, but reconstruction fidelity is often compromised, yielding visually plausible yet semantically inconsistent reconstructions. While prompt engineering with contextual optimization has been extensively explored in generative models, its potential for controlling perception-fidelity trade-offs in image compression remains largely under-explored. To address these challenges, we propose PO-CMC, a novel diffusion-based cross-modal image compression approach that introduces contextual prompt optimization to achieve efficient and perceptually faithful reconstruction. The proposed method comprises three synergistic components: an optimized image codec that produces a compact structural prior, a contextual prompt module that adaptively encodes semantic cues into compact textual embeddings, and a diffusion-based decoder that fuses the structural and semantic priors to reconstruct high-fidelity images. Extensive experiments show that PO-CMC achieves superior perceptual quality while maintaining comparable reconstruction fidelity, yielding an average BD-rate saving of 72.5 % and 79.8 % over VVC at equivalent LPIPS and DISTS levels, respectively.
Shenpeng Song, Zhimeng Huang, Junlong Gao, Chuanmin Jia, Siwei Ma 0001
DCC4
2025 STACO: Spatio-Temporal Adaptive Context Optimization for Neural Video Compression
abstract
This paper introduces the Spatio-Temporal Adaptive Context Optimization (STACO) method, which enhances the quality of contextual prediction across various resolutions, essential for subsequent compression. The STACO takes predicted contexts$C_t^{\{1,2,3\}}$as input and improves their quality by aligning them better with decoded features$f_{t}$, thus boosting coding efficiency. The STACO comprises Quality Perception Units and Consistency Synergy Modules, arranged in a hierarchical stacked architecture. This multi-scale design enables simultaneous processing of contexts at different spatial resolutions and facilitates information exchange through upsampling and downsampling. Enhanced contexts$\tilde{C}_{t}^{\{1,2,3\}}$are output after passing through residual connections, ensuring better alignment with reconstructed features. Using VTM-11.0 as anchor, the STACO significantly improves compression efficiency on common test condition (CTC) in HEVC, achieving an average BD-rate reduction of 17.99% for PSNR and 43.84% for MS-SSIM. By incorporating spatial quality mapping and temporal propagation, STACO offers a significant advancement in video compression.
Kexiang Feng, Shuhong Liao, Zhimeng Huang, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC4
2025 Rethinking Bjøntegaard Delta for Compression Efficiency Evaluation: Are we Calculating it Precisely and Reliably?
abstract
For decades, the Bjøntegaard Delta (BD) has been the metric for evaluating codec Rate-Distortion (R-D) performance. Yet, in most studies, BD is determined using just 4–5 R-D data points, could this be sufficient? As codecs and quality metrics advance, does the conventional BD estimation still hold up? Crucially, are the performance improvements of new codecs and tools genuine, or merely artifacts of estimation flaws? We address these concerns by reevaluating BD estimation. We have established a large-scale, high-precision R-D dataset to verify the accuracy of existing BD estimation algorithms. Moreover, we propose a robust method for high-precision BD estimation across diverse compression scenarios, enhanced by a reliability assessment to determine the probability distribution of BD values from R-D sample points. This approach both assesses the reliability of BD calculations and serves as a precise BD estimator. Our method's validity is confirmed through extensive testing on a dataset we constructed. Our findings advocate for the adoption of rigorous R-D sampling and reliability metrics in future compression research to ensure the validity and reliability of results. Our code and additional experimental details are publicly accessible at https://github.com/fgvfgfg564/BDCI.
Xinyu Hang, Shenpeng Song, Zhimeng Huang, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC4
2025 Image Coding for Machine with Visual-Language Mimic Feature Learning
abstract
This paper propose a Image Coding for Machine (ICM) framework with Visual-Language Mimic Feature Learning (VLM-ICM). VLM-ICM decouples the position and semantic information into language modality and extracts universal features from the input image. Language, inherently more semantically compact, helps reduce the bitrate. Meanwhile, the universal features in VLM-ICM, guided by the language at the decoder side, allow for flexible domain adaptation, thereby enhancing versatility and practicality.
Zhimeng Huang, Junlong Gao, Jiaqi Zhang 0007, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia
DCC7
2025 Dynamic Temporal Reference Aggregation for Neural Video Compression
abstract
Neural Video Compression (NVC) has advanced significantly in recent years, with improvements in inter prediction techniques. In inter prediction, most NVC approaches utilize pixel information or temporal features from neighboring frames as reference information, while using optical flow to represent motion information. In this paper, we introduce an innovative and efficient method for Dynamic Temporal Reference Aggregation (DTRA). The proposed DTRA consists of two components: Temporal Information Compensation (TIC) and Feature Level Motion Information Enhancement (MIE). The TIC module generates compensation information by leveraging long-term temporal information from the decoding buffer, enriching the semantic content of the reference features and enhancing their texture details. The MIE module refines the motion features at the encoder side and divides the motion information into multiple groups for diverse motion alignment at the decoder side, thereby improving the motion compensation. Extensive experiments demonstrate the effectiveness of the proposed method, achieving an average bitrate savings of 9.67% compared to state-of-the-art (SOTA) approaches.
Shuhong Liao, Kexiang Feng, Zhimeng Huang, Siwei Ma 0001, Chuanmin Jia
DCC7
2025 FAPC: Frequency-Based Adaptive Pixel Correction for Compressed Screen Content
abstract
Screen content is an important category of video. The statistical distribution of pixels in screen content exhibits substantial differences compared to camera-captured video, leading to different compression needs and challenges. Hitherto, most Screen Content Coding (SCC) tools are designed for block-based hybrid coding frameworks, which may not be suitable for emerging wavelet-based and learning-based coding frameworks. Consequently, a plug-and-play SCC tool independent of coding frameworks is lacking in the current video coding landscape. In this paper, an out-loop coding method, Frequency-based Adaptive Pixel Correction (FAPC), is proposed to improve the SCC performance for arbitrary codecs. First, a True Color Value (TCV) table is established based on the most frequently occurring pixel values. Then, the reconstructed pixels are corrected according to the TCV table. To realize precise pixel correction, an adaptive threshold derivation method is meticulously designed to control the pixel correction process. Furthermore, an inheritance coding strategy is proposed to reduce the overhead of parameter transmission. The proposed method has been integrated into three different coding frameworks. Simulation results demonstrate that the proposed method can achieve 10.66%, 3.55% and 10.67% luma component BD-BR gains on the three coding frameworks, respectively. These results prove the superiority and universality of the proposed method.
Zetian Song, Jiaqi Zhang 0007, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC3
2025 LL-ICM: Image Compression for Low-Level Machine Vision via Large Vision-Language Model
abstract
Image Compression for Machines (ICM) aims to compress images for machine vision tasks, while current methods mostly focus on the demands for high-level tasks. However, the quality of original images is usually not guaranteed in the real world, leading to even worse downstream task performance after compression. Thus, lowlevel (LL) restoration tasks should also be considered in ICM. In this paper, we propose the first ICM framework for LL machine vision tasks, namely LL-ICM, which optimizes the compression and LL processing performance simultaneously. Moreover, LL-ICM leverages large vision-language model (VLM) to solve different LL task within a single model, which is particularly useful when the distortion type of the original image is uncertain. As illustrated in Fig. 1(a), LL-ICM consists of a neural image codec and a VLM-based LL processing module. Given an original image with distortions, LL-ICM firstly compress it as$\hat{\mathbf{X}}$. Then, we extract a generalized feature F from$\hat{\mathbf{X}}$, which is then encoded as two representations, distortion type$\varphi$and caption$\sigma$. After that, the LL processing module receives$\hat{\mathbf{X}}$and its representations to generate the restored version of$\hat{\mathbf{X}}$, i.e.,$\hat{\mathbf{X}}_{\mathbf{H}}$.
Qi Zhang 0042, Chuanmin Jia, Shiqi Wang 0001
DCC3
2025 A Hardware-Friendly AVS3 Entropy Decoder Architecture for 8K Ultra-High-Definition Video
abstract
Logarithmic binary arithmetic coding (LBAC) is a entropy coding method first introduced in AVS2 and now used in the third generation of audio video coding standard (AVS3). While LBAC provides high coding efficiency, the data dependencies in AVS3 make it challenging to parallelize parsing syntax element. To enhance the throughput of AVS3 LBAC entropy decoding engine, a sophisticated hardware implementation architecture is meticulously designed in the paper. The proposed method decouples the bitstream decoding process into four distinct modules. To ensure the decoding efficiency and performance, an advanced parallel pipeline is carefully designed. Furthermore, we develop sub-branch state prediction mechanism and grouped decoding method to further improve the decoding efficiency. The proposed AVS3 entropy decoder has been implemented on the S10 FPGA System Board, and experimental results show that the proposed method could achieve real-time and seamless decoding efficiency for 8K AVS3 bitstreams, even at bitrates as high as 207.1Mbps.
Jiaqi Zhang 0007, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001
DCC3
2024 Neural Compression for AI Foundation Model Generated Images: Evaluation and Benchmark
abstract
We introduce a novel and challenging task within the AIGC era: coding for AIGI (AI-generated images). Specifically, we propose the first AIGI dataset called PKU-AIGI-500K, which is meticulously constructed based on five major foundation models with diverse prompts. Furthermore, We conduct extensive and systematic analysis of the essential characteristics of AIGC images. We thoroughly benchmark the rate-distortion performance and runtime complexity analysis of conventional and learned image coding solutions that are openly available, revealing new insights for emerging studies in AIGI compression. The main contributions of this paper can be summarized as follows: (i) we contribute and build the first AIGI dataset, PKU-AIGI-500K, containing 105+ prompts and 528k+ images based on five generative models. Additionally, we analyze the image features that may be beneficial for image compression, such as quality, texture, color, etc., providing novel insights into available solutions and paving the way for future research. (ii) Building upon the proposed PKU-AIGI-500K, we evaluate the compression efficiency using some popular traditional codecs and learning-based codecs to form a strong benchmark. We also observe that the learning-based models trained on natural images cannot achieve competitive performance on the AIGIs without fine-tuning, emphasizing the necessity of a sufficiently large dataset to advance the research of AIGI coding. (iii) The evaluation and benchmarking are accomplished as an AIGIs’ compression project.
Xunxu Duan, Hongbin Liu 0004, Li Zhang 0006, Chuanmin Jia
DCC4
2024 A Neural-network Enhanced Video Coding Framework beyond ECM
abstract
In this paper, a hybrid video compression framework is proposed that serves as a demonstrative showcase of deep learning-based approaches extending beyond the confines of traditional coding methodologies. The proposed hybrid framework is founded upon the Enhanced Compression Model (ECM), which is a further enhancement of the Versatile Video Coding (VVC) standard. We have augmented the latest ECM reference software with well-designed coding techniques, including block partitioning, deep learning-based loop filter, and the activation of block importance mapping (BIM) which was integrated but previously inactive within ECM, further enhancing coding performance. We evaluate the coding performance of the proposed framework with extensive experiments on the JVET dataset compared with ECM10.0 and VTM-11.0. Due to the testing environment and the coding complexity of the ECM, we did not conduct testing on Class A. The QPs are set as 22, 27, 32, 37, and 42. Compared with ECM-10.0, our method achieves 6.26%, 13.33%, and 12.33% BD-rate savings for the Y, U, and V components under random access (RA) configuration. The traditional hybrid coding framework combined with the three coding tools can further improve compression efficiency and has great potential for performance improvement.
Yanchen Zhao, Chuanmin Jia, Qizhe Wang, Yue Li 0015, Chaoyi Lin, Kai Zhang 0007, Li Zhang 0006, Siwei Ma 0001
DCC3
2024 A Dynamic Point Cloud Dataset for MPEG Point Cloud Compression and Performance Analysis
abstract
Recent years witnessed the development in MPEG point cloud compression (PCC). However, the exploration of inter-frame coding may be impeded due to the lack of dynamic point clouds (point cloud sequences). To promote the development of PCC technology, we propose Dynamic3D , a dynamic 3D point cloud dataset with high-quality real-captured 3D persons and objects. There are several appealing properties: 1) Dynamic scenes: It contains five sequences and each sequence comprises 600 frames with temporal variation; 2) Complex content: instead of a single person or object in the existing dataset from MPEG, our established dataset contains multiple persons or both person and objects; 3) Realistic capture: the color industrial cameras and infrared cameras are used for data acquisition. This dataset provides the vast exploration space for PCC, especially the elimination of temporal redundancy. Extensive simulations are conducted on this dataset by using the reference software of MPEG G-PCC and V-PCC, i.e., (GeS-TM and TMC2), delivering observations, analysis and opportunities for the future research of PCC.
Lili Zhao 0001, Qian Yin 0002, Lancao Ren, Lei Yang 0063, Chuanmin Jia, Siwei Ma 0001
DCC5
2023 Rate-Distortion Optimization for Cross Modal Compression
abstract
Recently, cross modal compression (CMC) is proposed to compress highly redundant visual data into a compact, common, human-comprehensible domain (such as text) to preserve semantic fidelity for semantic-related applications. However, CMC only achieves a certain level of semantic fidelity at a constant rate, and the model aims to optimize the probability of the ground truth text but not directly semantic fidelity. To tackle the problems, we propose a novel scheme named rate-distortion optimized CMC (RDO-CMC). Specifically, we model the text generation process as a Markov decision process and propose rate-distortion reward which is used in reinforcement learning to optimize text generation. In rate-distortion reward, the distortion measures both the semantic fidelity and naturalness of the encoded text. The rate for the text is estimated by the sum of the amount of information of all the tokens in the text since the amount of information of each token is a lower bound of coding bits. Experimentally, RDO-CMC effectively controls the rate in the CMC framework and achieves competitive performance on MSCOCO dataset.
Junlong Gao, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC2
2023 Learning to Compress Unmanned Aerial Vehicle (UAV) Captured Video: Benchmark and Analysis
abstract
In this paper, we propose to build a novel benchmark and neural video coding task named learning based Unmanned Aerial Vehicle (UAV) video coding. We collect the UAV videos with different content variations, including in-door and out-door scenes, object-scale variations and viewpoint distance, different climate condition etc. Then we encode those properly-selected videos using popular end-to-end optimized video codecs and conventional hybrid codecs, to form a comprehensive benchmark for learned drone video compression. We also provide a detailed analysis and envision the challenge of such task for future research. The main contributions of this paper are three folds. First, we construct a comprehensive benchmark for the task of drone video compression which consists of the rate-distortion (R-D) behavior of both hybrid and learned video codecs. To our knowledge, it is the first attempt in end-to-end optimized solution to compress drone videos. Second, we provide the review and analysis of the learned drone video compression schemes and further discuss the challenges of encoding UAV videos. Third, this benchmark and related research is accomplished as a milestone MPAI End-to-end Video (EEV) coding project. The proposed benchmark has constructed a solid baseline for compressing UAV videos and facilitates the future research works for related task.
Chuanmin Jia, Huifang Sun, Siwei Ma 0001, Wen Gao 0001
DCC1
2022 Rate Distortion Characteristic Modeling for Neural Image Compression
abstract
End-to-end optimized neural image compression (NIC) has obtained superior lossy compression performance recently. In this paper, we consider the problem of rate-distortion (R-D) characteristic analysis and modeling for NIC. We make efforts to formulate the essential mathematical functions to describe the R-D behavior of NIC using deep networks. Thus arbitrary bit-rate points could be elegantly realized by leveraging such model via a single trained network. We propose a plugin-in module to learn the relationship between the target bit-rate and the binary representation for the latent variable of auto-encoder. The proposed scheme resolves the problem of training distinct models to reach different points in the R-D space. Furthermore, we model the rate and distortion characteristic of NIC as a function of the coding parameter$\lambda$respectively. Our experiments show our proposed method is easy to adopt and realizes state-of-the-art continuous bit-rate coding performance, which implies that our approach would benefit the practical deployment of NIC.
Chuanmin Jia, Ziqing Ge, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC1
2022 Coarse-to-fine Prediction With Local and Nonlocal Correlations for Intra Coding
abstract
Recently many efforts have been devoted to learning non-linear predictions from neighboring samples with deep neural networks. However, existing methods mainly generate predictions with local reference samples, regardless of nonlocal self-similarity. In this paper, we aim to incorporate local and nonlocal correlations for intra prediction and propose a two-stage coarse-to-fine network (CTFN), which is integrated into VVC codec as an optional intra prediction mode. The prediction process of CTFN is decomposed into two stages. In the first stage, we train a set of networks to generate a coarse result with local reference samples. In the second stage, we extract sufficient features from nonlocal region using the coarse result as priors and transform the features into a fine prediction result. In particular, a patch-wise attention layer (PAL) is designed in the second stage that can fully explore nonlocal correlations in feature domain and assign weights to each nonlocal feature adaptively, as shown in Fig. 1. As such, the proposed CTFN can not only learn a non-linear mapping from local context, but also explicitly borrow similar features from nonlocal region in a weighted form. Different from image inpainting tasks, the patch synthesis problem is converted to patch matching problem with the CTFN, yielding more reliable predictions. More-over, we construct a classified dataset based on Pearson Correlation Coefficient for network training to better handle contents that are highly correlated. Experiments on VTM-11.0 show that the proposed network achieves 1.77% ED-rate reductions under all intra configuration, which outperforms the state-of-the-art methods.
Meng Lei, Xuewei Meng, Chuanmin Jia, Shanshe Wang, Zhipeng Cheng, Siwei Ma 0001
DCC3
2022 High-Order Intra Prediction for Future Video Coding
abstract
Intra prediction acts a significant role in removing the spatial redundancy in the hybrid coding framework. Versatile Video Coding (VVC) employs a set of angular intra modes to generate directional contents based on the linear projection hypothesis. However, the linear based predictor is not expressive enough for generating high-fidelity patterns with non-linear structure. To compensate for that, we propose a high-order intra prediction (HOIP) for future video coding in this paper. In particular, the HOIP is modeled by a quadratic extrapolation function. To be compatible with the present intra prediction mechanism, the quadratic function can be formulated by two angular intra modes. To reduce the encoding complexity, we further propose a search pruning strategy to find the most appropriate pair-wise modes, and it can be flexibly extended for higher coding performance. The extensive experimental results demonstrate the effectiveness of the proposed method. Up to 0.6% BD-rate saving is obtained with the moderate complexity increment.
Jiaqi Zhang 0007, Chuanmin Jia, Wen Gao 0001
DCC4
2022 Parametric Non-local In-loop Filter for Future Video Coding
abstract
In-loop filter has been comprehensively explored during the development of video coding standards to suppress compression artifacts. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take advantage of the image local similarity. Although some non-local based in-loop filters can make up for this short-coming, the unsupervised parameter selection scheme, which is widely used by non-local filters, limits the content adaptability. Given this, we propose a parametric non-local in-loop filter (PNLF) that fully considers the non-local characteristics and trains the filter coefficients based on the video content. In the filtering process, the reference samples based on the non-local similarity are first derived for each to-be-filtered sample. Then to-be-filtered samples are grouped into specific classes based on multiple features. For each class, filter coefficients are online trained in the encoder and transmitted to the decoder. Finally, the filtering process is conducted using the online-selected coefficients. Simulation results reveal that the proposed approach achieves 0.70%, 1.43%, and 2.09% bit-rate savings on average compared to VTM-11.0 under All Intra (AI), Random Access (RA), and Low-Delay B (LDB) configurations, respectively. The sequences used in the experiment include Class AI, A2, B, C, D, E, F, and SCC. Compared to the non-local structure-based filter (NLSF) [1], our proposed PNLF with fast block matching scheme [2] applied on B-frames and P-frames can achieve better performance gain with lower software and hardware complexity under RA and LDB configurations.
Xuewei Meng, Chuanmin Jia, Xinfeng Zhang 0001, Meng Lei, Shanshe Wang, Lin Li 0062, Siwei Ma 0001
DCC2
2022 Fast Partition Mode Decision via a Plug-in Fully Connected Network for Video Coding
abstract
Flexible coding unit partitioning such as quad-tree nested binary-tree and ternary-tree adopted by the emerging enhanced compression model (ECM) brings promising coding performance improvement. Meanwhile, the computational complexity increases dramatically, which may block the exploration and validation of new coding tools. This paper investigates a partition mode early pruning scheme via a fully connected network to reduce the encoding complexity for the ECM. In particular, we carefully select features and devise the fully connected network, which could seamlessly cooperate with the encoder, revealing promising learning and inference capability. Experimental results demonstrate that the proposed method achieves 15%~50% encoding time savings with moderate bit-rate increasing on the ECM, and the extra complexity regarding the fully connected network and feature extraction is negligible.
Jiaqi Zhang 0007, Meng Wang 0017, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC3
2022 Analysis on Compressed Domain: A Multi-Task Learning Approach
abstract
Image compression approaches based on deep learning have achieved remarkable success. Existing studies mainly focus on human vision and machine analysis tasks taking reconstructed images as input. However, those methods need images to be decoded before performing downstream visual tasks, which motivates us to explore how to directly conduct visual analysis using the compressed data without decoding. The overview of our proposed model is shown as Fig. 1(a). Specifically, a task-agnostic learning-based compression model is proposed, which effectively supports various compressed domain-based analytical tasks meanwhile reserves outstanding re-constructed perceptual quality compared with traditional and learning-based codecs. To obtain the extremely compacted data representation with essential semantic infor-mation, we take the help of the generative model on decoder part. Then, we propose a multi-task learning model which can directly obtain semantic information from the compressed visual data. The pipeline of the proposed model is detailedly illus-trated in Fig. 1(b). In addition, joint optimization strategy is adopted to achieve the best balance point among compression efficiency, reconstructed image quality, and the downstream visual tasks' performance. Experimental results verify that our proposed compressed domain-based multi-task analysis model outperforms the reconstructed image-based method on transmission efficiency, saving more than ten times of bit-rate consumption while preserving comparable visual analysis precision (i.e., classification and segmentation tasks) when compared with RGB image input models, which is evaluated on the CelebA-HO dataset.
Yuefeng Zhang, Chuanmin Jia, Jianhui Chang, Siwei Ma 0001
DCC2
2022 Interpretable Learned Image Compression: A Frequency Transform Decomposition Perspective
abstract
Image compression is a key problem in this age of information explosion. With the help of machine learning, recent studies have shown that learning-based image compression methods tend to surpass traditional codecs. Image compression can be split into three steps: transform, quantization, and entropy estimation. However, the transform step in traditional codecs lacks flexibility because of the strict mathematical premise while the transform in most learning-based codecs neglects its intrinsic interpretation. After observing compression degradation degree varies on different frequency bands as illustrated as Fig. 1(a), we propose an end-to-end compression model from the frequency perspective with a frequency-pyramid transform and a frequency-aware fusion module. The right of the Fig. 1(a) displays each frequency layer's component of the proposed model from low to high-frequency splits. Intuitively, we can infer that the low-frequency part contains the global structure while the high-frequency part gets finer details, satisfying the feature of human visual system (HVS). The proposed model are detailedly shown in Fig. 1(b) that independent probability estimation models are set for each frequency split. Extensive experiments are conducted to demonstrate that our model outperforms all traditional codecs (e.g., JPEG, JPEG2000, HEVC, and VVC) on MS-SSIM metric on both Kodak and CLIC2020 professional test datasets. Taking BPG-4:4:4 as the anchor, our proposed model achieves 11.6% BD-rate reduction under PSNR measurement, which is evaluated on the Kodak dataset.
Yuefeng Zhang, Chuanmin Jia, Siwei Ma 0001
DCC3
2021 Quad-Treea Based Sample Refinement Filter for Video Coding
abstract
In-loop filter is a crucial module in video coding, which can improve both subjective and object quality of reconstructed videos. In this paper, a new sample-based classification method is first proposed using features extracted from different stages of the existing in-loop filter process. Based on this method, an adaptive three-layer Quad-tree Based Sample Refinement Filter (QSRF) algorithm is designed to further improve the coding efficiency. Experimental results show that the proposed QSRF algorithm achieves 0.39%, 0.77% and 0.70% BD-rate savings for random access, lowdelay B and lowdelay P configurations compared to AVS3 reference software, respectively. Moreover, the proposed method can also improve visual quality of reconstructed videos significantly.
Yunrui Jian, Jiaqi Zhang 0007, Chuanmin Jia, Suhong Wang, Shanshe Wang, Siwei Ma 0001
DCC3
2021 Optimized Adaptive Loop Filter in Versatile Video Coding
abstract
In the Versatile Video Coding (VVC) standard, adaptive loop filter (ALF), including Geometry transformation-based Adaptive Loop Filter (GALF) and Cross Component Adaptive Loop Filter (CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is unfriendly to the software and hardware design. Therefore, we propose an optimized ALF framework, including the parallel design of GALF and CCALF, the adaptive parameter decision of GALF, and one-pass CCALF scheme by effectively estimating the CCALF filtering distortion without conducting filter operation. Compared to VTM-8.0, the proposed method can reduce the picture buffer access from 152 to 1 and achieve roughly 25% time-savings of the ALF module with negligible coding performance change under RA configuration. Some of the proposed methods have been adopted in the VVC reference software.
Xuewei Meng, Jiaqi Zhang 0007, Chuanmin Jia, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC3
2016 Structure-driven Adaptive Non-local Filter for High Efficiency Video Coding (HEVC)
abstract
Deblocking filter (DF) Is High Efficiency Video Coding (HEVC) is Only Applied to all Samples Adjacent to prediction units (PU), or transform units (TU), which actually exists two issues. The first one is that DF in HEVC does not fully exploit nonlocal similarity structure information in video. The second one is that DF is HEVC does not consider the inside pixels, which often suffer from quantization distrotion. To alleviate these issues, in this paper, a structure-driven adaptive non-local filter (SANF) Is Proposed By Simultaneously Enforcing The Intrinsic Local Sparsity And The Non-Local Self-Similarity Of Each Frame. Not only SANF deals with the boundary pixels, but also the inside area, which is able to effectively reduce block artifacts while enhancing the quality of the deblocked frames. Applying SANF to luma and chroma components after DF, simulation results demonstrate that the proposed SANF can save BD-rate reduction up to 10.3% with ALF off. For luma component, SANF achieves 4.1%. 3.3%, 4.4% BD-rate saving for all intra, low delay B and random access configurations, respectively with ALF off. furthermore, the performance with ALF on is also discussed.
Jian Zhang 0018, Chuanmin Jia, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001
DCC2