Wenhong Duan

dblp:327/8105 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-6835-3270ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 An Information-Guided Learned Framework for Free-View Image Coding
abstract
In this paper, we propose an information-guided learned image compression framework, which achieves efficient free-view image compression. Specifically, we model the inter-view correlations as inter-view prior based on the proposed feature transform module, which is used to guide the encoding and decoding process of different views. Specifically, we design the feature transform module to generate the inter-view prior. Different from prior information modeled in pixel domain, the inter-view prior in feature domain has higher dimensions and can provide richer and more correlated condition information, which can guide the encoding and decoding process of different view images effectively. Furthermore, we design a multi-prior fusion module, which can fuse the information from different reference views to achieve more accurate guidance. We evaluate the proposed framework by comparing the rate-distortion performance averaged over the BlendedMVS dataset. Extensive experiments demonstrate that the proposed framework can achieve the superior performance in compressing free-view images.
Wenhong Duan, Hui Yuan 0001, Siwei Ma 0001
DCC1
2026 Learned Image Compression via Local-to-Global Cross-Component Prior
abstract
Learned image compression (LIC) methods have shown promising results and achieved superior performance compared to traditional image compression methods. Due to the neglect of the utilization of cross-component correlations, there is still a potential for further performance improvement. In this paper, we first explore the inter-channel correlations of different color spaces and transform the image compression problem in RGB color space into that in YUV color space, which has cross-component prior information. We propose a novel image compression method that leverages local-to-global cross-component prior modeling, utilizing a cross-component attention mechanism to improve coding performance. First, we design the cross-component prior gate (CPG) to model the cross-component prior information based on attention mechanism. Inspired by common knowledge in data compression, luma component (Y) contains more details and textural/structural information compared to chroma components (UV). The proposed method can make full use of the cross-component guidance information from luma to chroma components to achieve effective image compression. Experimental results demonstrate that the proposed method can achieve superior performance compared to existing learned image compression methods. The proposed method can achieve 9.20% rate savings compared to the image compression standard Versatile Video Coding (VVC) Test Model (VTM-11.0) on Kodak dataset.
Wenhong Duan, Jiaye Fu, Li Song 0001, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.1
2026 SPC-NeRF: Spatial Predictive Compression for Voxel-Based Radiance Field
abstract
Representing the Neural Radiance Field (NeRF) with the Explicit Voxel Grid (EVG) is a promising direction for improving NeRFs. However, the EVG representation is not efficient for storage and transmission because of the tremendous memory cost. Existing methods for compressing EVG mainly inherit the methods designed for neural network compression, such as pruning and quantization, which do not take full advantage of the spatial correlation in the voxel grid. Inspired by the prosperous digital image compression techniques, this article proposes SPC-NeRF, a novel framework applying spatial predictive coding in EVG NeRF compression. The proposed framework can remove spatial redundancy efficiently for better compression performance. Our framework contains a progressive coding procedure to realize adaptive quantization precision according to the different importance of the voxels. Moreover, we model the coding bitrate of our framework and design a novel form of the loss function. With the loss function, we can jointly optimize the compression ratio and the rendering distortion to achieve higher coding efficiency. Extensive experiments demonstrate that our method can achieve 32% bit saving compared to the benchmark method VQRF on multiple representative test datasets, with comparable training time.
Zetian Song, Jiaqi Zhang 0007, Wenhong Duan, Yuhuai Zhang, Xinfeng Zhang 0001, Siwei Ma 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Lightweight Learning-Based In-Loop Filter for Real-Time Video Coding
abstract
In recent years, video coding tools based on neural networks have emerged continuously, showing remarkable coding performance. Especially the neural network-based loop filter tools are currently the hottest research direction in both standard development and academic research. Compared to other directions, such as neural network-based intra prediction and inter prediction, neural network-based loop filtering shows significant performance improvements. Moreover, as the final step in the coding loop, it is advantageous for hardware implementation. However, its disadvantages are also evident, as the computational complexity of neural network-based loop filtering tools is extremely high, often requiring hundreds or even thousands of kilo Multiply-Accumulate operations (kMACs) per pixel. Therefore, in this paper, we explore the possibility of applying neural network-based loop filtering in real-time codec. We propose a Lightweight Learning-Based In-Loop Filter (LLILF) which achieves significant improvement in coding performance with minimal complexity. It consists of only two convolutional layers and has 153 parameters. Experimental results show that the proposed method can achieve 1080p@30fps and 720p@60fps encoding on the SVT-AVS3 real-time codec. Additionally, the proposed network achieves an average BD-rate (VMAF) saving of 15.35% over SVT-AVS3 under Random Access (RA) configuration.
Yanchen Zhao, Wenhong Duan, Jiaqi Zhang 0007, Zhimeng Huang, Lin Li 0062, Siwei Ma 0001
ICME2
2024 Extreme Low Bitrate Image Compression System for Mobile Deployment
abstract
End-to-end image compression has achieved satisfactory results in recent studies. However, existing methods suffer from high complexity of complicated neural network computation and cannot be directly deployed on mobile devices due to the limitations of computing ability and storage. Therefore, considering the resource and computing ability constrains of the mobile devices, we make a trade-off in this paper between rate-distortion (R-D) performance, inference time, and model complexity. Then we design a novel lightweight perceptual image compression framework to alleviate the storage and complexity burden of mobile devices. Moreover, we design a hardware-friendly deployment scheme to apply the proposed compression framework on high-end mobile devices, which can achieve efficient image compression. Based on the above structures, we propose the first mobile system that achieves image compression on mobile devices. The supplementary material of our system demo is on https://sigport.org/documents/extreme-low-bitrate-Image-compression-system-mobile-deployment.
Wenhong Duan, Xianping Ma, Jianhui Chang, Shanshe Wang, Siwei Ma 0001, Chuanmin Jia
MMSP2
2024 Advanced Learning-Based Inter Prediction for Future Video Coding
abstract
In the fourth generation Audio Video coding Standard (AVS4), the Inter Prediction Filter (INTERPF) reduces discontinuities between prediction and adjacent reconstructed pixels in inter prediction. The paper proposes a low complexity learning-based inter prediction (LLIP) method to replace the traditional INTERPF. LLIP enhances the filtering process by leveraging a lightweight neural network model, where parameters can be exported for efficient inference. Specifically, we extract pixels and coordinates utilized by the traditional INTERPF to form the training dataset. Subsequently, we export the weights and biases of the trained neural network model and implement the inference process without any third-party dependency, enabling seamless integration into video codec without relying on Libtorch, thus achieving faster inference speed. Ultimately, we replace the traditional handcraft filtering parameters in INTERPF with the learned optimal filtering parameters. This practical solution makes the combination of deep learning encoding tools with traditional video encoding schemes more efficient. Experimental results show that our approach achieves 0.01%, 0.31%, and 0.25% coding gain for the Y, U, and V components under the random access (RA) configuration on average.
Yanchen Zhao, Wenhong Duan, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001
VCIP2
2023 Learned Image Compression Using Cross-Component Attention Mechanism
abstract
Learned image compression methods have achieved satisfactory results in recent years. However, existing methods are typically designed for RGB format, which are not suitable for YUV420 format due to the variance of different formats. In this paper, we propose an information-guided compression framework using cross-component attention mechanism, which can achieve efficient image compression in YUV420 format. Specifically, we design a dual-branch advanced information-preserving module (AIPM) based on the information-guided unit (IGU) and attention mechanism. On the one hand, the dual-branch architecture can prevent changes in original data distribution and avoid information disturbance between different components. The feature attention block (FAB) can preserve the important information. On the other hand, IGU can efficiently utilize the correlations between Y and UV components, which can further preserve the information of UV by the guidance of Y. Furthermore, we design an adaptive cross-channel enhancement module (ACEM) to reconstruct the details by utilizing the relations from different components, which makes use of the reconstructed Y as the textural and structural guidance for UV components. Extensive experiments show that the proposed framework can achieve the state-of-the-art performance in image compression for YUV420 format. More importantly, the proposed framework outperforms Versatile Video Coding (VVC) with 8.37% BD-rate reduction on common test conditions (CTC) sequences on average. In addition, we propose a quantization scheme for context model without model retraining, which can overcome the cross-platform decoding error caused by the floating-point operations in context model and provide a reference approach for the application of neural codec on different platforms.
Wenhong Duan, Zheng Chang 0002, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Li Song 0001, Wen Gao 0001
IEEE Trans. Image Process.1
2023 Differential Weight Quantization for Multi-Model Compression
abstract
Low bit-width quantization can effectively reduce the storage and computational costs of deep neural networks. Existing quantization methods are commonly designed for single model compression. For multi-model compression scenarios, multiple models for the same task or similar tasks need to be compressed simultaneously in multimedia tasks, such as compressing image super-resolution models for different scales and transferring of different models in multimedia. However, single model quantization methods do not consider the correlations among the weights of different models, which limits the further compression for the above multi-model compression scenarios. To sufficiently excavate the potential of compression on multi-model, we propose a novel quantization scheme for multi-model compression, namely differential weight quantization (DWQ), which focuses on the weights increment between the target model and the reference model. Specifically, DWQ is achieved by increment computation, increment quantization and fine-tuning, which utilizes the reference model to guide the subsequent quantization on the target model. Due to the correlations between the weights of different models, the distribution of weights increment is more centralized compared with original weights, which can achieve a higher compression ratio by lower bit representation on weights increment. Moreover, the progressive training method is proposed to accelerate the convergence and reduce quantization loss on the DWQ framework. Extensive experiments validate the effectiveness of DWQ based on weight-sharing and parameterized clipping activation (PACT) quantization technologies on multiple tasks. The proposed framework can achieve 2× compression improvement and reduce 30% computational complexity with comparable performance in the popular multimedia tasks.
Wenhong Duan, Zhenhua Liu 0003, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.1
2022 End-to-End Image Compression via Attention-Guided Information-Preserving Module
abstract
Deep learning-based end-to-end image compression has achieved significant compression performance in recent years. However, current learning-based image compression methods are designed considering the characteristic of RGB color space, which is not suitable for image compression in YUV 420 color space because of the variance between color formats. To achieve efficient image compression in YUV 420 color space, we propose an information-preserving compression framework using the attention mechanism. Specifically, we design an information-preserving module (IPM), where we utilize the dual-branch architecture to prevent changes in data distribution and propose the feature attention block (FAB) to preserve information. Furthermore, a cross-channel progressive enhancement (CPE) network is designed by taking advantage of the relations among different channels. Ex-perimental results show that the proposed framework outper-forms state-of-the-art compression standard Versatile Video Coding (VVC) with 2.52% BD-rate reduction on common test conditions (CTC) sequences on average.
Wenhong Duan, Chuanmin Jia, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001
ICME1