Peirong Ning

dblp:345/1573 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0006-6645-5130ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Image and video coding · 100%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding › image compression
learned image compression
1.422024
LLIC: Large Receptive Field Transform Coding With Adaptive Weights for Learned Image Compression · IEEE Trans. Multim. 2024
MLIC: Multi-Reference Entropy Model for Learned Image Compression · ACM Multimedia 2023
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
large kernel convolution
0.812024
LLIC: Large Receptive Field Transform Coding With Adaptive Weights for Learned Image Compression · IEEE Trans. Multim. 2024
Image and video coding
transform coding
0.812024
LLIC: Large Receptive Field Transform Coding With Adaptive Weights for Learned Image Compression · IEEE Trans. Multim. 2024
Image and video coding › image compression › learned image compression
entropy model
0.712023
MLIC: Multi-Reference Entropy Model for Learned Image Compression · ACM Multimedia 2023
Image and video coding
entropy coding
0.212024
LLIC: Large Receptive Field Transform Coding With Adaptive Weights for Learned Image Compression · IEEE Trans. Multim. 2024

Methods — techniques the papers use, named apart from their topics

non-local attention · 1.5bit allocation · 1.5adaptive weight generation · 1.5depthwise convolution · 0.8depth-wise convolution · 0.8checkerboard context capturing · 0.7attention mechanism · 0.7
YearPublicationVenuePosition
2024 Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing System
abstract
In streaming media services, video transcoding is a common practice to alleviate bandwidth demands. Unfortunately, traditional methods employing a uniform rate factor (RF) across all videos often result in significant inefficiencies. Content-adaptive encoding (CAE) techniques address this by dynamically adjusting encoding parameters based on video content characteristics. However, existing CAE methods are often tightly coupled with specific encoding strategies, leading to inflexibility. In this paper, we propose a model that predicts both RF-quality and RF-bitrate curves, which can be utilized to derive a comprehensive bitrate-quality curve. This approach facilitates flexible adjustments to the encoding strategy without necessitating model retraining. The model leverages codec features, content features, and anchor features to predict the bitrate-quality curve accurately. Additionally, we introduce an anchor suspension method to enhance prediction accuracy. Experiments confirm that the actual quality metric (VMAF) of the compressed video stays within ±1 of the target, achieving an accuracy of 99.14%. By incorporating our quality improvement strategy with the rate-quality curve prediction model, we conducted online A/B tests, obtaining both +0.107% improvements in video views and video completions and +0.064% app duration time. Our model has been deployed on the Xiaohongshu App.
Shibo Yin, Zhiyu Zhang 0010, Peirong Ning, Qiubo Chen, Guo Lu, Li Song 0001
VCIP3
2024 LLIC: Large Receptive Field Transform Coding With Adaptive Weights for Learned Image Compression
abstract
The effective receptive field (ERF) plays an important role in transform coding, which determines how much redundancy can be removed during transform and how many spatial priors can be utilized to synthesize textures during inverse transform. Existing methods rely on stacks of small kernels, whose ERFs remain insufficiently large, or heavy non-local attention mechanisms, which limit the potential of high-resolution image coding. To tackle this issue, we propose Large Receptive Field Transform Coding with Adaptive Weights for Learned Image Compression (LLIC). Specifically, for thefirsttime in the learned image compression community, we introducea fewlarge kernel-based depth-wise convolutions to reduce more redundancy while maintaining modest complexity. Due to the wide range of image diversity, we further propose a mechanism to augment convolution adaptability through the self-conditioned generation of weights. The large kernels cooperate with non-linear embedding and gate mechanisms for better expressiveness and lighter point-wise interactions. Our investigation extends to refined training methods that unlock the full potential of these large kernels. Moreover, to promote more dynamic inter-channel interactions, we introduce an adaptive channel-wise bit allocation strategy that autonomously generates channel importance factors in a self-conditioned manner. To demonstrate the effectiveness of the proposed transform coding, we align the entropy model to compare with existing transform methods and obtain models LLIC-STF, LLIC-ELIC, and LLIC-TCM. Extensive experiments demonstrate that our proposed LLIC models have significant improvements over the corresponding baselines and reduce the BD-Rate by$9.49\%, 9.47\%,\;\text{and}\; 10.94\%$on Kodak over VTM-17.0 Intra, respectively. Our LLIC models achieve state-of-the-art performances and better trade-offs between performance and complexity.
Wei Jiang 0031, Peirong Ning, Yongqi Zhai, Feng Gao 0014, Ronggang Wang
IEEE Trans. Multim.2
2023 MLIC: Multi-Reference Entropy Model for Learned Image Compression
abstract
Recently, learned image compression has achieved remarkable performance. The entropy model, which estimates the distribution of the latent representation, plays a crucial role in boosting rate-distortion performance. However, most entropy models only capture correlations in one dimension, while the latent representation contains channel-wise, local spatial, and global spatial correlations. To tackle this issue, we propose the Multi-Reference Entropy Model (MEM) and the advanced version, MEM+. These models capture the different types of correlations present in latent representation. Specifically, we first divide the latent representation into slices. When decoding the current slice, we use previously decoded slices as context and employ the attention map of the previously decoded slice to predict global correlations in the current slice. To capture local contexts, we introduce two enhanced checkerboard context capturing techniques that avoids performance degradation. Based on MEM and MEM+, we propose image compression models MLIC and MLIC+. Extensive experimental evaluations demonstrate that our MLIC and MLIC+ models achieve state-of-the-art performance, reducing BD-rate by 8.05% and 11.39% on the Kodak dataset compared to VTM-17.0 when measured in PSNR.
Wei Jiang 0031, Yongqi Zhai, Peirong Ning, Feng Gao 0014, Ronggang Wang
ACM Multimedia4