Meng Wang 0017

dblp:93/6765-17 · DBLP profile ↗
← Back
10ranked-venue papers in the field
4as first author
7since 2021 · last 2026
0000-0002-5655-1464ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 10 (4 first)
YearPublicationVenuePosition
2026 An Effective Template-Generated Video Compression Scheme by Exploiting Inter-Video Motion Correlation
abstract
Template-generated videos (TGVs), created by applying animation templates to static images, have become increasingly prevalent, producing massive user-generated content with highly consistent motion patterns. However, existing video compression schemes are designed to eliminate motion redundancy within individual videos, while overlooking the shared motion patterns widespread across TGVs. To address this limitation, we propose a novel compression scheme that effectively leverages inter-video motion priors to enhance the compression efficiency of TGVs. Specifically, the proposed scheme operates as a two-stage pipeline. In the first stage, high-quality motion priors are identified from a representative TGV based on spatial texture and prediction error. In the second stage, these motion priors are intelligently integrated to expand the motion representation space beyond the local candidate lists in Merge and AMVP modes, thereby enabling the codec to remove inter-video redundancy. Experimental results on the versatile video coding test model (VTM-23.0) demonstrate consistent coding gains across various compression scenarios for TGVs, achieving average BD-rate savings of$1.07 \%, 1.38 {\%}$, and 1.18% under low-delay P (LDP), low-delay B (LDB), and random access (RA) configurations, respectively.
Feng Xing, Yingwen Zhang, Meng Wang 0017, Hengyu Man, Shiqi Wang 0001, Xiaopeng Fan 0001
DCC3
2025 An Efficient Hidden Markov Model-Based Sample Adaptive Offset Mode Decision Algorithm for Versatile Video Coding
abstract
This paper proposes a highly efficient sample adaptive offset (SAO) mode decision algorithm. By leveraging both the directional correlations between the SAO and intra-prediction decisions, and the SAO decisions' spatial correlations, the SAO mode candidates are effectively pruned during the rate-distortion optimization process, accelerating the SAO encoding process with negligible BD-rate loss.
Feng Xing, Yingwen Zhang, Meng Wang 0017, Hengyu Man, Yongbing Zhang 0002, Shiqi Wang 0001, Xiaopeng Fan 0001
DCC3
2024 Extreme Image Compression Using Fine-tuned VQGANs
abstract
Recent advances in generative compression methods have demonstrated remarkable progress in enhancing the perceptual quality of compressed data, especially in scenarios with low bitrates. However, their efficacy and applicability to achieve extreme compression ratios (< 0.05 bpp) remain constrained. In this work, we propose a simple yet effective coding framework by introducing vector quantization (VQ)–based generative models into the image compression domain. The main insight is that the codebook learned by the VQGAN model yields a strong expressive capacity, facilitating efficient compression of continuous information in the latent space while maintaining reconstruction quality. Specifically, an image can be represented as VQ-indices by finding the nearest codeword, which can be encoded using lossless compression methods into bitstreams. We propose clustering a pre-trained large-scale codebook into smaller codebooks through the K-means algorithm, yielding variable bitrates and different levels of reconstruction quality within the coding framework. Furthermore, we introduce a transformer to predict lost indices and restore images in unstable environments. Extensive qualitative and quantitative experiments on various benchmark datasets demonstrate that the proposed framework outperforms state-of-the-art codecs in terms of perceptual quality-oriented metrics and human perception at extremely low bitrates (≤ 0.04 bpp). Remarkably, even with the loss of up to 20% of indices, the images can be effectively restored with minimal perceptual loss.
Qi Mao 0002, Tinghan Yang, Meng Wang 0017, Shiqi Wang 0001, Libiao Jin, Siwei Ma 0001
DCC5
2024 Leveraging Conv-Attention for Efficient and High-Quality JPEG AI Image Coding
abstract
In this paper, we present a Conv-Attention, a decoder-friendly attention mechanism, in an effort to advancing the practical application of the artificial intelligence-based image coding. More specifically, the proposed method is tailored for JPEG AI, which is the latest advanced neural-network based image coding standard. By identifying the obstacles by profiling the decoding complexity of JPEG AI, the attention module accounts for a significant proportion, which mainly attributes to the intricate network structure and involvement of less efficient operations. Conv-Attention model is composed with plain convolution and activation computations, equipping with sub-scaling and up-scaling design, such that the non-adjacent features can be well captured, leading to the reduction of decoding complexity and maintenance of the synthesis and attentive capability. Simulation results verify the effectiveness of the proposed method with JPEG AI reference software, wherein the decoding complexity is reduced by 80% with negligible coding performance loss. The proposed method was adopted in the 100th JPEG meeting.
Meng Wang 0017, Semih Esenlik, Zhaobin Zhang, Yaojun Wu 0001, Kai Zhang 0007, Li Zhang 0006, Shiqi Wang 0001
DCC1
2024 Performance Exploration of Jointly Rate-Distortion Optimized HEVC Intra Encoder
abstract
We extend the beam-search based joint rate-distortion optimization (BSJRDO) [1] to High Efficiency Video Coding (HEVC) and investigate the impact of different decisions on it. In BSJRDO, unlike the greedy search, which only maintains one locally optimal path at each decision stage, multiple paths are kept as candidates for future referencing.
Yingwen Zhang, Meng Wang 0017, Shiqi Wang 0001
DCC2
2022 Fast Partition Mode Decision via a Plug-in Fully Connected Network for Video Coding
abstract
Flexible coding unit partitioning such as quad-tree nested binary-tree and ternary-tree adopted by the emerging enhanced compression model (ECM) brings promising coding performance improvement. Meanwhile, the computational complexity increases dramatically, which may block the exploration and validation of new coding tools. This paper investigates a partition mode early pruning scheme via a fully connected network to reduce the encoding complexity for the ECM. In particular, we carefully select features and devise the fully connected network, which could seamlessly cooperate with the encoder, revealing promising learning and inference capability. Experimental results demonstrate that the proposed method achieves 15%~50% encoding time savings with moderate bit-rate increasing on the ECM, and the extra complexity regarding the fully connected network and feature extraction is negligible.
Jiaqi Zhang 0007, Meng Wang 0017, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC2
2021 Super Resolution for Compressed Screen Content Video
abstract
In this paper, we concentrate on the super-resolution (SR) of compressed screen content video, in an effort to address the real-world challenges by considering the underlying characteristics of screen content. Firstly, we propose a new dataset for the SR of screen content video with different distortion levels. Meanwhile, we design an efficient SR structure that could capture the characteristics of compressed screen content video and manipulate the inner-connections in consecutive compressed low-resolution frames, facilitating the high-quality recovery of the high-resolution counter-part. Moreover, we design a new loss function for network training to better remedy the compression distortion and perceptual distortion. Experimental results demonstrate the effectiveness and superiority of the proposed method.
Meng Wang 0017, Jizheng Xu, Li Zhang 0006, Shiqi Wang 0001
DCC1
2020 Sub-Sampled Cross-Component Prediction for Chroma Component Coding
abstract
Cross-component prediction, which takes advantage of inter-channel correlations, predicts the chroma block with the luma reconstructed block according to associated linear model. Instead of involving all available reference samples in building the linear model, in this paper, we propose a sub-sampled approach that utilizes at most four neighboring chroma samples and their corresponding down-sampled luma samples, leading to significantly reduced operations in the derivation of model parameters at both encoder and decoder. The proposed scheme is hardware friendly in terms of the overheads of memory access and clock cycles, and greatly benefits the practical implementations of the emerging video coding standard in real applications. Extensive experiments reveal that the proposed sub-sampled method provides simple operations and robust coding performance, leading to the adoption by Versatile Video Coding (VVC) Standard and the third generation Audio Video Coding Standard (AVS3).
Meng Wang 0017, Li Zhang 0006, Kai Zhang 0007, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC2
2020 Revisiting Local Texture Correlation for Rate-Distortion Optimized Intra Coding
abstract
In this paper, we focus on computationally modeling of the local texture correlations, in an effort to better explore the coding modes with higher priorities in the rate-distortion optimized intra coding. In particular, strong correlations and continuities of local texture with neighboring blocks have been revealed in our analysis, and empirical justifications provide us inspirations on the joint optimization of rate-distortion-complexity when angular modes become finer to adapt the local textures. We examine the philosophy with extensive experiments conducted for refining the intra full-RD list. The results show that better coding performance with on average 0.72% and 3.00% BD-Rate savings for the natural scene and screen content sequences can be achieved in AVS3 test model HPM-5.0 under all intra configuration, with negligible encoding and decoding time variations.
Meng Wang 0017, Li Zhang 0006, Hongbin Liu 0004, Jizheng Xu, Shiqi Wang 0001
DCC1
2019 Extended Quad-Tree Partitioning for Future Video Coding
abstract
The quad-tree plus binary-tree (QTBT) coding unit (CU) partitioning structure, which has been adopted to the next generation video coding standard, shows promising coding performance when compared with the conventional quad-tree structure in HEVC. In this paper, we propose the Extended Quad-tree (EQT) partitioning, which further extends the QTBT scheme and increases the partitioning exibility. More specifcally, EQT splits a parent CU into four sub-CUs of dierent sizes, which can adequately model the local image content that cannot be elaborately characterized with QTBT. Meanwhile, EQT partitioning allows the interleaving with BT partitioning for enhanced adaptability. Experimental results on the JEM7-QTBT-Only platform show that EQT brings better coding performance with 3.17%, 3.20% and 3.06% BD-Rate gains under random access, low-delay P and low-delay B configurations, respectively.
Meng Wang 0017, Li Zhang 0006, Kai Zhang 0007, Hongbin Liu 0004, Shiqi Wang 0001, Sam Kwong, Siwei Ma 0001
DCC1