Xuewei Meng

dblp:226/0982 · DBLP profile ↗
← Back
6ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0000-0001-8172-848XORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6 (2 first)
YearPublicationVenuePosition
2026 Virtual Reference Frame Synthesis for Video Coding via Local-Global Spatiotemporal Context Modeling
abstract
Inter prediction is a fundamental component of modern video coding, where the quality of reference frames critically affects motion compensation accuracy and overall compression efficiency. However, relying solely on reconstructed low-temporal-layer frames imposes significant limitations, as these frames often suffer from compression artifacts that degrade prediction quality. To overcome this limitation, we propose a Local-Global spatiotemporal context modeling-based virtual reference frame generation network (LGCM-Net) that synthesizes high-quality reference frames based on reconstructed frames, as shown in Fig. 1. The proposed network integrates hierarchical feature extraction with long-range dependency modeling, where QP-conditioned modulation is applied to shallow features to adapt them to quantization-induced quality variations, enabling temporally and structurally consistent reference generation closely aligned with the to-be-coded frame. Moreover, a coarse-to-fine multi-stage optical flow refinement mechanism is employed to progressively enhance motion accuracy, and a residual refiner further compensates remaining motion estimation errors and reconstruction artifacts to deliver a more accurate final prediction. The proposed method achieves$5.37 \%, 9.96 \%$, and 9.91% BD-rate reduction under the Random Access configuration in VVC reference Software (VTM-11.0_nnvc-10.0 w/o NN Coding tools) for the$\mathrm{Y}, \mathrm{U}$, and V components, respectively.
Yanchen Zhao, Xuewei Meng, Jiaqi Zhang 0007, Kai Zhang 0007, Siwei Ma 0001
DCC3
2026 Lightweight CNN-Based In-Loop Filtering for Video Coding with Hardware-Aware Optimizations
abstract
Neural network-based in-loop filtering significantly enhances video compression efficiency. However, high computational complexity hinders their deployment in real-time and ultra-high-definition scenarios. To address this, we propose a lightweight CNN-based in-loop filter for the luma component. In terms of model design, we utilize a U-Net-like architecture to learn the residual signal, incorporating depthwise separable 3 × 3 convolutions and 1 × 1 convolutions to reduce computational complexity, which results in a low complexity of only$37.707 \text{kMACs} /$pixel. For deployment optimization, we implement memory linearization to improve cache efficiency and combine blocked matrix multiplication with SIMD to maximize parallelism, ensuring cross-platform compatibility without third-party libraries. Experimental results on AVS4 EVM-0.9 (All-Intra) on a CPU platform show that the proposed method achieves BD-rate reductions of$1.36 \%, 0.35 \%$, and 0.34% for$\mathrm{Y}, \mathrm{U}$, and V components, respectively. Furthermore, the optimizations lead to a 91.5% reduction in decoding time, resulting in a decoding complexity of 7757% compared to the anchor.
Yanchen Zhao, Xuewei Meng, Jiaqi Zhang 0007, Haocheng Tang, Lin Li 0062, Siwei Ma 0001
DCC3
2025 Recurrent Intra Prediction Mode for Future Video Coding
abstract
Intra prediction is a crucial component of hybrid video coding framework due to its remarkable ability to reduce spatial redundancy in video signals. Unlike the single-mode based intra prediction in HEVC and VVC, intra fusion prediction methods, that combine the results of multiple angular prediction modes, were newly adopted by Enhanced Compression Model (ECM). However, intra fusion prediction over-relies on local spatial correlations and neglects potential texture similarities in non-adjacent regions. To overcome these limitations and elevate the accuracy of luma intra prediction, a Recurrent Intra Prediction Mode (RIPM) is proposed in this paper, which is composed of two sub-modules, i.e., Recurrent Intra Merge Mode (RIMM) and Recurrent Block Vector Substitution Module (RBVSM). RIMM utilizes the recurrent spatial texture information of the adjacent and non-adjacent spaces for adaptive mode derivation and prediction within the intra fusion prediction framework. RBVSM is a sophisticated mechanism for adaptive prediction mode selection and weight assignment during intra fusion prediction, resulting in enhanced coding performance with minimal impact on computational complexity. The proposed method, implemented on top of ECM-12.0, demonstrates a 0.095% BD-rate gain for the luma component under All Intra configuration, with negligible complexity increase. Currently, RIPM is under study in Exploration Experiments (EE) for ECM in JVET.
Jiaye Fu, Xuewei Meng, Siwei Ma 0001, Jiaqi Zhang 0007, Yao-Jen Chang, Vadim Seregin, Marta Karczewicz
DCC2
2022 Coarse-to-fine Prediction With Local and Nonlocal Correlations for Intra Coding
abstract
Recently many efforts have been devoted to learning non-linear predictions from neighboring samples with deep neural networks. However, existing methods mainly generate predictions with local reference samples, regardless of nonlocal self-similarity. In this paper, we aim to incorporate local and nonlocal correlations for intra prediction and propose a two-stage coarse-to-fine network (CTFN), which is integrated into VVC codec as an optional intra prediction mode. The prediction process of CTFN is decomposed into two stages. In the first stage, we train a set of networks to generate a coarse result with local reference samples. In the second stage, we extract sufficient features from nonlocal region using the coarse result as priors and transform the features into a fine prediction result. In particular, a patch-wise attention layer (PAL) is designed in the second stage that can fully explore nonlocal correlations in feature domain and assign weights to each nonlocal feature adaptively, as shown in Fig. 1. As such, the proposed CTFN can not only learn a non-linear mapping from local context, but also explicitly borrow similar features from nonlocal region in a weighted form. Different from image inpainting tasks, the patch synthesis problem is converted to patch matching problem with the CTFN, yielding more reliable predictions. More-over, we construct a classified dataset based on Pearson Correlation Coefficient for network training to better handle contents that are highly correlated. Experiments on VTM-11.0 show that the proposed network achieves 1.77% ED-rate reductions under all intra configuration, which outperforms the state-of-the-art methods.
Meng Lei, Xuewei Meng, Chuanmin Jia, Shanshe Wang, Zhipeng Cheng, Siwei Ma 0001
DCC2
2022 Parametric Non-local In-loop Filter for Future Video Coding
abstract
In-loop filter has been comprehensively explored during the development of video coding standards to suppress compression artifacts. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take advantage of the image local similarity. Although some non-local based in-loop filters can make up for this short-coming, the unsupervised parameter selection scheme, which is widely used by non-local filters, limits the content adaptability. Given this, we propose a parametric non-local in-loop filter (PNLF) that fully considers the non-local characteristics and trains the filter coefficients based on the video content. In the filtering process, the reference samples based on the non-local similarity are first derived for each to-be-filtered sample. Then to-be-filtered samples are grouped into specific classes based on multiple features. For each class, filter coefficients are online trained in the encoder and transmitted to the decoder. Finally, the filtering process is conducted using the online-selected coefficients. Simulation results reveal that the proposed approach achieves 0.70%, 1.43%, and 2.09% bit-rate savings on average compared to VTM-11.0 under All Intra (AI), Random Access (RA), and Low-Delay B (LDB) configurations, respectively. The sequences used in the experiment include Class AI, A2, B, C, D, E, F, and SCC. Compared to the non-local structure-based filter (NLSF) [1], our proposed PNLF with fast block matching scheme [2] applied on B-frames and P-frames can achieve better performance gain with lower software and hardware complexity under RA and LDB configurations.
Xuewei Meng, Chuanmin Jia, Xinfeng Zhang 0001, Meng Lei, Shanshe Wang, Lin Li 0062, Siwei Ma 0001
DCC1
2021 Optimized Adaptive Loop Filter in Versatile Video Coding
abstract
In the Versatile Video Coding (VVC) standard, adaptive loop filter (ALF), including Geometry transformation-based Adaptive Loop Filter (GALF) and Cross Component Adaptive Loop Filter (CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is unfriendly to the software and hardware design. Therefore, we propose an optimized ALF framework, including the parallel design of GALF and CCALF, the adaptive parameter decision of GALF, and one-pass CCALF scheme by effectively estimating the CCALF filtering distortion without conducting filter operation. Compared to VTM-8.0, the proposed method can reduce the picture buffer access from 152 to 1 and achieve roughly 25% time-savings of the ALF module with negligible coding performance change under RA configuration. Some of the proposed methods have been adopted in the VVC reference software.
Xuewei Meng, Jiaqi Zhang 0007, Chuanmin Jia, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
DCC1