Li Zhang 0136

dblp:89/5992-136 · DBLP profile ↗
← Back
6ranked-venue papers in the field
0as first author
6since 2021 · last 2025
ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6
YearPublicationVenuePosition
2025 Template Matching Based Motion Refinement on Subblock Merge Mode
abstract
The subblock merge mode, in which the current coding block is split into multiple subblocks for motion compensation but still inherits the motion at the coding block level, improves the accuracy of the inter prediction and reduces the signaling overhead of motion information at the same time. And thus, it was adopted into versatile video coding (VVC) due to its high efficiency and continually improved in the enhanced compression model (ECM). However, the motion used in subblock merge mode was inherited from the previously coded blocks and may not match well with the current coding block. Thus, to improve the accuracy of the motion for the subblock merge mode, it is proposed to apply template matching (TM) based motion refinement. For subblock temporal motion vector predictor (SbTMVP) candidates, it is proposed to refine both the motion displacement and subblock motion vectors (MVs) based on TM; for affine motion candidates, it is proposed to refine the affine model, including base MV and non-translation parameters, based on TM. The proposed method was implemented on top of ECM, and the experiment results show that by applying the proposed method, it achieves {−0.23%(Y), −0.23%(U), −0.20%(V)} and {−0.09%(Y), −0.36%(U), −0.01%(V)} BD-rate reduction on random access (RA) and low delay B (LDB) configurations, respectively. Due to the good trade-off between performance and complexity, the proposed method was adopted into ECM.
Jie Chen 0006, Ru-Ling Liao, Yan Ye 0003, Lei Zhao 0032, Kai Zhang 0007, Li Zhang 0136
DCC7
2025 Compressed Screen Content Image Enhancement with B-Spline Based Distortion Estimation
abstract
Screen content has emerged as a prominent medium in our increasingly connected world. However, compressed screen content images often suffer from unpleasant artifacts, significantly obstructing the comprehension of text and graphic regions. In this paper, we introduce a quality enhancement framework specifically designed for compressed screen content images. We first propose a dataset for enhancing the quality of screen content images affected by various levels of compression distortion, using state-of-the-art Versatile Video Coding with screen content coding techniques enabled. Given the unique characteristics of screen content images, our enhancement framework incorporates B-spline representation to mitigate the quality degradation caused by compression. Additionally, we focus on recovering distorted text by detecting text regions within the degraded image and generating a pristine textual map to guide the recovery process. Experimental results demonstrate that our proposed method effectively enhances the quality of reconstructed screen content images across different compression distortion levels, leading to the quantitative and qualitative improvement.
Yue Li 0015, Chaoyi Lin, Kai Zhang 0007, Li Zhang 0136
DCC5
2025 CCLOP: Cross-Component Enhanced LOP Filter for Video Coding
abstract
Recent exploration efforts in JVET (Joint Video Experts Team of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC29) has achieved progresses on neural network-based video coding (NNVC)11NNVC is also the name of the reference software for evaluating neural network-based video coding technologies in JVET. The project locates at https://vcgit.hhi.fraunhofer.de/jvet-ahg-nnvc/VVCSoftware_VTM. Latest version of NNVC features two normative deep tools, i.e., neural network-based intra prediction and neural network-based in-loop filtering. Specifically, the neural network-based filtering in NNVC supports three operating points, known as VLOP (very low-complexity operating point), LOP (low-complexity operating point), and HOP (high-complexity operating point). LOP filter receives more attention among these three due to its favorable performance-complexity trade-off. In this paper, we introduce CCLOP, a cross-component enhanced LOP filter. CCLOP builds upon LOP filter in NNVC but incorporates deep luma features for chroma filtering. We conduct extensive experiments to verify the effectiveness of CCLOP. Compared with NNVC-10, the latest reference software of NNVC, CCLOP achieves {-0.13%, −2.27%, −3.11%}, {-0.18%, −2.07%, −3.21%}, and {-0.03%, −1.81%, −2.51%} BD-rate changes on average for {Y, Cb, Cr} under random-access, low-delay, and all-intra configurations respectively, while maintaining the same complexity as existing LOP filter ([email protected] kMAC/pixel, [email protected] kMAC/pixel).
Yue Li 0015, Chaoyi Lin, Kai Zhang 0007, Li Zhang 0136
DCC5
2025 CD: Cool-Chic Video with Decoupled Representation
abstract
Neural compression methods often rely on highly expressive models to fit large datasets, resulting in significant decoding complexity. Overfitted codecs have been proposed as an alternative to reduce decoding complexity. However, these approaches typically lack flexibility in encoding configurations. To address this, we introduce CD, a neural video compression method that employs picture-wise overfitting. CD is built upon the Cool-chic video framework [1], but incorporates Decoupled representations for motion and residue. Additionally, we propose an effective training strategy for CD to further enhance its performance.
Yue Li 0015, Chaoyi Lin, Kai Zhang 0007, Li Zhang 0136
DCC5
2024 Inter Cross-Component Prediction Merge Mode for Video Coding beyond VVC
abstract
As an incubator of next generation video coding techniques beyond versatile video coding (VVC) capability, enhanced compression model (ECM) has been initiated by the Joint Video Exploration Team (JVET). This paper presents an Inter CCP merge mode to improve the coding performance for chroma inter coding. Experimental results show that Inter CCP merge mode provides an average Bjontegaard delta rate (BD-rate) change of 0.01%/-0.69%/-0.76% and -0.03%/-2.31%/-2.38% on Y/Cb/Cr components, respectively, compared with ECM-10.0 in RA/LDB configurations under the common test condition, with a negligible running time change.
Zhipin Deng, Kai Zhang 0007, Li Zhang 0136
DCC3
2024 Geometric Partitioning Mode with Affine Prediction in Video Coding
abstract
Geometric partitioning mode (GPM) splits a coding block into two partitions, which can be non-rectangular, separated by a straight splitting line. Two uni inter-predictions generated by translational motion compensation (TMC) for the two GPM partitions are blended to obtain the final prediction. With a promising coding gain, GPM has been adopted in versatile video coding (VVC). Beyond VVC, enhanced compression model (ECM) introduces several extensions on GPM, but GPM still cannot deal with affine motions well. This paper presents a method of GPM with affine prediction (GPM-affine). A GPM partition can be predicted by affine motion compensation (AMC) or TMC, indicated by a flag. A GPM partition predicted by AMC can be blended with the other GPM partition predicted by AMC, TMC, or intra-prediction. Experimental results show that GPM-affine provides an average luma BD-rate saving of 0.19% compared to ECM-10.0 in random access configurations under the common test condition, with a negligible running time change. On sequences with rich affine motions, 1% coding gain in average is observed. Currently, GPM-affine is under study in exploration experiments (EE) for ECM in JVET.
Kai Zhang 0007, Zhipin Deng, Li Zhang 0136
DCC3