Na Li 0015

dblp:18/3173-15 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 9 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 VP-JND: Visual Perception Assisted Deep Picture-Wise Just Noticeable Difference Prediction Model for Image Compression
abstract
The Picture-Wise Just Noticeable Difference (PW-JND) represents the visibility threshold of human vision when viewing distorted images. The PW-JND plays an important role in perceptual image processing and compression. However, predicting the PW-JND is challenging due to its dependence on image content, viewing conditions, and the viewer. In this paper, we propose a visual perception-assisted deep PW-JND (VP-JND) prediction model for image compression that combines data-driven methods with the perceptual mechanisms of human vision. First, we identify a correlation between PW-JND and conventional pixel-wise JND. Based on this observation, we design the VP-JND model, consisting of a pixel-wise JND model, a deep binary classifier (VP-JNDnet) and a binary block search algorithm for refining predictions. VP-JNDnet exploits the pixel-wise JND map of the original image to predict whether a compressed image is perceptually lossless. In addition, the model incorporates visual importance of content and regions by using a mixed attention module and calculating perceptual loss during training. Experimental results show that VP-JND achieved an average precision of 94.82% and a mean absolute difference of 3.92 in predicting the JPEG quality factor corresponding to the PW-JND on the MCL-JCI dataset, outperforming state-of-the-art JND models. When applied to perceptual lossless image coding, the predicted PW-JND enabled average bit rate savings of 89.35% for JPEG compression on MCL-JCI and 85.46%/41.13% for JPEG/BPG compression on KonJND-1k. These savings were relative to images compressed at the lowest distortion level. The source codes and trained models are publicly available at https://github.com/SYSU-Video/VP-JND.
Yun Zhang 0002, Shisheng Zhang, Na Li 0015, Chunling Fan, Raouf Hamzaoui
IEEE Trans. Circuits Syst. Video Technol.3
2026 Deep Learning-Based Joint Geometry and Attribute Up-Sampling for Large-Scale Colored Point Clouds
abstract
Colored point cloud comprising geometry and attribute components is one of the mainstream representations enabling realistic and immersive 3D applications. To generate large-scale and denser colored point clouds, we propose a deep learning-based Joint Geometry and Attribute Up-sampling (JGAU) method, which learns to model both geometry and attribute patterns and leverages the spatial attribute correlation. Firstly, we establish and release a large-scale dataset for colored point cloud up-sampling, named SYSU-PCUD, which has 121 large-scale colored point clouds with diverse geometry and attribute complexities in six categories and four sampling rates. Secondly, to improve the quality of up-sampled point clouds, we propose a deep learning-based JGAU framework to up-sample the geometry and attribute jointly. It consists of a geometry up-sampling network and an attribute up-sampling network, where the latter leverages the up-sampled auxiliary geometry to model neighborhood correlations of the attributes. Thirdly, we propose two coarse attribute up-sampling methods, Geometric Distance Weighted Attribute Interpolation (GDWAI) and Deep Learning-based Attribute Interpolation (DLAI), to generate coarsely up-sampled attributes for each point. Then, we propose an attribute enhancement module to refine the up-sampled attributes and generate high quality point clouds by further exploiting intrinsic attribute and geometry patterns. Extensive experiments show that Peak Signal-to-Noise Ratio (PSNR) achieved by the proposed JGAU are 33.90 dB, 32.10 dB, 31.10 dB, and 30.39 dB when up-sampling rates are $4\times $ , $8\times $ , $12\times $ , and $16\times $ , respectively. Compared to the state-of-the-art schemes, the JGAU achieves an average of 2.32 dB, 2.47 dB, 2.28 dB and 2.11 dB PSNR gains at four up-sampling rates, respectively, which are significant. The code is released with https://github.com/SYSU-Video/JGAU.
Yun Zhang 0002, Feifan Chen, Na Li 0015, Xu Wang 0006, Fen Miao, Sam Kwong
IEEE Trans. Image Process.3
2026 C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image Compression
abstract
Learned Image Compression (LIC) has achieved superior performance in recent years, of which the context entropy model is an important component. However, in the context entropy model, there is no deterministic correlation between neighboring channels, and it is difficult to capture inter-channel correlation as well as spatial correlation for further improving the performance. To address this issue, a Cubic-Checkerboard conTeXt entropy model (C-CTX) for LIC is proposed in this work, which is able to refer uniformly across the channel domain and maintain the correlations in the spatial domain. To make neighboring channels have more similar distribution, Cubic Checkerboard Mask (CCM) with channel- wise mask convolution is utilized to achieve uniform distribution in different domains and Channel Wise Re-Arrangement (CWRA) is performed in terms of entropy. Based on CCM and CWRA, two Feature Disentangle Modules (FDMs) are designed in C-CTX to project the context information within sub-spaces for catching spatial correlation and channel correlation separately. Extensive experimental evaluations show that our method outperforms the state-of-the-art works on six datasets, i.e., Kodak, Tecnick, CLIC'20, CLIC'21, CLIC'22, and JPEG-AI.
Shiyu Feng, Linwei Zhu, Yun Zhang 0002, Na Li 0015, Shiqi Wang 0001
IEEE Trans. Multim.4
2026 RegR-PCQA: Deep Learning Based Colored Point Cloud Quality Assessment Using 3D-to-2D Regularized Representation
abstract
Point Cloud Quality Assessment (PCQA) aims to accurately predict the visual quality of a point cloud, which is essential in optimizing and evaluating the point cloud compression, transmission and rendering. In this paper, we propose a deep learning based full reference PCQA using 3D-to-2D Regularized Representation (RegR-PCQA), where point clouds are projected to regularized 2D image representations and then measured with deep neural networks. Firstly, we propose a regularized representation module to project unstructured point clouds to 2D Regularized Geometry Images (RGIs) and Regularized Attribute Images (RAIs), which enhance the local adjacency and uniform distribution of points. An anchor matching is developed to build the correspondence of regularized images between the distorted and reference point clouds. Secondly, to exploit the visual features of the RGIs and RAIs, we propose a deep learning based two-branch PCQA network, in which vision transformer based Geometry Feature Extractor (GFE) extracts global structural features from RGIs and Convolutional Neural Network (CNN) based Attribute Feature Extractor (AFE) extracts local semantic features of the RAIs. Finally, based on the geometry and attribute features, the point cloud quality is predicted by the proposed quality regression module, where a spatial attention mechanism is exploited to assign different importance weights for the feature maps. Experimental results show that the Pearson Linear Correlation Coefficients (PLCC) achieved by the proposed RegR-PCQA are 0.8430, 0.9575, 0.7853 and 0.8576, respectively, on the SIAT-PCQD, SJTU-PCQA, WPC and WPC2.0 datasets, which are superior to the state-of-the-art PCQAs. Also, extensive experimental results on distortion types, sampling strategy and training rate show that the proposed RegR-PCQA achieves an excellent generalization.
Yun Zhang 0002, Mao Cui, Na Li 0015, Chunling Fan, Weisi Lin
IEEE Trans. Multim.3
2024 Neural Network Based Multi-Level In-Loop Filtering for Versatile Video Coding
abstract
To further improve the performance of Versatile Video Coding (VVC), a neural network based multi-level in-loop filtering framework for luma and chroma is presented in this letter, which includes Reference pixel Level (RL), Coding tree unit Level (CL), and Frame Level (FL). The neural network based filters in these levels can be flexibly enabled. In RL, the coding performance upper bound is analyzed and asymmetric convolution is designed. In CL, the pixels located at the bottom and rightmost have been assigned greater weights for loss calculation during training. In addition, the co-located luma is adopted in CL and FL chroma filtering for guiding chroma enhancement due to the high correlation between them. For the architecture of neural network, two input channel fusion schemes are combined to enjoy both of their benefits. Extensive experimental results show that the proposed multi-level in-loop filtering method can achieve 6.87%, 32.8%, and 36.9% bit rate reductions on average for Y, U, and V components under all intra configuration, which outperforms the state-of-the-art works.
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Wenhui Wu 0001, Shiqi Wang 0001, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2023 Perceptually Weighted Rate Distortion Optimization for Video-Based Point Cloud Compression
abstract
Dynamic point cloud is a volumetric visual data representing realistic 3D scenes for virtual reality and augmented reality applications. However, its large data volume has been the bottleneck of data processing, transmission, and storage, which requires effective compression. In this paper, we propose a Perceptually Weighted Rate-Distortion Optimization (PWRDO) scheme for Video-based Point Cloud Compression (V-PCC), which aims to minimize the perceptual distortion of reconstructed point cloud at the given bit rate. Firstly, we propose a general framework of perceptually optimized V-PCC to exploit visual redundancies in point clouds. Secondly, a multi-scale Projection based Point Cloud quality Metric (PPCM) is proposed to measure the perceptual quality of 3D point cloud. The PPCM model comprises 3D-to-2D patch projection, multi-scale structural distortion measurement, and fusion model. Approximations and simplifications of the proposed PPCM are also presented for both V-PCC integration and low complexity. Thirdly, based on the simplified PPCM model, we propose a PWRDO scheme with Lagrange multiplier adaptation, which is incorporated into the V-PCC to enhance the coding efficiency. Experimental results show that the proposed PPCM models can be used as standalone quality metrics, and they are able to achieve higher consistency with the human subjective scores than the state-of-the-art objective visual quality metrics. Also, compared with the latest V-PCC reference model, the proposed PWRDO-based V-PCC scheme achieves an average bit rate reduction of 13.52%, 8.16%, 10.56% and 9.54%, respectively, in terms of four objective visual quality metrics for point clouds. It is significantly superior to the state-of-the-art coding algorithms. The computational complexity of the proposed PWRDO increases by 1.71% and 0.05% on average to the V-PCC encoder and decoder, respectively, which is negligible. The source codes of the PPCM and PWRDO schemes are available at https://github.com/VVCodec/PPCM-PWRDO.
Yun Zhang 0002, Keqin Ding, Na Li 0015, Hanli Wang, Xiaoxia Huang 0004, C.-C. Jay Kuo
IEEE Trans. Image Process.3
2023 Deep Learning-Based Intra Mode Derivation for Versatile Video Coding
abstract
In intra coding, Rate Distortion Optimization (RDO) is performed to achieve the optimal intra mode from a pre-defined candidate list. The optimal intra mode is also required to be encoded and transmitted to the decoder side besides the residual signal, where lots of coding bits are consumed. To further improve the performance of intra coding in Versatile Video Coding (VVC) , an intelligent intra mode derivation method is proposed in this paper, termed as Deep Learning based Intra Mode Derivation (DLIMD) . In specific, the process of intra mode derivation is formulated as a multi-class classification task, which aims to skip the module of intra mode signaling for coding bits reduction. The architecture of DLIMD is developed to adapt to different quantization parameter settings and variable coding blocks including non-square ones, where only one single trained model is required. Different from the existing deep learning based classification problems, the hand-crafted features are also fed into intra mode derivation network besides the learned features from feature learning network. To compete with traditional methods, one additional binary flag is utilized in the video codec to indicate the selected scheme with RDO. Extensive experimental results reveal that the proposed method can achieve 2.28%, 1.74%, and 2.18% bit rate reduction on average for Y, U, and V components on the platform of VVC test model, which outperforms the state-of-the-art works.
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Joint Source-Channel Decoding of Polar Codes for HEVC-Based Video Streaming
abstract
Ultra High-Definition (UHD) and Virtual Reality (VR) video streaming over 5G networks are emerging, in which High-Efficiency Video Coding (HEVC) is used as source coding to compress videos more efficiently and polar code is used as channel coding to transmit bitstream reliably over an error-prone channel. In this article, a novel Joint Source-Channel Decoding (JSCD) of polar codes for HEVC-based video streaming is presented to improve the streaming reliability and visual quality. Firstly, a Kernel Density Estimation (KDE) fitting approach is proposed to estimate the positions of error channel decoded bits. Secondly, a modified polar decoder called R-SCFlip is designed to improve the channel decoding accuracy. Finally, to combine the KDE estimator and the R-SCFlip decoder together, the JSCD scheme is implemented in an iterative process. Extensive experimental results reveal that, compared to the conventional methods without JSCD, the error data-frame correction ratios are increased. Averagely, 1.07% and 1.11% Frame Error Ratio (FER) improvements have been achieved for Additive White Gaussian Noise (AWGN) and Rayleigh fading channels, respectively. Meanwhile, the qualities of the recovered videos are significantly improved. For the 2D videos, the average Peak Signal-to-Noise Ratio (PSNR) and Structural SIMilarity (SSIM) gains reach 14% and 34%, respectively. For the 360֯ videos, the average improvements in terms of Weighted-to-Spherically-uniform PSNR (WS-PSNR) and Voronoi-based Video Multimethod Assessment Fusion (VI-VMAF) reach 21% and 7%, respectively.
Jinzhi Lin, Yun Zhang 0002, Na Li 0015, Hongling Jiang
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Circular intra prediction for 360 degree video coding
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Jinyong Pi, Shiqi Wang 0001
J. Vis. Commun. Image Represent.3
2020 Sparse Representation-Based Intra Prediction for Lossless/Near Lossless Video Coding
abstract
In this paper, a novel intra prediction method is presented for lossless/near lossless High Efficiency Video Coding (HEVC), termed as Sparse Representation based Intra Prediction (SRIP). In specific, the existing Angular Intra Prediction (AIP) modes in HEVC are organized as a mode dictionary, which is utilized to sparsely represent the visual signal by minimizing the difference with respect to the ground truth. For the match of encoding and decoding, the sparse coefficients are also required to be encoded and transmitted to the decoder side. To further improve the coding performance, an additional binary flag is included in the video codec to indicate which strategy is finally adopted with the rate distortion optimization, i.e., SRIP or traditional AIP. Extensive experimental results reveal that the proposed method can achieve 0.36% bit rate saving on average in case of lossless scenario.
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Jinyong Pi, Xinju Wu
VCIP3
2019 Reinforcement learning based coding unit early termination algorithm for high efficiency video coding
Na Li 0015, Yun Zhang 0002, Linwei Zhu, Wenhan Luo, Sam Kwong
J. Vis. Commun. Image Represent.1
2019 Statistical Early Termination and Early Skip Models for Fast Mode Decision in HEVC INTRA Coding
abstract
In this article, statistical Early Termination (ET) and Early Skip (ES) models are proposed for fast Coding Unit (CU) and prediction mode decision in HEVC INTRA coding, in which three categories of ET and ES sub-algorithms are included. First, the CU ranges of the current CU are recursively predicted based on the texture and CU depth of the spatial neighboring CUs. Second, the statistical model based ET and ES schemes are proposed and applied to optimize the CU and INTRA prediction mode decision, in which the coding complexities over different decision layers are jointly minimized subject to acceptable rate-distortion degradation. Third, the mode correlations among the INTRA prediction modes are exploited to early terminate the full rate-distortion optimization in each CU decision layer. Extensive experiments are performed to evaluate the coding performance of each sub-algorithm and the overall algorithm. Experimental results reveal that the overall proposed algorithm can achieve 45.47% to 74.77%, and 58.09% on average complexity reduction, while the overall Bjøntegaard delta bit rate increase and Bjøntegaard delta peak signal-to-noise ratio degradation are 2.29% and −0.11 dB, respectively.
Yun Zhang 0002, Na Li 0015, Sam Kwong, Gangyi Jiang, Huanqiang Zeng
ACM Trans. Multim. Comput. Commun. Appl.2
2018 Effective Data Driven Coding Unit Size Decision Approaches for HEVC INTRA Coding
abstract
High Efficiency Video Coding (HEVC) INTRA coding improves compression efficiency by adopting advanced coding technologies, such as multi-level quad-tree block partitioning and up to 35-mode INTRA prediction. However, it significantly increases the coding complexity, memory access, and power consumption, which goes against its widely applications, especially for ultra-high definition and/or mobile video applications. To tackle this problem, we propose effective data driven coding unit (CU) size decision approaches for HEVC INTRA coding, which consists of two stages of support vector machine-based fast INTRA CU size decision schemes at four CU decision layers. At the first stage classification, a three output classifier with offline learning is developed to early terminate the CU size decision or early skip checking the current CU depth. As for the samples that neither early skipped nor early terminated, the second stage of binary classification, which learns online from previous coded frames, is proposed to further refine the CU size decision. Representative features for the CU size decision are explored at different decision layers and stages of classifications. Finally, the optimal parameters derived from the training data are achieved to reasonably allocate complexity among different CU layers at given total rate-distortion degradation constraint. Extensive experiments show that the proposed overall algorithm can achieve 27.95%–80.53% and 52.48% on average complexity reduction for the CU size decision as compared with the original HM16.7 model. Meanwhile, the average Bjonteggard delta peak-signal-to-noise ratio degradation is only −0.08 dB, which is negligible. The overall performance of the proposed algorithm outperforms the state-of-the-art benchmark schemes.
Yun Zhang 0002, Zhaoqing Pan, Na Li 0015, Xu Wang 0006, Gangyi Jiang, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2017 Instant coherent group motion filtering by group motion representations
Na Li 0015, Yun Zhang 0002, Wenhan Luo
Neurocomputing1
2016 Machine learning based fast H.264/AVC to HEVC transcoding exploiting block partition similarity
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong
J. Vis. Commun. Image Represent.3