EDBT 2026 Demo / reviewers in the wild / expert
Yun Zhang 0002
dblp:02/6428-2
· DBLP profile ↗
101ranked-venue papers
23as first author
39since 2021 · last 2026
0000-0001-9457-7801ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 20 first-author · 35 since 2021Computer networks · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point CloudsabstractColored point cloud becomes a fundamental representation in the realm of 3D vision. Effective Point Cloud Compression (PCC) is urgently needed due to the huge amount of data. In this paper, we propose an end-to-end Deep Joint Geometry and Attribute Compression (Deep-JGAC) method for dense colored point clouds. First, we propose a flexible Deep-JGAC framework, where the geometry and attribute encoders are compatible with either learning or non-learning encoders. Second, we propose an end-to-end deep residual self-attention-based geometry encoder to improve geometry coding efficiency, where a Hybrid Residual Self-attention Module (HRSM) is proposed to enhance geometry representation by considering its geometrical importance. Third, to solve the mismatch between the point cloud geometry and attribute caused by the geometry compression distortion, we propose an optimized re-colorization module to attach attribute to the geometrically distorted point cloud for attribute coding, which lowers the computational complexity. Extensive experimental results demonstrate that, in terms of the geometry quality metric D1-PSNR, the proposed Deep-JGAC achieves average Bjøntegaard Delta Bit Rate (BDBR) of -82.96%, -44.63%, -36.46%, -41.72%, and -31.16% compared to the G-PCC (Octree), G-PCC (Trisoup), V-PCC, GRASP, and PCGCv2, respectively. For the perceptual joint quality metric MS-GraphSIM, Deep-JGAC achieves an average BDBR of - 48.72%, -57.14%, -14.67% and -13.37% against G-PCC(Octree), IT-DL-PCC, V-PCC, and DeepPCC, respectively. In addition, the costs of encoding/decoding time are reduced by 32.8%/30.8%, 80.1%/81.8%, 97.2%/35.7%, 98.4%/92.3%, and 96.4%/99.6% on average compared to G-PCC (Octree), G-PCC (Trisoup), V-PCC, IT-DL-PCC and DeepPCC. The code and pre-trained models are available at https://github.com/SYSU-Video/Deep-JGAC. Yun Zhang 0002, Zixi Guo, Linwei Zhu, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | VP-JND: Visual Perception Assisted Deep Picture-Wise Just Noticeable Difference Prediction Model for Image CompressionabstractThe Picture-Wise Just Noticeable Difference (PW-JND) represents the visibility threshold of human vision when viewing distorted images. The PW-JND plays an important role in perceptual image processing and compression. However, predicting the PW-JND is challenging due to its dependence on image content, viewing conditions, and the viewer. In this paper, we propose a visual perception-assisted deep PW-JND (VP-JND) prediction model for image compression that combines data-driven methods with the perceptual mechanisms of human vision. First, we identify a correlation between PW-JND and conventional pixel-wise JND. Based on this observation, we design the VP-JND model, consisting of a pixel-wise JND model, a deep binary classifier (VP-JNDnet) and a binary block search algorithm for refining predictions. VP-JNDnet exploits the pixel-wise JND map of the original image to predict whether a compressed image is perceptually lossless. In addition, the model incorporates visual importance of content and regions by using a mixed attention module and calculating perceptual loss during training. Experimental results show that VP-JND achieved an average precision of 94.82% and a mean absolute difference of 3.92 in predicting the JPEG quality factor corresponding to the PW-JND on the MCL-JCI dataset, outperforming state-of-the-art JND models. When applied to perceptual lossless image coding, the predicted PW-JND enabled average bit rate savings of 89.35% for JPEG compression on MCL-JCI and 85.46%/41.13% for JPEG/BPG compression on KonJND-1k. These savings were relative to images compressed at the lowest distortion level. The source codes and trained models are publicly available at https://github.com/SYSU-Video/VP-JND. Yun Zhang 0002, Shisheng Zhang, Na Li 0015, Chunling Fan, Raouf Hamzaoui |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Deep Learning-Based Joint Geometry and Attribute Up-Sampling for Large-Scale Colored Point CloudsabstractColored point cloud comprising geometry and attribute components is one of the mainstream representations enabling realistic and immersive 3D applications. To generate large-scale and denser colored point clouds, we propose a deep learning-based Joint Geometry and Attribute Up-sampling (JGAU) method, which learns to model both geometry and attribute patterns and leverages the spatial attribute correlation. Firstly, we establish and release a large-scale dataset for colored point cloud up-sampling, named SYSU-PCUD, which has 121 large-scale colored point clouds with diverse geometry and attribute complexities in six categories and four sampling rates. Secondly, to improve the quality of up-sampled point clouds, we propose a deep learning-based JGAU framework to up-sample the geometry and attribute jointly. It consists of a geometry up-sampling network and an attribute up-sampling network, where the latter leverages the up-sampled auxiliary geometry to model neighborhood correlations of the attributes. Thirdly, we propose two coarse attribute up-sampling methods, Geometric Distance Weighted Attribute Interpolation (GDWAI) and Deep Learning-based Attribute Interpolation (DLAI), to generate coarsely up-sampled attributes for each point. Then, we propose an attribute enhancement module to refine the up-sampled attributes and generate high quality point clouds by further exploiting intrinsic attribute and geometry patterns. Extensive experiments show that Peak Signal-to-Noise Ratio (PSNR) achieved by the proposed JGAU are 33.90 dB, 32.10 dB, 31.10 dB, and 30.39 dB when up-sampling rates are $4\times $ , $8\times $ , $12\times $ , and $16\times $ , respectively. Compared to the state-of-the-art schemes, the JGAU achieves an average of 2.32 dB, 2.47 dB, 2.28 dB and 2.11 dB PSNR gains at four up-sampling rates, respectively, which are significant. The code is released with https://github.com/SYSU-Video/JGAU. Yun Zhang 0002, Feifan Chen, Na Li 0015, Xu Wang 0006, Fen Miao, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2026 | Rate-Reconfigurable Deep Point Cloud Compression With Perceptual Bit Allocation OptimizationabstractConventional end-to-end learning-based point cloud compression requires training multiple models to adapt to different target bit rates. Moreover, the rate difference between geometry and attribute components of point clouds is not well-considered. In this paper, we propose an end-to-end Rate-Reconfigurable Deep Point Cloud Compression (RR-DPCC) with on/off-line Perceptual Bit Allocation Optimization (PBAO-ON/OFF), which achieves arbitrary bit rate control with one trained deep model and high efficiency joint geometry and attribute coding. First, we propose the framework of the RR-DPCC using PBAO-ON/OFF, which includes Point Cloud Quality Assessment (PCQA) for perceptual quality measurement, PBAO-ON/OFF modules for bit allocation and RR-DPCC for high efficiency point cloud coding. Second, we propose a one-stream network of the RR-DPCC to encode the attribute and geometry of point clouds jointly. Moreover, in RR-DPCC, a bitrate reconfigurable module is proposed to encode multiple fine-grained bitrate points with one trained model and a rate allocation module is proposed to allocate bits between geometry and attribute. Third, we propose on/off-line PBAO algorithms to maximize the perceptual quality of the reconstructed point cloud, where the bits are properly allocated based on the importance of geometry and attribute. Meanwhile, rate-distortion models (R- $\alpha $ / $\beta $ and D- $\alpha $ / $\beta $ ) are derived for high accuracy rate control and bit allocation. Experimental results show that the proposed RR-DPCC achieves fine-grained bitrate control and allocation through a single trained model. When combined the proposed RR-DPCC with PBAO-ON, it reduces -6.56% and -18.68% bit rate on average as comparing with the state-of-the-art V-PCC and Deep Joint Geometry and Attribute Compression (Deep-JGAC), respectively. When combined with the PBAO-OFF, it achieves -4.90% and -15.34% bit rate reductions on average, and reduces 98.38%/22.05% and 53.75%/10.04% encoding/decoding time on average with respect to V-PCC and Deep-JGAC. Yun Zhang 0002, Lewen Fan, Zixi Guo, Xu Wang 0006, Xiaoxia Huang 0004, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2026 | C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image CompressionabstractLearned Image Compression (LIC) has achieved superior performance in recent years, of which the context entropy model is an important component. However, in the context entropy model, there is no deterministic correlation between neighboring channels, and it is difficult to capture inter-channel correlation as well as spatial correlation for further improving the performance. To address this issue, a Cubic-Checkerboard conTeXt entropy model (C-CTX) for LIC is proposed in this work, which is able to refer uniformly across the channel domain and maintain the correlations in the spatial domain. To make neighboring channels have more similar distribution, Cubic Checkerboard Mask (CCM) with channel- wise mask convolution is utilized to achieve uniform distribution in different domains and Channel Wise Re-Arrangement (CWRA) is performed in terms of entropy. Based on CCM and CWRA, two Feature Disentangle Modules (FDMs) are designed in C-CTX to project the context information within sub-spaces for catching spatial correlation and channel correlation separately. Extensive experimental evaluations show that our method outperforms the state-of-the-art works on six datasets, i.e., Kodak, Tecnick, CLIC'20, CLIC'21, CLIC'22, and JPEG-AI. Shiyu Feng, Linwei Zhu, Yun Zhang 0002, Na Li 0015, Shiqi Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | DT-JRD: Deep Transformer-Based Just Recognizable Difference Prediction Model for Video Coding for MachinesabstractJust Recognizable Difference (JRD) represents the minimum visual difference that is detectable by machine vision, which can be exploited to promote machine vision-oriented visual signal processing. In this paper, we propose a Deep Transformer-based JRD (DT-JRD) prediction model for Video Coding for Machines (VCM), where the accurately predicted JRD can be used to reduce the coding bit rate while maintaining the accuracy of machine tasks. Firstly, we model the JRD prediction as a multi-class classification and propose a DT-JRD prediction model that integrates an improved embedding, a content and distortion feature extraction, a multi-class classification, and a novel learning strategy. Secondly, inspired by the perception property that machine vision exhibits a similar response to distortions near JRD, we propose an asymptotic JRD loss by using Gaussian Distribution-based Soft Labels (GDSL), which significantly extends the number of training labels and relaxes classification boundaries. Finally, we propose a DT-JRD-based VCM to reduce the coding bits while maintaining the accuracy of object detection. Extensive experimental results demonstrate that the mean absolute error of the predicted JRD by the DT-JRD is 5.574, outperforming the state-of-the-art JRD prediction model by 13.1%. Coding experiments show that compared with the VVC, the DT-JRD-based VCM achieves an average of 29.58% bit rate reduction while maintaining the object detection accuracy. Yun Zhang 0002, Long Xu 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2026 | RegR-PCQA: Deep Learning Based Colored Point Cloud Quality Assessment Using 3D-to-2D Regularized RepresentationabstractPoint Cloud Quality Assessment (PCQA) aims to accurately predict the visual quality of a point cloud, which is essential in optimizing and evaluating the point cloud compression, transmission and rendering. In this paper, we propose a deep learning based full reference PCQA using 3D-to-2D Regularized Representation (RegR-PCQA), where point clouds are projected to regularized 2D image representations and then measured with deep neural networks. Firstly, we propose a regularized representation module to project unstructured point clouds to 2D Regularized Geometry Images (RGIs) and Regularized Attribute Images (RAIs), which enhance the local adjacency and uniform distribution of points. An anchor matching is developed to build the correspondence of regularized images between the distorted and reference point clouds. Secondly, to exploit the visual features of the RGIs and RAIs, we propose a deep learning based two-branch PCQA network, in which vision transformer based Geometry Feature Extractor (GFE) extracts global structural features from RGIs and Convolutional Neural Network (CNN) based Attribute Feature Extractor (AFE) extracts local semantic features of the RAIs. Finally, based on the geometry and attribute features, the point cloud quality is predicted by the proposed quality regression module, where a spatial attention mechanism is exploited to assign different importance weights for the feature maps. Experimental results show that the Pearson Linear Correlation Coefficients (PLCC) achieved by the proposed RegR-PCQA are 0.8430, 0.9575, 0.7853 and 0.8576, respectively, on the SIAT-PCQD, SJTU-PCQA, WPC and WPC2.0 datasets, which are superior to the state-of-the-art PCQAs. Also, extensive experimental results on distortion types, sampling strategy and training rate show that the proposed RegR-PCQA achieves an excellent generalization. Yun Zhang 0002, Mao Cui, Na Li 0015, Chunling Fan, Weisi Lin |
IEEE Trans. Multim. | 1 |
| 2026 | Temporal Consistency-Aware Dynamic Point Clouds Color Attribute EnhancementabstractDynamic point clouds, widely used in virtual reality and autonomous driving systems, often suffer from distortions due to quantization in the process of compression. These distortions significantly degrade the visual quality of dynamic point clouds, especially temporal inconsistency. To address this issue, a temporal consistency-aware dynamic point clouds color attribute enhancement method is proposed in this work. Specifically, a 3D Spatial-Temporal Search (STS) module is designed to adaptively search point cloud patches in the temporal domain for feature alignment. These matched patches are then individually fed into Single Frame Feature Extraction (SFFE) module that comprises of multi-head attention and graph convolution to exploit latent features of point cloud color attribute. In addition, to further capture both the spatial and temporal dependencies, a Convolutional Point cloud Long Short-Term Memory (Conv-PointLSTM) network is applied, which integrates convolution and max pooling with LSTM mechanism to facilitate the color attribute correspondents across the spatial-temporal latent features. Experimental results demonstrate that the proposed method can achieve 0.44 dB gains on average in terms of Peak Signal-to-Noise Ratio (PSNR) and 1.50%/5.31%bit rate reductions at the low/high bit rate, which outperforms the state-of-the-art works. The source code and trained models are available athttps://github.com/xu-coder-666/DPC. Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Hui Yuan 0001, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2026 | Global-Local Progressive Integration and Semantic-Aligned Quality Transfer for No-Reference Image Quality AssessmentabstractAccurate image quality assessment without reference signals presents a fundamental challenge in low-level visual tasks. In this article, we propose a global–local progressive integration model with three key contributions. (1) We develop a dual-feature extraction framework combining vision Transformer (ViT)-based global feature extractor and convolutional neural networks (CNNs)-based local feature extractor to capture image distortions at different granularities. (2) We propose a progressive feature integration scheme with multi-scale kernels to align global–local features, followed by channel-wise self-attention and spatial interaction for multi-grained representations. (3) We propose a semantic-aligned quality transfer (SAQT) method that extends the training data by assigning subjective quality scores to diverse image content. Experimental results demonstrate that our model yields 5.04% and 5.40% improvements in SROCC over the second-best state-of-the-art methods (i.e., DEIQT and CICI) for cross-authentic and cross-synthetic dataset generalization tests, respectively. Furthermore, the proposed SAQT method further yields 2.26% and 13.23% performance gains in evaluations on single-synthetic and cross-synthetic datasets. Code and dataset are available at: https://github.com/SYSU-Video/GlintIQA . Yun Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | CausalFi: Causality-Based Cross-Domain Human Activity Recognition With Wi-FiabstractWiFi-based human activity recognition (HAR) has demonstrated significant potential in diverse intelligent applications. However, sensitive to environmental factors, cross-domain WiFi channel state information (CSI) poses significant challenges in the generalization of HAR models across different environments. In this paper, the structural causal model (SCM) is introduced to model causal relationships among activities, latent variables, and CSI data, laying a solid foundation toward developing a domain-invariant model for WiFi-based HAR tasks. In this paper, we propose CausalFi, a novel framework that leverages causal inference to mitigate the confounding effects of latent variables, significantly improving generalization performance. Integrating novel feature selection and importance sampling algorithms as the condition and intervention operations, CausalFi can effectively identify the stable action-relevant features from WiFi CSI for activity recognition. Furthermore, a novel counterfactual style augmentation approach is proposed to increase the stylistic diversity of the training data, reducing the risk of biased data distributions even with limited training samples in source domains. We implement a prototype of CausalFi using commercial ASUS RT-AC86U WiFi devices and conduct extensive cross-domain experiments to validate the effectiveness of the proposed approach. With an average recognition accuracy of 92.6%, CausalFi significantly outperforms state-of-the-art baselines in complex cross-domain environments, confirming the practicality of our framework for real-world WiFi-based HAR applications. © 2026 IEEE. Yang Zhou 0051, Xiaoxia Huang 0004, Yun Zhang 0002, Yuguang Fang |
IEEE Trans. Netw. | 4 |
| 2025 | Learning Based Fast Coding Unit Decision for Video-Based Point Cloud Compression
Lewen Fan, Yun Zhang 0002 |
ICIG (1) | 2 |
| 2025 | Joint multi-dimensional dynamic attention and transformer for general image restoration
Huan Zhang 0008, Xu Zhang 0044, Nian Cai, Jianglei Di, Yun Zhang 0002 |
Comput. Vis. Image Underst. | 5 |
| 2025 | Enhancing 3D video watching experiences: Tackling compression and 3D warping distortions in synthesized view with perceptual guidance
Huan Zhang 0008, Xu Zhang 0044, Linwei Zhu, Yun Zhang 0002, Jiang-Zhong Cao, Bingo Wing-Kuen Ling |
Expert Syst. Appl. | 4 |
| 2025 | Multi-granular embedding optimization with spatial-channel adaptive tuning for perceptual image quality assessment
Yun Zhang 0002 |
Neurocomputing | 2 |
| 2025 | Texture-aware fast mode decision and complexity allocation for VVC based point cloud compressionabstractVideo-Based Point Cloud Compression (V-PCC) leverages Versatile Video Coding (VVC) to compress point clouds efficiently, yet suffers from high computational complexity that challenges real-time applications. To address this critical problem, we propose a texture-aware fast Coding Unit (CU) mode decision algorithm and a complexity allocation strategy for VVC-based V-PCC. By analyzing CU distributions and complexity characteristics, we introduce adaptive early termination thresholds that incorporate spatial, parent–child, and intra/inter CU correlations. Furthermore, we established a complexity allocation method by formulating and solving an optimization problem to determine relaxation factors for optimal complexity-efficiency trade-offs. Experimental results demonstrate that the proposed fast mode decision achieves an average of 33.89% and 44.59% complexity reductions compared to the anchor VVC-based V-PCC, which are better than those of the state-of-the-art fast mode decision schemes. Meanwhile, the average Bjónteggard Delta Bit Rate (BDBR) loss is 1.04% and 1.64%, which are negligible. Lewen Fan, Yun Zhang 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Geometry-Guided Latent Diffusion Model for Static Point Cloud Color Attribute Denoising
Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Signal Process. Lett. | 3 |
| 2025 | Multiscale Feature Importance-Based Bit Allocation for End-to-End Feature Coding for MachinesabstractFeature Coding for Machines (FCM) aims to compress intermediate features effectively for remote intelligent analytics, which is crucial for future intelligent visual applications. In this article, we propose a Multiscale Feature Importance-based Bit Allocation (MFIBA) for end-to-end FCM. First, we find that the importance of features for machine vision tasks varies with the scales, object size, and image instances. Based on this finding, we propose a Multiscale Feature Importance Prediction (MFIP) module to predict the importance weight for each scale of features. Second, we propose a task loss-rate model to establish the relationship between the task accuracy losses of using compressed features and the bit rate of encoding these features. Finally, we develop an MFIBA for end-to-end FCM, which is able to assign coding bits of multiscale features more reasonably based on their importance. Experimental results demonstrate that when combined with a retained Efficient Learned Image Compression (ELIC), the proposed MFIBA achieves an average of 38.202% bit-rate savings in object detection compared to the anchor ELIC. Moreover, the proposed MFIBA achieves an average of 17.212% and 36.492% feature bit-rate savings for instance segmentation and keypoint detection, respectively. When the proposed MFIBA is applied to the LIC-TCM, it achieves an average of 18.103%, 19.866%, and 19.597% bit-rate savings on three machine vision tasks, respectively, which validates the proposed MFIBA has good generalizability and adaptability to different machine vision tasks and FCM base codecs. Junle Liu, Yun Zhang 0002, Zixi Guo, Xiaoxia Huang 0004, Gangyi Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | RGB-D Data Compression via Bi-Directional Cross-Modal Prior Transfer and Enhanced Entropy ModelingabstractRGB-D data, being homogeneous cross-modal data, demonstrates significant correlations among data elements. However, current research focuses only on a uni-directional pattern of cross-modal contextual information, neglecting the exploration of bi-directional relationships in the compression field. Thus, we propose a joint RGB-D compression scheme, which is combined with Bi-Directional Cross-Modal Prior Transfer (Bi-CPT) modules and a Bi-Directional Cross-Modal Enhanced Entropy (Bi-CEE) model. The Bi-CPT module is designed for compact representations of cross-modal features, effectively eliminating spatial and modality redundancies at different granularity levels. In contrast to the traditional entropy models, our proposed Bi-CEE model not only achieves spatial-channel contextual adaptation through partitioning RGB and depth features but also incorporates information from other modalities as prior to enhance the accuracy of probability estimation for latent variables. Furthermore, this model enables parallel multi-stage processing to accelerate coding. Experimental results demonstrate the superiority of our proposed framework over the current compression scheme, outperforming both rate-distortion performance and downstream tasks, including surface reconstruction and semantic segmentation. The source code will be available at https://github.com/xyy7/Learning-based-RGB-D-Image-Compression . Yuyu Xu, Qiudan Zhang, Wenhui Wu 0001, Yun Zhang 0002, Xu Wang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Neural Network Based Multi-Level In-Loop Filtering for Versatile Video CodingabstractTo further improve the performance of Versatile Video Coding (VVC), a neural network based multi-level in-loop filtering framework for luma and chroma is presented in this letter, which includes Reference pixel Level (RL), Coding tree unit Level (CL), and Frame Level (FL). The neural network based filters in these levels can be flexibly enabled. In RL, the coding performance upper bound is analyzed and asymmetric convolution is designed. In CL, the pixels located at the bottom and rightmost have been assigned greater weights for loss calculation during training. In addition, the co-located luma is adopted in CL and FL chroma filtering for guiding chroma enhancement due to the high correlation between them. For the architecture of neural network, two input channel fusion schemes are combined to enjoy both of their benefits. Extensive experimental results show that the proposed multi-level in-loop filtering method can achieve 6.87%, 32.8%, and 36.9% bit rate reductions on average for Y, U, and V components under all intra configuration, which outperforms the state-of-the-art works. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Wenhui Wu 0001, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Colored Point Cloud Quality Assessment Using Complementary Features in 3D and 2D SpacesabstractPoint Cloud Quality Assessment (PCQA) plays an essential role in optimizing point cloud acquisition, encoding, transmission, and rendering for human-centric visual media applications. In this paper, we propose an objective PCQA model using Complementary Features from 3D and 2D spaces, called CF-PCQA, to measure the visual quality of colored point clouds. First, we develop four effective features in 3D space to represent the perceptual properties of colored point clouds, which include curvature, kurtosis, luminance distance and hue features of points in 3D space. Second, we project the 3D point cloud onto 2D planes using patch projection and extract a structural similarity feature of the projected 2D images in the spatial domain, as well as a sub-band similarity feature in the wavelet domain. Finally, we propose a feature selection and a learning model to fuse high dimensional features and predict the visual quality of the colored point clouds. Extensive experimental results show that the Pearson Linear Correlation Coefficients (PLCCs) of the proposed CF-PCQA were 0.9117, 0.9005, 0.9340 and 0.9826 on the SIAT-PCQD, SJTU-PCQA, WPC2.0 and ICIP2020 datasets, respectively. Moreover, statistical significance tests demonstrate that the CF-PCQA significantly outperforms the state-of-the-art PCQA benchmark schemes on the four datasets. Mao Cui, Yun Zhang 0002, Chunling Fan, Raouf Hamzaoui, Qinglan Li |
IEEE Trans. Multim. | 2 |
| 2024 | Learning to Predict Object-Wise Just Recognizable Distortion for Image and Video CompressionabstractJust Recognizable Distortion (JRD) refers to the minimum distortion that notably affects the recognition performance of a machine vision model. If a distortion added to images or videos falls within this JRD threshold, the degradation of the recognition performance will be unnoticeable. Based on this JRD property, it will be useful to Video Coding for Machine (VCM) to minimize the bit rate while maintaining the recognition performance of compressed images. In this study, we propose a deep learning-based JRD prediction model for image and video compression. We first construct a large image dataset of Object-Wise JRD (OW-JRD) containing 29,218 original images with 80 object categories, and each image was compressed into 64 distorted versions using Versatile Video Coding (VVC). Secondly, we analyze of the distribution of the OW-JRD, formulate JRD prediction as binary classification problems and propose a deep learning-based OW-JRD prediction framework. Thirdly, we propose a deep learning based binary OW-JRD predictor to predict whether an image object is still detectable or not under different compression levels. Also, we propose an error-tolerance strategy that corrects misclassifications from the binary classifier. Finally, extensive experiments on large JRD image datasets demonstrate that the Mean Absolute Errors (MAEs) of the predicted OW-JRD are 4.90 and 5.92 on different numbers of the classes, which is significantly better than the state-of-the-art JRD prediction model. Moreover, ablation studies on deep network structures, object sizes, features, data padding strategies and image/video coding schemes are presented to validate the effectiveness of the proposed JRD model. Yun Zhang 0002, Haoqin Lin, Jing Sun 0010, Linwei Zhu, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2024 | Quality Assessment for DIBR-Synthesized Views Based on Wavelet Transform and Gradient Magnitude SimilarityabstractTo drive upgrades of Depth-Image-Based Rendering (DIBR) algorithms, depth image refinement, etc., quality assessment models for DIBR-synthesized images in 3D video systems are developed. However, most of these models could not effectively evaluate distortion due to irregular stretching (e.g., crumbling), which is more complex and common than black holes and regular stretching (e.g., horizontal stretching) in synthesized images. To make an attempt at this issue, a new quality assessment method is proposed for DIBR views. First, feature point matching and affine transformation are adopted to remove and compensate for the global object shift between reference and synthesized view images. Second, multi-scale discrete wavelet transform is utilized to extract multi-scale structure distortion; gradient magnitude similarity is further integrated to highlight the distortion features; morphological open operation and median filtering are adopted to exclude perceptually unimportant features. Third, scores are obtained by standard deviation pooling on distortion feature maps for each wavelet scale and sub-band. Experimental results demonstrate that our proposed model outperforms the state-of-the-art handcrafted feature-based DIBR-synthesized image quality assessment models on IETR database, and performs the best on average on IETR and IRCCyN/IVC databases. The source code will be available athttps://github.com/House-yuyu/DIBR_IQA. Huan Zhang 0008, Dongsheng Zheng, Yun Zhang 0002, Jiang-Zhong Cao, Weisi Lin, Bingo Wing-Kuen Ling |
IEEE Trans. Multim. | 3 |
| 2024 | Dynamic Weighted Gradient Reversal Network for Visible-infrared Person Re-identificationabstractDue to intra-modality variations and cross-modality discrepancy, visible-infrared person re-identification (VI Re-ID) is an important and challenging task in intelligent video surveillance. The cross-modality discrepancy is mainly caused by the differences between visible images and infrared images, the inherent essence of which is heterogeneous. To alleviate this discrepancy, we propose a Dynamic Weighted Gradient Reversal Network (DGRNet) to enhance the learning of discriminative common representations by confusing the modality discrimination. In the proposed DGRNet, we design the gradient reversal model guiding adversarial training between identity classifier and modality discriminator to reduce the modality discrepancy of the same person in different modalities. Furthermore, we propose an optimization training method, that is, designing dynamic weight of gradient reversal to achieve optimal adversarial training, and dynamic weight has the ability to dynamically and adaptively evaluate the significance of target loss term, without involving hyper-parameter tuning. Extensive experiments were conducted on two public VI Re-ID datasets, SYSU-MM01 and RegDB. The experimental results show that the proposed DGRNet outperforms state-of-the-art methods and demonstrate the effectiveness of the DGRNet to learn more discriminative common representations for VI Re-ID. Chenghua Li, Zongze Li 0007, Jing Sun 0010, Yun Zhang 0002, Xiaoping Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Joint Graph Attention and Asymmetric Convolutional Neural Network for Deep Image CompressionabstractRecent deep image compression methods have achieved prominent progress by using nonlinear modeling and powerful representation capabilities of neural networks. However, most existing learning-based image compression approaches employ customized convolutional neural network (CNN) to utilize visual features by treating all pixels equally, neglecting the effect of local key features. Meanwhile, the convolutional filters in CNN usually express the local spatial relationship within the receptive field and seldom consider the long-range dependencies from distant locations. This results in the long-range dependencies of latent representations not being fully compressed. To address these issues, an end-to-end image compression method is proposed by integrating graph attention and asymmetric convolutional neural network (ACNN). Specifically, ACNN is used to strengthen the effect of local key features and reduce the cost of model training. Graph attention is introduced into image compression to address the bottleneck problem of CNN in modeling long-range dependencies. Meanwhile, regarding the limitation that existing attention mechanisms for image compression hardly share information, we propose a self-attention approach which allows information flow to achieve reasonable bit allocation. The proposed self-attention approach is in compliance with the perceptual characteristics of human visual system, as information can interact with each other via attention modules. Moreover, the proposed self-attention approach takes into account channel-level relationship and positional information to promote the compression effect of rich-texture regions. Experimental results demonstrate that the proposed method achieves state-of-the-art rate-distortion performances after being optimized by MS-SSIM compared to recent deep compression models on the benchmark datasets of Kodak and Tecnick. The project page with the source code can be found inhttps://mic.tongji.edu.cn. Zhisen Tang, Hanli Wang, Xiaokai Yi, Yun Zhang 0002, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | AFD-Former: A Hybrid Transformer With Asymmetric Flow Division for Synthesized View Quality EnhancementabstractRecently, CNN-based post-processing has shown great potential in Synthesized View Quality Enhancement (SVQE). However, due to the limited receptive field of convolution, it is ineffective in explicitly modeling long-range dependencies, which are critical to eliminate the distortion induced by Depth Image Based Rendering (DIBR) in synthesized views. Although transformers exhibit tremendous success at learning global contextual information, it is weak at extracting local texture information. To take full advantages of the CNN and transformer, we present a novel U-shaped hybrid transformer with asymmetric flow division to collaboratively capture global-local information for SVQE, termed as AFD-former. Specifically, the AFD-former utilizes the Transformer-CNN Block (TCB) as encoder and decoder, in which several Dynamic Hybrid Attention Blocks (DHABs) are designed to simultaneously model long-range interactions and retain texture details. Then, considering that the deeper layers of the U-shaped network play more roles in capturing global information while shallow layers more in extracting local information, an Asymmetric Flow Division Unit (AFDU) is embedded into each DHAB to assign different contributions of global-local contextual information to the transformer and CNN branches across different layers. Finally, a dynamic learnable modulator is incorporated into two branches to help model effectively feature representation learning. That can be viewed as the dynamic process of adjusting the weight for each channel of the input feature based on contextual cues. Extensive experiments demonstrate that the proposed AFD-former can significantly enhance perceptual quality of synthesized views with similar SVQE speed compared with the related state-of-the-art SVQE methods. The source code will be available athttps://github.com/House-yuyu/AFD-former. Xu Zhang 0044, Nian Cai, Huan Zhang 0008, Yun Zhang 0002, Jianglei Di, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Perceptually Weighted Rate Distortion Optimization for Video-Based Point Cloud CompressionabstractDynamic point cloud is a volumetric visual data representing realistic 3D scenes for virtual reality and augmented reality applications. However, its large data volume has been the bottleneck of data processing, transmission, and storage, which requires effective compression. In this paper, we propose a Perceptually Weighted Rate-Distortion Optimization (PWRDO) scheme for Video-based Point Cloud Compression (V-PCC), which aims to minimize the perceptual distortion of reconstructed point cloud at the given bit rate. Firstly, we propose a general framework of perceptually optimized V-PCC to exploit visual redundancies in point clouds. Secondly, a multi-scale Projection based Point Cloud quality Metric (PPCM) is proposed to measure the perceptual quality of 3D point cloud. The PPCM model comprises 3D-to-2D patch projection, multi-scale structural distortion measurement, and fusion model. Approximations and simplifications of the proposed PPCM are also presented for both V-PCC integration and low complexity. Thirdly, based on the simplified PPCM model, we propose a PWRDO scheme with Lagrange multiplier adaptation, which is incorporated into the V-PCC to enhance the coding efficiency. Experimental results show that the proposed PPCM models can be used as standalone quality metrics, and they are able to achieve higher consistency with the human subjective scores than the state-of-the-art objective visual quality metrics. Also, compared with the latest V-PCC reference model, the proposed PWRDO-based V-PCC scheme achieves an average bit rate reduction of 13.52%, 8.16%, 10.56% and 9.54%, respectively, in terms of four objective visual quality metrics for point clouds. It is significantly superior to the state-of-the-art coding algorithms. The computational complexity of the proposed PWRDO increases by 1.71% and 0.05% on average to the V-PCC encoder and decoder, respectively, which is negligible. The source codes of the PPCM and PWRDO schemes are available at https://github.com/VVCodec/PPCM-PWRDO. Yun Zhang 0002, Keqin Ding, Na Li 0015, Hanli Wang, Xiaoxia Huang 0004, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2023 | Deep Learning-Based Intra Mode Derivation for Versatile Video CodingabstractIn intra coding, Rate Distortion Optimization (RDO) is performed to achieve the optimal intra mode from a pre-defined candidate list. The optimal intra mode is also required to be encoded and transmitted to the decoder side besides the residual signal, where lots of coding bits are consumed. To further improve the performance of intra coding in Versatile Video Coding (VVC) , an intelligent intra mode derivation method is proposed in this paper, termed as Deep Learning based Intra Mode Derivation (DLIMD) . In specific, the process of intra mode derivation is formulated as a multi-class classification task, which aims to skip the module of intra mode signaling for coding bits reduction. The architecture of DLIMD is developed to adapt to different quantization parameter settings and variable coding blocks including non-square ones, where only one single trained model is required. Different from the existing deep learning based classification problems, the hand-crafted features are also fed into intra mode derivation network besides the learned features from feature learning network. To compete with traditional methods, one additional binary flag is utilized in the video codec to indicate the selected scheme with RDO. Extensive experimental results reveal that the proposed method can achieve 2.28%, 1.74%, and 2.18% bit rate reduction on average for Y, U, and V components on the platform of VVC test model, which outperforms the state-of-the-art works. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | End-To-End Depth Map Compression Framework Via Rgb-To-Depth Structure Priors LearningabstractIn this paper, we propose a novel framework to exploit and utilize the shared information inner RGB-D data for efficient depth map compression. Two main codecs, designed based on the existing end-to-end image compression network, are adopted for RGB image compression and enhanced depth image compression with RGB-to-Depth structure prior, respectively. In particular, we propose a Structure Prior Fusion (SPF) module to extract the structure information from both RGB and depth codecs at multi-scale feature levels and fuse the cross-modal feature to generate more efficient structure priors for depth compression. Extensive experiments show that the proposed framework can achieve competitive rate-distortion performance as well as RGB-D task-specific performance at depth map compression compared with the direct compression scheme. Zhuo Chen 0006, Yun Zhang 0002, Xu Wang 0006, Sam Kwong |
ICIP | 4 |
| 2022 | Texture-Aware Spherical Rotation for High Efficiency Omnidirectional Intra Video CodingabstractTo adapt to the existing video coding standards, omnidirectional videos are usually projected from Three-Dimensional (3D) sphere to Two-Dimensional (2D) plane. However, this projection will cause geometrical stretching distortion and boundary discontinuity, which may degrade coding efficiency. In this paper, we present a Spherical Rotation based Omnidirectional Video Coding (SROVC) method, which exploits the textural properties of omnidirectional videos with spherical rotation. Firstly, SROVC framework is presented and Full-traversal Spherical Rotation (FSR) is developed to derive the optimal rotation angle with frame-level Rate Distortion Optimization (RDO). Secondly, to achieve comparable coding gains and lower computational complexity when compared with FSR, a Texture-aware Spherical Rotation (TSR) method is proposed to predict the rotation angle. Finally, to further reduce complexity and maintain coding efficiency, a Group-oriented TSR (G-TSR) approach is presented, in which the group length is statistically determined. Extensive experiments demonstrate that the proposed TSR and G-TSR schemes can achieve bit rate reductions up to 4.38%, 0.94% and 0.89% on average for CubeMap Projection (CMP) based high efficiency omnidirectional video coding. Additionally, the TSR scheme achieves bit rate saving from 0.91% to 1.19% on average under three more CMP-based projection formats, and 1.89% for joint rotation of X, Y, and Z axes. Jinyong Pi, Yun Zhang 0002, Linwei Zhu, Jinzhi Lin, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Learning Based Just Noticeable Difference and Perceptual Quality Prediction Models for Compressed VideoabstractHuman visual system has a limitation of sensitivity in detecting small distortion in an image/video and the minimum perceptual threshold is so called Just Noticeable Difference (JND). JND modelling is challenging since it highly depends on visual contents and perceptual factors are not fully understood. In this paper, we propose deep learning based JND and perceptual quality prediction models, which are able to predict the Satisfied User Ratio (SUR) and Video Wise JND (VWJND) of compressed videos with different resolutions and coding parameters. Firstly, the SUR prediction is modeled as a regression problem that fits deep learning tools. Then, Video Wise Spatial SUR method (VW-SSUR) is proposed to predict the SUR value for compressed video, which mainly considers the spatial distortion. Thirdly, we further propose Video Wise Spatial-Temporal SUR (VW-STSUR) method to improve the SUR prediction accuracy by considering the spatial and temporal information. Two fusion schemes that fuse the spatial and temporal information in quality score level and in feature level, respectively, are investigated. Finally, key factors including key frame and patch selections, cross resolution prediction and complexity are analyzed. Experimental results demonstrate the proposed VW-SSUR method outperforms in both SUR and VWJND prediction as compared with the state-of-the-art schemes. Moreover, the proposed VW-STSUR further improves the accuracy as compared with the VW-SSUR and the conventional JND models, where the mean SUR prediction error is 0.049, and mean VWJND prediction error is 1.69 in quantization parameter and 0.84 dB in peak signal-to-noise ratio. Yun Zhang 0002, Huanhua Liu, You Yang 0002, Xiaoping Fan, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Deep Learning-Based Perceptual Video Quality Enhancement for 3D Synthesized ViewabstractDue to occlusion among views and temporal inconsistency in depth video, spatio-temporal distortion occurs in 3D synthesized video with depth image-based rendering. In this paper, we propose a deep Convolutional Neural Network (CNN)-based synthesized video denoising algorithm to reduce temporal flicker distortion and improve perceptual quality of 3D synthesized video. First, we analyze the spatio-temporal distortion, and model eliminating spatio-temporal distortion as a perceptual video denoising problem. Then, a deep learning-based synthesized video denoising network is proposed, in which a CNN-friendly spatio-temporal loss function is derived from a synthesized video quality metric and integrated with a single image denoising network architecture. Finally, specific schemes, i.e., specific Synthesized Video Denoising Networks (SynVD-Nets), and a general scheme, i.e., General SynVD-Net (GSynVD-Net), based on existing CNN-based denoising models, are developed to handle synthesized video with different distortion levels more effectively. Experimental results show that the proposed SynVD-Net and GSynVD-Net can outperform deep learning-based counterparts and conventional denoising methods, and significantly enhance perceptual quality of 3D synthesized video. Huan Zhang 0008, Yun Zhang 0002, Linwei Zhu, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Joint Source-Channel Decoding of Polar Codes for HEVC-Based Video StreamingabstractUltra High-Definition (UHD) and Virtual Reality (VR) video streaming over 5G networks are emerging, in which High-Efficiency Video Coding (HEVC) is used as source coding to compress videos more efficiently and polar code is used as channel coding to transmit bitstream reliably over an error-prone channel. In this article, a novel Joint Source-Channel Decoding (JSCD) of polar codes for HEVC-based video streaming is presented to improve the streaming reliability and visual quality. Firstly, a Kernel Density Estimation (KDE) fitting approach is proposed to estimate the positions of error channel decoded bits. Secondly, a modified polar decoder called R-SCFlip is designed to improve the channel decoding accuracy. Finally, to combine the KDE estimator and the R-SCFlip decoder together, the JSCD scheme is implemented in an iterative process. Extensive experimental results reveal that, compared to the conventional methods without JSCD, the error data-frame correction ratios are increased. Averagely, 1.07% and 1.11% Frame Error Ratio (FER) improvements have been achieved for Additive White Gaussian Noise (AWGN) and Rayleigh fading channels, respectively. Meanwhile, the qualities of the recovered videos are significantly improved. For the 2D videos, the average Peak Signal-to-Noise Ratio (PSNR) and Structural SIMilarity (SSIM) gains reach 14% and 34%, respectively. For the 360֯ videos, the average improvements in terms of Weighted-to-Spherically-uniform PSNR (WS-PSNR) and Voronoi-based Video Multimethod Assessment Fusion (VI-VMAF) reach 21% and 7%, respectively. Jinzhi Lin, Yun Zhang 0002, Na Li 0015, Hongling Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Circular intra prediction for 360 degree video coding
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Jinyong Pi, Shiqi Wang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Viewport Perception Based Blind Stereoscopic Omnidirectional Image Quality AssessmentabstractCompared with traditional 2D images, stereoscopic omnidirectional images (SOIs) usually have more complex perceptual factors due to the particularities of imaging and display, making the objective quality assessment of SOIs challenging. In this paper, we construct a large and diverse subjective SOIs database named as NBU-SOID for further research demand. And then, we propose a viewport perception based blind SOIs quality assessment (VP-BSOIQA) method by considering the impacts of viewport, user behavior and stereoscopic perception on human visual system, which is mainly composed of binocular perception model (BPM) and omnidirectional perception model (OPM). In the BPM, a binocular combination perception map is generated by the dimension reduction of stereopair and the weighting of binocular energy to reflect the binocular masking effect. In the OPM, several viewports are first created to ensure the consistency of evaluation objects. Then, the intra-viewport and inter-viewport weighting factors are designed with the common influences of visual attention and peripheral vision sensitivity to aggregate the novel multi-orientation structural features extracted from all potential viewports. Experimental results on the NBU-SOID and SOLID databases demonstrate that BPM and OPM can be robustly combined with the existing 2D image quality assessment (IQA) methods, thus averagely achieving 10.2% and 12.2% performance gain in terms of SRCC, respectively. In addition, the proposed VP-BSOIQA method outperforms the state-of-the-art blind IQA methods in predicting the quality of SOIs. Yubin Qi, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Subjective Quality Database and Objective Study of Compressed Point Clouds With 6DoF Head-Mounted DisplayabstractIn this paper, we focus on subjective and objective Point Cloud Quality Assessment (PCQA) in an immersive environment and study the effect of geometry and texture attributes in compression distortion. Using a Head-Mounted Display (HMD) with six degrees of freedom, we establish a subjective PCQA database, named SIAT Point Cloud Quality Database (SIAT-PCQD). Our database consists of 340 distorted point clouds compressed by the MPEG point cloud encoder with the combination of 20 sequences and 17 pairs of geometry and texture quantization parameters. The impact of distorted geometry and texture attributes is further discussed in this paper. Then, we propose two projection-based objective quality evaluation methods, i.e., a weighted view projection based model and a patch projection based model. Our subjective database and findings can be used in point cloud processing, transmission, and coding, especially for virtual reality applications. The subjective datasethttps://dx.doi.org/10.21227/ad8d-7r28http://codec.siat.ac.cn/video_download_siat-pcqd.htmlhasbeen released in the public repository. Xinju Wu, Yun Zhang 0002, Chunling Fan, Junhui Hou, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Deep Learning-Based Chroma Prediction for Intra Versatile Video CodingabstractColor images always exhibit a high correlation between luma and chroma components. Cross component linear model (CCLM) has been introduced to exploit such correlation for removing redundancy in the on-going video coding standard, i.e., versatile video coding (VVC). To further improve the coding performance, this paper presents a deep learning based intra chroma prediction method, termed as convolutional neural network based chroma prediction (CNNCP). More specifically, the process of chroma prediction is formulated to produce the colorful version from available information input. CNNCP includes two sub-networks for luma down-sampling and chroma prediction, which are jointly optimized to fully exploit spatial and cross component information. In addition, the outputs of CCLM are adopted as chroma initialization for performance enhancement, and the coding distortion level characterized by quantization parameter is fed into the network to release the negative affect from compression artifacts. To further improve the coding performance, the competition is performed between the conventional chroma prediction and CNNCP in terms of rate-distortion cost with a binary flag signalled. The learned CNNCP is incorporated into both video encoder and decoder. Extensive experimental results demonstrate that the proposed scheme can achieve 4.283%, 3.343%, and 4.634% bit rate savings for luma and two chroma components, compared with the VVC test model version 4.0 (VTM 4.0). Linwei Zhu, Yun Zhang 0002, Shiqi Wang 0001, Sam Kwong, Xin Jin 0002, Yu Qiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Online Learning-Based Multi-Stage Complexity Control for Live Video CodingabstractHigh Efficiency Video Coding (HEVC) can significantly improve the compression efficiency in comparison with the preceding H.264/Advanced Video Coding (AVC) but at the cost of extremely high computational complexity. Hence, it is challenging to realize live video applications on low-delay and power-constrained devices, such as the smart mobile devices. In this article, we propose an online learning-based multi-stage complexity control method for live video coding. The proposed method consists of three stages: multi-accuracy Coding Unit (CU) decision, multi-stage complexity allocation, and Coding Tree Unit (CTU) level complexity control. Consequently, the encoding complexity can be accurately controlled to correspond with the computing capability of the video-capable device by replacing the traditional brute-force search with the proposed algorithm, which properly determines the optimal CU size. Specifically, the multi-accuracy CU decision model is obtained by an online learning approach to accommodate the different characteristics of input videos. In addition, multi-stage complexity allocation is implemented to reasonably allocate the complexity budgets to each coding level. In order to achieve a good trade-off between complexity control and rate distortion (RD) performance, the CTU-level complexity control is proposed to select the optimal accuracy of the CU decision model. The experimental results show that the proposed algorithm can accurately control the coding complexity from 100% to 40%. Furthermore, the proposed algorithm outperforms the state-of-the-art algorithms in terms of both accuracy of complexity control and RD performance. Chao Huang 0008, Zongju Peng, Yong Xu 0001, Qiuping Jiang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 6 |
| 2021 | Cubemap-Based Perception-Driven Blind Quality Assessment for 360-degree Imagesabstractimage can be represented with different formats, such as the equirectangular projection (ERP) image, viewport images or spherical image, for its different processing procedures and applications. Accordingly, the 360-degree image quality assessment (360-IQA) can be performed on these different formats. However, the performance of 360-IQA with the ERP image is not equivalent with those with the viewport images or spherical image due to the over-sampling and the resulted obvious geometric distortion of ERP image. This imbalance problem brings challenge to ERP image based applications, such as 360-degree image/video compression and assessment. In this paper, we propose a new blind 360-IQA framework to handle this imbalance problem. In the proposed framework, cubemap projection (CMP) with six inter-related faces is used to realize the omnidirectional viewing of 360-degree image. A multi-distortions visual attention quality dataset for 360-degree images is firstly established as the benchmark to analyze the performance of objective 360-IQA methods. Then, the perception-driven blind 360-IQA framework is proposed based on six cubemap faces of CMP for 360-degree image, in which human attention behavior is taken into account to improve the effectiveness of the proposed framework. The cubemap quality feature subset of CMP image is first obtained, and additionally, attention feature matrices and subsets are also calculated to describe the human visual behavior. Experimental results show that the proposed framework achieves superior performances compared with state-of-the-art IQA methods, and the cross dataset validation also verifies the effectiveness of the proposed framework. In addition, the proposed framework can also be combined with new quality feature extraction method to further improve the performance of 360-IQA. All of these demonstrate that the proposed framework is effective in 360-IQA and has a good potential for future applications. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, You Yang 0002, Zongju Peng |
IEEE Trans. Image Process. | 4 |
| 2021 | Highly Efficient Multiview Depth Coding Based on Histogram Projection and Allowable Depth DistortionabstractMismatches between the precisions of representing the disparity, depth value and rendering position in 3D video systems cause redundancies in depth map representations. In this paper, we propose a highly efficient multiview depth coding scheme based on Depth Histogram Projection (DHP) and Allowable Depth Distortion (ADD) in view synthesis. Firstly, DHP exploits the sparse representation of depth maps generated from stereo matching to reduce the residual error from INTER and INTRA predictions in depth coding. We provide a mathematical foundation for DHP-based lossless depth coding by theoretically analyzing its rate-distortion cost. Then, due to the mismatch between depth value and rendering position, there is a many-to-one mapping relationship between them in view synthesis, which induces the ADD model. Based on this ADD model and DHP, depth coding with lossless view synthesis quality is proposed to further improve the compression performance of depth coding while maintaining the same synthesized video quality. Experimental results reveal that the proposed DHP based depth coding can achieve an average bit rate saving of 20.66% to 19.52% for lossless coding on Multiview High Efficiency Video Coding (MV-HEVC) with different groups of pictures. In addition, our depth coding based on DHP and ADD achieves an average depth bit rate reduction of 46.69%, 34.12% and 28.68% for lossless view synthesis quality when the rendering precision varies from integer, half to quarter pixels, respectively. We obtain similar gains for lossless depth coding on the 3D-HEVC, HEVC Intra coding and JPEG2000 platforms. Yun Zhang 0002, Linwei Zhu, Raouf Hamzaoui, Sam Kwong, Yo-Sung Ho |
IEEE Trans. Image Process. | 1 |
| 2020 | Content-aware Hybrid Equi-angular Cubemap Projection for Omnidirectional Video CodingabstractOmnidirectional video is required to be projected from the Three-Dimensional (3D) sphere to a Two-Dimensional (2D) plane before compression due to its spherical characteristics. Therefore, various projection formats have been proposed in recent years. However, these existing projection methods have problems of either oversampling or discontinuous boundary, which penalize the coding performance. Among them, Hybrid Equiangular Cubemap (HEC) projection has achieved significant coding gains by keeping boundary continuity when compared with Equi-Angular Cubemap (EAC) projection. However, the parameters of its mapping function are fixed and cannot adapt to the video contents, which results in non-uniform sampling in certain regions. To address this limitation, a projection method named Content-aware HEC (CHEC) is presented in this paper. In particular, these parameters of mapping function are adaptively achieved by minimizing the projection conversion distortion. Additionally, an omnidirectional video coding framework with adaptive parameters of mapping function is proposed to effectively improve the coding performance. Experimental results show that the proposed scheme achieves 8.57% and 0.11% bit rate reduction on average in terms of End-to-End Weighted to Spherically uniform Peak Signal to Noise Ratio (E2E WS-PSNR) when compared with Equi-Rectangular Projection (ERP) and HEC projections, respectively. Jinyong Pi, Yun Zhang 0002, Linwei Zhu, Xinju Wu, Xuemei Zhou |
VCIP | 2 |
| 2020 | Sparse Representation-Based Intra Prediction for Lossless/Near Lossless Video CodingabstractIn this paper, a novel intra prediction method is presented for lossless/near lossless High Efficiency Video Coding (HEVC), termed as Sparse Representation based Intra Prediction (SRIP). In specific, the existing Angular Intra Prediction (AIP) modes in HEVC are organized as a mode dictionary, which is utilized to sparsely represent the visual signal by minimizing the difference with respect to the ground truth. For the match of encoding and decoding, the sparse coefficients are also required to be encoded and transmitted to the decoder side. To further improve the coding performance, an additional binary flag is included in the video codec to indicate which strategy is finally adopted with the rate distortion optimization, i.e., SRIP or traditional AIP. Extensive experimental results reveal that the proposed method can achieve 0.36% bit rate saving on average in case of lossless scenario. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Jinyong Pi, Xinju Wu |
VCIP | 2 |
| 2020 | Salient object detection via reliability-based depth compactness and depth contrastabstractIt can be intuitively inferred that a high‐quality depth map can be used to quickly detect the salient region in stereo vision, implying that depth information plays an essential role in stereoscopic visual attention. However, existing methods generally use the depth map as an auxiliary cue to improve the saliency detection performance. In this study, the authors present an algorithm to directly detect the salient object from a high‐quality depth image. The proposed algorithm utilises a depth reliability indicator to assess the confidence of a depth image. Depth compactness, a novel feature that incorporates the depth reliability of the super‐pixels, is computed as a primary salient feature. Moreover, in order to enhance another salient feature (i.e. depth contrast), they develop a coarse background filtering method to suppress background interference. Experimental results demonstrate that the proposed method performs favourably against the popular depth‐aware saliency detection approaches at a lower computational cost. Yang Zhou 0052, Yun Zhang 0002, Haibing Yin |
IET Image Process. | 3 |
| 2020 | A novel deep neural network based approach for sparse code multiple access
Jinzhi Lin, Shengzhong Feng, Yun Zhang 0002, Zhile Yang, Yong Zhang 0001 |
Neurocomputing | 3 |
| 2020 | Machine learning based video coding optimizations: A survey
Yun Zhang 0002, Sam Kwong, Shiqi Wang 0001 |
Inf. Sci. | 1 |
| 2020 | Sparse Representation-Based Video Quality Assessment for Synthesized 3D VideosabstractThe temporal flicker distortion is one of the most annoying noises in synthesized virtual view videos when they are rendered by compressed multi-view video plus depth in Three Dimensional (3D) video system. To assess the synthesized view video quality and further optimize the compression techniques in 3D video system, objective video quality assessment which can accurately measure the flicker distortion is highly needed. In this paper, we propose a full reference sparse representation based video quality assessment method towards synthesized 3D videos. Firstly, a synthesized video, treated as a 3D volume data with spatial (X-Y) and temporal (T) domains, is reformed and decomposed as a number of spatially neighboring temporal layers, i.e., X-T or Y-T planes. Gradient features in temporal layers of the synthesized video and strong edges of depth maps are used as key features in detecting the location of flicker distortions. Secondly, dictionary learning and sparse representation for the temporal layers are then derived and applied to effectively represent the temporal flicker distortion. Thirdly, a rank pooling method is used to pool all the temporal layer scores and obtain the score for the flicker distortion. Finally, the temporal flicker distortion measurement is combined with the conventional spatial distortion measurement to assess the quality of synthesized 3D videos. Experimental results on synthesized video quality database demonstrate our proposed method is significantly superior to other state-of-the-art methods, especially on the view synthesis distortions induced from depth videos. Yun Zhang 0002, Huan Zhang 0008, Mei Yu 0001, Sam Kwong, Yo-Sung Ho |
IEEE Trans. Image Process. | 1 |
| 2020 | Deep Learning-Based Picture-Wise Just Noticeable Distortion Prediction Model for Image CompressionabstractPicture Wise Just Noticeable Difference (PW-JND), which accounts for the minimum difference of a picture that human visual system can perceive, can be widely used in perception-oriented image and video processing. However, the conventional Just Noticeable Difference (JND) models calculate the JND threshold for each pixel or sub-band separately, which may not reflect the total masking effect of a picture accurately. In this paper, we propose a deep learning based PW-JND prediction model for image compression. Firstly, we formulate the task of predicting PW-JND as a multi-class classification problem, and propose a framework to transform the multi-class classification problem to a binary classification problem solved by just one binary classifier. Secondly, we construct a deep learning based binary classifier named perceptually lossy/lossless predictor which can predict whether an image is perceptually lossy to another or not. Finally, we propose a sliding window based search strategy to predict PW-JND based on the prediction results of the perceptually lossy/lossless predictor. Experimental results show that the mean accuracy of the perceptually lossy/lossless predictor reaches 92%, and the absolute prediction error of the proposed PW-JND model is 0.79 dB on average, which shows the superiority of the proposed PW-JND model to the conventional JND models. Huanhua Liu, Yun Zhang 0002, Huan Zhang 0008, Chunling Fan, Sam Kwong, C.-C. Jay Kuo, Xiaoping Fan |
IEEE Trans. Image Process. | 2 |
| 2020 | Efficient In-Loop Filtering Based on Enhanced Deep Convolutional Neural Networks for HEVCabstractThe raw video data can be compressed much by the latest video coding standard, high efficiency video coding (HEVC). However, the block-based hybrid coding used in HEVC will incur lots of artifacts in compressed videos, the video quality will be severely influenced. To settle this problem, the in-loop filtering is used in HEVC to eliminate artifacts. Inspired by the success of deep learning, we propose an efficient in-loop filtering algorithm based on the enhanced deep convolutional neural networks (EDCNN) for significantly improving the performance of in-loop filtering in HEVC. Firstly, the problems of traditional convolutional neural networks models, including the normalization method, network learning ability, and loss function, are analyzed. Then, based on the statistical analyses, the EDCNN is proposed for efficiently eliminating the artifacts, which adopts three solutions, including a weighted normalization method, a feature information fusion block, and a precise loss function. Finally, the PSNR enhancement, PSNR smoothness, RD performance, subjective test, and computational complexity/GPU memory consumption are employed as the evaluation criteria, and experimental results show that when compared with the filter in HM16.9, the proposed in-loop filtering algorithm achieves an average of 6.45% BDBR reduction and 0.238 dB BDPSNR gains. Zhaoqing Pan, Xiaokai Yi, Yun Zhang 0002, Byeungwoo Jeon, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2020 | Generative Adversarial Network-Based Intra Prediction for Video CodingabstractIn this paper, a novel intra prediction method is proposed to improve the video coding performance, in which the generative adversarial network (GAN) is adopted to intelligently remove the spatial redundancy with the inference process. The proposed GAN-based method improves the prediction by exploiting more information and generating more flexible prediction patterns. In particular, the intra prediction is modeled as an inpainting task, which is accomplished with the GAN model to fill in the missing part by conditioning on the available reconstructed pixels. As such, the learned GAN model is incorporated into both video encoder and decoder, and the rate-distortion optimization is performed for the competition between GAN-based intra prediction and traditional angular-based intra prediction to achieve better coding performance. The proposed scheme is implemented into the high-efficiency video coding test model (HM 16.17) and the versatile video coding test model (VTM 1.1). The experimental results show that the proposed algorithm can achieve 6.6%, 7.5%, and 7.5% under HM 16.17 and 6.75%, 7.63%, and 7.65% under VTM 1.1 bit rate savings on average for luma and chroma components in the intra coding scenario. Linwei Zhu, Sam Kwong, Yun Zhang 0002, Shiqi Wang 0001, Xu Wang 0006 |
IEEE Trans. Multim. | 3 |
| 2020 | Frame-level Bit Allocation Optimization Based on Video Content Characteristics for HEVCabstractRate control plays an important role in high efficiency video coding (HEVC), and bit allocation is the foundation of rate control. The video content characteristics are significant for bit allocation, and modeling an accurate relationship between video content characteristics and bit allocation is essential for bit allocation optimization. Therefore, in this article, a video content characteristics–based frame-level optimal bit allocation algorithm is proposed for improving the rate distortion (RD) performance of HEVC. First, the number of search points of motion estimation is used to evaluate the motion activity of video content, and the relationship between the search points and bit allocation is modeled as the search-points model. Second, the grey level co-occurrence matrix and temporal perceptual information are used to evaluate the spatial and temporal texture complexity, and the relationship between the video content texture complexity and bit allocation is modeled as the texture-complexity model. Then, the search-points model and texture-complexity model are jointly employed to allocate the coding bits for the second and third layers of the HEVC hierarchical coding structure. Finally, the remaining coding bits of a group-of-pictures (GOP) are allocated to the first layer of HEVC coding structure. To evaluate the performance of the proposed algorithm, the RD performance and bitrate accuracy are used as evaluation criteria, and the experimental results show that when compared with the popularly used R-λ model–based bit allocation algorithm, the proposed algorithm achieves an average of -3.43% BDBR reduction and 0.13 dB BDPSNR gains with only 0.02% loss of bitrate accuracy. Zhaoqing Pan, Xiaokai Yi, Yun Zhang 0002, Hui Yuan 0001, Fu Lee Wang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Interactive Subjective Study on Picture-level Just Noticeable Difference of Compressed Stereoscopic ImagesabstractThe Just Noticeable Difference (JND) reveals the minimum distortion that the Human Visual System (HVS) can perceive. Traditional studies on JND mainly focus on background luminance adaptation and contrast masking. However, the HVS does not perceive visual content based on individual pixels or blocks, but on the entire image. In this work, we conduct an interactive subjective visual quality study on the Picture-level JND (PJND) of compressed stereo images. The study, which involves 48 subjects and 10 stereoscopic images compressed with H.265 intra coding and JPEG2000, includes two parts. In the first part, we determine the minimum distortion that the HVS can perceive against a pristine stereo image. In the second part, we explore the minimum distortion that each subject perceives against a distorted stereo image. Modeling the distribution of the PJND samples as Gaussian, we obtain their complementary cumulative distribution functions, which are known as Satisfied User Ratio (SUR) functions. Statistical analysis results demonstrate that the SUR is highly dependent on the image contents. The HVS is more sensitive to distortion in images with more texture details. The compressed stereoscopic images and the PJND samples are collected in a data set called SIAT-JSSI, which we release to the public. Chunling Fan, Yun Zhang 0002, Raouf Hamzaoui, Qingshan Jiang |
ICASSP | 2 |
| 2019 | SUR-Net: Predicting the Satisfied User Ratio Curve for Image Compression with Deep LearningabstractThe Satisfied User Ratio (SUR) curve for a lossy image compression scheme, e.g., JPEG, characterizes the probability distribution of the Just Noticeable Difference (JND) level, the smallest distortion level that can be perceived by a subject. We propose the first deep learning approach to predict such SUR curves. Instead of the direct approach of regressing the SUR curve itself for a given reference image, our model is trained on pairs of images, original and compressed. Relying on a Siamese Convolutional Neural Network (CNN), feature pooling, a fully connected regression-head, and transfer learning, we achieved a good prediction performance. Experiments on the MCL-JCI dataset showed a mean Bhattacharyya distance between the predicted and the original JND distributions of only 0.072. Chunling Fan, Hanhe Lin, Vlad Hosu, Yun Zhang 0002, Qingshan Jiang, Raouf Hamzaoui, Dietmar Saupe |
QoMEX | 4 |
| 2019 | Picture-level just noticeable difference for symmetrically and asymmetrically compressed stereoscopic images: Subjective quality assessment study and datasets
Chunling Fan, Yun Zhang 0002, Huan Zhang 0008, Raouf Hamzaoui, Qingshan Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Reinforcement learning based coding unit early termination algorithm for high efficiency video coding
Na Li 0015, Yun Zhang 0002, Linwei Zhu, Wenhan Luo, Sam Kwong |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Depth perceptual quality assessment for symmetrically and asymmetrically distorted stereoscopic 3D videos
Yun Zhang 0002, Xiangkai Liu, Huanhua Liu, Chunling Fan |
Signal Process. Image Commun. | 1 |
| 2019 | WLDISR: Weighted Local Sparse Representation-Based Depth Image Super-Resolution for 3D Video SystemabstractIn this paper, we propose a Weighted Local sparse representation based Depth Image Super-Resolution (WLDISR) schemes aiming at improving the Virtual View Image (VVI) quality of 3D video system. Different from color images, depth images are mainly used to provide geometrical information in synthesizing VVI. Due to the view synthesis characteristics difference between textural structures and smooth regions of depth images, we divide the depth images into edge and smooth patches and learn two local dictionaries, respectively. Meanwhile, the weight term is derived and incorporated explicitly in the cost function to denote different importance of edge structures and smooth regions to the VVI quality. Then, local sparse representation and weighted sparse representation are jointly used in both dictionary learning and reconstruction phases in depth image super-resolution. Based on different optimizations on learning and reconstruction modules, three WLDISR schemes, WLDISR-D, WLDISR-R, and WLDISR-ALL, are proposed. Experimental results on 3D sequences demonstrate that the proposed WLDISR-D, WLDISR-R, and WLDISR-ALL schemes can achieve more than 1.9-, 2.03-, and 2.16-dB gains on average, respectively, in terms of the VVIs' quality, as compared with the state-of-the-art schemes. In addition, the visual quality of VVIs is also improved. Huan Zhang 0008, Yun Zhang 0002, Hanli Wang, Yo-Sung Ho, Shengzhong Feng |
IEEE Trans. Image Process. | 2 |
| 2019 | Statistical Early Termination and Early Skip Models for Fast Mode Decision in HEVC INTRA CodingabstractIn this article, statistical Early Termination (ET) and Early Skip (ES) models are proposed for fast Coding Unit (CU) and prediction mode decision in HEVC INTRA coding, in which three categories of ET and ES sub-algorithms are included. First, the CU ranges of the current CU are recursively predicted based on the texture and CU depth of the spatial neighboring CUs. Second, the statistical model based ET and ES schemes are proposed and applied to optimize the CU and INTRA prediction mode decision, in which the coding complexities over different decision layers are jointly minimized subject to acceptable rate-distortion degradation. Third, the mode correlations among the INTRA prediction modes are exploited to early terminate the full rate-distortion optimization in each CU decision layer. Extensive experiments are performed to evaluate the coding performance of each sub-algorithm and the overall algorithm. Experimental results reveal that the overall proposed algorithm can achieve 45.47% to 74.77%, and 58.09% on average complexity reduction, while the overall Bjøntegaard delta bit rate increase and Bjøntegaard delta peak signal-to-noise ratio degradation are 2.29% and −0.11 dB, respectively. Yun Zhang 0002, Na Li 0015, Sam Kwong, Gangyi Jiang, Huanqiang Zeng |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Deep intensity guidance based compression artifacts reduction for depth map
Xu Wang 0006, Yun Zhang 0002, Lin Ma 0002, Sam Kwong, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | No-reference quality assessment of DIBR-synthesized videos by measuring temporal flickering
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yun Zhang 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Image processing for synthesis imaging of mingantu spectral radioheliograph (MUSER)
Long Xu 0001, Yihua Yan, Lin Ma 0002, Yun Zhang 0002 |
Multim. Tools Appl. | 4 |
| 2018 | Effective Data Driven Coding Unit Size Decision Approaches for HEVC INTRA CodingabstractHigh Efficiency Video Coding (HEVC) INTRA coding improves compression efficiency by adopting advanced coding technologies, such as multi-level quad-tree block partitioning and up to 35-mode INTRA prediction. However, it significantly increases the coding complexity, memory access, and power consumption, which goes against its widely applications, especially for ultra-high definition and/or mobile video applications. To tackle this problem, we propose effective data driven coding unit (CU) size decision approaches for HEVC INTRA coding, which consists of two stages of support vector machine-based fast INTRA CU size decision schemes at four CU decision layers. At the first stage classification, a three output classifier with offline learning is developed to early terminate the CU size decision or early skip checking the current CU depth. As for the samples that neither early skipped nor early terminated, the second stage of binary classification, which learns online from previous coded frames, is proposed to further refine the CU size decision. Representative features for the CU size decision are explored at different decision layers and stages of classifications. Finally, the optimal parameters derived from the training data are achieved to reasonably allocate complexity among different CU layers at given total rate-distortion degradation constraint. Extensive experiments show that the proposed overall algorithm can achieve 27.95%–80.53% and 52.48% on average complexity reduction for the CU size decision as compared with the original HM16.7 model. Meanwhile, the average Bjonteggard delta peak-signal-to-noise ratio degradation is only −0.08 dB, which is negligible. The overall performance of the proposed algorithm outperforms the state-of-the-art benchmark schemes. Yun Zhang 0002, Zhaoqing Pan, Na Li 0015, Xu Wang 0006, Gangyi Jiang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Residual Highway Convolutional Neural Networks for in-loop Filtering in HEVCabstractHigh efficiency video coding (HEVC) standard achieves half bit-rate reduction while keeping the same quality compared with AVC. However, it still cannot satisfy the demand of higher quality in real applications, especially at low bit rates. To further improve the quality of reconstructed frame while reducing the bitrates, a residual highway convolutional neural network (RHCNN) is proposed in this paper for in-loop filtering in HEVC. The RHCNN is composed of several residual highway units and convolutional layers. In the highway units, there are some paths that could allow unimpeded information across several layers. Moreover, there also exists one identity skip connection (shortcut) from the beginning to the end, which is followed by one small convolutional layer. Without conflicting with deblocking filter (DF) and sample adaptive offset (SAO) filter in HEVC, RHCNN is employed as a high-dimension filter following DF and SAO to enhance the quality of reconstructed frames. To facilitate the real application, we apply the proposed method to I frame, P frame, and B frame, respectively. For obtaining better performance, the entire quantization parameter (QP) range is divided into several QP bands, where a dedicated RHCNN is trained for each QP band. Furthermore, we adopt a progressive training scheme for the RHCNN where the QP band with lower value is used for early training and their weights are used as initial weights for QP band of higher values in a progressive manner. Experimental results demonstrate that the proposed method is able to not only raise the PSNR of reconstructed frame but also prominently reduce the bit-rate compared with HEVC reference software. Yongbing Zhang 0002, Xiangyang Ji, Yun Zhang 0002, Ruiqin Xiong, Qionghai Dai |
IEEE Trans. Image Process. | 4 |
| 2018 | Convolutional Neural Network-Based Synthesized View Quality Enhancement for 3D Video CodingabstractThe quality of synthesized view plays an important role in the three dimensional (3D) video system. In this paper, to further improve the coding efficiency, a convolutional neural network (CNN) based synthesized view quality enhancement method for 3D High Efficiency Video Coding (HEVC) is proposed. Firstly, the distortion elimination in synthesized view is formulated as an image restoration task with the aim to reconstruct the latent distortion free synthesized image. Secondly, the learned CNN models are incorporated into 3D HEVC codec to improve the view synthesis performance for both view synthesis optimization (VSO) and the final synthesized view, where the geometric and compression distortions are considered according to the specific characteristics of synthesized view. Thirdly, a new Lagrange multiplier in the rate-distortion (RD) cost function is derived to adapt the CNN based VSO process to embrace a better 3D video coding performance. Extensive experimental results show that the proposed scheme can efficiently eliminate the artifacts in the synthesized image, and reduce 25.9% and 11.7% bit rate in terms of peak-signal-to-noise ratio (PSNR) and structural similarity (SSIM) index, which significantly outperforms the state-of-theart methods. Linwei Zhu, Yun Zhang 0002, Shiqi Wang 0001, Hui Yuan 0001, Sam Kwong, Horace Ho-Shing Ip |
IEEE Trans. Image Process. | 2 |
| 2017 | Study of subjective and objective quality assessment for screen content imagesabstractIn this paper, we present the results of a recent large-scale subjective study of image quality on a collection of screen contents distorted by a variety of application-relevant processes. With the development of multi-device interactive multimedia applications, metrics to predict the visual quality of screen content images (SCIs) as perceived by subjects are becoming fundamentally important. For developing the objective image quality assessment (IQA) method, there is a need for large-scale public database with diversity of distorted types and scene contents, and available subjective scores of distorted SCIs. The resulting Immersive Media Laboratory screen content image quality database (IML-SCIQD) contains 1250 distorted SCIs from 25 reference SCIs with 10 distortion types. Each image was rated by 35 human observers, and the different mean opinion scores (DMOS) were obtained after data processing. The performance comparison of 17 state-of-the-arts, publicly available IQA algorithms are evaluated on the new database. The database will be available online in our project website. Xu Wang 0006, Yingying Zhu 0001, Yun Zhang 0002, Jianmin Jiang, Sam Kwong |
ICIP | 4 |
| 2017 | Multi-class ranking based most probable prediction unit selection for HEVC encodingabstractIn this paper, an incremental learning based multi-class Prediction Units (PUs) ranking approach is presented for High Efficiency Video Coding (HEVC) Rate-Distortion-Complexity (RDC) optimization. In particular, the process of PUs selection is formulated as a binary classification plus multi-class ranking task, and incremental learning is applied for classifier training to better exploit the information in the emerging training data. Furthermore, the proposed most probable PUs selection scheme is incorporated into a joint RDC optimization framework, where the complexity can be flexibly allocated targeting at minimizing computational cost under a constrained RD performance degradation. Experimental results demonstrate that the proposed approach can reduce 53.7% and 50.4% computational complexity on average under low delay P and random access configurations with ignorable RD performance degradation, which outperforms the state-of-the-art approaches in terms of RDC performance. Linwei Zhu, Sam Kwong, Yun Zhang 0002, Xu Wang 0006, Shiqi Wang 0001 |
VCIP | 3 |
| 2017 | Instant coherent group motion filtering by group motion representations
Na Li 0015, Yun Zhang 0002, Wenhan Luo |
Neurocomputing | 2 |
| 2017 | Stereoscopic image quality assessment by learning non-negative matrix factorization-based color visual characteristics and considering binocular interactions
Gangyi Jiang, Haiyong Xu, Mei Yu 0001, Ting Luo 0001, Yun Zhang 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | VideoSet: A large-scale compressed video quality dataset based on JND measurementabstract• A large-scale JND-based coded video quality dataset is presented. • The VideoSet contains 220 5-s sequences in four resolutions coded by H.264/AVC. • The subjective test procedure, JND data cleaning and properties are described. • The significance and implications of the VideoSet are discussed. • This work points out a clear path to data-driven perceptual coding. A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed in Lin et al. (2015). Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California in Jin et al. (2016) and Wang et al. (2016) [3]. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-s sequences in four resolutions (i.e., 1920 × 1080 , 1280 × 720 , 960 × 540 and 640 × 360 ). For each of the 880 video clips, we encode it using the H.264/AVC codec with QP = 1 , … , 51 and measure the first three JND points with 30 + subjects. The dataset is called the “VideoSet”, which is an acronym for “Video Subject Evaluation Test (SET)”. This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort (Wang et al., 2016 [4]). Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeong-Hoon Park, Shawmin Lei, Xin Zhou 0001, Man-On Pun, Xin Jin 0002, Ronggang Wang, Xu Wang 0006, Yun Zhang 0002, Jiwu Huang, Sam Kwong, C.-C. Jay Kuo |
J. Vis. Commun. Image Represent. | 11 |
| 2017 | Allowable depth distortion based fast mode decision and reference frame selection for 3D depth coding
Yun Zhang 0002, Zhaoqing Pan, Yang Zhou 0052, Linwei Zhu |
Multim. Tools Appl. | 1 |
| 2017 | Visual comfort prediction for stereoscopic image using stereoscopic visual saliency
Yang Zhou 0052, Yongjian He, Yun Zhang 0002 |
Multim. Tools Appl. | 4 |
| 2017 | Objective Video Quality Assessment Based on Perceptually Weighted Mean Squared ErrorabstractObject quality assessment for compressed video is critical to various video compression systems that are essential in the video delivery and storage. Although mean squared error (MSE) is computationally simple, it may not be accurate to reflect the perceptual quality of compressed videos, which are also affected dramatically by the characteristics of the human visual system (HVS), such as contrast sensitivity, visual attention, and masking effect. In this paper, a video quality metric is proposed based on perceptually weighted MSE. A low-pass filter is designed to model the contrast sensitivity of the HVS with the consideration of visual attention. The imperceptible distortion is adaptively removed in the salient and nonsalient regions. To quantitatively measure the masking effect, the randomness of video content is proposed in both the spatial and temporal domains. Since the masking effect highly depends on the regularity of structure and motion in the spatial and temporal directions, the video signal is modeled as a linear dynamic system, and the prediction error of future frames from previous frames is used as randomness to measure the significance of masking. The relation is investigated between MSE and perceptual quality scores across various contents, and a masking modulation model is proposed to compensate the impact of the masking effect on the MSE. The performance of the proposed quality metric is validated on three video databases with various compression distortions. The experimental results demonstrate that the proposed algorithm outperforms other benchmark quality metrics. Sudeng Hu, Lina Jin, Hanli Wang, Yun Zhang 0002, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Allowable depth distortion based depth filtering for 3D high efficiency video codingabstractDepth videos shall be efficiently compressed and transmitted to the client for view synthesis in Three-Dimensional (3D) video system. Since depth video may contain noise that reduce the coding efficiency, we propose a depth filtering algorithm for 3D depth coding, which exploits the Allowable Depth Distortion (ADD) in view synthesis and is able to improve the coding performance of the depth encoder. Firstly, the depth values has the same rendering position based on the ADD model are clustered. Then, the clustered depth are filtered and set to the optimal depth value for each group by minimizing the view synthesis error. The filtered depth videos are smoother and can be more effectively compressed by the existing 3D High Efficiency Video Coding (HEVC) depth encoder. Experimental results show that the proposed depth filtering method can assist the depth encoder achieve 5.87% bit rate reduction in terms of Bjonteggard Delta Bit Rate (BDBR) and 0.25dB quality gain in terms of Bjonteggard Delta Peak-Signal-to-Noise Ratio (BDPSNR) on average as compared with that of coding the original depth maps. Yun Zhang 0002, Linwei Zhu, Xiangkai Liu, Gangyi Jiang |
ISCAS | 1 |
| 2016 | Novel visibility threshold model for asymmetrically distorted stereoscopic imagesabstractExisting perceptual researches on stereoscopic images mainly focus on the threshold of whole image distortion, rather than the effect of texture feature on the so-called threshold of just-noticeable distortion. Obviously, it is unreasonable to use a single unified perception threshold for natural stereoscopic images as the texture complexity typically varies in different blocks of natural images. To solve this problem, we generated an asymmetrically distorted stereoscopic image database with different texture densities and conducted a large number of subjective experiments. A strong correlation between the asymmetrical visibility threshold and texture complexity was revealed from the subjective experiments. Finally, a nonlinear fitting model was designed to uncover this relationship, which can be applied to asymmetrical coding to control the perceived quality of stereoscopic images. Baozhen Du, Mei Yu 0001, Gangyi Jiang, Yun Zhang 0002, Feng Shao 0001, Zongju Peng, Tianzhi Zhu |
VCIP | 4 |
| 2016 | Content adaptive directional transform for high efficiency video codingabstractHEVC is an emerging new standard for digital video compression, which is regarded as a successor to H.264/AVC standard. It still belongs to block-based hybrid video coding framework. The block patterns range from 4×4 to 64×64 blocks, and DCT is extended from 4×4 to 32×32. 2D-DCT for image is performed along the vertical and horizontal directions, so it is good at the energy compaction of residual block with vertical or horizontal edges. However, the edges are usually neither horizontal nor vertical for most cases, such as neither vertical nor horizontal intra prediction, so directional transform was explored in the past several years. In this paper, a directional transform adaptive to image content is proposed. Firstly, the prediction residual blocks are collected from coding a number of video sequences with plenty of image content. Secondly, for each kind of image content, the residual blocks are clustered to form the given number of clusters. Thirdly, each cluster contributes a transform basis after Singular Value Decomposition (SVD). The experimental results in terms of PSNR gains demonstrate the efficiency of the proposed algorithm with the comparison with the standard HM software. Long Xu 0001, Lin Ma 0002, Yun Zhang 0002, Yihua Yan |
VCIP | 3 |
| 2016 | Early DIRECT mode decision based on all-zero block and rate distortion cost for multiview video codingabstractThe exhaustive variable‐block‐size mode decision can efficiently remove the redundancies among the multiview videos, while it also leads to significant increase of computational complexity in the multiview video coding (MVC) encoder, and the high encoding complexity becomes a bottleneck for the MVC encoder to achieve real‐time multimedia applications. To address this bottleneck, many fast mode decision methods have been proposed. However, most of them are only suitable for optimising the encoding complexity of the odd views of the MVC encoder. In this study, based on the property of the all‐zero block and rate distortion (RD) cost of the DIRECT mode as well as the correlations between the current macroblock (MB) and its spatial–temporal nearby MBs, an early DIRECT mode decision method is proposed for reducing the encoding complexity of the MVC. Experimental results show that the proposed method achieves 48.25 and 55.64% on average encoding time saving for the even and odd views, respectively, whereas the RD performance degradation is quite acceptable. In summary, the proposed method efficiently reduces the encoding complexity for the MVC encoder. Zhaoqing Pan, Yun Zhang 0002, Jianjun Lei 0001, Long Xu 0001, Xingming Sun |
IET Image Process. | 2 |
| 2016 | Fast reference frame selection based on content similarity for low complexity HEVC encoder
Zhaoqing Pan, Jianjun Lei 0001, Yun Zhang 0002, Xingming Sun, Sam Kwong |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Machine learning based fast H.264/AVC to HEVC transcoding exploiting block partition similarity
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | High-Efficiency 3D Depth Coding Based on Perceptual Quality of Synthesized VideoabstractIn 3D video systems, imperfect depth images often induce annoying temporal noise, e.g., flickering, to the synthesized video. However, the quality of synthesized view is usually measured with peak signal-to-noise ratio or mean squared error, which mainly focuses on pixelwise frame-by-frame distortion regardless of the obvious temporal artifacts. In this paper, a novel full reference synthesized video quality metric (SVQM) is proposed to measure the perceptual quality of the synthesized video in 3D video systems. Based on the proposed SVQM, an improved rate-distortion optimization (RDO) algorithm is developed with the target of minimizing the perceptual distortion of synthesized view at given bit rate. Then, the improved RDO algorithm is incorporated into the 3D High Efficiency Video Coding (3D-HEVC) software to improve the 3D depth video coding efficiency. Experimental results show that the proposed SVQM metric has better consistency with human perception on evaluating the synthesized view compared with the state-of-the-art image/video quality assessment algorithms. Meanwhile, this SVQM metric maintains low complexity and easy integration to the current video codec. In addition, the proposed SVQM-based depth coding scheme can achieve approximately 15.27% and 17.63% overall bit rate reduction or 0.42- and 0.46-dB gain in terms of SVQM quality score on average as compared with the latest 3D-HEVC reference model and the state-of-the-art depth coding algorithm, respectively. Yun Zhang 0002, Xiaoxiang Yang, Xiangkai Liu, Yongbing Zhang 0002, Gangyi Jiang, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2015 | Multi-task rank learning for image quality assessmentabstractIn practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each distortion type by using single-task learning, which lead to the poor generalization ability of the models as applied to practical image processing. There are often the underlying cross relatedness amongst these single-task learnings in IQA, which is ignored by the previous approaches. To solve this problem, we propose a multi-task learning framework to train IQA models simultaneously across individual tasks each of which concerns one distortion type. These relatedness can be therefore exploited to improve the generalization ability of IQA models from single-task learning. In addition, pairwise image quality rank instead of image quality rating is optimized in learning task. By mapping image quality rank to image quality rating, a novel no-reference (NR) IQA approach can be derived. The experimental results confirm that the proposed Multi-task Rank Learning based IQA (MRLIQ) approach is prominent among all state-of-the-art NR-IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yun Zhang 0002, Yihua Yan |
ICASSP | 7 |
| 2015 | Fast Transform Unit Depth Decision Based on Quantized Coefficients for HEVCabstractThe quad tree structure based Transform Unit (TU) helps high efficiency video coding to improve the coding efficiency. However, the achieved coding efficiency comes at the cost of the increased computational complexity. In this paper, based on the quantizated coefficients of the TU, we propose an early termination for the quad tree structure based TU encoding process. If the quantized coefficients of the luminance components are all zeros, the TU encoding process will be terminated. Experimental results show that the proposed method achieves about 55.13% on average TU encoding time saving, while the rate distortion performance degradation is negligible. Zhaoqing Pan, Jianjun Lei 0001, Yun Zhang 0002, Sam Kwong |
SMC | 3 |
| 2015 | Smooth View Quality Oriented Bit Allocation Optimization for 3D Video CodingabstractView level bit allocation is an fundamental optimization problem in multiview video plus depth (MVD) based 3D video coding (3DVC). In this paper, we propose a smooth view quality oriented view level bit allocation framework for MVD based 3DVC. The Cauchy-density based rate-distortion model of the texture video and depth map are employed to represent the rate distortion properties. The relationship between the distortion of synthesized view and quantization step size of texture videos and depth maps is approximately fitted as linear model. Final, the bit allocation problem is solved by convex optimization algorithms. Experimental results demonstrated that our proposed algorithm can achieve good performance with acceptable computational complexity comparing to the full search scheme. Xu Wang 0006, Sam Kwong, Wei Gao 0003, Yu Zhou 0027, Hui Yuan 0001, Yun Zhang 0002 |
SMC | 6 |
| 2015 | Binocular vision based objective quality assessment method for stereoscopic images
Gangyi Jiang, Junming Zhou, Mei Yu 0001, Yun Zhang 0002, Feng Shao 0001, Zongju Peng |
Multim. Tools Appl. | 4 |
| 2015 | View synthesis distortion elimination filter for depth video coding in 3D video broadcasting
Linwei Zhu, Yun Zhang 0002, Xu Wang 0006, Sam Kwong |
Multim. Tools Appl. | 2 |
| 2015 | View synthesis distortion model based frame level rate control optimization for multiview depth video coding
Xu Wang 0006, Sam Kwong, Hui Yuan 0001, Yun Zhang 0002, Zhaoqing Pan |
Signal Process. | 4 |
| 2015 | Low Complexity HEVC INTRA Coding for High-Quality Mobile Video CommunicationabstractINTRA video coding is essential for high quality mobile video communication and industrial video applications since it enhances video quality, prevents error propagation, and facilitates random access. The latest high-efficiency video coding (HEVC) standard has adopted flexible quad-tree-based block structure and complex angular INTRA prediction to improve the coding efficiency. However, these technologies increase the coding complexity significantly, which consumes large hardware resources, computing time and power cost, and is an obstacle for real-time video applications. To reduce the coding complexity and save power cost, we propose a fast INTRA coding unit (CU) depth decision method based on statistical modeling and correlation analyses. First, we analyze the spatial CU depth correlation with different textures and present effective strategies to predict the most probable depth range based on the spatial correlation among CUs. Since the spatial correlation may fail for image boundary and transitional areas between textural and smooth areas, we then present a statistical model-based CU decision approach in which adaptive early termination thresholds are determined and updated based on the rate-distortion (RD) cost distribution, video content, and quantization parameters (QPs). Experimental results show that the proposed method can reduce the complexity by about 56.76% and 55.61% on average for various sequences and configurations; meanwhile, the RD degradation is negligible. Yun Zhang 0002, Sam Kwong, Zhaoqing Pan, Hui Yuan 0001, Gangyi Jiang |
IEEE Trans. Ind. Informatics | 1 |
| 2015 | Compressed Image Quality Metric Based on Perceptually Weighted DistortionabstractObjective quality assessment for compressed images is critical to various image compression systems that are essential in image delivery and storage. Although the mean squared error (MSE) is computationally simple, it may not be accurate to reflect the perceptual quality of compressed images, which is also affected dramatically by the characteristics of human visual system (HVS), such as masking effect. In this paper, an image quality metric (IQM) is proposed based on perceptually weighted distortion in terms of the MSE. To capture the characteristics of HVS, a randomness map is proposed to measure the masking effect and a preprocessing scheme is proposed to simulate the processing that occurs in the initial part of HVS. Since the masking effect highly depends on the structural randomness, the prediction error from neighborhood with a statistical model is used to measure the significance of masking. Meanwhile, the imperceptible signal with high frequency could be removed by preprocessing with low-pass filters. The relation is investigated between the distortions before and after masking effect, and a masking modulation model is proposed to simulate the masking effect after preprocessing. The performance of the proposed IQM is validated on six image databases with various compression distortions. The experimental results show that the proposed algorithm outperforms other benchmark IQMs. Sudeng Hu, Lina Jin, Hanli Wang, Yun Zhang 0002, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 4 |
| 2015 | Subjective and Objective Video Quality Assessment of 3D Synthesized Views With Texture/Depth Compression DistortionabstractThe quality assessment for synthesized video with texture/depth compression distortion is important for the design, optimization, and evaluation of the multi-view video plus depth (MVD)-based 3D video system. In this paper, the subjective and objective studies for synthesized view assessment are both conducted. First, a synthesized video quality database with texture/depth compression distortion is presented with subjective scores given by 56 subjects. The 140 videos are synthesized from ten MVD sequences with different texture/depth quantization combinations. Second, a full reference objective video quality assessment (VQA) method is proposed concerning about the annoying temporal flicker distortion and the change of spatio-temporal activity in the synthesized video. The proposed VQA algorithm has a good performance evaluated on the entire synthesized video quality database, and is particularly prominent on the subsets which have significant temporal flicker distortion induced by depth compression and view synthesis process. Xiangkai Liu, Yun Zhang 0002, Sudeng Hu, Sam Kwong, C.-C. Jay Kuo, Qiang Peng |
IEEE Trans. Image Process. | 2 |
| 2015 | Machine Learning-Based Coding Unit Depth Decisions for Flexible Complexity Allocation in High Efficiency Video CodingabstractIn this paper, we propose a machine learning-based fast coding unit (CU) depth decision method for High Efficiency Video Coding (HEVC), which optimizes the complexity allocation at CU level with given rate-distortion (RD) cost constraints. First, we analyze quad-tree CU depth decision process in HEVC and model it as a three-level of hierarchical binary decision problem. Second, a flexible CU depth decision structure is presented, which allows the performances of each CU depth decision be smoothly transferred between the coding complexity and RD performance. Then, a three-output joint classifier consists of multiple binary classifiers with different parameters is designed to control the risk of false prediction. Finally, a sophisticated RD-complexity model is derived to determine the optimal parameters for the joint classifier, which is capable of minimizing the complexity in each CU depth at given RD degradation constraints. Comparative experiments over various sequences show that the proposed CU depth decision algorithm can reduce the computational complexity from 28.82% to 70.93%, and 51.45% on average when compared with the original HEVC test model. The Bjøntegaard delta peak signal-to-noise ratio and Bjøntegaard delta bit rate are -0.061 dB and 1.98% on average, which is negligible. The overall performance of the proposed algorithm outperforms those of the state-of-the-art schemes. Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Hui Yuan 0001, Zhaoqing Pan, Long Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Rate Distortion Optimized Inter-View Frame Level Bit Allocation Method for MV-HEVCabstractIn multi-view video coding, since inter-view prediction has been adopted as an important coding tool which could improve coding efficiency greatly, inter-view dependency is inevitable, i.e., the distortion of the reference view (RV) picture could be propagated to the non-reference view (NRV) pictures . Therefore, in order to achieve higher coding efficiency , the inter-view dependency must be taken into account for inter-view bit allocation. In this paper, the inter-view dependency is analyzed in detail, and a rate-distortion (RD) model for NRVs is derived by taking the distortion of RV into account. Based on the derived RD model, the inter-view bit allocation is represented as a mathematical problem with an analytic form, and is solved by a convex optimization (Lagrangian Multiplier) method. Experimental results demonstrate that the RD performance and the inter-view quality consistency of the proposed method is better than existing methods, while the complexity of the proposed method is comparable with the existing methods. Hui Yuan 0001, Sam Kwong, Xu Wang 0006, Wei Gao 0003, Yun Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2014 | Fast Coding Tree Unit depth decision for high efficiency video codingabstractHigh Efficiency Video Coding (HEVC) is the latest video coding standard, which adapts quadtree structure based Coding Tree Unit (CTU) to improve the coding efficiency. In HEVC encoding process, the CTU is recursively partitioned into coding units according to the quadtree depth. This technique increases the coding efficiency of HEVC, however, the achieved coding efficiency comes at the cost of high computational complexity. In this paper, we propose a fast C-TU quadtree depth decision algorithm to reduce the computational complexity of HEVC. Firstly, based on the best C-TU depth correlation among spatial and temporal neighboring CTUs, an early quadtree depth 0 decision algorithm is proposed. Then, according to the correlation between the prediction unit mode and the best CTU depth selection, a quadtree depth 3 skipped decision algorithm is proposed. Experimental results show that the proposed algorithm can achieve 40% on average encoding time saving, while maintaining a comparable rate-distortion performance. Zhaoqing Pan, Sam Kwong, Yun Zhang 0002, Jianjun Lei 0001, Hui Yuan 0001 |
ICIP | 3 |
| 2014 | Global and local exploitation for saliency using bag-of-wordsabstractThe guidance of attention helps human vision system to detect objects rapidly. In this study, the authors present a new saliency detection algorithm by using bag‐of‐words (BOW) representation. The authors regard salient regions as coming from globally rare features and regions locally differ from their surroundings. Our approach consists of three stages: first, calculate global rarity of visual words. A vocabulary, a group of visual words, is generated from the given image and a rarity factor for each visual word is introduced according to its occurrence. Second, calculate local contrast. Representations of local patch are achieved from the histograms of words. Then, local contrast is computed by the difference between the two BOW histograms of a patch and its surroundings. Finally, saliency is measured by the combination of global rarity and local patch contrast. We compare our model with the previous methods on natural images, and experimental results demonstrate good performance of our model and fair consistency with human eye fixations. Zhenzhu Zheng, Yun Zhang 0002, Luxin Yan |
IET Comput. Vis. | 2 |
| 2014 | Generalized Nash Bargaining Solution to Rate Control Optimization for Spatial Scalable Video CodingabstractRate control (RC) optimization is indispensable for scalable video coding (SVC) with respect to bitstream storage and video streaming usage. From the perspective of centralized resource allocation optimization, the inner-layer bit allocation problem is similar to the bargaining problem. Therefore, bargaining game theory can be employed to improve the RC performance for spatial SVC. In this paper, we propose a bargaining game based one-pass RC scheme for spatial H.264/SVC. In each spatial layer (SL), the encoding constraints, such as bit rates, buffer size are jointly modeled as resources in the inner-layer bit allocation bargaining game. The modified rate-distortion (R-D) model incorporated with the inter-layer coding information is investigated. Then the generalized Nash bargaining solution (NBS) is employed to achieve an optimal bit allocation solution. The bandwidth is allocated to the frames from the generalized NBS adaptively based on their own bargaining powers. Experimental results demonstrate that the proposed rate control algorithm achieves appealing image quality improvement and buffer smoothness. The average mismatch of our proposed algorithm is within the range of 0:19%2:63%. Xu Wang 0006, Sam Kwong, Long Xu 0001, Yun Zhang 0002 |
IEEE Trans. Image Process. | 4 |
| 2014 | Efficient Multiview Depth Coding Optimization Based on Allowable Depth Distortion in View SynthesisabstractDepth video is used as the geometrical information of 3D world scenes in 3D view synthesis. Due to the mismatch between the number of depth levels and disparity levels in the view synthesis, the relationship between depth distortion and rendering position error can be modeled as a many-to-one mapping function, in which different depth distortion values might be projected to the same geometrical distortion in the synthesized virtual view image. Based on this property, we present an allowable depth distortion (ADD) model for 3D depth map coding. Then, an ADD-based rate-distortion model is proposed for mode decision and motion/disparity estimation modules aiming at minimizing view synthesis distortion at a given bit rate constraint. In addition, an ADD-based depth bit reduction algorithm is proposed to further reduce the depth bit rate while maintaining the qualities of the synthesized images. Experimental results in intra depth coding show that the proposed overall algorithm achieves Bjontegaard delta peak signal-to-noise ratio gains of 1.58 and 2.68 dB on average for half and integer-pixel rendering precisions, respectively. In addition, the proposed algorithms are also highly efficient for inter depth coding when evaluated with different metrics. Yun Zhang 0002, Sam Kwong, Sudeng Hu, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2013 | Early termination for TZSearch in HEVC Motion EstimationabstractThe TZSearch algorithm was adopted in the high efficiency video coding reference software HM as a fast Motion Estimation (ME) algorithm for its excellent performance in reducing ME time and maintaining a comparable Rate Distortion (RD) performance. However, the multiple initial search point decision and the hybrid block matching search contribute a relatively high computational complexity to TZSearch. In this paper, based on the statistical analysis of the probability of median predictor to be selected as the final best point in the large Coding Units (CUs) (64×64, 32×32) and small CUs (16×16, 8×8) as well as the center-biased characteristic of the final best search point in ME process, we propose two early terminations for TZSearch. Experimental results show that the proposed early terminations can achieve 38.96% encoding time saving, while the RD performance degradation is quite acceptable. Zhaoqing Pan, Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Long Xu 0001 |
ICASSP | 2 |
| 2013 | View-spatial-temporal post-refinement for view synthesis in 3D video systems
Linwei Zhu, Yun Zhang 0002, Mei Yu 0001, Gangyi Jiang, Sam Kwong |
Signal Process. Image Commun. | 2 |
| 2013 | Rate-Distortion Optimized Rate Control for Depth Map-Based 3-D Video CodingabstractIn this paper, a novel rate control scheme with optimized bits allocation for the 3-D video coding is proposed. First, we investigate the R-D characteristics of the texture and depth map of the coded view, as well as the quality dependency between the virtual view and the coded view. Second, an optimal bit allocation scheme is developed to allocate target bits for both the texture and depth maps of different views. Meanwhile, a simplified model parameter estimation scheme is adopted to speed up the coding process. Finally, the experimental results on various 3-D video sequences demonstrate that the proposed algorithm achieves excellent R-D efficiency and bit rate accuracy compared to benchmark algorithms. Sudeng Hu, Sam Kwong, Yun Zhang 0002, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 3 |
| 2013 | Regional Bit Allocation and Rate Distortion Optimization for Multiview Depth Video Coding With View Synthesis Distortion ModelabstractIn this paper, we propose a view synthesis distortion model (VSDM) that establishes the relationship between depth distortion and view synthesis distortion for the regions with different characteristics: color texture area corresponding depth (CTAD) region and color smooth area corresponding depth (CSAD), respectively. With this VSDM, we propose regional bit allocation (RBA) and rate distortion optimization (RDO) algorithms for multiview depth video coding (MDVC) by allocating more bits on CTAD for rendering quality and fewer bits on CSAD for compression efficiency. Experimental results show that the proposed VSDM based RBA and RDO can improve the coding efficiency significantly for the test sequences. In addition, for the proposed overall MDVC algorithm that integrates VSDM based RBA and RDO, it achieves 9.99% and 14.51% bit rate reduction on average for the high and low bit rate, respectively. It can improve virtual view image quality 0.22 and 0.24 dB on average at the high and low bit rate, respectively, when compared with the original joint multiview video coding model. The RD performance comparisons using five different metrics also validate the effectiveness of the proposed overall algorithm. In addition, the proposed algorithms can be applied to both INTRA and INTER frames. Yun Zhang 0002, Sam Kwong, Long Xu 0001, Sudeng Hu, Gangyi Jiang, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2012 | A Universal Rate Control Scheme for Video TranscodingabstractVideo transcoding is proposed for the bitrate adaption, spatial and/or temporal resolutions adaption, and video format conversion. In video streaming application, it converts videos at server to the compatible versions demanded by networks or clients' devices, so that the videos can be delivered over networks and displayed in the clients' devices successfully. This paper provides a universal rate control scheme for various video transcoding purposes. First, a new rate-distortion (R-D) model is established theoretically for better representing the real R-D feature of transcoding. Second, a window-level rate control algorithm is proposed for providing smooth visual quality with compliant buffer constraint by utilizing the two-pass R-D model and a new proposed sliding window buffer control strategy. Finally, a universal rate control scheme for transcoding is developed based on the established R-D model and the proposed window-level rate control algorithm. The extensive experimental results demonstrate that as compared to other state-of-the-art rate control algorithms for transcoding, the proposed scheme can achieve more bit control accuracy with the average mismatch below 0.2%, and much more consistent visual quality with 0.1 dB-0.3 dB peak-to-signal noise ratio improvement in average, while with low computational complexity. Long Xu 0001, Sam Kwong, Hanli Wang, Yun Zhang 0002, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Considering binocular spatial sensitivity in stereoscopic image quality assessmentabstractDeveloping reliable and generic perceptual quality metrics is an challenging issue in three-dimensional (3D) visual signals processing, although many two dimensional (2D) image quality metrics have been proposed and work well on 2D images. In this paper, the binocular spatial sensitivity influenced by the binocular fusion and rivalry properties is considered in the quality measurement. Firstly, the binocular spatial sensitivity map is modeled to reflect the properties. Then, a framework of integration of binocular spatial sensitivity map into quality assessment is presented. Experimental results show that the proposed metric correlate well with human perception of quality on a dataset of 3D images and human subjective scores. Xu Wang 0006, Sam Kwong, Yun Zhang 0002 |
VCIP | 3 |
| 2011 | Subjective quality analyses of stereoscopic images in 3DTV systemabstractSubjective quality evaluation is the basis of quality evaluation of stereoscopic images. As the lack of a public and diverse testing database currently, in this paper, a symmetric stereoscopic images database is built. And then the subjective quality of stereoscopic images is analyzed from two aspects, one is the effects of JPEG, JPEG2000, H.264. The other is the comparisons between symmetric and asymmetric stereoscopic images from Gaussian blurring, white Gaussian noise, JPEG and JPEG2000, respectively. The results show three compressions are quite different in the subjective quality of symmetric stereoscopic images at different bitrates, and the comparisons between symmetric and asymmetric stereoscopic images investigate the properties of binocular fusion, binocular suppression, and binocular summation. Junming Zhou, Gangyi Jiang, Xiangying Mao, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Yun Zhang 0002 |
VCIP | 7 |
| 2010 | Stereoscopic Visual Attention Model for 3D Video
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003 |
MMM | 1 |
| 2010 | Depth perceptual region-of-interest based multiview video coding
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, You Yang 0002, Zongju Peng, Ken Chen 0003 |
J. Vis. Commun. Image Represent. | 1 |