VLDB 2026 Research / reviewers in the wild / expert
Zhaoqing Pan
dblp:127/0937
· DBLP profile ↗
46ranked-venue papers
18as first author
24since 2021 · last 2026
0000-0003-1390-399XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 16 first-author · 22 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Computer networks · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mining Temporal Redundancy Using Long Short-Term Motion Aggregation and Global-Local Decorrelation for Learned Video CompressionabstractThe conditional coding paradigm is widely used in learned video compression, which shows superior performance in capturing redundancies within a large context space. However, existing Conditional coding-based Learned Video Compression (C-LVC) methods ignore that the predicted motion vectors usually contain large uncertainty due to complex motions, occlusions, etc., which consequently decrease the accuracy of the generated temporal contexts. In addition, existing C-LVC methods have a weak ability to mine diverse dependencies within the context space, which are closely related to the coding efficiency. To address these issues, an efficient temporal redundancy mining method is proposed to improve the coding efficiency of C-LVC in this paper. To generate accurate temporal contexts, a Long Short-Term Motion Aggregation (LSTMA) model is proposed, in which an LSTMA-based motion estimation module is developed to capture both current and aggregated long short-term motion information to reduce the uncertainty of predicted motion vectors. Based on the dual motion information, an LSTMA-based temporal context mining module is developed to exploit the aggregated long short-term motion information and increase the accuracy of the generated temporal contexts. In order to fully eliminate spatial-temporal redundancies in a video, a Global-Local Information Decorrelation Module (GLIDM)-based context codec is proposed, in which the GLIDM is designed based on the visual state space block (namely vmamba), the residual block, and the squeeze-and-excitation block to effectively capture long-range, short-range spatial-temporal dependencies and channel-wise dependencies. Experimental results demonstrate that our proposed method can effectively improve the coding performance of C-LVC, and outperforms other state-of-the-art LVC methods. Zhaoqing Pan, Jianjun Lei 0001, Bo Peng 0007, Haoran Xie 0001, Fu Lee Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Multi-Scale Feature Compression via Multi-Receptive-Field Convolutional Neural Network for Machine VisionabstractMulti-scale feature compression is essential in machine vision tasks for reducing storage and transmission costs while maintaining task performance. However, existing multi-scale feature compression methods fail to effectively extract and aggregate the local and global correlations of multi-scale features, resulting in incomplete elimination of feature redundancies. Moreover, these multi-scale feature compression methods mainly rely on mean square error-based loss functions to optimize signal fidelity, but they fail to adequately preserve semantic information critical to machine vision tasks. To address these issues, a Multi-receptive-field Convolutional Neural Network (MCNN)-based multi-scale feature compression method is proposed in this paper, which not only achieves compact fusion of multi-scale features but also enhances the semantic fidelity of reconstructed features. To effectively eliminate feature redundancies, a multi-receptive-field-based feature fusion module is designed for capturing both local and global correlations in multi-scale features. To enhance the quality of the reconstructed features, a cosine similarity-based multi-fidelity loss function is developed by considering both signal and semantic fidelity. Extensive experiments on the object detection and instance segmentation tasks show that the proposed MCNN outperforms the state-of-the-art multi-scale feature compression methods in terms of compression efficiency. The source code of our proposed MCNN is publicly available athttps://github.com/NUIST-Videocoding/MCNN.git. Zhaoqing Pan, Haihang Wang, Tiesong Zhao, Haoran Xie 0001, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2026 | Deep Video Coding With Bit-Depth ScalabilityabstractDeep video coding techniques have achieved significant advancements, leading to enhanced compression performance. However, existing approaches are primarily optimized for 8-bit content, thereby limiting their effectiveness in scenarios with different bit-depths. In this paper, we propose a deep bit-depth scalable video codec (DB-SVC) that supports two-layer scalability for different bit-depths. First, we design a base layer (BL) for low bit-depth (LBD) videos, incorporating a dual-stage multi-scale feature extraction module (DFEM) to enhance compression efficiency while providing reference features for subsequent coding. Second, we introduce an inter-layer bit-depth enhancement module (IBEM) that refines the bit-depth of BL reconstructed frames by leveraging interlayer information, thus enhancing the reference quality without increasing coding overhead. Third, we design an enhancement layer (EL) tailored for high bit-depth (HBD) videos, employing a bit-depth residual compression (BRC) method to achieve a more accurate reconstruction of HBD videos. DB-SVC supports progressive decoding of LBD and HBD videos, accommodating diverse display requirements. Experimental results demonstrate that DB-SVC outperforms state-of-the-art codecs in LBD and HBD scenarios. At the same PSNR/MS-SSIM levels, DB-SVC achieves average bit-rate savings of 11.94%/53.35% for 8-bit videos and 55.98%/73.16% for 10-bit videos while comparing with VTM13.2, showcasing its superior compression performance. Zhaoqing Pan, Tiesong Zhao, Xiaoming Tao 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Evaluating Visual Quality of Autostereoscopic 3D Displays via a Multimodal Parameter Perception NetworkabstractEvaluating the visual quality of autostereoscopic 3D displays is crucial for quantifying their stereoscopic viewing experience and optimizing display performance. Existing quality evaluation methods primarily predict the visual quality of autostereoscopic 3D displays by indirectly learning display parameter information from image content. However, these methods fail to explicitly model the relationship between display parameters and visual quality, thereby limiting their prediction accuracy. To address this problem, a Multimodal Parameter Perception Network (MPPNet)-based visual quality assessment method is proposed in this paper, which treats display parameters as textual modalities to explicitly establish their relationship with visual quality. To effectively understand the semantic information of display parameter texts, a Contrastive Language-Image Pretraining (CLIP)-based adaptive text encoder is proposed to generate robust semantic representations by capturing both general and domain-specific semantic embeddings. In parallel, a hierarchical vision encoder is adopted to extract visual representations from display images, which simulates the human binocular perception by capturing multi-level visual features from the left and right views. To achieve comprehensive cross-modal interaction, a mamba-based cross-modal fusion module is proposed to fuse textual and visual representations of display parameters by capturing both shallow and deep correlations. Extensive experimental results demonstrate that the proposed MPPNet achieves state-of-the-art performance in evaluating the visual quality of autostereoscopic 3D displays. Haoran Xie 0001, Fu Lee Wang, Zhaoqing Pan |
ACM Multimedia | 5 |
| 2025 | Efficient Chroma Intra Prediction via Exemplar Colorization Network for Versatile Video CodingabstractChroma intra prediction aims to reduce chroma redundancies within a frame, which plays an important role in improving the coding efficiency of intra coding. Existing chroma intra prediction methods typically utilize the spatial relationship between the current luma block and its neighboring reference luma blocks to predict its chroma samples. However, the spatial properties of luma components differ from those of chroma components, which limits the accuracy of chroma intra prediction. To tackle this issue, an efficient Exemplar Colorization Network (ECNet)-based chroma intra prediction method is proposed in this paper, in which the colorization relationship between reference luma and chroma components is exploited to predict the chroma components for the current luma component. Inspired by the principle that semantic information in an image exhibits short-range continuity, a Spatial-consistency-based Colorization Transfer Network (SCTNet) is proposed, which builds and transfers colorization representations of neighboring reference blocks for chroma prediction. To improve the chroma prediction capability of SCTNet, a colorization learning module is developed to learn the robust mapping relationship from the luma component to the chroma component in a region-to-pixel manner, and a weight-adaptive reconstruction module is designed to adaptively utilize reference information from neighboring blocks to generate an initial prediction result. In addition, to further improve the accuracy of chroma intra prediction, a multi-reference-based chroma refinement network is proposed, which simultaneously uses the spatial information of neighboring reference chroma blocks and the current luma block to eliminate blocking and color-bleeding artifacts in the initial prediction result. Experimental results demonstrate that our proposed ECNet outperforms the state-of-the-art chroma intra prediction methods in terms of coding performance. Zhaoqing Pan, Jixing Chen, Bo Peng 0007, Jianjun Lei 0001, Fu Lee Wang, Nam Ling, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2025 | Hierarchical Uncertainty-Aware Salient Object Detection for $360 ^{\circ }$ Images via Bi-Projection Collaborative Learningabstract$360^{\circ }$salient object detection has recently received much attention for 3D scene perception owing to its omnidirectional field of view (FoV). The capability of recognizing salient objects of$360^{\circ }$images remains technically challenging due to severe spherical distortion. In this paper, we develop a hierarchical uncertainty-aware$360^{\circ }$image salient object detection methodology that explicitly explores the geometric and spatial complementary coherence of Tangent projection (TP) and Equirectangular projection (ERP) by a collaborative learning strategy. Concretely, to mitigate spherical distortion, we first intend to learn saliency-related features from less-distorted tangent images, in which a deformation-aware attention block is introduced to mitigate the geometric distortion caused by projecting a$360^{\circ }$image onto a 2D plane. However, the discrepancies among tangent images pose a new challenge to$360^{\circ }$image salient object detection. To tackle this issue and achieve accurate localization for salient objects of all sizes, we design a spatial-frequency saliency feature aggregation module to leverage fast Fourier convolution to capture global contextual information from ERP images, such that obtaining more representative saliency features. Moreover, a hierarchical uncertainty-aware bi-projection consistency learning module with strong local-global information embedding capabilities is constructed, which learns the geometric and spatial correlations between tangent images and ERP images via a collaborative learning strategy. Ultimately, salient object maps are produced for$360^{\circ }$images on the basis of the merged saliency features driven by the uncertainty. Extensive experiments show that our developed method improves${\mathrm{F}}_\beta ^{\sigma }$by an average of 31.67% compared to twenty existing advanced methods on the publicly available 360-SOD dataset. Qiudan Zhang, Kaiyu Ji, Xu Wang 0006, Zhaoqing Pan, Jianmin Jiang |
IEEE Trans. Multim. | 5 |
| 2024 | Spatial Similarity-Based Fast Mode Decision for VVC Chroma Intra CodingabstractThe Versatile Video Coding (VVC) is the latest video coding standard that includes advanced intra coding techniques such as the chroma separate tree and cross-component linear model to improve the prediction accuracy of chroma components. However, these sophisticated techniques inevitably increase the computational complexity of chroma intra coding. In this paper, we propose a fast chroma intra-prediction mode decision algorithm to reduce the computational complexity of the VVC encoder. Firstly, the mode selection correlation between the current chroma Coding Unit (CU) and its nearest-encoded chroma CU is analyzed. Then, based on the variation of the sum of modulus differences, an early mode decision is proposed for the current chroma CU to determine whether the best coding mode should be traversed or directly inherited from its nearest encoded chroma CU. The experimental results show that the proposed algorithm significantly reduces the computational complexity of chroma intra-prediction and achieves an average of 28% chroma intra encoding time saving with only 0.08-1.43% performance degradation in terms of BD-rate. Haihang Wang, Jixing Chen, Fu Lee Wang, Zhaoqing Pan |
VCIP | 6 |
| 2024 | Saliency Map-Guided End-to-End Image Coding for MachinesabstractExisting end-to-end image coding for machines (ICM) methods generally use joint training strategies to promote the compression efficiency for machine vision without considering the influence of different regions in the image. To encourage the image compression network to focus on the regions that are critical to the subsequent visual task, this paper proposes a saliency map-guided image compression network (SMIC-Net) for ICM. Specifically, a saliency map-guided transform module (SMTM) is proposed to improve the representation ability of image features for object detection task by exploring the semantic and structural information of the detected object. Besides, a saliency map-guided mean square error (SM-MSE) loss is designed to place more emphasis on the detected object regions. Experimental results demonstrate that the proposed SMIC-Net effectively promotes the compression efficiency for machine vision. Bo Peng 0007, Tianxiang Lin, Dengchao Jin, Zhaoqing Pan, Jianjun Lei 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | SWGNet: Step-Wise Reference Frame Generation Network for Multiview Video CodingabstractIn multiview video coding, the coding performance highly depends on the quality of the reference frames. In view of this, a step-wise reference frame generation network (SWGNet) is designed to improve the quality of the reference frame for efficient multiview video coding. In particular, a frame-level to block-level learning paradigm is proposed to step-wisely generate a high-quality reference frame. In the frame-level stage, by exploiting parallax correlations between temporal and inter-view references on the basis of image alignment, a parallax-guided frame-level synthesis module is proposed to generate an elementary reference frame. Then, in the block-level stage, a transformer-based block-level aggregation module is designed to further refine the texture details of the reference frame by modeling long-range dependencies among pixels. The proposed SWGNet is integrated into 3D-HEVC, and extensive experiments demonstrate that the proposed method achieves significant bitrate saving compared with 3D-HEVC. Jing Zhang 0017, Yonghong Hou, Zhaoqing Pan, Bo Peng 0007, Nam Ling, Jianjun Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | λ-Domain Rate Control via Wavelet-Based Residual Neural Network for VVC HDR Intra CodingabstractHigh dynamic range (HDR) video offers a more realistic visual experience than standard dynamic range (SDR) video, while introducing new challenges to both compression and transmission. Rate control is an effective technology to overcome these challenges, and ensure optimal HDR video delivery. However, the rate control algorithm in the latest video coding standard, versatile video coding (VVC), is tailored to SDR videos, and does not produce well coding results when encoding HDR videos. To address this problem, a data-driven λ -domain rate control algorithm is proposed for VVC HDR intra frames in this paper. First, the coding characteristics of HDR intra coding are analyzed, and a piecewise R- λ model is proposed to accurately determine the correlation between the rate (R) and the Lagrange parameter λ for HDR intra frames. Then, to optimize bit allocation at the coding tree unit (CTU)-level, a wavelet-based residual neural network (WRNN) is developed to accurately predict the parameters of the piecewise R- λ model for each CTU. Third, a large-scale HDR dataset is established for training WRNN, which facilitates the applications of deep learning in HDR intra coding. Extensive experimental results show that our proposed HDR intra frame rate control algorithm achieves superior coding results than the state-of-the-art algorithms. The source code of this work will be released at https://github.com/TJU-Videocoding/WRNN.git. Jianjun Lei 0001, Zhaoqing Pan, Bo Peng 0007, Haoran Xie 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Deep In-Loop Filtering via Multi-Domain Correlation Learning and Partition Constraint for Multiview Video CodingabstractThe deep learning-based in-loop filtering methods have greatly improved the coding efficiency for High Efficiency Video Coding (HEVC). However, directly applying these HEVC-orientated in-loop filtering methods to multiview video coding may not obtain satisfactory performance due to the characteristics of multiview video. In this paper, a deep in-loop filtering method based on multi-domain correlation learning and partition constraint network (MDP-Net) is proposed to boost the multiview video coding performance. To the best of our knowledge, this work is the first attempt at deep in-loop filtering for multiview video coding. Specifically, a multi-domain correlation learning module is presented to restore the high-frequency details of the distorted frame by exploring the multi-domain correlations. Besides, based on the block partition information generated in video coding, a partition-constrained reconstruction module is proposed to better attenuate the compression artifacts by designing a partition loss. Finally, the proposed MDP-Net is integrated into 3D-HEVC reference software, and the experimental results demonstrate that the proposed method achieves considerable performance improvement compared with 3D-HEVC. Bo Peng 0007, Renjie Chang, Zhaoqing Pan, Ge Li 0006, Nam Ling, Jianjun Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Reducing Background Induced Domain Shift for Adaptive Person Re-IdentificationabstractCross-domain person re-identification (Re-ID) is a challenging and important task in monitoring safety and procedure compliance of industrial work places. In this article, a novel method is proposed to reduce background induced domain shift for adaptive person Re-ID. Specifically, a foreground-background joint clustering module is proposed to extract discriminative foreground and background features and an attention-based feature disentanglement module is designed to reduce the interference of background with the extraction of discriminative foreground features. Experimental results on three widely used person Re-ID benchmarking datasets (Market-1501, DukeMTMC-reID, and MSMT17) have demonstrated that the proposed method achieves promising performance compared with the state-of-the-art methods. Jianjun Lei 0001, Tianyi Qin, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Haifeng Shen, Sam Kwong |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Learned Video Compression With Efficient Temporal Context LearningabstractIn contrast to image compression, the key of video compression is to efficiently exploit the temporal context for reducing the inter-frame redundancy. Existing learned video compression methods generally rely on utilizing short-term temporal correlations or image-oriented codecs, which prevents further improvement of the coding performance. This paper proposed a novel temporal context-based video compression network (TCVC-Net) for improving the performance of learned video compression. Specifically, a global temporal reference aggregation (GTRA) module is proposed to obtain an accurate temporal reference for motion-compensated prediction by aggregating long-term temporal context. Furthermore, in order to efficiently compress the motion vector and residue, a temporal conditional codec (TCC) is proposed to preserve structural and detailed information by exploiting the multi-frequency components in temporal context. Experimental results show that the proposed TCVC-Net outperforms public state-of-the-art methods in terms of both PSNR and MS-SSIM metrics. Dengchao Jin, Jianjun Lei 0001, Bo Peng 0007, Zhaoqing Pan, Li Li 0040, Nam Ling |
IEEE Trans. Image Process. | 4 |
| 2022 | Texture-Guided End-to-End Depth Map CompressionabstractEnd-to-end compression methods designed for the texture image have achieved excellent coding performances. Due to the characteristic differences between the depth map and the texture image, the texture-oriented methods have limitations in depth map compression. To address this problem, this paper proposes a texture-guided end-to-end depth map compression network (TDMC-Net). Specifically, the proposed TDMC-Net is mainly composed of the texture-guided transform module (TTM) which performs the nonlinear transform with providing the textual context to reduce the redundancy in depth feature, and a texture-guided conditional entropy model (TCEM) which is designed to improve the entropy model by introducing the texture conditional prior. Experimental results show that the proposed TDMC-Net boosts the depth coding efficiency by utilizing the texture information and achieves superior performance. Bo Peng 0007, Yuying Jing, Dengchao Jin, Xiangrui Liu, Zhaoqing Pan, Jianjun Lei 0001 |
ICIP | 5 |
| 2022 | Multiple Resolution Prediction With Deep Up-Sampling for Depth Video CodingabstractThe depth video contains large smooth contents with sharp edges. Since the deep learning-based color video orientated intra prediction methods pay no attention to the characteristics of depth video, they are unsuitable for optimizing the coding efficiency of depth video. In this paper, a multiple resolution prediction method with deep up-sampling is proposed to promote the coding efficiency of depth video. To efficiently encode the depth blocks of different complexity, the depth block is selectively encoded at different resolutions, including$\times 1$,$\times 1$/2, and$\times 1$/4 resolutions. If the block is encoded with a low-resolution (LR), the resolution of reconstructed LR depth block is recovered by an up-sampling network. To constrain the quality of both reconstructed high-resolution depth block and its synthesized view, a view synthesis distortion guidance mechanism is proposed for the up-sampling network. In addition, a distillation-based lightweight up-sampling network is proposed to reduce the computational complexity. Experimental results demonstrate that the proposed multiple resolution prediction method obtains an average of 10.84% BD-rate saving in comparison with 3D-HEVC. Ge Li 0006, Jianjun Lei 0001, Zhaoqing Pan, Bo Peng 0007, Nam Ling |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | TSAN: Synthesized View Quality Enhancement via Two-Stream Attention Network for 3D-HEVCabstractIn three-dimensional video system, the texture and depth videos are jointly encoded, and then the Depth Image Based Rendering (DIBR) is utilized to realize view synthesis. However, the compression distortion of texture and depth videos, as well as the disocclusion problem in DIBR degrade the visual quality of the synthesized view. To address this problem, a Two-stream Attention Network (TSAN)-based synthesized view quality enhancement method is proposed for 3D-High Efficiency Video Coding (3D-HEVC) in this article. First, the shortcomings of the view synthesis technique and traditional convolutional neural networks are analyzed. Then, based on these analyses, a TSAN with two information extraction streams is proposed for enhancing the quality of the synthesized view, in which the global information extraction stream learns the contextual information, and the local information extraction stream extracts the texture information from the rendered image. Third, a Multi-Scale Residual Attention Block (MSRAB) is proposed, which can efficiently detect features in different scales, and adaptively refine features by considering interdependencies among spatial dimensions. Extensive experimental results show that the proposed synthesized view quality enhancement method achieves significantly better performance than the state-of-the-art methods. Zhaoqing Pan, Jianjun Lei 0001, Nam Ling, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | RDEN: Residual Distillation Enhanced Network-Guided Lightweight Synthesized View Quality Enhancement for 3D-HEVCabstractIn the three-dimensional video system, the depth image-based rendering is a key technique for generating synthesized views, which provides audiences with depth perception and interactivity. However, the inaccuracy of depth information leads to geometrical rendering position errors, and the compression distortion of texture and depth videos degrades the quality of the synthesized views. Although existing quality enhancement methods can eliminate the distortions in the synthesized views, their huge computational complexity hinders their applications in real-time multimedia systems. To this end, a residual distillation enhanced network (RDEN)-guided lightweight synthesized view quality enhancement (SVQE) method is proposed to minimize holes and compression distortions in the synthesized views while reducing the model complexity. First, a rethinking on the deep-learning-based SVQE methods is performed. Then, a feature distillation attention block is proposed to effectively reduce the distortions in the synthesized views and make the model fulfill more real-time tasks, which is a lightweight and flexible feature extraction block using an information distillation mechanism and a lightweight multi-scale spatial attention mechanism. Third, a residual feature fusion block is proposed to improve the enhancement performance by using the feature fusion mechanism, which efficiently improves the feature extraction capability without introducing any additional parameters. Experimental results prove that the proposed RDEN efficiently improves the SVQE performance while consuming few computational complexities compared with the state-of-the-art SVQE methods. Zhaoqing Pan, Jianjun Lei 0001, Nam Ling, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | DACNN: Blind Image Quality Assessment via a Distortion-Aware Convolutional Neural NetworkabstractDeep neural networks have achieved great performance on blind Image Quality Assessment (IQA), but it is still challenging for using one network to accurately predict the quality of images with different distortions. In this paper, a Distortion-Aware Convolutional Neural Network (DACNN) is proposed for blind IQA, which works effectively for not only synthetically distorted images but also authentically distorted images. The proposed DACNN consists of a distortion aware module, a distortion fusion module, and a quality prediction module. In the distortion aware module, a Siamese network-based pretraining strategy is proposed to design a synthetic distortion-aware network for full learning the synthetic distortions, and an authentic distortion-aware network is used for extracting the authentic distortions. To efficiently fuse the learned distortion features, and make the network pay more attention to the essential features, a weight-adaptive fusion network is proposed to adaptively adjust the weight of each distortion. Finally, the quality prediction module is adopted to map the fused features to a quality score. Extensive experiments on four authentic IQA databases and four synthetic IQA databases have proved the effectiveness of the proposed DACNN. Zhaoqing Pan, Jianjun Lei 0001, Yuming Fang 0001, Xiao Shao, Nam Ling, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | C2FNet: A Coarse-to-Fine Network for Multi-View 3D Point Cloud GenerationabstractGeneration of a 3D model of an object from multiple views has a wide range of applications. Different parts of an object would be accurately captured by a particular view or a subset of views in the case of multiple views. In this paper, a novel coarse-to-fine network (C2FNet) is proposed for 3D point cloud generation from multiple views. C2FNet generates subsets of 3D points that are best captured by individual views with the support of other views in a coarse-to-fine way, and then fuses these subsets of 3D points to a whole point cloud. It consists of a coarse generation module where coarse point clouds are constructed from multiple views by exploring the cross-view spatial relations, and a fine generation module where the coarse point cloud features are refined under the guidance of global consistency in appearance and context. Extensive experiments on the benchmark datasets have demonstrated that the proposed method outperforms the state-of-the-art methods. Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Qingming Huang |
IEEE Trans. Image Process. | 5 |
| 2022 | Disparity-Aware Reference Frame Generation Network for Multiview Video CodingabstractMultiview video coding (MVC) aims to compress the multiview video through the elimination of video redundancies, where the quality of the reference frame directly affects the compression efficiency. In this paper, we propose a deep virtual reference frame generation method based on a disparity-aware reference frame generation network (DAG-Net) to transform the disparity relationship between different viewpoints and generate a more reliable reference frame. The proposed DAG-Net consists of a multi-level receptive field module, a disparity-aware alignment module, and a fusion reconstruction module. First, a multi-level receptive field module is designed to enlarge the receptive field, and extract the multi-scale deep features of the temporal and inter-view reference frames. Then, a disparity-aware alignment module is proposed to learn the disparity relationship, and perform disparity shift on the inter-view reference frame to align it with the temporal reference frame. Finally, a fusion reconstruction module is utilized to fuse the complementary information and generate a more reliable virtual reference frame. Experiments demonstrate that the proposed reference frame generation method achieves superior performance for multiview video coding. Jianjun Lei 0001, Zongqian Zhang, Zhaoqing Pan, Dong Liu 0002, Xiangrui Liu, Ying Chen 0011, Nam Ling |
IEEE Trans. Image Process. | 3 |
| 2022 | VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality AssessmentabstractGuided by the free-energy principle, generative adversarial networks (GAN)-based no-reference image quality assessment (NR-IQA) methods have improved the image quality prediction accuracy. However, the GAN cannot well handle the restoration task for the free-energy principle-guided NR-IQA methods, especially for the severely destroyed images, which results in that the quality reconstruction relationship between the distorted image and its restored image cannot be accurately built. To address this problem, a visual compensation restoration network (VCRNet)-based NR-IQA method is proposed, which uses a non-adversarial model to efficiently handle the distorted image restoration task. The proposed VCRNet consists of a visual restoration network and a quality estimation network. To accurately build the quality reconstruction relationship between the distorted image and its restored image, a visual compensation module, an optimized asymmetric residual block, and an error map-based mixed loss function, are proposed for increasing the restoration capability of the visual restoration network. For further addressing the NR-IQA problem of severely destroyed images, the multi-level restoration features which are obtained from the visual restoration network are used for the image quality estimation. To prove the effectiveness of the proposed VCRNet, seven representative IQA databases are used, and experimental results show that the proposed VCRNet achieves the state-of-the-art image quality prediction accuracy. The implementation of the proposed VCRNet has been released at https://github.com/NUIST-Videocoding/VCRNet. Zhaoqing Pan, Jianjun Lei 0001, Yuming Fang 0001, Xiao Shao, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2022 | MIEGAN: Mobile Image Enhancement via a Multi-Module Cascade Neural NetworkabstractVisual quality of images captured by mobile devices is often inferior to that of images captured by a Digital Single Lens Reflex (DSLR) camera. This paper presents a novel generative adversarial network-based mobile image enhancement method, referred to as MIEGAN. It consists of a novel multi-module cascade generative network and a novel adaptive multi-scale discriminative network. The multi-module cascade generative network is built upon a two-stream encoder, a feature transformer, and a decoder. In the two-stream encoder, a luminance-regularizing stream is proposed to help the network focus on low-light areas. In the feature transformation module, two networks effectively capture both global and local information of an image. To further assist the generative network to generate the high visual quality images, a multi-scale discriminator is used instead of a regular single discriminator to distinguish whether an image is fake or real globally and locally. To balance the global and local discriminators, an adaptive weight allocation is proposed. In addition, a contrast loss is proposed, and a new mixed loss function is developed to improve the visual quality of the enhanced images. Extensive experiments on the popular DSLR photo enhancement dataset and MIT-FiveK dataset have verified the effectiveness of the proposed MIEGAN. Zhaoqing Pan, Jianjun Lei 0001, Wanqing Li 0001, Nam Ling, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2021 | No-reference stereoscopic image quality assessment based on global and local content characteristics
Lili Shen, Xiongfei Chen, Zhaoqing Pan, Kefeng Fan, Jianjun Lei 0001 |
Neurocomputing | 3 |
| 2021 | A CNN-Based Fast Inter Coding Method for VVCabstractThe Versatile Video Coding (VVC) achieves superior coding efficiency as compared with the High Efficiency Video Coding (HEVC), while its excellent coding performance is at the cost of several high computational complexity coding tools, such as Quad-Tree plus Multi-type Tree (QTMT)-based Coding Units (CUs) and multiple inter prediction modes. To reduce the computational complexity of VVC, a CNN-based fast inter coding method is proposed in this paper. First, a multi-information fusion CNN (MF-CNN) model is proposed to early terminate the QTMT-based CU partition process by jointly using the multi-domain information. Then, a content complexity-based early Merge mode decision is proposed to skip the time-consuming inter prediction modes by considering the CU prediction residuals and the confidence of MF-CNN. Experimental results show that the proposed method reduces an average of 30.63% VVC encoding time, and the Bjøontegaard Delta Bit Rate (BDBR) increases about 3%. Zhaoqing Pan, Peihan Zhang, Bo Peng 0007, Nam Ling, Jianjun Lei 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Motion and disparity vectors early determination for texture video in 3D-HEVC
Zhaoqing Pan, Xiaokai Yi |
Multim. Tools Appl. | 1 |
| 2020 | Efficient In-Loop Filtering Based on Enhanced Deep Convolutional Neural Networks for HEVCabstractThe raw video data can be compressed much by the latest video coding standard, high efficiency video coding (HEVC). However, the block-based hybrid coding used in HEVC will incur lots of artifacts in compressed videos, the video quality will be severely influenced. To settle this problem, the in-loop filtering is used in HEVC to eliminate artifacts. Inspired by the success of deep learning, we propose an efficient in-loop filtering algorithm based on the enhanced deep convolutional neural networks (EDCNN) for significantly improving the performance of in-loop filtering in HEVC. Firstly, the problems of traditional convolutional neural networks models, including the normalization method, network learning ability, and loss function, are analyzed. Then, based on the statistical analyses, the EDCNN is proposed for efficiently eliminating the artifacts, which adopts three solutions, including a weighted normalization method, a feature information fusion block, and a precise loss function. Finally, the PSNR enhancement, PSNR smoothness, RD performance, subjective test, and computational complexity/GPU memory consumption are employed as the evaluation criteria, and experimental results show that when compared with the filter in HM16.9, the proposed in-loop filtering algorithm achieves an average of 6.45% BDBR reduction and 0.238 dB BDPSNR gains. Zhaoqing Pan, Xiaokai Yi, Yun Zhang 0002, Byeungwoo Jeon, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2020 | Frame-level Bit Allocation Optimization Based on Video Content Characteristics for HEVCabstractRate control plays an important role in high efficiency video coding (HEVC), and bit allocation is the foundation of rate control. The video content characteristics are significant for bit allocation, and modeling an accurate relationship between video content characteristics and bit allocation is essential for bit allocation optimization. Therefore, in this article, a video content characteristics–based frame-level optimal bit allocation algorithm is proposed for improving the rate distortion (RD) performance of HEVC. First, the number of search points of motion estimation is used to evaluate the motion activity of video content, and the relationship between the search points and bit allocation is modeled as the search-points model. Second, the grey level co-occurrence matrix and temporal perceptual information are used to evaluate the spatial and temporal texture complexity, and the relationship between the video content texture complexity and bit allocation is modeled as the texture-complexity model. Then, the search-points model and texture-complexity model are jointly employed to allocate the coding bits for the second and third layers of the HEVC hierarchical coding structure. Finally, the remaining coding bits of a group-of-pictures (GOP) are allocated to the first layer of HEVC coding structure. To evaluate the performance of the proposed algorithm, the RD performance and bitrate accuracy are used as evaluation criteria, and the experimental results show that when compared with the popularly used R-λ model–based bit allocation algorithm, the proposed algorithm achieves an average of -3.43% BDBR reduction and 0.13 dB BDPSNR gains with only 0.02% loss of bitrate accuracy. Zhaoqing Pan, Xiaokai Yi, Yun Zhang 0002, Hui Yuan 0001, Fu Lee Wang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Fast Coding Unit Decision for Intra Screen Content Coding Based on Ensemble LearningabstractThe Screen Content Coding (SCC) is an extension of High Efficiency Video Coding (HEVC), and it achieves significant improvement on compression ratio. However, the obtained coding efficiency is at the cost of high computational complexity. In this paper, to reduce the computation complexity, we propose to use an ensemble classifier for predicting the coding unit (CU) in intra-coding. Firstly, the L1-loss based linear support vector machine (SVM) is employed as basic classifier for its simplicity. Then, a bagging scheme is applied to train the linear classifiers and boost the prediction accuracy by ensemble learning. Compared with the reference software SCM-5.0, the proposed scheme can achieve 30% complexity reduction on average with only 1.64% bit rates increase. Yali Xue, Xu Wang 0006, Linwei Zhu, Zhaoqing Pan, Sam Kwong |
ICASSP | 4 |
| 2019 | A comprehensive search for expert classification methods in disease diagnosis and predictionabstractAbstract Healthcare data analysis is currently a challenging and crucial research issue for the development of a robust disease diagnosis and prediction system. Many specific and a few common methods have been discussed in the literature for healthcare data classification. The present study implements 32 classification methods of six categories (Bayes, function‐based, lazy, meta, rule‐based, and tree‐based) with the objective of searching the best and common categories and methods in healthcare data mining. The performance of each classification method has been evaluated based on analysis time, classification accuracy, precision, recall, F‐measure, area under the receiver operating characteristic curve, root mean square error, kappa coefficient, Kulczynski's measure, and Fowlkes–Mallows index and compared with more than 90 classification methods used in past studies. Seventeen healthcare datasets related to thyroid, cancer, skin disease, heart disease, hepatitis, lymphography, audiology, diabetes, surgery, arrhythmia, postsurvival, liver, and tumour have been used in the performance assessment of the classification methods. The tree‐based classification methods have a better performance (with an average classification accuracy of 79.92% and maximum accuracy of 99.50%; an analysis time of 3.91 s for the logistic model tree classifier) than the other methods. Furthermore, the association of datasets and classification methods has been discussed. Sunil Kr. Jha, Zhaoqing Pan, Ehsan Elahi 0006, Nilesh V. Patel |
Expert Syst. J. Knowl. Eng. | 2 |
| 2019 | Market impact analysis via deep learned architectures
Xiaodong Li 0007, Jingjing Cao, Zhaoqing Pan |
Neural Comput. Appl. | 3 |
| 2019 | Automatic Medical Image Registration Based on an Integrated Method Combining Feature and Area Information
Jiucheng Xie, Chi-Man Pun, Zhaoqing Pan, Hao Gao 0005, Baoyun Wang |
Neural Process. Lett. | 3 |
| 2019 | Machine Learning for Wireless Multimedia Data Security
Zhaoqing Pan, Ching-Nung Yang, Victor S. Sheng, Naixue Xiong, Weizhi Meng 0001 |
Secur. Commun. Networks | 1 |
| 2018 | Effective Data Driven Coding Unit Size Decision Approaches for HEVC INTRA CodingabstractHigh Efficiency Video Coding (HEVC) INTRA coding improves compression efficiency by adopting advanced coding technologies, such as multi-level quad-tree block partitioning and up to 35-mode INTRA prediction. However, it significantly increases the coding complexity, memory access, and power consumption, which goes against its widely applications, especially for ultra-high definition and/or mobile video applications. To tackle this problem, we propose effective data driven coding unit (CU) size decision approaches for HEVC INTRA coding, which consists of two stages of support vector machine-based fast INTRA CU size decision schemes at four CU decision layers. At the first stage classification, a three output classifier with offline learning is developed to early terminate the CU size decision or early skip checking the current CU depth. As for the samples that neither early skipped nor early terminated, the second stage of binary classification, which learns online from previous coded frames, is proposed to further refine the CU size decision. Representative features for the CU size decision are explored at different decision layers and stages of classifications. Finally, the optimal parameters derived from the training data are achieved to reasonably allocate complexity among different CU layers at given total rate-distortion degradation constraint. Extensive experiments show that the proposed overall algorithm can achieve 27.95%–80.53% and 52.48% on average complexity reduction for the CU size decision as compared with the original HM16.7 model. Meanwhile, the average Bjonteggard delta peak-signal-to-noise ratio degradation is only −0.08 dB, which is negligible. The overall performance of the proposed algorithm outperforms the state-of-the-art benchmark schemes. Yun Zhang 0002, Zhaoqing Pan, Na Li 0015, Xu Wang 0006, Gangyi Jiang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | CTU-Level Complexity Control for High Efficiency Video CodingabstractAmong the existing video-related applications, a large proportion have requirements for the scalability of the video coding complexity, such as live video chatting and video coding on power-limited mobile devices. Hence, the complexity control algorithms, which aim to make an effective and flexible tradeoff between coding complexity and rate-distortion (RD) performance, have a great practical value. In this paper, a novel complexity control scheme for high efficiency video coding (HEVC) is proposed by dynamically adjusting the depth range for each coding tree unit (CTU). To control the complexity accurately, a statistical model is proposed to estimate the coding complexity of each CTU. Then the complexity budget is allocated to each CTU proportionally to its estimated complexity. At last, the depth range is optimized for each CTU based on the allocated complexity and the probability that contains the actual maximum depth. Our method works well even if the ratio of target complexity to full complexity drops to 40%. The experimental results show that our proposed method outperforms other four state-of-the-art methods in terms of the RD performance, and has superior complexity control accuracy and complexity control stability compared with other one-pass complexity control strategies. Jia Zhang 0002, Sam Kwong, Tiesong Zhao, Zhaoqing Pan |
IEEE Trans. Multim. | 4 |
| 2018 | Adaptive Fractional-Pixel Motion Estimation Skipped Algorithm for Efficient HEVC Motion EstimationabstractHigh-Efficiency Video Coding (HEVC) efficiently addresses the storage and transmit problems of high-definition videos, especially for 4K videos. The variable-size Prediction Units (PUs)--based Motion Estimation (ME) contributes a significant compression rate to the HEVC encoder and also generates a huge computation load. Meanwhile, high-level encoding complexity prevents widespread adoption of the HEVC encoder in multimedia systems. In this article, an adaptive fractional-pixel ME skipped scheme is proposed for low-complexity HEVC ME. First, based on the property of the variable-size PUs--based ME process and the video content partition relationship among variable-size PUs, all inter-PU modes during a coding unit encoding process are classified into root-type PU mode and children-type PU modes. Then, according to the ME result of the root-type PU mode, the fractional-pixel ME of its children-type PU modes is adaptively skipped. Simulation results show that, compared to the original ME in HEVC reference software, the proposed algorithm reduces ME encoding time by an average of 63.22% while encoding efficiency performance is maintained. Zhaoqing Pan, Jianjun Lei 0001, Fu Lee Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2017 | Allowable depth distortion based fast mode decision and reference frame selection for 3D depth coding
Yun Zhang 0002, Zhaoqing Pan, Yang Zhou 0052, Linwei Zhu |
Multim. Tools Appl. | 2 |
| 2016 | Early DIRECT mode decision based on all-zero block and rate distortion cost for multiview video codingabstractThe exhaustive variable‐block‐size mode decision can efficiently remove the redundancies among the multiview videos, while it also leads to significant increase of computational complexity in the multiview video coding (MVC) encoder, and the high encoding complexity becomes a bottleneck for the MVC encoder to achieve real‐time multimedia applications. To address this bottleneck, many fast mode decision methods have been proposed. However, most of them are only suitable for optimising the encoding complexity of the odd views of the MVC encoder. In this study, based on the property of the all‐zero block and rate distortion (RD) cost of the DIRECT mode as well as the correlations between the current macroblock (MB) and its spatial–temporal nearby MBs, an early DIRECT mode decision method is proposed for reducing the encoding complexity of the MVC. Experimental results show that the proposed method achieves 48.25 and 55.64% on average encoding time saving for the even and odd views, respectively, whereas the RD performance degradation is quite acceptable. In summary, the proposed method efficiently reduces the encoding complexity for the MVC encoder. Zhaoqing Pan, Yun Zhang 0002, Jianjun Lei 0001, Long Xu 0001, Xingming Sun |
IET Image Process. | 1 |
| 2016 | Fast reference frame selection based on content similarity for low complexity HEVC encoder
Zhaoqing Pan, Jianjun Lei 0001, Yun Zhang 0002, Xingming Sun, Sam Kwong |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Fast Transform Unit Depth Decision Based on Quantized Coefficients for HEVCabstractThe quad tree structure based Transform Unit (TU) helps high efficiency video coding to improve the coding efficiency. However, the achieved coding efficiency comes at the cost of the increased computational complexity. In this paper, based on the quantizated coefficients of the TU, we propose an early termination for the quad tree structure based TU encoding process. If the quantized coefficients of the luminance components are all zeros, the TU encoding process will be terminated. Experimental results show that the proposed method achieves about 55.13% on average TU encoding time saving, while the rate distortion performance degradation is negligible. Zhaoqing Pan, Jianjun Lei 0001, Yun Zhang 0002, Sam Kwong |
SMC | 1 |
| 2015 | View synthesis distortion model based frame level rate control optimization for multiview depth video coding
Xu Wang 0006, Sam Kwong, Hui Yuan 0001, Yun Zhang 0002, Zhaoqing Pan |
Signal Process. | 5 |
| 2015 | Fast Mode Decision Using Inter-View and Inter-Component Correlations for Multiview Depth Video CodingabstractWith the development of three-dimensional (3-D) display technologies, 3-D video has attracted more and more interest. Multiview video plus depth (MVD) is one of the most popular representation formats of 3-D video. In MVD coding system, multiview depth video needs to be coded and transmitted in addition to the texture video. This paper presents a novel fast mode decision (FMD) method for odd views in multiview depth video coding. First, the inter-view and inter-component coding correlations are analyzed to provide efficient reference information. Then, with a view to the characteristics of different types of frames, different early termination strategies are proposed. For the nonanchor frame, the early termination criterion is based on the rate-distortion cost information of the even views and the coded block pattern information. For the anchor frame, the criterion is set stricter to maintain the coding accuracy. Experimental results show that the proposed method can reduce 78.07% coding time on average, without significant loss of video quality. Jianjun Lei 0001, Jing Sun 0010, Zhaoqing Pan, Sam Kwong, Jinhui Duan, Chunping Hou |
IEEE Trans. Ind. Informatics | 3 |
| 2015 | Low Complexity HEVC INTRA Coding for High-Quality Mobile Video CommunicationabstractINTRA video coding is essential for high quality mobile video communication and industrial video applications since it enhances video quality, prevents error propagation, and facilitates random access. The latest high-efficiency video coding (HEVC) standard has adopted flexible quad-tree-based block structure and complex angular INTRA prediction to improve the coding efficiency. However, these technologies increase the coding complexity significantly, which consumes large hardware resources, computing time and power cost, and is an obstacle for real-time video applications. To reduce the coding complexity and save power cost, we propose a fast INTRA coding unit (CU) depth decision method based on statistical modeling and correlation analyses. First, we analyze the spatial CU depth correlation with different textures and present effective strategies to predict the most probable depth range based on the spatial correlation among CUs. Since the spatial correlation may fail for image boundary and transitional areas between textural and smooth areas, we then present a statistical model-based CU decision approach in which adaptive early termination thresholds are determined and updated based on the rate-distortion (RD) cost distribution, video content, and quantization parameters (QPs). Experimental results show that the proposed method can reduce the complexity by about 56.76% and 55.61% on average for various sequences and configurations; meanwhile, the RD degradation is negligible. Yun Zhang 0002, Sam Kwong, Zhaoqing Pan, Hui Yuan 0001, Gangyi Jiang |
IEEE Trans. Ind. Informatics | 4 |
| 2015 | Machine Learning-Based Coding Unit Depth Decisions for Flexible Complexity Allocation in High Efficiency Video CodingabstractIn this paper, we propose a machine learning-based fast coding unit (CU) depth decision method for High Efficiency Video Coding (HEVC), which optimizes the complexity allocation at CU level with given rate-distortion (RD) cost constraints. First, we analyze quad-tree CU depth decision process in HEVC and model it as a three-level of hierarchical binary decision problem. Second, a flexible CU depth decision structure is presented, which allows the performances of each CU depth decision be smoothly transferred between the coding complexity and RD performance. Then, a three-output joint classifier consists of multiple binary classifiers with different parameters is designed to control the risk of false prediction. Finally, a sophisticated RD-complexity model is derived to determine the optimal parameters for the joint classifier, which is capable of minimizing the complexity in each CU depth at given RD degradation constraints. Comparative experiments over various sequences show that the proposed CU depth decision algorithm can reduce the computational complexity from 28.82% to 70.93%, and 51.45% on average when compared with the original HEVC test model. The Bjøntegaard delta peak signal-to-noise ratio and Bjøntegaard delta bit rate are -0.061 dB and 1.98% on average, which is negligible. The overall performance of the proposed algorithm outperforms those of the state-of-the-art schemes. Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Hui Yuan 0001, Zhaoqing Pan, Long Xu 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Fast Coding Tree Unit depth decision for high efficiency video codingabstractHigh Efficiency Video Coding (HEVC) is the latest video coding standard, which adapts quadtree structure based Coding Tree Unit (CTU) to improve the coding efficiency. In HEVC encoding process, the CTU is recursively partitioned into coding units according to the quadtree depth. This technique increases the coding efficiency of HEVC, however, the achieved coding efficiency comes at the cost of high computational complexity. In this paper, we propose a fast C-TU quadtree depth decision algorithm to reduce the computational complexity of HEVC. Firstly, based on the best C-TU depth correlation among spatial and temporal neighboring CTUs, an early quadtree depth 0 decision algorithm is proposed. Then, according to the correlation between the prediction unit mode and the best CTU depth selection, a quadtree depth 3 skipped decision algorithm is proposed. Experimental results show that the proposed algorithm can achieve 40% on average encoding time saving, while maintaining a comparable rate-distortion performance. Zhaoqing Pan, Sam Kwong, Yun Zhang 0002, Jianjun Lei 0001, Hui Yuan 0001 |
ICIP | 1 |
| 2013 | Early termination for TZSearch in HEVC Motion EstimationabstractThe TZSearch algorithm was adopted in the high efficiency video coding reference software HM as a fast Motion Estimation (ME) algorithm for its excellent performance in reducing ME time and maintaining a comparable Rate Distortion (RD) performance. However, the multiple initial search point decision and the hybrid block matching search contribute a relatively high computational complexity to TZSearch. In this paper, based on the statistical analysis of the probability of median predictor to be selected as the final best point in the large Coding Units (CUs) (64×64, 32×32) and small CUs (16×16, 8×8) as well as the center-biased characteristic of the final best search point in ME process, we propose two early terminations for TZSearch. Experimental results show that the proposed early terminations can achieve 38.96% encoding time saving, while the RD performance degradation is quite acceptable. Zhaoqing Pan, Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Long Xu 0001 |
ICASSP | 1 |
| 2013 | Multiview Coding Mode Decision With Hybrid Optimal Stopping ModelabstractIn a generic decision process, optimal stopping theory aims to achieve a good tradeoff between decision performance and time consumed, with the advantages of theoretical decision-making and predictable decision performance. In this paper, optimal stopping theory is employed to develop an effective hybrid model for the mode decision problem, which aims to theoretically achieve a good tradeoff between the two interrelated measurements in mode decision, as computational complexity reduction and rate-distortion degradation. The proposed hybrid model is implemented and examined with a multiview encoder. To support the model and further promote coding performance, the multiview coding mode characteristics, including predicted mode probability and estimated coding time, are jointly investigated with inter-view correlations. Exhaustive experimental results with a wide range of video resolutions reveal the efficiency and robustness of our method, with high decision accuracy, negligible computational overhead, and almost intact rate-distortion performance compared to the original encoder. Tiesong Zhao, Sam Kwong, Hanli Wang, Zhou Wang 0001, Zhaoqing Pan, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 5 |