Shuai Wan

dblp:63/3024 · DBLP profile ↗
← Back
92ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0001-8617-149XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 6 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Optimized Adaptive Loop Filter Based on the Refined Adaptation Parameter Sets in H.266/VVC
abstract
Adaptive loop filter (ALF), including luma ALF, chroma ALF, and cross-component adaptive loop filter (CCALF), has been adopted in H.266/versatile video coding (VVC). It can enhance the quality of reconstructed videos based on the Wiener filtering principle. In-depth analyses reveal that the efficiency of chroma ALF and CCALF is significantly lower than that of luma ALF, primarily due to the insufficient number of chroma-related filter sets in the adaptation parameter sets (APSs). To address this limitation, this paper focuses on the efficient design of ALF APSs. Specifically, we propose a luma-guided rate-distortion optimization (RDO) criterion for chroma components to improve the efficiency of the new ALF APS. Furthermore, an improved ALF APS list management is introduced to extend the lifetime of chroma-related filter sets in the ALF APS list.
Junyan Huo, Wenjie Zou, Fuzheng Yang 0001, Shuai Wan
DCC5
2026 Robustness-aware decoupling framework for adversarial detection in remote sensing images
Yuru Su, Shaohui Mei, Shuai Wan
Pattern Recognit.4
2026 Text and Non-Text Latent Feature Disentanglement for Screen Content Image Compression
abstract
With the growing prevalence of screen content images in multimedia communication, efficient compression has become increasingly crucial. Unlike natural scene images, screen content typically contains rich text regions that exhibit unique characteristics and low correlation with surrounding non-text elements. The intricate mixture of text and non-text within images poses significant challenges for existing learned compression networks, as the text and non-text features are severely entangled in the latent domain along the channel dimension, leading to compromised reconstruction quality and suboptimal entropy estimation. In this paper, we propose a novel Disentangled Image Compression Architecture (DICA) that enhances the analysis module and the entropy model of existing compression architectures to address these limitations. First, we introduce a Disentangled Analysis Module (DAM) by augmenting original analysis modules with an additional text approximation branch and a disentangling network. They work in concert to disentangle latent features into text and non-text classes along the channel dimension, resulting in a more structured feature distribution that better aligns with compression requirements. Second, we propose a Disentangled Channel-Conditional Entropy Model (DCEM) that efficiently leverages the feature distribution bias introduced by DAM, thereby further improving compression performance. Experimental results demonstrate that the proposed DICA, along with DAM and DCEM can be integrated into various channel-conditional compression backbones, significantly improving their performance in screen content compression—particularly in hard-to-compress text regions. When integrated with an advanced WACNN backbone, our method achieves a 13% overall BD-Rate gain and a 16% BD-Rate gain in text regions on the SIQAD dataset.
Hao Wang 0184, Junyan Huo, Fei Yang 0004, Shuai Wan, Gaoxing Chen, Luis Herranz, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Deep G-PCC Geometry Preprocessing via Joint Optimization With a Differentiable Codec Surrogate for Enhanced Compression Efficiency
abstract
Geometry-based point cloud compression (G-PCC), an international standard designed by MPEG, provides a generic framework for compressing diverse types of point clouds while ensuring interoperability across applications and devices. However, G-PCC underperforms compared to recent deep learning-based PCC methods despite its lower computational power consumption. To enhance the efficiency of G-PCC without sacrificing its interoperability or computational flexibility, we propose the first compression-oriented point cloud voxelization network jointly optimized with a differentiable G-PCC surrogate model. The surrogate model mimics the rate-distortion behavior of the non-differentiable G-PCC codec, enabling end-to-end gradient propagation. The versatile voxelization network adaptively transforms input point clouds using learning-based voxelization and effectively manipulates point clouds via global scaling, fine-grained pruning, and point-level editing for rate-distortion trade-off. During inference, only the lightweight voxelization network is prepended to the G-PCC encoder, requiring no modifications to the decoder, thus introducing no computational overhead for end users. Extensive experiments demonstrate a 38.84% average BD-rate reduction over G-PCC. By bridging classical codecs with deep learning, this work offers a practical pathway to enhance legacy compression standards while preserving their backward compatibility, making it ideal for real-world deployment.
Wanhao Ma, Wei Zhang 0072, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Image Process.3
2026 Camera Motion-Conditioned Motion Estimation for Neural Video Coding for Cloud Gaming
abstract
The rising popularity of cloud gaming highlights the importance of effective video compression for this domain. Despite the potential of neural video codecs to surpass traditional codecs, their application to cloud gaming videos generally achieves limited performance primarily due to two reasons: (1) most advanced neural video codecs are designed for natural videos with only a few optimized for cloud gaming content, and (2) the distinctive characteristics of cloud gaming videos including abrupt and large camera movements, coupled with repetitive textures, make the direct application of existing codecs suboptimal. To this end, this paper introduces an enhanced neural video codec for cloud gaming with camera motion-conditioned motion estimation. By leveraging the unique feature of cloud gaming, we effectively utilize the camera motion information to condition motion estimation with an attention mechanism, overcoming the pronounced challenges of motion estimation and facilitating the accurate learning of optical flow. Furthermore, we optimize the loss function of the motion estimation network by employing a comprehensive loss across multiple feature levels, ensuring the learned optical flow is multi-faceted and well-suited to subsequent motion compensation. Extensive experiments demonstrate the effectiveness of the proposed codec, providing improved motion estimation capabilities and superior rate-distortion performance. Moreover, the robustness of our codec to rapid camera movements is validated, which makes it highly suitable for cloud gaming scenarios.
Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz
IEEE Trans. Multim.6
2025 Customizing Image Codecs for Text-Rich Screen Content with Plugin Processing Networks
abstract
With the rapid growth of remote education, telemedicine, and cloud gaming, screen content images have become prevalent in these applications. They differ significantly from natural scene images, making learning-based image codecs optimized with natural scenes inefficient when compressing them. Through empirical analysis, we observe the textual region in screen content is not only hard to compress in itself but also impacts the compression efficiency of the non-textual region. To customize the image codecs to screen content without altering their parameters, we introduced plugin pre- and post-processing modules. Specifically, we designed a filtering network in the pre-processing module to remove compression-unfriendly information from textual regions and a restoration network in the post-processing module to recover it. Additionally, we implemented a multi-scale fuse approach to enhance the high-frequency details in images. Experiments on public datasets demonstrated that our plugin solution can be seamlessly integrated into learning-based image codecs, significantly improving compression performance.
Hao Wang 0184, Junyan Huo, Shuai Wan, Gaoxing Chen, Fuzheng Yang 0001
ICME3
2025 Octree-STCM: Octree-Based Spatio-Temporal Context Model for Lossless Geometry Compression of Dynamic Point Cloud
abstract
Deep learning approaches have demonstrated remarkable effectiveness in point cloud geometry compression. However, existing octree-based methods face limitations due to insufficient contextual utilization within temporal sequences of dynamic point clouds. This paper proposes a spatio-temporal context model under an octree structure to enhance lossless compression of dynamic point cloud geometry. Firstly, a context extraction module is employed to capture the intra-contexts based on spatial correlations and the inter-contexts based on temporal dependencies. Subsequently, a context network employing 3D convolutional layers and fully connected layers is designed to extract spatio-temporal features from various contexts. After the context features integration, a multilayer perceptron is used to approximate the probability distribution of the occupancy symbol. The derived probability distributions finally optimize the arithmetic coding efficiency. Experimental results demonstrate that the proposed method outperforms the state-of-the-art octree-based approaches across multiple benchmark datasets.
Zhecheng Wang 0002, Shuai Wan, Jianqiang Huang 0002
ICMR2
2025 Enhanced neural video compression for cloud gaming videos with aligned frame generation
abstract
The burgeoning popularity of cloud gaming makes it critical for efficient video compression to relieve the growing bandwidth pressure. While existing neural video coding approaches have demonstrated strong compression potential on natural videos, there is an absence of efficient neural codecs dedicated to gaming videos. To bridge this gap, in this paper, we propose an end-to-end neural video compression method designed specifically for cloud gaming videos. By effectively utilizing the unique camera motion information inherent to cloud gaming, the previous reconstructed frame is maximally aligned to the current frame through a learningbased module with multiple losses, which then replaces the previous reconstructed frame for optical flow estimation. By significantly reducing the displacement between two consecutive frames caused by camera motion, the motion estimation accuracy is enhanced, effectively handling the large and abrupt motion scenarios frequently present in gaming videos. Furthermore, the aligned tensor obtained in the previous step is used to enhance the latent prior of the entropy model, providing a superior temporal prior for coding. Extensive experimental results demonstrate the superior performance of our proposed method compared to one of the previous state-of-the-art approaches, DCVC-HEM, providing significant progress in end-to-end neural compression in cloud gaming videos
Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz
Expert Syst. Appl.6
2025 Curriculum learning-based slimmable cross-component prediction for video coding
Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Juil Sock, Fei Yang 0004, Luis Herranz
Neurocomputing2
2025 Rendering-Oriented 3D Point Cloud Attribute Compression Using Sparse Tensor-Based Transformer
abstract
The evolution of 3D visualization techniques has fundamentally transformed how we interact with digital content. At the forefront of this change is point cloud technology, offering an immersive experience that surpasses traditional 2D representations. However, the massive data size of point clouds presents significant challenges in data compression. Current methods for lossy point cloud attribute compression (PCAC) generally focus on reconstructing the original point clouds with minimal error. However, for point cloud visualization scenarios, the reconstructed point clouds with distortion still need to undergo a complex rendering process, which affects the final user-perceived quality. In this paper, we propose an end-to-end deep learning framework that seamlessly integrates PCAC with differentiable rendering, denoted as rendering-oriented PCAC (RO-PCAC), directly targeting the quality of rendered multiview images for viewing. In a differentiable manner, the impact of the rendering process on the reconstructed point clouds is taken into account. Moreover, we characterize point clouds as sparse tensors and propose a sparse tensor-based transformer, called SP-Trans. By aligning with the local density of the point cloud and utilizing an enhanced local attention mechanism, SP-Trans captures the intricate relationships within the point cloud, further improving feature analysis and synthesis within the framework. Extensive experiments demonstrate that the proposed RO-PCAC achieves state-of-the-art compression performance, compared to existing reconstruction-oriented methods, including traditional, learning-based, and hybrid methods. The code will be released athttps://github.com/net-F/RO-PCAC.git.
Xiao Huo, Junhui Hou, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Adaptive Enhanced Global Intra Prediction for Efficient Video Coding in Beyond VVC
abstract
Global intra prediction (GIP), including intra-block copy and template matching prediction (TMP), exploits the global correlation of the same image to improve the coding efficiency. In Beyond VVC, TMP uses template matching to determine the reference blocks for efficient prediction. There usually exists an error between the coding block and reference blocks, caused by the content mismatch or the coding distortion of the reference blocks. We propose an enhancement over the reference blocks, namely enhanced GIP (EGIP). Specifically, we design an enhanced filter according to the templates of the coding block and the reference blocks, with the reconstructed template of the coding block as the label for supervised learning. To support different enhancements, we design two types of inputs, i.e., EGIP based on neighboring samples (N-EGIP) and EGIP based on multiple hypothesis references (M-EGIP). Experimental results show that, based on enhanced compression model (ECM) version 8.0, N-EGIP achieves BD-rate reductions of 0.37%, 0.42%, and 0.40%, and M-EGIP brings 0.34%, 0.37%, and 0.34% BD-rate savings for Y, Cb, and Cr components, respectively. A higher coding gain, 0.46%, 0.54%, and 0.52% BD-rate savings, can be achieved by integrating N-EGIP and M-EGIP together. Owing to the coding gain and small complexity increase, the proposed EGIP has been adopted in the exploration of Beyond VVC and integrated into its reference software.
Junyan Huo, Yanzhuo Ma, Zhenyao Zhang, Hui Yuan 0001, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 RBFIM: Perceptual Quality Assessment for Compressed Point Clouds Using Radial Basis Function Interpolation
abstract
One of the main challenges in point cloud compression (PCC) is how to evaluate the perceived distortion so that the codec can be optimized for perceptual quality. Current standard practices in PCC highlight a primary issue: while single-feature metrics are widely used to assess compression distortion, the classic method of searching point-to-point nearest neighbors frequently fails to adequately build precise correspondences between point clouds, resulting in an ineffective capture of human perceptual features. To overcome the related limitations, we propose a novel assessment method called RBFIM, utilizing radial basis function (RBF) interpolation to convert discrete point features into a continuous feature function for the distorted point cloud. By substituting the geometry coordinates of the original point cloud into the feature function, we obtain the bijective sets of point features. This enables an establishment of precise corresponding features between distorted and original point clouds and significantly improves the accuracy of quality assessments. Moreover, this method avoids the complexity caused by bidirectional searches. Extensive experiments on multiple subjective quality datasets of compressed point clouds demonstrate that our RBFIM excels in addressing human perception tasks, thereby providing robust support for PCC optimization efforts.
Shuai Wan, Fuzheng Yang 0001, Mengting Yu, Junhui Hou
IEEE Trans. Multim.2
2025 Lightweight Deep Exemplar Colorization via Semantic Attention-Guided Laplacian Pyramid
abstract
Exemplar-based colorization aims to generate plausible colors for a grayscale image with the guidance of a color reference image. The main challenging problem is finding the correct semantic correspondence between the target image and the reference image. However, the colors of the object and background are often confused in the existing methods. Besides, these methods usually use simple encoder-decoder architectures or pyramid structures to extract features and lack appropriate fusion mechanisms, which results in the loss of high-frequency information or high complexity. To address these problems, this article proposes a lightweight semantic attention-guided Laplacian pyramid network (SAGLP-Net) for deep exemplar-based colorization, exploiting the inherent multi-scale properties of color representations. They are exploited through a Laplacian pyramid, and semantic information is introduced as high-level guidance to align the object and background information. Specially, a semantic guided non-local attention fusion module is designed to exploit the long-range dependency and fuse the local and global features. Moreover, a Laplacian pyramid fusion module based on criss-cross attention is proposed to fuse high frequency components in the large-scale domain. An unsupervised multi-scale multi-loss training strategy is further introduced for network training, which combines pixel loss, color histogram loss, total variance regularisation, and adversarial loss. Experimental results demonstrate that our colorization method achieves better subjective and objective performance with lower complexity than the state-of-the-art methods.
Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Marta Mrak, Juil Sock, Fei Yang 0004, Luis Herranz
IEEE Trans. Vis. Comput. Graph.2
2024 Adaptive Chroma Block Vector Derivation from Luma for Screen Content Coding
abstract
Intra Block Copy (IBC) and Intra Template Matching Prediction (IntraTMP) are two efficient algorithms to sufficiently exploit the correlation in the same picture. Block Vector (BV) is used to represent the displacement between the current block and its reference within the same picture. The BV information of luma can be employed to help the chroma coding efficiently. Based on this feature, an adaptive chroma prediction is proposed to derive the BV of the chroma block from the luma. Two strategies are designed to improve the coding performance, including multiple positions’ check and template-based BV refinement. Compared with Enhanced Compression Model (ECM) of beyond VVC, 0.43%, 0.35%, and 0.60% BD-rate savings for Y, Cb, and Cr components are achieved for Class F, and 2.23%, 2.31%, and 2.93% BD-rate savings are provided for Class TGM. We also integrated the proposed method into the VVC Test Model (VTM). A similar coding improvement can be observed. Due to the coding gain and low complexity, the proposed method has been adopted into the beyond VVC exploration and integrated into the latest version of ECM.
Junyan Huo, Xue Hao, Shuai Wan, Fuzheng Yang 0001
ICASSP3
2024 Learning-Based Video Compression with Continuously Variable Bitrate Coding
abstract
In this paper, we propose a learning-based video compression which can perform continuously variable bitrate coding. The proposed method generates feature transformation parameters through a conditional network according to the input spatial quality map. These parameters are then used to adaptively transform the intermediate features of the encoder, decoder, and spatiotemporal entropy model in the codec, thus enabling variable bitrate coding. Additionally, to improve the compression efficiency of the codec, we propose incorporating the quality map of the preceding frame into the hyperprior encoder and leveraging the temporal prior encoder. A multi-stage training strategy is employed to jointly train the codec with a multi-frame rate-distortion loss function. The experimental results demonstrate that the proposed method can achieve continuously variable bitrate adaptation while maintaining rate-distortion performance comparable to the fixed bitrate model. Furthermore, the proposed method also supports ROI-based compression.
Mingyi Yang, Xionghui Mao, Yujie Yin, Defa Wang, Shuai Wan, Fuzheng Yang 0001
ICIP6
2024 Feature Decoupling Based Adversarial Examples Detection Method for Remote Sensing Scene Classification
abstract
Deep Neural Networks (DNNs) have demonstrated remarkable effectiveness in remote sensing (RS) image processing. However, they remain vulnerable to adversarial examples, which are generated by adding tiny but purposeful perturbations to clean examples. Such vulnerabilities in critical applications like environmental monitoring and urban planning can lead to significant negative consequences. To mitigate the interference of adversarial examples on DNNs, in this paper, a feature decoupling based adversarial examples detection (FD-AED) method for RS images is proposed, where non-robust features are employed in the detection process. Specifically, a loss function is designed for the de-coupler to disentangle the features into robust and non-robust features. Non-robust features are particularly useful because they often contain subtle clues that distinguish between clean and adversarial examples. By focusing on these non-robust features, the adversarial example detector can more effectively capture the differences between clean and adversarial examples. Experimental results indicate that the proposed FD-AED method effectively decouples robust and non-robust features, achieving more precise and reliable detection of adversarial examples.
Yuru Su, Shaohui Mei, Shuai Wan
IGARSS3
2024 Improving Optimal Binarization with Update On-the-fly in G-PCC Entropy Coding: Probability Initialization and Adaptive Bounds Setting for Context Models
abstract
Geometry-based point cloud compression (G-PCC) uses Context-based Adaptive Binary Arithmetic Coding to encode the geometry and attribute information. The context information is built in context models for entropy coding. G-PCC adopts the Optimal Binarization with Update On-the-fly (OBUF) to reduce the number of context models. In the current design, however, the probability initialization for both fine-and coarse-grained contexts does not follow the principle of entropy continuation. Moreover, the mapping process to produce coarse-grained contexts is a combination of several fine-grained contexts, leading to an unstable update of probability for coarse-grained contexts, which affects the accuracy of the fine-grained context model in probability estimation.To address the underlying problems, we propose two approaches to improve OBUF: initializing the probabilities for fine-grained and coarse-grained contexts according to entropy continuation and setting the probability update upper and lower bounds for coarse-grained contexts adaptively. The experimental results demonstrate that the proposed technique is more consistent with the underlying principles of OBUF and significantly improves the performance of both octree-based and Trisoup-based geometry coding. Due to the theoretical consistency and outstanding performance, the proposed methods have been adopted into the state-of-the-art G-PCC.
Shidi Hao, Shuai Wan, Tengya Tian, Wei Zhang 0072, Fuzheng Yang 0001
ISCAS2
2024 Rate Control for Slimmable Video Codec using Multilayer Perceptron
abstract
As a flexible design for practical end-to-end video compression, slimmable video coding achieves variable rate coding with an adaptable complexity. In this paper, rate control using a multilayer perceptron is proposed to achieve accurate bitrate control for slimmable video coding. An average bitrate error of less than 0.3% is achieved on tested sequences at a slight decrease in the rate-distortion performance. Furthermore, the results show that there is no overflow or underflow in the buffer occupancy during the rate control process.
Defa Wang, Shuai Wan, Fei Yang 0004, Luis Herranz
ISCAS3
2024 Analysis of Relationship between Point Cloud Just Noticeable Difference and Attribute Quantization Parameters
abstract
Utilizing just noticeable difference(JND) thresholds to describe distortion visibility enables efficient transmission and storage of point clouds while maintaining high quality. Therefore, describing the relationship between point cloud JND and attribute quantization parameters (QP) is essential. This relationship is ,however, not clear regarding point cloud compression so far. We investigate the correlation between JND and attribute QP using Geometry-based Point Cloud Compression(G-PCC), which is the latest standard for point cloud compression. Moreover, we provide the corresponding datasets in this work. Using the K-means clustering algorithm, we partition the collected raw data into intervals and identify the peaks closest to the means in each sub-interval as the desired JND points. For human-content point clouds, although the number of JND varies, the QP values corresponding to different levels of JND show a regular distribution. The study focuses on the QP corresponding to the starting JND point(JNDS) and the ending JND point (JNDE).The findings indicate that the critical points for human-content point clouds are relatively consistent, with JNDSmainly concentrated at QP = 27 and JNDEat QP = 45. Conversely, perception in object-content point clouds is more influenced by their content, and the distribution of JND is not as uniform as in human-content point clouds. The experimental datasets and research results are accessible at the following link: https://github.com/ZhangChen2022/JND-attribute-QP-on-G-PCC.
Luqian Bai, Mengting Yu, Shuai Wan, Hejie Yang
VCIP4
2024 A slimmable framework for practical neural video compression
abstract
Deep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs.
Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz
Neurocomputing6
2024 Task-Switchable Pre-Processor for Image Compression for Multiple Machine Vision Tasks
abstract
Visual content is increasingly being processed by machines for various automated content analysis tasks instead of being consumed by humans. Despite the existence of several compression methods tailored for machine tasks, few consider real-world scenarios with multiple tasks. In this paper, we aim to address this gap by proposing a task-switchable pre-processor that optimizes input images specifically for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. The proposed task-switchable pre-processor adeptly maintains relevant semantic information based on the specific characteristics of different downstream tasks, while effectively suppressing irrelevant information to reduce bitrate. To enhance the processing of semantic information for diverse tasks, we leverage pre-extracted semantic features to modulate the pixel-to-pixel mapping within the pre-processor. By switching between different modulations, multiple tasks can be seamlessly incorporated into the system. Extensive experiments demonstrate the practicality and simplicity of our approach. It significantly reduces the number of parameters required for handling multiple tasks while still delivering impressive performance. Our method showcases the potential to achieve efficient and effective compression for machine vision tasks, supporting the evolving demands of real-world applications.
Mingyi Yang, Fei Yang 0004, Luka Murn, Marc Gorriz, Juil Sock, Shuai Wan, Fuzheng Yang 0001, Luis Herranz
IEEE Trans. Circuits Syst. Video Technol.6
2024 Chroma Intra Prediction With Lightweight Attention-Based Neural Networks
abstract
Neural networks can be successfully used for cross-component prediction in video coding. In particular, attention-based architectures are suitable for chroma intra prediction using luma information because of their capability to model relations between difierent channels. However, the complexity of such methods is still very high and should be further reduced, especially for decoding. In this paper, a cost-effective attention-based neural network is designed for chroma intra prediction. Moreover, with the goal of further improving coding performance, a novel approach is introduced to utilize more boundary information effectively. In addition to improving prediction, a simplification methodology is also proposed to reduce inference complexity by simplifying convolutions. The proposed schemes are integrated into H.266/Versatile Video Coding (VVC) pipeline, and only one additional binary block-level syntax flag is introduced to indicate whether a given block makes use of the proposed method. Experimental results demonstrate that the proposed scheme achieves up to −0.46%/−2.29%/−2.17% BD-rate reduction on Y/Cb/Cr components, respectively, compared with H.266/VVC anchor. Reductions in the encoding and decoding complexity of up to 22% and 61%, respectively, are achieved by the proposed scheme with respect to the previous attention-based chroma intra prediction method while maintaining coding performance.
Chengyi Zou, Shuai Wan, Tiannan Ji, Marc Gorriz, Marta Mrak, Luis Herranz
IEEE Trans. Circuits Syst. Video Technol.2
2024 Near-Lossless Compression of Point Cloud Attribute Using Quantization Parameter Cascading and Rate-Distortion Optimization
abstract
Near-lossless compression of point clouds is suitable for the application scenarios with low distortion tolerance and certain requirements on the rate. Near-lossless attribute compression usually adopts a level-of-detail structure, where the dependencies between the layers make it possible to improve the rate-distortion (R-D) performance by using different quantization parameters for different layers. In this work, a theoretical analysis of the dependencies between adjacent layers is carried out, based on which the dependent Distortion-Quantization and Rate-Quantization models are established for point cloud attribute compression. Then an algorithm for quantization parameter cascading based on R-D optimization is proposed and implemented for near-lossless compression of point cloud attributes. The experimental results show that the proposed method has a superior performance gain compared to state-of-the-art for the Hausdorff R-D performance. At the same time, the proposed method improves subjective quality and is well adapted to various categories of point clouds.
Shuai Wan, Zhecheng Wang 0002, Fuzheng Yang 0001
IEEE Trans. Multim.2
2024 Joint Rate-Distortion Optimization for Video Coding and Learning-Based In-Loop Filtering
abstract
Learning-based in-loop filters (ILFs) have recently been widely deployed in the video codec to remove compression artifacts and to obtain better-quality reconstructed videos. However, in the existing codec, the impact of the learning-based ILF is not considered in the Rate-Distortion optimization (RDO) process. With the learning-based ILF, the set of coding parameters selected by the conventional RDO process may no longer be the best one, and the best overall Rate-Distortion (R-D) performance can not be guaranteed. In this article, we propose a joint RDO (JRDO) for Video Coding and learning-based in-loop filtering, which incorporates the effect of the learning-based ILF on the reconstructed video into the RDO process, aiming to achieve the best overall R-D performance of the reconstructed video after in-loop filtering. Furthermore, to realize the proposed JRDO in a standardized video codec, we propose practical strategies to efficiently estimate the effect of learning-based ILF during the RDO process, i.e., efficiently estimate the distortion of the reconstructed block after in-loop filtering during the RDO process. Extensive experiments demonstrate that the proposed joint RDO is standard-compliant and can improve the R-D performance without increasing the decoding time. Besides, the superiority of joint RDO is achieved in various ILFs, indicating the generality of the proposed work.
Mingyi Yang, Junyan Huo, Xile Zhou, Wenhan Qiao, Shuai Wan, Hao Wang 0184, Fuzheng Yang 0001
IEEE Trans. Multim.5
2023 Efficient Super-Resolution for Compression Of Gaming Videos
abstract
Due to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming videos at low bitrates. Most state-of-the-art super-resolution networks generalize over a variety of RGB inputs and use a unified network architecture for frames of different levels of degradation, leading to high complexity and redundancy. Since games usually consist of a limited number of fixed scenarios, we specialize one model for each scenario and assign appropriate network capacities for different QPs to perform super-resolution under the guidance of reconstructed high-quality luma components. Experimental results show that our framework achieves a superior quality-complexity trade-off compared to the ESRnet baseline, saving at most 93.59% parameters while maintaining comparable performance. The compression efficiency compared to HEVC is also improved by more than 17% BD-rate gain.
Luka Murn, Luis Herranz, Fei Yang 0004, Marta Mrak, Wei Zhang 0072, Shuai Wan, Marc Gorriz
ICASSP7
2023 Semantic Preprocessor for Image Compression for Machines
abstract
Visual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision.
Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak
ICASSP6
2023 Adaptive Geometry Reconstruction for Geometry-based Point Cloud Compression
abstract
Since the geometry constitutes most of the bitrate and is used for attribute coding, it is crucial for geometry-based point cloud compression (G-PCC). However, the current research focuses on geometry coding while ignoring reconstruction. In G-PCC, the reconstructed points are located at the center of the quantization nodes, which may not match the surface of the point clouds. Therefore, we first estimate the normal direction. Then, the offset direction of the reconstructed point is determined by considering its adjacent points’ occupancy and attributes. Finally, the position of the reconstructed point is adjusted taking into account both the normal and offset directions. The method aids in both objective and subjective quality. Experimental results demonstrate that the proposed method outperforms the state-of-the-art G-PCC. It has significant performance gains in point-to-point and point-to-plane errors, 4.5% and 9.0% on average, respectively. It also has a minor performance gain in attribute coding.
Shuai Wan, Xiaobin Ding, Fuzheng Yang 0001, Zhecheng Wang 0002
ICME2
2023 Impact of Geometry and Attribute Distortion in Subjective Quality of Point Clouds for G-PCC
abstract
To better understand the effect of geometry and attribute distortion on the subjective quality of point clouds, we built a database with subjective quality scores for compressed point clouds. We analyzed the impact of geometry and attribute distortion to subjective quality. Our database, named the NWPU Point Cloud Quality Database (NWPU-PCQD), contains 817 point clouds with various geometry and attribute quantification distortions, which were subjectively evaluated using head-mounted displays (HMDs). Through correlation and significance analysis of the subjective quality scores, we observe that geometry and attribute quantization do not equally contribute to the overall subjective quality, where the effect of geometry quantization dominates. Moreover, there is an interaction between the geometry and attribute distortion. Our subjective database and findings are useful for point cloud processing, transmission, and compression.
Shuai Wan, Fuzheng Yang 0001
VCIP2
2023 The Complexity Optimization of Dependent Quantization in VVC By Delaying Starting Point
abstract
Compared with the video compression standard High Efficiency Video Coding (HEVC), the latest standard Versatile Video Coding (VVC) has achieved great improvement in coding efficiency due to the application of more sophisticated coding tools, which also results in high coding complexity. In this paper, to reduce the computational complexity of dependent quantization in VVC, an optimization method based on theoretical analysis of rate-distortion performance and statistical low-complexity model establishment is proposed to delay the starting point of the quantized path. Experimental results demonstrate that the proposed method can postpone the starting point to a position closer to the last non-zero quantized coefficient, reducing 3.57% encoding time, with a 0.42% increase in Bjφntegaard delta bit-rate (BD-BR) compared with VVenC 1.0.0 under medium preset.
Shixuan Feng, Luge Wang, Fuzheng Yang 0001, Shuai Wan
VCIP5
2023 Optimization of octree-based adaptive geometry quantization via up-sampling for G-PCC
abstract
To improve the reconstructed point cloud after adaptive geometry quantization in geometry-based point cloud compression, a least squares plane (LSP) projection-based up-sampling method and a quantization parameter (QP) decision method based on loss function are proposed. First, the LSP fitting is carried out to locate the interpolated point based on the nearest neighbors of the current node during decoding, enhancing both the subjective and objective quality of the reconstructed point cloud. Second, the QP decision for each node is based on the mean squared error between the original point cloud and the reconstructed point cloud. The experimental results show that the proposed methods achieve performance gains in terms of point-to-point and point-to-plane errors for geometry by 6.3% and 1.6%, respectively, and for attributes by 1.5%, 0.7%, and 0.5%. There also has been a significant improvement in subjective quality.
Shuai Wan, Xiaobin Ding, Zhecheng Wang 0002
VCIP2
2023 Multi-scale deep feature fusion based sparse dictionary selection for video summarization
Mingyang Ma 0004, Shuai Wan, Xiuxiu Han, Shaohui Mei
Signal Process. Image Commun.3
2023 Local Geometry-Based Intra Prediction for Octree-Structured Geometry Coding of Point Clouds
abstract
Point cloud compression (PCC) is crucial for efficient and flexible storage as well as feasible transmission of point clouds in practice. For geometry compression, one popular approach is the octree-based solution. The intra prediction mechanism utilizes the spatial correlation in the static point cloud to predict the occupancy bit of the octree node for entropy coding, reducing the spatial redundancy. In this study, two local geometry-based prediction methods are proposed following statistical and theoretical analyses: binary prediction, which outputs the binary state (i.e., occupied or unoccupied), and ternary prediction, which provides a third option other than occupied or unoccupied (i.e., not predicted). In comparison to the state-of-the-art, the proposed binary prediction offers the Bjontegaard delta rate (BD-rate) of −0.8% for lossy compression and the bits per input point (bpip) of 100.09% for lossless compression in average, respectively. The binary prediction reduces the computational complexity in terms of more than 20% decrease in decoding time. In particular, it also provides noticeable reduction of the memory usage during entropy coding. The proposed ternary prediction provides −1.2% BD-rate for lossy compression and 97.19% bpip for lossless compression in average, respectively, in comparison to the state-of-the-art. While achieving performance gain, it is considerably more computational efficient by saving about 18% decoding time. Due to these advantages, part of the proposed ternary prediction has been adopted by the ongoing MPEG standard of geometry-based point cloud compression (G-PCC).
Zhecheng Wang 0002, Shuai Wan
IEEE Trans. Circuits Syst. Video Technol.2
2023 Reconstruction-Assisted and Distance-Optimized Adversarial Training: A Defense Framework for Remote Sensing Scene Classification
abstract
Despite deep neural networks (DNNs) have been widely applied in remote sensing (RS) scene classification and achieved satisfying performance, the vulnerability of DNNs towards adversarial examples significantly degrades their performance. Moreover, the relatively limited labeled samples of RS scene classification make DNNs more likely to overfit, leading to weak generalizability and noise sensitivity. This may result in DNNs being more vulnerable to adversarial examples. Consequently, the defense of adversarial examples is of crucial importance to improve both the generalizability and robustness of DNNs in the RS scene classification task. However, few studies have been conducted on defense for RS scene classification, especially ignoring the intrinsic characteristics of RS images. In this paper, an effective defense framework for RS scene classification, named reconstruction-assisted and distance-optimized adversarial training (RDAT), is proposed to defend adversarial examples. In order to solve the problems caused by high interclass similarity, a distance-optimized (DO) strategy is designed for adversarial training to strengthen the learning of underfitting content, increase the interclass distance, and improve the robustness of the networks. Furthermore, in order to generate high quality samples for adversarial training, a reconstruction-assisted (RA) block is proposed to eliminate adversarial perturbations in adversarial examples. Specifically, in this block, by swin transformer (SwinT) block and multi-scale convolution (MSC) block, SwinT-MSC-UNet (SMUNet) is constructed to fully extract global and multi-scale local features to adapt to the characteristics of RS images with large variance of ground object scales. Extensive experiments on the benchmark datasets, i.e., UC Merced (UCM) and Aerial Image Dataset (AID), have demonstrate that the proposed RDAT can effectively resist multiple adversarial attacks and yield superior results than other defense methods for RS scene classification.
Yuru Su, Ge Zhang 0006, Shaohui Mei, Jiawei Lian, Ye Wang 0020, Shuai Wan
IEEE Trans. Geosci. Remote. Sens.6
2023 Adaptive Chroma Prediction Based on Luma Difference for H.266/VVC
abstract
Cross-component chroma prediction plays an important role in improving coding efficiency for H.266/VVC. We use the differences between reference samples and the predicted sample to design an attention model for chroma prediction, namely luma difference-based chroma prediction (LDCP). Specifically, the luma differences (LDs) between reference samples and the predicted sample are employed as the input of the attention model, which is designed as a softmax function to map LDs to chroma weights nonlinearly. Finally, a weighted chroma prediction is conducted based on the weights and chroma reference samples. To provide adaptive weights, the model parameter of the softmax function can be determined based on the template (T-LDCP) or offline learning (L-LDCP), respectively. Experimental results show that the T-LDCP achieves BD-rate reductions of 0.34%, 2.02%, and 2.34% for the Y, Cb, and Cr components, and the L-LDCP brings 0.32%, 2.06%, and 2.21% BD-rate savings for Y, Cb, and Cr components, respectively. The L-LDCP introduces slight encoding and decoding time increments, i.e., 2% and 1%, when integrated into the latest VVC test model version 18.0. Besides, the LDCP can be implemented by a pixel-level parallelization which is hardware-friendly.
Junyan Huo, Danni Wang, Hui Yuan 0001, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Image Process.4
2022 Unified Matrix Coding for NN Originated MIP in H.266/VVC
abstract
Matrix-based Intra Prediction (MIP) is an effective coding algorithm in H.266/Versatile Video Coding (VVC) which is originated by Neural Networks (NN). With the requirement of low complexity, MIP is conducted by a matrix-vector multiplication. To handle with the diversity of video content, 30 matrices are trained and stored to derive predicted samples. Since matrices from training are usually floating-point values, which should be avoided in H.266/VVC, two parameters, shift and offset, are introduced for each matrix to convert floating-point values to integers. This paper designs an efficient algorithm to determine the input vector of MIP, with which the range of the matrices can be minimized, and all matrices can be converted to integers with a unified shift and a unified offset. The proposed algorithm removes the matrix-dependent parameters for integer conversion and saves the memory for storing MIP parameters. Experimental results demonstrate that the proposed algorithm has a similar coding performance with VVC reference software. Due to the unified operation, memory reduction, and no coding loss, the proposed algorithm has been adopted into H.266/VVC.
Junyan Huo, Shuai Wan, Fuzheng Yang 0001
ICASSP4
2022 DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed Video
abstract
In this paper, we propose a deformable convolution-based generative adversarial network (DCNGAN) for perceptual quality enhancement of compressed videos. DCNGAN is also adaptive to the quantization parameters (QPs). Compared with optical flows, deformable convolutions are more effective and efficient to align frames. Deformable convolutions can operate on multiple frames, thus leveraging more temporal information, which is beneficial for enhancing the perceptual quality of compressed videos. Instead of aligning frames in a pairwise manner, the deformable convolution can process multiple frames simultaneously, which leads to lower computational complexity. Experimental results demonstrate that the proposed DCNGAN outperforms other state-of-the-art compressed video quality enhancement algorithms.
Saiping Zhang, Luis Herranz, Marta Mrak, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001
ICASSP5
2022 Towards Lightweight Neural Network-based Chroma Intra Prediction for Video Coding
abstract
In video compression the luma channel can be useful for predicting chroma channels (Cb, Cr), as has been demonstrated with the Cross-Component Linear Model (CCLM) used in Versatile Video Coding (VVC) standard. More recently, it has been shown that neural networks can even better capture the relationship among different channels. In this paper, a new attention-based neural network is proposed for cross-component intra prediction. With the goal to simplify neural network design, the new framework consists of four branches: boundary branch and luma branch for extracting features from reference samples, attention branch for fusing the first two branches, and prediction branch for computing the predicted chroma samples. The proposed scheme is integrated into VVC test model together with one additional binary block-level syntax flag which indicates whether a given block makes use of the proposed method. Experimental results demonstrate 0.31%/2.36%/2.00% BD-rate reductions on Y/Cb/Cr components, respectively, on top of the VVC Test Model (VTM) 7.0 which uses CCLM.
Chengyi Zou, Shuai Wan, Marta Mrak, Marc Gorriz, Luis Herranz, Tiannan Ji
ICIP2
2022 Logistic Regression Guided Coding of Single Child Mode for Point Cloud Geometry Compression
abstract
Geometry coding in geometry-based point cloud compression (G-PCC) is octree-structured, including a bitwise occupancy mode for a general case, and a single child mode for a node containing a single occupied child node. However, the current usage of the single child mode is limited because of the strict eligibility determination based on neighboring nodes. Context modeling is also missing for entropy coding of the coordinate index of the single occupied child node relative to the node. Guided by logistic regression (LR), this paper first proposes an algorithm to determine the eligibility of a node for the single child mode. Without resorting to the occupancy of the neighboring nodes, the proposed algorithm provides more opportunities for employing the single child mode. In addition, LR is also used in predicting the relative coordinate index of the single occupied child node. Based on the analysis of predicted results, we model contexts for the entropy coding of the single child mode. Experiments reveal that the proposed method improves the existing single child mode in G-PCC with overall coding gain in terms of bit per input point (bpip). Besides, the proposed method also saves coding time.
Zhecheng Wang 0002, Shuai Wan
PCS2
2022 Geometry Reconstruction for Spatial Scalability in Point Cloud Compression Based on the Prediction of Neighbours' Weights
abstract
Spatial scalability is a critical feature in geometrybased point cloud compression (G-PCC). The current design of geometry reconstructions for spatial scalability applies points in fixed positions (center of nodes) and ignores the connection of points in regions. This work analyses the correlation between neighbours' occupancy and locally optimal reconstruction points within a node using the Pearson Product Moment Correlation Coefficient (PPMCC). Then we propose a geometry reconstruction method based on predicting the neighbours' weights. Geometry reconstruction points are calculated by applying weights inverse to distance to different categories of neighbours (face neighbours, edge neighbours, corner neighbours). Compared to the state-of-the-art G-PCC, performance improvement of 1.03dB in D1-PSNR and 2.90dB in D2-PSNR, on average, can be observed using the proposed method. Meanwhile, a simplified method is available to satisfy different complexity requirements.
Shuai Wan
VCIP2
2022 Rate Controllable Learned Image Compression Based on RFL Model
abstract
In this paper, we propose a rate controllable image compression framework, Rate Controllable Variational Autoencoder (RC-VAE), based on the Rate-Feature-Level (RFL) model established through our exploration on the correlation among target rates, image features and quantization levels. Considering that, when meeting the same target rate, different images should be quantized in different levels, we focus on jointly utilizing the target rate and the extracted features of the image to predict the corresponding quantization level and propose the RFL model. Combining the proposed RFL model with a Hyperprior Continuously Variable Rate (HCVR) image compression network, we further propose the RC-VAE. By controlling information loss in quantization process, the RC-VAE can work at the target rate. Experimental results have demonstrated that one single RC-VAE model can adapt to multiple target rates with higher rate control accuracy and better R-D performance compared with the state-of-the-art rate controllable Image compression networks.
Saiping Zhang, Luge Wang, Xionghui Mao, Fuzheng Yang 0001, Shuai Wan
VCIP5
2022 Unified Cross-Component Linear Model in VVC Based on a Subset of Neighboring Samples
abstract
To compress industrial video content efficiently, H.266/Versatile Video Coding (VVC) introduces cross-component linear model (CCLM) prediction as a new coding tool, in which chroma components are predicted from the luma component based on a linear model. In this article, we propose a subset-based CCLM (S-CCLM), in which the model parameters are derived based on a subset of neighboring samples. To choose the most proper subset, we build the relationship between the prediction error and the geometric distance and resolve the optimal subset construction problem by minimizing the geometric distance. With the well-designed subset, a weight-guided parameter derivation algorithm is further proposed to improve the accuracy of the model parameters. The experimental results show that the proposed S-CCLM can achieve Bjontegaard delta bitrate (BD-rate) reductions of 0.14%, 0.64%, and 0.75% for the Y, Cb, and Cr components, respectively, when the number of samples in the subset,$N$, is 4 and BD-rate reductions of 0.22%, 0.80%, and 0.95% when$N$is 8. Given a small fixed$N$, fewer memory access operations are needed during the CCLM calculation, and a unified CCLM process can be achieved for coding blocks with different sizes and different modes. Due to its hardware-friendly architecture, the S-CCLM has been partially adopted by H.266/VVC.
Junyan Huo, Hongqing Du, Shuai Wan, Hui Yuan 0001, Yanzhuo Ma, Fuzheng Yang 0001
IEEE Trans. Ind. Informatics4
2022 Graph Convolutional Dictionary Selection With L₂, ₚ Norm for Video Summarization
abstract
Video Summarization (VS) has become one of the most effective solutions for quickly understanding a large volume of video data. Dictionary selection with self representation and sparse regularization has demonstrated its promise for VS by formulating the VS problem as a sparse selection task on video frames. However, existing dictionary selection models are generally designed only for data reconstruction, which results in the neglect of the inherent structured information among video frames. In addition, the sparsity commonly constrained by$L_{2,1}$norm is not strong enough, which causes the redundancy of keyframes, i.e., similar keyframes are selected. Therefore, to address these two issues, in this paper we propose a general framework called graph convolutional dictionary selection with$L_{2,p}$($0< p\leq 1$) norm (GCDS$_{2,p}$) for both keyframe selection and skimming based summarization. Firstly, we incorporate graph embedding into dictionary selection to generate the graph embedding dictionary, which can take the structured information depicted in videos into account. Secondly, we propose to use$L_{2,p}$($0< p\leq 1$) norm constrained row sparsity, in which$p$can be flexibly set for two forms of video summarization. For keyframe selection,$0< p< 1$can be utilized to select diverse and representative keyframes; and for skimming,$p=1$can be utilized to select key shots. In addition, an efficient iterative algorithm is devised to optimize the proposed model, and the convergence is theoretically proved. Experimental results including both keyframe selection and skimming based summarization on four benchmark datasets demonstrate the effectiveness and superiority of the proposed method.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Xian-Sheng Hua 0001, David Dagan Feng
IEEE Trans. Image Process.3
2021 DVC-P: Deep Video Compression with Perceptual Optimizations
abstract
Recent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual op-timizations (DVC-P), which aims at increasing perceptual quality of decoded videos. Our proposed DVC-P is based on Deep Video Compression (DVC) network, but improves it with perceptual optimizations. Specifically, a discriminator network and a mixed loss are employed to help our network trade off among distortion, perception and rate. Furthermore, nearest-neighbor interpolation is used to eliminate checkerboard artifacts which can appear in sequences encoded with DVC frameworks. Thanks to these two improvements, the perceptual quality of decoded sequences is improved. Experimental results demonstrate that, compared with the baseline DVC, our proposed method can generate videos with higher perceptual quality achieving 12.27% reduction in a perceptual BD- rate equivalent, on average.
Saiping Zhang, Marta Mrak, Luis Herranz, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001
VCIP5
2021 Intra prediction based on geometry padding for omnidirectional video coding
Shuai Wan
Multim. Tools Appl.2
2021 Similarity Based Block Sparse Subset Selection for Video Summarization
abstract
Video summarization (VS) is generally formulated as a subset selection problem where a set of representative keyframes or key segments is selected from an entire video frame set. Though many sparse subset selection based VS algorithms have been proposed in the past decade, most of them adopt linear sparse formulation in the explicit feature vector space of video frames, and don’t consider the local or global relationships among frames. In this paper, we first extend the conventional sparse subset selection for VS into kernel block sparse subset selection (KBS3) to utilize the advantage of kernel sparse coding and introduce a local inter-frame relationship through packing of frame blocks. Going a step further, we propose a similarity based block sparse subset selection (SB2S3) model by applying a specially designed transformation matrix on the KBS3 model in order to introduce a kind of global inter-frame relationship through the similarity. Finally, a greedy pursuit based algorithm is devised for the proposed NP-hard model optimization. The proposed SB2S3 has the following advantages: 1) through the similarity between each frame and any other frame, the global relationship among all frames can be considered; 2) through block sparse coding, the local relationship of adjacent frames is further considered; and 3) it has a wider application, since features can derive similarity, but not vice versa. It is believed that the effect of modeling such global and local relationships among frames in this paper, is similar to that of modeling the long-range and short-range dependencies among frames in deep learning based methods. Experimental results on three benchmark datasets have demonstrated that the proposed approach is superior to not only other sparse subset selection based VS methods but also most unsupervised deep-learning based VS methods.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng, Mohammed Bennamoun
IEEE Trans. Circuits Syst. Video Technol.3
2021 Keyframe Extraction From Laparoscopic Videos via Diverse and Weighted Dictionary Selection
abstract
Laparoscopic videos have been increasingly acquired for various purposes including surgical training and quality assurance, due to the wide adoption of laparoscopy in minimally invasive surgeries. However, it is very time consuming to view a large amount of laparoscopic videos, which prevents the values of laparoscopic video archives from being well exploited. In this paper, a dictionary selection based video summarization method is proposed to effectively extract keyframes for fast access of laparoscopic videos. Firstly, unlike the low-level feature used in most existing summarization methods, deep features are extracted from a convolutional neural network to effectively represent video frames. Secondly, based on such a deep representation, laparoscopic video summarization is formulated as a diverse and weighted dictionary selection model, in which image quality is taken into account to select high quality keyframes, and a diversity regularization term is added to reduce redundancy among the selected keyframes. Finally, an iterative algorithm with a rapid convergence rate is designed for model optimization, and the convergence of the proposed method is also analyzed. Experimental results on a recently released laparoscopic dataset demonstrate the clear superiority of the proposed methods. The proposed method can facilitate the access of key information in surgeries, training of junior clinicians, explanations to patients, and archive of case files.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, ZongYuan Ge, Vincent Lam, David Dagan Feng
IEEE J. Biomed. Health Informatics3
2021 Patch Based Video Summarization With Block Sparse Representation
abstract
In recent years, sparse representation has been successfully utilized for video summarization (VS). However, most of the sparse representation based VS methods characterize each video frame with global features. As a result, some important local details could be neglected by global features, which may compromise the performance of summarization. In this paper, we propose to partition each video frame into a number of patches and characterize each patch with global features. Instead of concatenating the features of each patch and utilizing conventional sparse representation, we formulate the VS problem with such video frame representation as block sparse representation by considering each video frame as a block containing a number of patches. By taking the reconstruction constraint into account, we devise a simultaneous version of block-based OMP (Orthogonal Matching Pursuit) algorithm, namely SBOMP, to solve the proposed model. The proposed model is further extended to a neighborhood based model which considers temporally adjacent frames as a super block. This is one of the first sparse representation based VS methods taking both spatial and temporal contexts into account with blocks. Experimental results on two widely used VS datasets have demonstrated that our proposed methods present clear superiority over existing sparse representation based VS methods and are highly comparable to some deep learning ones requiring supervision information for extra model training.
Shaohui Mei, Mingyang Ma 0004, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.3
2020 Video summarization via block sparse dictionary selection
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing3
2020 Rate-distortion-complexity optimization for x265
Saiping Zhang, Fuzheng Yang 0001, Shuai Wan
J. Vis. Commun. Image Represent.3
2019 Discriminative CNN Via Metric Learning for Hyperspectral Classification
abstract
Convolutional neural networks (CNNs) have been demonstrated to be capable of learning effective spatial-spectral features for hyperspectral classification. However, traditional CNNs are mainly trained using classification errors in decision domain. In this paper, a metric learning based training strategy is proposed to further enhance feature separability by training CNNs in feature domain as well as decision domain. Specifically, a metric learning loss function is designed to train CNNs in the second last fully connected feature layer, instead of the last fully connected decision layer. As a result, both within-class feature similarity and between-class feature separability can be enhanced even with a small amount of training samples. Experimental results over two benchmark hyperspectral data sets demonstrate that the proposed metric learning strategy is very effective to explore more discriminative features and its performance obviously outperforms several state-of-art CNNs for classification of hyperspectral images.
Zhongqi Tian, Zhi Zhang 0023, Shaohui Mei, Ruoqiao Jiang, Shuai Wan, Qian Du 0001
IGARSS5
2019 Robust video summarization using collaborative representation of adjacent frames
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
Multim. Tools Appl.3
2019 Temporal-Layer-Motivated Lambda Domain Picture Level Rate Control for Random-Access Configuration in H.265/HEVC
abstract
Rate control is a key technique for video communication systems. The aim of rate control is to transmit the best possible quality video sequences under various restrictions, such as channel bandwidth, buffer capacity, maximum time delay allowed for a given service, and so on. The λ domain rate control technique (λ-RC) has been integrated into the latest High Efficiency Video Coding standard (H.265/HEVC) test model, due to its accurate bit estimation and high rate-distortion performance. However, it is found that the λ-RC is not the optimal choice under the random-access configuration. When the random-access configuration is used, pictures are organized into temporal layers, where pictures in different layers are of different importance in terms of prediction. In this paper, a picture level lambda domain rate control technique for the randomaccess configuration in H.265/HEVC is proposed. The influence of temporal layers is effectively considered in the proposed algorithm referred to as TL-λ-PRC. Experimental results verify that the proposed TL-λ-PRC is efficient in coding performance and accurate in bit estimation. Compared with the λ-RC with the fixed ration bit allocation which has been implemented in the test model of H.265/HEVC (HM 14.0), TL-λ-PRC achieves an average reduction of 4.10% and 3.49% for slow motion and fast motion sequences, respectively, in BD-rate (Bjøntegaard-Delta bitrate) with more accurate bit estimation. The performance of different algorithms in terms of algorithm complexity and quality fluctuation are also carefully analyzed in this contribution.
Yanchao Gong, Shuai Wan, Kaifang Yang, Hong Ren Wu, Ying Liu 0026
IEEE Trans. Circuits Syst. Video Technol.2
2018 Video Summarization via Weighted Neighborhood Based Representation
abstract
The recent explosive growth of multimedia data has posed a new set of challenges in computer vision, and video summarization (VS) techniques are increasingly important to automatically summarize a large amount of multimedia data in an effective and efficient manner. Recent years have witnessed the rise and developments of sparse representation based approaches for VS. While the existing methods select keyframes according to the information contained in the single frame, and such a selection based solely on single-frame information may not be robust. Therefore, in this paper, the information of the single frame's neighborhood is taken into consideration, and different weights are assigned to these neighbouring frames. We formulate the VS problem as a weighted neighborhood based representation model, and design a greedy pursuit algorithm to extract keyframes. Experimental results on a benchmark dataset demonstrate that the proposed method can outperform the state of the arts.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Ah Chung Tsoi, David Dagan Feng
ICIP3
2018 Hyperspectral Classification Via Spatial Context Exploration with Multi-Scale CNN
abstract
Spatial context has shown to be very useful in hyperspectral image processing. Existing convolutional neural network (CNN)-based methods for hyperspectral classification explore spatial context by single-scale convolution kernels in 2D or 3D shapes. However, such single-scale convolution may not be capable to explore the complex spatial context in a hyperspectral image. In this paper, we propose a multi-scale CNN, MS-CNN to explore the spatial context in different extents, in which adaptive spatial neighborhood convolution kernels are used to simultaneously extract multiple spectral-spatial features from spatial context of pixels. These features obtained by different spatial kernels are then concatenated and fused for further feature extraction and classification. Experimental results show that the proposed adaptive spatial neighborhood convolution are more effective to explore spatial context than traditional single-scale spatial convolution and the performance of the proposed MS-CNN outperforms several state-of-art CNNs for classification of hyperspectral images.
Zhongqi Tian, Jingyu Ji, Shaohui Mei, Junhui Hou, Shuai Wan, Qian Du 0001
IGARSS5
2018 Stretching Schemes for Coding Frames of Panoramic Videos in Craster Parabolic Projection
abstract
Panoramic videos are spherical in nature, which further brings great challenges to deal with them. Usually they are projected to planar domain and processed as planar perspective videos. Craster parabolic projection (CPP), as a sphere-to-plane projection format, achieves approximately uniform sampling on the sphere. Without redundant pixels, it can store and represent panoramic videos effectively. However, frames in CPP format are no longer rectangular, which further violates the off-the-shelf video coding standards. In this paper, four stretching schemes are proposed for coding frames of panoramic videos in CPP. For introducing as few pixels as possible, strips are regard as the basic units. Strips in frames in different areas are stretched into rectangles in different sizes for coding. Spherical continuity, planar continuity, nearest-neighbour interpolation and Lanczos interpolation are considered in stretching respectively. Experimental results demonstrate that, compared with strips in Equi-rectangular projection (ERP) format, the proposed schemes can achieve BD-rate reductions up to 30.68% for Y, 32.68% for U and 34.13% for V, and that different schemes are well adapted for different strips.
Saiping Zhang, Mengpin Qiu, Fuzheng Yang 0001, Shuai Wan
VCIP5
2018 Perceptual video quality metric for compression artefacts: from two-dimensional to omnidirectional
abstract
In this study, the perceptual video quality (PVQ) for the vogue 360‐degree video with compression artefacts viewed on the visual reality (VR) head‐mounted display (HMD) is obtained for the first time by an elaborately designed subjective experiment. The characteristic of the PVQ of omnidirectional (360‐degree) video on the VR HMD and PC monitor is then investigated. The PVQ on the HMD is found to be linearly related to the video coding quality (VCQ) on the PC monitor. A quality evaluation model is then proposed based on the mapping formula and a new assessment parameter, where the impact of the video resolution and display device is involved. In this way, the traditional video quality assessment metrics for two‐dimensional video can be extended to assess the PVQ of the 360‐degree video viewed on the HMD. At last, the video structural similarity metric is taken as an example to predict the input of the proposed model, i.e. the VCQ of the 360‐degree video. Experimental results demonstrate that the proposed model can effectively and conveniently solve the challenge of PVQ assessment for the 360‐degree video with compression distortions on the HMD.
Wenjie Zou, Fuzheng Yang 0001, Shuai Wan
IET Image Process.3
2018 Event-Based Perceptual Quality Assessment for HTTP-Based Video Streaming With Playback Interruption
abstract
The popularity of HTTP-based video streaming services has been increasing in recent years. The quality of HTTP-based video streaming services is measured by the user's quality of experience (QoE). Perceptual quality (PQ), a strong indicator of the QoE, has been widely studied in the literature. Specifically, previous studies primarily focused on modeling the overall perceptual quality at the end of the viewing process. The specific influence of each event during the viewing process was neglected. Since the PQ is a function of time, which models the human perception of audiovisual services, this contribution focuses on investigating the time-varying feature of human perception. In this paper, the viewing timeline of subjects is divided into rebuffering and playback. An event-based perceptual quality assessment (EPQ) framework is introduced, which can assess the PQ of interrupted HTTP-based video streaming at any point in time. To understand the human perception of video streams with playback interruptions, a subjective quality assessment experiment was designed to obtain the PQ at a series of points in time. A between-subjects design was adopted in which each subject can view video content without repetition. Based on the experimental results, a PQ assessment model was proposed to predict the change in PQ (i.e., ΔPQ) after each rebuffering and playback. The overall PQ at any point in time is calculated as the summation of the ΔPQs of all the previous events. The experimental results demonstrated that the EPQ-based model outperforms the rebuffering components contained in the ITU-T Rec. P.1201.1, P.1201 Amd. 2 and P.1203.3 in our database.
Wenjie Zou, Fuzheng Yang 0001, Jiarun Song, Shuai Wan, Wei Zhang 0072, Hong Ren Wu
IEEE Trans. Multim.4
2017 Rate-Distortion Optimization for Video Coding under Given Computational Complexity
abstract
Rate-distortion optimization (RDO) is widely applied in video coding, which aims at minimizing the coding distortion under a target coding rate. Conventionally, RDO in video coding does not take into account the coding complexity. However, because of the diversity of video applications, the video encoders in different applications may have different requirements of or limitation on the computational complexity. Therefore, it is desirable for video encoders to perform RDO in flexible computational complexity. In this paper, we propose a novel RDO scheme under the given computational complexity for the latest H.265/HEVC standard. A model for prediction of the rate-distortion cost (RD cost) is first established based on a pre-searching process. Then according to the predicted RD cost, the rate-distortion-complexity (R-D-C) characteristics of different coding tree units (CTUs) are analyzed. Finally, the total complexity budget is properly allocated to different CTUs according to their R-D-C characteristics. Experimental results demonstrate that, compared with x265, the proposed algorithm can reduce, on average, the BD-rate by 18.8% under the same requirements of encoding speed.
Junkai Feng, Saiping Zhang, Fuzheng Yang 0001, Shuai Wan
DCC4
2017 Exploring the influence of feature representation for dictionary selection based video summarization
abstract
Dictionary selection based video summarization (VS) algorithms, in which keyframes are considered as a dictionary to reconstruct all the video frames, have been demonstrated to be effective and efficient for video summarization. It has been noticed that the feature representation of video plays a great impact of the performance of VS. In this paper, the influence of feature representation of video frames on the performance of dictionary selection-based VS is for the first time investigated. In addition to the traditional hand-crafted features used in VS, such as color histogram, the deep features learned through deep neural networks are firstly used to represent video frames for dictionary selection-based VS. The impact of dimensionality reduction to the high-dimensional deep learning features on VS is further discussed. Experimental results on a benchmark video dataset demonstrate that deep learning features are able to achieve better performance than traditional hand-crafted features for dictionary selection-based VS. Moreover, the dimensionality of deep learning features can be reduced to decrease the computational cost without the degradation of VS performance.
Mingyang Ma 0004, Shaohui Mei, Jingyu Ji, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICIP4
2017 Hyperspectral image super-resolution via convolutional neural network
abstract
Due to the tradeoff between spatial and spectral resolution in remote sensing imaging, hyperspectral images are often acquired with a relative low spatial resolution, which limits their applications in many areas. Inspired by recent achievements in convolutional neural network (CNN) based super resolution (SR), a novel CNN based framework is constructed for SR of hyperspectral images by considering both spatial context and spectral correlation. As a result, the spectral distortion incurred by directly applying traditional SR algorithms to hyperspectral images is alleviated. Experimental results on several benchmark hyperspectral datasets have demonstrated that higher quality of reconstruction and spectral fidelity can be achieved, compared to band-wise manner based algorithms.
Shaohui Mei, Xin Yuan 0002, Jingyu Ji, Shuai Wan, Junhui Hou, Qian Du 0001
ICIP4
2017 Content adaptive quantization parameter cascading for random-access structure in HEVC
abstract
In high-efficiency coding (HEVC), the random-access structure (RAS) is employed due to its high coding efficiency and “random-access” performance. Pictures in RAS are assigned to different temporal layers. And how to select the QP for each temporal layer in quantization parameter cascading (QPC) technique is critical for improving the coding efficiency of RAS. In order to further improve the coding efficiency of RAS, a QPC technique considering the video content characteristics (denoted as VC-QPC for short) is proposed. In VC-QPC, motion and texture complexities of a video are used for predicting the optimal quantization parameter values of temporal layers. Compared with the method in the test model of HEVC, i.e., HM14.0, the BD-rate of proposed VC-QPC is -5.60% for RAS.
Kaifang Yang, Shuai Wan, Yanchao Gong
ICIP2
2017 Nonlinear kernel sparse dictionary selection for video summarization
abstract
Sparse dictionary selection (SDS) has demonstrated to be an effective solution for keyframe based video summarization (VS), which generally assumes a linear relation among similar video frames. However, such a linear assumption is not always true for videos. In this paper, the nonlinearity among frames is taken into consideration and a nonlinear SDS model is formulated for VS, in which the nonlinearity is transformed to linearity by projecting a video to a high dimensional feature space induced by a kernel function. Moreover, a kernel simultaneous orthogonal matching pursuit (KSOMP) is proposed to solve the problem. In order to achieve an intuitive and flexible configuration of the VS process, an adaptive criterion is devised to produce video summaries with different lengths for different video content. Experimental results on benchmark video datasets demonstrate that the proposed algorithm outperforms several state-of-the-art VS algorithms.
Mingyang Ma 0004, Shaohui Mei, Junhui Hou, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICME4
2017 An efficient Lagrangian multiplier selection method based on temporal dependency for rate-distortion optimization in H.265/HEVC
Kaifang Yang, Shuai Wan, Yanchao Gong, Hong Ren Wu
Signal Process. Image Commun.2
2017 Rate-Distortion-Optimization-Based Quantization Parameter Cascading Technique for Random-Access Configuration in H.265/HEVC
abstract
The random-access configuration is employed in H.265/HEVC video coding to make inter-frame prediction more efficient. The coding efficiency under the random-access configuration is closely related to the quantization parameter cascading (QPC) technique which determines the QP for encoding pictures in different temporal layers. The present QPC technique for the random-access configuration in H.265/HEVC is not optimized. In this paper, first, a rate-distortion-optimization-based technique for QPC, referred to as RDO-QPC, is proposed for the random-access configuration in H.265/HEVC. Based on the results from RDO-QPC, a simplified QPC technique, referred to as SRDO-QPC, is proposed. The experimental results verify the efficiency of the proposed two QPC techniques. Compared with the original configuration of QPC in H.265/HEVC, RDO-QPC achieves an average gain of 0.19 dB in APSNR and an average reduction of 4.87% in ABR, while SRDO-QPC achieves an average gain of 0.17 dB in APSNR and an average reduction of 4.32% in ABR. These two QPC techniques can both be used in practical situations, meeting different requirements of computational complexity.
Yanchao Gong, Shuai Wan, Kaifang Yang, Bo Li 0089
IEEE Trans. Circuits Syst. Video Technol.2
2016 Adaptive quantization parameter cascading for random-access prediction in H.265/HEVC based on dependent R-D models
abstract
In H.265/HEVC, the random-access prediction structure helps to increase the coding efficiency and provides the temporal scalability. However, the quantization parameter cascading (QPC) strategy being used is not optimized in terms of the rate-distortion performance. This paper analyzes the rate and distortion dependency between temporal layers, and proposes a PSNR metric-based distortion model and a piecewise rate model to optimize the QPC. Experimental results demonstrate that the proposed algorithm achieves an average APSNR gain of 0.12dB with the maximum APSNR gain being 0.37dB. Meanwhile, an up to 8.7% BD-rate reduction with the average BD-rate reduction being 3.3% can be obtained.
Shuai Wan, Yanchao Gong, Kaifang Yang
ICIP2
2016 Detection and estimation of supra-threshold distortion levels of pictures based on just-noticeable difference
abstract
A subjective assessment method is described to determine picture quality levels in the supra-threshold region for processed images, with reference to their original counterparts, based on just-noticeable difference (JND) detection experiment. It has been found that the range of JND levels is dependent on picture contents and can be predicted as a function of texture masking factor computed in the pixel domain. The experimental data obtained also reveal that relationship of JND levels in the supra-threshold region and the MSE (mean squared error) can be approximated by a linear function whose slope is modeled as a function of edge and texture contrast masking factors. The model is devised to predict JND levels which provide subjective picture quality rating discernible by human viewers and can be used for visual quality regulated image/video coding, as well as evaluating the capacity of existing objective metrics in predicting picture quality and/or distortion relative to JND based quality/distortion rating categories.
Kaifang Yang, Shuai Wan, Hong Ren Wu, Weisi Lin, Damian M. Tan, Yanchao Gong, Leyi Xie
VCIP2
2016 Reduction of temporal distortion in video coding based on detection of just-noticeable temporal pumping artifact
Yanchao Gong, Shuai Wan, Kaifang Yang, Hong Ren Wu, Fuzheng Yang 0001, Bo Li 0089
Signal Process. Image Commun.2
2016 QoE Evaluation of Multimedia Services Based on Audiovisual Quality and User Interest
abstract
Quality of experience (QoE) has significant influence on whether or not a user will choose a service or product in the competitive era. For multimedia services, there are various factors in a communication ecosystem working together on users, which stimulate their different senses inducing multidimensional perceptions of the services, and inevitably increase the difficulty in measurement and estimation of the user's QoE. In this paper, a user-centric objective QoE evaluation model (QAVIC model for short) is proposed to estimate the user's overall QoE for audiovisual services, which takes account of perceptual audiovisual quality (QAV) and user interest in audiovisual content (IC) amongst influencing factors on QoE such as technology, content, context, and user in the communication ecosystem. To predict the user interest, a number of general viewing behaviors are considered to formulate the IC evaluation model. Subjective tests have been conducted for training and validation of the QAVIC model. The experimental results show that the proposed QAVIC model can estimate the user's QoE reasonably accurately using a 5-point scale absolute category rating scheme.
Jiarun Song, Fuzheng Yang 0001, Yicong Zhou, Shuai Wan, Hong Ren Wu
IEEE Trans. Multim.4
2015 A frame level metric for just noticeable temporal pumping artifact in videos encoded with the hierarchical prediction structure
abstract
At low bit-rates, video coding in the hierarchical prediction structure (HPS) using the quantization parameter cascading strategy will introduce the temporal pumping artifact (TPA). TPA presents itself as a stumbling visual effect and is caused by severe quality fluctuations among adjacent pictures. This paper proposes a frame level metric for just noticeable temporal pumping artifact (JNTPA) relying on the characteristics of temporal-spatial masking in the human visual system (HVS). The proposed metric quantitatively determines the TPA when it becomes visible. The experiments have demonstrated that the estimated JNTPA values for different videos using the proposed metric are in line with the HVS perception. An accurate estimation of the JNTPA can be used in video coding for avoiding the annoying TPA, and in perceptual quality evaluation for videos.
Yanchao Gong, Shuai Wan, Fuzheng Yang 0001, Hong Ren Wu, Bo Li 0089
ICIP2
2015 Onboard image selection for small-satellite based remote sensing mission
abstract
The contradiction that imaging system can acquire huge amount of image data while communication system can deliver only a very small part of them has become a bottleneck for the small-satellite based earth observation. In this paper, a novel onboard image selection strategy is designed to select most informative images that are acquired by the imaging system for transmission. Specifically, the image that is worst reconstructed by previously transmitted images, instead of the instantly acquired image, is selected since it possesses most distinguishing information to previously transmitted images. Experiment on simulated image sequence has demonstrated the effectiveness of the proposed onboard image selection algorithm.
Yihang Wang 0001, Shaohui Mei, Shuai Wan, Yi Wang 0068
IGARSS3
2015 A linear dependent rate-quantization model for scalable video enhancement layer encoding
abstract
In this paper, we propose a linear dependent rate-quantization model for video enhancement layers encoding in H.264/AVC based scalable video coding (SVC). It is noted that the proposed model is applicable for different scalable structures, such as temporal, quality, spatial and combined scalability. Leveraging the base layer information (such as bitrate and quantization parameter), proposed model can accurately predict the number of bit required for the enhancement layer encoding. Such linear model demonstrates the high accuracy for bitrate estimation at enhancement layers, with the average prediction accuracy over 94%. It has the noticeable improvement from the existing works, without requiring additional complexity increase. Meanwhile, proposed model is applied to do the rate control for enhancement layers encoding. Experimental results show that the average bitrate mismatch error can be significantly reduced compared with the existing algorithms.
Junhui Hou, Shuai Wan, Lap-Pui Chau
ISCAS2
2015 Video summarization via minimum sparse reconstruction
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Shuai Wan, Mingyi He, David Dagan Feng
Pattern Recognit.4
2015 Perceptual based SAO rate-distortion optimization method with a simplified JND model for H.265/HEVC
Kaifang Yang, Shuai Wan, Yanchao Gong, Hong Ren Wu
Signal Process. Image Commun.2
2014 Iterative keyframe selection by orthogonal subspace projection
abstract
Recent developments on sparse dictionary selection have demonstrated promising results for Video Summarization (VS). However, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly. In this paper, a selection matrix is proposed to model the VS problem, according to which the L0norm of this selection matrix is imposed to ensure sparsity directly. As a result, a computational efficient Orthogonal Subspace Projection (OSP) based Iterative Keyframe Selection (IKS) algorithm is proposed for VS. In addition, a Percentage Of Reconstruction (POR) criterion is proposed to provide an intuitive and flexible control of the length of final video summaries even without prior knowledge of a given video. Experimental results on a popular benchmark dataset demonstrate that our proposed algorithm outperforms the state-of-the-art methods.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Shuai Wan, David Dagan Feng
ICIP5
2014 An efficient algorithm to eliminate temporal pumping artifact in video coding with hierarchical prediction structure
Yanchao Gong, Shuai Wan, Kaifang Yang, Fuzheng Yang 0001
J. Vis. Commun. Image Represent.2
2013 A fast algorithm of bitstream extraction using distortion prediction based on simulated annealing
Kaifang Yang, Shuai Wan, Yanchao Gong
J. Vis. Commun. Image Represent.2
2012 Perception of Temporal Pumping Artifact in Video Coding with the Hierarchical Prediction Structure
abstract
The usage of the hierarchical prediction structure in video coding has introduced a special type of temporal noise, i.e., the temporal pumping artifact. This artifact presents itself as severe quality fluctuations among adjacent pictures and is quite annoying due to the pumping or stumbling effect in perception. In this paper the fundamental reason of perception of the temporal pumping artifact is analyzed. The key factors influencing perception of temporal pumping artifact are evaluated based on subjective experiments, in terms of amplitude, frequency and phase of quality fluctuations, respectively. The detailed analysis suggests how the temporal pumping artifact can be well alleviated or even eliminated through adjusting coding parameters.
Shuai Wan, Yanchao Gong, Fuzheng Yang 0001
ICME1
2012 Efficient bitstream extraction for scalable video based on simulated annealing
abstract
SUMMARY This paper presents an efficient method for bitstream extraction for scalable video based on simulated annealing. Following the same spirit as annealing, the proposed method searches for the optimized combination of quality layers for extraction through slowly reducing the simulated temperature according to the characteristics of frames in different temporal levels. Experimental results show that the proposed method provides an optimized performance, which is significantly higher than that of the basic extraction method. When compared with the quality layer‐based extraction method in the reference software model of H.264/SVC (Joint Scalable Video Model), the proposed method can achieve a similar rate‐distortion performance with a significantly reduced computational complexity. Furthermore, the proposed method can obtain a more smoothed video quality, which is always preferable by the end user. Copyright © 2011 John Wiley & Sons, Ltd.
Shuai Wan, Kaifang Yang, Haiyong Zhou
Concurr. Comput. Pract. Exp.1
2010 Frame-loss adaptive temporal pooling for video quality assessment
abstract
In this paper a frame-loss adaptive temporal pooling method for video quality assessment is proposed. Extensive subjective tests have been carried out to determine the duration of successive frames based on which steady quality judgment can be made by human observers. The resulting duration is applied to the determination of the length of Group of Frames (GOF), where a flexible algorithm is used to separate the input video into variable sized GOFs. Short-term temporal pooling is first performed for each of the GOF to get the GOF quality, where quality contribution of each frame is incorporated with the context and frame loss well taken into account. The video quality is then obtained by long-term temporal pooling of the GOF quality considering the fact that perceptual video quality is predominately determined by the worst parts of the video. Extensive experimental results have demonstrated the effectiveness of the proposed method both for regular and irregular frame loss.
Shuai Wan, Fuzheng Yang 0001, Chenglong Jiang
VCIP1
2010 Rate-Distortion Criterion Based Picture Padding for Arbitrary Resolution Video Coding Using H.264/MPEG-4 AVC
abstract
The video coding standard H.264/MPEG-4 AVC is designed based on the basic unit of macroblock (MB). In order to support the arbitrary resolution video coding when the picture does not contain an integral number of MBs, the H.264/MPEG-4 AVC encoder usually invokes a picture padding process to extend the boundaries of a picture by adding extra pixels. In this way, the video of an arbitrary resolution can be coded by an ordinary encoder and the resulting bitstream follows the H.264/MPEG-4 AVC syntax. From the perspective of coding efficiency and the concept of soft decision, we consider the issue of rate-distortion (R-D) based picture padding as an optimization problem, and present an iterative solution which integrates the optimal picture padding in the R-D optimization (RDO) process. Furthermore, to reduce the related computational burden, a heuristic solution is proposed based on comprehensive analyses on the influence of the padded pixels on the encoder performance. This solution takes the prediction errors and compression distortions of padded pixels as zero in the RDO process of the encoder, and automatically implements the concepts and techniques of template matching and spatial extrapolation to determine the padded pixels. Since it is carried out by directly utilizing the already existing algorithms of RDO, the proposed heuristic solution brings little additional computational complexity to the encoder. Based on the reference software for H.264/MPEG-4 AVC with its version of JM15.1, the effectiveness of the proposed method has been verified by extensive experiments.
Yilin Chang, Fuzheng Yang 0001, Shuai Wan
IEEE Trans. Circuits Syst. Video Technol.4
2010 No-Reference Quality Assessment for Networked Video via Primary Analysis of Bit Stream
abstract
A no-reference (NR) quality measure for networked video is introduced using information extracted from the compressed bit stream without resorting to complete video decoding. This NR video quality assessment measure accounts for three key factors which affect the overall perceived picture quality of networked video, namely, picture distortion caused by quantization, quality degradation due to packet loss and error propagation, and temporal effects of the human visual system. First, the picture quality in the spatial domain is measured, for each frame, relative to quantization under an error-free transmission condition. Second, picture quality is evaluated with respect to packet loss and the subsequent error propagation. The video frame quality in the spatial domain is, therefore, jointly determined by coding distortion and packet loss. Third, a pooling scheme is devised as the last step of the proposed quality measure to capture the perceived quality degradation in the temporal domain. The results obtained by performance evaluations using MPEG-4 coded video streams have demonstrated the effectiveness of the proposed NR video quality metric.
Fuzheng Yang 0001, Shuai Wan, Qingpeng Xie, Hong Ren Wu
IEEE Trans. Circuits Syst. Video Technol.2
2009 A Novel Blind Measurement of Blocking Artifacts for H.264/AVC Video
abstract
This paper presents a novel method for fast and quantified estimation of blocking artifacts for H.264 videos. Based on physiological characteristics of the human visual system (HVS), the proposed method takes into account the temporal blocky distortion between two successive frames and is well adapted to the videos in which the in-loop deblocking filter is applied. Furthermore, the proposed method is computationally efficient and does not require the original video sequence for reference. Extensive experiments have demonstrated the computational efficiency of the proposed method. Comparison results also show that the blocking artifacts evaluated by the proposed method agree well with the other objective assessment method such as PSNR, VQM and SSIM. Due to its accuracy and efficiency, the proposed method can be used as a quality metric alone or a factor in quality evaluation in real-time multimedia communication applications.
ZhaoLin Zhang, Haoshan Shi, Shuai Wan
ICIG3
2009 Frame layer rate control for H.264/AVC with hierarchical B-frames
Yilin Chang, Fuzheng Yang 0001, Shuai Wan, Sixin Lin, Lianhuan Xiong
Signal Process. Image Commun.4
2009 Lagrange multiplier selection in wavelet-based scalable video coding for quality scalability
Shuai Wan, Fuzheng Yang 0001, Ebroul Izquierdo
Signal Process. Image Commun.1
2007 Error Robustness Scheme for Scalable Video Based on the Concatenation of LDPC and Turbo Codes
abstract
In this paper, a novel approach for transmission of scalable video over wireless channel is proposed. The proposed approach jointly optimises the bit allocation between a wavelet-based scalable video coding framework and a forward error correction codes. The forward error correction codes is based on the serial concatenation of LDPC codes and turbo codes. Turbo codes shows good performance at high error rates region but LDPC outperforms turbo codes at low error rates. So the concatenation of LDPC and TC enhances the performance at both low and high signal to noise ratios. The scheme minimizes the reconstructed video distortion at the decoder subject to a constraint on the overall transmission bitrate budget. The minimization is achieved by exploiting the source rate distortion characteristics and the statistics of the available codes. Furthermore, an efficient decoding algorithm is proposed. Experimental results clearly demonstrate the superiority of the proposed approach over conventional forward error correction techniques.
Naeem Ramzan, Shuai Wan, Ebroul Izquierdo
ICIP (6)2
2007 Lagrange Multiplier Selection for 3-D Wavelet Based Scalable Video Coding
abstract
In this paper a thorough analysis on the theoretical rate distortion model and the rate distortion performance in an open-loop structure is conducted. A Lagrange multiplier selection for 3D wavelet based scalable video coding is then derived. The proposed Lagrange multiplier is adaptive with respect to the characteristics of video content. Furthermore, it is especially suitable for 3-D wavelet based scalable video coding where quantisation steps are unavailable. Extensive experimental results have demonstrated the effectiveness of the proposed Lagrange multiplier selection.
Fuzheng Yang 0001, Shuai Wan, Ebroul Izquierdo
ICIP (2)2
2007 An Efficient Joint Source-Channel Coding for Wavelet Based Scalable Video
abstract
A robust and efficient approach for scalable video transmission over wireless channels is presented. The proposed approach jointly optimizes source and channel coding in order to minimize the overall end-to-end distortion. In particular, the forward error correction method based on turbo codes is considered. Aiming at improving the overall performance of the underlying joint source-channel coding, the combination of the channel coding rate, interleaver and packet size are optimized for turbo codes subject to a constraint on the overall transmission bitrate budget. Experimental results show that the proposed approach outperforms conventional forward error correction techniques at all bit error rates, even in very adverse conditions.
Naeem Ramzan, Shuai Wan, Ebroul Izquierdo
ISCAS2
2007 Perceptually adaptive joint deringing-deblocking filtering for scalable video transmission over wireless networks
Shuai Wan, Marta Mrak, Naeem Ramzan, Ebroul Izquierdo
Signal Process. Image Commun.1
2007 Rate-Distortion Optimized Motion-Compensated Prediction for Packet Loss Resilient Video Coding
abstract
A rate-distortion optimized motion-compensated prediction method for robust video coding is proposed. Contrasting methods from the conventional literature, the proposed approach uses the expected reconstructed distortion after transmission, instead of the displaced frame difference in motion estimation. Initially, the end-to-end reconstructed distortion is estimated through arecursive per-pixel estimation algorithm. Then the total bit rate for motion-compensated encoding is predicted using a suitable rate distortion model. The results are fed into the Lagrangian optimization at the encoder to perform motion estimation. Here, the encoder automatically finds an optimized motion compensated prediction by estimating the best tradeoff between coding efficiency and end-to-end distortion. Finally, rate-distortion optimization is applied again to estimate the macroblock mode. This process uses previously selected optimized motion vectors and their corresponding reference frames. It also considers intraprediction. Extensive computer simulations in lossy channel environments were conducted to assess the performance of the proposed method. Selected results for both single and multiple reference frames settings are described. A comparative evaluation using other conventional techniques from the literature was also conducted. Furthermore, the effects of mismatches between the actual channel packet loss rate and the one assumed at the encoder side have been evaluated and reported in this paper.
Shuai Wan, Ebroul Izquierdo
IEEE Trans. Image Process.1
2006 End-to-End Rate-Distortion Optimized Motion Estimation
abstract
An end-to-end rate-distortion optimized motion estimation method for robust video coding in lossy networks is proposed. In this method the expected reconstructed distortion after transmission and the total bit rate for displaced frame difference are estimated at the encoder. The results are fed into the Lagrangian optimization at the encoder to perform motion estimation. Here the encoder automatically finds an optimized motion compensated prediction by estimating the best trade off between coding efficiency and end-to-end distortion. Computer simulations in lossy channel environments were conducted to assess the performance of the proposed method. A comparative evaluation using other conventional techniques from the literature was also conducted.
Shuai Wan, Ebroul Izquierdo, Fuzheng Yang 0001, Yilin Chang
ICIP1
2005 A novel objective no-reference metric for digital video quality assessment
abstract
A novel objective no-reference metric is proposed for video quality assessment of digitally coded videos containing natural scenes. Taking account of the temporal dependency between adjacent images of the videos and characteristics of the human visual system, the spatial distortion of an image is predicted using the differences between the corresponding translational regions of high spatial complexity in two adjacent images, which are weighted according to temporal activities of the video. The overall video quality is measured by pooling the spatial distortions of all images in the video. Experiments using reconstructed video sequences indicate that the objective scores obtained by the proposed metric agree well with the subjective assessment scores.
Fuzheng Yang 0001, Shuai Wan, Yilin Chang, Hong Ren Wu
IEEE Signal Process. Lett.2
2003 A no-reference video quality assessment method based on digital watermark
abstract
Video quality assessment in real time is a critical requirement for channel evaluation and codec optimization in mobile multimedia system. However, most proposed approaches for video quality assessment need reference sequences, which is quite impossible in real time multimedia communications especially in wireless and IP video services. This paper proposes a novel no-reference video quality assessment method. By comparison of the extracted watermark embedded in transmitted video with the copy of original watermark in the destination, we can assess all reconstructed video quality without any original reference video sequences. Simulation results conclusively demonstrate that our method offers no-reference video quality assessment with little channel overhead and can be applied to real time multimedia systems.
Fuzheng Yang 0001, Xin-dai Wang, Yilin Chang, Shuai Wan
PIMRC4