Xuekai Wei

dblp:211/0698 · DBLP profile ↗
← Back
52ranked-venue papers
9as first author
48since 2021 · last 2027
0000-0002-3761-1759ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Computer networks · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Semantic-guided multi-feature fusion for underwater image quality assessment
Huayan Pu, Jun Luo 0006, Jielu Yan, Weizhi Xian, Xuekai Wei, Mingliang Zhou 0001
Expert Syst. Appl.7
2026 Image quality assessment: Unifying spatial and frequency distribution discrepancy in deep feature domains via Rényi divergence
Dongzi Wang 0003, Weizhi Xian, Jielu Yan, Xuekai Wei, Mingliang Zhou 0001, Sam Kwong
Expert Syst. Appl.4
2026 GCA-DETR: Global-context-aware-based detection transformer
Zhenzhe Hechen, Mingliang Zhou 0001, Xuekai Wei, Sam Kwong
Inf. Sci.3
2025 Tokenphormer: Structure-aware Multi-token Graph Transformer for Node Classification
abstract
Graph Neural Networks (GNNs) are widely used in graph data mining tasks. Traditional GNNs follow a message passing scheme that can effectively utilize local and structural information. However, the phenomena of over-smoothing and over-squashing limit the receptive field in message passing processes. Graph Transformers were introduced to address these issues, achieving a global receptive field but suffering from the noise of irrelevant nodes and loss of structural information. Therefore, drawing inspiration from fine-grained token-based representation learning in Natural Language Processing (NLP), we propose the Structure-aware Multi-token Graph Transformer (Tokenphormer), which generates multiple tokens to effectively capture local and structural information and explore global information at different levels of granularity. Specifically, we first introduce the walk-token generated by mixed walks consisting of four walk types to explore the graph and capture structure and contextual information flexibly. To ensure local and global information coverage, we also introduce the SGPM-token (obtained through the Self-supervised Graph Pre-train Model, SGPM) and the hop-token, extending the length and density limit of the walk-token, respectively. Finally, these expressive tokens are fed into the Transformer model to learn node representations collaboratively. Experimental results demonstrate that the capability of the proposed Tokenphormer can achieve state-of-the-art performance on node classification tasks.
Zhaoqi Lu, Xuekai Wei, Rongqin Chen 0001, Shenghui Zhang, Pak Lon Ip, Leong Hou U
AAAI3
2025 Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual Inference
abstract
Existing full-reference image quality assessment (FR-IQA) methods often fail to capture the complex causal mechanisms that underlie human perceptual responses to image distortions, limiting their ability to generalize across diverse scenarios. In this paper, we propose an FR-IQA method based on abductive counterfactual inference to investigate the causal relationships between deep network features and perceptual distortions. First, we explore the causal effects of deep features on perception and integrate causal reasoning with feature comparison, constructing a model that effectively handles complex distortion types across different IQA scenarios. Second, the analysis of the perceptual causal correlations of our proposed method is independent of the backbone architecture and thus can be applied to a variety of deep networks. Through abductive counterfactual experiments, we validate the proposed causal relationships, confirming the model’s superior perceptual relevance and interpretability of quality scores. The experimental results demonstrate the robustness and effectiveness of the method, providing competitive quality predictions across multiple benchmarks. The source code is available at https://anonymous.4open.science/r/DeepCausalQuality-25BC.
Wenhao Shen, Mingliang Zhou 0001, Xuekai Wei, Yong Feng 0002, Huayan Pu, Weijia Jia 0001
CVPR4
2025 EEG Decoding and Visual Reconstruction via 3D Geometric with Nonstationarity Modelling
abstract
Electroencephalogram (EEG) signal processing has advanced in revealing the mechanisms of human visual perception, but existing methods often overlook two key EEG properties: (1) 3D geometric relationships between EEG electrodes, which reflects the ability to model the brain in stereoscopic terms; and (2) nonstationarity of EEG signals, which involves capturing the dynamic changes in the frequency spectrum. To address these limitations, we introduce the GeoCap framework in this paper. Specifically, to effectively model the 3D geometry of EEG electrodes, we propose spherical manifold encoding (SME), which represents EEG channels on a 3D spherical manifold. Additionally, drawing inspiration from the capacity of cumulative fractional derivatives to incorporate historical context, we introduce the discrete multiscale Caputo derivative (DMSCD) to more accurately capture the temporal dynamics of EEG signals across multiple scales. Experiments on EEG decoding and visual reconstruction tasks demonstrate that GeoCap outperforms other state-of-the-art methods.
Kaiwen Wei, Xuekai Wei, Jielu Yan
ICASSP4
2025 A Dynamic Learning Strategy for Dempster-Shafer Theory with Applications in Classification and Enhancement
abstract
Effective modelling of uncertain information is crucial for quantifying uncertainty. Dempster–Shafer evidence (DSE) theory is a widely recognized approach for handling uncertain information. However, current methods often neglect the inherent a priori information within data during modelling, and imbalanced data lead to insufficient attention to key information in the model. To address these limitations, this paper presents a dynamic learning strategy based on nonuniform splitting mechanism and Hilbert space mapping. First, the framework uses a nonuniform splitting mechanism to dynamically adjust the weights of data subsets and combines the diffusion factor to effectively incorporate the data a priori information, thereby flexibly addressing uncertainty and conflict. Second, the conflict in the information fusion process is reduced by Hilbert space mapping. Experimental results on multiple tasks show that the proposed method significantly outperforms state-of-the-art methods and effectively improves the performance of classification and low-light image enhancement (LLIE) tasks. The code is available at https://anonymous.4open.science/r/Third-ED16.
Mingliang Zhou 0001, Xuekai Wei, Weizhi Xian, Jielu Yan, Weijia Jia 0001
NeurIPS4
2025 Continuous reinforcement learning via advantage value difference reward shaping: A proximal policy optimization perspective
Xuekai Wei, Weizhi Xian, Jielu Yan, Leong Hou U, Yong Feng 0002, Zhaowei Shang, Mingliang Zhou 0001
Eng. Appl. Artif. Intell.2
2025 Blind Image Quality Assessment: Exploring Content Fidelity Perceptibility via Quality Adversarial Learning
Mingliang Zhou 0001, Wenhao Shen, Xuekai Wei, Jun Luo 0003, Fan Jia 0005, Xu Zhuang, Weijia Jia 0001
Int. J. Comput. Vis.3
2025 Hierarchical degradation-aware network for full-reference image quality assessment
Xuting Lan, Fan Jia 0005, Xu Zhuang, Xuekai Wei, Jun Luo 0006, Mingliang Zhou 0001, Sam Kwong
Inf. Sci.4
2025 A rate allocation model for VVC intercoding using a quality dependency
Heqiang Wang, Xuekai Wei, Mingliang Zhou 0001, Horace Ho-Shing Ip, Sam Kwong
Inf. Sci.2
2025 Image Quality Assessment: Exploring Joint Degradation Effect of Deep Network Features via Kernel Representation Similarity Analysis
abstract
Typically, deep network-based full-reference image quality assessment (FR-IQA) models compare deep features from reference and distorted images pairwise, overlooking correlations among features from the same source. We propose a dual-branch framework to capture the joint degradation effect among deep network features. The first branch uses kernel representation similarity analysis (KRSA), which compares feature self-similarity matrices via the mean absolute error (MAE). The second branch conducts pairwise comparisons via the MAE, and a training-free logarithmic summation of both branches derives the final score. Our approach contributes in three ways. First, integrating the KRSA with pairwise comparisons enhances the model's perceptual awareness. Second, our approach is adaptable to diverse network architectures. Third, our approach can guide perceptual image enhancement. Extensive experiments on 10 datasets validate our method's efficacy, demonstrating that perceptual deformation widely exists in diverse IQA scenarios and that measuring the joint degradation effect can discern appealing content deformations.
Xingran Liao, Xuekai Wei, Mingliang Zhou 0001, Hau-San Wong, Sam Kwong
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 A Semantic-Aware Detail Adaptive Network for Image Enhancement
abstract
Low-light images often suffer from varying degrees of visual degradation. Current methods for recovering image texture details fail to rely on the self-adaptive correlation texture direction of the image itself, which leads the network to be unable to address the local texture characteristics of different images. To address this challenge, we propose a semantic-aware detail adaptive network (SDANet) that fully considers the image detail information. The network divides low-light images into high-frequency and low-frequency parts. Learning different forms of noise through a novel total variation regularization module with adaptive weights ensures that the final high-frequency part adequately integrates the texture information of the image. Simultaneously, a detail-adaptive module is incorporated to restore finer details in the resulting image. SDANet not only effectively suppresses noise in real low-light images while considering texture details but also effectively addresses the degradation of visible information, and it performs better than other state-of-the-art methods. The code is available athttps://github.com/cheer79/SDANet.
Xuekai Wei, Mingliang Zhou 0001, Jielu Yan, Huayan Pu, Jun Luo 0003, Zhengguo Li
IEEE Trans. Circuits Syst. Video Technol.2
2025 No-Reference Image Quality Assessment: Exploring Intrinsic Distortion Characteristics via Generative Noise Estimation With Mamba
abstract
In the field of no-reference image quality assessment (NR-IQA), the visual masking effect has long been a challenging issue. Although existing methods attempt to alleviate the interference caused by masking by generating pseudoreference images, the quality of these images is often constrained by the accuracy and reconstruction capabilities of image restoration algorithms. This can introduce additional biases, thereby affecting the reliability of the evaluation results. To address this problem, we propose a novel generative “noise” estimation framework (GNE-Vim) that eliminates the need for pseudoreference images. Instead, it deeply decouples the distortion components from degraded images and performs quality-aware modelling of these components. During the training phase, the model leverages both reference images and distortion components to guide the learning of the true distortion distribution. In the inference phase, quality prediction is conducted directly on the basis of the decoupled distortion components, making the evaluation results more aligned with human subjective perception. The experimental results demonstrate that the proposed method achieves strong performance across datasets containing various types of distortions. The source code is publicly available at the following website: https://github.com/opencodelxt/GNE-Vim.
Xuting Lan, Weizhi Xian, Mingliang Zhou 0001, Jielu Yan, Xuekai Wei, Jun Luo 0006, Weijia Jia 0001, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.5
2025 Boundary-Aware Feature Fusion With Dual-Stream Attention for Remote Sensing Small Object Detection
abstract
Detecting small objects in remote sensing images poses significant challenges to the field of computer vision, primarily stemming from the complexity of backgrounds, limitations in pixel resolution, and information loss during the feature fusion process. While general object detection has significantly advanced in recent years, remote sensing small object detection remains an unsolved problem, with existing frameworks struggling to achieve high performance at small scales. In this article, we propose a novel framework called the boundary-aware feature fusion network (BAFNet), which significantly enhances the model’s ability to represent and locate small objects precisely within complex remote sensing scenarios. First, a dual-stream attention fusion module captures complementary foreground and background cues through bidirectional context modeling. Jointly attending to objects and their surroundings enhances discriminative power for distinguishing small objects. Additionally, we incorporate a boundary-aware branch to better preserve crucial detailed information vital for small-scale objects. This auxiliary component supervises the fusion of contextual semantics and spatial information, aiding in retaining critical boundary details that are prone to loss during cross-layer feature fusion. We conducted experiments on the challenging AI-TOD, VisDrone, DIOR, and LEVIR-Ship datasets. The results demonstrate the superiority of our approach over other state-of-the-art (SOTA) object detection methods, particularly in terms of precisely identifying small objects within remote sensing images. The code is available athttps://github.com/ooo1128/BAFNet.
Jingnan Song, Mingliang Zhou 0001, Jun Luo 0006, Huayan Pu, Yong Feng 0002, Xuekai Wei, Weijia Jia 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 GAANet: Graph Aggregation Alignment Feature Fusion for Multispectral Object Detection
abstract
Multispectral object detection has shown great promise in security and industrial applications. RGB images offer rich texture but are limited by lighting, whereas IR images excel in low light but lack texture. Current methods face challenges in accurately capturing information differences and achieving effective feature fusion across modalities. To address these issues, we propose a graph aggregation alignment network (GAANet) for multispectral object detection. GAANet consists of two key modules: the graph interaction fusion module (GIFM) and the information alignment module (IAM). GIFM uses graph representation learning to effectively process single-modality features, and the direct connection information flow mechanism guides and references low-level multimodal features, ensuring the global and comprehensive fusion of node information in the graph space. The results are then refined through the IAM for secondary calibration and alignment of corresponding local regions, ensuring accurate fusion. We also introduce an information reconstruction path (IRP) and reconstruction loss to prevent the loss of single-modality information due to multiple IAM calculations. GAANet achieves excellent fusion detection capability and significantly reduces the number of parameters, reducing the model size by 61.2% compared with that of representative baselines such as CALNet. GAANet achieves state-of-the-art results on the DroneVehicle, LLVIP, and FLIR datasets, with superior object detection accuracy. It also performs well on the unaligned DVTOD dataset, effectively capturing feature offsets across modalities through global graph perception.
Mingliang Zhou 0001, Zhaowei Shang, Xuekai Wei, Huayan Pu, Jun Luo 0006, Weijia Jia 0001
IEEE Trans. Ind. Informatics4
2025 COFNet: Contrastive Object-Aware Fusion Using Box-Level Masks for Multispectral Object Detection
abstract
Multispectral object detection, which combines RGB visible light and thermal infrared spectral information, has broad applications in complex environments and varying illumination conditions. However, existing methods face challenges in processing multispectral data, such as inconspicuous object features in spectral images and significant discrepancies between input modality spaces and output detection spaces. To address these issues, we propose an innovative multispectral object detection method that combines contrastive learning and a new cross-modal feature fusion module. We introduce a mask feature contrastive loss that maximizes the similarity between the box-level mask features and modal features while suppressing background responses, enabling effective representative alignment between the input and output spaces. Additionally, we propose a mask-guided attention fusion module that uses a predicted pseudo mask to guide the fusion of different modal features, enhancing object responses and reducing background noise interference. Our extensive experiments on several challenging multispectral datasets demonstrate that our proposed COFNet achieves state-of-the-art performance.
Mingliang Zhou 0001, Yunyao Li 0003, Guangchao Yang, Xuekai Wei, Huayan Pu, Jun Luo 0006, Weijia Jia 0001
IEEE Trans. Multim.4
2025 Sparse Reduced-Rank Fully Connected Layers with Its Applications in Detection and Classification
abstract
Fully connected (FC) layers play a significant role in deep neural networks (DNNs) models. Owing to the complexity of its parameters, an FC layer has sufficient capacity to manage high-dimensional tasks, so a large amount of memory and powerful computing capabilities become essential requirements. However, the large number of parameters in an FC layer greatly limits the practical application of this model. To address this problem, we apply matrix optimization to an FC layer. First, an added penalty term properly maintains the sparsity of the imposed weights. Second, a rank constraint is applied to the two components of the factorized weight matrix. Our compression algorithm can effectively reduce the number of required network parameters, which not only reduces the computational complexity of the network but also results in better generalizability on a test dataset. Finally, the effectiveness of the proposed method is verified in two different computer vision task domains. Experiments show that our sparse reduced-rank method achieves a better compression ratio with a lower accuracy loss relative to the competing approaches. The code is available at https://github.com/cheer79/Compress_FC .
Mingliang Zhou 0001, Xuekai Wei, Yong Feng 0002, Tao Xiang 0001, Bin Fang 0001, Zhaowei Shang, Fan Jia 0005, Xu Zhuang, Huayan Pu, Jun Luo 0003
ACM Trans. Multim. Comput. Commun. Appl.3
2025 DTSD: A Dual Teacher-Student-Based Discrimination Model for Anomaly Detection
abstract
The rapid development of computer vision technology for detecting anomalies in industrial products has received unprecedented attention. In this article, we propose a dual teacher–student-based discrimination (DTSD) model for anomaly detection, which combines the advantages of both embedding-based and reconstruction-based methods. First, the DTSD builds a dual teacher‒student architecture consisting of a pretrained teacher encoder with frozen parameters, a student encoder, and a student decoder. By distillation of knowledge from the teacher encoder, the two teacher‒student modules acquire the ability to capture both local and global anomaly patterns. Second, to address the issue of poor reconstruction quality faced by previous reconstruction-based approaches in some challenging cases, the model employs a feature bank that stores encoded features of normal samples. By incorporating template features from the feature bank, the student decoder receives explicit guidance to enhance the quality of reconstruction. Finally, a segmentation network is utilized to adaptively integrate multiscale anomaly information from the two teacher–student modules, thereby improving segmentation accuracy. Extensive experiments demonstrate that our method outperforms existing state-of-the-art approaches. The code of DTSD is publicly available at https://github.com/Math-Computer/DTSD .
Weizhi Xian, Xuekai Wei, Jielu Yan, Yueting Huang, Kunyin Guo, Weijia Jia 0001, Mingliang Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 VideoGNN: Video Representation Learning via Dynamic Graph Modelling
abstract
Graphs offer a flexible structure for vision tasks, with CNNs and Transformers conditioned as two specific cases of graph structures. In CNNs, the input images are treated as graphs where only neighboring patches are connected, whereas Transformers view images as fully connected graphs. To leverage the potential of graphs in video representation learning, effective graph generation and training methods are crucial. To this end, we propose VideoGNN, which represents the video as a discrete time dynamic graph and learns the dynamic graph efficiently. Given the multitude of frames in videos, we introduce an efficient graph generation module characterized by low complexity and high quality, facilitating the transformation of videos into dynamic graphs. Additionally, we introduce a dual-view graph neural network to capture spatial and temporal information from the generated dynamic graphs. Then, a sequential model is applied to capture the long-term temporal information and generate the final frame embeddings. Experiments demonstrate that VideoGNN can achieve competitive results in terms of graph quality assessment and video downstream tasks. The codes are available at https://github.com/Dodo-D-Caster/VideoGNN .
Mingliang Zhou 0001, Jun Luo 0006, Huayan Pu, Leong Hou U, Xuekai Wei, Weijia Jia 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2024 HFGNN: Efficient Graph Neural Networks Using Hub-Fringe Structures
abstract
Existing message passing-based and transformer-based graph neural networks (GNNs) cannot satisfy requirements for learning representative graph embeddings due to restricted receptive fields, redundant message passing, and reliance on fixed aggregations. These methods face scalability and expressivity limitations from intractable exponential growth or quadratic complexity, restricting interaction ranges and information coverage across large graphs. Motivated by the analysis of long-range graph structures, we introduce a novel Graph Neural Network called Hub-Fringe Graph Neural Network (HFGNN). Our Hub-Fringe structure, drawing inspiration from the graph indexing technique known as Hub Labeling, offers a straightforward and effective approach for learning scalable graph representations while ensuring comprehensive coverage of information. HFGNN leverages this structure to enable selective propagation of relevant embeddings through a carefully designed message function. Theoretical analysis is presented to show the expressivity and scalability of the proposed method. Empirically, HFGNN exceeds standard GNNs on tasks including classification and regression, especially for large, long-range graphs where scalability and coverage matter. Ablation studies further confirm the benefits of our hub-fringe based graph neural network, including improved expressivity and scalability. The source codes is available at https://github.com/nick12340/HFGNN.
Pak Lon Ip, Shenghui Zhang, Xuekai Wei, Tsz Nam Chan, Leong Hou U
ICDM3
2024 A Rate Control Scheme for VVC Intercoding Using a Linear Model
abstract
Versatile video coding (VVC) aims to achieve high compression but also issues like varying content/network conditions. Existing rate control (RC) methods struggle to achieve optimal quality under these complex scenarios. This paper proposes a novel RC scheme for VVC based on a linear model. The Lagrange minimization multiplier is introduced under bit budget constraints, allowing optimized bit allocation. RC optimization is formulated as a convex solution, and is derived into the optimal quantization parameter (QP) for RC. Experimental analysis demonstrates the proposed linear model-based RC algorithm performances are better compared to other state-of-the-art methods due to their use of a linear model and optimal QP determination.
Heqiang Wang, Xuekai Wei, Weizhi Xian, Jun Luo 0006, Huayan Pu, Zhigang Chu, Xin Wang 0051, Xueyong Xu, Chang Lu 0005, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.2
2024 Saliency and Depth-Aware Full Reference 360-Degree Image Quality Assessment
abstract
With the widespread adoption of virtual reality and 360-degree video, there is a pressing need for objective metrics to assess quality in this immersive panoramic format reliably. However, existing image quality assessment models developed for traditional fixed-viewpoint content do not fully consider the specific perceptual issues involved in 360-degree viewing. This paper proposes a 360-degree image full-reference quality assessment (FR-IQA) methodology based on a multi-channel architecture. The proposed 360-degree FR-IQA method further optimizes and identifies the distorted image quality using two easily obtained useful saliency and depth-aware image features. The convolutional neural network (CNN) is designed for training. Furthermore, the proposed method accounts for predicting user viewing behaviors within 360-degree images, which will further benefit the multi-channel CNN architecture and enable the weighted average pooling of the predicted FR-IQA scores. The performance is evaluated on publicly available databases to demonstrate the advantages brought by the proposed multi-channel model in performance evaluation and cross-database evaluation experiments, where it outperforms other state-of-the-art ones. Moreover, an ablation study exhibits good generalization ability and robustness.
Xuekai Wei, Qunyue Huang, Bin Fang 0001, Lei Ouyang, Weizhi Xian, Jun Luo 0003, Huayan Pu, Xueyong Xu, Chang Lu 0005, Hao Nan, Xu Liu 0006, Yachao Li 0001, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.1
2024 Anomaly Detection Integration-Framework for Network Services in Computer Education Systems
abstract
Public computer education systems provide students essential opportunities to enhance computer literacy and information skills. However, the widespread adoption of online education technology exposes the field to several critical security risks. Threats, such as malware infections, data breaches, and other network intrusions, are all challenging the security of education systems, posing potential hazards to students’ personal information and even the entire teaching environment. To spur further work into specialized anomaly detection techniques for computer education, this paper presents an anomaly detection framework tailored for network services in computer education environments to safeguard these systems. Specifically, the proposed approach learns from large-scale online educational traffic data to classify the security state into five alert levels, enabling more granular anomaly detection and analysis. To assess their detection performance, deep learning and traditional machine learning algorithms are implemented and compared for multi-class intrusion classification. The results show that the proposed framework provides an effective security solution to bolster the integrity and stability of computer education systems against evolving network threats, enhancing threat intelligence to inform proactive security by detecting and characterizing anomalies through multilevel classification.
Shouhong Yang, Xuekai Wei, Huayan Pu, Jun Luo 0006, Hong Yue, Fei Cheng 0001, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.5
2024 An End-to-End Video Coding Method via Adaptive Vision Transformer
abstract
Deep learning-based video coding methods have demonstrated superior performance compared to classical video coding standards in recent years. The vast majority of the existing deep video coding (DVC) networks are based on convolutional neural networks (CNNs), and their main drawback is that since CNNs are affected by the size of the receptive field, they cannot effectively handle long-range dependencies and local detail recovery. Therefore, how to better capture and process the overall structure as well as local texture information in the video coding task is the core issue. Notably, the transformer employs a self-attention mechanism that captures dependencies between any two positions in the input sequence without being constrained by distance limitations. This is an effective solution to the problem described above. In this paper, we propose end-to-end transformer-based adaptive video coding (TAVC). First, we compress the motion vector and residuals through a compression network built on the vision transformer (ViT) and design the motion compensation network based on ViT. Second, based on the requirement of video coding to adapt to different resolution inputs, we introduce a position encoding generator (PEG) as adaptive position encoding (APE) to maintain its translation invariance across different resolution video coding tasks. The experiment shows that for multiscale structural similarity index measurement (MS-SSIM) metrics, this method exhibits significant performance gaps compared to conventional engineering codecs, such as [Formula: see text], [Formula: see text], and VTM-15.2. We also achieved a good performance improvement compared to the CNN-based DVC methods. In the case of peak signal-to-noise ratio (PSNR) evaluation metrics, TAVC also achieves good performance.
Mingliang Zhou 0001, Zhaowei Shang, Huayan Pu, Jun Luo 0006, Xiaoxu Huang, Shilong Wang 0001, Huajun Cao, Xuekai Wei, Weizhi Xian
Int. J. Pattern Recognit. Artif. Intell.9
2024 TSAD: Two-Stage Separable Adversarial Distortion-Based Robust Watermarking Framework for Diffusion Tensor Imaging
abstract
Recent deep learning-based watermarking methods have achieved impressive results. However, they struggle with unknown distortions and often suffer from poor generalization, slow convergence, unstable training, and degraded visual quality in watermarked images. To address the above problems, this paper proposes a two-stage separable adversarial distortion (TSAD)-based robust watermarking algorithm for diffusion tensor imaging (DTI). The algorithm uses a noise-free end-to-end network in the first stage for learning and training DTI images. In the second stage, it fixes the watermark embedding network trained in the first stage, interacts the noise distortion network with the watermark extraction network to perform adversarial training for improving robustness. Experimental results show that our method achieves comparable or better robustness to seen distortions and better robustness to unseen distortions, along with enhanced stability, faster convergence, and improved visual quality in watermarked DTI images.
Zhi Li 0012, Zhangyu Liu, Hong Yue, Fei Cheng 0001, Qin Mao, Xuekai Wei, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.9
2024 Recent Advances in Rate Control: From Optimization to Implementation and Beyond
abstract
Video coding is a video compression technique that compresses the original video sequence to produce a smaller archive file or reduce the transmission bandwidth under constraints on the visual quality loss. Rate control (RC) plays a critical role in video coding. It can achieve stable stream output in practical applications, especially real-time video applications such as video conferencing or game live streaming. Most RC algorithms either directly or indirectly characterise the relationship between the bit rate (R) and quantisation (Q) and then allocate bits to every coding unit so as to guarantee the global bit rate and video quality level. This paper comprehensively reviews the classic RC technologies used in international video standards of past generations, analyses the mathematical models and implementation mechanisms of various schemes, and compares the performance of recent state-of-the-art RC algorithms. Finally, we discuss future directions and new application areas for RC methods. We hope that this review can help support the development, implementation, and application of RC for new video coding standards.
Xuekai Wei, Mingliang Zhou 0001, Heqiang Wang, Lei Chen 0093, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.1
2024 EF-DETR: A Lightweight Transformer-Based Object Detector With an Encoder-Free Neck
abstract
Object detection plays a key role in helping to enable industrial quality control and safety monitoring. This article introduces a lightweight and efficient transformer-based object detection network called the encoder-free DEtection TRansformer (EF-DETR). This novel architecture enhances the DETR model through a redesigned network structure, leading to improved accuracy in object detection and a more lightweight network. To address the issue of suboptimal object detection accuracy, especially for small objects in the DETR model, we introduce a multiscale feature extractor and a high-efficiency feature fusion module. These components facilitate the direct extraction of fine-grained features, thereby enabling effective object detection. Departing from the use of a high-complexity encoder structure, we explore the utilization of an encoder-free neck structure to reduce the network's computational complexity. In addition, to expedite convergence, denoising training is incorporated into the decoder. This article presents extensive experiments, and the EF-DETR demonstrates strong performance on the MS COCO2017 dataset compared to other popular models.
Jingnan Song, Mingliang Zhou 0001, Xuekai Wei, Huayan Pu, Jun Luo 0003, Weijia Jia 0001
IEEE Trans. Ind. Informatics4
2024 Deep Dual-Stream Convolutional Neural Networks for Cardiac Image Semantic Segmentation
abstract
Cardiac image segmentation is essential when applying biomedical informatics to improve industrial healthcare applications. To extract context and detailed information more efficiently and further improve cardiac image segmentation accuracy, we present a novel deep dual-stream convolutional neural network (CNN) for cardiac image semantic segmentation in this article. We use a body stream and a shape stream, respectively, in this method. First, in the body stream we propose integrating a gated fully fusion module to fuse multilevel features in the encoder and decoder paths. In addition, we integrate a feature aggregation module to extract the multiscale context. Second, in the shape stream, we propose using a gated shape CNN exploiting multilevel context to extract detailed information, such as boundary and shape features. Finally, we apply a multitask loss function to align the predicted masks with the ground truth labels. Our experiments on the public cardiac magnetic resonance image dataset show significant performance in the left and right ventricular cavities and myocardium compared to the state-of-the-art algorithms.
Hengqi Hu, Bin Fang 0001, Yuting Ran, Xuekai Wei, Weizhi Xian, Mingliang Zhou 0001, Sam Kwong
IEEE Trans. Ind. Informatics4
2024 Image Quality Assessment: Measuring Perceptual Degradation via Distribution Measures in Deep Feature Spaces
abstract
This study aims to develop advanced and training-free full-reference image quality assessment (FR-IQA) models based on deep neural networks. Specifically, we investigate measures that allow us to perceptually compare deep network features and reveal their underlying factors. We find that distribution measures enjoy advanced perceptual awareness and test the Wasserstein distance (WSD), Jensen-Shannon divergence (JSD), and symmetric Kullback-Leibler divergence (SKLD) measures when comparing deep features acquired from various pretrained deep networks, including the Visual Geometry Group (VGG) network, SqueezeNet, MobileNet, and EfficientNet. The proposed FR-IQA models exhibit superior alignment with subjective human evaluations across diverse image quality assessment (IQA) datasets without training, demonstrating the advanced perceptual relevance of distribution measures when comparing deep network features. Additionally, we explore the applicability of deep distribution measures in image super-resolution enhancement tasks, highlighting their potential for guiding perceptual enhancements. The code is available on website. (https://github.com/Buka-Xing/Deep-network-based-distribution-measures-for-full-reference-image-quality-assessment).
Xingran Liao, Xuekai Wei, Mingliang Zhou 0001, Zhengguo Li, Sam Kwong
IEEE Trans. Image Process.2
2024 Low-Light Enhancement Method Based on a Retinex Model for Structure Preservation
abstract
Enhancing low-light image visibility is a critical task in computer vision since it helps to improve input for high-level algorithms. High-quality images typically have clear structural information. In previous studies, due to the lack of proper structural guidance, restored images had some problems, such as unclear structural areas and overexposed or underexposed local areas. To address the above problems, in this paper, we introduce a coefficient of variation (COV) with excellent performance in maintaining structural information, and then we propose a low-light image enhancement method that utilizes the COV to extract structural information from images. First, we apply a traditional retinex model to estimate both reflectance and illumination. Second, we use the COV to indicate the degree of dispersion of the input sample, which enables us to obtain a robust structure-distinguishing weight map for low-light images. The weight map is adaptively divided to obtain a structural weight map, which is then used to enhance the gradient image. This process is applied before the reflectance layer of the retinex model. Finally, the result is obtained by using the block coordinate descent method. According to extensive experiments, outstanding results can be achieved by our proposed method in terms of both subjective and objective evaluation metrics in comparison with other state-of-the-art methods. The source code is available at our website.
Mingliang Zhou 0001, Xingtai Wu, Xuekai Wei, Tao Xiang 0001, Bin Fang 0001, Sam Kwong
IEEE Trans. Multim.3
2024 A Deep Retinex-Based Low-Light Enhancement Network Fusing Rich Intrinsic Prior Information
abstract
Images captured under low-light conditions are characterized by lower visual quality and perception levels than images obtained in better lighting scenarios. Studies focused on low-light enhancement techniques seek to address this dilemma. However, simple image brightening results in significant noise, blurring, and color distortion. In this paper, we present a low-light enhancement (LLE) solution that effectively synergizes Retinex theory with deep learning. Specifically, we construct an efficient image gradient map estimation module based on convolutional networks that can efficiently generate noise-free image gradient maps to assist with denoising. Second, to improve upon the traditional optimization model, we design a matrix-preserving optimization method (MPOM) coupled with deep learning modules, and it exhibits high speed and low memory consumption. Third, we incorporate image structure, image texture, and implicit prior information to optimize the enhancement process for low-light conditions and overcome prevailing limitations, such as oversmoothing, significant noise, and so forth. Through extensive experiments, we show that our approach has notable advantages over the existing methods and demonstrate superiority and effectiveness, surpassing the state-of-the-art methods by an average of 1.23 dB in PSNR for the LOL and VE-LOL datasets. The code for the proposed method is available in a public repository for open-source use: https://github.com/luxunL/DRNet .
Xuekai Wei, Xiaofeng Liao 0001, Fan Jia 0005, Xu Zhuang, Mingliang Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2024 SecEG: A Secure and Efficient Strategy against DDoS Attacks in Mobile Edge Computing
abstract
Application-layer distributed denial-of-service (DDoS) attacks incapacitate systems by using up their resources, causing service interruptions, financial losses, and more. Consequently, advanced deep-learning techniques are used to detect and mitigate these attacks in cloud infrastructures. However, in mobile edge computing (MEC), it becomes economically impractical to equip each node with defensive resources, as these resources may largely remain unused in edge devices. Furthermore, current methods are mainly concentrated on improving the accuracy of DDoS attack detection and saving CPU resources, neglecting the effective allocation of computational power for benign tasks under DDoS attacks. To address these issues, this paper introduces SecEG, a secure and efficient strategy against DDoS attacks for MEC that integrates container-based task isolation with lightweight online anomaly detection on edge nodes. More specifically, a new model is proposed to analyze resource contention dynamics between DDoS attacks and benign tasks. Subsequently, by employing periodic packet sampling and real-time attack intensity predicting, an autoencoder-based method is proposed to detect DDoS attacks. We leverage an efficient scheduling method to optimize the edge resource allocation and the service quality for benign users during DDoS attacks. When executed in the real-world edge environment, our experimental findings validate the efficacy of the proposed SecEG strategy. Compared to conventional methods, the service rate of benign requests increases by 23% under intense DDoS attacks, and the CPU resource is saved up to 35%.
Tianhui Meng, Jianxiong Guo, Xuekai Wei, Weijia Jia 0001
ACM Trans. Sens. Networks4
2024 A spatiotemporal and motion information extraction network for action recognition
Xianmin Wang, Mingliang Zhou 0001, Xuekai Wei, Xiaojun Ren, Xuemei Zong
Wirel. Networks4
2023 A Lightweight Multi-Scale Based Attention Network for Image Super-Resolution
abstract
In this paper, we propose a lightweight multi-scale based attention network (MBAN) for single-image super-resolution (SISR). First, a deep feature transform block (DFTB) is designed for multi-scale feature extraction; this block combines group convolution and improved channel attention (ICA) for performance purposes while remaining sufficiently lightweight. Second, a dual multi-scale attention block (DMAB) is proposed for long-range information interaction; this block employs different window sizes for self-attention (SA) and short connections between different branches to achieve multiscale attention interaction. Finally, our MBAN is constructed by cascaded multi-scale based attention blocks (MBABs) that perform detail restoration; these blocks simultaneously extract multi-scale local features and integrate multi-scale global features with the DFTBs and DMABs. Extensive experiments suggest the superiority of our MBAN over the state-of-the-art (SOTA) lightweight SR methods in terms of both quantitative metrics and visual quality.
Yanjie Yang, Jun Luo 0006, Huayan Pu, Mingliang Zhou 0001, Xuekai Wei, Taiping Zhang, Zhaowei Shang
IECON5
2023 A Two-Stage Three-Dimensional Attention Network for Lightweight Image Super-Resolution
abstract
In recent years, single image super-resolution (SISR) methods using convolutional neural networks (CNN) have achieved satisfactory performance. Nevertheless, the large model scale and the slow inference speed of these methods greatly limit the application scenarios. In this paper, we propose a two-stage three-dimensional attention network (ATTNet) for lightweight image super-resolution. First, we put forward the spatial feature encoder–decoder (SFE-D) with a spatial attention mechanism. Next, the channel transposed attention module (CTAM) with a channel self-attention mechanism is designed. Both the modules are used for fine feature extraction in the low-resolution stage. Finally, the content-based pixel recombination module (CPRM) is proposed to reconstruct the detailed content with a joint attention mechanism in the high-resolution stage. According to our experimental results, significant performance in terms of the quantitative metrics and the subjective visual quality can be achieved on average compared with the state-of-the-art lightweight SISR algorithms.
Lei Chen 0093, Yanjie Yang, Xu Zhuang, Qin Mao, Hong Yue, Xuekai Wei, Fei Cheng 0001, Xuemei Zong, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.7
2023 Full-Reference Image Quality Assessment via Low-Level and High-Level Feature Fusion
abstract
We propose a full-reference image quality assessment (FR-IQA) method by incorporating low-level and high-level image features. First, in contrast to the preexisting deep IQA methods, which only use the features extracted by the deep network, we not only use the image gradient to replace the low-level features in the first two stages of the deep network, but also combine them with the middle-stage features of the deep network to construct the new low-level features. The deep features of shallow layers contain some unwanted noise and further result in a decline in IQA performance. Second, we combine the global features extracted by the self-attention-based model with the semantic features extracted by the convolutional neural network to form the high-level features. Instead of directly using the self-attention-based model trained on the classification task, we first train a no-reference (NR) IQA regression model on a larger dataset and then use the global features of this NR-IQA model. The self-attention-based model can capture the internal connections of the image and is more effective in extracting global information due to its larger perceptual field. In the final pooling stage, we combine the average pooling and the standard deviation pooling to obtain the dispersion and concentration of the similarity maps for a more comprehensive description of quality. Experiments show that our FR-IQA method is able to obtain competitive results on three standard IQA datasets.
Chao Wu 0006, Xiaofeng Liao 0001, Hong Yue, Xueyong Xu, Xuekai Wei, Dingcheng Wu, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.5
2023 Toward Low-Latency and High-Quality Adaptive 360$^\circ$ Streaming
abstract
Advanced 360$^\circ$video streaming is essential for high-quality communication and industrial video applications, supporting novel and interactive visual experiences that promote consumption. However, due to the inherent fluctuations in communication networks, hardware resources, and power costs, a tradeoff needs to be achieved between occupying network bandwidth and guaranteeing 360$^\circ$visual quality. To address this issue, this article proposes a quality-aware global optimization solution that combines different types of sensory characteristics to further enhance the 360$^\circ$streaming efficiency in cellular networks. First, considering the characteristics of 360$^\circ$video, each frame is divided into three different regions based on a proposed field-of-view prediction method. Second, a new region-based rate-distortion model reflecting the bitrate-quality relationship is proposed. The divided regions are represented by different rate-distortion models. In addition, a rate-distortion model parameter update strategy with robustness to region changes is proposed to further guarantee transmission performance. Finally, we propose a globally optimized adaptive bitrate allocation algorithm to optimize 360$^\circ$mobile streaming, which uses both the rate-distortion models and viewpoint prediction results. Evaluation results indicate that the proposed method outperforms state-of-the-art approaches in terms of several quality of experience objectives under various network conditions.
Xuekai Wei, Mingliang Zhou 0001, Weijia Jia 0001
IEEE Trans. Ind. Informatics1
2023 Joint Decision Tree and Visual Feature Rate Control Optimization for VVC UHD Coding
abstract
In this paper, a joint decision tree and visual feature optimization rate control scheme for ultrahigh-definition (UHD) versatile video coding (VVC) is proposed. First, we design a new rate-distortion (R-D) model for UHD videos, and we establish a decision-tree-based multiclass classification scheme to improve the prediction accuracy of the R-D model by fully considering visual features. Second, based on the proposed R-D model, the globally optimal solution is obtained through convex optimization. Finally, we embed our algorithm into the latest VVC reference software, VTM 10.2. According to our experimental results, compared with the latest algorithm in VTM 10.2 and other state-of-the-art algorithms, our method can achieve significant bit rate reductions while maintaining a given peak signal-to-noise ratio (PSNR) or structural similarity index measure (SSIM).
Mingliang Zhou 0001, Xuekai Wei, Weijia Jia 0001, Sam Kwong
IEEE Trans. Image Process.2
2023 Vehicular Abandoned Object Detection Based on VANET and Edge AI in Road Scenes
abstract
Rapid processing of abandoned objects is one of the most important tasks in road maintenance. Abandoned object detection heavily relies on traditional object detection approaches at a fixed location. However, detection accuracy and range are still far from satisfactory. This study proposes an abandoned object detection approach based on vehicular ad-hoc networks (VANETs) and edge artificial intelligence (AI) in road scenes. We propose a vehicular detection architecture for abandoned objects to achieve task-based AI technology for large-scale road maintenance in mobile computing circumstances. To improve detection accuracy and reduce repeated detection rates in mobile computing, we propose a detection algorithm that combines a deep learning network and a deduplication module for high-frequency detection. Finally, we propose a location estimation approach for abandoned objects based on the World Geodetic System 1984 (WGS84) coordinate system and an affine projection model to accurately compute the positions of abandoned objects. Experimental results show that our proposed algorithm achieves an average accuracy of 99.57% and 53.11% on the two datasets, respectively. Additionally, our whole system achieves real-time detection and high-precision localization performance on real roads.
Gang Wang 0023, Mingliang Zhou 0001, Xuekai Wei, Guang Yang 0006
IEEE Trans. Intell. Transp. Syst.3
2023 Low-light Image Enhancement via a Frequency-based Model with Structure and Texture Decomposition
abstract
This article proposes a frequency-based structure and texture decomposition model in a Retinex-based framework for low-light image enhancement and noise suppression. First, we utilize the total variation-based noise estimation to decompose the observed image into low-frequency and high-frequency components. Second, we use a Gaussian kernel for noise suppression in the high-frequency layer. Third, we propose a frequency-based structure and texture decomposition method to achieve low-light enhancement. We extract texture and structure priors by using the high-frequency layer and a low-frequency layer, respectively. We present an optimization problem and solve it with the augmented Lagrange multiplier to generate a balance between structure and texture in the reflectance map. Our experimental results reveal that the proposed method can achieve superior performance in naturalness preservation and detail retention compared with state-of-the-art algorithms for low-light image enhancement. Our code is available on the following website. 1
Mingliang Zhou 0001, Hongyue Leng, Bin Fang 0001, Tao Xiang 0001, Xuekai Wei, Weijia Jia 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Global Optimization Solution for Dynamic Adaptive 360-Degree Streaming
abstract
This paper proposes a global optimization solution that combines different types of perceptual information to further improve transmission efficiency. First, a new rate-distortion (R-D) model is proposed to reflect the characteristics of 360degree video. Second, a globally optimized adaptive bitrate control algorithm is proposed using both the R-D models to adjust the bitrate for each tile in a segment. Finally, a new model parameter updating strategy that is robust to quality variations is proposed to further reduce prediction errors. Comparison results indicate that the proposed method can outperform state-of-the-art methods in terms of various quality of experience (QoE) objectives1.
Xuekai Wei, Mingliang Zhou 0001, Weijia Jia 0001
ICASSP1
2022 Contrastive distortion-level learning-based no-reference image-quality assessment
abstract
A contrastive distortion-level learning-based no-reference image-quality assessment (NR-IQA) framework is proposed in this study to further effectively model various distortion types with the same or different distortion levels. The proposed method aims to improve the prediction accuracy of NR-IQA. The proposed method consists of three parts: multiscale distortion-level representation learning, single-image NR-IQA, and a representation affinity module, which can reduce NR-IQA computational complexity while maintaining a low-distortion representation of high-distortion inputs. The proposed NR-IQA method aims to extract distributional features of samples in real distorted images and predict ambiguity based on distortion-level learning. Experimental results show that by comparing on many NR-IQA data sets the proposed method can outperform state-of-the-art methods.
Xuekai Wei, Jin Li 0002, Mingliang Zhou 0001, Xianmin Wang
Int. J. Intell. Syst.1
2022 A Hybrid Control Scheme for 360-Degree Dynamic Adaptive Video Streaming Over Mobile Devices
abstract
A 360-degree streaming system can provide immersive, interactive, and autonomous experiences surrounding the user by means of viewpoint changes to see different angles of a 360-degree video. However, due to the limited capacity and highly dynamic conditions of cellular networks, high-resolution 360-degree video playback over mobile devices often suffers from playback freezing, and bandwidth waste is inevitably incurred in delivering out-of-view video data. In this paper, a hybrid control scheme is presented for segment-level continuous bitrate selection and tile-level bitrate allocation for 360-degree streaming over mobile devices to increase users’ quality of experience. First, a deep reinforcement learning (RL) method is proposed to predict the segment bitrate and avoid playback freezing. Second, a viewpoint-prediction-map-based cooperative bargaining game theory is proposed for bitrate allocation optimization to choose a suitable bitrate for each tile to reduce unreasonable bandwidth waste. The proposed scheme is compared with state-of-the-art approaches under a wide variety of mobile network conditions with multiple viewpoint traces and 360-degree video contents. The experimental results indicate that the proposed method outperforms the compared state-of-the-art approaches in terms of various experimental objectives on mobile devices.
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Weijia Jia 0001
IEEE Trans. Mob. Comput.1
2021 Joint Reinforcement Learning and Game Theory Bitrate Control Method for 360-Degree Dynamic Adaptive Streaming
abstract
A joint reinforcement learning (RL) and game theory method is presented for segment-level continuous bitrate selection and tile-level bitrate allocation in tile-based 360-degree streaming to increase users’ quality of experience (QoE). First, a viewpoint prediction method based on single-user (SU) viewpoint traces and the saliency map (SM) model is presented to model viewing behaviours. Second, an RL method is proposed to predict segment bitrate and a cooperative bargaining game theory is proposed for bitrate allocation optimization to choose a suitable bitrate for every tile with the help of the viewpoint prediction map. Performance evaluation results indicate that the proposed method can outperform the state-of-the-art methods in terms of different QoE objectives.
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Tao Xiang 0001
ICASSP1
2021 QOE-Based Neural Live Streaming Method with Continuous Dynamic Adaptive Video Quality Control
abstract
In this paper, a quality of experience (QoE)-based neural live streaming method with dynamic adaptive video quality control is developed to improve streaming performance. First, the dynamic adaptive streaming issue is formulated as a Markov decision process (MDP) problem. Second, an reinforcement learning (RL)-based approach is proposed as an appropriate solution, where the client functions as an RL agent and the environment is made up of various networks. User QoE is the reward by mutual consideration of video quality and play-back state. Finally, to optimize the total reward, the RL algorithm chooses the required video quality for each video segment. Experimental results show that the proposed RL-based streaming algorithm outperforms state-of-the-art schemes in terms of both temporal and visual QoE metrics by a noticeable margin while guaranteeing application-level fairness when multiple clients share a bottlenecked network. The code is available on the following website: https://github.com/OpenCode007/ICME2021.
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Tao Xiang 0001
ICME1
2021 Reinforcement learning-based QoE-oriented dynamic adaptive streaming framework
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Shiqi Wang 0001, Guopu Zhu, Jingchao Cao
Inf. Sci.1
2021 Rate Control Method Based on Deep Reinforcement Learning for Dynamic Video Sequences in HEVC
abstract
Rate control (RC) plays a critical role in the transmission of high-quality video data under certain bandwidth restrictions in High Efficiency Video Coding (HEVC). Most current HEVC RC algorithms based on spatio-temporal information for rate-distortion (R-D) model parameters cannot effectively handle the cases with dynamic video sequences that contain fast moving objects, significant object occlusion or scene changes. In this paper, we propose an RC method based on deep reinforcement learning (DRL) for dynamic video sequences in HEVC to improve the coding efficiency. First, the rate control problem is formulated as a Markov decision process (MDP) problem. Second, with the MDP model, we develop a DRL-based algorithm to find the optimal quantization parameters (QPs) by training a deep neural network. The resulting intelligent agent selects the optimal RC strategy to reduce distortion, buffer and quality fluctuations by observing the current state of the encoder. The asynchronous advantage actor-critic (A3C) method is used to solve the MDP problem. Finally, the proposed DRL-based RC method is implemented in the newest video coding standard. Experimental results show that the proposed method offers substantially enhanced RC accuracy and consistently outperforms HEVC reference software and other state-of-the-art algorithms.
Mingliang Zhou 0001, Xuekai Wei, Sam Kwong, Weijia Jia 0001, Bin Fang 0001
IEEE Trans. Multim.2
2020 Global Rate-Distortion Optimization-Based Rate Control for HEVC HDR Coding
abstract
High dynamic range (HDR) video compression technology, which is capable of delivering a wider range of luminance and a larger colour gamut than standard dynamic range (SDR) technology, has been widely used in recent years in many fields, including industrial image processing, digital entertainment, and machine vision. Rate control (RC) is of paramount importance to HDR compression and transmission; accordingly, an RC scheme for HDR in High Efficiency Video Coding (HEVC) is proposed in this paper. First, considering the HDR characteristics, we propose an HDR-Visual Difference Predictor (VDP)-2-based rate-distortion (R-D) model to improve the coding performance. Second, we directly utilize $\lambda $ rather than the bit rate in the optimization process to obtain the optimal solution. Finally, we propose a new model parameter estimation method to further reduce the RC errors. According to our experimental results, significant bit rate reductions in terms of HDR-VDP-2, the Video Quality Metric (VQM) and the mean peak-signal-to-noise ratio (mPSNR) can be achieved on average compared with the state-of-the-art algorithm used in HM16.19.
Mingliang Zhou 0001, Xuekai Wei, Shiqi Wang 0001, Sam Kwong, Chi-Keung Fong, Peter Hon-Wah Wong, Wilson Y. F. Yuen
IEEE Trans. Circuits Syst. Video Technol.2
2020 Just Noticeable Distortion-Based Perceptual Rate Control in HEVC
abstract
In this paper, we propose a just noticeable distortion (JND)-based perceptual rate control method for high efficiency video coding (HEVC). First, the JND factor of a coding unit has been mathematically shown to be an approximation of the average pixel-level JND weight, which means that it can also be used as a weight for bitrate allocation. Second, rate-distortion (R-D) modelling is conducted based on the JND factor. Finally, the proposed R-D model is integrated into an existing rate control framework to improve the coding efficiency, and the proposed algorithm is implemented in the newest video coding standard. As the experimental results reveal, compared with HEVC reference software, our algorithm achieves significantly improved coding performance, subjective coding quality and bitrate accuracy.
Mingliang Zhou 0001, Xuekai Wei, Sam Kwong, Weijia Jia 0001, Bin Fang 0001
IEEE Trans. Image Process.2
2019 SSIM-Based Global Optimization for CTU-Level Rate Control in HEVC
abstract
In this paper, we propose a coding tree unit (CTU)-level rate control scheme from the perspective of SSIM-based rate-distortion optimization to improve the coding efficiency. First, we establish the SSIM-based rate-distortion model based on the divisive normalization scheme, which characterizes the relationship between the local visual quality and the coding bits. Then, the established model is applied to the CTU-level rate control and transformed into a global optimization problem solved by convex optimization. Finally, a new model parameter updating strategy for the CTU-level rate control is presented that is robust to scene variations. Our algorithm can achieve optimal CTU-level bit allocation given the bit-rate budget. The experimental results show that our algorithm substantially enhances the coding performance and consistently outperforms both the rate control scheme in the HEVC reference software and existing algorithms in terms of rate-perceptual distortion performance using different test configurations.
Mingliang Zhou 0001, Xuekai Wei, Shiqi Wang 0001, Sam Kwong, Chi-Keung Fong, Peter Hon-Wah Wong, Wilson Y. F. Yuen, Wei Gao 0003
IEEE Trans. Multim.2
2018 Cooperative Bargaining Game-Based Multiuser Bandwidth Allocation for Dynamic Adaptive Streaming Over HTTP
abstract
Dynamic adaptive streaming over HTTP (DASH) has emerged as an efficient technology for video streaming. For a DASH system, a most common case is that a limited server bandwidth is competed by multiusers. In order to improve user quality of experience (QoE) and guarantee fairness, we propose to use the game theory in a proxy server to allocate the bandwidth collaboratively for multiusers. By taking user buffer length, received video bit rates, video qualities, etc., into account, the bandwidth allocation problem is formulated as a cooperative bargaining problem and the Nash bargaining solution (NBS) is obtained by convex optimization. The requested bit rate of users will be rewritten as the proxy calculated bit rate (i.e., NBS) when the user requested bit rate is larger. Experimental results demonstrate that user QoE and fairness can be improved significantly, i.e., the delay frequency and duration are smaller, and the received video qualities are higher and more stable, when comparing the proposed method with existing methods.
Hui Yuan 0001, Xuekai Wei, Fuzheng Yang 0001, Jimin Xiao, Sam Kwong
IEEE Trans. Multim.2