Yongxu Liu 0001

dblp:67/9559-1 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
17since 2021 · last 2025
0009-0008-5719-1107ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Towards Syn-to-Real IQA: A Novel Perspective on Reshaping Synthetic Data Distributions
abstract
Blind Image Quality Assessment (BIQA) has advanced significantly through deep learning, but the scarcity of large-scale labeled datasets remains a challenge. While synthetic data offers a promising solution, models trained on existing synthetic datasets often show limited generalization ability. In this work, we make a key observation that representations learned from synthetic datasets often exhibit a discrete and clustered pattern that hinders regression performance: features of high-quality images cluster around reference images, while those of low-quality images cluster based on distortion types. Our analysis reveals that this issue stems from the distribution of synthetic data rather than model architecture. Consequently, we introduce a novel framework SynDR-IQA, which reshapes synthetic data distribution to enhance BIQA generalization. Based on theoretical derivations of sample diversity and redundancy's impact on generalization error, SynDR-IQA employs two strategies: distribution-aware diverse content upsampling, which enhances visual diversity while preserving content distribution, and density-aware redundant cluster downsampling, which balances samples by reducing the density of densely clustered areas. Extensive experiments across three cross-dataset settings (synthetic-to-authentic, synthetic-to-algorithmic, and synthetic-to-synthetic) demonstrate the effectiveness of our method. The code is available at https://github.com/Li-aobo/SynDR-IQA.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong
NeurIPS3
2025 Compressing Vision Transformer from the View of Model Property in Frequency Domain
Zhenyu Wang 0008, Xuemei Xie, Hao Luo 0004, Weisheng Dong, Yongxu Liu 0001, Fan Wang 0019, Guangming Shi
Int. J. Comput. Vis.7
2025 Semi-Supervised Graph Constraint Dual Classifier Network With Unknown Class Feature Learning for Hyperspectral Image Open-Set Classification
abstract
In view of the practical value of open datasets of hyperspectral images (HSIs), HSI open-set classification (OSC) has attracted more and more attention. Existing HSI OSC methods are usually based on learning labeled samples to identify unknown classes. However, due to the complex high-dimensional characteristics of HSIs and the limited number of labeled samples, the recognition of unknown classes based only on limited labeled samples often has low and unstable accuracy. To address this problem, we propose a semi-supervised graph constraint dual classifier network (SSGCDCN) that can achieve efficient and stable OSC by learning unknown class features and relationships among samples. First, a dual classifier consisting of a multi-classifier and multiple binary classifiers is constructed, which has the ability to discover the unknown class samples by assigning and enabling pseudo-labels to participate in model training to achieve unknown class feature learning. Then, to improve the classification accuracy of both known and unknown classes, a homogeneous graph constraint is imposed on SSGCDCN to learn the relationship information among samples (including labeled and unlabeled samples). This constraint can bring the features of similar samples closer while pushing apart features of dissimilar samples. Experiments evaluated on three datasets demonstrate that the proposed method can obtain superior OSC performance than other state-of-the-art classification methods.
Na Li 0040, Xiaopeng Song, Yongxu Liu 0001, Wenxiang Zhu, Chuang Li 0005, Wei-Tao Zhang, Yinghui Quan
IEEE Geosci. Remote. Sens. Lett.3
2025 Scene-Modulated High-Order Statistical Representation Learning for No-Reference Super-Resolution Image Quality Assessment
abstract
With the rapid development of single image super-resolution (SR) technology, there is an urgent need to develop a fair no reference Super-Resolution image Quality Assessment (SRQA) method. Existing no reference SRQA methods primarily concentrate on SR artifacts including structural distortion and texture distortion by extracting spatial features, but ignore the inductive bias of Deep Neural Network (DNN)-based SR models. As a result, they function effectively for interpolation-based and dictionary-based algorithms, but struggle to perform as effectively with DNN-based SR algorithms. We found that the visual content generated by DNN-based SR models under different inductive biases often carries a content-invariant model-specific style, which can be captured by the correlations between hierarchical representation channels. To that end, we propose a novel Scene-modulated High-order Statistical Representation network (SmHSR) built on a multi-scale over-complete transformation. We quantify the perceptual quality of SR images as the shift of high-order statistical properties in their multi-scale over-complete representation, where intra-channel statistics are used to capture spatial correlations and inter-channel statistics are used to capture the inductive bias of SR models. In addition, the scene information implicit in the deep over-complete representation is used to modulate the high-order statistical properties, which simulates the top-down regulation of cognition on perception. Under the modulation of scene information, SmHSR can learn more sophisticated scene-aware statistical representation. The MultiLayer Perceptron (MLP) is used to map the high-order statistical representation to an overall quality. We test our method on multiple SR image quality databases. Experimental results show that our method outperforms the state-of-the-art SRQA methods.
Yongwei Mao, Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong
IEEE Trans. Circuits Syst. Video Technol.3
2025 Incremental Multitask Contrastive Learning Network for End-to-End Few-Shot Open-Set Classification of Hyperspectral Images
abstract
Hyperspectral image open-set classification has gained increasing attention due to its practical significance. However, existing approaches face two major challenges: (1) poor and unstable classification performance under limited labeled samples, and (2) the lack of end-to-end open-set classification frameworks. To address these issues, we propose an Incremental Multi-Task Contrastive Learning Network (IMTCLN), which integrates four learning tasks to achieve end-to-end open-set classification under few-shot conditions through feature sharing and multi-task collaboration. First, we introduce an expanded class labeling method in the model’s output layer, enabling end-to-end open-set classification. Second, among the four learning tasks, the supervised classification task learns the mapping between known-class samples and their labels using limited labeled data. To enhance classification performance under few-shot conditions, we design a semi-supervised Euclidean contrastive learning task, which improves intra-class compactness and inter-class separability by modeling homogeneous and heterogeneous sample relationships. Additionally, for effective unknown-class recognition, we propose a supervised Mahalanobis contrastive learning task, optimizing the Mahalanobis distance among known classes to identify unknown-class samples. Finally, to further enhance classification stability, we introduce an incremental learning task, which leverages pseudo-labeled unknown-class samples to learn their discriminative features, enabling robust discrimination between known and unknown classes. Extensive experiments on three public datasets demonstrate that IMTCLN significantly outperforms existing methods, particularly under extremely limited labeled samples, showcasing superior open-set classification performance and stability.
Na Li 0040, Xiaopeng Song, Wenxiang Zhu, Yongxu Liu 0001, Chuang Li 0005, Yinghui Quan
IEEE Trans. Geosci. Remote. Sens.4
2025 Forgetting the Background: A Masking Approach for Enhanced Infrared Small-Target Detection
abstract
Infrared small-target detection (ISTD) in a single frame is an essential, yet challenging task due to its small size of targets, weak energy, and clutter background. Current methods either design complex network architectures to facilitate multilevel information interaction (e.g., DNA-Net and UIU-Net) or introduce structural texture priors to enhance feature discrimination (e.g., SRNet and CSRNet). However, both methods fail to explicitly distinguish or suppress the interference of complex background from infrared small targets, which makes them easy to “get lost” in clutter background with insufficient attention to the targets. In this work, we innovatively propose a novel background-masking approach (denoted as BGM) for ISTD. The proposed BGM aims to force the network to focus exclusively on the target by masking out irrelevant background information, thereby enhancing the network’s ability to detect weak and small infrared targets. Specifically, we present a new ISTD method that leverages a proxy training task with masking, enabling the network to simultaneously predict on both the original input and the masked data, where the background is randomly masked/forgotten. This strategy allows for a better concentration of the model on the shapeless targets rather than the cluttered background. The method is flexible with a simple U-shaped network without complicated manipulation and also computationally efficient without increasing the overall computational burden during inference. Extensive experiments demonstrate that our proposed BGM effectively enhances the detection performance of infrared small targets and achieves 70.8% mean intersection over union (mIoU) on IRSTD-1K. The source code would be available athttps://github.com/ZhihaoMa123/BGM
Yongxu Liu 0001, Wenxiang Zhu, Na Li 0040, Chuang Li 0005, Zhenyu Wang 0008, Wei Feng 0004, Junzheng Jiang, Yinghui Quan
IEEE Trans. Geosci. Remote. Sens.1
2025 HiCAL: Hierarchical Consistency-Based Active Learning for Drone-View Object Detection
abstract
The recent years have witnessed the great progress of drone-view object detection in both economic and military applications. Generally, the good performance of drone-view object detection requires a large amount of annotated data, which has imposed significant demands on human and material resources. To optimize the labelling expenses, previous work has introduced active learning to select the most valuable samples for annotation, and balances the annotation cost and model performance. However, existing active learning methods are primarily controlled by the “absolute” prediction of the model (e.g., the predicted categories for diversity, and the classification confidence for uncertainty). It would be highly misleading when the model outputs wrong prediction but with high confidence. This confident misleading is more severe in drone-view object detection as the targets are captured with varied viewpoints, illumination conditions, and possible occlusion. In this paper, we refresh the active learning with Perturbation Consistency Test (PCT), which transforms the absolute prediction into the relative error to address the situation where the absolute prediction is unreliable. The basic idea is to test the prediction consistency when the input samples are with/without perturbation, and regards the inconsistency as a measurement of the model’s resilience to guide the active selection. To this end, a Hierarchical Consistency-based Active Learning (HiCAL) is built, which constructs adversarially pair-wise inputs with hierarchical perturbation. The samples are perturbed with multi-granularity (i.e., pixel level, feature level, and object level) and afterwards, the entropy difference of the paired outputs before/after perturbation is calculated as the measurement. The samples with high difference are selected to follow a standard active learning loop. Experimental results show that HiCAL can achieve superior performance in different datasets and is easy to adapt to various types of object detectors. The code will be available on: https://github.com/zstar1003/HiCAL.
Yongxu Liu 0001, Qinghang Zhao, Jinjian Wu
IEEE Trans. Geosci. Remote. Sens.2
2024 Scaling and Masking: A New Paradigm of Data Sampling for Image and Video Quality Assessment
abstract
Quality assessment of images and videos emphasizes both local details and global semantics, whereas general data sampling methods (e.g., resizing, cropping or grid-based fragment) fail to catch them simultaneously. To address the deficiency, current approaches have to adopt multi-branch models and take as input the multi-resolution data, which burdens the model complexity. In this work, instead of stacking up models, a more elegant data sampling method (named as SAMA, scaling and masking) is explored, which compacts both the local and global content in a regular input size. The basic idea is to scale the data into a pyramid first, and reduce the pyramid into a regular data dimension with a masking strategy. Benefiting from the spatial and temporal redundancy in images and videos, the processed data maintains the multi-scale characteristics with a regular input size, thus can be processed by a single-branch model. We verify the sampling method in image and video quality assessment. Experiments show that our sampling method can improve the performance of current single-branch models significantly, and achieves competitive performance to the multi-branch models without extra model complexity. The source code will be available at https://github.com/Sissuire/SAMA.
Yongxu Liu 0001, Yinghui Quan, Guoyao Xiao, Jinjian Wu
AAAI1
2024 Bridging the Synthetic-to-Authentic Gap: Distortion-Guided Unsupervised Domain Adaptation for Blind Image Quality Assessment
abstract
The annotation of blind image quality assessment (BIQA) is labor-intensive and time-consuming, especially for authentic images. Training on synthetic data is expected to be beneficial, but synthetically trained models often suf-fer from poor generalization in real domains due to domain gaps. In this work, we make a key observation that introducing more distortion types in the synthetic dataset may not improve or even be harmful to generalizing au-thentic image quality assessment. To solve this challenge, we propose distortion-guided unsupervised domain adaptationfor BIQA (DGQA), a novel framework that leverages adaptive multi-domain selection via prior knowledge from distortion to match the data distribution between the source domains and the target domain, thereby reducing negative transfer from the outlier source domains. Extensive experiments on two cross-domain settings (synthetic distortion to authentic distortion and synthetic distortion to algorith-mic distortion) have demonstrated the effectiveness of our proposed DGQA. Besides, DGQA is orthogonal to existing model-based BIQA methods, and can be used in combi-nation with such models to improve performance with less training data.
Jinjian Wu, Yongxu Liu 0001, Leida Li
CVPR3
2024 Deep Multitask Learning with Graph Constraints for Hyperspectral Images Open-Set Classification
abstract
Existing methods for hyperspectral image classification (HSIC) typically assume a closed-set scenario, where all target types are known, and the classifier can assign only predefined classes to samples (pixels). However, in real remote sensing applications, open-set scenarios are common, where unknown classes exist. To address this problem, we propose a graph-constrained deep multi-task approach for open-set HSIC. Our method tackles the challenge of detecting unknown classes by integrating multiple-class classifiers and multiple binary classifiers. Additionally, to handle the limited labeled samples issue in HSIC, we propose utilizing homogeneous and heterogeneous graphs to constrain the two types of classifiers, thereby improving the accuracy of unknown class detection and known class classification. Experimental results on the Pavia University dataset demonstrate that our proposed method outperforms other closed-set and open-set classification methods significantly.
Na Li 0040, Xiaopeng Song, Yinghui Quan, Wenxiang Zhu, Yongxu Liu 0001
IGARSS5
2024 Global Feature and Semantic Information Extraction Network Based on Frozen SAM Encoder for Hyperspectral Image Classification
abstract
Nowadays, various types of foundational models have emerged, showcasing remarkable performance across a multitude of downstream tasks. However, in the domain of hyperspectral image classification (HSIC), substantial research is still required to effectively leverage the advantages of foundational models and adapt them to hyperspectral data. Consequently, we propose a HSIC algorithm based on a fixed-parameter SAM encoder. Specifically, the global feature extraction subnetwork integrates global patch information to obtain processed features. Subsequently, the semantic information extraction subnetwork is trained using cross-entropy to extract semantic features of categories, culminating in pixel-level classification. Experiments on two HSI datasets indicate that the proposed method can obtain better classification performance when compared with seven state-of-the-art methods.
Wenxiang Zhu, Deping Chen, Yinghui Quan, Liang Guo 0002, Yongxu Liu 0001, Na Li 0040
IGARSS5
2024 Self-Adaptive Global Feature Fusion Network With Spectral Prompt for Hyperspectral Image Classification
abstract
Nowadays, foundation models have demonstrated exceptional performance across numerous downstream tasks. However, the effective application of these models to hyperspectral image classification (HSIC) is challenged by the unique characteristics of hyperspectral data, including high dimensionality, high variability, and high spatial structure complexity. Therefore, methods need to be developed, which leverage the advantages of foundation models while addressing these challenges. First, a novel HSIC algorithm based on a frozen-parameter segment anything model (SAM) encoder, called SAGFFNet, is proposed. This framework represents the first attempt to use a frozen SAM encoder for global feature extraction and to use spectral dimension data as prompts, enabling precise global spatial-spectral feature extraction with the aid of spectral information. Second, by introducing the self-adaptive padding mechanism and the global feature extraction subnetwork (GFEsNet), the model is enabled to extract distinctive and discriminative features for each category from hyperspectral data through varying padding sizes, thereby enhancing the feature extraction and generalization capabilities of the foundation model. Subsequently, the spectral feature prompt subnetwork (SFPsNet) is designed to extract spectral feature information from samples of different classes as prompt features, assisting the framework in better understanding the global features extracted by GFEsNet. Finally, the semantic information decoder subnetwork (SIDsNet) is introduced as a semantic information decoder, achieving efficient fusion of global spatial-spectral features and spectral prompt features, which significantly improves classification performance. Experiments conducted on four hyperspectral image datasets show that the proposed method outperforms nine existing approaches in terms of classification accuracy.
Deping Chen, Wenxiang Zhu, Chuang Li 0005, Yongxu Liu 0001, Na Li 0040, Wei-Tao Zhang, Yinghui Quan
IEEE Trans. Geosci. Remote. Sens.4
2024 Blind Image Quality Assessment Based on Perceptual Comparison
abstract
Blind image quality assessment (BIQA) is a regression task with continuous label space, the feature space of which is expected to have a corresponding continuity in the target space. However, existing approaches typically learn quality score regression directly in an end-to-end fashion, which leaves networks susceptible to interference from task-agnostic information, and fails to capture the continuity of BIQA. In this work, by explicitly establishing inter-sample associations, a simple yet effective BIQA framework based on perceptual comparison is proposed to capture the continuity. To this end, besides the basic quality score regression, the relative quality scores between images are predicted to exploit the relative quality relationships between samples for optimizing the representation of image perceptual quality. In addition, based on the human perceptual characteristic, we derive a novel sample weighting strategy to dynamically adjust the weights for different samples in the network learning process for further improving the robustness of the model. The performances on both single-database and cross-database experiments achieve state-of-the-art, indicating the effectiveness of the proposed method. Besides, the proposed framework is model-agnostic, which can effectively improve the performance of the benchmark model with no extra inference cost.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.3
2023 Quality Assessment of UGC Videos Based on Decomposition and Recomposition
abstract
The prevalence of short-video applications imposes more requirements for video quality assessment (VQA). User-generated content (UGC) videos are captured under an unprofessional environment, thus suffering from various dynamic degradations, such as camera shaking. To cover the dynamic degradations, existing recurrent neural network-based UGC-VQA methods can only provide implicit modeling, which is unclear and difficult to analyze. In this work, we consider explicit motion representation for dynamic degradations, and propose a motion-enhanced UGC-VQA method based on decomposition and recomposition. In the decomposition stage, a dual-stream decomposition module is built, and VQA task is decomposed into single frame-based quality assessment problem and cross frames-based motion understanding. The dual streams are well grounded on the two-pathway visual system during perception, and require no extra UGC data due to knowledge transfer. Hierarchical features from shallow to deep layers are gathered to narrow the gaps from tasks and domains. In the recomposition stage, a progressively residual aggregation module is built to recompose features from the dual streams. Representations with different layers and pathways are interacted and aggregated in a progressive and residual manner, which keeps a good trade-off between representation deficiency and redundancy. Extensive experiments on UGC-VQA databases verify that our method achieves the state-of-the-art performance and keeps a good capability of generalization. The source code will be available inhttps://github.com/Sissuire/DSD-PRO.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.1
2022 Spatiotemporal Representation Learning for Blind Video Quality Assessment
abstract
Blind video quality assessment (BVQA) is of great importance for video-related applications, yet still challenging even in this deep learning era. The difficulty lies in the shortage of large-scale labeled data, thus making it hard to train a robust spatiotemporal encoder for BVQA. To relieve such difficulty, we first build a video dataset, which contains over 320K samples suffering from various compression and transmission artifacts. While manually annotating the dataset with subjective perception is much labor-intensive and time-consuming, we adopt reference-based VQA algorithms to weakly label the data automatically. We consider that single weak label is derived from single knowledge, which is deficient and incomplete for VQA. To alleviate the bias from single weak label (i.e., single knowledge) in the weakly labeled dataset, we propose HEterogeneous Knowledge Ensemble (HEKE) for spatiotemporal representation learning. Compared to learning from single knowledge, learning with HEKE is thought to achieve a lower infimum theoretically, and obtain richer representation. On the basis of the built dataset and the HEKE methodology, a feature encoder specific to BVQA is formed, and directly extract spatiotemporal representation from videos. Then, the video quality can be either acquired in a completely BVQA manner without ground truth, or via a finetuning-based regressor with labels. Extensive experiments on various VQA databases show that our BVQA model with the pretrained encoder achieves the state-of-the-art performance. More surprisingly, even trained on the synthetic data, our model still shows competitive performance on authentic databases. The data and source code will be available athttps://github.com/Sissuire/BVQA-HEKE.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.1
2022 Video Quality Assessment With Serial Dependence Modeling
abstract
Video quality assessment (VQA) is much more challenging than image quality assessment, due to the difficulty of modeling temporal influence among frames. Most of the existing VQA methods usually isolate each moment within the video (i.e., it neglects the sequential nature), leading to a large gap from the subjective perception. Recent research on neuroscience suggests a serially dependent perception (SDP) mechanism in the human visual system (HVS). Namely, the HVS tends to incorporate the recent past visual experience to predict the present perception. Inspired by the SDP, we suggest that the HVS prefers stable and continuous degradations in videos due to their predictability, and exhibits less tolerance to interrupted and unpredictable disturbances. Thus, we introduce a novel serial dependence modeling (SDM) framework for full-reference VQA in this paper. Firstly, the instantaneous degradation is measured on both the static appearance and motion information for each glimpse of scenes. Since motion plays an important role in videos, two types of structures are extracted for motion representation, namely, an explicit content-based 3D structure and an implicit feature-based 2D structure. Next, an assessment-directed long-short term memory (A-LSTM) is proposed to capture the serial dependence among instantaneous degradations. With the consideration of the perceptual effect from the previous moment on the current one, especially the effect from the perceptually worst moment, the serially dependent degradation is characterized. Finally, by mimicking the subjective rating for video-viewing, an attention-based quality decision procedure is presented to acquire the final video quality. Experimental results on publicly available VQA databases demonstrate that the proposed method maintains good consistency with the subjective perception.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Multim.1
2021 No-Reference Video Quality Assessment with Heterogeneous Knowledge Ensemble
abstract
Blind assessment of video quality is still challenging even in this deep learning era. The limited number of samples in existing databases is insufficient to learn a good feature extractor for video quality assessment (VQA), while manually labeling a larger database with subjective perception is very labor-intensive and time-consuming. To relieve such difficulty, we first collect 3589 high-quality video clips as the reference and build a large VQA dataset. The dataset contains more than 300K samples degraded by various distortion types due to compression and transmission error, and provides weak labels for each distorted sample with several full-reference VQA algorithms. To learn effective representation from the weakly labeled data, we alleviate the bias of single weak label (i.e., single knowledge) via learning from multiple heterogeneous knowledge. To this end, we propose a novel no-reference VQA (NR-VQA) method with HEterogeneous Knowledge Ensemble (HEKE). Comparing to learning from single knowledge, HEKE can theoretically reach a lower infimum, and learn richer representation due to the heterogeneity. Extensive experimental results show that the proposed HEKE outperforms existing NR-VQA methods, and achieves the state-of-the-art performance. The source code will be available at https://github.com/Sissuire/BVQA-HEKE.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia2
2019 Quality Assessment for Video With Degradation Along Salient Trajectories
abstract
With the rapid growth of digital video through the Internet, a reliable objective video-quality assessment (VQA) algorithm is in great demand for video management. Motion information plays a dominant role for video perception, and the human visual system (HVS) is able to track moving objects effectively with eye movement. Moreover, the middle temporal area of the brain is selective for moving objects with particular velocities. In other words, visual contents that are along the motion trajectories will automatically attract our attention for dedicated processing. Inspired by the motion-related process in the HVS, we suggest analyzing the degradation along attended motion trajectories for VQA. The characteristic of motion velocity along each trajectory is analyzed for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial-quality degradation from each frame, a novel full-reference assessor along salient trajectories (FAST) for VQA (which combines the spatial, temporal, and joint spatial-temporal quality degradations) is introduced. Experimental results on five publicly available VQA databases demonstrate that the proposed FAST VQA model performs consistently with the subjective perception. The source code of the proposed method is available at http://web.xidian.edu.cn/wjj/paper.html.
Jinjian Wu, Yongxu Liu 0001, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Multim.2
2018 Motion Trajectory based Spatial-Temporal Degradation Measurement for Video Quality Assessment
abstract
With the rapid growth of digital video through the Internet, a reliable video quality assessment (VQA) technology is greatly demanded for video management. Motion information plays a dominant role for video perception, however it is too difficult to be accurately analyzed for VQA. The human visual system (HVS) is highly adaptive to track moving objects with pursuit eye movement. Inspired by the motion process in the HVS, we suggest to analyze the degradation along attended motion trajectories for VQA. As a convenient representation of motion, optical flow is calculated for motion trajectory searching. Next, the degradation on the motion velocity along each trajectory is analyzed with the optical flow for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial quality degradation from each frame, a novel VQA model is introduced. Experimental results on the public available VQA databases demonstrate that the proposed VQA model performs highly consistency with the subjective perception.
Jinjian Wu, Yongxu Liu 0001, Guangming Shi
VCIP2
2017 Saliency change based reduced reference image quality assessment
abstract
The image quality assessment (IQA) technique, which aims to perform coherently with subjective perception, is useful in quality-orientated image processing systems. In this paper, we suggest to take the saliency change into account for reduced reference (RR) IQA model. Generally, a saliency region will attract more attention, and our human vision is more sensitive to quality degradation on such region. Inspired by this, saliency values are firstly used to highlight these sensitive regions, and a local saliency weighted histogram (LSWH) based on visual orientation pattern is generated for visual feature extraction. Next, strong distortion may change the saliency from the reference to the distorted images. Thus, the saliency of each visual orientation pattern is measured, and a global saliency based histogram (GSBH) is created. Finally, by combining the LSWH and GSBH, a novel IQA model for reduced reference is introduced. Experimental results on five publicly available databases demonstrate that the proposed model uses only several values (9 values) as reference information, and performs consistently with subjective perception.
Jinjian Wu, Yongxu Liu 0001, Guangming Shi, Weisi Lin
VCIP2