EDBT 2026 Demo / reviewers in the wild / expert
Yutao Liu 0002
dblp:64/9557-2
· DBLP profile ↗
29ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-3066-1884ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UQ-Bench: A Benchmark for Evaluating Multimodal LLMs on Underwater Image Quality AssessmentabstractDespite the rapid progress of multimodal large language models (MLLMs), their capacity for low-level visual perception in underwater environments remains underexplored. To address this gap, we present UQ-Bench, the first systematically designed benchmark for evaluating the ability of MLLMs to perceive and assess underwater image quality at the low-level visual attribute level. UQ-Bench comprises three components: (1) UW-Perception, a dataset of 3,000 underwater images paired with targeted questions on key degradations such as color cast, blur, contrast, and exposure, covering both global and local perceptual dimensions; (2) UW-Describe, a dataset of 500 images with expert-annotated gold-standard descriptions for assessing the accuracy of model-generated text; and (3) UW-Eval, an evaluation protocol employing human mean opinion scores (MOS) for quantitative quality assessment. To ensure rigorous and reproducible benchmarking, we propose a GPT-assisted evaluation framework that aligns model outputs with expert references and enables fine-grained analysis of distortion perception. Experimental results demonstrate that while MLLMs exhibit preliminary competence in underwater low-level visual tasks, they still fall short in capturing subtle degradations and achieving human-level consistency, highlighting the need for further advances in foundation models for marine vision. Jingchao Cao, Guo An, Feng Gao 0005, Ke Gu 0001, Yutao Liu 0002 |
AAAI | 5 |
| 2026 | FVNet: Harnessing Liquid Neural Dynamics for Lightweight Visual RepresentationabstractEfficient visual backbone design remains crucial for resource-constrained computer vision applications. Inspired by the adaptive continuous-time dynamics observed in biological neurons, we propose FVNet, a novel lightweight architecture that integrates liquid neural dynamics for efficient and dynamic visual feature extraction. Central to FVNet is the Fluid Temporal Flow Unit (FTFU), which employs continuous-time equations with learnable time constants to capture spatio-temporal dependencies adaptively. By further stacking these units in a Multi-Phase Fluid Block (MPFB), our model processes features across parallel temporal scales, enabling context-aware feature encoding without incurring excessive computational overhead. Through a discrete closed-form solution, FVNet achieves the representational power of continuous-time models while avoiding the instability and overhead of iterative numerical solvers. Extensive experiments on various vision tasks demonstrate that FVNet achieves superior performance and efficiency over existing state-of-the-art lightweight networks. Zhenzhe Hou, Xiaohui Chu, Runze Hu, Yutao Liu 0002 |
AAAI | 5 |
| 2025 | ERD: Encoder-Residual-Decoder Neural Network for Underwater Image EnhancementabstractIn underwater environments, the absorption and scattering of light often result in various types of degradation in captured images, including color cast, low contrast, low brightness, and blurriness. These undesirable effects pose significant challenges for both underwater photography and downstream tasks such as object detection, recognition, and navigation. To address these challenges, we propose a novel end-to-end underwater image enhancement (UIE) network via the multistage and mixed attention mechanism and a residual-based feature refinement module, called ERD. Specifically, our network includes an encoder stage for extracting features from input underwater images with channel, spatial, and patch attention modules to emphasize degraded channels and regions for restoration; a residual stage for further purification of informative features through sufficient feature learning; and a decoder stage for effective image reconstruction. Inspired by visual perception mechanism, we design the frequency domain loss and edge details loss to retain more high-frequency information and object details while ensuring that the enhanced image approximates the reference image in terms of color tone while preserving content and structure. To comprehensively evaluate our proposed UIE model, we also curated three additional underwater image datasets through online collection and generation using Cycle-GAN. Rigorous experiments conducted on a total of eight underwater image datasets demonstrate that the proposed ERD model outperforms state-of-the-art methods in enhancing both real-world and generated underwater images. Our code and datasets are available athttps://github.com/fansuregrin/ERD. Jingchao Cao, Wangzhen Peng, Yutao Liu 0002, Junyu Dong, Patrick Le Callet, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Multi-Scale Local and Global Feature Fusion for Blind Quality Assessment of Enhanced ImagesabstractImage enhancement plays a crucial role in computer vision by improving visual quality while minimizing distortion. Traditional methods enhance images through pixel value transformations, yet they often introduce new distortions. Recent advancements in deep learning-based techniques promise better results but challenge the preservation of image fidelity. Therefore, it is essential to evaluate the visual quality of enhanced images. However, existing quality assessment methods frequently encounter difficulties due to the unique distortions introduced by these enhancements, thereby restricting their effectiveness. To address these challenges, this paper proposes a novel blind image quality assessment (BIQA) method for enhanced natural images, termed multi-scale local feature fusion and global feature representation-based quality assessment (MLGQA). This model integrates three key components: a multi-scale Feature Attention Mechanism (FAM) for local feature extraction, a Local Feature Fusion (LFF) module for cross-scale feature synthesis, and a Global Feature Representation (GFR) module using Vision Transformers to capture global perceptual attributes. This synergistic framework effectively captures both fine-grained local distortions and broader global features that collectively define the visual quality of enhanced images. Furthermore, in the absence of a dedicated benchmark for enhanced natural images, we design the Natural Image Enhancement Database (NIED), a large-scale dataset consisting of 8,581 original images and 102,972 enhanced natural images generated through a wide array of traditional and deep learning-based enhancement techniques. Extensive experiments on NIED demonstrate that the proposed MLGQA model significantly outperforms current state-of-the-art BIQA methods in terms of both prediction accuracy and robustness. Jingchao Cao, Yutao Liu 0002, Feng Gao 0005, Ke Gu 0001, Guangtao Zhai, Junyu Dong, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Attention and Mamba-Driven Quality Assessment for Underwater ImagesabstractUnderwater imaging is essential in a variety of fields, including resource exploration, marine observation, and scientific research. However, the quality of underwater images is often compromised by environmental factors such as light scattering, absorption, and the presence of fog, leading to distortions such as color shifts, low contrast, and blurriness. To address these challenges, we propose a novel underwater image quality assessment (UIQA) method, the Attention and Mamba-driven Quality Index (AMQI). The AMQI model employs a multi-stage architecture designed to capture both local and global image features critical for underwater quality evaluation. First, a Shallow Feature Extractor (SFE) captures essential spatial details. Next, the Local Information Representation Network (LIR-Net), equipped with Channel Attention (CA) and Large Kernel-guided Spatial (LKS) mechanisms, enhances fine details and captures long-range dependencies to address underwater-specific distortions. The Global Information Representation Network (GIR-Net) further processes the features using a combination of the Visual State-Space Model (VSSM) and ResNet-50 to capture high-level semantic and contextual information. Finally, the Feature-Quality Mapping Network (FQM) converts the learned features into a quality score, ensuring precise predictions of image quality. Extensive experiments on the Underwater Image Quality Database (UIQD) demonstrate that AMQI outperforms current state-of-the-art IQA and UIQA models in terms of accuracy and correlation with human subjective evaluations. The model's robustness and generalization capabilities are further validated through detailed ablation studies and cross-database evaluations, showcasing its strong performance across diverse underwater environments. The source code is available athttps://github.com/ibaochao/AMQI. Jingchao Cao, Baochao Zhang, Yutao Liu 0002, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 3 |
| 2025 | Deep No-Reference Quality Assessment for Underwater Enhanced ImagesabstractThe goal of underwater image enhancement (UIE) is to boost the acquired underwater image quality, which increases the value of the underwater image significantly. However, without effective underwater enhanced image quality assessment (UEIQA) measures that benchmark the UIE, the process of UIE becomes driftless and the enhanced results of different UIE algorithms cannot be fairly compared. Toward this end, we in this work construct a dedicated UEIQA scheme on the basis of deep investigation of the underwater enhanced image characteristics. Specifically, in our proposed method, we respectively design deep neural networks to represent the unique attributes of the underwater enhanced image, such as color cast, local distortions, naturalness degree, sharpness, contrast, fog density, etc., that are highly correlated with the image quality. Then we introduce the Vision Transformer (ViT) to capture the dependencies among different image attributes and infer the image quality level. Extensive experiments conducted on three typical UEIQA databases, i.e., SOTA, UID2021 and SAUD, show that the proposed UEIQA model yields noteworthy higher prediction accuracy than the representative IQA and UEIQA metrics, e.g., achieving SRCC values of 0.891 ( vs. 0.749 in SAUD) and 0.933 ( vs. 0.798 in UID2021). The proposed UEIQA model will be released athttps://github.com/YT2015?tab=repositories. Yutao Liu 0002, Baochao Zhang, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 1 |
| 2025 | Mixed Attention and Channel Shift Transformer for Efficient Action RecognitionabstractThe practical use of the Transformer-based methods for processing videos is constrained by the high computing complexity. Although previous approaches adopt the spatiotemporal decomposition of 3D attention to mitigate the issue, they suffer from the drawback of neglecting the majority of visual tokens. This article presents a novel mixed attention operation that subtly fuses the random, spatial, and temporal attention mechanisms. The proposed random attention stochastically samples video tokens in a simple yet effective way, complementing other attention methods. Furthermore, since the attention operation concentrates on learning long-distance relationships, we employ the channel shift operation to encode short-term temporal characteristics. Our model can provide more comprehensive motion representations thanks to the amalgamation of these techniques. Experimental results show that the proposed method produces competitive action recognition results with low computational overhead on both large-scale and small-scale public video datasets. Xiusheng Lu, Yanbin Hao, Lechao Cheng, Sicheng Zhao, Yutao Liu 0002, Mingli Song |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Semi-Supervised Blind Image Quality Assessment through Knowledge Distillation and Incremental LearningabstractBlind Image Quality Assessment (BIQA) aims to simulate human assessment of image quality. It has a great demand for labeled data, which is often insufficient in practice. Some researchers employ unsupervised methods to address this issue, which is challenging to emulate the human subjective system. To this end, we introduce a unified framework that combines semi-supervised and incremental learning to address the mentioned issue. Specifically, when training data is limited, semi-supervised learning is necessary to infer extensive unlabeled data. To facilitate semi-supervised learning, we use knowledge distillation to assign pseudo-labels to unlabeled data, preserving analytical capability. To gradually improve the quality of pseudo labels, we introduce incremental learning. However, incremental learning can lead to catastrophic forgetting. We employ Experience Replay by selecting representative samples during multiple rounds of semi-supervised learning, to alleviate forgetting and ensure model stability. Experimental results show that the proposed approach achieves state-of-the-art performance across various benchmark datasets. After being trained on the LIVE dataset, our method can be directly transferred to the CSIQ dataset. Compared with other methods, it significantly outperforms unsupervised methods on the CSIQ dataset with a marginal performance drop (-0.002) on the LIVE dataset. In conclusion, our proposed method demonstrates its potential to tackle the challenges in real-world production processes. Wensheng Pan, Timin Gao, Yan Zhang 0109, Xiawu Zheng, Yunhang Shen, Ke Li 0015, Runze Hu, Yutao Liu 0002, Pingyang Dai |
AAAI | 8 |
| 2024 | Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise SensitivityabstractThe current state-of-the-art No-Reference Image Quality Assessment (NR-IQA) methods typically rely on feature extraction from upstream semantic backbone networks, assuming that all extracted features are relevant. However, we make a key observation that not all features are beneficial, and some may even be harmful, necessitating careful selection. Empirically, we find that many image pairs with small feature spatial distances can have vastly different quality scores, indicating that the extracted features may contain quality-irrelevant noise. To address this issue, we propose a Quality-Aware Feature Matching IQA Metric (QFM-IQM) that employs an adversarial perspective to remove harmful semantic noise features from the upstream task. Specifically, QFM-IQM enhances the semantic noise distinguish capabilities by matching image pairs with similar quality scores but varying semantic features as adversarial semantic noise and adaptively adjusting the upstream task’s features by reducing sensitivity to adversarial noise perturbation. Furthermore, we utilize a distillation framework to expand the dataset and improve the model’s generalization ability. Extensive experiments conducted on eight standard IQA datasets have demonstrated the effectiveness of our proposed QFM-IQM. Timin Gao, Runze Hu, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Jingyuan Zheng, Yunhang Shen, Ke Li 0015, Yutao Liu 0002, Pingyang Dai, Rongrong Ji |
ICML | 10 |
| 2024 | Integrating Global Context Contrast and Local Sensitivity for Blind Image Quality AssessmentabstractBlind Image Quality Assessment (BIQA) mirrors subjective made by human observers. Generally, humans favor comparing relative qualities over predicting absolute qualities directly. However, current BIQA models focus on mining the "local" context, i.e., the relationship between information among individual images and the absolute quality of the image, ignoring the "global" context of the relative quality contrast among different images in the training data. In this paper, we present the Perceptual Context and Sensitivity BIQA (CSIQA), a novel contrastive learning paradigm that seamlessly integrates "global” and "local” perspectives into the BIQA. Specifically, the CSIQA comprises two primary components: 1) A Quality Context Contrastive Learning module, which is equipped with different contrastive learning strategies to effectively capture potential quality correlations in the global context of the dataset. 2) A Quality-aware Mask Attention Module, which employs the random mask to ensure the consistency with visual local sensitivity, thereby improving the model’s perception of local distortions. Extensive experiments on eight standard BIQA datasets demonstrate the superior performance to the state-of-the-art BIQA methods. Runze Hu, Jingyuan Zheng, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Ke Li 0015, Yunhang Shen, Yutao Liu 0002, Pingyang Dai, Rongrong Ji |
ICML | 9 |
| 2024 | UIQI: A Comprehensive Quality Evaluation Index for Underwater ImagesabstractDue to the light absorption and scattering in waterbodies, acquired underwater images frequently suffer from color cast, blur, low contrast, noise, etc., which seriously degrade the image quality and affect their subsequent applications. Therefore, it is necessary to propose a reliable and practical underwater image quality assessment (IQA) model that can faithfully evaluate underwater image quality. To this end, in this article, we establish a novel quality assessment model for underwater images by in-depth analysis and characterization of multiple image properties. Specifically, we propose characterizing the image luminance, color cast, sharpness, contrast, fog density and noise to comprehensively describe the image quality to evaluate the underwater image quality more accurately. Dedicated features are elaborately investigated to characterize those quality-aware image properties. After feature extraction, we employ support vector regression (SVR) to integrate all the quality-aware features and regress them onto the underwater image quality score. Extensive tests performed on standard underwater image quality databases demonstrate the superior prediction performance of the proposed underwater IQA model to state-of-the-art congeneric quality assessment models. Yutao Liu 0002, Ke Gu 0001, Jingchao Cao, Shiqi Wang 0001, Guangtao Zhai, Junyu Dong, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2024 | Underwater Image Quality Assessment: Benchmark Database and Objective MethodabstractUnderwater image quality assessment (UIQA) plays a crucial role in monitoring and detecting the quality of acquired underwater images in underwater imaging systems. Currently, the investigation of UIQA encounters two major challenges. First, a lack of large-scale UIQA databases for benchmarking UIQA algorithms remains, which greatly restricts the development of UIQA research. The other limitation is that there is a shortage of effective UIQA methods that can faithfully predict underwater image quality. To alleviate these two challenges, in this paper, we first construct a large-scale UIQA database (UIQD). Specifically, UIQD contains a total of 5369 authentic underwater images that span abundant underwater scenes and typical quality degradation conditions. Extensive subjective experiments are executed to annotate the perceived quality of the underwater images in UIQD. Based on an in-depth analysis of underwater image characteristics, we further establish a novel baseline UIQA metric that integrates channel and spatial attention mechanisms and a transformer. Channel- and spatial attention modules are used to capture the image channel and local quality degradations, while the transformer module characterizes the image quality from a global perspective. Multilayer perception is employed to fuse the local and global feature representations and yield the image quality score. Extensive experiments conducted on UIQD demonstrate that the proposed UIQA model achieves superior prediction performance compared with the state-of-the-art UIQA and IQA methods. The proposed UIQD and UIQA models will be released athttps://github.com/YT2015?tab=repositories. Yutao Liu 0002, Baochao Zhang, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong |
IEEE Trans. Multim. | 1 |
| 2023 | Data-Efficient Image Quality Assessment with Attention-Panel DecoderabstractBlind Image Quality Assessment (BIQA) is a fundamental task in computer vision, which however remains unresolved due to the complex distortion conditions and diversified image contents. To confront this challenge, we in this paper propose a novel BIQA pipeline based on the Transformer architecture, which achieves an efficient quality-aware feature representation with much fewer data. More specifically, we consider the traditional fine-tuning in BIQA as an interpretation of the pre-trained model. In this way, we further introduce a Transformer decoder to refine the perceptual information of the CLS token from different perspectives. This enables our model to establish the quality-aware feature manifold efficiently while attaining a strong generalization capability. Meanwhile, inspired by the subjective evaluation behaviors of human, we introduce a novel attention panel mechanism, which improves the model performance and reduces the prediction uncertainty simultaneously. The proposed BIQA method maintains a light-weight design with only one layer of the decoder, yet extensive experiments on eight standard BIQA datasets (both synthetic and authentic) demonstrate its superior performance to the state-of-the-art BIQA methods, i.e., achieving the SRCC values of 0.875 (vs. 0.859 in LIVEC) and 0.980 (vs. 0.969 in LIVE). Checkpoints, logs and code will be available at https://github.com/narthchin/DEIQT. Guanyi Qin, Runze Hu, Yutao Liu 0002, Xiawu Zheng, Xiu Li 0001, Yan Zhang 0109 |
AAAI | 3 |
| 2023 | Toward a No-Reference Quality Metric for Camera-Captured ImagesabstractExisting no-reference (NR) image quality assessment (IQA) metrics are still not convincing for evaluating the quality of the camera-captured images. Toward tackling this issue, we, in this article, establish a novel NR quality metric for quantifying the quality of the camera-captured images reliably. Since the image quality is hierarchically perceived from the low-level preliminary visual perception to the high-level semantic comprehension in the human brain, in our proposed metric, we characterize the image quality by exploiting both the low-level image properties and the high-level semantics of the image. Specifically, we extract a series of low-level features to characterize the fundamental image properties, including the brightness, saturation, contrast, noiseness, sharpness, and naturalness, which are highly indicative of the camera-captured image quality. Correspondingly, the high-level features are designed to characterize the semantics of the image. The low-level and high-level perceptual features play complementary roles in measuring the image quality. To infer the image quality, we employ the support vector regression (SVR) to map all the informative features to a single quality score. Thorough tests conducted on two standard camera-captured image databases demonstrate the effectiveness of the proposed quality metric in assessing the image quality and its superiority over the state-of-the-art NR quality metrics. The source code of the proposed metric for camera-captured images is released at https://github.com/YT2015?tab=repositories. Runze Hu, Yutao Liu 0002, Ke Gu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Cybern. | 2 |
| 2021 | Precise No-Reference Image Quality Evaluation Based on Distortion IdentificationabstractThe difficulty of no-reference image quality assessment (NR IQA) often lies in the lack of knowledge about the distortion in the image, which makes quality assessment blind and thus inefficient. To tackle such issue, in this article, we propose a novel scheme for precise NR IQA, which includes two successive steps, i.e., distortion identification and targeted quality evaluation. In the first step, we employ the well-known Inception-ResNet-v2 neural network to train a classifier that classifies the possible distortion in the image into the four most common distortion types, i.e., Gaussian white noise (WN), Gaussian blur (GB), jpeg compression (JPEG), and jpeg2000 compression (JP2K). Specifically, the deep neural network is trained on the large-scale Waterloo Exploration database, which ensures the robustness and high performance of distortion classification. In the second step, after determining the distortion type of the image, we then design a specific approach to quantify the image distortion level, which can estimate the image quality specially and more precisely. Extensive experiments performed on LIVE, TID2013, CSIQ, and Waterloo Exploration databases demonstrate that (1) the accuracy of our distortion classification is higher than that of the state-of-the-art distortion classification methods, and (2) the proposed NR IQA method outperforms the state-of-the-art NR IQA methods in quantifying the image quality. Chenggang Yan 0001, Tong Teng, Yutao Liu 0002, Yongbing Zhang 0002, Haoqian Wang, Xiangyang Ji |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Depth Image Denoising Using Nuclear Norm and Learning Graph ModelabstractDepth image denoising is increasingly becoming the hot research topic nowadays, because it reflects the three-dimensional scene and can be applied in various fields of computer vision. But the depth images obtained from depth camera usually contain stains such as noise, which greatly impairs the performance of depth-related applications. In this article, considering that group-based image restoration methods are more effective in gathering the similarity among patches, a group-based nuclear norm and learning graph (GNNLG) model was proposed. For each patch, we find and group the most similar patches within a searching window. The intrinsic low-rank property of the grouped patches is exploited in our model. In addition, we studied the manifold learning method and devised an effective optimized learning strategy to obtain the graph Laplacian matrix, which reflects the topological structure of image, to further impose the smoothing priors to the denoised depth image. To achieve fast speed and high convergence, the alternating direction method of multipliers is proposed to solve our GNNLG. The experimental results show that the proposed method is superior to other current state-of-the-art denoising methods in both subjective and objective criterion. Chenggang Yan 0001, Zhisheng Li, Yongbing Zhang 0002, Yutao Liu 0002, Xiangyang Ji, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Unsupervised Blind Image Quality Evaluation via Statistical Measurements of Structure, Naturalness, and PerceptionabstractMost existing blind image quality assessment (BIQA) methods belong to supervised methods, which always need a large number of image samples and expensive subjective scores for training a quality prediction model. In this paper, we focus our attention on the unsupervised BIQA methods and put forward a novel unsupervised approach. The main idea of our method is to quantify the image quality degradation through measuring the structure, naturalness, and the perception quality variations of the distorted image from the pristine natural images. In specific, the structure variation is captured by the deviations of the image phase congruency and gradients distributions. The naturalness variation is characterized through the distributions variations of the locally mean subtracted and contrast normalized (MSCN) coefficients and the products of pairs of the adjacent MSCN coefficients. Compared with existing unsupervised methods, we initiatively introduce the perception quality measurement into the construction of unsupervised BIQA method, which is conducted by characterizing the prediction discrepancy between the image and its brain prediction based on the free-energy principle in the newly revealed brain theory. After feature extraction, we learn a pristine multivariate Gaussian (MVG) model with the extracted features from a set of pristine natural images. The quality of a new image is finally defined as the distance between its MVG model and the learned pristine MVG model. The extensive experiments conducted on LIVE, TID2013, CSIQ, Toyama, CID2013, and the Waterloo Exploration databases demonstrate that the proposed method achieves comparative prediction performance with the state-of-the-art BIQA methods. Yutao Liu 0002, Ke Gu 0001, Yongbing Zhang 0002, Xiu Li 0001, Guangtao Zhai, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Blind Image Quality Assessment by Natural Scene Statistics and Perceptual CharacteristicsabstractOpinion-unaware blind image quality assessment (OU BIQA) refers to establishing a blind quality prediction model without using the expensive subjective quality scores, which is a highly promising direction in the BIQA research. In this article, we focus on OU BIQA and propose a novel OU BIQA method. Specifically, in our proposed method, we deeply investigate the natural scene statistics (NSS) and the perceptual characteristics of the human brain for visual perception. Accordingly, a set of quality-aware NSS and perceptual characteristics-related features are designed to characterize the image quality effectively. For inferring the image quality, we learn a pristine multivariate Gaussian (MVG) model on a collection of pristine images, which serves as the reference information for quality evaluation. At last, the quality of a new given image is defined by measuring the divergence between its MVG model and the learned pristine MVG model. Thorough experiments performed on seven popular image databases demonstrate that the proposed OU BIQA method delivers superior performance to the state-of-the-art OU BIQA methods. The Matlab source code of the proposed method will be made publicly available at https://github.com/YT2015?tab=;repositories. Yutao Liu 0002, Ke Gu 0001, Xiu Li 0001, Yongbing Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Blind Quality Assessment of Camera Images Based on Low-Level and High-Level Statistical FeaturesabstractCamera images in reality are easily affected by various distortions, such as blur, noise, blockiness, and the like, which damage the quality of images. The complexity of distortions in camera images raises significant challenge for precisely predicting their perceptual quality. In this paper, we present an image quality assessment (IQA) approach that aims to solve this challenging problem to some extent. In the proposed method, we first extract the low-level and high-level statistical features, which can capture the quality degradations effectively. On the one hand, the first kind of statistical features are extracted from the locally mean subtracted and contrast normalized coefficients, which denote the low-level features in the early human vision. On the other hand, the recently proposed brain theory and neuroscience, especially the free-energy principle, reveal that the human brain tries to explain its encountered visual scenes through an inner creative model, with which the brain can produce the projection for the image. Then, the quality of perceptions can be reflected by the divergence between the image and its brain projection. Based on this, we extract the second type of features from the brain perception mechanism, which represent the high-level features. The low-level and high-level statistical features can play a complementary role in quality prediction. After feature extraction, we design a neural network to integrate all the features and convert them to the final quality score. Extensive tests performed on two real camera image datasets prove the validity of our method and its advantageous predicting ability over the competitive IQA models. Yutao Liu 0002, Ke Gu 0001, Shiqi Wang 0001, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Reduced-Reference Image Quality Assessment Based on Free-Energy Principle with Multi-Channel DecompositionabstractThe free-energy principle studied in brain theory and neuroscience accounts for the mechanism of perception and understanding in human brain, which is highly adapted for measuring the visual quality of perceptions. On the other hand, psychologists and neurologists report that different frequency and orientation components of one stimulus arouse different neurons in striate cortex. In this paper, a novel reduce-reference (RR) image quality assessment (IQA) metric based on free-energy principle in multi-channel is proposed, which is called MCFEM (Multi-Channel Free-Energy principle Metric). We first decompose the input reference image and distorted image via a two-level discrete Haar wavelet transform (DHWT). Next, the free-energy features of each subband images are computed based on sparse representation. Finally, an overall quality index is received through the support vector regressor (SVR). Extensive experimental comparisons on four (LIVE, CSIQ, TID2008 and TID2013) benchmark image databases demonstrate that the proposed method is highly competitive with the representative RR and no-reference models as well as full-reference ones. Wenhan Zhu, Guangtao Zhai, Yutao Liu 0002, Ning Lin, Xiaokang Yang 0001 |
MMSP | 3 |
| 2018 | Reduced-Reference Image Quality Assessment in Free-Energy Principle and Sparse RepresentationabstractThe free-energy principle in recent studies of brain theory and neuroscience models the perception and understanding of the outside scene as an active inference process, in which the brain tries to account for the visual scene with an internal generative model. Specifically, with the internal generative model, the brain yields corresponding predictions for its encountered visual scenes. Then, the discrepancy between the visual input and its brain prediction should be closely related to the quality of perceptions. On the other hand, sparse representation has been evidenced to resemble the strategy of the primary visual cortex in the brain for representing natural images. With the strong neurobiological support for sparse representation, in this paper, we approximate the internal generative model with sparse representation and propose an image quality metric accordingly, which is named FSI (free-energy principle and sparse representation-based index for image quality assessment). In FSI, the reference and distorted images are, respectively, predicted by the sparse representation at first. Then, the difference between the entropies of the prediction discrepancies is defined to measure the image quality. Experimental results on four large-scale image databases confirm the effectiveness of the FSI and its superiority over representative image quality assessment methods. The FSI belongs to reduced-reference methods, and it only needs a single number from the reference image for quality estimation. Yutao Liu 0002, Guangtao Zhai, Ke Gu 0001, Xianming Liu 0005, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Reduced Reference Image Quality Assessment Based on Entropy of Classified PrimitivesabstractThe human visual perception is a layered progressive process that brain assimilates visual information gradually, from primary information, structural information to detailed information. Recently, the visual primitives (atoms in the dictionary) extracted by sparse representation have been shown to be highly related to the layered progressive process of human visual perception. In this paper, the visual primitives are first classified into three categories: DCprimary, sketch and texture in terms of their inherent properties regarding tothe perceptual information. Then, we propose a novel reduced reference (RR) image quality assessment (IQA) metric using perceptual information represented by entropy of classified primitives (EoCP). Specifically, EoCP is a measurement of the distribution statistics of the visual primitives, which can represent the visual information. The differences of EoCPs between the reference image and its distorted version are calculated as features to characterize perceptual loss. The extracted features (only three scalars) are used to compute the quality score by a prediction function which is trained using support vector regression(SVR). Experimental results on LIVE, CSIQ and TID2013 image databases demonstrate that the proposed metric achieves high consistency with the human perception and show competitive performance with state-of-the-art IQA metrics. Zhaolin Wan, Yutao Liu 0002, Debin Zhao |
DCC | 2 |
| 2017 | Dynamic backlight scaling considering ambient luminance for mobile energy savingabstractThe mobile video playback involves many subsystems of the devices such as computing, rendering and displaying subsystems. Among all subsystems, the displaying subsystem accounts for at least 38% of all consumed power, and it can be up to 68% with the maximum backlight brightness. What is more, lots of people watch videos via mobile devices in various situations, where the ambient luminance condition is different. Therefore, how to save mobile energy and improve the Quality of Experience (QoE) in different situations become significant problems. In this paper, we try to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users' QoE. In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance and backlight level. We first study whether the continuous playback of backlight-scaled shots using the proposed scaling magnitude would cause flicker effect or not. Then motivated by the findings of these subject studies, we implement a Dynamic Backlight Scaling (DBS) strategy. The experiment results demonstrate that the DBS strategy can save more than 40% power at most and can also save 10% power even at a very high ambient luminance. Wei Sun 0029, Guangtao Zhai, Xiongkuo Min, Yutao Liu 0002, Siwei Ma 0001, Jing Liu 0002, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 4 |
| 2017 | Reduced reference stereoscopic image quality assessment based on entropy of classified primitivesabstractStereoscopic vision is a complex system which receives and integrates perceptual information from both monocular and binocular cues. In this paper, a novel reduced-reference stereoscopic image quality assessment scheme is proposed, based on the visual perceptual information measured by entropy of classified primitives (EoCP) and mutual information of classified primitives (MIoCP), named as DCprimary, sketch and texture primitives respectively, which is in accordance with the hierarchical progressive process of human visual perception. Specifically, EoCP of each-view image are calculated as monocular cue, and MIoCP between two-view images is derived as binocular cue. The Maximum (MAX) mechanism is applied to determine the perceptual information. The perceptual information differences between the original and distorted images are used to predict the stereoscopic image quality by support vector regression (SVR). Experimental results on LIVE phase II asymmetric database validate the proposed metric achieves significantly higher consistency with subjective ratings and outperforms state-of-the-art stereoscopic image quality assessment methods. Zhaolin Wan, Yutao Liu 0002, Debin Zhao |
ICME | 3 |
| 2017 | Quality assessment for real out-of-focus blurred images
Yutao Liu 0002, Ke Gu 0001, Guangtao Zhai, Xianming Liu 0005, Debin Zhao, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Perceptual image quality assessment combining free-energy principle and sparse representationabstractSince the purpose of objective image quality assessment is to be consistent with subjective image quality assessment as highly as possible, the understanding of the mechanisms of human visual system will certainly benefit the study of objective image quality assessment. Recent developments in brain theory and neuroscience, particularly the free-energy principle, account for the perception and understanding of visual scenes. As the free-energy principle conjectures, the brain tries to generate the corresponding prediction for its encountered scene by an internal generative model. On the other hand, sparse representation is evidenced to resemble the neural response properties of simple cells in the primary visual cortex. Conjunctively, in this paper, we suppose the prediction manner of the internal generative model in free-energy principle follows sparse representation and propose an image quality metric accordingly. Experiments on LIVE, TID2008 and CSIQ image databases demonstrate the effectiveness of the proposed image quality metric. Noteworthily, our metric needs little information (only a single scalar) of the reference image and is training-free. Yutao Liu 0002, Guangtao Zhai, Xianming Liu 0005, Debin Zhao |
ISCAS | 1 |
| 2016 | A reduced-reference quality assessment scheme for blurred imagesabstractIn this paper, we propose a reduced-reference scheme for evaluating the quality of blurred images under the theory of free-energy principle. Specifically, the free-energy principle indicates that the brain tries to account for the input image with an internal generative model and the discrepancy between the image and its model-explained version, which can be measured by free energy, is related to the image's perceptual quality. Accordingly, we define a visual distance between the blurred image and its original image in free energy to evaluate the quality of the blurred image. Therefore, the proposed quality scheme belongs to reduced-reference methods, which needs some information from the original image for quality assessment. Experimental results on public databases, LIVE, TID2013 and C-SIQ, demonstrate the proposed method works in high consistency with subjective assessment results and outperforms representative image quality assessment approaches. Zongxi Han, Guangtao Zhai, Yutao Liu 0002, Ke Gu 0001, Xinfeng Zhang 0001 |
VCIP | 3 |
| 2015 | Quality assessment for out-of-focus blurred imagesabstractDuring the process of image acquisition, images are often subject to out-of-focus or defocus blur because of the improper adjustment of the camera's focal length, this image blur will degrade the image quality. However, in the literature, image quality assessment (IQA) methods dedicated to evaluating the quality of images with out-of-focus blur remain few. Therefore, in this paper, we focus our attention on the quality assessment of images that suffer from out-of-focus blur and propose an objective quality assessment method accordingly. Concretely, we construct a dedicated out-of-focus blurred image dataset, which is composed of 150 images subjected to different degrees of out-of-focus blur and the mean opinion scores (MOSs). Then, we propose a specific objective quality metric for the blurred images, which combines image sharpness assessment and saliency-guided pooling strategy. Experimental results demonstrate the proposed metric highly correlates with human judgements of image quality. Yutao Liu 0002, Guangtao Zhai, Xianming Liu 0005, Debin Zhao |
VCIP | 1 |
| 2013 | Motion vector refinement for frame rate up conversion on 3D videoabstractWith the rapid development of digital video technology, frame rate up conversion is widely used. In this paper, a novel motion vector refinement method for frame rate up conversion on depth based 3D video is proposed. Our method involves two major stages in frame rate up conversion which are motion estimation and motion vector filtering. In the motion estimation process, the depth constraint to block matching algorithm is introduced in the bi-directional motion estimation method to obtain the motion vectors. In the motion vector filtering process, a depth-guided filter is designed to enhance the consistence of motions in the same depth plane. The refined motion vectors are used for frame interpolation. Experimental results show that the proposed method achieves 0.45 dB gain in terms of PSNR on average and improves the visual quality of the frame rate up-converted video. Yutao Liu 0002, Xiaopeng Fan 0001, Xinwei Gao, Yan Liu 0014, Debin Zhao |
VCIP | 1 |