Ke Gu 0001

dblp:13/9079-1 · DBLP profile ↗
← Back
141ranked-venue papers
36as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 107 · 23 first-author · 28 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 5 since 2021Systems, architecture and hardware · 10 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 3 since 2021Computer networks · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2026 UQ-Bench: A Benchmark for Evaluating Multimodal LLMs on Underwater Image Quality Assessment
abstract
Despite the rapid progress of multimodal large language models (MLLMs), their capacity for low-level visual perception in underwater environments remains underexplored. To address this gap, we present UQ-Bench, the first systematically designed benchmark for evaluating the ability of MLLMs to perceive and assess underwater image quality at the low-level visual attribute level. UQ-Bench comprises three components: (1) UW-Perception, a dataset of 3,000 underwater images paired with targeted questions on key degradations such as color cast, blur, contrast, and exposure, covering both global and local perceptual dimensions; (2) UW-Describe, a dataset of 500 images with expert-annotated gold-standard descriptions for assessing the accuracy of model-generated text; and (3) UW-Eval, an evaluation protocol employing human mean opinion scores (MOS) for quantitative quality assessment. To ensure rigorous and reproducible benchmarking, we propose a GPT-assisted evaluation framework that aligns model outputs with expert references and enables fine-grained analysis of distortion perception. Experimental results demonstrate that while MLLMs exhibit preliminary competence in underwater low-level visual tasks, they still fall short in capturing subtle degradations and achieving human-level consistency, highlighting the need for further advances in foundation models for marine vision.
Jingchao Cao, Guo An, Feng Gao 0005, Ke Gu 0001, Yutao Liu 0002
AAAI4
2026 A Lightweight Deep and Wide Network for Image-Based Detection of Industrial Waste Gas
abstract
Due to inadequate monitoring, key pollutants (e.g., PM2.5, VOCs, etc) very possibly leak into atmosphere, thus to endanger the long-term and short-term life safety of people that work and live in the environment. Therefore, it is imperative to effectively and efficiently detect the leakage of industrial waste gas, for the purpose of timely lowering the risk of pollution and explosions. To solve such a problem, we in this paper propose a new lightweight deep and wide network (LdwNet) for detecting the leakage of industrial waste gas from an image, which brings about the two main merits: 1) Compensating for the deficiencies of sensor-based detection methods, which can accurately detect the leakage of waste gas and even measure its concentrations but require to seek leakage sources beforehand; 2) Overcoming the shortcomings of image-based detection methods, which leverage DNN-based recognition technologies and usually suffer from low efficacy, low efficiency and high energy consumption during the model training and inference. To specify, the proposed LdwNet is developed by simulating human perception, motivated by the method which detects the leakage of industrial waste gas from surveillance images with the human observation and judgement. First, based on the inspiration that the human eyes are highly sensitive to horizontal and vertical stimuli, we construct a novel lightweight parallel-series-stripe (PS2) module to validly extract features with very few parameters. Second, to fully exploit deep and shallow features for fusing the global and local information, we extend the PS2 module as a backbone along both the deep and wide directions to build the multi-channel network. Third, to achieve effective, efficient and low-carbon detection in model running, we constraint the extended PS2 modules with parameter sharing to prodigiously reduce the model parameters and thus to make the proposed model ultra-lightweight. Experiments on the datasets of carbon particulate matters and ethylene leakage prove that our LdwNet with ten thousand parameters outperforms the state-of-the-art models with millions of parameters in detection accuracy and implementation cost, and this renders our proposed LdwNet more suitable for real industrial applications.
Ke Gu 0001, Hongyan Liu 0004, Jingchao Cao, Lai-Kuan Wong, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001, Weisi Lin, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.1
2026 Perceptual Quality Assessment of Low-Light Enhanced Images: A Multi-Annotated Subjective Dataset and a Multimodal Objective Method
abstract
Low-light Image Enhancement Algorithms (LIEAs) aim to improve the visibility and visual quality of images captured in low-light environments. However, none of the existing LIEAs can comprehensively restore all visual contents, which makes it inevitable for the Enhanced Low-light Images (ELIs) to have different degrees of distortion, thereby affecting the visual quality. Currently, there is little research focusing on the quality assessment of these ELIs, partly due to the lack of publicly available datasets. Moreover, existing quality assessment methods primarily focus on a single visual modality and fail to sufficiently exploit the structural information across multiple image attributes, consequently resulting in suboptimal prediction performance. To this end, this paper conducts a systematic study on both subjective and objective quality assessment of ELIs. Firstly, we construct the first Multi-annotated and multi-modal Low-light image Enhancement quality dataset (MLE), which contains 1,000 ELIs, along with subjective studies to obtain multiple attribute annotations, quality scores, and textual descriptions. Based on this, we further propose an Attribute-guided Vision-Language Graph Reasoning Network (AVGR-Net) for ELI quality prediction, which effectively integrates multi-attribute visual and textual information through cross-modal graph reasoning and alignment. Extensive data analysis and experimental results validate both the reliability of the MLE dataset and the superior performance of the AVGR-Net compared to state-of-the-art methods.
Bo Hu 0008, Leida Li, Ke Gu 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2026 Infrared Image Quality Estimation With Node-to-Graph Regression
abstract
By comparison with the commonly seen visible light images that can be effectively characterized within a Euclidean space, infrared images have non-Euclidean characteristics since their pixels contain rich thermal radiation information, such as heat distribution, surface temperature and thermal radiation. Considering the advantages of Graph Convolutional Networks (GCNs) in processing non-Euclidean data, this study proposes to introduce the GCNs to estimate the quality of infrared images by developing the Node-to-Graph Regression (NGR) model. To specify, the proposed NGR model is composed of two main steps, namely network establishment and network training. In the first step, following the classical researches of image quality estimation that include local distortion measurement followed by pooling for inferring the image quality score, this study captures the local distortion of the input infrared images by stacking up a set of Vision Graph (VSG) blocks to generate one node map, and then conducts the weighted pooling method on the node map to yield the graph output as the estimated quality score. In the second step, for enhancing the model's performance and generalization ability in the network training process, this study implements the node regression with the big data pre-training method to raise the local distortion extraction ability in a broad range of image scenarios and distortion intensities, and then performs the graph regression by using the knowledge distillation method to reduce the over-fitting risk. Using the largest-size infrared image quality evaluation database (I2QED), this study compared the proposed NGR model with three dozen mainstream and state-of-the-art competitors, and results showed that our proposed NGR model achieved the optimal performance.
Ke Gu 0001, Hongyan Liu 0004, Yubin Gao, Chen Wang 0019, Lai-Kuan Wong, Weisi Lin, Guangtao Zhai, Wenjun Zhang 0001, Daniel Thalmann
IEEE Trans. Multim.1
2026 Distortion-Sensitive Masked Autoencoder for Omnidirectional Video Quality Assessment
abstract
Omnidirectional Video Quality Assessment (OVQA) is a challenging task due to the limited availability of adequate numbers of training samples for learning representations of distortions on omnidirectional videos. The recent masked autoencoder (MAE) has shown promising performance in learning local and global representations in a self-supervised way, and can be used to attempt to mitigate the difficulty of having insufficient annotated samples to adequately train omnidirectional video quality prediction models. But the reconstruction tasks that MAE models are designed for do not pertain to predicting diverse perceptual distortions, especially those relevant to the task of OVQA. We have attempted to overcome these limitations to harness and apply the power of the MAE concept to the OVQA problem. Towards this purpose, we create a Distortion-Sensitive Masked AutoEncoder (DS-MAE) that is able to represent perceptual distortions on omnidirectional videos. DS-MAE extracts viewports from omnidirectional videos and employs a masked autoencoding module (MAM) and a knowledge replay module (KRM) to learn representations on each viewport. In the MAM, distorted patches from omnidirectional videos are masked, by replacing them with undistorted counterparts. The autoencoder is trained to reconstruct the masked distortions, imbuing them with the ability to represent diverse video degradations. The KRM extracts and stores content representations, which are then “replayed” to mitigate potential catastrophic forgetting of content during training of the DS-MAE. Finally, a simple OVQA model is constructed using the pre-trained DS-MAE across all viewports. The new model, called OmniVQA, was tested on three public OVQA datasets. The experimental results show that OmniVQA delivers competitive performance against all compared models.
Zongyao Hu, Lixiong Liu, Ke Gu 0001, Leida Li, Alan C. Bovik
IEEE Trans. Multim.3
2025 Mix-YOLONet: Deep Image Dehazing for Improving Object Detection
Xin Lim, Lai-Kuan Wong, Yuen Peng Loh, Ke Gu 0001, Weisi Lin
MMM (2)4
2025 Neurodynamics-Driven Model Predictive Control With Soft-Measurement for Desulfurization System
abstract
Sulfur dioxide (SO2) emissions are the main problems causing the air pollution and respiratory disease, which makes desulfurization important in the coal combustion and chemical industry processes. The wet-flue-gas desulfurization (WFGD) is currently an effective method to deal with this issue. However, the WFGD system is actually a complex process with multi-variable coupling, high nonlinearity and large time-delay, which brings great challenges for the traditional mechanism-based modeling and control. In this paper, a neurodynamics-driven model predictive control (NDMPC) with soft-measurement is proposed to predict and control the dynamics of SO2to improve the operational performance of the WFGD system. First, we design a self-organizing fuzzy neural network (SOFNN) as the soft-measuring model, which is driven by the data and neurodynamics to adaptively predict SO2in the WFGD system. Second, the designed SOFNN is considered as a predictive model that can give the predicted values for the future moments. Third, we also design a loss function where the one-step output of the predictive model and the control law are the independent variables. The resulting rolling optimization can give the optimal control law sequence by minimizing the loss function. Finally, simulation experiments on the practical data from the WFDG system show that the NDMPC outperforms the other methods in terms of the soft-measurement and control performance, especially generally reducing the SO2emission and economic cost by 60.25% and 22.27%, respectively.
Gongming Wang, Ke Gu 0001, Hong Chen 0025, Honggui Han, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.3
2025 Multi-Scale Local and Global Feature Fusion for Blind Quality Assessment of Enhanced Images
abstract
Image enhancement plays a crucial role in computer vision by improving visual quality while minimizing distortion. Traditional methods enhance images through pixel value transformations, yet they often introduce new distortions. Recent advancements in deep learning-based techniques promise better results but challenge the preservation of image fidelity. Therefore, it is essential to evaluate the visual quality of enhanced images. However, existing quality assessment methods frequently encounter difficulties due to the unique distortions introduced by these enhancements, thereby restricting their effectiveness. To address these challenges, this paper proposes a novel blind image quality assessment (BIQA) method for enhanced natural images, termed multi-scale local feature fusion and global feature representation-based quality assessment (MLGQA). This model integrates three key components: a multi-scale Feature Attention Mechanism (FAM) for local feature extraction, a Local Feature Fusion (LFF) module for cross-scale feature synthesis, and a Global Feature Representation (GFR) module using Vision Transformers to capture global perceptual attributes. This synergistic framework effectively captures both fine-grained local distortions and broader global features that collectively define the visual quality of enhanced images. Furthermore, in the absence of a dedicated benchmark for enhanced natural images, we design the Natural Image Enhancement Database (NIED), a large-scale dataset consisting of 8,581 original images and 102,972 enhanced natural images generated through a wide array of traditional and deep learning-based enhancement techniques. Extensive experiments on NIED demonstrate that the proposed MLGQA model significantly outperforms current state-of-the-art BIQA methods in terms of both prediction accuracy and robustness.
Jingchao Cao, Yutao Liu 0002, Feng Gao 0005, Ke Gu 0001, Guangtao Zhai, Junyu Dong, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.5
2025 Perceptual Information Fidelity for Quality Estimation of Industrial Images
abstract
Depending on high quality images, industrial vision technologies can basically oversee all the industrial production processes, such as workpiece processing and assembly automation, which play a highly significant role in promoting detection automation and production capacity in assembly lines. Unlike the natural scene images which consist of richer colors and natural lines, industrial images that cover complex industrial goods and equipment are made up of fewer colors, more regular shapes, massive graphic elements, etc., causing existing image processing methods for quality estimation, enhancement and monitoring to fail. Human beings usually play the part of the final receiver of an industrial image, so in the researches of image quality estimation, it is necessary to take the perception process of human eyes and brain to the input images into consideration. On this basis, we in this paper propose a novel perceptual information fidelity based image quality estimation model, abbreviated as PIF. Particularly, we first introduce a visual-cell low-pass filter and an optical-nerve noise model, which are separately inspired by the two processes: one is that an image in the form of optical signals arrives at the retina through the eye’s optical system to form the stimuli; the other is that the aforesaid stimuli in the form of electrical signals transfer to the human brain through the optical nerve. Second, we construct a novel image content-aware adjustor to optimize the above visual-cell low-pass filter and optical-nerve noise model. Third, we compare the two quantities of the information that is present in the clean image and how much of the information can be extracted from the lossy image to generate the overall quality score. Experiments on the two large-size industrial image quality databases demonstrate the excellent performance achieved by our proposed PIF model, with a remarkable performance gain over the existing state-of-the-art competitors.
Ke Gu 0001, Hongyan Liu 0004, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Attention and Mamba-Driven Quality Assessment for Underwater Images
abstract
Underwater imaging is essential in a variety of fields, including resource exploration, marine observation, and scientific research. However, the quality of underwater images is often compromised by environmental factors such as light scattering, absorption, and the presence of fog, leading to distortions such as color shifts, low contrast, and blurriness. To address these challenges, we propose a novel underwater image quality assessment (UIQA) method, the Attention and Mamba-driven Quality Index (AMQI). The AMQI model employs a multi-stage architecture designed to capture both local and global image features critical for underwater quality evaluation. First, a Shallow Feature Extractor (SFE) captures essential spatial details. Next, the Local Information Representation Network (LIR-Net), equipped with Channel Attention (CA) and Large Kernel-guided Spatial (LKS) mechanisms, enhances fine details and captures long-range dependencies to address underwater-specific distortions. The Global Information Representation Network (GIR-Net) further processes the features using a combination of the Visual State-Space Model (VSSM) and ResNet-50 to capture high-level semantic and contextual information. Finally, the Feature-Quality Mapping Network (FQM) converts the learned features into a quality score, ensuring precise predictions of image quality. Extensive experiments on the Underwater Image Quality Database (UIQD) demonstrate that AMQI outperforms current state-of-the-art IQA and UIQA models in terms of accuracy and correlation with human subjective evaluations. The model's robustness and generalization capabilities are further validated through detailed ablation studies and cross-database evaluations, showcasing its strong performance across diverse underwater environments. The source code is available athttps://github.com/ibaochao/AMQI.
Jingchao Cao, Baochao Zhang, Yutao Liu 0002, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong
IEEE Trans. Multim.5
2025 Air Pollution Monitoring by Integrating Local and Global Information in Self-Adaptive Multiscale Transform Domain
abstract
This paper proposed a novel image-based air pollution monitor (IAPM) by incorporating local and global information in the self-adaptive multiscale transform domain, so as to achieve the timely and effective leakage detection of typical air pollutants from a single image. To be specific, this paper first developed a screen-shaped module according to two significant findings in visual neuroscience, which include the high sensitivity of human eyes to horizontal and vertical stimuli and the center-surround inhibition, by designing and fusing the square module, horizontal strip module and vertical strip module parallelly for simulating the behaviour of human eyes to extract local features. Second, the learnable weights and proportional mapping were applied to incorporate the screen-shaped module and lightweight vision transformer as backbone, towards more richly exploiting and fusing local and global information just as the way a brain perceives external stimuli. Third, a new self-adaptive multiscale transform domain method was devised based on two motivations from the visual characteristics of multiscale perception and the brain characteristics of self-adaptive domain transform to modify the backbone by using the operations of pooling and pointwise convolution. Extensive experiments implemented on the datasets of carbon particulate matters and ethylene leakage confirmed the superior monitoring performance of the proposed IAPM model beyond the state-of-the-art (SOTA) peers by an accuracy gain of about 4%. Furthermore, the proposed IAPM model only required 0.089 GFLOPs and 0.15 million model parameters, remarkably outperforming SOTA competitors in computational efficiency and storage resources.
Ke Gu 0001, Hongyan Liu 0004, Bo Liu 0024, Junfei Qiao 0001, Weisi Lin, Wenjun Zhang 0001
IEEE Trans. Multim.1
2025 Deep No-Reference Quality Assessment for Underwater Enhanced Images
abstract
The goal of underwater image enhancement (UIE) is to boost the acquired underwater image quality, which increases the value of the underwater image significantly. However, without effective underwater enhanced image quality assessment (UEIQA) measures that benchmark the UIE, the process of UIE becomes driftless and the enhanced results of different UIE algorithms cannot be fairly compared. Toward this end, we in this work construct a dedicated UEIQA scheme on the basis of deep investigation of the underwater enhanced image characteristics. Specifically, in our proposed method, we respectively design deep neural networks to represent the unique attributes of the underwater enhanced image, such as color cast, local distortions, naturalness degree, sharpness, contrast, fog density, etc., that are highly correlated with the image quality. Then we introduce the Vision Transformer (ViT) to capture the dependencies among different image attributes and infer the image quality level. Extensive experiments conducted on three typical UEIQA databases, i.e., SOTA, UID2021 and SAUD, show that the proposed UEIQA model yields noteworthy higher prediction accuracy than the representative IQA and UEIQA metrics, e.g., achieving SRCC values of 0.891 ( vs. 0.749 in SAUD) and 0.933 ( vs. 0.798 in UID2021). The proposed UEIQA model will be released athttps://github.com/YT2015?tab=repositories.
Yutao Liu 0002, Baochao Zhang, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong
IEEE Trans. Multim.4
2024 KBY-Net: A Dual Learning Framework for Improving Object Detection in Rainy Weather Conditions
abstract
Rainy weather conditions significantly degrade image quality, posing a major challenge for object detection tasks. Conventional methods often address this issue through domain adaptation, or the "derain then detect" approach that utilizes image deraining as the preprocessing technique. This paper presents KBY-Net, a novel end-to-end Y-Net architecture that is built upon the YOLOv8 architecture and leverages multi-task learning for concurrent image restoration and object detection. First, KBY-Net incorporates a novel KBY-decoder designed for image deraining. This decoder leverages Cross Stage Partial (CSP) layer and kernel basis attention (KBA) module to improve feature representation. Second, KBY-Net adopted two innovative modules; a multi-Dconv head transposed attention (MDTA) module at the bottleneck and a multi-axis feature fusion (MFF) block at the neck of the Y-Net. The multi-DConv module empowers the model to capture long-range dependencies and complex representations, and the MFF block refines the extracted features – both contribute significantly to accurate object detection in challenging rainy scenes. Empirical evaluations on benchmark rainy datasets demonstrate that KBY-Net outperforms the state-ofthe-art object detection approaches by a significant margin both quantitatively and qualitatively
Zheng-Xian Keh, Lai-Kuan Wong, Yuen Peng Loh, Ke Gu 0001, Weisi Lin
MMAsia4
2024 Blind image quality index with high-level Semantic Guidance and low-level fine-grained Representation
Bo Hu 0008, Leida Li, Ke Gu 0001, Shuaijian Wang, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing4
2024 Coarse- and Fine-Grained Fusion Hierarchical Network for Hole Filling in View Synthesis
abstract
Depth image-based rendering (DIBR) techniques play an essential role in free-viewpoint videos (FVVs), which generate the virtual views from a reference 2D texture video and its associated depth information. However, the background regions occluded by the foreground in the reference view will be exposed in the synthesized view, resulting in obvious irregular holes in the synthesized view. To this end, this paper proposes a novel coarse and fine-grained fusion hierarchical network (CFFHNet) for hole filling, which fills the irregular holes produced by view synthesis using the spatial contextual correlations between the visible and hole regions. CFFHNet adopts recurrent calculation to learn the spatial contextual correlation, while the hierarchical structure and attention mechanism are introduced to guide the fine-grained fusion of cross-scale contextual features. To promote texture generation while maintaining fidelity, we equip CFFHNet with a two-stage framework involving an inference sub-network to generate the coarse synthetic result and a refinement sub-network for refinement. Meanwhile, to make the learned hole-filling model better adaptable and robust to the "foreground penetration" distortion, we trained CFFHNet by generating a batch of training samples by adding irregular holes to the foreground and background connection regions of high-quality images. Extensive experiments show the superiority of our CFFHNet over the current state-of-the-art DIBR methods. The source code will be available at https://github.com/wgc-vsfm/view-synthesis-CFFHNet.
Guangcheng Wang, Kui Jiang, Ke Gu 0001, Hongyan Liu 0004, Hantao Liu, Wenjun Zhang 0001
IEEE Trans. Image Process.3
2024 Perception-and-Cognition-Inspired Quality Assessment for Sonar Image Super-Resolution
abstract
Due to the light-independent imaging characteristics, sonar images play a crucial role in fields such as underwater detection and rescue. However, the resolution of sonar images is negatively correlated with the imaging distance. To overcome this limitation, Super-Resolution (SR) techniques have been introduced into sonar image processing. Nevertheless, it is not always guaranteed that SR maintains the utility of the image. Therefore, quantifying the utility of SR reconstructed Sonar Images (SRSIs) can facilitate their optimization and usage. Existing Image Quality Assessment (IQA) methods are inadequate for evaluating SRSIs as they fail to consider both the unique characteristics of sonar images and reconstruction artifacts while meeting task requirements. In this paper, we propose a Perception-and-Cognition-inspired quality Assessment method for Sonar image Super-resolution (PCASS). Our approach incorporates a hierarchical feature fusion-based framework inspired by the cognitive process in the human brain to comprehensively evaluate SRSIs' quality under object recognition tasks. Additionally, we select features at each level considering visual perception characteristics introduced by SR reconstruction artifacts such as texture abundance, contour details, and semantic information to measure image quality accurately. Importantly, our method does not require training data and is suitable for scenarios with limited available images. Experimental results validate its superior performance.
Boqin Cai, Sumei Zheng, Tiesong Zhao, Ke Gu 0001
IEEE Trans. Multim.5
2024 UIQI: A Comprehensive Quality Evaluation Index for Underwater Images
abstract
Due to the light absorption and scattering in waterbodies, acquired underwater images frequently suffer from color cast, blur, low contrast, noise, etc., which seriously degrade the image quality and affect their subsequent applications. Therefore, it is necessary to propose a reliable and practical underwater image quality assessment (IQA) model that can faithfully evaluate underwater image quality. To this end, in this article, we establish a novel quality assessment model for underwater images by in-depth analysis and characterization of multiple image properties. Specifically, we propose characterizing the image luminance, color cast, sharpness, contrast, fog density and noise to comprehensively describe the image quality to evaluate the underwater image quality more accurately. Dedicated features are elaborately investigated to characterize those quality-aware image properties. After feature extraction, we employ support vector regression (SVR) to integrate all the quality-aware features and regress them onto the underwater image quality score. Extensive tests performed on standard underwater image quality databases demonstrate the superior prediction performance of the proposed underwater IQA model to state-of-the-art congeneric quality assessment models.
Yutao Liu 0002, Ke Gu 0001, Jingchao Cao, Shiqi Wang 0001, Guangtao Zhai, Junyu Dong, Sam Kwong
IEEE Trans. Multim.2
2024 Underwater Image Quality Assessment: Benchmark Database and Objective Method
abstract
Underwater image quality assessment (UIQA) plays a crucial role in monitoring and detecting the quality of acquired underwater images in underwater imaging systems. Currently, the investigation of UIQA encounters two major challenges. First, a lack of large-scale UIQA databases for benchmarking UIQA algorithms remains, which greatly restricts the development of UIQA research. The other limitation is that there is a shortage of effective UIQA methods that can faithfully predict underwater image quality. To alleviate these two challenges, in this paper, we first construct a large-scale UIQA database (UIQD). Specifically, UIQD contains a total of 5369 authentic underwater images that span abundant underwater scenes and typical quality degradation conditions. Extensive subjective experiments are executed to annotate the perceived quality of the underwater images in UIQD. Based on an in-depth analysis of underwater image characteristics, we further establish a novel baseline UIQA metric that integrates channel and spatial attention mechanisms and a transformer. Channel- and spatial attention modules are used to capture the image channel and local quality degradations, while the transformer module characterizes the image quality from a global perspective. Multilayer perception is employed to fuse the local and global feature representations and yield the image quality score. Extensive experiments conducted on UIQD demonstrate that the proposed UIQA model achieves superior prediction performance compared with the state-of-the-art UIQA and IQA methods. The proposed UIQD and UIQA models will be released athttps://github.com/YT2015?tab=repositories.
Yutao Liu 0002, Baochao Zhang, Runze Hu, Ke Gu 0001, Guangtao Zhai, Junyu Dong
IEEE Trans. Multim.4
2024 Learning Label Semantics for Weakly Supervised Group Activity Recognition
abstract
Weakly supervised group activity recognition deals with the dependence on individual-level annotations during understanding scenes involving multiple individuals, which is a challenging task. Existing methods either take the trained detectors to extract individual features or utilize the attention mechanisms for partial context encoding, followed by integration to form the final group-level representations. However, the detectors require individual-level annotations during the training phase and have a mis-detection issue, and the partial contexts extracted immediately from the whole complex scene are too ambiguous without the guidance of concrete semantics. In this paper, we investigate the hierarchical structure inherent in group-level labels to extract the fine-grained semantics without using detectors for weakly supervised group activity recognition. A multi-hot encoding strategy combined with a semantic encoder is first adopted to get the label semantics embeddings. The semantic and visual scene information are then fused through a semantic decoder to obtain activity-specific features. Lastly, we employ the multi-label classification and integrate the scores of hierarchical activity labels. Experimental results show that our proposed method achieves the state-of-the-art performance on three benchmarks, and the accuracy on the Volleyball dataset exceeds the second-best method by 2%.
Lifang Wu, Ye Xiang, Ke Gu 0001, Ge Shi 0002
IEEE Trans. Multim.4
2024 Quality Assessment for Stitched Panoramic Images via Patch Registration and Bidimensional Feature Aggregation
abstract
Quality assessment for stitched panoramic images (SPIQA) is of great significance for the stitching algorithm optimization. By contrast, this task is much more challenging and arduous than traditional IQA task due to the high resolution of stitched panoramic images and the particularity and complexity of stitching distortions. For this task, we propose an effective method based on patch registration and bidimensional feature aggregation (PRBFA). First, inspired by the attention mechanism of the human visual system and the limited range of human vision, a soft patch segmentation and selection method is presented to determine the key patches in panoramic images to participate in the following patch matching and feature alignment stages, achieving patch registration between the panoramic image and the corresponding constituent images. Further, to fully simulate the human visual perception process from local viewport to panorama, the feature exploration is successively performed from local to global, which is also adaptive to the complexity of the distortions in stitched panoramic images. For performance testification, extensive experiments are conducted on the publicly released SPIQA database, the results of which prove the performance superiority of the PRBFA method.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Ke Gu 0001, Jinjian Wu
IEEE Trans. Multim.5
2023 Toward a No-Reference Quality Metric for Camera-Captured Images
abstract
Existing no-reference (NR) image quality assessment (IQA) metrics are still not convincing for evaluating the quality of the camera-captured images. Toward tackling this issue, we, in this article, establish a novel NR quality metric for quantifying the quality of the camera-captured images reliably. Since the image quality is hierarchically perceived from the low-level preliminary visual perception to the high-level semantic comprehension in the human brain, in our proposed metric, we characterize the image quality by exploiting both the low-level image properties and the high-level semantics of the image. Specifically, we extract a series of low-level features to characterize the fundamental image properties, including the brightness, saturation, contrast, noiseness, sharpness, and naturalness, which are highly indicative of the camera-captured image quality. Correspondingly, the high-level features are designed to characterize the semantics of the image. The low-level and high-level perceptual features play complementary roles in measuring the image quality. To infer the image quality, we employ the support vector regression (SVR) to map all the informative features to a single quality score. Thorough tests conducted on two standard camera-captured image databases demonstrate the effectiveness of the proposed quality metric in assessing the image quality and its superiority over the state-of-the-art NR quality metrics. The source code of the proposed metric for camera-captured images is released at https://github.com/YT2015?tab=repositories.
Runze Hu, Yutao Liu 0002, Ke Gu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Cybern.3
2023 Visibility and Distortion Measurement for No-Reference Dehazed Image Quality Assessment via Complex Contourlet Transform
abstract
Recently, most dehazed image quality assessment (DQA) methods have focused on estimating remaining haze and omitting distortion impact from the side effect of dehazing algorithms, which leads to their limited performance. Addressing this problem, we propose a method for learning both visibility and distortion-aware features no-reference (NR) dehazed image quality assessment (VDA-DQA). Visibility-aware features are exploited to characterize clarity optimization after dehazing, including the brightness-, contrast-, and sharpness-aware features extracted by the complex contourlet transform (CCT). Then, distortion-aware features are employed to measure the distortion artifacts of images, including the normalized histogram of the local binary pattern (LBP) from the reconstructed dehazed image and the statistics of the CCT subbands corresponding to the chroma and saturation map. Finally, all the above features are mapped into quality scores by support vector regression (SVR). Extensive experimental results on six public DQA datasets verify the superiority of the proposed VDA-DQA method in terms of consistency with subjective visual perception and outperform state-of-the-art methods.
Tuxin Guan, Ke Gu 0001, Hantao Liu, Yuhui Zheng, Xiaojun Wu 0001
IEEE Trans. Multim.3
2023 An Underwater Image Quality Assessment Metric
abstract
Various image enhancement algorithms are adopted to improve underwater images that often suffer from visual distortions. It is critical to assess the output quality of underwater images undergoing enhancement algorithms, and use the results to optimise underwater imaging systems. In our previous study, we created a benchmark for quality assessment of underwater image enhancement via subjective experiments. Building on the benchmark, this paper proposes a new objective metric that can automatically assess the output quality of image enhancement, namely UWEQM. By characterising specific underwater physics and relevant properties of the human visual system, image quality attributes are computed and combined to yield an overall metric. Experimental results show that the proposed UWEQM metric yields good performance in predicting image quality as perceived by human subjects.
Hantao Liu, Delu Zeng, Tao Xiang 0001, Leida Li, Ke Gu 0001
IEEE Trans. Multim.6
2023 Toward Visual Behavior and Attention Understanding for Augmented 360 Degree Videos
abstract
Augmented reality (AR) overlays digital content onto reality. In an AR system, correct and precise estimations of user visual fixations and head movements can enhance the quality of experience by allocating more computational resources for analyzing, rendering, and 3D registration on the areas of interest. However, there is inadequate research to help in understanding the visual explorations of the users when using an AR system or modeling AR visual attention. To bridge the gap between the saliency prediction on real-world scenes and on scenes augmented by virtual information, we construct the ARVR saliency dataset. The virtual reality (VR) technique is employed to simulate the real-world. Annotations of object recognition and tracking as augmented contents are blended into omnidirectional videos. The saliency annotations of head and eye movements for both original and augmented videos are collected and together constitute the ARVR dataset. We also design a model that is capable of solving the saliency prediction problem in AR. Local block images are extracted to simulate the viewport and offset the projection distortion. Conspicuous visual cues in the local block images are extracted to constitute the spatial features. The optical flow information is estimated as an important temporal feature. We also consider the interplay between virtual information and reality. The composition of the augmentation information is distinguished, and the joint effects of adversarial augmentation and complementary augmentation are estimated. The Markov chain is constructed with block images as graph nodes. In the determination of the edge weights, both the characteristics of the viewing behaviors and the visual saliency mechanisms are considered. The order of importance for block images is estimated through the state of equilibrium of the Markov chain. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method.
Yucheng Zhu, Xiongkuo Min, Dandan Zhu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Ke Gu 0001, Jiantao Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2022 Reference-Free DIBR-Synthesized Video Quality Metric in Spatial and Temporal Domains
abstract
Depth image-based rendering (DIBR) techniques play an important role in free viewpoint videos (FVVs), which have a wide range of applications including immersive entertainment, remote monitoring, education, etc. FVVs are usually synthesized by DIBR techniques in a “blind” environment (without a reference video). Thus, an effective reference-free synthesized video quality assessment (VQA) metric is vital. At present, many image quality assessment (IQA) algorithms for DIBR-synthesized images have been proposed, but limited researches have been concerned about the quality assessment of DIBR-synthesized videos. To this end, this paper proposes a novel reference-free VQA method for synthesized videos, which operates in Spatial and Temporal Domains, dubbed as STD. The design fundamental of the proposed STD metric considers the effects of two major distortions introduced by DIBR techniques on the visual quality of synthesized videos. First, considering the geometric distortion introduced by DIBR technologies can increase high-frequency contents of the synthesized frame, the influence of the geometric distortion on the visual quality of a synthesized video can be effectively evaluated by estimating high-frequency energies of each synthesized frame in spatial domain. Second, temporal inconsistency caused by DIBR techniques brings the temporal flicker distortion, which is one of the most annoying artifacts in DIBR-synthesized videos. In temporal domain, we quantify temporal inconsistency by measuring motion differences between consecutive frames. Specifically, optical flow method is first used to estimate the motion field between adjacent frames. Then, we calculate the structural similarity of adjacent optical flow fields and further adopt the structural similarity value to weight the pixel differences of adjacent optical flow fields. Experiments show that the above two features are able to well perceive the visual quality of DIBR-synthesized videos. Furthermore, since the two features are extracted from spatial and temporal domains, respectively, we integrate them using a linear weighting strategy to obtain our STD metric, which proves advantageous over two components and the competing state-of-the-art I/VQA methods. The source code is available athttps://github.com/wgc-vsfm/DIBR-video-quality-assessment.
Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Kui Jiang, Zheng He 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Omnidirectional Image Quality Assessment by Distortion Discrimination Assisted Multi-Stream Network
abstract
Omnidirectional image (OI) quality assessment is crucial to facilitate the development of virtual reality (VR) related technology. In this work, a distortion discrimination assisted multi-stream network is proposed for OI quality assessment. The multi-stream architecture is constructed by generating the viewport images received by the retina at one point to simulate the characteristics of humans perceiving VR contents. Additionally, the strategy of generating several viewport image sets from one OI is proposed for data augmentation. Furthermore, the facts that the human brain has the ability for both quality assessment and distortion type distinguishment, and the process of human brain handling two tasks exists information interaction inspire us to employ an auxiliary distortion discrimination task to facilitate the quality assessment task learning. Extensive experiments conducted on two public OI databases demonstrate the superiority of the proposed method to both traditional 2D quality metrics and existing metrics specific for OIs. Moreover, utilizing the assistant task is proven to be more effective than the single task learning for OI quality evaluation. Better generalization performance is also verified to be another valuable trait of the proposed method.
Yu Zhou 0009, Yanjing Sun, Leida Li, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Active Vision for Deep Visual Learning: A Unified Pooling Framework
abstract
Convolutional neural networks (CNNs) can be generally regarded as learning-based visual systems for computer vision tasks. By imitating the operating mechanism of the human visual system (HVS), CNNs can even achieve better results than human beings in some visual tasks. However, they are primary when compared to the HVS for the reason that the HVS has the ability of active vision to promptly analyze and adapt to specific tasks. In this article, a new unified pooling framework is proposed and a series of pooling methods are designed based on the framework to implement active vision to CNNs. In addition, an active selection pooling (ASP) is put forward to reorganize the existing and newly proposed pooling methods. The CNN models with an ASP tend to have a behavior of focus selection according to tasks during the training process, which acts extremely similar to the HVS.
Nan Guo 0006, Ke Gu 0001, Junfei Qiao 0001, Hantao Liu
IEEE Trans. Ind. Informatics2
2022 Single Image Super-Resolution Quality Assessment: A Real-World Dataset, Subjective Studies, and an Objective Metric
abstract
Numerous single image super-resolution (SISR) algorithms have been proposed during the past years to reconstruct a high-resolution (HR) image from its low-resolution (LR) observation. However, how to fairly compare the performance of different SISR algorithms/results remains a challenging problem. So far, the lack of comprehensive human subjective study on large-scale real-world SISR datasets and accurate objective SISR quality assessment metrics makes it unreliable to truly understand the performance of different SISR algorithms. We in this paper make efforts to tackle these two issues. Firstly, we construct a real-world SISR quality dataset (i.e., RealSRQ) and conduct human subjective studies to compare the performance of the representative SISR algorithms. Secondly, we propose a new objective metric, i.e., KLTSRQA, based on the Karhunen-Loéve Transform (KLT) to evaluate the quality of SISR images in a no-reference (NR) manner. Experiments on our constructed RealSRQ and the latest synthetic SISR quality dataset (i.e., QADS) have demonstrated the superiority of our proposed KLTSRQA metric, achieving higher consistency with human subjective scores than relevant existing NR image quality assessment (NR-IQA) metrics. The dataset and the code will be made available at https://github.com/Zhentao-Liu/RealSRQ-KLTSRQA.
Qiuping Jiang, Ke Gu 0001, Feng Shao 0001, Xinfeng Zhang 0001, Hantao Liu, Weisi Lin
IEEE Trans. Image Process.3
2022 Dynamic Backlight Scaling Considering Ambient Luminance for Mobile Videos on LCD Displays
abstract
The power consumption of mobile devices is always a concern for users due to the constraints on the size and weight of mobile devices. Up to now, mobile video traffic has accounted for the majority of total traffic, which implies that viewing mobile videos has been the major activity when using mobile devices. Among all the subsystems involving in mobile video playback, the display is the most power consuming subsystem. To reduce the power consumption, dynamic backlight scaling (DBS) technique is developed by adjusting the backlight magnitude when playing the mobile video. However, the convenience of mobile devices makes lots of people watch mobile videos in various luminance environments, which makes the existing DBS methods ineffective since ambient luminance varies greatly. In this paper, we propose a novel DBS strategy to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users’ quality of experience (QoE). In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance, and backlight luminance. We then investigate whether the continuous playback of backlight-scaled videos using the proposed scaling magnitude under various luminance environments would cause flicker effect or not. Motivated by the findings of these studies, we implement a novel DBS strategy for mobile energy saving which is suitable for various ambient luminance conditions. The experimental results demonstrate that the proposed DBS strategy can save more than 40 percent power at most and can save 10 percent power even at a very high ambient luminance condition. We also show that the proposed DBS strategy can be easily adapted to different user preferences and different devices, and can be conveniently integrated into practical applications.
Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Siwei Ma 0001, Xiaokang Yang 0001
IEEE Trans. Mob. Comput.4
2022 Combining Retargeting Quality and Depth Perception Measures for Quality Evaluation of Retargeted Stereopairs
abstract
Stereoscopic Image Retargeting (SIR) aims to adapt stereoscopic images and videos to 3D display devices with various aspect ratios by emphasizing the important content while retaining surrounding context with minimal visual distortion. To address the issue of SIR evaluation, this paper presents a new objective quality assessment method for retargeted stereopairs by combining image quality and depth perception measures. Specifically, the image quality measure is conducted between the source and retargeted intermediate views generated by the view synthesis method to characterize the geometric distortion and content loss of the retargeted stereopair, while several depth-aware features are extracted to measure the visual comfort/discomfort and depth sensation when human views a 3D scene. Then, the extracted features are integrated into an overall perceptual quality prediction. Experiment results on NBU SIRQA and SIRD databases verify the superiority of our method.
Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Zhenqi Fu, Xiangchao Meng, Ke Gu 0001, Yo-Sung Ho
IEEE Trans. Multim.6
2021 Improved deep CNNs based on Nonlinear Hybrid Attention Module for image classification
Nan Guo 0006, Ke Gu 0001, Junfei Qiao 0001, Jing Bi 0001
Neural Networks2
2021 Predicting the Quality of View Synthesis With Color-Depth Image Fusion
abstract
With the increasing prevalence of free-viewpoint video applications, virtual view synthesis has attracted extensive attention. In view synthesis, a new viewpoint is generated from the input color and depth images with a depth-image-based rendering (DIBR) algorithm. Current quality evaluation models for view synthesis typically operate on the synthesized images, i.e. after the DIBR process, which is computationally expensive. So a natural question is that can we infer the quality of DIBR-based synthesized images using the input color and depth images directly without performing the intricate DIBR operation. With this motivation, this paper presents a no-reference image quality prediction model for view synthesis via COlor-Depth Image Fusion, dubbed CODIF, where the actual DIBR is not needed. First, object boundary regions are detected from the color image, and a Wavelet-based image fusion method is proposed to imitate the interaction between color and depth images during the DIBR process. Then statistical features of the interactional regions and natural regions are extracted from the fused color-depth image to portray the influences of distortions in color/depth images on the quality of synthesized views. Finally, all statistical features are utilized to learn the quality prediction model for view synthesis. Extensive experiments on public view synthesis databases demonstrate the advantages of the proposed metric in predicting the quality of view synthesis, and it even suppresses the state-of-the-art post-DIBR view synthesis quality metrics.
Leida Li, Yipo Huang, Jinjian Wu, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Ensemble Meta-Learning for Few-Shot Soot Density Recognition
abstract
In each petrochemical plant around the world, the flare stack as a requisite facility produces a large amount of soot due to the incomplete combustion of flare gas, and this strongly endangers air quality and human health. Despite severe damages, the abovementioned abnormal conditions rarely occur, and, thus, only few-shot samples are available. To address such difficulty, in this article, we design an image-based flare soot density recognition network (FSDR-Net) via a new ensemble meta-learning technology. More particularly, we first train a deep convolutional neural network (CNN) by applying the model-agnostic meta-learning algorithm on a variety of learning tasks that are relevant to the flare soot recognition so as to obtain the general-purpose optimized initial parameters (GOIP). Second, for the new task of recognizing the flare soot density via only few-shot instances, a new ensemble is developed to selectively aggregate several predictions that are generated based on a wide range of learning rates and a small number of gradient steps. Results of experiments conducted on the density recognition of flare soot corroborate the superiority of our proposed FSDR-Net as compared with the popular and state-of-the-art deep CNNs.
Ke Gu 0001, Junfei Qiao 0001
IEEE Trans. Ind. Informatics1
2021 Semi-Reference Sonar Image Quality Assessment Based on Task and Visual Perception
abstract
In submarine and underwater detection tasks, conventional optical imaging and analysis methods are not universally applicable due to the limited penetration depth of visible light. Instead, sonar imaging has become a preferred alternative. However, the capture and transmission conditions in complicated and dynamic underwater environments inevitably lead to visual quality degradation of sonar images, which might also impede further recognition, analysis and understanding. To measure this quality decrease and provide a solid quality indicator for sonar image enhancement, we propose a task- and perception-oriented sonar image quality assessment (TPSIQA) method, in which a semi-reference (SR) approach is applied to adapt to the limited bandwidth of underwater communication channels. In particular, we exploit reduced visual features that are critical for both human perception of and object recognition in sonar images. The final quality indicator is obtained through ensemble learning, which aggregates an optimal subset of multiple base learners to achieve both high accuracy and a high generalization ability. In this way, we are able to develop a compact but generalized quality metric using a small database of sonar images. Experimental results demonstrate competitive performance, high efficiency, and strong robustness of our method compared to the latest available image quality metrics.
Ke Gu 0001, Tiesong Zhao, Gangyi Jiang, Patrick Le Callet
IEEE Trans. Multim.2
2021 Exploiting Local Degradation Characteristics and Global Statistical Properties for Blind Quality Assessment of Tone-Mapped HDR Images
abstract
Tone mapping operators (TMOs) are developed to convert a high dynamic range (HDR) image into a low dynamic range (LDR) one for display with the goal of preserving as much visual information as possible. However, image quality degradation is inevitable due to the dynamic range compression during the tone-mapping process. This accordingly raises an urgent demand for effective quality evaluation methods to select a high-quality tone-mapped image (TMI) from a set of candidates generated by distinct TMOs or the same TMO with different parameter settings. A key element to the success of TMI quality evaluation is to extract effective features that are highly consistent with human perception. Towards this end, this paper proposes a novel blind TMI quality metric by exploiting both local degradation characteristics and global statistical properties for feature extraction. Several image attributes including texture, structure, colorfulness and naturalness are considered either locally or globally. The extracted local and global features are aggregated into an overall quality via regression. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art blind quality models designed for synthetically distorted images (SDIs) and the blind quality models specifically developed for TMIs.
Xuejin Wang, Qiuping Jiang, Feng Shao 0001, Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.4
2021 PM₂.₅ Monitoring: Use Information Abundance Measurement and Wide and Deep Learning
abstract
This article devises a photograph-based monitoring model to estimate the real-time PM2.5concentrations, overcoming currently popular electrochemical sensor-based PM2.5monitoring methods’ shortcomings such as low-density spatial distribution and time delay. Combining the proposed monitoring model, the photographs taken by various camera devices (e.g., surveillance camera, automobile data recorder, and mobile phone) can widely monitor PM2.5concentration in megacities. This is beneficial to offering helpful decision-making information for atmospheric forecast and control, thus reducing the epidemic of COVID-19. To specify, the proposed model fuses Information Abundance measurement and Wide and Deep learning, dubbed as IAWD, for PM2.5monitoring. First, our model extracts two categories of features in a newly proposed DS transform space to measure the information abundance (IA) of a given photograph since the growth of PM2.5concentration decreases its IA. Second, to simultaneously possess the advantages of memorization and generalization, a new wide and deep neural network is devised to learn a nonlinear mapping between the above-mentioned extracted features and the groundtruth PM2.5concentration. Experiments on two recently established datasets totally including more than 100 000 photographs demonstrate the effectiveness of our extracted features and the superiority of our proposed IAWD model as compared to state-of-the-art relevant computing techniques.
Ke Gu 0001, Hongyan Liu 0004, Zhifang Xia, Junfei Qiao 0001, Weisi Lin, Daniel Thalmann
IEEE Trans. Neural Networks Learn. Syst.1
2020 An adaptive hybrid evolutionary immune multi-objective algorithm based on uniform distribution selection
Junfei Qiao 0001, Shengxiang Yang, Cuili Yang, Wenjing Li 0004, Ke Gu 0001
Inf. Sci.6
2020 Learning a Unified Blind Image Quality Metric via On-Line and Off-Line Big Training Instances
abstract
In this work, we resolve a big challenge that most current image quality metrics (IQMs) are unavailable across different image contents, especially simultaneously coping with natural scene (NS) images or screen content (SC) images. By comparison with existing works, this paper deploys on-line and off-line data for proposing a unified no-reference (NR) IQM, not only applied to different distortion types and intensities but also to various image contents including classical NS images and prevailing SC images. Our proposed NR IQM is developed with two data-driven learning processes following feature extraction, which is based on scene statistic models, free-energy brain principle, and human visual system (HVS) characteristics. In the first process, the scene statistic models and an image retrieve technique are combined, based on on-line and off-line training instances, to derive a novel loose classifier for retrieving clean images and helping to infer the image content. In the second process, the features extracted by incorporating the inferred image content, free-energy and low-level perceptual characteristics of the HVS are learned by utilizing off-line training samples to analyze the distortion types and intensities and thereby to predict the image quality. The two processes mentioned above depend on a gigantic quantity of training data, much exceeding the number of images applied to performance validation, and thus make our model's performance more reliable. Through extensive experiments, it has been validated that the proposed blind IQM is capable of simultaneously inferring the quality of NS and SC images, and it has attained superior performance as compared with popular and state-of-the-art IQMs on the subjective NS and SC image quality databases. The source code of our model will be released with the publication of the paper at https://kegu.netlify.com.
Ke Gu 0001, Junfei Qiao 0001, Qiuping Jiang, Weisi Lin, Daniel Thalmann
IEEE Trans. Big Data1
2020 Statistical and Structural Information Backed Full-Reference Quality Measure of Compressed Sonar Images
abstract
In sonar applications, important information such as distributions of minerals, underwater creatures has a high probability of being contained in sonar images. In many underwater applications such as underwater rescue and biometric tracking, it is necessary to send sonar images underwater for further analysis. Due to the bad conditions of underwater acoustic channel and current underwater acoustic communication technologies, sonar images very possibly suffer from several typical types of distortions. As far as we know, limited efforts have been made to gather meaningful sonar image databases and benchmark reliable objective quality model, so far. This paper develops a new objective sonar image quality predictor (SIQP), whose core is the combination of two features specific to a quality measure of sonar images. These two features, which come from statistical and structural information inspired by the characteristics of sonar images and the human visual system, reflect image quality from the global and detailed aspects. The performance comparison of the proposed metric with popular and prevailing quality evaluation models is conducted using a newly established sonar image quality database. The results of experiments show the superiority of our SIQP metric over the available quality evaluation models.
Ke Gu 0001, Weisi Lin, Fei Yuan 0001, En Cheng
IEEE Trans. Circuits Syst. Video Technol.2
2020 Blind Realistic Blur Assessment Based on Discrepancy Learning
abstract
Blur is one of the most common distortions that degrade natural images. This stimulates the blossom of sharpness assessment metrics. Existing sharpness metrics possess good performance for evaluating simulated blur, but are limited for the more common realistic blur that are introduced during image capture and processing in real life. To this end, we propose an effective Realistic Blur Assessment method (RBA) based on discrepancy learning. First, motivated by the fact that the distortion-free reference images are usually unavailable in practice, but the Human Visual System (HVS) can still accurately perceive image sharpness by quantifying the perceptual discrepancy between the distorted image and the hallucinated reference image in mind, we propose to train a discrepancy generation model to automatically generate the discrepancy map from the distorted image analogous to the HVS. This is achieved by using a deep neural network with rich training images. With the discrepancy map, two sharpness-aware features, i.e. sparse representation based entropy of primitive and content-guided variation of power, are then extracted to severally quantify spatial visual information amount and spectral power. Finally, the two features are integrated to produce the overall sharpness score. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-arts.
Leida Li, Yu Zhou 0009, Ke Gu 0001, Yuzhe Yang 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Unsupervised Blind Image Quality Evaluation via Statistical Measurements of Structure, Naturalness, and Perception
abstract
Most existing blind image quality assessment (BIQA) methods belong to supervised methods, which always need a large number of image samples and expensive subjective scores for training a quality prediction model. In this paper, we focus our attention on the unsupervised BIQA methods and put forward a novel unsupervised approach. The main idea of our method is to quantify the image quality degradation through measuring the structure, naturalness, and the perception quality variations of the distorted image from the pristine natural images. In specific, the structure variation is captured by the deviations of the image phase congruency and gradients distributions. The naturalness variation is characterized through the distributions variations of the locally mean subtracted and contrast normalized (MSCN) coefficients and the products of pairs of the adjacent MSCN coefficients. Compared with existing unsupervised methods, we initiatively introduce the perception quality measurement into the construction of unsupervised BIQA method, which is conducted by characterizing the prediction discrepancy between the image and its brain prediction based on the free-energy principle in the newly revealed brain theory. After feature extraction, we learn a pristine multivariate Gaussian (MVG) model with the extracted features from a set of pristine natural images. The quality of a new image is finally defined as the distance between its MVG model and the learned pristine MVG model. The extensive experiments conducted on LIVE, TID2013, CSIQ, Toyama, CID2013, and the Waterloo Exploration databases demonstrate that the proposed method achieves comparative prediction performance with the state-of-the-art BIQA methods.
Yutao Liu 0002, Ke Gu 0001, Yongbing Zhang 0002, Xiu Li 0001, Guangtao Zhai, Debin Zhao, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Blind Quality Metric of DIBR-Synthesized Images in the Discrete Wavelet Transform Domain
abstract
Free viewpoint video (FVV) has received considerable attention owing to its widespread applications in several areas such as immersive entertainment, remote surveillance and distanced education. Since FVV images are synthesized via a depth image-based rendering (DIBR) procedure in the "blind" environment (without reference images), a real-time and reliable blind quality assessment metric is urgently required. However, the existing image quality assessment metrics are insensitive to the geometric distortions engendered by DIBR. In this research, a novel blind method of DIBR-synthesized images is proposed based on measuring geometric distortion, global sharpness and image complexity. First, a DIBR-synthesized image is decomposed into wavelet subbands by using discrete wavelet transform. Then, the Canny operator is employed to detect the edges of the binarized low-frequency subband and high-frequency subbands. The edge similarities between the binarized low-frequency subband and high-frequency subbands are further computed to quantify geometric distortions in DIBR-synthesized images. Second, the log-energies of wavelet subbands are calculated to evaluate global sharpness in DIBR-synthesized images. Third, a hybrid filter combining the autoregressive and bilateral filters is adopted to compute image complexity. Finally, the overall quality score is derived to normalize geometric distortion and global sharpness by the image complexity. Experiments show that our proposed quality method is superior to the competing reference-free state-of-the-art DIBR-synthesized image quality models.
Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Leida Li, Zhifang Xia, Lifang Wu
IEEE Trans. Image Process.3
2020 Deep Dual-Channel Neural Network for Image-Based Smoke Detection
abstract
Smoke detection plays an important role in industrial safety warning systems and fire prevention. Due to the complicated changes in the shape, texture, and color of smoke, identifying the smoke from a given image still remains a substantial challenge, and this has accordingly aroused a considerable amount of research attention recently. To address the problem, we devise a new deep dual-channel neural network (DCNN) for smoke detection. In contrast to popular deep convolutional networks (e.g., Alex-Net, VGG-Net, Res-Net, and Dense-Net and the DNCNN that is specifically devoted to detecting smoke), our proposed end-to-end network is mainly composed of dual channels of deep subnetworks. In the first subnetwork, we sequentially connect multiple convolutional layers and max-pooling layers. Then, we selectively append the batch normalization layer to each convolutional layer for overfitting reduction and training acceleration. The first subnetwork is shown to be good at extracting the detailed information of smoke, such as texture. In the second subnetwork, in addition to the convolutional, batch normalization, and max-pooling layers, we further introduce two important components. One is the skip connection for avoiding the vanishing gradient and improving the feature propagation. The other is the global average pooling for reducing the number of parameters and mitigating the overfitting issue. The second subnetwork can capture the base information of smoke, such as contours. We finally deploy a concatenation operation to combine the aforementioned two deep subnetworks to complement each other. Based on the augmented data obtained by rotating the training images, our proposed DCNN can promptly and stably converge to the perfect performance. Experimental results conducted on the publicly available smoke detection database verify that the proposed DCNN has attained a very high detection rate that exceeds 99.5% on average, superior to state-of-the-art relevant competitors. Furthermore, our DCNN only employs approximately one-third of the parameters needed by the comparatively tested deep neural networks. The source code of DCNN will be released at https://kegu.netlify.com/.
Ke Gu 0001, Zhifang Xia, Junfei Qiao 0001, Weisi Lin
IEEE Trans. Multim.1
2020 ATMFN: Adaptive-Threshold-Based Multi-Model Fusion Network for Compressed Face Hallucination
abstract
Although tremendous strides have been recently made in face hallucination, exiting methods based on a single deep learning framework can hardly satisfactorily provide fine facial features from tiny faces under complex degradation. This article advocates an adaptive-threshold-based multi-model fusion network (ATMFN) for compressed face hallucination, which unifies different deep learning models to take advantages of their respective learning merits. First of all, we construct CNN-, GAN- and RNN-based underlying super-resolvers to produce candidate SR results. Further, the attention subnetwork is proposed to learn the individual fusion weight matrices capturing the most informative components of the candidate SR faces. Particularly, the hyper-parameters of the fusion matrices and the underlying networks are optimized together in an end-to-end manner to drive them for collaborative learning. Finally, a threshold-based fusion and reconstruction module is employed to exploit the candidates' complementarity and thus generate high-quality face images. Extensive experiments on benchmark face datasets and real-world samples show that our model outperforms the state-of-the-art SR methods in terms of quantitative indicators and visual effects. The code and configurations are released at https://github.com/kuihua/ATMFN.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Guangcheng Wang, Ke Gu 0001, Junjun Jiang
IEEE Trans. Multim.5
2020 Reduced Reference Stereoscopic Image Quality Assessment Using Sparse Representation and Natural Scene Statistics
abstract
An ideal quality assessment model should simulate the properties of the visual brain to be consistent with human evaluation. The visual brain appears to have both evolved to seek an efficient, decorrelated representation of image information and to “match” the statistics of the natural image. On one hand, the theoretical studies suggest that sparse representation resembles the strategy in the primary visual cortex of brain for representing natural images. On the other hand, the natural scene statistics have driven the evolution of human visual system and have also inspired the understanding and simulating of visual perception. Inspired by these observations, in this paper, we propose a novel reduced-reference stereoscopic image quality assessment metric using sparse representation and natural scene statistics to simulate the visual perception of the brain. Specifically, the distribution statistics of the classified visual primitives extracted by sparse representation are used to measure the visual information, which is closely related to the hierarchical progressive process of human visual perception. Particularly, the mutual information of classified primitives between two view images is derived as a binocular cue to simulate the binocular fusion process. The maximum mechanism that is applied to select the visual information is a pooling mechanism with which complex cells use the maximal stimuli from a group of simple cells during the transfer process in the primary visual cortex. The natural scene statistics of locally normalized luminance coefficients are used to evaluate the natural losses due to the presence of distortions. The differences of the visual information and the natural scene statistics between the original and distorted images are used to compute the quality score by a prediction function which is trained using support vector regression. Experimental results show that the proposed metric outperforms the state-of-the-art stereoscopic image quality assessment metrics on LIVE 3D IQA database and NBU-MDSID Phase-II database, and delivers competitive performance on Waterloo IVC 3D database.
Zhaolin Wan, Ke Gu 0001, Debin Zhao
IEEE Trans. Multim.2
2020 Blind Image Quality Assessment by Natural Scene Statistics and Perceptual Characteristics
abstract
Opinion-unaware blind image quality assessment (OU BIQA) refers to establishing a blind quality prediction model without using the expensive subjective quality scores, which is a highly promising direction in the BIQA research. In this article, we focus on OU BIQA and propose a novel OU BIQA method. Specifically, in our proposed method, we deeply investigate the natural scene statistics (NSS) and the perceptual characteristics of the human brain for visual perception. Accordingly, a set of quality-aware NSS and perceptual characteristics-related features are designed to characterize the image quality effectively. For inferring the image quality, we learn a pristine multivariate Gaussian (MVG) model on a collection of pristine images, which serves as the reference information for quality evaluation. At last, the quality of a new given image is defined by measuring the divergence between its MVG model and the learned pristine MVG model. Thorough experiments performed on seven popular image databases demonstrate that the proposed OU BIQA method delivers superior performance to the state-of-the-art OU BIQA methods. The Matlab source code of the proposed method will be made publicly available at https://github.com/YT2015?tab=;repositories.
Yutao Liu 0002, Ke Gu 0001, Xiu Li 0001, Yongbing Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Blind Quality Assessment for 3D-synthesized Images by Measuring Geometric Distortions and Image Complexity
abstract
Free viewpoint video (FVV), owing to its comprehensive applications in immersive entertainment, remote surveillance and distanced education, has received extensive attention and been regarded as a new important direction of video technology development. Depth image-based rendering (DIBR) technologies are employed to synthesize FVV images in the "blind" environment. Therefore, a real-time reliable blind quality assessment metric is urgently required. However, existing stste-of-art quality assessment methods are limited to estimate geometric distortions generated by DIBR. In this research, a novel blind quality metric, measuring Geometric Distortions and Image Complexity (GDIC), is proposed for DIBR-synthesized images. Firstly, a DIBR-synthesized image is decomposed into wavelet subbands by using discrete wavelet transform. Then, we adopt canny operator to capture the edge of wavelet subbands and compute the edge similarity between low-frequency subband and highfrequency subbands. The edge similarity is used to quantify geometric distortions in DIBR-synthesized images. Secondly, a hybrid filter combining the autoregressive and bilateral filter is adopted to compute image complexity. Finally, the overall quality score is calculated by normalizing geometric distortions via image complexity. Experiments show that our proposed GDIC is superior to prevailing image quality assessment metrics, which were intended for natural and DIBR-synthesized images.
Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Zhifang Xia
ICASSP3
2019 MC360IQA: The Multi-Channel CNN for Blind 360-Degree Image Quality Assessment
abstract
In this paper, we present a multi-channel convolution neural network (CNN) for blind 360-degree image quality assessment (MC360IQA). To be consistent with the visual content of 360-degree images seen in the VR device, our model adopts the viewport images as the input. Specifically, we project each 360-degree image into six viewport images to cover omnidirectional visual content. By rotating the longitude of the front view, we can project one omnidirectional image onto lots of different groups of viewport images, which is an efficient way to avoid overfitting. MC360IQA consists of two parts, multi-channel CNN and image quality regressor. Multi-channel CNN includes six parallel ResNet34 networks, which are used to extract the features of the corresponding six viewport images. Image quality regressor fuses the features and regresses them to final scores. The results show that our model achieves the best performance among the state-of-art full-reference (FR) and no-reference (NR) image quality assessment (IQA) models on the available 360-degree IQA database.
Wei Sun 0029, Weike Luo, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001, Ke Gu 0001, Siwei Ma 0001
ISCAS6
2019 Frequency-Domain Analysis Based Exploitation Of Color Channels For Color Image Demosaicking
abstract
Color-difference interpolation (CDI) has been a widely used technique for various color demosaicking methods. CDI-based methods perform interpolation in the color-difference domain assuming that the color-difference signal is a low-pass signal. Recently, a residual interpolation (RI) algorithm, which conducts interpolation in the residual domain, has been developed, and it assumes that the residual domain is flatter or smoother than the channel-difference domain. In this paper, we comprehensively show a frequency domain analysis of these assumptions and observe that it is image dependent and creates artifacts in the interpolated image. With this view, we propose an algorithm that uses the inter-color correlation as well as the residual smoothness among the different channel much better than the existing algorithms. Experimental results emphasize that the proposed algorithm atribute better performances the existing algorithms in terms of both visual and objective quality.
Sunil Prasad Jaiswal, Vinit Jakhetiya, Ke Gu 0001, Sharath Chandra Guntuku, Ashutosh Singla
VCIP3
2019 Blind image quality assessment based on joint log-contrast statistics
Qiaohong Li, Weisi Lin, Ke Gu 0001, Yabin Zhang 0002, Yuming Fang 0001
Neurocomputing3
2019 A Highly Efficient Blind Image Quality Assessment Metric of 3-D Synthesized Images Using Outlier Detection
abstract
With multitudes of image processing applications, image quality assessment (IQA) has become a prerequisite for obtaining maximally distinctive statistics from images. Despite the widespread research in this domain over several years, existing IQA algorithms have a number of key limitations concerning different image distortion types and algorithms' computational efficiency. Images that are synthesized using depth image-based rendering have applications in various disciplines, such as free viewpoint videos, which enable synthesis of novel realistic images in the referenceless environment. In the literature, very few no-reference (NR) quality assessment metrics of three-dimensional (3-D) synthesized images are proposed, and most of them are computationally expensive, which makes it difficult for them to be deployed in real-time applications. In this paper, we attribute the geometrically distorted pixels as outliers in 3-D synthesized images. This assumption is validated using the three $sigma$ rule-based robust outlyingness ratio. We propose a novel fast and accurate blind IQA metric of 3-D synthesized images using nonlinear median filtering since the median filtering has the capability of identifying and removing outliers. The advantages of the proposed algorithm are twofold. First, it uses a simple technique, i.e., median filtering, to capture the level of geometric and structural distortions (up to some extend). Second, the proposed algorithm has higher computational efficiency. Experiments show the superiority of the proposed NR IQA algorithm over existing state-of-the-art full-, reduced-, and NR IQA methods, in terms of both predicting accuracy and computational complexity.
Vinit Jakhetiya, Ke Gu 0001, Trisha Singhal, Sharath Chandra Guntuku, Zhifang Xia, Weisi Lin
IEEE Trans. Ind. Informatics2
2019 Reference-Free Quality Assessment of Sonar Images via Contour Degradation Measurement
abstract
Sonar imagery plays a significant role in oceanic applications since there is little natural light underwater, and light is irrelevant to sonar imaging. Sonar images are very likely to be affected by various distortions during the process of transmission via the underwater acoustic channel for further analysis. At the receiving end, the reference image is unavailable due to the complex and changing underwater environment and our unfamiliarity with it. To the best of our knowledge, one of the important usages of sonar images is target recognition on the basis of contour information. The contour degradation degree for a sonar image is relevant to the distortions contained in it. To this end, we developed a new no-reference contour degradation measurement for perceiving the quality of sonar images. The sparsities of a series of transform coefficient matrices, which are descriptive of contour information, are first extracted as features from the frequency and spatial domains. The contour degradation degree for a sonar image is then measured by calculating the ratios of extracted features before and after filtering this sonar image. Finally, a bootstrap aggregating (bagging)-based support vector regression module is learned to capture the relationship between the contour degradation degree and the sonar image quality. The results of experiments validate that the proposed metric is competitive with the state-of-the-art reference-based quality metrics and outperforms the latest reference-free competitors.
Ke Gu 0001, Weisi Lin, Zhifang Xia, Patrick Le Callet, En Cheng
IEEE Trans. Image Process.2
2019 Combining Local and Global Measures for DIBR-Synthesized Image Quality Evaluation
abstract
Depth-Image-Based-Rendering (DIBR) techniques are significant for three-dimensional (3D) video applications, e.g., 3D television and free viewpoint video (FVV). Unfortunately, the DIBR-synthesized image suffers from various distortions, which induce an annoying viewing experience for the entire FVV. Proposing a quality evaluator for DIBR-synthesized images is fundamental for the design of perceptual friendly FVV systems. Since the associated reference image is usually not accessible, full-reference (FR) methods cannot be directly applied for quality evaluation of the synthesized image. In addition, most traditional no-reference (NR) methods fail to effectively measure the specifically DIBR-related distortions. In this paper, we propose a novel NR quality evaluation method accounting for two categories of DIBR-related distortions, i.e., geometric distortions and sharpness. First, the disoccluded regions, as one of the most obvious geometric distortions, are captured by analyzing local similarity. Then, another typical geometric distortion (i.e., stretching) is detected and measured by calculating the similarity between it and its equal-size adjacent region. Second, considering the property of scale invariance, the global sharpness is measured as the distance between the distorted image and its downsampled version. Finally, the perceptual quality is estimated by linearly pooling the scores of two geometric distortions and sharpness together. Experimental results verify the superiority of the proposed method over the prevailing FR and NR metrics. More specifically, it is superior to all competing methods except APT in terms of effectiveness, but greatly outmatches APT in terms of implementation time.
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Guangtao Zhai
IEEE Trans. Image Process.3
2019 Objective Quality Evaluation of Dehazed Images
abstract
Vision-based intelligent systems like automatic driving or driving assistance can be improved by enhancing the visibility of the scenes captured in bad weather conditions. In particular, many image dehazing algorithms (DHAs) have been proposed to facilitate such applications in hazy weather. Contrary to the substantial progress of DHA developing, the quality evaluation of DHAs falls behind. Generally, DHAs can be evaluated qualitatively by human subjects or quantitatively by objective quality measures. Compared with the subjective evaluation which is time consuming and difficult to apply, objective measures with quantitative results are more needed in practical systems. But in the literature, very few measures are widely utilized, and even less measures correlate well with the overall dehazing quality (DHQ). In this paper, we study the DHQ evaluation using real hazy images systematically. We first construct a DHQ database, which is the largest of its kind so far and includes 1750 dehazed images generated from 250 real hazy images of various haze densities using seven representative DHAs. A subjective quality evaluation study is subsequently conducted on the DHQ database. Then, we propose an objective DHQ index (DHQI) by extracting and fusing three groups of features, including: 1) haze-removing features; 2) structure-preserving features; and 3) over-enhancement features, which have captured the most key aspects of dehazing. DHQI can be utilized to evaluate DHAs or optimize practical dehazing systems. Validations on the constructed DHQ database and three other databases with synthetic haze have verified the effectiveness of DHQI. Finally, we give an overview of the current DHA quality evaluation strategies, discuss their merits and demerits, and give some suggestions on systematic DHA quality evaluation. The DHQ database and the code of DHQI will be released to facilitate further research.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Xiaokang Yang 0001, Xin-Ping Guan
IEEE Trans. Intell. Transp. Syst.3
2019 Blind Quality Assessment of Camera Images Based on Low-Level and High-Level Statistical Features
abstract
Camera images in reality are easily affected by various distortions, such as blur, noise, blockiness, and the like, which damage the quality of images. The complexity of distortions in camera images raises significant challenge for precisely predicting their perceptual quality. In this paper, we present an image quality assessment (IQA) approach that aims to solve this challenging problem to some extent. In the proposed method, we first extract the low-level and high-level statistical features, which can capture the quality degradations effectively. On the one hand, the first kind of statistical features are extracted from the locally mean subtracted and contrast normalized coefficients, which denote the low-level features in the early human vision. On the other hand, the recently proposed brain theory and neuroscience, especially the free-energy principle, reveal that the human brain tries to explain its encountered visual scenes through an inner creative model, with which the brain can produce the projection for the image. Then, the quality of perceptions can be reflected by the divergence between the image and its brain projection. Based on this, we extract the second type of features from the brain perception mechanism, which represent the high-level features. The low-level and high-level statistical features can play a complementary role in quality prediction. After feature extraction, we design a neural network to integrate all the features and convert them to the final quality score. Extensive tests performed on two real camera image datasets prove the validity of our method and its advantageous predicting ability over the competitive IQA models.
Yutao Liu 0002, Ke Gu 0001, Shiqi Wang 0001, Debin Zhao, Wen Gao 0001
IEEE Trans. Multim.2
2019 Quality Evaluation of Image Dehazing Methods Using Synthetic Hazy Images
abstract
To enhance the visibility and usability of images captured in hazy conditions, many image dehazing algorithms (DHAs) have been proposed. With so many image DHAs, there is a need to evaluate and compare these DHAs. Due to the lack of the reference haze-free images, DHAs are generally evaluated qualitatively using real hazy images. But it is possible to perform quantitative evaluation using synthetic hazy images since the reference haze-free images are available and full-reference (FR) image quality assessment (IQA) measures can be utilized. In this paper, we follow this strategy and study DHA evaluation using synthetic hazy images systematically. We first build a synthetic haze removing quality (SHRQ) database. It consists of two subsets: regular and aerial image subsets, which include 360 and 240 dehazed images created from 45 and 30 synthetic hazy images using 8 DHAs, respectively. Since aerial imaging is an important application area of dehazing, we create an aerial image subset specifically. We then carry out subjective quality evaluation study on these two subsets. We observe that taking DHA evaluation as an exact FR IQA process is questionable, and the state-of-the-art FR IQA measures are not effective for DHA evaluation. Thus, we propose a DHA quality evaluation method by integrating some dehazing-relevant features, including image structure recovering, color rendition, and over-enhancement of low-contrast areas. The proposed method works for both types of images, but we further improve it for aerial images by incorporating its specific characteristics. Experimental results on two subsets of the SHRQ database validate the effectiveness of the proposed measures.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yucheng Zhu, Jiantao Zhou 0001, Guodong Guo, Xiaokang Yang 0001, Xin-Ping Guan, Wenjun Zhang 0001
IEEE Trans. Multim.3
2019 No-Reference Quality Evaluator of Transparently Encrypted Images
abstract
In past years, various encrypted algorithms have been proposed to fully or partially protect the multimedia content in view of practical applications. In the context of digital TV broadcasting, transparent encryption only protects partial content and fulfills both security and quality requirements. To date, only a few reference-based works have been reported to evaluate the quality of transparently encrypted images. However, these works are incapable of reference-unavailable conditions. In this paper, we conduct the first attempt that proposes a novel quality evaluator in the absence of reference images. The key strategy of the proposed metric lies in extracting features by considering the motivation of transparently encrypted images. Specifically, given that encrypted images prevent content from being easily recognized, several features, including correlation coefficient, information entropy, and intensity statistic, are preliminarily extracted to estimate visual recognizability. Meanwhile, considering that encrypted images are avoided since they are of extremely low quality, we also capture many features to measure the distortions on multiple quality-sensitive image attributes, such as naturalness, structure, and texture. Finally, the quality evaluator is built by bridging all extracted features and corresponding quality scores via a regression module. Experimental results demonstrate that the proposed method is superior to the mainstream no-reference quality evaluation methods designed for synthetically distorted images and possesses a close approximation to state-of-the-art reference-based methods designed for encrypted images.
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Hantao Liu
IEEE Trans. Multim.3
2018 A Blind Quality Measure for Industrial 2D Matrix Symbols Using Shallow Convolutional Neural Network
abstract
Industrial two-dimensional (2D) matrix symbols are ubiquitous throughout the automatic assembly lines. Most industrial 2D symbols are corrupted by various inevitable artifacts. State-of-the-art decoding algorithms are not able to directly handle low-quality symbols irrespective of problematic artifacts. Degraded symbols require appropriate preprocessing methods, such as morphology filtering, median filtering, or sharpening filtering, according to specific distortion type. In this paper, we first establish a database including 3000 industrial 2D symbols which are degraded by 6 types of distortions. Second, we utilize a shallow convolutional neural network (CNN) to identify the distortion type and estimate the quality grade for 2D symbols. Finally, we recommend an appropriate preprocessing method for low-quality symbol according to its distortion type and quality grade. Experimental results indicate that the proposed method outperforms state-of-the-art methods in terms of PLCC, SRCC and RMSE. It also promotes decoding efficiency at the cost of low extra time spent.
Zhaohui Che, Guangtao Zhai, Jing Liu 0002, Ke Gu 0001, Patrick Le Callet, Jiantao Zhou 0001, Xianming Liu 0005
ICIP4
2018 Adaptive Screen Content Image Enhancement Strategy using Layer-based Segmentation
abstract
The ubiquitous screen content images (SCIs) play a significant role in various scenarios currently. However, most SCIs captured by consumer devices are frequently corrupted with distortions, especially contrast distortion. Unlike the natural images, SCIs are composed of text, graphics and natural scene pictures so that traditional image enhancement methods are not suitable for these compound images. Therefore, we innovatively proposed an adaptive strategy for enhancing SCIs in this paper. Firstly, we devised a segmentation method to divide SCI into text and pictorial regions. Next, the famous guided image filter (GIF) with big and small kernel sizes served as unsharpness masking for processing different regions adaptively. For verifying performance, the proposed method was tested on recently prevalent SCI datasets including SIQAD, and Webpage Dataset. Experimental results indicate that the proposed approach outperforms state-of-the-art methods in most SCIs with flat background.
Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet, Xianming Liu 0005, Deming Zhai, Xiao Gu 0001
ISCAS3
2018 A Large-Scale Compressed 360-Degree Spherical Image Database: From Subjective Quality Evaluation to Objective Model Comparison
abstract
360-degree images/videos have been dramatically increasing in recent years. But the high resolution makes it difficult to be transported, compressed and stored, and thus constrains the development of 360-degree images/videos. Therefore, it is important to study how popular coding technologies influence the quality of 360-degree images. In this paper, we present a study on subjective assessment of compressed 360-degree images and investigate whether existing objective image quality assessment (IQA) methods can effectively evaluate the quality of compressed 360-degree images. We first construct the largest compressed 360-degree image database (CVIQD2018) including 16 source images and 528 compressed ones with three prevailing coding technologies. Then, we implement 16 full reference (FR) IQA metrics, which include 10 traditional IQA metrics for 2D images and 3 PSNR-based metrics for 360-degree images, as well as 5 no reference (NR) IQA metrics and calculate the correlation between each above metric and subjective assessment in terms of three commonly used performance indices. The experiment results reveal structure information, visual saliency information and compensation for geometric distortion are crucial for evaluating the quality of compressed 360-degree images.
Wei Sun 0029, Ke Gu 0001, Siwei Ma 0001, Wenhan Zhu, Guangtao Zhai
MMSP2
2018 Referenceless quality metric of multiply-distorted images based on structural degradation
Tao Dai 0001, Ke Gu 0001, Li Niu 0002, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia
Neurocomputing2
2018 Just Noticeable Difference for natural images using RMS contrast and feed-back mechanism
Vinit Jakhetiya, Weisi Lin, Sunil Prasad Jaiswal, Ke Gu 0001, Sharath Chandra Guntuku
Neurocomputing4
2018 Training-free referenceless camera image blur assessment via hypercomplex singular value decomposition
Lijuan Tang, Qiaohong Li, Leida Li, Ke Gu 0001, Jiansheng Qian
Multim. Tools Appl.4
2018 No-reference image quality assessment with center-surround based natural scene statistics
Jun Wu 0022, Zhaoqiang Xia, Huifang Li 0004, Kezheng Sun, Ke Gu 0001, Hong Lu 0008
Multim. Tools Appl.5
2018 Reduced-reference quality assessment of DIBR-synthesized images based on multi-scale edge intensity similarity
Yu Zhou 0009, Leida Li, Ke Gu 0001, Lijuan Tang
Multim. Tools Appl.4
2018 Saliency-induced reduced-reference quality index for natural scene and screen content images
Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Menghan Hu, Xiaokang Yang 0001
Signal Process.2
2018 Reduced-Reference Quality Assessment of Screen Content Images
abstract
The screen content images (SCIs) quality influences the user experience and the interactive performance of remote computing systems. With numerous approaches proposed to evaluate the quality of natural images, much less work has been dedicated to reduced-reference image quality assessment (RR-IQA) of SCIs. Here, we propose an RR-IQA method from the perspective of SCI visual perception. In particular, the quality of the distorted SCI is evaluated by comparing a set of extracted statistical features that consider both primary visual information and unpredictable uncertainty. A unique property that differentiates the proposed method from previous RR-IQA methods for natural images is the consideration of behaviors when human subjects view the screen content, which motivates us to establish the perceptual model according to the distinct properties of SCIs. Validations based on the screen content IQA database show that the proposed algorithm provides accurate predictions across a wide range of SCI distortions with negligible transmission overhead.
Shiqi Wang 0001, Ke Gu 0001, Xinfeng Zhang 0001, Weisi Lin, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Recurrent Air Quality Predictor Based on Meteorology- and Pollution-Related Factors
abstract
Air quality is currently arousing drastically increasing attention from the governments and populace all over the world. In this paper, we propose a heuristic recurrent air quality predictor (RAQP) to infer air quality. The RAQP exploits some key meteorology- and pollution-related variables to infer air pollutant concentrations (APCs), e.g. the fine particulate matter (PM2.5). It is natural that the meteorological factors and APCs at the current time have strong influences on air quality the next adjacent moment, that is to say, there exist high correlations between them. With this consideration, applying simple machine learners to the current meteorology- and pollution-related factors can reliably predict the air quality indices at a time later. However, owing to the nonlinear and chaotic reasons, the above correlations decline with the time interval enlarged. In such cases, it fails to forecast the air quality after several hours by only using simple machine learners and the current measurements of meteorology- and pollution-related variables. To solve the problem, our RAQP method recurrently applies the 1-h prediction model, which learns the current records of meteorology- and pollution-related factors to predict the air quality 1 h later, to then estimate the air quality after several hours. Via extensive experiments, results confirm that the RAQP predictor is superior to the relevant state-of-the-art techniques and nonrecurrent methods when applied to air quality prediction.
Ke Gu 0001, Junfei Qiao 0001, Weisi Lin
IEEE Trans. Ind. Informatics1
2018 A Prediction Backed Model for Quality Assessment of Screen Content and 3-D Synthesized Images
abstract
In this paper, we address problems associated with free-energy-principle-based image quality assessment (IQA) algorithms for objectively assessing the quality of Screen Content (SC) and three-dimensional (3-D) synthesized images and also propose a very fast and efficient IQA algorithm to address these issues. These algorithms separate an image into predicted and disorder residual parts and assume disorder residual part does not contribute much to the overall perceptual quality. These algorithms fail for quality estimation of SC images as information of textual regions in SC images are largely separated into the disorder residual part and less information in the predicted part and subsequently, given a negligible emphasis. However, this is in contrast with the characteristics of human vision. Since our eyes are well trained to detect text in daily life. So, our human vision has prior information about text regions and can sense small distortions in these regions. In this paper, we proposed a new reduced-reference IQA algorithm for SC images based upon a more perceptually relevant prediction model and distortion categorization, which overcomes problems with existing free-energy-principle-based predictors. From experiments, it is validated that the proposed model has a better capability of efficiently estimating the quality of SC images as compared to the recently developed reduced-reference IQA algorithms. We also applied the proposed algorithm to judge the quality of 3-D synthesized images and observed that it even achieves better performance than the full-reference IQA metrics specifically designed for the 3-D synthesized views.
Vinit Jakhetiya, Ke Gu 0001, Weisi Lin, Qiaohong Li, Sunil Prasad Jaiswal
IEEE Trans. Ind. Informatics2
2018 Model-Based Referenceless Quality Metric of 3D Synthesized Images Using Local Image Description
abstract
New challenges have been brought out along with the emerging of 3D-related technologies, such as virtual reality, augmented reality (AR), and mixed reality. Free viewpoint video (FVV), due to its applications in remote surveillance, remote education, and so on, based on the flexible selection of direction and viewpoint, has been perceived as the development direction of next-generation video technologies and has drawn a wide range of researchers' attention. Since FVV images are synthesized via a depth image-based rendering (DIBR) procedure in the "blind" environment (without reference images), a reliable real-time blind quality evaluation and monitoring system is urgently required. But existing assessment metrics do not render human judgments faithfully mainly because geometric distortions are generated by DIBR. To this end, this paper proposes a novel referenceless quality metric of DIBR-synthesized images using the autoregression (AR)-based local image description. It was found that, after the AR prediction, the reconstructed error between a DIBR-synthesized image and its AR-predicted image can accurately capture the geometry distortion. The visual saliency is then leveraged to modify the proposed blind quality metric to a sizable margin. Experiments validate the superiority of our no-reference quality method as compared with prevailing full-, reduced-, and no-reference models.
Ke Gu 0001, Vinit Jakhetiya, Junfei Qiao 0001, Xiaoli Li 0011, Weisi Lin, Daniel Thalmann
IEEE Trans. Image Process.1
2018 Optimizing Multistage Discriminative Dictionaries for Blind Image Quality Assessment
abstract
State-of-the-art algorithms for blind image quality assessment (BIQA) typically have two categories. The first category approaches extract natural scene statistics (NSS) as features based on the statistical regularity of natural images. The second category approaches extract features by feature encoding with respect to a learned codebook. However, several problems need to be addressed in existing codebook-based BIQA methods. First, the high-dimensional codebook-based features are memory-consuming and have the risk of over-fitting. Second, there is a semantic gap between the constructed codebook by unsupervised learning and image quality. To address these problems, we propose a novel codebook-based BIQA method by optimizing multistage discriminative dictionaries (MSDDs). To be specific, MSDDs are learned by performing the label consistent K-SVD (LC-KSVD) algorithm in a stage-by-stage manner. For each stage, a new quality consistency constraint called “quality-discriminative regularization” term is introduced and incorporated into the reconstruction error term to form a unified objective function, which can be effectively solved by LC-KSVD for discriminative dictionary learning. Then, the latter stage takes the reconstruction residual data in the former stage as input based on which LC-KSVD is repeatedly performed until the final stage is reached. Once the MSDDs are learned, multistage feature encoding is performed to extract feature codes. Finally, the feature codes are concatenated across all stages and aggregated over the entire image for quality prediction via regression. The proposed method has been evaluated on five databases and experimental results well confirm its superiority over existing relevant BIQA methods.
Qiuping Jiang, Feng Shao 0001, Weisi Lin, Ke Gu 0001, Gangyi Jiang, Huifang Sun
IEEE Trans. Multim.4
2018 Quality Assessment of DIBR-Synthesized Images by Measuring Local Geometric Distortions and Global Sharpness
abstract
Depth-image-based rendering (DIBR) is a fundamental technique in free viewpoint video, which is widely adopted to synthesize virtual viewpoints. The warping and rendering operations in DIBR generally introduce geometric distortions and sharpness change. The state-of-the-art quality indices are limited in dealing with such images since they are sensitive to geometric changes. In this paper, a new quality model for DIBR-synthesized view images is presented by measuring LOcal Geometric distortions in disoccluded regions and global Sharpness (LOGS). A disoccluded region detection method is first proposed using SIFT-flow-based warping. Then, the sizes and distortion strength of local disoccluded regions are combined to generate a score. Furthermore, a reblurring-based strategy is proposed to quantify the global sharpness. Finally, the overall quality score is calculated by pooling the scores of local disoccluded regions and global sharpness. Experiments on four public DIBR-synthesized image/video databases show the superiority of the proposed metric over the state-of-the-art quality models. The proposed method is further adopted for boosting the performances of existing quality metrics and benchmarking DIBR algorithms, both achieving very promising results.
Leida Li, Yu Zhou 0009, Ke Gu 0001, Weisi Lin, Shiqi Wang 0001
IEEE Trans. Multim.3
2018 Reduced-Reference Image Quality Assessment in Free-Energy Principle and Sparse Representation
abstract
The free-energy principle in recent studies of brain theory and neuroscience models the perception and understanding of the outside scene as an active inference process, in which the brain tries to account for the visual scene with an internal generative model. Specifically, with the internal generative model, the brain yields corresponding predictions for its encountered visual scenes. Then, the discrepancy between the visual input and its brain prediction should be closely related to the quality of perceptions. On the other hand, sparse representation has been evidenced to resemble the strategy of the primary visual cortex in the brain for representing natural images. With the strong neurobiological support for sparse representation, in this paper, we approximate the internal generative model with sparse representation and propose an image quality metric accordingly, which is named FSI (free-energy principle and sparse representation-based index for image quality assessment). In FSI, the reference and distorted images are, respectively, predicted by the sparse representation at first. Then, the difference between the entropies of the prediction discrepancies is defined to measure the image quality. Experimental results on four large-scale image databases confirm the effectiveness of the FSI and its superiority over representative image quality assessment methods. The FSI belongs to reduced-reference methods, and it only needs a single number from the reference image for quality estimation.
Yutao Liu 0002, Guangtao Zhai, Ke Gu 0001, Xianming Liu 0005, Debin Zhao, Wen Gao 0001
IEEE Trans. Multim.3
2018 Blind Quality Assessment Based on Pseudo-Reference Image
abstract
Traditional full-reference image quality assessment (IQA) metrics generally predict the quality of the distorted image by measuring its deviation from a perfect quality image called reference image. When the reference image is not fully available, the reduced-reference and no-reference IQA metrics may still be able to derive some characteristics of the perfect quality images, and then measure the distorted image's deviation from these characteristics. In this paper, contrary to the conventional IQA metrics, we utilize a new “reference” called pseudo-reference image (PRI) and a PRI-based blind IQA (BIQA) framework. Different from a traditional reference image, which is assumed to have a perfect quality, PRI is generated from the distorted image and is assumed to suffer from the severest distortion for a given application. Based on the PRI-based BIQA framework, we develop distortion-specific metrics to estimate blockiness, sharpness, and noisiness. The PRI-based metrics calculate the similarity between the distorted image's and the PRI's structures. An image suffering from severer distortion has a higher degree of similarity with the corresponding PRI. Through a two-stage quality regression after a distortion identification framework, we then integrate the PRI-based distortion-specific metrics into a general-purpose BIQA method named blind PRI-based (BPRI) metric. The BPRI metric is opinion-unaware (OU) and almost training-free except for the distortion identification process. Comparative studies on five large IQA databases show that the proposed BPRI model is comparable to the state-of-the-art opinion-aware- and OU-BIQA models. Furthermore, BPRI not only performs well on natural scene images, but also is applicable to screen content images. The MATLAB source code of BPRI and other PRI-based distortion-specific metrics will be publicly available.
Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Jing Liu 0002, Xiaokang Yang 0001, Chang Wen Chen
IEEE Trans. Multim.2
2018 Analysis of Structural Characteristics for Quality Assessment of Multiply Distorted Images
abstract
Perceptual image quality assessment (IQA) plays an important role in numerous applications, including image restoration, compression, enhancement, and others. Although many works have been conducted on individually distorted IQA problems and have achieved encouraging results, few studies have been conducted on multiple distorted (MD) IQA problems. Thus, limited progress has been made. In this paper, we propose a novel no reference image quality assessment (NR-IQA) method, named improved multiscale local binary pattern (IMLBP), for addressing multiply distorted IQA problems. The image structures are sensitive to image distortions, which motivates us to utilize the structural characteristics for overall image quality prediction. We improved the local binary pattern (LBP) by considering the human visual mechanism to better extract the structural information. The IMLBP contains two parts, the LBP and the radius difference LBP (DLBP). The DLBP reflects the values' changes in the radial direction. Specifically, when the radius value is small, the proposed descriptor is computed to represent microstructural information. Conversely, it represents macrostructural information when the radius becomes large. Moreover, to better mimick the human visual mechanism, the IMLBP is computed with the multiscale strategy and the operation is based on a patch unit whose size is proportional to the radius value. The frequency histogram of feature maps is transformed to feature vectors. Subsequently, a predictable function trained by the support vector regression is used to infer the overall quality score. Experimental results show that the proposed method outperforms most state-of-the-art IQA metrics on publicly available multiply distorted image databases.
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Nam Ling, Beichen Li 0002
IEEE Trans. Multim.3
2018 Blind Quality Index for Multiply Distorted Images Using Biorder Structure Degradation and Nonlocal Statistics
abstract
In the past decade, extensive image quality metrics have been proposed. The majority of them are tailored for the images that contain a specific type of distortion. However, in practice, the images are usually degraded by different types of distortions simultaneously. This poses great challenges to the existing quality metrics. Motivated by this, this paper proposes a no-reference quality index for the multiply distorted images using the biorder structure degradation and the nonlocal statistics. The design philosophy is inspired by the fact that the human visual system (HVS) is highly sensitive to the degradations of both the spatial contrast and the spatial distribution, which are prone to be changed by the joint effects of the multiple distortions. Specifically, the multiresolution representation of the image is first built by downsampling to simulate the hierarchical property of the HVS. Then, the structure degradation is calculated to measure the spatial contrast. Considering the fact that the human visual cortex has the separate mechanisms to perceive the first- and second-order structures, dubbed biorder structures, the degradations of biorder structures are calculated to account for the spatial contrast, producing the first group of the quality-aware features. Furthermore, the nonlocal self-similarity statistics is calculated to measure the spatial distribution, producing the second group of features. Finally, all the features are fed into the random forest regression model to learn the quality model for the multiply distorted images. Extensive experimental results conducted on the three public databases demonstrate the superiority of the proposed metric to the state-of-the-art metrics. Moreover, the proposed metric is also advantageous over the existing metrics in terms of the generalization ability.
Yu Zhou 0009, Leida Li, Jinjian Wu, Ke Gu 0001, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.4
2018 Learning a No-Reference Quality Assessment Model of Enhanced Images With Big Data
abstract
In this paper, we investigate into the problem of image quality assessment (IQA) and enhancement via machine learning. This issue has long attracted a wide range of attention in computational intelligence and image processing communities, since, for many practical applications, e.g., object detection and recognition, raw images are usually needed to be appropriately enhanced to raise the visual quality (e.g., visibility and contrast). In fact, proper enhancement can noticeably improve the quality of input images, even better than originally captured images, which are generally thought to be of the best quality. In this paper, we present two most important contributions. The first contribution is to develop a new no-reference (NR) IQA model. Given an image, our quality measure first extracts 17 features through analysis of contrast, sharpness, brightness and more, and then yields a measure of visual quality using a regression module, which is learned with big-data training samples that are much bigger than the size of relevant image data sets. The results of experiments on nine data sets validate the superiority and efficiency of our blind metric compared with typical state-of-the-art full-reference, reduced-reference and NA IQA methods. The second contribution is that a robust image enhancement framework is established based on quality optimization. For an input image, by the guidance of the proposed NR-IQA measure, we conduct histogram modification to successively rectify image brightness and contrast to a proper level. Thorough tests demonstrate that our framework can well enhance natural images, low-contrast images, low-light images, and dehazed images. The source code will be released at https://sites.google.com/site/guke198701/publications.
Ke Gu 0001, Dacheng Tao, Junfei Qiao 0001, Weisi Lin
IEEE Trans. Neural Networks Learn. Syst.1
2018 Evaluating Quality of Screen Content Images Via Structural Variation Analysis
abstract
With the quick development and popularity of computers, computer-generated signals have drastically invaded into our daily lives. Screen content image is a typical example, since it also includes graphic and textual images as components as compared with natural scene images which have been deeply explored, and thus screen content image has posed novel challenges to current researches, such as compression, transmission, display, quality assessment, and more. In this paper, we focus our attention on evaluating the quality of screen content images based on the analysis of structural variation, which is caused by compression, transmission, and more. We classify structures into global and local structures, which correspond to basic and detailed perceptions of humans, respectively. The characteristics of graphic and textual images, e.g., limited color variations, and the human visual system are taken into consideration. Based on these concerns, we systematically combine the measurements of variations in the above-stated two types of structures to yield the final quality estimation of screen content images. Thorough experiments are conducted on three screen content image quality databases, in which the images are corrupted during capturing, compression, transmission, etc. Results demonstrate the superiority of our proposed quality model as compared with state-of-the-art relevant methods.
Ke Gu 0001, Junfei Qiao 0001, Xiongkuo Min, Guanghui Yue 0001, Weisi Lin, Daniel Thalmann
IEEE Trans. Vis. Comput. Graph.1
2017 Reduced-reference quality metric for screen content image
abstract
With the prevalence of digital products like cellphone, tablet and personal computer, the screen content image (SCI) consisting of text, graphic, and natural scene picture becomes a significant media in various communication scenarios. Consequently, we proposed a reduced-reference quality metric dedicated for SCI. The main contribution includes 2 aspects: 1) we innovatively proposed a layer-based segmentation method to divide SCI into text layer and pictorial layer; 2) we designed respective quality metrics dedicated for text and pictorial layers with a novel pooling strategy considering human visual saliency for SCI. Furthermore, exhaustive experimental results indicate that the proposed metric is highly comparative compared with state-of-the-art full-reference quality metrics.
Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet
ICIP3
2017 Foveated nonlocal dual denoising
abstract
Recently developed dual domain image denoising (DDID) algorithm and its variants, such as dual domain filter (DDF), achieve remarkable results by combining bilateral filter with frequency-based method. However, this kind of algorithms require large patches to guarantee the denoising performance and most of them produce ringing artifacts due to the Gibbs phenomenon induced by high-contrast details. To address these issues, we propose a Foveated Nonlocal Dual Denoising (FNDD) algorithm by unifying foveated nonlocal means and frequency-based methods. In this way, the ability to preserve the high-contrast details is noticeably improved by exploiting foveated self-similarity (patch similarity) instead of pixel similarity, thus leading to void of artifacts. Moreover, we propose an entropy-based back projection step for compensating the detail loss to further improve the performance. Experimental results validate that FNDD significantly outperforms DDID in terms of both quantitative metrics and subjective visual quality under much smaller patches, and even achieves comparable results against state-of-the-art competitors.
Tao Dai 0001, Ke Gu 0001, Qingtao Tang, Kwok-Wai Hung, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia
ICIP2
2017 Blind quality assessment of multiply-distorted images based on structural degradation
abstract
It is known that images available usually undergo some stages of processing (e.g., acquisition, compression, transmission and display), and each stage may introduce certain type of distortion. Hence, images distorted by multiple types of distortions are common in real applications. Research in human visual perception has evidenced that the human visual system (HVS) is sensitive to image structural information. This fact inspires us to design a new blind/no-reference (NR) image quality assessment (IQA) method to evaluate the visual quality of multiply-distorted images based on structural degradation. Specifically, quality-aware features are extracted from both the first- and high-order image structures by local binary pattern (LBP) operators. Experimental results on two well-known multiply-distorted image databases demonstrate the outstanding performance of the proposed method.
Tao Dai 0001, Ke Gu 0001, Zhiya Xu, Qingtao Tang, Haoyi Liang, Yongbing Zhang 0002, Shutao Xia
ICIP2
2017 Using multiscale analysis for blind quality assessment of DIBR-synthesized images
abstract
In this paper we propose to blindly evaluate the quality of images synthesized based on a depth image-based rendering (DIBR) procedure. As an important branch in virtual reality (VR), superior DIBR techniques provide free viewpoints in many real applications such as remote surveillance and education, but few efforts have been made to measure the performance of DIBR methods (i.e. the quality of DIBR-synthesized images), especially in the condition of reference unavailable. To this aim, we put forward a new no-reference (NR) image quality assessment (IQA) model via multiscale analysis, dubbed as MSA. The design philosophy of our proposed MSA model is that the DIBR-introduced geometry distortions damage the self-similarity characteristic of natural images and the damage degrees present regular variations at distinct scales. Through systematically incorporating the measurements of the variations provided above, our MSA model can faithfully predict the quality of images generated using different DIBR technologies. Results of experiments demonstrate that the proposed blind MSA model has delivered noticeably better performance than state-of-the-art full-and no-reference IQA methods.
Ke Gu 0001, Junfei Qiao 0001, Patrick Le Callet, Zhifang Xia, Weisi Lin
ICIP1
2017 CVIQD: Subjective quality evaluation of compressed virtual reality images
abstract
The 360-degree spherical images/videos, also called Virtual Reality (VR) images/videos, can provide immersive experience of the real-world scenes in some specific systems. This makes it widely employed in concerts/sports events live and VR movies. However, it is difficult to transport, compress or store VR images/videos due to their high resolution. So it is significant to research how the popular coding technologies influence the quality of VR images. To this aim, this paper carries out subjective quality evaluation of compressed VR images and examines the correlation performance of popular objective quality measures in accordance with the aforesaid subjective ratings. We first establish a Compressed VR Image Quality Database (CVIQD), which includes five source VR images and associated 165 compressed images under three prevailing coding technologies. The Single-Stimulus (SS) method is exploited to collect the subjective scores from 20 inexperienced viewers. Next, we implement 10 classical and recent objective quality metrics on the CVIQD database and compute the correlation between each above quality metric and subjective assessment in terms of five commonly used performance indices. Experimental results reveal that multi-scale based MS-SSIM and ADD-SSIM models have lead to high correlation with human visual perception.
Wei Sun 0029, Ke Gu 0001, Guangtao Zhai, Siwei Ma 0001, Weisi Lin, Patrick Le Callet
ICIP2
2017 Perceptual evaluation of single-image super-resolution reconstruction
abstract
In recent years, single-image super-resolution (SR) reconstruction has aroused wide attention. Massive SR enhancement algorithms have been proposed. However, much less work has been down on the perceptual evaluation of SR enhanced images and the corresponding enhancement algorithms. In this work, we create a Super-resolution Reconstructed Image Database (SRID), which consists of images produced by two interpolation methods and six popular SR image enhancement algorithms at different amplification factors. Then, subjective experiment is conducted to collect the subjective scores by using the single-stimulus method. The performances of the SR image enhancement algorithms are then evaluated by the obtained subjective scores. Finally, the performances of the general-purpose no-reference (NR) image quality metrics are investigated on the SRID database. This study shows that it is difficult for the state-of-the-art NR image quality metrics to predict the quality of SR enhanced images.
Guangcheng Wang, Leida Li, Qiaohong Li, Ke Gu 0001, Zhaolin Lu, Jiansheng Qian
ICIP4
2017 No-reference quality assessment for JPEG compressed images
abstract
JPEG is a most commonly used standard of compression for digital images. Quality factor (Qfactor) for JPEG compressed image is actually a suitable indicator to the perceptual quality. However, the information of the compressor might be unknown due to various reasons. To evaluate the Qfactor, we recompress the formerly compressed image and measure the consistency between them. Then we define the fixed points (the points on the Qfactor-axis where the content of recompressed images are almost the same with that of directly compressed images) by following the Qfactor based specifications and form the image set. The quality of JPEG compressed images are measured by combining the estimated Qfactor with the features extracted from the image set. The experimental results confirm that the proposed image quality assessment technique, which is no-reference, is able to faithfully predict the visual quality of JPEG compressed images.
Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Wenhan Zhu
QoMEX3
2017 Portable information security display system via Spatial PsychoVisual Modulation
abstract
With the rapid development of visual media, people prefer to pay more attention to privacy protection in public situations. Currently, most existing researches on information security such as cryptography and steganography mainly concern transmission and yet little has been done to keep the information displayed on screens from reaching eyes of the bystanders. At the same time, the reporter just stands in front of the screen during traditional meetings. Limited time, screen area and report forms, which inevitably leads to limited information. As a result, we design a portable screen for assisting the reporter to present any information in a new form. In some public occasions, the reporter can show private content with important information to authorized audience while the others can not see that. In this paper, we propose a new Spatial PsychoVisual Modulation (SPVM) based solution to the privacy problem. This system uses two synchronized projectors with linear polarization filters and polarization glasses, and a camera with linear polarization filter the metallic screen. It can guarantee the system shows private information synchronously. We have implemented the system and experimental results demonstrate the effectiveness and robustness of the proposed information security display system.
Guangtao Zhai, Jia Wang 0004, Ke Gu 0001
VCIP4
2017 Internal generative mechanism inspired reduced reference image quality assessment with entropy of primitive
abstract
In this paper, we propose a novel reduced-reference (RR) image quality assessment (IQA) algorithm based on the internal generative mechanism, which suggests that the human visual system (HVS) can actively predict the primary visual information and avoid the uncertainty. Specifically, the explanation of the visual scene is formulated as the process of sparse representation. In particular, the entropy of primitive accounts for the primary visual information and the discrepancy between the image signal and its best sparse description is regarded as the uncertainty in perception. As such, the combined feature that can summarize the primary visual information and uncertainty in sparse domain is required to be transmitted in the RR-IQA framework. Comparative studies of the proposed reduced reference metric is conduced on both single and multiple distortion databases, and experimental results demonstrate that the proposed metric can achieve high correlation with the human perception by only sending ignorable additional information.
Shanshe Wang, Shiqi Wang 0001, Ke Gu 0001, Siwei Ma 0001, Wen Gao 0001
VCIP3
2017 Subjective quality assessment of animation images
abstract
In the past few decades, many attempts have been maken to evaluate the image quality assessment (IQA) of natural scene images. However, the IQA research of animation images (AIs) has been highly overlooked. In this article, we carry out in-depth study on perceptual quality assessment of AIs. As the lack of a public and diverse testing database currently, this paper builds a large-scale Animation Images Quality Assessment Database (AIQAD). This database totally includes 1050 distorted images derived from 30 source images by corrupting seven distortion types with multiple distortion levels. Then, a subjective experiment, which is the basic and accurate quality evaluation measurement, is conducted to obtain the mean opinion score (MOS) for each image. Furthermore, we also investigate the feasibility of utilizing existing mainstream full reference (FR) IQA metrics to solve the IQA problem of AIs. Experimental results demonstrate that existing mainstream FR IQA metrics merely achieve fair performance on the proposed database.
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001
VCIP3
2017 Visual attention analysis and prediction on human faces
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Jing Liu 0002, Shiqi Wang 0001, Xinfeng Zhang 0001, Xiaokang Yang 0001
Inf. Sci.3
2017 A generic denoising framework via guided principal component analysis
Tao Dai 0001, Zhiya Xu, Haoyi Liang, Ke Gu 0001, Qingtao Tang, Yisen Wang 0001, Weizhi Lu, Shutao Xia
J. Vis. Commun. Image Represent.4
2017 Quality assessment for real out-of-focus blurred images
Yutao Liu 0002, Ke Gu 0001, Guangtao Zhai, Xianming Liu 0005, Debin Zhao, Wen Gao 0001
J. Vis. Commun. Image Represent.2
2017 An efficient and effective blind camera image quality metric via modeling quaternion wavelet coefficients
Lijuan Tang, Leida Li, Kezheng Sun, Zhifang Xia, Ke Gu 0001, Jiansheng Qian
J. Vis. Commun. Image Represent.5
2017 No reference image blurriness assessment with local binary patterns
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Nam Ling
J. Vis. Commun. Image Represent.3
2017 Just-Noticeable Difference-Based Perceptual Optimization for JPEG Compression
abstract
The Quantization table in JPEG, which specifies the quantization scale for each discrete cosine transform (DCT) coefficient, plays an important role in image codec optimization. However, the generic quantization table design that is based on the characteristics of human visual system (HVS) cannot adapt to the variations of image content. In this letter, we propose a just-noticeable difference (JND) based quantization table derivation method for JPEG by optimizing the rate-distortion costs for all the frequency bands. To achieve better perceptual quality, the DCT domain JND-based distortion metric is utilized to model the stair distortion perceived by HVS. The rate-distortion cost for each band is derived by estimating the rate with the first-order entropy of quantized coefficients. Subsequently, the optimal quantization table is obtained by minimizing the total rate-distortion costs of all the bands. Extensive experimental results show that the quantization table generated by the proposed method achieves significant bit-rate savings compared with JPEG recommended quantization table and specifically developed quantization tables in terms of both objective and subjective evaluations.
Xinfeng Zhang 0001, Shiqi Wang 0001, Ke Gu 0001, Weisi Lin, Siwei Ma 0001, Wen Gao 0001
IEEE Signal Process. Lett.3
2017 No-Reference Quality Metric of Contrast-Distorted Images Based on Information Maximization
abstract
The general purpose of seeing a picture is to attain information as much as possible. With it, we in this paper devise a new no-reference/blind metric for image quality assessment (IQA) of contrast distortion. For local details, we first roughly remove predicted regions in an image since unpredicted remains are of much information. We then compute entropy of particular unpredicted areas of maximum information via visual saliency. From global perspective, we compare the image histogram with the uniformly distributed histogram of maximum information via the symmetric Kullback-Leibler divergence. The proposed blind IQA method generates an overall quality estimation of a contrast-distorted image by properly combining local and global considerations. Thorough experiments on five databases/subsets demonstrate the superiority of our training-free blind technique over state-of-the-art full- and no-reference IQA methods. Furthermore, the proposed model is also applied to amend the performance of general-purpose blind quality metrics to a sizable margin.
Ke Gu 0001, Weisi Lin, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Chang Wen Chen
IEEE Trans. Cybern.1
2017 No-Reference Quality Assessment of Screen Content Pictures
abstract
Recent years have witnessed a growing number of image and video centric applications on mobile, vehicular, and cloud platforms, involving a wide variety of digital screen content images. Unlike natural scene images captured with modern high fidelity cameras, screen content images are typically composed of fewer colors, simpler shapes, and a larger frequency of thin lines. In this paper, we develop a novel blind/no-reference (NR) model for accessing the perceptual quality of screen content pictures with big data learning. The new model extracts four types of features descriptive of the picture complexity, of screen content statistics, of global brightness quality, and of the sharpness of details. Comparative experiments verify the efficacy of the new model as compared with existing relevant blind picture quality assessment algorithms applied on screen content image databases. A regression module is trained on a considerable number of training samples labeled with objective visual quality predictions delivered by a high-performance full-reference method designed for screen content image quality assessment (IQA). This results in an opinion-unaware NR blind screen content IQA algorithm. Our proposed model delivers computational efficiency and promising performance. The source code of the new model will be available at: https://sites.google.com/site/guke198701/publications.
Ke Gu 0001, Jun Zhou 0007, Junfei Qiao 0001, Guangtao Zhai, Weisi Lin, Alan C. Bovik
IEEE Trans. Image Process.1
2017 Unified Blind Quality Assessment of Compressed Natural, Graphic, and Screen Content Images
abstract
Digital images in the real world are created by a variety of means and have diverse properties. A photographical natural scene image (NSI) may exhibit substantially different characteristics from a computer graphic image (CGI) or a screen content image (SCI). This casts major challenges to objective image quality assessment, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. To tackle this problem, we first construct a cross-content-type (CCT) database, which contains 1,320 distorted NSIs, CGIs, and SCIs, compressed using the high efficiency video coding (HEVC) intra coding method and the screen content compression (SCC) extension of HEVC. We then carry out a subjective experiment on the database in a well-controlled laboratory environment. Moreover, we propose a unified content-type adaptive (UCA) blind image quality assessment model that is applicable across content types. A key step in UCA is to incorporate the variations of human perceptual characteristics in viewing different content types through a multi-scale weighting framework. This leads to superior performance on the constructed CCT database. UCA is training-free, implying strong generalizability. To verify this, we test UCA on other databases containing JPEG, MPEG-2, H.264, and HEVC compressed images/videos, and observe that it consistently achieves competitive performance.
Xiongkuo Min, Kede Ma, Ke Gu 0001, Guangtao Zhai, Zhou Wang 0001, Weisi Lin
IEEE Trans. Image Process.3
2016 No-reference image quality assessment for photographic images of consumer device
abstract
In this paper we study common, camera-specific kinds of distortions and propose a no-reference image quality assessment algorithm for photographic images produced by consumer devices. Those real consumer-type images, being different from simulated-distortion images, are with realistic artifacts and quality ranges. We find that the state-of-the-art no-reference image quality assessment approaches do not perform well on those photographic images, and propose an approach that achieves high prediction performance on a dataset of consumer-centric images. The proposed method, with no need for the original image, is able to reveal camera-specific problems and differentiate consumer cameras.
Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Zhaohui Che
ICASSP3
2016 Visual saliency detection via image complexity feature
abstract
In this paper we propose a novel bottom-up visual saliency detection model by analysis of image complexity. Compared with existing works, we emphasize the important impact of image complexity on saliency detection. Inspired by the free energy theory, a hybrid parametric and non-parametric model is used to estimate the complexity of a visual signal. Taking the image complexity as a new feature, this paper constructs a heuristic framework to systematically combine two different types of saliency detection models, separately using local and global features, in order to predict human fixation points more accurately. In contrast to classical and modern models, our algorithm has achieved noticeably superior results. And furthermore, it is worthy to stress that the proposed saliency detection method can also help to facilitate the performance of image quality metrics on popular image databases.
Min Liu 0003, Ke Gu 0001, Guangtao Zhai, Patrick Le Callet
ICIP2
2016 Quality assessment of 3D synthesized images via disoccluded region discovery
abstract
Depth-Image-Based-Rendering (DIBR) is fundamental in free-viewpoint 3D video, which has been widely used to generate synthesized views from multi-view images. The majority of DIBR algorithms cause disoccluded regions, which are the areas invisible in original views but emerge in synthesized views. The quality of synthesized images is mainly contaminated by distortions in these disoccluded regions. Unfortunately, traditional image quality metrics are not effective for these synthesized images because they are sensitive to geometric distortions. To solve the problem, this paper proposes an objective quality evaluation method for 3D Synthesized images via Disoccluded Region Discovery (SDRD). A self-adaptive scale transform model is first adopted to preprocess the images on account of the impacts of view distance. Then disoccluded regions are detected by comparing the absolute difference between the preprocessed synthesized image and the warped image of preprocessed reference image. Furthermore, the disoccluded regions are weighted by a weighting function proposed to account for the varying sensitivities of human eyes to the size of disoccluded regions. Experiments conducted on IRCCyN/IVC DIBR image database demonstrate that the proposed SDRD method remarkably outperforms traditional 2D and existing DIBR-related quality metrics.
Yu Zhou 0009, Leida Li, Ke Gu 0001, Yuming Fang 0001, Weisi Lin
ICIP3
2016 Blind quality assessment of compressed images via pseudo structural similarity
abstract
Block-based compression causes severe pseudo structures. We find that the pseudo structures of images compressed by different levels show some degree of similarity. So we propose to evaluate the quality of compressed images via the similarity between pseudo structures of two images. To obtain a “reference” image, we introduce the most distorted image (MDI), which is derived from the distorted image and suffers from the highest degree of compression. The proposed pseudo structural similarity (PSS) model calculates the similarity between pseudo structures of the distorted image and MDI. Pseudo structures of the distorted image become similar to the MDI's under the condition of severe compression. Via comparative tests, the proposed PSS model, on one hand, is shown to be comparable to state-of-the-art competitors, and on the other hand, it is not only good at assessing natural scene images but also performs the best in the hotly-researched screen content image (SCI) database. It deserves to mention that PSS is able to boost the performance of mainstream general-purpose no-reference (NR) quality measures.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yuming Fang 0001, Xiaokang Yang 0001, Xiaolin Wu 0001, Jiantao Zhou 0001, Xianming Liu 0005
ICME3
2016 Quality assessment of contrast-altered images
abstract
In image / video systems, the contrast adjustment which manages to enhance the visual quality is nowadays an important research topic. Yet very limited efforts have been devoted to the exploration of image quality assessment (IQA) for contrast adjustment. To address the problem, this paper proposes a novel reduced-reference (RR) IQA metric with the integration of bottom-up and top-down strategies. The former one stems from the recently revealed free energy theory which tells that the human visual system always seeks to understand an input image by the uncertainty removal, while the latter one is towards using the symmetric K-L divergence to compare the histogram of the contrast-altered image with that of the reference image. The bottom-up and top-down strategies are lastly combined to derive the Reduced-reference Contrast-altered Image Quality Measure (RCIQM). A comparison with numerous existing IQA models is conducted on contrast related CID2013, CCID2014, CSIQ, TID2008 and TID2013 databases, and results validate the superiority of the proposed technique.
Min Liu 0003, Ke Gu 0001, Guangtao Zhai, Jiantao Zhou 0001, Weisi Lin
ISCAS2
2016 Blindly evaluating stereoscopic image quality with free-energy principle
abstract
Three-dimensional (3D) imaging technology has been growingly prevalent in today's world. But objective quality assessment of 3D images is a challenging task. In this paper, we propose a blind metric to predict the perceptual quality of stereopairs within the concept of free energy. On the basis of a psychological measure, the free energy is a principle telling where supervises more and attracts human attention. We believe that the “surprise” can account for the binocular rivalry and thus be used to predict the quality of stereopairs. We first evaluate the quality of the monoscopic image, then introduce the computation process of binocular rivalry's results for deciding the relative importance of the left and right views, and finally infer the overall quality score. Our algorithm is tested on the symmetric LIVE3D-I and asymmetric LIVE3D-II databases. Experimental results confirm that the proposed blind 3D IQA technique, without distortion identification, is able to faithfully predict the visual quality of stereopairs.
Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Min Liu 0003
ISCAS3
2016 Closing the gap: Visual quality assessment considering viewing conditions
abstract
Most of existing visual quality assessment algorithms are tested on standard databases that are created in controlled viewing conditions (e.g. display device, viewing distance and lighting). This implies that all the recoded subjective scores are only valid for the specific settings used in the database. However, with the prevalence of mobile devices, the practical viewing environments can significantly vary from moment to moment. It is our daily experience that the same image can look drastically different on dissimilar devices under changed viewing distance and/or lighting conditions. In other words, a gap exists between the eyes and the visual contents behind the screen in current research of quality assessment. Therefore, in this work, we perform subjective quality evaluation with varied actual viewing conditions. To make the research reproducible, we build a prototype system to record what the eyes really see from the screen and construct the viewing environment-changed image database. The database will be made available to the public. Meanwhile we design a dedicated effective environment-assessing algorithm. We believe that this work will benefit the research of visual quality assessment towards more practical applications.
Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Zhaohui Che
QoMEX3
2016 Quality assessment for dual-view display system
abstract
Spatial psychovisual modulation (SPVM) is a new information display technology, which aims to generate multiple visual percepts for different viewers on a single display simultaneously. After the proposal of SPVM, lots of efforts have been made and several applications (i.e., dual-view display system) have been implemented based on this technology. The dual-view display (DVD) system is considered as an effective digital image hiding system based on SPVM theory, but little work has been dedicated to the perceptual quality assessment of DVD system. Up to now, there is no clear and standard method to evaluate the performance of the dual-view display system. It is important for the viewers to see a clear and non-aliasing image when they are front of the screen. Therefore, in this paper, we will build a DVD database and carry out a subjective experiment to evaluate the performance of the DVD system, and then we investigate and analyze the performance of prevailing no-reference (NR) image quality metrics on the particular DVD system. We have a sufficient belief that this paper can supply the guideline for the performance on the DVD system and serve as a good testing bed for future research of SPVM technology.
Yuanchun Chen, Guangtao Zhai, Ke Gu 0001, Jia Wang 0004, Zhongpai Gao, Yucheng Zhu
VCIP4
2016 A reduced-reference quality assessment scheme for blurred images
abstract
In this paper, we propose a reduced-reference scheme for evaluating the quality of blurred images under the theory of free-energy principle. Specifically, the free-energy principle indicates that the brain tries to account for the input image with an internal generative model and the discrepancy between the image and its model-explained version, which can be measured by free energy, is related to the image's perceptual quality. Accordingly, we define a visual distance between the blurred image and its original image in free energy to evaluate the quality of the blurred image. Therefore, the proposed quality scheme belongs to reduced-reference methods, which needs some information from the original image for quality assessment. Experimental results on public databases, LIVE, TID2013 and C-SIQ, demonstrate the proposed method works in high consistency with subjective assessment results and outperforms representative image quality assessment approaches.
Zongxi Han, Guangtao Zhai, Yutao Liu 0002, Ke Gu 0001, Xinfeng Zhang 0001
VCIP4
2016 Exploiting neural models for no-reference image quality assessment
abstract
We propose an improved algorithm for no-reference image quality assessment (NR-IQA) using the convolutional neural network (CNN) and neural theory based saliency detection. Firstly, we extract non-overlapping patches from the input image. For each patch, we obtain the quality score by CNN network, which consists of seven layers and integrates feature learning and regression into image patch quality estimation. Considering that the patches attracting much attention take significant role in visual perception, an efficient technique based on free energy based neural model is used to detect the saliency map. This saliency map is then applied as a weighting mask to output the quality score of the whole image. Results of experiments show that our algorithm achieves state-of-the-art performance, as compared with the prevailing IQA methods.
Cenhui Pan, Yi Xu 0001, Yichao Yan, Ke Gu 0001, Xiaokang Yang 0001
VCIP4
2016 Transform-domain in-loop filter with block similarity for HEVC
abstract
In-loop filtering is an important technique in modern video coding standards. In this paper, we propose a transform-domain in-loop filter to further improve the compression performance of high efficiency video coding (HEVC) standard. The proposed method estimates block transform coefficients by adaptively fusing two prediction sources according to their uncertainties respectively. The first prediction is the block transform coefficients of compressed video frames, the uncertainty of which is related to quantization parameters. The second prediction is the weighted average of transform blocks in a neighborhood, and the weights are designed according to block similarity. Its uncertainty is estimated based on the coefficient variance. To optimize the filtering performance, the parameters utilized in the proposed in-loop filter are learned from compressed videos for each quantization parameter offline, and frame level flags are utilized to switch the proposed in-loop filter according to rate-distortion cost. Extensive experimental results show that the proposed in-loop filter can further improves the compression efficiency of HEVC.
Xinfeng Zhang 0001, Weisi Lin, Ke Gu 0001, Qiaohong Li, Shanshe Wang, Siwei Ma 0001
VCIP3
2016 Learning a blind quality evaluation engine of screen content images
Ke Gu 0001, Guangtao Zhai, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001
Neurocomputing1
2016 Color image quality assessment based on sparse representation and reconstruction residual
Leida Li, Wenhan Xia, Yuming Fang 0001, Ke Gu 0001, Jinjian Wu, Weisi Lin, Jiansheng Qian
J. Vis. Commun. Image Represent.4
2016 Blind quality index for camera images with natural scene statistics and patch-based sharpness assessment
Lijuan Tang, Leida Li, Ke Gu 0001, Xingming Sun, Jianying Zhang
J. Vis. Commun. Image Represent.3
2016 Joint Chroma Downsampling and Upsampling for Screen Content Image
abstract
Screen content images are originally captured in a full-chroma format. The chroma downsampling, which is commonly applied to the chroma component in screen content image representation and processing (e.g., YUV4:2:0 compression), will significantly degrade the image quality and create annoying artifacts such as blur and color shifting. To tackle this problem, in this paper we propose luma aware chroma downsampling and upsampling algorithms to jointly improve the quality of the chroma image reconstruction. Guided by the luma information, the chroma upsampling algorithm is proposed with the utilization of major color and index map representation. The geometric information-based linear mapping is developed to transfer the structure of luma to the interpolated chroma. Subsequently, the error sensitivity of the upsampling method is analyzed, and content dependent downsampling algorithm is presented to minimize the error sensitivity function. We further explore the applicability of the proposed scheme in the scenario of screen content compression, targeting at improving the decoded chroma image quality for display. Extensive experimental results demonstrate the viability and efficiency of the proposed scheme.
Shiqi Wang 0001, Ke Gu 0001, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 The Analysis of Image Contrast: From Quality Assessment to Automatic Enhancement
abstract
Proper contrast change can improve the perceptual quality of most images, but it has largely been overlooked in the current research of image quality assessment (IQA). To fill this void, we in this paper first report a new large dedicated contrast-changed image database (CCID2014), which includes 655 images and associated subjective ratings recorded from 22 inexperienced observers. We then present a novel reduced-reference image quality metric for contrast change (RIQMC) using phase congruency and statistics information of the image histogram. Validation of the proposed model is conducted on contrast related CCID2014, TID2008, CSIQ and TID2013 databases, and results justify the superiority and efficiency of RIQMC over a majority of classical and state-of-the-art IQA methods. Furthermore, we combine aforesaid subjective and objective assessments to derive the RIQMC based Optimal HIstogram Mapping (ROHIM) for automatic contrast enhancement, which is shown to outperform recently developed enhancement technologies.
Ke Gu 0001, Guangtao Zhai, Weisi Lin, Min Liu 0003
IEEE Trans. Cybern.1
2016 Saliency-Guided Quality Assessment of Screen Content Images
abstract
With the widespread adoption of multidevice communication, such as telecommuting, screen content images (SCIs) have become more closely and frequently related to our daily lives. For SCIs, the tasks of accurate visual quality assessment, high-efficiency compression, and suitable contrast enhancement have thus currently attracted increased attention. In particular, the quality evaluation of SCIs is important due to its good ability for instruction and optimization in various processing systems. Hence, in this paper, we develop a new objective metric for research on perceptual quality assessment of distorted SCIs. Compared to the classical MSE, our method, which mainly relies on simple convolution operators, first highlights the degradations in structures caused by different types of distortions and then detects salient areas where the distortions usually attract more attention. A comparison of our algorithm with the most popular and state-of-the-art quality measures is performed on two new SCI databases (SIQAD and SCD). Extensive results are provided to verify the superiority and efficiency of the proposed IQA technique.
Ke Gu 0001, Shiqi Wang 0001, Huan Yang 0001, Weisi Lin, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Multim.1
2016 Blind Quality Assessment of Tone-Mapped Images Via Analysis of Information, Naturalness, and Structure
abstract
High dynamic range (HDR) imaging techniques have been working constantly, actively, and validly in the fault detection and disease diagnosis in the astronomical and medical fields, and currently they have also gained much more attention from digital image processing and computer vision communities. While HDR imaging devices are starting to have friendly prices, HDR display devices are still out of reach of typical consumers. Due to the limited availability of HDR display devices, in most cases tone mapping operators (TMOs) are used to convert HDR images to standard low dynamic range (LDR) images for visualization. But existing TMOs cannot work effectively for all kinds of HDR images, with their performance largely depending on brightness, contrast, and structure properties of a scene. To accurately measure and compare the performance of distinct TMOs, in this paper develop an effective and efficient no-reference objective quality metric which can automatically assess LDR images created by different TMOs without access to the original HDR images. Our model is shown to be statistically superior to recent full- and no-reference quality measures on the existing tone-mapped image database and a new relevant database built in this work.
Ke Gu 0001, Shiqi Wang 0001, Guangtao Zhai, Siwei Ma 0001, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001, Wen Gao 0001
IEEE Trans. Multim.1
2016 Guided Image Contrast Enhancement Based on Retrieved Images in Cloud
abstract
We propose a guided image contrast enhancement framework based on cloud images, in which the context- sensitive and context-free contrast is jointly improved via solving a multi-criteria optimization problem. In particular, the context-sensitive contrast is improved by performing advanced unsharp masking on the input and edge-preserving filtered images, while the context-free contrast enhancement is achieved by the sigmoid transfer mapping. To automatically determine the contrast enhancement level, the parameters in the optimization process are estimated by taking advantages of the retrieved images with similar content. For the purpose of automatically avoiding the involvement of low-quality retrieved images as the guidance, a recently developed no-reference image quality metric is adopted to rank the retrieved images from the cloud. The image complexity from the free-energy-based brain theory and the surface quality statistics in salient regions are collaboratively optimized to infer the parameters. Experimental results confirm that the proposed technique can efficiently create visually-pleasing enhanced images which are better than those produced by the classical techniques in both subjective and objective comparisons.
Shiqi Wang 0001, Ke Gu 0001, Siwei Ma 0001, Weisi Lin, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Multim.2
2016 Fixation Prediction through Multimodal Analysis
abstract
In this article, we propose to predict human eye fixation through incorporating both audio and visual cues. Traditional visual attention models generally make the utmost of stimuli’s visual features, yet they bypass all audio information. In the real world, however, we not only direct our gaze according to visual saliency, but also are attracted by salient audio cues. Psychological experiments show that audio has an influence on visual attention, and subjects tend to be attracted by the sound sources. Therefore, we propose fusing both audio and visual information to predict eye fixation. In our proposed framework, we first localize the moving--sound-generating objects through multimodal analysis and generate an audio attention map. Then, we calculate the spatial and temporal attention maps using the visual modality. Finally, the audio, spatial, and temporal attention maps are fused to generate the final audiovisual saliency map. The proposed method is applicable to scenes containing moving--sound-generating objects. We gather a set of video sequences and collect eye-tracking data under an audiovisual test condition. Experiment results show that we can achieve better eye fixation prediction performance when taking both audio and visual cues into consideration, especially in some typical scenes in which object motion and audio are highly correlated.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2015 Perceptual screen content image quality assessment and compression
abstract
Compression of screen content has recently emerged as an active research topic due to the increasing demand in many applications such as wireless display and virtual desktop infrastructure. Screen content images (SCIs) exhibit different statistical properties in textual and pictorial regions, and the human visual system (HVS) also behaves differently when viewing the textual and pictorial regions in terms of the extent of visual field. Here we propose a perceptual SCI quality assessment approach that incorporates visual field adaptation and information content weighting. Furthermore, we propose a perceptual coding scheme in an attempt to optimize the HEVC Screen Content Coding encoder. Experimental results show that the proposed quality assessment method not only better predicts the perceptual quality of SCIs, but also leads to an effective way to optimize screen content coding schemes.
Shiqi Wang 0001, Ke Gu 0001, Kai Zeng 0003, Zhou Wang 0001, Weisi Lin
ICIP2
2015 Screen image quality assessment incorporating structural degradation measurement
abstract
Screen content is typically composed of computer generated text and graphics. The contents shown on the screen exhibit various unnatural properties, such as sharp edges and thin lines with few color variations. In this paper we design a novel structure-induced quality metric (SIQM) for assessing the screen image quality. The proposed SIQM works by weighting the benchmark structural similarity index (SSIM) with the structural degradation measurement that is computed using SSIM as well. Experimental results conducted on the newly released subjective quality database concerning screen images show that on one hand the proposed technique is superior to existing quality measures, and on the other hand our model is able to optimize screen video coding and thus introduce remarkable visual quality improvement.
Ke Gu 0001, Shiqi Wang 0001, Guangtao Zhai, Siwei Ma 0001, Weisi Lin
ISCAS1
2015 A general histogram modification framework for efficient contrast enhancement
abstract
In this paper we propose a new general histogram modification framework for contrast enhancement. The proposed model works with a hybrid transformation technique to improve image brightness and contrast based on an optional histogram matching in terms of reassigned probability distribution and S-shaped transfer mapping. Experimental results conducted on natural, dimmed, and tone-mapped images show that the proposed technique creates enhanced images efficiently with equivalent or superior visual quality to those produced by classical and state-of-the-art enhancement approaches.
Ke Gu 0001, Guangtao Zhai, Shiqi Wang 0001, Min Liu 0003, Jiantao Zhou 0001, Weisi Lin
ISCAS1
2015 Sparse Structural Similarity for Objective Image Quality Assessment
abstract
In this paper, a novel full-reference (FR) image quality assessment (IQA) metric based on sparse representation is proposed. Sparse representation has been widely applied in many applications such as image denoising and restoration. It is a high-efficiency way in representing sparse and redundant natural images. Also it has been shown to be highly related to the human visual perception, which is characterized by a set of responses of neurons in visual cortex. In this paper, the sparse representation is applied in decomposing natural images into multiple layers depending on the visual importance. Inspired by these observations, a novel IQA metric called sparse structural similarity is proposed by measuring the fidelity of the stimulation of visual cortices. Experimental results on public databases indicate that the proposed method is effective in predicting subjective evaluation and as compared to state-of-the-art FR-IQA methods.
Xiang Zhang 0004, Shiqi Wang 0001, Ke Gu 0001, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001
SMC3
2015 Visual attention on human face
abstract
Human faces are always the focus of visual attention since faces can provide plenty of information. Although some visual attention models incorporating face cues work better in scenes containing faces, no visual attention model is particularly designed for faces. On faces, many high-level factors will influence visual attention distribution. In practice, there are many visual communication systems in which faces occupy the scenes, such as video calls. Specific visual attention model designed for face images will be of great value in these circumstances. In this paper, we conduct research on visual attention analysis and modelling on human faces. To facilitate this research, we collect 120 face images and perform eye-tracking experiments with these images. Eye-movement data shows that detailed visual attention allocation exists on faces. Using face detection and facial landmark localization, we find that some facial features are highly effective for visual attention prediction. The performance of many visual attention models can be improved by incorporating those facial features.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001
VCIP3
2015 Fixation prediction through multimodal analysis
abstract
In this paper, we propose to predict human fixations by incorporating both audio and visual cues. Traditional visual attention models generally make the utmost of stimuli's visual features, while discarding all audio information. But in the real world, we human beings not only direct our gaze according to visual saliency but also may be attracted by some salient audio. Psychological experiments show that audio may have some influence on visual attention, and subjects tend to be attracted the sound sources. Therefore, we propose to fuse both audio and visual information to predict fixations. In our framework, we first localize the moving-sounding objects through multimodal analysis and generate an audio attention map, in which greater value denotes higher possibility of a position being the sound source. Then we calculate the spatial and temporal attention maps using only the visual modality. At last, the audio, spatial and temporal attention maps are fused, generating our final audio-visual saliency map. We gather a set of videos and collect eye-tracking data under audio-visual test conditions. Experiment results show that we can achieve better performance when considering both audio and visual cues.
Xiongkuo Min, Guangtao Zhai, Chunjia Hu, Ke Gu 0001
VCIP4
2015 Visual Saliency Detection With Free Energy Theory
abstract
Visual saliency can be thought of as the product of human brain activity. Most existing models were built upon local features or global features or both. Lately, a so-called free energy principle unifies several brain theories within one framework, and tells where easily surprise human viewers in a visual stimulus through a psychological measure. We believe that this “surprise” should be highly related to visual saliency, and thereby introduce a novel computational Free Energy inspired Saliency detection technique (FES). Our method computes the local entropy of the gap between an input image signal and its predicted counterpart that is reconstructed from the input one with a semi-parametric model. Experimental results prove that our algorithm predicts human fixation points accurately and is superior to classical/state-of-the-art competitors.
Ke Gu 0001, Guangtao Zhai, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Signal Process. Lett.1
2015 Automatic Contrast Enhancement Technology With Saliency Preservation
abstract
In this paper, we investigate the problem of image contrast enhancement. Most existing relevant technologies often suffer from the drawback of excessive enhancement, thereby introducing noise/artifacts and changing visual attention regions. One frequently used solution is manual parameter tuning, which is, however, impractical for most applications since it is labor intensive and time consuming. In this research, we find that saliency preservation can help produce appropriately enhanced images, i.e., improved contrast without annoying artifacts. We therefore design an automatic contrast enhancement technology with a complete histogram modification framework and an automatic parameter selector. This framework combines the original image, its histogram equalized product, and its visually pleasing version created by a sigmoid transfer function that was developed in our recent work. Then, a visual quality judging criterion is developed based on the concept of saliency preservation, which assists the automatic parameters selection, and finally properly enhanced image can be generated accordingly. We test the proposed scheme on Kodak and Video Quality Experts Group databases, and compare with the classical histogram equalization technique and its variations as well as state-of-the-art contrast enhancement approaches. The experimental results demonstrate that our technique has superior saliency preservation ability and outstanding enhancement effect.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.1
2015 No-Reference Image Sharpness Assessment in Autoregressive Parameter Space
abstract
In this paper, we propose a new no-reference (NR)/blind sharpness metric in the autoregressive (AR) parameter space. Our model is established via the analysis of AR model parameters, first calculating the energy- and contrast-differences in the locally estimated AR coefficients in a pointwise way, and then quantifying the image sharpness with percentile pooling to predict the overall score. In addition to the luminance domain, we further consider the inevitable effect of color information on visual perception to sharpness and thereby extend the above model to the widely used YIQ color space. Validation of our technique is conducted on the subsets with blurring artifacts from four large-scale image databases (LIVE, TID2008, CSIQ, and TID2013). Experimental results confirm the superiority and efficiency of our method over existing NR algorithms, the stateof-the-art blind sharpness/blurriness estimators, and classical full-reference quality evaluators. Furthermore, the proposed metric can be also extended to stereoscopic images based on binocular rivalry, and attains remarkably high performance on LIVE3D-I and LIVE3D-II databases.
Ke Gu 0001, Guangtao Zhai, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Image Process.1
2015 Using Free Energy Principle For Blind Image Quality Assessment
abstract
In this paper we propose a new no-reference (NR) image quality assessment (IQA) metric using the recently revealed free-energy-based brain theory and classical human visual system (HVS)-inspired features. The features used can be divided into three groups. The first involves the features inspired by the free energy principle and the structural degradation model. Furthermore, the free energy theory also reveals that the HVS always tries to infer the meaningful part from the visual stimuli. In terms of this finding, we first predict an image that the HVS perceives from a distorted image based on the free energy theory, then the second group of features is composed of some HVS-inspired features (such as structural information and gradient magnitude) computed using the distorted and predicted images. The third group of features quantifies the possible losses of “naturalness” in the distorted image by fitting the generalized Gaussian distribution to mean subtracted contrast normalized coefficients. After feature extraction, our algorithm utilizes the support vector machine based regression module to derive the overall quality score. Experiments on LIVE, TID2008, CSIQ, IVC, and Toyama databases confirm the effectiveness of our introduced NR IQA metric compared to the state-of-the-art.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Multim.1
2014 An efficient color image quality metric with local-tuned-global model
abstract
This paper investigates the problem of full-reference (FR) image quality assessment (IQA). In general, the ideal IQA metric should be effective and efficient, yet most of existing FR IQA methods cannot reach these two targets simultaneously. Under the supposition that the human visual perception to image quality depends on salient local distortion and global quality degradation, we introduce a novel effective and efficient local-tuned-global (LTG) model induced IQA metric. Extensive experiments are conducted on five publicly available subject-rated color image quality databases, including LIVE, TID2008, CSIQ, IVC and TID2013, to evaluate and compare our algorithm with classical and state-of-the-art FR IQA approaches. The proposed LTG is shown to work fast and outperform those competing methods.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
ICIP1
2014 Deep learning network for blind image quality assessment
abstract
Nowadays, blind image quality assessment (BIQA) has been intensively studied with machine learning, such as support vector machine (SVM) and k-means. Existing BIQA metrics, however, do not perform robust for various kinds of distortion types. We believe this problem is because those frequently used traditional machine learning techniques exploit shallow architectures, which only contain one single layer of nonlinear feature transformation, and thus cannot highly mimic the mechanism of human visual perception to image quality. The recent advance of deep neural network (DNN) can help to solve this problem, since the DNN is found to better capture the essential attributes of images. We in this paper therefore introduce a new Deep learning based Image Quality Index (DIQI) for blind quality assessment. Extensive studies are conducted on the new TID2013 database and confirm the effectiveness of our DIQI relative to classical full-reference and state-of-the-art reduced- and no-reference IQA approaches.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
ICIP1
2014 Details preservation inspired blind quality metric of tone mapping methods
abstract
High dynamic range (HDR) images are extremely meaningful, especially in the space and medical fields. For visualization of HDR images on standard low dynamic range (LDR) display devices, how to convert HDR to LDR images naturally becomes a valuable issue, which has aroused a variety of tone-mapping operators (TMOs). To compare different LDR images created by distinct TMOs, researchers have recently provided a subject-rated tone-mapped image database, and then developed a full-reference objective tone-mapped image quality index (TMQI) based on the measurement of multi-scale signal fidelity and statistical naturalness. Instead, the basic property of HDR images about details preservation is studied in this paper. With it, a natural inference is that higher-quality tone-mapped images are capable of displaying much more details. We therefore propose a blind quality metric by estimating the amount of details in images generated by darkening/brightening an original tone-mapped images. Experimental results on the above tone-mapped image database confirm that the proposed method, despite of no reference, is robust and statistically superior to the currently optimal full-reference TMQI algorithm, and remarkably outperforms state-of-the-art no-reference IQA metrics.
Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiaokang Yang 0001, Wenjun Zhang 0001
ISCAS1
2014 Visual attention data for image quality assessment databases
abstract
Images usually contain areas that particularly attract people's attention and visual attention is an important feature of human visual system (HVS). Visual attention had been shown to be effective in improving performance of existing image quality assessment (IQA) metrics. However, with the quick advancement of IQA research, the booming of open IQA databases calls for associated comprehensive and accurate visual attention dataset. Despite of the large number of existing computational attention/saliency models, the most accurate measure of human attention is still human based. In this research, we first conduct extensive eye tracking experiments for all the pristine images from the seven widely used IQA databases (LIVE, TID2008, CSIQ, Toyama, LIVE Multiply Distortion, IVC and A57 databases). Then we propose a gaze-duration adaptive weighting approach to generate saliency maps from the eye tracking data. When applied on the IQA databases, experimental results suggest that accuracy of benchmark quality metrics, e.g. PSNR and SSIM can be systematically improved, outperforming existing saliency datasets. Both the eye tracking data and the saliency maps in this research will be made publicly available at gvsp.sjtu.edu.cn.
Xiongkuo Min, Guangtao Zhai, Zhongpai Gao, Ke Gu 0001
ISCAS4
2014 Learning to integrate local and global features for a blind image quality measure
abstract
In this paper, we present a new algorithm for blind/no-reference image quality assessment (BIQA/NR-IQA). Most existing measures are “opinion-aware”, demanding human opinion scored images to map image features to them. The task of obtaining human scores of images is, however, commonly thought to be uneconomical, and thus we focus on “opinion free” (OF) quality metrics in this research. By integrating local and global features, this paper develops a learning-based BIQA approach with three steps by combining local and global features together. In the first step of extracting local features, we use the quality aware clustering with the centroid of each quality level trained by K-means, while we in the second step compute the global features based on the natural scene statistics. Finally, the third step uses the SVR to train a regression module from the above-mentioned local and global features to derive the overall image quality score. Experimental results on LIVE, TID2008, CSIQ, and TID2013 databases validate the effectiveness of our proposed metric (a general framework) as compared to popular no-, reduced- and full-reference IQA approaches.
Min Liu 0003, Guangtao Zhai, Ke Gu 0001, Xiaokang Yang 0001
SMARTCOMP3
2013 Subjective and objective quality assessment for images with contrast change
abstract
It is widely known that, for most natural images, appropriate contrast enhancement can usually lead to improved subjective quality. Despite of its importance to image processing, contrast change has largely been overlooked in the current research of image quality assessment (IQA). To fill this void, in this paper we first report a new and dedicated contrast-changed image database (CID2013). The CID2013 database is composed of four hundred contrast-changed images of fifteen original natural images and the mean opinion scores (MOSs) recorded from twenty-two inexperienced viewers. We then proposed a novel reduced-reference image quality metric for contrast-changed images (RIQMC) using entropies and order statistics of the image histograms. Experimental results on the CID2013, TID2008, and CSIQ databases demonstrate that the proposed RIQMC metric outperforms some mainstream image quality assessment methods.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Min Liu 0003
ICIP1
2013 No-reference image quality assessment metric by combining free energy theory and structural degradation model
abstract
In the research of image quality assessment (IQA), no-reference approaches are usually thought of as a big challenge since none of original image information is available. To tackle this problem, we propose a new no-reference image quality metric through combining two recently proposed reduced-reference IQA models, namely the free energy based distortion metric (FEDM) and the structural degradation model (SDM). In this work, it will be shown that there exists an approximate linear relationship between the original image information of the free energy feature and the structural degradation information. Based on this observation and the application of support vector machine (SVM) that is widely used in the current study of IQA, our newly developed No-reference Free energy and Structural degradation based Distortion Metric (NFSDM) is found to alleviate the dependance of original images, and has achieved remarkably well prediction accuracy, outperforming the most two full-reference IQA approaches PSNR/SSIM and several mainstream no-reference image quality metrics.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Longfei Liang
ICME1
2013 A new reduced-reference image quality assessment using structural degradation model
abstract
Image quality assessment (IQA) is an important research area in image processing. Reduced-reference (RR) IQA methods contained therein mainly aim to estimate image quality degradations with partial information about the reference image. Following the remarkable achievement of SSIM, structural information has been recognized as one key factor, and has aroused many image quality metrics so far. In this paper, we design a structural degradation model (SDM). Then, the quality score of an image is defined as a nonlinear combination, or SVM based integration, of distance between the structural degradation information of the original and distorted images. Accordingly, a new RR IQA approach using the SDM model is exploited. Experimental results on LIVE database are provided to justify the superior prediction accuracy performance of the proposed method as compared to three significant image quality metrics, PSNR, SSIM and FEDM.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
ISCAS1
2013 Self-adaptive scale transform for IQA metric
abstract
Recently, an increasing number of image quality assessment (IQA) algorithms have been developed based on multi-scale methods, such as MS-SSIM, IFC, VIF and IW-PSNR/SSIM. Inspired by the achievement of multi-scale type of IQA algorithms, this paper proposes a self-adaptive scale transform based IQA approach. Using image size and viewing distance as input variables, we construct a self-adaptive scale transform function to estimate the suitable scale transform coefficient for the following image quality metrics. Two of the most well-known full-reference IQA methods (PSNR and SSIM), and three publicly-available subjectrated image databases (LIVE, IVC and Toyama-MICT) with clear image size and viewing distance values are used as testing beds in this paper. Experimental results and comparative studies on different combinations of IQA methods and image databases suggest the effectiveness and the robustness of the proposed approach.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
ISCAS1
2013 Brightness preserving video contrast enhancement using S-shaped Transfer function
abstract
This paper presents an efficient perceptual model inspired efficient video contrast enhancement algorithm. We propose a S-shaped transfer function for image pixel values that effectively improves the perceived contrast while preserving brightness of the scene. The S-shaped transfer function has only one control parameter that can be adaptively chosen for different video contents, such as sports, cartoon, news, and landscape programs. Then, the input image brightness is further preserved, in order to maintain the perception of human visual system (HVS) to some special scenes, such as dark scene and seaside scene. Experiments and comparative study on VQEG Phase I test database demonstrate that the proposed S-shaped Transfer function based Brightness Preserving (STBP) contrast enhancement algorithm outperforms various histogram equalization based methods such as HE, DSIHE, RSIHE and WTHE, yet with much lower computational complexity.
Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiongkuo Min, Xiaokang Yang 0001, Wenjun Zhang 0001
VCIP1
2013 Adaptive high-frequency clipping for improved image quality assessment
abstract
It is widely known that the human visual system (HVS) applies multi-resolution analysis to the scenes we see. In fact, many of the best image quality metrics, e.g. MS-SSIM and IW-PSNR/SSIM are based on multi-scale models. However, in existing multi-scale type of image quality assessment (IQA) methods, the resolution levels are fixed. In this paper, we examine the problem of selecting optimal levels in the multi-resolution analysis to preprocess the image for perceptual quality assessment. According to the contrast sensitivity function (CSF) of the HVS, the sampling of visual information by the human eyes approximates a low-pass process. For images, the amount of information we can extract depends on the size of the image (or the object(s) inside) as well as the viewing distance. Therefore, we proposed a wavelet transform based adaptive high-frequency clipping (AHC) model to approximate the effective visual information that enters the HVS. After the high-frequency clipping, rather than processing separately on each level, we transform the filtered images back to their original resolutions for quality assessment. Extensive experimental results show that on various databases (LIVE, IVC, and Toyama-MICT), performance of existing image quality algorithms (PSNR and SSIM) can be substantially improved by applying the metrics to those AHC model processed images.
Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiaokang Yang 0001, Jun Zhou 0007, Wenjun Zhang 0001
VCIP1
2012 A new no-reference stereoscopic image quality assessment based on ocular dominance theory and degree of parallax
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
ICPR1
2012 Nonlinear additive model based saliency map weighting strategy for image quality assessment
abstract
Most state-of-the-art image quality metrics are based on the two-step approach: local distortion/fidelity measurement and pooling. During the pooling stage, many weighting strategies have been proposed incorporating properties of the distortion itself, various masking effects and visual attention. Recently, researchers have devoted great enthusiasm and effort to the improvement of image quality assessment using visual saliency models. In this research, it is noticed that visual saliency features of both the original image and the distorted one have impacts on the process of image quality assessment. To reduce the overlapping effects, a nonlinear additive model is proposed to integrate saliency features from the original and distorted images towards improved error weighting results. Our extensive experimental studies on four publicly available image databases (LIVE, TID2008, CSIQ and A57) indicate that the proposed improved nonlinear additive model based saliency map weighting strategy constantly leads to higher prediction accuracy for image quality assessment than traditional methods.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Li Chen 0021, Wenjun Zhang 0001
MMSP1
2012 Robust object tracking with bidirectional corner matching and trajectory smoothness algorithm
abstract
This paper proposes a novel method for robust object tracking. The method consists of three different components: a short term tracker, an object detector, and an online object model. For the short term tracker, we use an advanced Lucas Kanade tracker with bidirectional corner matching to capture object frame by frame. Meanwhile, statistical filtering and matching algorithm combined with haar-like feature random fern play as a detector to extract all possible object candidates in the current frame. Making use of trajectory information, the online object model decides the best target match among the candidates. And the model also trains the random fern feature adaptively online to better guide consecutive tracking. We demonstrate our method is robust to track an object in a long term and under large variations of view angle and lighting conditions. Moreover, our method is efficient to re-detect the object and keep tracking even after it's out of view or recover from heavy occlusion. To achieve state-of-the-art performance, it is highlighted that our method can be extended to multiple objects tracking application. Finally, comparisons with other state-of-the-art trackers are presented to show the robustness of our tracker.
Guoshan Wu, Yi Xu 0001, Xiaokang Yang 0001, Ke Gu 0001
MMSP5