Hantao Liu

dblp:81/5931 · DBLP profile ↗
← Back
102ranked-venue papers
9as first author
62since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 83 · 9 first-author · 46 since 2021Artificial intelligence and machine learning · 13 · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment
abstract
As super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches.
Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
AAAI8
2026 Cross-Modal Interaction for Multi-Dimensional AI-Generated Image Quality Assessment
Minghao Zou, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
QoMEX4
2026 Robust indoor localization via factor graph fusion with motion-aware regression, adaptive relocalization, and continuous map constraints
Yujin Kuang, Zhengdong Wang, Xiangyin Meng, Xiaoguo Zhang, Hantao Liu
Eng. Appl. Artif. Intell.7
2026 EHIN: Early-aware hierarchical interaction network for weakly-supervised referring image segmentation
Anqing Chen, Wanli Ma 0001, Weide Liu, Yakun Ju, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
Neurocomputing9
2026 MIQANet: A Novel Dual-Branch Deep Learning Framework for MRI Image Quality Assessment
abstract
Image quality assessment (IQA) algorithms have significantly advanced over the past two decades, primarily focusing on natural images. However, applying these methods directly to medical imaging often yields suboptimal performance due to inherent differences such as the structural complexity of medical images and the limited availability of annotated databases. In this study, we conduct a comprehensive evaluation of state-of-the-art IQA methods, including 29 traditional full-reference (FR), 4 traditional no-reference (NR), and 9 deep learning-based approaches, to assess their effectiveness in the context of medical imaging. Our evaluation is performed on a recently developed MRI image quality assessment benchmark, revealing critical performance gaps in existing methods. Building on these findings, we propose a novel dual-branch deep learning framework specifically designed for medical IQA (MIQANet). The proposed approach effectively combines global contextual information with local structural details, enhancing the model’s ability to capture subtle degradations and structural inconsistencies in MRI scans. Experiential results demonstrate the superiority of our approach over existing methods, providing valuable theoretical and practical insights for enhancing quality assessment of medical images.
Yueran Ma, Huasheng Wang, Jean-Yves Tanguy, Phillip Wardle, Elizabeth A. Krupinski, Padraig Corcoran, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.9
2026 KSIQA: A Knowledge-Sharing Model for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) aims to quantitatively measure human perception of visual quality without comparing a distorted image to a reference. Despite recent advances, existing NR-IQR approaches often demonstrate insufficient ability to capture perceptual cues in the absence of a reference, limiting their generalisability across diverse and complex real-world image degradations. These limitations hinder their ability to match the reliability of full-reference IQA (FR-IQA) counterparts. A key challenge, therefore, is to enable NR-IQA models to emulate the reference-aware reasoning exhibited by humans and FR-IQA methods. To address this challenge, we propose a novel NR-IQA model based on a knowledge-sharing (KS) strategy to simulate this capability and predict image quality more effectively. Specifically, we designate an FR-IQA model as the teacher and an NR-IQA model as the student. Unlike conventional knowledge distillation (KD), our proposed architecture enables the NR-IQA student and FR-IQA teacher to share a decoder rather than being independent models. Furthermore, the student model contains a Mental Imagery Generation (MIG) module to learn mental imagery as the reference. To fully exploit local and global information, we adopt a vision transformer (ViT) branch and a convolutional neural network branch for feature extraction (FE). Finally, a quality-aware regressor (QAR) combined with deep ordinal regression is constructed to infer the quality score. Experiments show that our proposed NR-IQA model, KSIQA, has class-leading performance against current no-reference (NR) techniques across widespread benchmark datasets.
Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Ying Chen 0011, Roger M. Whitaker, Walter Colombo, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.10
2025 Analysing and Predicting Radiologists' Expertise Using Eye-Tracking Data: Insights for Diagnostic Decision-Making
abstract
Radiologists’ search strategies and decision-making processes during chest X-ray diagnosis vary with their levels of expertise. Understanding these differences can inform training programmes and support the development of tools to enhance diagnostic accuracy. We hypothesize that eye-tracking data can reveal variations in expertise levels that serve as a predictor of radiologist expertise. To investigate this, we develop a database of 191 chest X-ray images with ground-truth annotations, including diagnostic decisions and eye movement patterns from 13 radiologists of varying levels of expertise. Statistical analyses reveal distinct diagnostic search patterns associated with different expertise levels. In addition, we propose a predictive framework to estimate expertise levels based on eye-tracking data. This study advances the understanding of expertise-driven differences in diagnostic search strategies and demonstrates the potential of eye-tracking data in enhancing training in clinical radiology.
Yueran Ma, Phillip Wardle, Gualtiero Colombo 0001, Padraig Corcoran, Hantao Liu
ICME10
2025 CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIP
abstract
Blind dehazed image quality assessment (BDQA), which aims to accurately predict the visual quality of dehazed images without any reference information, is essential for the evaluation, comparison, and optimization of image dehazing algorithms. Existing learning-based BDQA methods have achieved remarkable success, while the small scale of DQA datasets limits their performance. To address this issue, in this paper, we propose to adapt Contrastive Language-Image Pre-Training (CLIP), pre-trained on large-scale image-text pairs, to the BDQA task. Specifically, inspired by the fact that the human visual system understands images based on hierarchical features, we take global and local information of the dehazed image as the input of CLIP. To accurately map the input hierarchical information of dehazed images into the quality score, we tune both the vision branch and language branch of CLIP with prompt learning. Experimental results on two authentic DQA datasets demonstrate that our proposed approach, named CLIP-DQA, achieves more accurate quality predictions over existing BDQA methods. The code is available at https://github.com/JunFu1995/CLIP-DQA.
Yirui Zeng, Jun Fu 0007, Hadi Amirpour, Huasheng Wang, Guanghui Yue 0001, Hantao Liu, Ying Chen 0011, Wei Zhou 0021
ISCAS6
2025 Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng
MICCAI (15)3
2025 PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics Experiments
abstract
Visual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limitations of existing datasets: (1) limited annotation diversity, which limits the support for diverse vision tasks within a unified dataset; (2) insufficient coverage of domains, particularly a lack of datasets tailored for educational scenarios; and (3) a lack of explicit procedural guidance, with weak logical rules and insufficient representation of a structured task process. To address these gaps, we introduce PhysLab, the first dataset that captures students conducting complex physics experiments. The dataset includes four representative experiments that feature diverse scientific instruments and rich human-object interaction (HOI) patterns. PhysLab comprises 620 long-form videos and provides multi-granularity annotations that support a variety of vision tasks, including action recognition, object detection, HOI analysis, etc. We establish baselines and perform extensive evaluations to highlight key challenges in the parsing of procedural educational videos. We expect PhysLab to serve as a valuable resource for advancing comprehensive visual parsing, facilitating intelligent classroom systems, and fostering closer integration among computer vision, multimedia, and educational technologies. The dataset and the evaluation toolkit are publicly available at https://github.com/ZMH-SDUST/PhysLab.
Minghao Zou, Qingtian Zeng, Yongping Miao, Hantao Liu, Wei Zhou 0021
ACM Multimedia6
2025 Vision-based human action quality assessment: A systematic review
abstract
Human Action Quality Assessment (AQA), which aims to automatically evaluate the performance of actions executed by humans, is an emerging field of human action analysis. Although many review articles have been conducted for human action analysis fields such as action recognition and action prediction, there is a lack of up-to-date and systematic reviews related to AQA. This paper aims to provide a systematic literature review of existing papers on vision-based human AQA. This systematic review was rigorously conducted following the PRISMA guideline through the databases of Scopus , IEEE Xplore, and Web of Science in July 2024. Ninety-six research articles were selected for the final analysis after applying inclusion and exclusion criteria. This review presents an overview of various aspects of AQA, including existing applications, data acquisition methods, public datasets, state-of-the-art methods and evaluation metrics . We observe an increase in the number of studies in AQA since 2019, primarily due to the advent of deep learning methods and motion capture devices. We categorize these AQA methods into skeleton-based and video-based methods based on the data modality used. There are different evaluation metrics for various AQA tasks. SRC is the most commonly used evaluation metric, with fifty-six out of ninety-six selected papers using it to evaluate their models. Sports event scoring, surgical skill evaluation and rehabilitation assessment are the most popular three scenarios in this direction based on existing papers and there are more new scenarios being explored such as piano skill assessment. Furthermore, the existing challenges and future research directions are provided, which can be a helpful guide for researchers to explore AQA.
Huasheng Wang, Katarzyna Stawarz, Shiyin Li, Hantao Liu
Expert Syst. Appl.6
2025 Underwater Image Quality Evaluation: A Comprehensive Review
abstract
ABSTRACT Underwater image quality evaluation (UIQE) is crucial in improving image processing techniques and optimizing the design of the imaging system to obtain object information more accurately. However, existing UIQE methods are designed based on limited images or consider only a few natural scene statistics (NSS) metrics, lacking consideration for generalization across various underwater imaging applications. In this paper, an in‐depth review of the existing UIQE methods based on evaluation operations is provided, emphasizing the bias present when evaluating UIQE methods using individual metrics. To address this, a novel metric called quadrilateral datum evaluation (QDE) is designed for UIQE methods. It comprehensively considers robustness across different datasets, as well as correlation and ranking consistency with mean opinion scores (MOS). This is the first solution to measure an UIQE method from an all‐encompassing visual perspective. By using QDE, UIQE methods characterized by greater feature strength and small imbalance demonstrate good consistency and robustness across multiple aspects, providing a basis for the design of UIQE methods.
Mengjiao Shen, Jinyang Zhong, Hantao Liu, Can Pan
IET Image Process.4
2025 Applying cross-modal plasticity principles in auditory training applications
abstract
Research indicates that a significant number of individuals are in a suboptimal auditory health state , yet their auditory function can potentially be improved through auditory training. To raise awareness of auditory health issues, auditory training apps should provide effective yet accessible training methods alongside engaging mechanisms that motivate users to adopt and sustain auditory training habits, ultimately facilitating self-directed auditory health management. Current auditory training apps overlook the importance of user needs and motivation, as well as the relationship between the two, leading to low engagement and retention rates. We document the specific needs of both normal-hearing and mildly hearing-impaired users and analyze physiological data collected during auditory training tasks. Through lab experiment study, we provide quantitative evidence supporting the effectiveness of the auditory training method used in this study. The findings indicate that audiovisual-based auditory training contributes to improved auditory performance. Building on these insights, we develop an auditory training app prototype that integrates gamification and narrative design into the auditory training app prototype, and examine their impact on user motivation and engagement. Furthermore, based on the results, we propose design recommendations for future auditory training app development, emphasizing the need to align training effectiveness with user motivation and engagement strategies.
Qiqi Huang, Katarzyna Stawarz, Linqi Zhao, Shuya Yang, Wenyu Xie, Fanghao Song, Hantao Liu
Int. J. Hum. Comput. Stud.7
2025 Hierarchical boundary feature alignment network for video salient object detection
abstract
The deep learning based video salient object detection (VSOD) models have achieved great success in the past few years, however, these VSOD models still suffer from the following two problems: i) struggle in accurately predicting those pixels surrounding salient objects; ii) unaligned features of different scales lead to deviations in feature fusion . To tackle these problems, we propose a hierarchical boundary feature alignment network (HBFA). Specifically, the proposed HBFA consists of a temporal–spatial fusion module (TSM) and three decoding branches. TSM captures multi-scale spatiotemporal information. The two boundary feature branches are used to guide the whole network to pay more attention to the boundary of salient objects, while the feature alignment branch is capable of fusing the features from the internal and external branches while aligning features across different scales. Our extensive experiments show that the proposed method reaches a new state-of-the-art performance.
Amin Mao, Jiebin Yan, Yuming Fang 0001, Hantao Liu
J. Vis. Commun. Image Represent.4
2025 Towards Scalable and Efficient Full-Reference Omnidirectional Image Quality Assessment
abstract
Full-Reference (FR) image quality assessment (IQA) (FR-IQA) has achieved notable success due to its irreplaceable role in algorithm and system optimization; however, it has less been investigated in omnidirectional image quality assessment (OIQA). In this paper, we make an attempt to FR-OIQA considering the constraint of the computation budget, in which this issue is formulated as “quality perception from patch to sequence”,i.e.,Intra-PatchSequence degradation modeling andInter-PatchSequence similarity calculation (denoted by IPS$^{2}$). Specifically, IPS$^{2}$directly accepts local patches from the omnidirectional image (OI) in the format of Equirectangular Projection as input, avoiding other preprocessing operations, such as scan-path prediction and projection transformation. Subsequently, IPS$^{2}$uses a deep feature extractor to capture patch quality and then sends the patch- wise quality maps to the cross-patch similarity (CPS) module, which explicitly models intra-patch sequence degradation and inter-patch sequence similarity via self-attention. Finally, a quality regressor is used to aggregate these features of the CPS module and predict the global quality of the OI. The experimental results on a large-scale OIQA database show that the proposed IPS$^{2}$outperforms most state-of-the-art methods in quality prediction accuracy while offering substantial reductions in computational cost and model size.
Jiebin Yan, Zhihua Wang 0002, Yuming Fang 0001, Hantao Liu
IEEE Signal Process. Lett.5
2025 CLIP-DQA V2: Exploring CLIP for Dehazed Image Quality Assessment From a Fragment-Level Perspective
Yirui Zeng, Jun Fu 0007, Guanghui Yue 0001, Hantao Liu, Wei Zhou 0021
IEEE Signal Process. Lett.4
2025 Adaptive Spatiotemporal Graph Transformer Network for Action Quality Assessment
abstract
Long video action quality assessment (AQA) aims to evaluate the performance of long-term actions depicted in a video and produce an overall assessment for action quality. A video of long-term actions often contains more complicated temporal and spatial information than that of short-term actions. However, existing approaches that segment a video into individual clips for independent analysis potentially disrupt the narrative flow and diminish contextual details within and across clips, impeding comprehensive video understanding. To address this challenge, we propose an adaptive spatiotemporal graph transformer network (ASGTN) that combines multiple graph structures and transformer attention mechanisms to capture both local and global contextual information within and across clips in a long video. Specifically, the adaptive spatiotemporal graph (ASG) combines a spatial graph branch, designed to enrich the local nuanced spatiotemporal relations within an individual clip, and a temporal graph branch, tailored to dynamically learn the semantic context across different clips. Furthermore, a transformer encoder is integrated to amplify the global dependencies across clips in the entire video. This structure is designed to preserve narrative coherence and maintain essential contextual details in video-level features. Finally, we employ a level-focused decoder to predict the action quality score distribution. Experiments demonstrate that our model achieves state-of-the-art results on popular AQA datasets. Our code is available athttps://github.com/jiangliu5/ASGTN_AQA.
Huasheng Wang, Wei Zhou 0021, Katarzyna Stawarz, Padraig Corcoran, Ying Chen 0011, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.7
2025 Image Manipulation Quality Assessment
abstract
Image quality assessment (IQA) and its computational models play a vital role in modern computer vision applications. Research has traditionally focused on signal distortions arising during image compression and transmission, and their impact on perceived image quality. However, little attention is paid to image manipulation that alters an image using various filters. With the prevalence of image manipulation in real-life scenarios, it is critical to understand how humans perceive filter-altered images and to develop reliable IQA models capable of automatically assessing the quality of filtered images. In this paper, we build a new IQA database for filter-altered images, comprised of 360 images manipulated by various filters. To ensure the subjective IQA faithfully reflects human visual perception, we conduct a fully-controlled psychovisual experiment. Building upon the ground truth, we propose an innovative deep learning-based no-reference IQA (NR-IQA) model named IMQA that can accurately predict the perceived quality of filter-altered images. This model involves constructing an image filtering-aware module to learn discriminatory features for filter-altered images; and fuses these features with the representations generated by an image quality-aware module. Experimental results demonstrate the superior performance of the proposed IMQA model.
Xinbo Wu, Jianxun Lou, Wan'an Liu, Paul L. Rosin, Gualtiero Colombo 0001, Stuart M. Allen, Roger M. Whitaker, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.9
2025 Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality Assessment
abstract
Many super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods.
Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
IEEE Trans. Image Process.8
2025 Distortion-Induced Saliency Shifts in Video
abstract
Visual saliency modelling is of fundamental importance in modern video processing and its applications. Our previous eye-tracking study revealed that signal distortions caused by editing, compression, or transmission alter gaze patterns and consequently induce saliency shifts in both spatial and temporal domains. Saliency shifts provide crucial insights into viewers’ behavioural responses to video distortions, facilitating the perception-based optimisation of video algorithms. However, the spatio-temporal saliency shifts and their measurable effects on perception related applications remain largely unexplored. In this paper, we first investigate the measurement of distortion-induced saliency shifts (DSS) in videos and analyse DSS behaviours as functions of video content, time order and critical distortion disruption. Second, based on our findings, we construct three vision models to quantitatively simulate distinct DSS behaviours and integrate them into a comprehensive DSS behaviour model. Finally, we demonstrate that the computational DSS model can enhance emerging video technologies.
Xinbo Wu, Jianxun Lou, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu
IEEE Trans. Multim.6
2025 Chest X-Ray Visual Saliency Modeling: Eye-Tracking Dataset and Saliency Prediction Model
abstract
Radiologists' eye movements during medical image interpretation reflect their perceptual-cognitive processes of diagnostic decisions. The eye movement data can be modeled to represent clinically relevant regions in a medical image and potentially integrated into an artificial intelligence (AI) system for automatic diagnosis in medical imaging. In this article, we first conduct a large-scale eye-tracking study involving 13 radiologists interpreting 191 chest X-ray (CXR) images, establishing a best-of-its-kind CXR visual saliency benchmark. We then perform analysis to quantify the reliability and clinical relevance of saliency maps (SMs) generated for CXR images. We develop CXR image saliency prediction method (CXRSalNet), a novel saliency prediction model that leverages radiologists' gaze information to optimize the use of unlabeled CXR images, enhancing training and mitigating data scarcity. We also demonstrate the application of our CXR saliency model in enhancing the performance of AI-powered diagnostic imaging systems.
Jianxun Lou, Huasheng Wang, Xinbo Wu, John Cho Hui Ng, Kaveri A. Thakoor, Padraig Corcoran, Ying Chen 0011, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.9
2025 A Bioinspired Deep Learning Framework for Saliency-Based Image Quality Assessment
abstract
Advancements in deep learning have led to significant progress in no-reference (NR) image quality assessment (NR-IQA) for evaluating the perceived quality of digital images without relying on a reference. However, existing NR-IQA models remain suboptimal in handling complex and diverse natural images. Visual saliency constitutes a critical element for enhancing the reliability of NR-IQA, but the optimal use of saliency in deep learning-based NR-IQA has not heretofore been significantly explored. In this article, we present a novel method for integrating saliency in NR-IQA, which is motivated by the saliency-based visual search mechanism that different parts of the visual input are visited by the focus of attention (FOA) in the order of decreasing saliency. By dividing saliency into the high and low levels of FOA, we build a bioinspired deep neural network-BioSIQNet-based on a multitask learning (MTL) framework. The network architecture consists of two saliency-specific tasks and one primary image quality assessment (IQA) task. The low and high saliency (HS) are separately encoded and integrated into the early and deeper layers of the IQA network, respectively, analogous to the hierarchical processing in the visual cortex of the brain that allocates low attentional resources to process the simple patterns and high resources to learn intricate representations. We demonstrate that leveraging the synergy between visual attention and image quality perception and joint learning of these interconnected visual tasks can enhance the overall learning capabilities of the primary IQA model. Experiments validate the effectiveness of our proposed BioSIQNet for NR-IQA.
Huasheng Wang, Yueran Ma, Hongchen Tan, Xiaochang Liu, Ying Chen 0011, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.6
2024 Time-Interval Visual Saliency Prediction in Mammogram Reading
abstract
Radiologists’ eye movements during medical image interpretation reflect their perceptual-cognitive behaviour and correlate with diagnostic decisions. Previous study has shown the significance of gaze behaviour of different time intervals for the decision-making process. Being able to automatically predict the visual attention of radiologists for different reading phases would enhance the reliability and explainability of artificial intelligence (AI) in diagnostic imaging. In this paper, we investigate the time-interval visual saliency in mammogram reading. We propose a novel visual saliency prediction model based on deep learning, which predicts a sequence of time-interval saliency maps for an input mammogram. Experimental results demonstrate the efficacy of the proposed time-interval saliency model.
Jianxun Lou, Xinbo Wu, Hantao Liu
ICASSP5
2024 A Benchmark of Variance of Opinion Scores in Image Quality Assessment
abstract
Mean opinion score (MOS) has been used as the benchmark to measure the perceived quality of digital images. However, the usefulness of MOS diminishes when a substantial variation between individual opinions occurs. It is critical to measure the stimulus-driven variance of opinion scores (VOS) and scrutinise images that evoke a large VOS, and consequently, use VOS to inform our interpretation of MOS. In this paper, we create a VOS benchmark for individual differences in image quality assessment and analyse the importance of VOS classification as a function of distortion intensity, distortion type and scene content. In addition, a simple yet effective deep learning-based model is built, aiming to identify images with a large variation in viewers’ quality judgements.
Jianxun Lou, Xinbo Wu, Padraig Corcoran, Gualtiero Colombo 0001, Roger M. Whitaker, Hantao Liu
ICIP7
2024 Blind Quality Assessment of Panoramic Images Based on Multiple Viewport Sequences
abstract
With the development of virtual reality (VR) technology, panoramic image (PI), which is an important digital form of immersive multimedia, has drawn much attention from researchers. However, distortions are inevitably introduced in the process of processing, encoding and compression, which damages their quality and affects the user’s experience. Therefore, assessing the quality of panoramic images is urgent. In this paper, with the consideration of viewing behavior, we propose a novel blind panoramic image quality assessment model, which consists of three parts, viewport generation, feature extraction and quality prediction. Specifically, inspired by the viewing process of PI, we first generate multiple viewport sequences according to the real viewing trajectory and then extract multilevel features with a pre-trained backbone. The concatenated features are taken as the input of a recurrent neural network to evaluate the perceptual quality of PI. To validate the effectiveness of the proposed method, objective experiments are conducted on the public subjective panoramic image quality database. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods.
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Hantao Liu
ISCAS5
2024 TranSalNet+: Distortion-aware saliency prediction
abstract
Predicting the saliency of images affected by distortion is a challenging but emerging research problem. Given a distorted image, we wish to accurately predict saliency as perceived by humans. A recent distortion-aware saliency benchmark – the CUDAS database – reveals the inadequacy of existing saliency models in handling distorted images. In this paper, we devise a deep learning Distortion-Aware Saliency Module (DASM) that enables capturing saliency features related to image distortions, and integrates this module into a saliency prediction architecture. To achieve the high expressive capability of DASM using supervised learning, we create a dedicated dataset that draws upon a large-scale saliency dataset and machine-generated image quality assessments . Experimental results demonstrate the superior performance of the proposed model in predicting the saliency of distorted images.
Jianxun Lou, Xinbo Wu, Padraig Corcoran, Paul L. Rosin, Hantao Liu
Neurocomputing5
2024 RAD-IQMRI: A benchmark for MRI image quality assessment
abstract
Magnetic resonance imaging (MRI) is susceptible to visual artifacts that can degrade the perceptual image quality, potentially leading to inaccurate or inefficient diagnoses in clinical practice. It is critical to evaluate the perceptual image quality and build this technique into clinical solutions. In a previous study, an MRI database was created for image quality assessment (IQA), where various types of MRI artifacts with different degrees of degradation were simulated. Application specialists assessed the image quality; however, radiologists’ perception of MRI image quality remains unknown. To make IQA clinically relevant, in this paper we conduct a new subjective experiment where 13 radiologists rated the quality of images contained in the MRI database. Based on this subjective IQA benchmark named RAD-IQMRI, we evaluate the performance of state-of-the-art objective IQA models, providing insights into their application for MRI image quality assessment in clinical settings.
Yueran Ma, Jianxun Lou, Jean-Yves Tanguy, Padraig Corcoran, Hantao Liu
Neurocomputing5
2024 Fine-tuning coreference resolution for different styles of clinical narratives
abstract
OBJECTIVE: Coreference resolution (CR) is a natural language processing (NLP) task that is concerned with finding all expressions within a single document that refer to the same entity. This makes it crucial in supporting downstream NLP tasks such as summarization, question answering and information extraction. Despite great progress in CR, our experiments have highlighted a substandard performance of the existing open-source CR tools in the clinical domain. We set out to explore some practical solutions to fine-tune their performance on clinical data. METHODS: We first explored the possibility of automatically producing silver standards following the success of such an approach in other clinical NLP tasks. We designed an ensemble approach that leverages multiple models to automatically annotate co-referring mentions. Subsequently, we looked into other ways of incorporating human feedback to improve the performance of an existing neural network approach. We proposed a semi-automatic annotation process to facilitate the manual annotation process. We also compared the effectiveness of active learning relative to random sampling in an effort to further reduce the cost of manual annotation. RESULTS: Our experiments demonstrated that the silver standard approach was ineffective in fine-tuning the CR models. Our results indicated that active learning should also be applied with caution. The semi-automatic annotation approach combined with continued training was found to be well suited for the rapid transfer of CR models under low-resource conditions. The ensemble approach demonstrated a potential to further improve accuracy by leveraging multiple fine-tuned models. CONCLUSION: Overall, we have effectively transferred a general CR model to a clinical domain. Our findings based on extensive experimentation have been summarized into practical suggestions for rapid transferring of CR models across different styles of clinical narratives.
Yuxiang Liao, Hantao Liu, Irena Spasic
J. Biomed. Informatics2
2024 A task offloading strategy based on sequential waiting model in MEC
Xiulan Sun, Wenzao Li, Hantao Liu, Zhan Wen, Chengyu Wen
Multim. Tools Appl.3
2024 Vision-Language Consistency Guided Multi-Modal Prompt Learning for Blind AI Generated Image Quality Assessment
abstract
Recently, textual prompt tuning has shown inspirational performance in adapting Contrastive Language-Image Pre-training (CLIP) models to natural image quality assessment. However, such uni-modal prompt learning method only tunes the language branch of CLIP models. This is not enough for adapting CLIP models to AI generated image quality assessment (AGIQA) since AGIs visually differ from natural images. In addition, the consistency between AGIs and user input text prompts, which correlates with the perceptual quality of AGIs, is not investigated to guide AGIQA. In this letter, we propose vision-language consistency guided multi-modal prompt learning for blind AGIQA, dubbed CLIP-AGIQA. Specifically, we introduce learnable textual and visual prompts in language and vision branches of CLIP models, respectively. Moreover, we design a text-to-image alignment quality prediction task, whose learned vision-language consistency knowledge is used to guide the optimization of the above multi-modal prompts. Experimental results on two public AGIQA datasets demonstrate that the proposed method outperforms state-of-the-art quality assessment models.
Jun Fu 0007, Wei Zhou 0021, Qiuping Jiang, Hantao Liu, Guangtao Zhai
IEEE Signal Process. Lett.4
2024 Blind Image Quality Assessment via Adaptive Graph Attention
abstract
Recent advancements in blind image quality assessment (BIQA) are primarily propelled by deep learning technologies. While leveraging transformers can effectively capture long-range dependencies and contextual details in images, the significance of local information in image quality assessment can be undervalued. To address this challenging problem, we propose a novel feature enhancement framework tailored for BIQA. Specifically, we devise an Adaptive Graph Attention (AGA) module to simultaneously augment both local and contextual information. It not only refines the post-transformer features into an adaptive graph, facilitating local information enhancement, but also exploits interactions amongst diverse feature channels. The proposed technique can better reduce redundant information introduced during feature updates compared to traditional convolution layers, streamlining the self-updating process for feature maps. Experimental results show that our proposed model outperforms state-of-the-art BIQA models in predicting the perceived quality of images. The code of the model will be made publicly available.
Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.7
2024 Coarse- and Fine-Grained Fusion Hierarchical Network for Hole Filling in View Synthesis
abstract
Depth image-based rendering (DIBR) techniques play an essential role in free-viewpoint videos (FVVs), which generate the virtual views from a reference 2D texture video and its associated depth information. However, the background regions occluded by the foreground in the reference view will be exposed in the synthesized view, resulting in obvious irregular holes in the synthesized view. To this end, this paper proposes a novel coarse and fine-grained fusion hierarchical network (CFFHNet) for hole filling, which fills the irregular holes produced by view synthesis using the spatial contextual correlations between the visible and hole regions. CFFHNet adopts recurrent calculation to learn the spatial contextual correlation, while the hierarchical structure and attention mechanism are introduced to guide the fine-grained fusion of cross-scale contextual features. To promote texture generation while maintaining fidelity, we equip CFFHNet with a two-stage framework involving an inference sub-network to generate the coarse synthetic result and a refinement sub-network for refinement. Meanwhile, to make the learned hole-filling model better adaptable and robust to the "foreground penetration" distortion, we trained CFFHNet by generating a batch of training samples by adding irregular holes to the foreground and background connection regions of high-quality images. Extensive experiments show the superiority of our CFFHNet over the current state-of-the-art DIBR methods. The source code will be available at https://github.com/wgc-vsfm/view-synthesis-CFFHNet.
Guangcheng Wang, Kui Jiang, Ke Gu 0001, Hongyan Liu 0004, Hantao Liu, Wenjun Zhang 0001
IEEE Trans. Image Process.5
2024 Predicting Radiologists' Gaze With Computational Saliency Models in Mammogram Reading
abstract
Previous studies have shown that there is a strong correlation between radiologists' diagnoses and their gaze when reading medical images. The extent to which gaze is attracted by content in a visual scene can be characterised as visual saliency. There is a potential for the use of visual saliency in computer-aided diagnosis in radiology. However, little is known about what methods are effective for diagnostic images, and how these methods could be adapted to address specific applications in diagnostic imaging. In this study, we investigate 20 state-of-the-art saliency models including 10 traditional models and 10 deep learning-based models in predicting radiologists' visual attention while reading 196 mammograms. We found that deep learning-based models represent the most effective type of methods for predicting radiologists' gaze in mammogram reading; and that the performance of these saliency models can be significantly improved by transfer learning. In particular, an enhanced model can be achieved by pre-training the model on a large-scale natural image saliency dataset and then fine-tuning it on the target medical image dataset. In addition, based on a systematic selection of backbone networks and network architectures, we proposed a parallel multi-stream encoded model which outperforms the state-of-the-art approaches for predicting saliency of mammograms.
Jianxun Lou, Hanhe Lin, Philippa Young, Zelei Yang, Susan Cheng Shelmerdine, David Marshall 0001, Emiliano Spezi, Marco Palombo, Hantao Liu
IEEE Trans. Multim.10
2024 Going the Extra Mile in Face Image Quality Assessment: A Novel Database and Model
abstract
An accurate computational model for image quality assessment (IQA) benefits many vision applications, such as image filtering, image processing, and image generation. Although the study of face images is an important subfield in computer vision research, the lack of face IQA data and models limits the precision of current IQA metrics on face image processing tasks such as face superresolution, face enhancement, and face editing. To narrow this gap, in this article, we first introduce the largest annotated IQA database developed to date, which contains 20,000 human faces – an order of magnitude larger than all existing rated datasets of faces – of diverse individuals in highly varied circumstances. Based on the database, we further propose a novel deep learning model to accurately predict face image quality, which, for the first time, explores the use of generative priors for IQA. By taking advantage of rich statistics encoded in well pretrained off-the-shelf generative models, we obtain generative prior information and use it as latent references to facilitate blind IQA. The experimental results demonstrate both the value of the proposed dataset for face IQA and the superior performance of the proposed model.
Shaolin Su, Hanhe Lin, Vlad Hosu, Oliver Wiedemann, Jinqiu Sun, Yu Zhu 0004, Hantao Liu, Yanning Zhang 0001, Dietmar Saupe
IEEE Trans. Multim.7
2024 SSPNet: Predicting Visual Saliency Shifts
abstract
When images undergo quality degradation caused by editing, compression or transmission, their saliency tends to shift away from its original position. Saliency shifts indicate visual behaviour change and therefore contain vital information regarding perception of visual content and its distortions. Given a pristine image and its distorted format, we want to be able to detect saliency shifts induced by distortions. The resulting saliency shift map (SSM) can be used to identify the region and degree of visual distraction caused by distortions, and consequently to perceptually optimise image coding or enhancement algorithms. To this end, we first create a largest-of-its-kind eye-tracking database, comprising 60 pristine images and their associated 540 distorted formats viewed by 96 subjects. We then propose a computational model to predict the saliency shift map (SSM), utilising transformers and convolutional neural networks. Experimental results demonstrate that the proposed model is highly effective in detecting distortion-induced saliency shifts in natural images.
Huasheng Wang, Jianxun Lou, Xiaochang Liu, Hongchen Tan, Roger M. Whitaker, Hantao Liu
IEEE Trans. Multim.6
2023 Exploring Human Models of Innovation for Generative AI
Gualtiero Colombo 0001, Hantao Liu, Roger M. Whitaker
ICCC2
2023 Impact of Radiologist Experience on Medical Image Quality Perception
abstract
Low quality medical images can lead to inaccurate interpretation and diagnosis. Therefore, it is important to understand radiologists' perception of distortions in visual content. In this study, 12 radiologists with different degrees of experience scored MRI images of varying levels of quality. Statistical analyses were conducted to reveal the influence of the radiologists' experience on their perception of image quality. In scoring images of joints, brain and liver, the highly experienced radiologists gave significantly higher scores than less experienced radiologists. In scoring images of fetus and spine, there were no significant differences in scores between groups with different degrees of experience. No radiologists had expertise in breast images, and their experience in other anatomical areas did not significantly affect scoring of breast images. Overall, highly experienced radiologists gave higher scores for images with edge ghosting, plain ghosting or white noise than radiologists with less experience. The findings will provide a reference for determining or improving image quality standards in clinical practice.
Yueran Ma, Jean-Yves Tanguy, Padraig Corcoran, Hantao Liu
QoMEX5
2023 UAVs and Mobile Sensors Trajectories Optimization with Deep Learning Trained by Genetic Algorithm Towards Data Collection Scenario
Yuwen Pan, Yuanwang Yang, Hantao Liu, Wenzao Li
Mob. Networks Appl.3
2023 Deep Ordinal Regression Framework for No-Reference Image Quality Assessment
abstract
Due to the rapid development of deep learning techniques, no-reference image quality assessment (NR-IQA) has achieved significant improvement. NR-IQA aims to predict a real-valued variable for image quality, using the image in question as the sole input. Existing deep learning-based NR-IQA models are formulated as a regression problem and trained by minimising the mean squared error. The error measurement does not consider the relative ordering between different ratings on the quality scale, which consequently affects the efficacy of the model. To account for this problem, we reformulate NR-IQA learning as an ordinal regression problem and propose a simple yet effective framework using deep convolutional neural networks (DCNN) and Transformers. NR-IQA learning is achieved by a deep ordinal loss and using a soft ordinal inference to transform the predicted probabilities to a continuous variable for image quality. Experimental results demonstrate the superiority of our proposed NR-IQA model based on deep ordinal regression. In addition, this framework can be easily extended with various DCNN architectures to build advanced IQA models.
Huasheng Wang, Yulin Tu, Xiaochang Liu, Hongchen Tan, Hantao Liu
IEEE Signal Process. Lett.5
2023 Reduced-Reference Quality Assessment of Point Clouds via Content-Oriented Saliency Projection
abstract
Many dense 3D point clouds have been exploited to represent visual objects instead of traditional images or videos. To evaluate the perceptual quality of various point clouds, in this letter, we propose a novel and efficient Reduced-Reference quality metric for point clouds, which is based on Content-oriented sAliency Projection (RR-CAP). Specifically, we make the first attempt to simplify reference and distorted point clouds into projected saliency maps with a downsampling operation. Through this process, we tackle the issue of transmitting large-volume original point clouds to end-users for quality assessment. Then, motivated by the characteristics of the human visual system (HVS), the objective quality scores of distorted point clouds are produced by combining content-oriented similarity and statistical correlation measurements. Finally, extensive experiments are conducted on SJTU-PCQA and WPC databases. The experiment results demonstrate that our proposed algorithm outperforms existing reduced-reference and no-reference quality metrics, and significantly reduces the performance gap between state-of-the-art full-reference quality assessment methods. In addition, we show the performance variation of each proposed technical component by ablation tests.
Wei Zhou 0021, Guanghui Yue 0001, Ruizeng Zhang, Yipeng Qin, Hantao Liu
IEEE Signal Process. Lett.5
2023 A Perception-Aware Decomposition and Fusion Framework for Underwater Image Enhancement
abstract
This paper presents a perception-aware decomposition and fusion framework for underwater image enhancement (UIE). Specifically, a general structural patch decomposition and fusion (SPDF) approach is introduced. SPDF is built upon the fusion of two complementary pre-processed inputs in a perception-aware and conceptually independent image space. First, a raw underwater image is pre-processed to produce two complementary versions including a contrast-corrected image and a detail-sharpened image. Then, each of them is decomposed into three conceptually independent components, i.e., mean intensity, contrast, and structure, via structural patch decomposition (SPD). Afterwards, the corresponding components are fused using tailored strategies. The three components after fusion are finally integrated via inverting the decomposition to reconstruct a final enhanced underwater image. The main advantage of SPDF is that two complementary pre-processed images are fused in a perception-aware and conceptually independent image space and the fusions of different components can be performed separately without any interactions and information loss. Comprehensive comparisons on two benchmark datasets demonstrate that SPDF outperforms several state-of-the-art UIE algorithms qualitatively and quantitatively. Moreover, the effectiveness of SPDF is also verified on another two relevant tasks, i.e., low-light image enhancement and single image dehazing. The code will be made available soon.
Yaozu Kang, Qiuping Jiang, Chongyi Li, Wenqi Ren, Hantao Liu, Pengjun Wang
IEEE Trans. Circuits Syst. Video Technol.5
2023 EHNQ: Subjective and Objective Quality Evaluation of Enhanced Night-Time Images
abstract
Vision-based practical applications, such as consumer photography and automated driving systems, greatly rely on enhancing the visibility of images captured in night-time environments. For this reason, various image enhancement algorithms (EHAs) have been proposed. However, little attention has been given to the quality evaluation of enhanced night-time images. In this paper, we conduct the first dedicated exploration of the subjective and objective quality evaluation of enhanced night-time images. First, we build an enhanced night-time image quality (EHNQ) database, which is the largest of its kind so far. It includes 1,500 enhanced images generated from 100 real night-time images using 15 different EHAs. Subsequently, we perform a subjective quality evaluation and obtain subjective quality scores on the EHNQ database. Thereafter, we present an objective blind quality index for enhanced night-time images (BEHN). Enhanced night-time images usually suffer from inappropriate brightness and contrast, deformed structure, and unnatural colorfulness. In BEHN, we capture perceptual features that are highly relevant to these three types of corruptions, and we design an ensemble training strategy to map the extracted features into the quality score. Finally, we conduct extensive experiments on EHNQ and EAQA databases. The experimental and analysis results validate the performance of the proposed BEHN compared with the state-of-the-art approaches. Our EHNQ database is publicly available for download athttps://sites.google.com/site/xiangtaooo/.
Ying Yang 0019, Tao Xiang 0001, Shangwei Guo, Hantao Liu, Xiaofeng Liao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Study of Spatio-Temporal Modeling in Video Quality Assessment
abstract
Video quality assessment (VQA) has received remarkable attention recently. Most of the popular VQA models employ recurrent neural networks (RNNs) to capture the temporal quality variation of videos. However, each long-term video sequence is commonly labeled with a single quality score, with which RNNs might not be able to learn long-term quality variation well: What's the real role of RNNs in learning the visual quality of videos? Does it learn spatio-temporal representation as expected or just aggregating spatial features redundantly? In this study, we conduct a comprehensive study by training a family of VQA models with carefully designed frame sampling strategies and spatio-temporal fusion methods. Our extensive experiments on four publicly available in- the-wild video quality datasets lead to two main findings. First, the plausible spatio-temporal modeling module (i. e., RNNs) does not facilitate quality-aware spatio-temporal feature learning. Second, sparsely sampled video frames are capable of obtaining the competitive performance against using all video frames as the input. In other words, spatial features play a vital role in capturing video quality variation for VQA. To our best knowledge, this is the first work to explore the issue of spatio-temporal modeling in VQA.
Yuming Fang 0001, Zhaoqian Li, Jiebin Yan, Xiangjie Sui, Hantao Liu
IEEE Trans. Image Process.5
2023 Visibility and Distortion Measurement for No-Reference Dehazed Image Quality Assessment via Complex Contourlet Transform
abstract
Recently, most dehazed image quality assessment (DQA) methods have focused on estimating remaining haze and omitting distortion impact from the side effect of dehazing algorithms, which leads to their limited performance. Addressing this problem, we propose a method for learning both visibility and distortion-aware features no-reference (NR) dehazed image quality assessment (VDA-DQA). Visibility-aware features are exploited to characterize clarity optimization after dehazing, including the brightness-, contrast-, and sharpness-aware features extracted by the complex contourlet transform (CCT). Then, distortion-aware features are employed to measure the distortion artifacts of images, including the normalized histogram of the local binary pattern (LBP) from the reconstructed dehazed image and the statistics of the CCT subbands corresponding to the chroma and saturation map. Finally, all the above features are mapped into quality scores by support vector regression (SVR). Extensive experimental results on six public DQA datasets verify the superiority of the proposed VDA-DQA method in terms of consistency with subjective visual perception and outperform state-of-the-art methods.
Tuxin Guan, Ke Gu 0001, Hantao Liu, Yuhui Zheng, Xiaojun Wu 0001
IEEE Trans. Multim.4
2023 An Underwater Image Quality Assessment Metric
abstract
Various image enhancement algorithms are adopted to improve underwater images that often suffer from visual distortions. It is critical to assess the output quality of underwater images undergoing enhancement algorithms, and use the results to optimise underwater imaging systems. In our previous study, we created a benchmark for quality assessment of underwater image enhancement via subjective experiments. Building on the benchmark, this paper proposes a new objective metric that can automatically assess the output quality of image enhancement, namely UWEQM. By characterising specific underwater physics and relevant properties of the human visual system, image quality attributes are computed and combined to yield an overall metric. Experimental results show that the proposed UWEQM metric yields good performance in predicting image quality as perceived by human subjects.
Hantao Liu, Delu Zeng, Tao Xiang 0001, Leida Li, Ke Gu 0001
IEEE Trans. Multim.2
2023 Blind Dehazed Image Quality Assessment: A Deep CNN-Based Approach
abstract
Research on image dehazing has made the need for a suitable dehazed image quality assessment (DIQA) method even more urgent. The performance of existing DIQA methods heavily relies on handcrafted haze-related features. Since hazy images with uneven haze density distributions will result in uneven quality distributions after dehazing, the manually extracted feature expression is neither accurate nor robust. In this paper, we design a deep CNN-based DIQA method without a handcrafted feature requirement. Specifically, we propose a blind dehazed image quality assessment model (BDQM), which consists of three components: image preprocessing, a haze-related feature extraction network (HFNet), and an improved regression network (IRNet). In HFNet, we design a perceptual information enhancement (PIE) module to learn powerful feature representations and enhance network capability according to channel attention, multiscale convolution and residual concatenation. IRNet aims to aggregate all patch information for the quality prediction of the whole image, where the effect of inhomogeneous distortion from the dehazing procedure is attenuated via a specifically designed patch attention (PA) mechanism. Experimental results on benchmark datasets demonstrate the effectiveness and superiority of the proposed network architecture over state-of-the-art methods.
Tao Xiang 0001, Ying Yang 0019, Hantao Liu
IEEE Trans. Multim.4
2023 Semi-Supervised Authentically Distorted Image Quality Assessment With Consistency-Preserving Dual-Branch Convolutional Neural Network
abstract
Recently, convolutional neural networks (CNNs) have provided a favoured prospect for authentically distorted image quality assessment (IQA). For good performance, most existing CNN-based methods rely on a large amount of labeled data for training, which is time-consuming and cumbersome to collect. By simultaneously exploiting few labeled data and many unlabeled data, we make a pioneering attempt to propose a semi-supervised framework (termed SSLIQA) with consistency-preserving dual-branch CNN for authentically distorted IQA in this paper. The proposed SSLIQA introduces a consistency-preserving strategy and transfers two kinds of consistency knowledge from the teacher branch to the student branch. Concretely, SSLIQA utilizes the sample prediction consistency to train the student to mimic output activations of individual examples represented by the teacher. Considering that subjects often refer to previous analogous cases to make scoring decisions, SSLIQA computes the semantic relation among different samples in a batch and encourages the consistency of sample semantic relation between two branches to explore extra quality-related information. Benefiting from the consistency-preserving strategy, we can exploit numerous unlabeled data to improve network's effectiveness and generalization. Experimental results on three authentically distorted IQA databases show that the proposed SSLIQA is stably effective under different student-teacher combinations and different labeled-to-unlabeled data ratios. In addition, it points out a new way on how to achieve higher performance with a smaller network.
Guanghui Yue 0001, Leida Li, Tianwei Zhou, Hantao Liu, Tianfu Wang 0001
IEEE Trans. Multim.5
2023 Multimodal Sentiment Analysis With Image-Text Interaction Network
abstract
More and more users are getting used to posting images and text on social networks to share their emotions or opinions. Accordingly, multimodal sentiment analysis has become a research topic of increasing interest in recent years. Typically, there exist affective regions that evoke human sentiment in an image, which are usually manifested by corresponding words in people's comments. Similarly, people also tend to portray the affective regions of an image when composing image descriptions. As a result, the relationship between image affective regions and the associated text is of great significance for multimodal sentiment analysis. However, most of the existing multimodal sentiment analysis approaches simply concatenate features from image and text, which could not fully explore the interaction between them, leading to suboptimal results. Motivated by this observation, we propose a new image-text interaction network (ITIN) to investigate the relationship between affective image regions and text for multimodal sentiment analysis. Specifically, we introduce a cross-modal alignment module to capture region-word correspondence, based on which multimodal features are fused through an adaptive cross-modal gating module. Moreover, considering the complementary role of context information on sentiment analysis, we integrate the individual-modal contextual feature representations for achieving more reliable prediction. Extensive experimental results and comparisons on public datasets demonstrate that the proposed model is superior to the state-of-the-art methods.
Tong Zhu 0003, Leida Li, Jufeng Yang, Sicheng Zhao, Hantao Liu, Jiansheng Qian
IEEE Trans. Multim.5
2022 Predicting Radiologist Attention During Mammogram Reading with Deep and Shallow High-Resolution Encoding
abstract
Radiologists’ eye-movement during diagnostic image reading reflects their personal training and experience, which means that their diagnostic decisions are related to their perceptual processes. For training, monitoring, and performance evaluation of radiologists, it would be beneficial to be able to automatically predict the spatial distribution of the radiologist’s visual attention on the diagnostic images. The measurement of visual saliency is a well-studied area that allows for prediction of a person’s gaze attention. However, compared with the extensively studied natural image visual saliency (in free viewing tasks), the saliency for diagnostic images is less studied; there could be fundamental differences in eye-movement behaviours between these two domains. Most current saliency prediction models have been optimally developed for natural images, which could lead them to be less adept at predicting the visual attention of radiologists during the diagnosis. In this paper, we propose a method specifically for automatically capturing the visual attention of radiologists during mammogram reading. By adopting high-resolution image representations from both deep and shallow encoders, the proposed method avoids potential detail losses and achieves superior results on multiple evaluation metrics in a large mammogram eye-movement dataset.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Young Yang, Susan Cheng Shelmerdine, Hantao Liu
ICIP7
2022 Analysis of Video Quality Induced Spatio-Temporal Saliency Shifts
abstract
Human viewers’ eye movements reflect their perceptual responses to visual signals. Previous research has shown that distortions in videos cause spatio-temporal gaze shifts, which means gaze behaviour is related to video quality perception. It would be highly beneficial to understand gaze behaviour of viewing videos of varying perceived quality. However, little is known about the interactions between gaze, video content and distortions. In this paper, based on our eye-tracking database for video quality (SVQ160), we perform systematic analyses to reveal the impact of video content (VC) and time order (TO) on gaze shifts. Findings and quantitative methods for gaze behaviour can be used to develop advanced video quality metrics and video processing algorithms.
Xinbo Wu, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu
ICIP5
2022 Text's Armor: Optimized Local Adversarial Perturbation Against Scene Text Editing Attacks
abstract
Deep neural networks (DNNs) have shown their powerful capability in scene text editing (STE). With carefully designed DNNs, one can alter texts in a source image with other ones while maintaining their realistic look. However, such editing tools provide a great convenience for criminals to falsify documents or modify texts without authorization. In this paper, we propose to actively defeat text editing attacks by designing invisible "armors" for texts in the scene. We turn the adversarial vulnerability of DNN-based STE into strength and design local perturbations (i.e., "armors") specifically for texts using an optimized normalization strategy. Such local perturbations can effectively mislead STE attacks without affecting the perceptibility of scene background. To strengthen our defense capabilities, we systemically analyze and model STE attacks and provide a precise defense method to defeat attacks on different editing stages. We conduct both subjective and objective experiments to show the superior of our optimized local adversarial perturbation against state-of-the-art STE attacks. We also evaluate the portrait and landscape transferability of our perturbations.
Tao Xiang 0001, Hangcheng Liu, Shangwei Guo, Hantao Liu, Tianwei Zhang 0004
ACM Multimedia4
2022 Dual-stream Self-attention Network for Image Captioning
abstract
Self-attention based encoder-decoder models achieve dominant performance in image captioning. However, most existing image captioning models (ICMs) only focus on modeling the relation between spatial tokens, while channel-wise attention is neglected for getting visual representation. Considering that different channels of visual representation usually denote different visual objects, it may lead to poor performance in terms of object and attribute words in the captioning sentences generated by the ICMs. In this paper, we propose a novel dual-stream self-attention module (DSM) to alleviate the above issue. Specifically, we propose a parallel self-attention based module that simultaneously encodes visual information from the spatial and channel dimensions. Besides, to obtain channel-wise visual features effectively and efficiently, we introduce a group self-attention block with linear computational complexity. To validate the effectiveness of our model, we conduct extensive experiments on the standard IC benchmarks including MSCOCO and Flickr30k. Without bells and whistles, the proposed model performs new SOTAs containing 135.4 CIDEr score on MSCOCO and 70.8 CIDEr score on Flickr30k.
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Wenying Wen, Hantao Liu
VCIP5
2022 TranSalNet: Towards perceptually relevant visual saliency prediction
abstract
Convolutional neural networks (CNNs) have significantly advanced computational modelling for saliency prediction. However, accurately simulating the mechanisms of visual attention in the human cortex remains an academic challenge. It is critical to integrate properties of human vision into the design of CNN architectures, leading to perceptually more relevant saliency prediction. Due to the inherent inductive biases of CNN architectures, there is a lack of sufficient long-range contextual encoding capacity. This hinders CNN-based saliency models from capturing properties that emulate viewing behaviour of humans. Transformers have shown great potential in encoding long-range information by leveraging the self-attention mechanism. In this paper, we propose a novel saliency model that integrates transformer components to CNNs to capture the long-range contextual visual information. Experimental results show that the transformers provide added value to saliency prediction, enhancing its perceptual relevance in the performance. Our proposed saliency model using transformers has achieved superior results on public benchmarks and competitions for saliency prediction models. The source code of our proposed saliency model TranSalNet is available at: https://github.com/LJOVO/TranSalNet.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Dietmar Saupe, Hantao Liu
Neurocomputing5
2022 Study of Subjective and Objective Quality Assessment of Night-Time Videos
abstract
With the widespread usage of video capture devices and social media videos, videos are dominating the multimedia landscape. There is an emerging need for video quality assessment (VQA) that forms the backbone of advanced video systems. Night-time videos play an important role in user capturing, hence being able to accurately assess their quality is critical. However, the characteristics of night-time videos differ from those of general in-capture videos; and VQA algorithms that have been developed for general-purpose videos cannot accurately assess the quality of night-time videos. Research is needed to gain a better understanding of how humans perceive the quality of night-time videos, and use this new understanding to develop reliable VQA algorithms. To this end, we construct a large-scale night-time VQA database, namely Mobile In-capture Night-time Database for Video Quality (MIND-VQ), containing 1181 night-time videos, 435 subjects, and over 130000 opinion scores. We perform thorough analyses to reveal subjective quality assessment behaviors of night-time videos. Furthermore, we propose a new VQA model, namely Visibility-based Night-time Video Quality Assessment Network, VINIA. Spatial and temporal visibility-aware components are characterized to reflect properties of human perception of night-time VQA task. A series of experiments are conducted to compare our VINIA with other existing VQA algorithms using our new MIND-VQ database and other public VQA databases. Experimental results show that our subjective VQA database provides new insights and our new VINIA model achieves superior performance in accessing night-time video quality.
Xiaodi Guan, Fan Li 0003, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Blind Image Quality Assessment for Authentic Distortions by Intermediary Enhancement and Iterative Training
abstract
With the boom of deep neural networks, blind image quality assessment (BIQA) has achieved great processes. However, the current BIQA metrics are limited when evaluating low-quality images as compared to medium-quality and high-quality images, which restricts their applications in real world problems. In this paper, we first identify that two challenges caused by distribution shift and long-tailed distribution lead to the compromised performance on low-quality images. Then, we propose an intermediary enhancement-based bilateral network with iterative training strategy for solving these two challenges. Drawing on the experience of transitive transfer learning, the proposed metric adaptively introduces enhanced intermediary images to transfer more information to low-quality images for mitigating the distribution shift. Our metric also adopts an iterative training strategy to deal with the long-tailed distribution. This strategy decouples feature extraction and score regression for better representation learning and regressor training. It not only transfers the knowledge learned from the earlier stage to the latter stage, but also makes the model pay more attention to long-tailed low-quality images. We conduct extensive experiments on five authentically distorted image quality datasets. The results show that our metric significantly improves the evaluating performance on low-quality images and delivers state-of-the-art intra-dataset results. During generalization tests, our metric also achieves the best cross-dataset performance.
Tianshu Song, Leida Li, Pengfei Chen 0003, Hantao Liu, Jiansheng Qian
IEEE Trans. Circuits Syst. Video Technol.4
2022 Active Vision for Deep Visual Learning: A Unified Pooling Framework
abstract
Convolutional neural networks (CNNs) can be generally regarded as learning-based visual systems for computer vision tasks. By imitating the operating mechanism of the human visual system (HVS), CNNs can even achieve better results than human beings in some visual tasks. However, they are primary when compared to the HVS for the reason that the HVS has the ability of active vision to promptly analyze and adapt to specific tasks. In this article, a new unified pooling framework is proposed and a series of pooling methods are designed based on the framework to implement active vision to CNNs. In addition, an active selection pooling (ASP) is put forward to reorganize the existing and newly proposed pooling methods. The CNN models with an ASP tend to have a behavior of focus selection according to tasks during the training process, which acts extremely similar to the HVS.
Nan Guo 0006, Ke Gu 0001, Junfei Qiao 0001, Hantao Liu
IEEE Trans. Ind. Informatics4
2022 Single Image Super-Resolution Quality Assessment: A Real-World Dataset, Subjective Studies, and an Objective Metric
abstract
Numerous single image super-resolution (SISR) algorithms have been proposed during the past years to reconstruct a high-resolution (HR) image from its low-resolution (LR) observation. However, how to fairly compare the performance of different SISR algorithms/results remains a challenging problem. So far, the lack of comprehensive human subjective study on large-scale real-world SISR datasets and accurate objective SISR quality assessment metrics makes it unreliable to truly understand the performance of different SISR algorithms. We in this paper make efforts to tackle these two issues. Firstly, we construct a real-world SISR quality dataset (i.e., RealSRQ) and conduct human subjective studies to compare the performance of the representative SISR algorithms. Secondly, we propose a new objective metric, i.e., KLTSRQA, based on the Karhunen-Loéve Transform (KLT) to evaluate the quality of SISR images in a no-reference (NR) manner. Experiments on our constructed RealSRQ and the latest synthetic SISR quality dataset (i.e., QADS) have demonstrated the superiority of our proposed KLTSRQA metric, achieving higher consistency with human subjective scores than relevant existing NR image quality assessment (NR-IQA) metrics. The dataset and the code will be made available at https://github.com/Zhentao-Liu/RealSRQ-KLTSRQA.
Qiuping Jiang, Ke Gu 0001, Feng Shao 0001, Xinfeng Zhang 0001, Hantao Liu, Weisi Lin
IEEE Trans. Image Process.6
2022 Underwater Image Quality Assessment: Subjective and Objective Methods
abstract
Underwater image enhancement plays a critical role in marine industry. Various algorithms are applied to enhance underwater images, but their performance in terms of perceptual quality has been little studied. In this paper, we investigate five popular enhancement algorithms and their output image quality. To this end, we have created a benchmark, including images enhanced by different algorithms and ground truth image quality obtained by human perception experiments. We statistically analyse the impact of various enhancement algorithms on the perceived quality of underwater images. Also, the visual quality provided by these algorithms is evaluated objectively, aiming to inform the development of objective metrics for automatic assessment of the quality for underwater image enhancement. The image quality benchmark and its objective metric are made publicly available.
Shuangyin Liu, Delu Zeng, Hantao Liu
IEEE Trans. Multim.5
2021 A Metric For Quantifying Image Quality Induced Saliency Variation
abstract
Saliency plays an important role in the area of image quality assessment. Image distortions cause shift/redistribution of saliency from its original places. There is a need to be able to measure such distortion-included saliency variation (DSV), so that the use of saliency can be optimised for automated image quality assessment. Effort has been made in our previous study to build a benchmark for the measurement of DSV through subjective testing. In this paper, we demonstrate that exiting similarity measures are unhelpful for the quantification of DSV. Thus, we propose a new metric for DSV combining local and global measures using convex optimization. The experimental results show that our proposed metric can accurately quantify saliency variation.
Delu Zeng, Hantao Liu
ICIP4
2021 Study of Saccadic Eye Movements in Diagnostic Imaging
abstract
Eye movements reflect the visual process of humans’ perception and cognition. In the field of medical imaging, the diagnosis rendered by radiologists is closely related to their eye movements when reading radiological images. It is beneficial to study the eye movements of radiologists to improve the diagnostic performance. However, existing studies are mainly focused on the radiologists’ fixations but rarely on their saccade patterns. Moreover, these studies are almost based on limited datasets. In this paper, we present a quantitative study of the gaze behavior of radiologists from the perspective of saccade patterns on a large-scale dataset. The dataset comprises of the eye-tracking data of 10 expert radiologists reading 196 mammograms. By analyzing the saccade amplitude, direction, and bias of radiologists, we found that radiologists have specific saccade patterns in image reading and the saccade patterns are significantly affected by the different reading phases, working experience, and orientations of the mammograms.
Jianxun Lou, Philippa Young, Hantao Liu
ICIP5
2021 PRNet: A Progressive Recovery Network for Revealing Perceptually Encrypted Images
abstract
Perceptual encryption is an efficient way of protecting image content by only selectively encrypting a portion of significant data in plain images. Existing security analysis of perceptual encryption usually resorts to traditional cryptanalysis techniques, which require heavy manual work and strict prior knowledge of encryption schemes. In this paper, we introduce a new end-to-end method of analyzing the visual security of perceptually encrypted images, without any manual work or knowing any prior knowledge of the encryption scheme. Specifically, by leveraging convolutional neural networks (CNNs), we propose a progressive recovery network (PRNet) to recover visual content from perceptually encrypted images. Our PRNet is stacked with several dense attention recovery blocks (DARBs), where each DARB contains two branches: feature extraction branch and image recovery branch. These two branches cooperate to rehabilitate more detailed visual information and generate efficient feature representation via densely connected structure and dual-saliency mechanism. We conduct extensive experiments to demonstrate that PRNet works on different perceptual encryption schemes with different settings, and the results show that PRNet significantly outperforms the state-of-the-art CNN-based image restoration methods.
Tao Xiang 0001, Ying Yang 0019, Shangwei Guo, Hangcheng Liu, Hantao Liu
ACM Multimedia5
2021 TTL-IQA: Transitive Transfer Learning Based No-Reference Image Quality Assessment
abstract
Image quality assessment (IQA) based on deep learning faces the overfitting problem due to limited training samples available in existing IQA databases. Transfer learning is a plausible solution to the problem, in which the shared features derived from the large-scale Imagenet source domain could be transferred from the original recognition task to the intended IQA task. However, the Imagenet source domain and the IQA target domain as well as their corresponding tasks are not directly related. In this paper, we propose a new transitive transfer learning method for no-reference image quality assessment (TTL-IQA). First, the architecture of the multi-domain transitive transfer learning for IQA is developed to transfer the Imagenet source domain to the auxiliary domain, and then to the IQA target domain. Second, the auxiliary domain and the auxiliary task are constructed by a new generative adversarial network based on distortion translation (DT-GAN). Furthermore, a TTL network of the semantic features transfer (SFTnet) is proposed to optimize the shared features for the TTL-IQA. Experiments are conducted to evaluate the performance of the proposed method on various IQA databases, including the LIVE, TID2013, CSIQ, LIVE multiply distorted and LIVE challenge. The results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, our proposed method demonstrates a strong generalization ability.
Fan Li 0003, Hantao Liu
IEEE Trans. Multim.3
2020 CUID: A New Study Of Perceived Image Quality And Its Subjective Assessment
abstract
Research on image quality assessment (IQA) remains limited mainly due to our incomplete knowledge about human visual perception. Existing IQA algorithms have been designed or trained with insufficient subjective data with a small degree of stimulus variability. This has led to challenges for those algorithms to handle complexity and diversity of real-world digital content. Perceptual evidence from human subjects serves as a grounding for the development of advanced IQA algorithms. It is thus critical to acquire reliable subjective data with controlled perception experiments that faithfully reflect human behavioural responses to distortions in visual signals. In this paper, we present a new study of image quality perception where subjective ratings were collected in a controlled lab environment. We investigate how quality perception is affected by a combination of different categories of images and different types and levels of distortions. The database will be made publicly available to facilitate calibration and validation of IQA algorithms.
Lucie Lévêque, Kenneth Dasalla, Leida Li, Hantao Liu
ICIP8
2020 Deep Learning VS. Traditional Algorithms for Saliency Prediction of Distorted Images
abstract
Saliency has been widely studied in relation to image quality assessment (IQA). The optimal use of saliency in IQA metrics, however, is nontrivial and largely depends on whether saliency can be accurately predicted for images containing various distortions. Although tremendous progress has been made in saliency modelling, very little is known about whether and to what extent state-of-the-art methods are beneficial for saliency prediction of distorted images. In this paper, we analyse the ability of deep learning versus traditional algorithms in predicting saliency, based on an IQA-aware saliency benchmark, the SIQ288 database. Building off the variations in model performance, we make recommendations for model selections for IQA applications.
Hanhe Lin, Dietmar Saupe, Hantao Liu
ICIP5
2020 Deep feature importance awareness based no-reference image quality prediction
Fan Li 0003, Hantao Liu
Neurocomputing3
2020 Special Issue on Advances in Statistical Methods-based Visual Quality Assessment
Fei Zhou 0001, Wenming Yang, Xinbo Gao 0001, Hantao Liu, Rui Zhu 0006, Jing-Hao Xue
Signal Process. Image Commun.4
2020 A Metric for Video Blending Quality Assessment
abstract
We propose an objective approach to assess the quality of video blending. Blending is a fundamental operation in video editing, which can smooth the intensity changes of relevant regions. However blending also generates artefacts such as bleeding and ghosting. To assess the quality of the blended videos, our approach considers the illuminance consistency as a positive aspect while regard the artefacts as a negative aspect. Temporal coherence between frames is also considered. We evaluate our metric on a video blending dataset where the results of subjective evaluation are available. Experimental results validate the effectiveness of our proposed metric, and shows that this metric gives superior performance over existing video quality metrics.
Zhe Zhu, Hantao Liu, Jiaming Lu, Shi-Min Hu 0001
IEEE Trans. Image Process.2
2019 The Effect of Spatio-temporal Inconsistency on the Subjective Quality Evaluation of Omnidirectional Videos
abstract
With the development of immersive media technologies, omnidirectional video services have been launched in many fields. Conducting subjective quality evaluation research becomes a crucial step to benchmark and ensure the quality of omnidirectional video services. As omnidirectional videos record spherical visual scenes that are broader than the visual field of human eyes, the quality scores rated by different observers are based on individual spatio-temporal viewing experience. The potential spatio-temporal inconsistency between observers may impact the reliability of subjective quality evaluation and thus challenge existing experimental methodologies. In this paper, we focus on investigating the effect of spatial-temporal inconsistency on the subjective quality evaluation of omnidirectional videos. A systematic quality evaluation experiment was designed with various viewing methods involved. Experimental results showed that the spatio-temporal inconsistency has a significant impact on the reliability of subjective quality results and the impact is strongly determined by the viewing method. We intend to provide recommendations with respect to the subjective quality evaluation of omnidirectional videos.
Wei Zhang 0072, Wenjie Zou, Fuzheng Yang 0001, Lucie Lévêque, Hantao Liu
ICASSP5
2019 Subjective Assessment of Image Quality Induced Saliency Variation
abstract
Our previous study has shown that image distortions cause saliency distraction, and that visual saliency of a distorted image differs from that of its distortion-free reference. Being able to measure such distortion-induced saliency variation (DSV) significantly benefits algorithms for automated image quality assessment. Methods of quantifying DSV, however, remain unexplored due to the lack of a benchmark. In this paper, we build a benchmark for the measurement of DSV through a subjective study. Sixteen experts in computer vision were asked to compare saliency maps of distorted images to the corresponding saliency maps of the original images. All saliency maps were rendered from ground truth human fixations. A statistical analysis is performed to reveal the behaviours and properties of human assessment of the saliency variation. The benchmark is made publicly available to the research community.
Lucie Lévêque, Wei Zhang 0072, Hantao Liu
ICIP3
2019 An Eye-Tracking Database of Video Advertising
abstract
Reliably predicting where people look in images and videos remains challenging and requires substantial eye-tracking data to be collected and analysed for various applications. In this paper, we present an eye-tracking study where twenty-eight participants viewed forty still scenes of video advertising. First, we analyse human attentional behaviour based on gaze data. Then, we evaluate to what extent a machine - saliency model - can predict human behaviour. Experimental results show that there is a significant gap between human and machine in visual saliency. The resulting eye-tracking data would benefit the development of saliency models for video advertising or other relevant applications. The eye-tracking data are made publicly available to the research community.
Lucie Lévêque, Hantao Liu
ICIP2
2019 A Comparative Study of DNN-Based Models for Blind Image Quality Prediction
abstract
Recently, deep learning methods have gained substantial attention in the research community and have proven useful for blind image quality assessment (BIQA). Although previous study of deep neural networks (DNN) methods is presented, some novelty methods, which are recently proposed, are not summarized. In this paper, we provide a comparative study on the application of DNN methods for BIQA. First, we systematically analyze the existing DNN-based quality assessment methods. Then, we compare the predictive performance of various methods in synthetic and authentic databases, providing important information that can help understand the underlying properties between different methods. Finally, we describe some emerging challenges in designing and training DNN-based BIQA, along with few directions that are worth further investigations in the future.
Fan Li 0003, Hantao Liu
ICIP3
2019 International Comparison of Radiologists' Assessment of the Perceptual Quality of Medical Ultrasound Video
abstract
Telemedicine can provide timely and high-quality clinical health care from a distance, improving access to and delivery of medical services in resource-poor settings. It can also save lives in situations of emergency. In many circumstances, the success of telemedicine practice heavily relies on the transmission of medical videos over large distances. However, video communication systems are prone to distortion in visual signals, affecting the task performance and thus putting patients at risk. It is critical to understand how practitioners perceive the quality of visual media and use such knowledge to improve clinical practice in telemedicine. In this paper, we investigate the hypothesis that visual quality perception varies between clinicians who work in different practice settings. To evaluate this hypothesis, we performed a subjective experiment where French and Chinese radiologists were asked to rate the quality of ultrasound videos compressed using different compression configurations. The results show that the way the perceived quality changes with the compression configuration is consistent among the both settings studied, however, French radiologists were more bothered by the compression artifacts. The findings can help inform future studies to develop tailored telemedicine systems for specific settings or individuals.
Lucie Lévêque, Wei Zhang 0072, Hantao Liu
QoMEX3
2019 No-reference quality assessment for contrast-distorted images based on multifaceted statistical representation of structure
Yu Zhou 0009, Leida Li, Hancheng Zhu, Hantao Liu, Shiqi Wang 0001, Yao Zhao 0001
J. Vis. Commun. Image Represent.4
2019 A statistical evaluation of eye-tracking data of screening mammography: Effects of expertise and experience on image reading
Lucie Lévêque, Baptiste Vande Berg, Hilde Bosmans, Lesley Cockmartin, Machteld Keupers, Chantal Van Ongeval, Hantao Liu
Signal Process. Image Commun.7
2019 Pairwise-Comparison-Based Rank Learning for Benchmarking Image Restoration Algorithms
abstract
Image restoration has attracted substantial attention recently and many image restoration algorithms have been proposed for restoring latent clear images from degraded images. However, determining how to objectively evaluate the performances of these algorithms remains an open problem, which may hinder the further development of advanced image restoration techniques. Most image restoration-quality metrics are designed for specific restoration applications; hence, their generalization ability is limited. For benchmarking image restoration algorithms, the ranking of restored images that are generated via various algorithms, is the most heavily considered factor. Inspired by this, this paper presents a pairwise-comparison-based rank learning framework for benchmarking the performances of image restoration algorithms, which focuses on the relative quality ranking of restored images. Under the proposed framework, we further propose a general image restoration quality metric by integrating quality-aware features in both the spatial and frequency domains. The proposed metric exhibits good generalization performance, and it is applicable to various restoration applications. The results of extensive experiments that were conducted on eight public databases of five restoration scenarios demonstrate the superior performance of the proposed method over the existing quality metrics. Moreover, the proposed framework is used to improve the existing quality metrics for benchmarking image restoration algorithms and highly encouraging results are obtained.
Bo Hu 0008, Leida Li, Hantao Liu, Weisi Lin, Jiansheng Qian
IEEE Trans. Multim.3
2019 No-Reference Quality Evaluator of Transparently Encrypted Images
abstract
In past years, various encrypted algorithms have been proposed to fully or partially protect the multimedia content in view of practical applications. In the context of digital TV broadcasting, transparent encryption only protects partial content and fulfills both security and quality requirements. To date, only a few reference-based works have been reported to evaluate the quality of transparently encrypted images. However, these works are incapable of reference-unavailable conditions. In this paper, we conduct the first attempt that proposes a novel quality evaluator in the absence of reference images. The key strategy of the proposed metric lies in extracting features by considering the motivation of transparently encrypted images. Specifically, given that encrypted images prevent content from being easily recognized, several features, including correlation coefficient, information entropy, and intensity statistic, are preliminarily extracted to estimate visual recognizability. Meanwhile, considering that encrypted images are avoided since they are of extremely low quality, we also capture many features to measure the distortions on multiple quality-sensitive image attributes, such as naturalness, structure, and texture. Finally, the quality evaluator is built by bridging all extracted features and corresponding quality scores via a regression module. Experimental results demonstrate that the proposed method is superior to the mainstream no-reference quality evaluation methods designed for synthetically distorted images and possesses a close approximation to state-of-the-art reference-based methods designed for encrypted images.
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Hantao Liu
IEEE Trans. Multim.5
2018 On the Subjective Assessment of the Perceived Quality of Medical Images and Videos
abstract
Medical professionals are viewing an increasing number of images and videos in their clinical routine. However, various types of distortions can affect medical imaging data, and therefore impact the viewers' experienced quality and their clinical practice. Thus it is necessary to quantify this impact and understand how the viewers, i.e., medical experts, perceive the quality of (distorted) images and videos. In this paper, we present an up-to-date review of the methodologies used in the literature for the subjective quality assessment of medical images and videos and discuss their merits and drawbacks depending on the use case.
Lucie Lévêque, Hantao Liu, Sabina Barakovic, Jasmina Barakovic, Maria G. Martini, Meriem Outtas, Lu Zhang 0037, Asli Kumcu, Ljiljana Platisa, Rafael Rodrigues, António M. G. Pinheiro, Athanassios N. Skodras
QoMEX2
2018 A Saliency Dispersion Measure for Improving Saliency-Based Image Quality Metrics
abstract
Objective image quality metrics (IQMs) potentially benefit from the addition of visual saliency. However, challenges to optimizing the performance of saliency-based IQMs remain. A previous eye-tracking study has shown that gaze is concentrated in fewer places in images with highly salient features than in images lacking salient features. From this, it can be inferred that the former are more likely to benefit from adding a saliency term to an IQM. To understand whether these ideas still hold when using computational saliency instead of eye-tracking data, we first conducted a statistical evaluation using 15 state-of-the-art saliency models and 10 well-known IQMs. We then used the results to devise an algorithm, which adaptively incorporates saliency in IQMs for natural scenes, based on saliency dispersion. Experimental results demonstrate that this can give significant improvements.
Wei Zhang 0072, Ralph R. Martin, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.3
2018 A Comparative Study of Algorithms for Realtime Panoramic Video Blending
abstract
Unlike image blending algorithms, video blending algorithms have been little studied. In this paper, we investigate 6 popular blending algorithms-feather blending, multi-band blending, modified Poisson blending, mean value coordinate blending, multi-spline blending and convolution pyramid blending. We consider their application to blending realtime panoramic videos, a key problem in various virtual reality tasks. To evaluate the performances and suitabilities of the 6 algorithms for this problem, we have created a video benchmark with several videos captured under various conditions. We analyze the time and memory needed by the above 6 algorithms, for both CPU and GPU implementations (where readily parallelizable). The visual quality provided by these algorithms is also evaluated both objectively and subjectively. The video benchmark and algorithm implementations are publicly available1.
Zhe Zhu, Jiaming Lu, Minxuan Wang, Song-Hai Zhang, Ralph R. Martin, Hantao Liu, Shi-Min Hu 0001
IEEE Trans. Image Process.6
2017 Video quality perception in telesurgery
abstract
Telesurgery enables an expert surgeon to assist a remote surgeon during a surgical intervention, which benefits patient care in resource-poor settings. In reality, videos of surgical procedures are compressed and transmitted over large distances in real time and, therefore, are subject to a wide variety of distortions. These distortions degrade the quality of videos and potentially affect the performance of the surgeons. Very little work has been carried out on human perception of video quality in the context of telesurgery. In this paper, we investigate the impact of video compression on the perceived quality of surgical videos. We designed and performed a psychophysical experiment where surgeons rated the quality of surgical videos distorted with two different compression schemes at various compression ratios. Experimental results demonstrate that the impact of video content and compression strategy on the perceived quality is statistically significant.
Lucie Lévêque, Hantao Liu, Christine Cavaro-Ménard, Yongqiang Cheng 0001, Patrick Le Callet
MMSP2
2017 Learning picture quality from visual distraction: Psychophysical studies and computational models
Wei Zhang 0072, Hantao Liu
Neurocomputing2
2017 Study of Saliency in Objective Video Quality Assessment
abstract
Reliably predicting video quality as perceived by humans remains challenging and is of high practical relevance. A significant research trend is to investigate visual saliency and its implications for video quality assessment. Fundamental problems regarding how to acquire reliable eye-tracking data for the purpose of video quality research and how saliency should be incorporated in objective video quality metrics (VQMs) are largely unsolved. In this paper, we propose a refined methodology for reliably collecting eye-tracking data, which essentially eliminates bias induced by each subject having to view multiple variations of the same scene in a conventional experiment. We performed a large-scale eye-tracking experiment that involved 160 human observers and 160 video stimuli distorted with different distortion types at various degradation levels. The measured saliency was integrated into several best known VQMs in the literature. With the assurance of the reliability of the saliency data, we thoroughly assessed the capabilities of saliency in improving the performance of VQMs, and devised a novel approach for optimal use of saliency in VQMs. We also evaluated to what extent the state-of-the-art computational saliency models can improve VQMs in comparison to the improvement achieved by using "ground truth" eye-tracking data. The eye-tracking database is made publicly available to the research community.
Wei Zhang 0072, Hantao Liu
IEEE Trans. Image Process.2
2017 Toward a Reliable Collection of Eye-Tracking Data for Image Quality Research: Challenges, Solutions, and Applications
abstract
Image quality assessment potentially benefits from the addition of visual attention. However, incorporating aspects of visual attention in image quality models by means of a perceptually optimized strategy is largely unexplored. Fundamental challenges, such as how visual attention is affected by the concurrence of visual signals and their distortions; whether visual attention affected by distortion or that driven by the original scene only should be included in an image quality model; and how to select visual attention models for the image quality application context, remain. To shed light on the above unsolved issues, designing and performing eye-tracking experiments are essential. Collecting eye-tracking data for the purpose of image quality study is so far confronted with a bias due to the involvement of stimulus repetition. In this paper, we propose a new experimental methodology to eliminate such inherent bias. This allows obtaining reliable eye-tracking data with a large degree of stimulus variability. In fact, we first conducted 5760 eye movement trials that included 160 human observers freely viewing 288 images of varying quality. We then made use of the resulting eye-tracking data to provide insights into the optimal use of visual attention in image quality research. The new eye-tracking data are made publicly available to the research community.
Wei Zhang 0072, Hantao Liu
IEEE Trans. Image Process.2
2016 Benchmarking state-of-the-art visual saliency models for image quality assessment
abstract
A significant current research trend in image quality assessment is to investigate the added value of visual attention aspects. Previous approaches mainly focused on adopting a specific saliency model to improve a specific image quality metric (IQM). It is still not known yet which of the existing saliency models is generally applicable in IQMs; which of the IQMs can profit most/least from the addition of saliency; and how this improvement depends on the saliency model used and the IQM targeted. In this paper, a large-scale benchmark study is conducted to assess the capabilities and limitations of the state-of-the-art saliency models in the context of IQMs. The study provides guidance for the application of saliency models in IQMs, in terms of the effect of saliency model dependency, IQM dependency, and image distortion dependency.
Wei Zhang 0072, Xiaojie Zha, Hantao Liu
ICASSP4
2016 Saliency in objective video quality assessment: What is the ground truth?
abstract
Finding ways to be able to objectively and reliably assess video quality as would be perceived by humans has become a pressing concern in the multimedia community. To enhance the performance of video quality metrics (VQMs), a research trend is to incorporate visual saliency aspects. Existing approaches have focused on utilizing a computational saliency model to improve a VQM. Since saliency models still remain limited in predicting where people look in videos, the benefits of inclusion of saliency in VQMs may heavily depend on the accuracy of the saliency model used. To gain an insight into the actual added value of saliency in VQMs, ground truth saliency obtained from eye-tracking instead of computational saliency is an essential prerequisite. However, collecting eye-tracking data within the context of video quality is confronted with a bias due to the involvement of massive stimulus repetition. In this paper, we introduce a new experimental methodology to alleviate such potential bias and consequently, to be able to deliver reliable intended data. We recorded eye movements from 160 human observers while they freely viewed 160 video stimuli distorted with different distortion types at various degradation levels. We analyse the extent to which ground truth saliency as well as computational saliency actually benefit existing state of the art VQMs. Our dataset opens new challenges for saliency modelling in video quality research and helps better gauge progress in developing saliency-based VQMs.
Wei Zhang 0072, Hantao Liu
MMSP2
2016 SIQ288: A saliency dataset for image quality research
abstract
Saliency modelling for image quality research has been an active topic in multimedia over the last five years. Saliency aspects have been added to many image quality metrics (IQMs) to improve their performance in predicting perceived quality. However, challenges to optimising the performance of saliency-based IQMs remain. To make further progress, a better understanding of human attention deployment in relation to image quality through eye-tracking experimentation is indispensable. Collecting substantial eye-tracking data is often confronted with a bias due to the involvement of massive stimulus repetition that typically occurs in an image quality study. To mitigate this problem, we proposed a new experimental methodology with dedicated control mechanisms, which allows collecting more reliable eye-tracking data. We recorded 5760 trials of eye movements from 160 human observers. Our dataset consists of 288 images representing a large degree of variability in terms of scene content, distortion type as well as degradation level. We illustrate how saliency is affected by the variations of image quality. We also compare state of the art saliency models in terms of predicting where people look in both original and distorted scenes. Our dataset helps investigate the actual role saliency plays in judging image quality, and provides a benchmark for gauging saliency models in the context of image quality.
Wei Zhang 0072, Hantao Liu
MMSP2
2016 The Relative Impact of Ghosting and Noise on the Perceived Quality of MR Images
abstract
Magnetic resonance (MR) imaging is vulnerable to a variety of artifacts, which potentially degrade the perceived quality of MR images and, consequently, may cause inefficient and/or inaccurate diagnosis. In general, these artifacts can be classified as structured or unstructured depending on the correlation of the artifact with the original content. In addition, the artifact can be white or colored depending on the flatness of the frequency spectrum of the artifact. In current MR imaging applications, design choices allow one type of artifact to be traded off with another type of artifact. Hence, to support these design choices, the relative impact of structured versus unstructured or colored versus white artifacts on perceived image quality needs to be known. To this end, we conducted two subjective experiments. Clinical application specialists rated the quality of MR images, distorted with different types of artifacts at various levels of degradation. The results demonstrate that unstructured artifacts deteriorate quality less than structured artifacts, while colored artifacts preserve quality better than white artifacts.
Hantao Liu, Jos Koonen, Miha Fuderer, Ingrid Heynderickx
IEEE Trans. Image Process.1
2016 The Application of Visual Saliency Models in Objective Image Quality Assessment: A Statistical Evaluation
abstract
Advances in image quality assessment have shown the potential added value of including visual attention aspects in its objective assessment. Numerous models of visual saliency are implemented and integrated in different image quality metrics (IQMs), but the gain in reliability of the resulting IQMs varies to a large extent. The causes and the trends of this variation would be highly beneficial for further improvement of IQMs, but are not fully understood. In this paper, an exhaustive statistical evaluation is conducted to justify the added value of computational saliency in objective image quality assessment, using 20 state-of-the-art saliency models and 12 best-known IQMs. Quantitative results show that the difference in predicting human fixations between saliency models is sufficient to yield a significant difference in performance gain when adding these saliency models to IQMs. However, surprisingly, the extent to which an IQM can profit from adding a saliency model does not appear to have direct relevance to how well this saliency model can predict human fixations. Our statistical analysis provides useful guidance for applying saliency models in IQMs, in terms of the effect of saliency model dependence, IQM dependence, and image distortion dependence. The testbed and software are made publicly available to the research community.
Wei Zhang 0072, Ali Borji, Zhou Wang 0001, Patrick Le Callet, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.5
2015 Studying human behavioural responses to time-varying distortions for video quality assessment
abstract
Advances in video quality assessment have shown the added value of including temporal aspects of artifact perception in its objective metrics. The impact of quality variations over time and the associated implications for the assessment of overall video quality are so far not fully understood yet. To investigate the human behavioural responses to time-varying distortions and the relevance of such perception to overall quality judgements, a series of subjective experiments were conducted. In the major experiment, 30 human subjects scored the overall quality of 120 stimuli, which consist of 84 videos of various time-varying quality profiles and of their corresponding 36 constituent segments of spatio-temporally constant quality. Results show that the pattern of quality variations over time tends to affect the temporal summation strategy that delivers an overall quality.
Juan Vicente Talens-Noguera, Wei Zhang 0072, Hantao Liu
ICIP3
2015 The quest for the integration of visual saliency models in objective image quality assessment: A distraction power compensated combination strategy
abstract
Novel research on image quality metrics (IQMs) attempts to further improve their reliability by including visual attention aspects of the human visual system. Literature so far mainly focuses on the extension of a specific IQM with a specific visual saliency model. In this paper, we quest the integration of visual saliency models in IQMs, in terms of its statistical meaningfulness and combination strategy. In the first step an exhaustive evaluation is conducted by integrating twenty state-of-the-art saliency models into eight best-known IQMs for image quality assessment. It demonstrates linearly combining saliency and IQMs yields a statistically significant gain in performance. Based on the statistics, we revisit the combination strategy of saliency and IQMs and propose a new strategy taking into account the distraction power of local distortions. Results show that the proposed combination strategy consistently outperforms the conventionally used linear combination strategy.
Wei Zhang 0072, Juan Vicente Talens-Noguera, Hantao Liu
ICIP3
2014 Studying the added value of computational saliency in objective image quality assessment
abstract
Advances in image quality assessment have shown the potential added value of including visual attention aspects in objective quality metrics. Numerous models of visual saliency are implemented and integrated in different quality metrics; however, their ability of improving a metric's performance in predicting perceived image quality is not fully investigated. In this paper, we conduct an exhaustive comparison of 20 state-of-the-art saliency models in the context of image quality assessment. Experimental results show that adding computational saliency is beneficial to quality prediction in general terms. However, the amount of performance gain that can be obtained by adding saliency in quality metrics highly depends on the saliency model and on the metric.
Wei Zhang 0072, Ali Borji, Fuzheng Yang 0001, Ping Jiang 0001, Hantao Liu
VCIP5
2013 How Does Image Content Affect the Added Value of Visual Attention in Objective Image Quality Assessment?
abstract
Our previous research has demonstrated that adding natural scene saliency (NSS) obtained from eye-tracking data may improve an objective metric's performance in predicting perceived image quality. In this letter, we further investigate the image content dependency of this improvement. Results show that the variation in saliency between observers highly depends on image content, and that this variation predicts the extent to which a certain image may profit from adding saliency in the objective image quality assessment.
Hantao Liu, Ulrich Engelke, Junle Wang, Patrick Le Callet, Ingrid Heynderickx
IEEE Signal Process. Lett.1
2013 Comparative Study of Fixation Density Maps
abstract
Fixation density maps (FDM) created from eye tracking experiments are widely used in image processing applications. The FDM are assumed to be reliable ground truths of human visual attention and as such, one expects a high similarity between FDM created in different laboratories. So far, no studies have analyzed the degree of similarity between FDM from independent laboratories and the related impact on the applications. In this paper, we perform a thorough comparison of FDM from three independently conducted eye tracking experiments. We focus on the effect of presentation time and image content and evaluate the impact of the FDM differences on three applications: visual saliency modeling, image quality assessment, and image retargeting. It is shown that the FDM are very similar and that their impact on the applications is low. The individual experiment comparisons, however, are found to be significantly different, showing that inter-laboratory differences strongly depend on the experimental conditions of the laboratories. The FDM are publicly available to the research community.
Ulrich Engelke, Hantao Liu, Junle Wang, Patrick Le Callet, Ingrid Heynderickx, Hans-Jürgen Zepernick, Anthony J. Maeder
IEEE Trans. Image Process.2
2012 Towards an efficient model of visual saliency for objective image quality assessment
abstract
Based on “ground truth” eye-tracking data, earlier research [1] shows that adding natural scene saliency (NSS) can improve an objective metric's performance in predicting perceived image quality. To include NSS in a real-world implementation of an objective metric, a computational model instead of eye-tracking data is needed. Existing models of visual saliency are generally designed for a specific domain, and so, not applicable to image quality prediction. In this paper, we propose an efficient model for NSS, inspired by findings from our eye-tracking studies. Experimental results show that the proposed model sufficiently captures the saliency of the eye-tracking data, and applying the model to objective image quality metrics enhances their performance in the same manner as when including eye-tracking data.
Hantao Liu, Ingrid Heynderickx
ICASSP1
2011 Visual Attention in Objective Image Quality Assessment: Based on Eye-Tracking Data
abstract
Since the human visual system (HVS) is the ultimate assessor of image quality, current research on the design of objective image quality metrics tends to include an important feature of the HVS, namely, visual attention. Different metrics for image quality prediction have been extended with a computational model of visual attention, but the resulting gain in reliability of the metrics so far was variable. To better understand the basic added value of including visual attention in the design of objective metrics, we used measured data of visual attention. To this end, we performed two eye-tracking experiments: one with a free-looking task and one with a quality assessment task. In the first experiment, 20 observers looked freely to 29 unimpaired original images, yielding us so-called natural scene saliency (NSS). In the second experiment, 20 different observers assessed the quality of distorted versions of the original images. The resulting saliency maps showed some differences with the NSS, and therefore, we applied both types of saliency to four different objective metrics predicting the quality of JPEG compressed images. For both types of saliency the performance gain of the metrics improved, but to a larger extent when adding the NSS. As a consequence, we further integrated NSS in several state-of-the-art quality metrics, including three full-reference metrics and two no-reference metrics, and evaluated their prediction performance for a larger set of distortions. By doing so, we evaluated whether and to what extent the addition of NSS is beneficial to objective quality prediction in general terms. In addition, we address some practical issues in the design of an attention-based metric. The eye-tracking data are made available to the research community .
Hantao Liu, Ingrid Heynderickx
IEEE Trans. Circuits Syst. Video Technol.1
2010 Comparing two eye-tracking databases: The effect of experimental setup and image presentation time on the creation of saliency maps
abstract
Visual attention models are typically designed based on human gaze patterns recorded through eye tracking. In this paper, two similar eye tracking experiments from independent laboratories are presented, in which humans observed natural images under task-free condition. The resulting saliency maps are analysed with respect to two criteria; the consistency between the experiments and the impact of the image presentation time. It is shown, that the saliency maps between the experiments are strongly correlated independent of presentation time. It is further revealed that the presentation time can be reduced without substantially sacrificing the accuracy of the convergent saliency map. The results provide valuable insight into the similarity of saliency maps from independent laboratories and are highly beneficial for the creation of converging saliency maps at reduced experimental time and cost.
Ulrich Engelke, Hantao Liu, Hans-Jürgen Zepernick, Ingrid Heynderickx, Anthony J. Maeder
PCS2
2010 A No-Reference Metric for Perceived Ringing Artifacts in Images
abstract
A novel no-reference metric that can automatically quantify ringing annoyance in compressed images is presented. In the first step a recently proposed ringing region detection method extracts the regions which are likely to be impaired by ringing artifacts. To quantify ringing annoyance in these detected regions, the visibility of ringing artifacts is estimated, and is compared to the activity of the corresponding local background. The local annoyance score calculated for each individual ringing region is averaged over all ringing regions to yield a ringing annoyance score for the whole image. A psychovisual experiment is carried out to measure ringing annoyance subjectively and to validate the proposed metric. The performance of our metric is compared to existing alternatives in literature and shows to be highly consistent with subjective data.
Hantao Liu, Nick Klomp, Ingrid Heynderickx
IEEE Trans. Circuits Syst. Video Technol.1
2010 A Perceptually Relevant Approach to Ringing Region Detection
abstract
An efficient approach toward a no-reference ringing metric intrinsically exists of two steps: first detecting regions in an image where ringing might occur, and second quantifying the ringing annoyance in these regions. This paper presents a novel approach toward the first step: the automatic detection of regions visually impaired by ringing artifacts in compressed images. It is a no-reference approach, taking into account the specific physical structure of ringing artifacts combined with properties of the human visual system (HVS). To maintain low complexity for real-time applications, the proposed approach adopts a perceptually relevant edge detector to capture regions in the image susceptible to ringing, and a simple yet efficient model of visual masking to determine ringing visibility. The approach is validated with the results of a psychovisual experiment, and its performance is compared to existing alternatives in literature for ringing region detection. Experimental results show that our method is promising in terms of both reliability and computational efficiency.
Hantao Liu, Nick Klomp, Ingrid Heynderickx
IEEE Trans. Image Process.1
2009 Studying the added value of visual attention in objective image quality metrics based on eye movement data
abstract
Current research on image quality assessment tends to include visual attention in objective metrics to further enhance their performance. A variety of computational models of visual attention are implemented in different metrics, but their accuracy in representing human visual attention is not fully proved yet. Thus, to provide more accurate evidence on whether and to what extent visual attention can be beneficial for objective quality prediction, the use of ¿ground truth¿ visual attention data is highly desired. In this paper, the data of an eye-tracking experiment are integrated in two objective metrics well-known in literature. Experimental results demonstrate that there is indeed a gain in performance including visual attention in objective metrics. The amount of gain in performance tends to depend on the type of objective metric and image distortion.
Hantao Liu, Ingrid Heynderickx
ICIP1
2009 How to apply spatial saliency into objective metrics for JPEG compressed images?
abstract
This paper investigates how saliency obtained from eye-tracking data can be integrated into objective metrics for JPEG compressed images. The objective metrics used in this paper are both based on features, locally extracted from the images and serving as input to a neural network for the overall quality prediction. We compare various weighting functions to combine saliency with these objective metrics, taking into account the possible distraction due to artifacts that might affect the quality judgment. Experimental results indicate that including saliency into objective metrics in an appropriate way can further enhance their performance.
Judith Redi, Hantao Liu, Paolo Gastaldo, Rodolfo Zunino, Ingrid Heynderickx
ICIP2
2008 A no-reference perceptual blockiness metric
abstract
A novel no-reference blockiness metric that can automatically and perceptually quantify blocking artifacts of DCT coding is presented. The proposed metric is built upon the specific structure information of the artifact itself combined with the properties of the human visual system (HVS) by means of a simple and efficient model of visual masking. Investigations are conducted to reduce the additional cost introduced by the human vision model, without compromising its overall prediction ability. The proposed metric is validated through comparing its performance to state-of-the-art HVS model based blockiness metrics with respect to accuracy, reliability and computational complexity.
Hantao Liu, Ingrid Heynderickx
ICASSP1
2007 A Simplified Human Vision Model Applied to a Blocking Artifact Metric
Hantao Liu, Ingrid Heynderickx
CAIP1