Leida Li

dblp:92/6630 · also Lei-Da Li · DBLP profile ↗
← Back
207ranked-venue papers
31as first author
124since 2021 · last 2026
0000-0001-9069-8796ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 157 · 23 first-author · 98 since 2021Artificial intelligence and machine learning · 48 · 3 first-author · 37 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 2 since 2021Security and privacy · 8 · 2 first-authorComputer networks · 3 · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
abstract
Livestreaming has become increasingly prevalent in modern visual communication, where automatic camera quality tuning is essential for delivering superior user Quality of Experience (QoE). Such tuning requires accurate blind image quality assessment (BIQA) to guide parameter optimization decisions. Unfortunately, the existing BIQA models typically only predict an overall coarse-grained quality score, which cannot provide fine-grained perceptual guidance for precise camera parameter tuning. To bridge this gap, we first establish FGLive-10K, a comprehensive fine-grained BIQA database containing 10,185 high-resolution images captured under varying camera parameter configurations across diverse livestreaming scenarios. The dataset features 50,925 multi-attribute quality annotations and 19,234 fine-grained pairwise preference annotations. Based on FGLive-10K, we further develop TuningIQA, a fine-grained BIQA metric for livestreaming camera tuning, which integrates human-aware feature extraction and graph-based camera parameter fusion. Extensive experiments and comparisons demonstrate that TuningIQA significantly outperforms state-of-the-art BIQA methods in both score regression and fine-grained quality ranking, achieving superior performance when deployed for livestreaming camera tuning.
Xiangfei Sheng, Zhichao Duan 0002, Xiaofeng Pan, Yipo Huang, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI7
2026 Fine-grained Image Quality Assessment for Perceptual Image Restoration
abstract
Recent years have witnessed remarkable achievements in perceptual image restoration (IR), creating an urgent demand for accurate image quality assessment (IQA), which is essential for both performance comparison and algorithm optimization. Unfortunately, the existing IQA metrics exhibit inherent weakness for IR task, particularly when distinguishing fine-grained quality differences among restored images. To address this dilemma, we contribute the first-of-its-kind fine-grained image quality assessment dataset for image restoration, termed FGRestore, comprising 18,408 restored images across six common IR tasks. Beyond conventional scalar quality scores, FGRestore was also annotated with 30,886 fine-grained pairwise preferences. Based on FGRestore, a comprehensive benchmark was conducted on the existing IQA metrics, which reveal significant inconsistencies between score-based IQA evaluations and the fine-grained restoration quality. Motivated by these findings, we further propose FGResQ, a new IQA model specifically designed for image restoration, which features both coarse-grained score regression and fine-grained quality ranking. Extensive experiments and comparisons demonstrate that FGResQ significantly outperforms state-of-the-art IQA metrics.
Xiangfei Sheng, Xiaofeng Pan, Zhichao Yang 0013, Pengfei Chen 0003, Leida Li
AAAI5
2026 LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations
abstract
The increasing popularity of long Text-to-Image (T2I) generation has created an urgent need for automatic and interpretable models that can evaluate the image-text alignment in long prompt scenarios. However, the existing T2I alignment benchmarks predominantly focus on short prompt scenarios and only provide MOS or Likert scale annotations. This inherent limitation hinders the development of long T2I evaluators, particularly in terms of the interpretability of alignment. In this study, we contribute LongT2IBench, which comprises 14K long text-image pairs accompanied by graph-structured human annotations. Given the detail-intensive nature of long prompts, we first design a Generate-Refine-Qualify annotation protocol to convert them into textual graph structures that encompass entities, attributes, and relations. Through this transformation, fine-grained alignment annotations are achieved based on these granular elements. Finally, the graph-structed annotations are converted into alignment scores and interpretations to facilitate the design of T2I evaluation models. Based on LongT2IBench, we further propose LongT2IExpert, a LongT2I evaluator that enables multi-modal large language models (MLLMs) to provide both quantitative scores and structured interpretations through an instruction-tuning process with Hierarchical Alignment Chain-of-Thought (CoT). Extensive experiments and comparisons demonstrate the superiority of the proposed LongT2IExpert in alignment evaluation and interpretation.
Zhichao Yang 0013, Tianjiao Gu, Jianjie Wang, Feiyu Lin, Xiangfei Sheng, Pengfei Chen 0003, Leida Li
AAAI7
2026 HumanCrop-Thinker: An inference-driven framework with explicit thinking for explainable human-centric image cropping
Yipo Huang, Pengfei Chen 0003, Leida Li
Expert Syst. Appl.4
2026 HVS-inspired blind image quality index with prominent perception learning and multi-level progressive integration
Taiyang Chen, Bo Hu 0008, Chunyi Li 0001, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
Neurocomputing4
2026 Few-label blind image quality assessment via samples chosen from new and existing scenes
Deqiang Cheng 0001, Tianshu Song, Qiqi Kou, Leida Li
Pattern Recognit.5
2026 MIGF-Net: Multimodal interaction-guided fusion network for image aesthetics assessment
Yun Liu 0009, Zhipeng Wen, Leida Li, Peiguang Jing, Daoxin Fan
Pattern Recognit.3
2026 Learning Scene-Invariant Distribution for Generalizable Blind Image Quality Assessment
abstract
The inherent diversity of visual scenes poses a fundamental challenge in blind image quality assessment (BIQA), which has become a major obstacle to the model generalization. In this study, we found that human annotations for images with different visual scenes exhibit distinct quality distribution discrepancies. The existing BIQA models tend to overfit to such diversified distributions, which in turn leads to compromised model generalizability, especially when dealing with unseen scenes in the real-world scenario. Motivated by the above facts, this paper presents a generalizable BIQA model by learning Scene-INvariant Distribution, named SIND. Specifically, we propose a distribution alignment framework to alleviate the distribution discrepancy for quality regression models, which is achieved by automatically scaling and shifting the cross-scene distributions into a unified distribution. Then, the aligned unified distribution is leveraged to supervise the model training, achieving scene-invariant and quality-aware feature representation. In addition, a token-complementary patch reasoning network is designed to extract comprehensive quality-aware features from both the image overview and detail, achieving more accurate quality prediction. Extensive experiments for both image technical- and aesthetic-quality assessment tasks show the superiority of the proposed SIND model over the state-of-the-arts. Moreover, the proposed framework is model-agnostic and can enhance model generalizability without incurring extra inference costs. The proposed method won the championship in the NTIRE 2024 Portrait Quality Assessment Challenge. Codes will be available at https://github.com/ZachL1/SIND.
Yipo Huang, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.5
2026 Multimodal Emotion-Aligned Cognitive Networks for Image Aesthetic Assessment
abstract
Image aesthetic assessment (IAA) is a challenging task due to the subjectivity and abstraction of aesthetic perception. Psychological studies reveal that aesthetic experiences often trigger emotional responses, while comment texts directly reflect people’s expressions of aesthetics and emotions. However, existing multimodal IAA methods neglect the alignment between modalities. To address this, we propose a multimodal emotion-alignment cognitive network (MEC-Net) for IAA, employing strategies of emotion alignment, subjective–objective interaction, and multimodal fusion. First, an emotion alignment module is introduced to align image and text modalities using emotional stimuli, enhancing the consistency of heterogeneous modal features. Then, a subjective and objective representation module is proposed to extract multi-source information from text and images separately. Next, a subjective-objective interactive LSTM (SO-LSTM) is designed to capture the deep interaction between images and text in aesthetic understanding. Finally, an dynamic multimodal fusion (DMF) based on low-rank decomposition is proposed to integrate subjective, objective, and subjective-objective interactive modal features for aesthetic distribution prediction. Extensive experiments and qualitative analysis on image aesthetic benchmarks indicate that the proposed MEC-Net outperforms the state-of-the-art on three IAA tasks. Further, we increase emotion classification task-driven evaluation metrics to verify the strong generalizability of the proposed MEC-Net.
Xixi Nie, Shixin Huang, Jiawei Luo 0002, Xiaodan Zhang 0005, Leida Li, Hongchun Qu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 RAM-VQA: Restoration Assisted Multi-Modality Video Quality Assessment
abstract
Video Quality Assessment (VQA) strives to computationally emulate human perceptual judgments and has garnered significant attention given its widespread applicability. However, existing methodologies face two primary impediments: (1) limited proficiency in evaluating samples at quality extremes (e.g., severely degraded or near-perfect videos), and (2) insufficient sensitivity to nuanced quality variations arising from a misalignment with human perceptual mechanisms. Although vision-language models offer promising semantic understanding, their reliance on visual encoders pre-trained for high-level tasks often compromises their sensitivity to low-level distortions. To surmount these challenges, we propose the Restoration-Assisted Multi-modality VQA (RAM-VQA) framework. Uniquely, our approach leverages video restoration as a proxy to explicitly model distortion-sensitive features. The framework operates through two synergistic stages: a prompt learning stage that constructs a quality-aware textual space using triple-level references (degraded, restored, and pristine) derived from the restoration process, and a dual-branch evaluation stage that integrates semantic cues with technical quality indicators via spatio-temporal differential analysis. Extensive experiments demonstrate that RAM-VQA achieves state-of-the-art performance across diverse benchmarks, exhibiting superior capability in handling extreme-quality content while ensuring robust generalization.
Pengfei Chen 0003, Jiebin Yan, Rajiv Soundararajan, Giuseppe Valenzise, Leida Li
IEEE Trans. Image Process.6
2026 Perceptual Quality Assessment of Low-Light Enhanced Images: A Multi-Annotated Subjective Dataset and a Multimodal Objective Method
abstract
Low-light Image Enhancement Algorithms (LIEAs) aim to improve the visibility and visual quality of images captured in low-light environments. However, none of the existing LIEAs can comprehensively restore all visual contents, which makes it inevitable for the Enhanced Low-light Images (ELIs) to have different degrees of distortion, thereby affecting the visual quality. Currently, there is little research focusing on the quality assessment of these ELIs, partly due to the lack of publicly available datasets. Moreover, existing quality assessment methods primarily focus on a single visual modality and fail to sufficiently exploit the structural information across multiple image attributes, consequently resulting in suboptimal prediction performance. To this end, this paper conducts a systematic study on both subjective and objective quality assessment of ELIs. Firstly, we construct the first Multi-annotated and multi-modal Low-light image Enhancement quality dataset (MLE), which contains 1,000 ELIs, along with subjective studies to obtain multiple attribute annotations, quality scores, and textual descriptions. Based on this, we further propose an Attribute-guided Vision-Language Graph Reasoning Network (AVGR-Net) for ELI quality prediction, which effectively integrates multi-attribute visual and textual information through cross-modal graph reasoning and alignment. Extensive data analysis and experimental results validate both the reliability of the MLE dataset and the superior performance of the AVGR-Net compared to state-of-the-art methods.
Bo Hu 0008, Leida Li, Ke Gu 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2026 Exploring Cross-Modal Mutual Prompt Learning for Video Quality Assessment
abstract
Enhancing video quality assessment (VQA) through semantic information integration is a critical research focus. Recent research has employed the Contrastive Language-Image Pre-training (CLIP) model as a foundation to improve semantic perception. However, the image-text alignment inherent in these pre-trained Vision-Language (VL) models frequently results in suboptimal VQA performance. While prompt engineering has recently targeted the language component to address this alignment issue, the unique insights resided in visual analysis is still overlooked for further advancing VQA tasks. Additionally, seeking a trade-off between quality separability and domain invariance in VQA remains largely unresolved within the VL paradigm. In this paper, we introduce a novel cross-modal prompt-based approach to tackle these challenges. Specifically, we propose learnable prompts within the vision branch to foster synergy between visual and language modalities through a language-to-vision coupling function. The multi-view backbone is then carefully crafted with content enhancement and distortion-aware temporal modulation to ensure quality separability. The language prompts, derived from visual representations, are further supported by adaptive weighting mechanisms to optimize the balance between quality separability and domain invariance. Experimental results demonstrate the effectiveness of our proposed method over leading VQA models, showing significant improvements in generalization across diverse datasets. The source code for this work is publicly available athttps://github.com/cpf0079/CM2PL.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Jiebin Yan, Vinit Jakhetiya, Aladine Chetouani
IEEE Trans. Multim.2
2026 Distortion-Sensitive Masked Autoencoder for Omnidirectional Video Quality Assessment
abstract
Omnidirectional Video Quality Assessment (OVQA) is a challenging task due to the limited availability of adequate numbers of training samples for learning representations of distortions on omnidirectional videos. The recent masked autoencoder (MAE) has shown promising performance in learning local and global representations in a self-supervised way, and can be used to attempt to mitigate the difficulty of having insufficient annotated samples to adequately train omnidirectional video quality prediction models. But the reconstruction tasks that MAE models are designed for do not pertain to predicting diverse perceptual distortions, especially those relevant to the task of OVQA. We have attempted to overcome these limitations to harness and apply the power of the MAE concept to the OVQA problem. Towards this purpose, we create a Distortion-Sensitive Masked AutoEncoder (DS-MAE) that is able to represent perceptual distortions on omnidirectional videos. DS-MAE extracts viewports from omnidirectional videos and employs a masked autoencoding module (MAM) and a knowledge replay module (KRM) to learn representations on each viewport. In the MAM, distorted patches from omnidirectional videos are masked, by replacing them with undistorted counterparts. The autoencoder is trained to reconstruct the masked distortions, imbuing them with the ability to represent diverse video degradations. The KRM extracts and stores content representations, which are then “replayed” to mitigate potential catastrophic forgetting of content during training of the DS-MAE. Finally, a simple OVQA model is constructed using the pre-trained DS-MAE across all viewports. The new model, called OmniVQA, was tested on three public OVQA datasets. The experimental results show that OmniVQA delivers competitive performance against all compared models.
Zongyao Hu, Lixiong Liu, Ke Gu 0001, Leida Li, Alan C. Bovik
IEEE Trans. Multim.4
2026 AesPrompt: Zero-Shot Image Aesthetics Assessment With Multi-Granularity Aesthetic Prompt Learning
abstract
Recent years have witnessed increasing interest towards image aesthetics assessment (IAA), which predicts the aesthetic appeal of images by simulating human perception. The state-of-the-art IAA methods, despite their significant advancements, typically rely heavily on time-consuming and labor-intensive human annotation of aesthetic scores. Furthermore, they are subject to the generalization challenge, which is highly desired in real-world applications. Motivated by this, zero-shot image aesthetics assessment (ZIAA) is investigated to achieve robust model generalization without relying on manual aesthetic annotations, which remains largely underexplored. Specifically, a novel aesthetic prompt learning framework for ZIAA, dubbed AesPrompt, is presented in this paper. The key insight of AesPrompt is to emulate the human aesthetic perception process for learning aesthetic-oriented prompts in a multi-granularity manner. First, we first develop a new pseudo aesthetic distribution generation paradigm based on multi-LLM ensemble. Then, external knowledge of multi-granularity prompts encompassing image themes, emotions, and aesthetics is acquired. Through learning the multi-granularity aesthetic-oriented prompts, the proposed method achieves better generalization and interpretability. Extensive experiments on five IAA benchmarks demonstrate that AesPrompt consistently outperforms the state-of-the-art ZIAA methods across diverse-sourced images, covering natural images, artistic images, and artificial intelligence-generated images.
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Giuseppe Valenzise
IEEE Trans. Multim.2
2026 IBCL-VQA: Video Quality Assessment with In-Batch Contrastive Learning and Two-Phase Feature Fusion
abstract
Video Quality Assessment (VQA) technology is of significant importance for improving video transmission, storage, and processing. Although Convolutional Neural Networks (CNNs)-based and Transformer-based methods have achieved significant progress, they still suffer from some drawbacks. The previous methods treated video data as independent samples, thereby neglecting the close or distant relationships between different quality levels and consequently constraining the model’s discriminative ability; meanwhile, most existing fusion strategies utilize fixed architectures that lack the ability to adapt to the data for optimal integration, resulting in insufficient utilization of spatio-temporal information and limited expression capabilities of the fused features. To address the above issues, this article proposes a VQA method with In-Batch Contrastive Learning and Two-Phase Feature Fusion (IBCL-VQA). Firstly, the spatial features are extracted through two branches, which not only preserve the global semantics but also focus on the local regions. The temporal characteristics are obtained through a pre-trained video recognition model. Secondly, we propose an in-batch contrastive learning mechanism which, through the principles of maximizing intra-class similarity and minimizing inter-class similarity, combined with a dynamically adjusted penalty strategy, models the correlation between video quality levels. Thirdly, a two-phase feature fusion strategy, consisting of the Gated Spatio-temporal Attention Unit (GSTU) and the Adaptive Fusion Cell (AFC), is further proposed. The former achieves spatio-temporal feature fusion through dynamic weight allocation, and the latter adaptively integrates the features from the two branches based on data characteristics. Finally, a regression module outputs the quality score. Experimental results on five real-world VQA datasets demonstrate the superior performance of the IBCL-VQA. Furthermore, the strong generalizability is verified through cross-database testing. The code and pre-trained weights will be publicly available at: https://github.com/BoHu90/IBCL-VQA .
Bo Hu 0008, Leida Li, Lihuo He, Xinbo Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Asymmetric Hierarchical Difference-aware Interaction Network for Event-guided Motion Deblurring
abstract
Event cameras are bio-inspired sensors that are capable of capturing motion information with high temporal resolution, which show potential in aiding image motion deblurring recently. Most existing methods indiscriminately handle feature fusion of two modalities with symmetric unidirectional/bidirectional interactions at different-level layers in feature encoder, while ignoring the different dependencies between cross-modal hierarchical features. To tackle these limitations, we propose a novel Asymmetric Hierarchical Difference-aware Interaction Network (AHDINet) for event-based motion deblurring, which explores the complementarity of two modalities with differential dependence modeling of cross-modal hierarchical features. Thereby, an event-assisted edge complement module is designed to leverage event modality to enhance the edge details of the image features in low-level encoder stage, and an image-assisted semantic complement module is developed to transfer contextual semantics of image features to event branch in high-level encoder stage. Benefiting from the proposed differentiated interaction mode, the respective advantages of image and event modalities are fully exploited. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance.
Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
AAAI3
2025 AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated Images
abstract
With the advancement of AI-generated content technologies, AI-generated images (AGIs) have become increasingly influential in artistic creation and visual communication. However, the aesthetic quality of AGIs varies significantly due to technical limitations and the influence of user input, underscoring the urgent need for systematic aesthetic evaluation of AGIs. In addition, it is difficult to ensure the consistency of text-to-image, which compresses the application space of AGIs. To address these issues, a fine-grained dataset for Aesthetic and Alignment evaluation of AGIs (AGIAA-2K) is presented. This dataset contains 2,064 images generated using 172 well-designed prompts across six different AGI models, with each image annotated based on the subjective experiment. Then, the rationality of the dataset is verified by data analysis. Finally, the performances of the existing algorithms are evaluated in terms of image aesthetic assessment and text-to-image alignment of AGIs. The results demonstrate that these algorithms cannot effectively evaluate these two aspects. The AGIAA-2K is available at https://github.com/BoHu90/AGIAA-2K.
Bo Hu 0008, Nanxiang Li, Lihuo He, Wen Lu 0004, Leida Li, Xinbo Gao 0001
ICASSP5
2025 A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment
abstract
Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the perception and experience, which may seriously hinder its application in fields such as competitive sports. Up to now, few explorations focus on the quality assessment issue of wide-angle video. This deficiency primarily stems from the absence of a specialized dataset for wide-angle videos. To bridge this gap, we construct the first Multi-annotated and multi-modal Wide-angle Video quality assessment (MWV) dataset. Then, the performances of state-of-the-art video quality methods on the MWV dataset are investigated by inter-dataset testing and intra-dataset testing. Experimental results show that these methods impose significant limitations on their applicability.
Bo Hu 0008, Chunyi Li 0001, Lihuo He, Leida Li, Xinbo Gao 0001
ICASSP5
2025 MACA-VQA: Quality Assessment of UGC Videos via Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention Fusion
abstract
User-generated content (UGC) videos often exhibit complex distortions and diverse content, posing significant challenges for traditional video quality assessment (VQA) methods. Approaches that directly merge distortion and semantic information risk feature conflicts and the loss of details. In addition, simple concatenation of spatiotemporal features fails to capture vital interactions, limiting predictive accuracy. Motivated by these challenges, this paper proposes a Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention Fusion framework for VQA, named MACA-VQA. Specifically, a novel multi-level adaptive strategy progressively incorporates distortion information into each Transformer layer of the CLIP model, enabling layer-wise fusion of semantic and distortion features. Furthermore, a newly introduced cross-attention fusion mechanism dynamically integrates spatiotemporal features, capturing complex, multidimensional interactions. Extensive experiments demonstrate that MACA-VQA achieves state-of-the-art performance on multiple public datasets, validating its effectiveness and robustness in both intra-dataset and inter-dataset scenarios. The source code is available at https://github.com/BoHu90/MACA-VQA
Bo Hu 0008, Yimeng Zhao, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
ICME3
2025 Text-to-Image Diffusion Models are AI-Generated Image Quality Scorers
abstract
Despite the remarkable progress in text-to-image generation models, AI-Generated Images (AGIs) still face significant challenges, such as poor perception quality and text-image misalignment. Consequently, AI-Generated Image Quality Assessment (AGIQA), which aims to assess capabilities of generative models, has gained increasing attention. Diffusion models, pre-trained on large-scale text-image generative tasks, inherently encode rich prior knowledge regarding image quality and text-image alignment, which is crucial for AGIQA task. Inspired by this, unlike most existing CLIP-based methods, we present an initial exploration of diffusion model-based AGIQA, termed Diff-AGIQA. Specifically, to effectively harness pre-trained diffusion models, we introduce a simple yet effective strategy focusing on two key aspects: 1) feature selection, identifying the discriminative features within diffusion models; and 2) visual prompting, to extract more task-specific features while keeping the diffusion model frozen. Extensive experiments on AGIQA-1K, AGIQA-3K, AGIQA-20K, and RichHF-18K datasets demonstrate the potential of applying diffusion models to AGIQA task. Code is publicly available https://github.com/sxfly99/Diff-AGIQA.
Xiangfei Sheng, Weidong Zou, Pengfei Chen 0003, Leida Li
ICME6
2025 Towards Region-Adaptive Feature Disentanglement and Enhancement for Small Object Detection
abstract
Current feature fusion strategies often fail to adequately account for the influence of activation intensity across different scales on small object features, which impedes the effective detection of small objects. To address this limitation, we propose the Region-Adaptive Feature Disentanglement and Enhancement (RAFDE) strategy, which improves both downsampling and feature fusion by leveraging activation intensity variations at multiple scales. First, we introduce the Boundary Transitional Region-enhanced Downsampling (BTRD) module, which enhances boundary transitional regions containing both strongly and weakly activated features, thereby mitigating the loss of crucial boundary information for small objects. Second, we present the Regional-Adaptive Feature Fusion (RAFF) module, which adaptively disentangles and fuses co-activated and uni-activated regions from adjacent levels into the current level, effectively reducing the risk of small objects being overwhelmed. Extensive experiments on several public datasets demonstrate that the RAFDE strategy is highly effective and outperforms state-of-the-art methods. The code is available at https://github.com/b-yanchao/RAFDE.git.
Yanchao Bi, Yang Ning, Xiushan Nie, Xiankai Lu, Yongshun Gong, Leida Li
IJCAI6
2025 Low-light Image Enhancement Quality Assessment: A Real-World Dataset and An Objective Method
abstract
Low-light Image Enhancement (LIE) technology adaptively improves brightness while preserving texture details and suppressing noise artifacts, thereby reducing visual degradation caused by insufficient illumination. While deep learning-based image enhancement algorithms have made significant progress, a key gap remains in establishing standardized methods for fairly evaluating and comparing their performance. To bridge this gap, this paper systematically investigates enhanced low-light image quality assessment from both subjective and objective dimensions. First, we introduce a Real-world Low-light Image Enhancement quality assessment dataset (RLIE), which contains 1540 images from 154 scenarios, each with a subjective score given by the subjects. Based on this, we propose a low light enhanced image quality assessment method based on Multi-level Illumination Injection and Hierarchical Discrepancy Perception (MIIHDP). The core idea of this method is to hierarchically inject separated illumination information into the feature extraction process, then tailor the processing of difference information at different scales to obtain a more comprehensive representation. Finally, extensive statistical analyses demonstrate the rationality of the proposed RLIE dataset, and experimental results show the superior performance of the proposed MIIHDP compared with state-of-the-arts. Our dataset and code are released at: https://github.com/BoHu90/RLIE.
Chunyi Li 0001, Bo Hu 0008, Taiyang Chen, Leida Li, Lihuo He, Xinbo Gao 0001
ACM Multimedia4
2025 InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic Images
abstract
Aesthetic Image Cropping (AIC) aims to improve the visual appeal of images by removing redundant content while preserving attractive elements. Despite the encouraging progresses achieved in data-driven approaches, most existing models struggle to understand user intentions, particularly for diversified scenes with multiple subjects. Moreover, they can only provide cropping results without explanations, which further restricts their usability in real-world applications. Motivated by the above facts, we introduce InstructCrop : a multimodal large language model (MLLM)-based AIC framework, which can understand user instructions and provide explanatory reasons for cropping results. Specifically, we first build a multimodal Image Cropping Instruction Tuning (ICIT) dataset through a cost-effective paradigm by generating high-quality instruction tuning data based on the existing cropping datasets. Then, we embed dynamic domain knowledge into the cropping model by integrating cropping-aware experts of aesthetic assessment and composition classification. Finally, we adapt MLLMs to generate the cropping results and corresponding explanations. Quantitative and qualitative experiments on three benchmark datasets demonstrate that InstructCrop enables effective and interpretable image cropping, which aligns better with user intentions. Data and code are available at https://github.com/sxfly99/InstructCrop.
Xiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen 0003, Tong Zhu 0003, Leida Li
ACM Multimedia6
2025 Towards Syn-to-Real IQA: A Novel Perspective on Reshaping Synthetic Data Distributions
abstract
Blind Image Quality Assessment (BIQA) has advanced significantly through deep learning, but the scarcity of large-scale labeled datasets remains a challenge. While synthetic data offers a promising solution, models trained on existing synthetic datasets often show limited generalization ability. In this work, we make a key observation that representations learned from synthetic datasets often exhibit a discrete and clustered pattern that hinders regression performance: features of high-quality images cluster around reference images, while those of low-quality images cluster based on distortion types. Our analysis reveals that this issue stems from the distribution of synthetic data rather than model architecture. Consequently, we introduce a novel framework SynDR-IQA, which reshapes synthetic data distribution to enhance BIQA generalization. Based on theoretical derivations of sample diversity and redundancy's impact on generalization error, SynDR-IQA employs two strategies: distribution-aware diverse content upsampling, which enhances visual diversity while preserving content distribution, and density-aware redundant cluster downsampling, which balances samples by reducing the density of densely clustered areas. Extensive experiments across three cross-dataset settings (synthetic-to-authentic, synthetic-to-algorithmic, and synthetic-to-synthetic) demonstrate the effectiveness of our method. The code is available at https://github.com/Li-aobo/SynDR-IQA.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong
NeurIPS4
2025 Semi-SNN: Biological-inspired semi-supervised image classification with spiking neural networks
Biao Hou, Chuanfeng Ma, Leida Li, Hao Zhu 0009, Licheng Jiao
Neurocomputing5
2025 A two-stage strategy for brain-inspired unsupervised learning in spiking neural networks
Chuanfeng Ma, Biao Hou, Leida Li, Hao Zhu 0009, Dou Quan, Licheng Jiao
Neurocomputing5
2025 Multi-Modality Multi-Attribute Contrastive Pre-Training for Image Aesthetics Computing
abstract
In the Image Aesthetics Computing (IAC) field, most prior methods leveraged the off-the-shelf backbones pre-trained on the large-scale ImageNet database. While these pre-trained backbones have achieved notable success, they often overemphasize object-level semantics and fail to capture the high-level concepts of image aesthetics, which may only achieve suboptimal performances. To tackle this long-neglected problem, we propose a multi-modality multi-attribute contrastive pre-training framework, targeting at constructing an alternative to ImageNet-based pre-training for IAC. Specifically, the proposed framework consists of two main aspects. 1) We build a multi-attribute image description database with human feedback, leveraging the competent image understanding capability of the multi-modality large language model to generate rich aesthetic descriptions. 2) To better adapt models to aesthetic computing tasks, we integrate the image-based visual features with the attribute-based text features, and map the integrated features into different embedding spaces, based on which the multi-attribute contrastive learning is proposed for obtaining more comprehensive aesthetic representation. To alleviate the distribution shift encountered when transitioning from the general visual domain to the aesthetic domain, we further propose a semantic affinity loss to restrain the content information and enhance model generalization. Extensive experiments demonstrate that the proposed framework sets new state-of-the-arts for IAC tasks.
Yipo Huang, Leida Li, Pengfei Chen 0003, Haoning Wu 0001, Weisi Lin, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Cross-Modality Interactive Attention Network for AI-generated image quality assessment
Tianwei Zhou, Songbai Tan, Leida Li, Baoquan Zhao, Qiuping Jiang, Guanghui Yue 0001
Pattern Recognit.3
2025 Blind Quality Assessment of Wide-Angle Videos Based on Deformation Representation Learning and Multi-Dimensional Feature Fusion
abstract
Wide-angle videos shot with short-focus lenses often exhibit deformation distortions, which poses significant challenges for video quality assessment (VQA). Although current VQA methods focus primarily on video content and distortion perception, there has been little explicit research on the impact of deformation characteristics on the perception of wide-angle video quality. To this end, this paper makes the first attempt to construct a novel wide-angle video quality assessment method based on deformation representation learning and multi-dimensional feature fusion, termed DRLMF. Specifically, we first analyze the deformation distribution characteristics of wide-angle videos based on the deformation camera model. Based on this, a three-stream video perception and assessment network is proposed. The first branch extracts global semantics using the image encoder of CLIP. The second branch introduces an effective deformation region selection strategy and proposes an interpretable deformation representation learning module. This module leverages the perception advantages of convolutional neural networks (CNNs) in local distortions and considers the correlation between patch size and distortion perception. The third branch extracts motion features using an action recognition network. Finally, an effective multi-dimensional feature fusion module is proposed to integrate more refined and richer semantic, deformation, and motion features. Extensive experiments on wide-angle VQA datasets and standard video datasets show that the DRLMF outperforms the state-of-the-arts in terms of prediction monotonicity and accuracy. The codes will be available at https://github.com/BoHu90/DRLMF.
Bo Hu 0008, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Towards Explainable Image Aesthetics Assessment With Attribute-Oriented Critiques Generation
abstract
Compared with the unimodal image aesthetics assessment (IAA), multimodal IAA has demonstrated superior performance. This indicates that the critiques could provide rich aesthetics-aware semantic information, which also enhance the explainability of IAA models. However, images are not always accompanied with critiques in real-world situation, rendering multimodal IAA inapplicable in most cases. Therefore, it would be interesting to investigate whether we can generate aesthetic critiques to facilitate image aesthetic representation learning and enhance model explainability. Motivated by these facts, this paper presents an attribute-oriented Critiques Generation framework for explainable IAA, dubbed CG-IAA, which consists of three major components, i.e., Vision-Language Aesthetic Pretraining (VLAP), Multi-Attribute Experts Learning (MAEL) and Multimodal Aesthetics Prediction (MAP). Specifically, the vanilla CLIP is first finetuned on a multimodal IAA database. Considering that the aesthetic critiques typically consist of multiple attributes, a new multimodal IAA database which contains over 1 million critiques with up to four aesthetic attributes is constructed with the language model-based knowledge transfer. Then, CLIP-based multi-attribute experts are trained based on this database. Finally, the pretrained experts are utilized to generate aesthetic critiques for assisting unimodal image aesthetics prediction. Extensive experiments have been done on four popular IAA databases, and the results demonstrate the advantage of CG-IAA over the state-of-the-arts. Furthermore, CG-IAA features better explainability and generalization with the assistance of generated critiques. The source code is available athttps://github.com/sxfly99/CG-IAA.
Leida Li, Xiangfei Sheng, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong
IEEE Trans. Circuits Syst. Video Technol.1
2025 Scene-Modulated High-Order Statistical Representation Learning for No-Reference Super-Resolution Image Quality Assessment
abstract
With the rapid development of single image super-resolution (SR) technology, there is an urgent need to develop a fair no reference Super-Resolution image Quality Assessment (SRQA) method. Existing no reference SRQA methods primarily concentrate on SR artifacts including structural distortion and texture distortion by extracting spatial features, but ignore the inductive bias of Deep Neural Network (DNN)-based SR models. As a result, they function effectively for interpolation-based and dictionary-based algorithms, but struggle to perform as effectively with DNN-based SR algorithms. We found that the visual content generated by DNN-based SR models under different inductive biases often carries a content-invariant model-specific style, which can be captured by the correlations between hierarchical representation channels. To that end, we propose a novel Scene-modulated High-order Statistical Representation network (SmHSR) built on a multi-scale over-complete transformation. We quantify the perceptual quality of SR images as the shift of high-order statistical properties in their multi-scale over-complete representation, where intra-channel statistics are used to capture spatial correlations and inter-channel statistics are used to capture the inductive bias of SR models. In addition, the scene information implicit in the deep over-complete representation is used to modulate the high-order statistical properties, which simulates the top-down regulation of cognition on perception. Under the modulation of scene information, SmHSR can learn more sophisticated scene-aware statistical representation. The MultiLayer Perceptron (MLP) is used to map the high-order statistical representation to an overall quality. We test our method on multiple SR image quality databases. Experimental results show that our method outperforms the state-of-the-art SRQA methods.
Yongwei Mao, Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong
IEEE Trans. Circuits Syst. Video Technol.4
2025 Progressively Generated Text-Assisted Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate user perceptions to judge the aesthetic quality of images. Due to the high subjectivity of users and the complexity of image aesthetics, modeling IAQA solely at the image level is a compromise. Consequently, existing methods mainly focus on multimodal-based models and achieve effective performance. These methods explore aesthetic comments on images to characterize users and serve as auxiliary text information for multimodal modeling. Unfortunately, this may suffer from two limitations. One limitation is that aesthetic comments are often unavailable for an unknown image in the test phase, and another limitation is that the semantic information of these comments may be uncertain and fuzzy. Therefore, this paper proposes a progressively generated text-assisted image aesthetic quality assessment method, aiming to address the lack of aesthetic comments and the fuzziness of aesthetic judgments in these comments. Specifically, we first adopt a Multimodal Large Language Model (MLLM) to generate aesthetic comments on images by simulating user perceptions and utilize the generated comments to characterize their aesthetic perception to assist in the pre-training of our multimodal-based IAQA model. Then, we design an attribute prediction module to determine the attribute levels of aesthetic judgments and utilize text template construction to further generate explicit descriptions of image aesthetics. Finally, we leverage the generated attribute descriptions to further assist in training our IAQA model. By progressively generating textual auxiliary descriptions of aesthetics for images, the proposed model can gradually determine the aesthetic quality of the images. Massive experimental results indicate that the proposed method outperforms existing mainstream methods on multiple IAQA datasets.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Kunyang Sun, Leida Li
IEEE Trans. Fuzzy Syst.6
2025 CrossEI: Boosting Motion-Oriented Object Tracking With an Event Camera
abstract
With the differential sensitivity and high time resolution, event cameras can record detailed motion clues, which form a complementary advantage with frame-based cameras to enhance the object tracking, especially in challenging dynamic scenes. However, how to better match heterogeneous event-image data and exploit rich complementary cues from them still remains an open issue. In this paper, we align event-image modalities by proposing a motion adaptive event sampling method, and we revisit the cross-complementarities of event-image data to design a bidirectional-enhanced fusion framework. Specifically, this sampling strategy can adapt to different dynamic scenes and integrate aligned event-image pairs. Besides, we design an image-guided motion estimation unit for extracting explicit instance-level motions, aiming at refining the uncertain event clues to distinguish primary objects and background. Then, a semantic modulation module is devised to utilize the enhanced object motion to modify the image features. Coupled with these two modules, this framework learns both the high motion sensitivity of events and the full texture of images to achieve more accurate and robust tracking. The proposed method is easily embedded in existing tracking pipelines, and trained end-to-end. We evaluate it on four large benchmarks, i.e. FE108, VisEvent, FE240hz and CoeSot. Extensive experiments demonstrate our method achieves state-of-the-art performance, and large improvements are pointed as contributions by our sampling strategy and fusion concept.
Zhiwen Chen 0002, Jinjian Wu, Weisheng Dong, Leida Li, Guangming Shi
IEEE Trans. Image Process.4
2025 Diffusion Model-Based Visual Compensation Guidance and Visual Difference Analysis for No-Reference Image Quality Assessment
abstract
Existing free-energy guided No-Reference Image Quality Assessment (NR-IQA) methods continue to face challenges in effectively restoring complexly distorted images. The features guiding the main network for quality assessment lack interpretability, and efficiently leveraging high-level feature information remains a significant challenge. As a novel class of state-of-the-art (SOTA) generative model, the diffusion model exhibits the capability to model intricate relationships, enhancing image restoration effectiveness. Moreover, the intermediate variables in the denoising iteration process exhibit clearer and more interpretable meanings for high-level visual information guidance. In view of these, we pioneer the exploration of the diffusion model into the domain of NR-IQA. We design a novel diffusion model for enhancing images with various types of distortions, resulting in higher quality and more interpretable high-level visual information. Our experiments demonstrate that the diffusion model establishes a clear mapping relationship between image reconstruction and image quality scores, which the network learns to guide quality assessment. Finally, to fully leverage high-level visual information, we design two complementary visual branches to collaboratively perform quality evaluation. Extensive experiments are conducted on seven public NR-IQA datasets, and the results demonstrate that the proposed model outperforms SOTA methods for NR-IQA. The codes will be available at https://github.com/handsomewzy/DiffV2IQA.
Zhaoyang Wang 0003, Bo Hu 0008, Mingyang Zhang 0002, Jie Li 0001, Leida Li, Maoguo Gong, Xinbo Gao 0001
IEEE Trans. Image Process.5
2025 AI-Generated Image Quality Assessment Based on Task-Specific Prompt and Multi-Granularity Similarity
abstract
Recently, AI-generated images (AIGIs), synthesized based on initial textual prompts, have attracted widespread attention. However, due to limitations in current generation techniques, these images often exhibit degraded perceptual quality and semantic misalignment with the guiding prompts. Therefore, evaluating both perceptual quality and text-to-image alignment is essential for optimizing the performance of generative models. Existing methods design textual prompts solely based on the initial prompt for both perceptual and alignment quality tasks, and compute only coarse-grained similarity between the designed prompt and the generated image. However, such task-agnostic prompts overlook the distinctions between the perceptual and alignment quality tasks, and coarse-level similarity fails to capture semantic details, leading to suboptimal evaluation performance. To address these challenges, we propose a novel AIGI quality assessment framework, termed TPMS, which incorporates task-specific prompt and multi-granularity similarity computation. The task-specific prompt constructs dedicated prompts for perceptual and alignment quality respectively, allowing the model to capture distinct quality cues tailored to each evaluation task. Multi-granularity similarity measures the coarse-level similarity between the generated image and task-specific prompts to capture global quality characteristics, and the fine-level similarity between the generated image and the initial prompt to enhance semantic detail awareness. By integrating these two complementary similarities, TPMS enables precise and robust quality prediction. Extensive experiments on four widely-used AIGI quality benchmarks validate the effectiveness and superiority of the proposed framework.
Jili Xia, Lihuo He, Cheng Deng 0002, Leida Li, Xinbo Gao 0001
IEEE Trans. Image Process.4
2025 GAN-Guided Few-Shot Attention Network for Medical Images Fusion Quality Assessment
abstract
Medical image fusion (MIF) plays an important role in precision diagnostics and treatment planning management, and medical image fusion quality assessment (MIFQA) has an aggressive effect in improving MIF performance. However, obtaining medical reference images is difficult, and the significant demand for medical prior knowledge and reference images is an important challenge in the field of MIFQA. To address this issue, this paper proposes a two-stage model for MIFQA. In the first stage, we design a GAN-based Quality-aware Network called QANet. By fusing the radiologist's mean opinion score (MOS) with the source image, the model is guided to generate one reference images of each quality. Then, in the second stage, the reference images are fed into our proposed class attention siamese network (CASNet) based on class activation mapping (CAM) under few-shot learning to fully explore the information in limited reference images. It can enforce the model to focus on the key lesion area and effectively reduce the dependence of MIFQA on medical fused images. Finally, the quality score of the unlabeled fused image is predicted by calculating the distance with reference image. Experiments on home-made MIFQA dataset shows that our method can achieve results that are ahead of the state-of-the-art methods.
Lu Tang 0001, Nailong Hou, Leida Li, Chuangeng Tian, Guanyu Zhu
IEEE Trans. Medical Imaging5
2025 Visual-Language Multi-Task Blind Image Quality Assessment With Local Quality Weighting
abstract
The objective of blind image quality assessment (BIQA) is to develop a model capable of automatically evaluating image quality without requiring any reference knowledge. While multi-task learning has been widely utilized in BIQA, it has predominantly remained unimodal. This paper delves into the Visual-Language multi-task BIQA model, where distortion knowledge can be captured through image-text contrastive learning. Specifically, Visual-Language auxiliary tasks targeting distortion type and quality level are introduced, respectively, where both positive and negative image-text pairs are constructed for the target distorted image. Subsequently, image-text correspondences are learned in the embedding space while simultaneously evaluating image quality. Notably, in the auxiliary task learning, the proposed method not only brings the image and its corresponding positive text prompt closer but also pushes away the image from its negative text prompts, thereby facilitating the extraction of pertinent distortion features. In the quality assessment task, a patch-wise strategy is employed during the training phase. Differing from conventional BIQA methods, a novel NSS-guided quality weighting is introduced to gauge the correlation between patch quality and global quality, thereby enabling precise quality prediction. Extensive experiments are conducted on six IQA datasets, and the experimental results verify the superiority of the proposed method.
Jili Xia, Lihuo He, Bo Hu 0008, Leida Li, Xinbo Gao 0001
IEEE Trans. Multim.4
2025 Variational Multiple-Instance Learning With Embedding Correlation Modeling for Hyperspectral Target Detection
abstract
The hyperspectral target detection is widely concerned in geoscience and remote sensing due to the abundant spectral information in hyperspectral imagery. However, the detection performance is highly dependent on the high-quality target signature or pixel-level supervised signals, which are extremely challenging and costly. In this article, we propose a variational multiple-instance neural network with embedding correlation modeling (VMIL-ECM) for weakly supervised hyperspectral target detection, which relaxes the rigid target prior (e.g., target signatures and/or pixel-level annotations), and only region-level labels are required. VMIL-ECM explicitly models the location of the targets within the region as a latent variable under the nonindependent and identically distributed (non-i.i.d.) assumption to estimate the underlying ground-truth target locations. The expectation-maximization (EM) algorithm is employed to iteratively optimize the posterior distribution of latent variables and learn discriminative spectral features for the target detection. To fully utilize the contextual information within the hyperspectral region, a permutation-invariant transformer-based structure is devised to explore the embedding correlation among instances. Moreover, a dynamic thresholding strategy is adopted to produce the reliable fine-grained supervised signals. Extensive experiments on three simulated datasets and two real-field datasets are conducted to verify the effectiveness of VMIL-ECM, and the state-of-the-art performance has been achieved over the existing comparison methods. The code for the VMIL-ECM is publicly available at: https://github.com/BoYangXDU/VMIL-ECM.
Bo Yang 0047, Changzhe Jiao, Jinjian Wu, Leida Li
IEEE Trans. Neural Networks Learn. Syst.4
2025 Image Cropping with Content and Composition Attribute-aware Global Relation Reasoning
abstract
Image cropping aims to find visually pleasing content in an image, which will enhance its aesthetic quality. Existing image cropping approaches mainly emphasize the geometric properties of images, such as composition and layout, neglecting the rich aesthetic information available from the physical attributes (e.g., content and themes), and background information beyond the foreground in images. Consequently, this article proposes an image cropping method based on the content and composition attribute-aware global relation reasoning, which aims at guiding the generation of cropped sub-images by exploring critical attributes based on content and composition as well as global object correlations that affect aesthetics in images. Particularly, to comprehensively introduce aesthetic information into image cropping, we capture feature representations reinforced by content and composition attributes simultaneously. The feature representations can strengthen the visual aesthetics of cropped sub-images. To make the cropped sub-images amply contain more global information, we introduce a global relation reasoning branch in the proposed cropping module, which can fully exploit the dependency relationship between the foreground and background in images. Extensive experiments on image cropping benchmarks demonstrate that our approach is superior to state-of-the-art image cropping methods.
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001, Leida Li
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Motion Deblurring via Spatial-Temporal Collaboration of Frames and Events
abstract
Motion deblurring can be advanced by exploiting informative features from supplementary sensors such as event cameras, which can capture rich motion information asynchronously with high temporal resolution. Existing event-based motion deblurring methods neither consider the modality redundancy in spatial fusion nor temporal cooperation between events and frames. To tackle these limitations, a novel spatial-temporal collaboration network (STCNet) is proposed for event-based motion deblurring. Firstly, we propose a differential-modality based cross-modal calibration strategy to suppress redundancy for complementarity enhancement, and then bimodal spatial fusion is achieved with an elaborate cross-modal co-attention mechanism to weight the contributions of them for importance balance. Besides, we present a frame-event mutual spatio-temporal attention scheme to alleviate the errors of relying only on frames to compute cross-temporal similarities when the motion blur is significant, and then the spatio-temporal features from both frames and events are aggregated with the custom cross-temporal coordinate attention. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/STCNet.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Guangming Shi
AAAI4
2024 Bridging the Synthetic-to-Authentic Gap: Distortion-Guided Unsupervised Domain Adaptation for Blind Image Quality Assessment
abstract
The annotation of blind image quality assessment (BIQA) is labor-intensive and time-consuming, especially for authentic images. Training on synthetic data is expected to be beneficial, but synthetically trained models often suf-fer from poor generalization in real domains due to domain gaps. In this work, we make a key observation that introducing more distortion types in the synthetic dataset may not improve or even be harmful to generalizing au-thentic image quality assessment. To solve this challenge, we propose distortion-guided unsupervised domain adaptationfor BIQA (DGQA), a novel framework that leverages adaptive multi-domain selection via prior knowledge from distortion to match the data distribution between the source domains and the target domain, thereby reducing negative transfer from the outlier source domains. Extensive experiments on two cross-domain settings (synthetic distortion to authentic distortion and synthetic distortion to algorith-mic distortion) have demonstrated the effectiveness of our proposed DGQA. Besides, DGQA is orthogonal to existing model-based BIQA methods, and can be used in combi-nation with such models to improve performance with less training data.
Jinjian Wu, Yongxu Liu 0001, Leida Li
CVPR4
2024 Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention
abstract
Recently, Vision Transformer has achieved great success in recovering missing details in low-resolution sequences, i.e., the video super-resolution (VSR) task. Despite its su-periority in VSR accuracy, the heavy computational bur-den as well as the large memory footprint hinder the de-ployment of Transformer-based VSR models on constrained devices. In this paper, we address the above issue by proposing a novel feature-level masked processing frame-work: VSR with Masked Intra and inter-frame Attention (MIA-VSR). The core of MIA-VSR is leveraging feature-level temporal continuity between adjacent frames to re-duce redundant computations and make more rational use of previously enhanced SR features. Concretely, we propose an intra-frame and inter-frame attention block which takes the respective roles of past features and input features into consideration and only exploits previously enhanced fea-tures to provide supplementary information. In addition, an adaptive block-wise mask prediction module is developed to skip unimportant computations according to feature sim-ilarity between adjacent frames. We conduct detailed ab-lation studies to validate our contributions and compare the proposed method with recent state-of-the-art VSR approaches. The experimental results demonstrate that MIA-VSR improves the memory and computation efficiency over state-of-the-art methods, without trading off PSNR accuracy. The code is available at https://github.com/LabShuHangGU/MIA-VSR.
Leheng Zhang, Xiaorui Zhao, Keze Wang, Leida Li, Shuhang Gu
CVPR5
2024 Semantics-Aware Image Aesthetics Assessment using Tag Matching and Contrastive Ranking
abstract
The perception of image aesthetics is built upon the understanding of semantic content. However, how to evaluate the aesthetic quality of images with diversified semantic backgrounds remains challenging in image aesthetics assessment (IAA). To address the dilemma, this paper presents a semantics-aware image aesthetics assessment approach, which first analyzes the semantic content of images and then models the aesthetic distinctions among images from two perspectives, i.e., aesthetic attribute and aesthetic level. Concretely, we propose two strategies, dubbed tag matching and contrastive ranking, to extract knowledge pertaining to image aesthetics. The tag matching identifies the semantic category and the dominant aesthetic attributes based on predefined tag libraries. The contrastive ranking is designed to uncover the comparative relationships among images with different aesthetic levels but similar semantic backgrounds. In the process of contrastive ranking, the impact of long-tailed distribution of aesthetic data is also considered by balanced sampling and traversal contrastive learning. Extensive experiments and comparisons on three benchmark IAA databases demonstrate the superior performance of the proposed model in terms of both prediction accuracy and alleviating long-tailed effect. The code will be public at https://github.com/yzc-ippl/TMCR **REMOVE 2nd URL**://github.com/yzc-ippl/TMCR.
Zhichao Yang 0013, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong
ACM Multimedia2
2024 AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception
abstract
The highly abstract nature of image aesthetics perception (IAP) poses a significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma, resulting in MLLMs falling short of aesthetics perception capabilities. To address the above challenge, we first introduce a comprehensively annotated Aesthetic Multi-Modality Instruction Tuning (AesMMIT) dataset, which serves as the footstone for building multi-modality aesthetics foundation models. Specifically, to align MLLMs with human aesthetics perception, we construct a corpus-rich aesthetic critique database with 21,904 diverse-sourced images and 88K human natural language feedbacks, which are collected via progressive questions, ranging from coarse-grained aesthetic grades to fine-grained aesthetic descriptions. To ensure that MLLMs can handle diverse queries, we further prompt GPT to refine the aesthetic critiques and assemble the large-scale aesthetic instruction tuning dataset, i.e. AesMMIT, which consists of 409K multi-typed instructions to activate stronger aesthetic capabilities. Based on the AesMMIT database, we fine-tune the open-sourced general foundation models, achieving multi-modality Aesthetic Expert models, dubbed AesExpert. Extensive experiments demonstrate that the proposed AesExpert models deliver significantly better aesthetic perception performances than the state-of-the-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision. Project Page: https://yipoh.github.io/aes-expert/.
Yipo Huang, Xiangfei Sheng, Zhichao Yang 0013, Zhichao Duan 0002, Pengfei Chen 0003, Leida Li, Weisi Lin, Guangming Shi
ACM Multimedia7
2024 Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Leida Li
ACM Multimedia7
2024 Blind image quality index with high-level Semantic Guidance and low-level fine-grained Representation
Bo Hu 0008, Leida Li, Ke Gu 0001, Shuaijian Wang, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing3
2024 Weakly-supervised cloud detection and effective cloud removal for remote sensing images
Xiuhong Yang, Tiankun Gou, Zhiyong Lv, Leida Li, Haiyan Jin
J. Vis. Commun. Image Represent.4
2024 Aesthetic image cropping meets VLP: Enhancing good while reducing bad
Leida Li, Pengfei Chen 0003
J. Vis. Commun. Image Represent.2
2024 Quality-aware blind image motion deblurring
Tianshu Song, Leida Li, Jinjian Wu, Weisheng Dong, Deqiang Cheng 0001
Pattern Recognit.2
2024 Emotion-aware hierarchical interaction network for multimodal image aesthetics assessment
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001
Pattern Recognit.2
2024 Active Learning-Based Sample Selection for Label-Efficient Blind Image Quality Assessment
abstract
Despite the considerable effort devoted to high-generalizable blind image quality assessment (BIQA), the generalization performance of the state-of-the-art metrics remains limited when facing new visual scenes. A straightforward way to address the dilemma is labeling a great number of images from the new scene and subsequently training a new model, which is quite labor-intensive and cost-expensive. Hence, there is an urgent need to mitigate the dependency on labeled samples by designing a data-efficient BIQA algorithm. Motivated by the above facts, this paper presents an Active Learning-based IQA (AL-IQA) framework, which reduces the requirement for training samples by selecting representative images from two perspectives, including distortion and content. Specifically, in terms of distortion, we design distortion prompts and adopt Contrastive Language-Image Pre-Training (CLIP) to predict image distortion in a zero-shot manner. Then, we employ curriculum learning-inspired strategy to select samples with gradually increasing difficulty (measured by prediction uncertainty of CLIP), in order to facilitate model training. Meantime, in terms of content, we adopt distribution matching-based dataset distillation to distill unlabeled images into several high-density informative synthetic images. Then, feature distances between unlabeled images and distilled images are compared to identify images with the most representative content. Finally, Borda count is adopted to capture a consensus of both distortion and content through weighted counting, and prompt tuning is utilized for adapting the model to the IQA task. Extensive experiments are conducted on five IQA datasets, and the results demonstrate that the proposed AL-IQA not only effectively reduces the number of training samples but also achieves state-of-the-art prediction accuracy and generalization performance. The source code is available athttps://github.com/esnthere/AL-IQA.
Tianshu Song, Leida Li, Deqiang Cheng 0001, Pengfei Chen 0003, Jinjian Wu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Deep Pyramid Network for Low-Light Endoscopic Image Enhancement
abstract
Endoscopic images captured under low-light enclosed intestinal environment usually have poor visibility (manifested as uneven illumination and noise), affecting the work efficiency of physicians and the accuracy of lesion detection. To improve the image quality, the literature has reported many low-light image enhancement (LIE) methods. However, most methods do not perform well in handling the low-light endoscopic image enhancement (LEIE) task, usually bringing additional artifacts or amplifying noise. In this paper, we propose a novel deep pyramid enhancement network (DPENet) to enhance endoscopic images from both global and local perspectives. Specifically, considering the uneven illumination of endoscopic images, DPENet utilizes an image pyramid framework with three parallel branches to explore and integrate both global and local features at different scales. To suppress noise, DPENet sets multiple scale-space feature extraction blocks (SFEBs) in each branch. SFEB consists of a contextual feature extraction module (CFEM) and a spatial residual attention module (SRAM). CFEM mines contextual information to help the network understand semantic information while suppress the isolated noise. SRAM leverages the spatial attention mechanism to help the network adaptively focus on dim regions. Experimental results on a public dataset and our collected dataset show that DPENet is competent for the LEIE task with promising results, and outperforms 9 state-of-the-art LIE methods in both qualitative and quantitative aspects.
Guanghui Yue 0001, Runmin Cong, Tianwei Zhou, Leida Li, Tianfu Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 A Spatial-Temporal Video Quality Assessment Method via Comprehensive HVS Simulation
abstract
The quality of videos is the primary concern of video service providers. Built upon deep neural networks, video quality assessment (VQA) has rapidly progressed. Although existing works have introduced the knowledge of the human visual system (HVS) into VQA, there are still some limitations that hinder the full exploitation of HVS, including incomplete modeling with few HVS characteristics and insufficient connection among these characteristics. In this article, we present a novel spatial-temporal VQA method termed HVS-5M, wherein we design five modules to simulate five characteristics of HVS and create a bioinspired connection among these modules in a cooperative manner. Specifically, on the side of the spatial domain, the visual saliency module first extracts a saliency map. Then, the content-dependency and the edge masking modules extract the content and edge features, respectively, which are both weighted by the saliency map to highlight those regions that human beings may be interested in. On the other side of the temporal domain, the motion perception module extracts the dynamic temporal features. Besides, the temporal hysteresis module simulates the memory mechanism of human beings and comprehensively evaluates the video quality according to the fusion features from the spatial and temporal domains. Extensive experiments show that our HVS-5M outperforms the state-of-the-art VQA methods. Ablation studies are further conducted to verify the effectiveness of each module toward the proposed method. The source code is available at https://github.com/GZHU-DVL/HVS-5M.
Aoxiang Zhang, Yuan-Gen Wang, Weixuan Tang 0004, Leida Li, Sam Kwong
IEEE Trans. Cybern.4
2024 Learning Frame-Event Fusion for Motion Deblurring
abstract
Motion deblurring is a highly ill-posed problem due to the significant loss of motion information in the blurring process. Complementary informative features from auxiliary sensors such as event cameras can be explored for guiding motion deblurring. The event camera can capture rich motion information asynchronously with microsecond accuracy. In this paper, a novel frame-event fusion framework is proposed for event-driven motion deblurring (FEF-Deblur), which can sufficiently explore long-range cross-modal information interactions. Firstly, different modalities are usually complementary and also redundant. Cross-modal fusion is modeled as complementary-unique features separation-and-aggregation, avoiding the modality redundancy. Unique features and complementary features are first inferred with parallel intra-modal self-attention and inter-modal cross-attention respectively. After that, a correlation-based constraint is designed to act between unique and complementary features to facilitate their differentiation, which assists in cross-modal redundancy suppression. Additionally, spatio-temporal dependencies among neighboring inputs are crucial for motion deblurring. A recurrent cross attention is introduced to preserve inter-input attention information, in which the current spatial features and aggregated temporal features are attending to each other by establishing the long-range interaction between them. Extensive experiments on both synthetic and real-world motion deblurring datasets demonstrate our method outperforms state-of-the-art event-based and image/video-based methods. The code will be made publicly available.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.4
2024 Blind Image Quality Index With Cross-Domain Interaction and Cross-Scale Integration
abstract
With the assistance of Convolutional Neural Networks (CNNs), Image Quality Assessment (IQA) models have made great progress in evaluating both simulated distortion and authentic distortion. However, most of the existing IQA models only learn the features of distorted images, and thus do not make full use of the available feature representation of other domains. Furthermore, the common multi-scale fusion strategies are relatively simple, such as downsampling and concatenating, which further limits the prediction performance. To this end, we propose a novel blind image quality index with cross-domain interaction and cross-scale integration, which is designed based on the combination of CNN and Transformer. First, the hierarchical spatial-domain and gradient-domain representations are obtained through a typical CNN architecture. Then, based on the proposed gradient-query cross-attention, these two types of features are fully interacted in the Cross-Domain Interaction (CDI) module. To represent the distortion information more comprehensively, the Cross-Scale Integration (CSI) module is proposed to combine the information between different scales progressively. Finally, the quality score is obtained through a simple regression module. The experimental results on five public IQA databases of both simulated and authentic scenes show that the proposed model outperforms the compared state-of-the-art metrics. In addition, cross-database experiments show that the proposed model has strong generalization performance.
Bo Hu 0008, Leida Li, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2024 Coarse-to-Fine Image Aesthetics Assessment With Dynamic Attribute Selection
abstract
Image aesthetics assessment (IAA) is an interesting but challenging task, owing to the ineffable nature of human sense of beauty. The study of IAA has evolved from simple binary classification to more complex score regression and distribution prediction. It is effortless for people to perform aesthetic binary classification,i.e., aesthetically pleasing or not. However, further judgment on the fine-level scalar aesthetic score is complex and typically determined by aesthetic attributes presented in the image, such as content, lighting and color. Motivated by the above facts, this paper presents a Coarse-to-fine image Aesthetics assessment model guided by Dynamic Attribute Selection, dubbed CADAS. The underlying idea is to simulate the process of human aesthetic perception by performing coarse-to-fine aesthetic reasoning. Specifically, a hierarchical AttributeNet is first pre-trained by imitating the staged mechanism of human aesthetic experience, producing the candidate aesthetic attributes. Then, an AestheticNet is introduced to perform the coarse-level binary classification, based on which a confidence-based attribute selection strategy is designed to dynamically pick out the dominant aesthetic attributes from the candidate ones. Finally, a self-attention-based FusionNet is designed to explore the interaction between dominant aesthetic attributes and aesthetic features, producing the fine-level aesthetic prediction. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-arts. Furthermore, CADAS is also able to output the dominant aesthetic attributes in images, facilitating model explainability.
Yipo Huang, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Guangming Shi
IEEE Trans. Multim.2
2024 Blind Image Quality Assessment Based on Perceptual Comparison
abstract
Blind image quality assessment (BIQA) is a regression task with continuous label space, the feature space of which is expected to have a corresponding continuity in the target space. However, existing approaches typically learn quality score regression directly in an end-to-end fashion, which leaves networks susceptible to interference from task-agnostic information, and fails to capture the continuity of BIQA. In this work, by explicitly establishing inter-sample associations, a simple yet effective BIQA framework based on perceptual comparison is proposed to capture the continuity. To this end, besides the basic quality score regression, the relative quality scores between images are predicted to exploit the relative quality relationships between samples for optimizing the representation of image perceptual quality. In addition, based on the human perceptual characteristic, we derive a novel sample weighting strategy to dynamically adjust the weights for different samples in the network learning process for further improving the robustness of the model. The performances on both single-database and cross-database experiments achieve state-of-the-art, indicating the effectiveness of the proposed method. Besides, the proposed framework is model-agnostic, which can effectively improve the performance of the benchmark model with no extra inference cost.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.4
2024 Multi-Level Transitional Contrast Learning for Personalized Image Aesthetics Assessment
abstract
Personalized image aesthetics assessment (PIAA) is aimed at modeling the unique aesthetic preferences of individuals, based on which personalized aesthetic scores are predicted. People have different standards for image aesthetics, and accordingly, images rated at the same aesthetic level by different users explicitly reveal their aesthetic preferences. However, previous PIAA models treat each individual as an isolated optimization target, failing to take full advantage of the contrastive information among users. Further, although people's aesthetic preferences are unique, they still share some commonalities, meaning that PIAA models could be built on the basis of generic aesthetics. Motivated by the above facts, this article presents a Multi-level Transitional Contrast Learning (MTCL) framework for PIAA by transiting features from generic aesthetics to personalized aesthetics via contrastive learning. First, a generic image aesthetics assessment network is pre-trained to learn the common aesthetic features. Then, image sets rated to have the same aesthetic levels by different users are employed to learn the differentiated aesthetic features through multiple level-wise contrast learning based on the generic aesthetic features. Finally, a target user's PIAA model is built by integrating generic and differentiated aesthetic features. Extensive experiments on four benchmark PIAA databases demonstrate that the proposed MTCL model outperforms the state-of-the-arts.
Zhichao Yang 0013, Leida Li, Yuzhe Yang 0001, Weisi Lin
IEEE Trans. Multim.2
2024 Quality Assessment for Stitched Panoramic Images via Patch Registration and Bidimensional Feature Aggregation
abstract
Quality assessment for stitched panoramic images (SPIQA) is of great significance for the stitching algorithm optimization. By contrast, this task is much more challenging and arduous than traditional IQA task due to the high resolution of stitched panoramic images and the particularity and complexity of stitching distortions. For this task, we propose an effective method based on patch registration and bidimensional feature aggregation (PRBFA). First, inspired by the attention mechanism of the human visual system and the limited range of human vision, a soft patch segmentation and selection method is presented to determine the key patches in panoramic images to participate in the following patch matching and feature alignment stages, achieving patch registration between the panoramic image and the corresponding constituent images. Further, to fully simulate the human visual perception process from local viewport to panorama, the feature exploration is successively performed from local to global, which is also adaptive to the complexity of the distortions in stitched panoramic images. For performance testification, extensive experiments are conducted on the publicly released SPIQA database, the results of which prove the performance superiority of the PRBFA method.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Ke Gu 0001, Jinjian Wu
IEEE Trans. Multim.4
2024 Deep Shape-Texture Statistics for Completely Blind Image Quality Evaluation
abstract
Opinion-Unaware Blind Image Quality Assessment (OU-BIQA) models aim to predict image quality without training on reference images and subjective quality scores. Thereinto, image statistical comparison is a classic paradigm, while the performance is limited by the representation ability of visual descriptors. Deep features as visual descriptors have advanced IQA in recent research, but they are discovered to be highly texture-biased and lack shape-bias. On this basis, we find out that image shape and texture cues respond differently toward distortions, and the absence of either one results in an incomplete image representation. Therefore, to formulate a well-rounded statistical description for images, we utilize the shape-biased and texture-biased deep features produced by Deep Neural Networks (DNNs) simultaneously. More specifically, we design a Shape-Texture Adaptive Fusion (STAF) module to merge shape and texture information, based on which we formulate quality-relevant image statistics. The perceptual quality is quantified by the variant Mahalanobis distance between the inner and outer Deep Shape-Texture Statistics (DSTS), wherein the inner and outer statistics respectively describe the quality fingerprints of the distorted image and natural images. The proposed DSTS delicately utilizes shape-texture statistical relations between different data scales in the deep domain and achieves state-of-the-art (SOTA) quality prediction performance on images with artificial and authentic distortions.
Peilin Chen 0001, Hanwei Zhu, Keyan Ding, Leida Li, Shiqi Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Attribute-assisted Multimodal Network for Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) is challenging due to its highly abstract nature. Nowadays, people tend to share images and comment them on social networks, which can provide rich information for judging image aesthetics. As a result, user comments of an image can be jointly utilized to learn better feature representations for IAA. Previous researches have shown that aesthetic attributes are crucial factors in determining image aesthetic quality and influencing people’s aesthetic perception. Accordingly, when commenting an image, people usually give descriptions from the perspective of aesthetic attributes. Inspired by this, this paper presents a new Attribute-Assisted Multimodal network (AAM-Net) for image aesthetics assessment. Specifically, we propose a cross-modal attribute interaction module to explore the related aesthetic attribute semantics shared by an image and the corresponding aesthetic comments. Then, a cross-modal gate unit is introduced to further refine significant attribute semantics interactively. Finally, informative aesthetic features can be obtained for predicting image aesthetic distributions. Experimental results on two public multimodal IAA databases demonstrate the superiority of the proposed model over the state-of-the-art methods.
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo
ICME2
2023 BMI-Net: A Brain-inspired Multimodal Interaction Network for Image Aesthetic Assessment
abstract
Image aesthetic assessment (IAA) has drawn wide attention in recent years as more and more users post images and texts on the Internet to share their views. The intense subjectivity and complexity of IAA make it extremely challenging. Text triggers the subjective expression of human aesthetic experience based on human implicit memory, so incorporating the textual information and identifying the relationship with the image is of great importance for IAA. However, IAA with the image as input fails to fully consider subjectivity, while existing multimodal IAA ignores the interrelationship among modalities. To this end, we propose a brain-inspired multimodal interaction network (BMI-Net) that simulates how the association area of the cerebral cortex processes sensory stimuli. In particular, the knowledge integration LSTM (KI-LSTM) is proposed to learn the image-text interaction relation. The proposed scalable multimodal fusion (SMF) based on low-rank decomposition fuses image, text and interaction modalities to predict the aesthetic distribution. Extensive experiments show that the proposed BMI-Net outperforms existing state-of-the-art methods on three IAA tasks.
Xixi Nie, Bo Hu 0008, Xinbo Gao 0001, Leida Li, Xiaodan Zhang 0005, Bin Xiao 0002
ACM Multimedia4
2023 AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) aims at predicting the aesthetic quality of images. Recently, large pre-trained vision-language models, like CLIP, have shown impressive performances on various visual tasks. When it comes to IAA, a straightforward way is to finetune the CLIP image encoder using aesthetic images. However, this can only achieve limited success without considering the uniqueness of multimodal data in the aesthetics domain. People usually assess image aesthetics according to fine-grained visual attributes, e.g., color, light and composition. However, how to learn aesthetics-aware attributes from CLIP-based semantic space has not been addressed before. With this motivation, this paper presents a CLIP-based multi-attribute contrastive learning framework for IAA, dubbed AesCLIP. Specifically, AesCLIP consists of two major components, i.e., aesthetic attribute-based comment classification and attribute-aware learning. The former classifies the aesthetic comments into different attribute categories. Then the latter learns an aesthetic attribute-aware representation by contrastive learning, aiming to mitigate the domain shift from the general visual domain to the aesthetics domain. Extensive experiments have been done by using the pre-trained AesCLIP on four popular IAA databases, and the results demonstrate the advantage of AesCLIP over the state-of-the-arts. The source code will be public at https://github.com/OPPOMKLab/AesCLIP.
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong, Yuzhe Yang 0001, Liwu Xu, Guangming Shi
ACM Multimedia2
2023 Event-based Motion Deblurring with Modality-Aware Decomposition and Recomposition
abstract
Event camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events.
Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia3
2023 Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature Representation
abstract
Personalized image aesthetics assessment (PIAA) has gained increasing attention from researchers due to its ability to measure individual users' specific aesthetic experiences. However, most existing PIAA methods rely on holistic features or simplistic coding to characterize users' aesthetic preferences for images, and we believe that more rich explicit features are needed in modeling PIAA. Consequently, we propose an attribute-guided fine-grained feature-aware personalized image aesthetics assessment method, which can fully capture fine-grained features from multiple attributes to represent users' aesthetic preferences for images. To achieve this, we first build a fine-grained feature extraction (FFE) module to obtain the refined local features of image attributes to compensate for holistic features. The FFE module is then used to generate user-level features, which are combined with the image-level features to obtain user-preferred fine-grained feature representations. By training extensive users' PIAA tasks, the aesthetic distribution of most users can be transferred to the personalized scores of individual users. To enable our proposed model to learn more generalizable aesthetics among individual users, we incorporate the degree of dispersion between users' personalized scores and image aesthetic distribution as a coefficient in the loss function during model training. Experimental results on several PIAA databases show that our method outperforms existing mainstream PIAA methods, and can effectively infer users' personalized aesthetics of images.
Hancheng Zhu, Zhiwen Shao, Yong Zhou 0003, Guangcheng Wang, Pengfei Chen 0003, Leida Li
ACM Multimedia6
2023 Technical Quality-Assisted Image Aesthetics Quality Assessment
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Liwu Xu, Yuzhe Yang 0001
PRCV (11)2
2023 Reduced-reference image deblurring quality assessment based on multi-scale feature enhancement and aggregation
Bo Hu 0008, Shuaijian Wang, Xinbo Gao 0001, Leida Li, Ji Gan, Xixi Nie
Neurocomputing4
2023 Anchor-based knowledge embedding for image aesthetics assessment
Leida Li, Tianwu Zhi, Guangming Shi, Yuzhe Yang 0001, Liwu Xu, Yandong Guo
Neurocomputing1
2023 Transfer learning for just noticeable difference estimation
Yongwei Mao, Jinjian Wu, Leida Li, Weisheng Dong
Inf. Sci.4
2023 Adaptive Search-and-Training for Robust and Efficient Network Pruning
abstract
Both network pruning and neural architecture search (NAS) can be interpreted as techniques to automate the design and optimization of artificial neural networks. In this paper, we challenge the conventional wisdom of training before pruning by proposing a joint search-and-training approach to learn a compact network directly from scratch. Using pruning as a search strategy, we advocate three new insights for network engineering: 1) to formulate adaptive search as a cold start strategy to find a compact subnetwork on the coarse scale; and 2) to automatically learn the threshold for network pruning; 3) to offer flexibility to choose between efficiency and robustness. More specifically, we propose an adaptive search algorithm in the cold start by exploiting the randomness and flexibility of filter pruning. The weights associated with the network filters will be updated by ThreshNet, a flexible coarse-to-fine pruning method inspired by reinforcement learning. In addition, we introduce a robust pruning strategy leveraging the technique of knowledge distillation through a teacher-student network. Extensive experiments on ResNet and VGGNet have shown that our proposed method can achieve a better balance in terms of efficiency and accuracy and notable advantages over current state-of-the-art pruning methods in several popular datasets, including CIFAR10, CIFAR100, and ImageNet. The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/AST-NP.htm.
Xiaotong Lu, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Dynamic Expert-Knowledge Ensemble for Generalizable Video Quality Assessment
abstract
Despite the impressive progress of supervised methods in quality assessment for in- the-wild videos, models trained on one domain often fail to generalize well to others due to the domain shifts caused by distortion diversity and content variation. Domain generalizable video quality assessment (VQA) methods that can work across domains remain an open research challenge. Although combining more data following the mixed-domain training strategy can improve the generalization performance to a certain extent, the specific knowledge from each source domain, which could potentially be useful for improving unseen domain generalization, is ignored in this principle. Motivated by this, we propose a domain generalizable VQA method named Dynamic Ensemble of Expert-Knowledge (DEEK), a novel framework that dynamically exploits the expert-knowledge from each source domain to achieve a generalizable ensemble prediction. Specifically, based on the multiple experts each trained to specialize in a particular source domain, we aim to exploit complementary information provided by the expert-knowledge. We effectively train an ensemble model by proposing a quality-sensitive InfoNCE loss to regularize the collaborative training of all experts in the contrastive learning formulation, aiming to exploit complementary information provided by the expert-knowledge when forming the ensemble. By dynamically integrating the experts according to their relevances to the target data, these expert-knowledge could be leveraged for better generalization. Experiments on five VQA datasets verify that our approach outperforms the state-of-the-arts by large margins.
Pengfei Chen 0003, Leida Li, Haoliang Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.2
2023 ECSNet: Spatio-Temporal Feature Learning for Event Camera
abstract
The neuromorphic event cameras can efficiently sense the latent geometric structures and motion clues of a scene by generating asynchronous and sparse event signals. Due to the irregular layout of the event signals, how to leverage their plentiful spatio-temporal information for recognition tasks remains a significant challenge. Existing methods tend to treat events as dense image-like or point-serie representations. However, they either suffer from severe destruction on the sparsity of event data or fail to encode robust spatial cues. To fully exploit their inherent sparsity with reconciling the spatio-temporal information, we introduce a compact event representation, namely 2D-1T event cloud sequence (2D-1T ECS). We couple this representation with a novel light-weight spatio-temporal learning framework (ECSNet) that accommodates both object classification and action recognition tasks. The core of our framework is a hierarchical spatial relation module. Equipped with specially designed surface-event-based sampling unit and local event normalization unit to enhance the inter-event relation encoding, this module learns robust geometric features from the 2D event clouds. And we propose a motion attention module for efficiently capturing long-term temporal context evolving with the 1T cloud sequence. Empirically, the experiments show that our framework achieves par or even better state-of-the-art performance. Importantly, our approach cooperates well with the sparsity of event data without any sophisticated operations, hence leading to low computational costs and prominent inference speeds.
Zhiwen Chen 0002, Jinjian Wu, Junhui Hou, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.4
2023 Theme-Aware Visual Attribute Reasoning for Image Aesthetics Assessment
abstract
People usually assess image aesthetics according to visual attributes, e.g., interesting content, good lighting and vivid color, etc. Further, the perception of visual attributes depends on the image theme. Therefore, the inherent relationship between visual attributes and image theme is crucial for image aesthetics assessment (IAA), which has not been comprehensively investigated. With this motivation, this paper presents a new IAA model based on Theme-Aware Visual Attribute Reasoning (TAVAR). The underlying idea is to simulate the process of human perception in image aesthetics by performing bilevel reasoning. Specifically, a visual attribute analysis network and a theme understanding network are first pre-trained to extract aesthetic attribute features and theme features, respectively. Then, the first level Attribute-Theme Graph (ATG) is built to investigate the coupling relationship between visual attributes and image theme. Further, a flexible aesthetics network is introduced to extract general aesthetic features, based on which we built the second level Attribute-Aesthetics Graph (AAG) to mine the relationship between theme-aware visual attributes and aesthetic features, producing the final aesthetic prediction. Extensive experiments on four public IAA databases demonstrate the superiority of the proposed TAVAR model over the state-of-the-arts. Furthermore, TAVAR features better explainability due to the use of visual attributes.
Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.1
2023 Image Aesthetics Assessment With Attribute-Assisted Multimodal Memory Network
abstract
Image aesthetics assessment (IAA) has attracted growing interest in recent years but is still challenging due to its highly abstract nature. Nowadays, more and more people tend to comment images shared on the social networks, which can provide rich aesthetics-aware semantic information from different aspects. Therefore, user comments of an image can be exploited as supplementary information for enhancing aesthetic representation learning. Previous researches have demonstrated that aesthetic attributes make significant effect on image aesthetic quality and humans’ aesthetic perception. Typically, people are used to give comments on an image from the perspective of aesthetic attributes, based on which the aesthetic quality of images can be inferred. Motivated by this, this paper presents an Attribute-assisted Multimodal Memory Network (AMM-Net) for image aesthetics assessment, which utilizes aesthetic attributes to model the interactions between visual and textual modalities. Specifically, we design two memory networks to capture the attribute-aware information most related to the image and associated comments respectively. Further, with multiple memory hops, attribute semantics shared by the two modalities are refined and cross-modal interactions are enhanced progressively. Finally, more discriminative aesthetic representations can be obtained for IAA. The experimental results and comparisons on two public multimodal IAA datasets demonstrate the superiority of the proposed model over the state-of-the-art methods. The source code is available athttps://github.com/zhutong0219/AMM-Net.
Leida Li, Tong Zhu 0003, Pengfei Chen 0003, Yuzhe Yang 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2023 Quality Assessment of UGC Videos Based on Decomposition and Recomposition
abstract
The prevalence of short-video applications imposes more requirements for video quality assessment (VQA). User-generated content (UGC) videos are captured under an unprofessional environment, thus suffering from various dynamic degradations, such as camera shaking. To cover the dynamic degradations, existing recurrent neural network-based UGC-VQA methods can only provide implicit modeling, which is unclear and difficult to analyze. In this work, we consider explicit motion representation for dynamic degradations, and propose a motion-enhanced UGC-VQA method based on decomposition and recomposition. In the decomposition stage, a dual-stream decomposition module is built, and VQA task is decomposed into single frame-based quality assessment problem and cross frames-based motion understanding. The dual streams are well grounded on the two-pathway visual system during perception, and require no extra UGC data due to knowledge transfer. Hierarchical features from shallow to deep layers are gathered to narrow the gaps from tasks and domains. In the recomposition stage, a progressively residual aggregation module is built to recompose features from the dual streams. Representations with different layers and pathways are interacted and aggregated in a progressive and residual manner, which keeps a good trade-off between representation deficiency and redundancy. Extensive experiments on UGC-VQA databases verify that our method achieves the state-of-the-art performance and keeps a good capability of generalization. The source code will be available inhttps://github.com/Sissuire/DSD-PRO.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.3
2023 An Underwater Image Quality Assessment Metric
abstract
Various image enhancement algorithms are adopted to improve underwater images that often suffer from visual distortions. It is critical to assess the output quality of underwater images undergoing enhancement algorithms, and use the results to optimise underwater imaging systems. In our previous study, we created a benchmark for quality assessment of underwater image enhancement via subjective experiments. Building on the benchmark, this paper proposes a new objective metric that can automatically assess the output quality of image enhancement, namely UWEQM. By characterising specific underwater physics and relevant properties of the human visual system, image quality attributes are computed and combined to yield an overall metric. Experimental results show that the proposed UWEQM metric yields good performance in predicting image quality as perceived by human subjects.
Hantao Liu, Delu Zeng, Tao Xiang 0001, Leida Li, Ke Gu 0001
IEEE Trans. Multim.5
2023 Explainable and Generalizable Blind Image Quality Assessment via Semantic Attribute Reasoning
abstract
Blind image quality assessment (BIQA) that can directly evaluate image quality without perfect-quality reference has been a long-standing research topic. Although the existing BIQA models have achieved very encouraging performance, the lack of explainability and generalization ability limits their real-world applications to a great extent. People usually assess image quality according to semantic attributes, e.g., brightness, color, contrast, noise and sharpness. Furthermore, judgment on image quality is also impacted by the scene presented in the image. Therefore, the inherent relationship between semantic attributes and scenes is crucial for image quality assessment, which has rarely been explored yet. With this motivation, this paper presents a Semantic Attribute Reasoning based image QUality Evaluator (SARQUE). Specifically, we propose a two-stream network to predict semantic attributes and scene categories from distorted images. To investigate the inherent relationship between the semantic attributes and scene category, a semantic reasoning module is further proposed based on the graph convolution network (GCN), producing the final quality score. Extensive experiments conducted on five in-the-wild image quality databases demonstrate the superiority of the proposed SARQUE model over the state-of-the-arts. Furthermore, the proposed model features better explainability and generalization ability due to the use of semantic attributes.
Yipo Huang, Leida Li, Yuzhe Yang 0001, Yandong Guo
IEEE Trans. Multim.2
2023 Knowledge-Guided Blind Image Quality Assessment With Few Training Samples
abstract
Blind image quality assessment (BIQA) for in-the-wild images has achieved great progress by training advanced deep neural networks. However, the current BIQA models are suffering the generalization challenge, meaning that a well-trained BIQA model is still very limited in evaluating images with different distributions. Deep BIQA models are data-intensive, but the annotation of image quality labels is extremely expensive. To design a generalizable BIQA model with few training samples is highly desired. Motivated by the above fact, this paper presents a knowledge-guided BIQA (KG-IQA) framework by integrating domain knowledge from the human visual system (HVS) and natural scene statistics (NSS). Specifically, the quality-aware HVS and NSS features are first extracted as prior knowledge. Then, we embed the two types of knowledge into the conventional deep neural network by learning to predict the HVS and NSS features, producing the knowledge-enhanced quality features, based on which the final image quality score is obtained. We conduct extensive experiments and comparisons on five authentically distorted IQA datasets. The experimental results demonstrate that the introduction of knowledge greatly reduces the requirement on the amount of training images, and the proposed KG-IQA model achieves superior performance in terms of both prediction accuracy and generalization ability.
Tianshu Song, Leida Li, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi
IEEE Trans. Multim.2
2023 Context Region Identification Based Quality Assessment of 3D Synthesized Views
abstract
Perceptual quality assessment of 3D synthesized views is an open research problem in computer vision. Researchers across the globe have developed several algorithms to identify distortions. At the same time, the existing algorithms cannot quantify the context in which these distortions affect the overall perceptual quality. According to the recently proposed 3D view synthesis algorithm, the choice of context region for the disocclusion plays a vital role in predicting the quality of 3D views. The context region taken from the background of a view produces a perceptually better quality of 3D synthesized views than when the context region is taken from the foreground. With this view, the proposed algorithm aims to identify the context region and incorporate this information for the perceptual quality assessment of 3D synthesized views. We observed that the depth energy maps of the 3D synthesized views vary significantly with the change in the context region and subsequently can identify the context region. Hence, in this work, we propose a new and efficient quality assessment algorithm based upon the variation in the depth of 3D synthesized and reference views, giving two-fold advantages: 1. It can predict the quality based on whether the context region is foreground or not. 2. It is also able to suggest the possible location of distortions. We have proposed two new algorithms for both situations when the context region is foreground or not. The overall predicted score is the direct multiplication of the quality score estimated when the context region is foreground or not. When applied to the established benchmark dataset, the proposed technique performs satisfactorily with the PLCC of 0.7707 and 0.7572 of SRCC. Also, the proposed algorithm can work as a plug-in to improve the performance of the existing algorithms.
Sadbhawna Thakur, Vinit Jakhetiya, Badri N. Subudhi, Sunil Prasad Jaiswal, Leida Li, Weisi Lin
IEEE Trans. Multim.5
2023 Semi-Supervised Authentically Distorted Image Quality Assessment With Consistency-Preserving Dual-Branch Convolutional Neural Network
abstract
Recently, convolutional neural networks (CNNs) have provided a favoured prospect for authentically distorted image quality assessment (IQA). For good performance, most existing CNN-based methods rely on a large amount of labeled data for training, which is time-consuming and cumbersome to collect. By simultaneously exploiting few labeled data and many unlabeled data, we make a pioneering attempt to propose a semi-supervised framework (termed SSLIQA) with consistency-preserving dual-branch CNN for authentically distorted IQA in this paper. The proposed SSLIQA introduces a consistency-preserving strategy and transfers two kinds of consistency knowledge from the teacher branch to the student branch. Concretely, SSLIQA utilizes the sample prediction consistency to train the student to mimic output activations of individual examples represented by the teacher. Considering that subjects often refer to previous analogous cases to make scoring decisions, SSLIQA computes the semantic relation among different samples in a batch and encourages the consistency of sample semantic relation between two branches to explore extra quality-related information. Benefiting from the consistency-preserving strategy, we can exploit numerous unlabeled data to improve network's effectiveness and generalization. Experimental results on three authentically distorted IQA databases show that the proposed SSLIQA is stably effective under different student-teacher combinations and different labeled-to-unlabeled data ratios. In addition, it points out a new way on how to achieve higher performance with a smaller network.
Guanghui Yue 0001, Leida Li, Tianwei Zhou, Hantao Liu, Tianfu Wang 0001
IEEE Trans. Multim.3
2023 Pyramid Feature Aggregation for Hierarchical Quality Prediction of Stitched Panoramic Images
abstract
Panoramic image quality assessment (PIQA) is crucial to the successful application of technologies that can provide immersive visual experience. Stitching distortions are one of the main types of distortions that result in panoramic image degradation. However, most existing PIQA methods are general-purpose ones, which ignore the special characteristics of the stitching distortions caused by imperfect stitching algorithms. This results in unsatisfactory performance. To this end, we propose an effective stitched PIQA method, which consists of an imaginary reference generation (IRG) module and a hierarchical quality prediction (HQP) module. Among them, the IRG module is proposed to mimic the capability of the human visual system in imagining the raw version in the face of a degraded image. For the IRG module learning, we construct a large-scale database. The HQP module is presented to adapt to the particularity and complexity of stitching distortions, which is achieved by the pyramid feature aggregation. Extensive experiments and comparisons have been performed on the stitched PIQA database and the experimental results demonstrate the superiority of the proposed method in evaluating the quality of stitched panoramic images.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Jinjian Wu, Xinbo Gao 0001
IEEE Trans. Multim.4
2023 Multimodal Sentiment Analysis With Image-Text Interaction Network
abstract
More and more users are getting used to posting images and text on social networks to share their emotions or opinions. Accordingly, multimodal sentiment analysis has become a research topic of increasing interest in recent years. Typically, there exist affective regions that evoke human sentiment in an image, which are usually manifested by corresponding words in people's comments. Similarly, people also tend to portray the affective regions of an image when composing image descriptions. As a result, the relationship between image affective regions and the associated text is of great significance for multimodal sentiment analysis. However, most of the existing multimodal sentiment analysis approaches simply concatenate features from image and text, which could not fully explore the interaction between them, leading to suboptimal results. Motivated by this observation, we propose a new image-text interaction network (ITIN) to investigate the relationship between affective image regions and text for multimodal sentiment analysis. Specifically, we introduce a cross-modal alignment module to capture region-word correspondence, based on which multimodal features are fused through an adaptive cross-modal gating module. Moreover, considering the complementary role of context information on sentiment analysis, we integrate the individual-modal contextual feature representations for achieving more reliable prediction. Extensive experimental results and comparisons on public datasets demonstrate that the proposed model is superior to the state-of-the-art methods.
Tong Zhu 0003, Leida Li, Jufeng Yang, Sicheng Zhao, Hantao Liu, Jiansheng Qian
IEEE Trans. Multim.2
2023 Multimodal Emotion Classification With Multi-Level Semantic Reasoning Network
abstract
Nowadays, people are accustomed to posting images and associated text for expressing their emotions on social networks. Accordingly, multimodal sentiment analysis has drawn increasingly more attention. Most of the existing image-text multimodal sentiment analysis methods simply predict the sentiment polarity. However, the same sentiment polarity may correspond to quite different emotions, such as happiness vs. excitement and disgust vs. sadness. Therefore, sentiment polarity is ambiguous and may not convey the accurate emotions that people want to express. Psychological research has shown that objects and words are emotional stimuli and that semantic concepts can affect the role of stimuli. Inspired by this observation, this paper presents a new MUlti-Level SEmantic Reasoning network (MULSER) for fine-grained image-text multimodal emotion classification, which not only investigates the semantic relationship among objects and words respectively, but also explores the semantic relationship between regional objects and global concepts. For image modality, we first build graphs to extract objects and global representation, and employ a graph attention module to perform bilevel semantic reasoning. Then, a joint visual graph is built to learn the regional-global semantic relations. For text modality, we build a word graph and further apply graph attention to reinforce the interdependencies among words in a sentence. Finally, a cross-modal attention fusion module is proposed to fuse semantic-enhanced visual and textual features, based on which informative multimodal representations are obtained for fine-grained emotion classification. The experimental results on public datasets demonstrate the superiority of the proposed model over the state-of-the-art methods.
Tong Zhu 0003, Leida Li, Jufeng Yang, Sicheng Zhao, Xiao Xiao 0007
IEEE Trans. Multim.2
2023 Learning Personalized Image Aesthetics From Subjective and Objective Attributes
abstract
Due to the widespread popularity of social media, researchers have developed a strong interest in learning the personalized image aesthetics of online users. Personalized image aesthetics assessment (PIAA) aims to study the aesthetic preferences of individual users for images, which should be affected by the properties of both users and images. Existing PIAA approaches usually use the generic aesthetics learned from images as a prior model and adapt it to PIAA models through a small number of data annotated by individual users. However, the prior model merely learns the objective attributes of images, which is agnostic to the subjective attributes of users, complicating efficient learning of the personalized image aesthetics of individual users. Therefore, we propose a personalized image aesthetics assessment method that integrates the subjective attributes of users and objective attributes of images simultaneously. To characterize these two attributes jointly, an attribute extraction module is introduced to learn users’ personality traits and image aesthetic attributes. Then, an aesthetic prior model is built from numerous individual users’ annotated data, which leverages the personality traits of users and the aesthetic attributes of rated images as prior knowledge to model both the image aesthetic distribution and users’ residual scores relative to generic aesthetics simultaneously. Finally, a PIAA model is obtained by fine-tuning the aesthetic prior model with an individual user’s annotated data. Experiments demonstrate that the proposed method is superior to existing PIAA methods in learning individual users’ personalized image aesthetics.
Hancheng Zhu, Yong Zhou 0003, Leida Li, Yandong Guo
IEEE Trans. Multim.3
2022 Robust Depth Completion with Uncertainty-Driven Loss Functions
abstract
Recovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumulated outliers in the synthesized ground truth. In this work, we introduce uncertainty-driven loss functions to improve the robustness of depth completion and handle the uncertainty in depth completion. Specifically, we propose an explicit uncertainty formulation for robust depth completion with Jeffrey's prior. A parametric uncertain-driven loss is introduced and translated to new loss functions that are robust to noisy or missing data. Meanwhile, we propose a multiscale joint prediction model that can simultaneously predict depth and uncertainty maps. The estimated uncertainty map is also used to perform adaptive prediction on the pixels with high uncertainty, leading to a residual map for refining the completion results. Our method has been tested on KITTI Depth Completion Benchmark and achieved the state-of-the-art robustness performance in terms of MAE, IMAE, and IRMSE metrics.
Yufan Zhu, Weisheng Dong, Leida Li, Jinjian Wu, Xin Li 0005, Guangming Shi
AAAI3
2022 Personalized Image Aesthetics Assessment with Rich Attributes
abstract
Personalized image aesthetics assessment (PIAA) is challenging due to its highly subjective nature. People's aesthetic tastes depend on diversified factors, including image characteristics and subject characters. The existing PIAA databases are limited in terms of annotation diversity, especially the subject aspect, which can no longer meet the increasing demands of PIAA research. To solve the dilemma, we conduct so far, the most comprehensive subjective study of personalized image aesthetics and introduce a new Personalized image Aesthetics database with Rich Attributes (PARA), which consists of 31,220 images with annotations by 438 subjects. PARA features wealthy annotations, including 9 image-oriented objective attributes and 4 human-oriented subjective attributes. In addition, desensitized subject information, such as personality traits, is also provided to support study of PIAA and user portraits. A comprehensive analysis of the annotation data is provided and statistic study indicates that the aesthetic preferences can be mirrored by proposed subjective attributes. We also propose a conditional PIAA model by utilizing subject information as conditional prior. Experimental results indicate that the conditional PIAA model can outperform the control group, which is also the first attempt to demonstrate how image aesthetics and subject characters interact to produce the intricate personalized tastes on image aesthetics. We believe the database and the associated analysis would be useful for conducting next-generation PIAA study. The project page of PARA can be found at: https://cv-datasets.institutecv.com/#/data-sets.
Yuzhe Yang 0001, Liwu Xu, Leida Li, Nan Qie, Yandong Guo
CVPR3
2022 Uncertainty Learning in Kernel Estimation for Multi-stage Blind Image Super-Resolution
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
ECCV (18)5
2022 Self-feature Distillation with Uncertainty Modeling for Degraded Image Recognition
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
ECCV (24)5
2022 Starvqa: Space-Time Attention for Video Quality Assessment
abstract
Transformer based on self-attention mechanism is blooming in computer vision nowadays. However, its application to video quality assessment (VQA) has not been reported. Evaluating the quality of in-the-wild videos is challenging due to the unknown of pristine reference and shooting distortion. This paper presents a novel space-time attention network for the VQA problem, named StarVQA. StarVQA builds a Transformer by alternately concatenating the divided space-time attention. To adapt the Transformer architecture for training, StarVQA designs a vectorized regression loss by encoding the mean opinion score (MOS) to the probability vector and embedding a special vectorized label token as the learnable variable. To capture the long-range spatiotemporal dependencies of a video sequence, StarVQA encodes the space-time position information of each patch to the input of the Transformer. Various experiments are conducted on the de-facto in-the-wild video datasets, including LIVE-VQC, KoNViD-1k, LSVQ, and LSVQ-1080p. Experimental results demonstrate the superiority of StarVQA over the state-of-the-art. The source code is available at https://github.com/GZHU-DVL/StarVQA.
Fengchuang Xing, Yuan-Gen Wang, Hanpin Wang, Leida Li, Guopu Zhu
ICIP4
2022 Psychology Inspired Model for Hierarchical Image Aesthetic Attribute Prediction
abstract
Deep neural network has proved its effectiveness in image aesthetic quality assessment (IAQA), but still lacks reasonable interpretability. Aesthetic attributes provide rich intermediate-level information for understanding the underlying principles of image aesthetics, but has not been fully investigated. Psychological studies have shown that aesthetic experience involves hierarchical stages, i.e., human process image aesthetics following a staged information processing mechanism. Motivated by this, this paper presents a Hierarchical Image Aesthetic Attribute (HIAA) prediction model, aiming to imitate the staged mechanism of human aesthetic experience. Image aesthetic attributes are first divided into several hierarchical groups. Then, hierarchical features are extracted from the cascaded layers of the deep neural network to predict the aesthetic attributes in a group-wise manner. The overall image aesthetic score is also predicted by aggregating the hierarchical features. Experimental results demonstrate that the proposed HIAA model outperforms the state-of-the-arts in terms of both aesthetic attribute prediction and aesthetic score regression.
Leida Li, Jiachen Duan, Yuzhe Yang 0001, Liwu Xu, Yandong Guo
ICME1
2022 AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular Data
abstract
Dynamic Vision Sensor (DVS) is a compelling neuromorphic camera compared to conventional camera, but it suffers from fiercer noise. Due to the nature of irregular format and asynchronous readout, DVS data is always transformed into a regular tensor (e.g., 3D voxel or image) for deep learning method, which corrupts its own asynchronous properties. To maintain asynchronous, we establish an innovative asynchronous event denoise neural network, named AEDNet, which directly consumes the correlation of the irregular signal in spatial-temporal range without destroying its original structural property. Based on the property of continuation in temporal domain and discreteness in spatial domain, we decompose the DVS signal into two parts, i.e., temporal correlation and spatial affinity, and separately process these two parts. Our spatial feature embedding unit is a unique feature extraction module that extracts feature from event-level, which perfectly maintains its spatial-temporal correlation. To test effectiveness, we build a novel dataset named DVSCLEAN containing both simulated and real-world data. The experimental results of AEDNet achieve SOTA.
Huachen Fang, Jinjian Wu, Leida Li, Junhui Hou, Weisheng Dong, Guangming Shi
ACM Multimedia3
2022 Transductive Aesthetic Preference Propagation for Personalized Image Aesthetics Assessment
abstract
Personalized image aesthetics assessment (PIAA) aims at capturing individual aesthetic preference. Fine-tuning on personalized data has been proven to be effective in PIAA task. However, a fixed fine-tuning strategy may cause under/over-fitting on limited personal data and it also brings additional training cost. To alleviate these issues, we employ a meta learning-based Transductive Aesthetic Preference Propagation (TAPP-PIAA) algorithm under regression manner to substitute the fine-tuning strategy. Specifically, each user's data is regarded as a meta-task and spilt into support and query set. Then, we extract deep aesthetic features with a pre-trained generic image aesthetic assessment (GIAA) model. Next, we treat image features as graph nodes and their similarities as edge weights to construct an undirected nearest neighbor graph for inference. Instead of fine-tuning on support set, TAPP-PIAA propagates aesthetic preference from support to query set with a predefined propagation formula. Finally, to learn a generalizable aesthetic representation for various users, we optimize our TAPP-PIAA across different users with meta-learning framework. Experimental results indicate that our TAPP-PIAA can surpass the state-of-the-art methods on benchmark databases.
Yuzhe Yang 0001, Huaxiong Li, Haoxing Chen, Liwu Xu, Leida Li, Yandong Guo
ACM Multimedia6
2022 Learning for Motion Deblurring with Hybrid Frames and Events
abstract
Event camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia4
2022 Semantic Attribute Guided Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) measures the perceived beauty of images using a computational approach. People usually assess the aesthetics of an image according to semantic attributes, e.g., lighting, color, object emphasis, etc. However, the state-of-the-art IAA approaches usually follow the data-driven framework without considering the rich attributes contained in images. With this motivation, this paper presents a new semantic attribute guided IAA model, where the attention maps of semantic attributes are employed to enhance the representation ability of general aesthetic features for more effective aesthetics assessment. Specifically, we first design an attribute attention generation network to obtain the attention maps for different semantic attributes, which are utilized to weight the general aesthetic features, producing the semantic attribute-enhanced feature representations. Then, the Graph Convolutional Network (GCN) is employed to further investigate the inherent relationship among the enhanced aesthetic features, producing the final image aesthetics prediction. Extensive experiments and comparisons on three public IAA databases demonstrate the effectiveness of the proposed method.
Jiachen Duan, Pengfei Chen 0003, Leida Li, Jinjian Wu, Guangming Shi
VCIP3
2022 Blind image quality assessment based on progressive multi-task learning
Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi
Neurocomputing4
2022 Modeling content-attribute preference for personalized image esthetics assessment
Yuanyang Wang, Yihua Huang 0003, Xiumin Chen, Leida Li, Guangming Shi
Image Vis. Comput.4
2022 Lightweight multi-scale convolutional neural network for real time stereo matching
Yanbing Xue, Doudou Zhang, Leida Li, Shiyin Li
Image Vis. Comput.3
2022 Hierarchical discrepancy learning for image restoration quality assessment
Bo Hu 0008, Shuaijian Wang, Leida Li, Jiaxu Leng, Yuzhe Yang 0001, Xinbo Gao 0001
Signal Process.3
2022 SPIQ: A Self-Supervised Pre-Trained Model for Image Quality Assessment
abstract
Blind image quality assessment (BIQA) has witnessed a flourishing progress due to the rapid advances in deep learning technique. The vast majority of prior BIQA methods try to leverage models pre-trained on ImageNet to mitigate the data shortage problem. These well-trained models, however, can be sub-optimal when applied to BIQA task that varies considerably from the image classification domain. To address this issue, we make the first attempt to leverage the plentiful unlabeled data to conduct self-supervised pre-training for BIQA task. Based on the distorted images generated from the high-quality samples using the designed distortion augmentation strategy, the proposed pre-training is implemented by a feature representation prediction task. Specifically, patch-wise feature representations corresponding to a certain grid are integrated to make prediction for the representation of the patch below it. The prediction quality is then evaluated using a contrastive loss to capture quality-aware information for BIQA task. Experimental results conducted on KADID-10 k and KonIQ-10 k databases demonstrate that the learned pre-trained model can significantly benefit the existing learning based IQA models.
Pengfei Chen 0003, Leida Li, Qingbo Wu 0001, Jinjian Wu
IEEE Signal Process. Lett.2
2022 Blind Image Quality Index for Authentic Distortions With Local and Global Deep Feature Aggregation
abstract
Blind image quality assessment (BIQA) for authentic distortions is still a great challenge, even in today’s deep learning era. It has been widely acknowledged that local and global features are both indispensable for IQA, which play complementary roles. While combining local and global features is straightforward in traditional handcrafted feature-based IQA metrics, it is not an easy task in the deep learning framework. This is mainly due to the fact that deep neural networks typically require input images with a fixed size. Current metrics either resize the image or use local patches as input, which are problematic in that they cannot integrate local and global aspects as well as their interactions to achieve comprehensive quality evaluation. Motivated by the above facts, this paper presents a new BIQA metric for authentic distortions by aggregating local and global deep features in a Vision-Transformer framework. In the proposed metric, selective local regions and global content are simultaneously input for complementary feature extraction, and the Vision-Transformer is employed to build the relationship between different local patches and image quality. Self-attention mechanism is further adopted to explore the interaction between local and global deep features, producing the final image quality score. Extensive experiments on five authentically distorted IQA databases demonstrate that the proposed metric outperforms the state-of-the-arts in terms of both prediction performance and generalization ability.
Leida Li, Tianshu Song, Jinjian Wu, Weisheng Dong, Jiansheng Qian, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.1
2022 Spatiotemporal Representation Learning for Blind Video Quality Assessment
abstract
Blind video quality assessment (BVQA) is of great importance for video-related applications, yet still challenging even in this deep learning era. The difficulty lies in the shortage of large-scale labeled data, thus making it hard to train a robust spatiotemporal encoder for BVQA. To relieve such difficulty, we first build a video dataset, which contains over 320K samples suffering from various compression and transmission artifacts. While manually annotating the dataset with subjective perception is much labor-intensive and time-consuming, we adopt reference-based VQA algorithms to weakly label the data automatically. We consider that single weak label is derived from single knowledge, which is deficient and incomplete for VQA. To alleviate the bias from single weak label (i.e., single knowledge) in the weakly labeled dataset, we propose HEterogeneous Knowledge Ensemble (HEKE) for spatiotemporal representation learning. Compared to learning from single knowledge, learning with HEKE is thought to achieve a lower infimum theoretically, and obtain richer representation. On the basis of the built dataset and the HEKE methodology, a feature encoder specific to BVQA is formed, and directly extract spatiotemporal representation from videos. Then, the video quality can be either acquired in a completely BVQA manner without ground truth, or via a finetuning-based regressor with labels. Extensive experiments on various VQA databases show that our BVQA model with the pretrained encoder achieves the state-of-the-art performance. More surprisingly, even trained on the synthetic data, our model still shows competitive performance on authentic databases. The data and source code will be available athttps://github.com/Sissuire/BVQA-HEKE.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.3
2022 Blind Image Quality Assessment for Authentic Distortions by Intermediary Enhancement and Iterative Training
abstract
With the boom of deep neural networks, blind image quality assessment (BIQA) has achieved great processes. However, the current BIQA metrics are limited when evaluating low-quality images as compared to medium-quality and high-quality images, which restricts their applications in real world problems. In this paper, we first identify that two challenges caused by distribution shift and long-tailed distribution lead to the compromised performance on low-quality images. Then, we propose an intermediary enhancement-based bilateral network with iterative training strategy for solving these two challenges. Drawing on the experience of transitive transfer learning, the proposed metric adaptively introduces enhanced intermediary images to transfer more information to low-quality images for mitigating the distribution shift. Our metric also adopts an iterative training strategy to deal with the long-tailed distribution. This strategy decouples feature extraction and score regression for better representation learning and regressor training. It not only transfers the knowledge learned from the earlier stage to the latter stage, but also makes the model pay more attention to long-tailed low-quality images. We conduct extensive experiments on five authentically distorted image quality datasets. The results show that our metric significantly improves the evaluating performance on low-quality images and delivers state-of-the-art intra-dataset results. During generalization tests, our metric also achieves the best cross-dataset performance.
Tianshu Song, Leida Li, Pengfei Chen 0003, Hantao Liu, Jiansheng Qian
IEEE Trans. Circuits Syst. Video Technol.2
2022 Omnidirectional Image Quality Assessment by Distortion Discrimination Assisted Multi-Stream Network
abstract
Omnidirectional image (OI) quality assessment is crucial to facilitate the development of virtual reality (VR) related technology. In this work, a distortion discrimination assisted multi-stream network is proposed for OI quality assessment. The multi-stream architecture is constructed by generating the viewport images received by the retina at one point to simulate the characteristics of humans perceiving VR contents. Additionally, the strategy of generating several viewport image sets from one OI is proposed for data augmentation. Furthermore, the facts that the human brain has the ability for both quality assessment and distortion type distinguishment, and the process of human brain handling two tasks exists information interaction inspire us to employ an auxiliary distortion discrimination task to facilitate the quality assessment task learning. Extensive experiments conducted on two public OI databases demonstrate the superiority of the proposed method to both traditional 2D quality metrics and existing metrics specific for OIs. Moreover, utilizing the assistant task is proven to be more effective than the single task learning for OI quality evaluation. Better generalization performance is also verified to be another valuable trait of the proposed method.
Yu Zhou 0009, Yanjing Sun, Leida Li, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Generalizable No-Reference Image Quality Assessment via Deep Meta-Learning
abstract
Recently, researchers have shown great interest in using convolutional neural networks (CNNs) for no-reference image quality assessment (NR-IQA). Due to the lack of big training data, the efforts of existing metrics in optimizing CNN-based NR-IQA models remain limited. Furthermore, the diversity of distortions in images result in the generalization problem of NR-IQA models when trained with known distortions and tested on unseen distortions, which is an easy task for human. Hence, we propose a NR-IQA metric via deep meta-learning, which is highly generalizable in the face of unseen distortions. The fundamental idea is to learn the meta-knowledge shared by human when evaluating the quality of images with diversified distortions. Specifically, we define NR-IQA of different distortions as a series of tasks and propose a task selection strategy to build two task sets, which are characterized by synthetic to synthetic and synthetic to authentic distortions, respectively. Based on these two task sets, an optimization-based meta-learning is proposed to learn the generalized NR-IQA model, which can be directly used to evaluate the quality of images with unseen distortions. Extensive experiments demonstrate that our NR-IQA metric outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability.
Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.2
2022 Personalized Image Aesthetics Assessment via Meta-Learning With Bilevel Gradient Optimization
abstract
Typical image aesthetics assessment (IAA) is modeled for the generic aesthetics perceived by an "average" user. However, such generic aesthetics models neglect the fact that users' aesthetic preferences vary significantly depending on their unique preferences. Therefore, it is essential to tackle the issue for personalized IAA (PIAA). Since PIAA is a typical small sample learning (SSL) problem, existing PIAA models are usually built by fine-tuning the well-established generic IAA (GIAA) models, which are regarded as prior knowledge. Nevertheless, this kind of prior knowledge based on "average aesthetics" fails to incarnate the aesthetic diversity of different people. In order to learn the shared prior knowledge when different people judge aesthetics, that is, learn how people judge image aesthetics, we propose a PIAA method based on meta-learning with bilevel gradient optimization (BLG-PIAA), which is trained using individual aesthetic data directly and generalizes to unknown users quickly. The proposed approach consists of two phases: 1) meta-training and 2) meta-testing. In meta-training, the aesthetics assessment of each user is regarded as a task, and the training set of each task is divided into two sets: 1) support set and 2) query set. Unlike traditional methods that train a GIAA model based on average aesthetics, we train an aesthetic meta-learner model by bilevel gradient updating from the support set to the query set using many users' PIAA tasks. In meta-testing, the aesthetic meta-learner model is fine-tuned using a small amount of aesthetic data of a target user to obtain the PIAA model. The experimental results show that the proposed method outperforms the state-of-the-art PIAA metrics, and the learned prior model of BLG-PIAA can be quickly adapted to unseen PIAA tasks.
Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, Guangming Shi
IEEE Trans. Cybern.2
2022 Contrastive Self-Supervised Pre-Training for Video Quality Assessment
abstract
Video quality assessment (VQA) task is an ongoing small sample learning problem due to the costly effort required for manual annotation. Since existing VQA datasets are of limited scale, prior research tries to leverage models pre-trained on ImageNet to mitigate this kind of shortage. Nonetheless, these well-trained models targeting on image classification task can be sub-optimal when applied on VQA data from a significantly different domain. In this paper, we make the first attempt to perform self-supervised pre-training for VQA task built upon contrastive learning method, targeting at exploiting the plentiful unlabeled video data to learn feature representation in a simple-yet-effective way. Specifically, we implement this idea by first generating distorted video samples with diverse distortion characteristics and visual contents based on the proposed distortion augmentation strategy. Afterwards, we conduct contrastive learning to capture quality-aware information by maximizing the agreement on feature representations of future frames and their corresponding predictions in the embedding space. In addition, we further introduce distortion prediction task as an additional learning objective to push the model towards discriminating different distortion categories of the input video. Solving these prediction tasks jointly with the contrastive learning not only provides stronger surrogate supervision signals, but also learns the shared knowledge among the prediction tasks. Extensive experiments demonstrate that our approach sets a new state-of-the-art in self-supervised learning for VQA task. Our results also underscore that the learned pre-trained model can significantly benefit the existing learning based VQA models. Source code is available at https://github.com/cpf0079/CSPT.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.2
2022 From Whole Video to Frames: Weakly-Supervised Domain Adaptive Continuous-Time QoE Evaluation
abstract
Due to the rapid increase in video traffic and relatively limited delivery infrastructure, end users often experience dynamically varying quality over time when viewing streaming videos. The user quality-of-experience (QoE) must be continuously monitored to deliver an optimized service. However, modern approaches for continuous-time video QoE estimation require densely annotating the continuous-time QoE labels, which is labor-intensive and time-consuming. To cope with such limitations, we propose a novel weakly-supervised domain adaptation approach for continuous-time QoE evaluation, by making use of a small amount of continuously labeled data in the source domain and abundant weakly-labeled data (only containing the retrospective QoE labels) in the target domain. Specifically, given a pair of videos from source and target domains, effective spatiotemporal segment-level feature representation is first learned by a combination of 2D and 3D convolutional networks. Then, a multi-task prediction framework is developed to simultaneously achieve continuous-time and retrospective QoE predictions, where a quality attentive adaptation approach is investigated to effectively alleviate the domain discrepancy without hampering the prediction performance. This approach is enabled by explicitly attending to the video-level discrimination and segment-level transferability in terms of the domain discrepancy. Experiments on benchmark databases demonstrate that the proposed method significantly improves the prediction performance under the cross-domain setting.
Leida Li, Pengfei Chen 0003, Weisi Lin, Mai Xu, Guangming Shi
IEEE Trans. Image Process.1
2022 Seeking Subjectivity in Visual Emotion Distribution Learning
abstract
Visual Emotion Analysis (VEA), which aims to predict people's emotions towards different visual stimuli, has become an attractive research topic recently. Rather than a single label classification task, it is more rational to regard VEA as a Label Distribution Learning (LDL) problem by voting from different individuals. Existing methods often predict visual emotion distribution in a unified network, neglecting the inherent subjectivity in its crowd voting process. In psychology, the Object-Appraisal-Emotion model has demonstrated that each individual's emotion is affected by his/her subjective appraisal, which is further formed by the affective memory. Inspired by this, we propose a novel Subjectivity Appraise-and-Match Network (SAMNet) to investigate the subjectivity in visual emotion distribution. To depict the diversity in crowd voting process, we first propose the Subjectivity Appraising with multiple branches, where each branch simulates the emotion evocation process of a specific individual. Specifically, we construct the affective memory with an attention-based mechanism to preserve each individual's unique emotional experience. A subjectivity loss is further proposed to guarantee the divergence between different individuals. Moreover, we propose the Subjectivity Matching with a matching loss, aiming at assigning unordered emotion labels to ordered individual predictions in a one-to-one correspondence with the Hungarian algorithm. Extensive experiments and comparisons are conducted on public visual emotion distribution datasets, and the results demonstrate that the proposed SAMNet consistently outperforms the state-of-the-art methods. Ablation study verifies the effectiveness of our method and visualization proves its interpretability.
Jingyuan Yang 0002, Jie Li 0001, Leida Li, Xiumei Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.3
2022 Fine-Grained Image Quality Caption With Hierarchical Semantics Degradation
abstract
Blind image quality assessment (BIQA), which is capable of precisely and automatically estimating human perceived image quality with no pristine image for comparison, attracts extensive attention and is of wide applications. Recently, many existing BIQA methods commonly represent image quality with a quantitative value, which is inconsistent with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where humans are able to extract image contents with local to global semantics as they need. The mediocre quality value represents coarse or holistic image quality and fails to reflect degradation on hierarchical semantics. In this paper, to comply with human cognition, a novel quality caption model is inventively proposed to measure fine-grained image quality with hierarchical semantics degradation. Research on human visual system indicates there are hierarchy and reverse hierarchy correlations between hierarchical semantics. Meanwhile, empirical evidence shows that there are also bi-directional degradation dependencies between them. Thus, a novel bi-directional relationship-based network (BDRNet) is proposed for semantics degradation description, through adaptively exploring those correlations and degradation dependencies in a bi-directional manner. Extensive experiments demonstrate that our method outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability.
Wen Yang 0008, Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.4
2022 Video Quality Assessment With Serial Dependence Modeling
abstract
Video quality assessment (VQA) is much more challenging than image quality assessment, due to the difficulty of modeling temporal influence among frames. Most of the existing VQA methods usually isolate each moment within the video (i.e., it neglects the sequential nature), leading to a large gap from the subjective perception. Recent research on neuroscience suggests a serially dependent perception (SDP) mechanism in the human visual system (HVS). Namely, the HVS tends to incorporate the recent past visual experience to predict the present perception. Inspired by the SDP, we suggest that the HVS prefers stable and continuous degradations in videos due to their predictability, and exhibits less tolerance to interrupted and unpredictable disturbances. Thus, we introduce a novel serial dependence modeling (SDM) framework for full-reference VQA in this paper. Firstly, the instantaneous degradation is measured on both the static appearance and motion information for each glimpse of scenes. Since motion plays an important role in videos, two types of structures are extracted for motion representation, namely, an explicit content-based 3D structure and an implicit feature-based 2D structure. Next, an assessment-directed long-short term memory (A-LSTM) is proposed to capture the serial dependence among instantaneous degradations. With the consideration of the perceptual effect from the previous moment on the current one, especially the effect from the perceptually worst moment, the serially dependent degradation is characterized. Finally, by mimicking the subjective rating for video-viewing, an attention-based quality decision procedure is presented to acquire the final video quality. Experimental results on publicly available VQA databases demonstrate that the proposed method maintains good consistency with the subjective perception.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Multim.4
2021 A Circular-Structured Representation for Visual Emotion Distribution Learning
abstract
Visual Emotion Analysis (VEA) has attracted increasing attention recently with the prevalence of sharing images on social networks. Since human emotions are ambiguous and subjective, it is more reasonable to address VEA in a label distribution learning (LDL) paradigm rather than a single-label classification task. Different from other LDL tasks, there exist intrinsic relationships between emotions and unique characteristics within them, as demonstrated in psychological theories. Inspired by this, we propose a well-grounded circular-structured representation to utilize the prior knowledge for visual emotion distribution learning. To be specific, we first construct an Emotion Circle to unify any emotional state within it. On the proposed Emotion Circle, each emotion distribution is represented with an emotion vector, which is defined with three attributes (i.e., emotion polarity, emotion type, emotion intensity) as well as two properties (i.e., similarity, additivity). Besides, we design a novel Progressive Circular (PC) loss to penalize the dissimilarities between predicted emotion vector and labeled one in a coarse-to-fine manner, which further boosts the learning process in an emotion-specific way. Extensive experiments and comparisons are conducted on public visual emotion distribution datasets, and the results demonstrate that the proposed method outperforms the state-of-the-art methods.
Jingyuan Yang 0002, Jie Li 0001, Leida Li, Xiumei Wang 0002, Xinbo Gao 0001
CVPR3
2021 Unsupervised Curriculum Domain Adaptation for No-Reference Video Quality Assessment
abstract
During the last years, convolutional neural networks (C-NNs) have triumphed over video quality assessment (VQA) tasks. However, CNN-based approaches heavily rely on annotated data which are typically not available in VQA, leading to the difficulty of model generalization. Recent advances in domain adaptation technique makes it possible to adapt models trained on source data to unlabeled target data. However, due to the distortion diversity and content variation of the collected videos, the intrinsic subjectivity of VQA tasks hampers the adaptation performance. In this work, we propose a curriculum-style unsupervised domain adaptation to handle the cross-domain no-reference VQA problem. The proposed approach could be divided into two stages. In the first stage, we conduct an adaptation between source and target domains to predict the rating distribution for target samples, which can better reveal the subjective nature of VQA. From this adaptation, we split the data in target domain into confident and uncertain subdomains using the proposed uncertainty-based ranking function, through measuring their prediction confidences. In the second stage, by regarding samples in confident subdomain as the easy tasks in the curriculum, a fine-level adaptation is conducted between two subdomain-s to fine-tune the prediction model. Extensive experimental results on benchmark datasets highlight the superiority of the proposed method over the competing methods in both accuracy and speed. The source code is released at https://github.com/cpf0079/UCDA.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
ICCV2
2021 No-Reference Video Quality Assessment with Heterogeneous Knowledge Ensemble
abstract
Blind assessment of video quality is still challenging even in this deep learning era. The limited number of samples in existing databases is insufficient to learn a good feature extractor for video quality assessment (VQA), while manually labeling a larger database with subjective perception is very labor-intensive and time-consuming. To relieve such difficulty, we first collect 3589 high-quality video clips as the reference and build a large VQA dataset. The dataset contains more than 300K samples degraded by various distortion types due to compression and transmission error, and provides weak labels for each distorted sample with several full-reference VQA algorithms. To learn effective representation from the weakly labeled data, we alleviate the bias of single weak label (i.e., single knowledge) via learning from multiple heterogeneous knowledge. To this end, we propose a novel no-reference VQA (NR-VQA) method with HEterogeneous Knowledge Ensemble (HEKE). Comparing to learning from single knowledge, HEKE can theoretically reach a lower infimum, and learn richer representation due to the heterogeneity. Extensive experimental results show that the proposed HEKE outperforms existing NR-VQA methods, and achieves the state-of-the-art performance. The source code will be available at https://github.com/Sissuire/BVQA-HEKE.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia3
2021 Image Quality Caption with Attentive and Recurrent Semantic Attractor Network
abstract
In this paper, a novel quality caption model is inventively developed to assess the image quality with hierarchical semantics. Existing image quality assessment (IQA) methods usually represent image quality with a quantitative value, resulting in inconsistency with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where hierarchical semantics are extracted. The mediocre quality value fails to reflect degradations on hierarchical semantics. Therefore, a new IQA framework is proposed to describe the quality for needs-oriented cognition. A novel quality caption procedure is firstly introduced, in which the quality is represented as patterns of activation distributed across the diverse degradations on hierarchical semantics. Then, an attentive and recurrent semantic attractor network (ARSANet) is designed to activate the distributed patterns for image quality description. Experiments demonstrate that our method achieves superior performance and is highly compliant with human cognition.
Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia3
2021 Deep Maximum a Posterior Estimator for Video Denoising
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
Int. J. Comput. Vis.5
2021 Blind image quality prediction with hierarchical feature aggregation
Jinjian Wu, Wen Yang 0008, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
Inf. Sci.3
2021 Predicting the Quality of View Synthesis With Color-Depth Image Fusion
abstract
With the increasing prevalence of free-viewpoint video applications, virtual view synthesis has attracted extensive attention. In view synthesis, a new viewpoint is generated from the input color and depth images with a depth-image-based rendering (DIBR) algorithm. Current quality evaluation models for view synthesis typically operate on the synthesized images, i.e. after the DIBR process, which is computationally expensive. So a natural question is that can we infer the quality of DIBR-based synthesized images using the input color and depth images directly without performing the intricate DIBR operation. With this motivation, this paper presents a no-reference image quality prediction model for view synthesis via COlor-Depth Image Fusion, dubbed CODIF, where the actual DIBR is not needed. First, object boundary regions are detected from the color image, and a Wavelet-based image fusion method is proposed to imitate the interaction between color and depth images during the DIBR process. Then statistical features of the interactional regions and natural regions are extracted from the fused color-depth image to portray the influences of distortions in color/depth images on the quality of synthesized views. Finally, all statistical features are utilized to learn the quality prediction model for view synthesis. Extensive experiments on public view synthesis databases demonstrate the advantages of the proposed metric in predicting the quality of view synthesis, and it even suppresses the state-of-the-art post-DIBR view synthesis quality metrics.
Leida Li, Yipo Huang, Jinjian Wu, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Temporal Reasoning Guided QoE Evaluation for Mobile Live Video Broadcasting
abstract
Quality of experience (QoE) that serves as a direct evaluation of viewing experience from the end users is of vital importance for network optimization, and should be constantly monitored. Unlike existing video-on-demand streaming services, real-time interactivity is critical to the mobile live broadcasting experience for both broadcasters and their audiences. While existing QoE metrics that are validated on limited video contents and synthetic stall patterns have shown effectiveness in their trained QoE benchmarks, a common caveat is that they often encounter challenges in practical live broadcasting scenarios, where one needs to accurately understand the activity in the video with fluctuating QoE and figure out what is going to happen to support the real-time feedback to the broadcaster. In this paper, we propose a temporal relational reasoning guided QoE evaluation approach for mobile live video broadcasting, namely TRR-QoE, which explicitly attends to the temporal relationships between consecutive frames to achieve a more comprehensive understanding of the distortion-aware variation. In our design, video frames are first processed by deep neural network (DNN) to extract quality-indicative features. Afterwards, besides explicitly integrating features of individual frames to account for the spatial distortion information, multi-scale temporal relational information corresponding to diverse temporal resolutions are made full use of to capture temporal-distortion-aware variation. As a result, the overall QoE prediction could be derived by combining both aspects. The results of experiments conducted on a number of benchmark databases demonstrate the superiority of TRR-QoE over the representative state-of-the-art metrics.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Yabin Zhang 0002, Weisi Lin
IEEE Trans. Image Process.2
2021 Blind Image Quality Assessment With Active Inference
abstract
Blind image quality assessment (BIQA) is a useful but challenging task. It is a promising idea to design BIQA methods by mimicking the working mechanism of human visual system (HVS). The internal generative mechanism (IGM) indicates that the HVS actively infers the primary content (i.e., meaningful information) of an image for better understanding. Inspired by that, this paper presents a novel BIQA metric by mimicking the active inference process of IGM. Firstly, an active inference module based on the generative adversarial network (GAN) is established to predict the primary content, in which the semantic similarity and the structural dissimilarity (i.e., semantic consistency and structural completeness) are both considered during the optimization. Then, the image quality is measured on the basis of its primary content. Generally, the image quality is highly related to three aspects, i.e., the scene information (content-dependency), the distortion type (distortion-dependency), and the content degradation (degradation-dependency). According to the correlation between the distorted image and its primary content, the three aspects are analyzed and calculated respectively with a multi-stream convolutional neural network (CNN) based quality evaluator. As a result, with the help of the primary content obtained from the active inference and the comprehensive quality degradation measurement from the multi-stream CNN, our method achieves competitive performance on five popular IQA databases. Especially in cross-database evaluations, our method achieves significant improvements.
Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie, Guangming Shi, Weisi Lin
IEEE Trans. Image Process.3
2021 SOLVER: Scene-Object Interrelated Visual Emotion Reasoning Network
abstract
Visual Emotion Analysis (VEA) aims at finding out how people feel emotionally towards different visual stimuli, which has attracted great attention recently with the prevalence of sharing images on social networks. Since human emotion involves a highly complex and abstract cognitive process, it is difficult to infer visual emotions directly from holistic or regional features in affective images. It has been demonstrated in psychology that visual emotions are evoked by the interactions between objects as well as the interactions between objects and scenes within an image. Inspired by this, we propose a novel Scene-Object interreLated Visual Emotion Reasoning network (SOLVER) to predict emotions from images. To mine the emotional relationships between distinct objects, we first build up an Emotion Graph based on semantic concepts and visual features. Then, we conduct reasoning on the Emotion Graph using Graph Convolutional Network (GCN), yielding emotion-enhanced object features. We also design a Scene-Object Fusion Module to integrate scenes and objects, which exploits scene features to guide the fusion process of object features with the proposed scene-based attention mechanism. Extensive experiments and comparisons are conducted on eight public visual emotion datasets, and the results demonstrate that the proposed SOLVER consistently outperforms the state-of-the-art methods by a large margin. Ablation studies verify the effectiveness of our method and visualizations prove its interpretability, which also bring new insight to explore the mysteries in VEA. Notably, we further discuss SOLVER on three other potential datasets with extended experiments, where we validate the robustness of our method and notice some limitations of it.
Jingyuan Yang 0002, Xinbo Gao 0001, Leida Li, Xiumei Wang 0002, Jinshan Ding
IEEE Trans. Image Process.3
2021 Blind Quality Assessment for Tone-Mapped Images by Analysis of Gradient and Chromatic Statistics
abstract
A tone-mapped image (TMI) obtained from the corresponding high dynamic range (HDR) image induces artifacts and distortion, which might result in the loss of structure information and impaired color. By analyzing the visual characteristics of TMIs, this work proposes a robust blind visual quality evaluation method for TMIs by using gradient and chromatic statistics (VQGC). First, motivated by the perceptual mechanism that the human visual system (HVS) is sensitive to image structure variation, we employ the gradient features to measure structure degradation in TMIs. To predict structure distortion accurately, we compute the gradient magnitude and orientation to measure image structure variation, and the relative gradient magnitude and orientation are also computed to capture microstructure change. Second, the color invariance descriptors are utilized to capture the visual degradation of colorfulness by local binary pattern (LBP) on four chromatic feature maps. Finally, the gradient and chromatic features are combined together as the final quality-aware feature vector, which is applied to assess the perceptual quality of TMIs by support vector regression (SVR). Comparison experiments show that the performance of the proposed method is better than other existing blind quality assessment methods on public databases.
Yuming Fang 0001, Jiebin Yan, Rengang Du, Yifan Zuo 0001, Wenying Wen, Yan Zeng 0001, Leida Li
IEEE Trans. Multim.7
2021 Quality Evaluation for Image Retargeting With Instance Semantics
abstract
To meet the ever-increasing demand for devices with diversified displays, image retargeting has become a prevalent technique for adaptive image resizing. In practice, the retargeting operation inevitably causes impairments in the images; thus, image retargeting quality assessment (IRQA) is urgently needed and, can be used to guide algorithm optimization, selection and design. Unlike traditional image quality assessment, image retargeting introduces geometric distortions, which typically affect high-level image semantics. With this motivation, this paper presents a quality evaluation model for image retargeting based on INstance SEMantics (INSEM). Considering that the human visual system (HVS) perceives images highly dependent on apprehensible areas and that impairments in image retargeting mainly degrade the salient instances, an image instance is utilized as the basic semantic unit, and a top-down method is devised to extract instance-level semantic features for IRQA. In addition, taking into account the influence of semantic categories on the perception of retargeting quality, we further propose Semantic-based self-adaptive pooling (SSAP) to integrate instance-based semantic features. Finally, global features are incorporated to generate quality scores that are more consistent with people's perceptions. Extensive experiments and comparisons of three public databases, in terms of both intradatabase and cross-database settings, demonstrate the superiority of the proposed metric over state-of-the-art methods.
Leida Li, Jinjian Wu, Lin Ma 0002, Yuming Fang 0001
IEEE Trans. Multim.1
2021 Quality Index for View Synthesis by Measuring Instance Degradation and Global Appearance
abstract
Virtual view synthesis plays a vital role in the application of multi-view and free-viewpoint videos. Depth-image-based rendering (DIBR) is the most commonly used approach in view synthesis, and many DIBR algorithms have been proposed. However, how to evaluate the quality of DIBR-synthesized images and benchmark the DIBR algorithms are still very challenging, which may hinder the further development of the view synthesis technique. Hence, an effective quality metric for evaluating the distortions in view synthesis is urgently needed. With this motivation, this paper presents a quality index for view synthesis by simultaneously measuring local Instance DEgradation and global Appearance (IDEA). Due to the imperfection of rendering algorithms, local geometric distortions are easily introduced around instance contours, causing instance degradation, which is the dominant distortion in synthesized views. In this work, image instances are first detected and local instance degradation is measured based on discrete orthogonal moments. Meantime, we propose to measure the global appearance of synthesized images based on the superpixel representation. By integrating both local and global aspects of the distortions, a more accurate quality model is built for view synthesis. Extensive experiments and comparisons have demonstrated the superiority of the proposed method in evaluating the quality of DIBR-synthesized images and benchmarking the performance of view synthesis algorithms.
Leida Li, Yu Zhou 0009, Jinjian Wu, Fu Li 0002, Guangming Shi
IEEE Trans. Multim.1
2021 Probabilistic Undirected Graph Based Denoising Method for Dynamic Vision Sensor
abstract
Dynamic Vision Sensor (DVS) is a new type of neuromorphic event-based sensor, which has an innate advantage in capturing fast-moving objects. Due to the interference of DVS hardware itself and many external factors, noise is unavoidable in the output of DVS. Different from frame/image with structural data, the output of DVS is in the form of address-event representation (AER), which means that the traditional denoising methods cannot be used for the output (i.e., event stream) of the DVS. In this paper, we propose a novel event stream denoising method based on probabilistic undirected graph model (PUGM). The motion of objects always shows a certain regularity/trajectory in space and time, which reflects the spatio-temporal correlation between effective events in the stream. Meanwhile, the event stream of DVS is composed by the effective events and random noise. Thus, a probabilistic undirected graph model is constructed to describe such priori knowledge (i.e., spatio-temporal correlation). The undirected graph model is factorized into the product of the cliques energy function, and the energy function is defined to obtain the complete expression of the joint probability distribution. Better denoising effect means a higher probability (lower energy), which means the denoising problem can be transfered into energy optimization problem. Thus, the iterated conditional modes (ICM) algorithm is used to optimize the model to remove the noise. Experimental results on denoising show that the proposed algorithm can effectively remove noise events. Moreover, with the preprocessing of the proposed algorithm, the recognition accuracy on AER data can be remarkably promoted.
Jinjian Wu, Chuanwei Ma, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.3
2020 MetaIQA: Deep Meta-Learning for No-Reference Image Quality Assessment
abstract
Recently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training DCNNs heavily relies on massive annotated data. Unfortunately, IQA is a typical small sample problem. Therefore, most of the existing DCNN-based IQA metrics operate based on pre-trained networks. However, these pre-trained networks are not designed for IQA task, leading to generalization problem when evaluating different types of distortions. With this motivation, this paper presents a no-reference IQA metric based on deep meta-learning. The underlying idea is to learn the meta-knowledge shared by human when evaluating the quality of images with various distortions, which can then be adapted to unknown distortions easily. Specifically, we first collect a number of NR-IQA tasks for different distortions. Then meta-learning is adopted to learn the prior knowledge shared by diversified distortions. Finally, the quality prior model is fine-tuned on a target NR-IQA task for quickly obtaining the quality model. Extensive experiments demonstrate that the proposed metric outperforms the state-of-the-arts by a large margin. Furthermore, the meta-model learned from synthetic distortions can also be easily generalized to authentic distortions, which is highly desired in real-world applications of IQA metrics.
Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
CVPR2
2020 CUID: A New Study Of Perceived Image Quality And Its Subjective Assessment
abstract
Research on image quality assessment (IQA) remains limited mainly due to our incomplete knowledge about human visual perception. Existing IQA algorithms have been designed or trained with insufficient subjective data with a small degree of stimulus variability. This has led to challenges for those algorithms to handle complexity and diversity of real-world digital content. Perceptual evidence from human subjects serves as a grounding for the development of advanced IQA algorithms. It is thus critical to acquire reliable subjective data with controlled perception experiments that faithfully reflect human behavioural responses to distortions in visual signals. In this paper, we present a new study of image quality perception where subjective ratings were collected in a controlled lab environment. We investigate how quality perception is affected by a combination of different categories of images and different types and levels of distortions. The database will be made publicly available to facilitate calibration and validation of IQA algorithms.
Lucie Lévêque, Kenneth Dasalla, Leida Li, Hantao Liu
ICIP6
2020 Active Inference of GAN for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) is a challenging task. It is a promising idea to design NR-IQA algorithms by mimicking how human visual system (HVS) works. The internal generative mechanism (IGM) indicates that HVS actively infers the primary content of an image for better understanding. Inspired by that, a novel NR-IQA method with active inference is proposed in this paper. First, a generative adversarial network (GAN) is proposed to predict the primary content of a distorted image, in which two IGM-inspired constraints are considered during the optimization. Next, based on the correlation between the distorted image and its primary content, different degradations (i.e., the content/distortion-/structure-dependency degradation) are measured simultaneously with a multi-stream convolutional neural network (CNN) for NR-IQA. Benefit from the primary content obtained from GAN and the multiple degradations measurement of CNN, our method achieves the state-of-the-art on five public IQA databases.
Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie
ICME3
2020 RIRNet: Recurrent-In-Recurrent Network for Video Quality Assessment
abstract
Video quality assessment (VQA), which is capable of automatically predicting the perceptual quality of source videos especially when reference information is not available, has become a major concern for video service providers due to the growing demand for video quality of experience (QoE) by end users. While significant advances have been achieved from the recent deep learning techniques, they often lead to misleading results in VQA tasks given their limitations on describing 3D spatio-temporal regularities using only fixed temporal frequency. Partially inspired by psychophysical and vision science studies revealing the speed tuning property of neurons in visual cortex when performing motion perception (i.e., sensitive to different temporal frequencies), we propose a novel no-reference (NR) VQA framework named Recurrent-In-Recurrent Network (RIRNet) to incorporate this characteristic to prompt an accurate representation of motion perception in VQA task. By fusing motion information derived from different temporal frequencies in a more efficient way, the resulting temporal modeling scheme is formulated to quantify the temporal motion effect via a hierarchical distortion description. It is found that the proposed framework is in closer agreement with quality perception of the distorted videos since it integrates concepts from motion perception in human visual system (HVS), which is manifested in the designed network structure composed of low- and high- level processing. A holistic validation of our methods on four challenging video quality databases demonstrates the superior performances over the state-of-the-art methods.
Pengfei Chen 0003, Leida Li, Lei Ma 0003, Jinjian Wu, Guangming Shi
ACM Multimedia2
2020 No-reference quality assessment for live broadcasting videos in temporal and spatial domains
abstract
Nowadays, live broadcasting video has become increasingly popular and high‐quality live broadcasting video is highly needed. In practice, live broadcasting videos usually undergo several processing stages, which inevitably introduce multiple distortions, e. g. frame freezing and intensity mutation, causing the degraded quality of experience. However, little work has been done to the quality evaluation of live broadcasting videos, which may hinder the further development of more advanced live broadcasting video delivery systems. Motivated by this, this study presents a no‐reference quality evaluation model for live broadcasting videos (LBVQA) in temporal and spatial domains. In the temporal domain, statistic features are extracted to measure the frame freezing and intensity mutation, and the entropy‐based feature is extracted to describe the global jitter. In the spatial domain, blurring is measured based on phase coherence, and abnormal exposure ratio is calculated based on an adaptive threshold. Finally, all features are fed into a backpropagation neural network to train the quality prediction model. Experimental results on the Live Broadcasting Video Database demonstrate the advantages of the proposed metric over the state‐of‐the‐art image and video quality metrics.
Yipo Huang, Leida Li, Yu Zhou 0009, Bo Hu 0008
IET Image Process.2
2020 No-reference quality index of depth images based on statistics of edge profiles for view synthesis
Leida Li, Jinjian Wu, Shiqi Wang 0001, Guangming Shi
Inf. Sci.1
2020 Inferring Personality Traits from Attentive Regions of User Liked Images Via Weakly Supervised Dual Convolutional Network
Hancheng Zhu, Leida Li, Allen Tan
Neural Process. Lett.2
2020 On the use of a scanpath predictor and convolutional neural network for blind image quality assessment
Aladine Chetouani, Leida Li
Signal Process. Image Commun.2
2020 Subjective and objective quality assessment for image restoration: A critical survey
Bo Hu 0008, Leida Li, Jinjian Wu, Jiansheng Qian
Signal Process. Image Commun.2
2020 Perceptual quality assessment for multimodal medical image fusion
Lu Tang 0001, Chuangeng Tian, Leida Li, Bo Hu 0008
Signal Process. Image Commun.3
2020 Blind Quality Index of Depth Images Based on Structural Statistics for View Synthesis
abstract
The quality of depth images is crucial for virtual view synthesis. However, the quality assessment of depth images is still largely unexplored. This letter presents a blind quality metric of Depth image based on Structural Statistics (DSS). The design philosophy is inspired by the fact that structural distortion in the depth images usually leads to geometric distortion, which is the main cause for degraded quality of synthesized views. Specifically, the statistical features for shape and orientation are calculated based on discrete orthogonal moments and gradients, generating two groups of quality-aware features. Then, the quality model is built from the extracted statistical features using a regression module. The experimental results demonstrate the effectiveness of the proposed metric.
Yipo Huang, Leida Li, Hancheng Zhu, Bo Hu 0008
IEEE Signal Process. Lett.2
2020 Perceptual Quality Assessment for Screen Content Images by Spatial Continuity
abstract
In this paper, we propose an effective blind quality assessment method for screen content images (SCIs), called perceptual quality measure by spatial continuity (PQSC). With the center-surround mechanism in the human visual system (HVS), the proposed method extracts the statistical features on chromatic and textural variations in SCIs to measure the visual distortion. First, by considering the chromatic continuity between spatially adjacent pixels, photo-metric invariant chromatic descriptors are extracted as zero-order and first-order features. Second, motivated by the perceptual mechanism that the HVS is sensitive to image texture variation, we employ local ternary pattern operator to effectively depict the spatial continuity of texture. With these extracted chromatic and textural features, we further adopt histogram to compute the statistical chromatic and textural features. Support vector regression (SVR) is used to train the quality prediction model from visual features to human ratings. Experimental results on three public benchmark databases demonstrate that the performance of our method is superior to the current blind image quality assessment methods, even better than some full reference image quality assessment counterparts.
Yuming Fang 0001, Rengang Du, Yifan Zuo 0001, Wenying Wen, Leida Li
IEEE Trans. Circuits Syst. Video Technol.5
2020 Blind Realistic Blur Assessment Based on Discrepancy Learning
abstract
Blur is one of the most common distortions that degrade natural images. This stimulates the blossom of sharpness assessment metrics. Existing sharpness metrics possess good performance for evaluating simulated blur, but are limited for the more common realistic blur that are introduced during image capture and processing in real life. To this end, we propose an effective Realistic Blur Assessment method (RBA) based on discrepancy learning. First, motivated by the fact that the distortion-free reference images are usually unavailable in practice, but the Human Visual System (HVS) can still accurately perceive image sharpness by quantifying the perceptual discrepancy between the distorted image and the hallucinated reference image in mind, we propose to train a discrepancy generation model to automatically generate the discrepancy map from the distorted image analogous to the HVS. This is achieved by using a deep neural network with rich training images. With the discrepancy map, two sharpness-aware features, i.e. sparse representation based entropy of primitive and content-guided variation of power, are then extracted to severally quantify spatial visual information amount and spectral power. Finally, the two features are integrated to produce the overall sharpness score. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-arts.
Leida Li, Yu Zhou 0009, Ke Gu 0001, Yuzhe Yang 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Personality-Assisted Multi-Task Learning for Generic and Personalized Image Aesthetics Assessment
abstract
Traditional image aesthetics assessment (IAA) approaches mainly predict the average aesthetic score of an image. However, people tend to have different tastes on image aesthetics, which is mainly determined by their subjective preferences. As an important subjective trait, personality is believed to be a key factor in modeling individual's subjective preference. In this paper, we present a personality-assisted multi-task deep learning framework for both generic and personalized image aesthetics assessment. The proposed framework comprises two stages. In the first stage, a multi-task learning network with shared weights is proposed to predict the aesthetics distribution of an image and Big-Five (BF) personality traits of people who like the image. The generic aesthetics score of the image can be generated based on the predicted aesthetics distribution. In order to capture the common representation of generic image aesthetics and people's personality traits, a Siamese network is trained using aesthetics data and personality data jointly. In the second stage, based on the predicted personality traits and generic aesthetics of an image, an inter-task fusion is introduced to generate individual's personalized aesthetic scores on the image. The performance of the proposed method is evaluated using two public image aesthetics databases. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts in both generic and personalized IAA tasks.
Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Weisi Lin
IEEE Trans. Image Process.1
2020 Blind Quality Metric of DIBR-Synthesized Images in the Discrete Wavelet Transform Domain
abstract
Free viewpoint video (FVV) has received considerable attention owing to its widespread applications in several areas such as immersive entertainment, remote surveillance and distanced education. Since FVV images are synthesized via a depth image-based rendering (DIBR) procedure in the "blind" environment (without reference images), a real-time and reliable blind quality assessment metric is urgently required. However, the existing image quality assessment metrics are insensitive to the geometric distortions engendered by DIBR. In this research, a novel blind method of DIBR-synthesized images is proposed based on measuring geometric distortion, global sharpness and image complexity. First, a DIBR-synthesized image is decomposed into wavelet subbands by using discrete wavelet transform. Then, the Canny operator is employed to detect the edges of the binarized low-frequency subband and high-frequency subbands. The edge similarities between the binarized low-frequency subband and high-frequency subbands are further computed to quantify geometric distortions in DIBR-synthesized images. Second, the log-energies of wavelet subbands are calculated to evaluate global sharpness in DIBR-synthesized images. Third, a hybrid filter combining the autoregressive and bilateral filters is adopted to compute image complexity. Finally, the overall quality score is derived to normalize geometric distortion and global sharpness by the image complexity. Experiments show that our proposed quality method is superior to the competing reference-free state-of-the-art DIBR-synthesized image quality models.
Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Leida Li, Zhifang Xia, Lifang Wu
IEEE Trans. Image Process.4
2019 QoE Evaluation for Live Broadcasting Video
abstract
The great variations of videographic skills in shot environment, photographic apparatus, compression and processing protocols give rise to very complicated impairments in the live broadcasting videos, which can adversely impact the quality of experience (QoE) of end users. Evaluating QoE of these videos is of great significance. Given the fact that there is still no publicly available database that studies the combined effects of the distortions in the live broadcasting videos, we have built the Live Broadcasting Video Database (LBVD) with the associated QoE scores. Towards depicting the distortions in live broadcasting videos, totally 1013 videos were included in the database, with diversified, authentic distortions. A subjective evaluation of these videos is conducted, and the correlation results between the tested state-of-the-art objective metrics and the subjective QoE scores on this database reveal that further studies are in urgent need for a better objective QoE metric dedicated to the live broadcasting videos.
Pengfei Chen 0003, Leida Li, Yipo Huang, Fengfeng Tan
ICIP2
2019 Personality Driven Multi-task Learning for Image Aesthetic Assessment
abstract
With the prevalence of convolutional neural networks (CNNs), assessing the aesthetics of an image has gained great advances recently. Individual users often have different aesthetic preferences on images, which we believe are mainly affected by their personality traits. However, most of the current aesthetics models predict a generic aesthetic score based on handcrafted and/or learned feature representations, which are unified and thus cannot reflect the individual differences during image aesthetic rating. In this paper, we propose an end-to-end personality driven multi-task deep learning model to address this problem. Firstly, both image aesthetics and personality traits are learned from the proposed multi-task model. Then the personality features are employed to modulate the aesthetics features, producing the optimal generic image aesthetics scores. The experimental results on two public databases show that the proposed method is superior to the state-of-the-art approaches.
Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Allen Tan
ICME1
2019 Incremental Few-Shot Learning for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition has received increasing attention due to its important role in video surveillance applications. However, most existing methods are designed for a fixed set of attributes. They are unable to handle the incremental few-shot learning scenario, i.e. adapting a well-trained model to newly added attributes with scarce data, which commonly exists in the real world. In this work, we present a meta learning based method to address this issue. The core of our framework is a meta architecture capable of disentangling multiple attribute information and generalizing rapidly to new coming attributes. By conducting extensive experiments on the benchmark dataset PETA and RAP under the incremental few-shot setting, we show that our method is able to perform the task with competitive performances and low resource requirements.
Liuyu Xiang, Xiaoming Jin, Guiguang Ding, Jungong Han, Leida Li
IJCAI5
2019 Naturalness Preserved Image Aesthetic Enhancement with Perceptual Encoder Constraint
abstract
Typical supervised image enhancement pipeline is to minimize the distance between the enhanced image and the reference one. Pixel-wise and perceptual-wise loss functions could help to improve the general image quality, however are not very efficient in improving the image aesthetic quality. In this paper, we propose a novel Residual connected Dilated U-Net (RDU-Net) for improving the image aesthetic quality. By using different dilation rates, the RDU-Net can extract multiple receptive-field features and merge the maximum information from local to global, which are highly desired in image enhancement. Also, we propose an encoder constraint perceptual loss, which could teach the enhancement network to dig out the latent aesthetic factors and make the enhanced image more natural and aesthetically appealing. The proposed approach can alleviate the over-enhancement phenomenons. The experimental results show that the proposed perceptual loss function could give a steady back propagation and the proposed method outperforms the state-of-the-arts.
Leida Li, Yuzhe Yang 0001, Hancheng Zhu
ICMR1
2019 PDANet: Polarity-consistent Deep Attention Network for Fine-grained Visual Emotion Regression
abstract
Existing methods on visual emotion analysis mainly focus on coarse-grained emotion classification, i.e. assigning an image with a dominant discrete emotion category. However, these methods cannot well reflect the complexity and subtlety of emotions. In this paper, we study the fine-grained regression problem of visual emotions based on convolutional neural networks (CNNs). Specifically, we develop a Polarity-consistent Deep Attention Network (PDANet), a novel network architecture that integrates attention into a CNN with an emotion polarity constraint. First, we propose to incorporate both spatial and channel-wise attentions into a CNN for visual emotion regression, which jointly considers the local spatial connectivity patterns along each channel and the interdependency between different channels. Second, we design a novel regression loss, i.e. polarity-consistent regression (PCR) loss, based on the weakly supervised emotion polarity to guide the attention generation. By optimizing the PCR loss, PDANet can generate a polarity preserved attention map and thus improve the emotion regression performance. Extensive experiments are conducted on the IAPS, NAPS, and EMOTIC datasets, and the results demonstrate that the proposed PDANet outperforms the state-of-the-art approaches by a large margin for fine-grained visual emotion regression. Our source code is released at: https://github.com/ZizhouJia/PDANet.
Sicheng Zhao, Zizhou Jia, Hui Chen 0013, Leida Li, Guiguang Ding, Kurt Keutzer
ACM Multimedia4
2019 Structure-aware person search with self-attention and online instance aggregation matching
Cunyuan Gao, Rui Yao 0006, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu, Leida Li
Neurocomputing6
2019 No-reference image quality assessment with visual pattern degradation
Jinjian Wu, Man Zhang 0007, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
Inf. Sci.3
2019 Blind quality assessment of gamut-mapped images via local and global statistical analysis
Hao Cai 0004, Leida Li, Zili Yi, Minglun Gong
J. Vis. Commun. Image Represent.2
2019 No-reference quality assessment for contrast-distorted images based on multifaceted statistical representation of structure
Yu Zhou 0009, Leida Li, Hancheng Zhu, Hantao Liu, Shiqi Wang 0001, Yao Zhao 0001
J. Vis. Commun. Image Represent.2
2019 Fractional quaternion cosine transform and its application in color image copy-move forgery detection
Beijing Chen, Qingtang Su, Leida Li
Multim. Tools Appl.4
2019 Internal generative mechanism driven blind quality index for deblocked images
Bo Hu 0008, Leida Li, Jiansheng Qian
Multim. Tools Appl.2
2019 Blind quality index for tone-mapped images based on luminance partition
Pengfei Chen 0003, Leida Li, Xinfeng Zhang 0001, Shanshe Wang, Allen Tan
Pattern Recognit.2
2019 Towards a blind image quality evaluator using multi-scale second-order statistics
Hao Cai 0004, Leida Li, Zili Yi, Minglun Gong
Signal Process. Image Commun.2
2019 Quality assessment for view synthesis using low-level and mid-level structural representation
Yu Zhou 0009, Leida Li, Suiyi Ling, Patrick Le Callet
Signal Process. Image Commun.2
2019 Robust Localization of Interpolated Frames by Motion-Compensated Frame Interpolation Based on an Artifact Indicated Map and Tchebichef Moments
abstract
Motion-compensated frame interpolation (MCFI), a frame-interpolation technique to increase the motion continuity of low frame-rate video, can be utilized by counterfeiters for faking high bitrate video or splicing videos with different frame rates. For existing MCFI detectors, their performances are degraded under real-world scenarios such as H.264/AVC compression, noise, or blur. To address this issue, a robust MCFI detector is proposed to locate interpolated frames. By analyzing the distribution of residual energies within interpolated frames, we observe that there exist strong correlations between artifact regions and high residual energies. Thus, an artifact indicated map is introduced to select candidate artifact regions. Then, Tchebichef moments (TMs) are exploited to characterize the blurring effects or deformed structures among these regions. Specifically, the mean value of absolute high-order TMs of selected regions is used to model these temporal inconsistencies. Finally, a sliding window is adopted to locate interpolated frames, which are further refined by three post-processing operations. Chrominance information is also integrated with luminance information for robust identification of interpolated frames. Extensive experimental results show that compared with the state-of-the-art MCFI detectors, the proposed approach is more robust for compressed videos under various real-world scenarios.
Xiangling Ding, Ningbo Zhu, Leida Li, Yue Li 0016, Gaobo Yang
IEEE Trans. Circuits Syst. Video Technol.3
2019 No-Reference Quality Assessment for View Synthesis Using DoG-Based Edge Statistics and Texture Naturalness
abstract
View synthesis is a key technique in free-viewpoint video, which renders virtual views based on texture and depth images. The distortions in synthesized views come from two stages, i.e., the stage of the acquisition and processing of texture and depth images, and the rendering stage using depth-image-based-rendering (DIBR) algorithms. The existing view synthesis quality metrics are designed for the distortions caused by a single stage, which cannot accurately evaluate the quality of the entire view synthesis process. With the considerations that the distortions introduced by two stages both cause edge degradation and texture unnaturalness, and the Difference-of-Gaussian (DoG) representation is powerful in capturing image edge and texture characteristics by simulating the center-surrounding receptive fields of retinal ganglion cells of human eyes, this paper presents a no-reference quality index for Synthesized views using DoG-based Edge statistics and Texture naturalness (SET). To mimic the multi-scale property of the Human Visual System (HVS), DoG images are first calculated at multiple scales. Then the orientation selective statistics features and the texture naturalness features are calculated on the DoG images and the coarsest scale image, producing two groups of quality-aware features. Finally, the quality model is learnt from these features using the random forest regression model. Experimental results on two view synthesis image databases demonstrate that the proposed metric is advantageous over the relevant state-of-the-arts in dealing with the distortions in the whole view synthesis process.
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yuming Fang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2019 Pairwise-Comparison-Based Rank Learning for Benchmarking Image Restoration Algorithms
abstract
Image restoration has attracted substantial attention recently and many image restoration algorithms have been proposed for restoring latent clear images from degraded images. However, determining how to objectively evaluate the performances of these algorithms remains an open problem, which may hinder the further development of advanced image restoration techniques. Most image restoration-quality metrics are designed for specific restoration applications; hence, their generalization ability is limited. For benchmarking image restoration algorithms, the ranking of restored images that are generated via various algorithms, is the most heavily considered factor. Inspired by this, this paper presents a pairwise-comparison-based rank learning framework for benchmarking the performances of image restoration algorithms, which focuses on the relative quality ranking of restored images. Under the proposed framework, we further propose a general image restoration quality metric by integrating quality-aware features in both the spatial and frequency domains. The proposed metric exhibits good generalization performance, and it is applicable to various restoration applications. The results of extensive experiments that were conducted on eight public databases of five restoration scenarios demonstrate the superior performance of the proposed method over the existing quality metrics. Moreover, the proposed framework is used to improve the existing quality metrics for benchmarking image restoration algorithms and highly encouraging results are obtained.
Bo Hu 0008, Leida Li, Hantao Liu, Weisi Lin, Jiansheng Qian
IEEE Trans. Multim.2
2018 Image Quality Assessment Based Label Smoothing in Deep Neural Network Learning
abstract
For many computer vision problems, deep neural networks are trained and validated based on the assumption that the input images are pristine (i.e., artifact-free). However, digital images are subject to a wide range of distortions in real application scenarios, while the practical issues regarding image quality in high level visual information understanding have been largely ignored. In this paper, in view of the fact that most widely deployed deep learning models are susceptible to various image distortions, distorted images are involved for data augmentation in the deep neural network training process to learn a reliable model for practical applications. In particular, an image quality assessment based label smoothing method, which aims at regularizing the label distribution of training images, is further proposed to tune the objective functions in learning the neural network. Experimental results show that the proposed method is effective in dealing with both low and high quality images in the typical image classification task.
Zhuo Chen 0006, Weisi Lin, Shiqi Wang 0001, Long Xu 0001, Leida Li
ICASSP5
2018 Internal Generative Mechanism Driven Blind Quality Index for Deblocked Images
abstract
Image deblocking has been widely studied. However, the relevant quality evaluation of deblocked images remains an open problem. The deblocked images are usually contaminated by multiple distortions, typically blocking artifacts and blur. Although various quality metrics have been reported, they are not designed specially for deblocked images, so they cannot accurately predict the quality of deblocked images. To fill this gap, we propose a new quality metric for deblocked images. With the guidance of the internal generative mechanism (IG- M) theory, a deblocked image is first decomposed into two portions, i.e., the predicted and disorderly portions. Then the distortions in the predicted portion are evaluated. Specifically, the distortion-specific features are extracted to evaluate blocking artifacts and blur in the spatial domain, separately. The joint effect of blocking artifacts and blur is evaluated by extracting energy-based features in the Curvelet domain. Finally, all features are combined to train a random forest model for quality prediction of deblocked images. Experimental results conducted on a newly released DeBlocked Image Database (DBID) demonstrate that the proposed metric outperforms the existing relevant quality metrics.
Bo Hu 0008, Leida Li, Jiansheng Qian
ICIP2
2018 Perceptual quality evaluation for motion deblurring
abstract
Motion deblurring has been widely studied. However, the relevant quality evaluation of motion deblurred images remains an open problem. The motion deblurred images are usually contaminated by noise, ringing and residual blur (NRRB) simultaneously. Unfortunately, most of the existing quality metrics are not designed for multiply distorted images, so they are limited in predicting the quality of motion deblurred images. In this study, the authors propose a new quality metric for motion deblurred images by measuring NRRB. For a motion deblurred image, the noise level is first estimated. Then the ringing effect is measured by incorporating visual saliency model to adapt to the characteristic of the human visual system. A reblurring‐based method is proposed to extract similarity features between a motion deblurred image and its re‐blurred version for evaluating the residual blur. Finally, the overall quality score of a motion deblurred image is obtained by pooling the scores of noise, ringing and blur. Experimental results conducted on a motion deblurring database demonstrate that the proposed metric significantly outperforms the existing quality metrics. In addition, the proposed NRRB metric is used for improving the existing general‐purpose no‐reference metrics, and very encouraging results are achieved.
Bo Hu 0008, Leida Li, Jiansheng Qian
IET Comput. Vis.2
2018 Multiple-parameter fractional quaternion Fourier transform and its application in colour image encryption
abstract
In this study, by using the quaternion algebra, multiple‐parameter fractional quaternion Fourier transform (MPFrQFT) is proposed to generalise the conventional multiple‐parameter fractional Fourier transform (MPFrFT) to quaternion signal processing in a holistic manner. First, the new transform MPFrQFT and its inverse transform are defined. An efficient discrete implementation method of MPFrQFT is then proposed, in which the relationship between MPFrQFT and MPFrFT of four components is utilised for a quaternion signal. Finally, a new colour image encryption algorithm based on the proposed MPFrQFT and the double random phase encoding technique is proposed to evaluate the performance of the proposed MPFrQFT. Experimental results demonstrate that: (i) the computational time of the proposed implementation method is almost a half of the direct method's time; (ii) the proposed MPFrQFT‐based encryption algorithm has an overall better performance than eight compared algorithms in security test and robustness test: it is more secure than the compared frequency‐based algorithms due to the larger key space and the more sensitive key ‘transform orders’; it is also more robust than the compared spatial‐domain algorithms.
Beijing Chen, Leida Li, Dingcheng Wang, Xingming Sun
IET Image Process.4
2018 No-reference quality assessment of DIBR-synthesized videos by measuring temporal flickering
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yun Zhang 0002
J. Vis. Commun. Image Represent.2
2018 Training-free referenceless camera image blur assessment via hypercomplex singular value decomposition
Lijuan Tang, Qiaohong Li, Leida Li, Ke Gu 0001, Jiansheng Qian
Multim. Tools Appl.3
2018 A robust forgery detection algorithm for object removal by exemplar-based image inpainting
Dengyong Zhang, Zaoshan Liang, Gaobo Yang, Qingguo Li, Leida Li, Xingming Sun
Multim. Tools Appl.5
2018 Reduced-reference quality assessment of DIBR-synthesized images based on multi-scale edge intensity similarity
Yu Zhou 0009, Leida Li, Ke Gu 0001, Lijuan Tang
Multim. Tools Appl.3
2018 Evaluating attributed personality traits from scene perception probability
Hancheng Zhu, Leida Li, Sicheng Zhao
Pattern Recognit. Lett.2
2018 No Reference Quality Assessment for Screen Content Images With Both Local and Global Feature Representation
abstract
In this paper, we propose a novel no reference quality assessment method by incorporating statistical luminance and texture features (NRLT) for screen content images (SCIs) with both local and global feature representation. The proposed method is designed inspired by the perceptual property of the human visual system (HVS) that the HVS is sensitive to luminance change and texture information for image perception. In the proposed method, we first calculate the luminance map through the local normalization, which is further used to extract the statistical luminance features in global scope. Second, inspired by existing studies from neuroscience that high-order derivatives can capture image texture, we adopt four filters with different directions to compute gradient maps from the luminance map. These gradient maps are then used to extract the second-order derivatives by local binary pattern. We further extract the texture feature by the histogram of high-order derivatives in global scope. Finally, support vector regression is applied to train the mapping function from quality-aware features to subjective ratings. Experimental results on the public large-scale SCI database show that the proposed NRLT can achieve better performance in predicting the visual quality of SCIs than relevant existing methods, even including some full reference visual quality assessment methods.
Yuming Fang 0001, Jiebin Yan, Leida Li, Jinjian Wu, Weisi Lin
IEEE Trans. Image Process.3
2018 Quality Assessment of DIBR-Synthesized Images by Measuring Local Geometric Distortions and Global Sharpness
abstract
Depth-image-based rendering (DIBR) is a fundamental technique in free viewpoint video, which is widely adopted to synthesize virtual viewpoints. The warping and rendering operations in DIBR generally introduce geometric distortions and sharpness change. The state-of-the-art quality indices are limited in dealing with such images since they are sensitive to geometric changes. In this paper, a new quality model for DIBR-synthesized view images is presented by measuring LOcal Geometric distortions in disoccluded regions and global Sharpness (LOGS). A disoccluded region detection method is first proposed using SIFT-flow-based warping. Then, the sizes and distortion strength of local disoccluded regions are combined to generate a score. Furthermore, a reblurring-based strategy is proposed to quantify the global sharpness. Finally, the overall quality score is calculated by pooling the scores of local disoccluded regions and global sharpness. Experiments on four public DIBR-synthesized image/video databases show the superiority of the proposed metric over the state-of-the-art quality models. The proposed method is further adopted for boosting the performances of existing quality metrics and benchmarking DIBR algorithms, both achieving very promising results.
Leida Li, Yu Zhou 0009, Ke Gu 0001, Weisi Lin, Shiqi Wang 0001
IEEE Trans. Multim.1
2018 Blind Quality Index for Multiply Distorted Images Using Biorder Structure Degradation and Nonlocal Statistics
abstract
In the past decade, extensive image quality metrics have been proposed. The majority of them are tailored for the images that contain a specific type of distortion. However, in practice, the images are usually degraded by different types of distortions simultaneously. This poses great challenges to the existing quality metrics. Motivated by this, this paper proposes a no-reference quality index for the multiply distorted images using the biorder structure degradation and the nonlocal statistics. The design philosophy is inspired by the fact that the human visual system (HVS) is highly sensitive to the degradations of both the spatial contrast and the spatial distribution, which are prone to be changed by the joint effects of the multiple distortions. Specifically, the multiresolution representation of the image is first built by downsampling to simulate the hierarchical property of the HVS. Then, the structure degradation is calculated to measure the spatial contrast. Considering the fact that the human visual cortex has the separate mechanisms to perceive the first- and second-order structures, dubbed biorder structures, the degradations of biorder structures are calculated to account for the spatial contrast, producing the first group of the quality-aware features. Furthermore, the nonlocal self-similarity statistics is calculated to measure the spatial distribution, producing the second group of features. Finally, all the features are fed into the random forest regression model to learn the quality model for the multiply distorted images. Extensive experimental results conducted on the three public databases demonstrate the superiority of the proposed metric to the state-of-the-art metrics. Moreover, the proposed metric is also advantageous over the existing metrics in terms of the generalization ability.
Yu Zhou 0009, Leida Li, Jinjian Wu, Ke Gu 0001, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.2
2017 Perceptual evaluation of single-image super-resolution reconstruction
abstract
In recent years, single-image super-resolution (SR) reconstruction has aroused wide attention. Massive SR enhancement algorithms have been proposed. However, much less work has been down on the perceptual evaluation of SR enhanced images and the corresponding enhancement algorithms. In this work, we create a Super-resolution Reconstructed Image Database (SRID), which consists of images produced by two interpolation methods and six popular SR image enhancement algorithms at different amplification factors. Then, subjective experiment is conducted to collect the subjective scores by using the single-stimulus method. The performances of the SR image enhancement algorithms are then evaluated by the obtained subjective scores. Finally, the performances of the general-purpose no-reference (NR) image quality metrics are investigated on the SRID database. This study shows that it is difficult for the state-of-the-art NR image quality metrics to predict the quality of SR enhanced images.
Guangcheng Wang, Leida Li, Qiaohong Li, Ke Gu 0001, Zhaolin Lu, Jiansheng Qian
ICIP2
2017 Detection of image seam carving by using weber local descriptor and local binary patterns
Dengyong Zhang, Qingguo Li, Gaobo Yang, Leida Li, Xingming Sun
J. Inf. Secur. Appl.4
2017 An efficient and effective blind camera image quality metric via modeling quaternion wavelet coefficients
Lijuan Tang, Leida Li, Kezheng Sun, Zhifang Xia, Ke Gu 0001, Jiansheng Qian
J. Vis. Commun. Image Represent.2
2017 Detecting image seam carving with low scaling ratio using multi-scale spatial and spectral entropies
Dengyong Zhang, Ting Yin, Gaobo Yang, Leida Li, Xingming Sun
J. Vis. Commun. Image Represent.5
2017 Detecting video frame rate up-conversion based on frame-level analysis of average texture variation
Min Xia 0002, Gaobo Yang, Leida Li, Ran Li 0003, Xingming Sun
Multim. Tools Appl.3
2017 No-reference quality assessment of compressive sensing image recovery
Bo Hu 0008, Leida Li, Jinjian Wu, Shiqi Wang 0001, Lu Tang 0001, Jiansheng Qian
Signal Process. Image Commun.2
2017 Enhanced Just Noticeable Difference Model for Images With Pattern Complexity
abstract
The just noticeable difference (JND) in an image, which reveals the visibility limitation of the human visual system (HVS), is widely used for visual redundancy estimation in signal processing. To determine the JND threshold with the current schemes, the spatial masking effect is estimated as the contrast masking, and this cannot accurately account for the complicated interaction among visual contents. Research on cognitive science indicates that the HVS is highly adapted to extract the repeated patterns for visual content representation. Inspired by this, we formulate the pattern complexity as another factor to determine the total masking effect: the interaction is relatively straightforward with a limited masking effect in a regular pattern, and is complicated with a strong masking effect in an irregular pattern. From the orientation selectivity mechanism in the primary visual cortex, the response of each local receptive field can be considered as a pattern; therefore, in this paper, the orientation that each pixel presents is regarded as the fundamental element of a pattern, and the pattern complexity is calculated as the diversity of the orientation in a local region. Finally, considering both pattern complexity and luminance contrast, a novel spatial masking estimation function is deduced, and an improved JND estimation model is built. Experimental results on comparing with the latest JND models demonstrate the effectiveness of the proposed model, which performs highly consistent with the human perception. The source code of the proposed model is publicly available at http://web.xidian.edu.cn/wjj/en/index.html.
Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin, C.-C. Jay Kuo
IEEE Trans. Image Process.2
2017 No-Reference and Robust Image Sharpness Evaluation Based on Multiscale Spatial and Spectral Features
abstract
The human visual system exhibits multiscale characteristic when perceiving visual scenes. The hierarchical structures of an image are contained in its scale space representation, in which the image can be portrayed by a series of increasingly smoothed images. Inspired by this, this paper presents a no-reference and robust image sharpness evaluation (RISE) method by learning multiscale features extracted in both the spatial and spectral domains. For an image, the scale space is first built. Then sharpness-aware features are extracted in gradient domain and singular value decomposition domain, respectively. In order to take into account the impact of viewing distance on image quality, the input image is also down-sampled by several times, and the DCT-domain entropies are calculated as quality features. Finally, all features are utilized to learn a support vector regression model for sharpness prediction. Extensive experiments are conducted on four synthetically and two real blurred image databases. The experimental results demonstrate that the proposed RISE metric is superior to the relevant state-of-the-art methods for evaluating both synthetic and real blurring. Furthermore, the proposed metric is robust, which means that it has very good generalization ability.
Leida Li, Wenhan Xia, Weisi Lin, Yuming Fang 0001, Shiqi Wang 0001
IEEE Trans. Multim.1
2016 Aspect Ratio Similarity (ARS) for image retargeting quality assessment
abstract
During the past few years, there have been various kinds of content-aware image retargeting methods proposed for image resizing. However, the lack of effective objective retargeting quality metric limits the further development of image retargeting. Different from the traditional image quality assessment, the quality degradation of the retargeted images is mainly caused by the geometric changes due to retargeting. In this paper, we propose a practical approach to reveal the geometric changes during image retargeting, and design an Aspect Ratio Similarity (ARS) metric to predict the visual quality of the retargeted image. The experimental results on the widely used dataset show that the proposed metric outperforms the state of the arts.
Yabin Zhang 0002, Weisi Lin, Xinfeng Zhang 0001, Yuming Fang 0001, Leida Li
ICASSP5
2016 Quality assessment of 3D synthesized images via disoccluded region discovery
abstract
Depth-Image-Based-Rendering (DIBR) is fundamental in free-viewpoint 3D video, which has been widely used to generate synthesized views from multi-view images. The majority of DIBR algorithms cause disoccluded regions, which are the areas invisible in original views but emerge in synthesized views. The quality of synthesized images is mainly contaminated by distortions in these disoccluded regions. Unfortunately, traditional image quality metrics are not effective for these synthesized images because they are sensitive to geometric distortions. To solve the problem, this paper proposes an objective quality evaluation method for 3D Synthesized images via Disoccluded Region Discovery (SDRD). A self-adaptive scale transform model is first adopted to preprocess the images on account of the impacts of view distance. Then disoccluded regions are detected by comparing the absolute difference between the preprocessed synthesized image and the warped image of preprocessed reference image. Furthermore, the disoccluded regions are weighted by a weighting function proposed to account for the varying sensitivities of human eyes to the size of disoccluded regions. Experiments conducted on IRCCyN/IVC DIBR image database demonstrate that the proposed SDRD method remarkably outperforms traditional 2D and existing DIBR-related quality metrics.
Yu Zhou 0009, Leida Li, Ke Gu 0001, Yuming Fang 0001, Weisi Lin
ICIP2
2016 Color space identification from single images
abstract
In this paper, we focus on the problem of RGB color space identification from a single image. At the moment, RGB color spaces are widely adopted in photography for image producing. The problems with respect to color space identification, such as to get the consistent printing or displaying quality on screen devices and software applications and prevention of multimedia unauthorized usage(shown or printed by other device via gamut mapping), need to be concerned. Current techniques are all relying on EXchangeable Image File Format (EXIF) to extract color space information. In this paper, we use a two-dimensional non-causal regressive model to explore the image demosaicing properties in order to extract discriminative features and train them on SVM classifier for image color space detection without relying on EXIF. In our experiment, images in three different color spaces (sRGB, adobeRGB and pro PhotoRGB) are generated for color space identification task. The experimental results show that the proposed technique has an good performance on image color space identification.
Haoliang Li, Alex Chichung Kot, Leida Li
ISCAS3
2016 Perceptual evaluation of Compressive Sensing Image Recovery
abstract
Compressive sensing (CS) has been attracting tremendous attention in recent years. Extensive CS recovery algorithms have been proposed for effective image reconstruction. However, little work has been dedicated to the perceptual evaluation of CS image recovery algorithms and the corresponding recovered images. In this paper, we first build a Compressive Sensing Recovered Image Database (CSRID), which contains images generated by ten popular CS image recovery algorithms at different sensing rates. We then carry out a subjective experiment using the single-stimulus method to obtain the subjective qualities of the images. The subjective scores are then used to evaluate the performances of the CS image recovery algorithms. Finally, the performances of general-purpose no-reference (NR) quality metrics and image blur metrics are investigated on the CSRID database. Experimental results show that the state-of-the-art quality metrics are very limited in predicting the quality of CS recovered images.
Bo Hu 0008, Leida Li, Jiansheng Qian, Yuming Fang 0001
QoMEX2
2016 No-reference quality assessment of deblocked images
Leida Li, Yu Zhou 0009, Weisi Lin, Jinjian Wu, Xinfeng Zhang 0001, Beijing Chen
Neurocomputing1
2016 Reversible data hiding based on an adaptive pixel-embedding strategy and two-layer embedding
ShaoWei Weng, Jeng-Shyang Pan 0001, Leida Li
Inf. Sci.3
2016 Orientation selectivity based visual pattern for reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Guangming Shi, Leida Li, Yuming Fang 0001
Inf. Sci.4
2016 Detecting video frame-rate up-conversion based on periodic properties of edge-intensity
Gaobo Yang, Xingming Sun, Leida Li
J. Inf. Secur. Appl.4
2016 Color image quality assessment based on sparse representation and reconstruction residual
Leida Li, Wenhan Xia, Yuming Fang 0001, Ke Gu 0001, Jinjian Wu, Weisi Lin, Jiansheng Qian
J. Vis. Commun. Image Represent.1
2016 Blind quality index for camera images with natural scene statistics and patch-based sharpness assessment
Lijuan Tang, Leida Li, Ke Gu 0001, Xingming Sun, Jianying Zhang
J. Vis. Commun. Image Represent.2
2016 Perceptual quality evaluation for image defocus deblurring
Leida Li, Ya Yan, Yuming Fang 0001, Shiqi Wang 0001, Lu Tang 0001, Jiansheng Qian
Signal Process. Image Commun.1
2016 Visual structural degradation based reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Yuming Fang 0001, Leida Li, Guangming Shi, S. Issac Niwas
Signal Process. Image Commun.4
2016 No-Reference Image Blur Assessment Based on Discrete Orthogonal Moments
abstract
Blur is a key determinant in the perception of image quality. Generally, blur causes spread of edges, which leads to shape changes in images. Discrete orthogonal moments have been widely studied as effective shape descriptors. Intuitively, blur can be represented using discrete moments since noticeable blur affects the magnitudes of moments of an image. With this consideration, this paper presents a blind image blur evaluation algorithm based on discrete Tchebichef moments. The gradient of a blurred image is first computed to account for the shape, which is more effective for blur representation. Then the gradient image is divided into equal-size blocks and the Tchebichef moments are calculated to characterize image shape. The energy of a block is computed as the sum of squared non-DC moment values. Finally, the proposed image blur score is defined as the variance-normalized moment energy, which is computed with the guidance of a visual saliency model to adapt to the characteristic of human visual system. The performance of the proposed method is evaluated on four public image quality databases. The experimental results demonstrate that our method can produce blur scores highly consistent with subjective evaluations. It also outperforms the state-of-the-art image blur metrics and several general-purpose no-reference quality metrics.
Leida Li, Weisi Lin, Xuesong Wang 0001, Gaobo Yang, Khosro Bahrami, Alex Chichung Kot
IEEE Trans. Cybern.1
2016 Sparse Representation-Based Image Quality Index With Adaptive Sub-Dictionaries
abstract
Distortions cause structural changes in digital images, leading to degraded visual quality. Dictionary-based sparse representation has been widely studied recently due to its ability to extract inherent image structures. Meantime, it can extract image features with slightly higher level semantics. Intuitively, sparse representation can be used for image quality assessment, because visible distortions can cause significant changes to the sparse features. In this paper, a new sparse representation-based image quality assessment model is proposed based on the construction of adaptive sub-dictionaries. An overcomplete dictionary trained from natural images is employed to capture the structure changes between the reference and distorted images by sparse feature extraction via adaptive sub-dictionary selection. Based on the observation that image sparse features are invariant to weak degradations and the perceived image quality is generally influenced by diverse issues, three auxiliary quality features are added, including gradient, color, and luminance information. The proposed method is not sensitive to training images, so a universal dictionary can be adopted for quality evaluation. Extensive experiments on five public image quality databases demonstrate that the proposed method produces the state-of-the-art results, and it delivers consistently well performances when tested in different image quality databases.
Leida Li, Hao Cai 0004, Yabin Zhang 0002, Weisi Lin, Alex Chichung Kot, Xingming Sun
IEEE Trans. Image Process.1
2016 Backward Registration-Based Aspect Ratio Similarity for Image Retargeting Quality Assessment
abstract
During the past few years, there have been various kinds of content-aware image retargeting operators proposed for image resizing. However, the lack of effective objective retargeting quality assessment metrics limits the further development of image retargeting techniques. Different from traditional image quality assessment (IQA) metrics, the quality degradation during image retargeting is caused by artificial retargeting modifications, and the difficulty for image retargeting quality assessment (IRQA) lies in the alternation of the image resolution and content, which makes it impossible to directly evaluate the quality degradation like traditional IQA. In this paper, we interpret the image retargeting in a unified framework of resampling grid generation and forward resampling. We show that the geometric change estimation is an efficient way to clarify the relationship between the images. We formulate the geometric change estimation as a backward registration problem with Markov random field and provide an effective solution. The geometric change aims to provide the evidence about how the original image is resized into the target image. Under the guidance of the geometric change, we develop a novel aspect ratio similarity (ARS) metric to evaluate the visual quality of retargeted images by exploiting the local block changes with a visual importance pooling strategy. Experimental results on the publicly available MIT RetargetMe and CUHK data sets demonstrate that the proposed ARS can predict more accurate visual quality of retargeted images compared with the state-of-the-art IRQA metrics.
Yabin Zhang 0002, Yuming Fang 0001, Weisi Lin, Xinfeng Zhang 0001, Leida Li
IEEE Trans. Image Process.5
2016 Image Sharpness Assessment by Sparse Representation
abstract
Recent advances in sparse representation show that overcomplete dictionaries learned from natural images can capture high-level features for image analysis. Since atoms in the dictionaries are typically edge patterns and image blur is characterized by the spread of edges, an overcomplete dictionary can be used to measure the extent of blur. Motivated by this, this paper presents a no-reference sparse representation-based image sharpness index. An overcomplete dictionary is first learned using natural images. The blurred image is then represented using the dictionary in a block manner, and block energy is computed using the sparse coefficients. The sharpness score is defined as the variance-normalized energy over a set of selected high-variance blocks, which is achieved by normalizing the total block energy using the sum of block variances. The proposed method is not sensitive to training images, so a universal dictionary can be used to evaluate the sharpness of images. Experiments on six public image quality databases demonstrate the advantages of the proposed method.
Leida Li, Jinjian Wu, Haoliang Li, Weisi Lin, Alex Chichung Kot
IEEE Trans. Multim.1
2015 Detecting seam carving based image resizing using local binary patterns
Ting Yin, Gaobo Yang, Leida Li, Dengyong Zhang, Xingming Sun
Comput. Secur.3
2015 GridSAR: Grid strength and regularity for robust evaluation of blocking artifacts in JPEG images
Leida Li, Yu Zhou 0009, Jinjian Wu, Weisi Lin, Haoliang Li
J. Vis. Commun. Image Represent.1
2015 An efficient forgery detection algorithm for object removal by exemplar-based image inpainting
Zaoshan Liang, Gaobo Yang, Xiangling Ding, Leida Li
J. Vis. Commun. Image Represent.4
2015 Detection of seam carving-based video retargeting using forensics hash
abstract
Abstract Seam carving is a content‐aware multimedia retargeting technique to adaptively resize multimedia data for different display sizes. However, it can also be used to remove objects from digital object or video for malicious purposes. In this paper, a forensics hash‐based tampering detection and localization approach is proposed for seam carving‐based video retargeting. It extracts the invariant Speeded‐up Robust Feature points from every spatiotemporal image to represent the matching surface, and the relative position change of the neighboring matching surface is used to build the forensic hash in a compact and scalable way. Experimental results show that the proposed forensics approach can effectively estimate the exact amount and rough locations of deleted seam carving surfaces. It achieves desirable detection performance even when there are frames deleted. If the hash length is reasonably increased, it can estimate the rough location and exact amount of deleted frames. Moreover, the built forensics hash is of good robustness, scalability, and compactness. Copyright © 2014 John Wiley & Sons, Ltd.
Wei Fei, Gaobo Yang, Leida Li, Dengyong Zhang
Secur. Commun. Networks3
2015 Blurred Image Splicing Localization by Exposing Blur Type Inconsistency
abstract
In a tampered blurred image generated by splicing, the spliced region and the original image may have different blur types. Splicing localization in this image is a challenging problem when a forger uses some postprocessing operations as antiforensics to remove the splicing traces anomalies by resizing the tampered image or blurring the spliced region boundary. Such operations remove the artifacts that make detection of splicing difficult. In this paper, we overcome this problem by proposing a novel framework for blurred image splicing localization based on the partial blur type inconsistency. In this framework, after the block-based image partitioning, a local blur type detection feature is extracted from the estimated local blur kernels. The image blocks are classified into out-of-focus or motion blur based on this feature to generate invariant blur type regions. Finally, a fine splicing localization is applied to increase the precision of regions boundary. We can use the blur type differences of the regions to trace the inconsistency for the splicing localization. Our experimental results show the efficiency of the proposed method in the detection and the classification of the out-of-focus and motion blur types. For splicing localization, the result demonstrates that our method works well in detecting the inconsistency in the partial blur types of the tampered images. However, our method can be applied to blurred images only.
Khosro Bahrami, Alex Chichung Kot, Leida Li, Haoliang Li
IEEE Trans. Inf. Forensics Secur.3
2014 Learning Structural Regularity for Evaluating Blocking Artifacts in JPEG Images
abstract
Image degradation damages genuine visual structures and causes pseudo structures. Pseudo structures are usually present with regularities. This letter proposes a machine learning based blocking artifacts metric for JPEG images by measuring the regularities of pseudo structures. Image corner, block boundary and color change properties are used to differentiate the blocking artifacts. A support vector regression (SVR) model is adopted to learn the underlying relations between these features and perceived blocking artifacts. The blocking artifacts score of a test image is predicted using the trained model. Extensive experiments demonstrate the effectiveness of the method.
Leida Li, Weisi Lin, Hancheng Zhu
IEEE Signal Process. Lett.1
2014 Referenceless Measure of Blocking Artifacts by Tchebichef Kernel Analysis
abstract
This letter presents a Referenceless quality Measure of Blocking artifacts (RMB) using Tchebichef moments. It is based on the observation that Tchebichef kernels with different orders have varying abilities to capture blockiness. In a block manner, high-odd-order moments are computed to score the blocking artifacts. The blockiness scores are further weighted to incorporate the characteristic of Human Visual System (HVS), which is achieved by classifying the blocks into smooth and textured. Experimental results and comparisons demonstrate the advantage of the proposed method.
Leida Li, Hancheng Zhu, Gaobo Yang, Jiansheng Qian
IEEE Signal Process. Lett.1
2012 Detecting Removed Object from Video with Stationary Background
Leida Li, Gaobo Yang, Guozhang Hu
IWDW1
2012 Geometrically invariant image watermarking using Polar Harmonic Transforms
Leida Li, Shushang Li, Ajith Abraham, Jeng-Shyang Pan 0001
Inf. Sci.1
2010 Watermark Synchronization Based on Locally Most Stable Feature Points
Jiansheng Qian, Leida Li, Zhaolin Lu
ICCCI (2)2
2010 Forecasting Coal and Rock Dynamic Disaster Based on Adaptive Neuro-Fuzzy Inference System
Jianying Zhang, Jian Cheng 0004, Leida Li
ICCCI (2)3
2010 High capacity watermark embedding based on local invariant features
abstract
A novel robust image watermarking scheme is presented to embed a high capacity watermark into the feature point based characteristic regions. The watermark embedding positions are first determined by the scale-invariant feature transform (SIFT) based local circular regions. Then the binary watermark image is embedded by quantization in the Non-subsampled Contourlet Transform (NSCT) domain. In order to achieve rotation invariance, the watermark is embedded adaptively to the orientation of the region. Simulation results show that the proposed scheme can achieve high invisibility and it can efficiently resist traditional signal processing attacks and geometric attacks.
Leida Li, Jiansheng Qian, Jeng-Shyang Pan 0001
ICME1
2010 Rotation invariant watermark embedding based on scale-adapted characteristic regions
Leida Li, Xiaoping Yuan, Zhaolin Lu, Jeng-Shyang Pan 0001
Inf. Sci.1
2009 A New Histogram Based Image Watermarking Scheme Resisting Geometric Attacks
abstract
A histogram based image watermarking algorithm is proposed based on the invariant property of the statistic characteristic of an image for resisting geometric attacks. During embedding, a particular range of pixel values is first selected and divided into intervals according to a predefined step. Then the pixels within each interval are quantized to have the same value according to the watermark bit. During watermark extraction, the pixel values are divided into intervals using the same method as watermark embedding. A watermark bit can be extracted from each interval. Experimental results show that the proposed scheme is robust to rotation, scaling and cropping attacks, as well as some signal processing attacks.
Xinwei Li 0002, Leida Li, Hong-Xin Shen
IAS3
2008 Scale-Space Feature Based Image Watermarking in Contourlet Domain
Leida Li, Jeng-Shyang Pan 0001
IWDW1