Xiongkuo Min

dblp:139/6983 · DBLP profile ↗
← Back
269ranked-venue papers
17as first author
224since 2021 · last 2026
0000-0001-5693-0416ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 211 · 13 first-author · 175 since 2021Artificial intelligence and machine learning · 50 · 47 since 2021Computer networks · 19 · 1 first-author · 17 since 2021Systems, architecture and hardware · 16 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning
abstract
Face videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between visual and audio features, particularly in the mouth region. Several audio-aided face video restoration methods have been proposed, but they only focus on compression artifact removal. In this paper, we propose a General Audio-assisted face Video restoration Network (GAVN) to address various types of streaming video distortions via identity and temporal complementary learning. Specifically, GAVN first captures inter-frame temporal features in the low-resolution space to restore frames coarsely and save computational cost. Then, GAVN extracts intra-frame identity features in the high-resolution space with the assistance of audio signals and face landmarks to restore more facial details. Finally, the reconstruction module integrates temporal features and identity features to generate high-quality face videos. Experimental results demonstrate that GAVN outperforms the existing state-of-the-art methods on face video compression artifact removal, deblurring, and super-resolution.
Yuqin Cao, Wei Sun 0029, Xiaohong Liu 0001, Yulun Zhang 0001, Xiongkuo Min
AAAI6
2026 VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning
abstract
Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision.
Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min
AAAI9
2026 Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality Assessment
abstract
Reinforcement fine-tuning (RFT) is a proliferating paradigm for LMM training. Analogous to high-level reasoning tasks, RFT is similarly applicable to low-level vision domains, including image quality assessment (IQA). Existing RFT-based IQA methods typically use rule-based output rewards to verify the model's rollouts but provide no reward supervision for the "think” process, leaving its correctness and efficacy uncontrolled. Furthermore, these methods typically fine-tune directly on downstream IQA tasks without explicitly enhancing the model’s native low-level visual quality perception, which may constrain its performance upper bound. In response to these gaps, we propose the multi‐stage RFT IQA framework (Refine-IQA). In Stage-1, we build the Refine-Perception-20K dataset (with 12 main distortions, 20,907 locally-distorted images, and over 55K RFT samples) and design multi-task reward functions to strengthen the model’s visual quality perception. In Stage-2, targeting the quality scoring task, we introduce a probability difference reward involved strategy for "think" process supervision. The resulting Refine-IQA Series Models achieve outstanding performance on both perception and scoring tasks—and, notably, our paradigm activates a robust "think” (quality interpretating) capability that also attains exceptional results on the corresponding quality interpreting benchmark.
Ziheng Jia, Jiaying Qian, Zijian Chen 0001, Xiongkuo Min
AAAI5
2026 Scaling-up Perceptual Video Quality Assessment
abstract
The data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose OmniVQA, a framework designed to efficiently build high-quality, machine-dominated synthetic multi-modal instruction databases (MIDBs) for VQA. We then scale up to create OmniVQA-Chat-400K, the largest dataset in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we build the OmniVQA-MOS-20K dataset to enhance the model's quantitative quality rating capabilities. We then introduce a complementary training strategy that effectively leverages the knowledge from datasets for different tasks. Furthermore, we propose the OmniVQA-FG (fine-grain)-Benchmark to evaluate the fine-grained performance of models. Our results demonstrate that our models achieve state-of-the-art performance in both tasks.
Ziheng Jia, Xiaorong Zhu, Chunyi Li 0001, Jinliang Han, Xiaohong Liu 0001, Guangtao Zhai, Xiongkuo Min
AAAI8
2026 SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath Prediction
abstract
Scanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequate modeling and capturing of long-range temporal dependencies in fixation regions, and (2) suboptimal integration of spatial and temporal visual features, ultimately compromising prediction performance. To address these limitations, we propose a novel Dual-Temporal Modulated Diffusion model for Omnidirectional Images Scanpath Prediction, named SalDiff-DTM model, to effectively generate realistic human eye viewing trajectories. Specifically, to effectively model spatial relationships, we propose a novel Dual-Graph Convolutional Network (Dual-GCN) module that simultaneously captures semantic-level and image-level correlations. By integrating both local spatial details and global contextual information across the internal temporal dimension, this module achieves comprehensive and robust modeling of spatial relationships. To further enhance the modeling of temporal dependencies inherent in diverse fixation patterns, we introduce TABiMamba (Temporal-Aware BiLSTM-Mamba), a dedicated module that synergistically combines the contextual sensitivity of BiLSTM with the long-range sequence modeling capabilities of Mamba. This design facilitates deep information flow and context-aware sequential reasoning, thereby enabling high-fidelity capture of intricate temporal correlations. Inspired by the progressive refinement mechanism of diffusion models in various generative tasks, we propose a saliency-guided diffusion module that formulates the prediction problem as a conditional generative process, iteratively yielding accurate and perceptually plausible scanpaths. Extensive experiments demonstrate that SalDiff-DTM significantly outperforms state-of-the-art models, paving the way for future advancements in eye-tracking technologies and cognitive modeling.
Xiaohui Kong, Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min
AAAI5
2026 Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition
abstract
Yushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yushuo Zheng, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai
ACL (1)5
2026 DocIQ: A Benchmark Dataset and Feature Fusion Network for Document Image Quality Assessment
abstract
Document image quality assessment (DIQA) is an important component for various applications, including optical character recognition (OCR), document restoration, and the evaluation of document image processing systems. In this paper, we introduce a subjective DIQA dataset DIQA-5000. The DIQA-5000 dataset comprises 5,000 document images, generated by applying multiple document enhancement techniques to 500 real-world images with diverse distortions. Each enhanced image was rated by 15 subjects across three rating dimensions: overall quality, sharpness, and color fidelity. Furthermore, we propose a specialized no-reference DIQA model that exploits document layout features to maintain quality perception at reduced resolutions to lower computational cost. Recognizing that image quality is influenced by both low-level and high-level visual features, we designed a feature fusion module to extract and integrate multi-level features from document images. To generate multi-dimensional scores, our model employs independent quality heads for each dimension to predict score distributions, allowing it to learn distinct aspects of document image quality. Experimental results demonstrate that our method outperforms current state-of-the-art general-purpose IQA models on both DIQA-5000 and an additional document image dataset focused on OCR accuracy.
Fengjun Guo, Guangtao Zhai, Xiongkuo Min
ISCAS6
2026 Q-Agent: An MLLM-Driven Framework for Universal Visual Quality Assessment
Peihang Chen, Huiyu Duan, Zitong Xu, Yuqin Cao, Sijing Wu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
QoMEX13
2026 Assessing Personality Consistency in Large Language Models: A Psychometric Framework for Human-Centric Quality of Experience
Yitian Kou, Dandan Zhu 0001, Wei Sun 0029, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai
QoMEX5
2026 Towards versatile multimedia quality assessment for visual communications
Ziheng Jia, Chunyi Li 0001, Yingjie Zhou 0003, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
Sci. China Inf. Sci.6
2026 Enhancing blind video quality assessment with rich quality-aware features
Wei Sun 0029, Linhan Cao, Jun Jia, Xiongkuo Min, Guangtao Zhai
Expert Syst. Appl.6
2026 Beyond catastrophic forgetting: A continual learning-driven multi-modal fusion model for saliency prediction in dynamic scenes
Jiaqi Wang 0003, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
Expert Syst. Appl.5
2026 MI3S: A multimodal large language model assisted quality assessment framework for AI-generated talking heads
Yingjie Zhou 0003, Sijing Wu, Jun Jia, Yanwei Jiang, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
Inf. Process. Manag.8
2026 An All-in-One Quality Assessment Agent for 4D digital human: Bridging talking heads and animated human
Yingjie Zhou 0003, Farong Wen, Li Xu 0008, Yu Zhou 0016, Jiezhang Cao, Xiaohong Liu 0001, Xiongkuo Min, Yu Wang 0002, Guangtao Zhai
Inf. Process. Manag.10
2026 Preference-guided debiasing for no-reference enhancement image quality assessment
Shiqi Gao, Zitong Xu, Huiyu Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
Image Vis. Comput.5
2026 Developing Evolving Adaptability in Biological Intelligence: A Novel Biologically-Inspired Continual Learning Model for Video Saliency Prediction
abstract
In the era of deep learning, video saliency prediction task still remains major challenge due to the issue of catastrophic forgetting during feature learning. Most prior works commonly employ generative replay strategies to generate pseudo-samples from previous tasks, enabling them to recall the data distribution. However, scaling up generative replay to accommodate class-incremental and task-incremental settings poses challenges, as generated data with low quality can severely deteriorate performance. Additionally, existing advances mainly focus on preserving memory stability to alleviate catastrophic forgetting, but they remain difficult to flexibly adapt to incremental changes in dynamic scenes. To achieve a better balance between memory stability and learning plasticity, we propose a novel biologically-inspired continual learning (BICL) model tailored to effectively predict human attention in dynamic scenes while mitigate catastrophic forgetting. In particular, inspired by the function of the hippocampus in the human neural system, we elaborately design a visual saliency memory bank module to explicitly store and retrieve representative features from previous tasks. Furthermore, drawing inspiration from the Drosophila $\gamma$γMB system, we propose an active forgetting strategy equipped with multiple parallel adaptive learner modules, which can appropriately attenuate old memories in parameter distribution to enhance learning plasticity to adapt to new tasks, and accordingly to ensure compatibility among multiple learners. Notably, without compromising the performance of old tasks, our proposed model can achieve a better trade-off between memory stability and learning plasticity. Through extensive experiments on several benchmark datasets, our model not only enhances performance in task-incremental settings, but also potentially provides deep insights into neurological adaptive mechanisms.
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 DHQA-4D: A large-scale dataset and LMM-based metric for dynamic 4D digital human quality assessment
Sijing Wu, Yucheng Zhu, Huiyu Duan, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
Pattern Recognit.7
2026 UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
abstract
As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A/V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A/V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications. Code are available at https://github.com/charlotte9524/UNQA.
Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Long Ye, Weisi Lin, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2026 Subjective and Objective Quality Assessment of Display Content Videos
abstract
Display quality assessment plays a crucial role in evaluating the performance of display devices. However, existing video quality assessment methods primarily target compression-related distortions, failing to capture display-specific degradations including definition loss, color distortions, and motion artifacts that critically affect user subjective experiences during video playback. To address these limitations, we develop a specialized video dataset, namely Video Displaying Quality Assessment Dataset (VDQA), constructed using a DSLR camera with standardized parameter optimization of exposure settings (aperture, ISO sensitivity, and shutter speed). VDQA comprises 250 high-resolution video clips covering diverse content categories, providing a robust foundation for evaluating display devices across multiple quality dimensions. Additionally, we propose a deep learning-based model specifically designed for display quality assessment that employs three complementary pathways to independently evaluate definition, color fidelity, and motion quality. The model integrates Canny edge detection for explicit sharpness measurement, a color attention mechanism to enhance sensitivity to display color reproduction characteristics, and temporal modeling for motion artifact assessment. Experimental results demonstrate that the proposed model achieves superior performance in reflecting user subjective experiences for display content videos compared to state-of-the-art methods, with significant improvements in both color fidelity assessment and definition evaluation.
Fangfang Lu, Huiqun Yu, Kaiwei Zhang, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.6
2026 Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight Model
abstract
Surveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications.
Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.10
2026 AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
abstract
The rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I outputs. However, existing image quality assessment (IQA) methods are limited to providing global quality scores, failing to deliver fine-grained perceptual evaluations for structurally complex subjects like humans, which is a critical challenge considering the frequent anatomical and textural distortions in AI-generated human images (AGHIs). To address this gap, we introduce AGHI-QA, a large-scale benchmark specifically designed for quality assessment of AGHIs. The dataset comprises 4, 000 images generated from 400 carefully crafted text prompts using 10 state-of-the-art T2I models. We conduct a systematic subjective study to collect multidimensional annotations, including perceptual quality scores, text-image correspondence scores, visible and distorted body part labels. Based on AGHI-QA, we evaluate the strengths and weaknesses of current T2I methods in generating human images from multiple dimensions. Furthermore, we propose AGHI-Assessor, a novel quality metric that integrates the large multimodal model (LMM) with domain-specific human features for precise quality prediction and identification of visible and distorted body parts in AGHIs. Extensive experimental results demonstrate that AGHI-Assessor showcases state-of-the-art performance, significantly outperforming existing IQA methods in multidimensional quality assessment and surpassing leading LMMs in detecting structural distortions in AGHIs.
Sijing Wu, Wei Sun 0029, Yucheng Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.8
2026 Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model, and Training Strategy
abstract
The rapid development of multimodal large language models has resulted in remarkable advancements in visual perception and understanding, consolidating several tasks into a single visual question-answering framework. However, these models are prone to hallucinations, which limit their reliability as artificial intelligence systems. While this issue is extensively researched in natural language processing and image captioning, there remains a lack of investigation of hallucinations in Low-level Visual Perception and Understanding (HLPU), especially in the context of image quality assessment tasks. We consider that these hallucinations arise from an absence of clear self-awareness within the models. To address this issue, we first introduce the HLPU instruction database, the first instruction database specifically focused on hallucinations in low-level vision tasks. This database contains approximately 200K question-answer pairs and comprises four subsets, each covering different types of instructions. Subsequently, we propose the Self-Awareness Failure Elimination (SAFEQA) model, which utilizes image features, salient region features and quality features to improve the perception and comprehension abilities of the model in low-level vision tasks. Furthermore, we propose the Enhancing Self-Awareness Preference Optimization (ESA-PO) framework to increase the model’s awareness of knowledge boundaries, thereby mitigating the incidence of hallucination. Finally, we conduct comprehensive experiments on low-level vision tasks, with the results demonstrating that our proposed method significantly enhances self-awareness of the model in these tasks and reduces hallucinations. Notably, our proposed method improves both accuracy and self-awareness of the proposed model and outperforms close-source models in terms of various evaluation metrics. This research contributes to the advancement of self-awareness capabilities in multimodal large language models, particularly for low-level visual perception and understanding tasks.
Xiongkuo Min, Yuqin Cao, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2026 Quality Assessment and Distortion-Aware Saliency Prediction for AI-Generated Omnidirectional Images
abstract
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflecthumanfeedback for AI-generatedomnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research.
Huiyu Duan, Jing Liu 0002, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.6
2026 LMVQ: Label-Free Metric-Learning for General AI-Generated Video Quality Assessment
abstract
The recent rapid development of video generation technology has led to a significant demand for quality assessment of the latest AI-generated videos. However, current supervised approaches depend on expensive and quickly outdated human scores, and label-free methods overlook the general distortions of AI-generated videos. To address these limitations, we introduce LMVQ, a Label-free Metric-learning framework for general AI-generated Video Quality assessment of three dimensions, spatial, temporal, and alignment. The LMVQ is the first to introduce sample degradations specially designed for AIGC-specific distortions, and constructs a comprehensive training set through two complementary sample generation strategies. It then employs two synergistic modules, the Intra-Quality Token Transformer (IQ-Trans), which explicitly refines dimension-specific quality representations, and the Inter-Quality Mixture of Experts (IQ-MoE), which fuses interactions across multiple quality dimensions. Finally, a Multi-Proxy Metric-Learning (MPML) strategy aligns the learned representations with multi-dimensional quality scores and constrains the model to learn discriminative quality-aware representations. Extensive experiments on four public AIGC-VQA benchmarks show that MPML outperforms previous label-free methods by over 20%, and greatly narrows the gap with supervised methods. This provides a scalable, adaptive foundation for evaluating the ever-evolving quality of AI-generated videos.
Xinyue Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.6
2026 Future Fixation Sequence Prediction for Audio-Visual 360° Videos
abstract
Future fixation sequence prediction plays a crucial role in various aspects of virtual reality content production, transmission, rendering, and display. Accurate prediction of future fixation sequence can significantly enhance the quality of user experience, particularly in resource-constrained scenarios. In this paper, we present a novel framework for predicting future fixation sequence and achieves state-of-the-art performance. Specifically, the anti-projection-distortion FoV patch extraction algorithm is proposed to mitigate projection distortions. A comprehensive contextual representation is then constructed by integrating multiple data sources, including visual and audio information, historical fixation sequence, user identity, timestamp, and positional embeddings. The transformer-based predictor is proposed to perform the future fixation sequence prediction based on the integrated contextual representations. Additionally, we propose a framework that effectively utilizes saliency information as supervision and conduct saliency contrastive distillation during the training phase, eliminating the need for saliency data during inference. Overall, by integrating anti-projection-distortion and multimodal representations, along with key embeddings, a dedicated predictor, and contrastive distillation, our approach is designed to accurately predict future fixation sequences. Extensive experiments validate the effectiveness of our framework, demonstrating its superior performance in fixation prediction tasks.
Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Huiyu Duan, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Multi-Dimensional Quality Assessment for Single-Image-to-3D Contents: Dataset and Model
abstract
The rapid advancement of AI generation technologies has led to the widespread use of AI-generated multimedia content, including images, videos, and 3D contents, across various applications. While significant progress has been made in quality evaluation for 2D content, evaluating the quality of 3D content synthesized from single image remains an underexplored problem. To bridge this gap, we introduce the first comprehensive subjective evaluation database tailored for assessing the quality of 3D content generated from single image. Our database, named AIGC-SI23DCQA, includes three distinct categories of input images, i.e., realistic images, AI-generated images, and computer graphic (CG) images, with 100 images in each category. Using five representative single-image-to-3D algorithms, we produce 1,500 3D contents and collect 94,500 annotations across three quality dimensions, including texture fidelity, shape accuracy, and overall quality. Based on the constructed database, we first benchmark and evaluate the performance of existing quality assessment methods revealing their limitations in addressing this novel task. Thus, we further propose a novel objective quality assessment method, termed I3DQA, for effective single-image-to-3D content quality assessment. Specifically, I3DQA first extracts the reference features from the source image, and the multi-modal features from the generated 3D content, including the projected video, patches, and large-multimodal model (LMM) features. These features are integrated through symmetric transformer blocks, enabling effective quality-related feature fusion and score prediction. Extensive experiments demonstrate the superior performance of our method and validate the effectiveness of its components. This work provides a foundational resource and a robust framework for advancing research in this emerging field, and our database and model are released at https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Jing Liu 0002, Yun Liu 0009, Xiaohong Liu 0001, Jia Wang 0004, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai
IEEE Trans. Image Process.8
2026 MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional Images
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.6
2025 Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics
abstract
Textured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies focus on non-textured meshes, our work explores the complexities introduced by detailed texture patterns. We present a new dataset for textured mesh saliency, created through an innovative eye-tracking experiment in a six degrees of freedom (6-DOF) VR environment. This dataset addresses the limitations of previous studies by providing comprehensive eye-tracking data from multiple viewpoints, thereby advancing our understanding of human visual behavior and supporting more accurate and effective 3D content creation. Our proposed model predicts saliency maps for textured mesh surfaces by treating each triangular face as an individual unit and assigning a saliency density value to reflect the importance of each local surface region. The model incorporates a texture alignment module and a geometric extraction module, combined with an aggregation module to integrate texture and geometry for precise saliency prediction. We believe this approach will enhance the visual fidelity of geometric processing while ensuring computational efficiency, essential for real-time rendering and high-detail applications such as VR and gaming.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
AAAI3
2025 Redundancy Principles for MLLMs Benchmarks
abstract
Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinyu Fang, Chunyi Li 0001, Xiaohong Liu 0001, Xiongkuo Min, Haodong Duan, Kai Chen 0026, Guangtao Zhai
ACL (1)6
2025 FineVQ: Fine-Grained User Generated Content Video Quality Assessment
abstract
The rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA models generally only give an overall rating for a UGC video, which lacks fine-grained labels for serving video processing and recommendation applications. To address the challenges and promote the development of UGC videos, we establish the first large-scale Fine-grained Video quality assessment Database, termed FineVD, which comprises 6104 UGC videos with fine-grained quality scores and descriptions across multiple dimensions. Based on this database, we propose a Fine-grained Video Quality assessment (FineVQ) model to learn the fine-grained quality of UGC videos, with the capabilities of quality rating, quality scoring, and quality attribution. Extensive experimental results demonstrate that our proposed FineVQ can produce fine-grained video-quality results and achieve state-of-the-art performance on FineVD and other commonly used UGC-VQA datasets. Both FineVD and FineVQ are publicly available at: https://github.com/IntMeGroup/FineVQ.
Huiyu Duan, Qiang Hu 0003, Zitong Xu, Lu Liu 0005, Xiongkuo Min, Chunlei Cai, Tianxiao Ye, Xiaoyun Zhang 0001, Guangtao Zhai
CVPR7
2025 Image Quality Assessment: From Human to Machine Preference
abstract
Image Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference depends on downstream tasks such as segmentation and detection, rather than visual appeal. Considering the huge gap between human and machine visual systems, this paper proposes the topic: Image Quality Assessment for Machine Vision for the first time. Specifically, we (1) defined the subjective preferences of machines, including downstream tasks, test models, and evaluation metrics; (2) established the Machine Preference Database (MPD), which contains 2.25M fine-grained annotations and 30k reference/distorted image pair instances; (3) verified the performance of mainstream IQA algorithms on MPD. Experiments show that current IQA metrics are human-centric and cannot accurately characterize machine preferences. We sincerely hope that MPD can promote the evolution of IQA from human to machine preferences. Project page is on: https://github.com/lcysyzxdxc/MPD.
Chunyi Li 0001, Yuan Tian 0017, Xiaoyue Ling, Haodong Duan, Haoning Wu 0001, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Guo Lu, Weisi Lin, Guangtao Zhai
CVPR9
2025 AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
Huiyu Duan, Guangtao Zhai, Juntong Wang, Xiongkuo Min
CVPR5
2025 Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshes
abstract
Mesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the first to systematically capture the differences in saliency distribution under both textured and non-textured visual conditions. Furthermore, we introduce mesh Mamba, a unified saliency prediction model based on a state space model (SSM), designed to adapt across various mesh types. Mesh Mamba effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. More importantly, by sub-graph embedding and a bidirectional SSM, the model enables global context modeling for both local geometry and texture, preserving the topological structure and improving the understanding of visual details and structural complexity. Through extensive theoretical and empirical validation, our model not only improves performance across various mesh types but also demonstrates high scalability and versatility, particularly through cross validations of various visual features.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
CVPR3
2025 Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs
abstract
With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AI-generated content (AIGC), and computer graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of perceptual video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to the performance of human beings. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding.
Ziheng Jia, Haoning Wu 0001, Chunyi Li 0001, Zijian Chen 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai
CVPR9
2025 Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
abstract
Evaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval.
Tengchuan Kou, Shushi Wang, Chunyi Li 0001, Wei Sun 0029, Wei Wang 0213, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai
CVPR10
2025 Explore the Hallucination on Low-level Perception for MLLMs
abstract
The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models also exhibit hallucinations, which limit their reliability as AI systems, especially in tasks involving low-level visual perception and understanding. We believe that hallucinations stem from a lack of explicit self-awareness in these models, which directly impacts their overall performance. In this paper, we aim to define and evaluate the self-awareness of MLLMs in low-level visual perception and understanding tasks. To this end, we present QL-Bench, a benchmark settings to simulate human responses to low-level vision, investigating self-awareness in low-level visual perception through visual question answering related to low-level attributes such as clarity and lighting. Specifically, we construct the LLSAVisionQA dataset, comprising 2,990 single images and 1,999 image pairs, each accompanied by an open-ended question about its low-level features. Through the evaluation of 15 MLLMs, we demonstrate that while some models exhibit robust low-level visual capabilities, their self-awareness remains relatively underdeveloped. Notably, for the same model, simpler questions are often answered more accurately than complex ones. However, self-awareness appears to improve when addressing more challenging questions. We hope that our benchmark will motivate further research, particularly focused on enhancing the self-awareness of MLLMs in tasks involving low-level visual perception and understanding.
Haoning Wu 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai, Xiongkuo Min
ICASSP7
2025 HazeCLIP: Towards Language Guided Real-World Image Dehazing
abstract
Existing methods have achieved remarkable performance in image dehazing, particularly on synthetic datasets. However, they often struggle with real-world hazy images due to domain shift, limiting their practical applicability. This paper introduces HazeCLIP, a language-guided adaptation framework designed to enhance the real-world performance of pre-trained dehazing networks. Inspired by the Contrastive Language-Image Pre-training (CLIP) model’s ability to distinguish between hazy and clean images, we leverage it to evaluate dehazing results. Combined with a region-specific dehazing technique and tailored prompt sets, the CLIP model accurately identifies hazy areas, providing a high-quality, human-like prior that guides the fine-tuning process of pre-trained networks. Extensive experiments demonstrate that HazeCLIP achieves state-of-the-art performance in real-word image dehazing, evaluated through both visual quality and image quality assessment metrics. Codes are available at https://github.com/Troivyn/HazeCLIP.
Ruiyi Wang, Wenhao Li 0018, Xiaohong Liu 0001, Chunyi Li 0001, Xiongkuo Min, Guangtao Zhai
ICASSP6
2025 3DGCQA: A Quality Assessment Database for 3D AI-Generated Contents
abstract
Although 3D generated content (3DGC) offers advantages in reducing production costs and accelerating design timelines, its quality often falls short when compared to 3D professionally generated content. Common quality issues frequently affect 3DGC, highlighting the importance of timely and effective quality assessment. Such evaluations not only ensure a higher standard of 3DGCs for end-users but also provide critical insights for advancing generative technologies. To address existing gaps in this domain, this paper introduces a novel 3DGC quality assessment dataset, 3DGCQA, built using 7 representative Text-to-3D generation methods. During the dataset’s construction, 50 fixed prompts are utilized to generate contents across all methods, resulting in the creation of 313 textured meshes that constitute the 3DGCQA dataset. The visualization intuitively reveals the presence of 6 common distortion categories in the generated 3DGCs. To further explore the quality of the 3DGCs, subjective quality assessment is conducted by evaluators, whose ratings reveal significant variation in quality across different generation methods. Additionally, several objective quality assessment algorithms are tested on the 3DGCQA dataset. The results expose limitations in the performance of existing algorithms and underscore the need for developing more specialized quality assessment methods. To provide a valuable resource for future research and development in 3D content generation and quality assessment, the dataset has been open-sourced in https://github.com/zyj-2000/3DGCQA.
Yingjie Zhou 0003, Farong Wen, Jun Jia, Yanwei Jiang, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ICASSP7
2025 Information Density Principle for MLLM Benchmarks
abstract
With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and whether the test results meet their requirements. Therefore, we propose a critical principle of Information Density, which examines how much insight a benchmark can provide for the development of MLLMs. We characterize it from four key dimensions: (1) Fallacy, (2) Difficulty, (3) Redundancy, (4) Diversity. Through a comprehensive analysis of more than 10,000 samples, we measured the information density of 19 MLLM benchmarks. Experiments show that using the latest benchmarks in testing can provide more insight compared to previous ones, but there is still room for improvement in their information density. We hope this principle can promote the development and application of future MLLM benchmarks. Project page: https://github.com/lcysyzxdxc/bench4bench
Chunyi Li 0001, Xiaozhe Li, Yuan Tian 0017, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Haodong Duan, Kai Chen 0026, Guangtao Zhai
ICCV7
2025 FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching
Hongjiu Yu, Ying Chen 0011, Kai Li 0012, Xiongkuo Min, Huiyu Duan, Guangtao Zhai, Xu Liu 0006
ICCV7
2025 LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMs
abstract
Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual quality and text-image alignment. Given the high cost and inefficiency of manual evaluation, an automatic metric that aligns with human preferences is desirable. To this end, we present EvalMi-50K, a comprehensive dataset and benchmark for evaluating large-multimodal image generation, which features (i) comprehensive tasks, encompassing 2,100 extensive prompts across 20 fine-grained task dimensions, and (ii) large-scale human-preference annotations, including 100K mean-opinion scores (MOSs) and 50K question-answering (QA) pairs annotated on 50,400 images generated from 24 T2I models. Based on EvalMi-50K, we propose LMM4LMM, an LMM-based metric for evaluating large multimodal T2I generation from multiple dimensions including perception, text-image correspondence, and task-specific accuracy. Extensive experimental results show that LMM4LMM achieves state-of-the-art performance on EvalMi-50K, and exhibits strong generalization ability on other AI-generated image evaluation benchmark datasets, manifesting the generality of both the EvalMi-50K dataset and LMM4LMM metric. Both EvalMi-50K and LMM4LMM will be released at https://github.com/IntMeGroup/LMM4LMM.
Huiyu Duan, Juntong Wang, Guangtao Zhai, Xiongkuo Min
ICCV6
2025 Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
abstract
Speech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging digital human media. However, challenges persist regarding the quality of these talkers and AGTHs they generate, and comprehensive studies addressing these issues remain limited. To address this gap, this paper presents the largest AGTH quality assessment dataset THQA-10K to date, which selects 12 prominent T2I models and 14 advanced talkers to generate AGTHs for 14 prompts. After excluding instances where AGTH generation is unsuccessful, the THQA-10K dataset contains 10,457 AGTHs. Then, volunteers are recruited to subjectively rate the AGTHs and give the corresponding distortion categories. In our analysis for subjective experimental results, we evaluate the performance of talkers in terms of generalizability and quality, and also expose the distortions of existing AGTHs. Finally, an objective quality assessment method based on the first frame, Y-T slice and tone-lip consistency is proposed. Experimental results show that this method can achieve state-of-the-art (SOTA) performance in AGTH quality assessment. The work is released at https://github.com/zyj-2000/Talker.
Yingjie Zhou 0003, Jiezhang Cao, Farong Wen, Yanwei Jiang, Jun Jia, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ICCV8
2025 CDHQA: A Quality Assessment Database for Conversational Digital Human
Yingjie Zhou 0003, Yinghan Xia, Zhixiang Lu, Farong Wen, Yu Wang 0002, Yu Zhou 0016, Xiaohong Liu 0001, Xiongkuo Min, Jiezhang Cao, Guangtao Zhai
ICIG (3)11
2025 Exploring The Potential of Vision-Language Models for Pure-Image and Text-Guided-Image Saliency Prediction
abstract
We introduce VLSal, a saliency prediction framework that leverages Vision-Language Models (VLMs) to unify pure-image and text-guided-image saliency prediction tasks and achieve high performance in both. We extract visual features from the visual encoder and retrieve the corresponding visual token features from the language decoder, which serves as a natural feature fusion mechanism. These features are then processed through a U-Net-based saliency decoder to generate accurate saliency maps. To efficiently adapt the large-scale pretrained model, we apply Low-Rank Adaptation (LoRA) finetuning, reducing computational costs while preserving performance. Extensive experiments on benchmark datasets, including SALICON, MIT1003, and TIS, demonstrate that VLSal outperforms existing methods in both pure-image and text-guided-image saliency prediction.
Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICIP3
2025 A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
abstract
How to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, the precision and validity of which are still questionable. Furthermore, traditional benchmarks often utilize mostly natural-captured content rather than AIGIs to test the abilities of LMMs, leading to a noticeable gap for AIGIs. Therefore, we introduce **A-Bench** in this paper, a benchmark designed to diagnose *whether LMMs are masters at evaluating AIGIs*. Specifically, **A-Bench** is organized under two key principles: 1) Emphasizing both high-level semantic understanding and low-level visual quality perception to address the intricate demands of AIGIs. 2) Various generative models are utilized for AIGI creation, and various LMMs are employed for evaluation, which ensures a comprehensive validation scope. Ultimately, 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts. We hope that **A-Bench** will significantly enhance the evaluation process and promote the generation quality for AIGIs.
Haoning Wu 0001, Chunyi Li 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Zijian Chen 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai
ICLR6
2025 A Unified Inverse-Tone-Mapped HDR Video Quality Assessment Method across Two HDR Formats
abstract
High Dynamic Range Video Quality Assessment (HDR VQA) plays a pivotal role in Inverse Tone Mapping (ITM) research. Existing HDR VQA datasets and models mainly focus on a single HDR format, leading to poor generalization and limited application scope. To address this problem, this paper proposes a reference-free CONTrastive ITM-VQA (CONT-ITM-VQA) model via format transformation-based data augmentation and contrastive learning. Specifically, the format transformation-based data augmentation improves the model generalization, via applying Opto-electronic Transfer Function (OETF) transformations between the HDR formats; while the contrastive learning-based quality-related feature alignment aligns the quality features from different HDR formats of the same video to obtain more effective quality representations. It is worth noting that our method can be extended to other HDR-related quality assessment, not limited to ITM-HDR VQA. Experimental results demonstrate that our model closely mimics subjective judgments.
Leidong Fan, Xiongkuo Min, Qing Li 0029, Anjie Wang
ICME2
2025 SI23DCQA: Perceptual Quality Assessment of Single Image-to-3D Content
abstract
In recent years, significant efforts have been dedicated to advancing 3D content generation. However, existing quality assessment research predominantly focuses on evaluating Text-to-3D Content (T23DC) while ignoring Single Image-to-3D Content (SI23DC). In this paper, we establish the first Single Image-to-3D Content Quality Assessment (SI23DCQA) database to comprehensively study the perceptual quality of SI23DCs. The database contains 1500 SI23DCs, which are generated by 5 common SI23DC algorithms from 300 images including realistic images, AI generated images, and model rendered images. Afterward, we carry out a well-designed subjective experiment to collect subjective quality ratings for SI23DCs from three perspectives including overall, color, and shape. Additionally, a benchmark experiment is conducted with the state-of-the-art no reference image quality assessment (NR-IQA), no reference video quality assessment (NR-VQA), and no reference 3D quality assessment (NR-3DQA) and the experimental results show that current quality assessment methods are limited in evaluating the perceptual loss of SI23DCs. The database is released on https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
ICME5
2025 VIP-PCQA: A Multi-Modal Framework for No-reference Point Cloud Quality Assessment
abstract
Point clouds often suffer from geometric and color noise, as well as compression artifacts, during their production, storage, and transmission. Therefore, accurately and automatically evaluating the quality of point clouds is crucial for optimizing storage and compression strategies. This paper introduces the VIP-PCQA, a novel framework that combines Video, Image, and Point cloud modalities for no-reference Point Cloud Quality Assessment. The framework begins by rendering projection videos and normal images from point clouds, followed by sampling patches and computing statistical features related to color and geometry. Subsequently, a video encoder, two image encoders, and a point cloud encoder are employed to extract modality-specific features. Finally, these features are fused to regress the quality score. Experimental results on three publicly available benchmark databases demonstrate that VIP-PCQA achieves outstanding performance with excellent generalization capabilities. An ablation study further highlights the indispensable contribution of each modality to the framework’s success. The code is released on https://github.com/ZedFu/VIP-PCQA.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ICME5
2025 Perceiving Smoothness: Temporal Consistency Learning for Multi-Frame-Rate Video Quality Assessment
abstract
The video industry is continuously evolving, with videos featuring a wide range of frame rates. Video quality assessment (VQA) aims to automatically monitor and select the optimal frame rate for video communication systems. Existing VQA research has achieved impressive results in high frame rate (HFR) VQA tasks, but often lacks specific designs to address various frame rate distortions, such as smoothness distortions caused by frame rate variations, artifacts from the coupling of frame rate and compression, and confusion between low frame rate and slow motion. To address these challenges, we propose a VQA framework to perceive smoothness (PSVQA), which includes a novel frame-rate-driven feature processing module and a new feature fusion strategy. The module aggregates smoothness features from multi-scale temporal embeddings and incorporates frame rate guidance to resolve the discrepancies between temporal features and real perceptual experience. Furthermore, we combine spatial video features with temporal consistency features for quality modeling, optimizing the feature fusion module to enhance multi-frame-rate perception. Through extensive experiments on HFR and variable-frame-rate datasets, we validate the effectiveness of PSVQA.
Jinliang Han, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai
ICME2
2025 A Novel Framework for Realistic 3D Scene Regeneration with Graph of Thoughts
abstract
In embodied intelligence applications, highly realistic 3D scenes lay the foundation for perception and decision-making, while 3D scene regeneration creates more coherent and personalized virtual spaces, facilitating more efficient task adaptation and agent training. To address this, we propose a reasoning framework based on the Graph of Thoughts (GoT), which enhances the prompting capabilities of large language models (LLM) and integrates a synergistic mechanism of retrospective memory and feedback loops into the regeneration process. During the initial generation phase, we retain the Holodeck paradigm, combining LLM-driven scene design inferences with the spatial layout of 3D assets from Objaverse. In the regeneration phase, dynamic feedback loops trigger backtracking of reasoning memory to adjust relevant elements according to evolving requirements, while maintaining stability and consistency in unrelated elements, ensuring the scene’s overall coherence. We conduct both subjective and objective experiments to validate the effectiveness of this framework, demonstrating significant improvements in 3D scene generation.
Yitian Kou, Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
ICME4
2025 HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment
abstract
Image composition involves extracting a foreground object from one image and pasting it into another image through Image harmonization algorithms (IHAs), which aim to adjust the appearance of the foreground object to better match the background. Existing image quality assessment (IQA) methods may fail to align with human visual preference on image harmonization due to the insensitivity to minor color or light inconsistency. To address the issue and facilitate the advancement of IHAs, we introduce the first Image Quality Assessment Database for image Harmony evaluation (HarmonyIQAD), which consists of 1,350 harmonized images generated by 9 different IHAs, and the corresponding human visual preference scores. Based on this database, we propose a Harmony Image Quality Assessment (HarmonyIQA), to predict human visual preference for harmonized images. Extensive experiments show that HarmonyIQA achieves state-of-the-art performance on human visual preference evaluation for harmonized images, and also achieves competing results on traditional IQA tasks. Furthermore, cross-dataset evaluation also shows that HarmonyIQA exhibits better generalization ability than self-supervised learning-based IQA methods. The dataset and code are available at https://github.com/IntMeGroup/HarmonyIQA.
Zitong Xu, Huiyu Duan, Guangji Ma, Qingbo Wu 0001, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICME7
2025 LD2Scan: A Lightweight Dual-Temporal Constrained Scanpath Prediction Model for Omnidirectional Images
abstract
Predicting scanpaths in omnidirectional images (ODIs) is essential for simulating human gaze behaviors. However, current methods often struggle with long-term dependencies and exhibit high complexity, which limits their efficiency and scalability. To tackle these challenges, we propose LD2Scan, a lightweight diffusion-based model specifically designed for scanpath prediction in ODIs. It employs Efficient Equivariant (E4) convolution to enhance feature extraction from distorted ODIs while improving computational performance, thereby reducing resource demands. LD2Scan utilizes a dual-graph convolutional network (GCN) to enforce internal time constraints between fixations, integrating semantic-level GCN for sequential fixation modeling and image-level GCN to capture relationships across different images, enriching contextual information. We formulate the scanpath prediction issue as a conditional generation task, refining noisy scanpaths using features encoded by the dual-GCN and robust E4-processed features. Experimental results on several benchmark datasets demonstrate that LD2Scan outperforms existing methods in terms of both accuracy and efficiency.
Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
ICME5
2025 IllusionBench: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models
abstract
Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by state-of-the-art (SOTA) VLMs, revealing issues such as hallucinations and limited perceptual abilities. To address this gap, we introduce IllusionBench, a comprehensive visual illusion dataset that encompasses not only classic cognitive illusions but also real-world scene illusions. This dataset features 1,051 images, 5,548 question-answer pairs, and 1,051 golden text descriptions that address the presence, causes, and content of the illusions. We evaluate ten SOTA VLMs on this dataset using true-or-false, multiple-choice, and open-ended tasks. In addition to real-world illusions, we design trap illusions that resemble classical patterns but differ in reality, highlighting hallucination issues in SOTA models. The top-performing model, GPT-4o, achieves 80.59% accuracy on true-or-false tasks and 76.75% on multiple-choice questions, but still lags behind human performance. In the semantic description task, GPT-4o’s hallucinations on classical illusions result in low scores for trap illusions, even falling behind some open-source models. IllusionBench is, to the best of our knowledge, the largest and most comprehensive benchmark for visual illusions in VLMs to date.
Xinyi Wei, Xiaohong Liu 0001, Guangtao Zhai, Xiongkuo Min
ICME6
2025 CAP: An Advanced No-Reference Quality Assessment Method for AI-Generated 3D Meshes
abstract
The advent of generative AI has revolutionized 3D content design, significantly enhancing modelers’ efficiency. However, the quality of generated 3D content, particularly Generated Meshes (GMs), remains a critical concern. GMs pose unique challenges for quality assessment due to their complex geometry, detailed texture mapping, and distortions that differ from traditional meshes. Existing methods fail to address these GM-specific issues. To tackle this gap, we introduce a novel no-reference quality assessment method, CAP, which integrates CT-Slice, prompt Alignment, and Projections. CAP employs a six-face projection to capture external features and a CT-like slicing approach to extract internal quality features. Additionally, it leverages Contrastive Language-Image Pre-Training (CLIP) to measure the alignment between projection embeddings and prompts as a key quality indicator. Experimental results demonstrate that CAP effectively evaluates GM quality by combining internal, external, and alignment features. The code for this work has been open-sourced in https://github.com/zyj-2000/CAP.
Yingjie Zhou 0003, Farong Wen, Yanwei Jiang, Jun Jia, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ICME7
2025 ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
abstract
With the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the quality of experience (QoE) of egocentric spatial videos is crucial to ensure a high-quality viewing experience. However, the corresponding research is still lacking. In this paper, we use the concept of embodied experience to highlight this more immersive experience and study the new problem, i.e., embodied perceptual quality assessment for egocentric spatial videos. Specifically, we introduce the first Egocentric Spatial Video Quality Assessment Database (ESVQAD), which comprises 600 egocentric spatial videos captured using the Apple Vision Pro and their corresponding mean opinion scores (MOSs). Furthermore, we propose a novel multi-dimensional binocular feature fusion model, termed ESVQAnet, which integrates binocular spatial, motion, and semantic features to predict the overall perceptual quality. Experimental results demonstrate the ESVQAnet significantly outperforms 16 state-of-the-art VQA models on the embodied perceptual quality assessment task, and exhibits strong generalization capability on traditional VQA tasks. The database and code are available at https://github.com/IntMeGroup/ESVQA.
Xilei Zhu, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICME5
2025 AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
abstract
Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing audio-visual quality assessment methods struggle with unique distortions in AGAVs, such as unrealistic and inconsistent elements. To address this, we introduce AGAVQA-3k, the first large-scale AGAV quality assessment dataset, comprising $3,382$ AGAVs from $16$ VTA methods. AGAVQA-3k includes two subsets: AGAVQA-MOS, which provides multi-dimensional scores for audio quality, content consistency, and overall quality, and AGAVQA-Pair, designed for optimal AGAV pair selection. We further propose AGAV-Rater, a LMM-based model that can score AGAVs, as well as audio and music generated from text, across multiple dimensions, and selects the best AGAV generated by VTA methods to present to the user. AGAV-Rater achieves state-of-the-art performance on AGAVQA-3k, Text-to-Audio, and Text-to-Music datasets. Subjective tests also confirm that AGAV-Rater enhances VTA performance and user experience. The dataset and code is available at https://github.com/charlotte9524/AGAV-Rater.
Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai
ICML2
2025 A Dataset and Method for Assessing the Quality of Display Devices
abstract
This paper proposes a novel objective quality assessment method for display devices, aiming to comprehensively evaluate their overall quality, with a particular focus on the user’s subjective experience. By constructing a dedicated video dataset and integrating a no-reference video quality assessment (NR VQA) model, we directly learn spatiotemporal features related to quality from video frames and introduce a color attention mechanism to enhance the model’s sensitivity to color distortions. The model consists of feature extraction, quality regression, and quality pooling modules, capable of extracting quality-aware features from spatial, spatiotemporal, and color domains, and determining the final quality score using a temporal averaging pooling strategy. Experimental results demonstrate that the proposed model performs excellently in evaluating the overall quality of display devices, showing high practical applicability. The main contributions of this paper include constructing a video dataset, proposing a new model, and introducing a color attention mechanism, which effectively enhances the comprehensive evaluation of the subjective quality of display devices, filling the gap in display device quality assessment.
Haoyang Ni, Kaiwei Zhang, Fangfang Lu, Xiongkuo Min, Guangtao Zhai
ISCAS5
2025 Machine Vision Quality Assessment for Image Restoration
abstract
In recent years, substantial progress has been made in the realm of No-Reference Image Quality Assessment (NR-IQA) for image restoration, where performance has been predominantly evaluated using metrics such as BRISQUE [1] and Hyper-IQA [2]. However, these NR-IQA metrics assess the perceptual quality of images without considering their utility in specific machine vision tasks, such as object detection and semantic segmentation. In this paper, we propose a Machine Vision Quality Assessment (MVQA) framework for image restoration. Specifically, we introduce the weighted Alternative Free-response Operating Characteristic (wAFROC) [3] as a metric to assess the machine vision quality of three image restoration sub-tasks: dehazing, denoising, and superresolution. By accounting for both detection sensitivity and spatial localization, wAFROC provides a more comprehensive evaluation of image quality in machine vision contexts. Its effectiveness is validated through downstream tasks, specifically object detection and semantic segmentation. We construct an IQA dataset for image restoration to explore the impact of various image restoration algorithms on the accuracy of object detection and semantic segmentation algorithms. Extensive experimental results demonstrate that the MVQA framework, leveraging wAFROC, effectively predicts the influence of image quality on machine vision tasks, bridging the gap between perceptual IQA and task-specific quality requirements in machine vision applications.
Yiming Shi, Xiongkuo Min, Guangtao Zhai
ISCAS2
2025 Visual Saliency Prediction for Augmented Reality Videos
abstract
Augmented Reality (AR) is an emerging technology that allows users to perceive both virtual-world contents and real-world scenes simultaneously. It has numerous applications in industrial manufacturing, entertainment, gaming, education, etc. In AR environments, the visual confusion phenomenon caused by the overlay of augmented content and real backgrounds is evident, yet the understanding and research of visual saliency under the AR visual confusion condition remains limited. This paper primarily analyzes the interaction between real-world scenes and AR content, explores human visual saliency when using AR devices. First, we conduct a large-scale eye-tracking experiment based on a head-mounted AR device, and construct an AR saliency dataset containing 2160 videos, with corresponding collected eye movement data. Through qualitative analysis of the visual attention heat maps, we conclude that visual confusion significantly influences visual attention in AR video. Additionally, we quantitatively evaluate the performance of a series of classical saliency models and deep neural network saliency models on the dataset constructed in this project. For better predicting saliency in AR, we propose a general saliency prediction model, InternSal, which achieves state-of-the-art performance compared to other methods. The database and codes will be released to facilitate future research.
Zongyi Xie, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
ISCAS5
2025 LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and Benchmark
abstract
In recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly.
Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai
ISCAS6
2025 EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
abstract
The furnishing of multi-modal large language models (MLLMs) has led to the emergence of numerous benchmark studies, particularly those evaluating their perception and understanding capabilities. Among these, understanding image-evoked emotions aims to enhance MLLMs' empathy, with significant applications such as human-machine interaction and advertising recommendations. However, current evaluations of this MLLM capability remain coarse-grained, and a systematic and comprehensive assessment is still lacking. To this end, we introduce EEmo-Bench, a novel benchmark dedicated to the analysis of the evoked emotions in images across diverse content categories. Our core contributions include: 1) Regarding the diversity of the evoked emotions, we adopt an emotion ranking strategy and employ the Valence-Arousal-Dominance (VAD) as emotional attributes for emotional assessment. In line with this methodology, 1,960 images are collected and manually annotated. 2) We design four tasks to evaluate MLLMs' ability to capture the evoked emotions by single images and their associated attributes: Perception, Ranking, Description, and Assessment. Additionally, image-pairwise analysis is introduced to investigate the model's proficiency in performing joint and comparative analysis. In total, we collect 6,773 question-answer pairs and perform a thorough assessment on 19 commonly-used MLLMs. The results indicate that while some proprietary and large-scale open-source MLLMs achieve promising overall performance, the analytical capabilities in certain evaluation dimensions remain suboptimal. Our EEmo-Bench paves the path for further research aimed at enhancing the comprehensive perceiving and understanding capabilities of MLLMs concerning image-evoked emotions, which is crucial for machine-centric emotion perception and understanding. Our code and benchmark datasets are available at https://github.com/workerred/EEmo-Bench.
Lancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun 0029, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min
ACM Multimedia8
2025 Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and Method
abstract
With the rise of Text-to-Image (T2I) models, generating face images from text prompts has emerged as a prominent research area. However, evaluating the quality of these generated face images, particularly with respect to fine-grained facial attributes, remains a significant challenge. To address this, we introduce the Fine -grained Text-to-Face Image Quality Assessment (FineTFIQA) database, which is designed to evaluate the ability of T2I models to generate fine-grained face images. To the best of our knowledge, this database is the largest of its kind, containing 7,218 face images generated from 1,000 text prompts that cover 111 distinct facial attributes. A large group of subjects was invited to assess the quality of text-to-face images on four evaluation dimensions: perceptual quality, human likeness, attractiveness, and consistency. Additionally, we develop the Multi-Dimensional Text-to-Face Image Quality Assessment (MDTFIQA) method based on the Large Language Model (LLM), which combines both face image features and text features to evaluate generated images on all evaluation dimensions. Extensive experimental results demonstrate that traditional face image assessment methods and general image quality assessment methods are inadequate for accurately evaluating generated text-to-face images. Our method significantly outperforms these existing methods on all evaluation dimensions, proving to be an effective method for assessing the quality of generated text-to-face images.
Xiongkuo Min, Jinliang Han, Yuqin Cao, Sijing Wu, Yunze Dou, Guangtao Zhai
ACM Multimedia2
2025 VQA2: Visual Question Answering for Video Quality Assessment
abstract
The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field in low-level visual perception, focused initially on quantitative video quality scoring. However, driven by advances in LMMs, it is now progressing toward more holistic visual quality understanding tasks. Recent studies in the image domain have demonstrated that Visual Question Answering (VQA) can markedly enhance low-level visual quality evaluation. Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA² Instruction Dataset-the first visual question answering instruction dataset that focuses on video quality assessment. This dataset consists of 3 subsets and covers various video types, containing 157,755 instruction question-answer pairs. Then, leveraging this foundation, we present the VQA² series models. The VQA² series models interleave visual and motion tokens to enhance the perception of spatial-temporal quality details in videos. We conduct extensive experiments on video quality scoring and understanding tasks, and results demonstrate that the VQA² series models achieve excellent performance in both tasks. Notably, our final model, the VQA²-Assistant, exceeds the renowned GPT-4o in visual quality understanding tasks while maintaining strong competitiveness in quality scoring tasks. Our work provides a foundation and feasible approach for integrating low-level video quality assessment and understanding with LMMs.
Ziheng Jia, Jiaying Qian, Haoning Wu 0001, Wei Sun 0029, Chunyi Li 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai, Xiongkuo Min
ACM Multimedia10
2025 RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment
abstract
As camera-equipped robotic platforms become increasingly integrated into daily life, robotic-generated videos have begun to appear on streaming media platforms, enabling us to envision a future where humans and robots coexist. We innovatively propose the concept of Robotic-Generated Content (RGC) to term these videos generated from egocentric perspective of robots. The perceptual quality of RGC videos is critical in human-robot interaction scenarios, and RGC videos exhibit unique distortions and visual requirements that differ markedly from those of professionally-generated content (PGC) videos and user-generated content (UGC) videos. However, dedicated research on quality assessment of RGC videos is still lacking. To address this gap and to support broader robotic applications, we establish the first Robotic-Generated Content Database (RGCD), which contains a total of 2,100 videos drawn from three robot categories and sourced from diverse platforms. A subjective VQA experiment is conducted subsequently to assess human visual perception of robotic-generated videos. Finally, we conduct a benchmark experiment to evaluate the performance of 11 state-of-the-art VQA models on our database. Experimental results reveal significant limitations in existing VQA models when applied to complex, robotic-generated content, highlighting a critical need for RGC-specific VQA models. Our RGCD is publicly available at: https://github.com/IntMeGroup/RGC-VQA.
Jianing Jin, Jiangyong Ying, Huiyu Duan, Sijing Wu, Yushuo Zheng, Xiongkuo Min, Guangtao Zhai
ACM Multimedia8
2025 Towards a New Paradigm of Visual Signal Compression
abstract
Ultra-low bitrate image compression is a challenging and demand- ing topic. With the development of Large Multimodal Models (LMMs), a Cross Modality Compression (CMC) paradigm of Image-Text- Image has emerged. Compared with traditional codecs, this semantic- level compression can reduce image data size to 0.1% or even lower, which has strong potential applications. However, CMC has cer- tain defects in consistency with the original image and perceptual quality. To inspire insights into such a problem, we introduce CMC- Bench, a benchmark of the cooperative performance of Image-to- Text (I2T) and Text-to-Image (T2I) models for image compression. This benchmark covers 18,000 and 40,000 images respectively to verify 6 mainstream I2T and 12 T2I models, including 160,000 sub- jective preference scores annotated by human experts. At ultra-low bitrates, it proves that the combination of some I2T and T2I models has surpassed the most advanced visual signal codecs; meanwhile, it highlights where LMMs can be further optimized toward the compression task. We encourage LMM developers to participate in this test to promote the evolution of visual signal codec protocols.
Chunyi Li 0001, Xiele Wu, Haoning Wu 0001, Donghui Feng 0003, Guo Lu, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin
ACM Multimedia7
2025 Towards Explainable Partial-AIGC Image Quality Assessment
Jiaying Qian, Ziheng Jia, Guangtao Zhai, Xiongkuo Min
ACM Multimedia6
2025 DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
abstract
With the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticity. Current deepfake detection methods often depend on datasets with limited generation models and content diversity that fail to keep pace with the evolving complexity and increasing realism of the AI-generated content. Large multimodal models (LMMs), widely adopted in various vision tasks, have demonstrated strong zero-shot capabilities, yet their potential in deepfake detection remains largely unexplored. To bridge this gap, we present DFBench, a large-scale DeepFake Benchmark featuring (i) broad diversity, including 540,000 images across real, AI-edited, and AI-generated content, (ii) latest model, the fake images are generated by 12 state-of-the-art generation models, and (iii) bidirectional benchmarking and evaluating for both the detection accuracy of deepfake detectors and the evasion capability of generative models. Based on DFBench, we propose MoA-DF, Mixture of Agents for DeepFake detection, leveraging a combined probability strategy from multiple LMMs. MoA-DF achieves state-of-the-art performance, further proving the effectiveness of leveraging LMMs for deepfake detection. Database and codes are publicly available at https://github.com/IntMeGroup/DFBench.
Huiyu Duan, Juntong Wang, Ziheng Jia, Woo Yi Yang, Xiaorong Zhu, Jiaying Qian, Yuke Xing, Guangtao Zhai, Xiongkuo Min
ACM Multimedia11
2025 LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
abstract
The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.
Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Shiqi Gao, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ACM Multimedia10
2025 Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni Model
abstract
360° omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires specialized equipment, making ODI synthesis increasingly important. While common 2D image generation and editing methods are rapidly advancing, these models struggle to deliver satisfactory results when generating or editing ODIs due to the unique format and broad 360° Field-of-View (FoV) of ODIs. To bridge this gap, we construct Any2Omni , the first comprehensive ODI generation-editing dataset comprises 60,000+ training data covering diverse input conditions and up to 9 ODI generation and editing tasks. Built upon Any2Omni, we propose an Omni model for Omni-directional image generation and editing ( Omni 2), with the capability of handling various ODI generation and editing tasks under diverse input conditions using one model. Extensive experiments demonstrate the superiority and effectiveness of the proposed Omni2 model for both the ODI generation and editing tasks. Both the Any2Omni dataset and the Omni2 model are publicly available at: https://github.com/IntMeGroup/Omni2.
Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Lu Liu 0005, Zitong Xu, Guangji Ma, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ACM Multimedia8
2025 LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs
abstract
The rapid advancement in generative artificial intelligence have enabled the creation of 3D human faces (HFs) for applications including media production, virtual reality, security, healthcare, and game development, etc. However, assessing the quality and realism of these AI-generated 3D human faces remains a significant challenge due to the subjective nature of human perception and innate perceptual sensitivity to facial features. To this end, we conduct a comprehensive study on the quality assessment of AI-generated 3D human faces. We first introduce Gen3DHF, a large-scale benchmark comprising 2,000 videos of AI-Generated 3D Human Faces along with 4,000 Mean Opinion Scores (MOS) collected across two dimensions, i.e., quality and authenticity, 2,000 distortion-aware saliency maps and distortion descriptions. Based on Gen3DHF, we propose LMME3DHF, a Large Multimodal Model (LMM)-based metric for Evaluating 3DHF capable of quality and authenticity score prediction, distortion-aware visual question answering, and distortion-aware saliency prediction. Experimental results show that LMME3DHF achieves state-of-the-art performance, surpassing existing methods in both accurately predicting quality scores for AI-generated 3D human faces and effectively identifying distortion-aware salient regions and distortion types, while maintaining strong alignment with human perceptual judgments. Both the Gen3DHF database and the LMME3DHF will be released upon the publication.
Woo Yi Yang, Sijing Wu, Huiyu Duan, Guangtao Zhai, Xiongkuo Min
ACM Multimedia9
2025 Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation Metric
abstract
AI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git.
Wei Sun 0029, Xinyue Li 0001, Qihang Ge, Jun Jia, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai
ACM Multimedia11
2025 GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
abstract
The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained physical principles especially in geometric optics, remains underexplored. To address this gap, we introduce GOBench, the first benchmark to systematically evaluate MLLMs' ability across two tasks: 1) Generating Optically Authentic Imagery and 2) Understanding Underlying Optical Phenomena. We curate high-quality prompts of geometric optical scenarios and use MLLMs to construct the GOBench-Gen-1k dataset. We then organize subjective experiments to assess the generated imagery based on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, revealing MLLMs' generation flaws that violate optical principles. For the understanding task, we apply crafted evaluation instructions to test the optical understanding ability of eleven prominent MLLMs. The experimental results demonstrate that current models face significant challenges in both optical generation and understanding. The top-performing generative model, GPT-4o-Image, cannot perfectly complete all generation tasks, and the best-performing MLLM model, Gemini-2.5Pro, attains a mere 37.35% accuracy in optical understanding. Database and codes are publicly available at: https://github.com/aiben-ch/GOBench.
Xiaorong Zhu, Ziheng Jia, Haodong Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
ACM Multimedia6
2025 CompBench: Benchmarking and Comparing Image Generation with Large Multimodal Models
abstract
Recent advancements in large multimodal models (LMMs) have significantly enhanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, critical challenges in perceptual quality and text-image correspondence remain hindering the practicality of AI-generated images (AIGIs). Therefore, a reliable benchmark and automatic model for AIGI evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present CompBench, the largest dataset for benchmarking and comparing image generation, which features: (i) the largest AIGI pair comparison dataset, comprising 616,346 carefully curated image pairs generated by 24 state-of-the-art AIGI models annotated with 1.6M+ human annotations, enabling robust relative quality assessment through pairwise comparison, (ii) multi-dimensional pairwise comparison from perceptual and text-image correspondence perspectives across three difficulty levels, and (iii) bidirectional benchmarking and evaluating for both T2I generation models and AIGI comparison models. Based on CompBench, we propose LMM4Comp, a LMM-based evaluation metric that learns nuanced quality distinctions from multiple dimensions for pairwise comparison at both instance level and model level. Experiments demonstrate that LMM4Comp achieves state-of-the-art performance, highly aligning to human preference. Both of the CompBench dataset and LMM4Comp metric will be released at https://github.com/IntMeGroup/CompBench.
Huiyu Duan, Yuke Xing, Yiling Xu, Guangtao Zhai, Xiongkuo Min
MMSP6
2025 CVBench: Benchmarking and Comparing Video Generation with Large Multimodal Models
abstract
Large multimodal models (LMMs) have revolutionized both text-to-video (T2V) generation and video-to-text (V2T) interpretation. However, despite these advancements, issues such as imperfect perceptual quality and inconsistent text-video alignment continue to limit the practical deployment of AI-generated videos (AIGVs). Consequently, there is a pressing need for a reliable benchmark and automatic evaluation framework tailored for AIGVs. To this end, we propose CVBench, the largest and most comprehensive dataset for Comparative Video Benchmarking, including 60K video pairs generated by 30 state-of-the-art T2V models and 600K pairwise comparisons annotated with over 1.7 million human judgments from perspectives of both perceptual quality and text-video correspondence. This dataset enables bidirectional benchmarking and evaluation of both T2V generation models and V2T interpretation models. Based on CVBench, we propose VComp, a novel LMM-based evaluation metric that captures fine-grained quality differences from multiple perspectives for pairwise comparison at both the instance level and model level. Extensive experiments show that VComp achieves state-of-the-art alignment with human preferences. Both the CVBench dataset and VComp metric will be available at https://github.com/IntMeGroup/CVBench.
Huiyu Duan, Yuke Xing, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min
VCIP6
2025 MGM-DPO: Multi-Generative-Model Guided Preference Optimization for Advanced Text-to-Image Models
abstract
Text-to-image generation has rapidly progressed with autoregressive, diffusion, and distillation-based models, enabling high-fidelity and semantically relevant image synthesis from natural language prompts. Despite their success, these models still suffer from text–image misalignment, limited detail, and outputs that may violate human commonsense or aesthetic preferences. Existing human-feedback-based fine-tuning approaches partially alleviate these issues but primarily focus on earlier diffusion models (e.g., Stable Diffusion v1.4, v2.0, SDXL) and often exhibit over-optimization or under-optimization, leaving their effectiveness on state-of-the-art models unclear. In this work, we advance preference alignment for modern text-to-image models. We utilize EvalMi-50K, a large-scale dataset with human preference scores, to train our image quality assessment (IQA) model and use the predicted scores from the IQA model to reward the generation model. Building upon the dataset and IQA model, we propose Multi-Generative-Model Guided Diffusion Preference Optimization(MGM-DPO), which leverages diverse model outputs and a stable optimization strategy to align advanced diffusion models. Applied to Stable Diffusion 3.5 Large, MGM-DPO significantly improves fidelity, prompt alignment, and human preference alignment across multiple benchmarks, achieving state-of-the-art results among human-feedback-tuned models. Code is available at https://github.com/IntMeGroup/MGM-DPO.
Huiyu Duan, Guangtao Zhai, Xiongkuo Min
VCIP5
2025 Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai
Sci. China Inf. Sci.40
2025 A study on the user viewing experience of implanted advertisement videos based on visual saliency
Fangfang Lu, Yingjie Lian, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
Expert Syst. Appl.4
2025 CT-PCQA: A Convolutional Neural Network and Transformer combined Method for Point Cloud Quality Assessment
Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
Signal Process. Image Commun.4
2025 Energy-Efficient VR 360 Video Streaming in the IRS-Aided Rate-Splitting Multiple Access Network
abstract
Maximizing energy efficiency in VR 360 video transmission is essential for advancing VR applications. Motivated by this goal, we conduct a comprehensive study that integrates the characteristics of VR 360 video with beamforming strategies and intelligent reflecting surface (IRS) shifting techniques in the rate-splitting multiple access (RSMA) network. In the IRS-aided RSMA network, we propose a stable energy-efficient transmission (SEET) scheme aimed at minimizing the number of transmitted VR video chunks. The SEET scheme constructs a stable pre-transmission and playback flow, ensuring seamless and continuous display of the upcoming content without latency. We also propose a mixed-format-based chunk (MFC) method that simultaneously pre-transmits both 2D and 3D chunk frames to each user, further enhancing energy efficiency. We utilize an alternating optimization method to divide the original energy-efficient problem into three subproblems. To tackle the non-convex and NP-hard beamforming subproblem, we utilize the first-order Taylor expansion and then obtain the approximate transmission rates of common messages and private messages regarding the quadratic form of beamforming vectors. We then utilize quadratically constrained programming, fractional programming, and linear programming to obtain the near-optimal solutions for beamforming vectors, IRS phase shifts, and RSMA parameters, respectively. The final numerical results affirm that the proposed SEET scheme can notably minimize the beamforming power of the base station. Through the SEET scheme, the MFC method with the approximation method exhibits superior energy efficiency, outperforming existing transmission methods in terms of both energy utility and consumption by HMDs.
Qingqing Wu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Commun.4
2025 Time-Smooth Wireless Transmission of Probabilistic Slicing VR 360 Video in MISO-OFDM Systems
abstract
The multiple-input and single-output (MISO)-orthogonal frequency-division multiplexing (OFDM) systems afford low latency and high reliability for virtual reality (VR) 360 video in multi-user scenarios. Motivated by the goal of maintaining time-smoothness while holding acceptably low complexity, a crucial factor in VR video transmission, we conduct a comprehensive study that integrates the characteristics of VR video with the strategies for subcarrier assignment and power allocation. By analyzing the pre-transmitted tile-segments, the missing tile-segments, and the video frame structure, we propose two probabilistic slicing schemes (PSPs) to minimize the size of required tile-segments of VR video scenes. In time-smoothness maximization, the desired discrete encoding rate set, discrete subcarrier assignment, continuous power allocation, and fixed total power constraint make it a challenging mixed-integer nonlinear programming (MINLP) problem. Unlike the straightforward relaxation-recovery method, we firstly prove that a near-optimal recovered encoding rate is the discrete value closest to the optimal relaxed-continuous encoding rate. We then propose a Two-step Encoding Rate Maximization (TERM) method, including the relaxed-continuous sum-rate maximization and the discrete encoding rate recovery, to achieve the near-optimal subcarrier assignment and the power allocation with low complexity. Simulation results on real-world VR video dataset validate that the two PSPs can effectively minimize the number of transmitted tile-segments. The proposed TERM with PSPs can maintain time-smoothness of VR 360 video with an acceptably low level of complexity in MISO-OFDM systems.
Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Biqian Feng, Yucheng Zhu, Wenjun Zhang 0001
IEEE Trans. Commun.4
2025 Joint Luminance-Chrominance Learning for Image Debanding
abstract
Banding is a visually annoying artifact that frequently occurs along the chain of video acquisition, production, distribution, and display, showing a significant need for improvement in many fields. Thus far, efforts on banding removal are mainly knowledge-driven or merely learning on RGB space, which is either limited by domain knowledge or lacks the consideration for banding in chrominance channels. In this work, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner. Our debanding model is comprised of a luminance restoration network (LR-Net) and a chrominance restoration network (CR-Net). Each of them follows an encoder-decoder architecture, where a cascade of residual blocks is employed to exploit hierarchical non-local features in spatial dimensions for more powerful feature representation. Moreover, we investigate the characteristics of banding artifacts and apply specific loss functions to guide the debanding in different channels, thus boosting the restoration performance. Both qualitative and quantitative experiments show that our model significantly surpasses the existing method in terms of all 7 metrics. Ultimately, our network trained on simulated data exhibits good adaptiveness under various compression scenarios, which further demonstrates the effectiveness of the proposed model.
Zijian Chen 0001, Wei Sun 0029, Jun Jia, Ru Huang 0002, Fangfang Lu, Ying Chen 0011, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Study of Subjective and Objective Naturalness Assessment of AI-Generated Images
abstract
The proliferation of Artificial Intelligence-Generated Images (AIGIs) has greatly expanded the Image Naturalness Assessment (INA) problem. Different from early definitions that mainly focus on tone-mapped images with limited distortions (e.g., exposure, contrast, and color reproduction), INA on AI-generated images is especially challenging as it owns more diverse contents and could be affected by factors from multiple perspectives, including low-level technical distortions and high-level rationality distortions. In this paper, we take the first step to benchmark and assess the visual naturalness of AI-generated images. First, we construct the AI-Generated Image Naturalness (AGIN) dataset by conducting a large-scale subjective study to collect human opinions on the overall naturalness as well as perceptions from the technical quality and rationality perspectives. AGIN verifies several insights for the first time that naturalness is universally and disparately affected by both technical and rational distortions, while its manifestations vary with different generation tasks. Second, to automatically assess the naturalness of AIGIs that align with human opinions, we propose the Joint Objective Image Naturalness evaluaTor (JOINT). Specifically, JOINT imitates human reasoning in naturalness evaluation by jointly learning technical and rationality features with several specific designs to guide model behavior from respective perspectives. Experiments demonstrate that JOINT significantly outperforms existing methods for providing more subjectively consistent results on naturalness assessment. The dataset can be accessed athttps://github.com/zijianchen98/AGIN.
Zijian Chen 0001, Wei Sun 0029, Haoning Wu 0001, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 LCIQA: A Lightweight Contrastive-Learning Framework for Image Quality Assessment via Cross-Scale Consistency Minimization
abstract
Blind image quality assessment (BIQA), which functions without the need for a reference image, is a challenging yet essential task in various image processing systems and downstream vision applications, ranging from semantic recognition to image enhancement. Traditionally, numerous BIQA models have been developed using supervised learning methodologies, which rely heavily on the availability and quality of ground truth data. To improve the generalization capability and robustness of these models, recent studies have explored the application of contrastive learning, aiming to enhance the quality representation capacity of model backbones through a self-supervised approach. However, the training process for contrastive learning is computationally intensive, posing significant challenges in resource-constrained environments. To mitigate this issue, we propose a Lightweight Contrastive-learning-based IQA (LCIQA) framework, designed to be efficiently trained on a single GPU without relying on ground truth data. This framework maintains a fixed vision backbone and focuses on optimizing the parameters of subsequent IQA heads through contrastive learning. To accommodate a lightweight framework, we incorporate a quality task adapter to eliminate semantic biases introduced by the features extracted from the fixed-parameter backbone. A coarse-to-fine contrastive learning strategy is then employed to train the quality regression module. Extensive experiments demonstrate the superior performance of our model in terms of both accuracy and complexity. In addition, ablation studies validate the effectiveness of each component within the proposed framework.
Chenxi Feng, Xiongkuo Min, Long Ye, Yinghao Yang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 No-Reference Image Quality Assessment: Obtain MOS From Image Quality Score Distribution
abstract
Recent image quality assessment (IQA) methods typically focus on predicting the mean opinion score (MOS) of image quality, ignoring the image quality score distribution. This distribution provides valuable information beyond the MOS, including the standard deviation of opinion scores (SOS) and opinion scores at different quality levels. This paper introduces a novel no-reference IQA method that predicts the image quality score distribution to estimate the MOS. The proposed method consists of three modules: a visual feature extraction module, a graph convolutional module, and a MOS prediction module. In the visual feature extraction module, a convolutional neural network is designed to extract both first- and second-order visual features of images. The graph convolutional module employs a graph convolutional network (GCN)-based mapper to map these visual features to the image quality score distribution by exploring correlations between quality labels. The MOS is then derived from the predicted image quality score distribution in the MOS prediction module. We are the first to jointly train the method using both the MOS and the image quality score distribution, enabling it to learn richer subjective information and improve prediction performance. To address the lack of the ground-truth image quality score distribution in some IQA databases, we propose to use a SOS assumption to generate a Gaussian-based image quality score distribution that better reflects subjective perception. Additionally, we design appropriate loss functions for training. Experimental results demonstrate that our method effectively predicts both the image quality score distribution and the MOS, outperforming most state-of-the-art IQA methods.
Xiongkuo Min, Yuqin Cao, Xiaohong Liu 0001, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2025 LMM-VQA: Advancing Video Quality Assessment With Large Multimodal Models
abstract
The explosive growth of videos on streaming media platforms has underscored the urgent need for effective video quality assessment (VQA) algorithms to monitor and perceptually optimize the quality of streaming videos. However, VQA remains an extremely challenging task due to the diverse video content and the complex spatial and temporal distortions, thus necessitating more advanced methods to address these issues. Nowadays, large multimodal models (LMMs), such as GPT-4V, have exhibited strong capabilities for various visual understanding tasks, motivating us to leverage the powerful multimodal representation ability of LMMs to solve the VQA task. Therefore, we propose anLargeMulti-Modal basedVideoQuality Assessment (LMM-VQA) model, which introduces a novel spatiotemporal visual modeling strategy for quality-aware feature extraction. Specifically, we reformulate the quality regression problem into a question and answering (Q&A) task and construct Q&A prompts for VQA instruction tuning. Then, we design a spatiotemporal vision encoder to extract spatial and temporal features to represent the quality characteristics of videos, which are subsequently mapped into the language space by the spatiotemporal projector for modality alignment. Finally, the aligned visual tokens and the quality-inquired text tokens are aggregated as inputs for the large language model (LLM) to generate the quality score as well as the quality level. Extensive experiments demonstrate thatLMM-VQAachieves state-of-the-art performance across five VQA benchmarks, exhibiting an average improvement of 5% in generalization ability over existing methods. Furthermore, due to the advanced design of the spatiotemporal encoder and projector, LMM-VQA also performs exceptionally well on general video understanding tasks, further validating its effectiveness. Our code will be released at https://github.com/Sueqk/LMM-VQA.
Qihang Ge, Wei Sun 0029, Yu Zhang 0133, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.8
2025 Full-Reference and No-Reference Quality Assessment for Video Frame Interpolation
abstract
Video frame interpolation (VFI) synthesizes new frames from original video frames to produce high frame-rate videos and enhance their visual appeal. The quality of these interpolated frames significantly affects the perceptual experience of the synthesized video. Recent research in VFI has increasingly focused on perceptual quality of the interpolated frames and the overall video. However, most existing quality metrics do not align well with human perceptual experiences and often suffer from unnatural artifacts in the interpolated frames. Consequently, there is an urgent need for VFI video quality assessment (VFIVQA) methods to assess the quality of the synthesized videos. In this paper, we propose both a full-reference (FR) method and a no-reference (NR) method for VFIVQA. The FR method employs two feature extraction blocks to measure continuous frame changes, extracting flow features with short temporal spans and motion features with long temporal spans. By calculating multilevel similarities in the temporal dimension of 3D convolutional neural networks and fusing these similarity features, the quality score of the VFI video is obtained from the quality regression network. Since the flow feature extraction block does not utilize the reference VFI video, the proposed NR method consists solely of this feature block. Extensive validation on several VFIVQA datasets demonstrates that the proposed methods outperform state-of-the-art FR and NR methods.
Jinliang Han, Xiongkuo Min, Jun Jia, Xiaohong Liu 0001, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2025 Exploring Rich Subjective Quality Information for Image Quality Assessment in the Wild
abstract
Traditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method namedRichIQAto explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: 1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; 2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub.
Xiongkuo Min, Yuqin Cao, Guangtao Zhai, Wenjun Zhang 0001, Huifang Sun, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.1
2025 Who Is a Better Imitator: Subjective and Objective Quality Assessment of Animated Humans
abstract
Animated human (AH) have gained popularity due to their vivid appearance and smooth, natural movements. Various animation methods based on artificial intelligence (AI) have been introduced, which are viewed as “Imitators,” offering new solutions for designing AHs. However, the effectiveness of these AI-generated AHs varies significantly across different categories and within the same category, leading to visual distortions that adversely affect the viewer’s experience. Consequently, it is essential to evaluate the quality of AHs to provide reliable and objective indicators for their further development and to ensure the delivery of higher-quality AH videos to users. In this paper, the first Animated Human Quality Assessment (AHQA) dataset is constructed by selecting 6 advanced and popular imitators and 10 common actions to animate 20 AI-generated characters. The constructed dataset integrates different genders and age groups of character images, and two types of poses, standing and sitting, are selected, highlighting the comprehensiveness and diversity of the AHQA dataset. Subjective experiments reveal significant differences in the quality of AHs produced by different imitators. Finally, we propose a quality assessment method, VIP-QA, incorporating Video quality, Identity consistency, and Posture similarity for the AHQA dataset. Experimental results show that VIP-QA significantly outperforms existing assessment methods on multiple datasets by about 5%, more closely approximates human visual perception, and provides a valid objective metric for assessing imitators. All the work in this paper has been released at https://github.com/zyj-2000/Imitator.
Yingjie Zhou 0003, Jun Jia, Yanwei Jiang, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.6
2025 ScanDTM: A Novel Dual-Temporal Modulation Scanpath Prediction Model for Omnidirectional Images
abstract
Scanpath prediction for omnidirectional images aims to effectively simulate the human visual perception mechanism to generate dynamic realistic fixation trajectories. However, the majority of scanpath prediction methods for omnidirectional images are still in their infancy as they fail to accurately capture the time-dependency of viewing behavior and suffer from sub-optimal performance along with limited generalization capability. A desirable solution should achieve a better trade-off between prediction performance and generalization ability. To this end, we propose a novel dual-temporal modulation scanpath prediction (ScanDTM) model for omnidirectional images. Such a model is designed to effectively capture long-range time-dependencies between various fixation regions across both internal and external time dimensions, thereby generating more realistic scanpaths. In particular, we design a Dual Graph Convolutional Network (Dual-GCN) module comprising a semantic-level GCN and an image-level GCN. This module servers as a robust visual encoder that captures spatial relationships among various object regions within an image and fully utilizes similar images as complementary information to capture similarity relations across relevant images. Notably, the proposed Dual-GCN focuses on modeling temporal correlations from both local and global perspectives within the internal time dimension. Furthermore, drawing inspiration from the promising generalization capabilities of diffusion models across various generative tasks, we introduce a novel diffusion-guided saliency module. This module formulates the prediction issue as a conditional generative process for the saliency map, utilizing extracted semantic-level and image-level visual features as conditions. With the well-designed diffusion-guided saliency module, our proposed ScanDTM model acting as an external temporal modulator, we can progressively refine the generated scanpath from the noisy map. We conduct extensive experiments on several benchmark datasets, and the results demonstrate that our ScanDTM model significantly outperforms other competitors. Meanwhile, when applied to tasks such as saliency prediction and image quality assessment, our ScanDTM model consistently achieves superior generalization performance.
Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Blind Image Quality Assessment by Gaussian Mixture Distribution
abstract
In the field of image quality assessment (IQA), researchers have been studying the mean opinion score (MOS) of image quality for decades. They focus on developing IQA methods with the help of MOS without using the potential of the distribution of opinion scores (DOS). We find that the Gaussian mixture distribution (GMD) can more accurately describe the DOS of image quality on SJTU IQSD and KonIQ-10K databases compared to some traditional distributions. Therefore, this paper proposes a blind IQA method that predicts the MOS of image quality by learning the GMD-based image quality. The proposed method consists of a visual feature learning module and a GMD learning module. The visual feature learning module uses a multi-stage Swin Transformer model and a CLIP feature extractor to extract visual features from an image. The GMD learning module then maps the extracted visual features to the GMD-based image quality using a mixture density network, where the mean of the GMD represents the MOS of image quality. We not only use the MOS of image quality to train the proposed method, but also employ the DOS of image quality for auxiliary training to improve the prediction performance of the proposed method. To address the lack of DOS in some existing IQA databases, we introduce a pseudo DOS generation strategy to generate the DOS of image quality for training, which significantly improves the applicability of the proposed method. Numerous analyses show that the proposed method is superior to most state-of-the-art IQA methods in predicting both the MOS and the DOS, thus facilitating a deeper investigation into the DOS of image quality in IQA.
Xiongkuo Min, Yuqin Cao, Weisi Lin, Bu-Sung Lee, Guangtao Zhai
IEEE Trans. Image Process.2
2025 Advancing Zero-Shot Digital Human Quality Assessment Through Text-Prompted Evaluation
abstract
Digital humans have witnessed extensive applications in various domains, necessitating related quality assessment studies. However, there is a lack of comprehensive digital human quality assessment (DHQA) databases. To address this gap, we propose SJTU-H3D, a subjective quality assessment database specifically designed for full-body digital humans. It comprises 40 high-quality reference digital humans and 1,120 labeled distorted counterparts generated with seven types of distortions. The SJTU-H3D database can serve as a benchmark for DHQA research, allowing evaluation and refinement of processing algorithms. Further, we propose a zero-shot DHQA approach that focuses on no-reference (NR) scenarios to ensure generalization capabilities while mitigating database bias. Our method leverages semantic and distortion features extracted from projections, as well as geometry features derived from the mesh structure of digital humans. Specifically, we employ the Contrastive Language-Image Pre-training (CLIP) model to measure semantic affinity and incorporate the Naturalness Image Quality Evaluator (NIQE) model to capture low-level distortion information. Additionally, we utilize dihedral angles as geometry descriptors to extract mesh features. By aggregating these measures, we introduce the Digital Human Quality Index (DHQI), which demonstrates significant improvements in zero-shot performance. The DHQI can also serve as a robust baseline for DHQA tasks, facilitating advancements in the field. The database and the code are available at https://github.com/zzc-1998/SJTU-H3D.
Wei Sun 0029, Yingjie Zhou 0003, Haoning Wu 0001, Chunyi Li 0001, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin
IEEE Trans. Image Process.6
2025 Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional Videos
abstract
Virtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet.
Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Image Process.7
2025 How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and Model
abstract
Understanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency prediction in ODVs, the joint exploitation of these two modalities has been limited, primarily due to the absence of large-scale audio-visual saliency databases and comprehensive analyses. This paper comprehensively investigates audio-visual attention in ODVs from both subjective and objective perspectives. Specifically, we first introduce a new audio-visual saliency database for omnidirectional videos, termed AVS-ODV database, containing 162 ODVs and corresponding eye movement data collected from 60 subjects under three audio modes including mute, mono, and ambisonics. Based on the constructed AVS-ODV database, we perform an in-depth analysis of how audio influences visual attention in ODVs. To advance the research on audio-visual saliency prediction for ODVs, we further establish a new benchmark based on the AVS-ODV database by testing numerous state-of-the-art saliency models, including visual-only models and audio-visual models. In addition, given the limitations of current models, we propose an innovative omnidirectional audio-visual saliency prediction network (OmniAVS), which is built based on the U-Net architecture, and hierarchically fuses audio and visual features from the multimodal aligned embedding space. Extensive experimental results demonstrate that the proposed OmniAVS model outperforms other state-of-the-art models on both ODV AVS prediction and traditional AVS prediction tasks. The AVS-ODV database and the OmniAVS model are available at: https://github.com/IntMeGroup/AVS-ODV.
Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Xilei Zhu, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Image Process.7
2025 From Haziness to Clarity: A Novel Iterative Memory-Retrospective Emergence Model for Omnidirectional Image Saliency Prediction
abstract
To achieve saliency prediction in omnidirectional images (ODIs), the majority of prior works typically adopt the convolutional neural networks (CNNs)-based saliency models to extract semantic features to predict prominent regions in ODIs. Albeit achieving substantially performance gains, these works all employed purely visual computing paradigms and ignore to explore the nature of human visual attention mechanisms. In other words, existing saliency prediction works for ODIs are insufficient to capture the biological characteristics of the visual attention mechanism in the human brain. To establish a more explicit link between saliency prediction performance and brain-like visual attention mechanism, we simulate the mechanism of human retrospective memory in neuropsychology and propose IMRE model, a novel iterative memory-retrospective emergence model can predict and infer the salient features by recalling previously learned information. In IMRE model, we introduce four key modules to simulate the visual attention mechanism for predicting human fixations in the human brain. Firstly, the visual stimulus response module is designed to effectively extract semantic features and capture the intricate relationship between these features, acting as the human visual cortex. Secondly, the retrospective integration module serves to distill valuable information from a fuzzy memory ensemble, resembling the role of the basal ganglia in the neural system. Thirdly, the memory bank module explicitly records and stores subconscious response information and learned knowledge, acting like the hippocampus in neural system. Lastly, the prospective inference module accurately infers saliency maps from the refined useful information, resembling the role of the prefrontal cortex. During prediction, we utilize the introduced memory bank to retrieve and recall previously learned information, which simulates the process of memory emergence from haziness to clarity. Such a process aligns with the retrospective memory mechanism of the human brain. To validate the superiority of the proposed model in ODIs saliency prediction tasks, we conduct extensive experiments on two benchmark datasets. Experiments show impressive performances that IMRE model outperforms other state-of-the-art methods across all benchmark datasets. Importantly, experiments also highlight the IMRE model's ability to trace back to specific instances during prediction, thereby reducing model inference costs and enhancing interpretability.
Dandan Zhu 0001, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Image Process.3
2025 Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and Model
abstract
Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively. Specifically, the proposed method utilizes the projection videos of text-to-3D assets to extract 3D shape, texture and text-asset correspondence features, then fuses them to calculate the final three preference scores respectively. Extensive experimental results demonstrate the effectiveness of the proposed T23DAQA method in evaluating the quality of AI generated 3D asset, which is more consistent with human perception. To the best of our knowledge, this is the first work that studies the problem of text-guided 3D generation quality assessment, and our database and codes will be released to facilitate future research.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
IEEE Trans. Multim.5
2025 Quality-Guided Skin Tone Enhancement for Portrait Photography
abstract
In recent years, learning-based color and tone enhancement methods for photos have become increasingly popular. However, most learning-based image enhancement methods just learn a mapping from one distribution to another based on one dataset, lacking the ability to adjust images continuously and controllably. It is important to enable the learning-based enhancement models to adjust an image continuously, since in many cases we may want to get a slighter or stronger enhancement effect rather than one fixed adjusted result. In this paper, we propose a quality-guided image enhancement paradigm that enables image enhancement models to learn the distribution of images with various quality ratings. By learning this distribution, image enhancement models can associate image features with their corresponding perceptual qualities, which can be used to adjust images continuously according to different quality scores. To validate the effectiveness of our proposed method, a subjective quality assessment experiment is first conducted, focusing on skin tone adjustment in portrait photography. Guided by the subjective quality ratings obtained from this experiment, our method can adjust the skin tone corresponding to different quality requirements. Furthermore, an experiment conducted on 10 natural raw images corroborates the effectiveness of our model in situations with fewer subjects and fewer shots, and also demonstrates its general applicability to natural images.
Shiqi Gao, Huiyu Duan, Xinyue Li 0001, Yicong Peng, Qihang Xu, Yuanyuan Chang, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Multim.9
2025 Aggregate and Discriminate: Pseudo Clips-Guided Boundary Perception for Video Moment Retrieval
abstract
Video moment retrieval (VMR) aims to localize a video segment in an untrimmed video that is semantically relevant to a language query. The challenge of this task lies in effectively aligning the intricate and information-dense video modality with the succinctly summarized textual modality, and further localizing the starting and ending timestamps of the target moments. Previous works have attempted to achieve multi-granularity alignment of video and query in a coarse-to-fine manner, yet these efforts still fall short in addressing the inherent disparities in representation and information density between videos and queries, leading to modal misalignments. In this paper, we propose a progressive video moment retrieval framework, initially retrieving the most relevant and irrelevant video clips to the query as semantic guidance, thereby bridging the semantic gap between video modality and language modality. Futhermore, we introduce a pseudo clips guided aggregation module to aggregate densely relevant moment clips closer together and propose a discriminative boundary-enhanced decoder with the guidance of pseudo clips to push the semantically confusing proposals away. Extensive experiments on the Charades-STA, ActivityNet Captions and TACoS datasets demonstrate that our method outperforms existing methods.
Jing Liu 0002, Zongbing Zhang, Yuting Su 0001, Bing Yang 0003, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Multim.5
2025 Evaluating Point Cloud From Moving Camera Videos: A No-Reference Metric
abstract
Point cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as compression and downsampling distortions during the transmission process. To tackle the challenge of point cloud quality assessment (PCQA), many PCQA methods have been proposed to evaluate the visual quality levels of point clouds by assessing the rendered static 2D projections. Although such projectionbased PCQA methods achieve competitive performance with the assistance of mature image quality assessment (IQA) methods, they neglect that the 3D model is also perceived in a dynamic viewing manner, where the viewpoint is continually changed according to the feedback of the rendering device. Therefore, in this paper, we evaluate the point clouds from moving camera videos and explore the way of dealing with PCQA tasks via using video quality assessment (VQA) methods. First, we generate the captured videos by rotating the camera around the point clouds through several circular pathways. Then we extract both spatial and temporal quality-aware features from the selected key frames and the video clips through using trainable 2D-CNN and pretrained 3D-CNN models respectively. Finally, the visual quality of point clouds is represented by the video quality values. The experimental results reveal that the proposed method is effective for predicting the visual quality levels of the point clouds and even competitive with full-reference (FR) PCQA methods. The ablation studies further verify the rationality of the proposed framework and confirm the contributions made by the qualityaware features extracted via the dynamic viewing manner. The code is available athttps://github.com/zzc-1998/VQA_PC.
Wei Sun 0029, Yucheng Zhu, Xiongkuo Min, Wei Wu 0002, Ying Chen 0011, Guangtao Zhai
IEEE Trans. Multim.4
2025 Explain Vision Focus: Blending Human Saliency Into Synthetic Face Images
abstract
Synthetic faces have been extensively researched and applied in various fields, such as face parsing and recognition. Compared to real face images, synthetic faces engender more controllable and consistent experimental stimuli due to the ability to precisely merge expression animations onto the facial skeleton. Accordingly, we establish an eye-tracking database with 780 synthetic face images and fixation data collected from 22 participants. The use of synthetic images with consistent expressions ensures reliable data support for exploring the database and determining the following findings: (1) A correlation study between saliency intensity and facial movement reveals that the variation of attention distribution within facial regions is mainly attributed to the movement of the mouth. (2) A categorized analysis of different demographic factors demonstrates that the bias towards salient regions aligns with differences in some demographic categories of synthetic characters. In practice, inference of facial saliency distribution is commonly used to predict the regions of interest for facial video-related applications. Therefore, we propose a benchmark model that accurately predicts saliency maps, closely matching the ground truth annotations. This achievement is made possible by utilizing channel alignment and progressive summation for feature fusion, along with the incorporation of Sinusoidal Position Encoding. The ablation experiment also demonstrates the effectiveness of our proposed model. We hope that this paper will contribute to advancing the photorealism of generative digital humans.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Huiyu Duan, Guangtao Zhai
IEEE Trans. Multim.3
2025 Elevating Mesh Saliency in VR: Introducing a Novel Prediction Network and Dataset
abstract
In computer graphics, polygon meshes stand out as a popular representation providing effective delineation of delicate textures and complex geometries. When dealing with geometric processing tasks for critical regions of the mesh, it is necessary to consider the human visual perception related to saliency. Therefore, we establish a novel mesh saliency dataset, facilitated by a more comprehensive gathering pipeline of eye-tracking from subjects observing mesh models at arbitrary viewpoints in a virtual reality space with six degrees of freedom. Additionally, we propose a mesh saliency prediction model that accurately infers visual attention density maps for complex and irregular mesh surfaces. This model integrates surface curvature and triangular face shape information from multi-scale neighboring ranges as local geometric features, while also leveraging surface spatial positioning as a global feature. Our work aims to preserve critical areas and minimize visual loss in saliency-driven tasks such as mesh simplification, rendering, and texturing. We believe that our research can offer valuable insights for human-centered mesh computation applications.
Kaiwei Zhang, Mohan He, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model
abstract
In recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git .
Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.5
2025 MM-PCQA+: Advancing Multi-Modal Learning for Point Cloud Quality Assessment
abstract
The importance of visual quality in point clouds has been significantly underlined due to the rapid rise in 3D vision applications which aim to deliver affordable and superior user experiences. Reviewing the evolution of point cloud quality assessment (PCQA), it’s observed that visual quality evaluation typically employs single-modal data, either sourced from 2D projections or the 3D point clouds. The 2D projections possess abundant texture and semantic information while they are heavily reliant on viewpoints. In contrast, 3D point clouds are more reactive to geometric distortions and viewpoint-invariant. Consequently, to maximize the benefits of both point cloud and image modalities, we present an advanced no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA+) metric. Specifically, we divide the point clouds into sub-models to reflect local geometric distortions such as point shifting and down-sampling. Afterwards, we render the point clouds using a cube-like projection setup and sample the projections of interest using a point-visible-ratio for image feature extraction. In order to fulfill these objectives, the sub-models and projected images are encoded using point-based and image-based neural networks. Lastly, we implement symmetric cross-modal attention to amalgamate multi-modal quality-aware features. Experimental results demonstrate that our metric surpasses all state-of-the-art methods and significantly advances beyond previous no-reference PCQA methods.
Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Unified Approach to Mesh Saliency: Evaluating Textured and Non-Textured Meshes Through VR and Multifunctional Prediction
abstract
Mesh saliency aims to empower artificial intelligence with strong adaptability to highlight regions that naturally attract visual attention. Existing advances primarily emphasize the crucial role of geometric shapes in determining mesh saliency, but it remains challenging to flexibly sense the unique visual appeal brought by the realism of complex texture patterns. To investigate the interaction between geometric shapes and texture features in visual perception, we establish a comprehensive mesh saliency dataset, capturing saliency distributions for identical 3D models under both non-textured and textured conditions. Additionally, we propose a unified saliency prediction model applicable to various mesh types, providing valuable insights for both detailed modeling and realistic rendering applications. This model effectively analyzes the geometric structure of the mesh while seamlessly incorporating texture features into the topological framework, ensuring coherence throughout appearance-enhanced modeling. Through extensive theoretical and empirical validation, our approach not only enhances performance across different mesh types, but also demonstrates the model's scalability and generalizability, particularly through cross-validation of various visual features.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Vis. Comput. Graph.3
2025 ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
abstract
With the development of eXtended Reality (XR), photo capturing and display technology based on head-mounted displays (HMDs) have experienced significant advancements and gained considerable attention. Egocentric spatial images and videos are emerging as a compelling form of stereoscopic XR content. The assessment for the Quality of Experience (QoE) of XR content is important to ensure a high-quality viewing experience. Different from traditional 2D images, egocentric spatial images present challenges for perceptual quality assessment due to their special shooting, processing methods, and stereoscopic characteristics. However, the corresponding image quality assessment (IQA) research for egocentric spatial images is still lacking. In this paper, we establish the Egocentric Spatial Images Quality Assessment Database (ESIQAD), the first IQA database dedicated for egocentric spatial images as far as we know. Our ESIQAD includes 500 egocentric spatial images and the corresponding mean opinion scores (MOSs) under three display modes, including 2D display, 3D-window display, and 3D-immersive display. Based on our ESIQAD, we propose a novel mamba2-based multi-stage feature fusion model, termed ESIQAnet, which predicts the perceptual quality of egocentric spatial images under the three display modes. Specifically, we first extract features from multiple visual state space duality (VSSD) blocks, then apply cross attention to fuse binocular view information and use transposed attention to further refine the features. The multi-stage features are finally concatenated and fed into a quality regression network to predict the quality score. Extensive experimental results demonstrate that the ESIQAnet outperforms 22 state-of-the-art IQA models on the ESIQAD under all three display modes. The database and code are available at https://github.com/IntMeGroup/ESIQA.
Xilei Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Vis. Comput. Graph.4
2024 UniProcessor: A Text-Induced Unified Low-Level Image Processor
Huiyu Duan, Xiongkuo Min, Sijing Wu, Wei Shen 0002, Guangtao Zhai
ECCV (67)2
2024 GLARE: Low Light Image Enhancement via Generative Latent Feature Based Codebook Retrieval
Han Zhou 0003, Wei Dong 0011, Xiaohong Liu 0001, Shuaicheng Liu, Xiongkuo Min, Guangtao Zhai, Jun Chen 0005
ECCV (48)5
2024 A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital Humans
abstract
In an era where 3D Digital Humans (DHs) are becoming increasingly prevalent in fields like gaming, automotive, and the metaverse, the demand for high DH visual quality is rising. This paper presents the first-ever reduced-reference (RR) quality assessment metric tailored specifically for textured mesh DHs, aiming to optimize transmission systems and improve Quality of Experience (QoE) for viewers in resource-constrained environments. Four critical geometric curvature-related attributes and two texture-related indicators are computed, which are then statistically analyzed and utilized in a Support Vector Regression (SVR) model for robust and efficient quality prediction. Experimental results confirm that our method outperforms existing full-reference (FR) metrics, making it an invaluable tool for the future of 3D DHs in various applications. The code is available at https://github.com/zzc-1998/RR-DHQA.
Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ICASSP7
2024 SG-JND: Semantic-Guided Just Noticeable Distortion Predictor for Image Compression
abstract
Just noticeable distortion (JND), representing the threshold of distortion in an image that is minimally perceptible to the human visual system (HVS), is crucial for image compression algorithms to achieve a trade-off between transmission bit rate and image quality. However, traditional JND prediction methods only rely on pixel-level or sub-band level features, lacking the ability to capture the impact of image content on JND. To bridge this gap, we propose a Semantic-Guided JND (SG-JND) network to leverage semantic information for JND prediction. In particular, SG-JND consists of three essential modules: the image preprocessing module extracts semantic-level patches from images, the feature extraction module extracts multi-layer features by utilizing the cross-scale attention layers, and the JND prediction module regresses the extracted features into the final JND value. Experimental results show that SG-JND achieves the state-of-the-art performance on two publicly available JND datasets, which demonstrates the effectiveness of SG-JND and highlight the significance of incorporating semantic information in JND assessment.
Linhan Cao, Wei Sun 0029, Xiongkuo Min, Jun Jia, Zijian Chen 0001, Yucheng Zhu, Lizhou Liu, Qiubo Chen, Guangtao Zhai
ICIP3
2024 AIGCOIQA2024: Perceptual Quality Assessment of AI Generated Omnidirectional Images
abstract
[?]In recent years, the rapid advancement of Artificial Intelligence Generated Content (AIGC) has attracted widespread attention. Among the AIGC, AI generated omnidirectional images hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications, hence omnidirectional AIGC techniques have also been widely studied. AI-generated omnidirectional images exhibit unique distortions compared to natural omnidirectional images, however, there is no dedicated Image Quality Assessment (IQA) criteria for assessing them. This study addresses this gap by establishing a large-scale AI generated omnidirectional image IQA database named AIGCOIQA2024 and constructing a comprehensive benchmark. We first generate 300 omnidirectional images based on 5 AIGC models utilizing 25 text prompts. A subjective IQA experiment is conducted subsequently to assess human visual preferences from three perspectives including quality, comfortability, and correspondence. Finally, we conduct a benchmark experiment to evaluate the performance of state-of-the-art IQA models on our database. The AIGCOIQA2024 database is released to facilitate future research on https://github.com/IntMeGroup/AIGCOIQA.
Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICIP7
2024 Thqa: A Perceptual Quality Assessment Database for Talking Heads
abstract
In the realm of media technology, digital humans have gained prominence due to rapid advancements in computer technology. However, the manual modeling and control required for the majority of digital humans pose significant obstacles to efficient development. The speech-driven methods offer a novel avenue for manipulating the mouth shape and expressions of digital humans. Despite the proliferation of driving methods, the quality of many generated talking head (TH) videos remains a concern, impacting user visual experiences. To tackle this issue, this paper introduces the Talking Head Quality Assessment (THQA) database, featuring 800 TH videos generated through 8 diverse speechdriven methods. Extensive experiments affirm the THQA database’s richness in character and speech features. Subsequent subjective quality assessment experiments analyze correlations between scoring results and speech-driven methods, ages, and genders. In addition, experimental results show that mainstream image and video quality assessment methods have limitations for the THQA database, underscoring the imperative for further research to enhance TH video quality assessment. The THQA database is publicly accessible at https://github.com/zyj-2000/THQA.
Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Zhihua Wang 0002, Xiao-Ping Zhang 0002, Guangtao Zhai
ICIP5
2024 Q-Refine: A Perceptual Quality Refiner for AI-Generated Image
abstract
With the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited optimization capabilities for low-quality AIGIs but also brought negative optimization to high-quality AIGIs. To address this issue, a quality-award refiner named Q-Refine is proposed. Based on the preference of the Human Visual System (HVS), Q-Refine uses the Image Quality Assessment (IQA) metric to guide the refining process for the first time, and modify images of different qualities through three adaptive pipelines. Experimental data shows that for mainstream T2I models, Q-Refine can perform effective optimization to AIGIs of different qualities. It can be a general refiner to optimize AIGIs from both fidelity and aesthetic quality levels, thus expanding the application of the T2I generation models. The code is released on https://github.com/Q-Future/Q-Refine.
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai
ICME8
2024 Optimizing Projection-Based Point Cloud Quality Assessment with Human Preferred Viewpoints Selection
abstract
Viewpoint selection plays a pivotal role in projection-based point cloud quality assessment (PCQA). Generally speaking, sole reliance on a single projection fails to capture adequate quality information, leading to the prevalent use of multi-projection approaches. It is important to recognize that viewpoint selection is significantly influenced by human preferences and viewpoints that align with human predilections exert a greater impact on PCQA. Therefore, we introduce the first viewpoint selection database for PCQA, which comprises 405 distorted point clouds, accompanied by preferred viewpoints collected from humans. Then we propose a novel human preference index, devised from the Visible-Points Ratio and Visible-Color-Entropy Ratio, to guide the selection of viewpoints. Our experimental findings confirm that this human preference index correlates more closely with human preferences than traditional viewpoint selection settings. Moreover, the proposed PCQA method optimized with the human preference index demonstrates competitive performance as well.
Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Weisi Lin, Guangtao Zhai
ICME4
2024 Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
abstract
The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released.
Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ICML12
2024 FS-BAND: A Frequency-Sensitive Banding Detector
abstract
Banding artifact, as known as staircase-like contour, is a common quality annoyance that happens in compression, transmission, etc. scenarios, which largely affects the user’s quality of experience (QoE). The banding distortion typically appears as relatively small pixel-wise variations in smooth backgrounds, which is difficult to analyze in the spatial domain but easily reflected in the frequency domain. In this paper, we thereby study the banding artifact from the frequency aspect and propose a no-reference banding detection model to capture and evaluate banding artifacts, called the Frequency-Sensitive BANding Detector (FS-BAND). The proposed detector is able to generate a pixel-wise banding map with a perception correlated quality score. Experimental results show that the proposed FS-BAND method outperforms state-of-the-art image quality assessment (IQA) approaches with higher accuracy in banding classification task.
Zijian Chen 0001, Wei Sun 0029, Ru Huang 0002, Fangfang Lu, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005
ISCAS6
2024 Calculating Color Differences of Images via Siamese Neural Network
abstract
Recently, the color difference (CD) of standard dynamic range (SDR) images has attracted the attention of researchers. It is worth noting that due to the development of high dynamic range (HDR) image generation technology, the CD of the SDR and HDR images is also worth in-depth research. This is because the HDR image generated from an original SDR image may have changes in color. Some color changes can give people a comfortable impression, but this may also change the information originally expressed in the SDR image. Therefore, this paper researches the CD of the original SDR image and the generated HDR image, and proposes a network to predict the CDs of SDR and HDR image pairs. Specifically, we first build a SDR-HDR image CD dataset. The dataset contains 504 SDR and HDR image pairs, where HDR images are generated from the SDR images using five HDR image generation methods. Second, we propose a siamese neural network to predict the CDs of SDR and HDR image pairs, which consists of three parts: space conversion, feature extraction, and CD calculation. Finally, experiments prove that the proposed network has a superior ability to predict the CDs of SDR and HDR image pairs.
Xiongkuo Min, Xiaohong Liu 0001, Lei Sun 0009, Yonglin Luo, Zuowei Cao, Guangtao Zhai
ISCAS2
2024 PrefIQA: Human Preference Learning for AI-generated Image Quality Assessment
abstract
Despite recent advancements in generative models, the variation in image quality remains a significant concern. To tackle this issue, we propose PrefIQA, an effective human preference learning metric, which can better evaluate the quality of AI-generated images. PrefIQA consists of two units, namely Feature Extraction Unit and Feature Fusion Unit. In Feature Extraction Unit, we introduce a prompt-segmentation module to divide prompts into multiple phrases, enabling a more detailed evaluation of the alignment between images and texts. In Feature Fusion Unit, we introduce a modality-fusion module, which effectively mixes text features and image features to improve the overall performance. In the experiment part, extensive experiments are conducted, demonstrating that PrefIQA surpasses existing text-to-image alignment metrics. We believe that PrefIQA’s proposal would facilitate researches on AI-generated image quality assessment, and make a valuable contribution to the field of text-to-image generation.
Hengjian Gao, Kaiwei Zhang, Wei Sun 0029, Chunyi Li 0001, Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ISCAS7
2024 Multidimensional Similarity Fusion for Speech Quality Assessment
abstract
Since the perceptual quality of audio signals is easily to be affected by compression, transmission, noise adding, etc, it is of great significance to develop an effective audio quality assessment (AQA) method to measure end-user’s quality of experience. In this paper, we propose a full reference AQA model named Multidimensional Similarity Fusion for Audio Quality Assessment (MSF-AQA). We generalize the similarity-based image quality assessment methods for audio, then extract audio similarity features from multiple dimensions, and finally regress the multidimensional similarity features into the final quality score. The experimental results across three databases indicate that our MSF-AQA model outperforms the state-of-the-art AQA methods.
Xiongkuo Min, Yuqin Cao, Xiao-Ping Zhang 0002, Guangtao Zhai
ISCAS2
2024 DSA-QoE: Quality of Experience Evaluation for Streaming Video Based on Dual-Stage Attention
abstract
With the rapid development of streaming media technology, the real-time streaming video Quality of Experience (QoE) assessment has become an important objective for creating new Adaptive Bitrate (ABR) algorithms. The QoE prediction on the client side is challenging considering the sophisticated perception mechanisms of humans, especially the human attention behaviors over time. To address this issue, we propose a learnable model based on the dual-stage attention mechanism to precisely predict continuous QoE which is not covered by most of the current related works. Given the close relationship between the continuous and overall QoE, we use a unified framework to predict these two indices. We have conducted comparison experiments on 6 open datasets, and our model shows superior performance.
Ziheng Jia, Xiongkuo Min, Guangtao Zhai
ISCAS2
2024 Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment
abstract
With the rapid development of generative models, AI-Generated Content (AIGC) has exponentially increased in daily lives. Among them, Text-to-Video (T2V) generation has received widespread attention. Though many T2V models have been released for generating high perceptual quality videos, there is still lack of a method to evaluate the quality of these videos quantitatively. To solve this issue, we establish the largest-scale Text-to-Video Quality Assessment DataBase (T2VQA-DB) to date. The dataset is composed of 10,000 videos generated by 9 different T2V models, along with each video's corresponding mean opinion score. Based on T2VQA-DB, we propose a novel transformer-based model for subjective-aligned Text-to-Video Quality Assessment (T2VQA). The model extracts features from text-video alignment and video fidelity perspectives, then it leverages the ability of a large language model to give the prediction score. Experimental results show that T2VQA outperforms existing T2V metrics and SOTA video quality assessment models. Quantitative analysis indicates that T2VQA is capable of giving subjective-align predictions, validating its effectiveness. The dataset and code are available at https://github.com/QMME/T2VQA.
Tengchuan Kou, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Xiongkuo Min, Guangtao Zhai
ACM Multimedia6
2024 Large Multi-modality Model Assisted AI-Generated Image Quality Assessment
abstract
Traditional deep neural network (DNN)-based image quality assessment (IQA) models leverage convolutional neural networks (CNN) or Transformer to learn the quality-aware feature representation, achieving commendable performance on natural scene images. However, when applied to AI-Generated images (AGIs), these DNN-based IQA models exhibit subpar performance. This situation is largely due to the semantic inaccuracies inherent in certain AGIs caused by uncontrollable nature of the generation process. Thus, the capability to discern semantic content becomes crucial for assessing the quality of AGIs. Traditional DNN-based IQA models, constrained by limited parameter complexity and training data, struggle to capture complex fine-grained semantic features, making it challenging to grasp the existence and coherence of semantic content of the entire image. To address the shortfall in semantic content perception of current IQA models, we introduce a large Multi-modality model Assisted AI-Generated Image Quality Assessment (MA-AGIQA) model, which utilizes semantically informed guidance to sense semantic information and extract semantic vectors through carefully designed text prompts. Moreover, it employs a mixture of experts (MoE) structure to dynamically integrate the semantic information with the quality-aware features extracted by traditional DNN-based IQA models. Comprehensive experiments conducted on two AI-generated content datasets and two traditional IQA datasets show that MA-AGIQA achieves state-of-the-art performance, and demonstrate its superior generalization capabilities on assessing the quality of AGIs. The code is available at https://github.com/wangpuyi/MA-AGIQA.
Puyi Wang, Wei Sun 0029, Jun Jia, Yanwei Jiang, Xiongkuo Min, Guangtao Zhai
ACM Multimedia7
2024 LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM
abstract
Although large multi-modality models (LMMs) have seen extensive exploration and application in various quality assessment studies, their integration into Point Cloud Quality Assessment (PCQA) remains unexplored. Given LMMs' exceptional performance and robustness in low-level vision and quality assessment tasks, this study aims to investigate the feasibility of imparting PCQA knowledge to LMMs through text supervision. To achieve this, we transform quality labels into textual descriptions during the fine-tuning phase, enabling LMMs to derive quality rating logits from 2D projections of point clouds. To compensate for the loss of perception in the 3D domain, structural features are extracted as well. These quality logits and structural features are then combined and regressed into quality scores. Our experimental results affirm the effectiveness of our approach, showcasing a novel integration of LMMs into PCQA that enhances model understanding and assessment accuracy. We hope our contributions can inspire subsequent investigations into the fusion of LMMs with PCQA, fostering advancements in 3D visual quality analysis and beyond. The code is available at https://github.com/zzc-1998/LMM-PCQA.
Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Chaofeng Chen, Xiongkuo Min, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai
ACM Multimedia7
2024 Subjective and Objective Quality-of-Experience Assessment for 3D Talking Heads
abstract
In recent years, immersive communication has emerged as a compelling alternative to traditional video communication methods. One prospective avenue for immersive communication involves augmenting the user's immersive experience through the transmission of three-dimensional (3D) talking heads (THs). However, transmitting 3D THs poses significant challenges due to its complex and voluminous nature, often leading to pronounced distortion and a compromised user experience. Addressing this challenge, we introduce the 3D Talking Heads Quality Assessment (THQA-3D) dataset, comprising 1,000 sets of distorted and 50 original TH mesh sequences (MSs), to facilitate quality assessment in 3D TH transmission. A subjective experiment, characterized by a novel interactive approach, is conducted with recruited participants to assess the quality of MSs in THQA-3D dataset. Leveraging this dataset, we also propose a multimodal Quality-of-Experience (QoE) method incorporating a Large Quality Model (LQM). This method involves frontal projection of MSs and subsequent rendering into videos, with quality assessment facilitated by the LQM and a variable-length video memory filter (VVMF). Additionally, tone-lip coherence and silence detection techniques are employed to characterize audio-visual coherence in 3D MS streams. Experimental evaluation demonstrates the proposed method's superiority, achieving state-of-the-art performance on the THQA-3D dataset and competitiveness on other QoE datasets. Both the THQA-3D dataset and the QoE model have been publicly released at https://github.com/zyj-2000/THQA-3D
Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ACM Multimedia5
2024 GAIA: Rethinking Action Quality Assessment for AI-Generated Videos
abstract
Assessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current action quality assessment (AQA) algorithms predominantly focus on actions from real specific scenarios and are pre-trained with normative action features, thus rendering them inapplicable in AIGVs. To address these problems, we construct GAIA, a Generic AI-generated Action dataset, by conducting a large-scale subjective evaluation from a novel causal reasoning-based perspective, resulting in 971,244 ratings among 9,180 video-action pairs. Based on GAIA, we evaluate a suite of popular text-to-video (T2V) models on their ability to generate visually rational actions, revealing their pros and cons on different categories of actions. We also extend GAIA as a testbed to benchmark the AQA capacity of existing automatic evaluation methods. Results show that traditional AQA methods, action-related metrics in recent T2V benchmarks, and mainstream video quality methods perform poorly with an average SRCC of 0.454, 0.191, and 0.519, respectively, indicating a sizable gap between current models and human action perception patterns in AIGVs. Our findings underscore the significance of action quality as a unique perspective for studying AIGVs and can catalyze progress towards methods with enhanced capacities for AQA in AIGVs.
Zijian Chen 0001, Wei Sun 0029, Yuan Tian 0017, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005
NeurIPS8
2024 Perceptual Skin Tone Color Difference Measurement for Portrait Photography
abstract
In portrait photography, measuring the perceptual color differences (CDs) of skin tone is significant. Many studies have documented that the perception of skin tone is characteristically different from that of other colors. However, most existing CD measures are proposed based on psychophysical data of uniform color patches or natural images, and do not generalize well to the measurement of skin tone. In this paper, we construct the first large-scale portrait dataset for perceptual skin tone CD assessment and conduct psychophysical experiments to collect 160,000 perceptual CD judgments for 40,000 image triplets. Based on this dataset, we propose a deep skin tone CD measure for portrait photography. Extensive experiments demonstrate that our measure substantially outperforms existing CD measures on the problem of assessing skin tone CDs. The constructed dataset and code will be released to facilitate future research.
Shiqi Gao, Huiyu Duan, Qihang Xu, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
VCIP5
2024 ACIQA: A Dataset and Method for Assessing the Imaging Quality of Automotive Cameras
abstract
The imaging quality of automotive cameras is crucial in complex driving environments. Therefore, it is essential to conduct subjective experiments that can realistically reflect drivers’ evaluation of the imaging quality of automotive cameras in real traffic scenarios. To accurately assess the imaging quality of automotive cameras, this paper proposes a no-reference quality assessment method with quality scores that are highly consistent with human subjective perception. Initially, this study constructs a new image quality assessment dataset and then obtains the subjective scores of image quality through subjective experiments. The dataset is constructed by using a variety of realistic props to simulate scene elements that might be captured by an automotive camera and are captured using a wide range of cameras with different sensor types, lens focus, and viewing angles, resulting in a dataset of diverse images. The objective quality assessment method proposed in this paper consists of an object detection network and a multi-branch quality evaluation network. The object detection network is responsible for identifying and classifying scene elements, while the multi-branch quality evaluation network performs feature extraction and score regression on various types of elements to effectively evaluate the imaging quality of the automotive cameras. In the experiments, this no-reference quality assessment method is tested on our built dataset, and the results show that the proposed method exhibits the best performance compared with the state-of-the-art image quality assessment methods.
Haoyang Ni, Kaiwei Zhang, Ziheng Jia, Fangfang Lu, Xiongkuo Min, Guangtao Zhai
VCIP6
2024 End-to-end Prediction of Streaming Video Quality of Experience: Dataset and Approach
abstract
With the rapid development of video-on-demand (VOD) and real-time streaming video technologies, the accurate objective assessment of streaming video Quality of Experience (QoE) has become a focal point for optimizing streaming-related technologies. However, due to the inherent transmission distortions caused by poor Quality of Service (QoS) conditions in streaming videos, such as intermittent stalling, rebuffering, and drastic changes in video sharpness due to bitrate fluctuations, evaluating streaming video QoE presents numerous challenges. This paper introduces a large and diverse in-the-wild streaming video QoE evaluation dataset - the SJLIVE-1k dataset. This work addresses the limitations of corresponding datasets, which lack in-the-wild video sequences under real network conditions and whose amount of video content is insufficient. Furthermore, we propose an end-to-end objective QoE evaluation strategy that extracts video content and QoS features from the video itself without using any extra information. By implementing self-supervised contrastive learning as the "reminder" to bridge the gap between the different types of features, our approach achieves state-of-the-art results across three datasets. Our proposed dataset will be released to facilitate further research.
Ziheng Jia, Xiongkuo Min, Guangtao Zhai
VCIP2
2024 ReLI-QA: A Multidimensional Quality Assessment Dataset for Relighted Human Heads
abstract
Lighting conditions significantly affect the quality of both real and AI-generated images. Facial images are particularly sensitive to lighting due to their detailed nature and the importance of facial features in conveying identity. Poor lighting can easily obscure these critical details. To address this issue, various portrait relighting methods have been developed to adjust the lighting in improperly exposed images. However, these methods often encounter challenges such as overexposure, underexposure, and detail loss in the relighted portraits. Consequently, there is a need for effective quality assessment and control of relighted human heads (RHHs). In this study, one proposed simple baseline and three typical relighting methods are applied to six selected human head (HH) images, resulting in the creation of a quality assessment dataset named ReLI-QA, which comprises 840 RHHs. A multidimensional subjective quality assessment method based on visual guidance is proposed to accurately evaluate the visual quality of each RHH in the dataset. By analyzing the results of subjective experiments, the quality of RHHs is shown to be affected by multiple factors. Finally, based on ReLI-QA, some typical image quality assessment (IQA) methods are selected for benchmark experiments. The experimental results show the limitations of the existing methods in RHH quality assessment. The dataset and code for this research has been released at https://github.com/zyj-2000/ReLI-QA.
Yingjie Zhou 0003, Farong Wen, Jun Jia, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
VCIP5
2024 Perceptual video quality assessment: a survey
abstract
Abstract Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the advancement of Internet communication and cloud service technology, video content and traffic are growing exponentially, which further emphasizes the requirement for accurate and rapid assessment of video quality. Therefore, numerous subjective and objective video quality assessment studies have been conducted over the past two decades for both generic videos and specific videos such as streaming, user-generated content, 3D, virtual and augmented reality, high dynamic range, high frame rate, audio-visual, etc. This survey provides an up-to-date and comprehensive review of these video quality assessment studies. Specifically, we first review the subjective video quality assessment methodologies and databases, which are necessary for validating the performance of video quality metrics. Second, the objective video quality assessment measures for general purposes are categorized and surveyed according to the methodologies utilized in the quality measures. Third, we overview the objective video quality assessment measures for specific applications and emerging topics. Finally, the performance of the state-of-the-art video quality assessment measures is compared and analyzed. This survey provides a systematic overview of both classical works and recent progress in the realm of video quality assessment, which can help other researchers quickly access the field and conduct relevant research.
Xiongkuo Min, Huiyu Duan, Wei Sun 0029, Yucheng Zhu, Guangtao Zhai
Sci. China Inf. Sci.1
2024 Boosting power line inspection in bad weather: Removing weather noise with channel-spatial attention-based UNet
Yaocheng Li, Qinglin Qian, Huiyu Duan, Xiongkuo Min, Yongpeng Xu, Xiuchen Jiang
Multim. Tools Appl.4
2024 Analysis of Video Quality Datasets via Design of Minimalistic Video Quality Models
abstract
Blind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA models have been measured primarily on a few human-rated VQA datasets. Thus, it is crucial to gain a better understanding of existing VQA datasets in order to properly evaluate the current progress in BVQA. Towards this goal, we conduct a first-of-its-kind computational analysis of VQA datasets via designing minimalistic BVQA models. By minimalistic, we restrict our family of BVQA models to build only upon basic blocks: a video preprocessor (for aggressive spatiotemporal downsampling), a spatial quality analyzer, an optional temporal quality analyzer, and a quality regressor, all with the simplest possible instantiations. By comparing the quality prediction performance of different model variants on eight VQA datasets with realistic distortions, we find that nearly all datasets suffer from the easy dataset problem of varying severity, some of which even admit blind image quality assessment (BIQA) solutions. We additionally justify our claims by comparing our model generalization capabilities on these VQA datasets, and by ablating a dizzying set of BVQA design choices related to the basic building blocks. Our results cast doubt on the current progress in BVQA, and meanwhile shed light on good practices of constructing next-generation VQA datasets and models.
Wei Sun 0029, Wen Wen 0007, Xiongkuo Min, Long Lan, Guangtao Zhai, Kede Ma
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 BAND-2k: Banding Artifact Noticeable Database for Banding Detection and Quality Assessment
abstract
Banding, also known as staircase-like contours, frequently occurs in flat areas of images/videos processed by compression or quantization algorithms. As undesirable artifacts, banding destroys the original image structure, thus inevitably degrading users’ quality of experience (QoE). In this paper, we systematically investigate the banding image quality assessment (IQA) problem, aiming to detect the image banding artifacts and evaluate their perceptual visual quality. Considering that the existing image banding databases only contain limited content sources and banding generation methods, and lack perceptual quality labels (i.e. mean opinion scores), we first build the largest banding IQA database so far, namedBanding Artifact Noticeable Database (BAND-2k), which consists of 2,000 banding images generated by 15 compression and quantization schemes. A total of 23 workers participated in the subjective IQA experiment, yielding over 214,000 patch-level banding class labels and 44,371 reliable image-level quality rating scores. Subsequently, we develop an effective no-reference (NR) banding evaluator for banding detection and quality assessment by leveraging frequency characteristics of banding artifacts. To be more specific, a dual convolutional neural network (CNN) is employed to concurrently learn the feature representation from the high-frequency and low-frequency maps, thereby enhancing the ability to discern banding artifacts. The quality score of a banding image is generated by pooling the banding detection maps masked by the spatial frequency filters. The experimental results demonstrate that our banding evaluator achieves remarkably high accuracy in banding detection and also exhibits high SRCC and PLCC results with the perceptual quality labels, even without directly learning a regression model for banding quality evaluation. These findings unveil the strong correlations between the intensity of banding artifacts and the perceptual visual quality, thus validating the necessity of banding quality assessment. The BAND-2k database and the proposed banding evaluator are available at https://github.com/zijianchen98/BAND-2k.
Zijian Chen 0001, Wei Sun 0029, Jun Jia, Fangfang Lu, Jing Liu 0002, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.8
2024 Blind Image Quality Assessment: A Fuzzy Neural Network for Opinion Score Distribution Prediction
abstract
Image quality assessment (IQA) has always been a popular research topic. There have been many methods proposed for predicting image quality, also known as the mean opinion score (MOS). However, it is worth noting that different people may assign different opinion scores to the same image. Image quality described by all subjective opinion scores can express rich subjective information about the image, such as diversity and uncertainty, which cannot be accurately described by a single MOS. Therefore, this paper proposes a fuzzy neural network to predict the opinion score distribution (OSD) of image quality. The fuzzy neural network includes three sub-networks: a feature extraction network, a feature fuzzification network, and a fuzzy learning network. First, a novel network is designed to extract image features. The extracted features are then fuzzified by fuzzy theory to model the epistemic uncertainty in the feature extraction process. Finally, the OSD of image quality is predicted using the fuzzy learning network by learning the mapping from fuzzy features to fuzzy uncertainty when rating image quality. In addition, to train the proposed fuzzy neural network, we employ a new loss function based on the quantile and the cumulative density function. We experimentally validate the feasibility and superiority of the proposed method in two aspects. On the one hand, we demonstrate the performance of the proposed method in predicting the OSD of image quality on the SJTU IQSD and KonIQ-10K databases. On the other hand, we also prove the feasibility of the proposed method in predicting the MOS of image quality on several popular IQA databases, including CSIQ, TID2013, LIVE MD, and LIVE Challenge.
Xiongkuo Min, Yucheng Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2024 Continuous and Overall Quality of Experience Evaluation for Streaming Video Based on Rich Features Exploration and Dual-Stage Attention
abstract
With the rapid development of streaming media technology, the Quality of Experience (QoE) of streaming videos becomes crucial to optimize the video compression and transmission algorithms, such as adaptive bitrate (ABR). However, the complexity of human perceptual mechanisms, particularly in relation to temporal distortions, poses substantial challenges to effective QoE monitoring. In recent years, many efforts in video quality assessment (VQA) and video QoE evaluation have highlighted the influence of a broad spectrum of features—from Quality of Service (QoS) metrics to video content understanding—on viewer experience. On this basis, we believe that there is also a dynamic relationship among these features varying with the broadcasting content. Furthermore, research indicates a significant correlation between real-time and retrospective assessments of QoE for individual videos. In response to these insights, we introduce a novel approach leveraging a unified learnable network that incorporates dual-stage attention, the temporal and cross-feature attention, to accurately predict both continuous and overall QoE for streaming videos. The results of experiments conducted on several publicly available databases demonstrate the superiority of our proposed method over the state-of-the-art metrics.
Ziheng Jia, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2024 AGIQA-3K: An Open Database for AI-Generated Image Quality Assessment
abstract
With the rapid advancements of the text-to-image generative model, AI-generated images (AGIs) have been widely applied to entertainment, education, social media, etc. However, considering the large quality variance among different AGIs, there is an urgent need for quality models that are consistent with human subjective ratings. To address this issue, we extensively consider various popular AGI models, generated AGI through different prompts and model parameters, and collected subjective scores at the perceptual quality and text-to-image alignment, thus building the most comprehensive AGI subjective quality database AGIQA-3K so far. Furthermore, we conduct a benchmark experiment on this database to evaluate the consistency between the current Image Quality Assessment (IQA) model and human perception, while proposing StairReward that significantly improves the assessment performance of subjective text-to-image alignment. We believe that the fine-grained subjective scores in AGIQA-3K will inspire subsequent AGI quality models to fit human subjective perception mechanisms at both perception and alignment levels and to optimize the generation result of future AGI models. The database is released on https://github.com/lcysyzxdxc/AGIQA-3k-Database.
Chunyi Li 0001, Haoning Wu 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.5
2024 Un-Gaze: A Unified Transformer for Joint Gaze-Location and Gaze-Object Detection
abstract
This paper proposes an efficient and effective method for joint gaze location detection (GL-D) and gaze object detection (GO-D), i.e., gaze following detection. Current approaches frame GL-D and GO-D as two separate tasks, employing a multi-stage framework where human head crops must first be detected and then be fed into a subsequent GL-D sub-network, which is further followed by an additional object detector for GO-D. In contrast, we reframe the gaze following detection task as detecting human head locations and their gaze followings simultaneously, aiming at jointly detect human gaze location and gaze object in a unified and single-stage pipeline. To this end, we propose GTR, short for Gaze following detection TRansformer, streamlining the gaze following detection pipeline by eliminating all additional components, leading to the first unified paradigm that unites GL-D and GO-D in a fully end-to-end manner. GTR enables an iterative interaction between holistic semantics and human head features through a hierarchical structure, inferring the relations of salient objects and human gaze from the global image context and resulting in an impressive accuracy. Concretely, GTR achieves a 12.1 mAP gain ($\mathbf {25.1}\%$) on GazeFollowing and a 18.2 mAP gain ($\mathbf {43.3\%}$) on VideoAttentionTarget for GL-D, as well as a 19 mAP improvement ($\mathbf {45.2\%}$) on GOO-Real for GO-D. Meanwhile, unlike existing systems detecting gaze following sequentially due to the need for a human head as input, GTR has the flexibility to comprehend any number of people’s gaze followings simultaneously, resulting in high efficiency. Specifically, GTR introduces over a$\times 9$improvement in FPS and the relative gap becomes more pronounced as the human number grows.
Danyang Tu, Wei Shen 0002, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.4
2024 Synergetic Assessment of Quality and Aesthetic: Approach and Comprehensive Benchmark Dataset
abstract
Quantifications of image quality and aesthetic have been regarded as two independent fields in computer vision. Generally, image quality assessment aims at measuring image distortions and image aesthetic is judged by commonly established photography rules. However, either measuring image quality or aesthetic alone is not sufficient to qualitatively rank images. Therefore, this paper puts forward the synergetic assessment of quality and aesthetic to help understand the subjective human preferences of digital pictures more comprehensively. Specifically, considering that the images of existing benchmark datasets are only labeled with single attribute, we first establish a new dataset which contains 9042 real-world images with the corresponding human rated pair-wise quality-aesthetic scores. Previously, these images are only labeled with aesthetic score, and we evaluate the subjective quality score of them, so that it can make up the lack of image dataset with double attributes. Moreover, since the existing methods are mostly designed for individual attribute prediction. We then propose a two-stream learning network to assess both quality and aesthetic of images in parallel. This network follows the top-down perception mechanism which learns from both fined grained details and holistic image layout simultaneously. Furthermore, we introduce a Channel-Diversity loss, which can be deployed in grouped convolution operation, and can constrain channels to be mutually exclusive across the spatial dimensions. To some extent, this contributes to spotlight different local discriminative regions with a finer granularity. Finally, experiments demonstrate that our method outperforms the state-of-the-art methods on our established benchmark dataset and other benchmark datasets in terms of image quality and aesthetic assessment. We hope this paper could serve as a potent reference and be useful for future research on the study of image ranking. Both the benchmark dataset and the code will be publicly available to facilitate further research.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Zhongpai Gao, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.3
2024 How is Visual Attention Influenced by Text Guidance? Database and Model
abstract
The analysis and prediction of visual attention have long been crucial tasks in the fields of computer vision and image processing. In practical applications, images are generally accompanied by various text descriptions, however, few studies have explored the influence of text descriptions on visual attention, let alone developed visual saliency prediction models considering text guidance. In this paper, we conduct a comprehensive study on text-guided image saliency (TIS) from both subjective and objective perspectives. Specifically, we construct a TIS database named SJTU-TIS, which includes 1200 text-image pairs and the corresponding collected eye-tracking data. Based on the established SJTU-TIS database, we analyze the influence of various text descriptions on visual attention. Then, to facilitate the development of saliency prediction models considering text influence, we construct a benchmark for the established SJTU-TIS database using state-of-the-art saliency models. Finally, considering the effect of text descriptions on visual attention, while most existing saliency models ignore this impact, we further propose a text-guided saliency (TGSal) prediction model, which extracts and integrates both image features and text features to predict the image saliency under various text-description conditions. Our proposed model significantly outperforms the state-of-the-art saliency models on both the SJTU-TIS database and the pure image saliency databases in terms of various evaluation metrics. The SJTU-TIS database and the code of the proposed TGSal model will be released at: https://github.com/IntMeGroup/TGSal.
Xiongkuo Min, Huiyu Duan, Guangtao Zhai
IEEE Trans. Image Process.2
2024 Pixel-Learnable 3DLUT With Saturation-Aware Compensation for Image Enhancement
abstract
The 3D Lookup Table (3DLUT)-based methods are gaining popularity due to their satisfactory and stable performance in achieving automatic and adaptive real time image enhancement. In this paper, we present a new solution to the intractability in handling continuous color transformations of 3DLUT due to the lookup via three independent color channel coordinates in RGB space. Inspired by the inherent merits of the HSV color space, we separately enhance image intensity and color composition. The Transformer-based Pixel-Learnable 3D Lookup Table is proposed to undermine contouring artifacts, which enhances images in a pixel-wise manner with non-local information to emphasize the diverse spatially variant context. In addition, noticing the underestimation of composition color component, we develop the Saturation-Aware Compensation (SAC) module to enhance the under-saturated region determined by an adaptive SA map with Saturation-Interaction block, achieving well balance between preserving details and color rendition. Our approach can be applied to image retouching and tone mapping tasks with fairly good generality, especially in restoring localized regions with weak visibility. The performance in both theoretical analysis and comparative experiments manifests that the proposed solution is effective and robust.
Jing Liu 0002, Xiongkuo Min, Yuting Su 0001, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.3
2024 Unified Audio-Visual Saliency Model for Omnidirectional Videos With Spatial Audio
abstract
Spatial audio is a crucial component of omnidirectional videos (ODVs), which can provide an immersive experience by enabling viewers to perceive sound sources in all directions. However, most visual attention modeling works for ODVs focus only on visual cues, and audio modality is rather rarely considered. Additionally, the existing audio-visual saliency models for ODVs lack spatial audio location-awareness (i.e. sound source location-agnostic) and audio content attributes discriminability (i.e. audio content attributes-agnostic). To this end, we propose a novel audio-visual perception saliency (AVPS) model with spatial audio location-awareness and audio content attributes-adaptive to efficiently address the problem of fixation prediction in ODVs. Specifically, we first utilize the improved group equivariant convolutional neural network (G-CNN) with eidetic 3D LSTM (E3D-LSTM) to extract spatial-temporal visual features. Then we perceive sound source locations by computing the audio energy map (AEM) of the audio information in ODVs. Subsequently, we introduce SoundNet to extract audio features with multiple attributes. Finally, we develop an audio-visual feature fusion module to adaptively integrate spatial-temporal visual features and spatial auditory information to generate the final audio-visual saliency map. Extensive experiments in three audio modalities validate the effectiveness of the proposed model. Meanwhile, the performance of the proposed model is superior to the other 10 state-of-the-art saliency models.
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.5
2024 Hidden Barcode in Sub-Images with Invisible Locating Marker
abstract
The prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness.
Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.7
2024 GMS-3DQA: Projection-Based Grid Mini-patch Sampling for 3D Model Quality Assessment
abstract
Nowadays, most three-dimensional model quality assessment (3DQA) methods have been aimed at improving accuracy. However, little attention has been paid to the computational cost and inference time required for practical applications. Model-based 3DQA methods extract features directly from the 3D models, which are characterized by their high degree of complexity. As a result, many researchers are inclined towards utilizing projection-based 3DQA methods. Nevertheless, previous projection-based 3DQA methods directly extract features from multi-projections to ensure quality prediction accuracy, which calls for more resource consumption and inevitably leads to inefficiency. Thus, in this article, we address this challenge by proposing a no-reference (NR) projection-based G rid M ini-patch S ampling 3D Model Q uality A ssessment (GMS-3DQA) method. The projection images are rendered from six perpendicular viewpoints of the 3D model to cover sufficient quality information. To reduce redundancy and inference resources, we propose a multi-projection grid mini-patch sampling strategy (MP-GMS), which samples grid mini-patches from the multi-projections and forms the sampled grid mini-patches into one quality mini-patch map (QMM). The Swin-Transformer tiny backbone is then used to extract quality-aware features from the QMMs. The experimental results show that the proposed GMS-3DQA outperforms existing state-of-the-art NR-3DQA methods on the point cloud quality assessment databases for both accuracy and efficiency. The efficiency analysis reveals that the proposed GMS-3DQA requires far less computational resources and inference time than other 3DQA competitors. The code is available at https://github.com/zzc-1998/GMS-3DQA .
Wei Sun 0029, Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Zijian Chen 0001, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Subjective and Objective Quality Assessment for in-the-Wild Computer Graphics Images
abstract
Computer graphics images (CGIs) are artificially generated by means of computer programs and are widely perceived under various scenarios, such as games, streaming media, etc. In practice, the quality of CGIs consistently suffers from poor rendering during production, inevitable compression artifacts during the transmission of multimedia applications, and low aesthetic quality resulting from poor composition and design. However, few works have been dedicated to dealing with the challenge of computer graphics image quality assessment (CGIQA). Most image quality assessment (IQA) metrics are developed for natural scene images (NSIs) and validated on databases consisting of NSIs with synthetic distortions, which are not suitable for in-the-wild CGIs. To bridge the gap between evaluating the quality of NSIs and CGIs, we construct a large-scale in-the-wild CGIQA database consisting of 6,000 CGIs (CGIQA-6k) and carry out the subjective experiment in a well-controlled laboratory environment to obtain the accurate perceptual ratings of the CGIs. Then, we propose an effective deep learning–based no-reference (NR) IQA model by utilizing both distortion and aesthetic quality representation. Experimental results show that the proposed method outperforms all other state-of-the-art NR IQA methods on the constructed CGIQA-6k database and other CGIQA-related databases. The database is released at https://github.com/zzc-1998/CGIQA6K .
Wei Sun 0029, Yingjie Zhou 0003, Jun Jia, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.7
2023 MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos
abstract
User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing of UGC live videos, effective video quality assessment (VQA) tools are needed to monitor and perceptually optimize live streaming videos in the distributing process. In this paper, we address UGC Live VQA problems by constructing a first-of-a-kind subjective UGC Live VQA database and developing an effective evaluation tool. Concretely, 418 source UGC videos are collected in real live streaming scenarios and 3,762 compressed ones at different bit rates are generated for the subsequent subjective VQA experiments. Based on the built database, we develop a Multi-12imensional VQA (MD-VQA) evaluator to measure the visual quality of UGC live videos from semantic, distortion, and motion aspects respectively. Extensive experimental results show that MD-VQA achieves state-of-the-art performance on both our UGC Live VQA database and existing compressed UGC VQA databases.
Wei Wu 0002, Wei Sun 0029, Danyang Tu, Wei Lu 0021, Xiongkuo Min, Ying Chen 0011, Guangtao Zhai
CVPR6
2023 Perceptual Quality Assessment for Digital Human Heads
abstract
Digital humans are attracting more and more research interest during the last decade, the generation, representation, rendering, and animation of which have been put into large amounts of effort. However, the quality assessment of digital humans has fallen behind. Therefore, to tackle the challenge of digital human quality assessment issues, we propose the first large-scale quality assessment database for three-dimensional (3D) scanned digital human heads (DHHs). The constructed database consists of 55 reference DHHs and 1,540 distorted DHHs along with the subjective perceptual ratings. Then, a simple yet effective full-reference (FR) projection-based method is proposed to evaluate the visual quality of DHHs. The pretrained Swin Transformer tiny is employed for hierarchical feature extraction and the multi-head attention module is utilized for feature fusion. The experimental results reveal that the proposed method exhibits state-of-the-art performance among the mainstream FR metrics. The database is released at https://github.com/zzc-1998/DHHQA.
Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
ICASSP4
2023 Audio-Visual Saliency for Omnidirectional Videos
Xilei Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Li Chen 0021, Xiongkuo Min, Guangtao Zhai
ICIG (5)8
2023 Audio-Visual Quality Assessment for User Generated Content: Database and Method
abstract
With the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users’ Quality of Experience (QoE). However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the user’s QoE also depends on the accompanying audio signals. In this paper, we conduct the first study to address the problem of UGC audio and video quality assessment (AVQA). Specifically, we construct the first UGC AVQA database named the SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences, and conduct a user study to obtain the mean opinion scores of the A/V sequences. The content of the SJTU-UAV database is then analyzed from both the audio and video aspects to show the database characteristics. We also design a family of AVQA models, which fuse the popular VQA methods and audio features via support vector regressor (SVR). We validate the effectiveness of the proposed models on the three databases. The experimental results show that with the help of audio signals, the VQA models can evaluate the perceptual quality more accurately. The database will be released to facilitate further research.
Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Xiao-Ping Zhang 0002, Guangtao Zhai
ICIP2
2023 Geometry-Aware Video Quality Assessment for Dynamic Digital Human
abstract
Dynamic Digital Humans (DDHs) are 3D digital models that are animated using predefined motions and are inevitably bothered by noise/shift during the generation process and compression distortion during the transmission process, which needs to be perceptually evaluated. Usually, DDHs are displayed as 2D rendered animation videos and it is natural to adapt video quality assessment (VQA) methods to DDH quality assessment (DDH-QA) tasks. However, the VQA methods are highly dependent on viewpoints and less sensitive to geometry-based distortions. Therefore, in this paper, we propose a novel no-reference (NR) geometry-aware video quality assessment method for DDH-QA challenge. Geometry characteristics are described by the statistical parameters estimated from the DDHs’ geometry attribute distributions. Spatial and temporal features are acquired from the rendered videos. Finally, all kinds of features are integrated and regressed into quality values. Experimental results show that the proposed method achieves state-of-the-art performance on the DDH-QA database.
Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
ICIP4
2023 A No-Reference Quality Assessment Method for Digital Human Head
abstract
In recent years, digital humans have been widely applied in augmented/virtual reality (A/VR), where viewers are allowed to freely observe and interact with the volumetric content. However, the digital humans may be degraded with various distortions during the procedure of generation and transmission. Moreover, little effort has been put into the perceptual quality assessment of digital humans. Therefore, it is urgent to carry out objective quality assessment methods to tackle the challenge of digital human quality assessment (DHQA). In this paper, we develop a novel no-reference (NR) method based on Transformer to deal with DHQA in a multi-task manner. Specifically, the front 2D projections of the digital humans are rendered as inputs and the vision transformer (ViT) is employed for the feature extraction. Then we design a multi-task module to jointly classify the distortion types and predict the perceptual quality levels of digital humans. The experimental results show that the proposed method well correlates with the subjective ratings and outperforms the state-of-the-art quality assessment methods.
Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Xianghe Ma, Guangtao Zhai
ICIP4
2023 BH-VQA: Blind High Frame Rate Video Quality Assessment
abstract
High frame rate (HFR) videos can provide consumers with a more immersive viewing experience in motion-rich scenes. However, they also pose a great challenge for video compression and transmission due to the increase in frame rates. Therefore, it is very important to choose proper frame rates and bit rates to achieve a trade-off between transmission bandwidth and visual quality. In this paper, we propose a novel Blind HFR Video Quality Assessment (BH-VQA) model by exploring the efficient and effective motion representation from the deep neural network (DNN). Concretely, we first train a baseline VQA model (i.e. a backbone network and a regressor) on a large-scale VQA database to derive a powerful quality-aware feature extractor for the spatial and motion feature extraction. Then, the HFR video is split into a sequence of video clips and the spatial features of each video clip are extracted just using the first frame of the video clip. To capture temporal distortions caused by frame rate variations and object and camera motion, we calculate deep structural similarities between continuous frames of each video clip as the motion features. Finally, the temporal quality dependencies between video clips are learned through a gated recurrent unit (GRU) network to obtain the perceptual video quality score. Experimental results show that BH-VQA achieves the best performance on two publicly available HFR VQA databases. The code of BH-VQA will be released.
Wei Lu 0021, Wei Sun 0029, Danyang Tu, Xiongkuo Min, Guangtao Zhai
ICME5
2023 EEP-3DQA: Efficient and Effective Projection-Based 3D Model Quality Assessment
abstract
Currently, great numbers of efforts have been put into improving the effectiveness of 3D model quality assessment (3DQA) methods. However, little attention has been paid to the computational costs and inference time, which is also important for practical applications. Unlike 2D media, 3D models are represented by more complicated and irregular digital formats, such as point cloud and mesh. Thus it is normally difficult to perform an efficient module to extract quality-aware features of 3D models. In this paper, we address this problem from the aspect of projection-based 3DQA and develop a no-reference (NR) Efficient and Effective Projection-based 3D Model Quality Assessment (EEP-3DQA) method. The input projection images of EEP-3DQA are randomly sampled from the six perpendicular viewpoints of the 3D model and are further spatially downsampled by the grid-mini patch sampling strategy. Further, the lightweight Swin-Transformer tiny is utilized as the backbone to extract the quality-aware features. Finally, the proposed EEP-3DQA and EEP-3DQA-t (tiny version) achieve the best performance than the existing state-of-the-art NR-3DQA methods and even outperforms most full-reference (FR) 3DQA methods on the point cloud and mesh quality assessment databases while consuming less inference time than the compared 3DQA methods.
Wei Sun 0029, Yingjie Zhou 0003, Wei Lu 0021, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai
ICME6
2023 DDH-QA: A Dynamic Digital Humans Quality Assessment Database
abstract
In recent years, large amounts of effort have been put into pushing forward the real-world application of dynamic digital human (DDH). However, most current quality assessment research focuses on evaluating static 3D models and usually ignores motion distortions. Therefore, in this paper, we construct a large-scale dynamic digital human quality assessment (DDH-QA) database with diverse motion content as well as multiple distortions to comprehensively study the perceptual quality of DDHs. Both model-based distortion (noise, compression) and motion-based distortion (binding error, motion unnaturalness) are taken into consideration. Ten types of common motion are employed to drive the DDHs and a total of 800 DDHs are generated in the end. Afterward, we render the video sequences of the distorted DDHs as the evaluation media and carry out a well-controlled subjective experiment. Then a benchmark experiment is conducted with the state-of-the-art video quality assessment (VQA) methods and the experimental results show that existing VQA methods are limited in assessing the perceptual loss of DDHs. The database is available at https://github.com/zzc-1998/DDH-QA.
Yingjie Zhou 0003, Wei Sun 0029, Wei Lu 0021, Xiongkuo Min, Yu Wang 0002, Guangtao Zhai
ICME5
2023 MM-PCQA: Multi-Modal Learning for No-reference Point Cloud Quality Assessment
abstract
The visual quality of point clouds has been greatly emphasized since the ever-increasing 3D vision applications are expected to provide cost-effective and high-quality experiences for users. Looking back on the development of point cloud quality assessment (PCQA), the visual quality is usually evaluated by utilizing single-modal information, i.e., either extracted from the 2D projections or 3D point cloud. The 2D projections contain rich texture and semantic information but are highly dependent on viewpoints, while the 3D point clouds are more sensitive to geometry distortions and invariant to viewpoints. Therefore, to leverage the advantages of both point cloud and projected image modalities, we propose a novel no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA) metric. In specific, we split the point clouds into sub-models to represent local geometry distortions such as point shift and down-sampling. Then we render the point clouds into 2D image projections for texture feature extraction. To achieve the goals, the sub-models and projected images are encoded with point-based and image-based neural networks. Finally, symmetric cross-modal attention is employed to fuse multi-modal quality-aware information. Experimental results show that our approach outperforms all compared state-of-the-art methods and is far ahead of previous no-reference PCQA methods, which highlights the effectiveness of the proposed method. The code is available at https://github.com/zzc-1998/MM-PCQA.
Wei Sun 0029, Xiongkuo Min, Qiyuan Wang 0002, Guangtao Zhai
IJCAI3
2023 The Influence of Text-guidance on Visual Attention
abstract
Visual attention analysis and prediction have long been important tasks in computer vision and image processing. However, images often come along with various text descriptions in real applications, while the influence of these text-guidances on the visual saliency of corresponding images have rarely been studied. Therefore, in this paper, we mainly focus on the problem of whether and how the text-guidance influences the visual attention during image viewing, and perform subjective experiments, qualitative and quantitative comparisons as well as model evaluations on this new task. Specifically, we first conduct eye tracking experiments on 300 images under text-visual (TV) and visual (V) test conditions, respectively. Based on the subjective experiments, we perform qualitative and quantitative comparisons between the visual attention data collected under TV and V conditions, and conclude that the text-guidance can significantly influence the visual attention, especially when the text-described target is a non-salient object. Finally, we evaluate the existing saliency models on our database, and find that existing models cannot well handle this text-induced saliency prediction task. Our constructed database will be publicly available to facilitate future research.
Xiongkuo Min, Huiyu Duan, Guangtao Zhai
ISCAS2
2023 StableVQA: A Deep No-Reference Quality Assessment Model for Video Stability
abstract
Video shakiness is an unpleasant distortion of User Generated Content (UGC) videos, which is usually caused by the unstable hold of cameras. In recent years, many video stabilization algorithms have been proposed, yet no specific and accurate metric enables comprehensively evaluating the stability of videos. Indeed, most existing quality assessment models evaluate video quality as a whole without specifically taking the subjective experience of video stability into consideration. Therefore, these models cannot measure the video stability explicitly and precisely when severe shakes are present. In addition, there is no large-scale video database in public that includes various degrees of shaky videos with the corresponding subjective scores available, which hinders the development of Video Quality Assessment for Stability (VQA-S). To this end, we build a new database named StableDB that contains 1,952 diversely-shaky UGC videos, where each video has a Mean Opinion Score (MOS) on the degree of video stability rated by 34 subjects. Moreover, we elaborately design a novel VQA-S model named StableVQA, which consists of three feature extractors to acquire the optical flow, semantic, and blur features respectively, and a regression layer to predict the final stability score. Extensive experiments demonstrate that the StableVQA achieves a higher correlation with subjective opinions than the existing VQA-S models and generic VQA models. The database and codes are available at https://github.com/QMME/StableVQA.
Tengchuan Kou, Xiaohong Liu 0001, Wei Sun 0029, Jun Jia, Xiongkuo Min, Guangtao Zhai
ACM Multimedia5
2023 Perceptual Quality Assessment for Video Frame Interpolation
abstract
The quality of frames is significant for both research and application of video frame interpolation (VFI). In recent VFI studies, the methods of full-reference image quality assessment have generally been used to evaluate the quality of VFI frames. However, high frame rate reference videos, necessities for the full-reference methods, are difficult to obtain in most applications of VFI. To evaluate the quality of VFI frames without reference videos, a no-reference perceptual quality assessment method is proposed in this paper. This method is more compatible with VFI application and the evaluation scores from it are consistent with human subjective opinions. A new quality assessment dataset for VFI was constructed through subjective experiments firstly, to assess the opinion scores of interpolated frames. The dataset was created from triplets of frames extracted from high-quality videos using 9 state-of-the-art VFI algorithms. The proposed method evaluates the perceptual coherence of frames incorporating the original pair of VFI inputs. Specifically, the method applies a triplet network architecture, including three parallel feature pipelines, to extract the deep perceptual features of the interpolated frame as well as the original pair of frames. Coherence similarities of the two-way parallel features are jointly calculated and optimized as a perceptual metric. In the experiments, both full-reference and no-reference quality assessment methods were tested on the new quality dataset. The results show that the proposed method achieves the best performance among all compared quality assessment methods on the dataset.
Jinliang Han, Xiongkuo Min, Jun Jia, Lei Sun 0009, Zuowei Cao, Yonglin Luo, Guangtao Zhai
VCIP2
2023 Split-Conv: A Resource-efficient Compression Method for Image Quality Assessment Models
abstract
Blind Image Quality Assessment (BIQA) models based on deep neural networks (DNNs) have achieved state-of-the-art performance recently. However, the heavyweight architecture makes them hard to deploy on resource-constrained devices. Filter pruning is one of the most effective ways to compress the DNN model. However, most pruning methods need complete retraining after pruning which is too resource-consuming for IQA models with complex training processes. In this paper, we propose a resource-efficient structural pruning method for IQA models called Split-Conv, where the model only needs to be retrained on small IQA databases. Specifically, we split convolutional layers into sub-convolutional kernels and decorators, which can measure the importance of convolutional channels more precisely, thus reducing the requirement for retraining conditions. The experiments on several IQA models demonstrate the effectiveness of our IQA model pruning approach. For instance, DBCNN, StairIQA, and HyperIQA’s performances can be kept at the baseline level when 50% parameters are pruned with the model being retrained on the authentically distort IQA database containing only 586 images, and the performances of StairIQA and HyperIQA on KonIQ-10k database are still acceptable after 80% of the parameters have been pruned. Besides, due to Split-Conv’s capability to identify the important, useless, and harmful filters, the performances of some BIQA models can even be boosted after model compression with a proper pruning ratio.
Xiongkuo Min, Guangtao Zhai
VCIP2
2023 Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning
Dandan Zhu 0001, Xuan Shao, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
Appl. Intell.4
2023 Human attention based movie summarization: Dataset and baseline model
abstract
A movie summarization model can automatically edit a condensed version of a movie by selecting keyframes . Some previous works have proposed some movie summarizers based on traditional methods or recent neural networks and achieved some progress. Despite the demonstrated successes, there are some limitations: (1) previous works mainly resort to hand-crafted heuristics and most of them are unsupervised; (2) currently there is no publicly suitable dataset available for the supervised movie summarization; (3) existing works only focus on the movies themselves while neglecting the audiences, who have the most to say in which part of the movie is more attractive. To break through the aforementioned limitations, we establish a movie summarization dataset Movie50 and propose a novel human attention based annotation pipeline. Furthermore, we propose the A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better simulate human attention as well as exploit more plentiful information. The network is designed, trained end-to-end, and evaluated on the public dataset and our dataset. Extensive experiments demonstrate the superiority of the proposed method.
Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
Neurocomputing3
2023 Decoupled dynamic group equivariant filter for saliency prediction on omnidirectional image
Dandan Zhu 0001, Kaiwei Zhang, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
Neurocomputing5
2023 Perceptual quality assessment for fine-grained compressed images
Wei Sun 0029, Wei Wu 0002, Xiongkuo Min, Guangtao Zhai
J. Vis. Commun. Image Represent.5
2023 A Deep Learning-Based Multidimensional Aesthetic Quality Assessment Method for Mobile Game Images
abstract
Mobile games have played an increasingly significant role in people's leisure lives in recent years, thanks to the fast expansion of the gaming industry and the widespread use of mobile devices. The aesthetic quality of game pictures is a very important factor that attracts users' interest. However, evaluating the aesthetic quality of mobile game pictures is difficult since the painting styles of games vary greatly and the evaluation criteria are also diversified. In this article, we propose a multitask deep learning-based method, which is able to predict the aesthetic quality of mobile game images in multiple dimensions. The proposed model consists of two modules, a feature extraction module and a quality regression module. We extract quality-aware features from intermediate layers of the deep convolution neural network and then incorporate them into the final feature representation in the feature extraction module, allowing the model to fully use visual information from low to high levels. The quality regression module uses fully connected layers to map quality-aware features into quality scores across multiple dimensions. The multidimensional aesthetic quality scores are trained using a multitask learning approach, in which quality-aware features are shared across multiple dimensional quality prediction tasks. Finally, several key factors which help the proposed model perform better are analyzed. The experimental results indicate that our proposed method not only achieves the greatest performance on mobile game images, but also is applicable to natural scene images.
Tao Wang 0078, Wei Sun 0029, Wei Wu 0002, Ying Chen 0011, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai
IEEE Trans. Games5
2023 Image Quality Score Distribution Prediction via Alpha Stable Model
abstract
Based on potentially subjective and diverse image quality scores given by a group of subjects, we propose to predict the distribution of image quality scores rather than the mean opinion score (MOS) of image quality. Therefore, in this paper, we use an alpha stable model to parameterize the image quality score distribution (IQSD), and propose an objective method to predict the alpha-stable-model-based IQSD. First, the LIVE database is re-recorded. Specifically, we invite a large group of subjects (187 valid subjects) to evaluate the quality of all 808 images in the LIVE database, with their scores forming reliable IQSDs. All images in the LIVE database and their collected subjective quality scores form a new image quality assessment database, named the SJTU IQSD database. We then propose a framework and algorithm to predict the alpha-stable-model-based IQSD, in which quality features are extracted from the structural and natural statistical information of each image, and support vector regressors are trained to predict the alpha stable model parameters. Experiments carried out on the SJTU IQSD database verify the feasibility of using the alpha stable model to describe the IQSD, and the experimental results show that the alpha-stable-model-based IQSD can reflect a large amount of subjective information on image quality. We also prove that the objective alpha-stable-model-based IQSD prediction method is effective. The code and the SJTU IQSD database can be downloaded at ‘https://github.com/YixuanGao98/Image-Quality-Score-Distribution-Prediction-via-Alpha-Stable-Model.git’.
Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2023 Toward a No-Reference Quality Metric for Camera-Captured Images
abstract
Existing no-reference (NR) image quality assessment (IQA) metrics are still not convincing for evaluating the quality of the camera-captured images. Toward tackling this issue, we, in this article, establish a novel NR quality metric for quantifying the quality of the camera-captured images reliably. Since the image quality is hierarchically perceived from the low-level preliminary visual perception to the high-level semantic comprehension in the human brain, in our proposed metric, we characterize the image quality by exploiting both the low-level image properties and the high-level semantics of the image. Specifically, we extract a series of low-level features to characterize the fundamental image properties, including the brightness, saturation, contrast, noiseness, sharpness, and naturalness, which are highly indicative of the camera-captured image quality. Correspondingly, the high-level features are designed to characterize the semantics of the image. The low-level and high-level perceptual features play complementary roles in measuring the image quality. To infer the image quality, we employ the support vector regression (SVR) to map all the informative features to a single quality score. Thorough tests conducted on two standard camera-captured image databases demonstrate the effectiveness of the proposed quality metric in assessing the image quality and its superiority over the state-of-the-art NR quality metrics. The source code of the proposed metric for camera-captured images is released at https://github.com/YT2015?tab=repositories.
Runze Hu, Yutao Liu 0002, Ke Gu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Cybern.4
2023 Implicit Neural Representation Learning for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) super-resolution (SR) without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, implicit neural representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of the parametric model are predicted by a hypernetwork that operates on feature extraction using a convolution network. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding is deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high-frequency details. To verify the efficacy of our model, we conduct experiments on three HSI datasets (CAVE, NUS, and NTIRE2018). Experimental results show that the proposed model can achieve competitive reconstruction performance in comparison with the state-of-the-art methods. In addition, we provide an ablation study on the effect of individual components of our model. We hope this article could serve as a potent reference for future research.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Geosci. Remote. Sens.3
2023 Attention-Guided Neural Networks for Full-Reference and No-Reference Audio-Visual Quality Assessment
abstract
With the popularity of mobile Internet, audio and video (A/V) have become the main way for people to entertain and socialize daily. However, in order to reduce the cost of media storage and transmission, A/V signals will be compressed by service providers before they are transmitted to end-users, which inevitably causes distortions in the A/V signals and degrades the end-user's Quality of Experience (QoE). This motivates us to research the objective audio-visual quality assessment (AVQA). In the field of AVQA, most previous works only focus on single-mode audio or visual signals, which ignores that the perceptual quality of users depends on both audio and video signals. Therefore, we propose an objective AVQA architecture for multi-mode signals based on attentional neural networks. Specifically, we first utilize an attention prediction model to extract the salient regions of video frames. Then, a pre-trained convolutional neural network is used to extract short-time features of the salient regions and the corresponding audio signals. Next, the short-time features are fed into Gated Recurrent Unit (GRU) networks to model the temporal relationship between adjacent frames. Finally, the fully connected layers are utilized to fuse the temporal related features of A/V signals modeled by the GRU network into the final quality score. The proposed architecture is flexible and can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database and UnB-AVC Database demonstrate that our model outperforms the state-of-the-art AVQA methods. The code of the proposed method will be publicly available to promote the development of the field of AVQA.
Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai
IEEE Trans. Image Process.2
2023 Subjective and Objective Audio-Visual Quality Assessment for User Generated Content
abstract
In recent years, User Generated Content (UGC) has grown dramatically in video sharing applications. It is necessary for service-providers to use video quality assessment (VQA) to monitor and control users' Quality of Experience when watching UGC videos. However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the perceptual quality also depends on the accompanying audio signals. In this paper, we conduct a comprehensive study on UGC audio-visual quality assessment (AVQA) from both subjective and objective perspectives. Specially, we construct the first UGC AVQA database named SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences collected from the YFCC100m database. A subjective AVQA experiment is conducted on the database to obtain the mean opinion scores (MOSs) of the A/V sequences. To demonstrate the content diversity of the SJTU-UAV database, we give a detailed analysis of the SJTU-UAV database as well as other two synthetically-distorted AVQA databases and one authentically-distorted VQA database, from both the audio and video aspects. Then, to facilitate the development of AVQA fields, we construct a benchmark of AVQA models on the proposed SJTU-UAV database and other two AVQA databases, of which the benchmark models consist of AVQA models designed for synthetically distorted A/V sequences and AVQA models built through combining the popular VQA methods and audio features via support vector regressor (SVR). Finally, considering benchmark AVQA models perform poorly in assessing in-the-wild UGC videos, we further propose an effective AVQA model via jointly learning quality-aware audio and visual feature representations in the temporal domain, which is seldom investigated by existing AVQA models. Our proposed model outperforms the aforementioned benchmark AVQA models on the SJTU-UAV database and two synthetically distorted AVQA databases. The SJTU-UAV database and the code of the proposed model will be released to facilitate further research.
Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai
IEEE Trans. Image Process.2
2023 Blind Image Quality Assessment for Pathological Microscopic Image Under Screen and Immersion Scenarios
abstract
The high-quality pathological microscopic images are essential for physicians or pathologists to make a correct diagnosis. Image quality assessment (IQA) can quantify the visual distortion degree of images and guide the imaging system to improve image quality, thus raising the quality of pathological microscopic images. Current IQA methods are not ideal for pathological microscopy images due to their specificity. In this paper, we present deep learning-based blind image quality assessment model with saliency block and patch block for pathological microscopic images. The saliency block and patch block can handle the local and global distortions, respectively. To better capture the area of interest of pathologists when viewing pathological images, the saliency block is fine-tuned by eye movement data of pathologists. The patch block can capture lots of global information strongly related to image quality via the interaction between different image patches from different positions. The performance of the developed model is validated by the home-made Pathological Microscopic Image Quality Database under Screen and Immersion Scenarios (PMIQD-SIS) and cross-validated by the five public datasets. The results of ablation experiments demonstrate the contribution of the added blocks. The dataset and the corresponding code are publicly available at: https://github.com/mikugyf/PMIQD-SIS.
Yifei Guo, Menghan Hu, Xiongkuo Min, Yan Wang 0036, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001
IEEE Trans. Medical Imaging3
2023 Develop Then Rival: A Human Vision-Inspired Framework for Superimposed Image Decomposition
abstract
A single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a “develop-then-rival” process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. However, separating individual image views from a single superimposed image has been an important but challenging task in computer vision area for a long time. In this paper, we propose a human vision-inspired framework for single superimposed image decomposition. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods. The proposed method also achieves state-of-the-art results on related applications including single image reflection removal, single image rain removal, single image shadow removal, and illumination correction,etc., which validates the generalization of the framework.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Yuan Tian 0017, Jae-Hyun Jung, Xiaokang Yang 0001, Guangtao Zhai
IEEE Trans. Multim.3
2023 RIVIE: Robust Inherent Video Information Embedding
abstract
Imagine an interesting situation when watching a movie, we can scan the screen using our smartphones to get some extra information about this movie such as the cast, the release date, the movie's homepage, etc. Our prospect is a world where each video contains invisible information that can be delivered to us through mobile devices with cameras. This paper proposes the first deep learning-based information hiding method for videos to achieve information transmission from screens to cameras. Compared with hiding information in single images, the methods for videos need to maintain visual quality in both spatial and temporal domains. Furthermore, the training of video models builds on a large video dataset, which needs much more computational resources than training models for images. To reduce the computational complexity, we propose to simulate data on-the-fly to generate simulated sequences from single images. Then, we use the simulated data to train a spatio-temporal generator that hides information in videos while maintaining visual quality. During training, a temporal loss function based on the simulated data is exploited to ensure the temporal consistency of generated videos. After embedding, we use a decoder to recover the hidden information. To simulate the imaging pipeline from screens to cameras in the real world, we insert a distortion network between the generator and decoder. The distortion network is based on differentiable 3D rendering to cover possible distortions introduced in the procedure of camera imaging. Experimental results show that the hidden information in videos can be extracted by cameras without impacting the visual quality. Our work can be applied to many fields, such as advertisement, entertainment, and education.
Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Menghan Hu, Guangtao Zhai
IEEE Trans. Multim.4
2023 Blind Image Quality Assessment via Cross-View Consistency
abstract
Image quality assessment (IQA) is very important for both end-users and service-providers since a high-quality image can significantly improve the user's quality of experience (QoE). Most existing blind image quality assessment (BIQA) models were developed for synthetically distorted images, however, they perform poorly on in-the-wild images, which are widely existed in various practical applications. In this paper, a BIQA model is proposed that consists of a desirable self-supervised feature learning approach to mitigate the data shortage problem and learn comprehensive feature representations, and a self-attention-based feature fusion module to introduce self-attention mechanism. We develop the image quality assessment model under the framework of contrastive learning with multi views. Since human visual system perceives signals through multiple channels, the most important visual information should exist among all views of the channels. So we design the cross-view consistent information mining (CVC-IM) module to extract compact mutual information between different views. Color information and pseudo-reference image (PRI) of different distortion types are employed to formulate rich feature embeddings and preserve the quality-aware fidelity of learned representations. We employ the Transformer as the self-attention-based architecture to integrate feature embeddings. Extensive experiments show that our model achieves remarkable image quality assessment results on in-the-wild IQA datasets.
Yucheng Zhu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.4
2023 Toward Visual Behavior and Attention Understanding for Augmented 360 Degree Videos
abstract
Augmented reality (AR) overlays digital content onto reality. In an AR system, correct and precise estimations of user visual fixations and head movements can enhance the quality of experience by allocating more computational resources for analyzing, rendering, and 3D registration on the areas of interest. However, there is inadequate research to help in understanding the visual explorations of the users when using an AR system or modeling AR visual attention. To bridge the gap between the saliency prediction on real-world scenes and on scenes augmented by virtual information, we construct the ARVR saliency dataset. The virtual reality (VR) technique is employed to simulate the real-world. Annotations of object recognition and tracking as augmented contents are blended into omnidirectional videos. The saliency annotations of head and eye movements for both original and augmented videos are collected and together constitute the ARVR dataset. We also design a model that is capable of solving the saliency prediction problem in AR. Local block images are extracted to simulate the viewport and offset the projection distortion. Conspicuous visual cues in the local block images are extracted to constitute the spatial features. The optical flow information is estimated as an important temporal feature. We also consider the interplay between virtual information and reality. The composition of the augmentation information is distinguished, and the joint effects of adversarial augmentation and complementary augmentation are estimated. The Markov chain is constructed with block images as graph nodes. In the determination of the edge weights, both the characteristics of the viewing behaviors and the visual saliency mechanisms are considered. The order of importance for block images is estimated through the state of equilibrium of the Markov chain. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method.
Yucheng Zhu, Xiongkuo Min, Dandan Zhu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Ke Gu 0001, Jiantao Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 A Novel Lightweight Audio-visual Saliency Model for Videos
abstract
Audio information has not been considered an important factor in visual attention models regardless of many psychological studies that have shown the importance of audio information in the human visual perception system. Since existing visual attention models only utilize visual information, their performance is limited but also requires high-computational complexity due to the limited information available. To overcome these problems, we propose a lightweight audio-visual saliency (LAVS) model for video sequences. To the best of our knowledge, this article is the first trial to utilize audio cues for an efficient deep-learning model for the video saliency estimation. First, spatial-temporal visual features are extracted by the lightweight receptive field block (RFB) with the bidirectional ConvLSTM units. Then, audio features are extracted by using an improved lightweight environment sound classification model. Subsequently, deep canonical correlation analysis (DCCA) aims at capturing the correspondence between audio and spatial-temporal visual features, thus obtaining a spatial-temporal auditory saliency. Lastly, the spatial-temporal visual and auditory saliency are fused to obtain the audio-visual saliency map. Extensive comparative experiments and ablation studies validate the performance of the LAVS model in terms of effectiveness and complexity.
Dandan Zhu 0001, Xuan Shao, Qiangqiang Zhou, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Learning Invisible Markers for Hidden Codes in Offline-to-online Photography
abstract
QR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible codes/hyperlinks that can convey hidden information from offline to online. However, they require markers to locate invisible codes, which fails the purpose of invisible codes to be visible because of the markers. This paper proposes a novel invisible information hiding architecture for display/print-camera scenarios, consisting of hiding, locating, correcting, and recovery, where invisible markers are learned to make hidden codes truly invisible. We hide information in a sub-image rather than the entire image and include a localization module in the end-to-end framework. To achieve both high visual quality and high recovering robustness, an effective multi-stage training strategy is proposed. The experimental results show that the proposed method outperforms the state-of-the-art information hiding methods in both visual quality and robustness. In addition, the automatic localization of hidden codes significantly reduces the time of manually correcting geometric distortions for photos, which is a revolutionary innovation for information hiding in mobile applications.
Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
CVPR4
2022 End-to-End Human-Gaze-Target Detection with Transformers
abstract
In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture.
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
CVPR2
2022 Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
ECCV (4)2
2022 A Unified Two-Stage Model for Separating Superimposed Images
abstract
A single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a "develop-then-rival" process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. In this paper, we propose a human vision-inspired framework for separating superimposed images. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods.
Huiyu Duan, Xiongkuo Min, Wei Shen 0002, Guangtao Zhai
ICASSP2
2022 Surveillance Video Quality Assessment Based on Quality Related Retraining
abstract
Surveillance videos have been widely used in many vision-based systems, supporting intelligent tasks such as object detection and tracking. However, the quality of surveillance videos suffers from poor weather conditions and inevitable compression error, which may have a negative influence on the performance of such tasks. Therefore, accurately distinguishing distortions and predicting severity levels are crucial. In this paper, we propose a quality related retraining framework as well as a no-reference (NR) multi-task video quality assessment (VQA) model to tackle the challenge of surveillance videos quality assessment. The quality related retraining framework operates on a synthetic VQA database. The proposed NR VQA method utilizes both spatial and temporal information by using ResNet50 and SlowFast. Then multiple distortion detection heads are applied to predict the severity levels for corresponding distortions. The experimental results show that the proposed method gains competitive performance on the Video Surveillance Quality Assessment Dataset (VSQuAD). The ablation study further confirms the contributions of the quality related retraining framework, spatial information, and temporal information.
Wei Lu 0021, Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Guangtao Zhai
ICIP4
2022 Implicit Neural Representation Learning for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) super-resolution without additional auxiliary image remains a constant challenge due to its high-dimensional spectral patterns, where learning an effective spatial and spectral representation is a fundamental issue. Recently, Implicit Neural Representations (INRs) are making strides as a novel and effective representation, especially in the reconstruction task. Therefore, in this work, we propose a novel HSI reconstruction model based on INR which represents HSI by a continuous function mapping a spatial coordinate to its corresponding spectral radiance values. In particular, as a specific implementation of INR, the parameters of parametric model are predicted by a hypernetwork. It makes the continuous functions map the spatial coordinates to pixel values in a content-aware manner. Moreover, periodic spatial encoding are deeply integrated with the reconstruction procedure, which makes our model capable of recovering more high frequency details. Experimental results on CAVE, NUS, and NTIRE2018 datasets demonstrate the superiority of our model.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai
ICME3
2022 Human Attention Based Movie Summarization: Dataset and Baseline Model
abstract
The movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method.
Defang Zhao, Dandan Zhu 0001, Xiongkuo Min, Jiaomin Yue, Kaiwei Zhang, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
ICME3
2022 A No-Reference Deep Learning Quality Assessment Method for Super-Resolution Images Based on Frequency Maps
abstract
To support the application scenarios where high-resolution (HR) images are urgently needed, various single image super-resolution (SISR) algorithms are developed. However, SISR is an ill-posed inverse problem, which may bring artifacts like texture shift, blur, etc. to the reconstructed images, thus it is necessary to evaluate the quality of super-resolution images (SRIs). Note that most existing image quality assessment (IQA) methods were developed for synthetically distorted images, which may not work for SRIs since their distortions are more diverse and complicated. Therefore, in this paper, we propose a no-reference deep-learning image quality assessment method based on frequency maps because the artifacts caused by SISR algorithms are quite sensitive to frequency information. Specifically, we first obtain the high-frequency map (HM) and low-frequency map (LM) of SRI by using Sobel operator and piecewise smooth image approximation. Then, a two-stream network is employed to extract the quality-aware features of both frequency maps. Finally, the features are regressed into a single quality value using fully connected layers. The experimental results show that our method outperforms all compared IQA models on the selected three super-resolution quality assessment (SRQA) databases.
Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
ISCAS3
2022 SMESwin Unet: Merging CNN and Transformer for Medical Image Segmentation
Ziheng Wang 0004, Xiongkuo Min, Fangyu Shi, Ruinian Jin, Saida S. Nawrin, Ichen Yu, Ryoichi Nagatomi
MICCAI (5)2
2022 Saliency in Augmented Reality
abstract
With the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Danyang Tu, Jing Li 0026, Guangtao Zhai
ACM Multimedia3
2022 Image Quality Assessment: From Mean Opinion Score to Opinion Score Distribution
abstract
Recently, many methods have been proposed to predict the image quality which is generally described by the mean opinion score (MOS) of all subjective ratings given to an image. However, few efforts focus on predicting the opinion score distribution of the image quality ratings. In fact, the opinion score distribution reflecting subjective diversity, uncertainty, etc., can provide more subjective information about the image quality than a single MOS, which is worthy of in-depth study. In this paper, we propose a convolutional neural network based on fuzzy theory to predict the opinion score distribution of image quality. The proposed method consists of three main steps: feature extraction, feature fuzzification and fuzzy transfer. Specifically, we first use the pre-trained VGG16 without fully-connected layers to extract image features. Then, the extracted features are fuzzified by fuzzy theory, which is used to model epistemic uncertainty in the process of feature extraction. Finally, a fuzzy transfer network is used to predict the opinion score distribution of image quality by learning the mapping from epistemic uncertainty to the uncertainty existing in the image quality ratings. In addition, a new loss function is designed based on the subjective uncertainty of the opinion score distribution. Extensive experimental results prove the superior prediction performance of our proposed method.
Xiongkuo Min, Yucheng Zhu, Jing Li 0026, Xiao-Ping Zhang 0002, Guangtao Zhai
ACM Multimedia2
2022 A Deep Learning based No-reference Quality Assessment Model for UGC Videos
abstract
Quality assessment for User Generated Content (UGC) videos plays an important role in ensuring the viewing experience of end-users. Previous UGC video quality assessment (VQA) studies either use the image recognition model or the image quality assessment (IQA) models to extract frame-level features of UGC videos for quality regression, which are regarded as the sub-optimal solutions because of the domain shifts between these tasks and the UGC VQA task. In this paper, we propose a very simple but effective UGC VQA model, which tries to address this problem by training an end-to-end spatial feature extraction network to directly learn the quality-aware spatial feature representation from raw pixels of the video frames. We also extract the motion features to measure the temporal-related distortions that the spatial features cannot model. The proposed model utilizes very sparse frames to extract spatial features and dense frames (i.e. the video chunk) with a very low spatial resolution to extract motion features, which thereby has low computational complexity. With the better quality-aware features, we only use the simple multilayer perception layer (MLP) network to regress them into the chunk-level quality scores, and then the temporal average pooling strategy is adopted to obtain the video-level quality score. We further introduce a multi-scale quality fusion strategy to solve the problem of VQA across different spatial resolutions, where the multi-scale weights are obtained from the contrast sensitivity function of the human visual system. The experimental results show that the proposed model achieves the best performance on five popular UGC VQA databases, which demonstrates the effectiveness of the proposed model.
Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai
ACM Multimedia2
2022 A No-reference Quality Assessment Metric for Point Cloud Based on Captured Video Sequences
abstract
Point cloud is one of the most widely used digital formats of 3D models, the visual quality of which is quite sensitive to distortions such as downsampling, noise, and compression. To tackle the challenge of point cloud quality assessment (PCQA) in scenarios where reference is not available, we propose a no-reference quality assessment metric for colored point cloud based on captured video sequences. Specifically, three video sequences are obtained by rotating the camera around the point cloud through three specific orbits. The video sequences not only contain the static views but also include the multi-frame temporal information, which greatly helps understand the human perception of the point clouds. Then we modify the ResNet3D as the feature extraction model to learn the correlation between the capture videos and corresponding subjective quality scores. The experimental results show that our method outperforms most of the state-of-the-art full-reference and no-reference PCQA metrics, which validates the effectiveness of the proposed method.
Wei Sun 0029, Xiongkuo Min, Qiyuan Wang 0002, Guangtao Zhai
MMSP4
2022 A Full- Reference Quality Assessment Metric for Cartoon Images
abstract
Cartoon images are illustrations that are typically drawn, sometimes animated, in an unrealistic or semi-realistic style, which are widely applied in multimedia services. However, in some post-production processes as well as transmission systems, cartoon images are inevitably distorted by wrong color arrangement and compression. Therefore, it is urgent to carry out image quality assessment (IQA) metrics to automatically predict the perceptual quality levels of distorted cartoon images. Nevertheless, the existing mainstream IQA metrics are specially developed for natural scene images (NSIs). Due to the statistical difference in structure and color aspects between cartoon images and NSIs, the scores predicted by such metrics are often inconsistent with the human vision system (HVS) for cartoon images. To further improve the performance of cartoon image quality assessment (C-IQA) methods and provide guidance for practical applications, we propose a full-reference (FR) IQA method to tackle the challenge of C-IQA. Specifically, the proposed method extracts edge and texture features to analyze the structural error. Then the moment and entropy of various color spaces are computed to reflect color distortions. Then the features are regressed into quality scores with the assistance of a support vector regression (SVR) model. Experimental results show that our metric outperforms the mainstream FR-IQA metrics, which indicates that the proposed method is more capable of modeling the visual quality loss of cartoon images.
Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
MMSP4
2022 Subjective Quality Assessment for Images Generated by Computer Graphics
abstract
With the development of rendering techniques, computer graphics generated images (CGIs) have been widely used in practical application scenarios such as architecture design, video games, simulators, movies, etc. Different from natural scene images (NSIs), the distortions of CGIs are usually caused by poor rending settings and limited computation resources. What's more, some CGIs may also suffer from compression distortions in transmission systems like cloud gaming and stream media. However, limited work has been put forward to tackle the problem of computer graphics generated images' quality assessment (CG-IQA). Therefore, in this paper, we establish a large-scale subjective CG-IQA database to deal with the challenge of CG-IQA tasks. We collect 25,454 in-the-wild CGIs through previous databases and personal collection. After data cleaning, we carefully select 1,200 CGIs to conduct the subjective experiment. Several popular no-reference image quality assessment (NR-IQA) methods are tested on our database. The experimental results show that the handcrafted-based methods achieve low correlation with subjective judgment and deep learning-based methods obtain relatively better performance. The current NR-IQA models are not suitable for CG-IQA tasks and more effective models are urgently needed.
Tao Wang 0078, Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai
MMSP4
2022 Video-based Human-Object Interaction Detection from Tubelet Tokens
abstract
We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each token is learned by a selective attention mechanism to reduce redundant dependencies from others; 2) Expressiveness: each token is enabled to align with a semantic instance, i.e., an object or a human, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results show our method outperforms existing works by large margins, with a relative mAP gain of $16.14\%$ on VidHOI and a 2 points gain on CAD-120 as well as a $4 \times$ speedup.
Danyang Tu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Wei Shen 0002
NeurIPS3
2022 Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-Loop
abstract
No-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimization of man-made vision systems. Here we make one of the first attempts to examine the perceptual robustness of NR-IQA models. Under a Lagrangian formulation, we identify insightful connections of the proposed perceptual attack to previous beautiful ideas in computer vision and machine learning. We test one knowledge-driven and three data-driven NR-IQA methods under four full-reference IQA models (as approximations to human perception of just-noticeable differences). Through carefully designed psychophysical experiments, we find that all four NR-IQA models are vulnerable to the proposed perceptual attack. More interestingly, we observe that the generated counterexamples are not transferable, manifesting themselves as distinct design flows of respective NR-IQA methods. Source code are available at https://github.com/zwx8981/PerceptualAttack_BIQA.
Weixia Zhang, Dingquan Li, Xiongkuo Min, Guangtao Zhai, Guodong Guo, Xiaokang Yang 0001, Kede Ma
NeurIPS3
2022 MRIQA: Subjective Method and Objective Model for Magnetic Resonance Image Quality Assessment
abstract
Magnetic Resonance Imaging (MRI) is widely used for medical diagnosis, staging and follow-up of disease. However, MRI images may have artifacts due to various reasons such as patient movement or machine distortion, which may be unintentionally introduced during the procedure of medical image acquisition, processing, etc. These artifacts may affect the effectiveness of diagnosis or even cause false diagnosis. To solve this problem, we propose a general medical image quality assessment (MIQA) methodology, including subjective MIQA procedures and objective MIQA algorithms. We further apply this methodology to MRI images in this paper due to its widespread use in practical applications. We first establish a magnetic resonance imaging quality assessment (MRIQA) database, which contains 3809 MRI images. Then a subjective image quality assessment experiment is conducted by expert doctors according to the diagnostic value of these images, which split all MRI images into 1285 low quality images and 2524 high quality images. We then conduct a baseline deep learning experiment, and propose an attention based MIQANet model to automatically separate MRI images into high quality and low quality based on their diagnosis value. Our proposed method achieves a great quality assessment accuracy of 96.59%. The constructed MRIQA database and proposed MIQA model will be public available to further promote medical IQA research.
Fang Liu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
VCIP5
2022 Distinguishing Computer-Generated Images from Photographic Images: a Texture-Aware Deep Learning-Based Method
abstract
With the rapid development of computer graphics and generative models, computers are capable of generating images containing non-existent objects and scenes. Moreover, the computer-generated (CG) images may be indistinguishable from photographic (PG) images due to the strong representation ability of neural network and huge advancement of 3D rendering technologies. The abuse of such CG images may bring potential risks for personal property and social stability. Therefore, in this paper, we propose a dual-stream neural network to extract features enhanced by texture information to deal with the CG and PG image classification task. First, the input images are first converted to texture maps using the rotation-invariant uniform local binary patterns. Then we employ an attention-based texture-aware feature enhancement module to fuse the features extracted from each stage of the dual-stream neural network. Finally, the features are pooled and regressed into the predicted results by fully connected layers. The experimental results show that the proposed method achieves the best performance among all three popular CG and PG classification databases. The ablation study and cross-database validation experiments further confirm the effectiveness and generalization ability of the proposed algorithm.
Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
VCIP3
2022 A brief survey on adaptive video streaming quality assessment
Wei Zhou 0021, Xiongkuo Min, Qiuping Jiang
J. Vis. Commun. Image Represent.2
2022 Calculation of ophthalmic diagnostic parameters on a single eye image based on deep neural network
Xuefei Song, Xiongkuo Min, Huifang Zhou, Wei Sun 0029, Jia Wang 0004, Guangtao Zhai
Multim. Tools Appl.3
2022 No-Reference Quality Assessment for 3D Colored Point Cloud and Mesh Models
abstract
To improve the viewer’s Quality of Experience (QoE) and optimize computer graphics applications, 3D model quality assessment (3D-QA) has become an important task in the multimedia area. Point cloud and mesh are the two most widely used digital representation formats of 3D models, the visual quality of which is quite sensitive to lossy operations like simplification and compression. Therefore, many related studies such as point cloud quality assessment (PCQA) and mesh quality assessment (MQA) have been carried out to measure the visual quality of distorted 3D models. However, most previous studies utilize full-reference (FR) metrics, which indicates they can not predict the quality level in the absence of the reference 3D model. Furthermore, few 3D-QA metrics consider color information, which significantly restricts their effectiveness and scope of application. In this paper, we propose a no-reference (NR) quality assessment metric for colored 3D models represented by both point cloud and mesh. First, we project the 3D models from 3D space into quality-related geometry and color feature domains. Then, the 3D natural scene statistics (3D-NSS) and entropy are utilized to extract quality-aware features. Finally, a support vector regression (SVR) model is employed to regress the quality-aware features into visual quality scores. Our method is validated on the colored point cloud quality assessment database (SJTU-PCQA), the Waterloo point cloud assessment database (WPC), and the colored mesh quality assessment database (CMDM). The experimental results show that the proposed method outperforms most compared NR 3D-QA metrics with competitive computational resources and greatly reduces the performance gap with the state-of-the-art FR 3D-QA metrics. The code of the proposed model is publicly available now athttps://github.com/zzc-1998/NR-3DQA.
Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.3
2022 Viewing Behavior Supported Visual Saliency Predictor for 360 Degree Videos
abstract
In virtual reality (VR), correct and precise estimations of user’s visual fixations and head movements can enhance the quality of experience by allocating more computation resources for analysing and rendering on the areas of interest. However, there is insufficient research about understanding the visual exploration of users when modeling VR visual attention. To bridge the gap between the saliency prediction for traditional 2D content and omnidirectional content, we construct the visual attention dataset and propose the visual saliency prediction framework for panoramic videos. Around the instantaneous viewing behavior, we propose a traditional method to adapt 2D saliency models and design a CNN-based model to better predict visual saliency. In the proposed traditional model, mechanism of visual attention and viewing behaviors are considered in the computation of edge weights on graphs which are interpreted as Markov chains. The fraction of the visual attention that is diverted to each high-clarity vision (HCV) area is estimated through equilibrium distribution of this chain. We also propose the Graph-Based CNN model. The RGB channel and optical flow form the spatial-temporal units of HCVs, from which node feature vectors are extracted. Graph convolution is used to learn the mutual information between node feature vectors of HCVs and retain geometric information. Then feature vectors are aligned according to geometry structure of equirectangular format, and the feature decoder maps the aligned feature maps to the data distribution. We also construct the dynamic omnidirectional monocular (DOM) saliency dataset with 64 diverse videos evaluated by 28 people. The subjective results show that the instantaneous viewing behavior is important in the VR experience. Extensive experiments are conducted on the dataset and the results demonstrate the effectiveness of the proposed framework. The dataset will be released to facilitate the future studies related to visual saliency prediction for 360-degree contents.
Yucheng Zhu, Guangtao Zhai, Yiwei Yang 0007, Huiyu Duan, Xiongkuo Min, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 RIHOOP: Robust Invisible Hyperlinks in Offline and Online Photographs
abstract
In the era of multimedia and Internet, the quick response (QR) code helps people obtain information from offline to online quickly. However, the QR code is often limited in many scenarios because of its random and dull appearance. Therefore, this article proposes a novel approach to embed hyperlinks into common images, making the hyperlinks invisible for human eyes but detectable for mobile devices equipped with a camera. Our approach is an end-to-end neural network with an encoder to hide messages and a decoder to extract messages. To maintain the hidden message resilient to cameras, we build a distortion network between the encoder and the decoder to augment the encoded images. The distortion network uses differentiable 3-D rendering operations, which can simulate the distortion introduced by camera imaging in both printing and display scenarios. To maintain the visual attraction of the image with hyperlinks, a loss function conforming to the human visual system (HVS) is used to supervise the training of the encoder. Experimental results show that the proposed approach outperforms the previous work on both robustness and quality. Based on the proposed approach, many applications become possible, for example, "image hyperlinks" for advertisement on TV, website, or poster, and "invisible watermark" for copyright protection on digital resources or product packagings.
Jun Jia, Zhongpai Gao, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Cybern.5
2022 Confusing Image Quality Assessment: Toward Better Augmented Reality Experience
abstract
With the development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary value of AR is to promote the fusion of digital contents and real-world environments, however, studies on how this fusion will influence the Quality of Experience (QoE) of these two components are lacking. To achieve better QoE of AR, whose two layers are influenced by each other, it is important to evaluate its perceptual quality first. In this paper, we consider AR technology as the superimposition of virtual scenes and real scenes, and introduce visual confusion as its basic theory. A more general problem is first proposed, which is evaluating the perceptual quality of superimposed images, i.e., confusing image quality assessment. A ConFusing Image Quality Assessment (CFIQA) database is established, which includes 600 reference images and 300 distorted images generated by mixing reference images in pairs. Then a subjective quality perception experiment is conducted towards attaining a better understanding of how humans perceive the confusing images. Based on the CFIQA database, several benchmark models and a specifically designed CFIQA model are proposed for solving this problem. Experimental results show that the proposed CFIQA model achieves state-of-the-art performance compared to other benchmark models. Moreover, an extended ARIQA study is further conducted based on the CFIQA study. We establish an ARIQA database to better simulate the real AR application scenarios, which contains 20 AR reference images, 20 background (BG) reference images, and 560 distorted images generated from AR and BG references, as well as the correspondingly collected subjective quality ratings. Three types of full-reference (FR) IQA benchmark variants are designed to study whether we should consider the visual confusion when designing corresponding IQA algorithms. An ARIQA metric is finally proposed for better evaluating the perceptual quality of AR images. Experimental results demonstrate the good generalization ability of the CFIQA model and the state-of-the-art performance of the ARIQA model. The databases, benchmark models, and proposed metrics are available at: https://github.com/DuanHuiyu/ARIQA.
Huiyu Duan, Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Xiaokang Yang 0001, Patrick Le Callet
IEEE Trans. Image Process.2
2022 HazDesNet: An End-to-End Network for Haze Density Prediction
abstract
Vision-based intelligent systems such as driver assistance systems and transportation systems should take into account weather conditions. The presence of haze in images can be a critical threat to driving scenarios. Haze density measures the visibility and usability of hazy images captured in real-world conditions. The prediction of haze density can be valuable in various vision-based intelligent systems, especially in those systems deployed in outdoor environments. Haze density prediction is a challenging task since the haze and many scene contents have a lot in common in appearance. Existing methods generally utilize different priors and design complex handcrafted features to predict the visibility or haze density of the image. In this article, we propose a novel end-to-end convolutional neural network (CNN) based method to predict haze density, named as HazDesNet. Our HazDesNet takes a hazy image as input and predicts a pixel-level haze density map. The density map is then refined and smoothed, and the average of the refined map is calculated as the global haze density of the image. To verify the performance of HazDesNet, a subjective human study is performed to build a Human Perceptual Haze Density (HPHD) database, which includes 500 real-world hazy images and 100 synthetic hazy images, and the corresponding human-rated perceptual haze density scores. Experimental results show that our method achieves the best haze density prediction performance on our built HPHD database and existing databases. Besides the global quantitative results, our HazDesNet is capable of predicting a continuous, stable, fine, and high-resolution haze density map. We will make the database and code publicly available athttps://github.com/JiaheZhang/HazDesNet.
Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Jiantao Zhou 0001, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Dynamic Backlight Scaling Considering Ambient Luminance for Mobile Videos on LCD Displays
abstract
The power consumption of mobile devices is always a concern for users due to the constraints on the size and weight of mobile devices. Up to now, mobile video traffic has accounted for the majority of total traffic, which implies that viewing mobile videos has been the major activity when using mobile devices. Among all the subsystems involving in mobile video playback, the display is the most power consuming subsystem. To reduce the power consumption, dynamic backlight scaling (DBS) technique is developed by adjusting the backlight magnitude when playing the mobile video. However, the convenience of mobile devices makes lots of people watch mobile videos in various luminance environments, which makes the existing DBS methods ineffective since ambient luminance varies greatly. In this paper, we propose a novel DBS strategy to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users’ quality of experience (QoE). In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance, and backlight luminance. We then investigate whether the continuous playback of backlight-scaled videos using the proposed scaling magnitude under various luminance environments would cause flicker effect or not. Motivated by the findings of these studies, we implement a novel DBS strategy for mobile energy saving which is suitable for various ambient luminance conditions. The experimental results demonstrate that the proposed DBS strategy can save more than 40 percent power at most and can save 10 percent power even at a very high ambient luminance condition. We also show that the proposed DBS strategy can be easily adapted to different user preferences and different devices, and can be conveniently integrated into practical applications.
Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Siwei Ma 0001, Xiaokang Yang 0001
IEEE Trans. Mob. Comput.2
2022 SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient Memories
abstract
Deep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix.
Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet
IEEE Trans. Neural Networks Learn. Syst.6
2022 QoE Driven VR 360° Video Massive MIMO Transmission
abstract
Massive multiple-input and multiple-output (MIMO) enables ultra-high throughput and low latency for tile-based adaptive virtual reality (VR) 360° video transmission in wireless network. In this paper, we consider a massive MIMO system where multiple users in a single-cell theater watch an identical VR 360° video. Based on tile prediction, base station (BS) deliveries the tiles in predicted field of view (FoV) to users. By introducing practical supplementary transmission for missing tiles and unacceptable VR sickness, we propose the first stable transmission scheme for VR video. we formulate an integer non-linear programming (INLP) problem to maximize users’ average quality of experience (QoE) score. Moreover, we derive the achievable spectral efficiency (SE) expression of predictive tile groups and the approximately achievable SE expression of missing tile groups, respectively. Analytically, the overall throughput is related to the number of tile groups and the length of pilot sequences. By exploiting the relationship between the structure of viewport tiles and SE expression, we propose a multi-lattice multi-stream grouping method aimed at improving the overall throughput for VR video transmission. Moreover, we analyze the relationship between QoE objective and number of predictive tile. We transform the original INLP problem into an integer linear programming problem by setting the predictive tiles groups as some constants. With variable relaxation and recovery, we obtain the optimal average QoE. Extensive simulation results validate that the proposed algorithm effectively improves QoE.
Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Wenjun Zhang 0001, Zhi Ding 0001, Chengshan Xiao
IEEE Trans. Wirel. Commun.4
2021 Perceptual Quality Assessment for Recognizing True and Pseudo 4k Content
abstract
To meet the imperative demand for monitoring the quality of Ultra High-Definition (UHD) content in multimedia industries, we propose an efficient no-reference (NR) image quality assessment (IQA) metric to distinguish original and pseudo 4K contents and measure the quality of their quality in this paper. First, we establish a database including more than 3000 4K images composed of natural 4K images together with upscaled versions interpolated from 1080p and 720p images by fourteen algorithms. To improve computing efficiency, our model segments the input image and selects three representative patches by local variances. Then, we extract the histogram features and cut-off frequency features in the frequency domain as well as the natural scenes statistic (NSS) based features from the representative patches. Finally, we employ support vector regressor (SVR) to aggregate these extracted features as an overall quality metric to predict the quality score of the target image. Extensive experimental comparisons using seven common evaluation indicators demonstrate that the proposed model outperforms the competitive NR IQA methods and has a great ability to distinguish true and pseudo 4K images.
Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Xiaokang Yang 0001, Xiao-Ping Zhang 0002
ICASSP3
2021 Self-Conditioned Probabilistic Learning of Video Rescaling
abstract
Bicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a self-conditioned probabilistic framework for video rescaling to learn the paired downscaling and upscaling procedures simultaneously. During the training, we decrease the entropy of the information lost in the downscaling by maximizing its probability conditioned on the strong spatial-temporal prior information within the downscaled video. After optimization, the downscaled video by our framework preserves more meaningful information, which is beneficial for both the upscaling step and the downstream tasks, e.g., video action recognition task. We further extend the framework to a lossy video compression system, in which a gradient estimator for non-differential industrial lossy codecs is proposed for the end-to-end training of the whole system. Extensive experimental results demonstrate the superiority of our approach on video rescaling, video compression, and efficient action recognition tasks.
Yuan Tian 0017, Guo Lu, Xiongkuo Min, Zhaohui Che, Guangtao Zhai, Guodong Guo
ICCV3
2021 Deep Neural Networks For Full-Reference And No-Reference Audio-Visual Quality Assessment
abstract
In the field of audio and visual quality assessment, most of previous works only focused on the single-mode visual or audio signal. However, for multi-mode signals, such as video and the accompanying audio, the overall perceptual quality depends on both video and audio. In this paper, we proposed an objective audio-visual quality assessment (AVQA) architecture for multi-mode signals based on deep neural networks. We first use a pretrained convolutional neural network to extract features of the single video frames and the concurrent short audio segments. Then, the extracted features are fed into Gated Recurrent Unit networks for time sequence modeling. Finally, we utilize the fully connected layers to fuse the qualities of audio and visual signals into the final quality score. The proposed architecture can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database prove that our model outperforms the state-of-the-art AVQA methods.
Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai
ICIP2
2021 Muiqa: Image Quality Assessment Database And Algorithm For Medical Ultrasound Images
abstract
In the process of medical image acquisition, medical images may be blurred or ghosted due to machine noise, electromagnetic interference, man-made disturbance, etc. This can result in poor image quality and severely affect the diagnosis accuracy and confidence of doctors. IntraVascular UltraSound (IVUS) is an important supplementary method for the diagnosis of coronary angiography. IVUS images can be distorted for many reasons and some severe distortions can affect diagnosis confidence. However, existing manual medical image quality control method is extremely time-consuming and requires a lot of manpower. To solve this problem, we first construct an Medical UltraSound Image Quality Assessment (MUIQA) database, which consists of 10766 IVUS images with quality labels given by professional doctors. Then we propose a deep-learning network to automatically distinguish the low, medium and high level images from each other. We achieve good classification accuracy of 96.34% on the testing set.
Xiongkuo Min, Huiyu Duan, Yucheng Zhu, Guangtao Zhai
ICIP2
2021 Modeling Image Quality Score Distribution Using Alpha Stable Model
abstract
In recent years, image quality is generally described by a mean opinion score (MOS). However, we observe that an image’s quality ratings given by a group of subjects may not follow a Gaussian distribution and the image quality can not be fully described by a MOS. In this paper, we propose to describe the image quality using a parameterized distribution rather than a MOS, and an objective method is also proposed to predict the image quality score distribution (IQSD). Specifically, we selected 100 images from the LIVE database and invited a large group of subjects to evaluate the quality of these images. By analyzing the subjective quality ratings, we find that the IQSD can be well modeled by an alpha stable model and this model can reflect much more information than MOS. Therefore, we propose an algorithm to model the IQSD described by an alpha stable model. Features are extracted from images based on natural scene statistics and support vector regressors are trained to predict the IQSD described by an alpha stable model. We validate the proposed IQSD prediction model on the collected subjective quality ratings. Experimental results verify the effectiveness of the proposed algorithm in modeling the IQSD.
Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai
ICIP2
2021 Accurate Compensation Makes the World More Clear for the Visually Impaired
abstract
Visual impairment is one of the most serious social and public health problems in the world, therefore, it is of great theoretical and practical significance to study the image enhancement algorithms for the visually impaired, which is the basis for the development of assistive devices. In this paper, a general deep learning based image enhancement framework for the visually impaired is proposed, which can be used to enhance images to compensate for any visually impaired symptom that can be modeled. Take central vision loss as an example, we first model the central vision loss based on the contrast sensitivity function (CSF) specified by clinical indicator Pelli-Robson score and logMAR visual acuity, and then use the proposed framework to generate an image enhancement method aiming at compensating for the central vision loss. Both the simulation experiment and the patient experiment show the superiority of the proposed image enhancement method designed for the central vision loss, which also validates the effectiveness of the proposed framework.
Sijing Wu, Huiyu Duan, Xiongkuo Min, Danyang Tu, Guangtao Zhai
ICIP3
2021 Deep Audio-Visual Fusion Neural Network for Saliency Estimation
abstract
In this work, we propose a deep audio-visual fusion model to estimate the saliency of videos. The model extracts visual and audio features with two separate branches and fuses them to generate the saliency map. We design a novel temporal attention module to utilize the temporal information and a spatial feature pyramid module to fuse the spatial information. Then a multi-scale audio-visual fusion method is used to integrate different modalities. Furthermore, we propose a new dataset for audio-visual saliency estimation. The proposed dataset consists of 202 high quality video squences with a large range of motions, scenes and object types. Many of the videos have high audio-visual correspondence. Several experiments are conducted on different datasets. The results demonstrate that our model outperforms the previous state-of-the-art methods by a large margin and the proposed dataset can serve as a new benchmark for the audio-visual saliency estimation task.
Xiongkuo Min, Guangtao Zhai
ICIP2
2021 Attention Based Network For No-Reference UGC Video Quality Assessment
abstract
The quality assessment of user-generated content (UGC) videos is a challenging problem due to the absence of reference videos and their complex distortions. Traditional no-reference video quality assessment (NR-VQA) algorithms mainly target specific synthetic distortions. Less attention has been paid to authentic distortions in UGC videos, which are not distributed evenly in both the spatial and temporal domains. In this paper, we propose an end-to-end neural network model for UGC video quality assessment based on the attention mechanism. The key step in our approach is to embed the attention modules in the feature extraction network, which effectively extracts local distortion information. In addition, to exploit the temporal perception mechanism of the human visual system (HVS), the gated recurrent unit (GRU) and temporal pooling layer are integrated into the proposed model. We validate the proposed model on three public in-the-wild VQA databases: KoNViD-1k, CVD2014, and LIVE-Qualcomm. Experimental results demonstrate that the proposed method outperforms state-of-the-art NR-VQA models. The implementation of our method is released at https://github.com/qingshangithub/AB-VQA.
Fuwang Yi, Mianyi Chen, Wei Sun 0029, Xiongkuo Min, Yuan Tian 0017, Guangtao Zhai
ICIP4
2021 A No-Reference Evaluation Metric for Low-Light Image Enhancement
abstract
Low-light images, which are usually taken in dark or back-lighting conditions, are hard to perceive due to the low visibility and low contrast. To improve viewers’ Quality of Experience (QoE) and support the application of vision-based systems, various low-light image enhancement algorithms (LIEAs) have been proposed to lighten low-light images. However, some LIEAs may amplify the hidden distortions in the dark like noise and even further, introduce new distortions such as structural damage, color shift, etc, which severely affect the quality of light-enhanced images and need to be evaluated quantificationally. However, in the literature, few measures are proposed to assess the quality of light-enhanced images. Therefore, in this paper, we develop a no-reference low-light image enhancement evaluation (NLIEE) metric to predict the quality of light-enhanced images. The image quality is mainly assessed from four key aspects: light enhancement, color comparison, noise measurement, and structure evaluation. The experiment results show that NLIEE achieves the best performance among the general no-reference image quality assessment (NR IQA) models and quality descriptors for light enhancement.
Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
ICME3
2021 A Lightweight Saliency Prediction Model for Omnidirectional Images
abstract
At present, most high-performing saliency prediction models for omnidirectional images (ODIs) depend on deeper or wider convolutional neural networks (CNNs), benefiting from their superior feature representation capability but suffering from high computational costs. To address this issue, we propose a novel lightweight saliency prediction model to predict the eye fixations on ODIs. Specifically, our proposed model consists of three modules: a lightweight feature representation module, a supervised attention module, and a dynamic convolution aggregation module. Different from the existing saliency prediction models, our proposed model is the first to introduce the dynamic convolution into the saliency prediction and aggregate multiple parallel convolution kernels dynamically based on their attention. Such a dynamic convolution operation is not only computationally efficient (small kernel size), but also increases the feature representation capability since these convolution kernels are aggregated in a non-linear manner via attention. Experimental results on two benchmark datasets show that our model is lightweight and outperforms other state-of-the-art methods.
Dandan Zhu 0001, Yongqing Chen, Defang Zhao, Xiongkuo Min, Qiangqiang Zhou, Shaobo Yu, Guangtao Zhai, Xiaokang Yang 0001
ICME4
2021 Lavs: A Lightweight Audio-Visual Saliency Prediction Model
abstract
Audio information is essential for guiding human attention and visual perception, which has been verified by many comprehensive psychological studies. However, the audio modality has been rather neglected in modeling visual attention, most of the current visual attention models heavily depend on visual information. Additionally, current existing high-performing visual attention models rely on deeper convolution neural networks (CNNs), benefiting from their extraordinary feature learning ability but incurring high computational cost. To this end, we propose a novel lightweight audio-visual saliency (LAVS) model to efficiently address the problem of fixation prediction in videos. To the best of our knowledge, our proposed model constitutes the first attempt to exploit a lightweight network and combines the visual and audio cues to perform saliency estimation in videos. Specifically, our proposed model consists of four modules, which are spatial-temporal visual saliency estimation module, audio features extraction module, source sound localization module, and audio-visual saliency fusion module. Extensive experiments across datasets validate the effectiveness and real-time performance of the proposed LAVS model, which outperforms the other state-of-the-art methods.
Dandan Zhu 0001, Defang Zhao, Xiongkuo Min, Tian Han 0001, Qiangqiang Zhou, Shaobo Yu, Yongqing Chen, Guangtao Zhai, Xiaokang Yang 0001
ICME3
2021 A Multi-dimensional Aesthetic Quality Assessment Model for Mobile Game Images
abstract
With the development of the game industry and the popularization of mobile devices, mobile games have played an important role in people's entertainment life. The aesthetic quality of mobile game images determines the users' Quality of Experience (QoE) to a certain extent. In this paper, we propose a multi-task deep learning based method to evaluate the aesthetic quality of mobile game images in multiple dimensions (i.e. the fineness, color harmony, colorfulness, and overall quality). Specifically, we first extract the quality-aware feature representation through integrating the features from all intermediate layers of the convolution neural network (CNN) and then map these quality-aware features into the quality score space in each dimension via the quality regressor module, which consists of three fully connected (FC) layers. The proposed model is trained through a multi-task learning manner, where the quality-aware features are shared by different quality dimension prediction tasks, and the multi-dimensional quality scores of each image are regressed by multiple quality regression modules respectively. We further introduce an uncertainty principle to balance the loss of each task in the training stage. The experimental results show that our proposed model achieves the best performance on the Multi-dimensional Aesthetic assessment for Mobile Game image database (MAMG) among state-of-the-art image quality assessment (IQA) algorithms and aesthetic quality assessment (AQA) algorithms.
Tao Wang 0078, Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai
VCIP3
2021 Inter-Observer Visual Congruency in Video-Viewing
abstract
There are individual differences in human visual attention between observers when viewing the same scene. Inter-observer visual congruency (IOVC) describes the dispersion between different people's visual attention areas when they observe the same stimulus. Research on the IOVC of video is interesting but lacking. In this paper, we first introduce the measurement to calculate the IOVC of video. And an eye-tracking experiment is conducted in a realistic movie-watching environment to establish a movie scene dataset. Then we propose a method to predict the IOVC of video, which employs a dual-channel network to extract and integrate content and optical flow features. The effectiveness of the proposed prediction model is validated on our dataset. And the correlation between inter-observer congruency and video emotion is analyzed.
Jiaomin Yue, Dandan Zhu 0001, Xiongkuo Min, Xiao-Ping Zhang 0002, Guangtao Zhai
VCIP4
2021 A Full-Reference Quality Assessment Metric for Fine-Grained Compressed Images
abstract
Compressed image quality assessment (IQA) has been a crucial part of a wide range of image services such as storage and transmission. Due to the effect of different bit rates and compression methods, the compressed images usually have different levels of quality. Nowadays, the mainstream full-reference (FR) metrics are effective to predict the quality of compressed images at coarse-grained levels, however, they may perform poorly when quality differences of the compressed images are quite subtle. To better improve the Quality of Experience (QoE) and provide useful guidance for compression algorithms, we propose an FR-IQA metric for fine-grained compressed images, which estimates the image quality by analyzing the difference of structure and texture. Our metric is mainly validated on the fine-grained compression IQA (FGIQA) database and is tested on other commonly used compression IQA databases as well. The experimental results show that our metric outperforms mainstream FR-IQA metrics on the fine-grained compression IQA database and also obtains competitive performance on the coarse-grained compression IQA databases.
Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
VCIP3
2021 RANSP: Ranking attention network for saliency prediction on omnidirectional images
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Yucheng Zhu, Qiangqiang Zhou, Guangtao Zhai, Xiaokang Yang 0001
Neurocomputing3
2021 An Accurate and Efficient 1-D Barcode Detector for Medium of Deployment in IoT Systems
abstract
Camera-based 1-D barcode detectors have a lot of applications in Internet-of-Things (IoT) systems (e.g., retail, air travel, post and parcel services, and manufacturing). Based on the observation that 1-D barcodes always come with a 12-digit product code, this article proposes an end-to-end trainable and fully convoluted model that can detect and output accurate localization results of 1-D barcode and product code simultaneously. Our method uses dilated convolutions-based feature extractors which are then combined with systematic feature merging layers to create a U-shaped network. It predicts multichannel feature maps which later yield localization results after thresholding with a confidence map generated by the model and nonmaximum suppression. Furthermore, we use a Taylor series expansion-based criterion to rank and eliminate a subset of least important convolutional filters of the model, which further increases the inference speed to a great extent. This model can act as a preprocessing module for camera-based barcode decoders. Experimental results on the data set combined from public data sets and self-collected retail products data set obtained in more challenging environment conditions demonstrate that our strategy is effective in increasing the decoding rate of existing commercial barcode decoders efficiently.
Adnan Sharif, Guangtao Zhai, Jun Jia, Xiongkuo Min
IEEE Internet Things J.4
2021 Enhancing Decoding Rate of Barcode Decoders in Complex Scenes for IoT Systems
abstract
Camera-based multiclass (1-D and 2-D) barcode detectors that can help in decoding barcodes in different complex scenes have huge potential applications in situations where Internet of Things (IoT) is combined with artificial intelligence (AI) and augmented reality (AR). The decoding rate in such applications under real-life complex scenes is greatly affected by two major factors: first, we cannot accurately localize the barcodes, and second, we cannot decode the blur samples. In this article, we first propose a barcode localization algorithm that is capable of regressing four vertices of barcodes accurately. Our localization method comprises an anchor-free approach that outputs multiscale output prediction maps. These segmentation-like maps of each scale are then further divided into three types of maps (classification, centerness, and 8-D regression). Eight-dimensional localization result of barcodes along with classification result is then obtained after postprocessing. Second, we propose a conditional generative adversarial network-based model for deblurring blur QR codes. Extensive decoding experiments on a challenging complex scene data set show that our localization and deblurring methods can contribute to improving the decoding rate of existing barcode decoders.
Adnan Sharif, Guangtao Zhai, Xiongkuo Min, Jun Jia, Kashif Munir
IEEE Internet Things J.3
2021 Fine localization and distortion resistant detection of multi-class barcode in complex environments
Xiongkuo Min, Jun Jia, Zehao Zhu, Jia Wang 0004, Guangtao Zhai
Multim. Tools Appl.2
2021 Quality Assessment of Free-Viewpoint Videos by Quantifying the Elastic Changes of Multi-Scale Motion Trajectories
abstract
Virtual viewpoints synthesis is an essential process for many immersive applications including Free-viewpoint TV (FTV). A widely used technique for viewpoints synthesis is Depth-Image-Based-Rendering (DIBR) technique. However, such technique may introduce challenging non-uniform spatial-temporal structure-related distortions. Most of the existing state-of-the-art quality metrics fail to handle these distortions, especially the temporal structure inconsistencies observed during the switch of different viewpoints. To tackle this problem, an elastic metric and multi-scale trajectory based video quality metric (EM-VQM) is proposed in this paper. Dense motion trajectory is first used as a proxy for selecting temporal sensitive regions, where local geometric distortions might significantly diminish the perceived quality. Afterwards, the amount of temporal structure inconsistencies and unsmooth viewpoints transitions are quantified by calculating 1) the amount of motion trajectory deformations with elastic metric and, 2) the spatial-temporal structural dissimilarity. According to the comprehensive experimental results on two FTV video datasets, the proposed metric outperforms the state-of-the-art metrics designed for free-viewpoint videos significantly and achieves a gain of 12.86% and 16.75% in terms of median Pearson linear correlation coefficient values on the two datasets compared to the best one, respectively.
Suiyi Ling, Jing Li 0026, Zhaohui Che, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Image Process.4
2021 Video Frame Interpolation and Enhancement via Pyramid Recurrent Framework
abstract
Video frame interpolation aims to improve users' watching experiences by generating high-frame-rate videos from low-frame-rate ones. Existing approaches typically focus on synthesizing intermediate frames using high-quality reference images. However, the captured reference frames may suffer from inevitable spatial degradations such as motion blur, sensor noise, etc. Few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate and high-quality results from low-frame-rate degraded inputs. In this paper, we propose a unified optimization framework for video frame interpolation with spatial degradations. Specifically, we develop a frame interpolation module with a pyramid structure to cyclically synthesize high-quality intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates the recurrent module, thus can iteratively synthesize temporally smooth results. And the pyramid modules share weights across iterations, thus it does not expand the model's parameter size. Our model can be generalized to several applications such as up-converting the frame rate of videos with motion blur, reducing compression artifacts, and jointly super-resolving low-resolution videos. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods on various video frame interpolation and enhancement tasks.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min
IEEE Trans. Image Process.5
2021 Comparative Perceptual Assessment of Visual Signals Using Free Energy Features
abstract
In this paper, we put forward the concept of comparative perceptual quality assessment (C-PQA), which refers to the judgment of relative qualities of two visual signals of the same content, but subject to different types and levels of distortions. While it is straightforward for human observers to fulfill the CPQA task in daily lives, it remains a difficult challenge for the current research of perceptual quality assessment (PQA). Among the existing PQA algorithms, the full-reference (FR) and reducedreference (RR) methods both need prior knowledge of the original images while the no-reference (NR) algorithms usually work with a single input image. C-PQA is inherently different from those existing methods in that it takes an image pair as input and predicts their relative quality without using any knowledge about the original image. In this paper, we propose a brain theory inspired approach to C-PQA that emulates the process of comparing the relative quality of two visual stimuli as performed by the human visual system (HVS) within the framework of free energy minimization. The brain's internal generative models initialized on the inputs are then used to explain both images. During the internal generative modeling, a group of features are extracted and then integrated to determine the relative quality of two images. We designed a dedicated image database to test the proposed C-PQA algorithm. Experimental results show that the proposed method achieves up to 98% prediction accuracy in line with the subjective ratings, outperforming many state of the art PQA algorithms.
Guangtao Zhai, Yucheng Zhu, Xiongkuo Min
IEEE Trans. Multim.3
2021 Perceptual Quality Assessment of Low-light Image Enhancement
abstract
Low-light image enhancement algorithms (LIEA) can light up images captured in dark or back-lighting conditions. However, LIEA may introduce various distortions such as structure damage, color shift, and noise into the enhanced images. Despite various LIEAs proposed in the literature, few efforts have been made to study the quality evaluation of low-light enhancement. In this article, we make one of the first attempts to investigate the quality assessment problem of low-light image enhancement. To facilitate the study of objective image quality assessment (IQA), we first build a large-scale low-light image enhancement quality (LIEQ) database. The LIEQ database includes 1,000 light-enhanced images, which are generated from 100 low-light images using 10 LIEAs. Rather than evaluating the quality of light-enhanced images directly, which is more difficult, we propose to use the multi-exposure fused (MEF) image and stack-based high dynamic range (HDR) image as a reference and evaluate the quality of low-light enhancement following a full-reference (FR) quality assessment routine. We observe that distortions introduced in low-light enhancement are significantly different from distortions considered in traditional image IQA databases that are well-studied, and the current state-of-the-art FR IQA models are also not suitable for evaluating their quality. Therefore, we propose a new FR low-light image enhancement quality assessment (LIEQA) index by evaluating the image quality from four aspects: luminance enhancement, color rendition, noise evaluation, and structure preserving, which have captured the most key aspects of low-light enhancement. Experimental results on the LIEQ database show that the proposed LIEQA index outperforms the state-of-the-art FR IQA models. LIEQA can act as an evaluator for various low-light enhancement algorithms and systems. To the best of our knowledge, this article is the first of its kind comprehensive low-light image enhancement quality assessment study.
Guangtao Zhai, Wei Sun 0029, Xiongkuo Min, Jiantao Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Learning a Deep Agent to Predict Head Movement in 360-Degree Images
abstract
Virtual reality adequately stimulates senses to trick users into accepting the virtual environment. To create a sense of immersion, high-resolution images are required to satisfy human visual system, and low latency is essential for smooth operations, which put great demands on data processing and transmission. Actually, when exploring in the virtual environment, viewers only perceive the content in the current field of view. Therefore, if we can predict the head movements that are important behaviors of viewers, more processing resources can be allocated to the active field of view. In this article, we propose a model to predict the trajectory of head movement. Deep reinforcement learning is employed to mimic the decision making. In our framework, to characterize each state, features for viewport images are extracted by convolutional neural networks. In addition, the spherical coordinate maps and visited maps are generated for each viewport image, which facilitate the multiple dimensions of the state information by considering the impact of historical head movement and position information. To ensure the accurate simulation of visual behaviors during the watching of panoramas, we stipulate that the model imitates the behaviors of human demonstrators. To allow the model to generalize to more conditions, the intrinsic motivation is employed to guide the agent’s action toward reducing uncertainty, which can enhance robustness during the exploration. The experimental results demonstrate the effectiveness of the proposed stepwise head movement predictor.
Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Jiantao Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Blurry Video Frame Interpolation
abstract
Existing works reduce motion blur and up-convert frame rate through two separate ways, including frame deblurring and frame interpolation. However, few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate clear results from low-frame-rate blurry inputs. In this paper, we propose a blurry video frame interpolation method to reduce motion blur and up-convert frame rate simultaneously. Specifically, we develop a pyramid module to cyclically synthesize clear intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates a recurrent module, thus can iteratively synthesize temporally smooth results without significantly increasing the model size. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods. The source code and pre-trained model are available at https://github.com/laomao0/BIN.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min
CVPR5
2020 Identifying Children with Autism Spectrum Disorder Based on Gaze-Following
abstract
This paper presents a novel method to identify children with Autism Spectrum Disorder (ASD) based on the stimuli with gaze-following. Individuals with ASD are characterized by having atypical visual attention patterns, especially in social scenes. Gaze-following is considered to be a key element in understanding social scenarios, and it is reasonable to use stimuli with gaze-following to identify the children with ASD. Thus in this paper, we first construct a dataset of eye movements in gaze-following scenes for children with ASD (i.e., GazeFollow4ASD dataset), including 300 images with gaze-following information inside them and the corresponding eye movement data collected from 8 children with ASD and 10 healthy controls. We propose a novel deep neural network (DNN) model to extract discriminative features and classify children with ASD and healthy controls on single images. The proposed model shows the best performance among all compared methods on all datasets.
Yi Fang 0009, Huiyu Duan, Fangyu Shi, Xiongkuo Min, Guangtao Zhai
ICIP4
2020 Automatic Region Selection For Objective Sharpness Assessment Of Mobile Device Photos
abstract
Mobile devices are the source of a vast majority of digital photos today. Photos taken by mobile devices generally have fairly good visual quality. When evaluating high-quality mobile device photos, people have to manually zoom in to local regions to discern the subtle difference. Understandably, a global objective quality assessment method cannot perform well on such task. Therefore, local region selection is widely recognized as a prerequisite for the following quality evaluation. Clearly, subjective regions selection suffers from the drawbacks in terms of productivity, reproducibility and optimality. In this paper, we propose an automatic local region selection algorithm for sharpness measurement of mobile device photos. Specifically, local texture statistics, depth, saliency, as well as inter-pictures difference, are used as main features to select an optimal local region, in which the sharpness is then measured. For validation, we have built a largescale database for sharpness evaluation of mobile device photos, with 100 different scenes shot by several flagship mobile phones. The experimental results show that the performance of classic sharpness evaluation algorithms can be substantially improved with the region selected by the proposed algorithm.
Guangtao Zhai, Wenhan Zhu, Yucheng Zhu, Xiongkuo Min, Xiao-Ping Zhang 0002, Hua Yang 0001
ICIP5
2020 A Multiple Attributes Image Quality Database for Smartphone Camera Photo Quality Assessment
abstract
Smartphone is the superstar product in digital device market and the quality of smartphone camera photos (SCPs) is becoming one of the dominant considerations when consumers purchase smartphones. How to evaluate the quality of smartphone cameras and the taken photos is urgent issue to be solved. To bridge the gap between academic research accomplishment and industrial needs, in this paper, we establish a new Smartphone Camera Photo Quality Database (SCPQD2020) including 1800 images with 120 scenes taken by 15 smartphones. Exposure, color, noise and texture which are four dominant factors influencing the quality of SCP are evaluated in the subjective study, respectively. Ten popular no-reference (NR) image quality assessment (IQA) algorithms are tested and analyzed on our database. Experimental results demonstrate that the current objective models are not suitable for SCPs, and quality metrics having high correlation with human visual perception are highly needed.
Wenhan Zhu, Guangtao Zhai, Zongxi Han, Xiongkuo Min, Tao Wang 0078, Xiaokang Yang 0001
ICIP4
2020 Blind Stereoscopic Image Quality Assessment By Deep Neural Network Of Multi-Level Feature Fusion
abstract
In this paper, we propose an effective blind image quality assessment (BIQA) method for stereoscopic images by deep neural network (DNN) of multi-level feature fusion (MLFF) inspired by the multi-scale characteristics and binocular properties of the human visual system (HVS). Specifically, we firstly feed the left- and right-view images into a weight sharing convolutional neural network (CNN) for jointly feature extraction. To aggregate multi-level features, we concatenate the low-, middle-, and high-level feature maps of stereoscopic images to simulate the complicated visual interaction processing in the HVS. Two fully connected layers are used to build the nonlinear mapping from the highly abstract features to the quality scores of stereoscopic images. The experiments conducted on two public databases prove the validity of the proposed MLFF method.
Jiebin Yan, Yuming Fang 0001, Xiongkuo Min, Yiru Yao, Guangtao Zhai
ICME4
2020 Saliency Prediction on Omnidirectional Images with Brain-Like Shallow Neural Network
abstract
Deep feedforward convolutional neural networks (CNNs) perform well in the saliency prediction of omnidirectional images (ODIs), and have become the leading class of candidate models of the visual processing mechanism in the primate ventral stream. These CNNs have evolved from shallow network architecture to extremely deep and branching architecture to achieve superb performance in various vision tasks, yet it is unclear how brain-like they are. In particular, these deep feedforward CNNs are difficult to mapping to ventral stream structure of the brain visual system due to their vast number of layers and missing biologically-important connections, such as recurrence. To tackle this issue, some brain-like shallow neural networks are introduced. In this paper, we propose a novel brain-like network model for saliency prediction of head fixations on ODIs. Specifically, our proposed model consists of three modules: a CORnet-S module, a template feature extraction module and a ranking attention module (RAM). The CORnet-S module is a lightweight artificial neural network (ANN) with four anatomically mapped areas (V1, V2, V4 and IT) and it can simulate the visual processing mechanism of ventral visual stream in the human brain. The template features extraction module is introduced to extract attention maps of ODIs and provide guidance for the feature ranking in the following RAM module. The RAM module is used to rank and select features that are important for fine-grained saliency prediction. Extensive experiments have validated the effectiveness of the proposed model in predicting saliency maps of ODIs, and the proposed model outperforms other state-of-the-art methods with similar scale.
Dandan Zhu 0001, Yongqing Chen, Xiongkuo Min, Defang Zhao, Yucheng Zhu, Qiangqiang Zhou, Xiaokang Yang 0001, Tian Han 0001
ICPR3
2020 Perceptual image quality assessment: a survey
Guangtao Zhai, Xiongkuo Min
Sci. China Inf. Sci.2
2020 DevsNet: Deep Video Saliency Network using Short-term and Long-term Cues
Yuming Fang 0001, Chi Zhang 0027, Xiongkuo Min, Hanqin Huang, Yugen Yi, Guangtao Zhai, Chia-Wen Lin
Pattern Recognit.3
2020 How is Gaze Influenced by Image Transformations? Dataset and Model
abstract
Data size is the bottleneck for developing deep saliency models, because collecting eye-movement data is very time-consuming and expensive. Most of current studies on human attention and saliency modeling have used high-quality stereotype stimuli. In real world, however, captured images undergo various types of transformations. Can we use these transformations to augment existing saliency datasets? Here, we first create a novel saliency dataset including fixations of 10 observers over 1900 images degraded by 19 types of transformations. Second, by analyzing eye movements, we find that observers look at different locations over transformed versus original images. Third, we utilize the new data over transformed images, called data augmentation transformation (DAT), to train deep saliency models. We find that label-preserving DATs with negligible impact on human gaze boost saliency prediction, whereas some other DATs that severely impact human gaze degrade the performance. These label-preserving valid augmentation transformations provide a solution to enlarge existing saliency datasets. Finally, we introduce a novel saliency model based on generative adversarial networks (dubbed GazeGAN). A modified U-Net is utilized as the generator of the GazeGAN, which combines classic "skip connection" with a novel "center-surround connection" (CSC) module. Our proposed CSC module mitigates trivial artifacts while emphasizing semantic salient regions, and increases model nonlinearity, thus demonstrating better robustness against transformations. Extensive experiments and comparisons indicate that GazeGAN achieves state-of-the-art performance over multiple datasets. We also provide a comprehensive comparison of 22 saliency models on various transformed scenes, which contributes a new robustness benchmark to saliency community. Our code and dataset are available at.
Zhaohui Che, Ali Borji, Guangtao Zhai, Xiongkuo Min, Guodong Guo, Patrick Le Callet
IEEE Trans. Image Process.4
2020 A Metric for Light Field Reconstruction, Compression, and Display Quality Evaluation
abstract
Owning to the recorded light ray distributions, light field contains much richer information and provides possibilities of some enlightening applications, and it has becoming more and more popular. To facilitate the relevant applications, many light field processing techniques have been proposed recently. These operations also bring the loss of visual quality, and thus there is need of a light field quality metric to quantify the visual quality loss. To reduce the processing complexity and resource consumption, light fields are generally sparsely sampled, compressed, and finally reconstructed and displayed to the users. We consider the distortions introduced in this typical light field processing chain, and propose a full-reference light field quality metric. Specifically, we measure the light field quality from three aspects: global spatial quality based on view structure matching, local spatial quality based on near-edge mean square error, and angular quality based on multi-view quality analysis. These three aspects have captured the most common distortions introduced in light field processing, including global distortions like blur and blocking, local geometric distortions like ghosting and stretching, and angular distortions like flickering and sampling. Experimental results show that the proposed method can estimate light field quality accurately, and it outperforms the state-of-the-art quality metrics which may be effective for light field.
Xiongkuo Min, Jiantao Zhou 0001, Guangtao Zhai, Patrick Le Callet, Xiaokang Yang 0001, Xin-Ping Guan
IEEE Trans. Image Process.1
2020 Study of Subjective and Objective Quality Assessment of Audio-Visual Signals
abstract
The topics of visual and audio quality assessment (QA) have been widely researched for decades, yet nearly all of this prior work has focused only on single-mode visual or audio signals. However, visual signals rarely are presented without accompanying audio, including heavy-bandwidth video streaming applications. Moreover, the distortions that may separately (or conjointly) afflict the visual and audio signals collectively shape user-perceived quality of experience (QoE). This motivated us to conduct a subjective study of audio and video (A/V) quality, which we then used to compare and develop A/V quality measurement models and algorithms. The new LIVE-SJTU Audio and Video Quality Assessment (A/V-QA) Database includes 336 A/V sequences that were generated from 14 original source contents by applying 24 different A/V distortion combinations on them. We then conducted a subjective A/V quality perception study on the database towards attaining a better understanding of how humans perceive the overall combined quality of A/V signals. We also designed four different families of objective A/V quality prediction models, using a multimodal fusion strategy. The different types of A/V quality models differ in both the unimodal audio and video quality prediction models comprising the direct signal measurements and in the way that the two perceptual signal modes are combined. The objective models are built using both existing state-of-the-art audio and video quality prediction models and some new prediction models, as well as quality-predictive features delivered by a deep neural network. The methods of fusing audio and video quality predictions that are considered include simple product combinations as well as learned mappings. Using the new subjective A/V database as a tool, we validated and tested all of the objective A/V quality prediction models. We will make the database publicly available to facilitate further research.
Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Mylène C. Q. Farias, Alan C. Bovik
IEEE Trans. Image Process.1
2020 A Multimodal Saliency Model for Videos With High Audio-Visual Correspondence
abstract
Audio information has been bypassed by most of current visual attention prediction studies. However, sound could have influence on visual attention and such influence has been widely investigated and proofed by many psychological studies. In this paper, we propose a novel multi-modal saliency (MMS) model for videos containing scenes with high audio-visual correspondence. In such scenes, humans tend to be attracted by the sound sources and it is also possible to localize the sound sources via cross-modal analysis. Specifically, we first detect the spatial and temporal saliency maps from the visual modality by using a novel free energy principle. Then we propose to detect the audio saliency map from both audio and visual modalities by localizing the moving-sounding objects using cross-modal kernel canonical correlation analysis, which is first of its kind in the literature. Finally we propose a new two-stage adaptive audiovisual saliency fusion method to integrate the spatial, temporal and audio saliency maps to our audio-visual saliency map. The proposed MMS model has captured the influence of audio, which is not considered in the latest deep learning based saliency models. To take advantages of both deep saliency modeling and audio-visual saliency modeling, we propose to combine deep saliency models and the MMS model via a later fusion, and we find that an average of 5% performance gain is obtained. Experimental results on audio-visual attention databases show that the introduced models incorporating audio cues have significant superiority over state-of-the-art image and video saliency models which utilize a single visual modality.
Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Xiao-Ping Zhang 0002, Xiaokang Yang 0001, Xin-Ping Guan
IEEE Trans. Image Process.1
2020 The Prediction of Saliency Map for Head and Eye Movements in 360 Degree Images
abstract
By recording the whole scene around the capturer, virtual reality (VR) techniques can provide viewers the sense of presence. To provide a satisfactory quality of experience, there should be at least 60 pixels per degree, so the resolution of panoramas should reach 21600 × 10800. The huge amount of data will put great demands on data processing and transmission. However, when exploring in the virtual environment, viewers only perceive the content in the current field of view (FOV). Therefore if we can predict the head and eye movements which are important behaviors of viewer, more processing resources can be allocated to the active FOV. But conventional saliency prediction methods are not fully adequate for panoramic images. In this paper, a new panorama-oriented model, to predict head and eye movements, is proposed. Due to the superiority of computation in the spherical domain, the spherical harmonics are employed to extract features at different frequency bands and orientations. Related low- and high-level features including the rare components in the frequency domain and color domain, the difference between center vision and peripheral vision, visual equilibrium, person and car detection, and equator bias are extracted to estimate the saliency. To predict head movements, visual mechanisms including visual uncertainty and equilibrium are incorporated, and the graphical model and functional representation for the switch of head orientation are established. Extensive experimental results on the publicly available database demonstrate the effectiveness of our methods.
Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Jiantao Zhou 0001
IEEE Trans. Multim.3
2019 LPHD: A Large-Scale Head Pose Dataset for RGB Images
abstract
Head pose estimation has attracted many research interest in recent years. With the advent of deep learning, it is possible to predict the head pose accurately from the RGB images without the help of facial landmarks or depth information. However, existing head pose datasets often lack large pose head images, which extremely limits the development of head pose estimation algorithms. In this paper, we build the largescale head pose dataset (LHPD) including more than 140,000 images with the diverse and accurate head poses. The LHPD dataset includes the head images recorded from different shooting angles between the camera and the human body for the first time, which greatly expands the range of head pose compared to previous datasets. Therefore, the range of head pose can cover +/-90° for each Euler angle. The accurate and reliable head pose annotation is labeled by the motion capture system and careful calibration procedures. We then propose a head pose estimation method through fine-tuning the ResNet on the LHPD dataset when using the Euclidean distance of quaternions as the loss function. The results show that our method achieves better performance than current state-of-theart algorithms.
Wei Sun 0029, Yezhao Fan, Xiongkuo Min, Shihao Peng, Siwei Ma 0001, Guangtao Zhai
ICME3
2019 Video-Based Early ASD Detection via Temporal Pyramid Networks
abstract
Autism spectrum disorder (ASD) is a brain-based disorder characterized by social deficits and repetitive behaviors, high-rising in children. In this work, we first build a ASD video dataset, and then introduce the one glimpse early ASD detection (O-GAD) network, an effective and efficient end-to-end deep architecture for video-based early ASD detection. Our network can take arbitrary-length videos as input, detecting ASD typical actions and determining if repetitive behaviours appeared only at one glimpse. The experimental results show that our method outperforms other state-of-the-art video content analysis methods on this task in terms of mAP. Moreover, we conduct extensive ablation experiments to demonstrate the effectiveness and rationality of the designed network structure.
Yuan Tian 0017, Xiongkuo Min, Guangtao Zhai
ICME2
2019 MC360IQA: The Multi-Channel CNN for Blind 360-Degree Image Quality Assessment
abstract
In this paper, we present a multi-channel convolution neural network (CNN) for blind 360-degree image quality assessment (MC360IQA). To be consistent with the visual content of 360-degree images seen in the VR device, our model adopts the viewport images as the input. Specifically, we project each 360-degree image into six viewport images to cover omnidirectional visual content. By rotating the longitude of the front view, we can project one omnidirectional image onto lots of different groups of viewport images, which is an efficient way to avoid overfitting. MC360IQA consists of two parts, multi-channel CNN and image quality regressor. Multi-channel CNN includes six parallel ResNet34 networks, which are used to extract the features of the corresponding six viewport images. Image quality regressor fuses the features and regresses them to final scores. The results show that our model achieves the best performance among the state-of-art full-reference (FR) and no-reference (NR) image quality assessment (IQA) models on the available 360-degree IQA database.
Wei Sun 0029, Weike Luo, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001, Ke Gu 0001, Siwei Ma 0001
ISCAS3
2019 A dataset of eye movements for the children with autism spectrum disorder
abstract
Social difficulties are the hallmark features of Autism Spectrum Disorder (ASD) and can lead to atypical visual attention towards stimuli. Eye movements encode rich information about attention and psychological factors of an individual, which could help to characterize the traits of ASD. Learning atypical eye movements of the individuals with ASD towards various stimuli is important and has many application scenarios. However, due to the lack of open datasets, research in this sense is still limited. In this work, we present an open dataset of eye movements of children with Autism Spectrum Disorder. It consists of 300 natural scene images and the corresponding eye movement data collected from 14 children with ASD and 14 healthy controls. In particular, fixation maps and scanpaths are available in the dataset. Based on this dataset, researchers could analyze the visual traits of children with ASD and design specialized visual attention models to promote research in related fields, as well as design specialized models to identify the individuals with ASD. The dataset can be accessed in http://doi.org/10.5281/zenodo.2647418
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Zhaohui Che, Yi Fang 0009, Xiaokang Yang 0001, Jesús Gutiérrez 0001, Patrick Le Callet
MMSys3
2019 EMBDN: An Efficient Multiclass Barcode Detection Network for Complicated Environments
abstract
This article presents a novel method for efficient barcodes detection in real and complicated environments using a convolutional neural network (CNN)-based model. The method is developed as a preprocess-module of existing decoders to enhance decoding rates. Our method is trained as an end-to-end model to determine accurate locations of four barcode vertexes. Our method consists of four modules: 1) base net module; 2) region proposals generator; 3) classification and regression module; and 4) distortion removal module. The feature of barcodes extracted from the base net is fed to the next module. Region proposals are generated and selected as region of interest (ROI). Then the ROI are forward propagated to the classification and regression module to determine the positions and shapes of the barcodes. Finally, the distortion removal module is used to remove the geometric distortion according to regression parameters acquired from the previous step. The accurate position and distorted barcodes shape can be determined and corrected by our method. We validate our method on a challenging large-scale dataset in experiments. Compared with the previous methods, our method provides an end-to-end solution to determine accurate locations of barcode vertexes, which shows an excellent performance on detection accuracy. In addition, our method can enhance decoding rate through distortion removal.
Jun Jia, Guangtao Zhai, Zhongpai Gao, Zehao Zhu, Xiongkuo Min, Xiaokang Yang 0001, Guodong Guo
IEEE Internet Things J.6
2019 Objective Quality Evaluation of Dehazed Images
abstract
Vision-based intelligent systems like automatic driving or driving assistance can be improved by enhancing the visibility of the scenes captured in bad weather conditions. In particular, many image dehazing algorithms (DHAs) have been proposed to facilitate such applications in hazy weather. Contrary to the substantial progress of DHA developing, the quality evaluation of DHAs falls behind. Generally, DHAs can be evaluated qualitatively by human subjects or quantitatively by objective quality measures. Compared with the subjective evaluation which is time consuming and difficult to apply, objective measures with quantitative results are more needed in practical systems. But in the literature, very few measures are widely utilized, and even less measures correlate well with the overall dehazing quality (DHQ). In this paper, we study the DHQ evaluation using real hazy images systematically. We first construct a DHQ database, which is the largest of its kind so far and includes 1750 dehazed images generated from 250 real hazy images of various haze densities using seven representative DHAs. A subjective quality evaluation study is subsequently conducted on the DHQ database. Then, we propose an objective DHQ index (DHQI) by extracting and fusing three groups of features, including: 1) haze-removing features; 2) structure-preserving features; and 3) over-enhancement features, which have captured the most key aspects of dehazing. DHQI can be utilized to evaluate DHAs or optimize practical dehazing systems. Validations on the constructed DHQ database and three other databases with synthetic haze have verified the effectiveness of DHQI. Finally, we give an overview of the current DHA quality evaluation strategies, discuss their merits and demerits, and give some suggestions on systematic DHA quality evaluation. The DHQ database and the code of DHQI will be released to facilitate further research.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Xiaokang Yang 0001, Xin-Ping Guan
IEEE Trans. Intell. Transp. Syst.1
2019 Quality Evaluation of Image Dehazing Methods Using Synthetic Hazy Images
abstract
To enhance the visibility and usability of images captured in hazy conditions, many image dehazing algorithms (DHAs) have been proposed. With so many image DHAs, there is a need to evaluate and compare these DHAs. Due to the lack of the reference haze-free images, DHAs are generally evaluated qualitatively using real hazy images. But it is possible to perform quantitative evaluation using synthetic hazy images since the reference haze-free images are available and full-reference (FR) image quality assessment (IQA) measures can be utilized. In this paper, we follow this strategy and study DHA evaluation using synthetic hazy images systematically. We first build a synthetic haze removing quality (SHRQ) database. It consists of two subsets: regular and aerial image subsets, which include 360 and 240 dehazed images created from 45 and 30 synthetic hazy images using 8 DHAs, respectively. Since aerial imaging is an important application area of dehazing, we create an aerial image subset specifically. We then carry out subjective quality evaluation study on these two subsets. We observe that taking DHA evaluation as an exact FR IQA process is questionable, and the state-of-the-art FR IQA measures are not effective for DHA evaluation. Thus, we propose a DHA quality evaluation method by integrating some dehazing-relevant features, including image structure recovering, color rendition, and over-enhancement of low-contrast areas. The proposed method works for both types of images, but we further improve it for aerial images by incorporating its specific characteristics. Experimental results on two subsets of the SHRQ database validate the effectiveness of the proposed measures.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yucheng Zhu, Jiantao Zhou 0001, Guodong Guo, Xiaokang Yang 0001, Xin-Ping Guan, Wenjun Zhang 0001
IEEE Trans. Multim.1
2019 Multi-Channel Decomposition in Tandem With Free-Energy Principle for Reduced-Reference Image Quality Assessment
abstract
The visual quality of perceptions is highly correlated with the mechanisms of the human brain and visual system. Recently, the free-energy principle, which has been widely researched in brain theory and neuroscience, is introduced to quantize the perception, action, and learning in human brain. In the field of image quality assessment (IQA), on one hand, the free-energy principle can resort to the internal generative model to simulate the visual stimulus of the human beings. On the other hand, abundant psychological and neurobiological studies reveal that different frequency and orientation components of one visual stimulus arouse different neurons in the striate cortex, and the striate cortex processes visual information in the cerebral cortex. Motivated by these two aspects, a novel reduce-reference IQA metric called the multi-channel free-energy based reduced-reference quality metric is proposed in this paper. First, a two-level discrete Haar wavelet transform is used to decompose the input reference and distorted images. Next, to simulate the generative model in the human brain, the sparse representation is leveraged to extract the free-energy-based features in subband images. Finally, the overall quality metric is obtained through the support vector regressor. Extensive experimental comparisons on four benchmark image quality databases (LIVE, CSIQ, TID2008, and TID2013) demonstrate that the proposed method is highly competitive with the representative reduced-reference and classical full-reference models.
Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Menghan Hu, Jing Liu 0002, Guodong Guo, Xiaokang Yang 0001
IEEE Trans. Multim.3
2018 Learning to Predict where the Children with Asd Look
abstract
As is known to us, people with Autism Spectrum Disorder (ASD) have atypical visual attention towards stimuli. Learning the visual attention of people especially, children, with ASD contribute to related research in the field of medicine and psychology. In this paper, we first construct a saliency prediction for children with autism (SPCA) database, which is the first of its kind and consists of 500 images and the corresponding eye tracking data collected from 13 different children with ASD. We compare the performance of five state-of-the-art deep neural networks (DNN)-based saliency prediction approaches with their original networks and the fine-tuned networks on our database. We predict the atypical visual attention of children with ASD for the first time and get the best saliency prediction results for individuals with ASD so far.
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yi Fang 0009, Zhaohui Che, Xiaokang Yang 0001, Cheng Zhi, Hua Yang 0001
ICIP3
2018 Perceptual Quality Assessment of Omnidirectional Images
abstract
Omnidirectional images and videos can provide immersive experience of real-world scenes in Virtual Reality (VR) environment. We present a perceptual omnidirectional image quality assessment (IQA) study in this paper since it is extremely important to provide a good quality of experience under the VR environment. We first establish an omnidirectional IQA (OIQA) database, which includes 16 source images and 320 distorted images degraded by 4 commonly encountered distortion types, namely JPEG compression, JPEG2000 compression, Gaussian blur and Gaussian noise. Then a subjective quality evaluation study is conducted on the OIQA database in the VR environment. Considering that humans can only see a part of the scene at one movement in the VR environment, visual attention becomes extremely important. Thus we also track head and eye movement data during the quality rating experiments. The original and distorted omnidirectional images, subjective quality ratings, and the head and eye movement data together constitute the OIQA database. State-of-the-art full-reference (FR) IQA measures are tested on the OIQA database, and some new observations different from traditional IQA are made. The OIQA database will be released to facilitate further research.
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yucheng Zhu, Yi Fang 0009, Xiaokang Yang 0001
ISCAS3
2018 Saliency-induced reduced-reference quality index for natural scene and screen content images
Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Menghan Hu, Xiaokang Yang 0001
Signal Process.1
2018 The prediction of head and eye movement for 360 degree images
Yucheng Zhu, Guangtao Zhai, Xiongkuo Min
Signal Process. Image Commun.3
2018 Blind Quality Assessment Based on Pseudo-Reference Image
abstract
Traditional full-reference image quality assessment (IQA) metrics generally predict the quality of the distorted image by measuring its deviation from a perfect quality image called reference image. When the reference image is not fully available, the reduced-reference and no-reference IQA metrics may still be able to derive some characteristics of the perfect quality images, and then measure the distorted image's deviation from these characteristics. In this paper, contrary to the conventional IQA metrics, we utilize a new “reference” called pseudo-reference image (PRI) and a PRI-based blind IQA (BIQA) framework. Different from a traditional reference image, which is assumed to have a perfect quality, PRI is generated from the distorted image and is assumed to suffer from the severest distortion for a given application. Based on the PRI-based BIQA framework, we develop distortion-specific metrics to estimate blockiness, sharpness, and noisiness. The PRI-based metrics calculate the similarity between the distorted image's and the PRI's structures. An image suffering from severer distortion has a higher degree of similarity with the corresponding PRI. Through a two-stage quality regression after a distortion identification framework, we then integrate the PRI-based distortion-specific metrics into a general-purpose BIQA method named blind PRI-based (BPRI) metric. The BPRI metric is opinion-unaware (OU) and almost training-free except for the distortion identification process. Comparative studies on five large IQA databases show that the proposed BPRI model is comparable to the state-of-the-art opinion-aware- and OU-BIQA models. Furthermore, BPRI not only performs well on natural scene images, but also is applicable to screen content images. The MATLAB source code of BPRI and other PRI-based distortion-specific metrics will be publicly available.
Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Jing Liu 0002, Xiaokang Yang 0001, Chang Wen Chen
IEEE Trans. Multim.1
2018 Evaluating Quality of Screen Content Images Via Structural Variation Analysis
abstract
With the quick development and popularity of computers, computer-generated signals have drastically invaded into our daily lives. Screen content image is a typical example, since it also includes graphic and textual images as components as compared with natural scene images which have been deeply explored, and thus screen content image has posed novel challenges to current researches, such as compression, transmission, display, quality assessment, and more. In this paper, we focus our attention on evaluating the quality of screen content images based on the analysis of structural variation, which is caused by compression, transmission, and more. We classify structures into global and local structures, which correspond to basic and detailed perceptions of humans, respectively. The characteristics of graphic and textual images, e.g., limited color variations, and the human visual system are taken into consideration. Based on these concerns, we systematically combine the measurements of variations in the above-stated two types of structures to yield the final quality estimation of screen content images. Thorough experiments are conducted on three screen content image quality databases, in which the images are corrupted during capturing, compression, transmission, etc. Results demonstrate the superiority of our proposed quality model as compared with state-of-the-art relevant methods.
Ke Gu 0001, Junfei Qiao 0001, Xiongkuo Min, Guanghui Yue 0001, Weisi Lin, Daniel Thalmann
IEEE Trans. Vis. Comput. Graph.3
2017 Dynamic backlight scaling considering ambient luminance for mobile energy saving
abstract
The mobile video playback involves many subsystems of the devices such as computing, rendering and displaying subsystems. Among all subsystems, the displaying subsystem accounts for at least 38% of all consumed power, and it can be up to 68% with the maximum backlight brightness. What is more, lots of people watch videos via mobile devices in various situations, where the ambient luminance condition is different. Therefore, how to save mobile energy and improve the Quality of Experience (QoE) in different situations become significant problems. In this paper, we try to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users' QoE. In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance and backlight level. We first study whether the continuous playback of backlight-scaled shots using the proposed scaling magnitude would cause flicker effect or not. Then motivated by the findings of these subject studies, we implement a Dynamic Backlight Scaling (DBS) strategy. The experiment results demonstrate that the DBS strategy can save more than 40% power at most and can also save 10% power even at a very high ambient luminance.
Wei Sun 0029, Guangtao Zhai, Xiongkuo Min, Yutao Liu 0002, Siwei Ma 0001, Jing Liu 0002, Jiantao Zhou 0001, Xianming Liu 0005
ICME3
2017 Visual attention analysis and prediction on human faces
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Jing Liu 0002, Shiqi Wang 0001, Xinfeng Zhang 0001, Xiaokang Yang 0001
Inf. Sci.1
2017 Unified Blind Quality Assessment of Compressed Natural, Graphic, and Screen Content Images
abstract
Digital images in the real world are created by a variety of means and have diverse properties. A photographical natural scene image (NSI) may exhibit substantially different characteristics from a computer graphic image (CGI) or a screen content image (SCI). This casts major challenges to objective image quality assessment, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. To tackle this problem, we first construct a cross-content-type (CCT) database, which contains 1,320 distorted NSIs, CGIs, and SCIs, compressed using the high efficiency video coding (HEVC) intra coding method and the screen content compression (SCC) extension of HEVC. We then carry out a subjective experiment on the database in a well-controlled laboratory environment. Moreover, we propose a unified content-type adaptive (UCA) blind image quality assessment model that is applicable across content types. A key step in UCA is to incorporate the variations of human perceptual characteristics in viewing different content types through a multi-scale weighting framework. This leads to superior performance on the constructed CCT database. UCA is training-free, implying strong generalizability. To verify this, we test UCA on other databases containing JPEG, MPEG-2, H.264, and HEVC compressed images/videos, and observe that it consistently achieves competitive performance.
Xiongkuo Min, Kede Ma, Ke Gu 0001, Guangtao Zhai, Zhou Wang 0001, Weisi Lin
IEEE Trans. Image Process.1
2016 Blind quality assessment of compressed images via pseudo structural similarity
abstract
Block-based compression causes severe pseudo structures. We find that the pseudo structures of images compressed by different levels show some degree of similarity. So we propose to evaluate the quality of compressed images via the similarity between pseudo structures of two images. To obtain a “reference” image, we introduce the most distorted image (MDI), which is derived from the distorted image and suffers from the highest degree of compression. The proposed pseudo structural similarity (PSS) model calculates the similarity between pseudo structures of the distorted image and MDI. Pseudo structures of the distorted image become similar to the MDI's under the condition of severe compression. Via comparative tests, the proposed PSS model, on one hand, is shown to be comparable to state-of-the-art competitors, and on the other hand, it is not only good at assessing natural scene images but also performs the best in the hotly-researched screen content image (SCI) database. It deserves to mention that PSS is able to boost the performance of mainstream general-purpose no-reference (NR) quality measures.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yuming Fang 0001, Xiaokang Yang 0001, Xiaolin Wu 0001, Jiantao Zhou 0001, Xianming Liu 0005
ICME1
2016 Visual attention analysis and prediction on human faces with mole
abstract
Nowadays visual attention has been applied to many research and application problems. Different algorithms from low level to high level have been developed to detect the saliency map. For images with human face, high-level factors like mole may influence the visual attention. To investigate visual attention on human face with mole, we construct a Visual Attention database for Faces with Mole (VAFM) that contains face images, fixation density maps (FDM), landmark points as well as eye tracking data. Then we build visual attention model for face images with mole combining low-level saliency algorithms and high-level feature. Compared with the traditional low-level saliency algorithms, the proposed model perform better on our dataset.
Qianqian Wei, Guangtao Zhai, Chunjia Hu, Xiongkuo Min
VCIP4
2016 Fixation Prediction through Multimodal Analysis
abstract
In this article, we propose to predict human eye fixation through incorporating both audio and visual cues. Traditional visual attention models generally make the utmost of stimuli’s visual features, yet they bypass all audio information. In the real world, however, we not only direct our gaze according to visual saliency, but also are attracted by salient audio cues. Psychological experiments show that audio has an influence on visual attention, and subjects tend to be attracted by the sound sources. Therefore, we propose fusing both audio and visual information to predict eye fixation. In our proposed framework, we first localize the moving--sound-generating objects through multimodal analysis and generate an audio attention map. Then, we calculate the spatial and temporal attention maps using the visual modality. Finally, the audio, spatial, and temporal attention maps are fused to generate the final audiovisual saliency map. The proposed method is applicable to scenes containing moving--sound-generating objects. We gather a set of video sequences and collect eye-tracking data under an audiovisual test condition. Experiment results show that we can achieve better eye fixation prediction performance when taking both audio and visual cues into consideration, especially in some typical scenes in which object motion and audio are highly correlated.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2015 A hierarchical saliency detection approach for bokeh images
abstract
Bokeh is a popular photograph technique that aesthetically highlights the image subject by properly blurring the background contents. While a large number of bokeh images can be found in our albums, the impacts of bokeh on visual saliency has not been studied yet. Our study shows that traditional saliency models do not perform well on bokeh images in that the foreground/background cannot be efficiently differentiated. Therefore, in this paper we propose a hierarchical saliency model for bokeh images through combining local sharpness measure and foreground saliency detection. More specifically, first we use local sharpness feature as a clue to locate the foreground objects in bokeh image so as to get the first level saliency map. Then we compute the second level saliency map from the blurry background using a robust saliency model. In the third step, we can generate the final saliency map by unequally weighted pooling. In order to evaluate the models' performance quantitatively, we build up a bokeh image saliency database. We test the proposed model against traditional models on the bokeh image database. The results indicate that the proposed model systematically outperforms all traditional saliency models.
Zhaohui Che, Guangtao Zhai, Xiongkuo Min
MMSP3
2015 Visual attention on human face
abstract
Human faces are always the focus of visual attention since faces can provide plenty of information. Although some visual attention models incorporating face cues work better in scenes containing faces, no visual attention model is particularly designed for faces. On faces, many high-level factors will influence visual attention distribution. In practice, there are many visual communication systems in which faces occupy the scenes, such as video calls. Specific visual attention model designed for face images will be of great value in these circumstances. In this paper, we conduct research on visual attention analysis and modelling on human faces. To facilitate this research, we collect 120 face images and perform eye-tracking experiments with these images. Eye-movement data shows that detailed visual attention allocation exists on faces. Using face detection and facial landmark localization, we find that some facial features are highly effective for visual attention prediction. The performance of many visual attention models can be improved by incorporating those facial features.
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001
VCIP1
2015 Fixation prediction through multimodal analysis
abstract
In this paper, we propose to predict human fixations by incorporating both audio and visual cues. Traditional visual attention models generally make the utmost of stimuli's visual features, while discarding all audio information. But in the real world, we human beings not only direct our gaze according to visual saliency but also may be attracted by some salient audio. Psychological experiments show that audio may have some influence on visual attention, and subjects tend to be attracted the sound sources. Therefore, we propose to fuse both audio and visual information to predict fixations. In our framework, we first localize the moving-sounding objects through multimodal analysis and generate an audio attention map, in which greater value denotes higher possibility of a position being the sound source. Then we calculate the spatial and temporal attention maps using only the visual modality. At last, the audio, spatial and temporal attention maps are fused, generating our final audio-visual saliency map. We gather a set of videos and collect eye-tracking data under audio-visual test conditions. Experiment results show that we can achieve better performance when considering both audio and visual cues.
Xiongkuo Min, Guangtao Zhai, Chunjia Hu, Ke Gu 0001
VCIP1
2014 Dual-view medical image visualization based on spatial-temporal psychovisual modulation
abstract
Medical imaging technologies such as magnetic resonance imaging (MRI) and computerized tomographic (CT) are used to diagnose a wide range of medical diseases. Medical images are generated by detecting density differences between different tissues in the body. Multiple medical image visualization is of critical importance to diagnosis. This paper introduces a dual-view medical image visualization prototype based on spatial-temporal psychovisual modulation (STPVM). Temporal psychovisual modulation (TPVM) enables a single display to generate multiple visual content for different viewers. Spatial psychovisual modulation (SPVM) extends the idea of TPVM to spatial domain. STPVM combines TPVM and SPVM by exploiting both temporal and spatial redundancy of modern displays. Based on STPVM technology, one display can present even more images simultaneously. In this demo, two kinds of medical images e.g. T1 and T2 weighted MRI images, are presented simultaneously. Physicians can switch between either image by just moving the eye fixations. Since T1 and T2 are shown simultaneously and are aligned on the screen, it is more convenient for the physicians to get different information of the same spot from the T1 and T2 images. The developed demo is useful for physicians during surgery navigation and effectively reduces the burden of mental transfer.
Zhongpai Gao, Guangtao Zhai, Chunjia Hu, Xiongkuo Min
ICIP4
2014 Information security display system based on Spatial Psychovisual Modulation
abstract
Privacy protection is of increasing importance in this era of information explosion. This paper introduces an information security display system based on the idea of Spatial Psycho-visual Modulation (SPVM). With the rapid advance of modern manufacturing techniques, display devices now support very high pixel density (e.g. the retina display of Apple). Meanwhile the human visual system (HVS) cannot distinguish image signals with spatial frequency above a threshold, as predicted by the contrast sensitivity function (CSF). Therefore, it is now possible for us to devise a type of information security display using the mismatch between resolutions of modern display devices and the HVS. Given the desired visual stimuli for both the bystanders and the authorized users, we propose a method to design display signals accordingly. We select polarization as a way to effectively differentiate the bystanders and the authorized users. Operationally, applying complementary polarization to light emitted from different spatial sections of the display could in itself be a challenge. Fortunately, the development of stereoscopic display technologies has made available a lot of polarization based spatial multiplexing type of display devices. Hardware of the information security display system is based on a polarization based stereoscopic screen made by LG. Software of the information security display system is written in C++ with SDKs of DirectX and etc. Kinect is also included into our system to enhance the experience of human-computer interaction. Extended experimental results will be given in this paper to justify the effectiveness and robustness of the system. The developed system serves both as a proof-of-concept of the SPVM method, as well as a test bed for future research of SPVM based display technology.
Chunjia Hu, Guangtao Zhai, Zhongpai Gao, Xiongkuo Min
ICME4
2014 Influence of compression artifacts on visual attention
abstract
Visual attention is an important function of the human visual system (HVS). In the long term research of visual attention, various computational models have been proposed with encouraging results. However, most of those work were conducted on images with ideal visual quality. In practice, outputs of most visual communication systems contain different levels of artifacts, e.g. noise, blurring, blockiness and etc. Therefore, it is interesting to investigate the impacts of artifacts on visual attention. In this paper, we question into the problem of how the widely encountered JPEG compression artifacts affect visual attention. We designed eye-tracking experiments on images with different levels of compression and viewing time and quantitatively compared the recorded eye movement data. We found that compression level does have impacts on visual attention, and yet this influence can be negligible for low levels of compression. For high levels of compression, the visual artifacts alter visual attention in a systematic way. Dependence of the influence on viewing duration was also analyzed and it was observed that too short or too long viewing time reduces the impact of compression artifacts on visual attention.
Xiongkuo Min, Guangtao Zhai, Zhongpai Gao, Chunjia Hu
ICME1
2014 Information security display system based on temporal psychovisual modulation
abstract
This paper introduces an information security display system using temporal psychovisual modulation (TPVM). TPVM was proposed as a new information display technology using the interplay of signal processing, optoelectronics and psychophysics. Since the human visual system cannot detect quick temporal changes above the flicker fusion frequency (about 60 Hz) and yet modern display technologies offer much higher refresh rates, there is a chance for a single display to simultaneously serve different contents to multiple observers. A TPVM display broadcasts a set of images called atom frames at a high speed, and those atom frames are then weighted by liquid crystal (LC) shutter based viewing devices that are synchronized with the display before entering the human visual system and fusing into the desired visual stimuli. And through different viewing devices, people can see different information. In this work, we develop a TPVM based information security display prototype. There are two kinds of viewers, those authorized viewers with the viewing devices who can see the secret information and those unauthorized viewers (bystanders) without the viewing devices who only see mask/disguise images. The prototype is built on a 120 Hz LCD screen with synchronized LC shutter glasses that were originally developed for stereoscopic display. The system is written in C++ language with SDKs of Nvidia 3D Vision, DirectX, CEGUI, MuPDF and etc. We also added human-computer interaction support of the system using Kinect. The information security display system developed in this work serves as a proof-of-concept of the TPVM paradigm, as well as a testbed for future research of TPVM technology.
Zhongpai Gao, Guangtao Zhai, Xiongkuo Min
ISCAS3
2014 Visual attention data for image quality assessment databases
abstract
Images usually contain areas that particularly attract people's attention and visual attention is an important feature of human visual system (HVS). Visual attention had been shown to be effective in improving performance of existing image quality assessment (IQA) metrics. However, with the quick advancement of IQA research, the booming of open IQA databases calls for associated comprehensive and accurate visual attention dataset. Despite of the large number of existing computational attention/saliency models, the most accurate measure of human attention is still human based. In this research, we first conduct extensive eye tracking experiments for all the pristine images from the seven widely used IQA databases (LIVE, TID2008, CSIQ, Toyama, LIVE Multiply Distortion, IVC and A57 databases). Then we propose a gaze-duration adaptive weighting approach to generate saliency maps from the eye tracking data. When applied on the IQA databases, experimental results suggest that accuracy of benchmark quality metrics, e.g. PSNR and SSIM can be systematically improved, outperforming existing saliency datasets. Both the eye tracking data and the saliency maps in this research will be made publicly available at gvsp.sjtu.edu.cn.
Xiongkuo Min, Guangtao Zhai, Zhongpai Gao, Ke Gu 0001
ISCAS1
2014 Demo: DLP based anti-piracy display system
abstract
Camcorder piracy has great impact on the movie industry. Although there are many methods to prevent recording in theatre, no recognized technology satisfies the need of defeating camcorder piracy as well as having no effect on the audience. To realize anti-piracy, we uses a new paradigm of information display technology, called temporal psychovisual modulation (TPVM). TPVM exploits the difference in image formation mechanisms of human eyes and imaging sensors. Based on this difference, we build a prototype system on the platform of DLP® LightCrafter 4500™ which features high speed pattern display. The display system serves as a proof-of-concept of anti-piracy system.
Zhongpai Gao, Guangtao Zhai, Xiaolin Wu 0001, Xiongkuo Min, Chunjia Hu
VCIP4
2014 DLP based anti-piracy display system
abstract
Camcorder piracy has great impact on the movie industry. Although there are many methods to prevent recording in theatre, no recognized technology satisfies the need of defeating camcorder piracy as well as having no effect on the audience. This paper presents a new projector display technique to defeat camcorder piracy in the theatre using a new paradigm of information display technology, called temporal psychovisual modulation (TPVM). TPVM exploits the difference in image formation mechanisms of human eyes and imaging sensors. The images formed in human vision is continuous integration of the light field while discrete sampling is used in digital video acquisition which has "blackout" period in each sampling cycle. Based on this difference, we can decompose a movie into a set of display frames and broadcast them out at high speed so that the audience can not notice any disturbance, while the video frames captured by camcorder will contain highly objectionable artifacts. The proposed prototype system built on the platform of DLP® LightCrafter 4500™ serves as a proof-of-concept of anti-piracy system.
Zhongpai Gao, Guangtao Zhai, Xiaolin Wu 0001, Xiongkuo Min, Cheng Zhi
VCIP4
2014 Information security display via uncrowded window
abstract
With the booming of visual media, people pay more and more attention to privacy protection in public environments. Most existing research on information security such as cryptography and steganography is mainly concerned about transmission and yet little has been done to prevent the information displayed on screens from reaching eyes of the bystanders. This "security of the last foot (SOLF)" problem, if left without being taken care of, will inevitably lead to the total failure of a trustable information communication system. To deal with the SOLF problem, for the application of text-reading, we proposed an eye tracking based solution using the newly revealed concept of uncrowded window from vision research. The theory of uncrowded window suggests that human vision can only effectively recognize objects inside a small window. Object features outside the window may still be detectable but the feature detection results cannot be efficiently combined properly and therefore those objects will not be recognizable. We use eye-tracker to locate fixation points of the authorized reader in real time, and only the area inside the uncrowded window displays the private information we want to protect. A number of dummy windows with fake messages are displayed around the real uncrowded window as diversions. And without the precise knowledge about the fixations of the authorized reader, the chance for bystanders to capture the private message from those surrounding area and the dummy windows is very low. Meanwhile, since the authorized reader can only read within the uncrowded window, detrimental impact of those dummy windows is almost negligible. The proposed prototype system was written in C++ with SDKs of Direct3D, Tobii Gaze SDK, CEGUI, MuPDF, OpenCV and etc. Extended demonstration of the system will be provided to show that the proposed method is an effective solution to SOLF problem of information communication and display.
Zhongpai Gao, Guangtao Zhai, Jiantao Zhou 0001, Xiongkuo Min, Chunjia Hu
VCIP4
2013 Brightness preserving video contrast enhancement using S-shaped Transfer function
abstract
This paper presents an efficient perceptual model inspired efficient video contrast enhancement algorithm. We propose a S-shaped transfer function for image pixel values that effectively improves the perceived contrast while preserving brightness of the scene. The S-shaped transfer function has only one control parameter that can be adaptively chosen for different video contents, such as sports, cartoon, news, and landscape programs. Then, the input image brightness is further preserved, in order to maintain the perception of human visual system (HVS) to some special scenes, such as dark scene and seaside scene. Experiments and comparative study on VQEG Phase I test database demonstrate that the proposed S-shaped Transfer function based Brightness Preserving (STBP) contrast enhancement algorithm outperforms various histogram equalization based methods such as HE, DSIHE, RSIHE and WTHE, yet with much lower computational complexity.
Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiongkuo Min, Xiaokang Yang 0001, Wenjun Zhang 0001
VCIP4