Huiyu Duan

dblp:205/7136 · DBLP profile ↗
← Back
71ranked-venue papers
9as first author
66since 2021 · last 2026
0000-0002-6519-4067ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 59 · 8 first-author · 55 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
abstract
Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor robotics, etc. To bridge this gap, we introduce GeoX-Bench, a comprehensive Benchmark designed to explore and evaluate the capabilities of LMMs in cross-view Geo-localization and pose estimation. Specifically, GeoX-Bench contains 10,859 panoramic-satellite image pairs spanning 128 cities in 49 countries, along with corresponding 755,976 question-answering (QA) pairs. Among these, 42,900 QA pairs are designated for benchmarking, while the remaining are intended to enhance the capabilities of LMMs. Based on GeoX-Bench, we evaluate the capabilities of 25 state-of-the-art LMMs on cross-view geo-localization and pose estimation tasks, and further explore the empowered capabilities of instruction-tuning. Our benchmark demonstrate that while current LMMs achieve impressive performance in geo-localization tasks, their effectiveness declines significantly on the more complex pose estimation tasks, highlighting a critical area for future improvement, and instruction-tuning LMMs on the training data of GeoX-Bench can significantly improve the cross-view geo-sense abilities.
Yushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li 0001, Jing Liu 0002, Xiaohong Liu 0001, Guangtao Zhai
AAAI3
2026 Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition
abstract
Yushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yushuo Zheng, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai
ACL (1)2
2026 Q-Agent: An MLLM-Driven Framework for Universal Visual Quality Assessment
Peihang Chen, Huiyu Duan, Zitong Xu, Yuqin Cao, Sijing Wu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
QoMEX2
2026 Preference-guided debiasing for no-reference enhancement image quality assessment
Shiqi Gao, Zitong Xu, Huiyu Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
Image Vis. Comput.4
2026 DHQA-4D: A large-scale dataset and LMM-based metric for dynamic 4D digital human quality assessment
Sijing Wu, Yucheng Zhu, Huiyu Duan, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai
Pattern Recognit.4
2026 AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
abstract
The rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I outputs. However, existing image quality assessment (IQA) methods are limited to providing global quality scores, failing to deliver fine-grained perceptual evaluations for structurally complex subjects like humans, which is a critical challenge considering the frequent anatomical and textural distortions in AI-generated human images (AGHIs). To address this gap, we introduce AGHI-QA, a large-scale benchmark specifically designed for quality assessment of AGHIs. The dataset comprises 4, 000 images generated from 400 carefully crafted text prompts using 10 state-of-the-art T2I models. We conduct a systematic subjective study to collect multidimensional annotations, including perceptual quality scores, text-image correspondence scores, visible and distorted body part labels. Based on AGHI-QA, we evaluate the strengths and weaknesses of current T2I methods in generating human images from multiple dimensions. Furthermore, we propose AGHI-Assessor, a novel quality metric that integrates the large multimodal model (LMM) with domain-specific human features for precise quality prediction and identification of visible and distorted body parts in AGHIs. Extensive experimental results demonstrate that AGHI-Assessor showcases state-of-the-art performance, significantly outperforming existing IQA methods in multidimensional quality assessment and surpassing leading LMMs in detecting structural distortions in AGHIs.
Sijing Wu, Wei Sun 0029, Yucheng Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.7
2026 Quality Assessment and Distortion-Aware Saliency Prediction for AI-Generated Omnidirectional Images
abstract
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflecthumanfeedback for AI-generatedomnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research.
Huiyu Duan, Jing Liu 0002, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.2
2026 Future Fixation Sequence Prediction for Audio-Visual 360° Videos
abstract
Future fixation sequence prediction plays a crucial role in various aspects of virtual reality content production, transmission, rendering, and display. Accurate prediction of future fixation sequence can significantly enhance the quality of user experience, particularly in resource-constrained scenarios. In this paper, we present a novel framework for predicting future fixation sequence and achieves state-of-the-art performance. Specifically, the anti-projection-distortion FoV patch extraction algorithm is proposed to mitigate projection distortions. A comprehensive contextual representation is then constructed by integrating multiple data sources, including visual and audio information, historical fixation sequence, user identity, timestamp, and positional embeddings. The transformer-based predictor is proposed to perform the future fixation sequence prediction based on the integrated contextual representations. Additionally, we propose a framework that effectively utilizes saliency information as supervision and conduct saliency contrastive distillation during the training phase, eliminating the need for saliency data during inference. Overall, by integrating anti-projection-distortion and multimodal representations, along with key embeddings, a dedicated predictor, and contrastive distillation, our approach is designed to accurately predict future fixation sequences. Extensive experiments validate the effectiveness of our framework, demonstrating its superior performance in fixation prediction tasks.
Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, Huiyu Duan, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Multi-Dimensional Quality Assessment for Single-Image-to-3D Contents: Dataset and Model
abstract
The rapid advancement of AI generation technologies has led to the widespread use of AI-generated multimedia content, including images, videos, and 3D contents, across various applications. While significant progress has been made in quality evaluation for 2D content, evaluating the quality of 3D content synthesized from single image remains an underexplored problem. To bridge this gap, we introduce the first comprehensive subjective evaluation database tailored for assessing the quality of 3D content generated from single image. Our database, named AIGC-SI23DCQA, includes three distinct categories of input images, i.e., realistic images, AI-generated images, and computer graphic (CG) images, with 100 images in each category. Using five representative single-image-to-3D algorithms, we produce 1,500 3D contents and collect 94,500 annotations across three quality dimensions, including texture fidelity, shape accuracy, and overall quality. Based on the constructed database, we first benchmark and evaluate the performance of existing quality assessment methods revealing their limitations in addressing this novel task. Thus, we further propose a novel objective quality assessment method, termed I3DQA, for effective single-image-to-3D content quality assessment. Specifically, I3DQA first extracts the reference features from the source image, and the multi-modal features from the generated 3D content, including the projected video, patches, and large-multimodal model (LMM) features. These features are integrated through symmetric transformer blocks, enabling effective quality-related feature fusion and score prediction. Extensive experiments demonstrate the superior performance of our method and validate the effectiveness of its components. This work provides a foundational resource and a robust framework for advancing research in this emerging field, and our database and model are released at https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Jing Liu 0002, Yun Liu 0009, Xiaohong Liu 0001, Jia Wang 0004, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai
IEEE Trans. Image Process.2
2026 TFFN: Three-Branch Feature Fusion Network for Stereoscopic Omnidirectional Image Quality Assessment
abstract
Stereoscopic omnidirectional image (SOI) has both omnidirectional and stereoscopic perception features. Many previous models have proved the viewport characteristics and stereoscopic visual features are crucial for quality perception of SOI. However, effective monocular and binocular visual features extraction and fusion are difficult due to the size of SOI and inaccuracy of feature representation. In this paper, we proposed a three-branch feature fusion network (TFFN) by fusing two-stream binocular visual features and the important monocular features based on the viewport perspective. The hierarchical fusion module is first designed to fuse effective binocular visual features from different semantic scales, and the pseudo-difference information extraction module is built to obtain the accuracy monocular visual features to complement the binocular visual features. Finally, the above monocular and binocular visual features are fused together to measure the quality of SOI. The comparison experiments are conducted on three public datasets and the analysis of the results demonstrate the effectiveness of the proposed method.
Yun Liu 0009, Daoxin Fan, Huiyu Duan, Peiguang Jing, Guanghui Yue 0001, Guangtao Zhai
IEEE Trans. Multim.4
2025 FineVQ: Fine-Grained User Generated Content Video Quality Assessment
abstract
The rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA models generally only give an overall rating for a UGC video, which lacks fine-grained labels for serving video processing and recommendation applications. To address the challenges and promote the development of UGC videos, we establish the first large-scale Fine-grained Video quality assessment Database, termed FineVD, which comprises 6104 UGC videos with fine-grained quality scores and descriptions across multiple dimensions. Based on this database, we propose a Fine-grained Video Quality assessment (FineVQ) model to learn the fine-grained quality of UGC videos, with the capabilities of quality rating, quality scoring, and quality attribution. Extensive experimental results demonstrate that our proposed FineVQ can produce fine-grained video-quality results and achieve state-of-the-art performance on FineVD and other commonly used UGC-VQA datasets. Both FineVD and FineVQ are publicly available at: https://github.com/IntMeGroup/FineVQ.
Huiyu Duan, Qiang Hu 0003, Zitong Xu, Lu Liu 0005, Xiongkuo Min, Chunlei Cai, Tianxiao Ye, Xiaoyun Zhang 0001, Guangtao Zhai
CVPR1
2025 AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
Huiyu Duan, Guangtao Zhai, Juntong Wang, Xiongkuo Min
CVPR2
2025 FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching
Hongjiu Yu, Ying Chen 0011, Kai Li 0012, Xiongkuo Min, Huiyu Duan, Guangtao Zhai, Xu Liu 0006
ICCV8
2025 F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
abstract
Artificial intelligence generative models exhibit remarkable capabilities in content creation, particularly in face image generation, customization, and restoration. However, current AI-generated faces (AIGFs) often fall short of human preferences due to unique distortions, unrealistic details, and unexpected identity shifts, underscoring the need for a comprehensive quality evaluation framework for AIGFs. To address this need, we introduce FaceQ, a large-scale, comprehensive database of AI-generated Face images with fine-grained Quality annotations reflecting human preferences. The FaceQ database comprises 12,255 images generated by 29 models across three tasks: (1) face generation, (2) face customization, and (3) face restoration. It includes 32,742 mean opinion scores (MOSs) from 180 annotators, assessed across multiple dimensions: quality, authenticity, identity (ID) fidelity, and text-image correspondence. Using the FaceQ database, we establish F-Bench, a benchmark for comparing and evaluating face generation, customization, and restoration models, highlighting strengths and weaknesses across various prompts and evaluation dimensions. Additionally, we assess the performance of existing image quality assessment (IQA), face quality assessment (FQA), AI-generated content image quality assessment (AIGCIQA), and preference evaluation metrics, manifesting that these standard metrics are relatively ineffective in evaluating authenticity, ID fidelity, and text-image correspondence. The FaceQ database will be publicly available upon publication.
Lu Liu 0005, Huiyu Duan, Qiang Hu 0003, Chunlei Cai, Tianxiao Ye, Huayu Liu, Xiaoyun Zhang 0001, Guangtao Zhai
ICCV2
2025 LMM4LMM: Benchmarking and Evaluating Large-Multimodal Image Generation With LMMs
abstract
Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual quality and text-image alignment. Given the high cost and inefficiency of manual evaluation, an automatic metric that aligns with human preferences is desirable. To this end, we present EvalMi-50K, a comprehensive dataset and benchmark for evaluating large-multimodal image generation, which features (i) comprehensive tasks, encompassing 2,100 extensive prompts across 20 fine-grained task dimensions, and (ii) large-scale human-preference annotations, including 100K mean-opinion scores (MOSs) and 50K question-answering (QA) pairs annotated on 50,400 images generated from 24 T2I models. Based on EvalMi-50K, we propose LMM4LMM, an LMM-based metric for evaluating large multimodal T2I generation from multiple dimensions including perception, text-image correspondence, and task-specific accuracy. Extensive experimental results show that LMM4LMM achieves state-of-the-art performance on EvalMi-50K, and exhibits strong generalization ability on other AI-generated image evaluation benchmark datasets, manifesting the generality of both the EvalMi-50K dataset and LMM4LMM metric. Both EvalMi-50K and LMM4LMM will be released at https://github.com/IntMeGroup/LMM4LMM.
Huiyu Duan, Juntong Wang, Guangtao Zhai, Xiongkuo Min
ICCV2
2025 Exploring The Potential of Vision-Language Models for Pure-Image and Text-Guided-Image Saliency Prediction
abstract
We introduce VLSal, a saliency prediction framework that leverages Vision-Language Models (VLMs) to unify pure-image and text-guided-image saliency prediction tasks and achieve high performance in both. We extract visual features from the visual encoder and retrieve the corresponding visual token features from the language decoder, which serves as a natural feature fusion mechanism. These features are then processed through a U-Net-based saliency decoder to generate accurate saliency maps. To efficiently adapt the large-scale pretrained model, we apply Low-Rank Adaptation (LoRA) finetuning, reducing computational costs while preserving performance. Extensive experiments on benchmark datasets, including SALICON, MIT1003, and TIS, demonstrate that VLSal outperforms existing methods in both pure-image and text-guided-image saliency prediction.
Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICIP2
2025 SI23DCQA: Perceptual Quality Assessment of Single Image-to-3D Content
abstract
In recent years, significant efforts have been dedicated to advancing 3D content generation. However, existing quality assessment research predominantly focuses on evaluating Text-to-3D Content (T23DC) while ignoring Single Image-to-3D Content (SI23DC). In this paper, we establish the first Single Image-to-3D Content Quality Assessment (SI23DCQA) database to comprehensively study the perceptual quality of SI23DCs. The database contains 1500 SI23DCs, which are generated by 5 common SI23DC algorithms from 300 images including realistic images, AI generated images, and model rendered images. Afterward, we carry out a well-designed subjective experiment to collect subjective quality ratings for SI23DCs from three perspectives including overall, color, and shape. Additionally, a benchmark experiment is conducted with the state-of-the-art no reference image quality assessment (NR-IQA), no reference video quality assessment (NR-VQA), and no reference 3D quality assessment (NR-3DQA) and the experimental results show that current quality assessment methods are limited in evaluating the perceptual loss of SI23DCs. The database is released on https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
ICME2
2025 VIP-PCQA: A Multi-Modal Framework for No-reference Point Cloud Quality Assessment
abstract
Point clouds often suffer from geometric and color noise, as well as compression artifacts, during their production, storage, and transmission. Therefore, accurately and automatically evaluating the quality of point clouds is crucial for optimizing storage and compression strategies. This paper introduces the VIP-PCQA, a novel framework that combines Video, Image, and Point cloud modalities for no-reference Point Cloud Quality Assessment. The framework begins by rendering projection videos and normal images from point clouds, followed by sampling patches and computing statistical features related to color and geometry. Subsequently, a video encoder, two image encoders, and a point cloud encoder are employed to extract modality-specific features. Finally, these features are fused to regress the quality score. Experimental results on three publicly available benchmark databases demonstrate that VIP-PCQA achieves outstanding performance with excellent generalization capabilities. An ablation study further highlights the indispensable contribution of each modality to the framework’s success. The code is released on https://github.com/ZedFu/VIP-PCQA.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ICME3
2025 HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment
abstract
Image composition involves extracting a foreground object from one image and pasting it into another image through Image harmonization algorithms (IHAs), which aim to adjust the appearance of the foreground object to better match the background. Existing image quality assessment (IQA) methods may fail to align with human visual preference on image harmonization due to the insensitivity to minor color or light inconsistency. To address the issue and facilitate the advancement of IHAs, we introduce the first Image Quality Assessment Database for image Harmony evaluation (HarmonyIQAD), which consists of 1,350 harmonized images generated by 9 different IHAs, and the corresponding human visual preference scores. Based on this database, we propose a Harmony Image Quality Assessment (HarmonyIQA), to predict human visual preference for harmonized images. Extensive experiments show that HarmonyIQA achieves state-of-the-art performance on human visual preference evaluation for harmonized images, and also achieves competing results on traditional IQA tasks. Furthermore, cross-dataset evaluation also shows that HarmonyIQA exhibits better generalization ability than self-supervised learning-based IQA methods. The dataset and code are available at https://github.com/IntMeGroup/HarmonyIQA.
Zitong Xu, Huiyu Duan, Guangji Ma, Qingbo Wu 0001, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICME2
2025 ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
abstract
With the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the quality of experience (QoE) of egocentric spatial videos is crucial to ensure a high-quality viewing experience. However, the corresponding research is still lacking. In this paper, we use the concept of embodied experience to highlight this more immersive experience and study the new problem, i.e., embodied perceptual quality assessment for egocentric spatial videos. Specifically, we introduce the first Egocentric Spatial Video Quality Assessment Database (ESVQAD), which comprises 600 egocentric spatial videos captured using the Apple Vision Pro and their corresponding mean opinion scores (MOSs). Furthermore, we propose a novel multi-dimensional binocular feature fusion model, termed ESVQAnet, which integrates binocular spatial, motion, and semantic features to predict the overall perceptual quality. Experimental results demonstrate the ESVQAnet significantly outperforms 16 state-of-the-art VQA models on the embodied perceptual quality assessment task, and exhibits strong generalization capability on traditional VQA tasks. The database and code are available at https://github.com/IntMeGroup/ESVQA.
Xilei Zhu, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICME2
2025 Visual Saliency Prediction for Augmented Reality Videos
abstract
Augmented Reality (AR) is an emerging technology that allows users to perceive both virtual-world contents and real-world scenes simultaneously. It has numerous applications in industrial manufacturing, entertainment, gaming, education, etc. In AR environments, the visual confusion phenomenon caused by the overlay of augmented content and real backgrounds is evident, yet the understanding and research of visual saliency under the AR visual confusion condition remains limited. This paper primarily analyzes the interaction between real-world scenes and AR content, explores human visual saliency when using AR devices. First, we conduct a large-scale eye-tracking experiment based on a head-mounted AR device, and construct an AR saliency dataset containing 2160 videos, with corresponding collected eye movement data. Through qualitative analysis of the visual attention heat maps, we conclude that visual confusion significantly influences visual attention in AR video. Additionally, we quantitatively evaluate the performance of a series of classical saliency models and deep neural network saliency models on the dataset constructed in this project. For better predicting saliency in AR, we propose a general saliency prediction model, InternSal, which achieves state-of-the-art performance compared to other methods. The database and codes will be released to facilitate future research.
Zongyi Xie, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
ISCAS2
2025 RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment
abstract
As camera-equipped robotic platforms become increasingly integrated into daily life, robotic-generated videos have begun to appear on streaming media platforms, enabling us to envision a future where humans and robots coexist. We innovatively propose the concept of Robotic-Generated Content (RGC) to term these videos generated from egocentric perspective of robots. The perceptual quality of RGC videos is critical in human-robot interaction scenarios, and RGC videos exhibit unique distortions and visual requirements that differ markedly from those of professionally-generated content (PGC) videos and user-generated content (UGC) videos. However, dedicated research on quality assessment of RGC videos is still lacking. To address this gap and to support broader robotic applications, we establish the first Robotic-Generated Content Database (RGCD), which contains a total of 2,100 videos drawn from three robot categories and sourced from diverse platforms. A subjective VQA experiment is conducted subsequently to assess human visual perception of robotic-generated videos. Finally, we conduct a benchmark experiment to evaluate the performance of 11 state-of-the-art VQA models on our database. Experimental results reveal significant limitations in existing VQA models when applied to complex, robotic-generated content, highlighting a critical need for RGC-specific VQA models. Our RGCD is publicly available at: https://github.com/IntMeGroup/RGC-VQA.
Jianing Jin, Jiangyong Ying, Huiyu Duan, Sijing Wu, Yushuo Zheng, Xiongkuo Min, Guangtao Zhai
ACM Multimedia3
2025 DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
abstract
With the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticity. Current deepfake detection methods often depend on datasets with limited generation models and content diversity that fail to keep pace with the evolving complexity and increasing realism of the AI-generated content. Large multimodal models (LMMs), widely adopted in various vision tasks, have demonstrated strong zero-shot capabilities, yet their potential in deepfake detection remains largely unexplored. To bridge this gap, we present DFBench, a large-scale DeepFake Benchmark featuring (i) broad diversity, including 540,000 images across real, AI-edited, and AI-generated content, (ii) latest model, the fake images are generated by 12 state-of-the-art generation models, and (iii) bidirectional benchmarking and evaluating for both the detection accuracy of deepfake detectors and the evasion capability of generative models. Based on DFBench, we propose MoA-DF, Mixture of Agents for DeepFake detection, leveraging a combined probability strategy from multiple LMMs. MoA-DF achieves state-of-the-art performance, further proving the effectiveness of leveraging LMMs for deepfake detection. Database and codes are publicly available at https://github.com/IntMeGroup/DFBench.
Huiyu Duan, Juntong Wang, Ziheng Jia, Woo Yi Yang, Xiaorong Zhu, Jiaying Qian, Yuke Xing, Guangtao Zhai, Xiongkuo Min
ACM Multimedia2
2025 HVEval: Towards Unified Evaluation of Human-Centric Video Generation and Understanding
abstract
Human-centric videos play a significant role in the pervasive video content of modern life. However, the capabilities of text-to-video (T2V) generation models and video-to-text (V2T) understanding models for human-centric videos remain largely unexplored. To this end, we present HVEval, the first comprehensive evaluation dataset focusing on human-centric videos, which consists of 20,000 videos, 60k MOS annotations across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), and 20k category-specific Q&A pairs. Based on the HVEval dataset, this paper aims to answer three questions: (1) can today's T2V models effectively generate human-centric videos following the given prompts? (2) how effective are today's V2T LMMs in understanding and evaluating human-centric videos? (3) are current VQA metrics good enough for evaluating human-centric videos? Comprehensive evaluations of 24 T2V models, 20 LMMs, and 18 VQA metrics reveal their limitations in fine-grained text-controlled generation and human-aligned perception and understanding, highlighting the significant potential of our dataset and benchmarks to advance research in human-centric video generation and understanding.
Sijing Wu, Huiyu Duan, Yanwei Jiang, Yucheng Zhu, Guangtao Zhai
ACM Multimedia3
2025 FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment
abstract
Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA. The code and dataset will be released at: https://github.com/wsj-sjtu/FVQ.
Sijing Wu, Ziwen Xu, Huiyu Duan, Wei Sun 0029, Guangtao Zhai
ACM Multimedia5
2025 LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
abstract
The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.
Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Shiqi Gao, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ACM Multimedia2
2025 Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni Model
abstract
360° omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires specialized equipment, making ODI synthesis increasingly important. While common 2D image generation and editing methods are rapidly advancing, these models struggle to deliver satisfactory results when generating or editing ODIs due to the unique format and broad 360° Field-of-View (FoV) of ODIs. To bridge this gap, we construct Any2Omni , the first comprehensive ODI generation-editing dataset comprises 60,000+ training data covering diverse input conditions and up to 9 ODI generation and editing tasks. Built upon Any2Omni, we propose an Omni model for Omni-directional image generation and editing ( Omni 2), with the capability of handling various ODI generation and editing tasks under diverse input conditions using one model. Extensive experiments demonstrate the superiority and effectiveness of the proposed Omni2 model for both the ODI generation and editing tasks. Both the Any2Omni dataset and the Omni2 model are publicly available at: https://github.com/IntMeGroup/Omni2.
Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Lu Liu 0005, Zitong Xu, Guangji Ma, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ACM Multimedia2
2025 LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs
abstract
The rapid advancement in generative artificial intelligence have enabled the creation of 3D human faces (HFs) for applications including media production, virtual reality, security, healthcare, and game development, etc. However, assessing the quality and realism of these AI-generated 3D human faces remains a significant challenge due to the subjective nature of human perception and innate perceptual sensitivity to facial features. To this end, we conduct a comprehensive study on the quality assessment of AI-generated 3D human faces. We first introduce Gen3DHF, a large-scale benchmark comprising 2,000 videos of AI-Generated 3D Human Faces along with 4,000 Mean Opinion Scores (MOS) collected across two dimensions, i.e., quality and authenticity, 2,000 distortion-aware saliency maps and distortion descriptions. Based on Gen3DHF, we propose LMME3DHF, a Large Multimodal Model (LMM)-based metric for Evaluating 3DHF capable of quality and authenticity score prediction, distortion-aware visual question answering, and distortion-aware saliency prediction. Experimental results show that LMME3DHF achieves state-of-the-art performance, surpassing existing methods in both accurately predicting quality scores for AI-generated 3D human faces and effectively identifying distortion-aware salient regions and distortion types, while maintaining strong alignment with human perceptual judgments. Both the Gen3DHF database and the LMME3DHF will be released upon the publication.
Woo Yi Yang, Sijing Wu, Huiyu Duan, Guangtao Zhai, Xiongkuo Min
ACM Multimedia4
2025 CompBench: Benchmarking and Comparing Image Generation with Large Multimodal Models
abstract
Recent advancements in large multimodal models (LMMs) have significantly enhanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, critical challenges in perceptual quality and text-image correspondence remain hindering the practicality of AI-generated images (AIGIs). Therefore, a reliable benchmark and automatic model for AIGI evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present CompBench, the largest dataset for benchmarking and comparing image generation, which features: (i) the largest AIGI pair comparison dataset, comprising 616,346 carefully curated image pairs generated by 24 state-of-the-art AIGI models annotated with 1.6M+ human annotations, enabling robust relative quality assessment through pairwise comparison, (ii) multi-dimensional pairwise comparison from perceptual and text-image correspondence perspectives across three difficulty levels, and (iii) bidirectional benchmarking and evaluating for both T2I generation models and AIGI comparison models. Based on CompBench, we propose LMM4Comp, a LMM-based evaluation metric that learns nuanced quality distinctions from multiple dimensions for pairwise comparison at both instance level and model level. Experiments demonstrate that LMM4Comp achieves state-of-the-art performance, highly aligning to human preference. Both of the CompBench dataset and LMM4Comp metric will be released at https://github.com/IntMeGroup/CompBench.
Huiyu Duan, Yuke Xing, Yiling Xu, Guangtao Zhai, Xiongkuo Min
MMSP2
2025 Mixed-Reference Quality Assessment for Novel View Synthesis Scenes
Huiyu Duan, Dong Zhang 0006, Yongming Han, Guangtao Zhai
PRCV (8)3
2025 Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture Generation
abstract
The Audio-to-3D-Gesture (A2G) task exhibits significant potential across domains including virtual reality, computer graphics, and 3D animation production. However, current evaluation metrics, such as Fréchet Gesture Distance or Beat Constancy, fail at reflecting the human preference of the generated 3D gestures. To cope with this problem, exploring human preference and an objective quality assessment metric for AI-generated 3D human gestures is becoming increasingly significant. In this paper, we introduce the Ges-QA dataset, which includes 1,400 samples with multidimensional scores for gesture quality and audio-gesture consistency. Moreover, we collect binary classification labels to determine whether the generated gestures match the emotions of the audio. Equipped with our Ges-QA dataset, we propose a multi-modal transformer-based neural network with 3 branches for video, audio and 3D skeleton modalities, which can score A2G contents in multiple dimensions. Comparative experimental results and ablation studies demonstrate that Ges-QAer yields state-of-the-art performance on our dataset.
Zhilin Gao, Sijing Wu, Yuqin Cao, Huiyu Duan, Guangtao Zhai
VCIP5
2025 CVBench: Benchmarking and Comparing Video Generation with Large Multimodal Models
abstract
Large multimodal models (LMMs) have revolutionized both text-to-video (T2V) generation and video-to-text (V2T) interpretation. However, despite these advancements, issues such as imperfect perceptual quality and inconsistent text-video alignment continue to limit the practical deployment of AI-generated videos (AIGVs). Consequently, there is a pressing need for a reliable benchmark and automatic evaluation framework tailored for AIGVs. To this end, we propose CVBench, the largest and most comprehensive dataset for Comparative Video Benchmarking, including 60K video pairs generated by 30 state-of-the-art T2V models and 600K pairwise comparisons annotated with over 1.7 million human judgments from perspectives of both perceptual quality and text-video correspondence. This dataset enables bidirectional benchmarking and evaluation of both T2V generation models and V2T interpretation models. Based on CVBench, we propose VComp, a novel LMM-based evaluation metric that captures fine-grained quality differences from multiple perspectives for pairwise comparison at both the instance level and model level. Extensive experiments show that VComp achieves state-of-the-art alignment with human preferences. Both the CVBench dataset and VComp metric will be available at https://github.com/IntMeGroup/CVBench.
Huiyu Duan, Yuke Xing, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min
VCIP2
2025 MGM-DPO: Multi-Generative-Model Guided Preference Optimization for Advanced Text-to-Image Models
abstract
Text-to-image generation has rapidly progressed with autoregressive, diffusion, and distillation-based models, enabling high-fidelity and semantically relevant image synthesis from natural language prompts. Despite their success, these models still suffer from text–image misalignment, limited detail, and outputs that may violate human commonsense or aesthetic preferences. Existing human-feedback-based fine-tuning approaches partially alleviate these issues but primarily focus on earlier diffusion models (e.g., Stable Diffusion v1.4, v2.0, SDXL) and often exhibit over-optimization or under-optimization, leaving their effectiveness on state-of-the-art models unclear. In this work, we advance preference alignment for modern text-to-image models. We utilize EvalMi-50K, a large-scale dataset with human preference scores, to train our image quality assessment (IQA) model and use the predicted scores from the IQA model to reward the generation model. Building upon the dataset and IQA model, we propose Multi-Generative-Model Guided Diffusion Preference Optimization(MGM-DPO), which leverages diverse model outputs and a stable optimization strategy to align advanced diffusion models. Applied to Stable Diffusion 3.5 Large, MGM-DPO significantly improves fidelity, prompt alignment, and human preference alignment across multiple benchmarks, achieving state-of-the-art results among human-feedback-tuned models. Code is available at https://github.com/IntMeGroup/MGM-DPO.
Huiyu Duan, Guangtao Zhai, Xiongkuo Min
VCIP3
2025 A robust self-training algorithm based on relative node graph
Jikui Wang, Huiyu Duan, Cuihong Zhang, Feiping Nie 0001
Appl. Intell.2
2025 Situation-adaptive neural network for fast pre-computing image enhancement
Xinyue Li 0001, Huiyu Duan, Jia Wang 0004, Xiaohong Liu 0001, Guangtao Zhai
Sci. China Inf. Sci.2
2025 Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai
Sci. China Inf. Sci.38
2025 Energy-Efficient VR 360 Video Streaming in the IRS-Aided Rate-Splitting Multiple Access Network
abstract
Maximizing energy efficiency in VR 360 video transmission is essential for advancing VR applications. Motivated by this goal, we conduct a comprehensive study that integrates the characteristics of VR 360 video with beamforming strategies and intelligent reflecting surface (IRS) shifting techniques in the rate-splitting multiple access (RSMA) network. In the IRS-aided RSMA network, we propose a stable energy-efficient transmission (SEET) scheme aimed at minimizing the number of transmitted VR video chunks. The SEET scheme constructs a stable pre-transmission and playback flow, ensuring seamless and continuous display of the upcoming content without latency. We also propose a mixed-format-based chunk (MFC) method that simultaneously pre-transmits both 2D and 3D chunk frames to each user, further enhancing energy efficiency. We utilize an alternating optimization method to divide the original energy-efficient problem into three subproblems. To tackle the non-convex and NP-hard beamforming subproblem, we utilize the first-order Taylor expansion and then obtain the approximate transmission rates of common messages and private messages regarding the quadratic form of beamforming vectors. We then utilize quadratically constrained programming, fractional programming, and linear programming to obtain the near-optimal solutions for beamforming vectors, IRS phase shifts, and RSMA parameters, respectively. The final numerical results affirm that the proposed SEET scheme can notably minimize the beamforming power of the base station. Through the SEET scheme, the MFC method with the approximation method exhibits superior energy efficiency, outperforming existing transmission methods in terms of both energy utility and consumption by HMDs.
Qingqing Wu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Commun.3
2025 Multi-Task Guided No-Reference Omnidirectional Image Quality Assessment With Feature Interaction
abstract
Omnidirectional image quality assessment (OIQA) has become an increasingly vital problem in recent years. Most previous no-reference OIQA methods only extract local features from the distorted viewports, or extract global features from the entire distorted image, lacking the interaction and fusion between local and global features. Moreover, the lack of reference information also limits their performance. Thus, we propose a no-reference OIQA model which consists of three novel modules, including a bidirectional pseudo-reference module, a Mamba-based global feature extraction module, and a multi-scale local-global feature aggregation module. Specifically, by considering the image distortion degradation process, a bidirectional pseudo-reference module capturing the error maps on viewports is first constructed to refine the multi-scale local visual features, which can supply rich quality degradation reference information without the reference image. To well complement the local features, the VMamba module is adopted to extract the representative multi-scale global visual features. Inspired by human hierarchical visual perception characteristics, a novel multi-scale aggregation module is built to strengthen the feature interaction and effective fusion which can extract deep semantic information. Finally, motivated by the multi-task managing mechanism of human brain, a multi-task learning module is introduced to assist the main quality assessment task by digging the hidden information in compression type and distortion degree. Extensive experimental results demonstrate that our proposed method achieves the state-of-the-art performance on the no-reference OIQA task compared to other models.
Yun Liu 0009, Huiyu Duan, Yu Zhou 0009, Daoxin Fan, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.3
2025 Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional Videos
abstract
Virtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet.
Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Image Process.2
2025 How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and Model
abstract
Understanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency prediction in ODVs, the joint exploitation of these two modalities has been limited, primarily due to the absence of large-scale audio-visual saliency databases and comprehensive analyses. This paper comprehensively investigates audio-visual attention in ODVs from both subjective and objective perspectives. Specifically, we first introduce a new audio-visual saliency database for omnidirectional videos, termed AVS-ODV database, containing 162 ODVs and corresponding eye movement data collected from 60 subjects under three audio modes including mute, mono, and ambisonics. Based on the constructed AVS-ODV database, we perform an in-depth analysis of how audio influences visual attention in ODVs. To advance the research on audio-visual saliency prediction for ODVs, we further establish a new benchmark based on the AVS-ODV database by testing numerous state-of-the-art saliency models, including visual-only models and audio-visual models. In addition, given the limitations of current models, we propose an innovative omnidirectional audio-visual saliency prediction network (OmniAVS), which is built based on the U-Net architecture, and hierarchically fuses audio and visual features from the multimodal aligned embedding space. Extensive experimental results demonstrate that the proposed OmniAVS model outperforms other state-of-the-art models on both ODV AVS prediction and traditional AVS prediction tasks. The AVS-ODV database and the OmniAVS model are available at: https://github.com/IntMeGroup/AVS-ODV.
Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Xilei Zhu, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Image Process.2
2025 Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and Model
abstract
Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively. Specifically, the proposed method utilizes the projection videos of text-to-3D assets to extract 3D shape, texture and text-asset correspondence features, then fuses them to calculate the final three preference scores respectively. Extensive experimental results demonstrate the effectiveness of the proposed T23DAQA method in evaluating the quality of AI generated 3D asset, which is more consistent with human perception. To the best of our knowledge, this is the first work that studies the problem of text-guided 3D generation quality assessment, and our database and codes will be released to facilitate future research.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
IEEE Trans. Multim.2
2025 Quality-Guided Skin Tone Enhancement for Portrait Photography
abstract
In recent years, learning-based color and tone enhancement methods for photos have become increasingly popular. However, most learning-based image enhancement methods just learn a mapping from one distribution to another based on one dataset, lacking the ability to adjust images continuously and controllably. It is important to enable the learning-based enhancement models to adjust an image continuously, since in many cases we may want to get a slighter or stronger enhancement effect rather than one fixed adjusted result. In this paper, we propose a quality-guided image enhancement paradigm that enables image enhancement models to learn the distribution of images with various quality ratings. By learning this distribution, image enhancement models can associate image features with their corresponding perceptual qualities, which can be used to adjust images continuously according to different quality scores. To validate the effectiveness of our proposed method, a subjective quality assessment experiment is first conducted, focusing on skin tone adjustment in portrait photography. Guided by the subjective quality ratings obtained from this experiment, our method can adjust the skin tone corresponding to different quality requirements. Furthermore, an experiment conducted on 10 natural raw images corroborates the effectiveness of our model in situations with fewer subjects and fewer shots, and also demonstrates its general applicability to natural images.
Shiqi Gao, Huiyu Duan, Xinyue Li 0001, Yicong Peng, Qihang Xu, Yuanyuan Chang, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Multim.2
2025 Learning to Generate Realistic Images for Bit-Depth Enhancement via Camera Imaging Processing
abstract
With the prevalence of advanced displays devices, many attempts have been successfully made in bit-depth enhancement (BDE) to restore the low bit-depth (LBD) images to visually pleasant high bit-depth (HBD) images. However, most methods are still far from satisfactory when addressing real-world LBD images owing to their heavy dependence on LBD-HBD data pairs through direct pixel quantization. Therefore, in this paper, we propose a novel network dubbed RealGAN to generate real-world LBD images by simulating the complex quantization procedure in camera imaging process. Particularly, we design a two-mode differentiable quantization block embedded in the synthesis network facilitating adaptively simulation of the complicated quantization distortions. Furthermore, a simple residual group network is proposed in order to learn the distribution of degradation and non-linear processing in the Image Signal Processing (ISP) pipeline. In the absence of paired HBD and LBD data, the synthesis model is trained end-to-end within the generative adversarial framework using non-paired LBD and HBD images. Finally, we demonstrate that a series of BDE models can benefit from the proposed synthetic dataset and exhibit improved visual quality with sharper edges and finer textures on real-world scenes compared with the original versions trained on directly quantized LBD-HBD pairs.
Jing Liu 0002, Huiyu Duan, Yuting Su 0001, Guangtao Zhai
IEEE Trans. Multim.3
2025 Explain Vision Focus: Blending Human Saliency Into Synthetic Face Images
abstract
Synthetic faces have been extensively researched and applied in various fields, such as face parsing and recognition. Compared to real face images, synthetic faces engender more controllable and consistent experimental stimuli due to the ability to precisely merge expression animations onto the facial skeleton. Accordingly, we establish an eye-tracking database with 780 synthetic face images and fixation data collected from 22 participants. The use of synthetic images with consistent expressions ensures reliable data support for exploring the database and determining the following findings: (1) A correlation study between saliency intensity and facial movement reveals that the variation of attention distribution within facial regions is mainly attributed to the movement of the mouth. (2) A categorized analysis of different demographic factors demonstrates that the bias towards salient regions aligns with differences in some demographic categories of synthetic characters. In practice, inference of facial saliency distribution is commonly used to predict the regions of interest for facial video-related applications. Therefore, we propose a benchmark model that accurately predicts saliency maps, closely matching the ground truth annotations. This achievement is made possible by utilizing channel alignment and progressive summation for feature fusion, along with the incorporation of Sinusoidal Position Encoding. The ablation experiment also demonstrates the effectiveness of our proposed model. We hope that this paper will contribute to advancing the photorealism of generative digital humans.
Kaiwei Zhang, Dandan Zhu 0001, Xiongkuo Min, Huiyu Duan, Guangtao Zhai
IEEE Trans. Multim.4
2025 ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
abstract
With the development of eXtended Reality (XR), photo capturing and display technology based on head-mounted displays (HMDs) have experienced significant advancements and gained considerable attention. Egocentric spatial images and videos are emerging as a compelling form of stereoscopic XR content. The assessment for the Quality of Experience (QoE) of XR content is important to ensure a high-quality viewing experience. Different from traditional 2D images, egocentric spatial images present challenges for perceptual quality assessment due to their special shooting, processing methods, and stereoscopic characteristics. However, the corresponding image quality assessment (IQA) research for egocentric spatial images is still lacking. In this paper, we establish the Egocentric Spatial Images Quality Assessment Database (ESIQAD), the first IQA database dedicated for egocentric spatial images as far as we know. Our ESIQAD includes 500 egocentric spatial images and the corresponding mean opinion scores (MOSs) under three display modes, including 2D display, 3D-window display, and 3D-immersive display. Based on our ESIQAD, we propose a novel mamba2-based multi-stage feature fusion model, termed ESIQAnet, which predicts the perceptual quality of egocentric spatial images under the three display modes. Specifically, we first extract features from multiple visual state space duality (VSSD) blocks, then apply cross attention to fuse binocular view information and use transposed attention to further refine the features. The multi-stage features are finally concatenated and fed into a quality regression network to predict the quality score. Extensive experimental results demonstrate that the ESIQAnet outperforms 22 state-of-the-art IQA models on the ESIQAD under all three display modes. The database and code are available at https://github.com/IntMeGroup/ESIQA.
Xilei Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Vis. Comput. Graph.3
2024 UniProcessor: A Text-Induced Unified Low-Level Image Processor
Huiyu Duan, Xiongkuo Min, Sijing Wu, Wei Shen 0002, Guangtao Zhai
ECCV (67)1
2024 AIGCOIQA2024: Perceptual Quality Assessment of AI Generated Omnidirectional Images
abstract
[?]In recent years, the rapid advancement of Artificial Intelligence Generated Content (AIGC) has attracted widespread attention. Among the AIGC, AI generated omnidirectional images hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications, hence omnidirectional AIGC techniques have also been widely studied. AI-generated omnidirectional images exhibit unique distortions compared to natural omnidirectional images, however, there is no dedicated Image Quality Assessment (IQA) criteria for assessing them. This study addresses this gap by establishing a large-scale AI generated omnidirectional image IQA database named AIGCOIQA2024 and constructing a comprehensive benchmark. We first generate 300 omnidirectional images based on 5 AIGC models utilizing 25 text prompts. A subjective IQA experiment is conducted subsequently to assess human visual preferences from three perspectives including quality, comfortability, and correspondence. Finally, we conduct a benchmark experiment to evaluate the performance of state-of-the-art IQA models on our database. The AIGCOIQA2024 database is released to facilitate future research on https://github.com/IntMeGroup/AIGCOIQA.
Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICIP2
2024 PrefIQA: Human Preference Learning for AI-generated Image Quality Assessment
abstract
Despite recent advancements in generative models, the variation in image quality remains a significant concern. To tackle this issue, we propose PrefIQA, an effective human preference learning metric, which can better evaluate the quality of AI-generated images. PrefIQA consists of two units, namely Feature Extraction Unit and Feature Fusion Unit. In Feature Extraction Unit, we introduce a prompt-segmentation module to divide prompts into multiple phrases, enabling a more detailed evaluation of the alignment between images and texts. In Feature Fusion Unit, we introduce a modality-fusion module, which effectively mixes text features and image features to improve the overall performance. In the experiment part, extensive experiments are conducted, demonstrating that PrefIQA surpasses existing text-to-image alignment metrics. We believe that PrefIQA’s proposal would facilitate researches on AI-generated image quality assessment, and make a valuable contribution to the field of text-to-image generation.
Hengjian Gao, Kaiwei Zhang, Wei Sun 0029, Chunyi Li 0001, Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
ISCAS5
2024 MMHead: Towards Fine-grained Multi-modal 3D Facial Animation
abstract
3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation, especially text-guided 3D facial animation is rarely explored due to the lack of multi-modal 3D facial animation dataset. To fill this gap, we first construct a large-scale multi-modal 3D facial animation dataset, MMHead, which consists of 49 hours of 3D facial motion sequences, speech audios, and rich hierarchical text annotations. Each text annotation contains abstract action and emotion descriptions, fine-grained facial and head movements (i.e., expression and head pose) descriptions, and three possible scenarios that may cause such emotion. Concretely, we integrate five public 2D portrait video datasets, and propose an automatic pipeline to 1) reconstruct 3D facial motion sequences from monocular videos; and 2) obtain hierarchical text annotations with the help of AU detection and ChatGPT. Based on the MMHead dataset, we establish benchmarks for two new tasks: text-induced 3D talking head animation and text-to-3D facial motion generation. Moreover, a simple but efficient VQ-VAE-based method named MM2Face is proposed to unify the multi-modal information and generate diverse and plausible 3D facial motions, which achieves competitive results on both benchmarks. Extensive experiments and comprehensive analysis demonstrate the significant potential of our dataset and benchmarks in promoting the development of multi-modal 3D facial animation. The dataset will be released at: https://wsj-sjtu.github.io/MMHead/.
Sijing Wu, Yichao Yan, Huiyu Duan, Ziwei Liu 0002, Guangtao Zhai
ACM Multimedia4
2024 Perceptual Skin Tone Color Difference Measurement for Portrait Photography
abstract
In portrait photography, measuring the perceptual color differences (CDs) of skin tone is significant. Many studies have documented that the perception of skin tone is characteristically different from that of other colors. However, most existing CD measures are proposed based on psychophysical data of uniform color patches or natural images, and do not generalize well to the measurement of skin tone. In this paper, we construct the first large-scale portrait dataset for perceptual skin tone CD assessment and conduct psychophysical experiments to collect 160,000 perceptual CD judgments for 40,000 image triplets. Based on this dataset, we propose a deep skin tone CD measure for portrait photography. Extensive experiments demonstrate that our measure substantially outperforms existing CD measures on the problem of assessing skin tone CDs. The constructed dataset and code will be released to facilitate future research.
Shiqi Gao, Huiyu Duan, Qihang Xu, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
VCIP2
2024 MVBind: Self-Supervised Music Recommendation for Videos via Embedding Space Binding
abstract
Recent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which can significantly influence the emotions of the viewers. However, at present, the background music of short videos is generally chosen by the video producer, and there is a lack of automatic music recommendation methods for short videos. This paper introduces MVBind, an innovative Music-Video embedding space Binding model for cross-modal retrieval. MVBind operates as a self-supervised approach, acquiring inherent knowledge of intermodal relationships directly from data, without the need of manual annotations. Additionally, to compensate the lack of a corresponding musical-visual pair dataset for short videos, we construct a dataset, SVM-10K (Short Video with Music-10K), which mainly consists of meticulously selected short videos. On this dataset, MVBind manifests significantly improved performance compared to other baseline methods. The database and code are available at: https://github.com/IntMeGroup/MVBind.
Jiajie Teng, Huiyu Duan, Yucheng Zhu, Sijing Wu, Guangtao Zhai
VCIP2
2024 Perceptual video quality assessment: a survey
abstract
Abstract Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the advancement of Internet communication and cloud service technology, video content and traffic are growing exponentially, which further emphasizes the requirement for accurate and rapid assessment of video quality. Therefore, numerous subjective and objective video quality assessment studies have been conducted over the past two decades for both generic videos and specific videos such as streaming, user-generated content, 3D, virtual and augmented reality, high dynamic range, high frame rate, audio-visual, etc. This survey provides an up-to-date and comprehensive review of these video quality assessment studies. Specifically, we first review the subjective video quality assessment methodologies and databases, which are necessary for validating the performance of video quality metrics. Second, the objective video quality assessment measures for general purposes are categorized and surveyed according to the methodologies utilized in the quality measures. Third, we overview the objective video quality assessment measures for specific applications and emerging topics. Finally, the performance of the state-of-the-art video quality assessment measures is compared and analyzed. This survey provides a systematic overview of both classical works and recent progress in the realm of video quality assessment, which can help other researchers quickly access the field and conduct relevant research.
Xiongkuo Min, Huiyu Duan, Wei Sun 0029, Yucheng Zhu, Guangtao Zhai
Sci. China Inf. Sci.2
2024 Boosting power line inspection in bad weather: Removing weather noise with channel-spatial attention-based UNet
Yaocheng Li, Qinglin Qian, Huiyu Duan, Xiongkuo Min, Yongpeng Xu, Xiuchen Jiang
Multim. Tools Appl.3
2024 How is Visual Attention Influenced by Text Guidance? Database and Model
abstract
The analysis and prediction of visual attention have long been crucial tasks in the fields of computer vision and image processing. In practical applications, images are generally accompanied by various text descriptions, however, few studies have explored the influence of text descriptions on visual attention, let alone developed visual saliency prediction models considering text guidance. In this paper, we conduct a comprehensive study on text-guided image saliency (TIS) from both subjective and objective perspectives. Specifically, we construct a TIS database named SJTU-TIS, which includes 1200 text-image pairs and the corresponding collected eye-tracking data. Based on the established SJTU-TIS database, we analyze the influence of various text descriptions on visual attention. Then, to facilitate the development of saliency prediction models considering text influence, we construct a benchmark for the established SJTU-TIS database using state-of-the-art saliency models. Finally, considering the effect of text descriptions on visual attention, while most existing saliency models ignore this impact, we further propose a text-guided saliency (TGSal) prediction model, which extracts and integrates both image features and text features to predict the image saliency under various text-description conditions. Our proposed model significantly outperforms the state-of-the-art saliency models on both the SJTU-TIS database and the pure image saliency databases in terms of various evaluation metrics. The SJTU-TIS database and the code of the proposed TGSal model will be released at: https://github.com/IntMeGroup/TGSal.
Xiongkuo Min, Huiyu Duan, Guangtao Zhai
IEEE Trans. Image Process.3
2023 Audio-Visual Saliency for Omnidirectional Videos
Xilei Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Li Chen 0021, Xiongkuo Min, Guangtao Zhai
ICIG (5)3
2023 The Influence of Text-guidance on Visual Attention
abstract
Visual attention analysis and prediction have long been important tasks in computer vision and image processing. However, images often come along with various text descriptions in real applications, while the influence of these text-guidances on the visual saliency of corresponding images have rarely been studied. Therefore, in this paper, we mainly focus on the problem of whether and how the text-guidance influences the visual attention during image viewing, and perform subjective experiments, qualitative and quantitative comparisons as well as model evaluations on this new task. Specifically, we first conduct eye tracking experiments on 300 images under text-visual (TV) and visual (V) test conditions, respectively. Based on the subjective experiments, we perform qualitative and quantitative comparisons between the visual attention data collected under TV and V conditions, and conclude that the text-guidance can significantly influence the visual attention, especially when the text-described target is a non-salient object. Finally, we evaluate the existing saliency models on our database, and find that existing models cannot well handle this text-induced saliency prediction task. Our constructed database will be publicly available to facilitate future research.
Xiongkuo Min, Huiyu Duan, Guangtao Zhai
ISCAS3
2023 Develop Then Rival: A Human Vision-Inspired Framework for Superimposed Image Decomposition
abstract
A single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a “develop-then-rival” process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. However, separating individual image views from a single superimposed image has been an important but challenging task in computer vision area for a long time. In this paper, we propose a human vision-inspired framework for single superimposed image decomposition. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods. The proposed method also achieves state-of-the-art results on related applications including single image reflection removal, single image rain removal, single image shadow removal, and illumination correction,etc., which validates the generalization of the framework.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Yuan Tian 0017, Jae-Hyun Jung, Xiaokang Yang 0001, Guangtao Zhai
IEEE Trans. Multim.1
2022 End-to-End Human-Gaze-Target Detection with Transformers
abstract
In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture.
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
CVPR3
2022 Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002
ECCV (4)3
2022 A Unified Two-Stage Model for Separating Superimposed Images
abstract
A single superimposed image containing two image views causes visual confusion for both human vision and computer vision. Human vision needs a "develop-then-rival" process to decompose the superimposed image into two individual images, which effectively suppresses visual confusion. In this paper, we propose a human vision-inspired framework for separating superimposed images. We first propose a network to simulate the development stage, which tries to understand and distinguish the semantic information of the two layers of a single superimposed image. To further simulate the rivalry activation/suppression process in human brains, we carefully design a rivalry stage, which incorporates the original mixed input (superimposed image), the activated visual information (outputs of the development stage) together, and then rivals to get images without ambiguity. Experimental results show that our novel framework effectively separates the superimposed images and significantly improves the performance with better output quality compared with state-of-the-art methods.
Huiyu Duan, Xiongkuo Min, Wei Shen 0002, Guangtao Zhai
ICASSP1
2022 Saliency in Augmented Reality
abstract
With the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency.
Huiyu Duan, Wei Shen 0002, Xiongkuo Min, Danyang Tu, Jing Li 0026, Guangtao Zhai
ACM Multimedia1
2022 MRIQA: Subjective Method and Objective Model for Magnetic Resonance Image Quality Assessment
abstract
Magnetic Resonance Imaging (MRI) is widely used for medical diagnosis, staging and follow-up of disease. However, MRI images may have artifacts due to various reasons such as patient movement or machine distortion, which may be unintentionally introduced during the procedure of medical image acquisition, processing, etc. These artifacts may affect the effectiveness of diagnosis or even cause false diagnosis. To solve this problem, we propose a general medical image quality assessment (MIQA) methodology, including subjective MIQA procedures and objective MIQA algorithms. We further apply this methodology to MRI images in this paper due to its widespread use in practical applications. We first establish a magnetic resonance imaging quality assessment (MRIQA) database, which contains 3809 MRI images. Then a subjective image quality assessment experiment is conducted by expert doctors according to the diagnostic value of these images, which split all MRI images into 1285 low quality images and 2524 high quality images. We then conduct a baseline deep learning experiment, and propose an attention based MIQANet model to automatically separate MRI images into high quality and low quality based on their diagnosis value. Our proposed method achieves a great quality assessment accuracy of 96.59%. The constructed MRIQA database and proposed MIQA model will be public available to further promote medical IQA research.
Fang Liu 0001, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
VCIP3
2022 Viewing Behavior Supported Visual Saliency Predictor for 360 Degree Videos
abstract
In virtual reality (VR), correct and precise estimations of user’s visual fixations and head movements can enhance the quality of experience by allocating more computation resources for analysing and rendering on the areas of interest. However, there is insufficient research about understanding the visual exploration of users when modeling VR visual attention. To bridge the gap between the saliency prediction for traditional 2D content and omnidirectional content, we construct the visual attention dataset and propose the visual saliency prediction framework for panoramic videos. Around the instantaneous viewing behavior, we propose a traditional method to adapt 2D saliency models and design a CNN-based model to better predict visual saliency. In the proposed traditional model, mechanism of visual attention and viewing behaviors are considered in the computation of edge weights on graphs which are interpreted as Markov chains. The fraction of the visual attention that is diverted to each high-clarity vision (HCV) area is estimated through equilibrium distribution of this chain. We also propose the Graph-Based CNN model. The RGB channel and optical flow form the spatial-temporal units of HCVs, from which node feature vectors are extracted. Graph convolution is used to learn the mutual information between node feature vectors of HCVs and retain geometric information. Then feature vectors are aligned according to geometry structure of equirectangular format, and the feature decoder maps the aligned feature maps to the data distribution. We also construct the dynamic omnidirectional monocular (DOM) saliency dataset with 64 diverse videos evaluated by 28 people. The subjective results show that the instantaneous viewing behavior is important in the VR experience. Extensive experiments are conducted on the dataset and the results demonstrate the effectiveness of the proposed framework. The dataset will be released to facilitate the future studies related to visual saliency prediction for 360-degree contents.
Yucheng Zhu, Guangtao Zhai, Yiwei Yang 0007, Huiyu Duan, Xiongkuo Min, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Confusing Image Quality Assessment: Toward Better Augmented Reality Experience
abstract
With the development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary value of AR is to promote the fusion of digital contents and real-world environments, however, studies on how this fusion will influence the Quality of Experience (QoE) of these two components are lacking. To achieve better QoE of AR, whose two layers are influenced by each other, it is important to evaluate its perceptual quality first. In this paper, we consider AR technology as the superimposition of virtual scenes and real scenes, and introduce visual confusion as its basic theory. A more general problem is first proposed, which is evaluating the perceptual quality of superimposed images, i.e., confusing image quality assessment. A ConFusing Image Quality Assessment (CFIQA) database is established, which includes 600 reference images and 300 distorted images generated by mixing reference images in pairs. Then a subjective quality perception experiment is conducted towards attaining a better understanding of how humans perceive the confusing images. Based on the CFIQA database, several benchmark models and a specifically designed CFIQA model are proposed for solving this problem. Experimental results show that the proposed CFIQA model achieves state-of-the-art performance compared to other benchmark models. Moreover, an extended ARIQA study is further conducted based on the CFIQA study. We establish an ARIQA database to better simulate the real AR application scenarios, which contains 20 AR reference images, 20 background (BG) reference images, and 560 distorted images generated from AR and BG references, as well as the correspondingly collected subjective quality ratings. Three types of full-reference (FR) IQA benchmark variants are designed to study whether we should consider the visual confusion when designing corresponding IQA algorithms. An ARIQA metric is finally proposed for better evaluating the perceptual quality of AR images. Experimental results demonstrate the good generalization ability of the CFIQA model and the state-of-the-art performance of the ARIQA model. The databases, benchmark models, and proposed metrics are available at: https://github.com/DuanHuiyu/ARIQA.
Huiyu Duan, Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Xiaokang Yang 0001, Patrick Le Callet
IEEE Trans. Image Process.1
2021 Muiqa: Image Quality Assessment Database And Algorithm For Medical Ultrasound Images
abstract
In the process of medical image acquisition, medical images may be blurred or ghosted due to machine noise, electromagnetic interference, man-made disturbance, etc. This can result in poor image quality and severely affect the diagnosis accuracy and confidence of doctors. IntraVascular UltraSound (IVUS) is an important supplementary method for the diagnosis of coronary angiography. IVUS images can be distorted for many reasons and some severe distortions can affect diagnosis confidence. However, existing manual medical image quality control method is extremely time-consuming and requires a lot of manpower. To solve this problem, we first construct an Medical UltraSound Image Quality Assessment (MUIQA) database, which consists of 10766 IVUS images with quality labels given by professional doctors. Then we propose a deep-learning network to automatically distinguish the low, medium and high level images from each other. We achieve good classification accuracy of 96.34% on the testing set.
Xiongkuo Min, Huiyu Duan, Yucheng Zhu, Guangtao Zhai
ICIP3
2021 Accurate Compensation Makes the World More Clear for the Visually Impaired
abstract
Visual impairment is one of the most serious social and public health problems in the world, therefore, it is of great theoretical and practical significance to study the image enhancement algorithms for the visually impaired, which is the basis for the development of assistive devices. In this paper, a general deep learning based image enhancement framework for the visually impaired is proposed, which can be used to enhance images to compensate for any visually impaired symptom that can be modeled. Take central vision loss as an example, we first model the central vision loss based on the contrast sensitivity function (CSF) specified by clinical indicator Pelli-Robson score and logMAR visual acuity, and then use the proposed framework to generate an image enhancement method aiming at compensating for the central vision loss. Both the simulation experiment and the patient experiment show the superiority of the proposed image enhancement method designed for the central vision loss, which also validates the effectiveness of the proposed framework.
Sijing Wu, Huiyu Duan, Xiongkuo Min, Danyang Tu, Guangtao Zhai
ICIP2
2020 Identifying Children with Autism Spectrum Disorder Based on Gaze-Following
abstract
This paper presents a novel method to identify children with Autism Spectrum Disorder (ASD) based on the stimuli with gaze-following. Individuals with ASD are characterized by having atypical visual attention patterns, especially in social scenes. Gaze-following is considered to be a key element in understanding social scenarios, and it is reasonable to use stimuli with gaze-following to identify the children with ASD. Thus in this paper, we first construct a dataset of eye movements in gaze-following scenes for children with ASD (i.e., GazeFollow4ASD dataset), including 300 images with gaze-following information inside them and the corresponding eye movement data collected from 8 children with ASD and 10 healthy controls. We propose a novel deep neural network (DNN) model to extract discriminative features and classify children with ASD and healthy controls on single images. The proposed model shows the best performance among all compared methods on all datasets.
Yi Fang 0009, Huiyu Duan, Fangyu Shi, Xiongkuo Min, Guangtao Zhai
ICIP2
2019 A dataset of eye movements for the children with autism spectrum disorder
abstract
Social difficulties are the hallmark features of Autism Spectrum Disorder (ASD) and can lead to atypical visual attention towards stimuli. Eye movements encode rich information about attention and psychological factors of an individual, which could help to characterize the traits of ASD. Learning atypical eye movements of the individuals with ASD towards various stimuli is important and has many application scenarios. However, due to the lack of open datasets, research in this sense is still limited. In this work, we present an open dataset of eye movements of children with Autism Spectrum Disorder. It consists of 300 natural scene images and the corresponding eye movement data collected from 14 children with ASD and 14 healthy controls. In particular, fixation maps and scanpaths are available in the dataset. Based on this dataset, researchers could analyze the visual traits of children with ASD and design specialized visual attention models to promote research in related fields, as well as design specialized models to identify the individuals with ASD. The dataset can be accessed in http://doi.org/10.5281/zenodo.2647418
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Zhaohui Che, Yi Fang 0009, Xiaokang Yang 0001, Jesús Gutiérrez 0001, Patrick Le Callet
MMSys1
2019 Predicting the visual saliency of the people with VIMS
abstract
As is known to us, visually induced motion sickness (VIMS) is often experienced in a virtual environment. Learning the visual attention of people with VIMS contributes to related research in the field of virtual reality (VR) content design and psychology. In this paper, we first construct a saliency prediction for people with VIMS (SPPV) database, which is the first of its kind. The database consists of 80 omnidirectional images and the corresponding eye tracking data collected from 30 individuals. We analyze the performance of five state-of-the-art deep neural networks (DNN)-based saliency prediction algorithms with their original networks and the fine-tuned networks on our database. We predict the atypical visual attention of people with VIMS for the first time and obtain relatively good saliency prediction results for VIMS controls so far.
Guangtao Zhai, Huiyu Duan
VCIP3
2018 Learning to Predict where the Children with Asd Look
abstract
As is known to us, people with Autism Spectrum Disorder (ASD) have atypical visual attention towards stimuli. Learning the visual attention of people especially, children, with ASD contribute to related research in the field of medicine and psychology. In this paper, we first construct a saliency prediction for children with autism (SPCA) database, which is the first of its kind and consists of 500 images and the corresponding eye tracking data collected from 13 different children with ASD. We compare the performance of five state-of-the-art deep neural networks (DNN)-based saliency prediction approaches with their original networks and the fine-tuned networks on our database. We predict the atypical visual attention of children with ASD for the first time and get the best saliency prediction results for individuals with ASD so far.
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yi Fang 0009, Zhaohui Che, Xiaokang Yang 0001, Cheng Zhi, Hua Yang 0001
ICIP1
2018 Perceptual Quality Assessment of Omnidirectional Images
abstract
Omnidirectional images and videos can provide immersive experience of real-world scenes in Virtual Reality (VR) environment. We present a perceptual omnidirectional image quality assessment (IQA) study in this paper since it is extremely important to provide a good quality of experience under the VR environment. We first establish an omnidirectional IQA (OIQA) database, which includes 16 source images and 320 distorted images degraded by 4 commonly encountered distortion types, namely JPEG compression, JPEG2000 compression, Gaussian blur and Gaussian noise. Then a subjective quality evaluation study is conducted on the OIQA database in the VR environment. Considering that humans can only see a part of the scene at one movement in the VR environment, visual attention becomes extremely important. Thus we also track head and eye movement data during the quality rating experiments. The original and distorted omnidirectional images, subjective quality ratings, and the head and eye movement data together constitute the OIQA database. State-of-the-art full-reference (FR) IQA measures are tested on the OIQA database, and some new observations different from traditional IQA are made. The OIQA database will be released to facilitate further research.
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yucheng Zhu, Yi Fang 0009, Xiaokang Yang 0001
ISCAS1