VLDB 2026 Research / reviewers in the wild / expert
Chunyi Li 0001
dblp:192/6758-1
· DBLP profile ↗
45ranked-venue papers
11as first author
45since 2021 · last 2026
0009-0007-0634-1710ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 10 first-author · 34 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 17 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling-up Perceptual Video Quality AssessmentabstractThe data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose OmniVQA, a framework designed to efficiently build high-quality, machine-dominated synthetic multi-modal instruction databases (MIDBs) for VQA. We then scale up to create OmniVQA-Chat-400K, the largest dataset in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we build the OmniVQA-MOS-20K dataset to enhance the model's quantitative quality rating capabilities. We then introduce a complementary training strategy that effectively leverages the knowledge from datasets for different tasks. Furthermore, we propose the OmniVQA-FG (fine-grain)-Benchmark to evaluate the fine-grained performance of models. Our results demonstrate that our models achieve state-of-the-art performance in both tasks. Ziheng Jia, Xiaorong Zhu, Chunyi Li 0001, Jinliang Han, Xiaohong Liu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 4 |
| 2026 | GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal ModelsabstractLarge multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor robotics, etc. To bridge this gap, we introduce GeoX-Bench, a comprehensive Benchmark designed to explore and evaluate the capabilities of LMMs in cross-view Geo-localization and pose estimation. Specifically, GeoX-Bench contains 10,859 panoramic-satellite image pairs spanning 128 cities in 49 countries, along with corresponding 755,976 question-answering (QA) pairs. Among these, 42,900 QA pairs are designated for benchmarking, while the remaining are intended to enhance the capabilities of LMMs. Based on GeoX-Bench, we evaluate the capabilities of 25 state-of-the-art LMMs on cross-view geo-localization and pose estimation tasks, and further explore the empowered capabilities of instruction-tuning. Our benchmark demonstrate that while current LMMs achieve impressive performance in geo-localization tasks, their effectiveness declines significantly on the more complex pose estimation tasks, highlighting a critical area for future improvement, and instruction-tuning LMMs on the training data of GeoX-Bench can significantly improve the cross-view geo-sense abilities. Yushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li 0001, Jing Liu 0002, Xiaohong Liu 0001, Guangtao Zhai |
AAAI | 4 |
| 2026 | Towards versatile multimedia quality assessment for visual communications
Ziheng Jia, Chunyi Li 0001, Yingjie Zhou 0003, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
Sci. China Inf. Sci. | 3 |
| 2026 | Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
Xunchu Zhou, Xiaohong Liu 0001, Yudong Zhang 0001, Tengchuan Kou, Chunyi Li 0001, Haoning Wu 0001, Guangtao Zhai |
Int. J. Comput. Vis. | 7 |
| 2026 | HVS-inspired blind image quality index with prominent perception learning and multi-level progressive integration
Taiyang Chen, Bo Hu 0008, Chunyi Li 0001, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2025 | Redundancy Principles for MLLMs BenchmarksabstractZicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinyu Fang, Chunyi Li 0001, Xiaohong Liu 0001, Xiongkuo Min, Haodong Duan, Kai Chen 0026, Guangtao Zhai |
ACL (1) | 4 |
| 2025 | Image Quality Assessment: From Human to Machine PreferenceabstractImage Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference depends on downstream tasks such as segmentation and detection, rather than visual appeal. Considering the huge gap between human and machine visual systems, this paper proposes the topic: Image Quality Assessment for Machine Vision for the first time. Specifically, we (1) defined the subjective preferences of machines, including downstream tasks, test models, and evaluation metrics; (2) established the Machine Preference Database (MPD), which contains 2.25M fine-grained annotations and 30k reference/distorted image pair instances; (3) verified the performance of mainstream IQA algorithms on MPD. Experiments show that current IQA metrics are human-centric and cannot accurately characterize machine preferences. We sincerely hope that MPD can promote the evolution of IQA from human to machine preferences. Project page is on: https://github.com/lcysyzxdxc/MPD. Chunyi Li 0001, Yuan Tian 0017, Xiaoyue Ling, Haodong Duan, Haoning Wu 0001, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Guo Lu, Weisi Lin, Guangtao Zhai |
CVPR | 1 |
| 2025 | Towards All-in-One Medical Image Re-IdentificationabstractMedical image re-identification (MedReID) is underexplored so far, despite its critical applications in personalized healthcare and privacy protection. In this paper, we introduce a thorough benchmark and a unified model for this problem. First, to handle various medical modalities, we propose a novel Continuous Modality-based Parameter Adapter (ComPA). ComPA condenses medical content into a continuous modality representation and dynamically adjusts the modality-agnostic model with modalityspecific parameters at runtime. This allows a single model to adaptively learn and process diverse modality data. Furthermore, we integrate medical priors into our model by aligning it with a bag of pre-trained medical foundation models, in terms of the differential features. Compared to single-image feature, modeling the inter-image difference better fits the re-identification problem, which involves discriminating multiple images. We evaluate the proposed model against 25 foundation models and 8 large multimodal language models across 11 image datasets, demonstrating consistently superior performance. Additionally, we deploy the proposed MedReID technique to two realworld applications, i.e., history-augmented personalized diagnosis and medical privacy protection. Codes and model is available at https://github.com/tianyuan168326/All-inOne-MedReID-Pytorch. Yuan Tian 0017, Kaiyuan Ji, Rongzhao Zhang, Yankai Jiang 0003, Chunyi Li 0001, Xiaosong Wang 0001, Guangtao Zhai |
CVPR | 5 |
| 2025 | Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image DehazingabstractExisting real-world image dehazing methods primarily attempt to fine-tune pre-trained models or adapt their inference procedures, thus heavily relying on the pre-trained models and associated training data. Moreover, restoring heavily distorted information under dense haze requires generative diffusion models, whose potential in de-hazing remains underutilized partly due to their lengthy sampling processes. To address these limitations, we introduce a novel hazing-dehazing pipeline consisting of a Realistic Hazy Image Generation framework (HazeGen) and a Diffusion-based Dehazing framework (DiffDehaze). Specifically, HazeGen harnesses robust generative diffusion priors of real-world hazy images embedded in a pre-trained text-to-image diffusion model. By employing specialized hybrid training and blended sampling strategies, HazeGen produces realistic and diverse hazy images as high-quality training data for DiffDehaze. To alleviate the inefficiency and fidelity concerns associated with diffusion-based methods, DiffDehaze adopts an Accelerated Fidelity-Preserving Sampling process (AccSamp). The core of AccSamp is the Tiled Statistical Alignment Operation (AlignOp), which can provide a clean and faithful dehazing estimate within a small fraction of sampling steps to reduce complexity and enable effective fidelity guidance. Extensive experiments demonstrate the superior dehazing performance and visual quality of our approach over existing methods. The code is available at https://github.com/ruiyi-w/Learning-Hazing-to-Dehazing. Ruiyi Wang, Yushuo Zheng, Chunyi Li 0001, Shuaicheng Liu, Guangtao Zhai, Xiaohong Liu 0001 |
CVPR | 4 |
| 2025 | Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsabstractWith the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AI-generated content (AIGC), and computer graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of perceptual video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to the performance of human beings. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding. Ziheng Jia, Haoning Wu 0001, Chunyi Li 0001, Zijian Chen 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
CVPR | 4 |
| 2025 | Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision ContentabstractEvaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval. Tengchuan Kou, Shushi Wang, Chunyi Li 0001, Wei Sun 0029, Wei Wang 0213, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai |
CVPR | 4 |
| 2025 | Bidirectional Reference Image Quality Assessment via Content-Quality Correlation ModelingabstractThe emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images. To tackle these issues, a novel Bidirectional Reference Image Quality Assessment (BRIQA) method is proposed, centering on leveraging bidirectional reference images and content-quality correlation modeling. First, triplets of content-aligned low-quality and content-non-aligned high-quality reference images are generated using two easily accessible approaches. To prevent the extraction of redundant information, two feature extractors pretrained through unsupervised contrastive learning are utilized to independently extract content and quality features for the triplet images. Then, an attention-mixer is introduced to further mine quality difference information and enhance content feature. Finally, a content-quality correlation modeler is proposed to model the relationship between quality differences and visual contents. Experimental results on benchmark datasets demonstrate that the BRIQA outperforms existing state-of-the-art methods. Bo Hu 0008, Wenzhi Chen, Chunyi Li 0001, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001 |
ICASSP | 3 |
| 2025 | A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality AssessmentabstractWide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the perception and experience, which may seriously hinder its application in fields such as competitive sports. Up to now, few explorations focus on the quality assessment issue of wide-angle video. This deficiency primarily stems from the absence of a specialized dataset for wide-angle videos. To bridge this gap, we construct the first Multi-annotated and multi-modal Wide-angle Video quality assessment (MWV) dataset. Then, the performances of state-of-the-art video quality methods on the MWV dataset are investigated by inter-dataset testing and intra-dataset testing. Experimental results show that these methods impose significant limitations on their applicability. Bo Hu 0008, Chunyi Li 0001, Lihuo He, Leida Li, Xinbo Gao 0001 |
ICASSP | 3 |
| 2025 | HazeCLIP: Towards Language Guided Real-World Image DehazingabstractExisting methods have achieved remarkable performance in image dehazing, particularly on synthetic datasets. However, they often struggle with real-world hazy images due to domain shift, limiting their practical applicability. This paper introduces HazeCLIP, a language-guided adaptation framework designed to enhance the real-world performance of pre-trained dehazing networks. Inspired by the Contrastive Language-Image Pre-training (CLIP) model’s ability to distinguish between hazy and clean images, we leverage it to evaluate dehazing results. Combined with a region-specific dehazing technique and tailored prompt sets, the CLIP model accurately identifies hazy areas, providing a high-quality, human-like prior that guides the fine-tuning process of pre-trained networks. Extensive experiments demonstrate that HazeCLIP achieves state-of-the-art performance in real-word image dehazing, evaluated through both visual quality and image quality assessment metrics. Codes are available at https://github.com/Troivyn/HazeCLIP. Ruiyi Wang, Wenhao Li 0018, Xiaohong Liu 0001, Chunyi Li 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 4 |
| 2025 | Information Density Principle for MLLM BenchmarksabstractWith the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and whether the test results meet their requirements. Therefore, we propose a critical principle of Information Density, which examines how much insight a benchmark can provide for the development of MLLMs. We characterize it from four key dimensions: (1) Fallacy, (2) Difficulty, (3) Redundancy, (4) Diversity. Through a comprehensive analysis of more than 10,000 samples, we measured the information density of 19 MLLM benchmarks. Experiments show that using the latest benchmarks in testing can provide more insight compared to previous ones, but there is still room for improvement in their information density. We hope this principle can promote the development and application of future MLLM benchmarks. Project page: https://github.com/lcysyzxdxc/bench4bench Chunyi Li 0001, Xiaozhe Li, Yuan Tian 0017, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Haodong Duan, Kai Chen 0026, Guangtao Zhai |
ICCV | 1 |
| 2025 | Semantics Versus Identity: A Divide-and-Conquer Approach Towards Adjustable Medical Image De-Identification
Yuan Tian 0017, Rongzhao Zhang, Zijian Chen 0001, Yankai Jiang 0003, Chunyi Li 0001, Fang Yan 0002, Qiang Hu 0003, Xiaosong Wang 0001, Guangtao Zhai |
ICCV | 6 |
| 2025 | A-Bench: Are LMMs Masters at Evaluating AI-generated Images?abstractHow to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, the precision and validity of which are still questionable. Furthermore, traditional benchmarks often utilize mostly natural-captured content rather than AIGIs to test the abilities of LMMs, leading to a noticeable gap for AIGIs. Therefore, we introduce **A-Bench** in this paper, a benchmark designed to diagnose *whether LMMs are masters at evaluating AIGIs*. Specifically, **A-Bench** is organized under two key principles: 1) Emphasizing both high-level semantic understanding and low-level visual quality perception to address the intricate demands of AIGIs. 2) Various generative models are utilized for AIGI creation, and various LMMs are employed for evaluation, which ensures a comprehensive validation scope. Ultimately, 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts. We hope that **A-Bench** will significantly enhance the evaluation process and promote the generation quality for AIGIs. Haoning Wu 0001, Chunyi Li 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Zijian Chen 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ICLR | 3 |
| 2025 | LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and BenchmarkabstractIn recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly. Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ISCAS | 8 |
| 2025 | VQA2: Visual Question Answering for Video Quality AssessmentabstractThe advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field in low-level visual perception, focused initially on quantitative video quality scoring. However, driven by advances in LMMs, it is now progressing toward more holistic visual quality understanding tasks. Recent studies in the image domain have demonstrated that Visual Question Answering (VQA) can markedly enhance low-level visual quality evaluation. Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA² Instruction Dataset-the first visual question answering instruction dataset that focuses on video quality assessment. This dataset consists of 3 subsets and covers various video types, containing 157,755 instruction question-answer pairs. Then, leveraging this foundation, we present the VQA² series models. The VQA² series models interleave visual and motion tokens to enhance the perception of spatial-temporal quality details in videos. We conduct extensive experiments on video quality scoring and understanding tasks, and results demonstrate that the VQA² series models achieve excellent performance in both tasks. Notably, our final model, the VQA²-Assistant, exceeds the renowned GPT-4o in visual quality understanding tasks while maintaining strong competitiveness in quality scoring tasks. Our work provides a foundation and feasible approach for integrating low-level video quality assessment and understanding with LMMs. Ziheng Jia, Jiaying Qian, Haoning Wu 0001, Wei Sun 0029, Chunyi Li 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 6 |
| 2025 | Low-light Image Enhancement Quality Assessment: A Real-World Dataset and An Objective MethodabstractLow-light Image Enhancement (LIE) technology adaptively improves brightness while preserving texture details and suppressing noise artifacts, thereby reducing visual degradation caused by insufficient illumination. While deep learning-based image enhancement algorithms have made significant progress, a key gap remains in establishing standardized methods for fairly evaluating and comparing their performance. To bridge this gap, this paper systematically investigates enhanced low-light image quality assessment from both subjective and objective dimensions. First, we introduce a Real-world Low-light Image Enhancement quality assessment dataset (RLIE), which contains 1540 images from 154 scenarios, each with a subjective score given by the subjects. Based on this, we propose a low light enhanced image quality assessment method based on Multi-level Illumination Injection and Hierarchical Discrepancy Perception (MIIHDP). The core idea of this method is to hierarchically inject separated illumination information into the feature extraction process, then tailor the processing of difference information at different scales to obtain a more comprehensive representation. Finally, extensive statistical analyses demonstrate the rationality of the proposed RLIE dataset, and experimental results show the superior performance of the proposed MIIHDP compared with state-of-the-arts. Our dataset and code are released at: https://github.com/BoHu90/RLIE. Chunyi Li 0001, Bo Hu 0008, Taiyang Chen, Leida Li, Lihuo He, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2025 | Towards a New Paradigm of Visual Signal CompressionabstractUltra-low bitrate image compression is a challenging and demand- ing topic. With the development of Large Multimodal Models (LMMs), a Cross Modality Compression (CMC) paradigm of Image-Text- Image has emerged. Compared with traditional codecs, this semantic- level compression can reduce image data size to 0.1% or even lower, which has strong potential applications. However, CMC has cer- tain defects in consistency with the original image and perceptual quality. To inspire insights into such a problem, we introduce CMC- Bench, a benchmark of the cooperative performance of Image-to- Text (I2T) and Text-to-Image (T2I) models for image compression. This benchmark covers 18,000 and 40,000 images respectively to verify 6 mainstream I2T and 12 T2I models, including 160,000 sub- jective preference scores annotated by human experts. At ultra-low bitrates, it proves that the combination of some I2T and T2I models has surpassed the most advanced visual signal codecs; meanwhile, it highlights where LMMs can be further optimized toward the compression task. We encourage LMM developers to participate in this test to promote the evolution of visual signal codec protocols. Chunyi Li 0001, Xiele Wu, Haoning Wu 0001, Donghui Feng 0003, Guo Lu, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 1 |
| 2025 | AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated ContentabstractAI-based image enhancement techniques have been widely adopted in various visual applications, significantly improving the perceptual quality of user-generated content (UGC). However, the lack of specialized quality assessment models has become a significant limiting factor in this field, limiting user experience and hindering the advancement of enhancement methods. While perceptual quality assessment methods have shown strong performance on UGC and AIGC individually, their effectiveness on AI-enhanced UGC (AI-UGC) which blends features from both-remains largely unexplored. To address this gap, we construct AU-IQA, a benchmark dataset comprising 4,800 AI-UGC images produced by three representative enhancement types which include super-resolution, low-light enhancement, and denoising. On this dataset, we further evaluate a range of existing quality assessment models, including traditional IQA methods and large multimodal models. Finally, we provide a comprehensive analysis of how well current approaches perform in assessing the perceptual quality of AI-UGC. The access link to the AU-IQA is https://github.com/WNNGGU/AU-IQA-Dataset. Shushi Wang, Chunyi Li 0001, Han Zhou 0003, Wei Dong 0011, Jun Chen 0005, Guangtao Zhai, Xiaohong Liu 0001 |
ACM Multimedia | 2 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 30 |
| 2025 | MISC: Ultra-Low Bitrate Image Semantic Compression Driven by Large Multimodal ModelabstractWith the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, all existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. During recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC. Chunyi Li 0001, Guo Lu, Donghui Feng 0003, Haoning Wu 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Advancing Zero-Shot Digital Human Quality Assessment Through Text-Prompted EvaluationabstractDigital humans have witnessed extensive applications in various domains, necessitating related quality assessment studies. However, there is a lack of comprehensive digital human quality assessment (DHQA) databases. To address this gap, we propose SJTU-H3D, a subjective quality assessment database specifically designed for full-body digital humans. It comprises 40 high-quality reference digital humans and 1,120 labeled distorted counterparts generated with seven types of distortions. The SJTU-H3D database can serve as a benchmark for DHQA research, allowing evaluation and refinement of processing algorithms. Further, we propose a zero-shot DHQA approach that focuses on no-reference (NR) scenarios to ensure generalization capabilities while mitigating database bias. Our method leverages semantic and distortion features extracted from projections, as well as geometry features derived from the mesh structure of digital humans. Specifically, we employ the Contrastive Language-Image Pre-training (CLIP) model to measure semantic affinity and incorporate the Naturalness Image Quality Evaluator (NIQE) model to capture low-level distortion information. Additionally, we utilize dihedral angles as geometry descriptors to extract mesh features. By aggregating these measures, we introduce the Digital Human Quality Index (DHQI), which demonstrates significant improvements in zero-shot performance. The DHQI can also serve as a robust baseline for DHQA tasks, facilitating advancements in the field. The database and the code are available at https://github.com/zzc-1998/SJTU-H3D. Wei Sun 0029, Yingjie Zhou 0003, Haoning Wu 0001, Chunyi Li 0001, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Image Process. | 5 |
| 2025 | Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified ModelabstractIn recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git . Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | MM-PCQA+: Advancing Multi-Modal Learning for Point Cloud Quality AssessmentabstractThe importance of visual quality in point clouds has been significantly underlined due to the rapid rise in 3D vision applications which aim to deliver affordable and superior user experiences. Reviewing the evolution of point cloud quality assessment (PCQA), it’s observed that visual quality evaluation typically employs single-modal data, either sourced from 2D projections or the 3D point clouds. The 2D projections possess abundant texture and semantic information while they are heavily reliant on viewpoints. In contrast, 3D point clouds are more reactive to geometric distortions and viewpoint-invariant. Consequently, to maximize the benefits of both point cloud and image modalities, we present an advanced no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA+) metric. Specifically, we divide the point clouds into sub-models to reflect local geometric distortions such as point shifting and down-sampling. Afterwards, we render the point clouds using a cube-like projection setup and sample the projections of interest using a point-visible-ratio for image feature extraction. In order to fulfill these objectives, the sub-models and projected images are encoded using point-based and image-based neural networks. Lastly, we implement symmetric cross-modal attention to amalgamate multi-modal quality-aware features. Experimental results demonstrate that our metric surpasses all state-of-the-art methods and significantly advances beyond previous no-reference PCQA methods. Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Quality Assessment in the Era of Large Models: A SurveyabstractQuality assessment, which evaluates the visual quality level of multimedia experiences, has garnered significant attention from researchers and has evolved substantially through dedicated efforts. Before the advent of large models, quality assessment typically relied on small expert models tailored for specific tasks. While these smaller models are effective at handling their designated tasks and predicting quality levels, they often lack explainability and robustness. With the advancement of large models, which align more closely with human cognitive and perceptual processes, many researchers are now leveraging the prior knowledge embedded in these large models for quality assessment tasks. This emergence of quality assessment within the context of large models motivates us to provide a comprehensive review focusing on two key aspects: (1) the assessment of large models and (2) the role of large models in assessment tasks. We begin by reflecting on the historical development of quality assessment. Subsequently, we move to detailed discussions of related works concerning quality assessment in the era of large models. Finally, we offer insights into the future progression and potential pathways for quality assessment in this new era. We hope that this survey will enable a rapid understanding of the development of quality assessment in the era of large models and inspire further advancements in the field. Yingjie Zhou 0003, Chunyi Li 0001, Baixuan Zhao, Xiaohong Liu 0001, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation ModelsabstractMulti-modality large language models (MLLMs), as represented by GPT-4V, have introduced a paradigm shift for visual perception and understanding tasks, that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the identification of low-level visual attributes (e.g., clarity, brightness) to the evaluation on image quality, there's still an imperative to further improve the accuracy of MLLMs to substantially alleviate human burdens. To address this, we collect the first dataset consisting of human natural language feedback on low-level vision. Each feedback offers a comprehensive description of an image's low-level visual attributes, culminating in an overall quality assessment. The constructed Q-Pathway dataset includes 58K detailed human feedbacks on 18,973 multi-sourced images with diverse low-level appearance. To ensure MLLMs can adeptly handle diverse queries, we further propose a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct. Experimental results indicate that the Q-Instruct consistently elevates various low-level visual capabilities across multiple base models. We anticipate that our datasets can pave the way for a future that foundation models can assist humans on low-level visual tasks. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Kaixin Xu, Chunyi Li 0001, Jingwen Hou, Guangtao Zhai, Geng Xue, Wenxiu Sun, Qiong Yan, Weisi Lin |
CVPR | 8 |
| 2024 | Towards Open-Ended Visual Quality Comparison
Haoning Wu 0001, Hanwei Zhu, Erli Zhang 0001, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Xiaohong Liu 0001, Guangtao Zhai, Shiqi Wang 0001, Weisi Lin |
ECCV (3) | 7 |
| 2024 | A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital HumansabstractIn an era where 3D Digital Humans (DHs) are becoming increasingly prevalent in fields like gaming, automotive, and the metaverse, the demand for high DH visual quality is rising. This paper presents the first-ever reduced-reference (RR) quality assessment metric tailored specifically for textured mesh DHs, aiming to optimize transmission systems and improve Quality of Experience (QoE) for viewers in resource-constrained environments. Four critical geometric curvature-related attributes and two texture-related indicators are computed, which are then statistically analyzed and utilized in a Support Vector Regression (SVR) model for robust and efficient quality prediction. Experimental results confirm that our method outperforms existing full-reference (FR) metrics, making it an invaluable tool for the future of 3D DHs in various applications. The code is available at https://github.com/zzc-1998/RR-DHQA. Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 3 |
| 2024 | Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionabstractThe rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding**. To address this gap, we present **Q-Bench**, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment. **_a)_** To evaluate the low-level **_perception_** ability, we construct the **LLVisionQA** dataset, consisting of 2,990 diverse-sourced images, each equipped with a human-asked question focusing on its low-level attributes. We then measure the correctness of MLLMs on answering these questions. **_b)_** To examine the **_description_** ability of MLLMs on low-level information, we propose the **LLDescribe** dataset consisting of long expert-labelled *golden* low-level text descriptions on 499 images, and a GPT-involved comparison pipeline between outputs of MLLMs and the *golden* descriptions. **_c)_** Besides these two tasks, we further measure their visual quality **_assessment_** ability to align with human opinion scores. Specifically, we design a softmax-based strategy that enables MLLMs to predict *quantifiable* quality scores, and evaluate them on various existing image quality assessment (IQA) datasets. Our evaluation across the three abilities confirms that MLLMs possess preliminary low-level visual skills. However, these skills are still unstable and relatively imprecise, indicating the need for specific enhancements on MLLMs towards these abilities. We hope that our benchmark can encourage the research community to delve deeper to discover and enhance these untapped potentials of MLLMs. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Chunyi Li 0001, Wenxiu Sun, Qiong Yan, Guangtao Zhai, Weisi Lin |
ICLR | 7 |
| 2024 | Q-Refine: A Perceptual Quality Refiner for AI-Generated ImageabstractWith the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited optimization capabilities for low-quality AIGIs but also brought negative optimization to high-quality AIGIs. To address this issue, a quality-award refiner named Q-Refine is proposed. Based on the preference of the Human Visual System (HVS), Q-Refine uses the Image Quality Assessment (IQA) metric to guide the refining process for the first time, and modify images of different qualities through three adaptive pipelines. Experimental data shows that for mainstream T2I models, Q-Refine can perform effective optimization to AIGIs of different qualities. It can be a general refiner to optimize AIGIs from both fidelity and aesthetic quality levels, thus expanding the application of the T2I generation models. The code is released on https://github.com/Q-Future/Q-Refine. Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
ICME | 1 |
| 2024 | Optimizing Projection-Based Point Cloud Quality Assessment with Human Preferred Viewpoints SelectionabstractViewpoint selection plays a pivotal role in projection-based point cloud quality assessment (PCQA). Generally speaking, sole reliance on a single projection fails to capture adequate quality information, leading to the prevalent use of multi-projection approaches. It is important to recognize that viewpoint selection is significantly influenced by human preferences and viewpoints that align with human predilections exert a greater impact on PCQA. Therefore, we introduce the first viewpoint selection database for PCQA, which comprises 405 distorted point clouds, accompanied by preferred viewpoints collected from humans. Then we propose a novel human preference index, devised from the Visible-Points Ratio and Visible-Color-Entropy Ratio, to guide the selection of viewpoints. Our experimental findings confirm that this human preference index correlates more closely with human preferences than traditional viewpoint selection settings. Moreover, the proposed PCQA method optimized with the human preference index demonstrates competitive performance as well. Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Weisi Lin, Guangtao Zhai |
ICME | 6 |
| 2024 | Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsabstractThe explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released. Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ICML | 6 |
| 2024 | PrefIQA: Human Preference Learning for AI-generated Image Quality AssessmentabstractDespite recent advancements in generative models, the variation in image quality remains a significant concern. To tackle this issue, we propose PrefIQA, an effective human preference learning metric, which can better evaluate the quality of AI-generated images. PrefIQA consists of two units, namely Feature Extraction Unit and Feature Fusion Unit. In Feature Extraction Unit, we introduce a prompt-segmentation module to divide prompts into multiple phrases, enabling a more detailed evaluation of the alignment between images and texts. In Feature Fusion Unit, we introduce a modality-fusion module, which effectively mixes text features and image features to improve the overall performance. In the experiment part, extensive experiments are conducted, demonstrating that PrefIQA surpasses existing text-to-image alignment metrics. We believe that PrefIQA’s proposal would facilitate researches on AI-generated image quality assessment, and make a valuable contribution to the field of text-to-image generation. Hengjian Gao, Kaiwei Zhang, Wei Sun 0029, Chunyi Li 0001, Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ISCAS | 4 |
| 2024 | PAPS-OVQA: Projection-Aware Patch Sampling for Omnidirectional Video Quality AssessmentabstractIn immersive multimedia systems, the perceptual quality model of omnidirectional video is indispensable. However, to cope with its resolution that is several times higher than ordinary video, the existing omnidirectional video quality assessment (OVQA) models require extremely high computational complexity and usually need to transcode the projection into a certain format. Therefore, to assess the perceptual quality of omnidirectional video effectively, we propose Projection-Aware Patch Sampling (PAPS)-OVQA to process its three common projection formats simultaneously while resizing high-resolution video into patches sampled from uniform grids and finally apply Fragment Attention Network (FANet) to perform quality regression. As a result, we avoid the overhead computational cost of projection transcoding and reduce the complexity of the quality model greatly. Experimental data show that PAPS-OVQA guarantees good performance while retaining high efficiency under different projection formats. Chunyi Li 0001, Haoning Wu 0001, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ISCAS | 1 |
| 2024 | T2I-Scorer: Quantitative Evaluation on Text-to-Image Generation via Fine-Tuned Large Multi-Modal ModelsabstractText-to-image (T2I) generation is a pivotal and core interest within the realm of AI content generation. Amid the swift advancements of both open-source (such as Stable Diffusion) and proprietary (for example, DALLE, MidJourney) T2I models, there is a notable absence of a comprehensive and robust quantitative framework for evaluating their output quality. Traditional methods of quality assessment overlook the textual prompts when judging images; meanwhile, the advent of large multi-modal models (LMMs) introduces the capability to incorporate text prompts in evaluations, yet the challenge of fine-tuning these models for precise T2I quality assessment remains unresolved. In our study, we introduce the T2I-Scorer, a novel two-stage training methodology aimed at fine-tuning LMMs for T2I evaluation. For the first stage, we collect 397K GPT-4V-labeled question-answer pairs related to T2I evaluation. Termed as T2I-ITD, the pseudo-labeled dataset is analyzed and examined by human, and used for instruction tuning to improve the LMM's low-level quality perception. The first stage model, T2I-Scorer-IT, has reached superior accuracy on T2I evaluation than all kinds of existing T2I metrics under zero-shot settings. For the second stage, we define an explicit multi-task training scheme to further align the LMM with human opinion scores, and the fine-tuned T2I-Scorer can reach state-of-the-art accuracy on both image quality and image-text alignment perspectives with significant improvements. We anticipate the proposed metrics can serve as a reliable metric to gauge the ability of T2I generation models in the future. We will make code, data, and weights publicly available. Haoning Wu 0001, Xiele Wu, Chunyi Li 0001, Chaofeng Chen, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
ACM Multimedia | 3 |
| 2024 | Subjective-Aligned Dataset and Metric for Text-to-Video Quality AssessmentabstractWith the rapid development of generative models, AI-Generated Content (AIGC) has exponentially increased in daily lives. Among them, Text-to-Video (T2V) generation has received widespread attention. Though many T2V models have been released for generating high perceptual quality videos, there is still lack of a method to evaluate the quality of these videos quantitatively. To solve this issue, we establish the largest-scale Text-to-Video Quality Assessment DataBase (T2VQA-DB) to date. The dataset is composed of 10,000 videos generated by 9 different T2V models, along with each video's corresponding mean opinion score. Based on T2VQA-DB, we propose a novel transformer-based model for subjective-aligned Text-to-Video Quality Assessment (T2VQA). The model extracts features from text-video alignment and video fidelity perspectives, then it leverages the ability of a large language model to give the prediction score. Experimental results show that T2VQA outperforms existing T2V metrics and SOTA video quality assessment models. Quantitative analysis indicates that T2VQA is capable of giving subjective-align predictions, validating its effectiveness. The dataset and code are available at https://github.com/QMME/T2VQA. Tengchuan Kou, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 4 |
| 2024 | G-Refine: A General Quality Refiner for Text-to-Image Generation
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Tengchuan Kou, Chaofeng Chen, Lei Bai 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 1 |
| 2024 | LMM-PCQA: Assisting Point Cloud Quality Assessment with LMMabstractAlthough large multi-modality models (LMMs) have seen extensive exploration and application in various quality assessment studies, their integration into Point Cloud Quality Assessment (PCQA) remains unexplored. Given LMMs' exceptional performance and robustness in low-level vision and quality assessment tasks, this study aims to investigate the feasibility of imparting PCQA knowledge to LMMs through text supervision. To achieve this, we transform quality labels into textual descriptions during the fine-tuning phase, enabling LMMs to derive quality rating logits from 2D projections of point clouds. To compensate for the loss of perception in the 3D domain, structural features are extracted as well. These quality logits and structural features are then combined and regressed into quality scores. Our experimental results affirm the effectiveness of our approach, showcasing a novel integration of LMMs into PCQA that enhances model understanding and assessment accuracy. We hope our contributions can inspire subsequent investigations into the fusion of LMMs with PCQA, fostering advancements in 3D visual quality analysis and beyond. The code is available at https://github.com/zzc-1998/LMM-PCQA. Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Chaofeng Chen, Xiongkuo Min, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 4 |
| 2024 | AGIQA-3K: An Open Database for AI-Generated Image Quality AssessmentabstractWith the rapid advancements of the text-to-image generative model, AI-generated images (AGIs) have been widely applied to entertainment, education, social media, etc. However, considering the large quality variance among different AGIs, there is an urgent need for quality models that are consistent with human subjective ratings. To address this issue, we extensively consider various popular AGI models, generated AGI through different prompts and model parameters, and collected subjective scores at the perceptual quality and text-to-image alignment, thus building the most comprehensive AGI subjective quality database AGIQA-3K so far. Furthermore, we conduct a benchmark experiment on this database to evaluate the consistency between the current Image Quality Assessment (IQA) model and human perception, while proposing StairReward that significantly improves the assessment performance of subjective text-to-image alignment. We believe that the fine-grained subjective scores in AGIQA-3K will inspire subsequent AGI quality models to fit human subjective perception mechanisms at both perception and alignment levels and to optimize the generation result of future AGI models. The database is released on https://github.com/lcysyzxdxc/AGIQA-3k-Database. Chunyi Li 0001, Haoning Wu 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | GMS-3DQA: Projection-Based Grid Mini-patch Sampling for 3D Model Quality AssessmentabstractNowadays, most three-dimensional model quality assessment (3DQA) methods have been aimed at improving accuracy. However, little attention has been paid to the computational cost and inference time required for practical applications. Model-based 3DQA methods extract features directly from the 3D models, which are characterized by their high degree of complexity. As a result, many researchers are inclined towards utilizing projection-based 3DQA methods. Nevertheless, previous projection-based 3DQA methods directly extract features from multi-projections to ensure quality prediction accuracy, which calls for more resource consumption and inevitably leads to inefficiency. Thus, in this article, we address this challenge by proposing a no-reference (NR) projection-based G rid M ini-patch S ampling 3D Model Q uality A ssessment (GMS-3DQA) method. The projection images are rendered from six perpendicular viewpoints of the 3D model to cover sufficient quality information. To reduce redundancy and inference resources, we propose a multi-projection grid mini-patch sampling strategy (MP-GMS), which samples grid mini-patches from the multi-projections and forms the sampled grid mini-patches into one quality mini-patch map (QMM). The Swin-Transformer tiny backbone is then used to extract quality-aware features from the QMMs. The experimental results show that the proposed GMS-3DQA outperforms existing state-of-the-art NR-3DQA methods on the point cloud quality assessment databases for both accuracy and efficiency. The efficiency analysis reveals that the proposed GMS-3DQA requires far less computational resources and inference time than other 3DQA competitors. The code is available at https://github.com/zzc-1998/GMS-3DQA . Wei Sun 0029, Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Zijian Chen 0001, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | A Real-Time Blind Quality-of-Experience Assessment Metric for HTTP Adaptive StreamingabstractIn today’s Internet, HTTP Adaptive Streaming (HAS) is the mainstream standard for video streaming, which switches the bitrate of the video content based on an Adaptive BitRate (ABR) algorithm. An effective Quality of Experience (QoE) assessment metric can provide crucial feedback to an ABR algorithm. However, predicting such real-time QoE on the client side is challenging. The QoE prediction requires high consistency with the Human Visual System (HVS), low latency, and blind assessment, which are difficult to realize together. To address this challenge, we analyzed various characteristics of HAS systems and propose a non-uniform sampling metric to reduce time complexity. Furthermore, we design an effective QoE metric that integrates resolution and rebuffering time as the Quality of Service (QoS), as well as spatiotemporal output from a deep neural network and specific switching events as content information. These reward and penalty features are regressed into quality scores with a Support Vector Regression (SVR) model. Experimental results show that the accuracy of our metric outperforms the mainstream blind QoE metrics by 0.3, and its computing time is only 60% of the video playback, indicating that the proposed metric is capable of providing real-time guidance to ABR algorithms and improving the overall performance of HAS. The QoE model is released on https://github.com/lcysyzxdxc/ASPECT. Chunyi Li 0001, May Lim, Abdelhak Bentaleb, Roger Zimmermann |
ICME | 1 |
| 2022 | A Full- Reference Quality Assessment Metric for Cartoon ImagesabstractCartoon images are illustrations that are typically drawn, sometimes animated, in an unrealistic or semi-realistic style, which are widely applied in multimedia services. However, in some post-production processes as well as transmission systems, cartoon images are inevitably distorted by wrong color arrangement and compression. Therefore, it is urgent to carry out image quality assessment (IQA) metrics to automatically predict the perceptual quality levels of distorted cartoon images. Nevertheless, the existing mainstream IQA metrics are specially developed for natural scene images (NSIs). Due to the statistical difference in structure and color aspects between cartoon images and NSIs, the scores predicted by such metrics are often inconsistent with the human vision system (HVS) for cartoon images. To further improve the performance of cartoon image quality assessment (C-IQA) methods and provide guidance for practical applications, we propose a full-reference (FR) IQA method to tackle the challenge of C-IQA. Specifically, the proposed method extracts edge and texture features to analyze the structural error. Then the moment and entropy of various color spaces are computed to reflect color distortions. Then the features are regressed into quality scores with the assistance of a support vector regression (SVR) model. Experimental results show that our metric outperforms the mainstream FR-IQA metrics, which indicates that the proposed method is more capable of modeling the visual quality loss of cartoon images. Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
MMSP | 1 |