EDBT 2026 Demo / reviewers in the wild / expert
Wei Sun 0029
dblp:09/5042-29
· DBLP profile ↗
101ranked-venue papers
11as first author
93since 2021 · last 2026
0000-0001-8162-1949ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 7 first-author · 74 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 18 since 2021Computer networks · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio-Assisted Face Video Restoration with Temporal and Identity Complementary LearningabstractFace videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between visual and audio features, particularly in the mouth region. Several audio-aided face video restoration methods have been proposed, but they only focus on compression artifact removal. In this paper, we propose a General Audio-assisted face Video restoration Network (GAVN) to address various types of streaming video distortions via identity and temporal complementary learning. Specifically, GAVN first captures inter-frame temporal features in the low-resolution space to restore frames coarsely and save computational cost. Then, GAVN extracts intra-frame identity features in the high-resolution space with the assistance of audio signals and face landmarks to restore more facial details. Finally, the reconstruction module integrates temporal features and identity features to generate high-quality face videos. Experimental results demonstrate that GAVN outperforms the existing state-of-the-art methods on face video compression artifact removal, deblurring, and super-resolution. Yuqin Cao, Wei Sun 0029, Xiaohong Liu 0001, Yulun Zhang 0001, Xiongkuo Min |
AAAI | 3 |
| 2026 | VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningabstractVideo quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision. Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 2 |
| 2026 | A Scalable Benchmark Test Suite for Dynamic Multi-objective Optimization with a Changing Number of ObjectivesabstractDynamic multi-objective optimization with a changing number of objectives has recently attracted increasing attention due to its relevance to real-world problems whose evaluation criteria may evolve over time. However, existing benchmark test suites for this problem setting suffer from a fundamental limitation: when the number of objectives changes, the objective functions themselves also change implicitly. This makes it difficult to isolate and evaluate an algorithm's capability to handle dynamics in the number of objectives alone. In this paper, we analyze this issue in detail and show that several theoretical properties claimed in prior studies rely on an assumption that is violated by commonly used test suites. To address this problem, we propose a scalable benchmark test suite in which the objective functions are fixed throughout the optimization process, while the number of active objectives changes over time. Our benchmark is constructed by defining a maximum-objective problem and dynamically selecting subsets of objectives. To avoid degeneracy issues in classical DTLZ and WFG problems, we adopt Minus-DTLZ and Minus-WFG formulations, in which all objectives are mutually conflicting. Extensive benchmark studies using representative algorithms from the literature demonstrate the usefulness and flexibility of the proposed test suite. Zhiyun Xiao, Shaojiang Wang, Wei Sun 0029 |
PPSN (2) | 6 |
| 2026 | Q-Agent: An MLLM-Driven Framework for Universal Visual Quality Assessment
Peihang Chen, Huiyu Duan, Zitong Xu, Yuqin Cao, Sijing Wu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
QoMEX | 11 |
| 2026 | Assessing Personality Consistency in Large Language Models: A Psychometric Framework for Human-Centric Quality of Experience
Yitian Kou, Dandan Zhu 0001, Wei Sun 0029, Kaiwei Zhang, Xiongkuo Min, Guangtao Zhai |
QoMEX | 3 |
| 2026 | LEIQ-Assessor: Multi-Dimensional Quality Assessment of Low-Light Enhanced Images via Multi-Task Learning
Wei Sun 0029, Yanwei Jiang, Dandan Zhu 0001, Jinqiu Sang, Jikai Xu, Weixia Zhang, Guangtao Zhai |
QoMEX | 1 |
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 35 |
| 2026 | Enhancing blind video quality assessment with rich quality-aware features
Wei Sun 0029, Linhan Cao, Jun Jia, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 1 |
| 2026 | MI3S: A multimodal large language model assisted quality assessment framework for AI-generated talking heads
Yingjie Zhou 0003, Sijing Wu, Jun Jia, Yanwei Jiang, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
Inf. Process. Manag. | 6 |
| 2026 | DHQA-4D: A large-scale dataset and LMM-based metric for dynamic 4D digital human quality assessment
Sijing Wu, Yucheng Zhu, Huiyu Duan, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Pattern Recognit. | 6 |
| 2026 | UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual ContentabstractAs multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A/V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A/V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications. Code are available at https://github.com/charlotte9524/UNQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Long Ye, Weisi Lin, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Subjective and Objective Quality Assessment of Display Content VideosabstractDisplay quality assessment plays a crucial role in evaluating the performance of display devices. However, existing video quality assessment methods primarily target compression-related distortions, failing to capture display-specific degradations including definition loss, color distortions, and motion artifacts that critically affect user subjective experiences during video playback. To address these limitations, we develop a specialized video dataset, namely Video Displaying Quality Assessment Dataset (VDQA), constructed using a DSLR camera with standardized parameter optimization of exposure settings (aperture, ISO sensitivity, and shutter speed). VDQA comprises 250 high-resolution video clips covering diverse content categories, providing a robust foundation for evaluating display devices across multiple quality dimensions. Additionally, we propose a deep learning-based model specifically designed for display quality assessment that employs three complementary pathways to independently evaluate definition, color fidelity, and motion quality. The model integrates Canny edge detection for explicit sharpness measurement, a color attention mechanism to enhance sensitivity to display color reproduction characteristics, and temporal modeling for motion artifact assessment. Experimental results demonstrate that the proposed model achieves superior performance in reflecting user subjective experiences for display content videos compared to state-of-the-art methods, with significant improvements in both color fidelity assessment and definition evaluation. Fangfang Lu, Huiqun Yu, Kaiwei Zhang, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight ModelabstractSurveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications. Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human ImagesabstractThe rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I outputs. However, existing image quality assessment (IQA) methods are limited to providing global quality scores, failing to deliver fine-grained perceptual evaluations for structurally complex subjects like humans, which is a critical challenge considering the frequent anatomical and textural distortions in AI-generated human images (AGHIs). To address this gap, we introduce AGHI-QA, a large-scale benchmark specifically designed for quality assessment of AGHIs. The dataset comprises 4, 000 images generated from 400 carefully crafted text prompts using 10 state-of-the-art T2I models. We conduct a systematic subjective study to collect multidimensional annotations, including perceptual quality scores, text-image correspondence scores, visible and distorted body part labels. Based on AGHI-QA, we evaluate the strengths and weaknesses of current T2I methods in generating human images from multiple dimensions. Furthermore, we propose AGHI-Assessor, a novel quality metric that integrates the large multimodal model (LMM) with domain-specific human features for precise quality prediction and identification of visible and distorted body parts in AGHIs. Extensive experimental results demonstrate that AGHI-Assessor showcases state-of-the-art performance, significantly outperforming existing IQA methods in multidimensional quality assessment and surpassing leading LMMs in detecting structural distortions in AGHIs. Sijing Wu, Wei Sun 0029, Yucheng Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | LMVQ: Label-Free Metric-Learning for General AI-Generated Video Quality AssessmentabstractThe recent rapid development of video generation technology has led to a significant demand for quality assessment of the latest AI-generated videos. However, current supervised approaches depend on expensive and quickly outdated human scores, and label-free methods overlook the general distortions of AI-generated videos. To address these limitations, we introduce LMVQ, a Label-free Metric-learning framework for general AI-generated Video Quality assessment of three dimensions, spatial, temporal, and alignment. The LMVQ is the first to introduce sample degradations specially designed for AIGC-specific distortions, and constructs a comprehensive training set through two complementary sample generation strategies. It then employs two synergistic modules, the Intra-Quality Token Transformer (IQ-Trans), which explicitly refines dimension-specific quality representations, and the Inter-Quality Mixture of Experts (IQ-MoE), which fuses interactions across multiple quality dimensions. Finally, a Multi-Proxy Metric-Learning (MPML) strategy aligns the learned representations with multi-dimensional quality scores and constrains the model to learn discriminative quality-aware representations. Extensive experiments on four public AIGC-VQA benchmarks show that MPML outperforms previous label-free methods by over 20%, and greatly narrows the gap with supervised methods. This provides a scalable, adaptive foundation for evaluating the ever-evolving quality of AI-generated videos. Xinyue Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | FreeQR: Free Lunch for Aesthetic QR Codes Emerging From the Latent Space in Diffusion ModelsabstractIn the modern digital age, Quick Response (QR) codes serve as a critical interface for bridging the physical and virtual worlds, widely utilized in multimedia applications. However, traditional binary QR codes often lack the visual appeal desired in contexts. Aesthetic QR codes address this limitation by enabling the customization of QR code patterns to enhance visual attractiveness while retaining compatibility with standard QR decoders. Previous works have explored the use of diffusion models for generating such codes but often require extensive training of ControlNets and face challenges in maintaining scannability. To address these issues, we present FreeQR, a streamlined and effective approach that enables the stable generation of QR code images with diffusion models. Our methodology involves the strategic fusion between the specific channel in the latent space of the denoising process with the noised latent representations of the QR blueprint image at corresponding timesteps. This ensures that the generated images adhere to the brightness distribution required for effective scanning while achieving a balance between aesthetics and functionality. Additionally, we introduce gradient guidance based on scanning errors directly in the latent space, enabling the generation of scannable QR codes in seconds without additional model parameters. Experimental results demonstrate that FreeQR significantly enhances the aesthetics and scannability of QR codes compared to existing methods, making it a lightweight and efficient solution for multimedia applications. Yiwei Yang 0007, Jun Jia, Zheyuan Liu 0011, Zhongpai Gao, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 6 |
| 2025 | Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsabstractWith the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AI-generated content (AIGC), and computer graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of perceptual video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to the performance of human beings. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding. Ziheng Jia, Haoning Wu 0001, Chunyi Li 0001, Zijian Chen 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai |
CVPR | 7 |
| 2025 | Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision ContentabstractEvaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval. Tengchuan Kou, Shushi Wang, Chunyi Li 0001, Wei Sun 0029, Wei Wang 0213, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai |
CVPR | 5 |
| 2025 | A-Bench: Are LMMs Masters at Evaluating AI-generated Images?abstractHow to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, the precision and validity of which are still questionable. Furthermore, traditional benchmarks often utilize mostly natural-captured content rather than AIGIs to test the abilities of LMMs, leading to a noticeable gap for AIGIs. Therefore, we introduce **A-Bench** in this paper, a benchmark designed to diagnose *whether LMMs are masters at evaluating AIGIs*. Specifically, **A-Bench** is organized under two key principles: 1) Emphasizing both high-level semantic understanding and low-level visual quality perception to address the intricate demands of AIGIs. 2) Various generative models are utilized for AIGI creation, and various LMMs are employed for evaluation, which ensures a comprehensive validation scope. Ultimately, 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts. We hope that **A-Bench** will significantly enhance the evaluation process and promote the generation quality for AIGIs. Haoning Wu 0001, Chunyi Li 0001, Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Zijian Chen 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ICLR | 5 |
| 2025 | Perceiving Smoothness: Temporal Consistency Learning for Multi-Frame-Rate Video Quality AssessmentabstractThe video industry is continuously evolving, with videos featuring a wide range of frame rates. Video quality assessment (VQA) aims to automatically monitor and select the optimal frame rate for video communication systems. Existing VQA research has achieved impressive results in high frame rate (HFR) VQA tasks, but often lacks specific designs to address various frame rate distortions, such as smoothness distortions caused by frame rate variations, artifacts from the coupling of frame rate and compression, and confusion between low frame rate and slow motion. To address these challenges, we propose a VQA framework to perceive smoothness (PSVQA), which includes a novel frame-rate-driven feature processing module and a new feature fusion strategy. The module aggregates smoothness features from multi-scale temporal embeddings and incorporates frame rate guidance to resolve the discrepancies between temporal features and real perceptual experience. Furthermore, we combine spatial video features with temporal consistency features for quality modeling, optimizing the feature fusion module to enhance multi-frame-rate perception. Through extensive experiments on HFR and variable-frame-rate datasets, we validate the effectiveness of PSVQA. Jinliang Han, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICME | 3 |
| 2025 | AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality AssessmentabstractMany video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing audio-visual quality assessment methods struggle with unique distortions in AGAVs, such as unrealistic and inconsistent elements. To address this, we introduce AGAVQA-3k, the first large-scale AGAV quality assessment dataset, comprising $3,382$ AGAVs from $16$ VTA methods. AGAVQA-3k includes two subsets: AGAVQA-MOS, which provides multi-dimensional scores for audio quality, content consistency, and overall quality, and AGAVQA-Pair, designed for optimal AGAV pair selection. We further propose AGAV-Rater, a LMM-based model that can score AGAVs, as well as audio and music generated from text, across multiple dimensions, and selects the best AGAV generated by VTA methods to present to the user. AGAV-Rater achieves state-of-the-art performance on AGAVQA-3k, Text-to-Audio, and Text-to-Music datasets. Subjective tests also confirm that AGAV-Rater enhances VTA performance and user experience. The dataset and code is available at https://github.com/charlotte9524/AGAV-Rater. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICML | 4 |
| 2025 | LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and BenchmarkabstractIn recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly. Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ISCAS | 2 |
| 2025 | EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion AssessmentabstractThe furnishing of multi-modal large language models (MLLMs) has led to the emergence of numerous benchmark studies, particularly those evaluating their perception and understanding capabilities. Among these, understanding image-evoked emotions aims to enhance MLLMs' empathy, with significant applications such as human-machine interaction and advertising recommendations. However, current evaluations of this MLLM capability remain coarse-grained, and a systematic and comprehensive assessment is still lacking. To this end, we introduce EEmo-Bench, a novel benchmark dedicated to the analysis of the evoked emotions in images across diverse content categories. Our core contributions include: 1) Regarding the diversity of the evoked emotions, we adopt an emotion ranking strategy and employ the Valence-Arousal-Dominance (VAD) as emotional attributes for emotional assessment. In line with this methodology, 1,960 images are collected and manually annotated. 2) We design four tasks to evaluate MLLMs' ability to capture the evoked emotions by single images and their associated attributes: Perception, Ranking, Description, and Assessment. Additionally, image-pairwise analysis is introduced to investigate the model's proficiency in performing joint and comparative analysis. In total, we collect 6,773 question-answer pairs and perform a thorough assessment on 19 commonly-used MLLMs. The results indicate that while some proprietary and large-scale open-source MLLMs achieve promising overall performance, the analytical capabilities in certain evaluation dimensions remain suboptimal. Our EEmo-Bench paves the path for further research aimed at enhancing the comprehensive perceiving and understanding capabilities of MLLMs concerning image-evoked emotions, which is crucial for machine-centric emotion perception and understanding. Our code and benchmark datasets are available at https://github.com/workerred/EEmo-Bench. Lancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun 0029, Wei Zhou 0021, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 4 |
| 2025 | VQA2: Visual Question Answering for Video Quality AssessmentabstractThe advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field in low-level visual perception, focused initially on quantitative video quality scoring. However, driven by advances in LMMs, it is now progressing toward more holistic visual quality understanding tasks. Recent studies in the image domain have demonstrated that Visual Question Answering (VQA) can markedly enhance low-level visual quality evaluation. Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA² Instruction Dataset-the first visual question answering instruction dataset that focuses on video quality assessment. This dataset consists of 3 subsets and covers various video types, containing 157,755 instruction question-answer pairs. Then, leveraging this foundation, we present the VQA² series models. The VQA² series models interleave visual and motion tokens to enhance the perception of spatial-temporal quality details in videos. We conduct extensive experiments on video quality scoring and understanding tasks, and results demonstrate that the VQA² series models achieve excellent performance in both tasks. Notably, our final model, the VQA²-Assistant, exceeds the renowned GPT-4o in visual quality understanding tasks while maintaining strong competitiveness in quality scoring tasks. Our work provides a foundation and feasible approach for integrating low-level video quality assessment and understanding with LMMs. Ziheng Jia, Jiaying Qian, Haoning Wu 0001, Wei Sun 0029, Chunyi Li 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 5 |
| 2025 | FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality AssessmentabstractFace video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA. The code and dataset will be released at: https://github.com/wsj-sjtu/FVQ. Sijing Wu, Ziwen Xu, Huiyu Duan, Wei Sun 0029, Guangtao Zhai |
ACM Multimedia | 6 |
| 2025 | Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation MetricabstractAI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git. Wei Sun 0029, Xinyue Li 0001, Qihang Ge, Jun Jia, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 2 |
| 2025 | Subjective and Objective Quality-of-Experience Evaluation Study for Live Video StreamingabstractIn recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users’ satisfaction and overall experience, plays a critical role for media service providers to optimize large-scale live compression and transmission strategies to achieve perceptually optimal rate-distortion trade-off. Although many QoE metrics for video-on-demand (VoD) have been proposed, there remain significant challenges in developing QoE metrics for live video streaming. To bridge this gap, we conduct a comprehensive study of subjective and objective QoE evaluations for live video streaming. For the subjective QoE study, we introduce the first live video streaming QoE dataset, TaoLive QoE, which consists of 42 source videos collected from real live broadcasts and 1, 155 corresponding distorted ones degraded due to a variety of streaming distortions, including conventional streaming distortions such as compression, stalling, as well as live streaming-specific distortions like frame skipping, variable frame rate, etc. Subsequently, a human study was conducted to derive subjective QoE scores of videos in the TaoLive QoE dataset. For the objective QoE study, we benchmark existing QoE models on the TaoLive QoE dataset as well as publicly available QoE datasets for VoD scenarios, highlighting that current models struggle to accurately assess video QoE, particularly for live content. Hence, we propose an end-to-end QoE evaluation model, Tao-QoE, which integrates multi-scale semantic features and optical flow-based motion features to predicting a retrospective QoE score, eliminating reliance on statistical quality of service (QoS) features. Extensive experiments demonstrate that Tao-QoE outperforms other models on the TaoLive QoE dataset and five publicly available QoE datasets, showcasing the effectiveness and feasibility of Tao-QoE. Zehao Zhu, Wei Sun 0029, Jun Jia, Jia Wang 0004, Guangtao Zhai |
VCIP | 2 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 15 |
| 2025 | A study on the user viewing experience of implanted advertisement videos based on visual saliency
Fangfang Lu, Yingjie Lian, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 3 |
| 2025 | CT-PCQA: A Convolutional Neural Network and Transformer combined Method for Point Cloud Quality Assessment
Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Signal Process. Image Commun. | 3 |
| 2025 | Joint Luminance-Chrominance Learning for Image DebandingabstractBanding is a visually annoying artifact that frequently occurs along the chain of video acquisition, production, distribution, and display, showing a significant need for improvement in many fields. Thus far, efforts on banding removal are mainly knowledge-driven or merely learning on RGB space, which is either limited by domain knowledge or lacks the consideration for banding in chrominance channels. In this work, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner. Our debanding model is comprised of a luminance restoration network (LR-Net) and a chrominance restoration network (CR-Net). Each of them follows an encoder-decoder architecture, where a cascade of residual blocks is employed to exploit hierarchical non-local features in spatial dimensions for more powerful feature representation. Moreover, we investigate the characteristics of banding artifacts and apply specific loss functions to guide the debanding in different channels, thus boosting the restoration performance. Both qualitative and quantitative experiments show that our model significantly surpasses the existing method in terms of all 7 metrics. Ultimately, our network trained on simulated data exhibits good adaptiveness under various compression scenarios, which further demonstrates the effectiveness of the proposed model. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Ru Huang 0002, Fangfang Lu, Ying Chen 0011, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Study of Subjective and Objective Naturalness Assessment of AI-Generated ImagesabstractThe proliferation of Artificial Intelligence-Generated Images (AIGIs) has greatly expanded the Image Naturalness Assessment (INA) problem. Different from early definitions that mainly focus on tone-mapped images with limited distortions (e.g., exposure, contrast, and color reproduction), INA on AI-generated images is especially challenging as it owns more diverse contents and could be affected by factors from multiple perspectives, including low-level technical distortions and high-level rationality distortions. In this paper, we take the first step to benchmark and assess the visual naturalness of AI-generated images. First, we construct the AI-Generated Image Naturalness (AGIN) dataset by conducting a large-scale subjective study to collect human opinions on the overall naturalness as well as perceptions from the technical quality and rationality perspectives. AGIN verifies several insights for the first time that naturalness is universally and disparately affected by both technical and rational distortions, while its manifestations vary with different generation tasks. Second, to automatically assess the naturalness of AIGIs that align with human opinions, we propose the Joint Objective Image Naturalness evaluaTor (JOINT). Specifically, JOINT imitates human reasoning in naturalness evaluation by jointly learning technical and rationality features with several specific designs to guide model behavior from respective perspectives. Experiments demonstrate that JOINT significantly outperforms existing methods for providing more subjectively consistent results on naturalness assessment. The dataset can be accessed athttps://github.com/zijianchen98/AGIN. Zijian Chen 0001, Wei Sun 0029, Haoning Wu 0001, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | LMM-VQA: Advancing Video Quality Assessment With Large Multimodal ModelsabstractThe explosive growth of videos on streaming media platforms has underscored the urgent need for effective video quality assessment (VQA) algorithms to monitor and perceptually optimize the quality of streaming videos. However, VQA remains an extremely challenging task due to the diverse video content and the complex spatial and temporal distortions, thus necessitating more advanced methods to address these issues. Nowadays, large multimodal models (LMMs), such as GPT-4V, have exhibited strong capabilities for various visual understanding tasks, motivating us to leverage the powerful multimodal representation ability of LMMs to solve the VQA task. Therefore, we propose anLargeMulti-Modal basedVideoQuality Assessment (LMM-VQA) model, which introduces a novel spatiotemporal visual modeling strategy for quality-aware feature extraction. Specifically, we reformulate the quality regression problem into a question and answering (Q&A) task and construct Q&A prompts for VQA instruction tuning. Then, we design a spatiotemporal vision encoder to extract spatial and temporal features to represent the quality characteristics of videos, which are subsequently mapped into the language space by the spatiotemporal projector for modality alignment. Finally, the aligned visual tokens and the quality-inquired text tokens are aggregated as inputs for the large language model (LLM) to generate the quality score as well as the quality level. Extensive experiments demonstrate thatLMM-VQAachieves state-of-the-art performance across five VQA benchmarks, exhibiting an average improvement of 5% in generalization ability over existing methods. Furthermore, due to the advanced design of the spatiotemporal encoder and projector, LMM-VQA also performs exceptionally well on general video understanding tasks, further validating its effectiveness. Our code will be released at https://github.com/Sueqk/LMM-VQA. Qihang Ge, Wei Sun 0029, Yu Zhang 0133, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Advancing Zero-Shot Digital Human Quality Assessment Through Text-Prompted EvaluationabstractDigital humans have witnessed extensive applications in various domains, necessitating related quality assessment studies. However, there is a lack of comprehensive digital human quality assessment (DHQA) databases. To address this gap, we propose SJTU-H3D, a subjective quality assessment database specifically designed for full-body digital humans. It comprises 40 high-quality reference digital humans and 1,120 labeled distorted counterparts generated with seven types of distortions. The SJTU-H3D database can serve as a benchmark for DHQA research, allowing evaluation and refinement of processing algorithms. Further, we propose a zero-shot DHQA approach that focuses on no-reference (NR) scenarios to ensure generalization capabilities while mitigating database bias. Our method leverages semantic and distortion features extracted from projections, as well as geometry features derived from the mesh structure of digital humans. Specifically, we employ the Contrastive Language-Image Pre-training (CLIP) model to measure semantic affinity and incorporate the Naturalness Image Quality Evaluator (NIQE) model to capture low-level distortion information. Additionally, we utilize dihedral angles as geometry descriptors to extract mesh features. By aggregating these measures, we introduce the Digital Human Quality Index (DHQI), which demonstrates significant improvements in zero-shot performance. The DHQI can also serve as a robust baseline for DHQA tasks, facilitating advancements in the field. The database and the code are available at https://github.com/zzc-1998/SJTU-H3D. Wei Sun 0029, Yingjie Zhou 0003, Haoning Wu 0001, Chunyi Li 0001, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2025 | Weakly Supervised Referring Video Object Segmentation With Object-Centric Pseudo-GuidanceabstractReferring video object segmentation (RVOS) is an emerging task for multimodal video comprehension while the expensive annotating process of object masks restricts the scalability and diversity of RVOS datasets. To relax the dependency on expensive mask annotations and take advantage from large-scale partially annotated data, in this paper, we explore a novel extended RVOS task, namely weakly supervised referring video object segmentation (WRVOS), which employs multiple weak supervision sources, including object points and bounding boxes. Correspondingly, we propose a unified WRVOS framework. Specifically, an object-centric pseudo mask generation method is introduced to provide effective shape priors for the pseudo guidance of spatial object location. Then, a pseudo-guided optimization strategy is proposed to effectively optimize the object outlines in terms of spatial location and projection density with a multi-stage online learning strategy. Furthermore, a multimodal cross-frame level set evolution method is proposed to iteratively refine the object boundaries considering both temporal consistency and cross-modal interactions. Extensive experiments are conducted on four publicly available RVOS datasets, including A2D Sentences, J-HMDB Sentences, Ref-DAVIS, and Ref-YoutubeVOS. Performance comparison shows that the proposed method achieves state-of-the-art performance in both point-supervised and box-supervised settings. Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 4 |
| 2025 | Evaluating Point Cloud From Moving Camera Videos: A No-Reference MetricabstractPoint cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as compression and downsampling distortions during the transmission process. To tackle the challenge of point cloud quality assessment (PCQA), many PCQA methods have been proposed to evaluate the visual quality levels of point clouds by assessing the rendered static 2D projections. Although such projectionbased PCQA methods achieve competitive performance with the assistance of mature image quality assessment (IQA) methods, they neglect that the 3D model is also perceived in a dynamic viewing manner, where the viewpoint is continually changed according to the feedback of the rendering device. Therefore, in this paper, we evaluate the point clouds from moving camera videos and explore the way of dealing with PCQA tasks via using video quality assessment (VQA) methods. First, we generate the captured videos by rotating the camera around the point clouds through several circular pathways. Then we extract both spatial and temporal quality-aware features from the selected key frames and the video clips through using trainable 2D-CNN and pretrained 3D-CNN models respectively. Finally, the visual quality of point clouds is represented by the video quality values. The experimental results reveal that the proposed method is effective for predicting the visual quality levels of the point clouds and even competitive with full-reference (FR) PCQA methods. The ablation studies further verify the rationality of the proposed framework and confirm the contributions made by the qualityaware features extracted via the dynamic viewing manner. The code is available athttps://github.com/zzc-1998/VQA_PC. Wei Sun 0029, Yucheng Zhu, Xiongkuo Min, Wei Wu 0002, Ying Chen 0011, Guangtao Zhai |
IEEE Trans. Multim. | 2 |
| 2025 | Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified ModelabstractIn recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git . Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | MM-PCQA+: Advancing Multi-Modal Learning for Point Cloud Quality AssessmentabstractThe importance of visual quality in point clouds has been significantly underlined due to the rapid rise in 3D vision applications which aim to deliver affordable and superior user experiences. Reviewing the evolution of point cloud quality assessment (PCQA), it’s observed that visual quality evaluation typically employs single-modal data, either sourced from 2D projections or the 3D point clouds. The 2D projections possess abundant texture and semantic information while they are heavily reliant on viewpoints. In contrast, 3D point clouds are more reactive to geometric distortions and viewpoint-invariant. Consequently, to maximize the benefits of both point cloud and image modalities, we present an advanced no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA+) metric. Specifically, we divide the point clouds into sub-models to reflect local geometric distortions such as point shifting and down-sampling. Afterwards, we render the point clouds using a cube-like projection setup and sample the projections of interest using a point-visible-ratio for image feature extraction. In order to fulfill these objectives, the sub-models and projected images are encoded using point-based and image-based neural networks. Lastly, we implement symmetric cross-modal attention to amalgamate multi-modal quality-aware features. Experimental results demonstrate that our metric surpasses all state-of-the-art methods and significantly advances beyond previous no-reference PCQA methods. Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital HumansabstractIn an era where 3D Digital Humans (DHs) are becoming increasingly prevalent in fields like gaming, automotive, and the metaverse, the demand for high DH visual quality is rising. This paper presents the first-ever reduced-reference (RR) quality assessment metric tailored specifically for textured mesh DHs, aiming to optimize transmission systems and improve Quality of Experience (QoE) for viewers in resource-constrained environments. Four critical geometric curvature-related attributes and two texture-related indicators are computed, which are then statistically analyzed and utilized in a Support Vector Regression (SVR) model for robust and efficient quality prediction. Experimental results confirm that our method outperforms existing full-reference (FR) metrics, making it an invaluable tool for the future of 3D DHs in various applications. The code is available at https://github.com/zzc-1998/RR-DHQA. Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 5 |
| 2024 | SG-JND: Semantic-Guided Just Noticeable Distortion Predictor for Image CompressionabstractJust noticeable distortion (JND), representing the threshold of distortion in an image that is minimally perceptible to the human visual system (HVS), is crucial for image compression algorithms to achieve a trade-off between transmission bit rate and image quality. However, traditional JND prediction methods only rely on pixel-level or sub-band level features, lacking the ability to capture the impact of image content on JND. To bridge this gap, we propose a Semantic-Guided JND (SG-JND) network to leverage semantic information for JND prediction. In particular, SG-JND consists of three essential modules: the image preprocessing module extracts semantic-level patches from images, the feature extraction module extracts multi-layer features by utilizing the cross-scale attention layers, and the JND prediction module regresses the extracted features into the final JND value. Experimental results show that SG-JND achieves the state-of-the-art performance on two publicly available JND datasets, which demonstrates the effectiveness of SG-JND and highlight the significance of incorporating semantic information in JND assessment. Linhan Cao, Wei Sun 0029, Xiongkuo Min, Jun Jia, Zijian Chen 0001, Yucheng Zhu, Lizhou Liu, Qiubo Chen, Guangtao Zhai |
ICIP | 2 |
| 2024 | Thqa: A Perceptual Quality Assessment Database for Talking HeadsabstractIn the realm of media technology, digital humans have gained prominence due to rapid advancements in computer technology. However, the manual modeling and control required for the majority of digital humans pose significant obstacles to efficient development. The speech-driven methods offer a novel avenue for manipulating the mouth shape and expressions of digital humans. Despite the proliferation of driving methods, the quality of many generated talking head (TH) videos remains a concern, impacting user visual experiences. To tackle this issue, this paper introduces the Talking Head Quality Assessment (THQA) database, featuring 800 TH videos generated through 8 diverse speechdriven methods. Extensive experiments affirm the THQA database’s richness in character and speech features. Subsequent subjective quality assessment experiments analyze correlations between scoring results and speech-driven methods, ages, and genders. In addition, experimental results show that mainstream image and video quality assessment methods have limitations for the THQA database, underscoring the imperative for further research to enhance TH video quality assessment. The THQA database is publicly accessible at https://github.com/zyj-2000/THQA. Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Zhihua Wang 0002, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 3 |
| 2024 | Optimizing Projection-Based Point Cloud Quality Assessment with Human Preferred Viewpoints SelectionabstractViewpoint selection plays a pivotal role in projection-based point cloud quality assessment (PCQA). Generally speaking, sole reliance on a single projection fails to capture adequate quality information, leading to the prevalent use of multi-projection approaches. It is important to recognize that viewpoint selection is significantly influenced by human preferences and viewpoints that align with human predilections exert a greater impact on PCQA. Therefore, we introduce the first viewpoint selection database for PCQA, which comprises 405 distorted point clouds, accompanied by preferred viewpoints collected from humans. Then we propose a novel human preference index, devised from the Visible-Points Ratio and Visible-Color-Entropy Ratio, to guide the selection of viewpoints. Our experimental findings confirm that this human preference index correlates more closely with human preferences than traditional viewpoint selection settings. Moreover, the proposed PCQA method optimized with the human preference index demonstrates competitive performance as well. Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Chunyi Li 0001, Haoning Wu 0001, Weisi Lin, Guangtao Zhai |
ICME | 3 |
| 2024 | DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models
Yiwei Yang 0007, Zheyuan Liu 0011, Jun Jia, Zhongpai Gao, Wei Sun 0029, Xiaohong Liu 0001, Guangtao Zhai |
IJCAI | 6 |
| 2024 | FS-BAND: A Frequency-Sensitive Banding DetectorabstractBanding artifact, as known as staircase-like contour, is a common quality annoyance that happens in compression, transmission, etc. scenarios, which largely affects the user’s quality of experience (QoE). The banding distortion typically appears as relatively small pixel-wise variations in smooth backgrounds, which is difficult to analyze in the spatial domain but easily reflected in the frequency domain. In this paper, we thereby study the banding artifact from the frequency aspect and propose a no-reference banding detection model to capture and evaluate banding artifacts, called the Frequency-Sensitive BANding Detector (FS-BAND). The proposed detector is able to generate a pixel-wise banding map with a perception correlated quality score. Experimental results show that the proposed FS-BAND method outperforms state-of-the-art image quality assessment (IQA) approaches with higher accuracy in banding classification task. Zijian Chen 0001, Wei Sun 0029, Ru Huang 0002, Fangfang Lu, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005 |
ISCAS | 2 |
| 2024 | PrefIQA: Human Preference Learning for AI-generated Image Quality AssessmentabstractDespite recent advancements in generative models, the variation in image quality remains a significant concern. To tackle this issue, we propose PrefIQA, an effective human preference learning metric, which can better evaluate the quality of AI-generated images. PrefIQA consists of two units, namely Feature Extraction Unit and Feature Fusion Unit. In Feature Extraction Unit, we introduce a prompt-segmentation module to divide prompts into multiple phrases, enabling a more detailed evaluation of the alignment between images and texts. In Feature Fusion Unit, we introduce a modality-fusion module, which effectively mixes text features and image features to improve the overall performance. In the experiment part, extensive experiments are conducted, demonstrating that PrefIQA surpasses existing text-to-image alignment metrics. We believe that PrefIQA’s proposal would facilitate researches on AI-generated image quality assessment, and make a valuable contribution to the field of text-to-image generation. Hengjian Gao, Kaiwei Zhang, Wei Sun 0029, Chunyi Li 0001, Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ISCAS | 3 |
| 2024 | Large Multi-modality Model Assisted AI-Generated Image Quality AssessmentabstractTraditional deep neural network (DNN)-based image quality assessment (IQA) models leverage convolutional neural networks (CNN) or Transformer to learn the quality-aware feature representation, achieving commendable performance on natural scene images. However, when applied to AI-Generated images (AGIs), these DNN-based IQA models exhibit subpar performance. This situation is largely due to the semantic inaccuracies inherent in certain AGIs caused by uncontrollable nature of the generation process. Thus, the capability to discern semantic content becomes crucial for assessing the quality of AGIs. Traditional DNN-based IQA models, constrained by limited parameter complexity and training data, struggle to capture complex fine-grained semantic features, making it challenging to grasp the existence and coherence of semantic content of the entire image. To address the shortfall in semantic content perception of current IQA models, we introduce a large Multi-modality model Assisted AI-Generated Image Quality Assessment (MA-AGIQA) model, which utilizes semantically informed guidance to sense semantic information and extract semantic vectors through carefully designed text prompts. Moreover, it employs a mixture of experts (MoE) structure to dynamically integrate the semantic information with the quality-aware features extracted by traditional DNN-based IQA models. Comprehensive experiments conducted on two AI-generated content datasets and two traditional IQA datasets show that MA-AGIQA achieves state-of-the-art performance, and demonstrate its superior generalization capabilities on assessing the quality of AGIs. The code is available at https://github.com/wangpuyi/MA-AGIQA. Puyi Wang, Wei Sun 0029, Jun Jia, Yanwei Jiang, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 2 |
| 2024 | LMM-PCQA: Assisting Point Cloud Quality Assessment with LMMabstractAlthough large multi-modality models (LMMs) have seen extensive exploration and application in various quality assessment studies, their integration into Point Cloud Quality Assessment (PCQA) remains unexplored. Given LMMs' exceptional performance and robustness in low-level vision and quality assessment tasks, this study aims to investigate the feasibility of imparting PCQA knowledge to LMMs through text supervision. To achieve this, we transform quality labels into textual descriptions during the fine-tuning phase, enabling LMMs to derive quality rating logits from 2D projections of point clouds. To compensate for the loss of perception in the 3D domain, structural features are extracted as well. These quality logits and structural features are then combined and regressed into quality scores. Our experimental results affirm the effectiveness of our approach, showcasing a novel integration of LMMs into PCQA that enhances model understanding and assessment accuracy. We hope our contributions can inspire subsequent investigations into the fusion of LMMs with PCQA, fostering advancements in 3D visual quality analysis and beyond. The code is available at https://github.com/zzc-1998/LMM-PCQA. Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Wei Sun 0029, Chaofeng Chen, Xiongkuo Min, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai |
ACM Multimedia | 5 |
| 2024 | Subjective and Objective Quality-of-Experience Assessment for 3D Talking HeadsabstractIn recent years, immersive communication has emerged as a compelling alternative to traditional video communication methods. One prospective avenue for immersive communication involves augmenting the user's immersive experience through the transmission of three-dimensional (3D) talking heads (THs). However, transmitting 3D THs poses significant challenges due to its complex and voluminous nature, often leading to pronounced distortion and a compromised user experience. Addressing this challenge, we introduce the 3D Talking Heads Quality Assessment (THQA-3D) dataset, comprising 1,000 sets of distorted and 50 original TH mesh sequences (MSs), to facilitate quality assessment in 3D TH transmission. A subjective experiment, characterized by a novel interactive approach, is conducted with recruited participants to assess the quality of MSs in THQA-3D dataset. Leveraging this dataset, we also propose a multimodal Quality-of-Experience (QoE) method incorporating a Large Quality Model (LQM). This method involves frontal projection of MSs and subsequent rendering into videos, with quality assessment facilitated by the LQM and a variable-length video memory filter (VVMF). Additionally, tone-lip coherence and silence detection techniques are employed to characterize audio-visual coherence in 3D MS streams. Experimental evaluation demonstrates the proposed method's superiority, achieving state-of-the-art performance on the THQA-3D dataset and competitiveness on other QoE datasets. Both the THQA-3D dataset and the QoE model have been publicly released at https://github.com/zyj-2000/THQA-3D Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 3 |
| 2024 | GAIA: Rethinking Action Quality Assessment for AI-Generated VideosabstractAssessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current action quality assessment (AQA) algorithms predominantly focus on actions from real specific scenarios and are pre-trained with normative action features, thus rendering them inapplicable in AIGVs. To address these problems, we construct GAIA, a Generic AI-generated Action dataset, by conducting a large-scale subjective evaluation from a novel causal reasoning-based perspective, resulting in 971,244 ratings among 9,180 video-action pairs. Based on GAIA, we evaluate a suite of popular text-to-video (T2V) models on their ability to generate visually rational actions, revealing their pros and cons on different categories of actions. We also extend GAIA as a testbed to benchmark the AQA capacity of existing automatic evaluation methods. Results show that traditional AQA methods, action-related metrics in recent T2V benchmarks, and mainstream video quality methods perform poorly with an average SRCC of 0.454, 0.191, and 0.519, respectively, indicating a sizable gap between current models and human action perception patterns in AIGVs. Our findings underscore the significance of action quality as a unique perspective for studying AIGVs and can catalyze progress towards methods with enhanced capacities for AQA in AIGVs. Zijian Chen 0001, Wei Sun 0029, Yuan Tian 0017, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005 |
NeurIPS | 2 |
| 2024 | Perceptual video quality assessment: a surveyabstractAbstract Perceptual video quality assessment plays a vital role in the field of video processing due to the existence of quality degradations introduced in various stages of video signal acquisition, compression, transmission and display. With the advancement of Internet communication and cloud service technology, video content and traffic are growing exponentially, which further emphasizes the requirement for accurate and rapid assessment of video quality. Therefore, numerous subjective and objective video quality assessment studies have been conducted over the past two decades for both generic videos and specific videos such as streaming, user-generated content, 3D, virtual and augmented reality, high dynamic range, high frame rate, audio-visual, etc. This survey provides an up-to-date and comprehensive review of these video quality assessment studies. Specifically, we first review the subjective video quality assessment methodologies and databases, which are necessary for validating the performance of video quality metrics. Second, the objective video quality assessment measures for general purposes are categorized and surveyed according to the methodologies utilized in the quality measures. Third, we overview the objective video quality assessment measures for specific applications and emerging topics. Finally, the performance of the state-of-the-art video quality assessment measures is compared and analyzed. This survey provides a systematic overview of both classical works and recent progress in the realm of video quality assessment, which can help other researchers quickly access the field and conduct relevant research. Xiongkuo Min, Huiyu Duan, Wei Sun 0029, Yucheng Zhu, Guangtao Zhai |
Sci. China Inf. Sci. | 3 |
| 2024 | Analysis of Video Quality Datasets via Design of Minimalistic Video Quality ModelsabstractBlind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applications. As an experimental field, the improvements of BVQA models have been measured primarily on a few human-rated VQA datasets. Thus, it is crucial to gain a better understanding of existing VQA datasets in order to properly evaluate the current progress in BVQA. Towards this goal, we conduct a first-of-its-kind computational analysis of VQA datasets via designing minimalistic BVQA models. By minimalistic, we restrict our family of BVQA models to build only upon basic blocks: a video preprocessor (for aggressive spatiotemporal downsampling), a spatial quality analyzer, an optional temporal quality analyzer, and a quality regressor, all with the simplest possible instantiations. By comparing the quality prediction performance of different model variants on eight VQA datasets with realistic distortions, we find that nearly all datasets suffer from the easy dataset problem of varying severity, some of which even admit blind image quality assessment (BIQA) solutions. We additionally justify our claims by comparing our model generalization capabilities on these VQA datasets, and by ablating a dizzying set of BVQA design choices related to the basic building blocks. Our results cast doubt on the current progress in BVQA, and meanwhile shed light on good practices of constructing next-generation VQA datasets and models. Wei Sun 0029, Wen Wen 0007, Xiongkuo Min, Long Lan, Guangtao Zhai, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | BAND-2k: Banding Artifact Noticeable Database for Banding Detection and Quality AssessmentabstractBanding, also known as staircase-like contours, frequently occurs in flat areas of images/videos processed by compression or quantization algorithms. As undesirable artifacts, banding destroys the original image structure, thus inevitably degrading users’ quality of experience (QoE). In this paper, we systematically investigate the banding image quality assessment (IQA) problem, aiming to detect the image banding artifacts and evaluate their perceptual visual quality. Considering that the existing image banding databases only contain limited content sources and banding generation methods, and lack perceptual quality labels (i.e. mean opinion scores), we first build the largest banding IQA database so far, namedBanding Artifact Noticeable Database (BAND-2k), which consists of 2,000 banding images generated by 15 compression and quantization schemes. A total of 23 workers participated in the subjective IQA experiment, yielding over 214,000 patch-level banding class labels and 44,371 reliable image-level quality rating scores. Subsequently, we develop an effective no-reference (NR) banding evaluator for banding detection and quality assessment by leveraging frequency characteristics of banding artifacts. To be more specific, a dual convolutional neural network (CNN) is employed to concurrently learn the feature representation from the high-frequency and low-frequency maps, thereby enhancing the ability to discern banding artifacts. The quality score of a banding image is generated by pooling the banding detection maps masked by the spatial frequency filters. The experimental results demonstrate that our banding evaluator achieves remarkably high accuracy in banding detection and also exhibits high SRCC and PLCC results with the perceptual quality labels, even without directly learning a regression model for banding quality evaluation. These findings unveil the strong correlations between the intensity of banding artifacts and the perceptual visual quality, thus validating the necessity of banding quality assessment. The BAND-2k database and the proposed banding evaluator are available at https://github.com/zijianchen98/BAND-2k. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Fangfang Lu, Jing Liu 0002, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Continuous and Overall Quality of Experience Evaluation for Streaming Video Based on Rich Features Exploration and Dual-Stage AttentionabstractWith the rapid development of streaming media technology, the Quality of Experience (QoE) of streaming videos becomes crucial to optimize the video compression and transmission algorithms, such as adaptive bitrate (ABR). However, the complexity of human perceptual mechanisms, particularly in relation to temporal distortions, poses substantial challenges to effective QoE monitoring. In recent years, many efforts in video quality assessment (VQA) and video QoE evaluation have highlighted the influence of a broad spectrum of features—from Quality of Service (QoS) metrics to video content understanding—on viewer experience. On this basis, we believe that there is also a dynamic relationship among these features varying with the broadcasting content. Furthermore, research indicates a significant correlation between real-time and retrospective assessments of QoE for individual videos. In response to these insights, we introduce a novel approach leveraging a unified learnable network that incorporates dual-stage attention, the temporal and cross-feature attention, to accurately predict both continuous and overall QoE for streaming videos. The results of experiments conducted on several publicly available databases demonstrate the superiority of our proposed method over the state-of-the-art metrics. Ziheng Jia, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | AGIQA-3K: An Open Database for AI-Generated Image Quality AssessmentabstractWith the rapid advancements of the text-to-image generative model, AI-generated images (AGIs) have been widely applied to entertainment, education, social media, etc. However, considering the large quality variance among different AGIs, there is an urgent need for quality models that are consistent with human subjective ratings. To address this issue, we extensively consider various popular AGI models, generated AGI through different prompts and model parameters, and collected subjective scores at the perceptual quality and text-to-image alignment, thus building the most comprehensive AGI subjective quality database AGIQA-3K so far. Furthermore, we conduct a benchmark experiment on this database to evaluate the consistency between the current Image Quality Assessment (IQA) model and human perception, while proposing StairReward that significantly improves the assessment performance of subjective text-to-image alignment. We believe that the fine-grained subjective scores in AGIQA-3K will inspire subsequent AGI quality models to fit human subjective perception mechanisms at both perception and alignment levels and to optimize the generation result of future AGI models. The database is released on https://github.com/lcysyzxdxc/AGIQA-3k-Database. Chunyi Li 0001, Haoning Wu 0001, Wei Sun 0029, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Un-Gaze: A Unified Transformer for Joint Gaze-Location and Gaze-Object DetectionabstractThis paper proposes an efficient and effective method for joint gaze location detection (GL-D) and gaze object detection (GO-D), i.e., gaze following detection. Current approaches frame GL-D and GO-D as two separate tasks, employing a multi-stage framework where human head crops must first be detected and then be fed into a subsequent GL-D sub-network, which is further followed by an additional object detector for GO-D. In contrast, we reframe the gaze following detection task as detecting human head locations and their gaze followings simultaneously, aiming at jointly detect human gaze location and gaze object in a unified and single-stage pipeline. To this end, we propose GTR, short for Gaze following detection TRansformer, streamlining the gaze following detection pipeline by eliminating all additional components, leading to the first unified paradigm that unites GL-D and GO-D in a fully end-to-end manner. GTR enables an iterative interaction between holistic semantics and human head features through a hierarchical structure, inferring the relations of salient objects and human gaze from the global image context and resulting in an impressive accuracy. Concretely, GTR achieves a 12.1 mAP gain ($\mathbf {25.1}\%$) on GazeFollowing and a 18.2 mAP gain ($\mathbf {43.3\%}$) on VideoAttentionTarget for GL-D, as well as a 19 mAP improvement ($\mathbf {45.2\%}$) on GOO-Real for GO-D. Meanwhile, unlike existing systems detecting gaze following sequentially due to the need for a human head as input, GTR has the flexibility to comprehend any number of people’s gaze followings simultaneously, resulting in high efficiency. Specifically, GTR introduces over a$\times 9$improvement in FPS and the relative gap becomes more pronounced as the human number grows. Danyang Tu, Wei Shen 0002, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Hidden Barcode in Sub-Images with Invisible Locating MarkerabstractThe prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness. Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | GMS-3DQA: Projection-Based Grid Mini-patch Sampling for 3D Model Quality AssessmentabstractNowadays, most three-dimensional model quality assessment (3DQA) methods have been aimed at improving accuracy. However, little attention has been paid to the computational cost and inference time required for practical applications. Model-based 3DQA methods extract features directly from the 3D models, which are characterized by their high degree of complexity. As a result, many researchers are inclined towards utilizing projection-based 3DQA methods. Nevertheless, previous projection-based 3DQA methods directly extract features from multi-projections to ensure quality prediction accuracy, which calls for more resource consumption and inevitably leads to inefficiency. Thus, in this article, we address this challenge by proposing a no-reference (NR) projection-based G rid M ini-patch S ampling 3D Model Q uality A ssessment (GMS-3DQA) method. The projection images are rendered from six perpendicular viewpoints of the 3D model to cover sufficient quality information. To reduce redundancy and inference resources, we propose a multi-projection grid mini-patch sampling strategy (MP-GMS), which samples grid mini-patches from the multi-projections and forms the sampled grid mini-patches into one quality mini-patch map (QMM). The Swin-Transformer tiny backbone is then used to extract quality-aware features from the QMMs. The experimental results show that the proposed GMS-3DQA outperforms existing state-of-the-art NR-3DQA methods on the point cloud quality assessment databases for both accuracy and efficiency. The efficiency analysis reveals that the proposed GMS-3DQA requires far less computational resources and inference time than other 3DQA competitors. The code is available at https://github.com/zzc-1998/GMS-3DQA . Wei Sun 0029, Haoning Wu 0001, Yingjie Zhou 0003, Chunyi Li 0001, Zijian Chen 0001, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Subjective and Objective Quality Assessment for in-the-Wild Computer Graphics ImagesabstractComputer graphics images (CGIs) are artificially generated by means of computer programs and are widely perceived under various scenarios, such as games, streaming media, etc. In practice, the quality of CGIs consistently suffers from poor rendering during production, inevitable compression artifacts during the transmission of multimedia applications, and low aesthetic quality resulting from poor composition and design. However, few works have been dedicated to dealing with the challenge of computer graphics image quality assessment (CGIQA). Most image quality assessment (IQA) metrics are developed for natural scene images (NSIs) and validated on databases consisting of NSIs with synthetic distortions, which are not suitable for in-the-wild CGIs. To bridge the gap between evaluating the quality of NSIs and CGIs, we construct a large-scale in-the-wild CGIQA database consisting of 6,000 CGIs (CGIQA-6k) and carry out the subjective experiment in a well-controlled laboratory environment to obtain the accurate perceptual ratings of the CGIs. Then, we propose an effective deep learning–based no-reference (NR) IQA model by utilizing both distortion and aesthetic quality representation. Experimental results show that the proposed method outperforms all other state-of-the-art NR IQA methods on the constructed CGIQA-6k database and other CGIQA-related databases. The database is released at https://github.com/zzc-1998/CGIQA6K . Wei Sun 0029, Yingjie Zhou 0003, Jun Jia, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | MD-VQA: Multi-Dimensional Quality Assessment for UGC Live VideosabstractUser-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing of UGC live videos, effective video quality assessment (VQA) tools are needed to monitor and perceptually optimize live streaming videos in the distributing process. In this paper, we address UGC Live VQA problems by constructing a first-of-a-kind subjective UGC Live VQA database and developing an effective evaluation tool. Concretely, 418 source UGC videos are collected in real live streaming scenarios and 3,762 compressed ones at different bit rates are generated for the subsequent subjective VQA experiments. Based on the built database, we develop a Multi-12imensional VQA (MD-VQA) evaluator to measure the visual quality of UGC live videos from semantic, distortion, and motion aspects respectively. Extensive experimental results show that MD-VQA achieves state-of-the-art performance on both our UGC Live VQA database and existing compressed UGC VQA databases. Wei Wu 0002, Wei Sun 0029, Danyang Tu, Wei Lu 0021, Xiongkuo Min, Ying Chen 0011, Guangtao Zhai |
CVPR | 3 |
| 2023 | Perceptual Quality Assessment for Digital Human HeadsabstractDigital humans are attracting more and more research interest during the last decade, the generation, representation, rendering, and animation of which have been put into large amounts of effort. However, the quality assessment of digital humans has fallen behind. Therefore, to tackle the challenge of digital human quality assessment issues, we propose the first large-scale quality assessment database for three-dimensional (3D) scanned digital human heads (DHHs). The constructed database consists of 55 reference DHHs and 1,540 distorted DHHs along with the subjective perceptual ratings. Then, a simple yet effective full-reference (FR) projection-based method is proposed to evaluate the visual quality of DHHs. The pretrained Swin Transformer tiny is employed for hierarchical feature extraction and the multi-head attention module is utilized for feature fusion. The experimental results reveal that the proposed method exhibits state-of-the-art performance among the mainstream FR metrics. The database is released at https://github.com/zzc-1998/DHHQA. Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
ICASSP | 3 |
| 2023 | Agglomerative Transformer for Human-Object Interaction DetectionabstractWe propose an agglomerative Transformer (AGER) that enables Transformer-based human-object interaction (HOI) detectors to flexibly exploit extra instance-level cues in a single-stage and end-to-end manner for the first time. AGER acquires instance tokens by dynamically clustering patch tokens and aligning cluster centers to instances with textual guidance, thus enjoying two benefits: 1) Integrality: each instance token is encouraged to contain all discriminative feature regions of an instance, which demonstrates a significant improvement in the extraction of different instance-level cues and subsequently leads to a new state-of-the-art performance of HOI detection with 36.75 mAP on HICO-Det. 2) Efficiency: the dynamical clustering mechanism allows AGER to generate instance tokens jointly with the feature learning of the Transformer encoder, eliminating the need of an additional object detector or instance decoder in prior methods, thus allowing the extraction of desirable extra cues for HOI detection in a single-stage and end-to-end pipeline. Concretely, AGER reduces GFLOPs by 8.5% and improves FPS by 36%, even compared to a vanilla DETR-like pipeline without extra cue extraction. The code will be available at https://github.com/six6607/AGER.git. Danyang Tu, Wei Sun 0029, Guangtao Zhai, Wei Shen 0002 |
ICCV | 2 |
| 2023 | Audio-Visual Quality Assessment for User Generated Content: Database and MethodabstractWith the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users’ Quality of Experience (QoE). However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the user’s QoE also depends on the accompanying audio signals. In this paper, we conduct the first study to address the problem of UGC audio and video quality assessment (AVQA). Specifically, we construct the first UGC AVQA database named the SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences, and conduct a user study to obtain the mean opinion scores of the A/V sequences. The content of the SJTU-UAV database is then analyzed from both the audio and video aspects to show the database characteristics. We also design a family of AVQA models, which fuse the popular VQA methods and audio features via support vector regressor (SVR). We validate the effectiveness of the proposed models on the three databases. The experimental results show that with the help of audio signals, the VQA models can evaluate the perceptual quality more accurately. The database will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 3 |
| 2023 | Hierarchical Feature Fusion Transformer for No-Reference Image Quality AssessmentabstractRecently, increasing interest has been drawn in Transformer-based models for No-reference Image Quality Assessment (NR-IQA), especially for the hybrid approach. The hybrid approach tend to apply Transformer to aggregate quality information from feature maps extracted by Convolutional Neural Networks (CNN). However, existing methods cannot fully utilize the information of hierarchical features extracted by the deep neural network, resulting in the limited performance of image quality evaluation. In this work, we propose a novel Hierarchical Feature Fusion Transformer for NR-IQA (HiFFTiq), which is able to effectively exploit complementary strengths of features extracted by different layers. Further, we propose a new Uniform Partition Pooling (UPP) which can reduce the resolution of input features via uniform partitions and can well retain the quality-related information compared to the traditional pooling method Sliding Window Pooling (SWP). The results of experiment demonstrate that HiFFTiq leads to improvements of performance over the state-of-the-art methods on three large scale NR-IQA datasets. Zesheng Wang 0004, Wei Wu 0002, Wei Sun 0029, Ying Chen 0011, Kai Li 0012, Guangtao Zhai |
ICIP | 4 |
| 2023 | Geometry-Aware Video Quality Assessment for Dynamic Digital HumanabstractDynamic Digital Humans (DDHs) are 3D digital models that are animated using predefined motions and are inevitably bothered by noise/shift during the generation process and compression distortion during the transmission process, which needs to be perceptually evaluated. Usually, DDHs are displayed as 2D rendered animation videos and it is natural to adapt video quality assessment (VQA) methods to DDH quality assessment (DDH-QA) tasks. However, the VQA methods are highly dependent on viewpoints and less sensitive to geometry-based distortions. Therefore, in this paper, we propose a novel no-reference (NR) geometry-aware video quality assessment method for DDH-QA challenge. Geometry characteristics are described by the statistical parameters estimated from the DDHs’ geometry attribute distributions. Spatial and temporal features are acquired from the rendered videos. Finally, all kinds of features are integrated and regressed into quality values. Experimental results show that the proposed method achieves state-of-the-art performance on the DDH-QA database. Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
ICIP | 3 |
| 2023 | A No-Reference Quality Assessment Method for Digital Human HeadabstractIn recent years, digital humans have been widely applied in augmented/virtual reality (A/VR), where viewers are allowed to freely observe and interact with the volumetric content. However, the digital humans may be degraded with various distortions during the procedure of generation and transmission. Moreover, little effort has been put into the perceptual quality assessment of digital humans. Therefore, it is urgent to carry out objective quality assessment methods to tackle the challenge of digital human quality assessment (DHQA). In this paper, we develop a novel no-reference (NR) method based on Transformer to deal with DHQA in a multi-task manner. Specifically, the front 2D projections of the digital humans are rendered as inputs and the vision transformer (ViT) is employed for the feature extraction. Then we design a multi-task module to jointly classify the distortion types and predict the perceptual quality levels of digital humans. The experimental results show that the proposed method well correlates with the subjective ratings and outperforms the state-of-the-art quality assessment methods. Yingjie Zhou 0003, Wei Sun 0029, Xiongkuo Min, Xianghe Ma, Guangtao Zhai |
ICIP | 3 |
| 2023 | BH-VQA: Blind High Frame Rate Video Quality AssessmentabstractHigh frame rate (HFR) videos can provide consumers with a more immersive viewing experience in motion-rich scenes. However, they also pose a great challenge for video compression and transmission due to the increase in frame rates. Therefore, it is very important to choose proper frame rates and bit rates to achieve a trade-off between transmission bandwidth and visual quality. In this paper, we propose a novel Blind HFR Video Quality Assessment (BH-VQA) model by exploring the efficient and effective motion representation from the deep neural network (DNN). Concretely, we first train a baseline VQA model (i.e. a backbone network and a regressor) on a large-scale VQA database to derive a powerful quality-aware feature extractor for the spatial and motion feature extraction. Then, the HFR video is split into a sequence of video clips and the spatial features of each video clip are extracted just using the first frame of the video clip. To capture temporal distortions caused by frame rate variations and object and camera motion, we calculate deep structural similarities between continuous frames of each video clip as the motion features. Finally, the temporal quality dependencies between video clips are learned through a gated recurrent unit (GRU) network to obtain the perceptual video quality score. Experimental results show that BH-VQA achieves the best performance on two publicly available HFR VQA databases. The code of BH-VQA will be released. Wei Lu 0021, Wei Sun 0029, Danyang Tu, Xiongkuo Min, Guangtao Zhai |
ICME | 2 |
| 2023 | EEP-3DQA: Efficient and Effective Projection-Based 3D Model Quality AssessmentabstractCurrently, great numbers of efforts have been put into improving the effectiveness of 3D model quality assessment (3DQA) methods. However, little attention has been paid to the computational costs and inference time, which is also important for practical applications. Unlike 2D media, 3D models are represented by more complicated and irregular digital formats, such as point cloud and mesh. Thus it is normally difficult to perform an efficient module to extract quality-aware features of 3D models. In this paper, we address this problem from the aspect of projection-based 3DQA and develop a no-reference (NR) Efficient and Effective Projection-based 3D Model Quality Assessment (EEP-3DQA) method. The input projection images of EEP-3DQA are randomly sampled from the six perpendicular viewpoints of the 3D model and are further spatially downsampled by the grid-mini patch sampling strategy. Further, the lightweight Swin-Transformer tiny is utilized as the backbone to extract the quality-aware features. Finally, the proposed EEP-3DQA and EEP-3DQA-t (tiny version) achieve the best performance than the existing state-of-the-art NR-3DQA methods and even outperforms most full-reference (FR) 3DQA methods on the point cloud and mesh quality assessment databases while consuming less inference time than the compared 3DQA methods. Wei Sun 0029, Yingjie Zhou 0003, Wei Lu 0021, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai |
ICME | 2 |
| 2023 | DDH-QA: A Dynamic Digital Humans Quality Assessment DatabaseabstractIn recent years, large amounts of effort have been put into pushing forward the real-world application of dynamic digital human (DDH). However, most current quality assessment research focuses on evaluating static 3D models and usually ignores motion distortions. Therefore, in this paper, we construct a large-scale dynamic digital human quality assessment (DDH-QA) database with diverse motion content as well as multiple distortions to comprehensively study the perceptual quality of DDHs. Both model-based distortion (noise, compression) and motion-based distortion (binding error, motion unnaturalness) are taken into consideration. Ten types of common motion are employed to drive the DDHs and a total of 800 DDHs are generated in the end. Afterward, we render the video sequences of the distorted DDHs as the evaluation media and carry out a well-controlled subjective experiment. Then a benchmark experiment is conducted with the state-of-the-art video quality assessment (VQA) methods and the experimental results show that existing VQA methods are limited in assessing the perceptual loss of DDHs. The database is available at https://github.com/zzc-1998/DDH-QA. Yingjie Zhou 0003, Wei Sun 0029, Wei Lu 0021, Xiongkuo Min, Yu Wang 0002, Guangtao Zhai |
ICME | 3 |
| 2023 | MM-PCQA: Multi-Modal Learning for No-reference Point Cloud Quality AssessmentabstractThe visual quality of point clouds has been greatly emphasized since the ever-increasing 3D vision applications are expected to provide cost-effective and high-quality experiences for users. Looking back on the development of point cloud quality assessment (PCQA), the visual quality is usually evaluated by utilizing single-modal information, i.e., either extracted from the 2D projections or 3D point cloud. The 2D projections contain rich texture and semantic information but are highly dependent on viewpoints, while the 3D point clouds are more sensitive to geometry distortions and invariant to viewpoints. Therefore, to leverage the advantages of both point cloud and projected image modalities, we propose a novel no-reference Multi-Modal Point Cloud Quality Assessment (MM-PCQA) metric. In specific, we split the point clouds into sub-models to represent local geometry distortions such as point shift and down-sampling. Then we render the point clouds into 2D image projections for texture feature extraction. To achieve the goals, the sub-models and projected images are encoded with point-based and image-based neural networks. Finally, symmetric cross-modal attention is employed to fuse multi-modal quality-aware information. Experimental results show that our approach outperforms all compared state-of-the-art methods and is far ahead of previous no-reference PCQA methods, which highlights the effectiveness of the proposed method. The code is available at https://github.com/zzc-1998/MM-PCQA. Wei Sun 0029, Xiongkuo Min, Qiyuan Wang 0002, Guangtao Zhai |
IJCAI | 2 |
| 2023 | StableVQA: A Deep No-Reference Quality Assessment Model for Video StabilityabstractVideo shakiness is an unpleasant distortion of User Generated Content (UGC) videos, which is usually caused by the unstable hold of cameras. In recent years, many video stabilization algorithms have been proposed, yet no specific and accurate metric enables comprehensively evaluating the stability of videos. Indeed, most existing quality assessment models evaluate video quality as a whole without specifically taking the subjective experience of video stability into consideration. Therefore, these models cannot measure the video stability explicitly and precisely when severe shakes are present. In addition, there is no large-scale video database in public that includes various degrees of shaky videos with the corresponding subjective scores available, which hinders the development of Video Quality Assessment for Stability (VQA-S). To this end, we build a new database named StableDB that contains 1,952 diversely-shaky UGC videos, where each video has a Mean Opinion Score (MOS) on the degree of video stability rated by 34 subjects. Moreover, we elaborately design a novel VQA-S model named StableVQA, which consists of three feature extractors to acquire the optical flow, semantic, and blur features respectively, and a regression layer to predict the final stability score. Extensive experiments demonstrate that the StableVQA achieves a higher correlation with subjective opinions than the existing VQA-S models and generic VQA models. The database and codes are available at https://github.com/QMME/StableVQA. Tengchuan Kou, Xiaohong Liu 0001, Wei Sun 0029, Jun Jia, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 3 |
| 2023 | Perceptual quality assessment for fine-grained compressed images
Wei Sun 0029, Wei Wu 0002, Xiongkuo Min, Guangtao Zhai |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | A Deep Learning-Based Multidimensional Aesthetic Quality Assessment Method for Mobile Game ImagesabstractMobile games have played an increasingly significant role in people's leisure lives in recent years, thanks to the fast expansion of the gaming industry and the widespread use of mobile devices. The aesthetic quality of game pictures is a very important factor that attracts users' interest. However, evaluating the aesthetic quality of mobile game pictures is difficult since the painting styles of games vary greatly and the evaluation criteria are also diversified. In this article, we propose a multitask deep learning-based method, which is able to predict the aesthetic quality of mobile game images in multiple dimensions. The proposed model consists of two modules, a feature extraction module and a quality regression module. We extract quality-aware features from intermediate layers of the deep convolution neural network and then incorporate them into the final feature representation in the feature extraction module, allowing the model to fully use visual information from low to high levels. The quality regression module uses fully connected layers to map quality-aware features into quality scores across multiple dimensions. The multidimensional aesthetic quality scores are trained using a multitask learning approach, in which quality-aware features are shared across multiple dimensional quality prediction tasks. Finally, several key factors which help the proposed model perform better are analyzed. The experimental results indicate that our proposed method not only achieves the greatest performance on mobile game images, but also is applicable to natural scene images. Tao Wang 0078, Wei Sun 0029, Wei Wu 0002, Ying Chen 0011, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
IEEE Trans. Games | 2 |
| 2023 | Attention-Guided Neural Networks for Full-Reference and No-Reference Audio-Visual Quality AssessmentabstractWith the popularity of mobile Internet, audio and video (A/V) have become the main way for people to entertain and socialize daily. However, in order to reduce the cost of media storage and transmission, A/V signals will be compressed by service providers before they are transmitted to end-users, which inevitably causes distortions in the A/V signals and degrades the end-user's Quality of Experience (QoE). This motivates us to research the objective audio-visual quality assessment (AVQA). In the field of AVQA, most previous works only focus on single-mode audio or visual signals, which ignores that the perceptual quality of users depends on both audio and video signals. Therefore, we propose an objective AVQA architecture for multi-mode signals based on attentional neural networks. Specifically, we first utilize an attention prediction model to extract the salient regions of video frames. Then, a pre-trained convolutional neural network is used to extract short-time features of the salient regions and the corresponding audio signals. Next, the short-time features are fed into Gated Recurrent Unit (GRU) networks to model the temporal relationship between adjacent frames. Finally, the fully connected layers are utilized to fuse the temporal related features of A/V signals modeled by the GRU network into the final quality score. The proposed architecture is flexible and can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database and UnB-AVC Database demonstrate that our model outperforms the state-of-the-art AVQA methods. The code of the proposed method will be publicly available to promote the development of the field of AVQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Image Process. | 3 |
| 2023 | Subjective and Objective Audio-Visual Quality Assessment for User Generated ContentabstractIn recent years, User Generated Content (UGC) has grown dramatically in video sharing applications. It is necessary for service-providers to use video quality assessment (VQA) to monitor and control users' Quality of Experience when watching UGC videos. However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the perceptual quality also depends on the accompanying audio signals. In this paper, we conduct a comprehensive study on UGC audio-visual quality assessment (AVQA) from both subjective and objective perspectives. Specially, we construct the first UGC AVQA database named SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences collected from the YFCC100m database. A subjective AVQA experiment is conducted on the database to obtain the mean opinion scores (MOSs) of the A/V sequences. To demonstrate the content diversity of the SJTU-UAV database, we give a detailed analysis of the SJTU-UAV database as well as other two synthetically-distorted AVQA databases and one authentically-distorted VQA database, from both the audio and video aspects. Then, to facilitate the development of AVQA fields, we construct a benchmark of AVQA models on the proposed SJTU-UAV database and other two AVQA databases, of which the benchmark models consist of AVQA models designed for synthetically distorted A/V sequences and AVQA models built through combining the popular VQA methods and audio features via support vector regressor (SVR). Finally, considering benchmark AVQA models perform poorly in assessing in-the-wild UGC videos, we further propose an effective AVQA model via jointly learning quality-aware audio and visual feature representations in the temporal domain, which is seldom investigated by existing AVQA models. Our proposed model outperforms the aforementioned benchmark AVQA models on the SJTU-UAV database and two synthetically distorted AVQA databases. The SJTU-UAV database and the code of the proposed model will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Image Process. | 3 |
| 2023 | Blind Image Quality Assessment via Cross-View ConsistencyabstractImage quality assessment (IQA) is very important for both end-users and service-providers since a high-quality image can significantly improve the user's quality of experience (QoE). Most existing blind image quality assessment (BIQA) models were developed for synthetically distorted images, however, they perform poorly on in-the-wild images, which are widely existed in various practical applications. In this paper, a BIQA model is proposed that consists of a desirable self-supervised feature learning approach to mitigate the data shortage problem and learn comprehensive feature representations, and a self-attention-based feature fusion module to introduce self-attention mechanism. We develop the image quality assessment model under the framework of contrastive learning with multi views. Since human visual system perceives signals through multiple channels, the most important visual information should exist among all views of the channels. So we design the cross-view consistent information mining (CVC-IM) module to extract compact mutual information between different views. Color information and pseudo-reference image (PRI) of different distortion types are employed to formulate rich feature embeddings and preserve the quality-aware fidelity of learned representations. We employ the Transformer as the self-attention-based architecture to integrate feature embeddings. Extensive experiments show that our model achieves remarkable image quality assessment results on in-the-wild IQA datasets. Yucheng Zhu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Surveillance Video Quality Assessment Based on Quality Related RetrainingabstractSurveillance videos have been widely used in many vision-based systems, supporting intelligent tasks such as object detection and tracking. However, the quality of surveillance videos suffers from poor weather conditions and inevitable compression error, which may have a negative influence on the performance of such tasks. Therefore, accurately distinguishing distortions and predicting severity levels are crucial. In this paper, we propose a quality related retraining framework as well as a no-reference (NR) multi-task video quality assessment (VQA) model to tackle the challenge of surveillance videos quality assessment. The quality related retraining framework operates on a synthetic VQA database. The proposed NR VQA method utilizes both spatial and temporal information by using ResNet50 and SlowFast. Then multiple distortion detection heads are applied to predict the severity levels for corresponding distortions. The experimental results show that the proposed method gains competitive performance on the Video Surveillance Quality Assessment Dataset (VSQuAD). The ablation study further confirms the contributions of the quality related retraining framework, spatial information, and temporal information. Wei Lu 0021, Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Guangtao Zhai |
ICIP | 3 |
| 2022 | A No-Reference Deep Learning Quality Assessment Method for Super-Resolution Images Based on Frequency MapsabstractTo support the application scenarios where high-resolution (HR) images are urgently needed, various single image super-resolution (SISR) algorithms are developed. However, SISR is an ill-posed inverse problem, which may bring artifacts like texture shift, blur, etc. to the reconstructed images, thus it is necessary to evaluate the quality of super-resolution images (SRIs). Note that most existing image quality assessment (IQA) methods were developed for synthetically distorted images, which may not work for SRIs since their distortions are more diverse and complicated. Therefore, in this paper, we propose a no-reference deep-learning image quality assessment method based on frequency maps because the artifacts caused by SISR algorithms are quite sensitive to frequency information. Specifically, we first obtain the high-frequency map (HM) and low-frequency map (LM) of SRI by using Sobel operator and piecewise smooth image approximation. Then, a two-stream network is employed to extract the quality-aware features of both frequency maps. Finally, the features are regressed into a single quality value using fully connected layers. The experimental results show that our method outperforms all compared IQA models on the selected three super-resolution quality assessment (SRQA) databases. Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
ISCAS | 2 |
| 2022 | A Deep Learning based No-reference Quality Assessment Model for UGC VideosabstractQuality assessment for User Generated Content (UGC) videos plays an important role in ensuring the viewing experience of end-users. Previous UGC video quality assessment (VQA) studies either use the image recognition model or the image quality assessment (IQA) models to extract frame-level features of UGC videos for quality regression, which are regarded as the sub-optimal solutions because of the domain shifts between these tasks and the UGC VQA task. In this paper, we propose a very simple but effective UGC VQA model, which tries to address this problem by training an end-to-end spatial feature extraction network to directly learn the quality-aware spatial feature representation from raw pixels of the video frames. We also extract the motion features to measure the temporal-related distortions that the spatial features cannot model. The proposed model utilizes very sparse frames to extract spatial features and dense frames (i.e. the video chunk) with a very low spatial resolution to extract motion features, which thereby has low computational complexity. With the better quality-aware features, we only use the simple multilayer perception layer (MLP) network to regress them into the chunk-level quality scores, and then the temporal average pooling strategy is adopted to obtain the video-level quality score. We further introduce a multi-scale quality fusion strategy to solve the problem of VQA across different spatial resolutions, where the multi-scale weights are obtained from the contrast sensitivity function of the human visual system. The experimental results show that the proposed model achieves the best performance on five popular UGC VQA databases, which demonstrates the effectiveness of the proposed model. Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
ACM Multimedia | 1 |
| 2022 | A No-reference Quality Assessment Metric for Point Cloud Based on Captured Video SequencesabstractPoint cloud is one of the most widely used digital formats of 3D models, the visual quality of which is quite sensitive to distortions such as downsampling, noise, and compression. To tackle the challenge of point cloud quality assessment (PCQA) in scenarios where reference is not available, we propose a no-reference quality assessment metric for colored point cloud based on captured video sequences. Specifically, three video sequences are obtained by rotating the camera around the point cloud through three specific orbits. The video sequences not only contain the static views but also include the multi-frame temporal information, which greatly helps understand the human perception of the point clouds. Then we modify the ResNet3D as the feature extraction model to learn the correlation between the capture videos and corresponding subjective quality scores. The experimental results show that our method outperforms most of the state-of-the-art full-reference and no-reference PCQA metrics, which validates the effectiveness of the proposed method. Wei Sun 0029, Xiongkuo Min, Qiyuan Wang 0002, Guangtao Zhai |
MMSP | 3 |
| 2022 | A Full- Reference Quality Assessment Metric for Cartoon ImagesabstractCartoon images are illustrations that are typically drawn, sometimes animated, in an unrealistic or semi-realistic style, which are widely applied in multimedia services. However, in some post-production processes as well as transmission systems, cartoon images are inevitably distorted by wrong color arrangement and compression. Therefore, it is urgent to carry out image quality assessment (IQA) metrics to automatically predict the perceptual quality levels of distorted cartoon images. Nevertheless, the existing mainstream IQA metrics are specially developed for natural scene images (NSIs). Due to the statistical difference in structure and color aspects between cartoon images and NSIs, the scores predicted by such metrics are often inconsistent with the human vision system (HVS) for cartoon images. To further improve the performance of cartoon image quality assessment (C-IQA) methods and provide guidance for practical applications, we propose a full-reference (FR) IQA method to tackle the challenge of C-IQA. Specifically, the proposed method extracts edge and texture features to analyze the structural error. Then the moment and entropy of various color spaces are computed to reflect color distortions. Then the features are regressed into quality scores with the assistance of a support vector regression (SVR) model. Experimental results show that our metric outperforms the mainstream FR-IQA metrics, which indicates that the proposed method is more capable of modeling the visual quality loss of cartoon images. Chunyi Li 0001, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
MMSP | 3 |
| 2022 | Subjective Quality Assessment for Images Generated by Computer GraphicsabstractWith the development of rendering techniques, computer graphics generated images (CGIs) have been widely used in practical application scenarios such as architecture design, video games, simulators, movies, etc. Different from natural scene images (NSIs), the distortions of CGIs are usually caused by poor rending settings and limited computation resources. What's more, some CGIs may also suffer from compression distortions in transmission systems like cloud gaming and stream media. However, limited work has been put forward to tackle the problem of computer graphics generated images' quality assessment (CG-IQA). Therefore, in this paper, we establish a large-scale subjective CG-IQA database to deal with the challenge of CG-IQA tasks. We collect 25,454 in-the-wild CGIs through previous databases and personal collection. After data cleaning, we carefully select 1,200 CGIs to conduct the subjective experiment. Several popular no-reference image quality assessment (NR-IQA) methods are tested on our database. The experimental results show that the handcrafted-based methods achieve low correlation with subjective judgment and deep learning-based methods obtain relatively better performance. The current NR-IQA models are not suitable for CG-IQA tasks and more effective models are urgently needed. Tao Wang 0078, Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
MMSP | 3 |
| 2022 | Video-based Human-Object Interaction Detection from Tubelet TokensabstractWe present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each token is learned by a selective attention mechanism to reduce redundant dependencies from others; 2) Expressiveness: each token is enabled to align with a semantic instance, i.e., an object or a human, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results show our method outperforms existing works by large margins, with a relative mAP gain of $16.14\%$ on VidHOI and a 2 points gain on CAD-120 as well as a $4 \times$ speedup. Danyang Tu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Wei Shen 0002 |
NeurIPS | 2 |
| 2022 | Distinguishing Computer-Generated Images from Photographic Images: a Texture-Aware Deep Learning-Based MethodabstractWith the rapid development of computer graphics and generative models, computers are capable of generating images containing non-existent objects and scenes. Moreover, the computer-generated (CG) images may be indistinguishable from photographic (PG) images due to the strong representation ability of neural network and huge advancement of 3D rendering technologies. The abuse of such CG images may bring potential risks for personal property and social stability. Therefore, in this paper, we propose a dual-stream neural network to extract features enhanced by texture information to deal with the CG and PG image classification task. First, the input images are first converted to texture maps using the rotation-invariant uniform local binary patterns. Then we employ an attention-based texture-aware feature enhancement module to fuse the features extracted from each stage of the dual-stream neural network. Finally, the features are pooled and regressed into the predicted results by fully connected layers. The experimental results show that the proposed method achieves the best performance among all three popular CG and PG classification databases. The ablation study and cross-database validation experiments further confirm the effectiveness and generalization ability of the proposed algorithm. Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
VCIP | 2 |
| 2022 | Deep Fusion of Spectral-Spatial Priors for Cropland Segmentation in Remote Sensing ImagesabstractCropland segmentation is one of the critical techniques in the agriculture Remote Sensing (RS). Although the Deep Learning (DL) methods have achieved remarkable performance in the natural vision, the cropland segmentation of RS images still suffers from cropland adhesion due to the interference from the surrounding environment and the cropland cover. To tackle this problem, this letter proposes a two stage DL method with spectral-spatial priors. In the first stage, the Multi-feature Extraction Module (MEM) is designed to predict the boundary, an important spatial prior of the cropland. In the second stage, the spatial prior is further fused with the spectral prior by MEMs to get accurate cropland prediction. To evaluate the effectiveness and robustness of the proposed method, we construct a data set called Jiaxiang Cropland Set (JCS) and propose a region level evaluation indicator namely the Plot Mean Intersection over Union (PMIoU). The experiment results on the JCS demonstrate that the proposed method is both qualitatively and quantitatively competitive compared with the state-of-the-art methods. Laifeng Huang, Bin Sun 0001, Wei Sun 0029, Shutao Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Calculation of ophthalmic diagnostic parameters on a single eye image based on deep neural network
Xuefei Song, Xiongkuo Min, Huifang Zhou, Wei Sun 0029, Jia Wang 0004, Guangtao Zhai |
Multim. Tools Appl. | 5 |
| 2022 | No-Reference Quality Assessment for 3D Colored Point Cloud and Mesh ModelsabstractTo improve the viewer’s Quality of Experience (QoE) and optimize computer graphics applications, 3D model quality assessment (3D-QA) has become an important task in the multimedia area. Point cloud and mesh are the two most widely used digital representation formats of 3D models, the visual quality of which is quite sensitive to lossy operations like simplification and compression. Therefore, many related studies such as point cloud quality assessment (PCQA) and mesh quality assessment (MQA) have been carried out to measure the visual quality of distorted 3D models. However, most previous studies utilize full-reference (FR) metrics, which indicates they can not predict the quality level in the absence of the reference 3D model. Furthermore, few 3D-QA metrics consider color information, which significantly restricts their effectiveness and scope of application. In this paper, we propose a no-reference (NR) quality assessment metric for colored 3D models represented by both point cloud and mesh. First, we project the 3D models from 3D space into quality-related geometry and color feature domains. Then, the 3D natural scene statistics (3D-NSS) and entropy are utilized to extract quality-aware features. Finally, a support vector regression (SVR) model is employed to regress the quality-aware features into visual quality scores. Our method is validated on the colored point cloud quality assessment database (SJTU-PCQA), the Waterloo point cloud assessment database (WPC), and the colored mesh quality assessment database (CMDM). The experimental results show that the proposed method outperforms most compared NR 3D-QA metrics with competitive computational resources and greatly reduces the performance gap with the state-of-the-art FR 3D-QA metrics. The code of the proposed model is publicly available now athttps://github.com/zzc-1998/NR-3DQA. Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Dynamic Backlight Scaling Considering Ambient Luminance for Mobile Videos on LCD DisplaysabstractThe power consumption of mobile devices is always a concern for users due to the constraints on the size and weight of mobile devices. Up to now, mobile video traffic has accounted for the majority of total traffic, which implies that viewing mobile videos has been the major activity when using mobile devices. Among all the subsystems involving in mobile video playback, the display is the most power consuming subsystem. To reduce the power consumption, dynamic backlight scaling (DBS) technique is developed by adjusting the backlight magnitude when playing the mobile video. However, the convenience of mobile devices makes lots of people watch mobile videos in various luminance environments, which makes the existing DBS methods ineffective since ambient luminance varies greatly. In this paper, we propose a novel DBS strategy to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users’ quality of experience (QoE). In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance, and backlight luminance. We then investigate whether the continuous playback of backlight-scaled videos using the proposed scaling magnitude under various luminance environments would cause flicker effect or not. Motivated by the findings of these studies, we implement a novel DBS strategy for mobile energy saving which is suitable for various ambient luminance conditions. The experimental results demonstrate that the proposed DBS strategy can save more than 40 percent power at most and can save 10 percent power even at a very high ambient luminance condition. We also show that the proposed DBS strategy can be easily adapted to different user preferences and different devices, and can be conveniently integrated into practical applications. Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Siwei Ma 0001, Xiaokang Yang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2021 | Deep Neural Networks For Full-Reference And No-Reference Audio-Visual Quality AssessmentabstractIn the field of audio and visual quality assessment, most of previous works only focused on the single-mode visual or audio signal. However, for multi-mode signals, such as video and the accompanying audio, the overall perceptual quality depends on both video and audio. In this paper, we proposed an objective audio-visual quality assessment (AVQA) architecture for multi-mode signals based on deep neural networks. We first use a pretrained convolutional neural network to extract features of the single video frames and the concurrent short audio segments. Then, the extracted features are fed into Gated Recurrent Unit networks for time sequence modeling. Finally, we utilize the fully connected layers to fuse the qualities of audio and visual signals into the final quality score. The proposed architecture can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database prove that our model outperforms the state-of-the-art AVQA methods. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICIP | 3 |
| 2021 | Attention Based Network For No-Reference UGC Video Quality AssessmentabstractThe quality assessment of user-generated content (UGC) videos is a challenging problem due to the absence of reference videos and their complex distortions. Traditional no-reference video quality assessment (NR-VQA) algorithms mainly target specific synthetic distortions. Less attention has been paid to authentic distortions in UGC videos, which are not distributed evenly in both the spatial and temporal domains. In this paper, we propose an end-to-end neural network model for UGC video quality assessment based on the attention mechanism. The key step in our approach is to embed the attention modules in the feature extraction network, which effectively extracts local distortion information. In addition, to exploit the temporal perception mechanism of the human visual system (HVS), the gated recurrent unit (GRU) and temporal pooling layer are integrated into the proposed model. We validate the proposed model on three public in-the-wild VQA databases: KoNViD-1k, CVD2014, and LIVE-Qualcomm. Experimental results demonstrate that the proposed method outperforms state-of-the-art NR-VQA models. The implementation of our method is released at https://github.com/qingshangithub/AB-VQA. Fuwang Yi, Mianyi Chen, Wei Sun 0029, Xiongkuo Min, Yuan Tian 0017, Guangtao Zhai |
ICIP | 3 |
| 2021 | A No-Reference Evaluation Metric for Low-Light Image EnhancementabstractLow-light images, which are usually taken in dark or back-lighting conditions, are hard to perceive due to the low visibility and low contrast. To improve viewers’ Quality of Experience (QoE) and support the application of vision-based systems, various low-light image enhancement algorithms (LIEAs) have been proposed to lighten low-light images. However, some LIEAs may amplify the hidden distortions in the dark like noise and even further, introduce new distortions such as structural damage, color shift, etc, which severely affect the quality of light-enhanced images and need to be evaluated quantificationally. However, in the literature, few measures are proposed to assess the quality of light-enhanced images. Therefore, in this paper, we develop a no-reference low-light image enhancement evaluation (NLIEE) metric to predict the quality of light-enhanced images. The image quality is mainly assessed from four key aspects: light enhancement, color comparison, noise measurement, and structure evaluation. The experiment results show that NLIEE achieves the best performance among the general no-reference image quality assessment (NR IQA) models and quality descriptors for light enhancement. Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
ICME | 2 |
| 2021 | A Multi-dimensional Aesthetic Quality Assessment Model for Mobile Game ImagesabstractWith the development of the game industry and the popularization of mobile devices, mobile games have played an important role in people's entertainment life. The aesthetic quality of mobile game images determines the users' Quality of Experience (QoE) to a certain extent. In this paper, we propose a multi-task deep learning based method to evaluate the aesthetic quality of mobile game images in multiple dimensions (i.e. the fineness, color harmony, colorfulness, and overall quality). Specifically, we first extract the quality-aware feature representation through integrating the features from all intermediate layers of the convolution neural network (CNN) and then map these quality-aware features into the quality score space in each dimension via the quality regressor module, which consists of three fully connected (FC) layers. The proposed model is trained through a multi-task learning manner, where the quality-aware features are shared by different quality dimension prediction tasks, and the multi-dimensional quality scores of each image are regressed by multiple quality regression modules respectively. We further introduce an uncertainty principle to balance the loss of each task in the training stage. The experimental results show that our proposed model achieves the best performance on the Multi-dimensional Aesthetic assessment for Mobile Game image database (MAMG) among state-of-the-art image quality assessment (IQA) algorithms and aesthetic quality assessment (AQA) algorithms. Tao Wang 0078, Wei Sun 0029, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
VCIP | 2 |
| 2021 | A Full-Reference Quality Assessment Metric for Fine-Grained Compressed ImagesabstractCompressed image quality assessment (IQA) has been a crucial part of a wide range of image services such as storage and transmission. Due to the effect of different bit rates and compression methods, the compressed images usually have different levels of quality. Nowadays, the mainstream full-reference (FR) metrics are effective to predict the quality of compressed images at coarse-grained levels, however, they may perform poorly when quality differences of the compressed images are quite subtle. To better improve the Quality of Experience (QoE) and provide useful guidance for compression algorithms, we propose an FR-IQA metric for fine-grained compressed images, which estimates the image quality by analyzing the difference of structure and texture. Our metric is mainly validated on the fine-grained compression IQA (FGIQA) database and is tested on other commonly used compression IQA databases as well. The experimental results show that our metric outperforms mainstream FR-IQA metrics on the fine-grained compression IQA database and also obtains competitive performance on the coarse-grained compression IQA databases. Wei Sun 0029, Xiongkuo Min, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai |
VCIP | 2 |
| 2021 | Perceptual Quality Assessment of Low-light Image EnhancementabstractLow-light image enhancement algorithms (LIEA) can light up images captured in dark or back-lighting conditions. However, LIEA may introduce various distortions such as structure damage, color shift, and noise into the enhanced images. Despite various LIEAs proposed in the literature, few efforts have been made to study the quality evaluation of low-light enhancement. In this article, we make one of the first attempts to investigate the quality assessment problem of low-light image enhancement. To facilitate the study of objective image quality assessment (IQA), we first build a large-scale low-light image enhancement quality (LIEQ) database. The LIEQ database includes 1,000 light-enhanced images, which are generated from 100 low-light images using 10 LIEAs. Rather than evaluating the quality of light-enhanced images directly, which is more difficult, we propose to use the multi-exposure fused (MEF) image and stack-based high dynamic range (HDR) image as a reference and evaluate the quality of low-light enhancement following a full-reference (FR) quality assessment routine. We observe that distortions introduced in low-light enhancement are significantly different from distortions considered in traditional image IQA databases that are well-studied, and the current state-of-the-art FR IQA models are also not suitable for evaluating their quality. Therefore, we propose a new FR low-light image enhancement quality assessment (LIEQA) index by evaluating the image quality from four aspects: luminance enhancement, color rendition, noise evaluation, and structure preserving, which have captured the most key aspects of low-light enhancement. Experimental results on the LIEQ database show that the proposed LIEQA index outperforms the state-of-the-art FR IQA models. LIEQA can act as an evaluator for various low-light enhancement algorithms and systems. To the best of our knowledge, this article is the first of its kind comprehensive low-light image enhancement quality assessment study. Guangtao Zhai, Wei Sun 0029, Xiongkuo Min, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationabstractIn this work, we present a novel data-driven method for robust 6DoF object pose estimation from a single RGBD image. Unlike previous methods that directly regressing pose parameters, we tackle this challenging task with a keypoint-based approach. Specifically, we propose a deep Hough voting network to detect 3D keypoints of objects and then estimate the 6D pose parameters within a least-squares fitting manner. Our method is a natural extension of 2D-keypoint approaches that successfully work on RGB based 6DoF estimation. It allows us to fully utilize the geometric constraint of rigid objects with the extra depth information and is easy for a network to learn and optimize. Extensive experiments were conducted to demonstrate the effectiveness of 3D-keypoint detection in the 6D pose estimation task. Experimental results also show our method outperforms the state-of-the-art methods by large margins on several benchmarks. Code and video are available at https://github.com/ethnhe/PVN3D.git. Yisheng He, Wei Sun 0029, Jianran Liu, Haoqiang Fan, Jian Sun 0001 |
CVPR | 2 |
| 2019 | LPHD: A Large-Scale Head Pose Dataset for RGB ImagesabstractHead pose estimation has attracted many research interest in recent years. With the advent of deep learning, it is possible to predict the head pose accurately from the RGB images without the help of facial landmarks or depth information. However, existing head pose datasets often lack large pose head images, which extremely limits the development of head pose estimation algorithms. In this paper, we build the largescale head pose dataset (LHPD) including more than 140,000 images with the diverse and accurate head poses. The LHPD dataset includes the head images recorded from different shooting angles between the camera and the human body for the first time, which greatly expands the range of head pose compared to previous datasets. Therefore, the range of head pose can cover +/-90° for each Euler angle. The accurate and reliable head pose annotation is labeled by the motion capture system and careful calibration procedures. We then propose a head pose estimation method through fine-tuning the ResNet on the LHPD dataset when using the Euclidean distance of quaternions as the loss function. The results show that our method achieves better performance than current state-of-theart algorithms. Wei Sun 0029, Yezhao Fan, Xiongkuo Min, Shihao Peng, Siwei Ma 0001, Guangtao Zhai |
ICME | 1 |
| 2019 | MC360IQA: The Multi-Channel CNN for Blind 360-Degree Image Quality AssessmentabstractIn this paper, we present a multi-channel convolution neural network (CNN) for blind 360-degree image quality assessment (MC360IQA). To be consistent with the visual content of 360-degree images seen in the VR device, our model adopts the viewport images as the input. Specifically, we project each 360-degree image into six viewport images to cover omnidirectional visual content. By rotating the longitude of the front view, we can project one omnidirectional image onto lots of different groups of viewport images, which is an efficient way to avoid overfitting. MC360IQA consists of two parts, multi-channel CNN and image quality regressor. Multi-channel CNN includes six parallel ResNet34 networks, which are used to extract the features of the corresponding six viewport images. Image quality regressor fuses the features and regresses them to final scores. The results show that our model achieves the best performance among the state-of-art full-reference (FR) and no-reference (NR) image quality assessment (IQA) models on the available 360-degree IQA database. Wei Sun 0029, Weike Luo, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001, Ke Gu 0001, Siwei Ma 0001 |
ISCAS | 1 |
| 2018 | A Large-Scale Compressed 360-Degree Spherical Image Database: From Subjective Quality Evaluation to Objective Model Comparisonabstract360-degree images/videos have been dramatically increasing in recent years. But the high resolution makes it difficult to be transported, compressed and stored, and thus constrains the development of 360-degree images/videos. Therefore, it is important to study how popular coding technologies influence the quality of 360-degree images. In this paper, we present a study on subjective assessment of compressed 360-degree images and investigate whether existing objective image quality assessment (IQA) methods can effectively evaluate the quality of compressed 360-degree images. We first construct the largest compressed 360-degree image database (CVIQD2018) including 16 source images and 528 compressed ones with three prevailing coding technologies. Then, we implement 16 full reference (FR) IQA metrics, which include 10 traditional IQA metrics for 2D images and 3 PSNR-based metrics for 360-degree images, as well as 5 no reference (NR) IQA metrics and calculate the correlation between each above metric and subjective assessment in terms of three commonly used performance indices. The experiment results reveal structure information, visual saliency information and compensation for geometric distortion are crucial for evaluating the quality of compressed 360-degree images. Wei Sun 0029, Ke Gu 0001, Siwei Ma 0001, Wenhan Zhu, Guangtao Zhai |
MMSP | 1 |
| 2018 | Fast MPEG-CDVS Encoder With GPU-CPU Hybrid ComputingabstractThe compact descriptors for visual search (CDVS) standard from ISO/IEC moving pictures experts group has succeeded in enabling the interoperability for efficient and effective image retrieval by standardizing the bitstream syntax of compact feature descriptors. However, the intensive computation of a CDVS encoder unfortunately hinders its widely deployment in industry for large-scale visual search. In this paper, we revisit the merits of low complexity design of CDVS core techniques and present a very fast CDVS encoder by leveraging the massive parallel execution resources of graphics processing unit (GPU). We elegantly shift the computation-intensive and parallel-friendly modules to the state-of-the-arts GPU platforms, in which the thread block allocation as well as the memory access mechanism are jointly optimized to eliminate performance loss. In addition, those operations with heavy data dependence are allocated to CPU for resolving the extra but non-necessary computation burden for GPU. Furthermore, we have demonstrated the proposed fast CDVS encoder can work well with those convolution neural network approaches which enables to leverage the advantages of GPU platforms harmoniously, and yield significant performance improvements. Comprehensive experimental results over benchmarks are evaluated, which has shown that the fast CDVS encoder using GPU-CPU hybrid computing is promising for scalable visual search. Ling-Yu Duan, Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Jianxiong Yin, Simon See, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | CVIQD: Subjective quality evaluation of compressed virtual reality imagesabstractThe 360-degree spherical images/videos, also called Virtual Reality (VR) images/videos, can provide immersive experience of the real-world scenes in some specific systems. This makes it widely employed in concerts/sports events live and VR movies. However, it is difficult to transport, compress or store VR images/videos due to their high resolution. So it is significant to research how the popular coding technologies influence the quality of VR images. To this aim, this paper carries out subjective quality evaluation of compressed VR images and examines the correlation performance of popular objective quality measures in accordance with the aforesaid subjective ratings. We first establish a Compressed VR Image Quality Database (CVIQD), which includes five source VR images and associated 165 compressed images under three prevailing coding technologies. The Single-Stimulus (SS) method is exploited to collect the subjective scores from 20 inexperienced viewers. Next, we implement 10 classical and recent objective quality metrics on the CVIQD database and compute the correlation between each above quality metric and subjective assessment in terms of five commonly used performance indices. Experimental results reveal that multi-scale based MS-SSIM and ADD-SSIM models have lead to high correlation with human visual perception. Wei Sun 0029, Ke Gu 0001, Guangtao Zhai, Siwei Ma 0001, Weisi Lin, Patrick Le Callet |
ICIP | 1 |
| 2017 | GPU Based fast MPEG-CDVS encoderabstractThe compact descriptors for visual search (CDVS) standard from ISO/IEC Moving Picture Experts Group (MPEG) has received increasing attentions due to its effectiveness in mobile visual search related applications. To explore the efficiency of CDVS in real-time applications, we implement the first optimized CDVS encoder based on the graphics processing unit (GPU). In particular, the CDVS feature extraction is implemented in a CPU and GPU collaborative architecture, and the most of the computation-intensive operations are transferred to GPU platform, which achieves significant speedup for CDVS compared with CPU implementation. Extensive evaluations on standard datasets provide promising results of the proposed scheme for real applications scenarios. Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Ling-Yu Duan |
ICIP | 1 |
| 2017 | Dynamic backlight scaling considering ambient luminance for mobile energy savingabstractThe mobile video playback involves many subsystems of the devices such as computing, rendering and displaying subsystems. Among all subsystems, the displaying subsystem accounts for at least 38% of all consumed power, and it can be up to 68% with the maximum backlight brightness. What is more, lots of people watch videos via mobile devices in various situations, where the ambient luminance condition is different. Therefore, how to save mobile energy and improve the Quality of Experience (QoE) in different situations become significant problems. In this paper, we try to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users' QoE. In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance and backlight level. We first study whether the continuous playback of backlight-scaled shots using the proposed scaling magnitude would cause flicker effect or not. Then motivated by the findings of these subject studies, we implement a Dynamic Backlight Scaling (DBS) strategy. The experiment results demonstrate that the DBS strategy can save more than 40% power at most and can also save 10% power even at a very high ambient luminance. Wei Sun 0029, Guangtao Zhai, Xiongkuo Min, Yutao Liu 0002, Siwei Ma 0001, Jing Liu 0002, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 1 |