VLDB 2026 Research / reviewers in the wild / expert
Yuqin Cao
dblp:312/7606
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-5087-6559ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio-Assisted Face Video Restoration with Temporal and Identity Complementary LearningabstractFace videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between visual and audio features, particularly in the mouth region. Several audio-aided face video restoration methods have been proposed, but they only focus on compression artifact removal. In this paper, we propose a General Audio-assisted face Video restoration Network (GAVN) to address various types of streaming video distortions via identity and temporal complementary learning. Specifically, GAVN first captures inter-frame temporal features in the low-resolution space to restore frames coarsely and save computational cost. Then, GAVN extracts intra-frame identity features in the high-resolution space with the assistance of audio signals and face landmarks to restore more facial details. Finally, the reconstruction module integrates temporal features and identity features to generate high-quality face videos. Experimental results demonstrate that GAVN outperforms the existing state-of-the-art methods on face video compression artifact removal, deblurring, and super-resolution. Yuqin Cao, Wei Sun 0029, Xiaohong Liu 0001, Yulun Zhang 0001, Xiongkuo Min |
AAAI | 1 |
| 2026 | Q-Agent: An MLLM-Driven Framework for Universal Visual Quality Assessment
Peihang Chen, Huiyu Duan, Zitong Xu, Yuqin Cao, Sijing Wu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
QoMEX | 6 |
| 2026 | UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual ContentabstractAs multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A/V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A/V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications. Code are available at https://github.com/charlotte9524/UNQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Long Ye, Weisi Lin, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight ModelabstractSurveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications. Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model, and Training StrategyabstractThe rapid development of multimodal large language models has resulted in remarkable advancements in visual perception and understanding, consolidating several tasks into a single visual question-answering framework. However, these models are prone to hallucinations, which limit their reliability as artificial intelligence systems. While this issue is extensively researched in natural language processing and image captioning, there remains a lack of investigation of hallucinations in Low-level Visual Perception and Understanding (HLPU), especially in the context of image quality assessment tasks. We consider that these hallucinations arise from an absence of clear self-awareness within the models. To address this issue, we first introduce the HLPU instruction database, the first instruction database specifically focused on hallucinations in low-level vision tasks. This database contains approximately 200K question-answer pairs and comprises four subsets, each covering different types of instructions. Subsequently, we propose the Self-Awareness Failure Elimination (SAFEQA) model, which utilizes image features, salient region features and quality features to improve the perception and comprehension abilities of the model in low-level vision tasks. Furthermore, we propose the Enhancing Self-Awareness Preference Optimization (ESA-PO) framework to increase the model’s awareness of knowledge boundaries, thereby mitigating the incidence of hallucination. Finally, we conduct comprehensive experiments on low-level vision tasks, with the results demonstrating that our proposed method significantly enhances self-awareness of the model in these tasks and reduces hallucinations. Notably, our proposed method improves both accuracy and self-awareness of the proposed model and outperforms close-source models in terms of various evaluation metrics. This research contributes to the advancement of self-awareness capabilities in multimodal large language models, particularly for low-level visual perception and understanding tasks. Xiongkuo Min, Yuqin Cao, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality AssessmentabstractMany video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing audio-visual quality assessment methods struggle with unique distortions in AGAVs, such as unrealistic and inconsistent elements. To address this, we introduce AGAVQA-3k, the first large-scale AGAV quality assessment dataset, comprising $3,382$ AGAVs from $16$ VTA methods. AGAVQA-3k includes two subsets: AGAVQA-MOS, which provides multi-dimensional scores for audio quality, content consistency, and overall quality, and AGAVQA-Pair, designed for optimal AGAV pair selection. We further propose AGAV-Rater, a LMM-based model that can score AGAVs, as well as audio and music generated from text, across multiple dimensions, and selects the best AGAV generated by VTA methods to present to the user. AGAV-Rater achieves state-of-the-art performance on AGAVQA-3k, Text-to-Audio, and Text-to-Music datasets. Subjective tests also confirm that AGAV-Rater enhances VTA performance and user experience. The dataset and code is available at https://github.com/charlotte9524/AGAV-Rater. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICML | 1 |
| 2025 | Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodabstractWith the rise of Text-to-Image (T2I) models, generating face images from text prompts has emerged as a prominent research area. However, evaluating the quality of these generated face images, particularly with respect to fine-grained facial attributes, remains a significant challenge. To address this, we introduce the Fine -grained Text-to-Face Image Quality Assessment (FineTFIQA) database, which is designed to evaluate the ability of T2I models to generate fine-grained face images. To the best of our knowledge, this database is the largest of its kind, containing 7,218 face images generated from 1,000 text prompts that cover 111 distinct facial attributes. A large group of subjects was invited to assess the quality of text-to-face images on four evaluation dimensions: perceptual quality, human likeness, attractiveness, and consistency. Additionally, we develop the Multi-Dimensional Text-to-Face Image Quality Assessment (MDTFIQA) method based on the Large Language Model (LLM), which combines both face image features and text features to evaluate generated images on all evaluation dimensions. Extensive experimental results demonstrate that traditional face image assessment methods and general image quality assessment methods are inadequate for accurately evaluating generated text-to-face images. Our method significantly outperforms these existing methods on all evaluation dimensions, proving to be an effective method for assessing the quality of generated text-to-face images. Xiongkuo Min, Jinliang Han, Yuqin Cao, Sijing Wu, Yunze Dou, Guangtao Zhai |
ACM Multimedia | 4 |
| 2025 | Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture GenerationabstractThe Audio-to-3D-Gesture (A2G) task exhibits significant potential across domains including virtual reality, computer graphics, and 3D animation production. However, current evaluation metrics, such as Fréchet Gesture Distance or Beat Constancy, fail at reflecting the human preference of the generated 3D gestures. To cope with this problem, exploring human preference and an objective quality assessment metric for AI-generated 3D human gestures is becoming increasingly significant. In this paper, we introduce the Ges-QA dataset, which includes 1,400 samples with multidimensional scores for gesture quality and audio-gesture consistency. Moreover, we collect binary classification labels to determine whether the generated gestures match the emotions of the audio. Equipped with our Ges-QA dataset, we propose a multi-modal transformer-based neural network with 3 branches for video, audio and 3D skeleton modalities, which can score A2G contents in multiple dimensions. Comparative experimental results and ablation studies demonstrate that Ges-QAer yields state-of-the-art performance on our dataset. Zhilin Gao, Sijing Wu, Yuqin Cao, Huiyu Duan, Guangtao Zhai |
VCIP | 4 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 19 |
| 2025 | No-Reference Image Quality Assessment: Obtain MOS From Image Quality Score DistributionabstractRecent image quality assessment (IQA) methods typically focus on predicting the mean opinion score (MOS) of image quality, ignoring the image quality score distribution. This distribution provides valuable information beyond the MOS, including the standard deviation of opinion scores (SOS) and opinion scores at different quality levels. This paper introduces a novel no-reference IQA method that predicts the image quality score distribution to estimate the MOS. The proposed method consists of three modules: a visual feature extraction module, a graph convolutional module, and a MOS prediction module. In the visual feature extraction module, a convolutional neural network is designed to extract both first- and second-order visual features of images. The graph convolutional module employs a graph convolutional network (GCN)-based mapper to map these visual features to the image quality score distribution by exploring correlations between quality labels. The MOS is then derived from the predicted image quality score distribution in the MOS prediction module. We are the first to jointly train the method using both the MOS and the image quality score distribution, enabling it to learn richer subjective information and improve prediction performance. To address the lack of the ground-truth image quality score distribution in some IQA databases, we propose to use a SOS assumption to generate a Gaussian-based image quality score distribution that better reflects subjective perception. Additionally, we design appropriate loss functions for training. Experimental results demonstrate that our method effectively predicts both the image quality score distribution and the MOS, outperforming most state-of-the-art IQA methods. Xiongkuo Min, Yuqin Cao, Xiaohong Liu 0001, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Exploring Rich Subjective Quality Information for Image Quality Assessment in the WildabstractTraditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method namedRichIQAto explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: 1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; 2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub. Xiongkuo Min, Yuqin Cao, Guangtao Zhai, Wenjun Zhang 0001, Huifang Sun, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Blind Image Quality Assessment by Gaussian Mixture DistributionabstractIn the field of image quality assessment (IQA), researchers have been studying the mean opinion score (MOS) of image quality for decades. They focus on developing IQA methods with the help of MOS without using the potential of the distribution of opinion scores (DOS). We find that the Gaussian mixture distribution (GMD) can more accurately describe the DOS of image quality on SJTU IQSD and KonIQ-10K databases compared to some traditional distributions. Therefore, this paper proposes a blind IQA method that predicts the MOS of image quality by learning the GMD-based image quality. The proposed method consists of a visual feature learning module and a GMD learning module. The visual feature learning module uses a multi-stage Swin Transformer model and a CLIP feature extractor to extract visual features from an image. The GMD learning module then maps the extracted visual features to the GMD-based image quality using a mixture density network, where the mean of the GMD represents the MOS of image quality. We not only use the MOS of image quality to train the proposed method, but also employ the DOS of image quality for auxiliary training to improve the prediction performance of the proposed method. To address the lack of DOS in some existing IQA databases, we introduce a pseudo DOS generation strategy to generate the DOS of image quality for training, which significantly improves the applicability of the proposed method. Numerous analyses show that the proposed method is superior to most state-of-the-art IQA methods in predicting both the MOS and the DOS, thus facilitating a deeper investigation into the DOS of image quality in IQA. Xiongkuo Min, Yuqin Cao, Weisi Lin, Bu-Sung Lee, Guangtao Zhai |
IEEE Trans. Image Process. | 3 |
| 2025 | Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional VideosabstractVirtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet. Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 3 |
| 2024 | Multidimensional Similarity Fusion for Speech Quality AssessmentabstractSince the perceptual quality of audio signals is easily to be affected by compression, transmission, noise adding, etc, it is of great significance to develop an effective audio quality assessment (AQA) method to measure end-user’s quality of experience. In this paper, we propose a full reference AQA model named Multidimensional Similarity Fusion for Audio Quality Assessment (MSF-AQA). We generalize the similarity-based image quality assessment methods for audio, then extract audio similarity features from multiple dimensions, and finally regress the multidimensional similarity features into the final quality score. The experimental results across three databases indicate that our MSF-AQA model outperforms the state-of-the-art AQA methods. Xiongkuo Min, Yuqin Cao, Xiao-Ping Zhang 0002, Guangtao Zhai |
ISCAS | 3 |
| 2023 | Audio-Visual Quality Assessment for User Generated Content: Database and MethodabstractWith the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users’ Quality of Experience (QoE). However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the user’s QoE also depends on the accompanying audio signals. In this paper, we conduct the first study to address the problem of UGC audio and video quality assessment (AVQA). Specifically, we construct the first UGC AVQA database named the SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences, and conduct a user study to obtain the mean opinion scores of the A/V sequences. The content of the SJTU-UAV database is then analyzed from both the audio and video aspects to show the database characteristics. We also design a family of AVQA models, which fuse the popular VQA methods and audio features via support vector regressor (SVR). We validate the effectiveness of the proposed models on the three databases. The experimental results show that with the help of audio signals, the VQA models can evaluate the perceptual quality more accurately. The database will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 1 |
| 2023 | Attention-Guided Neural Networks for Full-Reference and No-Reference Audio-Visual Quality AssessmentabstractWith the popularity of mobile Internet, audio and video (A/V) have become the main way for people to entertain and socialize daily. However, in order to reduce the cost of media storage and transmission, A/V signals will be compressed by service providers before they are transmitted to end-users, which inevitably causes distortions in the A/V signals and degrades the end-user's Quality of Experience (QoE). This motivates us to research the objective audio-visual quality assessment (AVQA). In the field of AVQA, most previous works only focus on single-mode audio or visual signals, which ignores that the perceptual quality of users depends on both audio and video signals. Therefore, we propose an objective AVQA architecture for multi-mode signals based on attentional neural networks. Specifically, we first utilize an attention prediction model to extract the salient regions of video frames. Then, a pre-trained convolutional neural network is used to extract short-time features of the salient regions and the corresponding audio signals. Next, the short-time features are fed into Gated Recurrent Unit (GRU) networks to model the temporal relationship between adjacent frames. Finally, the fully connected layers are utilized to fuse the temporal related features of A/V signals modeled by the GRU network into the final quality score. The proposed architecture is flexible and can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database and UnB-AVC Database demonstrate that our model outperforms the state-of-the-art AVQA methods. The code of the proposed method will be publicly available to promote the development of the field of AVQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Image Process. | 1 |
| 2023 | Subjective and Objective Audio-Visual Quality Assessment for User Generated ContentabstractIn recent years, User Generated Content (UGC) has grown dramatically in video sharing applications. It is necessary for service-providers to use video quality assessment (VQA) to monitor and control users' Quality of Experience when watching UGC videos. However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the perceptual quality also depends on the accompanying audio signals. In this paper, we conduct a comprehensive study on UGC audio-visual quality assessment (AVQA) from both subjective and objective perspectives. Specially, we construct the first UGC AVQA database named SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences collected from the YFCC100m database. A subjective AVQA experiment is conducted on the database to obtain the mean opinion scores (MOSs) of the A/V sequences. To demonstrate the content diversity of the SJTU-UAV database, we give a detailed analysis of the SJTU-UAV database as well as other two synthetically-distorted AVQA databases and one authentically-distorted VQA database, from both the audio and video aspects. Then, to facilitate the development of AVQA fields, we construct a benchmark of AVQA models on the proposed SJTU-UAV database and other two AVQA databases, of which the benchmark models consist of AVQA models designed for synthetically distorted A/V sequences and AVQA models built through combining the popular VQA methods and audio features via support vector regressor (SVR). Finally, considering benchmark AVQA models perform poorly in assessing in-the-wild UGC videos, we further propose an effective AVQA model via jointly learning quality-aware audio and visual feature representations in the temporal domain, which is seldom investigated by existing AVQA models. Our proposed model outperforms the aforementioned benchmark AVQA models on the SJTU-UAV database and two synthetically distorted AVQA databases. The SJTU-UAV database and the code of the proposed model will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Image Process. | 1 |
| 2021 | Deep Neural Networks For Full-Reference And No-Reference Audio-Visual Quality AssessmentabstractIn the field of audio and visual quality assessment, most of previous works only focused on the single-mode visual or audio signal. However, for multi-mode signals, such as video and the accompanying audio, the overall perceptual quality depends on both video and audio. In this paper, we proposed an objective audio-visual quality assessment (AVQA) architecture for multi-mode signals based on deep neural networks. We first use a pretrained convolutional neural network to extract features of the single video frames and the concurrent short audio segments. Then, the extracted features are fed into Gated Recurrent Unit networks for time sequence modeling. Finally, we utilize the fully connected layers to fuse the qualities of audio and visual signals into the final quality score. The proposed architecture can be applied to both full-reference and no-reference AVQA. Experimental results on the LIVE-SJTU Database prove that our model outperforms the state-of-the-art AVQA methods. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Guangtao Zhai |
ICIP | 1 |