Weixia Zhang

dblp:196/3124 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-3634-2630ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning
abstract
Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision.
Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min
AAAI3
2026 LEIQ-Assessor: Multi-Dimensional Quality Assessment of Low-Light Enhanced Images via Multi-Task Learning
Wei Sun 0029, Yanwei Jiang, Dandan Zhu 0001, Jinqiu Sang, Jikai Xu, Weixia Zhang, Guangtao Zhai
QoMEX6
2026 QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu
QoMEX36
2026 Robust low-light image enhancement in the wild via data synthesis and generative diffusion prior
Zhihua Wang 0002, Qinghua Lin, Weixia Zhang, Wei Zhou 0021
Pattern Recognit.4
2026 Unveiling the Modular Design of Text-to-Video Quality Assessment
abstract
Artificial intelligence generated content (AIGC) are reshaping digital media creation, with text-to-video (T2V) generation emerging as one of its most powerful and widely-used techniques. Despite rapid progress, videos generated by T2V models still suffer from issues such as unrealistic spatial details, temporal inconsistencies, and content misalignment with the input textual prompts. It is of high importance to develop computational video quality assessment (VQA) models for T2V videos to ensure a favorable quality-of-experience (QoE) for end users. Towards comprehensive quality evaluation of AI-generated videos, modern T2V quality assessment (T2V QA) models typically integrate multiple modules that excel in capturing different and complementary quality-aware features. In this paper, we categorize the constituent modules of modern T2V QA models into four types: base quality evaluators, spatial perception modules, temporal perception modules, and text-video alignment modules. Within this framework, we systematically evaluate and compare the relative strengths and weaknesses of candidate models. Through experiments on multiple T2V datasets, we verify that the top-performing models from this competition demonstrate very competitive performance against existing VQA methods. The resulting framework is agnostic to specific architectural designs and can be continuously refined by integrating advancements from each of its constituent modules, making it well-suited for adapting T2V QA methods to the fast-evolving T2V generation techniques. Our source code is available at https://github.com/CH053N0N3/mineBVQA.
Bingkun Zheng, Weixia Zhang, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Evaluating Perceptual Color Preferences in Smartphone Photography: Dataset and Challenges
abstract
International audience
Zhihua Wang 0002, Weixia Zhang, Wei Zhou 0021, Xiaohong Liu 0001, Guangtao Zhai, Patrick Le Callet
ACM Multimedia2
2025 Learning from Model Rankings Improves Blind Super-Resolution Image Quality Assessment
abstract
Image super-resolution (SR) aims to generate a high-resolution (HR) image from a low-resolution (LR) input. Traditionally, full-reference image quality assessment (FR-IQA) models have been widely used to evaluate the perceptual quality of super-resolved images, relying on pristine reference images as the gold standard. However, in real-world SR applications, such reference images are often unavailable, posing challenges for the use of FR-IQA. While blind image quality assessment (BIQA) models can assess the perceptual quality of super-resolved images without requiring a reference, there remains a lack of comprehensive studies evaluating the effectiveness of existing BIQA models for real-world SR tasks. This dilemma can largely be attributed to the high cost of subjective testing required to collect sufficient human quality annotations, which in turn hinders the development of effective SR-IQA models. In this study, we tackle this challenge with a data-efficient approach. We first generate super-resolved images from LR inputs using state-of-the-art real-world SR methods. Then, we use the maximum differentiation competition (MAD) to select a diverse set of images for subjective testing, allowing us to efficiently gather human preferences and assess the alignment between BIQA model predictions and human judgments. The resulting global ranking of SR methods not only indicates the relative performance of recent real-world SR models, but also gives us an opportunity to develop a new BIQA model tailored for real-world SR-IQA. By utilizing the global rankings of SR algorithms as prior knowledge, we can refine pretrained BIQA models using vast amounts of super-resolved images without any supervisory signal. Experimental results show that our approach substantially enhances IQA performance for real-world SR while preserving robust predictive accuracy across various distortion scenarios. The dataset and the code are available at https://github.com/cschenjunlin/SR-IQA-SMC25.
Junlin Chen, Peibei Cao, Guangtao Zhai, Xiaokang Yang 0001, Weixia Zhang
SMC5
2024 A Comparative Study of Perceptual Quality Metrics For Audio-Driven Talking Head Videos
abstract
The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code is available at https://github.com/zwx8981/ADTH-QA.
Weixia Zhang, Chengguang Zhu, Jingnan Gao, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001
ICIP1
2024 Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
abstract
The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released.
Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ICML3
2024 Task-Specific Normalization for Continual Learning of Blind Image Quality Models
abstract
In this paper, we present a simple yet effective continual learning method for blind image quality assessment (BIQA) with improved quality prediction accuracy, plasticity-stability trade-off, and task-order/-length robustness. The key step in our approach is to freeze all convolution filters of a pre-trained deep neural network (DNN) for an explicit promise of stability, and learn task-specific normalization parameters for plasticity. We assign each new IQA dataset (i.e., task) a prediction head, and load the corresponding normalization parameters to produce a quality score. The final quality estimate is computed by a weighted summation of predictions from all heads with a lightweight K -means gating mechanism. Extensive experiments on six IQA datasets demonstrate the advantages of the proposed method in comparison to previous training techniques for BIQA.
Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Image Process.1
2023 Blind Image Quality Assessment via Vision-Language Correspondence: A Multitask Learning Perspective
abstract
We aim at advancing blind image quality assessment (BIQA), which predicts the human perception of image quality without any reference information. We develop a general and automated multitask learning scheme for BIQA to exploit auxiliary knowledge from other tasks, in a way that the model parameter sharing and the loss weighting are determined automatically. Specifically, we first describe all candidate label combinations (from multiple tasks) using a textual template, and compute the joint probability from the cosine similarities of the visual-textual embeddings. Predictions of each task can be inferred from the joint distribution, and optimized by carefully designed loss functions. Through comprehensive experiments on learning three tasks - BIQA, scene classification, and distortion type identification, we verify that the proposed BIQA method 1) benefits from the scene classification and distortion type identification tasks and outperforms the state-of-the-art on multiple IQA datasets, 2) is more robust in the group maximum differentiation competition, and 3) realigns the quality annotations from different IQA datasets more effectively. The source code is available at https://github.com/zwx8981/LIQE.
Weixia Zhang, Guangtao Zhai, Ying Wei 0001, Xiaokang Yang 0001, Kede Ma
CVPR1
2023 Llieformer: A Low-Light Image Enhancement Transformer Network with a Degraded Restoration Model
abstract
Low-light image enhancement aims at improving human perception or the effectiveness of computer vision tasks of images taken in dark. The low-light images are usually seriously lack in visual information. To tackle this problem, we propose a general Low-light Image Enhancement Transformer Network (LLIEFormer) with a degraded restoration model in this paper. The network of LLIEFormer synthesizes the advantages of Transformer to extract global information and convolutional neural networks to capture local details. We conduct extensive experiments on various low-illumination enhanced datasets including PairL1.6K and FiveK to demonstrate the effectiveness of our method. The results show that our LLIEFormer has better performance and wider applicability than other advanced methods. Our code will be available at https://github.com/xunpengyi/LLIEFormer.
Xunpeng Yi, Yizhen Zhao, Jia Yan 0006, Weixia Zhang
ICIP5
2023 Continual Learning for Blind Image Quality Assessment
abstract
The explosive growth of image data facilitates the fast development of image processing and computer vision methods for emerging visual applications, meanwhile introducing novel distortions to processed images. This poses a grand challenge to existing blind image quality assessment (BIQA) models, which are weak at adapting to subpopulation shift. Recent work suggests training BIQA methods on the combination of all available human-rated IQA datasets. However, this type of approach is not scalable to a large number of datasets and is cumbersome to incorporate a newly created dataset as well. In this paper, we formulate continual learning for BIQA, where a model learns continually from a stream of IQA datasets, building on what was learned from previously seen data. We first identify five desiderata in the continual setting with three criteria to quantify the prediction accuracy, plasticity, and stability, respectively. We then propose a simple yet effective continual learning method for BIQA. Specifically, based on a shared backbone network, we add a prediction head for a new dataset and enforce a regularizer to allow all prediction heads to evolve with new data while being resistant to catastrophic forgetting of old data. We compute the overall quality score by a weighted summation of predictions from all heads. Extensive experiments demonstrate the promise of the proposed continual learning method in comparison to standard training techniques for BIQA, with and without experience replay. We made the code publicly available at https://github.com/zwx8981/BIQA_CL.
Weixia Zhang, Dingquan Li, Chao Ma 0004, Guangtao Zhai, Xiaokang Yang 0001, Kede Ma
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Learning a Blind Quality Evaluator for UGC Videos in Perceptually Relevant Domains
abstract
The absence of reference videos with pristine quality and the complexity of authentic distortions pose great challenges for blind video quality assessment (BVQA) towards user-generated content (UGC) videos. Although it is straightfor-ward to leverage transfer learning techniques to learn effective BVQA models, it is nontrivial to explore how to bridge the domain shifts for better video representation learning. In this work, we propose to transfer meaningful knowledge from perceptually relevant domains, i.e., image quality assessment (IQA) with authentic distortions and video classification with rich motion patterns. We develop a promising strategy to use both groups of data to learn the feature extractors. We train the proposed model on the target VQA databases using a mixed list-wise ranking loss function. Extensive experiments on six VQA databases demonstrate that our method performs very competitively under both individual database and mixed database training settings. Codes and models are available at https://github.com/zwx8981/BVQA-2021.
Bowen Li 0018, Weixia Zhang, Jiu Jiang, Guangtao Zhai, Xianpei Wang
ICME2
2022 Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-Loop
abstract
No-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimization of man-made vision systems. Here we make one of the first attempts to examine the perceptual robustness of NR-IQA models. Under a Lagrangian formulation, we identify insightful connections of the proposed perceptual attack to previous beautiful ideas in computer vision and machine learning. We test one knowledge-driven and three data-driven NR-IQA methods under four full-reference IQA models (as approximations to human perception of just-noticeable differences). Through carefully designed psychophysical experiments, we find that all four NR-IQA models are vulnerable to the proposed perceptual attack. More interestingly, we observe that the generated counterexamples are not transferable, manifesting themselves as distinct design flows of respective NR-IQA methods. Source code are available at https://github.com/zwx8981/PerceptualAttack_BIQA.
Weixia Zhang, Dingquan Li, Xiongkuo Min, Guangtao Zhai, Guodong Guo, Xiaokang Yang 0001, Kede Ma
NeurIPS1
2022 Blindly Assess Quality of In-the-Wild Videos via Quality-Aware Pre-Training and Motion Perception
abstract
Perceptual quality assessment of the videos acquired in the wilds is of vital importance for quality assurance of video services. The inaccessibility of reference videos with pristine quality and the complexity of authentic distortions pose great challenges for this kind of blind video quality assessment (BVQA) task. Although model-based transfer learning is an effective and efficient paradigm for the BVQA task, it remains to be a challenge to explorewhatandhowto bridge the domain shifts for better video representation. In this work, we propose to transfer knowledge from image quality assessment (IQA) databases with authentic distortions and large-scale action recognition with rich motion patterns. We rely on both groups of data to learn the feature extractor and use a mixed list-wise ranking loss function to train the entire model on the target VQA databases. Extensive experiments on six benchmarking databases demonstrate that our method performs very competitively under both individual database and mixed databases training settings. We also verify the rationality of each component of the proposed method and explore a simple ensemble trick for further improvement.
Bowen Li 0018, Weixia Zhang, Guangtao Zhai, Xianpei Wang
IEEE Trans. Circuits Syst. Video Technol.2
2021 Learning to predict the quality of distorted-then-compressed images via a deep neural network
Bowen Li 0018, Weixia Zhang, Hongtai Yao, Xianpei Wang
J. Vis. Commun. Image Represent.3
2021 Language-Guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
abstract
The emerging vision-and-language navigation (VLN) problem aims at learning to navigate an agent to the target location in unseen photo-realistic environments according to the given language instruction. The main challenges of VLN arise mainly from two aspects: first, the agent needs to attend to the meaningful paragraphs of the language instruction corresponding to the dynamically-varying visual environments; second, during the training process, the agent usually imitate the expert demonstrations, i.e., the shortest-path to the target location specified by associated language instructions. Due to the discrepancy of action selection between training and inference, the agent solely on the basis of imitation learning does not perform well. Existing VLN approaches address this issue by sampling the next action from its predicted probability distribution during the training process. This allows the agent to explore diverse routes from the environments, yielding higher success rates. Nevertheless, without being presented with the golden shortest navigation paths during the training process, the agent may arrive at the target location through an unexpected longer route. To overcome these challenges, we design a cross-modal grounding module, which is composed of two complementary attention mechanisms, to equip the agent with a better ability to track the correspondence between the textual and visual modalities. We then propose to recursively alternate the learning schemes of imitation and exploration to narrow the discrepancy between training and inference. We further exploit the advantages of both these two learning schemes via adversarial learning. Extensive experimental results on the Room-to-Room (R2R) benchmark dataset demonstrate that the proposed learning scheme is generalized and complementary to prior arts. Our method performs well against state-of-the-art approaches in terms of effectiveness and efficiency.
Weixia Zhang, Chao Ma 0004, Qi Wu 0001, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Uncertainty-Aware Blind Image Quality Assessment in the Laboratory and Wild
abstract
Performance of blind image quality assessment (BIQA) models has been significantly boosted by end-to-end optimization of feature engineering and quality regression. Nevertheless, due to the distributional shift between images simulated in the laboratory and captured in the wild, models trained on databases with synthetic distortions remain particularly weak at handling realistic distortions (and vice versa). To confront the cross-distortion-scenario challenge, we develop a unified BIQA model and an approach of training it for both synthetic and realistic distortions. We first sample pairs of images from individual IQA databases, and compute a probability that the first image of each pair is of higher quality. We then employ the fidelity loss to optimize a deep neural network for BIQA over a large number of such image pairs. We also explicitly enforce a hinge constraint to regularize uncertainty estimation during optimization. Extensive experiments on six IQA databases show the promise of the learned method in blindly assessing image quality in the laboratory and wild. In addition, we demonstrate the universality of the proposed training strategy by using it to improve existing BIQA models.
Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Image Process.1
2020 Learning To Blindly Assess Image Quality In The Laboratory And Wild
abstract
Computational models for blind image quality assessment (BIQA) are typically trained in well-controlled laboratory environments with limited generalizability to realistically distorted images. Similarly, BIQA models optimized for images captured in the wild cannot adequately handle synthetically distorted images. To face the cross-distortion-scenario challenge, we develop a BIQA model and an approach of training it on multiple IQA databases (of different distortion scenarios) simultaneously. A key step in our approach is to create and combine image pairs within individual databases as the training set, which effectively bypasses the issue of perceptual scale realignment. We compute a continuous quality annotation for each pair from the corresponding human opinions, indicating the probability of one image having better perceptual quality. We train a deep neural network for BIQA over the training set of massive image pairs by minimizing the fidelity loss. Experiments on six IQA databases demonstrate that the optimized model by the proposed training strategy is effective in blindly assessing image quality in the laboratory and wild, outperforming previous BIQA methods by a large margin.
Weixia Zhang, Kede Ma, Guangtao Zhai, Xiaokang Yang 0001
ICIP1
2020 Blind Image Quality Assessment Using a Deep Bilinear Convolutional Neural Network
abstract
We propose a deep bilinear model for blind image quality assessment that works for both synthetically and authentically distorted images. Our model constitutes two streams of deep convolutional neural networks (CNNs), specializing in two distortion scenarios separately. For synthetic distortions, we first pre-train a CNN to classify the distortion type and the level of an input image, whose ground truth label is readily available at a large scale. For authentic distortions, we make use of a pre-train CNN (VGG-16) for the image classification task. The two feature sets are bilinearly pooled into one representation for a final quality prediction. We fine-tune the whole network on the target databases using a variant of stochastic gradient descent. The extensive experimental results show that the proposed model achieves state-of-the-art performance on both synthetic and authentic IQA databases. Furthermore, we verify the generalizability of our method on the large-scale Waterloo Exploration Database, and demonstrate its competitiveness using the group maximum differentiation competition methodology.
Weixia Zhang, Kede Ma, Jia Yan 0006, Dexiang Deng, Zhou Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Hierarchical Features Fusion for Image Aesthetics Assessment
abstract
Image aesthetics assessment is an interesting yet challenging topic which can be applied on numerous scenarios such as high quality image retrieval or recommendation systems. We propose a hierarchical features fusion aesthetic assessment (HFFAA) model for this task. HFFAA is a two-stream convolutional neural network (CNN) which is composed of two branches with heterogeneous and complementary aesthetic perceptual abilities. HFFAA learns the mapping from deep image representation into their ground-truth aesthetic labels (good or bad) in an end-to-end fashion. Extensive experiments demonstrate that the proposed model achieves superior performance on two widely evaluated public benchmark databases, i.e., CUHKPQ and AVA. We also validate the rationality of the designs of HFFAA through a series of ablation experiments.
Weixia Zhang, Guangtao Zhai, Xiaokang Yang 0001, Jia Yan 0006
ICIP1
2016 Blind Image Quality Assessment Based on Natural Redundancy Statistics
Jia Yan 0006, Weixia Zhang, Tianpeng Feng
ACCV (4)2