VLDB 2026 Research / reviewers in the wild / expert
Zhihua Wang 0002
dblp:66/6802-2
· DBLP profile ↗
38ranked-venue papers
12as first author
37since 2021 · last 2026
0000-0002-4398-536XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 19 · 8 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Token-Wise Attention-Guided Semantic Quality Assessment for Compressed Visual Features
Shien Ke, Changsheng Gao, Hadi Amirpour, Zhihua Wang 0002, Xiaoyan Sun 0001 |
QoMEX | 4 |
| 2026 | InvJND: Just Noticeable Difference Estimation via Deep Invertible Network
Qiuping Jiang, Zhihua Wang 0002, Shiqi Wang 0001, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
Int. J. Comput. Vis. | 4 |
| 2026 | Towards stable cross-domain depression recognition under missing modalities
Jiuyi Chen, Mingkui Tan, Haifeng Lu, Qiuna Xu, Zhihua Wang 0002, Runhao Zeng, Xiping Hu |
Pattern Recognit. | 5 |
| 2026 | Robust low-light image enhancement in the wild via data synthesis and generative diffusion prior
Zhihua Wang 0002, Qinghua Lin, Weixia Zhang, Wei Zhou 0021 |
Pattern Recognit. | 1 |
| 2026 | Test-time adaptive vision-language alignment for zero-shot group activity recognition
Runhao Zeng, Wenfu Peng, Xionglin Zhu, Ronghao Zhang, Zhihua Wang 0002 |
Pattern Recognit. | 6 |
| 2026 | Amplitude exchanging network for unsupervised underwater image enhancement
Runhao Zeng, Xionglin Zhu, Wenfu Peng, Jiezhang Cao, Zhihua Wang 0002, Qiuping Jiang |
Pattern Recognit. | 6 |
| 2026 | Latent Fingerprint Quality Assessment for Criminal Investigations: A Benchmark Dataset and MethodabstractFingerprint biometrics plays a crucial role in biometric identification, especially in applications such as criminal investigations. Although recent progress in recognition methodology has significantly enhanced automated fingerprint recognition, these systems still rely heavily on the quality of the input fingerprints. In criminal investigations, fingerprints are often of low quality due to their incidental deposition from natural oils and sweat, rather than being deliberately captured under controlled conditions. This degradation can significantly impact usability and identification accuracy, underscoring the need for effective Fingerprint Quality Assessment (FQA) methods. In this paper, we establish the Crime Scene Fingerprints quality assessment Dataset (CSFD-10k), the largest dataset of its kind, containing 11,500 fingerprint images from real criminal investigations. Of these, 10,000 samples are assigned Mean Opinion Scores (MOSs) for correlation testing, while the remaining 1,500 are labeled based on matching performance for generalizability testing. All labels are provided by frontline criminal police officers. Using this dataset, we propose a deep neural network-based Dual-Branch FQA (DB-FQA) framework that integrates image-level and edge-level features. The DB-FQA enhances ridge details by transforming raw grayscale fingerprints into edge maps using the Logical/Linear operator. A dual-branch network processes both the raw fingerprint and the edge map, and the Multi-scale Adaptive Cross feature Fusion (MACF) module fuses these features, guided by the edge map to highlight quality-related regions of interest. Extensive experiments demonstrate the robustness and superiority of our proposed method, offering substantial support for forensic fingerprint biometrics. The code and dataset are available at https://github.com/wzhsysu/FIQA. Chao Huang 0008, Ye Zhang 0017, Peibei Cao, Zhihua Wang 0002, Yang Yu 0014, Xiaochun Cao |
IEEE Trans. Image Process. | 6 |
| 2026 | Harnessing Multi-Modal Large Language Models for Measuring and Interpreting Color DifferencesabstractThe accurate measurement of perceptual color differences (CDs) between two images plays an important role in modern smartphone photography. Although traditional CD metrics provide numerical scores to quantify color variations, they often lack the ability to offer intuitive insights or explanations that reflect the factors behind these differences in a way that aligns with human perception and reasoning. Here, we present CD-Reasoning, an innovative method designed not merely to compute numerical CD scores but also to provide a detailed rationale for the observed CDs between images. This method surpasses simple numerical quantification, delivering a more profound and explanatory analysis that bridges quantitative assessments with the qualitative reasoning characteristic of human perception. The development of the CD-Reasoning model begins with the compilation of a multi-modal CD dataset dubbed M-SPCD based on the existing SPCD, where we collect textual descriptions that detail the quantification of CDs across seven pivotal attributes: white balance, brightness contrast, color contrast, overall brightness, overall color, shadow detail, and highlight detail. Utilizing the newly curated M-SPCD dataset, we enhance the capabilities of cutting-edge Multimodal Large Language Models (MLLMs) to not only accurately assess numerical CD scores but also to provide in-depth reasoning that explains the CDs between two images. Extensive experiments demonstrate that the proposed CD-Reasoning not only achieves superior accuracy compared to state-of-the-art CD metrics but also significantly exceeds leading MLLMs in CD interpreting. Source codes will be available at https://github.com/LongYu-LY/CD-Reasoning. Zhihua Wang 0002, Qiuping Jiang, Chao Huang 0008, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2026 | PyUIE: A Coarse-to-Fine Deep Pyramid Network for Underwater Image EnhancementabstractUnderwater images often suffer from color distortion, reduced contrast, and blurriness due to light refraction, absorption, and scattering. In this paper, we propose a coarse-to-fine deepPyramid network forUnderwaterImageEnhancement (PyUIE). Specifically, PyUIE begins by decomposing the input image into high- and low-frequency components using a Laplacian pyramid. The low-frequency residual, which primarily contains lighting and color information, is processed with a lightweight deterministic color mapping network to correct global illumination and color distortions. Concurrently, the high-frequency components containing the fine details are enhanced in a coarse-to-fine manner, such that each higher scale is guided by the reconstruction from the adjacent lower scale. This hierarchical strategy effectively mitigates the risk of over-enhancement by avoiding excessive modifications to the high-frequency components. Additionally, we implement a multi-scale supervised training strategy, enabling the model to learn and reconstruct features across multiple scales, which enhances its ability to capture diverse details and improves its generalization and robustness. Extensive experiments demonstrate that our method successfully restores fine details and small structures in underwater images while producing vivid and visually appealing colors, thereby outperforming existing enhancement methods in both qualitative and quantitative evaluations. The code is available athttps://github.com/ttttllt/PyUIE.git. Wenchao Jiang, Yingqing Tan, Zhenxuan Qiu, Zhihua Wang 0002, Yang Yu 0014, Qiuping Jiang |
IEEE Trans. Multim. | 4 |
| 2025 | Multi-view Evidential Learning-based Medical Image SegmentationabstractMedical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limited embedded knowledge scope. Vision foundation models have been demonstrated to be effective in extracting generalizable knowledge, but they cannot extract domain-specific knowledge without fine-tuning. In this work, we propose a novel multi-view evidential learning-based framework, which can extract both domain-specific and generalizable knowledge from multi-view features by combining the advantages of traditional and vision foundation models. Specifically, a novel multi-view state space model (MV-SSM) is designed to extract task-related knowledge while removing redundant information within multi-view features. The proposed MV-SSM utilizes Mamba, a state space model, to model cross-view contextual dependencies between domain-specific and generalizable features. Additionally, evidential learning is adopted to quantify the segmentation uncertainty of the model for boundary. In special, variational Dirichlet is introduced to characterize the distribution of the result probabilities, parameterized with collected evidence to quantify uncertainty. As a result, the model can reduce the segmentation uncertainties of boundaries by optimizing the parameters of the Dirichlet distribution. Experimental results on three datasets show that our method obtains superior segmentation performance. Chao Huang 0008, Yushu Shi, Wai Keung Wong, Chengliang Liu 0003, Wei Wang 0169, Zhihua Wang 0002, Jie Wen 0001 |
AAAI | 6 |
| 2025 | Enhancing Low-Light Images: A Synthetic Data Perspective on Practical and Generalizable SolutionsabstractRecently, deep neural networks (DNNs) have emerged as the leading approach for low-light image enhancement (LLIE). However, training these models generally requires large-scale paired datasets, which are challenging to obtain due to the labor-intensive and time-consuming nature of real-world data collection. To alleviate this issue, synthetic data are often combined with real-captured data for training. However, most existing low-light image synthesis methods are simply performed in the sRGB domain using Gamma correction or manual adjustments via Lightroom, which fail to incorporate the physical imaging prior through the image signal processing (ISP) pipeline and thus result in limited dataset size and degradation space. Consequently, LLIE methods trained on such data often exhibit some drawbacks in the results, such as inaccurate white balance and abnormal enhancement artifacts, which limit their practicality and generalizability. In this paper, we propose a practical low-light image synthesis pipeline capable of generating unlimited paired training data. Our pipeline starts with a reverse ISP model that converts sRGB images back to the unprocessed RAW domain, where we then simulate low-light degradation, noise degradation, and white balance adjustments. Finally, the degraded RAW images are processed through a forward ISP model to produce low-light sRGB images. The pipeline further employs multiple tone mapping curves and color correction matrices (CCMs) to expand the degradation space. Hence, trained with our proposed synthetic data, existing state-of-the-art (SOTA) LLIE deep models are expected to improve their performance. Extensive experiments across various datasets demonstrate that our synthetic data can indeed effectively enhance existing LLIE deep models, improving both their practicality and generalizability. Qinghua Lin, Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
AAAI | 3 |
| 2025 | Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionabstractKehua Feng, Keyan Ding, Tan Hongzhi, Kede Ma, Zhihua Wang, Shuangquan Guo, Cheng Yuzhou, Ge Sun, Guozhou Zheng, Qiang Zhang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma, Zhihua Wang 0002, Shuangquan Guo, Yuzhou Cheng, Guozhou Zheng, Qiang Zhang 0026, Huajun Chen |
ACL (1) | 5 |
| 2025 | ICAA-Mamba: Vision Mamba for Image Color Aesthetics AssessmentabstractImage Color Aesthetics Assessment (ICAA) focuses on evaluating the aesthetic quality of color composition within images. This task involves analyzing and quantifying the visual appeal of color arrangements, taking into account factors such as harmony, contrast, and balance, to provide an objective assessment of color aesthetics. In this paper, we investigate the application of the State Space Model (Mamba) to the ICAA task, with a focus on exploring the perceptual capabilities of vision Mamba. To this end, we propose a novel framework that employs Mamba as the backbone to extract informative patterns from images, leveraging its global receptive field and linear complexity with respect to input length. Instead of relying solely on a final-layer feature, followed by fully connected layers for quality prediction, we exploit a comprehensive set of multi-scale features, thereby constructing a richer global representation of color information. To fuse these multi-scale features effectively, we introduce a Weighted Feature Fusion Module (WFFM), which adaptively assigns weights to emphasize salient color-related information. We conduct extensive experiments on two benchmark datasets, ICAA17K and SPAQ, successfully demonstrating its effectiveness for the ICAA task. The code is available at https://github.com/qinghaoya/icaa-mamba. Qinghao Xie, Wenchao Jiang, Zhihua Wang 0002, Qiuping Jiang |
ICASSP | 3 |
| 2025 | Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple AnnotatorsabstractExisting deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this challenge, we propose a Deep opinion-Unaware BIQA model by learning and adapting from Multiple Annotators, termed DUBMA, thereby eliminating the need for human annotations. Specifically, we first generate a large-scale set of distorted image pairs and then assign relative quality rankings using existing full-reference IQA models. The resulting dataset is subsequently employed for training our DUBMA. Due to the inherent discrepancies between synthetic and real-world distortions, a domain shift may occur. To address this, we propose an outlier-robust unsupervised domain adaptation approach leveraging optimal transport. This strategy effectively reduces the gap between synthetic and real-world distortion domains, thereby boosting the model’s adaptability and overall performance. Extensive experiments show that DUBMA outperforms existing opinion-unaware BIQA methods in terms of prediction accuracy across multiple datasets. Zhihua Wang 0002, Xuelin Liu, Jiebin Yan, Jie Wen 0001, Wei Wang 0169, Chao Huang 0008 |
IJCAI | 1 |
| 2025 | Evaluating Perceptual Color Preferences in Smartphone Photography: Dataset and ChallengesabstractInternational audience Zhihua Wang 0002, Weixia Zhang, Wei Zhou 0021, Xiaohong Liu 0001, Guangtao Zhai, Patrick Le Callet |
ACM Multimedia | 1 |
| 2025 | Towards Scalable and Efficient Full-Reference Omnidirectional Image Quality AssessmentabstractFull-Reference (FR) image quality assessment (IQA) (FR-IQA) has achieved notable success due to its irreplaceable role in algorithm and system optimization; however, it has less been investigated in omnidirectional image quality assessment (OIQA). In this paper, we make an attempt to FR-OIQA considering the constraint of the computation budget, in which this issue is formulated as “quality perception from patch to sequence”,i.e.,Intra-PatchSequence degradation modeling andInter-PatchSequence similarity calculation (denoted by IPS$^{2}$). Specifically, IPS$^{2}$directly accepts local patches from the omnidirectional image (OI) in the format of Equirectangular Projection as input, avoiding other preprocessing operations, such as scan-path prediction and projection transformation. Subsequently, IPS$^{2}$uses a deep feature extractor to capture patch quality and then sends the patch- wise quality maps to the cross-patch similarity (CPS) module, which explicitly models intra-patch sequence degradation and inter-patch sequence similarity via self-attention. Finally, a quality regressor is used to aggregate these features of the CPS module and predict the global quality of the OI. The experimental results on a large-scale OIQA database show that the proposed IPS$^{2}$outperforms most state-of-the-art methods in quality prediction accuracy while offering substantial reductions in computational cost and model size. Jiebin Yan, Zhihua Wang 0002, Yuming Fang 0001, Hantao Liu |
IEEE Signal Process. Lett. | 3 |
| 2025 | Toward Dimension-Enriched Underwater Image Quality AssessmentabstractThe absorption and scattering of light in the water medium naturally impair the quality of underwater images, leading to multiple degradation effects including color casts, reduced visibility, and blurriness. Underwater Image Enhancement (UIE) techniques strive to mitigate these issues, yet the efficacy of different UIE algorithms remains highly variable. This variability underscores the necessity for an objective quality metric capable of precisely assessing the visual quality of underwater images. Traditional quality metrics, which primarily rely on a single score to depict the overall quality level, are insufficiently comprehensive to describe the complex degradation characteristics intrinsic to underwater environments and the multi-dimensional nature of underwater image quality. To address this issue, we construct the first UIE quality evaluation dataset with multi-dimensional quality annotations, broadening the subjective labels from a single overall quality score to multiple specific degradation-related scores. The dataset is known as an enhanced version of our previous Subjectively Annotated UIE Benchmark Dataset (SAUD) and is called SAUD2.0 hereinafter. Based on the SAUD2.0 dataset, we also introduce a Multi-stream COllaborative LEarning network (MCOLE) tailored for quality evaluation of enhanced underwater images. MCOLE capitalizes on the multi-dimensional quality annotations within SAUD2.0, facilitating the training of three specialized networks focused on extracting distinct sets of features: color, visibility, and semantic. These extracted features are then interacted and cohesively merged for quality prediction. Comprehensive experiments conducted on two benchmark datasets reveal that the proposed MCOLE outperforms current underwater image quality metrics. These results clearly validate the efficacy of exploring the multi-dimensional nature of underwater image quality and integrating such multi-dimensional quality annotations into underwater image quality evaluation. Our dataset and code are available athttps://github.com/0117Tzx/MCOLE. Qiuping Jiang, Xiao Yi, Li Ouyang, Jingchun Zhou, Zhihua Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Deep Underwater Image Quality Assessment With Explicit Degradation Awareness EmbeddingabstractUnderwater Image Quality Assessment (UIQA) is currently an area of intensive research interest. Existing deep learning-based UIQA models always learn a deep neural network to directly map the input degraded underwater image into a final quality score via end-to-end training. However, a wide variety of image contents or distortion types may correspond to the same quality score, making it challenging to train such a deep model merely with a single subjective quality score as supervision. An intuitive idea to solve this problem is to exploit more detailed degradation-aware information as supplementary guidance to facilitate model learning. In this paper, we devise a novel deep UIQA model with Explicit Degradation Awareness embedding, i.e., EDANet. To train the EDANet, a two-stage training strategy is adopted. First, a tailored Degradation Information Discovery subnetwork (DIDNet) is pre-trained to infer a residual map between the input degraded underwater image and its pseudoreference counterpart. The inferred residual map explicitly characterizes the local degradation of the input underwater image. The intermediate feature representations on the decoder side of DIDNet are then embedded into the Degradation-guided Quality Evaluation subnetwork (DQENet), which significantly enhances the feature characterization capability with higher degradation awareness for quality prediction. The superiority of our EDANet against 18 state-of-the-art methods has been well demonstrated by extensive comparisons on two benchmark datasets. The source code of our EDANet is available at https://github.com/yia-yuese/EDANet. Qiuping Jiang, Yuese Gu, Zongwei Wu, Chongyi Li, Huan Xiong, Feng Shao 0001, Zhihua Wang 0002 |
IEEE Trans. Image Process. | 7 |
| 2025 | A Lesion-Fusion Neural Network for Multi-View Diabetic Retinopathy GradingabstractAs the most common complication of diabetes, diabetic retinopathy (DR) is one of the main causes of irreversible blindness. Automatic DR grading plays a crucial role in early diagnosis and intervention, reducing the risk of vision loss in people with diabetes. In these years, various deep-learning approaches for DR grading have been proposed. Most previous DR grading models are trained using the dataset of single-field fundus images, but the entire retina cannot be fully visualized in a single field of view. There are also problems of scattered location and great differences in the appearance of lesions in fundus images. To address the limitations caused by incomplete fundus features, and the difficulty in obtaining lesion information. This work introduces a novel multi-view DR grading framework, which solves the problem of incomplete fundus features by jointly learning fundus images from multiple fields of view. Furthermore, the proposed model combines multi-view inputs such as fundus images and lesion snapshots. It utilizes heterogeneous convolution blocks (HCB) and scalable self-attention classes (SSAC), which enhance the ability of the model to obtain lesion information. The experimental results show that our proposed method performs better than the benchmark methods on the large-scale dataset. Xiaoling Luo 0001, Qihao Xu, Zhihua Wang 0002, Chao Huang 0008, Chengliang Liu 0003, Xiaopeng Jin, Jianguo Zhang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | M2Trans: Multi-Modal Regularized Coarse-to-Fine Transformer for Ultrasound Image Super-ResolutionabstractUltrasound image super-resolution (SR) aims to transform low-resolution images into high-resolution ones, thereby restoring intricate details crucial for improved diagnostic accuracy. However, prevailing methods relying solely on image modality guidance and pixel-wise loss functions struggle to capture the distinct characteristics of medical images, such as unique texture patterns and specific colors harboring critical diagnostic information. To overcome these challenges, this paper introduces the Multi-Modal Regularized Coarse-to-fine Transformer (M2Trans) for Ultrasound Image SR. By integrating the text modality, we establish joint image-text guidance during training, leveraging the medical CLIP model to incorporate richer priors from text descriptions into the SR optimization process, enhancing detail, structure, and semantic recovery. Furthermore, we propose a novel coarse-to-fine transformer comprising multiple branches infused with self-attention and frequency transforms to efficiently capture signal dependencies across different scales. Extensive experimental results demonstrate significant improvements over state-of-the-art methods on benchmark datasets, including CCA-US, US-CASE, and our newly created dataset MMUS1K, with a minimum improvement of 0.17dB, 0.30dB, and 0.28dB in terms of PSNR. Zhangkai Ni, Runyu Xiao, Wenhan Yang, Hanli Wang, Zhihua Wang 0002, Lihua Xiang |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Dataset and Metric for Quality Assessment of HDR Tone Mapping: Detail Visibility, Color Naturalness, and Overall QualityabstractTone-Mapping Operators (TMOs) aim at converting high dynamic range (HDR) images into standard dynamic range (SDR) ones that are suitable for being displayed on standard screens. As the visual quality of tone-mapped image (TMI) is paramount, conducting quality assessment of TMIs becomes crucial. Despite the growing body of research on TMI quality assessment, the existing metrics are often limited to a narrow selection of hand-picked examples generated by a restricted range of TMOs. Consequently, their ability of generalizing to the wide array of TMIs encountered in practical scenarios remains unclear. Moreover, the quality degradation in practical TMIs can be intricate, diverse, and complex. To overcome these limitations, we construct so far the largest subjective-annotated TMI quality assessment dataset which comprises a total number of 14,000 TMIs generated by applying 20 representative TMOs to 700 HDR images. The dataset is accompanied by subjective scores that encompass multiple quality dimensions, i.e., TMI quality dataset in terms of Detail visibility, Color naturalness, and overall Quality (TDCQ). In addition, we also design a multi-branch deep neural network tailored to characterize the multi-dimensional quality perception of TMIs, i.e., Color naturalness-, Detail visibility-aware TMI Quality (CDTIQ) metric, allowing for a comprehensive and multifaceted quality assessment of TMIs. Through extensive experiments, we demonstrate the superiority of our proposed metric, showcasing a higher correlation with subjective rating results compared to other relevant no-reference image quality metrics. Qiuping Jiang, Xiwen Li, Zhihua Wang 0002, Guangtao Zhai |
IEEE Trans. Multim. | 4 |
| 2024 | Multiscale Sliced Wasserstein Distances as Perceptual Color Difference Measures
Zhihua Wang 0002, Leon Wang, Tsein-I Liu, Yuming Fang 0001, Qilin Sun 0001, Kede Ma |
ECCV (53) | 2 |
| 2024 | Thqa: A Perceptual Quality Assessment Database for Talking HeadsabstractIn the realm of media technology, digital humans have gained prominence due to rapid advancements in computer technology. However, the manual modeling and control required for the majority of digital humans pose significant obstacles to efficient development. The speech-driven methods offer a novel avenue for manipulating the mouth shape and expressions of digital humans. Despite the proliferation of driving methods, the quality of many generated talking head (TH) videos remains a concern, impacting user visual experiences. To tackle this issue, this paper introduces the Talking Head Quality Assessment (THQA) database, featuring 800 TH videos generated through 8 diverse speechdriven methods. Extensive experiments affirm the THQA database’s richness in character and speech features. Subsequent subjective quality assessment experiments analyze correlations between scoring results and speech-driven methods, ages, and genders. In addition, experimental results show that mainstream image and video quality assessment methods have limitations for the THQA database, underscoring the imperative for further research to enhance TH video quality assessment. The THQA database is publicly accessible at https://github.com/zyj-2000/THQA. Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Zhihua Wang 0002, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 6 |
| 2024 | Multimodal Representation Distribution Learning for Medical Image Segmentation
Chao Huang 0008, Weichao Cai, Qiuping Jiang, Zhihua Wang 0002 |
IJCAI | 4 |
| 2024 | CD-iNet: Deep Invertible Network for Perceptual Image Color Difference Measurement
Zhihua Wang 0002, Keshuo Xu, Keyan Ding, Qiuping Jiang, Yifan Zuo 0001, Zhangkai Ni, Yuming Fang 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Rethinking and Conceptualizing Just Noticeable Difference Estimation by Residual LearningabstractThe human visual system (HVS) cannot perceive the pixel intensity change below a certain threshold which is also known as the just noticeable difference (JND). Conventional JND prediction models mainly follow a two-step pipeline by first modeling the diverse masking effects based on the findings of the HVS and then fusing the results of different masking effect models into an overall JND map. However, due to the insufficient understanding of the HVS properties at the current stage, it is difficult to devise accurate computational models to characterize the complex masking effects. Moreover, the reasonability of the manually designed fusion schemes also lacks justification. In this work, we rethink the JND estimation problem from a fresh perspective by conceptualizing the JND as the difference map between the pristine image and its corresponding Critical Perceptual Lossless (CPL) counterpart. Building on this insight, we introduce a deep residual learning framework called ResJND to learn the discrepancies between the pristine image and its CPL counterpart, aiming to predict JND map implicitly. To support the training of our proposed ResJND model, we construct a dedicated CPL image dataset called CPL-Set which comprises a collection of pristine images and their corresponding CPL images selected by thorough subjective experiments. Comprehensive experiments have conclusively shown that our ResJND model excels at accurately predicting the JND map. Additionally, it demonstrates superior performance in associated applications, such as JND-guided noise injection, JND-guided image compression, and distortion visibility prediction. Codes are available at: https://github.com/Knife646/ResJND. Qiuping Jiang, Zhihua Wang 0002, Shiqi Wang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Adaptive Structure and Texture Similarity Metric for Image Quality Assessment and OptimizationabstractObjective Image Quality Assessment (IQA) aims to design computational models that can automatically predict the perceived quality of images. The state-of-the-art full-reference IQA metric – Deep Image Structure and Texture Similarity (DISTS), neglects the fact that natural images often consist of local structure and texture, and requires supervised training on the annotated dataset. In this article, we introduce multiple adaptive strategies to improve DISTS, resulting in an opinion-unaware IQA metric, named A-DISTS. Specifically, A-DISTS first uses the dispersion index as a statistical feature to adaptively localize structure and texture regions at different scales. Second, it adaptively assigns the spatial weights between local structure and texture similarity measurements according to the estimated structure or texture probability maps. Finally, it calculates the entropy of image representation to adaptively weigh the importance of each feature map. As a result, A-DISTS is adapted to local image content and does not require any training. The experimental results demonstrated that the proposed metric correlates well with human rating in the standard and algorithm-dependent IQA databases, and exhibits competitive performance in the optimization tasks of single image super-resolution, motion deblurring, and multi-distortion removal. Keyan Ding, Rijin Zhong, Zhihua Wang 0002, Yang Yu 0014, Yuming Fang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Perception-Driven Deep Underwater Image Enhancement Without Paired SupervisionabstractUnderwater image enhancement (UIE) aims to improve the visual quality of raw underwater images. Current UIE algorithms primarily train a deep neural network (DNN) on synthetic datasets or datasets with pseudo labels by minimizing the reconstruction loss between enhanced images and ground truth images. However, there is a domain gap between synthetic and real-world underwater images, and the widely used$\ell _{1}$or$\ell _{2}$loss tends to overlook the importance of human perception, resulting in unsatisfactory perceptual quality of the final enhanced results. In this paper, we propose an unsupervised perception-driven DNN called PDD-Net for generalizable UIE. Instead of relying on paired images for training, we resort to an unsupervised generative adversarial network (GAN) with a large-scale set of easily available natural images as the target domain. This enables training on larger image sets collected from various domains while avoiding over-fitted to any specific data generation protocol. Additionally, to make the visual quality of enhanced underwater images more in line with human perception, we pre-train a DNN-based pairwise quality ranking (PQR) model based on which a PQR loss is formulated to progressively guides the enhancement of raw underwater image toward the higher quality direction. In addition, we introduce a global attention module (GAM) that integrates modulation and attention mechanisms to enable capturing rich global and local information, leading to improvements in both brightness and contrast. Extensive experiments demonstrate that our proposed PDD-Net exhibits excellent generalization capabilities and outperforms existing methods in terms of both visual perception quality and quantitative indicators across different datasets. Qiuping Jiang, Yaozu Kang, Zhihua Wang 0002, Wenqi Ren, Chongyi Li |
IEEE Trans. Multim. | 3 |
| 2023 | Learning a Deep Color Difference Metric for Photographic ImagesabstractMost well-established and widely used color difference (CD) metrics are handcrafted and subjectcalibrated against uniformly colored patches, which do not generalize well to photographic images characterized by natural scene complexities. Constructing CD formulae for photo-graphic images is still an active research topic in imaging/illumination, vision science, and color science communities. In this paper, we aim to learn a deep CD metric for photographic images with four desirable properties. First, it well aligns with the observations in vision science that color and form are linked inextricably in visual cortical processing. Second, it is a proper metric in the mathematical sense. Third, it computes accurate CDs between photographic images, differing mainly in color appearances. Fourth, it is robust to mild geometric distortions (e.g., translation or due to parallax), which are often present in photographic images of the same scene captured by different digital cameras. We show that all four properties can be satisfied at once by learning a multi-scale autoregressive normalizing flow for feature transform, followed by the Euclidean distance which is linearly proportional to the human perceptual CD. Quantitative and qualitative experiments on the large-scale SPCD dataset demonstrate the promise of the learned CD metric. Source code is available at https://github.com/haoychen3/CD-Flow. Zhihua Wang 0002, Yang Yang 0003, Qilin Sun 0001, Kede Ma |
CVPR | 2 |
| 2023 | Measuring Perceptual Color Differences of Smartphone PhotographsabstractMeasuring perceptual color differences (CDs) is of great importance in modern smartphone photography. Despite the long history, most CD measures have been constrained by psychophysical data of homogeneous color patches or a limited number of simplistic natural photographic images. It is thus questionable whether existing CD measures generalize in the age of smartphone photography characterized by greater content complexities and learning-based image signal processors. In this article, we put together so far the largest image dataset for perceptual CD assessment, in which the photographic images are 1) captured by six flagship smartphones, 2) altered by Photoshop, 3) post-processed by built-in filters of the smartphones, and 4) reproduced with incorrect color profiles. We then conduct a large-scale psychophysical experiment to gather perceptual CDs of 30,000 image pairs in a carefully controlled laboratory environment. Based on the newly established dataset, we make one of the first attempts to construct an end-to-end learnable CD formula based on a lightweight neural network, as a generalization of several previous metrics. Extensive experiments demonstrate that the optimized formula outperforms 33 existing CD measures by a large margin, offers reasonable local CD maps without the use of dense supervision, generalizes well to homogeneous color patch data, and empirically behaves as a proper metric in the mathematical sense. Our dataset and code are publicly available at https://github.com/hellooks/CDNet. Zhihua Wang 0002, Keshuo Xu, Yang Yang 0201, Jianlei Dong, Shuhang Gu, Lihao Xu, Yuming Fang 0001, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Toward a blind image quality evaluator in the wild by learning beyond human opinion scoresabstractNowadays, most existing blind image quality assessment (BIQA) models i n t h e w i l d heavily rely on human ratings, which are extraordinarily labor-expensive to collect. Here, we propose an o p i n i o n − f r e e BIQA method that learns from multiple annotators to assess the perceptual quality of images captured in the wild. Specifically, we first synthesize distorted images based on the pristine counterparts. We then randomly assemble a set of image pairs from the synthetic images, and use a group of IQA models to assign pseudo-binary labels for each pair indicating which image has higher quality as the supervisory signal. Based on the newly established pseudo-labeled dataset, we train a deep neural network (DNN)-based BIQA model to rank the perceptual quality, optimized for consistency with the binary rank labels. Since there exists domain shift, e.g., distortion shift and content shift, between the synthetic and in-the-wild images, we leverage two ways to alleviate this issue. First, the simulated distortions should be similar to authentic distortions as much as possible. Second, an unsupervised domain adaptation (UDA) module is further applied to encourage learning domain-invariant features between two domains. Extensive experiments demonstrate the effectiveness of our proposed o p i n i o n − f r e e BIQA model, yielding SOTA performance in terms of correlation with human opinion scores, as well as gMAD competition. Our code is available at: https://github.com/wangzhihua520/OF_BIQA . Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
Pattern Recognit. | 1 |
| 2023 | Corrigendum to 'Toward a Blind Image Quality Evaluator in the Wild by Learning beyond Human Opinion Scores' Pattern Recognition. Volume 137 (2023) 109296
Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
Pattern Recognit. | 1 |
| 2023 | Deep Blind Image Quality Assessment Powered by Online Hard Example MiningabstractRecently, blind image quality assessment (BIQA) models based on deep neural networks (DNNs) have achieved impressive performance on existing datasets. However, due to the intrinsic imbalance property of the training set, not all distortions or images are handled equally well. Online hard example mining (OHEM) is a promising way to alleviate this issue. Inspired by the recent finding that network pruning disproportionately hampers the model's memorization of a tractable subset, atypical, low-quality, long-tailed samples, that are hard-to-memorize during training and easily “forgotten” during pruning, we propose an effective “plug-and-play” OHEM pipeline, especially for generalizable deep BIQA. Specifically, we train two parallel weight-sharing branches simultaneously, where one is full model and other is a “self-competitor” generated from the full model online by network pruning. Then, we leverage the prediction disagreement between the full model and its pruned variant (i.e., the self-competitor) to expose easily “forgettable” samples, which are therefore regarded as the hard ones. We then enforce the prediction consistency between the full model and its pruned variant to implicitly put more focus on these hard samples, which benefits the full model to recover forgettable information introduced by pruning. Extensive experiments across multiple datasets and BIQA models demonstrate that the proposed OHEM can improve the model performance and generalizability as measured by correlation numbers and group maximum differentiation (gMAD) competition. Our code are available at:https://github.com/wangzhihua520/IQA_with_OHEM Zhihua Wang 0002, Qiuping Jiang, Shanshan Zhao 0001, Wensen Feng, Weisi Lin |
IEEE Trans. Multim. | 1 |
| 2022 | A Database of Visual Color Differences of Modern Smartphone PhotographyabstractMeasures for visual color differences (CDs) are pivotal in hardware and software upgrading of modern smartphone photography. Towards this goal, we construct currently the largest database for visual CDs of smartphone photography. Our database consists of 15, 335 natural images 1) captured by six latest flagship smartphones, 2) altered by Photoshop®, 3) post-processed by built-in filters of smartphones, and 4) reproduced with incorrect color profiles. Moreover, we conduct a large-scale psychophysical experiment to gather visual CDs of 30, 000 image pairs from 20 human subjects in a well-designed laboratory environment. Last, we apply our human-rated database to compare a total of 27 classical and recent CD metrics. We show that existing metrics are limited in assessing CDs of smartphone photography, and point out promising future directions of learning-based CD metrics. Keshuo Xu, Zhihua Wang 0002, Yang Yang 0201, Jianlei Dong, Lihao Xu, Yuming Fang 0001, Kede Ma |
ICIP | 2 |
| 2022 | Active Fine-Tuning From gMAD Examples Improves Blind Image Quality AssessmentabstractThe research in image quality assessment (IQA) has a long history, and significant progress has been made by leveraging recent advances in deep neural networks (DNNs). Despite high correlation numbers on existing IQA datasets, DNN-based models may be easily falsified in the group maximum differentiation (gMAD) competition. Here we show that gMAD examples can be used to improve blind IQA (BIQA) methods. Specifically, we first pre-train a DNN-based BIQA model using multiple noisy annotators, and fine-tune it on multiple synthetically distorted images, resulting in a "top-performing" baseline model. We then seek pairs of images by comparing the baseline model with a set of full-reference IQA methods in gMAD. The spotted gMAD examples are most likely to reveal the weaknesses of the baseline, and suggest potential ways for refinement. We query human quality annotations for the selected images in a well-controlled laboratory environment, and further fine-tune the baseline on the combination of human-rated images from gMAD and existing databases. This process may be iterated, enabling active fine-tuning from gMAD examples for BIQA. We demonstrate the feasibility of our active learning scheme on a large-scale unlabeled image set, and show that the fine-tuned quality model achieves improved generalizability in gMAD, without destroying performance on previously seen databases. Zhihua Wang 0002, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Troubleshooting Blind Image Quality Models in the WildabstractRecently, the group maximum differentiation competition (gMAD) has been used to improve blind image quality assessment (BIQA) models, with the help of full-reference metrics. When applying this type of approach to troubleshoot "best-performing" BIQA models in the wild, we are faced with a practical challenge: it is highly nontrivial to obtain stronger competing models for efficient failure-spotting. Inspired by recent findings that difficult samples of deep models may be exposed through network pruning, we construct a set of "self-competitors," as random ensembles of pruned versions of the target model to be improved. Diverse failures can then be efficiently identified via self-gMAD competition. Next, we fine-tune both the target and its pruned variants on the human-rated gMAD set. This allows all models to learn from their respective failures, preparing themselves for the next round of self-gMAD competition. Experimental results demonstrate that our method efficiently troubleshoots BIQA models in the wild with improved generalizability. Zhihua Wang 0002, Haotao Wang, Tianlong Chen 0001, Zhangyang Wang, Kede Ma |
CVPR | 1 |
| 2021 | Non-spike timing-dependent plasticity learning mechanism for memristive neural networks
Zhihua Wang 0002, Ruihan Hu, Qi Wu 0003 |
Appl. Intell. | 3 |
| 2019 | Towards a Hybrid BCI Gaming Paradigm Based on Motor Imagery and SSVEPabstractBrain-computer interfaces (BCIs) not only can allow individuals to voluntarily control external devices, helping to restore lost motor functions of the disabled, but can also be used by healthy users for entertainment and gaming applications. In this study, we proposed a hybrid BCI paradigm to explore a feasible and natural way to play games by using electroencephalogram (EEG) signals in a practical environment. In this paradigm, we combined motor imagery (MI) and steady-state visually evoked potentials (SSVEPs) to generate multiple commands. A classic game, Tetris, was chosen as the control object. The novelty of this study includes the effective usage of a “dwell time” approach and fusion rules to design BCI games. To demonstrate the feasibility of the proposed hybrid paradigm, ten subjects were chosen to participate in online control experiments. The experimental results showed that all subjects successfully completed the predefined tasks with high accuracy. This proposed hybrid BCI paradigm could potentially provide those who suffer disability or paralysis with additional entertainment options, such as brain-actuated games, that could improve their happiness and quality of life.Abbreviations: BCI: brain-computer interface; EEG: electroencephalogram; MI: motor imagery; SSVEP: steady-state visually evoked potential; ERP: event-related potential; SMR: sensorimotor rhythm; VEP: visual evoked potential; TCP/IP: transmission control protocol/internet protocol; GUI: graphical user interface; ERD/ERS: event-related desynchronization/synchronization; CIC: control intention classifier; LRC: left/right classifier; CSP: common spatial pattern; LDA: linear discriminant analysis; ROC: receiver operating characteristic; TPR: true positive rate; FPR: false positive rate; CCA: canonical correlation analysis. Zhihua Wang 0002, Yang Yu 0014, Ming Xu 0022, Yadong Liu 0001, Erwei Yin, Zongtan Zhou |
Int. J. Hum. Comput. Interact. | 1 |