VLDB 2026 Research / reviewers in the wild / expert
Dingquan Li
dblp:207/2000
· DBLP profile ↗
23ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-5549-9027ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack Against No-Reference Image Quality Assessment ModelsabstractNo-Reference Image Quality Assessment (NR-IQA) models play an important role in various real-world applications. Recently, adversarial attacks against NR-IQA models have attracted increasing attention, as they provide valuable insights for revealing model vulnerabilities and guiding robust system design. Some effective attacks have been proposed against NR-IQA models in white-box settings, where the attacker has full access to the target model. However, these attacks often suffer from poor transferability to unknown target models in more realistic black-box scenarios, where the target model is inaccessible. This work makes the first attempt to address the challenge of low transferability in attacking NR-IQA models by proposing a transferable Signed Ensemble Gaussian black-box Attack (SEGA). The main idea is to approximate the gradient of the target model by applying Gaussian smoothing to source models and ensembling their smoothed gradients. To ensure the imperceptibility of adversarial perturbations, SEGA further removes inappropriate perturbations using a specially designed perturbation filter mask. Experimental results demonstrate the superior transferability of SEGA, validating its effectiveness in enabling successful transfer-based black-box attacks against NR-IQA models. Yujia Liu 0005, Dingquan Li, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Hierarchical Attention Networks for Lossless Point Cloud Attribute CompressionabstractIn this paper, we propose a deep hierarchical attention context model for lossless attribute compression of point clouds, leveraging a multi-resolution spatial structure and residual learning. A simple and effective Level of Detail (LoD) structure is introduced to yield a coarse-to-fine representation. To enhance efficiency, points within the same refinement level are encoded in parallel, sharing a common context point group. By hierarchically aggregating information from neighboring points, our attention model learns contextual dependencies across varying scales and densities, enabling comprehensive feature extraction. We also adopt normalization for position coordinates and attributes to achieve scale-invariant compression. Additionally, we segment the point cloud into multiple slices to facilitate parallel processing, further optimizing time complexity. Experimental results demonstrate that the proposed method offers better coding performance than the latest G-PCC for color and reflectance attributes while maintaining more efficient encoding and decoding run-times. Yueru Chen, Wei Zhang 0072, Dingquan Li, Jing Wang 0115, Ge Li 0002 |
DCC | 3 |
| 2025 | Adaptive Semantic Compression: Compatible Bitstream for Scalable Human-Machine Perception Sample AdaptionabstractWith the development of visual analysis models, collaborative image compression for machine and human perception has brought new challenges to the optimization of algorithms. Existing optimization algorithms achieve this target through meticulously designed model structures and bitstream design. However, the difference in bitstream design makes it incompatible with trained and existing decoders, hindering its practicality. In this paper, we proposed the Adaptive Semantic Compression (ASC) framework to fine-tune pre-trained codec on individual samples to obtain scalable bitstreams in an intuitive yet effective way. First, to improve the efficiency of application in machine perception, we proposed the Latent Semantic Contraction (LSC) method to fine-tune the latent code while preserving the machine task performance of the decoded image. Second, to further optimize human perception, we proposed the Spatial-frequency Decoder Adaptation (SFDA) module. By compensating for distortion in the spatial and frequency domains, SFDA improves the humane perception quality of the reconstructed image. The bitstreams composed of LSC and SFDA can be decoded by existing decoders to reconstruct images, thus fully exploiting the performance of the existing model. We implemented our algorithm on different pre-trained compression models and verified the flexibility and compatibility on various test images. Experimental results show that the LSC module can save 24.97% to 29.10% of bitrates with machine perception performance. Furthermore, the application of SFDA brings a 3.16% gain in the BD-Rate with PSNR, up to 15.69%, compared to LSC. Dingquan Li, Guoqing Xiang, Jinchang Xu, Shanghang Zhang |
ICME | 2 |
| 2025 | A Norm Regularization Training Strategy for Robust Image Quality Assessment Models
Yujia Liu 0005, Chenxi Yang 0004, Dingquan Li, Tingting Jiang 0001, Tiejun Huang 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Low-Overhead Compression-Aware Channel Filtering for Hyperspectral Image CompressionabstractBoth traditional and learning-based hyperspectral image compression methods suffer from significant quality loss at high compression ratios. To address this, we propose a low-overhead, compression-aware channel filtering method. The encoder derives channel filters via Least Squares Regression between lossy compressed and original images. The bitstream, containing the compressed image and filters, is sent to the decoder, where the filters enhance image quality. This simple, compression-aware approach is compatible with any existing framework, enhancing quality while introducing only a negligible increase in bitstream size and decoding time, thereby achieving low overhead. Experimental results show consistent rate-distortion gains, reducing compression rates by 10.51% to 39.81% on the GF-5 dataset with minimal decoding and storage overhead. Wei Zhang 0072, Jiayao Xu, Yueru Chen, Dingquan Li, Wen Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Lightweight Spectral Super-Resolution Network for Hyperspectral Image CompressionabstractThe growing use of hyperspectral images demands efficient compression techniques to handle their extensive spectral data. However, current methods are constrained by their inability to adapt to high bit depth and effectively utilize the spectral characteristics, leading to suboptimal compression ratios. This paper presents a novel hyperspectral compression framework that employs a lightweight spectral super-resolution network to address these limitations. The proposed approach divides the hyperspectral image into two sub-images, comprising two distinct groups of bands: a base image consisting of anchor bands and a supplementary image comprising non-anchor bands. The base image is compressed losslessly using a conventional codec, thereby ensuring the preservation of essential information. In contrast, the supplementary image is compressed efficiently by overfitting a lightweight super-resolution network to predict the non-anchor bands during encoding. The optimized network parameters are encoded as side information to ensure high-quality spectral super-resolution during decoding. Experimental results on the ARAD hyperspectral image dataset demonstrate that our approach significantly outperforms state-of-the-art methods, effectively meeting the demand for efficient hyperspectral image compression while maintaining acceptable processing speeds. Wei Zhang 0072, Pengpeng Yu, Yueru Chen, Dingquan Li, Wen Gao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | LOD-PCAC: Level-of-Detail-Based Deep Lossless Point Cloud Attribute CompressionabstractPoint cloud attribute compression is a challenging issue in efficiently compressing large volumes of attributes. Despite notable advancements in lossy point cloud compression using deep learning, progress in lossless compression remains limited. Some methods have employed octree- or voxel-based partitioning techniques derived from geometric compression, achieving success on dense point clouds. However, these voxel-based approaches struggle with sparse or unevenly distributed point clouds, leading to performance degradation. In this work, we introduce a novel framework for learning-based lossless point cloud attribute compression, named LOD-PCAC, which leverages a Level-of-Detail (LOD) structure to ensure density-robust compression. Specifically, the input point cloud is divided into multiple detail levels, and vertices from these levels are selected to construct a Reference Set as context, which effectively captures multi-level information. Then we propose the Bit-level Residual Coder for efficient attribute compression. Instead of directly compressing attributes, our method first predicts attribute values and organizes the residual bits into a Bit Matrix as another context, simplifying predictions and fully exploiting channel correlations. Finally, a neural network with specialized encoders processes the context to estimate the probability of each residual bit. Experimental results demonstrate that the proposed method outperforms both traditional and learning-based approaches across various point clouds, exhibiting strong generalization across datasets and robustness to varying densities. Wenbo Zhao 0004, Wei Gao 0003, Dingquan Li, Jing Wang 0115 |
IEEE Trans. Image Process. | 3 |
| 2024 | Defense Against Adversarial Attacks on No-Reference Image Quality Models with Gradient Norm RegularizationabstractThe task of No-Reference Image Quality Assessment (NR-IQA) is to estimate the quality score of an input image without additional information. NR-IQA models play a crucial role in the media industry, aiding in performance evaluation and optimization guidance. However, these models are found to be vulnerable to adversarial attacks, which introduce imperceptible perturbations to input images, re-sulting in significant changes in predicted scores. In this paper, we propose a defense method to improve the stability in predicted scores when attacked by small perturbations, thus enhancing the adversarial robustness of NR-IQA models. To be specific, we present theoretical evidence showing that the magnitude of score changes is related to the g 1 norm of the model's gradient with respect to the input image. Building upon this theoretical foundation, we propose a norm regularization training strategy aimed at reducing the g 1 norm of the gradient, thereby boosting the robustness of NR-IQA models. Experiments conducted on four NR-IQA baseline models demonstrate the effectiveness of our strategy in reducing score changes in the presence of adversarial attacks. To the best of our knowledge, this work marks the first attempt to defend against adversarial attacks on NR-IQA models. Our study offers valuable insights into the adversarial robustness of NR-IQA models and provides a foundation for future research in this area. Yujia Liu 0005, Chenxi Yang 0004, Dingquan Li, Jianhao Ding, Tingting Jiang 0001 |
CVPR | 3 |
| 2024 | Lightweight super resolution network for point cloud geometry compressionabstractWe present an approach for compressing point cloud geometry by leveraging a lightweight super-resolution network. It involves decomposing a point cloud into a base point cloud and the interpolation patterns for reconstructing the original point cloud. While the base point cloud can be efficiently compressed using any lossless codec, such as Geometry-based Point Cloud Compression, a distinct strategy is employed for handling the interpolation patterns. Rather than directly compressing the interpolation patterns, a lightweight super-resolution network is utilized to learn this information through overfitting. Subsequently, the network parameter is transmitted to assist in point cloud reconstruction at the decoder side. Our approach differentiates itself from lookup table-based methods, allowing us to obtain more accurate interpolation patterns by accessing a broader range of neighboring voxels at an acceptable computational cost. Experiments on MPEG Cat1 (Solid) and Cat2 datasets demonstrate the remarkable compression performance achieved by our method. Wei Zhang 0072, Dingquan Li, Ge Li 0002, Wen Gao 0001 |
DCC | 2 |
| 2024 | Exploring Vulnerabilities of No-Reference Image Quality Assessment Models: A Query-Based Black-Box MethodabstractNo-Reference Image Quality Assessment (NR-IQA) aims to predict image quality scores consistent with human perception without relying on pristine reference images, serving as a crucial component in various visual tasks. Ensuring the robustness of NR-IQA methods is vital for reliable comparisons of different image processing techniques and consistent user experiences in recommendations. The attack methods for NR-IQA provide a powerful instrument to test the robustness of NR-IQA. However, current attack methods of NR-IQA heavily rely on the gradient of the NR-IQA model, leading to limitations when the gradient information is unavailable. In this paper, we present a pioneering query-based black box attack against NR-IQA methods. We propose the concept of score boundary and leverage an adaptive iterative approach with multiple score boundaries. Meanwhile, the initial attack directions are also designed to leverage the characteristics of the Human Visual System (HVS). Experiments show our method outperforms all compared state-of-the-art attack methods and is far ahead of previous black-box methods. The effective NR-IQA model DBCNN suffers a Spearman’s rank-order correlation coefficient (SROCC) decline of 0.6381 attacked by our method, revealing the vulnerability of NR-IQA models to black-box attacks. The proposed attack method also provides a potent tool for further exploration into NR-IQA robustness. Chenxi Yang 0004, Yujia Liu 0005, Dingquan Li, Tingting Jiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Hierarchical Prior-Based Super Resolution for Point Cloud Geometry CompressionabstractThe Geometry-based Point Cloud Compression (G-PCC) has been developed by the Moving Picture Experts Group to compress point clouds efficiently. Nevertheless, in its lossy mode, the reconstructed point cloud by G-PCC often suffers from noticeable distortions due to naïve geometry quantization (i.e., grid downsampling). This paper proposes a hierarchical prior-based super resolution method for point cloud geometry compression. The content-dependent hierarchical prior is constructed at the encoder side, which enables coarse-to-fine super resolution of the point cloud geometry at the decoder side. A more accurate prior generally yields improved reconstruction performance, albeit at the cost of increased bits required to encode this piece of side information. Our experiments on the MPEG Cat1A dataset demonstrate substantial Bjøntegaard-delta bitrate savings, surpassing the performance of the octree-based and trisoup-based G-PCC v14. We provide our implementations for reproducible research at https://github.com/lidq92/mpeg-pcc-tmc13. Dingquan Li, Kede Ma, Jing Wang 0115, Ge Li 0002 |
IEEE Trans. Image Process. | 1 |
| 2023 | Personalized Image Generation for Color Vision Deficiency PopulationabstractApproximately, 350 million people, a proportion of 8%, suffer from color vision deficiency (CVD). While image generation algorithms have been highly successful in synthesizing high-quality images, CVD populations are unintentionally excluded from target users and have difficulties understanding the generated images as normal viewers do. Although a straightforward baseline can be formed by combining generation models and recolor compensation methods as the post-processing, the CVD friendliness of the result images is still limited since the input image content of recolor methods is not CVD-oriented and will be fixed during the recolor compensation process. Besides, the CVD populations can not be fully served since the varying degrees of CVD are often neglected in recoloring methods. Instead, we propose a personalized CVD-friendly image generation algorithm with two key characteristics: (i) generating CVD-oriented images aligned with the needs of CVD populations; (ii) generating continuous personalized images for people with various CVD degrees through disentangling the color representation based on a triple-latent structure. Quantitative and qualitative experiments indicate our proposed image generation model can generate practical and compelling results compared to the normal generation model and combination baselines on several datasets. The code is available at: https://github.com/Jiangshuyi0V0/CVD-GAN.git Shuyi Jiang, Daochang Liu, Dingquan Li, Chang Xu 0002 |
ICCV | 3 |
| 2023 | Continual Learning for Blind Image Quality AssessmentabstractThe explosive growth of image data facilitates the fast development of image processing and computer vision methods for emerging visual applications, meanwhile introducing novel distortions to processed images. This poses a grand challenge to existing blind image quality assessment (BIQA) models, which are weak at adapting to subpopulation shift. Recent work suggests training BIQA methods on the combination of all available human-rated IQA datasets. However, this type of approach is not scalable to a large number of datasets and is cumbersome to incorporate a newly created dataset as well. In this paper, we formulate continual learning for BIQA, where a model learns continually from a stream of IQA datasets, building on what was learned from previously seen data. We first identify five desiderata in the continual setting with three criteria to quantify the prediction accuracy, plasticity, and stability, respectively. We then propose a simple yet effective continual learning method for BIQA. Specifically, based on a shared backbone network, we add a prediction head for a new dataset and enforce a regularizer to allow all prediction heads to evolve with new data while being resistant to catastrophic forgetting of old data. We compute the overall quality score by a weighted summation of predictions from all heads. Extensive experiments demonstrate the promise of the proposed continual learning method in comparison to standard training techniques for BIQA, with and without experience replay. We made the code publicly available at https://github.com/zwx8981/BIQA_CL. Weixia Zhang, Dingquan Li, Chao Ma 0004, Guangtao Zhai, Xiaokang Yang 0001, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Deep Geometry Post-Processing for Decompressed Point CloudsabstractPoint cloud compression plays a crucial role in reducing the huge cost of data storage and transmission. However, distortions can be introduced into the decompressed point clouds due to quantization. In this paper, we propose a novel learning-based post-processing method to enhance the decompressed point clouds. Specifically, a voxelized point cloud is first divided into small cubes. Then, a 3D convolutional network is proposed to predict the occupancy probability for each location of a cube. We leverage both local and global contexts by generating multi-scale probabilities. These probabilities are progressively summed to predict the results in a coarse-to-fine manner. Finally, we obtain the geometry-refined point clouds based on the predicted probabilities. Different from previous methods, we deal with decompressed point clouds with huge variety of distortions using a single model. Experimental results show that the proposed method can significantly improve the quality of the decompressed point clouds, achieving 9.30dB BDPSNR gain on three representative datasets on average. Ge Li 0002, Dingquan Li, Yurui Ren, Wei Gao 0003, Thomas H. Li |
ICME | 3 |
| 2022 | No-reference Image Quality Assessment via Non-local Dependency ModelingabstractIn this paper, we propose a no-reference image quality assessment method based on non-local features learned by a graph neural network (GNN). The proposed quality assessment framework is rooted in the view that the human visual system perceives image quality with long-dependency constructed among different regions, inspiring us to explore the non-local interactions in quality prediction. Instead of relying on convolutional neural network (CNN) based quality assessment methods that primarily focus on local field features, the GNN aiming for non-local quality perception facilitates modeling such long-dependency. In particular, we first adopt superpixel segmentation for the graph nodes construction. Subsequently, a spatial attention module is proposed to integrate the long- and short-range dependencies among the nodes of the whole image. The learned non-local features are finally combined with the local features extracted by the pre-trained CNN, achieving superior performance to the features utilized individually. Experimental results on intra-dataset and cross-dataset settings verify our proposed method's effectiveness and advanced generalization capability. Source codes are publicly accessible at https://github.com/SuperBruceJia/NLNet-IQA for scientific reproducible research. Shuyue Jia, Baoliang Chen, Dingquan Li, Shiqi Wang 0001 |
MMSP | 3 |
| 2022 | Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-LoopabstractNo-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimization of man-made vision systems. Here we make one of the first attempts to examine the perceptual robustness of NR-IQA models. Under a Lagrangian formulation, we identify insightful connections of the proposed perceptual attack to previous beautiful ideas in computer vision and machine learning. We test one knowledge-driven and three data-driven NR-IQA methods under four full-reference IQA models (as approximations to human perception of just-noticeable differences). Through carefully designed psychophysical experiments, we find that all four NR-IQA models are vulnerable to the proposed perceptual attack. More interestingly, we observe that the generated counterexamples are not transferable, manifesting themselves as distinct design flows of respective NR-IQA methods. Source code are available at https://github.com/zwx8981/PerceptualAttack_BIQA. Weixia Zhang, Dingquan Li, Xiongkuo Min, Guangtao Zhai, Guodong Guo, Xiaokang Yang 0001, Kede Ma |
NeurIPS | 2 |
| 2022 | Near-lossless Point Cloud Geometry Compression Based on Adaptive Residual CompensationabstractPoint cloud compression (PCC) is a crucial enabler for immersive multimedia applications since point cloud is one of the most primitive forms for representing 3D scenes and objects. Recently, some approaches are proposed to improve the average reconstruction quality of octree-based Geometry-based Point Cloud Compression (G-PCC). However, it is noticed that these approaches suffer considerable loss in terms of point-to-point (D1) Hausdorff distance when compared to G-PCC (octree). Here we introduce a near-lossless point cloud geometry compression method based on adaptive residual compensation by adding and removing points with large errors. It allows controlling of D1 Hausdorff (D1h) distance and maintains a great improvement in average reconstruction performance over G-PCC. Experimental results verify the effectiveness of our method, where our method achieves an average of 78.5% D1 and 11.4% D1h Bjontegaard-delta bitrate savings over the octree-based G-PCC on solid point clouds of the MPEG Cat1A dataset. Dingquan Li, Jing Wang 0115, Ge Li 0002 |
VCIP | 1 |
| 2021 | Reproducibility Companion Paper: Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality AssessmentabstractThis companion paper supports the experimental replication of the paper "Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality Assessment'' presented at ACM Multimedia 2020. We provide the software package for replicating the implementation of the "Norm-in-Norm'' loss and the corresponding "LinearityIQA'' model used in the original paper. This paper contains the guidelines to reproduce all the experimental results of the original paper. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001, Vajira Thambawita |
ACM Multimedia | 1 |
| 2021 | Unified Quality Assessment of in-the-Wild Videos with Mixed Datasets Training
Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
Int. J. Comput. Vis. | 1 |
| 2020 | Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality AssessmentabstractCurrently, most image quality assessment (IQA) models are supervised by the MAE or MSE loss with empirically slow convergence. It is well-known that normalization can facilitate fast convergence. Therefore, we explore normalization in the design of loss functions for IQA. Specifically, we first normalize the predicted quality scores and the corresponding subjective quality scores. Then, the loss is defined based on the norm of the differences between these normalized values. The resulting "Norm-in-Norm" loss encourages the IQA model to make linear predictions with respect to subjective quality scores. After training, the least squares regression is applied to determine the linear mapping from the predicted quality to the subjective quality. It is shown that the new loss is closely connected with two common IQA performance criteria (PLCC and RMSE). Through theoretical analysis, it is proved that the embedded normalization makes the gradients of the loss function more stable and more predictable, which is conducive to the faster convergence of the IQA model. Furthermore, to experimentally verify the effectiveness of the proposed loss, it is applied to solve a challenging problem: quality assessment of in-the-wild images. Experiments on two relevant datasets (KonIQ-10k and CLIVE) show that, compared to MAE or MSE loss, the new loss enables the IQA model to converge about 10 times faster and the final model achieves better performance. The proposed model also achieves state-of-the-art prediction performance on this challenging problem. For reproducible scientific research, our code is publicly available at \urlhttps://github.com/lidq92/LinearityIQA. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
ACM Multimedia | 1 |
| 2019 | Quality Assessment of In-the-Wild VideosabstractQuality assessment of in-the-wild videos is a challenging problem because of the absence of reference videos and shooting distortions. Knowledge of the human visual system can help establish methods for objective quality assessment of in-the-wild videos. In this work, we show two eminent effects of the human visual system, namely, content-dependency and temporal-memory effects, could be used for this purpose. We propose an objective no-reference video quality assessment method by integrating both effects into a deep neural network. For content-dependency, we extract features from a pre-trained image classification neural network for its inherent content-aware property. For temporal-memory effects, long-term dependencies, especially the temporal hysteresis, are integrated into the network with a gated recurrent unit and a subjectively-inspired temporal pooling layer. To validate the performance of our method, experiments are conducted on three publicly available in-the-wild video quality assessment databases: KoNViD-1k, CVD2014, and LIVE-Qualcomm, respectively. Experimental results demonstrate that our proposed method outperforms five state-of-the-art methods by a large margin, specifically, 12.39%, 15.71%, 15.45%, and 18.09% overall performance improvements over the second-best method VBLIINDS, in terms of SROCC, KROCC, PLCC and RMSE, respectively. Moreover, the ablation study verifies the crucial role of both the content-aware features and the modeling of temporal-memory effects. The PyTorch implementation of our method is released at https://github.com/lidq92/VSFA. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
ACM Multimedia | 1 |
| 2019 | Which Has Better Visual Quality: The Clear Blue Sky or a Blurry Animal?abstractImage content variation is a typical and challenging problem in no-reference image-quality assessment (NR-IQA). This work pays special attention to the impact of image content variation on NR-IQA methods. To better analyze this impact, we focus on blur-dominated distortions to exclude the impacts of distortion-type variations. We empirically show that current NR-IQA methods are inconsistent with human visual perception when predicting the relative quality of image pairs with different image contents. In view of deep semantic features of pretrained image classification neural networks always containing discriminative image content information, we put forward a new NR-IQA method based on semantic feature aggregation (SFA) to alleviate the impact of image content variation. Specifically, instead of resizing the image, we first crop multiple overlapping patches over the entire distorted image to avoid introducing geometric deformations. Then, according to an adaptive layer selection procedure, we extract deep semantic features by leveraging the power of a pretrained image classification model for its inherent content-aware property. After that, the local patch features are aggregated using several statistical structures. Finally, a linear regression model is trained for mapping the aggregated global features to image-quality scores. The proposed method, SFA, is compared with nine representative blur-specific NR-IQA methods, two general-purpose NR-IQA methods, and two extra full-reference IQA methods on Gaussian blur images (with and without Gaussian noise/JPEG compression) and realistic blur images from multiple databases, including LIVE, TID2008, TID2013, MLIVE1, MLIVE2, BID, and CLIVE. Experimental results show that SFA is superior to the state-of-the-art NR methods on all seven databases. It is also verified that deep semantic features play a crucial role in addressing image content variation, and this provides a new perspective for NR-IQA. Dingquan Li, Tingting Jiang 0001, Weisi Lin, Ming Jiang 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Exploiting High-Level Semantics for No-Reference Image Quality Assessment of Realistic Blur ImagesabstractTo guarantee a satisfying Quality of Experience (QoE) for consumers, it is required to measure image quality efficiently and reliably. The neglect of the high-level semantic information may result in predicting a clear blue sky as bad quality, which is inconsistent with human perception. Therefore, in this paper, we tackle this problem by exploiting the high-level semantics and propose a novel no-reference image quality assessment method for realistic blur images. Firstly, the whole image is divided into multiple overlapping patches. Secondly, each patch is represented by the high-level feature extracted from the pre-trained deep convolutional neural network model. Thirdly, three different kinds of statistical structures are adopted to aggregate the information from different patches, which mainly contain some common statistics i.e., the mean & standard deviation, quantiles and moments). Finally, the aggregated features are fed into a linear regression model to predict the image quality. Experiments show that, compared with low-level features, high-level features indeed play a more critical role in resolving the aforementioned challenging problem for quality estimation. Besides, the proposed method significantly outperforms the state-of-the-art methods on two realistic blur image databases and achieves comparable performance on two synthetic blur image databases. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
ACM Multimedia | 1 |