EDBT 2026 Demo / reviewers in the wild / expert
Zhou Wang 0001
dblp:28/2295-1
· DBLP profile ↗
156ranked-venue papers
23as first author
30since 2021 · last 2026
0000-0003-4413-4441ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 147 · 20 first-author · 26 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AR-COMPQ: A Dual-Path Perceptual Quality Assessment for Augmented Reality Images Composition
Mahzar Eisapour, Zhou Wang 0001 |
QoMEX | 2 |
| 2026 | Structural Similarity in Deep Features: Unified Image Quality Assessment Robust to Geometrically Disparate ReferenceabstractImage Quality Assessment (IQA) with references plays an important role in optimizing and evaluating computer vision tasks. Traditional methods assume that all pixels of the reference and test images are fully aligned. Such Aligned-Reference IQA (AR-IQA) approaches fail to address many real-world problems with various geometric deformations between the two images. Although significant effort has been made to attack Geometrically-Disparate-Reference IQA (GDR-IQA) problem, it has been addressed in a task-dependent fashion, for example, by dedicated designs for image super-resolution and retargeting, or by assuming the geometric distortions to be small that can be countered by translation-robust filters or by explicit image registrations. Here we rethink this problem and propose a unified, non-training-based Deep Structural Similarity (DeepSSIM) approach to address the above problems in a single framework, which assesses structural similarity of deep features in a simple but efficient way and uses an attention calibration strategy to alleviate attention deviation. The proposed method, without application-specific design, achieves state-of-the-art performance on AR-IQA datasets and meanwhile shows strong robustness to various GDR-IQA test cases. Interestingly, our test also shows the effectiveness of DeepSSIM as an optimization tool for training image super-resolution, enhancement and restoration, implying an even wider generalizability. Tiesong Zhao, Zhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Perceptual Geometry Distortion Assessment of Compressed 3D MeshesabstractThe evaluation of perceptual quality in 3D mesh compression, particularly for Video-based Dynamic Mesh Coding (V-DMC), is challenged by the scarcity of subject-rated datasets and the high computational cost of full-mesh decoding and the sophisticated visual feature extraction steps. To bridge this gap, we first introduce a novel V-DMC distortion dataset, comprising 16 high-quality original meshes and 400 compressed, textureless variants. We conducted a subjective quality assessment study with 30 participants using the Double Stimulus Impairment Scale (DSIS) method to collect reliable Mean Opinion Scores (MOS). We then propose streamMQ, the first-of-its-kind no-reference, bitstream-layer model for perceptual quality assessment of V-DMC compressed meshes. By extracting key geometric features such as quantization parameters and triangle count directly from the compressed bitstream, streamMQ predicts perceptual quality without full decoding. Experimental evaluation and comparison with state-of-the-art methods demonstrate that streamMQ achieves highly competitive quality assessment performance at tiny fractions of computational and storage costs, facilitating real-time and low-storage application environments. The dataset and source code will be made publicly available at https://github.com/HFL01/QDU-GDM. Fanglin Hou, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Attention Redundancy Reduction for Image Super-ResolutionabstractTransformer-based models have demonstrated great promises in single image super-resolution (SISR), but our investigations find significant redundancy in terms of high mutual information across the attention maps, which is associated with reduced efficiency and degraded performance of SOTA models. To address the problem, here we propose a low redundancy attention network (LRAN). First, to mitigate the redundancy among heads, we introduce in the self-attention computation a multi-element mechanism, which allows for the incorporation of various types of self-attention, thus increasing inter-head diversity. Second, to address the redundancy among blocks, we propose the encapsulated architecture, in which enhanced local perception unit and gated multi-layer perceptron are designed to capture local information. Specifically, this architecture incorporates a single self-attention layer between several MLP layers. Subsequently, the proposed gated multi-layer perceptron significantly enhances the SR quality. Extensive experiments demonstrate that LRAN outperforms SOTA models in the task of lightweight SR, achieving a better trade-off between quality and speed. For instance, the proposed LRAN-light surpasses SwinIR-light by 0.32dB PSNR in $\times 4$ SR on Urban100, while running $\times 4$ faster. Yican Liu, Delu Zeng, Zhou Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Perceptual Quality Assessment of Trisoup-Lifting Encoded 3D Point CloudsabstractNo-reference bitstream-layer point cloud quality assessment (PCQA) can be deployed without full decoding at any network node to achieve real-time quality monitoring. In this work, we develop the first PCQA model dedicated to Trisoup-Lifting encoded 3D point clouds by analyzing bitstreams without full decoding. Specifically, we investigate the relationship among texture bitrate per point (TBPP), texture complexity (TC) and texture quantization parameter (TQP) while geometry encoding is lossless. Subsequently, we estimate TC by utilizing TQP and TBPP. Then, we establish a texture distortion evaluation model based on TC, TBPP and TQP. Ultimately, by integrating this texture distortion model with a geometry attenuation factor, a function of trisoupNodeSizeLog2 (tNSL), we acquire a comprehensive NR bitstream-layer PCQA model named streamPCQ-TL. In addition, this work establishes a database named WPC6.0, the first PCQA database dedicated to Trisoup-Lifting encoding mode, encompassing 400 distorted point clouds with 4 geometry multiplied by 5 texture distortion levels. Experiment results on M-PCCD, ICIP2020 and the proposed WPC6.0 database suggest that the proposed streamPCQ-TL model exhibits robust and notable performance in contrast to existing advanced PCQA metrics, particularly in terms of computational cost. Juncheng Long, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Wei Gao 0003, Jiarun Song, Zhou Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | HybridMQA: Exploring Geometry-Texture Interactions for Colored Mesh Quality AssessmentabstractMesh quality assessment (MQA) models play a critical role in the design, optimization, and evaluation of mesh operation systems in a wide variety of applications. Current MQA models, whether model-based methods using topology-aware features or projection-based approaches working on rendered 2D projections, often fail to capture the intricate interactions between texture and 3D geometry. We introduce HybridMQA, a first-of-its-kind hybrid full-reference colored MQA framework that integrates model-based and projection-based approaches, capturing complex interactions between textural information and 3D structures for enriched quality representations. Our method employs graph learning to extract detailed 3D representations, which are then projected to 2D using a novel feature rendering process that precisely aligns them with colored projections. This enables the exploration of geometry-texture interactions via cross-attention, producing comprehensive mesh quality representations. Extensive experiments demonstrate HybridMQA’s superior performance across diverse datasets, highlighting its ability to effectively leverage geometry-texture interactions for a thorough understanding of mesh quality. Our project website is available at https://arshafiee.github.io/hybridmqa/. Armin Shafiee Sarvestani, Sheyang Tang, Zhou Wang 0001 |
CVPR | 3 |
| 2025 | Controllable Data Generation with Hierarchical Neural RepresentationsabstractImplicit Neural Representations (INRs) represent data as continuous functions using the parameters of a neural network, where data information is encoded in the parameter space. Therefore, modeling the distribution of such parameters is crucial for building generalizable INRs. Existing approaches learn a joint distribution of these parameters via a latent vector to generate new data, but such a flat latent often fails to capture the inherent hierarchical structure of the parameter space, leading to entangled data semantics and limited control over the generation process. Here, we propose a Controllable Hierarchical Implicit Neural Representation (CHINR) framework, which explicitly models conditional dependencies across layers in the parameter space. Our method consists of two stages: In Stage-1, we construct a Layers-of-Experts (LoE) network, where each layer modulates distinct semantics through a unique latent vector, enabling disentangled and expressive representations. In Stage-2, we introduce a Hierarchical Conditional Diffusion Model (HCDM) to capture conditional dependencies across layers, allowing for controllable and hierarchical data generation at various semantic granularities. Extensive experiments across different modalities demonstrate that CHINR improves generalizability and offers flexible hierarchical control over the generated content. Sheyang Tang, Jiayan Qiu, Zhou Wang 0001 |
ICML | 4 |
| 2025 | DQP-PCQA: Deep Quantization Parameters Bring New Insight to Point Cloud Quality AssessmentabstractWith the rapid development of immersive multimedia technology, the growing demand for high-quality visual experiences has driven the emergence of point cloud quality assessment (PCQA). While current deep learning-based PCQA models have achieved breakthroughs in performance, problems such as high computational complexity and limited model generalization ability still need to be solved. In this study, focusing on compression distortion, we analyzed and verified that the compression quantization parameter (QP) can be used as a key feature for predicting perceptual quality. Based on this, a novel no-reference point cloud perceptual quality assessment metric, DQP-PCQA, is proposed. Unlike existing PCQA models that only use mean opinion score (MOS) as a supervisory label, this study proposes a multi-objective constrained optimization scheme that adds geometric quantization parameter (GQP) and texture quantization parameter (TQP) as auxiliary supervisory labels to help the model can learn robust perceptual features that take into account both subjective quality and objective distortion. We conducted comparative experiments with other advanced PCQA models on several mainstream PCQA datasets. The results show that the DQP-PCQA model achieves fast convergence speed, excellent and stable performance, low complexity and strong generalization. Further migration experiments show that after applying our proposed method to other advanced PCQA models, the performance of the improved model is further improved. Our discovery provides new insight for PCQA research. To facilitate future reproducible research, the source code will be publicly released at https://github.com/Dds46/DQP-PCQA. Dongshuai Duan, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Energy-Adaptive Bitstream-Layer Model for Perceptual Quality Assessment of V-PCC Encoded 3D Point CloudsabstractThe scope of point cloud (PC) applications is expanding. We propose a no-reference bitstream-layer quality assessment model that eliminates the need for full decoding of the PC, providing quality evaluation scores during the V-PCC decoding process. Specifically, we illustrate the relationship between content diversity (CD) and perceptual coding distortion in lossless geometric coding. Subsequently, we model attribute distortion by predicting CD using transform energy (TE) and texture quantization parameter (TQP). By combining the geometric distortion model with geometry quantization parameters (GQP) and the attribute distortion model, we derive comprehensive quality prediction results. Our experimental results on four PC databases (WPC2.0, M-PCCD, VSENSE VVDB and VSENSE VVDB2) show that the proposed energy-adaptive bitstream-layer model (EABL) delivers competitive quality prediction performance in comparison with existing full-reference, reduced-reference and no-reference PC quality assessment models that require full decoding, and meanwhile exhibits large speed advantage. The source code will be made publicly available for repeatability research at https://github.com/arthas-sws/EABL_model. Wusi Sang, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Perceptual Quality Assessment of 360° Images Based on Generative Scanpath RepresentationabstractDespite substantial efforts dedicated to the design of heuristic models for omnidirectional (i.e., 360°) image quality assessment (OIQA), a conspicuous gap remains due to the lack of consideration for the diversity of viewing behaviors that leads to the varying perceptual quality of 360° images. Two critical aspects underline this oversight: the neglect of viewing conditions that significantly sway user gaze patterns and the overreliance on a single viewport sequence from the 360° image for quality inference. To address these issues, we introduce a unique generative scanpath representation (GSR) for effective quality inference of 360° images, which aggregates varied perceptual experiences of multi-hypothesis users under a predefined viewing condition. More specifically, given a viewing condition characterized by the starting point of viewing and exploration time, a set of scanpaths consisting of dynamic visual fixations can be produced using an apt scanpath generator. Following this vein, we use the scanpaths to convert the 360° image into the unique GSR, which provides a global overview of gazed-focused contents derived from scanpaths. As such, the quality inference of the 360° image is swiftly transformed to that of GSR. We then propose an efficient OIQA computational framework by learning the quality maps of GSR. Comprehensive experimental results validate that the predictions of the proposed framework are highly consistent with human perception in the spatiotemporal domain, especially in the challenging context of locally distorted 360° images under varied viewing conditions. The code will be released at https://github.com/xiangjieSui/GSR. Xiangjie Sui, Hanwei Zhu, Xuelin Liu, Yuming Fang 0001, Shiqi Wang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Omnidirectional Image Quality Captioning: A Large-Scale Database and a New ModelabstractThe fast growing application of omnidirectional images calls for effective approaches for omnidirectional image quality assessment (OIQA). Existing OIQA methods have been developed and tested on homogeneously distorted omnidirectional images, but it is hard to transfer their success directly to the heterogeneously distorted omnidirectional images. In this paper, we conduct the largest study so far on OIQA, where we establish a large-scale database called OIQ-10K containing 10,000 omnidirectional images with both homogeneous and heterogeneous distortions. A comprehensive psychophysical study is elaborated to collect human opinions for each omnidirectional image, together with the spatial distributions (within local regions or globally) of distortions, and the head and eye movements of the subjects. Furthermore, we propose a novel multitask-derived adaptive feature-tailoring OIQA model named IQCaption360, which is capable of generating a quality caption for an omnidirectional image in a manner of textual template. Extensive experiments demonstrate the effectiveness of IQCaption360, which outperforms state-of-the-art methods by a significant margin on the proposed OIQ-10K database. The OIQ-10K database and the related source codes are available at https://github.com/WenJuing/IQCaption360. Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Junjie Chen 0008, Wenhui Jiang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | Human Perception-Guided Meta-Training for Few-Shot NeRFabstractNeural Radiance Fields (NeRF) have recently achieved impressive results in synthesizing novel views. However, with the few-shot setup, the NeRF is prone to overfitting to the supervision views, which causes artifacts when rendering novel views. To this end, we design a human-perception-guided meta-training framework to enhance the novel views without needing daunting annotation efforts. Specifically, we first analyze that the few-shot NeRF suffers from quality fluctuation, distant view degradation, and geometrical inconsistency. Then, we design a human perception guider to evaluate and select views for enhancement. Further, a human perception score modulated hyper-restoration module is designed to handle different degradation. Finally, we adopt meta-learning to fit enhanced views to avoid any impairment to the supervision views. Besides, we can directly embed any NeRF model into the framework for training and enhancement. Experimental results on two widely used benchmark datasets, LLFF and DTU, demonstrate our superiority to the state-of-the-art methods. Bingyin Li, Sheyang Tang, Li Yu 0003, Zhou Wang 0001 |
ICASSP | 5 |
| 2024 | Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori EstimationabstractOver the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individual models are often biased toward certain types of image content or distortions, depending on the design principle and process. An intuitive idea is to harness the strengths and mitigate the weaknesses of each IQA model, by fusing the scores of multiple models into a stronger one. Here we make one of the first attempts to seek an optimal solution for the idea and propose a general framework for unsupervised IQA score fusion using deep Maximum a Posteriori (MAP) estimation. The proposed model conducts fine-grained uncertainty estimation at the score level to increase the accuracy and reduce the uncertainty in fused predictions. Comprehensive experiments demonstrate the superiority of the proposed model over individual IQA models and other fusion methods. It also exhibits an interesting capability of rejecting "bad" models in the fusion process. Zhongling Wang, Raymond Zhou, Shahrukh Athar, Wenbo Yang 0001, Zhou Wang 0001 |
ICASSP | 5 |
| 2024 | Perceptual Crack Detection for Rendered 3D Textured MeshesabstractRecent years have witnessed many advancements in the applications of 3D textured meshes. As the demand continues to rise, evaluating the perceptual quality of this new type of media content becomes crucial for quality assurance and optimization purposes. Different from traditional image quality assessment, crack is an annoying artifact specific to rendered 3D meshes that severely affects their perceptual quality. In this work, we make one of the first attempts to propose a novel Perceptual Crack Detection (PCD) method for detecting and localizing crack artifacts in rendered meshes. Specifically, motivated by the characteristics of the human visual system (HVS), we adopt contrast and Laplacian measurement modules to characterize crack artifacts and differentiate them from other undesired artifacts. Extensive experiments on large-scale public datasets of 3D textured meshes demonstrate effectiveness and efficiency of the proposed PCD method in correct localization and detection of crack artifacts. Moreover, to quantify the performance of the proposed detection method and validate its effectiveness, we propose a simple yet effective weighting mechanism to incorporate the resulting crack map into classical quality assessment (QA) models, which creates significant performance improvement in predicting the perceptual image quality when tested on public datasets of static 3D textured meshes. A software release of the proposed method is publicly available at: https://github.com/arshafiee/crack-detection-VVM Armin Shafiee Sarvestani, Wei Zhou 0021, Zhou Wang 0001 |
QoMEX | 3 |
| 2024 | Perceptual Depth Quality Assessment of Stereoscopic Omnidirectional ImagesabstractDepth perception plays an essential role in the viewer experience for immersive virtual reality (VR) visual environments. However, previous research investigations in the depth quality of 3D/stereoscopic images are rather limited, and in particular, are largely lacking for 3D viewing of 360-degree omnidirectional content. In this work, we make one of the first attempts to develop an objective quality assessment model named depth quality index (DQI) for efficient no-reference (NR) depth quality assessment of stereoscopic omnidirectional images. Motivated by the perceptual characteristics of the human visual system (HVS), the proposed DQI is built upon multi-color-channel, adaptive viewport selection, and interocular discrepancy features. Experimental results demonstrate that the proposed method outperforms state-of-the-art image quality assessment (IQA) and depth quality assessment (DQA) approaches in predicting the perceptual depth quality when tested using both single-viewport and omnidirectional stereoscopic image databases. Furthermore, we demonstrate that combining the proposed depth quality model with existing IQA methods significantly boosts the performance in predicting the overall quality of 3D omnidirectional images. Wei Zhou 0021, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesabstractScanpath prediction for 360° images aims to produce dynamic gaze behaviors based on the human visual perception mechanism. Most existing scanpath prediction methods for 360° images do not give a complete treatment of the time-dependency when predicting human scanpath, resulting in inferior performance and poor generalizability. In this paper, we present a scanpath prediction method for 360° images by designing a novel Deep Markov Model (DMM) architecture, namely ScanDMM. We propose a semantics-guided transition function to learn the nonlinear dynamics of time-dependent attentional landscape. Moreover, a state initialization strategy is proposed by considering the starting point of viewing, enabling the model to learn the dynamics with the correct “launcher”. We further demonstrate that our model achieves state-of-the-art performance on four 360° image databases, and exhibit its generalizability by presenting two applications of applying scanpath prediction models to other visual tasks - saliency detection and image quality assessment, expecting to provide profound insights into these fields. Xiangjie Sui, Yuming Fang 0001, Hanwei Zhu, Shiqi Wang 0001, Zhou Wang 0001 |
CVPR | 5 |
| 2023 | Blind Omnidirectional Image Quality Assessment: Integrating Local Statistics and Global SemanticsabstractOmnidirectional image quality assessment (OIQA) aims to predict the perceptual quality of omnidirectional images that cover the whole 180×360° viewing range of the visual environment. Here we propose a blind/no-reference OIQA method named Local Statistics and Global Semantics metric (LSGS) that bridges the gap between low-level statistics and high-level semantics of omnidirectional images. Specifically, statistic and semantic features are extracted in separate paths from multiple local viewports and the hallucinated global omnidirectional image, respectively. A quality regression along with a weighting process is then followed that maps the extracted quality-aware features to a perceptual quality prediction. Experimental results demonstrate that the proposed LSGS method offers highly competitive performance against state-of-the-art methods. Wei Zhou 0021, Zhou Wang 0001 |
ICIP | 2 |
| 2023 | Degraded Reference Image Quality AssessmentabstractIn practical media distribution systems, visual content usually undergoes multiple stages of quality degradation along the delivery chain, but the pristine source content is rarely available at most quality monitoring points along the chain to serve as a reference for quality assessment. As a result, full-reference (FR) and reduced-reference (RR) image quality assessment (IQA) methods are generally infeasible. Although no-reference (NR) methods are readily applicable, their performance is often not reliable. On the other hand, intermediate references of degraded quality are often available, e.g., at the input of video transcoders, but how to make the best use of them in proper ways has not been deeply investigated. Here we make one of the first attempts to establish a new paradigm named degraded-reference IQA (DR IQA). Specifically, by using a two-stage distortion pipeline we lay out the architectures of DR IQA and introduce a 6-bit code to denote the choices of configurations. We construct the first large-scale databases dedicated to DR IQA and have made them publicly available. We make novel observations on distortion behavior in multi-stage distortion pipelines by comprehensively analyzing five multiple distortion combinations. Based on these observations, we develop novel DR IQA models and make extensive comparisons with a series of baseline models derived from top-performing FR and NR models. The results suggest that DR IQA may offer significant performance improvement in multiple distortion environments, thereby establishing DR IQA as a valid IQA paradigm that is worth further exploration. Shahrukh Athar, Zhou Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Bitstream-Based Perceptual Quality Assessment of Compressed 3D Point CloudsabstractWith the increasing demand of compressing and streaming 3D point clouds under constrained bandwidth, it has become ever more important to accurately and efficiently determine the quality of compressed point clouds, so as to assess and optimize the quality-of-experience (QoE) of end users. Here we make one of the first attempts developing a bitstream-based no-reference (NR) model for perceptual quality assessment of point clouds without resorting to full decoding of the compressed data stream. Specifically, we first establish a relationship between texture complexity and the bitrate and texture quantization parameters based on an empirical rate-distortion model. We then construct a texture distortion assessment model upon texture complexity and quantization parameters. By combining this texture distortion model with a geometric distortion model derived from Trisoup geometry encoding parameters, we obtain an overall bitstream-based NR point cloud quality model named streamPCQ. Experimental results show that the proposed streamPCQ model demonstrates highly competitive performance when compared with existing classic full-reference (FR) and reduced-reference (RR) point cloud quality assessment methods with a fraction of computational cost. Honglei Su, Qi Liu 0029, Hui Yuan 0001, Huan Yang 0001, Zhenkuan Pan 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 7 |
| 2023 | Perceptual Quality Assessment of Colored 3D Point Cloudsabstract3D point clouds have found a wide variety of applications in multimedia processing, remote sensing, and scientific computing. Although most point cloud processing systems are developed to improve viewer experiences, little work has been dedicated to perceptual quality assessment of 3D point clouds. In this work, we build a new 3D point cloud database, namely the Waterloo Point Cloud (WPC) database. In contrast to existing datasets consisting of small-scale and low-quality source content of constrained viewing angles, the WPC database contains 20 high quality, realistic, and omni-directional source point clouds and 740 diversely distorted point clouds. We carry out a subjective quality assessment experiment over the database in a controlled lab environment. Our statistical analysis suggests that existing objective point cloud quality assessment (PCQA) models only achieve limited success in predicting subjective quality ratings. We propose a novel objective PCQA model based on an attention mechanism and a variant of information content-weighted structural similarity, which significantly outperforms existing PCQA models. The database has been made publicly available at https://github.com/qdushl/Waterloo-Point-Cloud-Database. Qi Liu 0029, Honglei Su, Zhengfang Duanmu, Wentao Liu 0001, Zhou Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Deep Image DebandingabstractBanding or false contour is an annoying visual artifact whose impact negatively degrades the perceptual quality of visual content. Since users are increasingly expecting better visual quality from such content and banding leads to deteriorated quality-of-experience, the area of banding removal or debanding has taken paramount importance. Existing debanding approaches are mostly knowledge-driven, while data-driven debanding approaches remain surprisingly missing. In this work, we construct a large-scale dataset of 51,490 pairs of corresponding pristine and banded image patches, which enables us to make one of the first attempts at developing a deep learning based banding artifact removal method for images that we name deep debanding network (deepDeband). We also develop a bilateral weighting scheme that fuses patch-level debanding results to full-size images. Extensive performance evaluation shows that deepDeband is successful at greatly reducing banding artifacts in images, outperforming existing methods both quantitatively and visually. The proposed algorithm and dataset are made publicly available.1 Raymond Zhou, Shahrukh Athar, Zhongling Wang, Zhou Wang 0001 |
ICIP | 4 |
| 2022 | Quality Assessment of Image Super-Resolution: Balancing Deterministic and Statistical FidelityabstractThere has been a growing interest in developing image super-resolution (SR) algorithms that convert low-resolution (LR) to higher resolution images, but automatically evaluating the visual quality of super-resolved images remains a challenging problem. Here we look at the problem of SR image quality assessment (SR IQA) in a two-dimensional (2D) space of deterministic fidelity (DF) versus statistical fidelity (SF). This allows us to better understand the advantages and disadvantages of existing SR algorithms, which produce images at different clusters in the 2D space of (DF, SF). Specifically, we observe an interesting trend from more traditional SR algorithms that are typically inclined to optimize for DF while losing SF, to more recent generative adversarial network (GAN) based approaches that by contrast exhibit strong advantages in achieving high SF but sometimes appear weak at maintaining DF. Furthermore, we propose an uncertainty weighting scheme based on content-dependent sharpness and texture assessment that merges the two fidelity measures into an overall quality prediction named the Super Resolution Image Fidelity (SRIF) index, which demonstrates superior performance against state-of-the-art IQA models when tested on subject-rated datasets. Wei Zhou 0021, Zhou Wang 0001 |
ACM Multimedia | 2 |
| 2022 | Target Detection Algorithm Based on Feature Optimization and Sample Equalization
Chao Li 0053, Fangzheng Huang, Zhaoxian Yang, Zhou Wang 0001, Dayan Ban |
WASA (2) | 4 |
| 2022 | Hierarchical Semantic Risk Minimization for Large-Scale ClassificationabstractHierarchical structures of labels usually exist in large-scale classification tasks, where labels can be organized into a tree-shaped structure. The nodes near the root stand for coarser labels, while the nodes close to leaves mean the finer labels. We label unseen samples from the root node to a leaf node, and obtain multigranularity predictions in the hierarchical classification. Sometimes, we cannot obtain a leaf decision due to uncertainty or incomplete information. In this case, we should stop at an internal node, rather than going ahead rashly. However, most existing hierarchical classification models aim at maximizing the percentage of correct predictions, and do not take the risk of misclassifications into account. Such risk is critically important in some real-world applications, and can be measured by the distance between the ground truth and the predicted classes in the class hierarchy. In this work, we utilize the semantic hierarchy to define the classification risk and design an optimization technique to reduce such risk. By defining the conservative risk and the precipitant risk as two competing risk factors, we construct the balanced conservative/precipitant semantic (BCPS) risk matrix across all nodes in the semantic hierarchy with user-defined weights to adjust the tradeoff between two kinds of risks. We then model the classification process on the semantic hierarchy as a sequential decision-making task. We design an algorithm to derive the risk-minimized predictions. There are two modules in this model: 1) multitask hierarchical learning and 2) deep reinforce multigranularity learning. The first one learns classification confidence scores of multiple levels. These scores are then fed into deep reinforced multigranularity learning for obtaining a global risk-minimized prediction with flexible granularity. Experimental results show that the proposed model outperforms state-of-the-art methods on seven large-scale classification datasets with the semantic tree. Yu Wang 0106, Zhou Wang 0001, Qinghua Hu, Yucan Zhou, Honglei Su |
IEEE Trans. Cybern. | 2 |
| 2022 | Learning-Based Quality Assessment for Image Super-ResolutionabstractImage Super-Resolution (SR) techniques improve visual quality by enhancing the spatial resolution of images. Quality evaluation metrics play a critical role in comparing and optimizing SR algorithms, but current metrics achieve only limited success, largely due to the lack of large-scale quality databases, which are essential for learning accurate and robust SR quality metrics. In this work, we first build a large-scale SR image database using a novel semi-automatic labeling approach, which allows us to label a large number of images with manageable human workload. The resulting SR Image quality database with Semi-Automatic Ratings (SISAR), so far the largest of SR-IQA database, contains 12 600 images of 100 natural scenes. We train an end-to-end Deep Image SR Quality (DISQ) model by employing two-stream Deep Neural Networks (DNNs) for feature extraction, followed by a feature fusion network for quality prediction. Experimental results demonstrate that the proposed method outperforms state-of-the-art metrics and achieves promising generalization performance in cross-database tests. The SISAR database and DISQ model will be made publicly available to facilitate reproducible research. Tiesong Zhao, Yuting Lin 0006, Zhou Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | A Bayesian Quality-of-Experience Model for Adaptive Streaming VideosabstractThe fundamental conflict between the enormous space of adaptive streaming videos and the limited capacity for subjective experiment casts significant challenges to objective Quality-of-Experience (QoE) prediction. Existing objective QoE models either employ pre-defined parametrization or exhibit complex functional form, achieving limited generalization capability in diverse streaming environments. In this study, we propose an objective QoE model, namely, the Bayesian streaming quality index (BSQI), to integrate prior knowledge on the human visual system and human annotated data in a principled way. By analyzing the subjective characteristics towards streaming videos from a corpus of subjective studies, we show that a family of QoE functions lies in a convex set. Using a variant of projected gradient descent, we optimize the objective QoE model over a database of training videos. The proposed BSQI demonstrates strong prediction accuracy in a broad range of streaming conditions, evident by state-of-the-art performance on four publicly available benchmark datasets and a novel analysis-by-synthesis visual experiment. Zhengfang Duanmu, Wentao Liu 0001, Diqi Chen, Zhou Wang 0001, Yizhou Wang 0001, Wen Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Capturing Banding in Images: Database Construction and Objective AssessmentabstractWith the fast technology advancement and the accelerated growth of high-quality image and video production and services, banding or false contour has become a frequently observed artifact in images, creating annoying negative impact on the visual quality-of-experience (QoE) of end users. Nevertheless, thorough investigations on the causes of banding, and effective and efficient methods to detect and reduce banding are largely lacking. This work targets at capturing and quantifying banding artifacts in images. We construct the first of its kind large-scale public database, consisting of 1,250 images with segmented banding regions and 169,501 image patches with class labels. We also develop a deep neural net-work based no-reference deep banding index (DBI), which not only produces an overall banding assessment of a given image, but also creates a banding map that indicates the variation of banding across the image space. Our experiments show that the proposed DBI method achieves accurate banding prediction with low computational cost. The database and the proposed algorithm are made publicly available1. Akshay Kapoor, Jatin Sapra, Zhou Wang 0001 |
ICASSP | 3 |
| 2021 | Autoencoder for Vibrotactile Signal CompressionabstractVibrotactile signals contain rich haptic information about textured surfaces but their large data volume makes it a challenging task to transmit such signals to remote locations to create immersive and realistic user experiences. Inspired by the recent success of deep neural network (DNN) based autoencoder, we make the first attempt to apply autoencoder for lossy compression of haptic vibrotactile signals, where a convolutional neural network (CNN) and a rate-distortion (RD) function are used as the transform and cost functions, respectively. Performance comparisons with state-of-the-art methods using both peak signal-to-noise ratio (PSNR) and perceptually motivated spectral temporal similarity (ST-SIM) measures show that the proposed end-to-end vibrotactile autoencoder (EVA) is highly competitive at preserving signal quality while keeping the data rate low. Rania Hassen, Zhou Wang 0001 |
ICASSP | 3 |
| 2021 | Real Versus Fake 4k - Authentic Resolution AssessmentabstractIn recent years, the native 4K/Ultra High Definition (UHD) resolution has been trending towards the new normal of video content creation and distribution, but the practical pipelines of video acquisition, production and delivery often involve downscaling stages where the spatial resolution drops below the 4K level. Even though the video may be upscaled back to 4K/UHD resolution later, the content has lost its authentic resolution. This work aims at authentic resolution assessment (ARA). We first construct a database of over 10,000 real and fake 4K/UHD images. We then develop a two-stage ARA (TSARA) approach that classifies a video frame to have real or fake 4K resolution, where the first stage classifies local patches using a convolutional neural network (CNN), and the second stage aggregates local assessment into a global image level decision using logistical regression. Experimental results show that the proposed approach achieves high accuracy at low computational cost, and outperforms state-of-the-art no-reference (NR) image quality assessment (IQA) and image sharpness assessment (ISA) models. The built database and the proposed method are made publicly available1. Rishi Rajesh Shah, Vyas Anirudh Akundy, Zhou Wang 0001 |
ICASSP | 3 |
| 2021 | Image Super-Resolution Quality Assessment: Structural Fidelity Versus Statistical NaturalnessabstractSingle image super-resolution (SISR) algorithms reconstruct high-resolution (HR) images with their low-resolution (LR) counterparts. It is desirable to develop image quality assessment (IQA) methods that can not only evaluate and compare SISR algorithms, but also guide their future development. In this paper, we assess the quality of SISR generated images in a two-dimensional (2D) space of structural fidelity versus statistical naturalness. This allows us to observe the behaviors of different SISR algorithms as a tradeoff in the 2D space. Specifically, SISR methods are traditionally designed to achieve high structural fidelity but often sacrifice statistical naturalness, while recent generative adversarial network (GAN) based algorithms tend to create more natural-looking results but lose significantly on structural fidelity. Furthermore, such a 2D evaluation can be easily fused to a scalar quality prediction. Interestingly, we find that a simple linear combination of a straightforward local structural fidelity and a global statistical naturalness measures produce surprisingly accurate predictions of SISR image quality when tested using public subject-rated SISR image datasets. Code of the proposed SFSN model is publicly available at https://github.con/weizhou-geek/SFSN. Wei Zhou 0021, Zhou Wang 0001, Zhibo Chen 0001 |
QoMEX | 2 |
| 2020 | Perceptual Quality Assessment of Smartphone PhotographyabstractAs smartphones become people's primary cameras to take photos, the quality of their cameras and the associated computational photography modules has become a de facto standard in evaluating and ranking smartphones in the consumer market. We conduct so far the most comprehensive study of perceptual quality assessment of smartphone photography. We introduce the Smartphone Photography Attribute and Quality (SPAQ) database, consisting of 11,125 pictures taken by 66 smartphones, where each image is attached with so far the richest annotations. Specifically, we collect a series of human opinions for each image, including image quality, image attributes (brightness, colorfulness, contrast, noisiness, and sharpness), and scene category labels (animal, cityscape, human, indoor scene, landscape, night scene, plant, still life, and others) in a well-controlled laboratory environment. The exchangeable image file format (EXIF) data for all images are also recorded to aid deeper analysis. We also make the first attempts using the database to train blind image quality assessment (BIQA) models constructed by baseline and multi-task deep neural networks. The results provide useful insights on how EXIF data, image attributes and high-level semantics interact with image quality, how next-generation BIQA models can be designed, and how better computational photography systems can be optimized on mobile devices. The database along with the proposed BIQA models are available at https://github.com/h4nwei/SPAQ. Yuming Fang 0001, Hanwei Zhu, Yan Zeng 0001, Kede Ma, Zhou Wang 0001 |
CVPR | 5 |
| 2020 | Perceptual Colour Difference Uniformity in High Dynamic Range and Wide Colour GamutabstractPerceptual uniformity is a highly desirable property of colour spaces or colour difference measures where equal level in colour value difference should result in equal perceptual difference. Designing colour spaces or colour difference measures of perceptual uniformity is a long standing problem in colour science. This has become increasingly important with the growing popularity of high dynamic range (HDR) and wide colour gamut (WCG) cameras, content and displays. We design an efficient testing framework to evaluate perceptual uniformity by subjective just noticeable difference (JND) measurement at a wide range of luminance levels followed by coefficient of variation (CV) computation. We carry out subjective testing on RGB, xyY, L*a*b*, YCbCr, CIECAM02-UCS and ICtCp colour spaces and ΔE2000 metric in ITU-R BT 2020 colour gamut across a wide range of luminance levels (from 0.01 to 500 nits) using a professional HDR/WCG display in a carefully controlled dark testing environment. Our results suggest that on average, the ICtCp space performs the best in the current test, but is still distant from achieving perceptual uniformity. Thilan Costa, Vincent C. Gaudet, Edward R. Vrscay, Zhou Wang 0001 |
ICIP | 4 |
| 2020 | FocusLiteNN: High Efficiency Focus Quality Assessment for Digital Pathology
Zhongling Wang, Mahdi S. Hosseini, Adyn Miles, Konstantinos N. Plataniotis, Zhou Wang 0001 |
MICCAI (5) | 5 |
| 2020 | Group Maximum Differentiation Competition: Model Comparison with Few SamplesabstractIn many science and engineering fields that require computational models to predict certain physical quantities, we are often faced with the selection of the best model under the constraint that only a small sample set can be physically measured. One such example is the prediction of human perception of visual quality, where sample images live in a high dimensional space with enormous content variations. We propose a new methodology for model comparison named group maximum differentiation (gMAD) competition. Given multiple computational models, gMAD maximizes the chances of falsifying a "defender" model using the rest models as "attackers". It exploits the sample space to find sample pairs that maximally differentiate the attackers while holding the defender fixed. Based on the results of the attacking-defending game, we introduce two measures, aggressiveness and resistance, to summarize the performance of each model at attacking other models and defending attacks from other models, respectively. We demonstrate the gMAD competition using three examples-image quality, image aesthetics, and streaming video quality-of-experience. Although these examples focus on visually discriminable quantities, the gMAD methodology can be extended to many other fields, and is especially useful when the sample space is large, the physical measurement is expensive and the cost of computational prediction is low. Kede Ma, Zhengfang Duanmu, Zhou Wang 0001, Qingbo Wu 0001, Wentao Liu 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | PEA265: Perceptual Assessment of Video Compression ArtifactsabstractThe most widely used video encoders share a common hybrid coding framework that includes block-based motion estimation/compensation and block-based transform coding. Despite their high coding efficiency, the encoded videos often exhibit visually annoying artifacts, denoted as Perceivable Encoding Artifacts (PEAs), which significantly degrade the visual Quality-of-Experience (QoE) of end users. To monitor and improve visual QoE, it is crucial to develop subjective and objective measures that can identify and quantify various types of PEAs. In this work, we make the first attempt to build a large-scale subject-labeled database composed of H.265/HEVC compressed videos containing various PEAs. The database, namely the PEA265, includes 4 types of spatial PEAs (i.e. blurring, blocking, ringing and color bleeding) and 2 types of temporal PEAs (i.e. flickering and floating). Each containing at least 60,000 image or video patches with positive and negative labels. Based on the PEA265 database, we develop and optimize Convolutional Neural Networks (CNNs) to objectively recognize different types of PEAs. Experiments show that our architecture is capable of identifying the 6 types of PEAs with an accuracy over 86%. To further demonstrate its application, we explore the relationship between collected PEA intensities and subjective quality scores of compressed videos. A quality metric is consequently proposed with superior performance in terms of correlation to Mean Opinion Score (MOS) values. We believe that the PEA265 database and our findings will benefit the future development of video quality assessment methods and perceptually motivated video encoders. Liqun Lin, Shiqi Yu 0004, Tiesong Zhao, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Blind Image Quality Assessment Using a Deep Bilinear Convolutional Neural NetworkabstractWe propose a deep bilinear model for blind image quality assessment that works for both synthetically and authentically distorted images. Our model constitutes two streams of deep convolutional neural networks (CNNs), specializing in two distortion scenarios separately. For synthetic distortions, we first pre-train a CNN to classify the distortion type and the level of an input image, whose ground truth label is readily available at a large scale. For authentic distortions, we make use of a pre-train CNN (VGG-16) for the image classification task. The two feature sets are bilinearly pooled into one representation for a final quality prediction. We fine-tune the whole network on the target databases using a variant of stochastic gradient descent. The extensive experimental results show that the proposed model achieves state-of-the-art performance on both synthetic and authentic IQA databases. Furthermore, we verify the generalizability of our method on the large-scale Waterloo Exploration Database, and demonstrate its competitiveness using the group maximum differentiation competition methodology. Weixia Zhang, Kede Ma, Jia Yan 0006, Dexiang Deng, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Characterizing Generalized Rate-Distortion Performance of Video Coding: An Eigen Analysis ApproachabstractRate-distortion (RD) theory is at the heart of lossy data compression. Here we aim to model the generalized RD (GRD) trade-off between the visual quality of a compressed video and its encoding profiles (e.g., bitrate and spatial resolution). We first define the theoretical functional space W of the GRD function by analyzing its mathematical properties. We show that W is a convex set in a Hilbert space, inspiring a computational model of the GRD function, and a method of estimating model parameters from sparse measurements. To demonstrate the feasibility of our idea, we collect a large-scale database of real-world GRD functions, which turn out to live in a low-dimensional subspace of W. Combining the GRD reconstruction framework and the learned low-dimensional space, we create a low-parameter eigen GRD method to accurately estimate the GRD function of a source video content from only a few queries. Experimental results on the database show that the learned GRD method significantly outperforms state-of-the-art empirical RD estimation methods both in accuracy and efficiency. Last, we demonstrate the promise of the proposed model in video codec comparison. Zhengfang Duanmu, Wentao Liu 0001, Kede Ma, Zhou Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Modeling Generalized Rate-Distortion FunctionsabstractMany multimedia applications require precise understanding of the rate-distortion characteristics measured by the function relating visual quality to media attributes, for which we term it the generalized rate-distortion (GRD) function. In this study, we explore the GRD behavior of compressed digital videos in a two-dimensional space of bitrate and resolution. Our analysis on a large-scale video dataset reveals that empirical parametric models are systematically biased while exhaustive search methods require excessive computation time to depict the GRD surfaces. By exploiting the properties that all GRD functions share, we develop an Robust Axial-Monotonic Clough-Tocher (RAMCT) interpolation method to model the GRD function. This model allows us to accurately reconstruct the complete GRD function of a source video content from a moderate number of measurements. To further reduce the computational cost, we present a novel sampling scheme based on a probabilistic model and an information measure. The proposed sampling method constructs a sequence of quality queries by minimizing the overall informativeness in the remaining samples. Experimental results show that the proposed algorithm significantly outperforms state-of-the-art approaches in accuracy and efficiency. Finally, we demonstrate the usage of the proposed model in three applications: rate-distortion curve prediction, per-title encoding profile generation, and video encoder comparison. Zhengfang Duanmu, Wentao Liu 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Perceptual Evaluation for Multi-Exposure Image Fusion of Dynamic ScenesabstractA common approach to high dynamic range (HDR) imaging is to capture multiple images of different exposures followed by multi-exposure image fusion (MEF) in either radiance or intensity domain. A predominant problem of this approach is the introduction of the ghosting artifacts in dynamic scenes with camera and object motion. While many MEF methods (often referred to as deghosting algorithms) have been proposed for reduced ghosting artifacts and improved visual quality, little work has been dedicated to perceptual evaluation of their deghosting results. Here we first construct a database that contains 20 multiexposure sequences of dynamic scenes and their corresponding fused images by nine MEF algorithms. We then carry out a subjective experiment to evaluate fused image quality, and find that none of existing objective quality models for MEF provides accurate quality predictions. Motivated by this, we develop an objective quality model for MEF of dynamic scenes. Specifically, we divide the test image into static and dynamic regions, measure structural similarity between the image and the corresponding sequence in the two regions separately, and combine quality measurements of the two regions into an overall quality score. Experimental results show that the proposed method significantly outperforms the state-of-the-art. In addition, we demonstrate the promise of the proposed model in parameter tuning of MEF methods.1. Yuming Fang 0001, Hanwei Zhu, Kede Ma, Zhou Wang 0001, Shutao Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Guided Learning for Fast Multi-Exposure Image FusionabstractWe propose a fast multi-exposure image fusion (MEF) method, namely MEF-Net, for static image sequences of arbitrary spatial resolution and exposure number. We first feed a low-resolution version of the input sequence to a fully convolutional network for weight map prediction. We then jointly upsample the weight maps using a guided filter. The final image is computed by a weighted fusion. Unlike conventional MEF methods, MEF-Net is trained end-to-end by optimizing the perceptually calibrated MEF structural similarity (MEF-SSIM) index over a database of training sequences at full resolution. Across an independent set of test sequences, we find that the optimized MEF-Net achieves consistent improvement in visual quality for most sequences, and runs 10 to 1000 times faster than state-of-the-art methods. The code is made publicly available at. Kede Ma, Zhengfang Duanmu, Hanwei Zhu, Yuming Fang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Perceptual Quality Assessment of UHD-HDR-WCG VideosabstractHigh Dynamic Range (HDR) Wide Color Gamut (WCG) Ultra High Definition (4K/UHD) content has become increasingly popular recently. Due to the increased data rate, novel video compression methods have been developed to maintain the quality of the videos being delivered to consumers under bandwidth constraints. This has led to new challenges for the development of objective Video Quality Assessment (VQA) models, which are traditionally designed without sufficient calibration and validation based on subjective quality assessment of UHD-HDR-WCG videos. The large performance variations between different consumer HDR TVs, and between consumer HDR TVs and professional HDR reference displays used for content production, further complicates the task of acquiring reliable subjective data that faithfully reflects the impact of compression on UHD-HDR-WCG videos. In this work, we construct a first-of-its-kind video database composed of PQ-encoded UHD-HDR-WCG content, which is subsequently compressed by H.264 and HEVC encoders. We carry out a subjective study on a professional 4K-HDR reference display in a controlled lab environment. We also benchmark representative Full Reference (FR) and No-Reference (NR) objective VQA models against the subjective data to evaluate their performance on compressed UHD-HDR-WCG video content. The database will be made available to the public, subject to content copyright constraints. Shahrukh Athar, Thilan Costa, Kai Zeng 0003, Zhou Wang 0001 |
ICIP | 4 |
| 2019 | Perceptual Quality Assessment of 3d Point CloudsabstractThe real-world applications of 3D point clouds have been growing rapidly in recent years, but effective approaches and datasets to assess the quality of 3D point clouds are largely lacking. In this work, we construct so far the largest 3D point cloud database with diverse source content and distortion patterns, and carry out a comprehensive subjective user study. We construct 20 high quality, realistic, and omni-directional point clouds of diverse contents. We then apply downsampling, Gaussian noise, and three types of compression algorithms to create 740 distorted point clouds. Based on the database, we carry out a subjective experiment to evaluate the quality of distorted point clouds, and perform a point cloud encoder comparison. Our statistical analysis find that existing point cloud quality assessment models are limited in predicting subjective quality ratings. The database will be made publicly available to facilitate future research. Honglei Su, Zhengfang Duanmu, Wentao Liu 0001, Qi Liu 0029, Zhou Wang 0001 |
ICIP | 5 |
| 2018 | Geometric Transformation Invariant Image Quality Assessment Using Convolutional Neural NetworksabstractMost existing full-reference (FR) image quality assessment (IQA) models assume that the reference and distorted images are perfectly aligned, and fail dramatically when the assumption does not hold. In this study, we first show that pre-registration, especially feature-based (as opposed to area-based) registration, is effective at reducing the performance drop of FR-IQA models. However, registration is an expensive process that often slows down the speed of the IQA algorithms by several orders of magnitude. This motivates us to construct an end-to-end convolutional neural network (CNN) for direct image quality prediction, which contains built-in invariance to geometric distortions. Our results show that when the training images are augmented by their geometrically transformed versions, the learned network performs at a high level without image registration, resulting in a fast and effective approach for geometric transformation invariant IQA. Kede Ma, Zhengfang Duanmu, Zhou Wang 0001 |
ICASSP | 3 |
| 2018 | Temporal Motion Smoothness and the Impact of Frame Rate Variation on Video QualityabstractThere has been a strong recent trend to improve the perceptual quality-of-experience of viewers by expanding the spatial resolution, dynamic range, color gamut, and frame rate of videos. Conceptually, increasing video frame rate should create a benefit of smoother perception of motion. However, how to measure motion smoothness is not a well resolved problem. In this study, we measure the smoothness of motion by examining the local phase correlation of complex wavelet coefficients along the temporal direction. Our experiments based on subjective-rated databases show that this novel measure provides a new means to capture the impact of frame rate on video quality, and demonstrates strong promise at improving the performance of objective video quality assessment models. Rasoul Mohammadi Nasiri, Zhengfang Duanmu, Zhou Wang 0001 |
ICIP | 3 |
| 2018 | End-to-End Blind Quality Assessment of Compressed Videos Using Deep Neural NetworksabstractBlind video quality assessment (BVQA) algorithms are traditionally designed with a two-stage approach - a feature extraction stage that computes typically hand-crafted spatial and/or temporal features, and a regression stage working in the feature space that predicts the perceptual quality of the video. Unlike the traditional BVQA methods, we propose a Video Multi-task End-to-end Optimized neural Network (V-MEON) that merges the two stages into one, where the feature extractor and the regressor are jointly optimized. Our model uses a multi-task DNN framework that not only estimates the perceptual quality of the test video but also provides a probabilistic prediction of its codec type. This framework allows us to train the network with two complementary sets of labels, both of which can be obtained at low cost. The training process is composed of two steps. In the first step, early convolutional layers are pre-trained to extract spatiotemporal quality-related features with the codec classification subtask. In the second step, initialized with the pre-trained feature extractor, the whole network is jointly optimized with the two subtasks together. An additional critical step is the adoption of 3D convolutional layers, which creates novel spatiotemporal features that lead to a significant performance boost. Experimental results show that the proposed model clearly outperforms state-of-the-art BVQA methods.The source code of V-MEON is available at https://ece.uwaterloo.ca/~zduanmu/acmmm2018bvqa. Wentao Liu 0001, Zhengfang Duanmu, Zhou Wang 0001 |
ACM Multimedia | 3 |
| 2018 | Quality-of-Experience for Adaptive Streaming Videos: An Expectation Confirmation Theory Motivated ApproachabstractThe dynamic adaptive streaming over HTTP (DASH) provides an inter-operable solution to overcome volatile network conditions, but how the human visual quality-ofexperience (QoE) changes with time-varying video quality is not well-understood. Here, we build a large-scale video database of time-varying quality and design a series of subjective experiments to investigate how humans respond to compression level, spatial and temporal resolution adaptations. Our path-analytic results show that quality adaptations influence the QoE by modifying the perceived quality of subsequent video segments. Specifically, the quality deviation introduced by quality adaptations is asymmetric with respect to the adaptation direction, which is further influenced by other factors such as compression level and content. Furthermore, we propose an objective QoE model by integrating the empirical findings from our subjective experiments and the expectation confirmation theory (ECT). Experimental results show that the proposed ECT-QoE model is in close agreement with subjective opinions and significantly outperforms existing QoE models. The video database together with the code are available online at https://ece.uwaterloo.ca/~zduanmu/tip2018ectqoe/. Zhengfang Duanmu, Kede Ma, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Deep Blur Mapping: Exploiting High-Level Semantics by Deep Neural NetworksabstractThe human visual system excels at detecting local blur of visual images, but the underlying mechanism is not well understood. Traditional views of blur such as reduction in energy at high frequencies and loss of phase coherence at localized features have fundamental limitations. For example, they cannot well discriminate flat regions from blurred ones. Here we propose that high-level semantic information is critical in successfully identifying local blur. Therefore, we resort to deep neural networks that are proficient at learning high-level features and propose the first end-to-end local blur mapping algorithm based on a fully convolutional network. By analyzing various architectures with different depths and design philosophies, we empirically show that high-level features of deeper layers play a more important role than low-level features of shallower layers in resolving challenging ambiguities for this task. We test the proposed method on a standard blur detection benchmark and demonstrate that it significantly advances the state-of-the-art (ODS F-score of 0.853). Furthermore, we explore the use of the generated blur maps in three applications, including blur region segmentation, blur degree estimation, and blur magnification. Kede Ma, Huan Fu, Tongliang Liu, Zhou Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2018 | End-to-End Blind Image Quality Assessment Using Deep Neural NetworksabstractWe propose a multi-task end-to-end optimized deep neural network (MEON) for blind image quality assessment (BIQA). MEON consists of two sub-networks-a distortion identification network and a quality prediction network-sharing the early layers. Unlike traditional methods used for training multi-task networks, our training process is performed in two steps. In the first step, we train a distortion type identification sub-network, for which large-scale training samples are readily available. In the second step, starting from the pre-trained early layers and the outputs of the first sub-network, we train a quality prediction sub-network using a variant of the stochastic gradient descent method. Different from most deep neural networks, we choose biologically inspired generalized divisive normalization (GDN) instead of rectified linear unit as the activation function. We empirically demonstrate that GDN is effective at reducing model parameters/layers while achieving similar quality prediction performance. With modest model complexity, the proposed MEON index achieves state-of-the-art performance on four publicly available benchmarks. Moreover, we demonstrate the strong competitiveness of MEON against state-of-the-art BIQA models using the group maximum differentiation competition methodology. Kede Ma, Wentao Liu 0001, Kai Zhang 0008, Zhengfang Duanmu, Zhou Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 5 |
| 2017 | Quality assessment of images undergoing multiple distortion stagesabstractIn practical media distribution systems, visual content often undergoes multiple stages of quality degradations along the delivery chain between the source and destination. By contrast, current image quality assessment (IQA) models are typically validated on image databases with a single distortion stage. In this work, we construct two large-scale image databases that are composed of more than 2 million images undergoing multiple stages of distortions and examine how state-of-the-art IQA algorithms behave over distortion stages. Our results suggest that the performance of existing IQA models degrades rapidly with distortion stages, especially when the distortion types of different stages vary. We also find that full-reference and no-reference frameworks, though both readily applicable, have major drawbacks at predicting the quality of images at middle distortion stages. However, when the quality level of the previous stage is accessible, significantly improved quality prediction performance may be achieved. This study points out a new avenue of degraded-reference IQA research that is both practically desirable and technically challenging. Shahrukh Athar, Abdul Rehman 0001, Zhou Wang 0001 |
ICIP | 3 |
| 2017 | Perceptual quality assessment of HDR deghosting algorithmsabstractHigh dynamic range (HDR) imaging techniques aim to extend the dynamic range of images that cannot be well captured using conventional camera sensors. A common practice is to take a stack of pictures with different exposure levels and fuse them to produce a final image with more details. However, a small displacement between images caused by either camera or scene motion would void the benefits and cause the so-called ghosting artifacts. Over the past decade, many HDR deghosting algorithms have been proposed, but little work has been dedicated to evaluate HDR deghosting results either subjectively or objectively. In this work, we present a comprehensive subjective study for HDR deghosting. Specifically, we create a database that contains 20 dynamic image sequences and their corresponding deghosting results by 9 deghosting algorithms. A subjective user study is then carried out to evaluate the perceptual quality of deghosted images. The experimental results demonstrate the performance and limitations of existing HDR deghosting algorithm as well as no-reference image quality assessment models. In the future, we will make the database available to the public. Yuming Fang 0001, Hanwei Zhu, Kede Ma, Zhou Wang 0001 |
ICIP | 4 |
| 2017 | A database for perceptual evaluation of image aestheticsabstractObjective image aesthetics assessment (IAA) is attracting an increasing amount of attention in recent years. One of the most critical issues that hampers IAA research is the lack of publicly available and reliable image databases that can be used to train and test IAA features and models, especially those databases that offer continuous-valued subjective opinion scores. In this work, we construct a Waterloo IAA database containing more than 1,000 images, and carry out a lab-controlled subjective user study. There are several unique and desirable features of the new database as compared to existing ones - It helps us better understand the level of diversity of subject opinions; it provides continuous-valued IAA scores approximately evenly distributed from poor to excellent aesthetics levels; it also allows us to test the effectiveness of various aesthetics features on predicting continuous aesthetics scores. Using the new database as a benchmark, we test more than 1,000 IAA features. The results indicate that existing features are still weak at aesthetics estimation, and the effectiveness of aesthetics features are content dependent. Therefore, understanding and assessing image aesthetics remain a major challenge for future research. The database will be made publicly available. Wentao Liu 0001, Zhou Wang 0001 |
ICIP | 2 |
| 2017 | Quality assessment of multi-view-plus-depth imagesabstractMulti-view-plus-depth (MVD) representation has gained significant attention recently as a means to encode 3D scenes, allowing for intermediate views to be synthesized on-the-fly at the display site through depth-image-based-rendering (DIBR). Automatic quality assessment of MVD images/videos is critical for the optimal design of MVD image/video coding and transmission schemes. Most existing image quality assessment (IQA) and video quality assessment (VQA) methods are applicable only after the DIBR process. Such post-DIBR measures are valuable in assessing the overall system performance, but are difficult to be directly employed in the encoder optimization process in MVD image/video coding. Here we make one of the first attempts to develop a perceptual pre-DIBR IQA approach for MVD images by employing an information content weighted approach that balances between local quality measures of texture and depth images. Experiment results show that the proposed approach achieves competitive performance when compared with state-of-the-art IQA algorithms applied post-DIBR. Jiheng Wang, Shiqi Wang 0001, Kai Zeng 0003, Zhou Wang 0001 |
ICME | 4 |
| 2017 | Quality-of-Experience of Adaptive Video Streaming: Exploring the Space of AdaptationsabstractWith the remarkable growth of adaptive streaming media applications, especially the wide usage of dynamic adaptive streaming schemes over HTTP (DASH), it becomes ever more important to understand the perceptual quality-of-experience (QoE) of end users, who may be constantly experiencing adaptations (switchings) of video bitrate, spatial resolution, and frame-rate from one time segment to another in a scale of a few seconds. This is a sophisticated and challenging problem, for which existing visual studies provide very limited guidance. Here we build a new adaptive streaming video database and carry out a series of subjective experiments to understand human QoE behaviors in this multi-dimensional adaptation space. Our study leads to several useful findings. First, our path-analytic results show that quality deviation introduced by quality adaptation is asymmetric with respect to the adaptation direction (positive or negative), and is further influenced by the intensity of quality change (intensity), dimension of adaptation (type), intrinsic video quality (level), content, and the interactions between them. Second, we find that for the same intensity of quality adaptation, a positive adaptation occurred in the low-quality range has more impact on QoE, suggesting an interesting Weber's law effect; while such phenomenon is reversed for a negative adaptation. Third, existing objective video quality assessment models are very limited in predicting time-varying video quality. Zhengfang Duanmu, Kede Ma, Zhou Wang 0001 |
ACM Multimedia | 3 |
| 2017 | SSIM-Motivated Two-Pass VBR Coding for HEVCabstractWe propose a structural similarity (SSIM)motivated two-pass variable bit rate control algorithm for High Efficiency Video Coding. Given a bit rate budget, the available bits are optimally allocated at group of pictures (GoP), frame, and coding unit (CU) levels by hierarchically constructing a perceptually uniform space with an SSIM-inspired divisive normalization mechanism. The Lagrange multiplier λ, which controls the tradeoff between perceptual distortion and bit rate, is adopted as the GoP level complexity measure. To derive λ, Laplacian distribution-based rate and perceptual distortion models are established after the first pass encoding, and the target bits are dynamically allocated by maintaining a uniform Lagrange multiplier level for each GoP through λ equalization. Within each GoP, rate control is further performed at frame and CU levels based on SSIM-inspired divisive normalization, aiming to transform the prediction residuals into a perceptually uniform space. Experiments show that the proposed scheme achieves high accuracy rate control and superior rate-SSIM performance, which is further verified by subjective visual testing. Shiqi Wang 0001, Abdul Rehman 0001, Kai Zeng 0003, Jiheng Wang, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Waterloo Exploration Database: New Challenges for Image Quality Assessment ModelsabstractThe great content diversity of real-world digital images poses a grand challenge to image quality assessment (IQA) models, which are traditionally designed and validated on a handful of commonly used IQA databases with very limited content variation. To test the generalization capability and to facilitate the wide usage of IQA techniques in real-world applications, we establish a large-scale database named the Waterloo Exploration Database, which in its current state contains 4744 pristine natural images and 94 880 distorted images created from them. Instead of collecting the mean opinion score for each image via subjective testing, which is extremely difficult if not impossible, we present three alternative test criteria to evaluate the performance of IQA models, namely, the pristine/distorted image discriminability test, the listwise ranking consistency test, and the pairwise preference consistency test (P-test). We compare 20 well-known IQA models using the proposed criteria, which not only provide a stronger test in a more challenging testing environment for existing models, but also demonstrate the additional benefits of using the proposed database. For example, in the P-test, even for the best performing no-reference IQA model, more than 6 million failure cases against the model are "discovered" automatically out of over 1 billion test pairs. Furthermore, we discuss how the new database may be exploited using innovative approaches in the future, to reveal the weaknesses of existing IQA models, to provide insights on how to improve the models, and to shed light on how the next-generation IQA models may be developed. The database and codes are made publicly available at: https://ece.uwaterloo.ca/~k29ma/exploration/. Kede Ma, Zhengfang Duanmu, Qingbo Wu 0001, Zhou Wang 0001, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2017 | dipIQ: Blind Image Quality Assessment by Learning-to-Rank Discriminable Image PairsabstractObjective assessment of image quality is fundamentally important in many image processing tasks. In this paper, we focus on learning blind image quality assessment (BIQA) models, which predict the quality of a digital image with no access to its original pristine-quality counterpart as reference. One of the biggest challenges in learning BIQA models is the conflict between the gigantic image space (which is in the dimension of the number of image pixels) and the extremely limited reliable ground truth data for training. Such data are typically collected via subjective testing, which is cumbersome, slow, and expensive. Here, we first show that a vast amount of reliable training data in the form of quality-discriminable image pairs (DIPs) can be obtained automatically at low cost by exploiting large-scale databases with diverse image content. We then learn an opinion-unaware BIQA (OU-BIQA, meaning that no subjective opinions are used for training) model using RankNet, a pairwise learning-to-rank (L2R) algorithm, from millions of DIPs, each associated with a perceptual uncertainty level, leading to a DIP inferred quality (dipIQ) index. Extensive experiments on four benchmark IQA databases demonstrate that dipIQ outperforms the state-of-the-art OU-BIQA models. The robustness of dipIQ is also significantly improved as confirmed by the group MAximum Differentiation competition method. Furthermore, we extend the proposed framework by learning models with ListNet (a listwise L2R algorithm) on quality-discriminable image lists (DIL). The resulting DIL inferred quality index achieves an additional performance gain. Kede Ma, Wentao Liu 0001, Tongliang Liu, Zhou Wang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2017 | Robust Multi-Exposure Image Fusion: A Structural Patch Decomposition ApproachabstractWe propose a simple yet effective structural patch decomposition approach for multi-exposure image fusion (MEF) that is robust to ghosting effect. We decompose an image patch into three conceptually independent components: signal strength, signal structure, and mean intensity. Upon fusing these three components separately, we reconstruct a desired patch and place it back into the fused image. This novel patch decomposition approach benefits MEF in many aspects. First, as opposed to most pixel-wise MEF methods, the proposed algorithm does not require post-processing steps to improve visual quality or to reduce spatial artifacts. Second, it handles RGB color channels jointly, and thus produces fused images with more vivid color appearance. Third and most importantly, the direction of the signal structure component in the patch vector space provides ideal information for ghost removal. It allows us to reliably and efficiently reject inconsistent object motions with respect to a chosen reference image without performing computationally expensive motion estimation. We compare the proposed algorithm with 12 MEF methods on 21 static scenes and 12 deghosting schemes on 19 dynamic scenes (with camera and object motion). Extensive experimental results demonstrate that the proposed algorithm not only outperforms previous MEF algorithms on static scenes but also consistently produces high quality fused images with little ghosting artifacts for dynamic scenes. Moreover, it maintains a lower computational cost compared with the state-of-the-art deghosting schemes. Kede Ma, Hui Li 0029, Hongwei Yong, Zhou Wang 0001, Deyu Meng, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2017 | Unified Blind Quality Assessment of Compressed Natural, Graphic, and Screen Content ImagesabstractDigital images in the real world are created by a variety of means and have diverse properties. A photographical natural scene image (NSI) may exhibit substantially different characteristics from a computer graphic image (CGI) or a screen content image (SCI). This casts major challenges to objective image quality assessment, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. To tackle this problem, we first construct a cross-content-type (CCT) database, which contains 1,320 distorted NSIs, CGIs, and SCIs, compressed using the high efficiency video coding (HEVC) intra coding method and the screen content compression (SCC) extension of HEVC. We then carry out a subjective experiment on the database in a well-controlled laboratory environment. Moreover, we propose a unified content-type adaptive (UCA) blind image quality assessment model that is applicable across content types. A key step in UCA is to incorporate the variations of human perceptual characteristics in viewing different content types through a multi-scale weighting framework. This leads to superior performance on the constructed CCT database. UCA is training-free, implying strong generalizability. To verify this, we test UCA on other databases containing JPEG, MPEG-2, H.264, and HEVC compressed images/videos, and observe that it consistently achieves competitive performance. Xiongkuo Min, Kede Ma, Ke Gu 0001, Guangtao Zhai, Zhou Wang 0001, Weisi Lin |
IEEE Trans. Image Process. | 5 |
| 2017 | Perceptual Depth Quality in Distorted Stereoscopic ImagesabstractSubjective and objective measurement of the perceptual quality of depth information in symmetrically and asymmetrically distorted stereoscopic images is a fundamentally important issue in stereoscopic 3D imaging that has not been deeply investigated. Here, we first carry out a subjective test following the traditional absolute category rating protocol widely used in general image quality assessment research. We find this approach problematic, because monocular cues and the spatial quality of images have strong impact on the depth quality scores given by subjects, making it difficult to single out the actual contributions of stereoscopic cues in depth perception. To overcome this problem, we carry out a novel subjective study where depth effect is synthesized at different depth levels before various types and levels of symmetric and asymmetric distortions are applied. Instead of following the traditional approach, we ask subjects to identify and label depth polarizations, and a depth perception difficulty index (DPDI) is developed based on the percentage of correct and incorrect subject judgements. We find this approach highly effective at quantifying depth perception induced by stereo cues and observe a number of interesting effects regarding image content dependency, distortion-type dependence, and the impact of symmetric versus asymmetric distortions. Furthermore, we propose a novel computational model for DPDI prediction. Our results show that the proposed model, without explicitly identifying image distortion types, leads to highly promising DPDI prediction performance. We believe that these are useful steps toward building a comprehensive understanding on 3D quality-of-experience of stereoscopic images. Jiheng Wang, Shiqi Wang 0001, Kede Ma, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Asymmetrically Compressed Stereoscopic 3D Videos: Quality Assessment and Rate-Distortion Performance EvaluationabstractObjective quality assessment of stereoscopic 3D video is challenging but highly desirable, especially in the application of stereoscopic video compression and transmission, where useful quality models are missing, that can guide the critical decision making steps in the selection of mixed-resolution coding, asymmetric quantization, and pre- and post-processing schemes. Here we first carry out subjective quality assessment experiments on two databases that contain various asymmetrically compressed stereoscopic 3D videos obtained from mixed-resolution coding, asymmetric transform-domain quantization coding, their combinations, and the multiple choices of postprocessing techniques. We compare these asymmetric stereoscopic video coding schemes with symmetric coding methods and verify their potential coding gains. We observe a strong systematic bias when using direct averaging of 2D video quality of both views to predict 3D video quality. We then apply a binocular rivalry inspired model to account for the prediction bias, leading to a significantly improved full reference quality prediction model of stereoscopic videos. The model allows us to quantitatively predict the coding gain of different variations of asymmetric video compression, and provides new insight on the development of high efficiency 3D video coding schemes. Jiheng Wang, Shiqi Wang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Blind Image Quality Assessment Based on Rank-Order Regularized RegressionabstractBlind image quality assessment (BIQA) aims to estimate the subjective quality of a query image without access to the reference image. Existing learning-based methods typically train a regression function by minimizing the average error between subjective opinion scores and model predictions. However, minimizing average error does not necessarily lead to correct quality rank-orders between the test images, which is a highly desirable property of image quality models. In this paper, we propose a novel rank-order regularized regression model to address this problem. The key idea is to introduce a pairwise rank-order constraint into the maximum margin regression framework, aiming to better preserve the correct perceptual preference. To the best of our knowledge, this is the first attempt to incorporate rank-order constraints into margin-based quality regression model. By combing with a new local spatial structure feature, we achieve highly consistent quality prediction with human perception. Experimental results show that the proposed method outperforms many state-of-the-art BIQA metrics on popular publicly available IQA databases (i.e., LIVE-II, TID2013, VCL@FER, LIVEMD, and ChallengeDB). Qingbo Wu 0001, Hongliang Li 0001, Zhou Wang 0001, Fanman Meng, Bing Luo 0003, Wei Li 0110, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2016 | Group MAD Competition? A New Methodology to Compare Objective Image Quality ModelsabstractObjective image quality assessment (IQA) models aim to automatically predict human visual perception of image quality and are of fundamental importance in the field of image processing and computer vision. With an increasing number of IQA models proposed, how to fairly compare their performance becomes a major challenge due to the enormous size of image space and the limited resource for subjective testing. The standard approach in literature is to compute several correlation metrics between subjective mean opinion scores (MOSs) and objective model predictions on several well-known subject-rated databases that contain distorted images generated from a few dozens of source images, which however provide an extremely limited representation of real-world images. Moreover, most IQA models developed on these databases often involve machine learning and/or manual parameter tuning steps to boost their performance, and thus their generalization capabilities are questionable. Here we propose a novel methodology to compare IQA models. We first build a database that contains 4,744 source natural images, together with 94,880 distorted images created from them. We then propose a new mechanism, namely group MAximum Differentiation (gMAD) competition, which automatically selects subsets of image pairs from the database that provide the strongest test to let the IQA models compete with each other. Subjective testing on the selected subsets reveals the relative performance of the IQA models and provides useful insights on potential ways to improve them. We report the gMAD competition results between 16 well-known IQA models, but the framework is extendable, allowing future IQA models to be added into the competition. Kede Ma, Qingbo Wu 0001, Zhou Wang 0001, Zhengfang Duanmu, Hongwei Yong, Hongliang Li 0001, Lei Zhang 0006 |
CVPR | 3 |
| 2016 | Objective quality assessment of tone-mapped videosabstractWith the fast advances in video acquisition, computational imaging, and display technologies, there has been a growing interest in high dynamic range (HDR) videos. Tone mapping operators (TMOs) that convert HDR content to low dynamic range (LDR) ones provide a practically useful solution for the visualization of HDR videos on standard LDR displays, where the user experience highly depends on the performance of the TMOs being used. Without an appropriate perceptual quality measure, different TMOs cannot be compared. Subjective experiments may be a reliable solution, but is time consuming, expensive, and difficult to be embedded into optimization processes. Here we make one of the first attempts to develop an objective quality assessment model for tone-mapped videos that incorporates structural fidelity, statistical naturalness and memory effect. Validation using subject-rated tone-mapped videos show that the proposed method is well-correlated with subjective scores. Hojatollah Yeganeh, Shiqi Wang 0001, Kai Zeng 0003, Mahzar Eisapour, Zhou Wang 0001 |
ICIP | 5 |
| 2016 | Quality-of-experience of streaming video: Interactions between presentation quality and playback stallingabstractNetwork streaming video services have been growing explosively in the past decade, but how to measure and assure the video quality-of-experience (QoE) of end consumers is still an open problem. Poor presentation quality and playback stalling have been identified as the most dominant factors that degrade user QoE. Although both factors have been studied individually, little is known about the interactions between them. In this work, we first construct a streaming video database that contains compressed videos at different distortion levels and with different stalling patterns. We then carry out a subjective test to evaluate the QoE of the videos. The results reveal some interesting dependency between presentation quality and playback stalling. Specifically, playback stalling always causes QoE degradation, but the strength of such degradation depends on the presentation quality when the stalling event occurs. Kai Zeng 0003, Hojatollah Yeganeh, Zhou Wang 0001 |
ICIP | 3 |
| 2016 | Quality-of-experience prediction for streaming videoabstractWith the rapid growth of streaming media applications, there has been a strong demand of objective models that can predict end users' quality-of-experience (QoE) when watching the video being streamed to their display devices. Existing methods typically use bitrate and global statistics of stalling events as the QoE indicators. This is problematic for two reasons. First, using the same bitrate to encode different video content could result in drastically different presentation QoE. Second, the interactions between presentation visual quality and playback stalling are not accounted for. Here we propose a novel QoE prediction approach that takes into consideration the instantaneous quality degradation due to perceptual video presentation impairment, the playback stalling events caused by imperfect network delivery, and the instantaneous interactions between presentation quality and playback stalling. The proposed algorithm demonstrates strong promise when tested using a subject-rated video streaming QoE database. Zhengfang Duanmu, Abdul Rehman 0001, Kai Zeng 0003, Zhou Wang 0001 |
ICME | 4 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 29 |
| 2016 | Adaptive Quantization Parameter Cascading in HEVC Hierarchical CodingabstractThe state-of-the-art High Efficiency Video Coding (HEVC) standard adopts a hierarchical coding structure to improve its coding efficiency. This allows for the quantization parameter cascading (QPC) scheme that assigns quantization parameters (Qps) to different hierarchical layers in order to further improve the rate-distortion (RD) performance. However, only static QPC schemes have been suggested in HEVC test model, which are unable to fully explore the potentials of QPC. In this paper, we propose an adaptive QPC scheme for an HEVC hierarchical structure to code natural video sequences characterized by diversified textures, motions, and encoder configurations. We formulate the adaptive QPC scheme as a non-linear programming problem and solve it in a scientifically sound way with a manageable low computational overhead. The proposed model addresses a generic Qp assignment problem of video coding. Therefore, it also applies to group-of-picture-level, frame-level and coding unit-level Qp assignments. Comprehensive experiments have demonstrated that the proposed QPC scheme is able to adapt quickly to different video contents and coding configurations while achieving noticeable RD performance enhancement over all static and adaptive QPC schemes under comparison as well as HEVC default frame-level rate control. We have also made valuable observations on the distributions of adaptive QPC sets in the videos of different types of contents, which provide useful insights on how to further improve static QPC schemes. Tiesong Zhao, Zhou Wang 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 2 |
| 2016 | The Application of Visual Saliency Models in Objective Image Quality Assessment: A Statistical EvaluationabstractAdvances in image quality assessment have shown the potential added value of including visual attention aspects in its objective assessment. Numerous models of visual saliency are implemented and integrated in different image quality metrics (IQMs), but the gain in reliability of the resulting IQMs varies to a large extent. The causes and the trends of this variation would be highly beneficial for further improvement of IQMs, but are not fully understood. In this paper, an exhaustive statistical evaluation is conducted to justify the added value of computational saliency in objective image quality assessment, using 20 state-of-the-art saliency models and 12 best-known IQMs. Quantitative results show that the difference in predicting human fixations between saliency models is sufficient to yield a significant difference in performance gain when adding these saliency models to IQMs. However, surprisingly, the extent to which an IQM can profit from adding a saliency model does not appear to have direct relevance to how well this saliency model can predict human fixations. Our statistical analysis provides useful guidance for applying saliency models in IQMs, in terms of the effect of saliency model dependence, IQM dependence, and image distortion dependence. The testbed and software are made publicly available to the research community. Wei Zhang 0072, Ali Borji, Zhou Wang 0001, Patrick Le Callet, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Data rate and dynamic range compression of medical images: Which one goes first?abstractAdvances in the field of medical imaging have led to an immense increase in the volume of images being acquired. A fast growing application is to enable physicians to access image data remotely from any viewing device. This casts new challenges for data rate compression. Meanwhile medical images typically have High Dynamic Range (HDR), which needs to be transformed to Low Dynamic Range (LDR) through a so-called “windowing” operation in order for them to be viewed on standard displays to best visualize specific types of content such as tissues or bone structures. This leads to a basic question: Should data compression be performed before windowing or vice versa? Answering this question needs domain knowledge and also requires comparing HDR and LDR images in terms of objective measures, which has only recently become possible. In this paper, we compare the two alternative schemes by using a recently proposed structural fidelity measure. Our study suggests that data compression followed by windowing delivers better performance than the other alternative. Shahrukh Athar, Hojatollah Yeganeh, Zhou Wang 0001 |
ICIP | 3 |
| 2015 | Perceptual evaluation of single image dehazing algorithmsabstractImages captured in outdoor scenes often suffer from poor visibility and color shift due to the presence of haze. Although many algorithms have been proposed to remove the haze, not much effort has been made on quality assessment of dehazed images. In this paper, we first build a database that contains 25 hazy images as well as dehazed images created by eight dehazing algorithms. A subjective user study is then carried out based on the database, from which we have several useful findings. First, considerable agreement between human subjects on the perceived quality of hazy and dehazed images is observed. Second, not a single dehazing algorithm performs the best for all test images. Third, existing objective image quality assessment (IQA) models are very limited in providing proper quality predictions of dehazed images. Kede Ma, Wentao Liu 0001, Zhou Wang 0001 |
ICIP | 3 |
| 2015 | Multi-exposure image fusion: A patch-wise approachabstractWe propose a patch-wise approach for multi-exposure image fusion (MEF). A key step in our approach is to decompose each color image patch into three conceptually independent components: signal strength, signal structure and mean intensity. Upon processing the three components separately based on patch strength and exposedness measures, we uniquely reconstruct a color image patch and place it back into the fused image. Unlike most pixel-wise MEF methods in the literature, the proposed algorithm does not require significant pre/postprocessing steps to improve visual quality or to reduce spatial artifacts. Moreover, the novel patch decomposition allows us to handle RGB color channels jointly and thus produces fused images with more vivid color appearances. Extensive experiments demonstrate the superiority of the proposed algorithm both qualitatively and quantitatively. Kede Ma, Zhou Wang 0001 |
ICIP | 2 |
| 2015 | Perceptual screen content image quality assessment and compressionabstractCompression of screen content has recently emerged as an active research topic due to the increasing demand in many applications such as wireless display and virtual desktop infrastructure. Screen content images (SCIs) exhibit different statistical properties in textual and pictorial regions, and the human visual system (HVS) also behaves differently when viewing the textual and pictorial regions in terms of the extent of visual field. Here we propose a perceptual SCI quality assessment approach that incorporates visual field adaptation and information content weighting. Furthermore, we propose a perceptual coding scheme in an attempt to optimize the HEVC Screen Content Coding encoder. Experimental results show that the proposed quality assessment method not only better predicts the perceptual quality of SCIs, but also leads to an effective way to optimize screen content coding schemes. Shiqi Wang 0001, Ke Gu 0001, Kai Zeng 0003, Zhou Wang 0001, Weisi Lin |
ICIP | 4 |
| 2015 | Quality prediction of asymmetrically compressed stereoscopic videosabstractObjective quality assessment of stereoscopic 3D video is a challenging problem. We carry out a subjective test on symmetrically and asymmetrically compressed stereoscopic videos followed by different levels of low-pass filtering. We observe a strong systematic bias when using direct averaging of 2D video quality of both views to predict 3D video quality. We use a binocular rivalry inspired model to account for the prediction bias, leading to significantly improved quality estimation of stereoscopic videos. The model allows us to quantitatively predict the potential coding gain of asymmetric video compression, and provides new insight on the development of high efficiency 3D video coding schemes. Jiheng Wang, Shiqi Wang 0001, Zhou Wang 0001 |
ICIP | 3 |
| 2015 | A highly efficient method for blind image quality assessmentabstractBlind image quality assessment (BIQA) has attracted a great deal of attention due to the increasing demand in industry and the promising recent progress in academia. To bridge the gap between academic research accomplishment and industrial needs, high efficiency BIQA approaches that allow for real-time computation are highly desirable. In this paper, we propose a novel BIQA method by selecting statistical features extracted from binary patterns of local image structures. This allows us to largely reduce the feature space to eventually one dimension. Somewhat surprisingly, such a single feature, faster-than-real-time approach named local pattern statistics index (LPSI) exhibits impressive generalization ability across different distortion types and achieves competitive quality prediction performance in comparison with state-of-the-art approaches on public databases such as LIVE II and TID2008. Qingbo Wu 0001, Zhou Wang 0001, Hongliang Li 0001 |
ICIP | 2 |
| 2015 | Perceptual quality assessment of high frame rate videoabstractHigh frame rate video has been a hot topic in the past few years driven by a strong need in the entertainment and gaming industry. Nevertheless, progress on perceptual quality assessment of high frame rate video remains limited, making it difficult to evaluate the exact perceptual gain by switching from low to high frame rates. In this work, we first conduct a subjective quality assessment experiment on a database that contains videos compressed at different frame rates, quantization levels and spatial resolutions. We then carry out a series of analysis on the subjective data to investigate the impact of frame rate on perceived video quality and its interplay with quantization level, spatial resolution, spatial complexity, and motion complexity. We observe that perceived video quality generally increases with frame rate, but the gain saturates at high rates. Such gain also depends on the interactions between quantization level, spatial resolution, and spatial and motion complexities. Rasoul Mohammadi Nasiri, Jiheng Wang, Abdul Rehman 0001, Shiqi Wang 0001, Zhou Wang 0001 |
MMSP | 5 |
| 2015 | SSIM-inspired two-pass rate control for High Efficiency Video CodingabstractWe propose a perceptual two-pass rate control scheme for High Efficiency Video Coding (HEVC). The target bits are optimally allocated by hierarchically constructing a perceptual uniform space derived based on an SSIM-inspired divisive normalization mechanism for each group of pictures (GoP), each frame, and each coding unit (CU). The Lagrange multiplier λ, which controls the trade-off between perceptual distortion and bit rate, is adopted as the GoP level complexity measure. After the first pass compression, Laplacian based rate and perceptual distortion models are established to adaptively derive λ, and the target bits are dynamically allocated by maintaining an uniform Lagrange multiplier level through λ equalization. Within each GoP, rate control is further performed at frame and CU levels in the perceptually uniform space. Extensive simulations verify that, the proposed scheme can achieve high accuracy rate control and superior rate-SSIM performance. Shiqi Wang 0001, Abdul Rehman 0001, Kai Zeng 0003, Zhou Wang 0001 |
MMSP | 4 |
| 2015 | Depth perception of distorted stereoscopic imagesabstractHow to measure the perceptual quality of depth information in stereoscopic natural images, especially images undergoing different types of symmetric and asymmetric distortions, is a fundamentally important issue that is not well understood. In this paper, we present two of our recent subjective studies on depth quality. The first one follows the absolute category rating (ACR) protocol that is widely used in general image quality assessment research. We find that traditional approaches such as ACR is problematic in this scenario because monocular cues and the spatial quality of images have strong impacts on the depth quality scores given by subjects, making it difficult to single out the actual contributions of stereoscopic cues in depth perception. To overcome this problem, we carry out the second subjective study where depth effect is synthesized at different depth levels before various types and levels of symmetric and asymmetric distortions are applied. Instead of following the traditional approach, we ask subjects to identify and label depth polarizations, and a Depth Perception Difficulty Index (DPDI) is developed based on the percentage of correct and incorrect subject judgements. We find this approach highly effective at quantifying depth perception induced by stereo cues and observe a number of interesting effects regarding image content dependency, distortion type dependency, and the impacts of symmetric versus asymmetric distortions. We believe that these are useful steps towards building comprehensive 3D quality-of-experience models for stereoscopic images. Jiheng Wang, Shiqi Wang 0001, Zhou Wang 0001 |
MMSP | 3 |
| 2015 | No-Reference Quality Assessment of Contrast-Distorted Images Based on Natural Scene StatisticsabstractContrast distortion is often a determining factor in human perception of image quality, but little investigation has been dedicated to quality assessment of contrast-distorted images without assuming the availability of a perfect-quality reference image. In this letter, we propose a simple but effective method for no-reference quality assessment of contrast distorted images based on the principle of natural scene statistics (NSS). A large scale image database is employed to build NSS models based on moment and entropy features. The quality of a contrast-distorted image is then evaluated based on its unnaturalness characterized by the degree of deviation from the NSS models. Support vector regression (SVR) is employed to predict human mean opinion score (MOS) from multiple NSS features as the input. Experiments based on three publicly available databases demonstrate the promising performance of the proposed method. Yuming Fang 0001, Kede Ma, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001, Guangtao Zhai |
IEEE Signal Process. Lett. | 3 |
| 2015 | A Patch-Structure Representation Method for Quality Assessment of Contrast Changed ImagesabstractContrast is a fundamental attribute of images that plays an important role in human visual perception of image quality. With numerous approaches proposed to enhance image contrast, much less work has been dedicated to automatic quality assessment of contrast changed images. Existing approaches rely on global statistics to estimate contrast quality. Here we propose a novel local patch-based objective quality assessment method using an adaptive representation of local patch structure, which allows us to decompose any image patch into its mean intensity, signal strength and signal structure components and then evaluate their perceptual distortions in different ways. A unique feature that differentiates the proposed method from previous contrast quality models is the capability to produce a local contrast quality map, which predicts local quality variations over space and may be employed to guide contrast enhancement algorithms. Validations based on four publicly available databases show that the proposed patch-based contrast quality index (PCQI) method provides accurate predictions on the human perception of contrast variations. Shiqi Wang 0001, Kede Ma, Hojatollah Yeganeh, Zhou Wang 0001, Weisi Lin |
IEEE Signal Process. Lett. | 4 |
| 2015 | Objective Quality Assessment for Multiexposure Multifocus Image FusionabstractThere has been a growing interest in image fusion technologies, but how to objectively evaluate the quality of fused images has not been fully understood. Here, we propose a method for objective quality assessment of multiexposure multifocus image fusion based on the evaluation of three key factors of fused image quality: 1) contrast preservation; 2) sharpness; and 3) structure preservation. Subjective experiments are conducted to create an image fusion database, based on which, performance evaluation shows that the proposed fusion quality index correlates well with subjective scores, and gives a significant improvement over the existing fusion quality measures. Rania Hassen, Zhou Wang 0001, Magdy M. A. Salama |
IEEE Trans. Image Process. | 2 |
| 2015 | High Dynamic Range Image Compression by Optimizing Tone Mapped Image Quality IndexabstractTone mapping operators (TMOs) aim to compress high dynamic range (HDR) images to low dynamic range (LDR) ones so as to visualize HDR images on standard displays. Most existing TMOs were demonstrated on specific examples without being thoroughly evaluated using well-designed and subject-validated image quality assessment models. A recently proposed tone mapped image quality index (TMQI) made one of the first attempts on objective quality assessment of tone mapped images. Here, we propose a substantially different approach to design TMO. Instead of using any predefined systematic computational structure for tone mapping (such as analytic image transformations and/or explicit contrast/edge enhancement), we directly navigate in the space of all images, searching for the image that optimizes an improved TMQI. In particular, we first improve the two building blocks in TMQI—structural fidelity and statistical naturalness components—leading to a TMQI-II metric. We then propose an iterative algorithm that alternatively improves the structural fidelity and statistical naturalness of the resulting image. Numerical and subjective experiments demonstrate that the proposed algorithm consistently produces better quality tone mapped images even when the initial images of the iteration are created by the most competitive TMOs. Meanwhile, these results also validate the superiority of TMQI-II over TMQI. Kede Ma, Hojatollah Yeganeh, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Perceptual Quality Assessment for Multi-Exposure Image FusionabstractMulti-exposure image fusion (MEF) is considered an effective quality enhancement technique widely adopted in consumer electronics, but little work has been dedicated to the perceptual quality assessment of multi-exposure fused images. In this paper, we first build an MEF database and carry out a subjective user study to evaluate the quality of images generated by different MEF algorithms. There are several useful findings. First, considerable agreement has been observed among human subjects on the quality of MEF images. Second, no single state-of-the-art MEF algorithm produces the best quality for all test images. Third, the existing objective quality models for general image fusion are very limited in predicting perceived quality of MEF images. Motivated by the lack of appropriate objective models, we propose a novel objective image quality assessment (IQA) algorithm for MEF images based on the principle of the structural similarity approach and a novel measure of patch structural consistency. Our experimental results on the subjective database show that the proposed model well correlates with subjective judgments and significantly outperforms the existing IQA models for general image fusion. Finally, we demonstrate the potential application of the proposed model by automatically tuning the parameters of MEF algorithms. Kede Ma, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Objective Quality Assessment for Color-to-Gray Image ConversionabstractColor-to-gray (C2G) image conversion is the process of transforming a color image into a grayscale one. Despite its wide usage in real-world applications, little work has been dedicated to compare the performance of C2G conversion algorithms. Subjective evaluation is reliable but is also inconvenient and time consuming. Here, we make one of the first attempts to develop an objective quality model that automatically predicts the perceived quality of C2G converted images. Inspired by the philosophy of the structural similarity index, we propose a C2G structural similarity (C2G-SSIM) index, which evaluates the luminance, contrast, and structure similarities between the reference color image and the C2G converted image. The three components are then combined depending on image type to yield an overall quality measure. Experimental results show that the proposed C2G-SSIM index has close agreement with subjective rankings and significantly outperforms existing objective quality metrics for C2G conversion. To explore the potentials of C2G-SSIM, we further demonstrate its use in two applications: 1) automatic parameter tuning for C2G conversion algorithms and 2) adaptive fusion of C2G converted images. Kede Ma, Tiesong Zhao, Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Surface Reconstruction in Gradient-Field Domain Using Compressed SensingabstractSurface reconstruction from measurements of spatial gradient is an important computer vision problem with applications in photometric stereo and shape-from-shading. In the case of morphologically complex surfaces observed in the presence of shadowing and transparency artifacts, a relatively large dense gradient measurements may be required for accurate surface reconstruction. Consequently, due to hardware limitations of image acquisition devices, situations are possible in which the available sampling density might not be sufficiently high to allow for recovery of essential surface details. In this paper, the above problem is resolved by means of derivative compressed sensing (DCS). DCS can be viewed as a modification of the classical CS, which is particularly suited for reconstructions involving image/surface gradients. In DCS, a standard CS setting is augmented through incorporation of additional constraints arising from some intrinsic properties of potential vector fields. We demonstrate that using DCS results in reduction in the number of measurements as compared with the standard (dense) sampling, while producing estimates of higher accuracy and smaller variability as compared with CS-based estimates. The results of this study are further supported by a series of numerical experiments. Oleg V. Michailovich, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Quality Prediction of Asymmetrically Distorted Stereoscopic 3D ImagesabstractObjective quality assessment of distorted stereoscopic images is a challenging problem, especially when the distortions in the left and right views are asymmetric. Existing studies suggest that simply averaging the quality of the left and right views well predicts the quality of symmetrically distorted stereoscopic images, but generates substantial prediction bias when applied to asymmetrically distorted stereoscopic images. In this paper, we first build a database that contains both single-view and symmetrically and asymmetrically distorted stereoscopic images. We then carry out a subjective test, where we find that the quality prediction bias of the asymmetrically distorted images could lean toward opposite directions (overestimate or underestimate), depending on the distortion types and levels. Our subjective test also suggests that eye dominance effect does not have strong impact on the visual quality decisions of stereoscopic images. Furthermore, we develop an information content and divisive normalization-based pooling scheme that improves upon structural similarity in estimating the quality of single-view images. Finally, we propose a binocular rivalry-inspired multi-scale model to predict the quality of stereoscopic images from that of the single-view images. Our results show that the proposed model, without explicitly identifying image distortion types, successfully eliminates the prediction bias, leading to significantly improved quality prediction of the stereoscopic images. Jiheng Wang, Abdul Rehman 0001, Kai Zeng 0003, Shiqi Wang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | Objective Quality Assessment of Interpolated Natural ImagesabstractImage interpolation techniques that create high-resolution images from low-resolution (LR) images are widely used in real world applications, but how to evaluate the quality of interpolated images is not a well-resolved issue. Subjective assessment methods are useful and reliable, but are also slow and expensive. Here, we propose an objective method to assess the quality of an interpolated natural image using the available LR image as a reference. Our method adopts a natural scene statistics (NSS) framework, where image quality degradation is gauged by the deviation of its statistical features from the NSS models trained upon high-quality natural images. Two distortion measures are proposed, namely, interpolated natural image distortion (IND) and weighted IND. Validations by subjective tests show that the proposed approach performs statistically equivalent or sometimes better than an average human subject. Moreover, we demonstrate the potential application of the proposed method in parameter tuning of image interpolation algorithms. Hojatollah Yeganeh, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Adaptive windowing for optimal visualization of medical images based on normalized information distanceabstractThere has been a growing recent interest of applying Kol-mogorov complexity and its related normalized information distance (NID) measures in real-world problems, but their application in the field of medical image processing remains limited. In this work we attempt to incorporate NID in the design of windowing operators for optimal visualization of high dynamic range (HDR) medical images, where predefined intensity interval of interest needs to be mapped to match the low dynamic range (LDR) of standard displays. By approximating NID using a Shannon entropy based method, we are able to optimize parametric windowing operators to maximize the information similarity between the HDR image and the LDR image after mapping. Experimental results demonstrate promising performance of the proposed approach. Nima Nikvand, Hojatollah Yeganeh, Zhou Wang 0001 |
ICASSP | 3 |
| 2014 | High dynamic range image tone mapping by optimizing tone mapped image quality indexabstractAn active research topic in recent years is to design tone mapping operators (TMOs) that convert high dynamic range (H-DR) to low dynamic range (LDR) images, so that HDR images can be visualized on standard displays. Nevertheless, most existing work has been done in the absence of a well-established and subject-validated image quality assessment (IQA) model, without which fair comparisons and further improvement are difficult. Recently, a tone mapped image quality index (TMQI) was proposed, which has shown to have good correlation with subjective evaluations of tone mapped images. Here we propose a substantially different approach to design TMO, where instead of using any pre-defined systematic computational structure (such as image transformation or contrast/edge enhancement) for tone mapping, we navigate in the space of all images, searching for the image that optimizes TMQI. The navigation involves an iterative process that alternately improves the structural fidelity and statistical naturalness of the resulting image, which are the two fundamental building blocks in TMQI. Experiments demonstrate the superior performance of the proposed method. Kede Ma, Hojatollah Yeganeh, Kai Zeng 0003, Zhou Wang 0001 |
ICME | 4 |
| 2014 | Quality prediction of asymmetrically distorted stereoscopic images from single viewsabstractObjective quality assessment of distorted stereoscopic images is a challenging problem. Existing studies suggest that simply averaging the quality of the left- and right-views well predicts the quality of symmetrically distorted stereoscopic images, but generates substantial prediction bias when applied to asymmetrically distorted stereoscopic images. In this study, we first carry out a subjective test, where we find that the prediction bias could lean towards opposite directions, largely depending on the distortion types. We then develop an information-content and divisive normalization based pooling scheme that improves upon SSIM in estimating the quality of single view images. Finally, we propose a binocular rivalry inspired model to predict the quality of stereoscopic images based on that of the single view images. Our results show that the proposed model, without explicitly identifying image distortion types, successfully eliminates the prediction bias, leading to significantly improved quality prediction of stereoscopic images. Jiheng Wang, Kai Zeng 0003, Zhou Wang 0001 |
ICME | 3 |
| 2014 | Video Saliency Incorporating Spatiotemporal Cues and Uncertainty WeightingabstractWe propose a novel algorithm to detect visual saliency from video signals by combining both spatial and temporal information and statistical uncertainty measures. The main novelty of the proposed method is twofold. First, separate spatial and temporal saliency maps are generated, where the computation of temporal saliency incorporates a recent psychological study of human visual speed perception. Second, the spatial and temporal saliency maps are merged into one using a spatiotemporally adaptive entropy-based uncertainty weighting approach. The spatial uncertainty weighing incorporates the characteristics of proximity and continuity of spatial saliency, while the temporal uncertainty weighting takes into account the variations of background motion and local contrast. Experimental results show that the proposed spatiotemporal uncertainty weighting algorithm significantly outperforms state-of-the-art video saliency detection models. Yuming Fang 0001, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | High dynamic range image tone mapping by maximizing a structural fidelity measureabstractTone mapping operators (TMOs) that convert high dynamic range (HDR) images to standard low dynamic range (LDR) images are highly desirable for the visualization of these images on standard displays. Although many existing TMOs produce visually appealing images, it is until recently validated objective measures that can assess their quality have been proposed. Without such objective measures, the design of traditional TMOs can only be based on intuitive ideas, lacking clear goals for further improvement. In this paper, we propose a substantially different tone mapping approach, where instead of explicitly designing a new computational structure for TMO, we search in the space of images to find better quality images in terms of a recent objective measure that can assess the structural fidelity between two images of different dynamic ranges. Specifically, starting from any initial image, the proposed algorithm moves the image along the gradient ascent direction and stops until it converges to a maximal point. Our experiments show that the proposed algorithm reliably produces better quality images upon a number of state-of-the-art TMOs. Hojatollah Yeganeh, Zhou Wang 0001 |
ICASSP | 2 |
| 2013 | Video saliency incorporating spatiotemporal cues and uncertainty weightingabstractWe propose a method to detect visual saliency from video signals by combing both spatial and temporal information and statistical uncertainty measures. The main novelty of the proposed method is twofold. First, separate spatial and temporal saliency maps are generated, where the computation of temporal saliency incorporates a recent psychological study of human visual speed perception, where the perceptual prior probability distribution of the speed of motion is measured through a series of psychovisual experiments. Second, the spatial and temporal saliency maps are merged into one using a spatiotemporally adaptive entropy-based uncertainty weighting approach. Experimental results show that the proposed method significantly outperforms state-of-the-art video saliency detection models. Yuming Fang 0001, Zhou Wang 0001, Weisi Lin |
ICME | 2 |
| 2013 | Image classification based on complex wavelet structural similarity
Abdul Rehman 0001, Yang Gao 0022, Jiheng Wang, Zhou Wang 0001 |
Signal Process. Image Commun. | 4 |
| 2013 | Image Sharpness Assessment Based on Local Phase CoherenceabstractSharpness is an important determinant in visual assessment of image quality. The human visual system is able to effortlessly detect blur and evaluate sharpness of visual images, but the underlying mechanism is not fully understood. Existing blur/sharpness evaluation algorithms are mostly based on edge width, local gradient, or energy reduction of global/local high frequency content. Here we understand the subject from a different perspective, where sharpness is identified as strong local phase coherence (LPC) near distinctive image features evaluated in the complex wavelet transform domain. Previous LPC computation is restricted to be applied to complex coefficients spread in three consecutive dyadic scales in the scale-space. Here we propose a flexible framework that allows for LPC computation in arbitrary fractional scales. We then develop a new sharpness assessment algorithm without referencing the original image. We use four subject-rated publicly available image databases to test the proposed algorithm, which demonstrates competitive performance when compared with state-of-the-art algorithms. Rania Hassen, Zhou Wang 0001, Magdy M. A. Salama |
IEEE Trans. Image Process. | 2 |
| 2013 | Perceptual Video Coding Based on SSIM-Inspired Divisive NormalizationabstractWe propose a perceptual video coding framework based on the divisive normalization scheme, which is found to be an effective approach to model the perceptual sensitivity of biological vision, but has not been fully exploited in the context of video coding. At the macroblock (MB) level, we derive the normalization factors based on the structural similarity (SSIM) index as an attempt to transform the discrete cosine transform domain frame residuals to a perceptually uniform space. We further develop an MB level perceptual mode selection scheme and a frame level global quantization matrix optimization method. Extensive simulations and subjective tests verify that, compared with the H.264/AVC video coding standard, the proposed method can achieve significant gain in terms of rate-SSIM performance and provide better visual quality. Shiqi Wang 0001, Abdul Rehman 0001, Zhou Wang 0001, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Objective Quality Assessment of Tone-Mapped ImagesabstractTone-mapping operators (TMOs) that convert high dynamic range (HDR) to low dynamic range (LDR) images provide practically useful tools for the visualization of HDR images on standard LDR displays. Different TMOs create different tone-mapped images, and a natural question is which one has the best quality. Without an appropriate quality measure, different TMOs cannot be compared, and further improvement is directionless. Subjective rating may be a reliable evaluation method, but it is expensive and time consuming, and more importantly, is difficult to be embedded into optimization frameworks. Here we propose an objective quality assessment algorithm for tone-mapped images by combining: 1) a multiscale signal fidelity measure on the basis of a modified structural similarity index and 2) a naturalness measure on the basis of intensity statistics of natural images. Validations using independent subject-rated image databases show good correlations between subjective ranking score and the proposed tone-mapped image quality index (TMQI). Furthermore, we demonstrate the extended applications of TMQI using two examples-parameter tuning for TMOs and adaptive fusion of multiple tone-mapped images. Hojatollah Yeganeh, Zhou Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Multiview Coding Mode Decision With Hybrid Optimal Stopping ModelabstractIn a generic decision process, optimal stopping theory aims to achieve a good tradeoff between decision performance and time consumed, with the advantages of theoretical decision-making and predictable decision performance. In this paper, optimal stopping theory is employed to develop an effective hybrid model for the mode decision problem, which aims to theoretically achieve a good tradeoff between the two interrelated measurements in mode decision, as computational complexity reduction and rate-distortion degradation. The proposed hybrid model is implemented and examined with a multiview encoder. To support the model and further promote coding performance, the multiview coding mode characteristics, including predicted mode probability and estimated coding time, are jointly investigated with inter-view correlations. Exhaustive experimental results with a wide range of video resolutions reveal the efficiency and robustness of our method, with high decision accuracy, negligible computational overhead, and almost intact rate-distortion performance compared to the original encoder. Tiesong Zhao, Sam Kwong, Hanli Wang, Zhou Wang 0001, Zhaoqing Pan, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 4 |
| 2012 | Gradient-based surface reconstruction using compressed sensingabstractSurface reconstruction from measurements of spatial gradient is an important computer vision problem with applications in photometric stereo and shape-from-shading. In the case of morphologically complex surfaces observed in the presence of shadowing and transparency artifacts, a relatively large number of gradient measurements may be required for accurate surface reconstruction. Consequently, due to hardware limitations of image acquisition devices, situations are possible in which the available sampling density might not be sufficiently high to allow for recovery of essential surface details. In this paper, the above problem is resolved by means of derivative compressed sensing (DCS). DCS can be viewed as a modification of the classical compressed sensing (CS), which is particularly suited for reconstructions involving image/surface gradients. We demonstrate that using DCS results in substantial data savings as compared to the standard (dense) sampling, while producing estimates of higher accuracy and smaller variability, as compared to CS-base estimates. The results of this study are further supported by a series of numerical experiments. Oleg V. Michailovich, Zhou Wang 0001 |
ICIP | 3 |
| 2012 | Objective quality assessment for image super-resolution: A natural scene statistics approachabstractThere has been an increasing number of image super-resolution (SR) algorithms proposed recently to create images with higher spatial resolution from low-resolution (LR) images. Nevertheless, how to evaluate the performance of such SR and interpolation algorithms remains an open problem. Subjective assessment methods are useful and reliable, but are expensive, time-consuming, and difficult to be embedded into the design and optimization procedures of SR and interpolation algorithms. Here we make one of the first attempts to develop an objective quality assessment method of a given resolution-enhanced image using the available LR image as a reference. Our algorithm follows the philosophy behind the natural scene statistics (NSS) approach. Specifically, we build statistical models of frequency energy falloff and spatial continuity based on high quality natural images and use the departures from such models to quantify image quality degradations. Subjective experiments have been carried out that verify the effectiveness of the proposed approach. Hojatollah Yeganeh, Zhou Wang 0001 |
ICIP | 3 |
| 2012 | 3D-SSIM for video quality assessmentabstractEffective and efficient objective video quality assessment (VQA) methods are highly desirable in modern visual communication systems for performance evaluation, quality control and resource allocation purposes. Simple VQA algorithms may be developed by direct extensions of still image quality assessment (IQA) approaches on a frame-by-frame basis. Advanced VQA methods take into account the temporal correlation and motion information contained in video signals but often lead to significantly increased computational complexity. Here we use a different approach to examine a video signal by considering it as a three-dimensional (3D) volume image. Specifically, we propose a 3D structural similarity (3D-SSIM) approach, which first creates a 3D quality map by applying SSIM evaluations within local 3D blocks, and then use local information content and local distortion based weighting methods to pool the quality map into a single quality measure. The resulting 3D-SSIM algorithm is computationally efficient and demonstrates highly competitive performance in comparison with state-of-the-art VQA algorithms when tested using four publicly available video quality databases. Kai Zeng 0003, Zhou Wang 0001 |
ICIP | 2 |
| 2012 | SSIM-Inspired Perceptual Video Coding for HEVCabstractRecent advances in video capturing and display technologies, along with the exponentially increasing demand of video services, challenge the video coding research community to design new algorithms able to significantly improve the compression performance of the current H.264/AVC standard. This target is currently gaining evidence with the standardization activities in the High Efficiency Video Coding (HEVC) project. The distortion models used in HEVC are mean squared error (MSE) and sum of absolute difference (SAD). However, they are widely criticized for not correlating well with perceptual image quality. The structural similarity (SSIM) index has been found to be a good indicator of perceived image quality. Meanwhile, it is computationally simple compared with other state-of-the-art perceptual quality measures and has a number of desirable mathematical properties for optimization tasks. We propose a perceptual video coding method to improve upon the current HEVC based on an SSIM-inspired divisive normalization scheme as an attempt to transform the DCT domain frame prediction residuals to a perceptually uniform space before encoding. Based on the residual divisive normalization process, we define a distortion model for mode selection and show that such a divisive normalization strategy largely simplifies the subsequent perceptual rate-distortion optimization procedure. We further adjust the divisive normalization factors based on local content of the video frame. Experiments show that the proposed scheme can achieve significant gain in terms of rate-SSIM performance when compared with HEVC. Abdul Rehman 0001, Zhou Wang 0001 |
ICME | 2 |
| 2012 | SSIM-Motivated Rate-Distortion Optimization for Video CodingabstractWe propose a rate-distortion optimization (RDO) scheme based on the structural similarity (SSIM) index, which was found to be a better indicator of perceived image quality than mean-squared error, but has not been fully exploited in the context of image and video coding. At the frame level, an adaptive Lagrange multiplier selection method is proposed based on a novel reduced-reference statistical SSIM estimation algorithm and a rate model that combines the side information with the entropy of the transformed residuals. At the macroblock level, the Lagrange multiplier is further adjusted based on an information theoretical approach that takes into account both the motion information content and perceptual uncertainty of visual speed perception. Finally, the mode for H.264/AVC coding is selected by the SSIM index and the adjusted Lagrange multiplier. Extensive experiments show that the proposed scheme can achieve significantly better rate-SSIM performance and provide better visual quality than conventional RDO coding schemes. Shiqi Wang 0001, Abdul Rehman 0001, Zhou Wang 0001, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | On the Mathematical Properties of the Structural Similarity IndexabstractSince its introduction in 2004, the structural similarity (SSIM) index has gained widespread popularity as a tool to assess the quality of images and to evaluate the performance of image processing algorithms and systems. There has been also a growing interest of using SSIM as an objective function in optimization problems in a variety of image processing applications. One major issue that could strongly impede the progress of such efforts is the lack of understanding of the mathematical properties of the SSIM measure. For example, some highly desirable properties such as convexity and triangular inequality that are possessed by the mean squared error may not hold. In this paper, we first construct a series of normalized and generalized (vector-valued) metrics based on the important ingredients of SSIM. We then show that such modified measures are valid distance metrics and have many useful properties, among which the most significant ones include quasi-convexity, a region of convexity around the minimizer, and distance preservation under orthogonal or unitary transformations. The groundwork laid here extends the potentials of SSIM in both theoretical development and practical applications. Dominique Brunet, Edward R. Vrscay, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Reduced-Reference Image Quality Assessment by Structural Similarity EstimationabstractReduced-reference image quality assessment (RR-IQA) provides a practical solution for automatic image quality evaluations in various applications where only partial information about the original reference image is accessible. Here we propose an RR-IQA method by estimating the structural similarity (SSIM) index, which is a widely used full-reference (FR) image quality measure shown to be a good indicator of perceptual image quality. Specifically, we extract statistical features from a multi-scale, multi-orientation divisive normalization transform and develop a distortion measure by following the philosophy in the construction of SSIM. We found an interesting linear relationship between the FR SSIM measure and our RR estimate when the image distortion type is fixed. A regression-bydiscretization method is then applied to normalize our measure across image distortion types. We use six publiclyavailable subject-rated databases to test the proposed RR-SSIM method, which shows strong correlations with both SSIM and subjective quality evaluations. Finally, we introduce the novel idea of partially repairing an image using RR features and use deblurring as an example to demonstrate its application. Abdul Rehman 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Image Deblurring Using Derivative Compressed Sensing for Optical Imaging ApplicationabstractThe problem of reconstruction of digital images from their blurred and noisy measurements is unarguably one of the central problems in imaging sciences. Despite its ill-posed nature, this problem can often be solved in a unique and stable manner, provided appropriate assumptions on the nature of the images to be recovered. In this paper, however, a more challenging setting is considered, in which accurate knowledge of the blurring operator is lacking, thereby transforming the reconstruction problem at hand into a problem of blind deconvolution. As a specific application, the current presentation focuses on reconstruction of short-exposure optical images measured through atmospheric turbulence. The latter is known to give rise to random aberrations in the optical wavefront, which are in turn translated into random variations of the point spread function of the optical system in use. A standard way to track such variations involves using adaptive optics. Thus, for example, the Shack-Hartmann interferometer provides measurements of the optical wavefront through sensing its partial derivatives. In such a case, the accuracy of wavefront reconstruction is proportional to the number of lenslets used by the interferometer and, hence, to its complexity. Accordingly, in this paper, we show how to minimize the above complexity through reducing the number of the lenslets while compensating for undersampling artifacts by means of derivative compressed sensing. Additionally, we provide empirical proof that the above simplification and its associated solution scheme result in image reconstructions, whose quality is comparable to the reconstructions obtained using conventional (dense) measurements of the optical wavefront. Oleg V. Michailovich, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Polyview Fusion: A Strategy to Enhance Video-Denoising AlgorithmsabstractWe propose a simple but effective strategy that aims to enhance the performance of existing video denoising algorithms, i.e., polyview fusion (PVF). The idea is to denoise the noisy video as a 3-D volume using a given base 2-D denoising algorithm but applied from multiple views (front, top, and side views). A fusion algorithm is then designed to merge the resulting multiple denoised videos into one, so that the visual quality of the fused video is improved. Extensive tests using a variety of base video-denoising algorithms show that the proposed PVF method leads to surprisingly significant and consistent gain in terms of both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) performance, particularly at high noise levels, where the improvement over state-of-the-art denoising algorithms is often more than 2 dB in PSNR. Kai Zeng 0003, Zhou Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | SSIM-inspired image denoising using sparse representationsabstractPerceptual image quality assessment (IQA) and sparse signal representation have recently emerged as high-impact research topics in the field of image processing. Here we make one of the first attempts to incorporate the structural similarity (SSIM) index, a promising IQA measure, into the framework of optimal sparse signal representation and approximation. In particular, we introduce a novel image denoising scheme where a modified orthogonal matching pursuit algorithm is proposed for finding the best sparse coefficient vector in maximum-SSIM sense for a given set of linearly independent atoms. Furthermore, a gradient descent algorithm is developed to achieve SSIM-optimal compromise in combining the input and sparse dictionary reconstructed images. Our experimental results show that the proposed method achieves better SSIM performance and provide better visual quality than least square optimal denoising methods. Abdul Rehman 0001, Zhou Wang 0001, Dominique Brunet, Edward R. Vrscay |
ICASSP | 2 |
| 2011 | Rate-SSIM optimization for video codingabstractThe structural similarity (SSIM) index has been found to be a good indicator of perceived image quality. In this paper, we propose a rate-SSIM optimization scheme for mode se lection in H.264/AVC video coding. To derive the Lagrange multiplier based on the properties of input sequences, a novel reduced-reference statistical SSIM model and a source-side information combined rate model are established. The proposed method is fully standard-compatible. Experimental results demonstrate that, compared with conventional rate distortion optimization coding schemes, the proposed scheme can achieve better rate-SSIM performance and provide better visual quality. Shiqi Wang 0001, Abdul Rehman 0001, Zhou Wang 0001, Siwei Ma 0001, Wen Gao 0001 |
ICASSP | 3 |
| 2011 | CW-SSIM based image classificationabstractComplex wavelet structural similarity (CW-SSIM) index has been proposed as a promising image similarity measure that is robust to small geometric distortions such as translation, scaling and rotation of images, but how to make the best use of it in image classification problems has not been deeply investigated. In this paper, we propose a novel “feature-extraction free” image classification algorithm based on CW-SSIM and use handwritten digit recognition as an example to demonstrate it. First, a CW-SSIM based unsupervised clustering method is used to divide the training images into clusters and to pick a representative image for each cluster. A supervised learning method based on support vector machines is then employed to maximize the classification accuracy based on CW-SSIM values between an input image and the representative images. Our experiments show that such a conceptually simple image classification method, which does not involve any registration, intensity normalization or sophisticated feature extraction processes, and does not rely on any modeling of the image patterns or distortion processes, achieves competitive performance with reduced computational complexity. Yang Gao 0022, Abdul Rehman 0001, Zhou Wang 0001 |
ICIP | 3 |
| 2011 | SSIM-based non-local means image denoisingabstractPerceptually inspired image processing has been an emerging field of study in recent years. Here we make one of the first efforts to incorporate the structural similarity (SSIM) index, a successful perceptual image quality assessment measure, into the framework of non-local means (NLM) image denoising, which is a state-of-the-art method that delivers superior desnoising performance. Specifically, a denoised image patch is obtained by weighted averaging of neighboring patches, where the similarity between patches as well as the weights assigned to the patches are determined based on an estimation of SSIM. A two-stage approach is proposed for robust SSIM estimation in the presence of noise. Moreover, motivated by the ideas behind SSIM, we adjust the contrast and mean of each patch before feeding it into the weighted averaging process. Our experimental results show that the proposed SSIM-based NLM algorithm achieves better SSIM and PSNR performance and provides better visual quality than least square based NLM method. Abdul Rehman 0001, Zhou Wang 0001 |
ICIP | 2 |
| 2011 | SSIM-inspired divisive normalization for perceptual video codingabstractWe propose a perceptual video coding framework based on an SSIM-inspired divisive normalization scheme as an attempt to transform the DCT domain frame prediction residuals to a perceptually uniform space before coding. Based on the residual divisive normalization process, we define a distortion model for mode selection and show that such a divisive normalization strategy largely simplifies the subsequent perceptual rate-distortion optimization procedure. Experiments demonstrate that the proposed scheme can achieve significant gain in terms of rate-SSIM performance in comparison with H.264/AVC. Shiqi Wang 0001, Abdul Rehman 0001, Zhou Wang 0001, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 3 |
| 2011 | Information Content Weighting for Perceptual Image Quality AssessmentabstractMany state-of-the-art perceptual image quality assessment (IQA) algorithms share a common two-stage structure: local quality/distortion measurement followed by pooling. While significant progress has been made in measuring local image quality/distortion, the pooling stage is often done in ad-hoc ways, lacking theoretical principles and reliable computational models. This paper aims to test the hypothesis that when viewing natural images, the optimal perceptual weights for pooling should be proportional to local information content, which can be estimated in units of bit using advanced statistical models of natural images. Our extensive studies based upon six publicly-available subject-rated image databases concluded with three useful findings. First, information content weighting leads to consistent improvement in the performance of IQA algorithms. Second, surprisingly, with information content weighting, even the widely criticized peak signal-to-noise-ratio can be converted to a competitive perceptual quality measure when compared with state-of-the-art algorithms. Third, the best overall performance is achieved by combining information content weighting with multiscale structural similarity measures. Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2010 | No-reference image sharpness assessment based on local phase coherence measurementabstractSharpness is one of the most determining factors in the perceptual assessment of image quality. Objective image sharpness measures may play important roles in the design and optimization of visual perception-based auto-focus systems and image enhancement, restoration and compression algorithms. Here we propose a new sharpness measure where sharpness is identified as strong local phase coherence evaluated in the complex wavelet transform domain. Our test using the LIVE blur database shows that the proposed algorithm correlates well with subjective quality evaluations. An additional advantage of our approach is that other image distortions such as compression, median filtering and noise contamination that may affect perceptual sharpness can also be detected. Rania Hassen, Zhou Wang 0001, Magdy M. A. Salama |
ICASSP | 2 |
| 2010 | Temporal motion smoothness measurement for reduced-reference video quality assessmentabstractReduced-reference (RR) video quality measures aim to predict the perceptual quality of distorted video signals using only partial information about the reference video. Existing RR video quality assessment models are mostly designed and/or trained for specific applications such as lossy compression, where the detectable distortion types are often fixed and limited. Here we propose a novel approach that measures temporal motion smoothness of a video sequence by examining the temporal variations of local phase structures in the complex wavelet transform domain. We show that the proposed measure can detect a wide range of well-known practical distortions, including noise contamination, blurring, line or frame jittering, and frame dropping. In addition, the proposed algorithm does not require a costly motion estimation process and has a low RR data rate, making it much easier to be adopted in real-world visual communication applications. Kai Zeng 0003, Zhou Wang 0001 |
ICASSP | 2 |
| 2010 | Generic image similarity based on Kolmogorov complexityabstractImage similarity measurement is a fundamental and common issue in a broad range of problems in image processing, compression, communication, recognition and retrieval. Existing image similarity measures are limited to restricted application environments. The theory of Kolmogorov complexity and the related normalized information distance (NID) measure provide an attractive theoretic framework for generic image similarity that is applicable to any scenario. While this is appealing, the difficulty lies in the implementation due to the non-computable nature of Kolmogorov complexity. In this paper, we propose a practical framework to approximate NID, where the key is to find the shortest program within a set of potential transformations that convert one image to another and vice versa. As one of the initial attempts in this new and promising research direction, our preliminary experimental work demonstrates the wider applicability of the proposed approach than existing methods. Nima Nikvand, Zhou Wang 0001 |
ICIP | 2 |
| 2010 | Reduced-reference SSIM estimationabstractThe structural similarity (SSIM) index has been shown to be a good perceptual image quality predictor. In many real-world applications such as network visual communications, however, SSIM is not applicable because its computation requires full access to the original image. Here we propose a reduced-reference approach that estimates SSIM with only partial information about the original image. Specifically, we extract statistical features from a multi-scale, multi-orientation divisive normalization transform and develop a distortion measure by following the philosophy analogous to that in the construction of SSIM. We found an interesting linear relationship between our reduced-reference SSIM estimate and full-reference SSIM when the image distortion type is fixed. A regression-by-discretization method is then applied to normalize our measure between image distortion types. We use the LIVE database to test the proposed distortion measure, which shows strong correlations with both SSIM and subjective evaluations. We also demonstrate how our reduced-reference features may be employed to partially repair a distorted image. Abdul Rehman 0001, Zhou Wang 0001 |
ICIP | 2 |
| 2010 | Objective assessment of tone mapping algorithmsabstractThere has been a growing interest in recent years to develop tone mapping algorithms that can convert high dynamic range (HDR) to low dynamic range (LDR) images, so that they can be visualized on standard displays. With a number of tone mapping algorithms proposed, a natural question is which one gives the best performance. Although subjective assessment methods provide useful references, they are expensive and time-consuming, and are difficult to be embedded into the design stage of tone mapping algorithms for optimization and parameter tuning purposes. This paper focuses on objective assessment of tone mapping operators. Inspired by the success of the structural similarity index method for image quality assessment, we propose a new objective assessment algorithm that creates multi-scale similarity maps between HDR and LDR images. Our experiments show that the proposed method correlates well with subjective rankings of existing tone mapping operators. Furthermore, we demonstrate how the proposed algorithm can be employed in an existing tone mapping algorithm for optimal parameter tuning. Hojatollah Yeganeh, Zhou Wang 0001 |
ICIP | 2 |
| 2010 | Quality-aware video based on robust embedding of intra- and inter-frame reduced-reference featuresabstractWith the rapid development of network visual communications, there is an urgent need of effective and efficient video quality assessment (VQA) methods for quality control and resource allocation purposes. In this paper, a spatial and temporal reduced-reference (RR) VQA measure is combined with a robust video watermarking approach, leading to a quality-aware video (QAV) system. At the sender side, both intra- and inter-frame RR features are calculated from the original video based on statistical models of natural video. This is followed by error control coding to improve robustness. The encoded features are then embedded invisibly into the same video signal using a robust angle quantization index modulation based watermarking method in 3D discrete cosine transform domain. At the receiver side, the RR features are extracted and decoded from the distorted video and employed to predict the perceptual degradation of the video signal. Experimental results demonstrate the applicability of the proposed approach to a wide range of distortion types and levels. Kai Zeng 0003, Zhou Wang 0001 |
ICIP | 2 |
| 2010 | CW-SSIM kernel based random forest for image classificationabstractComplex wavelet structural similarity (CW-SSIM) index has been proposed as a powerful image similarity metric that is robust to translation, scaling and rotation of images, but how to employ it in image classification applications has not been deeply investigated. In this paper, we incorporate CW-SSIM as a kernel function into a random forest learning algorithm. This leads to a novel image classification approach that does not require a feature extraction or dimension reduction stage at the front end. We use hand-written digit recognition as an example to demonstrate our algorithm. We compare the performance of the proposed approach with random forest learning based on other kernels, including the widely adopted Gaussian and the inner product kernels. Empirical evidences show that the proposed method is superior in its classification power. We also compared our proposed approach with the direct random forest method without kernel and the popular kernel-learning method support vector machine. Our test results based on both simulated and realworld data suggest that the proposed approach works superior to traditional methods without the feature selection procedure. Guangzhe Fan, Zhou Wang 0001, Jiheng Wang |
VCIP | 2 |
| 2009 | Multi-sensor image registration based-on local phase coherenceabstractThe major challenges in automatic multi-sensor image registration are the inconsistency in intensity or contrast patterns, and the existence of partial or missing information between images. Here we propose a novel image registration method based on local phase coherence features, which are insensitive to changes in intensity or contrast. Furthermore, a new objective function based on weighted mutual information is proposed, where less weight is given to the objects that have no correspondence between images. The proposed method has been tested on both synthetic and medical images and evaluated based on registration accuracy. Our experiments demonstrate good performance of the proposed approach with missing or partial data, with significant changes in contrast, and with the presence of noise. Rania Hassen, Zhou Wang 0001, Magdy M. A. Salama |
ICIP | 2 |
| 2009 | Quantifying color image distortions based on adaptive spatio-chromatic signal decompositionsabstractWe describe a framework for quantifying color image distortion based on an adaptive signal decomposition. Specifically, local blocks of the image error are decomposed using a set of spatio-chromatic basis functions that are adapted to the spatial and color structure of the original image. The adaptive functions are chosen to isolate specific distortions such as luminance, hue, and saturation changes. These adaptive basis functions are used to augment a generic orthonormal basis, and the overall distortion is computed from the weighted sum of the coefficients of the resulting overcomplete decomposition, with smaller weights chosen for the adaptive terms. A set of preliminary experiments show that the proposed distortion measure is consistent with human perception of color images subjected to a variety of different common distortions. The framework may be easily extended to include any form of continuous spatio-chromatic distortion. Umesh Rajashekar, Zhou Wang 0001, Eero P. Simoncelli |
ICIP | 2 |
| 2009 | Complex Wavelet Structural Similarity: A New Image Similarity IndexabstractWe introduce a new measure of image similarity called the complex wavelet structural similarity (CW-SSIM) index and show its applicability as a general purpose image similarity index. The key idea behind CW-SSIM is that certain image distortions lead to consistent phase changes in the local wavelet coefficients, and that a consistent phase shift of the coefficients does not change the structural content of the image. By conducting four case studies, we have demonstrated the superiority of the CW-SSIM index against other indices (e.g., Dice, Hausdorff distance) commonly used for assessing the similarity of a given pair of images. In addition, we show that the CW-SSIM index has a number of advantages. It is robust to small rotations and translations. It provides useful comparisons even without a preprocessing image registration step, which is essential for other indices. Moreover, it is computationally less expensive. Mehul P. Sampat, Zhou Wang 0001, Shalini Gupta, Alan C. Bovik, Mia K. Markey |
IEEE Trans. Image Process. | 2 |
| 2008 | Contextually adaptive signal representation using conditional principal component analysisabstractThe conventional method of generating a basis that is optimally adapted (in MSE) for representation of an ensemble of signals is Principal Component Analysis (PCA). A more ambitious modern goal is the construction of bases that are adapted to individual signal instances. Here we develop a new framework for instance-adaptive signal representation by exploiting the fact that many real-world signals exhibit local self-similarity. Specifically, we decompose the signal into multiscale subbands, and then represent local blocks of each subband using basis functions that are linearly derived from the surrounding context. The linear mappings that generate these basis functions are learned sequentially, with each one optimized to account for as much variance as possible in the local blocks. We apply this methodology to learning a coarse-to-fine representation of images within a multi-scale basis, demonstrating that the adaptive basis can account for significantly more variance than a PCA basis of the same dimensionality. Rosa M. Figueras i Ventura, Umesh Rajashekar, Zhou Wang 0001, Eero P. Simoncelli |
ICASSP | 3 |
| 2008 | General-purpose reduced-reference image quality assessment based on perceptually and statistically motivated image representationabstractDivisive normalization has been recognized as a successful approach to model the perceptual sensitivity of biological vision. It also provides a useful image representation that is well-matched to the statistical properties of natural images. Here we propose a reduced- reference image quality assessment method in the divisive normalization transform domain, where the quality of an image is evaluated based on a set of reduced-reference features extracted from a divisive normalization representation of the image. The proposed method is general-purpose, in the sense that no assumption is made about the types of distortions occurred in the image being evaluated. The proposed method is trained and tested using the LIVE database and demonstrates good performance for a wide range of distortions. Zhou Wang 0001 |
ICIP | 2 |
| 2007 | Quality-Aware VideoabstractDevelopment in network visual communications has emphasized on the need of objective, reliable and easy-to-use video quality assessment (VQA) systems. This paper introduces a novel idea of quality-aware video (QAV), in which extracted features about the original video sequence are invisibly embedded into the same video data. When such a QAV sequence is distributed over an error-prone network, a network user who receives it can decode the hidden messages and use them to evaluate the quality degradations between the original and the received video sequences. Our first implementation of QAV employs (1) a novel reduced-reference VQA method based on a statistical model of natural video, and (2) a 3D discrete cosine transform-based data hiding algorithm. The proposed approach does not assume any prior knowledge about image distortions, and the simulation results demonstrate its potentials to be generalized for different types and degrees of image distortions. Basavaraj Hiremath, Zhou Wang 0001 |
ICIP (3) | 3 |
| 2007 | Video Quality Assessment by Incorporating a Motion Perception ModelabstractMotion is one of the most important types of information contained in natural video, but direct use of motion information in the design of video quality assessment algorithms has not been deeply investigated. Here we propose to incorporate a recent motion perception model in an information theoretic framework. This allows us to estimate both the motion information content and the perceptual uncertainty in video signals. Improved video quality assessment algorithms are obtained by incorporating the model as spatiotemporal weighting factors, where the weight increases with the information content and decreases with the perceptual uncertainty. The proposed approach is validated using the Video Quality Experts Group Phase I test dataset. Zhou Wang 0001 |
ICIP (2) | 2 |
| 2007 | Perceptual Image Coding Based on a Maximum of Minimal Structural Similarity CriterionabstractPerceptual image coding algorithms typically impose perceptual modeling in a preprocessing stage. A perceptual normalization model is often used to transform the original image signal into a perceptually uniform space, in which all the transform coefficients have equal perceptual importance. Standard coding schemes are then applied uniformly to all coefficients. Here we use a different approach, in which we iteratively reallocates the available bits over the image space based on amaximumofminimalstructuralsimilaritycriterion. We demonstrate the proposed method by incorporating it with the bitplane coding scheme in the set partitioning in hierarchical trees algorithm. Zhou Wang 0001, Xinli Shang |
ICIP (2) | 1 |
| 2007 | Palmprint Verification using Complex Wavelet TransformabstractPalmprint is a unique and reliable biometric characteristic with high usability. With the increasing demand of automatic palmprint authentication systems, the development of accurate and robust palmprint verification algorithms has been attracting a lot of interests. The relative translation, rotation and distortion between two palmprint images will introduce much error in palmprint matching. However, an accurate registration of palmprint images is too time-consuming. In this paper, we propose a modified complex wavelet structural similarity index (CW-SSIM) to compute the matching score and hence identify the input palmprint. Since CW-SSIM is robust to translation, small rotation and distortion, a fast rough alignment of palmprint images is sufficient. CW-SSIM is also insensitive to luminance and contrast changes. Our experimental results show that the proposed scheme outperforms the state-of-the-art methods by achieving a higher genuine acceptance rate and a lower false acceptance rate simultaneously. Lei Zhang 0006, Zhenhua Guo 0001, Zhou Wang 0001, David Zhang 0001 |
ICIP (2) | 3 |
| 2007 | Facial Range Image Matching Using the ComplexWavelet Structural Similarity MetricabstractWe propose a novel 3D face recognition algorithm based on facial range image matching using the complex wavelet structural similarity metric (CW-SSIM) metric. Compared with many existing 3D surface matching methods, CW-SSIM is computationally efficient and is robust to small geometrical distortions. Using a data set that contains 360 3D face models of 12 subjects, we tested the performance of the proposed method and compared it with existing 3D surface matching based face recognition algorithms. Verification and identification performance of each algorithm was evaluated by means of the receiver operating characteristic curve and the cumulative match characteristic curve. Among the algorithms tested, the proposed algorithm based on the CW-SSIM resulted in the best overall performance with an equal error rate of 9.13% and a rank 1 recognition rate of 98.6%, significantly better than all the other algorithms. Besides the introduction of a novel approach for 3D face recognition, this is also the first attempt to expand the application scope of complex wavelet domain similarity measure to range image matching in general Shalini Gupta, Mehul P. Sampat, Mia K. Markey, Alan C. Bovik, Zhou Wang 0001 |
WACV | 5 |
| 2006 | Measuring Intra- and Inter-Observer Agreement in Identifying and Localizing Structures in Medical ImagesabstractInter- and intra-observer variability exists in any measurements made on medical images. There are two sources of variability. The first occurs when the observers identify and localize the object of interest, and the second happens when the observers make appropriate measurement on the object of interest. A number of statistical methods are available to quantify the degree of agreement between measurements made by different observers. However, little has been done to develop metrics for quantifying the variability in identifying and localizing the objects of interest prior to measurement. In this paper, we propose to use the complex wavelet structural similarity index (CW-SSIM) method to measure the variability in identifying and localizing structures on images. Performance comparisons using simulated images as well as real mammography images demonstrate the effectiveness and robustness of the CW-SSIM method. Mehul P. Sampat, Zhou Wang 0001, Mia K. Markey, Gary J. Whitman, Tanya W. Stephens, Alan C. Bovik |
ICIP | 2 |
| 2006 | Quality-aware imagesabstractWe propose the concept of quality-aware image, in which certain extracted features of the original (high-quality) image are embedded into the image data as invisible hidden messages. When a distorted version of such an image is received, users can decode the hidden messages and use them to provide an objective measure of the quality of the distorted image. To demonstrate the idea, we build a practical quality-aware image encoding, decoding and quality analysis system, which employs: 1) a novel reduced-reference image quality assessment algorithm based on a statistical model of natural images and 2) a previously developed quantization watermarking-based data hiding technique in the wavelet transform domain. Zhou Wang 0001, Guixing Wu, Hamid R. Sheikh, Eero P. Simoncelli, En-Hui Yang, Alan C. Bovik |
IEEE Trans. Image Process. | 1 |
| 2005 | Translation Insensitive Image Similarity in Complex Wavelet DomainabstractWe propose a complex wavelet domain image similarity measure, which is simultaneously insensitive to luminance change, contrast change and spatial translation. The key idea is to make use of the fact that these image distortions lead to consistent magnitude and/or phase changes of local wavelet coefficients. Since small scaling and rotation of images can be locally approximated by translation, the proposed measure also shows robustness to spatial scaling and rotation when these geometric distortions are small relative to the size of the wavelet filters. Compared with previous methods, the proposed measure is computationally efficient, and can evaluate the similarity of two images without a precise registration process at the front end. Zhou Wang 0001, Eero P. Simoncelli |
ICASSP (2) | 1 |
| 2005 | An adaptive linear system framework for image distortion analysisabstractWe describe a framework for decomposing the distortion between two images into a linear combination of components. Unlike conventional linear bases such as those in Fourier or wavelet decompositions, a subset of the components in our representation are not fixed, but are adaptively computed from the input images. We show that this framework is a generalization of a number of existing image comparison approaches. As an example of a specific implementation, we select the components based on the structural similarity principle, separating the overall image distortions into non-structural distortions (those that do not change the structures of the objects in the scene) and the remaining structural distortions. We demonstrate that the resulting measure is effective in predicting image distortions as perceived by human observers. Zhou Wang 0001, Eero P. Simoncelli |
ICIP (3) | 1 |
| 2004 | Video quality assessment based on structural distortion measurement
Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
Signal Process. Image Commun. | 1 |
| 2004 | Image quality assessment: from error visibility to structural similarityabstractObjective methods for assessing perceptual image quality traditionally attempted to quantify the visibility of errors (differences) between a distorted image and a reference image using a variety of known properties of the human visual system. Under the assumption that human visual perception is highly adapted for extracting structural information from a scene, we introduce an alternative complementary framework for quality assessment based on the degradation of structural information. As a specific example of this concept, we develop a Structural Similarity Index and demonstrate its promise through a set of intuitive examples, as well as comparison to both subjective ratings and state-of-the-art objective methods on a database of images compressed with JPEG and JPEG2000. Zhou Wang 0001, Alan C. Bovik, Hamid R. Sheikh, Eero P. Simoncelli |
IEEE Trans. Image Process. | 1 |
| 2003 | Local Phase Coherence and the Perception of Blur
Zhou Wang 0001, Eero P. Simoncelli |
NIPS | 1 |
| 2003 | Foveation scalable video coding with automatic fixation selectionabstractImage and video coding is an optimization problem. A successful image and video coding algorithm delivers a good tradeoff between visual quality and other coding performance measures, such as compression, complexity, scalability, robustness, and security. In this paper, we follow two recent trends in image and video coding research. One is to incorporate human visual system (HVS) models to improve the current state-of-the-art of image and video coding algorithms by better exploiting the properties of the intended receiver. The other is to design rate scalable image and video codecs, which allow the extraction of coded visual information at continuously varying bit rates from a single compressed bitstream. Specifically, we propose a foveation scalable video coding (FSVC) algorithm which supplies good quality-compression performance as well as effective rate scalability. The key idea is to organize the encoded bitstream to provide the best decoded video at an arbitrary bit rate in terms of foveated visual quality measurement. A foveation-based HVS model plays an important role in the algorithm. The algorithm is adaptable to different applications, such as knowledge-based video coding and video communications over time-varying, multiuser and interactive networks. Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
IEEE Trans. Image Process. | 1 |
| 2002 | Foveated multipoint videoconferencing at low bit ratesabstractMultipoint videoconferencing (MPVC) involves three or more participants engaged in video communication over a network. A video server combines the video streams from each participant and then broadcasts the resulting stream to all participants. In this paper, we propose to use foveation, which is non-uniform resolution representation of an image reflecting the sampling in the retina, to reduce the bandwidth requirements of MPVC. We develop foveated MPVC algorithms for variable and constant bit rate MPVC. We show that foveated MPVC can provide considerable bit rate savings, and for the same bit rate, provide improvement in subjective quality. Hamid R. Sheikh, Shizhong Liu, Zhou Wang 0001, Alan C. Bovik |
ICASSP | 3 |
| 2002 | Why is image quality assessment so difficult?abstractImage quality assessment plays an important role in various image processing applications. A great deal of effort has been made in recent years to develop objective image quality metrics that correlate with perceived quality measurement. Unfortunately, only limited success has been achieved. In this paper, we provide some insights on why image quality assessment is so difficult by pointing out the weaknesses of the error sensitivity based framework, which has been used by most image quality assessment approaches in the literature. Furthermore, we propose a new philosophy in designing image quality metrics: The main function of the human eyes is to extract structural information from the viewing field, and the human visual system is highly adapted for this purpose. Therefore, a measurement of structural distortion should be a good approximation of perceived image distortion. Based on the new philosophy, we implemented a simple but effective image quality indexing algorithm, which is very promising as shown by our current results. Zhou Wang 0001, Alan C. Bovik, Ligang Lu |
ICASSP | 1 |
| 2002 | No-reference perceptual quality assessment of JPEG compressed imagesabstractHuman observers can easily assess the quality of a distorted image without examining the original image as a reference. By contrast, designing objective No-Reference (NR) quality measurement algorithms is a very difficult task. Currently, NR quality assessment is feasible only when prior knowledge about the types of image distortion is available. This research aims to develop NR quality measurement algorithms for JPEG compressed images. First, we established a JPEG image database and subjective experiments were conducted on the database. We show that Peak Signal-to-Noise Ratio (PSNR), which requires the reference images, is a poor indicator of subjective quality. Therefore, tuning an NR measurement model towards PSNR is not an appropriate approach in designing NR quality metrics. Furthermore, we propose a computational and memory efficient NR quality assessment model for JPEG images. Subjective test results are used to train the model, which achieves good quality prediction performance. Hamid R. Sheikh, Zhou Wang 0001, Alan C. Bovik |
ICIP (1) | 2 |
| 2002 | Generalized bitplane-by-bitplane shift method for JPEG2000 ROI codingabstractOne interesting feature of the new JPEG2000 image coding standard is support of region of interest (ROI) coding using the maximum shift (Maxshift) method, which allows for arbitrarily shaped ROI image compression without shape coding or explicitly transmitting any shape information to the decoder. The major disadvantage of the Maxshift method is that it cannot adjust the scaling value which determines the degree of relative importance between the ROI and the background wavelet coefficients. The bitplane-by-bitplane shift (BbBShift) method was introduced to support both arbitrary ROI shape and arbitrary scaling without shape coding. We propose a generalized BbBShift (GBbBShift) method, which delivers much more flexibility than both Maxshift and BbBShift for "degree-of-interest" adjustment of the ROI with insignificant effect on coding efficiency and computational complexity. Experiments show that it can provide significantly better visual quality than Maxshift at low bit rates. GBbBShift is not compliant with the current JPEG2000 definitions. In order to use it, a new ROI coding mode would need to be added to the standard. Zhou Wang 0001, Serene Banerjee, Brian L. Evans, Alan C. Bovik |
ICIP (3) | 1 |
| 2002 | Video quality assessment using structural distortion measurementabstractObjective image/video quality measures play important roles in various image/video processing applications, such as compression, communication, printing, analysis, registration, restoration and enhancement. Most proposed quality assessment approaches in the literature are error sensitivity-based methods. We follow a new philosophy in designing image/video quality metrics, which uses structural distortion as an estimation of perceived visual distortion. We develop a new approach for video quality assessment. Experiments on the video quality experts group (VQEG) test data set shows that the new quality measure has higher correlation with subjective quality measurement than the proposed methods in VQEG's Phase I tests for full-reference video quality assessment. Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
ICIP (3) | 1 |
| 2002 | Full-reference video quality assessment considering structural distortion and no-reference quality evaluation of MPEG videoabstractThere has been an increasing need recently to develop objective quality measurement techniques that can predict perceived video quality automatically. This paper introduces two video quality assessment models. The first one requires the original video as a reference and is a structural distortion measurement based approach, which is different from traditional error sensitivity based methods. Experiments on the video quality experts group (VQEG) test data set show that the new quality measure has higher correlation with subjective quality evaluation than the proposed methods in VQEG's Phase I tests for full-reference video quality assessment. The second model is designed for quality estimation of compressed MPEG video stream without referring to the original video sequence. Preliminary experimental results show that it correlates well with our full-reference quality assessment model. Ligang Lu, Zhou Wang 0001, Alan C. Bovik, Jack Kouloheris |
ICME (1) | 2 |
| 2002 | A universal image quality indexabstractWe propose a new universal objective image quality index, which is easy to calculate and applicable to various image processing applications. Instead of using traditional error summation methods, the proposed index is designed by modeling any image distortion as a combination of three factors: loss of correlation, luminance distortion, and contrast distortion. Although the new index is mathematically defined and no human visual system model is explicitly employed, our experiments on various image distortion types indicate that it performs significantly better than the widely used distortion metric mean squared error. Demonstrative images and an efficient MATLAB implementation of the algorithm are available online at http://anchovy.ece.utexas.edu//spl sim/zwang/research/quality_index/demo.html. Zhou Wang 0001, Alan C. Bovik |
IEEE Signal Process. Lett. | 1 |
| 2002 | Bitplane-by-bitplane shift (BbBShift) - a suggestion for JPEG2000 region of interest image codingabstractThe JPEG2000 image coding standard defines two kinds of region of interest (ROI) coding methods-the general scaling based method and the maximum shift (maxshift) method. The former requires shape coding of the ROIs, which leads to increased complexity of codec implementations and limits the choice of ROI shapes (currently, only rectangle and ellipse shapes are defined). The latter allows for arbitrarily shaped ROI coding without explicitly transmitting any shape information to the decoder, but does not have the flexibility to select an arbitrary scaling value to define the relative importance of the ROI and the background wavelet coefficients. We propose a bitplane-by-bitplane shift (BbBShift) method, which supports both arbitrary ROI shape and arbitrary scaling without shape coding. Zhou Wang 0001, Alan C. Bovik |
IEEE Signal Process. Lett. | 1 |
| 2002 | Image information restoration based on long-range correlationabstractA new class of image information restoration algorithms, virtually different from traditional techniques, is proposed. In comparison with other approaches, our methods use not only the information in local areas, but also that in the remote regions of the image. The methods originate from the idea that there exists abundant long-range correlation within natural images and the human vision system, composed of our eyes and brains, can sufficiently utilize such types of information redundancy to implement the functions of image interpretation, representation, restoration, enhancement and error concealment. Our general approach can be summarized as five basic steps: fetching, searching, matching, competing and recovering. The experimental results on several practical applications show that our methods perform substantially better than many other state-of-the-art methods. David Zhang 0001, Zhou Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Rate scalable video coding using a foveation-based human visual system modelabstractRecently, there have been two interesting trends in image and video coding research. One is to use human visual system (HVS) models to improve the current state-of-the-art coding algorithms by better exploiting the properties of the intended receiver. The other is to design rate-scalable video codecs, which allow the extraction of coded visual information at continuously varying bit rates from a single compressed bitstream. We follow these two trends and propose a foveation scalable video coding (FSVC) algorithm, which supplies good quality-compression performance as well as effective rate scalability to support simple and precise bit rate control. A foveation-based HVS model plays a key role in the algorithm. The algorithm is amenable to the inclusion of various HVS models and adaptable to different video communication applications. Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
ICASSP | 1 |
| 2001 | Wavelet-based foveated image quality measurement for region of interest image codingabstractRegion of interest (ROI) image and video compression techniques have been widely used in visual communication applications in an effort to deliver good quality images and videos at limited bandwidths. Most image quality metrics have been developed for uniform resolution images. These metrics are not appropriate for the assessment of ROI coded images, where space-variant resolution is necessary. The spatial resolution of the human visual system (HVS) is highest around the point of fixation and decreases rapidly with increasing eccentricity. Since the ROIs are usually the regions "fixated" by human eyes, the foveation property of the HVS supplies a natural approach for guiding the design of ROI image quality measurement algorithms. We have developed an objective quality metric for ROI coded images in the wavelet transform domain. This metric can serve to mediate the compression and enhancement of ROI coded images and videos. We show its effectiveness by applying it to an embedded foveated image coding system. Zhou Wang 0001, Alan C. Bovik, Ligang Lu |
ICIP (2) | 1 |
| 2001 | Adaptive Frame Prediction for Foveation Scalable Video CodingabstractEmbedded rate scalable video coding allows for the extraction of coded visual information at continuously varying bit rates from a single compressed bitstream. This is a very attractive feature for many multimedia communication applications. Motion estimation (ME) /motion compensation (MC) techniques are widely employed in various video coding systems to reduce temporal information redundancy. One of the major challenging problems in ME/MC based rate scalable video coding is how to generate the prediction frame from the previous frame to match the current frame. This problem is more difficult in rate scalable coding than in fixed rate coding because the decoding data rate is unavailable to the encoder. We propose an adaptive frame prediction scheme for foveation scalable video coding (FSVC), which is a new video coding algorithm that combines a foveation-based human visual system (HVS) model with a wavelet-based rate scalable coding algorithm. The new frame prediction algorithm provides an adaptive mechanism to control the prediction errors while reduce error propagation. Ligang Lu, Zhou Wang 0001, Alan C. Bovik, Jack Kouloheris |
ICME | 2 |
| 2001 | Embedded foveation image codingabstractThe human visual system (HVS) is highly space-variant in sampling, coding, processing, and understanding. The spatial resolution of the HVS is highest around the point of fixation (foveation point) and decreases rapidly with increasing eccentricity. By taking advantage of this fact, it is possible to remove considerable high-frequency information redundancy from the peripheral regions and still reconstruct a perceptually good quality image. Great success has been obtained previously by a class of embedded wavelet image coding algorithms, such as the embedded zerotree wavelet (EZW) and the set partitioning in hierarchical trees (SPIHT) algorithms. Embedded wavelet coding not only provides very good compression performance, but also has the property that the bitstream can be truncated at any point and still be decoded to recreate a reasonably good quality image. In this paper, we propose an embedded foveation image coding (EFIC) algorithm, which orders the encoded bitstream to optimize foveated visual quality at arbitrary bit-rates. A foveation-based image quality metric, namely, foveated wavelet image quality index (FWQI), plays an important role in the EFIC system. We also developed a modified SPIHT algorithm to improve the coding efficiency. Experiments show that EFIC integrates foveation filtering with foveated image coding and demonstrates very good coding performance and scalability in terms of foveated image quality measurement. Zhou Wang 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 1 |
| 2000 | Blind Measurement of Blocking Artifacts in ImagesabstractThe objective measurement of blocking artifacts plays an important role in the design, optimization, and assessment of image and video coding systems. We propose a new approach that can blindly measure blocking artifacts in images without reference to the originals. The key idea is to model the blocky image as a non-blocky image interfered with a pure blocky signal. The task of the blocking effect measurement algorithm is then to detect and evaluate the power of the blocky signal. The proposed approach has the flexibility to integrate human visual system features such as the luminance and the texture masking effects. Zhou Wang 0001, Alan C. Bovik, Brian L. Evans |
ICIP | 1 |
| 2000 | Hybrid image coding based on partial fractal mapping
Zhou Wang 0001, David Zhang 0001, Yinglin Yu |
Signal Process. Image Commun. | 1 |
| 1998 | Restoration of impulse noise corrupted images using long-range correlationabstractWe present a new algorithm that can remove impulse noise from corrupted images while preserving details. The algorithm is fundamentally different from the traditional methods in that it can utilize information not just of a local window centered about the corrupted pixel, but also of some remote regions in the image. Computer simulations indicate that our algorithm outperforms many existing techniques. Zhou Wang 0001, David Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 1998 | A novel approach for reduction of blocking effects in low-bit-rate image compressionabstractIn this letter we propose a new approach for removing blocking artifacts in reconstructed block-encoded images. The key of the approach is using piecewise similarity within different parts of the image as a priori to give reasonable modifications to the block boundary pixels. This makes our approach different from traditional ones, which are often developed by applying some kinds of smoothing constraints on local regions. Experimental results show that our approach well achieves enhanced decoding for JPEG-decoded images both objectively and subjectively. Zhou Wang 0001, David Zhang 0001 |
IEEE Trans. Commun. | 1 |
| 1998 | Best neighborhood matching: an information loss restoration technique for block-based image coding systemsabstractImperfect transmission of block-coded images often results in lost blocks. A novel error concealment method called best neighborhood matching (BNM) is presented by using a special kind of information redundancy-blockwise similarity within the image. The proposed algorithm can utilize the information of not only neighboring pixels, but also remote regions in the image. Very good restoration results are obtained by experiments. Zhou Wang 0001, Yinglin Yu, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 1997 | Dynamic fractal transform with applications to image data compression
Zhou Wang 0001, Yinglin Yu |
J. Comput. Sci. Technol. | 1 |