EDBT 2026 Demo / reviewers in the wild / expert
Alan C. Bovik
dblp:b/ACBovik · also Al Bovik, Alan Bovik, Alan Conrad Bovik
· DBLP profile ↗
473ranked-venue papers
18as first author
91since 2021 · last 2026
0000-0001-6067-710XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 411 · 11 first-author · 85 since 2021Artificial intelligence and machine learning · 52 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 2 first-authorSystems, architecture and hardware · 4Human-computer interaction and ubiquitous computing · 3Computer networks · 2Security and privacy · 2Databases, data management, data science and information retrieval · 2Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Non-Aligned Reference Image Quality Assessment for Novel View SynthesisabstractEvaluating the perceptual quality of Novel View Synthesis (NVS) images remains a key challenge, particularly in the absence of pixel-aligned ground truth references. Full-Reference Image Quality Assessment (FR-IQA) methods fail under misalignment, while No-Reference (NR-IQA) methods struggle with generalization. In this work, we introduce a Non-Aligned Reference (NAR-IQA) framework tailored for NVS, where it is assumed that the reference view shares partial scene content but lacks pixel-level alignment. We constructed a large-scale image dataset containing synthetic distortions targeting Temporal Regions of Interest (TROI) to train our NAR-IQA model. Our model is built on a contrastive learning framework that incorporates LoRA-enhanced DINOv2 embeddings and is guided by supervision from existing IQA methods. We train exclusively on synthetically generated distortions, deliberately avoiding overfitting to specific real NVS samples and thereby enhancing the model’s generalization capability. Our model outperforms state-of-the-art FR-IQA, NR-IQA, and NAR-IQA methods, achieving robust performance on both aligned and non-aligned references. We also conducted a novel user study to gather data on human preferences when viewing non-aligned references in NVS. We find strong correlation between our proposed quality prediction model and the collected subjective ratings. For dataset, and code, please visit our project page: https://stootaghaj.github.io/nova-project/ Abhijay Ghildyal, Rajesh Sureddi, Nabajeet Barman, Saman Zad Tootaghaj, Alan C. Bovik |
WACV | 5 |
| 2026 | BrightRate: Quality Assessment for User-Generated HDR VideosabstractHigh Dynamic Range (HDR) videos offer superior luminance and color fidelity as compared to Standard Dynamic Range (SDR) content. The rapid growth of User-Generated Content (UGC) on platforms such as YouTube, Instagram, and TikTok has brought a significant increase in the volumes of streamed and shared UGC videos. This newer category of videos brings new challenges to the development of effective No-Reference (NR) video quality assessment (VQA) models specialized to HDR UGC, because of the extreme variety and severities of distortions, arising from diverse capture, editing, and processing outcomes. Towards addressing this issue, we introduce BrightVQ, a sizeable new psychometric data resource. It is the first large-scale subjective video quality database dedicated to the quality modelling of HDR UGC videos. BrightVQ comprises 2,100 videos, on which we collected 73,794 perceptual quality ratings. Using this dataset, we also developed BrightRate, a novel video quality prediction model designed to capture both UGC-specific distortions coexisting with HDR-specific artifacts. Extensive experimental results demonstrate that BrightRate achieves state-of-the-art performance across HDR databases. Shreshth Saini, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
WACV | 6 |
| 2026 | Blind S-3D VR picture quality prediction using trivariate brightness, color, and disparity statistics
Ajay Kumar Reddy Poreddy, Balasubramanyam Appina, Priyanka Kokil, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2026 | Quality Prediction of Embedded and Overlaid Text in User-Generated Visual ContentabstractUser-generated visual content (UGC) now occupies a significant fraction of internet traffic, and billions of UGC videos and pictures are uploaded daily. Among these, short-form video content now accounts for most of the videos consumed by online users. Given the popularity of short-form UGC content, being able to control the perceptual quality of UGC videos has emerged as an important problem. Visual UGC is subject to myriad types, severity, and combinations of distortions. While UGC video quality has been closely studied, the quality and legibility of text that is overlaid or embedded in short-form UGC videos has received relatively low attention. However, being able to accurately predict text quality in images is important, since it both impacts the overall perception of the content it is embedded in, as well as the messages being conveyed. It is also beneficial for applications involving image or video text recognition which can affect visual search and content identification. Analyzing the quality of text embedded in pictures or videos is a hard problem, since perception of it is commingled with the surrounding visual content. Our work, which greatly extends our early report on text legibility prediction, contributes to both the psychophysics of embedded text quality as well as to computational models of its perception. We have created two subjective datasets-designated as the LIVE-COCO Text Legibility (LIVE-COCO-TL) Database (a modification of COCO-Text), and the LIVE-YouTube Text-in-Video Quality (LIVE-YT-TVQ) Database. LIVE-COCO-TL contains 74,440 text patches with legibility annotations, while LIVE-YT-TVQ contains $\sim ~19$ K subjective quality ratings on 405 videos and 641 text patches extracted from them. We build models that predict embedded or overlaid text legibility and text quality, as well as a multi-task model that simultaneously predicts the overall quality of videos with embedded or overlaid and local text quality. We are making the databases and all models freely available at https://live.ece.utexas.edu/research/LIVE_YouTube_Text_Quality_Assessment/index.html. Maniratnam Mandal, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2026 | HoloQA: Full Reference Video Quality Assessor of Rendered Human Avatars in Virtual RealityabstractWe present HoloQA, a new state-of-the-art Full Reference Video Quality Assessment (VQA) model that was designed using principles of visual neuroscience, information theory, and self-supervised deep learning to accurately predict the quality of rendered digital human avatars in Virtual Reality (VR) and Augmented Reality (AR) systems. The growing adoption of VR/AR applications that aim to transmit digital human avatars over bandwidth-limited video networks has driven the need for VQA algorithms that better account for the kinds of distortions that reduce the quality of rendered and viewed avatars. As we will show, standard VQA models often fail to capture distortions unique to the rendering, transmission, and compression of videos containing human avatars. Towards solving this difficult problem, we adopt a multi-level Mixture-of-Experts approach. This involves computing distortion-aware perceptual features and high-level content-aware deep features that capture semantic attributes of human body avatars. The high-level features are computed using a self-supervised, pre-trained deep learning network. We show that HoloQA is able to achieve state-of-the-art performance on the recently introduced LIVE-Meta Rendered Human Avatar VQA database, demonstrating its efficacy in predicting the quality of rendered human avatars in VR. Furthermore, we demonstrate the competitive performance of HoloQA on other digital human avatar databases and on another synthetically generated video quality use case: cloud gaming. The code associated with this work will be made available on https://github.com/avinabsaha/HologramQAGitHub. Avinab Saha, Yu-Chih Chen, Christian Häne, Jean-Charles Bazin, Ioannis Katsavounidis, Alexandre Chapiro, Alan C. Bovik |
IEEE Trans. Image Process. | 7 |
| 2026 | Joint Quality Assessment and Example-Guided Tone Mapping by Disentangling Picture Appearance From ContentabstractThe deep learning revolution has strongly impacted low-level image processing tasks such as style/domain transfer, enhancement/restoration, and visual quality assessments. Despite often being treated separately, the aforementioned tasks share a common theme of understanding, editing, or enhancing the appearance of input images without modifying the underlying content. We leverage this observation to develop a novel disentangled representation learning method that decomposes inputs into content and appearance features. The model is trained in a self-supervised manner and we use the learned features to develop a new quality prediction model named DisQUE. We demonstrate through extensive evaluations that DisQUE achieves state-of-the-art accuracy across quality prediction tasks and distortion types. Moreover, we demonstrate that the same features may also be used for image processing tasks such as HDR tone mapping, where the desired output characteristics may be tuned using example input-output pairs. Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Hassene Tmar, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2026 | Distortion-Sensitive Masked Autoencoder for Omnidirectional Video Quality AssessmentabstractOmnidirectional Video Quality Assessment (OVQA) is a challenging task due to the limited availability of adequate numbers of training samples for learning representations of distortions on omnidirectional videos. The recent masked autoencoder (MAE) has shown promising performance in learning local and global representations in a self-supervised way, and can be used to attempt to mitigate the difficulty of having insufficient annotated samples to adequately train omnidirectional video quality prediction models. But the reconstruction tasks that MAE models are designed for do not pertain to predicting diverse perceptual distortions, especially those relevant to the task of OVQA. We have attempted to overcome these limitations to harness and apply the power of the MAE concept to the OVQA problem. Towards this purpose, we create a Distortion-Sensitive Masked AutoEncoder (DS-MAE) that is able to represent perceptual distortions on omnidirectional videos. DS-MAE extracts viewports from omnidirectional videos and employs a masked autoencoding module (MAM) and a knowledge replay module (KRM) to learn representations on each viewport. In the MAM, distorted patches from omnidirectional videos are masked, by replacing them with undistorted counterparts. The autoencoder is trained to reconstruct the masked distortions, imbuing them with the ability to represent diverse video degradations. The KRM extracts and stores content representations, which are then “replayed” to mitigate potential catastrophic forgetting of content during training of the DS-MAE. Finally, a simple OVQA model is constructed using the pre-trained DS-MAE across all viewports. The new model, called OmniVQA, was tested on three public OVQA datasets. The experimental results show that OmniVQA delivers competitive performance against all compared models. Zongyao Hu, Lixiong Liu, Ke Gu 0001, Leida Li, Alan C. Bovik |
IEEE Trans. Multim. | 5 |
| 2025 | PIT-QMM: A Large Multimodal Model for No-Reference Point Cloud Quality AssessmentabstractLarge Multimodal Models (LMMs) have recently enabled considerable advances in the realm of image and video quality assessment, but this progress has yet to be fully explored in the domain of 3D assets. We are interested in using these models to conduct No-Reference Point Cloud Quality Assessment (NR-PCQA), where the aim is to automatically evaluate the perceptual quality of a point cloud in absence of a reference. We begin with the observation that different modalities of data – text descriptions, 2D projections, and 3D point cloud views – provide complementary information about point cloud quality. We then construct PIT-QMM, a novel LMM for NR-PCQA that is capable of consuming text, images and point clouds end-to-end to predict quality scores. Extensive experimentation shows that our proposed method outperforms the state-of-the-art by significant margins on popular benchmarks with fewer training iterations. We also demonstrate that our framework enables distortion localization and identification, which paves a new way forward for model explainability and interactivity. Code and datasets are available at https://www.github.com/shngt/pit-qmm. Gregoire Phillips, Alan C. Bovik |
ICIP | 3 |
| 2025 | GeoScaler: Geometry and Rendering-Aware Downsampling of 3D Mesh TexturesabstractHigh-resolution texture maps are necessary to accurately represent real-world objects with 3D meshes. The large sizes of textures can bottleneck the real-time rendering of high-quality virtual 3D scenes on devices that have low computational budgets and limited memory. Downsampling the texture maps directly addresses the issue, albeit at the cost of visual fidelity. Traditionally, downsampling of texture maps is performed using methods such as bicubic interpolation and the Lanczos algorithm. These methods ignore the geometric layout of the mesh and its UV parameterization and also do not account for the rendering process used to obtain the final visualization that the users will experience. Towards filling these gaps, we introduce GeoScaler, which is a method of downsampling texture maps of 3D meshes while incorporating geometric cues and maximizing the visual fidelity of the rendered views of the textured meshes. We show that the textures generated by GeoScaler deliver significantly better quality rendered images compared to those generated by traditional downsampling methods. Sai Karthikey Pentapati, Anshul Rai, Arkady Ten, Chaitanya Atluru, Alan C. Bovik |
ICIP | 5 |
| 2025 | CHUG: Crowdsourced User-Generated HDR Video Quality DatasetabstractHigh Dynamic Range (HDR) videos enhance visual experiences with superior brightness, contrast, and color depth. The surge of User-Generated Content (UGC) on platforms like YouTube and TikTok introduces unique challenges for HDR video quality assessment (VQA) due to diverse capture conditions, editing artifacts, and compression distortions. Existing HDR-VQA datasets primarily focus on professionally generated content (PGC), leaving a gap in understanding real-world UGC-HDR degradations. To address this, we introduce CHUG: Crowdsourced User-Generated HDR Video Quality Dataset, the first large-scale subjective study on UGC-HDR quality. CHUG comprises 856 UGC-HDR source videos, transcoded across multiple resolutions and bitrates to simulate real-world scenarios, totaling 5,992 videos. A large-scale study via Amazon Mechanical Turk collected 211,848 perceptual ratings. CHUG provides a benchmark for analyzing UGC-specific distortions in HDR videos. We anticipate CHUG will advance No-Reference (NR) HDR-VQA research by offering a large-scale, diverse, and real-world UGC dataset. The dataset is publicly available at: https://shreshthsaini.github.io/CHUG/. Shreshth Saini, Alan C. Bovik, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli |
ICIP | 2 |
| 2025 | Triqa: Image Quality Assessment by Contrastive Pretraining on Ordered Distortion TripletsabstractImage Quality Assessment (IQA) models aim to predict perceptual image quality in alignment with human judgments. No-Reference (NR) IQA remains particularly challenging due to the absence of a reference image. While deep learning has significantly advanced this field, a major hurdle in developing NR-IQA models is the limited availability of subjectively labeled data. Most existing deep learning-based NR-IQA approaches rely on pre-training on large-scale datasets before fine-tuning for IQA tasks. To further advance progress in this area, we propose a novel approach that constructs a custom dataset using a limited number of reference content images and introduces a no-reference IQA model that incorporates both content and quality features for perceptual quality prediction. Specifically, we train a quality-aware model using contrastive triplet-based learning, enabling efficient training with fewer samples while achieving strong generalization performance across publicly available datasets. Our repository is available at https://github.com/rajeshsureddi/triqa.1 Rajesh Sureddi, Saman Zad Tootaghaj, Nabajeet Barman, Alan C. Bovik |
ICIP | 4 |
| 2025 | LGDM: Latent Guidance in Diffusion Models for Perceptual EvaluationsabstractDespite recent advancements in latent diffusion models that generate high-dimensional image data and perform various downstream tasks, there has been little exploration into perceptual consistency within these models on the task of No-Reference Image Quality Assessment (NR-IQA). In this paper, we hypothesize that latent diffusion models implicitly exhibit perceptually consistent local regions within the data manifold. We leverage this insight to guide on-manifold sampling using perceptual features and input measurements. Specifically, we propose Perceptual Manifold Guidance (PMG), an algorithm that utilizes pretrained latent diffusion models and perceptual quality features to obtain perceptually consistent multi-scale and multi-timestep feature maps from the denoising U-Net. We empirically demonstrate that these hyperfeatures exhibit high correlation with human perception in IQA tasks. Our method can be applied to any existing pretrained latent diffusion model and is straightforward to integrate. To the best of our knowledge, this paper is the first work on guiding diffusion model with perceptual features for NR-IQA. Extensive experiments on IQA datasets show that our method, LGDM, achieves state-of-the-art performance, underscoring the superior generalization capabilities of diffusion models for NR-IQA tasks. Shreshth Saini, Ru-Ling Liao, Alan C. Bovik |
ICML | 4 |
| 2025 | Rectified CFG++ for Flow Based ModelsabstractClassifier‑free guidance (CFG) is the workhorse for steering large diffusion models toward text‑conditioned targets, yet its naïve application to rectified flow (RF) based models provokes severe off–manifold drift, yielding visual artifacts, text misalignment, and brittle behaviour. We present Rectified-CFG++, an adaptive predictor–corrector guidance that couples the deterministic efficiency of rectified flows with a geometry‑aware conditioning rule. Each inference step first executes a conditional RF update that anchors the sample near the learned transport path, then applies a weighted conditional correction that interpolates between conditional and unconditional velocity fields. We prove that the resulting velocity field is marginally consistent and that its trajectories remain within a bounded tubular neighbourhood of the data manifold, ensuring stability across a wide range of guidance strengths. Extensive experiments on large‑scale text‑to‑image models (Flux, Stable Diffusion 3/3.5, Lumina) show that Rectified-CFG++ consistently outperforms standard CFG on benchmark datasets such as MS‑COCO, LAION‑Aesthetic, and T2I‑CompBench. Project page: https://rectified-cfgpp.github.io/. Shreshth Saini, Alan C. Bovik |
NeurIPS | 3 |
| 2025 | A Subjective Video Quality Dataset for Comparative Evaluation of HDR and SDR
Cheng-Han Lee, Yixu Chen, Zaixi Shang, Hai Wei, Alan C. Bovik |
PCS | 6 |
| 2025 | Hierarchical Neural Surfaces for 3D Mesh Compression
Sai Karthikey Pentapati, Gregoire Phillips, Alan C. Bovik |
PCS | 3 |
| 2025 | Mesh Compression with Quantized Neural Displacement FieldsabstractAbstract Implicit neural representations (INRs) have been successfully used to compress a variety of 3D surface representations such as Signed Distance Functions (SDFs), voxel grids, and also other forms of structured data such as images, videos, and audio. However, these methods have been limited in their application to unstructured data such as 3D meshes and point clouds. This work presents a simple yet effective method that extends the usage of INRs to compress 3D triangle meshes. Our method encodes a displacement field that refines the coarse version of the 3D mesh surface to be compressed using a small neural network. Once trained, the neural network weights occupy much lower memory than the displacement field or the original surface. We show that our method is capable of preserving intricate geometric textures and demonstrates state‐of‐the‐art performance for compression ratios ranging from 4x to 380x (See Figure 1 for an example). Sai Karthikey Pentapati, Gregoire Phillips, Alan C. Bovik |
Comput. Graph. Forum | 3 |
| 2025 | Estimating the resize parameter in end-to-end learned image compression
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Lukas Krasula, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2025 | Understanding, detecting, and removing perceptual banding artifacts in compressed videosabstractBanding artifacts, or false contouring, are a common compression impairment that often appears on large smooth regions of encoded videos and images. These staircase-like color bands can be very noticeable and annoying, even on otherwise high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we study this artifact, by first analyzing the perceptual and encoding aspects of banding artifacts, then propose a new distortion-specific no-reference video quality algorithm for predicting banding artifacts, inspired by perceptual models. The proposed banding detector can generate a pixel-wise banding visibility map, and output overall banding severity scores at both the frame and video levels. Furthermore, we propose a deep learning based approach to improve the overall perceptual quality of compressed videos by joint debanding and compression artifact removal. Our experimental results show that the proposed banding detector delivers better consistency with subjective evaluations, and is able to detect different perceptual severity levels of bands. The debanding experiments also show that the proposed algorithm outperforms recent debanding models both visually and quantitatively. The code is available at https://github.com/google/bband-adaband and https://github.com/vztu/DebandingNet . Zhengzhong Tu, Chia-Ju Chen, Jessie Lin, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
Signal Process. Image Commun. | 7 |
| 2025 | Cut-FUNQUE: An objective quality model for compressed tone-mapped High Dynamic Range videos
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Hassene Tmar, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2025 | Perceptually-Guided VR Style TransferabstractVirtual reality (VR) makes it possible to provide immersive multimedia content composed of omnidirectional videos (ODVs). Towards enabling more immersive and satisfying VR content, methods are needed to manipulate VR scenes, taking into account perceptual factors related to viewers' quality of experience (QoE). For example, style transfer methods can be applied to VR content, allowing users to create artistic or surreal effects in their immersive environments. Here, we study perceptual factors that affect the sensation of stylized immersiveness, including color dynamics and spatio-temporal consistency. To do this, we introduce an immersiveness sensitivity model of luminance and color perception, and use it to measure the color dynamics and spatio-temporal consistency of stylized VR contents. We subsequently use this model to construct a perceptually-guided VR style transfer model called VR Style Transfer GAN (VRST-GAN). VRST-GAN learns to transfer a desired style into VR to enhance immersiveness by considering color dynamics while preserving spatio-temporal consistency. We demonstrate the effectiveness of VRST-GAN via qualitative and quantitative experiments. We also develop a VR Immersiveness Predictor (VR-IP) that is able to predict the sensation of immersiveness using the perceptual model. In our experiments, VR-IP predicts immersiveness with an accuracy of 91%. Seonghwa Choi, Jungwoo Huh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2025 | Constructing Per-Shot Bitrate Ladders Using Visual Information FidelityabstractVideo service providers need their delivery systems to be able to adapt to network conditions, user preferences, display settings, and other factors. HTTP Adaptive Streaming (HAS) offers dynamic switching between different video representations to simultaneously enhance bandwidth consumption and users' streaming experiences. Per-shot encoding, pioneered by Netflix, optimizes the encoding parameters on each scene or shot. The Dynamic Optimizer (DO) uses the Video Multi-Method Assessment Fusion (VMAF) perceptual video quality prediction engine to deliver high-quality videos at reduced bitrates. Here we develop a perceptually optimized method of constructing optimal per-shot bitrate and quality ladders, using an ensemble of low-level features and Visual Information Fidelity (VIF) features. During inference, our method predicts the bitrate or quality ladder of a source video without any compression or quality estimation. We compare the performance of our model against other content-adaptive bitrate ladder prediction methods, a fixed bitrate ladder, and reference bitrate ladders constructed via exhaustive encoding using Bjøntegaard-delta (BD) metrics. Our proposed method shows excellent gains in bitrate and quality against the fixed bitrate ladder and only small losses against the reference bitrate ladder, while providing significant computational advantages. Krishna Srikar Durbha, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2025 | Subjective and Objective Analysis of Indian Social Media Video QualityabstractWe conducted a large-scale subjective study of the perceptual quality of User-Generated Mobile Video Content on a set of mobile-originated videos obtained from the Indian social media platform ShareChat. The content viewed by volunteer human subjects under controlled laboratory conditions has the benefit of culturally diversifying the existing corpus of User-Generated Content (UGC) video quality datasets. There is a great need for large and diverse UGC-VQA datasets, given the explosive global growth of the visual internet and social media platforms. This is particularly true in regard to videos obtained by smartphones, especially in rapidly emerging economies like India. ShareChat provides a safe and cultural community oriented space for users to generate and share content in their preferred Indian languages and dialects. Our subjective quality study, which is based on this data, supplies much needed cultural, visual, and language diversification to the overall shareable corpus of video quality data. We expect that this new data resource will also allow for the development of systems that can predict the perceived visual quality of Indian social media videos, and in this context, control scaling and compression protocols for streaming, provide better user recommendations, and guide content analysis and processing. We demonstrate the value of the new data resource by conducting a study of leading blind video quality models on it, including a simple new model, called MoEVA, which deploys a mixture of experts to predict video quality. Both the new LIVE-ShareChat Database and sample source code for MoEVA are being made freely available to the research community at https://github.com/sandeep-sm/LIVE-SC. Sandeep Mishra, Mukul Jha, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2025 | No-Reference Image Quality Assessment Leveraging GenAI ImagesabstractIn recent years, deep learning-based methods have made significant progress on the image quality assessment problem; however, challenges remain arising from the lack of annotated, real-world training data and consequent poor generalization ability. Towards addressing these challenges, we propose a no-reference image quality assessment (NR-IQA) method based on generative AI (GenAI) images. Specifically, we use GenAI images as reference images, employing a cold diffusion model to generate distorted images of four different distortion types, and we label these distorted images using a full-reference model, thereby making it possible to construct a large-scale pre-training dataset. We use this resource generation method to facilitate NR-IQA model building. We deploy a Multi-scale Cross Attention Block (MCAB) and a Scale Simple Attention Module (SSAM) to enhance feature representation by extracting multi-scale feature information from both the channel and spatial dimensions that are predictive of image quality. Extensive experiments on eight public databases demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. A public release of all the codes associated with this work will be made available on GitHub. Qingbing Sang, Qian Li 0060, Lixiong Liu, Zhaohong Deng, Xiaojun Wu 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2025 | Subjective and Objective Quality Assessment of Banding Artifacts on Compressed VideosabstractAlthough there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos. Noticeable banding artifacts can severely impact the perceptual quality of videos viewed on a high-end HDTV or high-resolution screen. Hence, there is a pressing need for a systematic investigation of the banding video quality assessment problem for advanced video codecs. Given that the existing publicly available datasets for studying banding artifacts are limited to still picture data only, which cannot account for temporal banding dynamics, we have created a first-of-a-kind open video dataset, dubbed LIVE-YT-Banding, which consists of 160 videos generated by four different compression parameters using the AV1 video codec. A total of 7,200 subjective opinions are collected from a cohort of 45 human subjects. To demonstrate the value of this new resources, we tested and compared a variety of models that detect banding occurrences, and measure their impact on perceived quality. Among these, we introduce an effective and efficient new no-reference (NR) video quality evaluator which we call CBAND. CBAND leverages the properties of the learned statistics of natural images expressed in the embeddings of deep neural networks. Our experimental results show that the perceptual banding prediction performance of CBAND significantly exceeds that of previous state-of-the-art models, and is also orders of magnitude faster. Moreover, CBAND can be employed as a differentiable loss function to optimize video debanding models. The LIVE-YT-Banding database, code, and pre-trained model are all publically available at https://github.com/uniqzheng/CBAND. Qi Zheng 0004, Li-Heng Chen, Chenlong He, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik, Yibo Fan, Zhengzhong Tu |
IEEE Trans. Image Process. | 7 |
| 2024 | A Real-World Satellite Video Subjective QOE DatabaseabstractIn the rapidly growing streaming service market, including satellite options, Internet Service Providers (ISPs) face the challenge of continually optimizing network performance to deliver superior video streaming quality, which is vital to optimize customer satisfaction. This pressing need has sparked a drive towards developing advanced Quality of Experience (QoE) prediction models, which are essential in enhancing streaming protocols and guaranteeing smooth viewing experiences for users. However, the efficacy of these models hinges on the availability of extensive, diverse datasets. To fill this critical data void, our study introduces the publicly available LIVE-Viasat Real-World Satellite QoE Database, with 179 videos from real-world streaming, encompassing a range of distortions. Enhanced by a study with 54 participants providing detailed QoE feedback, our work not only provides a rich analysis of the determinants of subjective QoE but also delves into how various streaming impairments influence user behavior, thereby offering a more holistic understanding of user satisfaction. Zaixi Shang, Alan C. Bovik, Jae Won Chung, David Lerner |
ICIP | 3 |
| 2024 | Subjective Portrait Region Cropping On Landscape Video StudyabstractWith the rise of mobile video consumption, adapting videos to non-traditional aspect ratios poses challenges for existing content. The use of static cropping and border padding often compromises visual quality, while warping may distort a video’s intended meaning. Here we advocate for a more effective approach - cropping significant regions within video frames in a temporal manner, while minimizing distortion and preserving essential content. However, the lack of a large-scale database devoted to informing these tasks impedes progress in this direction. Addressing this gap, we introduce the LIVE-YouTube Video Cropping (LIVE-YT VC) database, featuring 1800 videos labeled by 90 human subjects. Sourced from the YouTube-UGC and LSVQ databases, this collection serves as the largest subjective video portrait region cropping database. We evaluate our methodology using the SmartVidCrop [1] algorithm, establishing a benchmark for future research. Our contributions offer a crucial resource for advancing video aspect ratio transformation, ensuring that mobile-friendly video content retains its quality and meaning. The details of accessing the dataset have been provided in the supplementary material. Cheng-Han Lee, Maniratnam Mandal, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
ICIP | 6 |
| 2024 | Legit: Text Legibility For User-Generated MediaabstractUser-generated content (UGC) is ubiquitous across the internet as a result of billions of videos and images being uploaded each day. All kinds of UGC media are affected by natural distortions, occurring both during and after capture, which are inherently diverse and commingled. These distortions have different perceptual effects based on the media content. Given recent dramatic increases in the consumption of short-form content, the analysis and control of their perceptual quality has become an important problem. Regardless of the content, many UGC videos have overlaid and embedded texts in them, which are visually salient. Hence text quality has a significant impact on the global perception of video or image quality and needs to be studied. One of the most important factors in perceptual text quality in user-generated media is legibility, which has been studied very little in the context of computer vision. Predicting text legibility can also help in text recognition applications such as image search or document identification. This work aims at modeling text legibility using computer vision techniques and thus studying the relationship between text quality and legibility. We propose a modified dataset variant of COCO-Text [1] and a model for predicting text legibility for both handwritten and machine-generated texts. We also demonstrate how models trained to predict text legibility can help in the prediction of text (perceptual) quality. The dataset and models can be accessed here https://live.ece.utexas.edu/research/Quality/index.htm. Maniratnam Mandal, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 4 |
| 2024 | YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animalsabstract3D generation guided by text-to-image diffusion models enables the creation of visually compelling assets. However previous methods explore generation based on image or text. The boundaries of creativity are limited by what can be expressed through words or the images that can be sourced. We present YouDream, a method to generate high-quality anatomically controllable animals. YouDream is guided using a text-to-image diffusion model controlled by 2D views of a 3D pose prior. Our method is capable of generating novel imaginary animals that previous text-to-3D generative methods are unable to create. Additionally, our method can preserve anatomic consistency in the generated animals, an area where prior approaches often struggle. Moreover, we design a fully automated pipeline for generating commonly observed animals. To circumvent the need for human intervention to create a 3D pose, we propose a multi-agent LLM that adapts poses from a limited library of animal 3D poses to represent the desired animal. A user study conducted on the outcomes of YouDream demonstrates the preference of the animal models generated by our method over others. Visualizations and code are available at https://youdream3d.github.io/. Sandeep Mishra, Oindrila Saha, Alan C. Bovik |
NeurIPS | 3 |
| 2024 | Bitrate Ladder Construction Using Visual Information FidelityabstractRecently proposed perceptually optimized per-title video encoding methods provide better BD-rate savings than fixed bitrate-ladder approaches that have been employed in the past. However, a disadvantage of per-title encoding is that it requires significant time and energy to compute bitrate ladders. Over the past few years, a variety of methods have been proposed to construct optimal bitrate ladders including using low-level features to predict cross-over bitrates, optimal resolutions for each bitrate, predicting visual quality, etc. Here, we deploy features drawn from Visual Information Fidelity (VIF) (VIF features) extracted from uncompressed videos to predict the visual quality (VMAF) of compressed videos. We present multiple VIF feature sets extracted from different scales and subbands of a video to tackle the problem of bitrate ladder construction. Comparisons are made against a fixed bitrate ladder and a bitrate ladder obtained from exhaustive encoding using Bjontegaard delta metrics. Krishna Srikar Durbha, Hassene Tmar, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik |
PCS | 5 |
| 2024 | A FUNQUE Approach to the Quality Assessment of Compressed HDR VideosabstractRecent years have seen steady growth in the popularity and availability of High Dynamic Range (HDR) content, particularly videos, streamed over the internet. As a result, assessing the subjective quality of HDR videos, which are generally subjected to compression, is of increasing importance. In particular, we target the task of full-reference quality assessment of compressed HDR videos. The state-of-the-art (SOTA) approach HDRMAX involves augmenting off-the-shelf video quality models, such as VMAF, with features computed on nonlinearly transformed video frames. However, HDRMAX increases the computational complexity of models like VMAF. Here, we show that an efficient class of video quality prediction models named FUNQUE+ achieves SOTA accuracy. This shows that the FUNQUE+ models are flexible alternatives to VMAF that achieve higher HDR video quality prediction accuracy at lower computational cost. Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik |
PCS | 4 |
| 2024 | 3D-PSSIM: Projective Structural Similarity for 3D Mesh Quality Assessment Robust to Topological IrregularitiesabstractDespite acceleration in the use of 3D meshes, it is difficult to find effective mesh quality assessment algorithms that can produce predictions highly correlated with human subjective opinions. Defining mesh quality features is challenging due to the irregular topology of meshes, which are defined on vertices and triangles. To address this, we propose a novel 3D projective structural similarity index ( 3D- PSSIM) for meshes that is robust to differences in mesh topology. We address topological differences between meshes by introducing multi-view and multi-layer projections that can densely represent the mesh textures and geometrical shapes irrespective of mesh topology. It also addresses occlusion problems that occur during projection. We propose visual sensitivity weights that capture the perceptual sensitivity to the degree of mesh surface curvature. 3D- PSSIM computes perceptual quality predictions by aggregating quality-aware features that are computed in multiple projective spaces onto the mesh domain, rather than on 2D spaces. This allows 3D- PSSIM to determine which parts of a mesh surface are distorted by geometric or color impairments. Experimental results show that 3D- PSSIM can predict mesh quality with high correlation against human subjective judgments, across the presence of noise, even when there are large topological differences, outperforming existing mesh quality assessment models. Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Weisi Lin, Alan C. Bovik |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Learned fractional downsampling network for adaptive video streaming
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Chao Chen 0006, Alan C. Bovik |
Signal Process. Image Commun. | 6 |
| 2024 | HDR-ChipQA: No-reference quality assessment on High Dynamic Range videos
Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Sriram Sethuraman, Alan C. Bovik |
Signal Process. Image Commun. | 6 |
| 2024 | FAVER: Blind quality prediction of variable frame rate videos
Qi Zheng 0004, Zhengzhong Tu, Pavan C. Madhusudana, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan |
Signal Process. Image Commun. | 5 |
| 2024 | Subjective and Objective Analysis of Streamed Gaming VideosabstractThe rising popularity of online User-Generated-Content (UGC) in the form of streamed and shared videos, has hastened the development of perceptual Video Quality Assessment (VQA) models, which can be used to help optimize their delivery. Gaming videos, which are a relatively new type of UGC videos, are created when skilled and casual gamers post videos of their gameplay. These kinds of screenshots of UGC gameplay videos have become extremely popular on major streaming platforms like YouTube and Twitch. Synthetically-generated gaming content presents challenges to existing VQA algorithms, including those based on natural scene/video statistics models. Synthetically generated gaming content presents different statistical behavior than naturalistic videos. A number of studies have been directed towards understanding the perceptual characteristics of professionally generated gaming videos arising in gaming video streaming, online gaming, and cloud gaming. However, little work has been done on understanding the quality of UGC gaming videos, and how it can be characterized and predicted. Towards boosting the progress of gaming video VQA model development, we conducted a comprehensive study of subjective and objective VQA models on UGC gaming videos. To do this, we created a novel UGC gaming video resource, called the LIVE-YouTube Gaming video quality (LIVE-YT-Gaming) database, comprised of 600 real UGC gaming videos. We conducted a subjective human study on this data, yielding 18,600 human quality ratings recorded by 61 human subjects. We also evaluated a number of state-of-the-art (SOTA) VQA models on the new database, including a new one, called GAME-VQP, based on both natural video statistics and CNN-learned features. To help support work in this field, we are making the new LIVE-YT-Gaming Database, along with code for GAME-VQP, publicly available through the link:https://live.ece.utexas.edu/research/LIVE-YT-Gaming/index.html. Xiangxu Yu, Zhenqiang Ying, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Games | 6 |
| 2024 | Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual RealityabstractWe study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We also study the ability of video quality models to predict human judgments. As streaming human avatar videos in VR or AR become increasingly common, the need for more advanced human avatar video compression protocols will be required to address the tradeoffs between faithfully transmitting high-quality visual representations while adjusting to changeable bandwidth scenarios. During transmission over the internet, the perceived quality of compressed human avatar videos can be severely impaired by visual artifacts. To optimize trade-offs between perceptual quality and data volume in practical workflows, video quality assessment (VQA) models are essential tools. However, there are very few VQA algorithms developed specifically to analyze human body avatar videos, due, at least in part, to the dearth of appropriate and comprehensive datasets of adequate size. Towards filling this gap, we introduce the LIVE-Meta Rendered Human Avatar VQA Database, which contains 720 human avatar videos processed using 20 different combinations of encoding parameters, labeled by corresponding human perceptual quality judgments that were collected in six degrees of freedom VR headsets. To demonstrate the usefulness of this new and unique video resource, we use it to study and compare the performances of a variety of state-of-the-art Full Reference and No Reference video quality prediction models, including a new model called HoloQA. As a service to the research community, we publicly releases the metadata of the new database at https://live.ece.utexas.edu/research/LIVE-Meta-rendered-human-avatar/index.html. Yu-Chih Chen, Avinab Saha, Alexandre Chapiro, Christian Häne, Jean-Charles Bazin, Stefano Zanetti, Ioannis Katsavounidis, Alan C. Bovik |
IEEE Trans. Image Process. | 9 |
| 2024 | HDR or SDR? A Subjective and Objective Study of Scaled and Compressed VideosabstractWe conducted a large-scale study of human perceptual quality judgments of High Dynamic Range (HDR) and Standard Dynamic Range (SDR) videos subjected to scaling and compression levels and viewed on three different display devices. While conventional expectations are that HDR quality is better than SDR quality, we have found subject preference of HDR versus SDR depends heavily on the display device, as well as on resolution scaling and bitrate. To study this question, we collected more than 23,000 quality ratings from 67 volunteers who watched 356 videos on OLED, QLED, and LCD televisions, and among many other findings, observed that HDR videos were often rated as lower quality than SDR videos at lower bitrates, particularly when viewed on LCD and QLED displays. Since it is of interest to be able to measure the quality of videos under these scenarios, e.g. to inform decisions regarding scaling, compression, and SDR vs HDR, we tested several well-known full-reference and no-reference video quality models on the new database. Towards advancing progress on this problem, we also developed a novel no-reference model called HDRPatchMAX, that uses a contrast-based analysis of classical and bit-depth features to predict quality more accurately than existing metrics. Joshua P. Ebenezer, Zaixi Shang, Yixu Chen, Hai Wei, Sriram Sethuraman, Alan C. Bovik |
IEEE Trans. Image Process. | 7 |
| 2024 | Dual-Stream Complex-Valued Convolutional Network for Authentic Dehazed Image Quality AssessmentabstractEffectively evaluating the perceptual quality of dehazed images remains an under-explored research issue. In this paper, we propose a no-reference complex-valued convolutional neural network (CV-CNN) model to conduct automatic dehazed image quality evaluation. Specifically, a novel CV-CNN is employed that exploits the advantages of complex-valued representations, achieving better generalization capability on perceptual feature learning than real-valued ones. To learn more discriminative features to analyze the perceptual quality of dehazed images, we design a dual-stream CV-CNN architecture. The dual-stream model comprises a distortion-sensitive stream that operates on the dehazed RGB image, and a haze-aware stream on a novel dark channel difference image. The distortion-sensitive stream accounts for perceptual distortion artifacts, while the haze-aware stream addresses the possible presence of residual haze. Experimental results on three publicly available dehazed image quality assessment (DQA) databases demonstrate the effectiveness and generalization of our proposed CV-CNN DQA model as compared to state-of-the-art no-reference image quality assessment algorithms. Tuxin Guan, Yuhui Zheng, Xiaojun Wu 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2024 | Convex Hull Prediction for Adaptive Video Streaming by Recurrent LearningabstractAdaptive video streaming relies on the construction of efficient bitrate ladders to deliver the best possible visual quality to viewers under bandwidth constraints. The traditional method of content dependent bitrate ladder selection requires a video shot to be pre-encoded with multiple encoding parameters to find the optimal operating points given by the convex hull of the resulting rate-quality curves. However, this pre-encoding step is equivalent to an exhaustive search process over the space of possible encoding parameters, which causes significant overhead in terms of both computation and time expenditure. To reduce this overhead, we propose a deep learning based method of content aware convex hull prediction. We employ a recurrent convolutional network (RCN) to implicitly analyze the spatiotemporal complexity of video shots in order to predict their convex hulls. A two-step transfer learning scheme is adopted to train our proposed RCN-Hull model, which ensures sufficient content diversity to analyze scene complexity, while also making it possible to capture the scene statistics of pristine source videos. Our experimental results reveal that our proposed model yields better approximations of the optimal convex hulls, and offers competitive time savings as compared to existing approaches. On average, the pre-encoding time was reduced by 53.8% by our method, while the average Bjøntegaard delta bitrate (BD-rate) of the predicted convex hulls against ground truth was 0.26%, and the mean absolute deviation of the BD-rate distribution was 0.57%. Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2024 | A Study of Subjective and Objective Quality Assessment of HDR VideosabstractAs compared to standard dynamic range (SDR) videos, high dynamic range (HDR) content is able to represent and display much wider and more accurate ranges of brightness and color, leading to more engaging and enjoyable visual experiences. HDR also implies increases in data volume, further challenging existing limits on bandwidth consumption and on the quality of delivered content. Perceptual quality models are used to monitor and control the compression of streamed SDR content. A similar strategy should be useful for HDR content, yet there has been limited work on building HDR video quality assessment (VQA) algorithms. One reason for this is a scarcity of high-quality HDR VQA databases representative of contemporary HDR standards. Towards filling this gap, we created the first publicly available HDR VQA database dedicated to HDR10 videos, called the Laboratory for Image and Video Engineering (LIVE) HDR Database. It comprises 310 videos from 31 distinct source sequences processed by ten different compression and resolution combinations, simulating bitrate ladders used by the streaming industry. We used this data to conduct a subjective quality study, gathering more than 20,000 human quality judgments under two different illumination conditions. To demonstrate the usefulness of this new psychometric data resource, we also designed a new framework for creating HDR quality sensitive features, using a nonlinear transform to emphasize distortions occurring in spatial portions of videos that are enhanced by HDR, e.g., having darker blacks and brighter whites. We apply this new method, which we call HDRMAX, to modify the widely-deployed Video Multimethod Assessment Fusion (VMAF) model. We show that VMAF+HDRMAX provides significantly elevated performance on both HDR and SDR videos, exceeding prior state-of-the-art model performance. The database is now accessible at: https://live.ece.utexas.edu/research/LIVEHDR/LIVEHDR_index.html. The model will be made available at a later date at: https://live.ece.utexas.edu//research/Quality/index_algorithms.htm. Zaixi Shang, Joshua P. Ebenezer, Abhinau Kumar Venkataramanan, Hai Wei, Sriram Sethuraman, Alan C. Bovik |
IEEE Trans. Image Process. | 7 |
| 2024 | Subjective Quality Assessment of Compressed Tone-Mapped High Dynamic Range VideosabstractHigh Dynamic Range (HDR) videos are able to represent wider ranges of contrasts and colors than Standard Dynamic Range (SDR) videos, giving more vivid experiences. Due to this, HDR videos are expected to grow into the dominant video modality of the future. However, HDR videos are incompatible with existing SDR displays, which form the majority of affordable consumer displays on the market. Because of this, HDR videos must be processed by tone-mapping them to reduced bit-depths to service a broad swath of SDR-limited video consumers. Here, we analyze the impact of tone-mapping operators on the visual quality of streaming HDR videos. To this end, we built the first large-scale subjectively annotated open-source database of compressed tone-mapped HDR videos, containing 15,000 tone-mapped sequences derived from 40 unique HDR source contents. The videos in the database were labeled with more than 750,000 subjective quality annotations, collected from more than 1,600 unique human observers. We demonstrate the usefulness of the new subjective database by benchmarking objective models of visual quality on it. We envision that the new LIVE Tone-Mapped HDR (LIVE-TMHDR) database will enable significant progress on HDR video tone mapping and quality assessment in the future. To this end, we make the database freely available to the community at https://live.ece.utexas.edu/research/LIVE_TMHDR/index.html. Abhinau Kumar Venkataramanan, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2024 | One Transform to Compute Them All: Efficient Fusion-Based Full-Reference Video Quality AssessmentabstractThe Visual Multimethod Assessment Fusion (VMAF) algorithm has recently emerged as a state-of-the-art approach to video quality prediction, that now pervades the streaming and social media industry. However, since VMAF requires the evaluation of a heterogeneous set of quality models, it is computationally expensive. Given other advances in hardware-accelerated encoding, quality assessment is emerging as a significant bottleneck in video compression pipelines. Towards alleviating this burden, we propose a novel Fusion of Unified Quality Evaluators (FUNQUE) framework, by enabling computation sharing and by using a transform that is sensitive to visual perception to boost accuracy. Further, we expand the FUNQUE framework to define a collection of improved low-complexity fused-feature models that advance the state-of-the-art of video quality performance with respect to both accuracy, by 4.2% to 5.3%, and computational efficiency, by factors of 3.8 to 11 times!. Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2024 | MWFormer: Multi-Weather Image Restoration Using Degradation-Aware TransformersabstractRestoring images captured under adverse weather conditions is a fundamental task for many computer vision applications. However, most existing weather restoration approaches are only capable of handling a specific type of degradation, which is often insufficient in real-world scenarios, such as rainy-snowy or rainy-hazy weather. Towards being able to address these situations, we propose a multi-weather Transformer, or MWFormer for short, which is a holistic vision Transformer that aims to solve multiple weather-induced degradations using a single, unified architecture. MWFormer uses hyper-networks and feature-wise linear modulation blocks to restore images degraded by various weather types using the same set of learned parameters. We first employ contrastive learning to train an auxiliary network that extracts content-independent, distortion-aware feature embeddings that efficiently represent predicted weather types, of which more than one may occur. Guided by these weather-informed predictions, the image restoration Transformer adaptively modulates its parameters to conduct both local and global feature processing, in response to multiple possible weather. Moreover, MWFormer allows for a novel way of tuning, during application, to either a single type of weather restoration or to hybrid weather restoration without any retraining, offering greater controllability than existing methods. Our experimental results on multi-weather restoration benchmarks show that MWFormer achieves significant performance improvements compared to existing state-of-the-art methods, without requiring much computational cost. Moreover, we demonstrate that our methodology of using hyper-networks can be integrated into various network architectures to further boost their performance. The code is available at: https://github.com/taco-group/MWFormer. Ruoxi Zhu, Zhengzhong Tu, Alan C. Bovik, Yibo Fan |
IEEE Trans. Image Process. | 4 |
| 2024 | On the generation of adversarial examples for image quality assessment
Qingbing Sang, Hongguo Zhang, Lixiong Liu, Xiaojun Wu 0001, Alan C. Bovik |
Vis. Comput. | 5 |
| 2023 | Re-IQA: Unsupervised Learning for Image Quality Assessment in the WildabstractAutomatic Perceptual Image Quality Assessment is a challenging problem that impacts billions of internet, and social media users daily. To advance research in this field, we propose a Mixture of Experts approach to train two separate encoders to learn high-level content and low-level image quality features in an unsupervised setting. The unique novelty of our approach is its ability to generate low-level representations of image quality that are complementary to high-level features representing image content. We refer to the framework used to train the two encoders as Re-IQA. For Image Quality Assessment in the Wild, we deploy the complementary low and high-level image representations obtained from the Re-IQA framework to train a linear regression model, which is used to map the image representations to the ground truth quality scores, refer Figure 1. Our method achieves state-of-the-art performance on multiple large-scale image quality assessment databases containing both real and synthetic distortions, demonstrating how deep neural networks can be trained in an unsupervised setting to produce perceptually relevant representations. We conclude from our experiments that the low and high-level features obtained are indeed complementary and positively impact the performance of the linear regressor. A public release of all the codes associated with this work will be made available on GitHub. Avinab Saha, Sandeep Mishra, Alan C. Bovik |
CVPR | 3 |
| 2023 | Pik-Fix: Restoring and Colorizing Old PhotosabstractRestoring and inpainting the visual memories that are present, but often impaired, in old photos remains an intriguing but unsolved research topic. Decades-old photos often suffer from severe and commingled degradation such as cracks, defocus, and color-fading, which are difficult to treat individually and harder to repair when they interact. Deep learning presents a plausible avenue, but the lack of large-scale datasets of old photos makes addressing this restoration task very challenging. Here we present a novel reference-based end-to-end learning framework that is able to both repair and colorize old, degraded pictures. Our proposed framework consists of three modules: a restoration sub-network that conducts restoration from degradations, a similarity network that performs color histogram matching and color transfer, and a colorization subnet that learns to predict the chroma elements of images conditioned on chromatic reference signals. The overall system makes uses of color histogram priors from reference images, which greatly reduces the need for large-scale training data. We have also created a first-of-a-kind public dataset of real old photos that are paired with ground truth "pristine" photos that have been manually restored by PhotoShop experts. We conducted extensive experiments on this dataset and synthetic datasets, and found that our method significantly outperforms previous state-of-the-art models using both qualitative comparisons and quantitative measurements. The code is available at https://github.com/DerrickXuNu/Pik-Fix. Runsheng Xu, Zhengzhong Tu, Yuanqi Du, Zibo Meng, Jiaqi Ma 0003, Alan C. Bovik, Hongkai Yu |
WACV | 8 |
| 2023 | GAMIVAL: Video Quality Prediction on Mobile Cloud Gaming ContentabstractThe mobile cloud gaming industry has been rapidly growing over the last decade. When streaming gaming videos are transmitted to customers' client devices from cloud servers, algorithms that can monitor distorted video quality without having any reference video available are desirable tools. However, creating No-Reference Video Quality Assessment (NR VQA) models that can accurately predict the quality of streaming gaming videos rendered by computer graphics engines is a challenging problem, since gaming content generally differs statistically from naturalistic videos, often lacks detail, and contains many smooth regions. Until recently, the problem has been further complicated by the lack of adequate subjective quality databases of mobile gaming content. We have created a new gaming-specific NR VQA model called the Gaming Video Quality Evaluator (GAMIVAL), which combines and leverages the advantages of spatial and temporal gaming distorted scene statistics models, a neural noise model, and deep semantic features. Using a support vector regression (SVR) as a regressor, GAMIVAL achieves superior performance on the new LIVE-Meta Mobile Cloud Gaming (LIVE-Meta MCG) video quality database. Yu-Chih Chen, Avinab Saha, Chase Davis, Rahul Gowda, Ioannis Katsavounidis, Alan C. Bovik |
IEEE Signal Process. Lett. | 8 |
| 2023 | CONVIQT: Contrastive Video Quality EstimatorabstractPerceptual video quality assessment (VQA) is an integral component of many streaming and video sharing platforms. Here we consider the problem of learning perceptually relevant video quality representations in a self-supervised manner. Distortion type identification and degradation level determination is employed as an auxiliary task to train a deep learning model containing a deep Convolutional Neural Network (CNN) that extracts spatial features, as well as a recurrent unit that captures temporal information. The model is trained using a contrastive loss and we therefore refer to this training framework and resulting model as CONtrastive VIdeo Quality EstimaTor (CONVIQT). During testing, the weights of the trained model are frozen, and a linear regressor maps the learned features to quality scores in a no-reference (NR) setting. We conduct comprehensive evaluations of the proposed model against leading algorithms on multiple VQA databases containing wide ranges of spatial and temporal distortions. We analyze the correlations between model predictions and ground-truth quality ratings, and show that CONVIQT achieves competitive performance when compared to state-of-the-art NR-VQA models, even though it is not trained on those databases. Our ablation experiments demonstrate that the learned representations are highly robust and generalize well across synthetic and realistic distortions. Our results indicate that compelling representations with perceptual bearing can be obtained using self-supervised learning. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2023 | Helping Visually Impaired People Take Better Quality PicturesabstractPerception-based image analysis technologies can be used to help visually impaired people take better quality pictures by providing automated guidance, thereby empowering them to interact more confidently on social media. The photographs taken by visually impaired users often suffer from one or both of two kinds of quality issues: technical quality (distortions), and semantic quality, such as framing and aesthetic composition. Here we develop tools to help them minimize occurrences of common technical distortions, such as blur, poor exposure, and noise. We do not address the complementary problems of semantic quality, leaving that aspect for future work. The problem of assessing, and providing actionable feedback on the technical quality of pictures captured by visually impaired users is hard enough, owing to the severe, commingled distortions that often occur. To advance progress on the problem of analyzing and measuring the technical quality of visually impaired user-generated content (VI-UGC), we built a very large and unique subjective image quality and distortion dataset. This new perceptual resource, which we call the LIVE-Meta VI-UGC Database, contains 40K real-world distorted VI-UGC images and 40K patches, on which we recorded 2.7M human perceptual quality judgments and 2.7M distortion labels. Using this psychometric resource we also created an automatic limited vision picture quality and distortion predictor that learns local-to-global spatial quality relationships, achieving state-of-the-art prediction performance on VI-UGC pictures, significantly outperforming existing picture quality models on this unique class of distorted picture data. We also created a prototype feedback system that helps to guide users to mitigate quality issues and take better quality pictures, by creating a multi-task learning framework. The dataset and models can be accessed at: https://github.com/mandal-cv/visimpaired. Maniratnam Mandal, Deepti Ghadiyaram, Danna Gurari, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2023 | Self-Supervised Learning of Perceptually Optimized Block Motion Estimates for Video CompressionabstractBlock based motion estimation is integral to inter prediction processes performed in hybrid video codecs. Prevalent block matching based methods that are used to compute block motion vectors (MVs) rely on computationally intensive search procedures. They also suffer from the aperture problem, which tends to worsen as the block size is reduced. Moreover, the block matching criteria used in typical codecs do not account for the resulting levels of perceptual quality of the motion compensated pictures that are created upon decoding. Towards achieving the elusive goal of perceptually optimized motion estimation, we propose a search-free block motion estimation framework using a multi-stage convolutional neural network, which is able to conduct motion estimation on multiple block sizes simultaneously, using a triplet of frames as input. This composite block translation network (CBT-Net) is trained in a self-supervised manner on a large database that we created from publicly available uncompressed video content. We deploy the multi-scale structural similarity (MS-SSIM) loss function to optimize the perceptual quality of the motion compensated predicted frames. Our experimental results highlight the computational efficiency of our proposed model relative to conventional block matching based motion estimation algorithms, for comparable prediction errors. Further, when used to perform inter prediction in AV1, the MV predictions of the perceptually optimized model result in average Bjøntegaard-delta rate (BD-rate) improvements of -1.73% and -1.31% with respect to the MS-SSIM and Video Multi-Method Assessment Fusion (VMAF) quality metrics, respectively, as compared to the block matching based motion estimation system employed in the SVT-AV1 encoder. Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2023 | Study of Subjective and Objective Quality Assessment of Mobile Cloud Gaming VideosabstractWe present the outcomes of a recent large-scale subjective study of Mobile Cloud Gaming Video Quality Assessment (MCG-VQA) on a diverse set of gaming videos. Rapid advancements in cloud services, faster video encoding technologies, and increased access to high-speed, low-latency wireless internet have all contributed to the exponential growth of the Mobile Cloud Gaming industry. Consequently, the development of methods to assess the quality of real-time video feeds to end-users of cloud gaming platforms has become increasingly important. However, due to the lack of a large-scale public Mobile Cloud Gaming Video dataset containing a diverse set of distorted videos with corresponding subjective scores, there has been limited work on the development of MCG-VQA models. Towards accelerating progress towards these goals, we created a new dataset, named the LIVE-Meta Mobile Cloud Gaming (LIVE-Meta-MCG) video quality database, composed of 600 landscape and portrait gaming videos, on which we collected 14,400 subjective quality ratings from an in-lab subjective study. Additionally, to demonstrate the usefulness of the new resource, we benchmarked multiple state-of-the-art VQA algorithms on the database. The new database will be made publicly available on our website: https://live.ece.utexas.edu/research/LIVE-Meta-Mobile-Cloud-Gaming/index.html. Avinab Saha, Yu-Chih Chen, Chase Davis, Rahul Gowda, Ioannis Katsavounidis, Alan C. Bovik |
IEEE Trans. Image Process. | 8 |
| 2022 | MAXIM: Multi-Axis MLP for Image ProcessingabstractRecent progress on Transformers and multilayer perceptron (MLP) models provide new network architectural designs for computer vision tasks. Although these models proved to be effective in many vision tasks such as image recognition, there remain challenges in adapting them for lowlevel vision. The inflexibility to support high-resolution images and limitations of local attention are perhaps the main bottlenecks. In this work, we present a multi-axis MLP based architecture called MAXIM, that can serve as an efficient and flexible general-purpose vision backbone for image processing tasks. MAXIM uses a UNet-shaped hierarchical structure and supports long-range interactions enabled by spatially-gated MLPs. Specifically, MAXIM contains two MLP-based building blocks: a multi-axis gated MLP that allows for efficient and scalable spatial mixing of local and global visual cues, and a cross-gating block, an alternative to cross-attention, which accounts for cross-feature conditioning. Both these modules are exclusively based on MLPs, but also benefit from being both global and ‘fully-convolutional’, two properties that are desirable for image processing. Our extensive experimental results show that the proposed MAXIM model achieves state-of-the-art performance on more than ten benchmarks across a range of image processing tasks, including denoising, deblurring, de raining, dehazing, and enhancement while requiring fewer or comparable numbers of parameters and FLOPs than competitive models. The source code and trained models will be available at https://github.com/google-research/maxim. Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li |
CVPR | 6 |
| 2022 | MaxViT: Multi-axis Vision Transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li |
ECCV (24) | 6 |
| 2022 | Telepresence Video Quality Assessment
Zhenqiang Ying, Deepti Ghadiyaram, Alan C. Bovik |
ECCV (37) | 3 |
| 2022 | No-Reference Quality Assessment of Variable Frame-Rate Videos Using Temporal Bandpass StatisticsabstractRecent advances in mobile devices and cloud computing techniques have made it possible to capture, process, and share high resolution, high frame rate (HFR) videos across the Internet nearly instantaneously. Being able to monitor and control the quality of these streamed videos can enable the de-livery of many enjoyable content and perceptually optimized rate control. However, the development of no-reference (NR) VQA algorithms targeting frame rate variations has been little studied. Here, we propose a first-of-a-kind blind VQA model for evaluating HFR videos, which we dub the Framerate-Aware Videos Evaluator w/o Reference (FAVER). FAVER uses extended models of spatial natural scene statistics that encompass space-time wavelet-decomposed video signals, to conduct efficient frame rate sensitive quality prediction. Our extensive experiments on several HFR video quality datasets show that FAVER outperforms other blind VQA algorithms at a reasonable computational cost. The code will be released on https://github.com/uniqzheng/HFR-BVQA. Qi Zheng 0004, Zhengzhong Tu, Yibo Fan, Xiaoyang Zeng, Alan C. Bovik |
ICASSP | 5 |
| 2022 | Subjective Assessment Of High Dynamic Range Videos Under Different Ambient ConditionsabstractHigh Dynamic Range (HDR) videos can represent a much greater range of brightness and color than Standard Dynamic Range (SDR) videos and are rapidly becoming an industry standard. HDR videos have more challenging capture, transmission, and display requirements than legacy SDR videos. With their greater bit depth, advanced electro-optical transfer functions, and wider color gamuts, comes the need for video quality algorithms that are specifically designed to predict the quality of HDR videos. Towards this end, we present the first publicly released large-scale subjective study of HDR videos. We study the effect of distortions such as compression and aliasing on the quality of HDR videos. We also study the effect of ambient illumination on perceptual quality of HDR videos by conducting the study in both a dark lab environment and a brighter living-room environment. A total of 66 subjects participated in the study and more than 20,000 opinion scores were collected, which makes this the largest in-lab study of HDR video quality ever. We anticipate that the dataset will be a valuable resource for researchers to develop better models of perceptual quality for HDR videos. Zaixi Shang, Joshua P. Ebenezer, Alan C. Bovik, Hai Wei, Sriram Sethuraman |
ICIP | 3 |
| 2022 | Funque: Fusion of Unified Quality EvaluatorsabstractFusion-based quality assessment has emerged as a powerful method for developing high-performance quality models from quality models that individually achieve lower performances. A prominent example of such an algorithm is VMAF, which has been widely adopted as an industry standard for video quality prediction along with SSIM. In addition to advancing the state-of-the-art, it is imperative to alleviate the computational burden presented by the use of a heterogeneous set of quality models. In this paper, we unify "atom" quality models by computing them on a common transform domain that accounts for the Human Visual System, and we propose FUNQUE, a quality model that fuses unified quality evaluators. We demonstrate that in comparison to the state-of-the-art, FUNQUE offers significant improvements in both correlation against subjective scores and efficiency, due to computation sharing. Abhinau Kumar Venkataramanan, Cosmin Stejerean, Alan C. Bovik |
ICIP | 3 |
| 2022 | Blind Video Quality Assessment via Space-Time Slice StatisticsabstractUser-generated contents (UGC) have gained increased attention in the video quality community recently. Perceptual video quality assessment (VQA) of UGC videos is of great significance for content providers to monitor, process, and deliver massive numbers of UGC videos. Blind video quality prediction of UGC videos is challenging since complex mixtures of spatial and temporal distortions contribute to the overall perceptual quality. In this paper, we develop a simple, effective, and efficient blind VQA framework (STS-QA) based on the statistical analysis of space-time slices (STS) of videos. Specifically, we extract spatio-temporal statistical features along different orientations of video STS, that capture directional global motion, then train a shallow quality predictor. The proposed framework can be used to easily extend any existing video/image quality model to account for temporal or motion regularities. Our experimental results on three publicly available UGC databases demonstrate that our proposed STS-QA model can significantly boost prediction performance compared to baselines. The code will be released at: https://github.com/uniqzheng/STS_BVQA. Qi Zheng 0004, Zhengzhong Tu, Zhijian Hao, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan |
ICIP | 5 |
| 2022 | Learning to compress videos without computing motion
Meixu Chen, Todd Richard Goodall, Anjul Patney, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2022 | Making Video Quality Assessment Models Sensitive to Frame Rate DistortionsabstractWe consider the problem of capturing distortions arising from changes in frame rate as part of Video Quality Assessment (VQA). Variable frame rate (VFR) videos have become much more common, and streamed videos commonly range from 30 frames per second (fps) up to 120 fps. VFR-VQA offers unique challenges in terms of distortion types as well as in making non-uniform comparisons of reference and distorted videos having different frame rates. The majority of current VQA models require compared videos to be of the same frame rate, but are unable to adequately account for frame rate artifacts. The recently proposed Generalized Entropic Difference (GREED) VQA model succeeds at this task, using natural video statistics models of entropic differences of temporal band-pass coefficients, delivering superior performance on predicting video quality changes arising from frame rate distortions. Here we propose a simple fusion framework, whereby temporal features from GREED are combined with existing VQA models, towards improving model sensitivity towards frame rate distortions. We find through extensive experiments that this feature fusion significantly boosts model performance on both HFR/VFR datasets as well as fixed frame rate (FFR) VQA databases. Our results suggest that employing efficient temporal representations can result much more robust and accurate VQA models when frame rate variations can occur. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 5 |
| 2022 | Completely Blind Video Quality EvaluatorabstractAutomatic video quality assessment of user-generated content (UGC) has gained increased interest recently, due to the ubiquity of shared video clips uploaded and circulated on social media platforms across the globe. Most existing video quality models developed for this vast content are trained on large numbers of samples labeled during large-scale subjective studies, which are often fail to exhibit adequate generalization abilities on unseen data. Moreover, large labeled video quality datasets are not always available for every scenario, and may not address the coincident evaluation of social videos and the distortions that afflict them. Because of this, it is also desirable to develop opinion-unaware, “completely blind” video quality models, that are free of training, yet can compete with existing learning-based models. Here we propose such a model called VIQE (VIdeo Quality Evaluator), which we designed based on a comprehensive analysis of patch- and frame-wise video statistics, as well as of space-time statistical regularities of videos. The statistical features desired from the analysis capture complementary predictive aspects of perceptual quality, which are aggregated to obtain final video quality scores. Extensive experiments on recent large-scale video quality databases demonstrate that VIQE is even competitive with state-of-the-art opinion-aware models. The source code is being made available athttps://github.com/uniqzheng/Complete-Blind-VQA. Qi Zheng 0004, Zhengzhong Tu, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan |
IEEE Signal Process. Lett. | 4 |
| 2022 | FOVQA: Blind Foveated Video Quality AssessmentabstractPrevious blind or No Reference (NR) Image / video quality assessment (IQA/VQA) models largely rely on features drawn from natural scene statistics (NSS), but under the assumption that the image statistics are stationary in the spatial domain. Several of these models are quite successful on standard pictures. However, in Virtual Reality (VR) applications, foveated video compression is regaining attention, and the concept of space-variant quality assessment is of interest, given the availability of increasingly high spatial and temporal resolution contents and practical ways of measuring gaze direction. Distortions from foveated video compression increase with increased eccentricity, implying that the natural scene statistics are space-variant. Towards advancing the development of foveated compression / streaming algorithms, we have devised a no-reference (NR) foveated video quality assessment model, called FOVQA, which is based on new models of space-variant natural scene statistics (NSS) and natural video statistics (NVS). Specifically, we deploy a space-variant generalized Gaussian distribution (SV-GGD) model and a space-variant asynchronous generalized Gaussian distribution (SV-AGGD) model of mean subtracted contrast normalized (MSCN) coefficients and products of neighboring MSCN coefficients, respectively. We devise a foveated video quality predictor that extracts radial basis features, and other features that capture perceptually annoying rapid quality fall-offs. We find that FOVQA achieves state-of-the-art (SOTA) performance on the new 2D LIVE-FBT-FCVR database, as compared with other leading Foveated IQA / VQA models. we have made our implementation of FOVQA available at: https://live.ece.utexas.edu/research/Quality/FOVQA.zip. Yize Jin, Anjul Patney, Richard Webb, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2022 | Video Quality Model of Compression, Resolution and Frame Rate Adaptation Based on Space-Time RegularitiesabstractBeing able to accurately predict the visual quality of videos subjected to various combinations of dimension reduction protocols is of high interest to the streaming video industry, given rapid increases in frame resolutions and frame rates. In this direction, we have developed a video quality predictor that is sensitive to spatial, temporal, or space-time subsampling combined with compression. Our predictor is based on new models of space-time natural video statistics (NVS). Specifically, we model the statistics of divisively normalized difference between neighboring frames that are relatively displaced. In an extensive empirical study, we found that those paths of space-time displaced frame differences that provide maximal regularity against our NVS model generally align best with motion trajectories. Motivated by this, we built a new video quality prediction engine that extracts NVS features that represent how space-time directional regularities are disturbed by space-time distortions. Based on parametric models of these regularities, we compute features that are used to train a regressor that can accurately predict perceptual quality. As a stringent test of the new model, we apply it to the difficult problem of predicting the quality of videos subjected not only to compression, but also to downsampling in space and/or time. We show that the new quality model achieves state-of-the-art (SOTA) prediction performance on the new ETRI-LIVE Space-Time Subsampled Video Quality (STSVQ) and also on the AVT-VQDB-UHD-1 database. Dae Yeol Lee, Hyunsuk Ko, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2022 | A Subjective and Objective Study of Space-Time Subsampled Video QualityabstractVideo dimensions are continuously increasing to provide more realistic and immersive experiences to global streaming and social media viewers. However, increments in video parameters such as spatial resolution and frame rate are inevitably associated with larger data volumes. Transmitting increasingly voluminous videos through limited bandwidth networks in a perceptually optimal way is a current challenge affecting billions of viewers. One recent practice adopted by video service providers is space-time resolution adaptation in conjunction with video compression. Consequently, it is important to understand how different levels of space-time subsampling and compression affect the perceptual quality of videos. Towards making progress in this direction, we constructed a large new resource, called the ETRI-LIVE Space-Time Subsampled Video Quality (ETRI-LIVE STSVQ) database, containing 437 videos generated by applying various levels of combined space-time subsampling and video compression on 15 diverse video contents. We also conducted a large-scale human study on the new dataset, collecting about 15,000 subjective judgments of video quality. We provide a rate-distortion analysis of the collected subjective scores, enabling us to investigate the perceptual impact of space-time subsampling at different bit rates. We also evaluated and compare the performance of leading video quality models on the new database. The new ETRI-LIVE STSVQ database is being made freely available at (https://live.ece.utexas.edu/research/ETRI-LIVE_STSVQ/index.html). Dae Yeol Lee, Somdyuti Paul, Christos G. Bampis, Hyunsuk Ko, Seyoon Jeong, Blake Homan, Alan C. Bovik |
IEEE Trans. Image Process. | 8 |
| 2022 | Image Quality Assessment Using Contrastive LearningabstractWe consider the problem of obtaining image quality representations in a self-supervised manner. We use prediction of distortion type and degree as an auxiliary task to learn features from an unlabeled image dataset containing a mixture of synthetic and realistic distortions. We then train a deep Convolutional Neural Network (CNN) using a contrastive pairwise objective to solve the auxiliary problem. We refer to the proposed training framework and resulting deep IQA model as the CONTRastive Image QUality Evaluator (CONTRIQUE). During evaluation, the CNN weights are frozen and a linear regressor maps the learned representations to quality scores in a No-Reference (NR) setting. We show through extensive experiments that CONTRIQUE achieves competitive performance when compared to state-of-the-art NR image quality models, even without any additional fine-tuning of the CNN backbone. The learned representations are highly robust and generalize well across images afflicted by either synthetic or authentic distortions. Our results suggest that powerful quality representations with perceptual relevance can be obtained without requiring large labeled subjective image quality datasets. The implementations used in this paper are available at https://github.com/pavancm/CONTRIQUE. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2022 | Study of the Subjective and Objective Quality of High Motion Live Streaming VideosabstractVideo livestreaming is gaining prevalence among video streaming service s, especially for the delivery of live, high motion content such as sport ing events. The quality of the se livestreaming videos can be adversely affected by any of a wide variety of events, including capture artifacts, and distortions incurred during coding and transmission. High motion content can cause or exacerbate many kinds of distortion, such as motion blur and stutter. Because of this, the development of objective Video Quality Assessment (VQA) algorithms that can predict the perceptual quality of high motion, live streamed videos is greatly desired. Important resources for developing these algorithms are appropriate databases that exemplify the kinds of live streaming video distortions encountered in practice. Towards making progress in this direction, we built a video quality database specifically designed for live streaming VQA research. The new video database is called the Laboratory for Image and Video Engineering (LIVE) Livestream Database. The LIVE Livestream Database includes 315 videos of 45 source sequences from 33 original contents impaired by 6 types of distortions. We also performed a subjective quality study using the new database, whereby more than 12,000 human opinions were gathered from 40 subjects. We demonstrate the usefulness of the new resource by performing a holistic evaluation of the performance of current state-of-the-art (SOTA) VQA models. We envision that researchers will find the dataset to be useful for the development, testing, and comparison of future VQA models. The LIVE Livestream database is being made publicly available for these purposes at https://live.ece. utexas.edu/research/LIVE_APV_Study/apv_index.html. Zaixi Shang, Joshua P. Ebenezer, Hai Wei, Sriram Sethuraman, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2021 | Patch-VQ: 'Patching Up' the Video Quality ProblemabstractNo-reference (NR) perceptual video quality assessment (VQA) is a complex, unsolved, and important problem for social and streaming media applications. Efficient and accurate video quality predictors are needed to monitor and guide the processing of billions of shared, often imperfect, user-generated content (UGC). Unfortunately, current NR models are limited in their prediction capabilities on real-world, "in-the-wild" UGC video data. To advance progress on this problem, we created the largest (by far) subjective video quality dataset, containing 38,811 real-world distorted videos and 116,433 space-time localized video patches (‘v-patches’), and 5.5M human perceptual quality annotations. Using this, we created two unique NR-VQA models: (a) a local-to-global region-based NR VQA architecture (called PVQ) that learns to predict global video quality and achieves state-of-the-art performance on 3 UGC datasets, and (b) a first-of-a-kind space-time video quality mapping engine (called PVQ Mapper) that helps localize and visualize perceptual distortions in space and time. The entire dataset and prediction models are freely available at https://live.ece.utexas.edu/research.php. Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, Alan C. Bovik |
CVPR | 4 |
| 2021 | Improved Intra Mode Coding Beyond Av1abstractIn AOMedia Video 1 (AV1), directional intra prediction modes are applied to model local texture patterns that present certain directionality. Each intra prediction direction is represented with a nominal mode index and a delta angle. The delta angle is entropy coded using shared context between luma and chroma, and the context is derived using the associated nominal mode. In this paper, two methods are proposed to further reduce the signaling cost of delta angles: cross-component delta angle coding, and context-adaptive delta angle coding, whereby the cross-component and spatial correlation of the delta angles are explored, respectively. The proposed methods were implemented on top of a recent version of libaom. Experimental results show that the proposed cross-component delta angle coding achieved average 0.4% BD-rate reduction with 4% encoding time saving over all intra configurations. By combining both methods, an average 1.2% BD-rate reduction is achieved. Yize Jin, Xin Zhao 0003, Shan Liu 0001, Alan C. Bovik |
ICASSP | 5 |
| 2021 | Regression or classification? New methods to evaluate no-reference picture and video quality modelsabstractVideo and image quality assessment has long been projected as a regression problem, which requires predicting a continuous quality score given an input stimulus. However, recent efforts have shown that accurate quality score regression on real-world user-generated content (UGC) is a very challenging task. To make the problem more tractable, we propose two new methods - binary, and ordinal classification - as alternatives to evaluate and compare no-reference quality models at coarser levels. Moreover, the proposed new tasks convey more practical meaning on perceptually optimized UGC transcoding, or for preprocessing on media processing platforms. We conduct a comprehensive benchmark experiment of popular no-reference quality models on recent in-the-wild picture and video quality datasets, providing reliable baselines for both evaluation methods to support further studies. We hope this work promotes coarse-grained perceptual modeling and its applications to efficient UGC processing. Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICASSP | 7 |
| 2021 | A Foveated Video Quality Assessment Model Using Space-Variant Natural Scene StatisticsabstractIn Virtual Reality (VR) systems, head mounted displays (HMDs) are widely used to present VR contents. When displaying immersive (360° video) scenes, greater challenges arise due to limitations of computing power, frame rate, and transmission bandwidth. To address these problems, a variety of foveated video compression and streaming methods have been proposed, which seek to exploit the nonuniform sampling density of the retinal photoreceptors and ganglion cells, which decreases rapidly with increasing eccentricity. Creating foveated immersive video content leads to the need for specialized foveated video quality pridictors. Here we propose a No-Reference (NR or blind) method which we call “Space-Variant BRISQUE (SV-BRISQUE),” which is based on a new space-variant natural scene statistics model. When tested on a large database of foveated, compression-distorted videos along with human opinions of them, we found that our new model algorithm achieves state of the art (SOTA) performance with correlation 0.88 / 0.90 (PLCC / SROCC) against human subjectivity. Y. Jin, Todd Richard Goodall, Anjul Patney, R. Webb, Alan C. Bovik |
ICIP | 5 |
| 2021 | Video Quality Assessment of User Generated Content: A Benchmark Study and a New ModelabstractRecent years have witnessed an explosion of user-generated content (UGC) shared and streamed over the Internet. Accordingly, there is a great need for accurate video quality assessment (VQA) models for consumer or UGC videos to monitor, control, and optimize this vast content. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading blind VQA (BVQA) models. Besides, we also created a new fusion-based BVQA model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at a lower computational cost. We believe our reliable and reproducible benchmark will facilitate further research on deep learning-based BVQA modeling. An implementation of VIDEVAL has been made available online1.1https://github.com/vztu/VIDEVAL_release Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 6 |
| 2021 | A Temporal Statistics Model For UGC Video Quality PredictionabstractBlind video quality assessment of user-generated content (UGC) has become a trending and challenging problem. Previous studies have shown the efficacy of natural scene statistics for capturing spatial distortions. The exploration of temporal video statistics on UGC, however, is relatively limited. Here we propose the first general, effective and efficient temporal statistics model accounting for temporal- or motion-related distortions for UGC video quality assessment, by analyzing regularities in the temporal bandpass domain. The proposed temporal model can serve as a plug-in module to boost existing no-reference video quality predictors that lack motion-relevant features. Our experimental results on recent large-scale UGC video databases show that the proposed model can significantly improve the performances of existing methods, at a very reasonable computational expense. Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 6 |
| 2021 | A Progressive Architecture for Learned Fractional DownsamplingabstractIn many image and video processing applications, the ability to resize by a fractional factor, such as from 1080p to 720p, is essential. However, conventional CNN layers can only be used to alter the resolution of their inputs with integer scale factors. In this paper, we propose a downsampling network architecture that progressively reconstructs residuals at different scales. In particular, the aforementioned problem is solved by combining an upsampling sub-network and a downsampling subnetwork, both with integer scale factor. As an application, we apply the proposed downsampling network to an adaptive bitrate video streaming scenario. We extensively evaluate with different video codecs and upsampling algorithms to show the generality of our model. Our experimental results show that improvements in coding efficiency over the conventional Lanczos downsampling and state-of-the-art methods are attained, measured in different perceptual video quality models on large-resolution test videos. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Alan C. Bovik |
PCS | 5 |
| 2021 | MOVI-Codec: Deep Video Compression without MotionabstractTraditional video codecs follow the predictive coding architecture of motion-compensated prediction and residual transform coding. Inspired by recent advances in deep learning, we propose a new deep learning video compression architecture that does not require motion estimation, which is the most expensive component in traditional video codecs. Our network consists of three components: a Displacement Calculation Unit (DCU), a Displacement Compression Network (DCN), and a Frame Reconstruction Network (FRN). The DCU exploits displaced frame differences as motion information, thus removing the need for motion estimation found in hybrid codecs. DCN utilizes an RNN-based network to learn temporal dependencies between frames. In the FRN, a new version of the UNet model, called LSTM-UNet is proposed and utilized to learn space-time differential representations of the videos. Our experimental results show that our compression model, MOtionless VIdeo Codec (MOVI-Codec), learns how to efficiently compress videos without computing motion and outperforms the video coding standard H.264 and exceeds the performance of the modern global standard HEVC codec as measured by MS-SSIM, especially on higher resolution videos. Meixu Chen, Anjul Patney, Alan C. Bovik |
PCS | 3 |
| 2021 | Evaluating Foveated Video Quality Using Entropic DifferencingabstractVirtual Reality is regaining attention due to recent advancements in hardware technology. Immersive images / videos are becoming widely adopted to carry omnidirectional visual information. However, due to the requirements for higher spatial and temporal resolution of real video data, immersive videos require significantly larger bandwidth consumption. To reduce stresses on bandwidth, foveated video compression is regaining popularity, whereby the space-variant spatial resolution of the retina is exploited. Towards advancing the progress of foveated video compression, we propose a full reference (FR) foveated image quality assessment algorithm, which we call foveated entropic differencing (FED), which employs the natural scene statistics of bandpass responses by applying differences of local entropies weighted by a foveation-based error sensitivity function. We evaluate the proposed algorithm by measuring the correlations of the predictions that FED makes against human judgements on the newly created 2D and 3D LIVE-FBT-FCVR databases for Virtual Reality (VR). The performance of the proposed algorithm yields state-of-the-art as compared with other existing full reference algorithms. Software for FED has been made available at: http://live.ece.utexas.edu/research/Quality/FED.zip Yize Jin, Anjul Patney, Alan C. Bovik |
PCS | 3 |
| 2021 | High Frame Rate Video Quality Assessment using VMAF and Entropic DifferencesabstractThe popularity of streaming videos with live, high-action content has led to an increased interest in High Frame Rate (HFR) videos. In this work we address the problem of frame rate dependent Video Quality Assessment (VQA) when the videos to be compared have different frame rate and compression factor. The current VQA models such as VMAF have superior correlation with perceptual judgments when videos to be compared have same frame rates and contain conventional distortions such as compression, scaling etc. However this framework requires additional pre-processing step when videos with different frame rates need to be compared, which can potentially limit its overall performance. Recently, Generalized Entropic Difference (GREED) VQA model was proposed to account for artifacts that arise due to changes in frame rate, and showed superior performance on the LIVE-YT-HFR database which contains frame rate dependent artifacts such as judder, strobing etc. In this paper we propose a simple extension, where the features from VMAF and GREED are fused in order to exploit the advantages of both models. We show through various experiments that the proposed fusion framework results in more efficient features for predicting frame rate dependent video quality. We also evaluate the fused feature set on standard non-HFR VQA databases and obtain superior performance than both GREED and VMAF, indicating the combined feature set captures complimentary perceptual quality information. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
PCS | 5 |
| 2021 | Assessment of Subjective and Objective Quality of Live Streaming Sports VideosabstractVideo live streaming is gaining prevalence among video streaming services, especially for the delivery of popular sporting events. Many objective Video Quality Assessment (VQA) models have been developed to predict the perceptual quality of videos. Appropriate databases that exemplify the distortions encountered in live streaming videos are important to designing and learning objective VQA models. Towards making progress in this direction, we built a video quality database specifically designed for live streaming VQA research. The new video database is called the Laboratory for Image and Video Engineering (LIVE) Live stream Database. The LIVE Livestream Database includes 315 videos of 45 contents impaired by 6 types of distortions. We also performed a subjective quality study using the new database, whereby more than 12,000 human opinions were gathered from 40 subjects. We demonstrate the usefulness of the new resource by performing a holistic evaluation of the performance of current state-of-the-art (SOTA) VQA models. The LIVE Livestream database is being made publicly available for these purposes at https://live.ece.utexas.edu/research/LIVE_APV_Study/apv_index.html. Zaixi Shang, Joshua P. Ebenezer, Alan C. Bovik, Hai Wei, Sriram Sethuraman |
PCS | 3 |
| 2021 | Efficient User-Generated Video Quality PredictionabstractBlind video quality assessment of user-generated content (UGC) has become a trending, challenging, unsolved problem. Accurate and efficient video quality predictors suitable for this content are thus in great demand to achieve intelligent analysis and processing of UGC videos. However, previous video quality models are either incapable or inefficient for predicting the quality of complex, diverse UGC videos in practical applications. Here we introduce an effective and efficient video quality model for UGC content, which we dub the Rapid and Accurate Video Quality Evaluator (RAPIQUE), which we show performs comparably to state-of-the-art models but with orders-of-magnitude faster runtime. Our experimental results on recent large-scale UGC video quality databases show that RAPIQUE delivers top performances on all datasets at a considerably lower computational expense. An implementation of RAPIQUE is online: https://github.com/vztu/RAPIQUE. Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
PCS | 6 |
| 2021 | Completely blind image quality assessment via contourlet energy statisticsabstractAbstract An aim of completely blind image quality assessment (BIQA) is to develop algorithms which can grade image quality without any prior knowledge of the images. Here, a new contourlet energy statistics based completely on blind opinion‐unaware BIQA (OU‐BIQA) method is proposed, which can predict the perceptual severity of a range of image distortion types without requiring any prior knowledge. According to the energy distribution of the contourlet sub‐bands of natural images in log‐domain, the lower‐scale sub‐band energy can be predicted by the corresponding higher‐scale sub‐band energies of distorted images. A quality model is then constructed by quantifying the difference between predicted energy and realistic energy. Meanwhile, an effective method for adjusting and compensating an undesired distortion is integrated into the quality model. Experimental results show that the proposed new method outperforms state‐of‐the‐art OU‐BIQA models on relevant portions of TID2013 database, and is competitive on the LIVE IQA database. Moreover, the proposed model is very fast, suggesting a real‐time solution to high‐performance BIQA. Tuxin Guan, Yuhui Zheng, Bo Jin 0001, Xiaojun Wu 0001, Alan C. Bovik |
IET Image Process. | 6 |
| 2021 | Perceptual Monocular Depth Estimation
Janice Pan, Alan C. Bovik |
Neural Process. Lett. | 2 |
| 2021 | Blind image quality assessment in the contourlet domain
Tuxin Guan, Yuhui Zheng, Xiaochun Zhong, Xiaojun Wu 0001, Alan C. Bovik |
Signal Process. Image Commun. | 6 |
| 2021 | On visual masking estimation for adaptive quantization using steerable filters
Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2021 | Towards Perceptually Optimized Adaptive Video Streaming-A Realistic Quality of Experience DatabaseabstractMeasuring Quality of Experience (QoE) and integrating these measurements into video streaming algorithms is a multi-faceted problem that fundamentally requires the design of comprehensive subjective QoE databases and objective QoE prediction models. To achieve this goal, we have recently designed the LIVE-NFLX-II database, a highly-realistic database which contains subjective QoE responses to various design dimensions, such as bitrate adaptation algorithms, network conditions and video content. Our database builds on recent advancements in content-adaptive encoding and incorporates actual network traces to capture realistic network variations on the client device. The new database focuses on low bandwidth conditions which are more challenging for bitrate adaptation algorithms, which often must navigate tradeoffs between rebuffering and video quality. Using our database, we study the effects of multiple streaming dimensions on user experience and evaluate video quality and quality of experience models and analyze their strengths and weaknesses. We believe that the tools introduced here will help inspire further progress on the development of perceptually-optimized client adaptation and video streaming strategies. The database is publicly available at http://live.ece.utexas.edu/research/LIVE_NFLX_II/live_nflx_plus.html. Christos G. Bampis, Zhi Li 0001, Ioannis Katsavounidis, Te-Yuan Huang, Chaitanya Ekanadham, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2021 | ProxIQA: A Proxy Approach to Perceptual Optimization of Learned Image Compressionabstract(p = 1,2) norms has largely dominated the measurement of loss in neural networks due to their simplicity and analytical properties. However, when used to assess the loss of visual information, these simple norms are not very consistent with human perception. Here, we describe a different "proximal" approach to optimize image analysis networks against quantitative perceptual models. Specifically, we construct a proxy network, broadly termed ProxIQA, which mimics the perceptual model while serving as a loss layer of the network. We experimentally demonstrate how this optimization framework can be applied to train an end-to-end optimized image compression network. By building on top of an existing deep image compression model, we are able to demonstrate a bitrate reduction of as much as 31% over MSE optimization, given a specified perceptual quality (VMAF) level. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2021 | Perceptual Video Quality Prediction Emphasizing Chroma DistortionsabstractMeasuring the quality of digital videos viewed by human observers has become a common practice in numerous multimedia applications, such as adaptive video streaming, quality monitoring, and other digital TV applications. Here we explore a significant, yet relatively unexplored problem: measuring perceptual quality on videos arising from both luma and chroma distortions from compression. Toward investigating this problem, it is important to understand the kinds of chroma distortions that arise, how they relate to luma compression distortions, and how they can affect perceived quality. We designed and carried out a subjective experiment to measure subjective video quality on both luma and chroma distortions, introduced both in isolation as well as together. Specifically, the new subjective dataset comprises a total of 210 videos afflicted by distortions caused by varying levels of luma quantization commingled with different amounts of chroma quantization. The subjective scores were evaluated by 34 subjects in a controlled environmental setting. Using the newly collected subjective data, we were able to demonstrate important shortcomings of existing video quality models, especially in regards to chroma distortions. Further, we designed an objective video quality model which builds on existing video quality algorithms, by considering the fidelity of chroma channels in a principled way. We also found that this quality analysis implies that there is room for reducing bitrate consumption in modern video codecs by creatively increasing the compression factor on chroma channels. We believe that this work will both encourage further research in this direction, as well as advance progress on the ultimate goal of jointly optimizing luma and chroma compression in modern video encoders. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2021 | ChipQA: No-Reference Video Quality Prediction via Space-Time ChipsabstractWe propose a new model for no-reference video quality assessment (VQA). Our approach uses a new idea of highly-localized space-time (ST) slices called Space-Time Chips (ST Chips). ST Chips are localized cuts of video data along directions that implicitly capture motion. We use perceptually-motivated bandpass and normalization models to first process the video data, and then select oriented ST Chips based on how closely they fit parametric models of natural video statistics. We show that the parameters that describe these statistics can be used to reliably predict the quality of videos, without the need for a reference video. The proposed method implicitly models ST video naturalness, and deviations from naturalness. We train and test our model on several large VQA databases, and show that our model achieves state-of-the-art performance at reduced cost, without requiring motion computation. Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Sriram Sethuraman, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2021 | Subjective and Objective Quality Assessment of 2D and 3D Foveated Video Compression in Virtual RealityabstractIn Virtual Reality (VR), the requirements of much higher resolution and smooth viewing experiences under rapid and often real-time changes in viewing direction, leads to significant challenges in compression and communication. To reduce the stresses of very high bandwidth consumption, the concept of foveated video compression is being accorded renewed interest. By exploiting the space-variant property of retinal visual acuity, foveation has the potential to substantially reduce video resolution in the visual periphery, with hardly noticeable perceptual quality degradations. Accordingly, foveated image / video quality predictors are also becoming increasingly important, as a practical way to monitor and control future foveated compression algorithms. Towards advancing the development of foveated image / video quality assessment (FIQA / FVQA) algorithms, we have constructed 2D and (stereoscopic) 3D VR databases of foveated / compressed videos, and conducted a human study of perceptual quality on each database. Each database includes 10 reference videos and 180 foveated videos, which were processed by 3 levels of foveation on the reference videos. Foveation was applied by increasing compression with increased eccentricity. In the 2D study, each video was of resolution 7680×3840 and was viewed and quality-rated by 36 subjects, while in the 3D study, each video was of resolution 5376×5376 and rated by 34 subjects. Both studies were conducted on top of a foveated video player having low motion-to-photon latency (~50ms). We evaluated different objective image and video quality assessment algorithms, including both FIQA / FVQA algorithms and non-foveated algorithms, on our so called LIVE-Facebook Technologies Foveation-Compressed Virtual Reality (LIVE-FBT-FCVR) databases. We also present a statistical evaluation of the relative performances of these algorithms. The LIVE-FBT-FCVR databases have been made publicly available and can be accessed at https://live.ece.utexas.edu/research/LIVEFBTFCVR/index.html. Yize Jin, Meixu Chen, Todd Richard Goodall, Anjul Patney, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2021 | VR Sickness Versus VR Presence: A Statistical Prediction ModelabstractAlthough it is well-known that the negative effects of VR sickness, and the desirable sense of presence are important determinants of a user's immersive VR experience, there remains a lack of definitive research outcomes to enable the creation of methods to predict and/or optimize the trade-offs between them. Most VR sickness assessment (VRSA) and VR presence assessment (VRPA) studies reported to date have utilized simple image patterns as probes, hence their results are difficult to apply to the highly diverse contents encountered in general, real-world VR environments. To help fill this void, we have constructed a large, dedicated VR sickness/presence (VR-SP) database, which contains 100 VR videos with associated human subjective ratings. Using this new resource, we developed a statistical model of spatio-temporal and rotational frame difference maps to predict VR sickness. We also designed an exceptional motion feature, which is expressed as the correlation between an instantaneous change feature and averaged temporal features. By adding additional features (visual activity, content features) to capture the sense of presence, we use the new data resource to explore the relationship between VRSA and VRPA. We also show the aggregate VR-SP model is able to predict VR sickness with an accuracy of 90% and VR presence with an accuracy of 75% using the new VR-SP dataset. Woojae Kim, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2021 | ST-GREED: Space-Time Generalized Entropic Differences for Frame Rate Dependent Video Quality PredictionabstractWe consider the problem of conducting frame rate dependent video quality assessment (VQA) on videos of diverse frame rates, including high frame rate (HFR) videos. More generally, we study how perceptual quality is affected by frame rate, and how frame rate and compression combine to affect perceived quality. We devise an objective VQA model called Space-Time GeneRalized Entropic Difference (GREED) which analyzes the statistics of spatial and temporal band-pass video coefficients. A generalized Gaussian distribution (GGD) is used to model band-pass responses, while entropy variations between reference and distorted videos under the GGD model are used to capture video quality variations arising from frame rate changes. The entropic differences are calculated across multiple temporal and spatial subbands, and merged using a learned regressor. We show through extensive experiments that GREED achieves state-of-the-art performance on the LIVE-YT-HFR Database when compared with existing VQA models. The features used in GREED are highly generalizable and obtain competitive performance even on standard, non-HFR VQA databases. The implementation of GREED has been made available online: https://github.com/pavancm/GREED. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2021 | UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated ContentabstractRecent years have witnessed an explosion of user-generated content (UGC) videos shared and streamed over the Internet, thanks to the evolution of affordable and reliable consumer capture devices, and the tremendous popularity of social media platforms. Accordingly, there is a great need for accurate video quality assessment (VQA) models for UGC/consumer videos to monitor, control, and optimize this vast content. Blind quality prediction of in-the-wild videos is quite challenging, since the quality degradations of UGC videos are unpredictable, complicated, and often commingled. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading no-reference/blind VQA (BVQA) features and models on a fixed evaluation architecture, yielding new empirical insights on both subjective video quality studies and objective VQA model design. By employing a feature selection strategy on top of efficient BVQA models, we are able to extract 60 out of 763 statistical features used in existing methods to create a new fusion-based model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between VQA performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at considerably lower computational cost than other leading models. Our study protocol also defines a reliable benchmark for the UGC-VQA problem, which we believe will facilitate further research on deep learning-based VQA modeling, as well as perceptually-optimized efficient UGC video processing, transcoding, and streaming. To promote reproducible research and public evaluation, an implementation of VIDEVAL has been made available online: https://github.com/vztu/VIDEVAL. Zhengzhong Tu, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2021 | Predicting the Quality of Compressed Videos With Pre-Existing DistortionsabstractBecause of the increasing ease of video capture, many millions of consumers create and upload large volumes of User-Generated-Content (UGC) videos to social and streaming media sites over the Internet. UGC videos are commonly captured by naive users having limited skills and imperfect techniques, and tend to be afflicted by mixtures of highly diverse in-capture distortions. These UGC videos are then often uploaded for sharing onto cloud servers, where they are further compressed for storage and transmission. Our paper tackles the highly practical problem of predicting the quality of compressed videos (perhaps during the process of compression, to help guide it), with only (possibly severely) distorted UGC videos as references. To address this problem, we have developed a novel Video Quality Assessment (VQA) framework that we call 1stepVQA (to distinguish it from two-step methods that we discuss). 1stepVQA overcomes limitations of Full-Reference, Reduced-Reference and No-Reference VQA models by exploiting the statistical regularities of both natural videos and distorted videos. We also describe a new dedicated video database, which was created by applying a realistic VMAF-Guided perceptual rate distortion optimization (RDO) criterion to create realistically compressed versions of UGC source videos, which typically have pre-existing distortions. We show that 1stepVQA is able to more accurately predict the quality of compressed videos, given imperfect reference videos, and outperforms other VQA models in this scenario. Xiangxu Yu, Neil Birkbeck, Yilin Wang 0001, Christos G. Bampis, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2020 | From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture QualityabstractBlind or no-reference (NR) perceptual picture quality prediction is a difficult, unsolved problem of great consequence to the social and streaming media industries that impacts billions of viewers daily. Unfortunately, popular NR prediction models perform poorly on real-world distorted pictures. To advance progress on this problem, we introduce the largest (by far) subjective picture quality database, containing about 40, 000 real-world distorted pictures and 120, 000 patches, on which we collected about 4M human judgments of picture quality. Using these picture and patch quality labels, we built deep region-based architectures that learn to produce state-of-the-art global picture quality predictions as well as useful local picture quality maps. Our innovations include picture quality prediction architectures that produce global-to-local inferences as well as local-to-global inferences (via feedback). The dataset and source code are available at https: //live.ece.utexas.edu/research.php. Zhenqiang Ying, Praful Gupta, Dhruv Mahajan 0001, Deepti Ghadiyaram, Alan C. Bovik |
CVPR | 6 |
| 2020 | Adversarial Video Compression Guided by Soft Edge DetectionabstractWe propose a video compression framework using conditional Generative Adversarial Networks (GANs). We rely on two encoders: one that deploys a standard video codec and another one which generates low-level soft edge maps. For decoding, we use a standard video decoder as well as a decoder that is trained using a conditional GAN. Recent "deep" approaches to video compression require multiple videos to pre-train generative networks that conduct interpolation. By contrast, our scheme trains a generative decoder that requires only a small number of key frames and edge maps taken from a single video, without any interpolation. Experiments on two video datasets demonstrate that the proposed GAN-based compression engine is a promising alternative to traditional video codec approaches that can achieve higher quality reconstructions for very low bitrates. Jin Soo Park, Christos G. Bampis, Jaeseong Lee 0003, Mia K. Markey, Alexandros G. Dimakis, Alan C. Bovik |
ICASSP | 7 |
| 2020 | BBAND INDEX: A NO-REFERENCE BANDING ARTIFACT PREDICTORabstractBanding artifact, or false contouring, is a common video compression impairment that tends to appear on large flat regions in encoded videos. These staircase-shaped color bands can be very noticeable in high-definition videos. Here we study this artifact, and propose a new distortion-specific no-reference video quality model for predicting banding artifacts, called the Blind BANding Detector (BBAND index). BBAND is inspired by human visual models. The proposed detector can generate a pixel-wise banding visibility map and output a banding severity score at both the frame and video levels. Experimental results show that our proposed method outperforms state-of-the-art banding detection algorithms and delivers better consistency with subjective evaluations. Zhengzhong Tu, Jessie Lin, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
ICASSP | 5 |
| 2020 | A Comparative Evaluation Of Temporal Pooling Methods For Blind Video Quality AssessmentabstractMany objective video quality assessment (VQA) algorithms include a key step of temporal pooling of frame-level quality scores. However, less attention has been paid to studying the relative efficiencies of different pooling methods on noreference (blind) VQA. Here we conduct a large-scale comparative evaluation to assess the capabilities and limitations of multiple temporal pooling strategies on blind VQA of usergenerated videos. The study yields insights and general guidance regarding the application and selection of temporal pooling models. In addition, we also propose an ensemble pooling model built on top of high-performing temporal pooling models. Our experimental results demonstrate the relative efficacies of the evaluated temporal pooling models, using several popular VQA algorithms evaluated on two recent largescale natural video quality databases. Conclusively, we also provide an empirical recipe for applying temporal pooling of frame-based quality predictions. Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 6 |
| 2020 | Video Quality Model for Space-Time Resolution AdaptationabstractDelivering voluminous amounts of video data through limited bandwidth channels is a challenge affecting billions of viewers. Accordingly, it is becoming more important to understand the perceptual effects that arise from various dimension reduction methodologies. Towards this direction, we propose a new video quality model that predicts the perceptual quality of videos undergoing varying levels of spatio-temporal subsampling and compression. The new model is established upon the natural statistics principle of videos, which leverage the fact that pristine videos obey statistical regularities that are disturbed by distortions. We found that there exist space-time paths between video frames that best preserve the statistical regularity inherent in the spatial structure of the video frames. The distribution features extracted from frame differences displaced in the direction of these paths correlate more highly with human subjective quality opinions than those from non-displaced frame differences. Given that non-displaced frame differences are widely utilized in video quality models, the improved efficiency of spatially and/or temporally displaced (possibly by more than one frame) frame differences, is an important finding that may significantly elevate the success of studies on temporal features and video quality. Dae Yeol Lee, Hyunsuk Ko, Alan C. Bovik |
IPAS | 4 |
| 2020 | No-Reference Video Quality Assessment Using Space-Time ChipsabstractWe propose a new prototype model for no-reference video quality assessment (VQA) based on the natural statistics of space-time chips of videos. Space-time chips (ST-chips) are a new, quality-aware feature space which we define as space-time localized cuts of video data in directions that are determined by the local motion flow. We use parametrized distribution fits to the bandpass histograms of space-time chips to characterize quality, and show that the parameters from these models are affected by distortion and can hence be used to objectively predict the quality of videos. Our prototype method, which we call ChipQA-0, is agnostic to the types of distortion affecting the video, and is based on identifying and quantifying deviations from the expected statistics of natural, undistorted ST-chips in order to predict video quality. We train and test our resulting model on several large VQA databases and show that our model achieves high correlation against human judgments of video quality and is competitive with state-of-the-art models. Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Alan C. Bovik |
MMSP | 5 |
| 2020 | Optimizing Video Quality Estimation Across ResolutionsabstractMany algorithms have been developed to evaluate the perceptual quality of images and videos, based on models of picture statistics and visual perception. These algorithms attempt to capture user experience better than simple metrics like the peak signal-to-noise ratio (PSNR) and are widely utilized on streaming service platforms and in social networking applications to improve users' Quality of Experience. The growing demand for high-resolution streams and rapid increases in user-generated content (UGC) sharpens interest in the computation involved in carrying out perceptual quality measurements. In this direction, we propose a suite of methods to efficiently predict the structural similarity index (SSIM) of high-resolution videos distorted by scaling and compression, from computations performed at lower resolutions. We show the effectiveness of our algorithms by testing on a large corpus of videos and on subjective data. Abhinau Kumar Venkataramanan, Chengyang Wu, Alan C. Bovik |
MMSP | 3 |
| 2020 | Seeing Through the Clouds With DeepWaterMapabstractWe present our next-generation surface water mapping model, DeepWaterMapV2, which uses improved model architecture, data set, and a training setup to create surface water maps at lower cost, with higher precision and recall. We designed DeepWaterMapV2 to be memory efficient for large inputs. Unlike earlier models, our new model is able to process a full Landsat scene in one-shot and without dividing the input into tiles. DeepWaterMapV2 is robust against a variety of natural and artificial perturbations in the input, such as noise, different sensor characteristics, and small clouds. Our model can even “see” through the clouds without relying on any active sensor data, in cases where the clouds do not fully obstruct the scene. Although we trained the model on Landsat-8 images only, it also supports data from a variety of other Earth observing satellites, including Landsat-5, Landsat-7, and Sentinel-2, without any further training or calibration. Our code and trained model are available at https://github.com/isikdogan/deepwatermap. Furkan Isikdogan, Alan C. Bovik, Paola Passalacqua |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Video quality assessment using space-time slice mappings
Lixiong Liu, Tianshu Wang 0003, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2020 | Blind S3D image quality prediction using classical and non-classical receptive field models
Lixiong Liu, Jiufa Zhang, Michele A. Saad, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2020 | Learning to Distort Images Using Generative Adversarial NetworksabstractModeling image and video distortions is an important, but difficult problem of great consequence to numerous and diverse image processing and computer vision applications. While many statistical models have been proposed to synthesize different types of image noise, real-world distortions are far more difficult to emulate. Toward advancing progress on this interesting problem, we consider distortion generation as an image-to-image transformation problem, and solve it via a data-driven approach. Specifically, we use a conditional generative adversarial network (cGAN) which we train to learn four kinds of realistic distortions. We experimentally demonstrate that the learned model can produce the perceptual characteristics of several types of distortion. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2020 | Capturing Video Frame Rate Variations via Entropic DifferencingabstractHigh frame rate videos are increasingly getting popular in recent years, driven by the strong requirements of the entertainment and streaming industries to provide high quality of experiences to consumers. To achieve the best trade-offs between the bandwidth requirements and video quality in terms of frame rate adaptation, it is imperative to understand the effects of frame rate on video quality. In this direction, we devise a novel statistical entropic differencing method based on a Generalized Gaussian Distribution model expressed in the spatial and temporal band-pass domains, which measures the difference in quality between reference and distorted videos. The proposed design is highly generalizable and can be employed when the reference and distorted sequences have different frame rates. Our proposed model correlates very well with subjective scores in the recently proposed LIVE-YT-HFR database and achieves state of the art performance when compared with existing methodologies. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 5 |
| 2020 | Adaptive Debanding FilterabstractBanding artifacts, which manifest as staircase-like color bands on pictures or video frames, is a common distortion caused by compression of low-textured smooth regions. These false contours can be very noticeable even on high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we consider banding artifact removal as a visual enhancement problem, and accordingly, we solve it by applying a form of content-adaptive smoothing filtering followed by dithered quantization, as a post-processing module. The proposed debanding filter is able to adaptively smooth banded regions while preserving image edges and details, yielding perceptually enhanced gradient rendering with limited bit-depths. Experimental results show that our proposed debanding filter outperforms state-of-the-art false contour removing algorithms both visually and quantitatively. Zhengzhong Tu, Jessie Lin, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 5 |
| 2020 | Blind Noisy Image Quality Assessment Using Sub-Band KurtosisabstractNoise that afflicts natural images, regardless of the source, generally disturbs the perception of image quality by introducing a high-frequency random element that, when severe, can mask image content. Except at very low levels, where it may play a purpose, it is annoying. There exist significant statistical differences between distortion-free natural images and noisy images that become evident upon comparing the empirical probability distribution histograms of their discrete wavelet transform (DWT) coefficients. The DWT coefficients of low- or no-noise natural images have leptokurtic, peaky distributions with heavy tails; while noisy images tend to be platykurtic with less peaky distributions and shallower tails. The sample kurtosis is a natural measure of the peakedness and tail weight of the distributions of random variables. Here, we study the efficacy of the sample kurtosis of image wavelet coefficients as a feature driving, an extreme learning machine which learns to map kurtosis values into perceptual quality scores. The model is trained and tested on five types of noisy images, including additive white Gaussian noise, additive Gaussian color noise, impulse noise, masked noise, and high-frequency noise from the LIVE, CSIQ, TID2008, and TID2013 image quality databases. The experimental results show that the trained model has better quality evaluation performance on noisy images than existing blind noise assessment models, while also outperforming general-purpose blind and full-reference image quality assessment methods. Chenwei Deng, Shuigen Wang, Alan C. Bovik, Guang-Bin Huang, Baojun Zhao |
IEEE Trans. Cybern. | 3 |
| 2020 | Day and Night-Time Dehazing by Local Airlight EstimationabstractWe introduce an effective fusion-based technique to enhance both day-time and night-time hazy scenes. When inverting the Koschmieder light transmission model, and by contrast with the common implementation of the popular dark-channel DehazeHeCVPR2009, we estimate the airlight on image patches and not on the entire image. Local airlight estimation is adopted because, under night-time conditions, the lighting generally arises from multiple localized artificial sources, and is thus intrinsically non-uniform. Selecting the sizes of the patches is, however, non-trivial. Small patches are desirable to achieve fine spatial adaptation to the atmospheric light, but large patches help improve the airlight estimation accuracy by increasing the possibility of capturing pixels with airlight appearance (due to severe haze). For this reason, multiple patch sizes are considered to generate several images, that are then merged together. The discrete Laplacian of the original image is provided as an additional input to the fusion process to reduce the glowing effect and to emphasize the finest image details. Similarly, for day-time scenes we apply the same principle but use a larger patch size. For each input, a set of weight maps are derived so as to assign higher weights to regions of high contrast, high saliency and small saturation. Finally the derived inputs and the normalized weight maps are blended in a multi-scale fashion using a Laplacian pyramid decomposition. Extensive experimental results demonstrate the effectiveness of our approach as compared with recent techniques, both in terms of computational efficiency and the quality of the outputs. Cosmin Ancuti, Codruta O. Ancuti, Christophe De Vleeschouwer, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2020 | Dynamic Receptive Field Generation for Full-Reference Image Quality AssessmentabstractMost full-reference image quality assessment (FR-IQA) methods advanced to date have been holistically designed without regard to the type of distortion impairing the image. However, the perception of distortion depends nonlinearly on the distortion type. Here we propose a novel FR-IQA framework that dynamically generates receptive fields responsive to distortion type. Our proposed method-dynamic receptive field generation based image quality assessor (DRF-IQA)-separates the process of FR-IQA into two streams: 1) dynamic error representation and 2) visual sensitivity-based quality pooling. The first stream generates dynamic receptive fields on the input distorted image, implemented by a trained convolutional neural network (CNN), then the generated receptive field profiles are convolved with the distorted and reference images, and differenced to produce spatial error maps. In the second stream, a visual sensitivity map is generated. The visual sensitivity map is used to weight the spatial error map. The experimental results show that the proposed model achieves state-of-the-art prediction accuracy on various open IQA databases. Woojae Kim, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2020 | Quality Prediction on Deep Generative ImagesabstractIn recent years, deep neural networks have been utilized in a wide variety of applications including image generation. In particular, generative adversarial networks (GANs) are able to produce highly realistic pictures as part of tasks such as image compression. As with standard compression, it is desirable to be able to automatically assess the perceptual quality of generative images to monitor and control the encode process. However, existing image quality algorithms are ineffective on GAN generated content, especially on textured regions and at high compressions. Here we propose a new "naturalness"-based image quality predictor for generative images. Our new GAN picture quality predictor is built using a multi-stage parallel boosting system based on structural similarity features and measurements of statistical similarity. To enable model development and testing, we also constructed a subjective GAN image quality database containing (distorted) GAN images and collected human opinions of them. Our experimental results indicate that our proposed GAN IQA model delivers superior quality predictions on the generative image datasets, as well as on traditional image quality datasets. Hyunsuk Ko, Dae Yeol Lee, Seunghyun Cho, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2020 | Study of Subjective and Objective Quality Assessment of Audio-Visual SignalsabstractThe topics of visual and audio quality assessment (QA) have been widely researched for decades, yet nearly all of this prior work has focused only on single-mode visual or audio signals. However, visual signals rarely are presented without accompanying audio, including heavy-bandwidth video streaming applications. Moreover, the distortions that may separately (or conjointly) afflict the visual and audio signals collectively shape user-perceived quality of experience (QoE). This motivated us to conduct a subjective study of audio and video (A/V) quality, which we then used to compare and develop A/V quality measurement models and algorithms. The new LIVE-SJTU Audio and Video Quality Assessment (A/V-QA) Database includes 336 A/V sequences that were generated from 14 original source contents by applying 24 different A/V distortion combinations on them. We then conducted a subjective A/V quality perception study on the database towards attaining a better understanding of how humans perceive the overall combined quality of A/V signals. We also designed four different families of objective A/V quality prediction models, using a multimodal fusion strategy. The different types of A/V quality models differ in both the unimodal audio and video quality prediction models comprising the direct signal measurements and in the way that the two perceptual signal modes are combined. The objective models are built using both existing state-of-the-art audio and video quality prediction models and some new prediction models, as well as quality-predictive features delivered by a deep neural network. The methods of fusing audio and video quality predictions that are considered include simple product combinations as well as learned mappings. Using the new subjective A/V database as a tool, we validated and tested all of the objective A/V quality prediction models. We will make the database publicly available to facilitate further research. Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Mylène C. Q. Farias, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2020 | Speeding Up VP9 Intra Encoder With Hierarchical Deep Learning-Based Partition PredictionabstractIn VP9 video codec, the sizes of blocks are decided during encoding by recursively partitioning 64×64 superblocks using rate-distortion optimization (RDO). This process is computationally intensive because of the combinatorial search space of possible partitions of a superblock. Here, we propose a deep learning based alternative framework to predict the intra-mode superblock partitions in the form of a four-level partition tree, using a hierarchical fully convolutional network (H-FCN). We created a large database of VP9 superblocks and the corresponding partitions to train an H-FCN model, which was subsequently integrated with the VP9 encoder to reduce the intra-mode encoding time. The experimental results establish that our approach speeds up intra-mode encoding by 69.7% on average, at the expense of a 1.71% increase in the Bjøntegaard-Delta bitrate (BD-rate). While VP9 provides several built-in speed levels which are designed to provide faster encoding at the expense of decreased rate-distortion performance, we find that our model is able to outperform the fastest recommended speed level of the reference VP9 encoder for the good quality intra encoding configuration, in terms of both speedup and BD-rate. Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2020 | Quality Measurement of Images on Mobile Streaming Interfaces Deployed at ScaleabstractWith the growing use of smart cellular devices for entertainment purposes, audio and video streaming services now offer an increasingly wide variety of popular mobile applications that offer portable and accessible ways to consume content. The user interfaces of these applications have become increasingly visual in nature, and are commonly loaded with dense multimedia content such as thumbnail images, animated GIFs, and short videos. To efficiently render these and to aid rapid download to the client display, it is necessary to compress, scale and color subsample them. These operations introduce distortions, reducing the appeal of the application. It is desirable to be able to automatically monitor and govern the visual qualities of these small images, which are usually small images. However, while there exists a variety of high-performing image quality assessment (IQA) algorithms, none have been designed for this particular use case. This kind of content often has unique characteristics, such as overlaid graphics, intentional brightness, gradients, text, and warping. We describe a study we conducted on the subjective and objective quality of images embedded in the displayed user interfaces of mobile streaming applications. We created a database of typical "billboard" and "thumbnail" images viewed on such services. Using the collected data, we studied the effects of compression, scaling and chroma-subsampling on perceived quality by conducting a subjective study. We also evaluated the performance of leading picture quality prediction models on the new database. We report some surprising results regarding algorithm performance, and find that there remains ample scope for future model development. Zeina Sinno, Anush K. Moorthy, Jan De Cock, Zhi Li 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2020 | A Local Flatness Based Variational Approach to RetinexabstractA topic of continued interest in Retinex over the years has been finding ways to implement it with computational models of improved accuracy and efficiency. We have devised a new approach to digitally implementing the Retinex using a local deviation based variational model. The new model leads to improvements in the computed image quality with respect to illumination correction and image enhancement. Several contributions are made: 1) a new prior constraint, which we call local flatness, is proposed, and a new measure of Local Deviation (LD) is developed to quantify the degree of local illumination flatness; 2) a variational problem is defined and the solution is found by a logical sequence of steps; 3) discrete implementation of the variational solution is shown to effectively estimate and remove uneven illumination, yielding an accurate recovered image. Unlike other physical prior based variational Retinex models, which use the L2 norm of the illumination gradient to enforce smoothness of illumination, our LD prior selectively imposes local flatness on illumination by calculating the deviation between the estimated illumination surface to a reference plane. In the experiments, pseudo ground truth images are created by superimposing uneven illumination on real scenes, providing an effective way to objectively assess algorithm performance. The experimental results show that our method can reconstruct more accurate recovered images than other state-of-the-art methods, while maintaining good contrast. Fengying Xie, Rui Zhang 0069, Zhiguo Jiang 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2020 | A Unified Probabilistic Formulation of Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has been attracting considerable attention in recent years due to the explosive growth of digital photography in Internet and social networks. The IAA problem is inherently challenging, owning to the ineffable nature of the human sense of aesthetics and beauty, and its close relationship to understanding pictorial content. Three different approaches to framing and solving the problem have been posed: binary classification, average score regression and score distribution prediction. Solutions that have been proposed have utilized different types of aesthetic labels and loss functions to train deep IAA models. However, these studies ignore the fact that the three different IAA tasks are inherently related. Here, we reveal that the use of the different types of aesthetic labels can be developed within the same statistical framework, which we use to create a unified probabilistic formulation of all the three IAA tasks. This unified formulation motivates the use of an efficient and effective loss function for training deep IAA models to conduct different tasks. We also discuss the problem of learning from a noisy raw score distribution which hinders network performance. We then show that by fitting the raw score distribution to a more stable and discriminative score distribution, we are able to train a single model which is able to obtain highly competitive performance on all three IAA tasks. Extensive qualitative analysis and experimental results on image aesthetic benchmarks validate the superior performance afforded by the proposed formulation. The source code is available at. Hui Zeng 0001, Zisheng Cao, Lei Zhang 0006, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2019 | Automated Segmentation of Nanoparticles in BF TEM Images by U-Net Binarization and Branch and Bound
Sahar F. Zafar, Tuomas Eerola, Heikki Kälviäinen, Alan C. Bovik |
CAIP (1) | 5 |
| 2019 | Optimal Feature Selection for Blind Super-resolution Image Quality EvaluationabstractThe visual quality of images resulting from Super Resolution (SR) techniques is predicted with blind image quality assessment (BIQA) models trained on a database(s) of human rated distorted images and associated human subjective opinion scores. Such opinion-aware (OA) methods need a large amount of training samples with associated human subjective scores, which are scarce in the field of SR. By contrast, opinion distortion unaware (ODU) methods do not need human subjective scores for training. This paper presents an opinion-unaware BIQA measure of super resolved images based on optimally extracted perceptual features. This set of features was selected using a floating forward search whose objective function is the correlation with human judgment. The proposed BIQA method does not need any distorted images nor subjective quality scores for training, yet the experiments demonstrate its superior quality-prediction performance relative to state-of-the-art opinion-unaware BIQA methods, and that it is competitive to state-of-the-art opinion-aware BIQA methods. Juan Berón, Hernán Darío Benítez, Alan C. Bovik |
ICASSP | 3 |
| 2019 | Spatio-Temporal Measures Of NaturalnessabstractToday, a wide variety of casual users of diverse videographic skills and styles capture a large portion of all videos, often equipped with uncertain hands and operating under difficult lighting conditions. These videos are taken with many types of camera devices having different characteristics, resulting in a wide range and diversity of video qualities. These are the kinds of videos shared on YouTube, Snapchat, and Face-book. Being able to predict the quality of these videos is an important goal for a variety of invested practitioners, including camera designers, cloud engineers, and users who could be directed to recapture videos of poor quality. In nearly every instance, a high quality reference video is not available, hence blind video quality predictors are of the greatest interest. Towards advancing this area, we have studied the spatiotemporal statistic of a wide variety of natural videos, constructed new directional temporal statistical models of videos, and studied whether measures of directional spatio-temporal naturalness can be developed that are predictive of quality. Zeina Sinno, Alan C. Bovik |
ICIP | 2 |
| 2019 | Cloud Detection in Satellite Images Based on Natural Scene Statistics and Gabor FeaturesabstractCloud detection is an important task in remote sensing (RS) image processing. Numerous cloud detection algorithms have been developed. However, most existing methods suffer from the weakness of omitting small and thin clouds, and from an inability to discriminate clouds from photometrically similar regions, such as buildings and snow. Here, we derive a novel cloud detection algorithm for optical RS images, whereby test images are separated into three classes: thick clouds, thin clouds, and noncloudy. First, a simple linear iterative clustering algorithm is adopted that is able to segment potential clouds, including small clouds. Then, a natural scene statistics model is applied to the superpixels to distinguish between clouds and surface buildings. Finally, Gabor features are computed within each superpixel and a support vector machine is used to distinguish clouds from snow regions. The experimental results indicate that the proposed model outperforms state-of-the-art methods for cloud detection. Chenwei Deng, Zhen Li 0017, Shuigen Wang, Linbo Tang, Alan C. Bovik |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2019 | Image Statistic Models Characterize Well Log Image QualityabstractAssessing the image quality of well logs is essential to ensure the accuracy of their digitization and subsequent processing. Currently, the suitability of well logs for information retrieval is solely determined on the basis of subjective judgments of their image quality by human experts. The success of natural scene statistics (NSS)-based models that are used to conduct no-reference (NR) quality assessment of photographic images motivates us to try to exploit them to characterize the quality of nonphotographic images, such as well logs. Accordingly, we develop a scheme to characterize the quality of a well log as “acceptable” or “unacceptable” for subsequent processing based on the natural image quality evaluator (NIQE), a successful NR image quality assessment model based on the NSS. Our experimental results show that the objective quality scores thus obtained can be reliably used to eliminate well logs of inferior quality from the processing pipeline, which can serve as a beneficial step to reduce the human hours spent in examining well logs and to improve the rate of information retrieval as well as the accuracy of retrieved information. Source code for the trained well log image quality predictor is available at https://github.com/Somdyuti2/Well_log_IQA. Somdyuti Paul, Alan C. Bovik |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Spatiotemporal Feature Integration and Model Fusion for Full Reference Video Quality AssessmentabstractThe recently developed video multi-method assessment fusion (VMAF) framework integrates multiple quality-aware features to accurately predict the video quality. However, the VMAF does not yet exploit important principles of temporal perception that are relevant to the perceptual video distortion measurement. Here, we propose two improvements to the VMAF framework, called spatiotemporal VMAF and ensemble VMAF, which leverage perceptually-motivated space-time features that are efficiently calculated at multiple scales. We also conducted a large subjective video study, which we have found to be an excellent resource for training our feature-based approaches. In rigorous experiments, we found that the proposed algorithms demonstrate the state-of-the-art performance on multiple video applications. The compared algorithms will be made available as a part of the open source package in https://github.com/Netflix/vmaf. Christos G. Bampis, Zhi Li 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A Subjective and Objective Study of Stalling Events in Mobile Streaming VideosabstractOver-the-top mobile adaptive video streaming is invariably influenced by volatile network conditions, which can cause playback interruptions (stalling or rebuffering events) and bitrate fluctuations, thereby impairing users' quality of experience (QoE). Video quality assessment models that can accurately predict users' QoE under such volatile network conditions are rapidly gaining attention, since these methods could enable more efficient design of quality control protocols for media-driven services such as YouTube, Amazon, Netflix, and many others. However, the development of improved QoE prediction models requires data sets of videos afflicted with diverse stalling events that have been labeled with ground-truth subjective opinion scores. Toward this end, we have created a new mobile video quality database that we call LIVE Mobile Stall Video Database-II. Our database contains a total of 174 videos afflicted with distortions caused by 26 different stalling patterns. We describe the way we simulated the diverse stalling events to create a corpus of distorted videos, and we detail the human study we conducted to obtain continuous-time subjective scores from 54 subjects. We also present the outcomes of our comprehensive analysis of the impact of several factors that influence subjective QoE, and report the performance of existing QoE-prediction models on our data set. We are making the database (videos, subjective data, and video metadata) publicly available in order to help the advance state-of-the-art research on user-centric mobile network planning and management. The database may be accessed at http://live.ece.utexas.edu/research/LIVEStallStudy/liveMobile.html. Deepti Ghadiyaram, Janice Pan, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Study of Subjective Quality and Objective Blind Quality Prediction of Stereoscopic VideosabstractWe present a new subjective and objective study on full high-definition (HD) stereoscopic (3D or S3D) video quality. In subjective study, we constructed an S3D video dataset with 12 pristine and 288 test videos, and the test videos are generated by applying the H.264 and H.265 compression, blur and frame freeze artifacts. We also propose a no reference (NR) objective video quality assessment (QA) algorithm that relies on measurements of the statistical dependencies between the motion and disparity subband coefficients of S3D videos. Inspired by the Generalized Gaussian Distribution (GGD) approach in liu2011statistical, we model the joint statistical dependencies between the motion and disparity components as following a Bivariate Generalized Gaussian Distribution (BGGD). We estimate the BGGD model parameters (α,β) and the coherence measure (Ψ) from the eigenvalues of the sample covariance matrix (M) of the BGGD. In turn, we model the BGGD parameters of pristine S3D videos using a Multivariate Gaussian (MVG) distribution. The likelihood of a test video's MVG model parameters coming from the pristine MVG model is computed and shown to play a key role in the overall quality estimation. We also estimate the global motion content of each video by averaging the SSIM scores between pairs of successive video frames. To estimate the test S3D video's spatial quality, we apply the popular 2D NR unsupervised NIQE image QA model on a frame-by-frame basis on both views. The overall quality of a test S3D video is finally computed by pooling the test S3D video's likelihood estimates, global motion strength and spatial quality scores. The proposed algorithm, which is 'completely blind' (requiring no reference videos or training on subjective scores) is called the Motion and Disparity based 3D video quality evaluator (MoDi3D). We show that MoDi3D delivers competitive performance over a wide variety of datasets including the IRCCYN dataset, the WaterlooIVC Phase I dataset, the LFOVIA dataset and our proposed LFOVIAS3DPh2 S3D video dataset. Balasubramanyam Appina, Dendi Sathya Veera Reddy, K. Manasa, Sumohana S. Channappayya, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2019 | Detecting and Mapping Video ImpairmentsabstractAutomatically identifying the locations and severities of video artifacts without the advantage of an original reference video is a difficult task. We present a novel approach to conducting no-reference artifact detection in digital videos, implemented as an efficient and unique dual-path (parallel) excitatory/inhibitory neural network that uses a simple discrimination rule to define a bank of accurate distortion detectors. The learning engine is distortion-sensitized by pre-processing each video using a statistical image model. The overall system is able to produce full-resolution space-time distortion maps for visualization, as well as providing global distortion detection decisions that represent the state of the art in performance. Our model, which we call the Video Impairment Mapper (VIDMAP), produces a first-of-a-kind full resolution map of artifact detection probabilities. The current realization of this system is able to accurately detect and map eight of the most important artifact categories encountered during streaming video source inspection: aliasing, video encoding corruptions, quantization, contours/banding, combing, compression, dropped frames, and upscaling artifacts. We show that it is either competitive with or significantly outperforms the previous state-of-the-art on the whole-image artifact detection task. A software release of VIDMAP that has been trained to detect and map these artifacts is available online: http://live.ece.utexas.edu/research/quality/VIDMAP release.zip for public use and evaluation. Todd Richard Goodall, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2019 | Predicting Detection Performance on Security X-Ray Images as a Function of Image QualityabstractDeveloping methods to predict how image quality affects the task performance is a topic of great interest in many applications. While such studies have been performed in the medical imaging community, little work has been reported in the security X-ray imaging literature. In this paper, we develop models that predict the effect of image quality on the detection of the improvised explosive device components by bomb technicians in images taken using portable X-ray systems. Using a newly developed NIST-LIVE X-Ray Task Performance Database, we created a set of objective algorithms that predict bomb technician detection performance based on the measures of image quality. Our basic measures are traditional image quality indicators (IQIs) and perceptually relevant natural scene statistics (NSS)-based measures that have been extensively used in visible light image quality prediction algorithms. We show that these measures are able to quantify the perceptual severity of degradations and can predict the performance of expert bomb technicians in identifying threats. Combining NSS- and IQI-based measures yields even better task performance prediction than either of these methods independently. We also developed a new suite of statistical task prediction models that we refer to as quality inspectors of X-ray images (QUIX); we believe this is the first NSS-based model for security X-ray images. We also show that QUIX can be used to reliably predict conventional IQI metric values on the distorted X-ray images. Praful Gupta, Zeina Sinno, Jack L. Glover, Nicholas G. Paulter Jr., Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2019 | Large-Scale Study of Perceptual Video QualityabstractThe great variations of videographic skills in videography, camera designs, compression and processing protocols, communication and bandwidth environments, and displays leads to an enormous variety of video impairments. Current noreference (NR) video quality models are unable to handle this diversity of distortions. This is true in part because available video quality assessment databases contain very limited content, fixed resolutions, were captured using a small number of camera devices by a few videographers and have been subjected to a modest number of distortions. As such, these databases fail to adequately represent real world videos, which contain very different kinds of content obtained under highly diverse imaging conditions and are subject to authentic, complex and often commingled distortions that are difficult or impossible to simulate. As a result, NR video quality predictors tested on real-world video data often perform poorly. Towards advancing NR video quality prediction, we have constructed a largescale video quality assessment database containing 585 videos of unique content, captured by a large number of users, with wide ranges of levels of complex, authentic distortions. We collected a large number of subjective video quality scores via crowdsourcing. A total of 4776 unique participants took part in the study, yielding more than 205000 opinion scores, resulting in an average of 240 recorded human opinions per video. We demonstrate the value of the new resource, which we call the LIVE Video Quality Challenge Database (LIVE-VQC for short), by conducting a comparison of leading NR video quality predictors on it. This study is the largest video quality assessment study ever conducted along several key dimensions: number of unique contents, capture devices, distortion types and combinations of distortions, study participants, and recorded subjective scores. The database is available for download on this link: http://live.ece.utexas.edu/research/LIVEVQC/index.html. Zeina Sinno, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2019 | Predicting the Quality of Images Compressed After Distortion in Two StepsabstractIn a typical communication pipeline, images undergo a series of processing steps that can cause visual distortions before being viewed. Given a high quality reference image, a reference (R) image quality assessment (IQA) algorithm can be applied after compression or transmission. However, the assumption of a high quality reference image is often not fulfilled in practice, thus contributing to less accurate quality predictions when using stand-alone R IQA models. This is particularly common on social media, where hundreds of billions of usergenerated photos and videos containing diverse, mixed distortions are uploaded, compressed, and shared annually on sites like Facebook, YouTube, and Snapchat. The qualities of the pictures that are uploaded to these sites vary over a very wide range. While this is an extremely common situation, the problem of assessing the qualities of compressed images against their precompressed, but often severely distorted (reference) pictures has been little studied. Towards ameliorating this problem, we propose a novel two-step image quality prediction concept that combines NR with R quality measurements. Applying a first stage of NR IQA to determine the possibly degraded quality of the source image yields information that can be used to quality-modulate the R prediction to improve its accuracy. We devise a simple and efficient weighted product model of R and NR stages, which combines a pre-compression NR measurement with a post-compression R measurement. This first-of-a-kind two-step approach produces more reliable objective prediction scores. We also constructed a new, first-of-a-kind dedicated database specialized for the design and testing of two-step IQA models. Using this new resource, we show that twostep approaches yield outstanding performance when applied to compressed images whose original, pre-compression quality covers a wide range of realistic distortion types and severities. The two-step concept is versatile as it can use any desired R and NR components. We are making the source code of a particularly efficient model that we call 2stepQA publicly available at https://github.com/xiangxuyu/2stepQA. We are also providing the dedicated new two-step database free of charge at http://live.ece.utexas.edu/research/twostep/index.html. Xiangxu Yu, Christos G. Bampis, Praful Gupta, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2018 | Orthogonally-Divergent Fisheye Stereo
Janice Pan, Martin Mueller, Tarek Lahlou, Alan C. Bovik |
ACIVS | 4 |
| 2018 | Second Order Natural Scene Statistics Model of Blind Image Quality AssessmentabstractThe univariate statistics of bandpass-filtered images provide powerful features that drive many successful image quality assessment (IQA) algorithms. Bivariate Natural Scene Statistics (NSS), which model the joint statistics of multiple bandpass image samples also provide potentially powerful features to assess the perceptual quality of images, by capturing both image and distortion correlations. Here, we make the first attempt to use bivariate NSS features to build a model of no-reference image quality prediction. We show that our bivariate model outperforms existing state of the art image quality predictors. Zeina Sinno, Constantine Caramanis, Alan C. Bovik |
ICASSP | 3 |
| 2018 | Enhancing Temporal Quality Measurements in a Globally Deployed Streaming Video Quality PredictorabstractMost successful perceptual video quality assessment models are either frame-based, or perform spatiotemporal filtering or motion estimation to model the temporal aspects of video distortions. While good results are obtained on video quality databases, their increased computational complexity often causes video quality engineers to instead rely on simpler image-based quality algorithms. Towards balancing demands between prediction accuracy and compute efficiency, Netflix developed the Video Multi-method Assessment Fusion (VMAF) Framework, an efficient feature-based system that combines multiple perception-based elementary image measurements to produce video quality predictions. However, the current version of VMAF only weakly captures temporal video features which are sensitive to perceptual temporal video distortions. To this end, we propose an enhanced model we call SpatioTemporal VMAF (ST- VMAF) that incorporates temporal features that are easy to compute. We demonstrate the improved performance of ST- VMAF on many subjective video databases. The proposed model will be made available as part of the open source package in https://github.com/Netflix/vmaf. Christos G. Bampis, Zhi Li 0001, Alan C. Bovik |
ICIP | 3 |
| 2018 | Multivariate Statistics for Blind Image Quality ApplicationsabstractMany existing no-reference image quality approaches exploit only univariate statistical models of bandpass image coefficients, thereby neglecting the higher-order correlations that occur between adjacent coefficients which may be modified by distortion. We modeled multivariate natural image statistics in the spatial domain as following a Multivariate Generalized Gaussian distribution that provides useful information regarding the type and severity of distortions in image signals. We have also found that gaussianity assumptions when estimating local variances are violated in the presence of distortions, hence we estimate local energies using a generalized Gaussian distribution. To this end, we propose Multivariate-Generalized Contrast Normalization (MV-GCN): a multivariate approach which integrates a generalized contrast normalization step. We demonstrate the potential of our model in various image quality applications. Praful Gupta, Christos G. Bampis, Alan C. Bovik |
ICIP | 3 |
| 2018 | Large Scale Subjective Video Quality StudyabstractMost of today's video quality assessment (VQA) databases contain very limited content and distortion diversities and fail to adequately represent real world video impairments. This is in part because conducting subjective studies in the lab is slow, inefficient and expensive process. Crowdsourcing quality scores is a more scalable solution. However given that viewers operate under innumerable viewing conditions (in-cluding display resolutions, viewing distances, internet connection speeds) and because they are not closely supervised, multiple technical challenges arise. We carefully designed a framework in Amazon Mechanical Turk (AMT) to address the many technical issues that are faced. We launched the largest available VQA study, collecting more than 205000 opinion scores provided by more than 4700 unique participants. We have verified that our framework provided us with results that are highly consistent with the ones obtained in a lab environment under controlled conditions. Zeina Sinno, Alan C. Bovik |
ICIP | 2 |
| 2018 | Blind Image Quality Assessment with a Probabilistic Quality RepresentationabstractMost existing blind image quality assessment (BIQA) methods learn a regression model to predict scalar quality scores. Such a scheme ignores the fact that an image will receive divergent subjective scores from different subjects, which cannot be adequately represented by a single scalar number. This is particularly true on complex, real-world distorted images. However, the more informative score distributions are unavailable in existing image quality assessment (IQA) databases and can be potentially noisy when limited number of opinions are collected on each image. This paper proposes a probabilistic quality representation (PQR) and employs a more robust loss function to train deep BIQA models. Using a very straightforward implementation, the proposed method is shown to not only speed up the convergence of deep model training, but also greatly improve the quality prediction accuracy relative to scalar quality score regression methods under the same setting. The source code is available at https://github.com/HuiZeng/BIQA_Toolbox. Hui Zeng 0001, Lei Zhang 0006, Alan C. Bovik |
ICIP | 3 |
| 2018 | A Simple Prediction Fusion Improves Data-driven Full-Reference Video Quality Assessment ModelsabstractWhen developing data-driven video quality assessment algorithms, the size of the available ground truth subjective data may hamper the generalization capabilities of the trained models. Nevertheless, if the application context is known a priori, leveraging data-driven approaches for video quality prediction can deliver promising results. Towards achieving highperforming video quality prediction for compression and scaling artifacts, Netflix developed the Video Multi-method Assessment Fusion (VMAF) Framework, a full-reference prediction system which uses a regression scheme to integrate multiple perceptionmotivated features to predict video quality. However, the current version of VMAF does not fully capture temporal video features relevant to temporal video distortions. To achieve this goal, we developed Ensemble VMAF (E-VMAF): a video quality predictor that combines two models: VMAF and predictions based on entropic differencing features calculated on video frames and frame differences. We demonstrate the improved performance of E-VMAF on various subjective video databases. The proposed model will become available as part of the open source package in https://github. com/Netflix/vmaf. Christos G. Bampis, Alan C. Bovik, Zhi Li 0001 |
PCS | 2 |
| 2018 | Detecting Source Video Artifacts with Supervised Sparse FiltersabstractA variety of powerful picture quality predictors are available that rely on neuro-statistical models of distortion perception. We extend these principles to video source inspection, by coupling spatial divisive normalization with a filterbank tuned for artifact detection, implemented in an augmented sparse functional form. We call this method the Video Impairment Detection by SParse Error CapTure (VIDSPECT). We configure VIDSPECT to create state-of-the-art detectors of two kinds of commonly encountered source video artifacts: upscaling and combing. The system detects upscaling, identifies upscaling type, and predicts the native video resolution. It also detects combing artifacts arising from interlacing. Our approach is simple, highly generalizable, and yields better accuracy than competing methods. A software release of VIDSPECT is available online: http://live.ece.utexas.edu/research/quality/VIDSPECT release.zip for public use and evaluation. Todd Richard Goodall, Alan C. Bovik |
PCS | 2 |
| 2018 | Quality Assessment of Thumbnail and Billboard Images on Mobile DevicesabstractObjective image quality assessment (IQA) research entails developing algorithms that predict human judgments of picture quality. Validating performance entails evaluating algorithms under conditions similar to where they are deployed. Hence, creating image quality databases representative of target use cases is an important endeavor. Here we present a database that relates to quality assessment of billboard images commonly displayed on mobile devices. Billboard images are a subset of thumbnail images, that extend across a display screen, representing things like album covers, banners, or frames or artwork. We conducted a subjective study of the quality of billboard images distorted by processes like compression, scaling and chroma-subsampling, and compared high-performance quality prediction models on the images and subjective data. Zeina Sinno, Anush K. Moorthy, Jan De Cock, Zhi Li 0001, Alan C. Bovik |
PCS | 5 |
| 2018 | Eye Movement Pattern Modeling and Visual Comfort Viewing S3D ImagesabstractStereoscopic-3D (S3D) displays are widely used but present problems related to experiences of visual discomfort for human vision. One aspect of this issue is the movement of the gaze point within different depth fields. Here we aim to analyze the relationship between eye movement patterns and visual comfort experienced when viewing S3D images. Rather than simply labeling eye movement data according to categories such as gaze, saccade and so on, we depoly nonparametric Bayesian method to analyze and cluster several eye movement patterns, and to relate them to visual comfort. The results are relevant to the prediction of visual comfort assessment in S3D images by automatic algorithms. Jun Zhou 0007, Xiao Gu 0001, Shoucheng Zhu, Alan C. Bovik |
VCIP | 5 |
| 2018 | Learning a River Network Extractor Using an Adaptive Loss FunctionabstractWe have created a deep-learning-based river network extraction model, called DeepRiver, that learns the characteristics of rivers from synthetic data and generalizes them to natural data. To train this model, we created a very large database of exemplary synthetic local channel segments, including channel intersections. Our model uses a special loss function that automatically shifts the focus to the hardest-to-learn parts of an input image. This adaptive loss function makes it possible to learn to detect river centerlines, including the centerlines at junctions and bifurcations. DeepRiver learns to separate between rivers and oceans, and therefore, it is able to reliably extract rivers in coastal regions. The model produces maps of river centerlines, which have the potential to be quite useful for analyzing the properties of river networks. Furkan Isikdogan, Alan C. Bovik, Paola Passalacqua |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Feature-based prediction of streaming video QoE: Distortions, stalling and memory
Christos G. Bampis, Alan C. Bovik |
Signal Process. Image Commun. | 2 |
| 2018 | Video quality assessment accounting for temporal visual masking of local flicker
Lark Kwon Choi, Alan C. Bovik |
Signal Process. Image Commun. | 2 |
| 2018 | Generalized Gaussian scale mixtures: A model for wavelet coefficients of natural images
Praful Gupta, Anush K. Moorthy, Rajiv Soundararajan, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2018 | Perceptual quality evaluation of synthetic pictures distorted by compression and transmission
Debarati Kundu, Lark Kwon Choi, Alan C. Bovik, Brian L. Evans |
Signal Process. Image Commun. | 3 |
| 2018 | In-Capture Mobile Video Distortions: A Study of Subjective Behavior and Objective AlgorithmsabstractDigital videos often contain visual distortions that are introduced by the camera's hardware or processing software during the capture process. These distortions often detract from a viewer's quality of experience. Understanding how human observers perceive the visual quality of digital videos is of great importance to camera designers. Thus, the development of automatic objective methods that accurately quantify the impact of visual distortions on perception has greatly accelerated. Video quality algorithm design and verification require realistic databases of distorted videos and human judgments of them. However, most current publicly available video quality databases have been created under highly controlled conditions using graded, simulated, and post-capture distortions (such as jitter and compression artifacts) on high-quality videos. The commercial plethora of hand-held mobile video capture devices produces videos often afflicted by a variety of complex distortions generated during the capturing process. These in-capture distortions are not well-modeled by the synthetic, post-capture distortions found in existing VQA databases. Toward overcoming this limitation, we designed and created a new database that we call the LIVE-Qualcomm mobile in-capture video quality database, comprising a total of 208 videos, which model six common in-capture distortions. We also conducted a subjective quality assessment study using this database, in which each video was assessed by 39 unique subjects. Furthermore, we evaluated several top-performing no-reference IQA and VQA algorithms on the new database and studied how real-world in-capture distortions challenge both human viewers as well as automatic perceptual quality prediction models. The new database is freely available at: http://live.ece.utexas.edu/research/incaptureDatabase/index.html. Deepti Ghadiyaram, Janice Pan, Alan C. Bovik, Anush K. Moorthy, Prasanjit Panda, Kai-Chieh Yang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Recurrent and Dynamic Models for Predicting Streaming Video Quality of ExperienceabstractStreaming video services represent a very large fraction of global bandwidth consumption. Due to the exploding demands of mobile video streaming services, coupled with limited bandwidth availability, video streams are often transmitted through unreliable, low-bandwidth networks. This unavoidably leads to two types of major streaming-related impairments: compression artifacts and/or rebuffering events. In streaming video applications, the end-user is a human observer; hence being able to predict the subjective Quality of Experience (QoE) associated with streamed videos could lead to the creation of perceptually optimized resource allocation strategies driving higher quality video streaming services. We propose a variety of recurrent dynamic neural networks that conduct continuous-time subjective QoE prediction. By formulating the problem as one of time-series forecasting, we train a variety of recurrent neural networks and non-linear autoregressive models to predict QoE using several recently developed subjective QoE databases. These models combine multiple, diverse neural network inputs, such as predicted video quality scores, rebuffering measurements, and data related to memory and its effects on human behavioral responses, using them to predict QoE on video streams impaired by both compression artifacts and rebuffering events. Instead of finding a single time-series prediction model, we propose and evaluate ways of aggregating different models into a forecasting ensemble that delivers improved results with reduced forecasting variance. We also deploy appropriate new evaluation metrics for comparing time-series predictions in streaming applications. Our experimental results demonstrate improved prediction performance that approaches human performance. An implementation of this work can be found at https://github.com/christosbampis/NARX_QoE_release. Christos G. Bampis, Zhi Li 0001, Ioannis Katsavounidis, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2018 | Learning a Continuous-Time Streaming Video QoE ModelabstractOver-the-top adaptive video streaming services are frequently impacted by fluctuating network conditions that can lead to rebuffering events (stalling events) and sudden bitrate changes. These events visually impact video consumers' quality of experience (QoE) and can lead to consumer churn. The development of models that can accurately predict viewers' instantaneous subjective QoE under such volatile network conditions could potentially enable the more efficient design of quality-control protocols for media-driven services, such as YouTube, Amazon, Netflix, and so on. However, most existing models only predict a single overall QoE score on a given video and are based on simple global video features, without accounting for relevant aspects of human perception and behavior. We have created a QoE evaluator, called the time-varying QoE Indexer, that accounts for interactions between stalling events, analyzes the spatial and temporal content of a video, predicts the perceptual video quality, models the state of the client-side data buffer, and consequently predicts continuous-time quality scores that agree quite well with human opinion scores. The new QoE predictor also embeds the impact of relevant human cognitive factors, such as memory and recency, and their complex interactions with the video content being viewed. We evaluated the proposed model on three different video databases and attained standout QoE prediction performance. Deepti Ghadiyaram, Janice Pan, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2018 | Modeling the Perceptual Quality of Immersive Images Rendered on Head Mounted Displays: Resolution and CompressionabstractWe develop a model that expresses the joint impact of spatial resolution s and JPEG compression quality factor qf on immersive image quality. The model is expressed as the product of optimized exponential functions of these factors. The model is tested on a subjective database of immersive image contents rendered on a head mounted display (HMD). High Pearson correlation and Spearman correlation (> 0.95) and small relative root mean squared error (< 5.6%) are achieved between the model predictions and the subjective quality judgements. The immersive ground-truth images along with the rest of the database are made available for future research and comparisons. Mingkai Huang, Qiu Shen, Zhan Ma 0001, Alan C. Bovik, Praful Gupta, Rongbing Zhou, Xun Cao |
IEEE Trans. Image Process. | 4 |
| 2018 | Deep Visual Discomfort Predictor for Stereoscopic 3D ImagesabstractMost prior approaches to the problem of stereoscopic 3D (S3D) visual discomfort prediction (VDP) have focused on the extraction of perceptually meaningful handcrafted features based on models of visual perception and of natural depth statistics. Towards advancing performance on this problem, we have developed a deep learning based VDP model named Deep Visual Discomfort Predictor (DeepVDP). DeepVDP uses a convolutional neural network (CNN) to learn features that are highly predictive of experienced visual discomfort. Since a large amount of reference data is needed to train a CNN, we develop a systematic way of dividing S3D image into local regions defined as patches, and model a patch-based CNN using two sequential training steps. Since it is very difficult to obtain human opinions on each patch, instead a proxy ground-truth label that is generated by an existing S3D visual discomfort prediction algorithm called 3D-VDP is assigned to each patch. These proxy ground-truth labels are used to conduct the first stage of training the CNN. In the second stage, the automatically learned local abstractions are aggregated into global features via a feature aggregation layer. The learned features are iteratively updated via supervised learning on subjective 3D discomfort scores, which serve as ground-truth labels on each S3D image. The patchbased CNN model that has been pretrained on proxy groundtruth labels is subsequently retrained on true global subjective scores. The global S3D visual discomfort scores predicted by the trained DeepVDP model achieve state-of-the-art performance as compared to previous VDP algorithms. Heeseok Oh, Sewoong Ahn, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2018 | Towards a Closed Form Second-Order Natural Scene Statistics ModelabstractPrevious work on natural scene statistics (NSS)-based image models has focused primarily on characterizing the univariate bandpass statistics of single pixels. These models have proven to be powerful tools driving a variety of computer vision and image/video processing applications, including depth estimation, image quality assessment, and image denoising, among others. Multivariate NSS models descriptive of the joint distributions of spatially separated bandpass image samples have, however, received relatively little attention. Here, we develop a closed form bivariate spatial correlation model of bandpass and normalized image samples that completes an existing 2D joint generalized Gaussian distribution model of adjacent bandpass pixels. Our model is built using a set of diverse, high-quality naturalistic photographs, and as a control, we study the model properties on white noise. We also study the way the model fits are affected when the images are modified by common distortions. Zeina Sinno, Constantine Caramanis, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2017 | Towards automated quality curation of video collections from a realistic perspectiveabstractWe investigate the use of automated Video Quality Assessment (VQA) algorithms to evaluate digital video collections. These algorithms are driven by well-defined natural scene statistics (NSS), which capture the behavior of natural distortion-free videos. Because human vision has adapted to these real-world statistics over the course of evolution, quality predictions delivered by these NSS-based VQA algorithms correlate well with human opinions of quality. In particular, we expect these algorithms to accurately predict quality on sizable and diverse video collections. To test this hypothesis, we gathered a testbed of video clips that represent a larger video art collection. Next, we conducted a human study in which users scored the quality of the clips. Enabled by the human study, we trained three VQA algorithms (Video BLIINDS, BRISQUE, and VIIDEO) using our testbed collection to assess a real-world digital video art collection from our university museum. Two of the algorithms provided good automatic predictions of the quality of the videos. These same algorithms also highlighted limitations that arise when assessing artistic collections. We present current research progress and discuss future directions for testbed and algorithm improvement. Our ongoing effort furthers the field of Computational Archival Science by applying computational models of human perception to video appraisal and preservation tasks. Todd Richard Goodall, Maria Esteva, Sandra Sweat, Alan C. Bovik |
IEEE BigData | 4 |
| 2017 | Subjective and objective quality assessment of Mobile Videos with In-Capture distortionsabstractWe designed and created a new video database that models a variety of complex distortions generated during the video capturing process on hand-held mobile capturing devices. We describe the content and characteristics of the new database, which we call the LIVE Mobile In-Capture Video Quality Database. It comprises a total of 208 videos that were captured using eight different smart-phones and were affected by six common in-capture distortions. We also conducted a subjective video quality assessment study using this data, wherein each video was assessed by 36 unique subjects. We evaluated several top-performing No-Reference IQA and VQA algorithms on the new database and find insights on how real-world in-capture distortions challenge both human subjects as well as automatic perceptual quality prediction models. Deepti Ghadiyaram, Janice Pan, Alan C. Bovik, Anush K. Moorthy, Prasanjit Panda, Kai-Chieh Yang |
ICASSP | 3 |
| 2017 | Statistics of natural fused image distortionsabstractThe capability to automatically evaluate the quality of long wave infrared (LWIR) and visible light images has the potential to play an important role in determining and controlling the quality of a resulting fused LWIR-visible image. Extensive work has been conducted on studying the statistics of natural LWIR and visible light images. Nonetheless, there has been little work done on analyzing the statistics of fused images and associated distortions. In this paper, we study the natural scene statistics (NSS) of fused images and how they are affected by several common types of distortions, including blur, white noise, JPEG compression, and non-uniformity (NU). Based on the results of a separate subjective study on the quality of pristine and degraded fused images, we propose an opinion-aware (OA) fused image quality analyzer, whose relative predictions with respect to other state-of-the-art metrics correlate better with human perceptual evaluations. David Eduardo Moreno-Villamarin, Hernán Darío Benítez, Alan C. Bovik |
ICASSP | 3 |
| 2017 | Image quality assessment to enhance infrared face recognitionabstractAutomatic quality evaluation of infrared images has not been researched as extensively as for images of the visible spectrum. Moreover, there is a lack of studies on the influence of degradation of image quality on the performance of computer vision tasks operating on thermal images. Here, we quantify the impact of common image distortions on infrared face recognition, and present a method for aggregating perceptual quality-aware features to improve the identification rates. We use Natural Scene Statistics (NSS) to detect degradation of infrared images, and to adapt the face recognition algorithm to the quality of the test image. The proposed approach applied to a face identification algorithm based on thermal signatures yielded an improvement of rank one recognition rates between 11% and 19%. These results confirm the relevance of image quality assessment for improving biometric identification systems that use thermal images. Camilo G. Rodriguez Pulecio, Hernán Darío Benítez, Alan C. Bovik |
ICIP | 3 |
| 2017 | Blind Quality Assessment of Fused WorldView-3 Images by Using the Combinations of Pansharpening and Hypersharpening ParadigmsabstractWorldView 3 (WV-3) is the first commercially deployed super-spectral, very high-resolution (HR) satellite. However, the resolution of the short-wave infrared (SWIR) bands is much lower than that of the other bands. In this letter, we describe four different approaches, which are combinations of pansharpening and hypersharpening methods, to generate HR SWIR images. Since there are no ground truth HR SWIR images, we also propose a new picture quality predictor to assess hypersharpening performance, without the need for reference images. We describe extensive experiments using actual WV-3 images that demonstrate that some approaches can yield better performance than others, as measured by the proposed blind image quality assessment model of hypersharpened SWIR images. Chiman Kwan, Bence Budavari, Alan C. Bovik, Giovanni Marchisio |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Visual discomfort prediction on stereoscopic 3D images without explicit disparities
Jun Zhou 0007, Jun Sun 0005, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2017 | Binocular spatial activity and reverse saliency driven no-reference stereopair quality assessment
Lixiong Liu, Che-Chun Su, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2017 | Learning quality assessment of retargeted images
Bo Yan 0001, Bahetiyaer Bare, Ke Li 0010, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2017 | SpEED-QA: Spatial Efficient Entropic Differencing for Image and Video QualityabstractMany image and video quality assessment (I/VQA) models rely on data transformations of image/video frames, which increases their programming and computational complexity. By comparison, some of the most popular I/VQA models deploy simple spatial bandpass operations at a couple of scales, making them attractive for efficient implementation. Here we design reduced-reference image and video quality models of this type that are derived from the high-performance reduced reference entropic differencing (RRED) I/VQA models. A new family of I/VQA models, which we call the spatial efficient entropic differencing for quality assessment (SpEED-QA) model, relies on local spatial operations on image frames and frame differences to compute perceptually relevant image/video quality features in an efficient way. Software for SpEED-QA is available at: http://live.ece.utexas.edu/research/Quality/SpEED_Demo.zip. Christos G. Bampis, Praful Gupta, Rajiv Soundararajan, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2017 | Continuous Prediction of Streaming Video QoE Using Dynamic NetworksabstractStreaming video data accounts for a large portion of mobile network traffic. Given the throughput and buffer limitations that currently affect mobile streaming, compression artifacts and rebuffering events commonly occur. Being able to predict the effects of these impairments on perceived video quality of experience (QoE) could lead to improved resource allocation strategies enabling the delivery of higher quality video. Toward this goal, we propose a first of a kind continuous QoE prediction engine. Prediction is based on a nonlinear autoregressive model with exogenous outputs. Our QoE prediction model is driven by three QoE-aware inputs: An objective measure of perceptual video quality, rebuffering-aware information, and a QoE memory descriptor that accounts for recency. We evaluate our method on a recent QoE dataset containing continuous time subjective scores. Christos G. Bampis, Zhi Li 0001, Alan C. Bovik |
IEEE Signal Process. Lett. | 3 |
| 2017 | Single-Scale Fusion: An Effective Approach to Merging ImagesabstractDue to its robustness and effectiveness, multi-scale fusion (MSF) based on the Laplacian pyramid decomposition has emerged as a popular technique that has shown utility in many applications. Guided by several intuitive measures (weight maps) the MSF process is versatile and straightforward to be implemented. However, the number of pyramid levels increases with the image size, which implies sophisticated data management and memory accesses, as well as additional computations. Here, we introduce a simplified formulation that reduces MSF to only a single level process. Starting from the MSF decomposition, we explain both mathematically and intuitively (visually) a way to simplify the classical MSF approach with minimal loss of information. The resulting single-scale fusion (SSF) solution is a close approximation of the MSF process that eliminates important redundant computations. It also provides insights regarding why MSF is so effective. While our simplified expression is derived in the context of high dynamic range imaging, we show its generality on several well-known fusion-based applications, such as image compositing, extended depth of field, medical imaging, and blending thermal (infrared) images with visible light. Besides visual validation, quantitative evaluations demonstrate that our SSF strategy is able to yield results that are highly competitive with traditional MSF approaches. Codruta O. Ancuti, Cosmin Ancuti, Christophe De Vleeschouwer, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2017 | Study of Temporal Effects on Subjective Video Quality of ExperienceabstractHTTP adaptive streaming is being increasingly deployed by network content providers, such as Netflix and YouTube. By dividing video content into data chunks encoded at different bitrates, a client is able to request the appropriate bitrate for the segment to be played next based on the estimated network conditions. However, this can introduce a number of impairments, including compression artifacts and rebuffering events, which can severely impact an end-user's quality of experience (QoE). We have recently created a new video quality database, which simulates a typical video streaming application, using long video sequences and interesting Netflix content. Going beyond previous efforts, the new database contains highly diverse and contemporary content, and it includes the subjective opinions of a sizable number of human subjects regarding the effects on QoE of both rebuffering and compression distortions. We observed that rebuffering is always obvious and unpleasant to subjects, while bitrate changes may be less obvious due to content-related dependencies. Transient bitrate drops were preferable over rebuffering only on low complexity video content, while consistently low bitrates were poorly tolerated. We evaluated different objective video quality assessment algorithms on our database and found that objective video quality models are unreliable for QoE prediction on videos suffering from both rebuffering events and bitrate changes. This implies the need for more general QoE models that take into account objective quality models, rebuffering-aware information, and memory. The publicly available video content as well as metadata for all of the videos in the new database can be found at http://live.ece.utexas.edu/research/LIVE_NFLXStudy/nflx_index.html. Christos G. Bampis, Zhi Li 0001, Anush K. Moorthy, Ioannis Katsavounidis, Anne Aaron, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2017 | Graph-Driven Diffusion and Random Walk Schemes for Image SegmentationabstractWe propose graph-driven approaches to image segmentation by developing diffusion processes defined on arbitrary graphs. We formulate a solution to the image segmentation problem modeled as the result of infectious wavefronts propagating on an image-driven graph where pixels correspond to nodes of an arbitrary graph. By relating the popular Susceptible - Infected - Recovered epidemic propagation model to the Random Walker algorithm, we develop the Normalized Random Walker and a lazy random walker variant. The underlying iterative solutions of these methods are derived as the result of infections transmitted on this arbitrary graph. The main idea is to incorporate a degree-aware term into the original Random Walker algorithm in order to account for the node centrality of every neighboring node and to weigh the contribution of every neighbor to the underlying diffusion process. Our lazy random walk variant models the tendency of patients or nodes to resist changes in their infection status. We also show how previous work can be naturally extended to take advantage of this degreeaware term which enables the design of other novel methods. Through an extensive experimental analysis, we demonstrate the reliability of our approach, its small computational burden and the dimensionality reduction capabilities of graph-driven approaches. Without applying any regular grid constraint, the proposed graph clustering scheme allows us to consider pixellevel, node-level approaches and multidimensional input data by naturally integrating the importance of each node to the final clustering or segmentation solution. A software release containing implementations of this work and supplementary material can be found at: http://cvsp.cs.ntua.gr/research/GraphClustering/. Christos G. Bampis, Petros Maragos, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2017 | No-Reference Quality Assessment of Screen Content PicturesabstractRecent years have witnessed a growing number of image and video centric applications on mobile, vehicular, and cloud platforms, involving a wide variety of digital screen content images. Unlike natural scene images captured with modern high fidelity cameras, screen content images are typically composed of fewer colors, simpler shapes, and a larger frequency of thin lines. In this paper, we develop a novel blind/no-reference (NR) model for accessing the perceptual quality of screen content pictures with big data learning. The new model extracts four types of features descriptive of the picture complexity, of screen content statistics, of global brightness quality, and of the sharpness of details. Comparative experiments verify the efficacy of the new model as compared with existing relevant blind picture quality assessment algorithms applied on screen content image databases. A regression module is trained on a considerable number of training samples labeled with objective visual quality predictions delivered by a high-performance full-reference method designed for screen content image quality assessment (IQA). This results in an opinion-unaware NR blind screen content IQA algorithm. Our proposed model delivers computational efficiency and promising performance. The source code of the new model will be available at: https://sites.google.com/site/guke198701/publications. Ke Gu 0001, Jun Zhou 0007, Junfei Qiao 0001, Guangtao Zhai, Weisi Lin, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2017 | Quality Assessment of Perceptual Crosstalk on Two-View Auto-Stereoscopic DisplaysabstractCrosstalk is one of the most severe factors affecting the perceived quality of stereoscopic 3D images. It arises from a leakage of light intensity between multiple views, as in auto-stereoscopic displays. Well-known determinants of crosstalk include the co-location contrast and disparity of the left and right images, which have been dealt with in prior studies. However, when a natural stereo image that contains complex naturalistic spatial characteristics is viewed on an auto-stereoscopic display, other factors may also play an important role in the perception of crosstalk. Here, we describe a new way of predicting the perceived severity of crosstalk, which we call the Binocular Perceptual Crosstalk Predictor (BPCP). BPCP uses measurements of three complementary 3D image properties (texture, structural duplication, and binocular summation) in combination with two well-known factors (co-location contrast and disparity) to make predictions of crosstalk on two-view auto-stereoscopic displays. The new BPCP model includes two masking algorithms and a binocular pooling method. We explore a new masking phenomenon that we call duplicated structure masking, which arises from structural correlations between the original and distorted objects. We also utilize an advanced binocular summation model to develop a binocular pooling algorithm. Our experimental results indicate that BPCP achieves high correlations against subjective test results, improving upon those delivered by previous crosstalk prediction models. Jongyoo Kim, Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2017 | No-Reference Quality Assessment of Tone-Mapped HDR PicturesabstractBeing able to automatically predict digital picture quality, as perceived by human observers, has become important in many applications where humans are the ultimate consumers of displayed visual information. Standard dynamic range (SDR) images provide 8 b/color/pixel. High dynamic range (HDR) images, which are usually created from multiple exposures of the same scene, can provide 16 or 32 b/color/pixel, but must be tonemapped to SDR for display on standard monitors. Multi-exposure fusion techniques bypass HDR creation, by fusing the exposure stack directly to SDR format while aiming for aesthetically pleasing luminance and color distributions. Here, we describe a new no-reference image quality assessment (NR IQA) model for HDR pictures that is based on standard measurements of the bandpass and on newly conceived differential natural scene statistics (NSS) of HDR pictures. We derive an algorithm from the model which we call the HDR IMAGE GRADient-based Evaluator. NSS models have previously been used to devise NR IQA models that effectively predict the subjective quality of SDR images, but they perform significantly worse on tonemapped HDR content. Toward ameliorating this we make here the following contributions: 1) we design HDR picture NR IQA models and algorithms using both standard space-domain NSS features as well as novel HDR-specific gradient-based features that significantly elevate prediction performance; 2) we validate the proposed models on a large-scale crowdsourced HDR image database; and 3) we demonstrate that the proposed models also perform well on legacy natural SDR images. The software is available at: http://live.ece.utexas.edu/research/Quality/higradeRelease.zip. Debarati Kundu, Deepti Ghadiyaram, Alan C. Bovik, Brian L. Evans |
IEEE Trans. Image Process. | 3 |
| 2017 | Large-Scale Crowdsourced Study for Tone-Mapped HDR PicturesabstractMeasuring digital picture quality, as perceived by human observers, is increasingly important in many applications in which humans are the ultimate consumers of visual information. Standard dynamic range (SDR) images provide 8 b/color/pixel. High dynamic range (HDR) images, usually created from multiple exposures of the same scene, can provide 16 or 32 b/color/pixel, but need to be tonemapped to SDR for display on standard monitors. Multiexposure fusion (MEF) techniques bypass HDR creation by fusing an exposure stack directly to SDR images to achieve aesthetically pleasing luminance and color distributions. Many HDR and MEF databases have a relatively small number of images and human opinion scores, obtained under stringently controlled conditions, thereby limiting realistic viewing. Moreover, many of these databases are intended to compare tone-mapping algorithms, rather than being specialized for developing and comparing image quality assessment models. To overcome these challenges, we conducted a massively crowdsourced online subjective study. The primary contributions described in this paper are: 1) the new ESPL-LIVE HDR Image Database that we created containing diverse images obtained by tone-mapping operators and MEF algorithms, with and without post-processing; 2) a large-scale subjective study that we conducted using a crowdsourced platform to gather more than 300 000 opinion scores on 1811 images from over 5000 unique observers; and 3) a detailed study of the correlation performance of the state-of-the-art no-reference image quality assessment algorithms against human opinion scores of these images. The database is available at http://signal.ece.utexas.edu/%7Edebarati/HDRDatabase.zip. Debarati Kundu, Deepti Ghadiyaram, Alan C. Bovik, Brian L. Evans |
IEEE Trans. Image Process. | 3 |
| 2017 | Predicting the Quality of Fused Long Wave Infrared and Visible Light ImagesabstractThe capability to automatically evaluate the quality of long wave infrared (LWIR) and visible light images has the potential to play an important role in determining and controlling the quality of a resulting fused LWIR-visible light image. Extensive work has been conducted on studying the statistics of natural LWIR and visible images. Nonetheless, there has been little work done on analyzing the statistics of fused LWIR and visible images and associated distortions. In this paper, we analyze five multi-resolution-based image fusion methods in regards to several common distortions, including blur, white noise, JPEG compression, and non-uniformity. We study the natural scene statistics of fused images and how they are affected by these kinds of distortions. Furthermore, we conducted a human study on the subjective quality of pristine and degraded fused LWIR-visible images. We used this new database to create an automatic opinion-distortion-unaware fused image quality model and analyzer algorithm. In the human study, 27 subjects evaluated 750 images over five sessions each. We also propose an opinion-aware fused image quality analyzer, whose relative predictions with respect to other state-of-the-art models correlate better with human perceptual evaluations than competing methods. An implementation of the proposed fused image quality measures can be found at https://github.com/ujemd/NSS-of-LWIR-and-Vissible-Images. Also, the new database can be found at http://bit.ly/2noZlbQ. David Eduardo Moreno-Villamarin, Hernán Darío Benítez, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2017 | Enhancement of Visual Comfort and Sense of Presence on Stereoscopic 3D ImagesabstractConventional stereoscopic 3D (S3D) displays do not provide accommodation depth cues of the 3D image or video contents being viewed. The sense of content depths is thus limited to cues supplied by motion parallax (for 3D video), stereoscopic vergence cues created by presenting left and right views to the respective eyes, and other contextual and perspective depth cues. The absence of accommodation cues can induce two kinds of accommodation vergence mismatches (AVM) at the fixation and peripheral points, which can result in severe visual discomfort. With the aim of alleviating discomfort arising from AVM, we propose a new visual comfort enhancement approach for processing S3D visual signals to deliver a more comfortable 3D viewing experience at the display. This is accomplished via an optimization process whereby a predictive indicator of visual discomfort is minimized, while still aiming to maintain the viewer's sense of 3D presence by performing a suitable parallax shift, and by directed blurring of the signal. Our processing framework is defined on 3D visual coordinates that reflect the nonuniform resolution of retinal sensors and that uses a measure of 3D saliency strength. An appropriate level of blur that corresponds to the degree of parallax shift is found, making it possible to produce synthetic accommodation cues implemented using a perceptively relevant filter. By this method, AVM, the primary contributor to the discomfort felt when viewing S3D images, is reduced. We show via a series of subjective experiments that the proposed approach improves visual comfort while preserving the sense of 3D presence. Heeseok Oh, Jongyoo Kim, Jinwoo Kim 0005, Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2017 | Melanoma Classification on Dermoscopy Images Using a Neural Network Ensemble ModelabstractWe develop a novel method for classifying melanocytic tumors as benign or malignant by the analysis of digital dermoscopy images. The algorithm follows three steps: first, lesions are extracted using a self-generating neural network (SGNN); second, features descriptive of tumor color, texture and border are extracted; and third, lesion objects are classified using a classifier based on a neural network ensemble model. In clinical situations, lesions occur that are too large to be entirely contained within the dermoscopy image. To deal with this difficult presentation, new border features are proposed, which are able to effectively characterize border irregularities on both complete lesions and incomplete lesions. In our model, a network ensemble classifier is designed that combines back propagation (BP) neural networks with fuzzy neural networks to achieve improved performance. Experiments are carried out on two diverse dermoscopy databases that include images of both the xanthous and caucasian races. The results show that classification accuracy is greatly enhanced by the use of the new border features and the proposed classifier model. Fengying Xie, Haidi Fan, Zhiguo Jiang 0001, Rusong Meng, Alan C. Bovik |
IEEE Trans. Medical Imaging | 6 |
| 2016 | Night-time dehazing by fusionabstractWe introduce an effective technique to enhance night-time hazy scenes. Our technique builds on multi-scale fusion approach that use several inputs derived from the original image. Inspired by the dark-channel [1] we estimate night-time haze computing the airlight component on image patch and not on the entire image. We do this since under night-time conditions, the lighting generally arises from multiple artificial sources, and is thus intrinsically non-uniform. Selecting the size of the patches is non-trivial, since small patches are desirable to achieve fine spatial adaptation to the atmospheric light, this might also induce poor light estimates and reduced chance of capturing hazy pixels. For this reason, we deploy multiple patch sizes, each generating one input to a multiscale fusion process. Moreover, to reduce the glowing effect and emphasize the finest details, we derive a third input. For each input, a set of weight maps are derived so as to assign higher weights to regions of high contrast, high saliency and small saturation. Finally the derived inputs and the normalized weight maps are blended in a multi-scale fashion using a Laplacian pyramid decomposition. The experimental results demonstrate the effectiveness of our approach compared with recent techniques both in terms of computational efficiency and quality of the outputs. Cosmin Ancuti, Codruta O. Ancuti, Christophe De Vleeschouwer, Alan C. Bovik |
ICIP | 4 |
| 2016 | Projective non-negative matrix factorization for unsupervised graph clusteringabstractWe develop an unsupervised graph clustering and image segmentation algorithm based on non-negative matrix factorization. We consider arbitrarily represented visual signals (in 2D or 3D) and use a graph embedding approach for image or point cloud segmentation. We extend a Projective Non-negative Matrix Factorization variant to include local spatial relationships over the image graph. By using properly defined region features, one can apply our method of unsupervised graph clustering for object and image segmentation. To demonstrate this, we apply our ideas on many graph based segmentation tasks such as 2D pixel and super-pixel segmentation and 3D point cloud segmentation. Finally, we show results comparable to those achieved by the only existing work in pixel based texture segmentation using Nonnegative Matrix Factorization, deploying a simple yet effective extension that is parameter free. We provide a detailed convergence proof of our spatially regularized method and various demonstrations as supplementary material. This novel work brings together graph clustering with image segmentation. Christos G. Bampis, Petros Maragos, Alan C. Bovik |
ICIP | 3 |
| 2016 | Multi-scale underwater descatteringabstractUnderwater images suffer from severe perceptual/visual degradation, due to the dense and non-uniform medium, causing scattering and attenuation of the propagated light that is sensed. Typical restoration methods rely on the popular Dark Channel Prior to estimate the light attenuation factor, and subtract the back-scattered light influence to invert the underwater imaging model. However, as a consequence of using approximate and global estimates of the back-scattered light, most existing single-image underwater descattering techniques perform poorly when restoring non-uniformly illuminated scenes. To mitigate this problem, we introduce a novel approach that estimates the back-scattered light locally, based on the observation of a neighborhood around the pixel of interest. To circumvent issue related to selection of the neighborhood size, we propose to fuse the images obtained over both small and large neighborhoods, each capturing distinct features from the input image. In addition, the Laplacian of the original image is provided as a third input to the fusion process, to enhance texture details in the reconstructed image. These three derived inputs are seamlessly blended via a multi-scale fusion approach, using saliency, contrast, and saturation metrics to weight each input. We perform an extensive qualitative and quantitative evaluation against several specialized techniques. In addition to its simplicity, our method outperforms the previous art on extreme underwater cases of artificial ambient illumination and high water turbidity. Cosmin Ancuti, Codruta O. Ancuti, Christophe De Vleeschouwer, Rafael García, Alan C. Bovik |
ICPR | 5 |
| 2016 | Blind image quality assessment by relative gradient statistics and adaboosting neural network
Lixiong Liu, Qingjie Zhao, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2016 | Blind Picture Upscaling Ratio PredictionabstractNatural scene statistics are well studied in the context of picture quality assessment and have been used in a wide variety of top-performing picture quality prediction models. Upscaling artifacts have been measured with regards to quality impairment using these kinds of models. However, the assessment and classification of subtle, less discriminable upscaling artifacts remains an unsolved problem. The nearly imperceptible artifacts pertaining to the extent and type of upscaling have not been predicted using natural scene statistics (NSS)-based models. We develop an accurate model for predicting the upscaling ratio applied to any natural image. By decomposing an input image frame using an orthogonal filter bank and locally normalizing the resulting responses, we show that the local energy terms can be used to predict the upscaling ratio. In fact, a simple linear regressor can be trained on these energy measurements; hence, no hyperparameter tuning is necessary. We compare the proposed model with other no-reference models using real-world data contained in the Netflix collection. Todd Richard Goodall, Ioannis Katsavounidis, Zhi Li 0001, Anne Aaron, Alan C. Bovik |
IEEE Signal Process. Lett. | 5 |
| 2016 | Massive Online Crowdsourced Study of Subjective and Objective Picture QualityabstractMost publicly available image quality databases have been created under highly controlled conditions by introducing graded simulated distortions onto high-quality photographs. However, images captured using typical real-world mobile camera devices are usually afflicted by complex mixtures of multiple distortions, which are not necessarily well-modeled by the synthetic distortions found in existing databases. The originators of existing legacy databases usually conducted human psychometric studies to obtain statistically meaningful sets of human opinion scores on images in a stringently controlled visual environment, resulting in small data collections relative to other kinds of image analysis databases. Toward overcoming these limitations, we designed and created a new database that we call the LIVE In the Wild Image Quality Challenge Database, which contains widely diverse authentic image distortions on a large number of images captured using a representative variety of modern mobile devices. We also designed and implemented a new online crowdsourcing system, which we have used to conduct a very large-scale, multi-month image quality assessment (IQA) subjective study. Our database consists of over 350 000 opinion scores on 1162 images evaluated by over 8100 unique human observers. Despite the lack of control over the experimental environments of the numerous study participants, we demonstrate excellent internal consistency of the subjective data set. We also evaluate several top-performing blind IQA algorithms on it and present insights on how the mixtures of distortions challenge both end users as well as automatic perceptual quality prediction models. The new database is available for public use at http://live.ece.utexas.edu/research/ChallengeDB/index.html. Deepti Ghadiyaram, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2016 | Tasking on Natural Statistics of Infrared ImagesabstractNatural scene statistics (NSSs) provide powerful, perceptually relevant tools that have been successfully used for image quality analysis of visible light images. Since NSS capture statistical regularities that arise from the physical world, they are relevant to long wave infrared (LWIR) images, which differ from visible light images mainly by the wavelengths captured at the imaging sensors. We show that NSS models of bandpass LWIR images are similar to those of visible light images, but with different parameterizations. Using this difference, we exploit the power of NSS to successfully distinguish between LWIR images and visible light images. In addition, we study distortions unique to LWIR and find directional models useful for detecting the halo effect, simple bandpass models useful for detecting hotspots, and combinations of these models useful for measuring the degree of non-uniformity present in many LWIR images. For local distortion identification and measurement, we also describe a method for generating distortion maps using NSS features. To facilitate our evaluation, we analyze the NSS of LWIR images under pristine and distorted conditions, using four databases, each captured with a different IR camera. Predicting human performance for assessing distortion and quality in LWIR images is critical for task efficacy. We find that NSS features improve human targeting task performance prediction. Furthermore, we conducted a human study on the perceptual quality of noise-and blur-distorted LWIR images and create a new blind image quality predictor for IR images. Todd Richard Goodall, Alan C. Bovik, Nicholas G. Paulter Jr. |
IEEE Trans. Image Process. | 2 |
| 2016 | A Completely Blind Video Integrity OracleabstractConsiderable progress has been made toward developing still picture perceptual quality analyzers that do not require any reference picture and that are not trained on human opinion scores of distorted images. However, there do not yet exist any such completely blind video quality assessment (VQA) models. Here, we attempt to bridge this gap by developing a new VQA model called the video intrinsic integrity and distortion evaluation oracle (VIIDEO). The new model does not require the use of any additional information other than the video being quality evaluated. VIIDEO embodies models of intrinsic statistical regularities that are observed in natural vidoes, which are used to quantify disturbances introduced due to distortions. An algorithm derived from the VIIDEO model is thereby able to predict the quality of distorted videos without any external knowledge about the pristine source, anticipated distortions, or human judgments of video quality. Even with such a paucity of information, we are able to show that the VIIDEO algorithm performs much better than the legacy full reference quality measure MSE on the LIVE VQA database and delivers performance comparable with a leading human judgment trained blind VQA model. We believe that the VIIDEO algorithm is a significant step toward making real-time monitoring of completely blind video quality possible. Anish Mittal, Michele A. Saad, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2016 | Stereoscopic 3D Visual Discomfort Prediction: A Dynamic Accommodation and Vergence Interaction ModelabstractThe human visual system perceives 3D depth following sensing via its binocular optical system, a series of massively parallel processing units, and a feedback system that controls the mechanical dynamics of eye movements and the crystalline lens. The process of accommodation (focusing of the crystalline lens) and binocular vergence is controlled simultaneously and symbiotically via cross-coupled communication between the two critical depth computation modalities. The output responses of these two subsystems, which are induced by oculomotor control, are used in the computation of a clear and stable cyclopean 3D image from the input stimuli. These subsystems operate in smooth synchronicity when one is viewing the natural world; however, conflicting responses can occur when viewing stereoscopic 3D (S3D) content on fixed displays, causing physiological discomfort. If such occurrences could be predicted, then they might also be avoided (by modifying the acquisition process) or ameliorated (by changing the relative scene depth). Toward this end, we have developed a dynamic accommodation and vergence interaction (DAVI) model that successfully predicts visual discomfort on S3D images. The DAVI model is based on the phasic and reflex responses of the fast fusional vergence mechanism. Quantitative models of accommodation and vergence mismatches are used to conduct visual discomfort prediction. Other 3D perceptual elements are included in the proposed method, including sharpness limits imposed by the depth of focus and fusion limits implied by Panum's fusional area. The DAVI predictor is created by training a support vector machine on features derived from the proposed model and on recorded subjective assessment results. The experimental results are shown to produce accurate predictions of experienced visual discomfort. Heeseok Oh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2015 | Scene statistics of authentically distorted images in perceptually relevant color spaces for blind image quality assessmentabstractCurrent top-performing blind image quality assessment (IQA) models rely on benchmark databases comprising of singly distorted images, thereby learning image features that are only adequate to predict human perceived visual quality on such inauthentic distortions. Furthermore, the underlying image features of these models are often extracted from the achromatic luminance channel and could sometimes fail to account for the loss of their perceived quality that might potentially be distinctly captured in a different image modality. In this work, we propose a novel IQA model that focuses on the natural scene statistics of images afflicted with complex mixtures of unknown, authentic distortions. We derive several feature maps in different perceptually relevant color spaces and extract a large number of image features from them. We demonstrate the remarkable competence of our features in improving the automatic perceptual quality prediction on images containing both synthetic and authentic distortions. Deepti Ghadiyaram, Alan C. Bovik |
ICIP | 2 |
| 2015 | 3D visual discomfort predictor based on neural activity statisticsabstractVisual discomfort assessment (VDA) on stereoscopic images is of fundamental importance for making decisions regarding visual fatigue caused by unnatural binocular alignment. Nevertheless, no solid framework exists to quantify this discomfort using models of the responses of visual neurons. Binocular vision is realized by means of neural mechanisms that subserve the sensorimotor control of eye movements. We propose a neuronal model-based framework called Neural 3D Visual Discomfort Predictor (N3D-VDP) that automatically predicts the level of visual discomfort experienced when viewing stereoscopic 3D (S3D) images. The N3D-VDP model extracts features derived by estimating the neural activity associated with the processing of binocular disparities. In this regard we deploy a model of disparity processing in the extra-striate middle temporal (MT) region of occipital lobe. We compare the performance of N3D-VDP with other recent VDA algorithms using correlations against reported subjective visual discomfort, and show that N3D-VDP is statistically superior to the other methods. Heeseok Oh, Jongyoo Kim, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 4 |
| 2015 | Automatic Channel Network Extraction From Remotely Sensed Images by Singularity AnalysisabstractThe quantitative analysis of channel networks plays an important role in river studies. To provide a quantitative representation of channel networks, we propose anew method that extracts channels from remotely sensed images and estimates their widths. Our fully automated method is based on a recently proposed multiscale singularity index that strongly responds to curvilinear structures but weakly responds to edges. The algorithm produces a channel map using a single image where water and nonwater pixels have contrast, such as a Landsat near-infrared band image or a water index defined on multiple bands. The proposed method provides a robust alternative to the procedures that are used in the remote sensing of fluvial geomorphology and makes the classification and analysis of channel networks easier. The source code of the algorithm is available at http://live.ece.utexas. edu/research/cne/. Furkan Isikdogan, Alan C. Bovik, Paola Passalacqua |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Motion silencing of flicker distortions on naturalistic videos
Lark Kwon Choi, Lawrence K. Cormack, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2015 | A spatiotemporal weighted dissimilarity-based method for video saliency detection
Lijuan Duan, Honggang Qi, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2015 | Closed-Form Correlation Model of Oriented Bandpass Natural ImagesabstractMost prevalent statistical models of natural images characterize only the univariate distributions of divisively normalized bandpass responses or wavelet-like decompositions of them. However, the higher-order dependencies between spatially neighboring responses are not yet well understood. Towards filling this gap, we propose a new closed-form spatial-oriented correlation model that captures statistical regularities between perceptually decomposed natural image luminance samples. We validate the new correlation model on a variety of natural images. Experimental results demonstrate the robustness of the new correlation model across image content. A software release that implements the new closed-form spatial-oriented correlation model is available at http://live.ece.utexas.edu/research/3dnss/bicorr_release.zip. Che-Chun Su, Lawrence K. Cormack, Alan C. Bovik |
IEEE Signal Process. Lett. | 3 |
| 2015 | Referenceless Prediction of Perceptual Fog Density and Perceptual Image DefoggingabstractWe propose a referenceless perceptual fog density prediction model based on natural scene statistics (NSS) and fog aware statistical features. The proposed model, called Fog Aware Density Evaluator (FADE), predicts the visibility of a foggy scene from a single image without reference to a corresponding fog-free image, without dependence on salient objects in a scene, without side geographical camera information, without estimating a depth-dependent transmission map, and without training on human-rated judgments. FADE only makes use of measurable deviations from statistical regularities observed in natural foggy and fog-free images. Fog aware statistical features that define the perceptual fog density index derive from a space domain NSS model and the observed characteristics of foggy images. FADE not only predicts perceptual fog density for the entire image, but also provides a local fog density index for each patch. The predicted fog density using FADE correlates well with human judgments of fog density taken in a subjective study on a large foggy image database. As applications, FADE not only accurately assesses the performance of defogging algorithms designed to enhance the visibility of foggy images, but also is well suited for image defogging. A new FADE-based referenceless perceptual image defogging, dubbed DEnsity of Fog Assessment-based DEfogger (DEFADE) achieves better results for darker, denser foggy images as well as on standard foggy images than the state of the art defogging methods. A software release of FADE and DEFADE is available online for public use: http://live.ece.utexas.edu/research/fog/index.html. Lark Kwon Choi, Jaehee You, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2015 | Toward Naturalistic 2D-to-3D ConversionabstractNatural scene statistics (NSSs) models have been developed that make it possible to impose useful perceptually relevant priors on the luminance, colors, and depth maps of natural scenes. We show that these models can be used to develop 3D content creation algorithms that can convert monocular 2D videos into statistically natural 3D-viewable videos. First, accurate depth information on key frames is obtained via human annotation. Then, both forward and backward motion vectors are estimated and compared to decide the initial depth values, and a compensation process is applied to further improve the depth initialization. Then, the luminance/chrominance and initial depth map are decomposed by a Gabor filter bank. Each subband of depth is modeled to produce a NSS prior term. The statistical color-depth priors are combined with the spatial smoothness constraint in the depth propagation target function as a prior regularizing term. The final depth map associated with each frame of the input 2D video is optimized by minimizing the target function over all subbands. In the end, stereoscopic frames are rendered from the color frames and their associated depth maps. We evaluated the quality of the generated 3D videos using both subjective and objective quality assessment methods. The experimental results obtained on various sequences show that the presented method outperforms several state-of-the-art 2D-to-3D conversion methods. Xun Cao, Ke Lu 0002, Qionghai Dai, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2015 | Transfer Function Model of Physiological Mechanisms Underlying Temporal Visual Discomfort Experienced When Viewing Stereoscopic 3D ImagesabstractWhen viewing 3D images, a sense of visual comfort (or lack of) is developed in the brain over time as a function of binocular disparity and other 3D factors. We have developed a unique temporal visual discomfort model (TVDM) that we use to automatically predict the degree of discomfort felt when viewing stereoscopic 3D (S3D) images. This model is based on physiological mechanisms. In particular, TVDM is defined as a second-order system capturing relevant neuronal elements of the visual pathway from the eyes and through the brain. The experimental results demonstrate that the TVDM transfer function model produces predictions that correlate highly with the subjective visual discomfort scores contained in the large public databases. The transfer function analysis also yields insights into the perceptual processes that yield a stable S3D image. Taewan Kim 0002, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2015 | Disparity Estimation on Stereo MammogramsabstractWe consider the problem of depth estimation on digital stereo mammograms. Being able to elucidate 3D information from stereo mammograms is an important precursor to conducting 3D digital analysis of data from this promising new modality. The problem is generally much harder than the classic stereo matching problem on visible light images of the natural world, since nearly all of the 3D structural information of interest exists as complex network of multilayered, heavily occluded curvilinear structures. Toward addressing this difficult problem, we formulate a new stereo model that minimizes a global energy functional to densely estimate disparity on stereo mammogram images, by introducing a new singularity index as a constraint to obtain better estimates of disparity along critical curvilinear structures. Curvilinear structures, such as vasculature and spicules, are particularly salient structures in the breast, and being able to accurately position them in 3D is a valuable goal. Experiments on synthetic images with known ground truth and on real stereo mammograms highlight the advantages of the proposed stereo model over the canonical stereo model. S. M. Gautam, Alan C. Bovik, Mia K. Markey |
IEEE Trans. Image Process. | 2 |
| 2015 | 3D Visual Discomfort Predictor: Analysis of Disparity and Neural Activity StatisticsabstractBeing able to predict the degree of visual discomfort that is felt when viewing stereoscopic 3D (S3D) images is an important goal toward ameliorating causative factors, such as excessive horizontal disparity, misalignments or mismatches between the left and right views of stereo pairs, or conflicts between different depth cues. Ideally, such a model should account for such factors as capture and viewing geometries, the distribution of disparities, and the responses of visual neurons. When viewing modern 3D displays, visual discomfort is caused primarily by changes in binocular vergence while accommodation in held fixed at the viewing distance to a flat 3D screen. This results in unnatural mismatches between ocular fixations and ocular focus that does not occur in normal direct 3D viewing. This accommodation vergence conflict can cause adverse effects, such as headaches, fatigue, eye strain, and reduced visual ability. Binocular vision is ultimately realized by means of neural mechanisms that subserve the sensorimotor control of eye movements. Realizing that the neuronal responses are directly implicated in both the control and experience of 3D perception, we have developed a model-based neuronal and statistical framework called the 3D visual discomfort predictor (3D-VDP)that automatically predicts the level of visual discomfort that is experienced when viewing S3D images. 3D-VDP extracts two types of features: 1) coarse features derived from the statistics of binocular disparities and 2) fine features derived by estimating the neural activity associated with the processing of horizontal disparities. In particular, we deploy a model of horizontal disparity processing in the extrastriate middle temporal region of occipital lobe. We compare the performance of 3D-VDP with other recent discomfort prediction algorithms with respect to correlation against recorded subjective visual discomfort scores,and show that 3D-VDP is statistically superior to the other methods. Jincheol Park, Heeseok Oh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2015 | Oriented Correlation Models of Distorted Natural Images With Application to Natural Stereopair Quality EvaluationabstractIn recent years, bandpass statistical models of natural, photographic images of the world have been used with great success to solve highly diverse problems involving image representation, image repair, image quality assessment (IQA), and image compression. One missing element has been a reliable and generic model of spatial image correlation that reflects the distributions of oriented and relatively oriented spatial structures. We have developed such a model for bandpass pristine images and have generalized it here to also capture the spatial correlation structure of bandpass distorted images. The model applies well to both luminance and depth images. As a demonstration of the usefulness of the generalized model, we develop a new no-reference stereoscopic/3D IQA framework, dubbed stereoscopic/3D blind image naturalness quality index, which utilizes both univariate and generalized bivariate natural scene statistics (NSS) models. We first validate the robustness and effectiveness of these novel bivariate and correlation NSS features extracted from distorted stereopairs, then demonstrate that they are predictive of distortion severity. Our experimental results show that the resulting 3D image quality predictor based in part on the new model outperforms state-of-the-art full- and no-reference 3D IQA algorithms on both symmetrically and asymmetrically distorted stereoscopic image pairs. Che-Chun Su, Lawrence K. Cormack, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2015 | A Feature-Enriched Completely Blind Image Quality EvaluatorabstractExisting blind image quality assessment (BIQA) methods are mostly opinion-aware. They learn regression models from training images with associated human subjective scores to predict the perceptual quality of test images. Such opinion-aware methods, however, require a large amount of training samples with associated human subjective scores and of a variety of distortion types. The BIQA models learned by opinion-aware methods often have weak generalization capability, hereby limiting their usability in practice. By comparison, opinion-unaware methods do not need human subjective scores for training, and thus have greater potential for good generalization capability. Unfortunately, thus far no opinion-unaware BIQA method has shown consistently better quality prediction accuracy than the opinion-aware methods. Here, we aim to develop an opinion-unaware BIQA method that can compete with, and perhaps outperform, the existing opinion-aware methods. By integrating the features of natural image statistics derived from multiple cues, we learn a multivariate Gaussian model of image patches from a collection of pristine natural images. Using the learned multivariate Gaussian model, a Bhattacharyya-like distance is used to measure the quality of each image patch, and then an overall quality score is obtained by average pooling. The proposed BIQA method does not need any distorted sample images nor subjective quality scores for training, yet extensive experiments demonstrate its superior quality-prediction performance to the state-of-the-art opinion-aware BIQA methods. The MATLAB source code of our algorithm is publicly available at www.comp.polyu.edu.hk/~cslzhang/IQA/ILNIQE/ILNIQE.htm. Lin Zhang 0014, Lei Zhang 0006, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2014 | New bivariate statistical model of natural image correlationsabstractWe perform bivariate statistical analysis and modeling of the joint distributions of spatially adjacent sub-band responses for both luminance/chrominance and range data in natural scenes. In particular, we introduce a multivariate generalized Gaussian distribution and an exponentiated sine function to model the underlying statistics and correlations. The experimental results show that the bivariate statistics relating spatially adjacent pixels in both 2D color images and range maps are well described by the proposed models. We validate the robustness of the proposed bivariate models using a multi-variate statistical hypothesis test, and further demonstrate their effectiveness with application to a prototype depth estimation algorithm. Che-Chun Su, Lawrence K. Cormack, Alan C. Bovik |
ICASSP | 3 |
| 2014 | Adaptive video transmission with subjective quality constraintsabstractWe conducted a subjective study wherein we found that viewers' Quality of Experience (QoE) was strongly correlated with the empirical cumulative distribution function (eCDF) of the predicted video quality. Based on this observation, we propose a rate-adaptation algorithm that can incorporate QoE constraints on the empirical cumulative quality distribution per user. Simulation results show that the proposed technique can reduce network resource consumption by 29% over conventional average-quality maximized rate-adaptation algorithms. Chao Chen 0006, Gustavo de Veciana, Alan C. Bovik, Robert W. Heath Jr. |
ICIP | 4 |
| 2014 | Assessment of video naturalness using time-frequency statisticsabstractSuccessful video quality analysers make use of a reference video to compare against or by training on a database of human rated distorted videoes as priors, both of which are either not available or difficult to obtain in many practical scenarios. Although efforts have been made towards designing still picture quality analyzers that are `completely blind' and do not require any prior training on, or exposure to, distorted images or human opinions of them [1], we are attempting to fill an important but challenging gap by designing a `completely blind' video naturalness analyser. The principle of this new approach is based on the regularties observed in time-frequency relationships of natural vidoes across time. Our experimental results on the LIVE VQA (video quality assessment) database [2] show that, even with no prior knowledge, the new VQA algorithm performs better than the full reference (FR) quality measure PSNR. The approach is very lean in computational expense which makes it a very good candidate for real time signal processing applications. Anish Mittal, Michele A. Saad, Alan C. Bovik |
ICIP | 3 |
| 2014 | A hierarchical Bayesian-map approach to computational imagingabstractWe present a novel approach to inverse problems in imaging based on a Hierarchical Bayesian-MAP (HB-MAP) formulation. In this paper we specifically focus on the difficult and basic inverse problem of multi-sensor (tomographic) imaging wherein the source image of interest is viewed from multiple directions by independent sensors. We employ a Probabilistic Graphical Modeling extension of the Compound Gaussian (CG) distribution as a global image prior into a Hierarchical Bayesian inference procedure. We first demonstrate the performance of the algorithm on Monte-Carlo trials followed by empirical data involving natural (optical) images. We demonstrate how our algorithm outperforms many of the previous approaches in the literature including Filtered Back-projection (FBP) and a variety of state-of-the-art compressive sensing (CS) algorithms. Raghu G. Raj, Alan C. Bovik |
ICIP | 2 |
| 2014 | Delivery quality score model for Internet videoabstractThe vast majority of today's internet video services are consumed over-the-top (OTT) via reliable streaming (HTTP via TCP), where the primary noticeable delivery-related impairments are startup delay and stalling. In this paper we introduce an objective model called the delivery quality score (DQS) model, to predict user's QoE in the presence of such impairments. We describe a large subjective study that we carried out to tune and validate this model. Our experiments demonstrate that the DQS model correlates highly with the subjective data and that it outperforms other emerging models. Hojatollah Yeganeh, Roman C. Kordasiewicz, Michael Gallant, Deepti Ghadiyaram, Alan C. Bovik |
ICIP | 5 |
| 2014 | Binocular mismatch induced by luminance discrepancies on stereoscopic imagesabstractLuminance discrepancies between image pairs occur owing to inconsistent parameters between stereoscopic camera devices and from imperfect capture conditions. Such discrepancies induce binocular mismatches and affect the visual comfort that is felt by viewers, as well as their ability to fuse stereoscopic. To better understand and observe this effect, we built a stereoscopic images database of 240 luminance discrepancy images and 30 natural images with subjective scores of visual discomfort and fusion difficulty. Two features, binocular contrast and luminance similarity were extracted to analyze the relationship between the subjective scores and the luminance discrepancies. Structural dissimilarity and average luminance are used to predict the effects of binocular mismatches. The experimental results show that the combination of binocular contrast, structural dissimilarity and average luminance exhibits high consistency with subjective scores of visual discomfort, fusion difficulty and overall binocular mismatches in terms of Spearman's Rank Ordered Correlation Coefficient. Jun Zhou 0007, Jun Sun 0005, Alan C. Bovik |
ICME | 4 |
| 2014 | No-reference image blur index based on singular value curve
Qingbing Sang, Huixin Qi, Xiaojun Wu 0001, Alan C. Bovik |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | No-reference image quality assessment in curvelet domain
Lixiong Liu, Hongping Dong, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2014 | No-reference image quality assessment based on spatial and spectral entropies
Lixiong Liu, Hua Huang 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2014 | Blind image quality assessment using a reciprocal singular value curve
Qingbing Sang, Xiaojun Wu 0001, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2014 | C-DIIVINE: No-reference image quality assessment based on local magnitude and phase statistics of natural scenes
Yi Zhang 0033, Anush K. Moorthy, Damon M. Chandler, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2014 | Face Detection on Distorted Images Augmented by Perceptual Quality-Aware FeaturesabstractMotivated by the proliferation of low-cost digital cameras in mobile devices being deployed in automated surveillance networks, we study the interaction between perceptual image quality and a classic computer vision task of face detection. We quantify the degradation in performance of a popular and effective face detector when human-perceived image quality is degraded by distortions commonly occurring in capture, storage, and transmission of facial images, including noise, blur, and compression. It is observed that, within a certain range of perceived image quality, a modest increase in image quality can drastically improve face detection performance. These results can be used to guide resource or bandwidth allocation in acquisition or communication/delivery systems that are associated with face detection tasks. A new set of features, called qualHOG, are proposed for robust facedetection that augments face-indicative Histogram of Oriented Gradients (HOG) features with perceptual quality-aware spatial Natural Scene Statistics (NSS) features. Face detectors trained on these new features provide statistically significant improvement in tolerance to image distortions over a strong baseline. Distortiondependent and distortion-unaware variants of the face detectors are proposed and evaluated on a large database of face images representing a wide range of distortions. A biased variant of the training algorithm is also proposed that further enhances the robustness of these face detectors. To facilitate this research, we created a new distorted face database (DFD), containing face and non-face patches from images impaired by a variety of common distortion types and levels. This new data set and relevant code are available for download and further experimentation at www.live.ece.utexas.edu/research/Quality/index.htm. Suriya Gunasekar, Joydeep Ghosh, Alan C. Bovik |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Modeling the Time - Varying Subjective Quality of HTTP Video Streams With Rate AdaptationsabstractNewly developed hypertext transfer protocol (HTTP)-based video streaming technologies enable flexible rate-adaptation under varying channel conditions. Accurately predicting the users' quality of experience (QoE) for rate-adaptive HTTP video streams is thus critical to achieve efficiency. An important aspect of understanding and modeling QoE is predicting the up-to-the-moment subjective quality of a video as it is played, which is difficult due to hysteresis effects and nonlinearities in human behavioral responses. This paper presents a Hammerstein-Wiener model for predicting the time-varying subjective quality (TVSQ) of rate-adaptive videos. To collect data for model parameterization and validation, a database of longer duration videos with time-varying distortions was built and the TVSQs of the videos were measured in a large-scale subjective study. The proposed method is able to reliably predict the TVSQ of rate adaptive videos. Since the Hammerstein-Wiener model has a very simple structure, the proposed method is suitable for online TVSQ prediction in HTTP-based streaming. Chao Chen 0006, Lark Kwon Choi, Gustavo de Veciana, Constantine Caramanis, Robert W. Heath Jr., Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2014 | Saliency Prediction on Stereoscopic VideosabstractWe describe a new 3D saliency prediction model that accounts for diverse low-level luminance, chrominance, motion, and depth attributes of 3D videos as well as high-level classifications of scenes by type. The model also accounts for perceptual factors, such as the nonuniform resolution of the human eye, stereoscopic limits imposed by Panum's fusional area, and the predicted degree of (dis) comfort felt, when viewing the 3D video. The high-level analysis involves classification of each 3D video scene by type with regard to estimated camera motion and the motions of objects in the videos. Decisions regarding the relative saliency of objects or regions are supported by data obtained through a series of eye-tracking experiments. The algorithm developed from the model elements operates by finding and segmenting salient 3D space-time regions in a video, then calculating the saliency strength of each segment using measured attributes of motion, disparity, texture, and the predicted degree of visual discomfort experienced. The saliency energy of both segmented objects and frames are weighted using models of human foveation and Panum's fusional area yielding a single predictor of 3D saliency. Haksub Kim, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2014 | 3D Visual Activity Assessment Based on Natural Scene StatisticsabstractOne of the most challenging ongoing issues in the field of 3D visual research is how to perceptually quantify object and surface visualizations that are displayed within a virtual 3D space between a human eye and 3D display. To seek an effective method of quantification, it is necessary to measure various elements related to the perception of 3D objects at different depths. We propose a new framework for quantifying 3D visual information that we call 3D visual activity (3DVA), which utilizes natural scene statistics measured over 3D visual coordinates. We account for important aspects of 3D perception by carrying out a 3D coordinate transform reflecting the nonuniform sampling resolution of the eye and the process of stereoscopic fusion. The 3DVA utilizes the empirical distortions of wavelet coefficients to a parametric generalized Gaussian probability distribution model and a set of 3D perceptual weights. We conducted a series of simulations that demonstrate the effectiveness of the 3DVA for quantifying the statistical dynamics of visual 3D space with respect to disparity, motion, texture, and color. A successful example application is also provided, whereby 3DVA is applied to the problem of predicting visual fatigue experienced when viewing 3D displays. Kwanghyun Lee, Anush K. Moorthy, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2014 | No-Reference Sharpness Assessment of Camera-Shaken Images by Analysis of Spectral StructureabstractThe tremendous explosion of image-, video-, and audio-enabled mobile devices, such as tablets and smart-phones in recent years, has led to an associated dramatic increase in the volume of captured and distributed multimedia content. In particular, the number of digital photographs being captured annually is approaching 100 billion in just the U.S. These pictures are increasingly being acquired by inexperienced, casual users under highly diverse conditions leading to a plethora of distortions, including blur induced by camera shake. In order to be able to automatically detect, correct, or cull images impaired by shake-induced blur, it is necessary to develop distortion models specific to and suitable for assessing the sharpness of camera-shaken images. Toward this goal, we have developed a no-reference framework for automatically predicting the perceptual quality of camera-shaken images based on their spectral statistics. Two kinds of features are defined that capture blur induced by camera shake. One is a directional feature, which measures the variation of the image spectrum across orientations. The second feature captures the shape, area, and orientation of the spectral contours of camera shaken images. We demonstrate the performance of an algorithm derived from these features on new and existing databases of images distorted by camera shake. Taegeun Oh, Jincheol Park, Kalpana Seshadrinathan, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2014 | Blind Prediction of Natural Video QualityabstractWe propose a blind (no reference or NR) video quality evaluation model that is nondistortion specific. The approach relies on a spatio-temporal model of video scenes in the discrete cosine transform domain, and on a model that characterizes the type of motion occurring in the scenes, to predict video quality. We use the models to define video statistics and perceptual features that are the basis of a video quality assessment (VQA) algorithm that does not require the presence of a pristine video to compare against in order to predict a perceptual quality score. The contributions of this paper are threefold. 1) We propose a spatio-temporal natural scene statistics (NSS) model for videos. 2) We propose a motion model that quantifies motion coherency in video scenes. 3) We show that the proposed NSS and motion coherency models are appropriate for quality assessment of videos, and we utilize them to design a blind VQA algorithm that correlates highly with human judgments of quality. The proposed algorithm, called video BLIINDS, is tested on the LIVE VQA database and on the EPFL-PoliMi video database and shown to perform close to the level of top performing reduced and full reference VQA algorithms. Michele A. Saad, Alan C. Bovik, Christophe Charrier |
IEEE Trans. Image Process. | 2 |
| 2014 | Blind Image Quality Assessment Using Joint Statistics of Gradient Magnitude and Laplacian FeaturesabstractBlind image quality assessment (BIQA) aims to evaluate the perceptual quality of a distorted image without information regarding its reference image. Existing BIQA models usually predict the image quality by analyzing the image statistics in some transformed domain, e.g., in the discrete cosine transform domain or wavelet domain. Though great progress has been made in recent years, BIQA is still a very challenging task due to the lack of a reference image. Considering that image local contrast features convey important structural information that is closely related to image perceptual quality, we propose a novel BIQA model that utilizes the joint statistics of two types of commonly used local contrast features: 1) the gradient magnitude (GM) map and 2) the Laplacian of Gaussian (LOG) response. We employ an adaptive procedure to jointly normalize the GM and LOG features, and show that the joint statistics of normalized GM and LOG features have desirable properties for the BIQA task. The proposed model is extensively evaluated on three large-scale benchmark databases, and shown to deliver highly competitive performance with state-of-the-art BIQA models, as well as with some well-known full reference image quality assessment models. Wufeng Xue, Xuanqin Mou, Lei Zhang 0006, Alan C. Bovik, Xiangchu Feng |
IEEE Trans. Image Process. | 4 |
| 2014 | Gradient Magnitude Similarity Deviation: A Highly Efficient Perceptual Image Quality IndexabstractIt is an important task to faithfully evaluate the perceptual quality of output images in many applications, such as image compression, image restoration, and multimedia streaming. A good image quality assessment (IQA) model should not only deliver high quality prediction accuracy, but also be computationally efficient. The efficiency of IQA metrics is becoming particularly important due to the increasing proliferation of high-volume visual data in high-speed networks. We present a new effective and efficient IQA model, called gradient magnitude similarity deviation (GMSD). The image gradients are sensitive to image distortions, while different local structures in a distorted image suffer different degrees of degradations. This motivates us to explore the use of global variation of gradient based local quality map for overall image quality prediction. We find that the pixel-wise gradient magnitude similarity (GMS) between the reference and distorted images combined with a novel pooling strategy-the standard deviation of the GMS map-can predict accurately perceptual image quality. The resulting GMSD algorithm is much faster than most state-of-the-art IQA methods, and delivers highly competitive prediction accuracy. MATLAB source code of GMSD can be downloaded at http://www4.comp.polyu.edu.hk/~cslzhang/IQA/GMSD/GMSD.htm. Wufeng Xue, Lei Zhang 0006, Xuanqin Mou, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2014 | Multimodal Interactive Continuous Scoring of Subjective 3D Video Quality of ExperienceabstractPeople experience a variety of 3D visual programs, such as 3D cinema, 3D TV and 3D games, making it necessary to deploy reliable methodologies for predicting each viewer's subjective experience. We propose a new methodology that we call multimodal interactive continuous scoring of quality (MICSQ). MICSQ is composed of a device interaction process between the 3D display and a separate device (PC, tablet, etc.) used as an assessment tool, and a human interaction process between the subject(s) and the separate device. The scoring process is multimodal, using aural and tactile cues to help engage and focus the subject(s) on their tasks by enhancing neuroplasticity. Recorded human responses to 3D visualizations obtained via MICSQ correlate highly with measurements of spatial and temporal activity in the 3D video content. We have also found that 3D quality of experience (QoE) assessment results obtained using MICSQ are more reliable over a wide dynamic range of content than obtained by the conventional single stimulus continuous quality evaluation (SSCQE) protocol. Moreover, the wireless device interaction process makes it possible for multiple subjects to assess 3D QoE simultaneously in a large space such as a movie theater, at different viewing angles and distances. We conducted a series of interesting 3D experiments showing the accuracy and versatility of the new system, while yielding new findings on visual comfort in terms of disparity, motion and an interesting relation between the naturalness and depth of field (DOF) of a stereo camera. Taewan Kim 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Multim. | 4 |
| 2013 | A dynamic system model of time-varying subjective quality of video streams over HTTPabstractNewly developed HTTP-based video streaming technology enables flexible rate-adaptation in varying channel conditions. The users' Quality of Experience (QoE) of rate-adaptive HTTP video streams, however, is not well understood. Therefore, designing QoE-optimized rate-adaptive video streaming algorithms remains a challenging task. An important aspect of understanding and modeling QoE is to be able to predict the up-to-the-moment subjective quality of video as it is played. We propose a dynamic system model to predict the time-varying subjective quality (TVSQ) of rate-adaptive videos that is transported over HTTP. For this purpose, we built a video database and measured TVSQ via a subjective study. A dynamic system model is developed using the database and the measured human data. We show that the proposed model can effectively predict the TVSQ of rate-adaptive videos in an online manner, which is necessary to be able to conduct QoE-optimized online rate-adaptation for HTTP-based video streaming. Chao Chen 0006, Lark Kwon Choi, Gustavo de Veciana, Constantine Caramanis, Robert W. Heath Jr., Alan C. Bovik |
ICASSP | 6 |
| 2013 | Visually Lossless H.264 Compression of Natural VideosabstractWe performed a systematic evaluation of ‘visually lossless’ (VL) threshold selection for H.264/AVC (Advanced Video Coding) compressed natural videos spanning a wide range of content and motion. A psychovisual study was conducted using a two alternative forced choice task design, where by a series of reference vs. compressed video pairs were displayed to the subjects, where bit rates were varied to achieve a spread in the amount of compression. A statistical analysis was conducted on these data to estimate the VL threshold. Based on the visual thresholds estimated from the observed human ratings, we learn a mapping from ‘perceptually relevant’ statistical video features that capture visual lossless-ness, to statistically determined VL threshold. Using this VL threshold, we derive an H.264 compressibility index. This new Compressibility Index is shown to correlate well with human subjective judgments of VL thresholds. We have also made the code for compressibility index available online (Moorthy, A.K. and Bovik, A.C. (2010). H.264 Visually Lossless Compressibility Index (HVLCI), Software Release. http://live.ece.utexas.edu/research/quality/hvlci.zip.) for its use in practical applications and facilitate future research in this area. Anish Mittal, Anush K. Moorthy, Alan C. Bovik |
Comput. J. | 3 |
| 2013 | Passive Three Dimensional Face Recognition Using Iso-Geodesic Contours and Procrustes Analysis
Sina Jahanbin, Rana Jahanbin, Alan C. Bovik |
Int. J. Comput. Vis. | 3 |
| 2013 | Automatic Prediction of Perceptual Image and Video QualityabstractFinding ways to monitor and control the perceptual quality of digital visual media has become a pressing concern as the volume being transported and viewed continues to increase exponentially. This paper discusses the principles and methods of modern algorithms for automatically predicting the quality of visual signals. By casting the problem as analogous to assessing the efficacy of a visual communication system, it is possible to divide the quality assessment problem into understandable modeling subproblems. Along the way, we will visit models of natural images and videos, of visual perception, and a broad spectrum of applications. Alan C. Bovik |
Proc. IEEE | 1 |
| 2013 | Automatic segmentation of dermoscopy images using self-generating neural networks seeded by genetic algorithm
Fengying Xie, Alan C. Bovik |
Pattern Recognit. | 2 |
| 2013 | Full-reference quality assessment of stereopairs accounting for rivalry
Ming-Jun Chen, Che-Chun Su, Do-Kyoung Kwon, Lawrence K. Cormack, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2013 | Perceptually optimized blind repair of natural images
Anush K. Moorthy, Anish Mittal, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2013 | Subjective evaluation of stereoscopic image quality
Anush K. Moorthy, Che-Chun Su, Anish Mittal, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2013 | Making a "Completely Blind" Image Quality AnalyzerabstractAn important aim of research on the blind image quality assessment (IQA) problem is to devise perceptual models that can predict the quality of distorted images with as little prior knowledge of the images or their distortions as possible. Current state-of-the-art “general purpose” no reference (NR) IQA algorithms require knowledge about anticipated distortions in the form of training examples and corresponding human opinion scores. However we have recently derived a blind IQA model that only makes use of measurable deviations from statistical regularities observed in natural images, without training on human-rated distorted images, and, indeed without any exposure to distorted images. Thus, it is “completely blind.” The new IQA model, which we call the Natural Image Quality Evaluator (NIQE) is based on the construction of a “quality aware” collection of statistical features based on a simple and successful space domain natural scene statistic (NSS) model. These features are derived from a corpus of natural, undistorted images. Experimental results show that the new index delivers performance comparable to top performing NR IQA models that require training on large databases of human opinions of distorted images. A software release is available at http://live.ece.utexas.edu/research/quality/niqe_release.zip. Anish Mittal, Rajiv Soundararajan, Alan C. Bovik |
IEEE Signal Process. Lett. | 3 |
| 2013 | A Steerable, Multiscale Singularity IndexabstractWe propose a new steerable, multiscale ratio index for detecting impulse singularities in signals of arbitrary dimensionality. For example, it responds strongly to curvilinear masses (ridges) in images, but minimally to step discontinuities. The ratio index employs directional derivatives of gaussians, making it naturally steerable and scalable. Experiments on real images demonstrate the efficacy of the index for detecting multiscale curvilinear structures. A software version of the index can be downloaded from: http://live.ece.utexas.edu/research/SingularityIndex/SingularityIndex.zip. S. M. Gautam, Alan C. Bovik, Mia K. Markey |
IEEE Signal Process. Lett. | 2 |
| 2013 | A Markov Decision Model for Adaptive Scheduling of Stored Scalable VideosabstractWe propose two scheduling algorithms that seek to optimize the quality of scalably coded videos that have been stored at a video server before transmission. The first scheduling algorithm is derived from a Markov decision process (MDP) formulation developed here. We model the dynamics of the channel as a Markov chain and reduce the problem of dynamic video scheduling to a tractable Markov decision problem over a finite-state space. Based on the MDP formulation, a near-optimal scheduling policy is computed that minimizes the mean square error. Using insights taken from the development of the optimal MDP-based scheduling policy, the second proposed scheduling algorithm is an online scheduling method that only requires easily measurable knowledge of the channel dynamics, and is thus viable in practice. Simulation results show that the performance of both scheduling algorithms is close to a performance upper bound also derived in this paper. Chao Chen 0006, Robert W. Heath Jr., Alan C. Bovik, Gustavo de Veciana |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Video Quality Assessment by Reduced Reference Spatio-Temporal Entropic DifferencingabstractWe present a family of reduced reference video quality assessment (QA) models that utilize spatial and temporal entropic differences. We adopt a hybrid approach of combining statistical models and perceptual principles to design QA algorithms. A Gaussian scale mixture model for the wavelet coefficients of frames and frame differences is used to measure the amount of spatial and temporal information differences between the reference and distorted videos, respectively. The spatial and temporal information differences are combined to obtain the spatio-temporal-reduced reference entropic differences. The algorithms are flexible in terms of the amount of side information required from the reference that can range between a single scalar per frame and the entire reference information. The spatio-temporal entropic differences are shown to correlate quite well with human judgments of quality, as demonstrated by experiments on the LIVE video quality assessment database. Rajiv Soundararajan, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | No-Reference Quality Assessment of Natural StereopairsabstractWe develop a no-reference binocular image quality assessment model that operates on static stereoscopic images. The model deploys 2D and 3D features extracted from stereopairs to assess the perceptual quality they present when viewed stereoscopically. Both symmetric- and asymmetric-distorted stereopairs are handled by accounting for binocular rivalry using a classic linear rivalry model. The NSS features are used to train a support vector machine model to predict the quality of a tested stereopair. The model is tested on the LIVE 3D Image Quality Database, which includes both symmetric- and asymmetric-distorted stereoscopic 3D images. The experimental results show that our proposed model significantly outperforms the conventional 2D full-reference QA algorithms applied to stereopairs, as well as the 3D full-reference IQA algorithms on asymmetrically distorted stereopairs. Ming-Jun Chen, Lawrence K. Cormack, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2013 | Visually Weighted Compressive Sensing: Measurement and ReconstructionabstractCompressive sensing (CS) makes it possible to more naturally create compact representations of data with respect to a desired data rate. Through wavelet decomposition, smooth and piecewise smooth signals can be represented as sparse and compressible coefficients. These coefficients can then be effectively compressed via the CS. Since a wavelet transform divides image information into layered blockwise wavelet coefficients over spatial and frequency domains, visual improvement can be attained by an appropriate perceptually weighted CS scheme. We introduce such a method in this paper and compare it with the conventional CS. The resulting visual CS model is shown to deliver improved visual reconstructions. Hyungkeuk Lee, Heeseok Oh, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2013 | Video Quality Pooling Adaptive to Perceptual Distortion SeverityabstractIt is generally recognized that severe video distortions that are transient in space and/or time have a large effect on overall perceived video quality. In order to understand this phenomena, we study the distribution of spatio-temporally local quality scores obtained from several video quality assessment (VQA) algorithms on videos suffering from compression and lossy transmission over communication channels. We propose a content adaptive spatial and temporal pooling strategy based on the observed distribution. Our method adaptively emphasizes "worst" scores along both the spatial and temporal dimensions of a video sequence and also considers the perceptual effect of large-area cohesive motion flow such as egomotion. We demonstrate the efficacy of the method by testing it using three different VQA algorithms on the LIVE Video Quality database and the EPFL-PoliMI video quality database. Jincheol Park, Kalpana Seshadrinathan, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2013 | Color and Depth Priors in Natural ImagesabstractNatural scene statistics have played an increasingly important role in both our understanding of the function and evolution of the human vision system, and in the development of modern image processing applications. Because range (egocentric distance) is arguably the most important thing a visual system must compute (from an evolutionary perspective), the joint statistics between image information (color and luminance) and range information are of particular interest. It seems obvious that where there is a depth discontinuity, there must be a higher probability of a brightness or color discontinuity too. This is true, but the more interesting case is in the other direction--because image information is much more easily computed than range information, the key conditional probabilities are those of finding a range discontinuity given an image discontinuity. Here, the intuition is much weaker; the plethora of shadows and textures in the natural environment imply that many image discontinuities must exist without corresponding changes in range. In this paper, we extend previous work in two ways--we use as our starting point a very high quality data set of coregistered color and range values collected specifically for this purpose, and we evaluate the statistics of perceptually relevant chromatic information in addition to luminance, range, and binocular disparity information. The most fundamental finding is that the probabilities of finding range changes do in fact depend in a useful and systematic way on color and luminance changes; larger range changes are associated with larger image changes. Second, we are able to parametrically model the prior marginal and conditional distributions of luminance, color, range, and (computed) binocular disparity. Finally, we provide a proof of principle that this information is useful by showing that our distribution models improve the performance of a Bayesian stereo algorithm on an independent set of input images. To summarize, we show that there is useful information about range in very low-level luminance and color information. To a system sensitive to this statistical information, it amounts to an additional (and only recently appreciated) depth cue, and one that is trivial to compute from the image data. We are confident that this information is robust, in that we go to great effort and expense to collect very high quality raw data. Finally, we demonstrate the practical utility of these findings by using them to improve the performance of a Bayesian stereo algorithm. Che-Chun Su, Lawrence K. Cormack, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2012 | The multilinear compound Gaussian distributionabstractWe introduce a novel generalization of the compound Gaussian (CG) (or Gaussian Scale Mixture [1]) distribution which extends the Gaussian component of the CG model to a multilinear distribution. The resulting model, which we call the Multilinear Compound Gaussian (MCG) distribution, subsumes both GSM [1] and the previously developed MICA [3-4] distributions as complementary special cases; thereby allowing us to model a richer class of stochastic phenomena. First we derive the structural characterization of the MCG distribution and develop some of its important theoretical properties. Thereafter we describe a parameter estimation algorithm for learning this model from sample data, and then deploy this for modeling textures, including natural (i.e. optical) and SAR images. Our simulation results demonstrate how, for each case, we obtain improved performance over the CG model; thus indicating the versatility of the MCG model in accurately modeling various natural phenomena of interest. Raghu G. Raj, Alan C. Bovik |
ICASSP | 2 |
| 2012 | Optimizing 3D image display using the stereoacuity functionabstractWe develop an algorithm that predicts the best presentation of a stereo 3D image in the sense of viewers' preference. The algorithm operates in three steps. First, the 3D image is classified as either a “foreground dominant” or “background dominant” image. Next, for “foreground dominant” images, a model of the stereoacuity function is used to optimize the perceptual 3D resolution; for “background dominant” images, the nearest surface is placed in the 3D plane of the display screen. A human study was conducted to assess the algorithm and showed that the proposed model produced 3D images which had the best 3D quality scores among several candidate algorithms. Ming-Jun Chen, Do-Kyoung Kwon, Lawrence K. Cormack, Alan C. Bovik |
ICIP | 4 |
| 2012 | A new singularity indexabstractWe propose a new ratio index for the detection of impulse-like singularities in signals of arbitrary dimensionality. We show that the new singularity index responds strongly to singularities that are like impulses or smoothed impulses in cross section. For example, it responds strongly to curvilinear masses (ridges) in images, while responding minimally to edge-like singularities. The ratio index employs directional derivatives of gaussians, which makes the index naturally scalable. S. M. Gautam, Alan C. Bovik, Mia K. Markey |
ICIP | 2 |
| 2012 | Blind Image Quality Assessment Without Human Training Using Latent Quality FactorsabstractWe propose a highly unsupervised, training free, no reference image quality assessment (IQA) model that is based on the hypothesis that distorted images have certain latent characteristics that differ from those of “natural” or “pristine” images. These latent characteristics are uncovered by applying a “topic model” to visual words extracted from an assortment of pristine and distorted images. For the latent characteristics to be discriminatory between pristine and distorted images, the choice of the visual words is important. We extract quality-aware visual words that are based on natural scene statistic features [1]. We show that the similarity between the probability of occurrence of the different topics in an unseen image and the distribution of latent topics averaged over a large number of pristine natural images yields a quality measure. This measure correlates well with human difference mean opinion scores on the LIVE IQA database [2]. Anish Mittal, S. M. Gautam, Joydeep Ghosh, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2012 | Optimizing Multiscale SSIM for Compression via MLDSabstractA crucial step in the assessment of an image compression method is the evaluation of the perceived quality of the compressed images. Typically, researchers ask observers to rate perceived image quality directly and use these rating measures, averaged across observers and images, to assess how image quality degrades with increasing compression. These ratings in turn are used to calibrate and compare image quality assessment algorithms intended to predict human perception of image degradation. There are several drawbacks to using such omnibus measures. First, the interpretation of the rating scale is subjective and may differ from one observer to the next. Second, it is easy to overlook compression artifacts that are only present in particular kinds of images. In this paper, we use a recently developed method for assessing perceived image quality, maximum likelihood difference scaling (MLDS), and use it to assess the performance of a widely-used image quality assessment algorithm, multiscale structural similarity (MS-SSIM). MLDS allows us to quantify supra-threshold perceptual differences between pairs of images and to examine how perceived image quality, estimated through MLDS, changes as the compression rate is increased. We apply the method to a wide range of images and also analyze results for specific images. This approach circumvents the limitations inherent in the use of rating methods, and allows us also to evaluate MS-SSIM for different classes of visual image. We show how the data collected by MLDS allow us to recalibrate MS-SSIM to improve its performance. Christophe Charrier, Kenneth Knoblauch, Laurence T. Maloney, Alan C. Bovik, Anush K. Moorthy |
IEEE Trans. Image Process. | 4 |
| 2012 | No-Reference Image Quality Assessment in the Spatial DomainabstractWe propose a natural scene statistic-based distortion-generic blind/no-reference (NR) image quality assessment (IQA) model that operates in the spatial domain. The new model, dubbed blind/referenceless image spatial quality evaluator (BRISQUE) does not compute distortion-specific features, such as ringing, blur, or blocking, but instead uses scene statistics of locally normalized luminance coefficients to quantify possible losses of "naturalness" in the image due to the presence of distortions, thereby leading to a holistic measure of quality. The underlying features used derive from the empirical distribution of locally normalized luminances and products of locally normalized luminances under a spatial natural scene statistic model. No transformation to another coordinate frame (DCT, wavelet, etc.) is required, distinguishing it from prior NR IQA approaches. Despite its simplicity, we are able to show that BRISQUE is statistically better than the full-reference peak signal-to-noise ratio and the structural similarity index, and is highly competitive with respect to all present-day distortion-generic NR IQA algorithms. BRISQUE has very low computational complexity, making it well suited for real time applications. BRISQUE features may be used for distortion-identification as well. To illustrate a new practical application of BRISQUE, we describe how a nonblind image denoising algorithm can be augmented with BRISQUE in order to perform blind image denoising. Results show that BRISQUE augmentation leads to performance improvements over state-of-the-art methods. A software release of BRISQUE is available online: http://live.ece.utexas.edu/research/quality/BRISQUE_release.zip for public use and evaluation. Anish Mittal, Anush K. Moorthy, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2012 | Blind Image Quality Assessment: A Natural Scene Statistics Approach in the DCT DomainabstractWe develop an efficient, general-purpose, blind/noreference image quality assessment (NR-IQA) algorithm using a natural scene statistics (NSS) model of discrete cosine transform (DCT) coefficients. The algorithm is computationally appealing, given the availability of platforms optimized for DCT computation. The approach relies on a simple Bayesian inference model to predict image quality scores given certain extracted features. The features are based on an NSS model of the image DCT coefficients. The estimated parameters of the model are utilized to form features that are indicative of perceptual quality. These features are used in a simple Bayesian inference approach to predict quality scores. The resulting algorithm, which we name BLIINDS-II, requires minimal training and adopts a simple probabilistic model for score prediction. Given the extracted features from a test image, the quality score that maximizes the probability of the empirically determined inference model is chosen as the predicted quality score of that image. When tested on the LIVE IQA database, BLIINDS-II is shown to correlate highly with human judgments of quality, at a level that is competitive with the popular SSIM index. Michele A. Saad, Alan C. Bovik, Christophe Charrier |
IEEE Trans. Image Process. | 2 |
| 2012 | RRED Indices: Reduced Reference Entropic Differencing for Image Quality AssessmentabstractWe study the problem of automatic "reduced-reference" image quality assessment (QA) algorithms from the point of view of image information change. Such changes are measured between the reference- and natural-image approximations of the distorted image. Algorithms that measure differences between the entropies of wavelet coefficients of reference and distorted images, as perceived by humans, are designed. The algorithms differ in the data on which the entropy difference is calculated and on the amount of information from the reference that is required for quality computation, ranging from almost full information to almost no information from the reference. A special case of these is algorithms that require just a single number from the reference for QA. The algorithms are shown to correlate very well with subjective quality scores, as demonstrated on the Laboratory for Image and Video Engineering Image Quality Assessment Database and the Tampere Image Database. Performance degradation, as the amount of information is reduced, is also studied. Rajiv Soundararajan, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2011 | Temporal hysteresis model of time varying subjective video qualityabstractVideo quality assessment (QA) continues to be an important area of research due to the overwhelming number of applications where videos are delivered to humans. In particular, the problem of temporal pooling of quality sores has received relatively little attention. We observe a hysteresis effect in the subjective judgment of time-varying video quality based on measured behavior in a subjective study. Based on our analysis of the subjective data, we propose a hysteresis temporal pooling strategy for QA algorithms. Applying this temporal strategy to pool scores from PSNR, SSIM and MOVIE produces markedly improved subjective quality prediction. Kalpana Seshadrinathan, Alan C. Bovik |
ICASSP | 2 |
| 2011 | RRED indices: Reduced reference entropic differencing framework for image quality assessmentabstractWe study the problem of automatic "reduced reference" image quality assessment algorithms from the point of view of image information change. Algorithms that measure differences between the entropies of wavelet coefficients of reference and distorted images are designed. A family of algorithms are presented, each differing in the amount of data on which information change is predicted and ranging from al most full reference to almost no reference. A special case of this are algorithms that require just a single number from the reference for quality assessment. The algorithms are shown to correlate very well with subjective quality scores as demonstrated on the LIVE Image Quality Assessment Database. Rajiv Soundararajan, Alan C. Bovik |
ICASSP | 2 |
| 2011 | Calibrating MS-SSIM for compression distortions using MLDSabstractIn this paper, we describe a recently developed method for assessing perceived image quality, Maximum Likelihood Difference Scaling (MLDS), and use it to assess the performance of MS-SSIM on compression distored images. MLDS allows us to quantify supra-threshold perceptual differences between pairs of images and to examine how perceived image quality, estimated through MLDS, changes as the compression rate is increased. We show how the data collected by MLDS allows us to recalibrate MS-SSIM to improve its performance. Christophe Charrier, Kenneth Knoblauch, Laurence T. Maloney, Alan C. Bovik |
ICIP | 4 |
| 2011 | Adaptive policies for real-time video transmission: A Markov decision process frameworkabstractWe study the problem of adaptive video data scheduling over wireless channels. We prove that, under certain assumptions, adaptive video scheduling can be reduced to a Markov decision process over a finite state space. Therefore, the scheduling policy can be optimized via standard stochastic control techniques using a Markov decision formulation. Simulation results show that significant performance improvement can be achieved over heuristic transmission schemes. Chao Chen 0006, Robert W. Heath Jr., Alan C. Bovik, Gustavo de Veciana |
ICIP | 3 |
| 2011 | Optimal image transmission over Visual Sensor NetworksabstractIn this paper, we propose a methodology for optimal image transmission over a VSNs (Visual Sensor Networks) via cross-layer optimization. Toward this goal, we control the compression ratio of a captured image and network parameters such as source rate, flow rate and routing path. In particular, since this scheme is based on distributed optimization, we can avoid energy concentration in a specific node such as CH (Cluster Head) which increases the network lifetime. In the simulation, we demonstrate the network adaptation procedure over a randomly deployed VSNs and evaluate the quality of the transmitted image using the SSIM (Structural Similarity) index. Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 3 |
| 2011 | Spatio-temporal quality pooling accounting for transient severe impairments and egomotionabstractWith the increasing popularity of video applications, the reliable measurement of perceived video quality has increased in importance. We study methods for pooling video quality scores over space and time. The method accounts for localized severe impairments of the signal which exhibit significant influence on the subjective impression of the overall signal quality. It also accounts for the effect of camera motion (egomotion) on perceived quality. The method arrived at is tested on the LIVE Video Quality Database and is shown to perform quite well. Jincheol Park, Kalpana Seshadrinathan, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 4 |
| 2011 | DCT statistics model-based blind image quality assessmentabstractWe propose an efficient, general-purpose, distortion-agnostic, blind/no-reference image quality assessment (NR-IQA) algorithm based on a natural scene statistics model of discrete cosine transform (DCT) coefficients. The algorithm is computationally appealing, given the availability of platforms optimized for DCT computation. We propose a generalized parametric model of the extracted DCT coefficients. The parameters of the model are utilized to predict image quality scores. The resulting algorithm, which we name BLIINDS-II, requires minimal training and adopts a simple probabilistic model for score prediction. When tested on the LIVE IQA database, BLIINDS-II is shown to correlate highly with human visual perception of quality, at a level that is even competitive with the powerful full-reference SSIM index. Michele A. Saad, Alan C. Bovik, Christophe Charrier |
ICIP | 2 |
| 2011 | Natural scene statistics of color and rangeabstractColor and depth play important roles in natural scenes and in vision, and their perception is related. Extensive work has been conducted on studying the luminance statistics of natural scenes; however, there is very little work done on analyzing the statistics between luminance and range in natural scenes, not to mention color and range. In this paper, we present the LIVE Color+3D Database, which contains 12 sets of color images with corresponding ground truth range maps in a high-definition resolution of 1280×720. We examined the statistical distribution of range gradients conditioned on the Gabor responses of the color images, as well as the variations of statistical measures of range gradients with changes in the Gabor responses. The analysis results show that the distributions of range gradients conditioned on the Gabor responses have very similar exponential shapes for both luminance and chrominance channels. Moreover, we also found that the depth difference between neighboring pixels increases as the corresponding magnitudes of the Gabor responses rise. Che-Chun Su, Alan C. Bovik, Lawrence K. Cormack |
ICIP | 2 |
| 2011 | Visual quality assessment algorithms: what does the future hold?
Anush K. Moorthy, Alan C. Bovik |
Multim. Tools Appl. | 2 |
| 2011 | Automatic prediction of perceptual quality of multimedia signals - a survey
Kalpana Seshadrinathan, Alan C. Bovik |
Multim. Tools Appl. | 2 |
| 2011 | Evaluation of temporal variation of video quality in packet loss networks
Changhoon Yim, Alan C. Bovik |
Signal Process. Image Commun. | 2 |
| 2011 | Visual Conspicuity Index: Spatial Dissimilarity, Distance, and Central BiasabstractWe propose an image conspicuity index that combines three factors: spatial dissimilarity, spatial distance and central bias. The dissimilarity between image patches is evaluated in a reduced dimensional principal component space and is inversely weighted by the spatial separations between patches. An additional weighting mechanism is deployed that reflects the bias of human fixations towards the image center. The method is tested on three public image datasets and a video clip to evaluate its performance. The experimental results indicate highly competitive performance despite the simple definition of the proposed index. The conspicuity maps generated are more consistent with human fixations than prior state-of-the-art models when tested on color image datasets. This is demonstrated using both receiver operator characteristics (ROC) analysis and the Kullback-Leibler distance metric. The method should prove useful for such diverse image processing tasks as quality assessment, segmentation, search, or compression. The high performance and relative simplicity of the conspicuity index relative to other much more complex models suggests that it may find wide usage. Lijuan Duan, Chunpeng Wu, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2011 | Perceptually Scalable Extension of H.264abstractWe propose a novel visual scalable video coding (VSVC) framework, named VSVC H.264/AVC. In this approach, the non-uniform sampling characteristic of the human eye is used to modify scalable video coding (SVC) H.264/AVC. We exploit the visibility of video content and the scalability of the video codec to achieve optimal subjective visual quality given limited system resources. To achieve the largest coding gain with controlled perceptual quality degradation, a perceptual weighting scheme is deployed wherein the compressed video is weighted as a function of visual saliency and of the non-uniform distribution of retinal photoreceptors. We develop a resource allocation algorithm emphasizing both efficiency and fairness by controlling the size of the salient region in each quality layer. Efficiency is emphasized on the low quality layer of the SVC. The bits saved by eliminating perceptual redundancy in regions of low interest are allocated to lower block-level distortions in salient regions. Fairness is enforced on the higher quality layers by enlarging the size of the salient regions. The simulation results show that the proposed VSVC framework significantly improves the subjective visual quality of compressed videos. Hojin Ha, Jincheol Park, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Passive Multimodal 2-D+3-D Face Recognition Using Gabor Features and Landmark DistancesabstractWe introduce a novel multimodal framework for face recognition based on local attributes calculated from range and portrait image pairs. Gabor coefficients are computed at automatically detected landmark locations and combined with powerful anthropometric features defined in the form of geodesic and Euclidean distances between pairs of fiducial points. We make the pragmatic assumption that the 2-D and 3-D data is acquired passively (e.g., via stereo ranging) with perfect registration between the portrait data and the range data. Statistical learning approaches are evaluated independently to reduce the dimensionality of the 2-D and 3-D Gabor coefficients and the anthropometric distances. Three parallel face recognizers that result from applying the best performing statistical learning schemes are fused at the match score-level to construct a unified multimodal (2-D+3-D) face recognition system with boosted performance. Performance of the proposed algorithm is evaluated on a large public database of range and portrait image pairs and found to perform quite well. Sina Jahanbin, Hyohoon Choi, Alan C. Bovik |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2011 | Statistical Modeling of 3-D Natural Scenes With Application to Bayesian StereopsisabstractWe studied the empirical distributions of luminance, range and disparity wavelet coefficients using a coregistered database of luminance and range images. The marginal distributions of range and disparity are observed to have high peaks and heavy tails, similar to the well-known properties of luminance wavelet coefficients. However, we found that the kurtosis of range and disparity coefficients is significantly larger than that of luminance coefficients. We used generalized Gaussian models to fit the empirical marginal distributions. We found that the marginal distribution of luminance coefficients have a shape parameter p between 0.6 and 0.8, while range and disparity coefficients have much smaller parameters p < 0.32, corresponding to a much higher peak. We also examined the conditional distributions of luminance, range and disparity coefficients. The magnitudes of luminance and range (disparity) coefficients show a clear positive correlation, which means, at a location with larger luminance variation, there is a higher probability of a larger range (disparity) variation. We also used generalized Gaussians to model the conditional distributions of luminance and range (disparity) coefficients. The values of the two shape parameters (p,s) reflect the observed luminance-range (disparity) dependency. As an example of the usefulness of luminance statistics conditioned on range statistics, we modified a well-known Bayesian stereo ranging algorithm using our natural scene statistics models, which improved its performance. Yang Liu 0030, Lawrence K. Cormack, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2011 | Blind Image Quality Assessment: From Natural Scene Statistics to Perceptual QualityabstractOur approach to blind image quality assessment (IQA) is based on the hypothesis that natural scenes possess certain statistical properties which are altered in the presence of distortion, rendering them un-natural; and that by characterizing this un-naturalness using scene statistics, one can identify the distortion afflicting the image and perform no-reference (NR) IQA. Based on this theory, we propose an (NR)/blind algorithm-the Distortion Identification-based Image Verity and INtegrity Evaluation (DIIVINE) index-that assesses the quality of a distorted image without need for a reference image. DIIVINE is based on a 2-stage framework involving distortion identification followed by distortion-specific quality assessment. DIIVINE is capable of assessing the quality of a distorted image across multiple distortion categories, as against most NR IQA algorithms that are distortion-specific in nature. DIIVINE is based on natural scene statistics which govern the behavior of natural images. In this paper, we detail the principles underlying DIIVINE, the statistical features extracted and their relevance to perception and thoroughly evaluate the algorithm on the popular LIVE IQA database. Further, we compare the performance of DIIVINE against leading full-reference (FR) IQA algorithms and demonstrate that DIIVINE is statistically superior to the often used measure of peak signal-to-noise ratio (PSNR) and statistically equivalent to the popular structural similarity index (SSIM). A software release of DIIVINE has been made available online: "http://live.ece.utexas.edu/research/quality/DIIVINE_release.zip" xmlns:xlink="http://www.w3.org/1999/xlink">http://live.ece.utexas.edu/research/quality/DIIVINE_release.zip for public use and evaluation. Anush K. Moorthy, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2011 | Quality Assessment of Deblocked ImagesabstractWe study the efficiency of deblocking algorithms for improving visual signals degraded by blocking artifacts from compression. Rather than using only the perceptually questionable PSNR, we instead propose a block-sensitive index, named PSNR-B, that produces objective judgments that accord with observations. The PSNR-B modifies PSNR by including a blocking effect factor. We also use the perceptually significant SSIM index, which produces results largely in agreement with PSNR-B. Simulation results show that the PSNR-B results in better performance for quality assessment of deblocked images than PSNR and a well-known blockiness-specific index. Changhoon Yim, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2011 | Cross-Layer Optimization for Downlink Wavelet Video TransmissionabstractCross-layer optimization for efficient multimedia communications is an important emerging issue towards providing better quality-of-service (QoS) over capacity-limited wireless channels. This paper presents a cross-layer optimization approach that operates between the application and physical layers to achieve high fidelity downlink video transmission by optimizing with respect to a quality criterion termed “visual entropy” using Lagrangian relaxation. By utilizing the natural layered structure of wavelet coding, an optimal level of power allocation is determined, which permits the throughput of visual entropy to be maximized over a multi-cell environment. A theoretical approach to optimization using the Shannon capacity and the Karush-Kuhn-Tucker (KKT) conditions is explored when coupling the application with the physical layers. Simulations show that the throughput gain for cross-layer optimization by visual entropy is increased by nearly 80% at the cell boundary as compared with peak signal-to-noise ratio (PSNR). Hyungkeuk Lee, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Multim. | 3 |
| 2011 | Blind Image Quality Assessment Using a General Regression Neural NetworkabstractWe develop a no-reference image quality assessment (QA) algorithm that deploys a general regression neural network (GRNN). The new algorithm is trained on and successfully assesses image quality, relative to human subjectivity, across a range of distortion types. The features deployed for QA include the mean value of phase congruency image, the entropy of phase congruency image, the entropy of the distorted image, and the gradient of the distorted image. Image quality estimation is accomplished by approximating the functional relationship between these features and subjective mean opinion scores using a GRNN. Our experimental results show that the new method accords closely with human subjective judgment. Alan C. Bovik, Xiaojun Wu 0001 |
IEEE Trans. Neural Networks | 2 |
| 2010 | Natural scene statistics at stereo fixationsabstractWe conducted eye tracking experiments on naturalistic stereo images presented through a haploscope, and found that fixated luminance contrast and luminance gradient were generally higher than randomly selected luminance contrast and luminance gradient, which agrees with previous literatures. However we also found that the fixated disparity contrast and disparity gradient were generally lower than randomly selected disparity contrast and disparity gradient. We discuss the implications of this remarkable result. Yang Liu 0030, Lawrence K. Cormack, Alan C. Bovik |
ETRA | 3 |
| 2010 | Fast structural similarity index algorithmabstractThe development of real-time image quality assessment algorithms is an important direction on which little research has focused. This paper presents a design of real-time implementable full-reference image quality algorithms based on the SSIM index and multi-scale SSIM (MS-SSIM) index. The proposed algorithms, which modify SSIM/MS-SSIM to achieve speed, are tested on the LIVE image quality database and shown to yield performance commensurate with SSIM and MS-SSIM but with much lower computational complexity. Ming-Jun Chen, Alan C. Bovik |
ICASSP | 2 |
| 2010 | Statistics of natural image distortionsabstractNatural scene statistics (NSS) are an active area of research. Although there exist elegant models for NSS, the statistics of natural image distortions have received little attention. In this paper we study distorted image statistics (DIS) for natural scenes. We demonstrate that each distortion affects the statistics of natural images in a characteristic way and it is possible to parameterize this characteristic. We show that not only are DIS different for different distortions, but by such parametrization it is also possible to build a classifier that can classify a given image into a particular distortion category solely on the basis of DIS, with high accuracy. Applications of such categorization are of considerable scope and include DIS-based quality assessment and blind image distortion correction. Anush K. Moorthy, Alan C. Bovik |
ICASSP | 2 |
| 2010 | Temporal pooling of video quality estimates using perceptual motion modelsabstractEmerging multimedia applications have increased the need for video quality measurement. Motion is critical to this task, but is complicated owing to a variety of object movements and movement of the camera. Here, we categorize the various motion situations and deploy appropriate perceptual models to each category. We use these models to create a new approach to objective video quality assessment. Performance evaluation on the Laboratory for Image and Video Engineering (LIVE) Video Quality Database shows competitive performance compared to the leading contemporary VQA algorithms. Kwanghyun Lee, Jincheol Park, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 4 |
| 2010 | A two-stage framework for blind image quality assessmentabstractMost present day no-reference/blind image quality assessment (NR IQA) algorithms are distortion specific - i.e., they assume that the distortion affecting the image is known. Here we propose a novel two stage framework for distortion-independent blind image quality assessment based on natural scene statistics (NSS). The proposed framework is modular in that it can be extended beyond the distortion-pool considered here, and each module proposed can be replaced by better-performing ones in the future. We describe a 4-distortion demonstration of the proposed framework and show that it performs competitively with the full-reference peak-signal-to-noise-ratio on the LIVE IQA database. A software release of the proposed index has been made available online: http://live.ece.utexas.edu/research/quality/BIQI_4D_release.zip. Anush K. Moorthy, Alan C. Bovik |
ICIP | 2 |
| 2010 | Snakules: Snakes that seek spicules on mammographyabstractWe present a new method called “snakules” for the annotation of spicules on mammography. Snakules employs parametric open-ended snakes that are deployed in a region around a suspect spiculated mass location that has been identified by a radiologist or a computer-aided detection (CADe) algorithm. The set of convergent snakules deform, grow and adapt to the true spicules in the image, by an attractive process of curve evolution and motion that optimizes the local matching energy. Our results from an initial observer study involving an experienced radiologist demonstrate the strong potential of the method as an image analysis technique to improve the specificity of CADe algorithms. S. M. Gautam, Alan C. Bovik, Mia K. Markey |
ICIP | 2 |
| 2010 | A fast Multilinear ICA algorithmabstractWe extend our previous work on Multilinear Independent Component Analysis (MICA) by introducing a Fast-MICA algorithm that demonstrates the same improvement over classical ICA as the original MICA algorithm [1] while improving the computational speed by two polynomial orders of magnitude. Apart for enabling a faster determination of the multilinear structure of image patch probability density, this new approach opens up, for the first time, the possibility of computing a novel non-stationarity index based on the relative change in mutual information. We demonstrate the performance of our Fast-MICA algorithm together with an illustration of our novel non-stationarity index. Raghu G. Raj, Alan C. Bovik |
ICIP | 2 |
| 2010 | Natural DCT statistics approach to no-reference image quality assessmentabstractGeneral-purpose no-reference image quality assessment approaches still lag the advances in full-reference methods. Most no-reference methods are either distortion specific (i.e. they quantify one or more distortions such as blur, blockiness, or ringing), or they train a learning machine based on a large number of features. In this approach, we propose a discrete cosine transform (DCT) statistics-based support vector machine (SVM) approach based on only 3 features in the DCT domain. The approach extracts a very small number of features and is entirely in the DCT domain, making it computationally convenient. The results are shown to correlate highly with human visual perception of quality. Michele A. Saad, Alan C. Bovik, Christophe Charrier |
ICIP | 2 |
| 2010 | Anthropometric 3D Face Recognition
Shalini Gupta, Mia K. Markey, Alan C. Bovik |
Int. J. Comput. Vis. | 3 |
| 2010 | Perceptual Video Processing: Seeing the Future
Alan C. Bovik |
Proc. IEEE | 1 |
| 2010 | Content-partitioned structural similarity index for image quality assessment
Alan C. Bovik |
Signal Process. Image Commun. | 2 |
| 2010 | A Two-Step Framework for Constructing Blind Image Quality IndicesabstractPresent day no-reference/no-reference image quality assessment (NR IQA) algorithms usually assume that the distortion affecting the image is known. This is a limiting assumption for practical applications, since in a majority of cases the distortions in the image are unknown. We propose a new two-step framework for no-reference image quality assessment based on natural scene statistics (NSS). Once trained, the framework does not require any knowledge of the distorting process and the framework is modular in that it can be extended to any number of distortions. We describe the framework for blind image quality assessment and a version of this framework-the blind image quality index (BIQI) is evaluated on the LIVE image quality assessment database. A software release of BIQI has been made available online: http://live.ece.utexas.edu/research/quality/BIQI_release.zip. Anush K. Moorthy, Alan C. Bovik |
IEEE Signal Process. Lett. | 2 |
| 2010 | A DCT Statistics-Based Blind Image Quality IndexabstractAbstract—The development of general-purpose no-reference approaches to image quality assessment still lags recent advances in full-reference methods. Additionally, most no-reference or blind approaches are distortion-specific, meaning they assess only a specific type of distortion assumed present in the test image (such as blockiness, blur, or ringing). This limits their application domain. Other approaches rely on training a machine learning algorithm. These methods however, are only as effective as the features used to train their learning machines. Towards ameliorating this we introduce the BLIINDS index (BLind Image Integrity Notator using DCT Statistics) which is a no-reference approach to image quality assessment that does not assume a specific type of distortion of the image. It is based on predicting image quality based on observing the statistics of local discrete cosine transform coefficients, and it requires only minimal training. The method is shown to correlate highly with human perception of quality. Index Terms—Anisotropy, discrete cosine transform, kurtosis, natural scene statistics, no-reference quality assessment. I. Michele A. Saad, Alan C. Bovik, Christophe Charrier |
IEEE Signal Process. Lett. | 2 |
| 2010 | Perceptually Unequal Packet Loss Protection by Weighting Saliency and Error PropagationabstractWe describe a method for achieving perceptually minimal video distortion over packet-erasure networks using perceptually unequal loss protection (PULP). There are two main ingredients in the algorithm. First, a perceptual weighting scheme is employed wherein the compressed video is weighted as a function of the nonuniform distribution of retinal photoreceptors. Secondly, packets are assigned temporal importance within each group of pictures (GOP), recognizing that the severity of error propagation increases with elapsed time within a GOP. Using both frame-level perceptual importance and GOP-level hierarchical importance, the PULP algorithm seeks efficient forward error correction assignment that balances efficiency and fairness by controlling the size of identified salient region(s) relative to the channel state. PULP demonstrates robust performance and significantly improved subjective and objective visual quality in the face of burst packet losses. Hojin Ha, Jincheol Park, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Efficient Video Quality Assessment Along Temporal TrajectoriesabstractWe propose a new video quality assessment (VQA) algorithm - the motion compensated structural similarity index - that assesses not only spatial quality but also quality along temporal trajectories. Drawing inspiration from the motion-compensated approach followed for video compression, we propose a motion-compensated approach to temporal quality assessment. The proposed algorithm is computationally efficient as compared to other VQA algorithms that utilize motion information from extracted optical flow and correlates well with human perception of quality. In order to exemplify the utility of the algorithm in a practical setting, we evaluate the quality of H.264/AVC compressed videos. Efficiency of computation is enabled by the novel motion-vector re-use concept. Anush K. Moorthy, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Wireless Video Quality Assessment: A Study of Subjective Scores and Objective AlgorithmsabstractEvaluating the perceptual quality of video is of tremendous importance in the design and optimization of wireless video processing and transmission systems. In an endeavor to emulate human perception of quality, various objective video quality assessment (VQA) algorithms have been developed. However, the only subjective video quality database that exists on which these algorithms can be tested is dated and does not accurately reflect distortions introduced by present generation encoders and/or wireless channels. In order to evaluate the performance of VQA algorithms for the specific task of H.264 advanced video coding compressed video transmission over wireless networks, we conducted a subjective study involving 160 distorted videos. Various leading full reference VQA algorithms were tested for their correlation with human perception. The data from the paper has been made available to the research community, so that further research on new VQA algorithms and on the general area of VQA may be carried out. Anush K. Moorthy, Kalpana Seshadrinathan, Rajiv Soundararajan, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Efficient Stereoscopic Ranging via Stochastic Sampling of Match QualityabstractWe present an efficient method that computes dense stereo correspondences by stochastically sampling match quality values. Nonexhaustive sampling facilitates the use of quality metrics that take unique values at noninteger disparities. Depth estimates are iteratively refined with a stochastic cooperative search by perturbing the estimates, sampling match quality, and reweighting and aggregating the perturbations. The approach gains significant efficiencies when applied to video, where initial estimates are seeded using information from the previous pair in a novel application of the Z-buffering algorithm. This significantly reduces the number of search iterations required. We present a quantitative accuracy evaluation wherein the proposed method outperforms a microcanonical annealing approach by Barnard and a cooperative approach by Zitnick and Kanade , while using fewer match quality evaluations than either. The approach is shown to have more attractive memory usage and scaling than alternatives based on exhaustive sampling. Thayne Richard Coffman, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2010 | Unequal Power Allocation for JPEG Transmission Over MIMO SystemsabstractWith the introduction of multiple transmit and receive antennas in next generation wireless systems, real-time image and video communication are expected to become quite common, since very high data rates will become available along with improved data reliability. New joint transmission and coding schemes that explore advantages of multiple antenna systems matched with source statistics are expected to be developed. Based on this idea, we present an unequal power allocation scheme for transmission of JPEG compressed images over multiple-input multiple-output systems employing spatial multiplexing. The JPEG-compressed image is divided into different quality layers, and different layers are transmitted simultaneously from different transmit antennas using unequal transmit power, with a constraint on the total transmit power during any symbol period. Results show that our unequal power allocation scheme provides significant image quality improvement as compared to different equal power allocations schemes, with the peak-signal-to-noise-ratio gain as high as 14 dB at low signal-to-noise-ratios. Muhammad F. Sabir, Alan C. Bovik, Robert W. Heath Jr. |
IEEE Trans. Image Process. | 2 |
| 2010 | Motion Tuned Spatio-Temporal Quality Assessment of Natural VideosabstractThere has recently been a great deal of interest in the development of algorithms that objectively measure the integrity of video signals. Since video signals are being delivered to human end users in an increasingly wide array of applications and products, it is important that automatic methods of video quality assessment (VQA) be available that can assist in controlling the quality of video being delivered to this critical audience. Naturally, the quality of motion representation in videos plays an important role in the perception of video quality, yet existing VQA algorithms make little direct use of motion information, thus limiting their effectiveness. We seek to ameliorate this by developing a general, spatio-spectrally localized multiscale framework for evaluating dynamic video fidelity that integrates both spatial and temporal (and spatio-temporal) aspects of distortion assessment. Video quality is evaluated not only in space and time, but also in space-time, by evaluating motion quality along computed motion trajectories. Using this framework, we develop a full reference VQA algorithm for which we coin the term the MOtion-based Video Integrity Evaluation index, or MOVIE index. It is found that the MOVIE index delivers VQA scores that correlate quite closely with human subjective judgment, using the Video Quality Expert Group (VQEG) FRTV Phase 1 database as a test bed. Indeed, the MOVIE index is found to be quite competitive with, and even outperform, algorithms developed and submitted to the VQEG FRTV Phase 1 study, as well as more recent VQA algorithms tested on this database. Kalpana Seshadrinathan, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2010 | Study of Subjective and Objective Quality Assessment of VideoabstractWe present the results of a recent large-scale subjective study of video quality on a collection of videos distorted by a variety of application-relevant processes. Methods to assess the visual quality of digital videos as perceived by human observers are becoming increasingly important, due to the large number of applications that target humans as the end users of video. Owing to the many approaches to video quality assessment (VQA) that are being developed, there is a need for a diverse independent public database of distorted videos and subjective scores that is freely available. The resulting Laboratory for Image and Video Engineering (LIVE) Video Quality Database contains 150 distorted videos (obtained from ten uncompressed reference videos of natural scenes) that were created using four different commonly encountered distortion types. Each video was assessed by 38 human subjects, and the difference mean opinion scores (DMOS) were recorded. We also evaluated the performance of several state-of-the-art, publicly available full-reference VQA algorithms on the new database. A statistical evaluation of the relative performance of these algorithms is also presented. The database has a dedicated web presence that will be maintained as long as it remains relevant and the data is available online. Kalpana Seshadrinathan, Rajiv Soundararajan, Alan C. Bovik, Lawrence K. Cormack |
IEEE Trans. Image Process. | 3 |
| 2010 | Snakules: A Model-Based Active Contour Algorithm for the Annotation of Spicules on MammographyabstractWe have developed a novel, model-based active contour algorithm, termed "snakules", for the annotation of spicules on mammography. At each suspect spiculated mass location that has been identified by either a radiologist or a computer-aided detection (CADe) algorithm, we deploy snakules that are converging open-ended active contours also known as snakes. The set of convergent snakules have the ability to deform, grow and adapt to the true spicules in the image, by an attractive process of curve evolution and motion that optimizes the local matching energy. Starting from a natural set of automatically detected candidate points, snakules are deployed in the region around a suspect spiculated mass location. Statistics of prior physical measurements of spiculated masses on mammography are used in the process of detecting the set of candidate points. Observer studies with experienced radiologists to evaluate the performance of snakules demonstrate the potential of the algorithm as an image analysis technique to improve the specificity of CADe algorithms and as a CADe prompting tool. S. M. Gautam, Alan C. Bovik, J. David Giese, Mehul P. Sampat, Gary J. Whitman, Tamara Miner Haygood, Tanya W. Stephens, Mia K. Markey |
IEEE Trans. Medical Imaging | 2 |
| 2009 | Range image quality assessment by Structural SimilarityabstractWe propose a new quality metric for range images that is based on the multi-scale structural similarity (MS-SSIM) index. The new metric operates in a manner to SSIM but allows for special handling of missing data. We demonstrate its utility by reevaluating the set of stereo algorithms evaluated in the Middlebury stereo vision page http://vision.middlebury.edu/stereo/. The new algorithm which we term Range SSIM (R-SSIM) index possesses features that make it an attractive choice for assessing the quality of range images. William S. Malpica, Alan C. Bovik |
ICASSP | 2 |
| 2009 | Estimation and analysis of urban traffic flowabstractThis paper describes methods for extracting traffic flow information from urban traffic scenes. The ultimate goal is to collect a macroscopic view of traffic flow information in a fully automatic and segmentation-free way. First, traffic flow is calculated by optical flow estimation. Then, traffic flow regions are defined by the initial traffic flow, and further analysis is performed only in the defined traffic flow regions. Basic statistics of the traffic flow vectors are studied. It is shown that traffic flow computed by optical flow estimation effectively captures traffic scene activity. Also, the statistics of traffic flow vectors contain meaningful and interesting characteristics. An example application demonstrates the applicability and potential uses of the statistics. Joonsoo Lee, Alan C. Bovik |
ICIP | 2 |
| 2009 | Optimal power allocation for minimizing visual distortion over MIMO communication systemsabstractA recent dynamic increase in demand for wireless multimedia services has greatly accelerated the research on cross layer optimization techniques for transmitting multimedia data over wireless channel. In this paper, we explore a novel theoretical approach for joint optimization between the rate distortion (RD) of H.264/AVC video and the link-capacity of MIMO parallel subchannels. We obtain the optimal power level of subchannels through an optimization problem to minimize total visual distortion. In the simulation results, compared to the water filling (WF) method, the proposed scheme provides better results in aspects of visual quality in the face of sum rate loss. Jincheol Park, Uk Jang, Taegeun Oh, Sanghoon Lee 0001, Alan C. Bovik |
ICIP | 5 |
| 2009 | Active, Foveated, Uncalibrated Stereovision
James Monaco, Alan C. Bovik, Lawrence K. Cormack |
Int. J. Comput. Vis. | 2 |
| 2009 | Joint Source-Channel Distortion Modeling for MPEG-4 VideoabstractMultimedia communication has become one of the main applications in commercial wireless systems. Multimedia sources, mainly consisting of digital images and videos, have high bandwidth requirements. Since bandwidth is a valuable resource, it is important that its use should be optimized for image and video communication. Therefore, interest in developing new joint source-channel coding (JSCC) methods for image and video communication is increasing. Design of any JSCC scheme requires an estimate of the distortion at different source coding rates and under different channel conditions. The common approach to obtain this estimate is via simulations or operational rate-distortion curves. These approaches, however, are computationally intensive and, hence, not feasible for real-time coding and transmission applications. A more feasible approach to estimate distortion is to develop models that predict distortion at different source coding rates and under different channel conditions. Based on this idea, we present a distortion model for estimating the distortion due to quantization and channel errors in MPEG-4 compressed video streams at different source coding rates and channel bit error rates. This model takes into account important aspects of video compression such as transform coding, motion compensation, and variable length coding. Results show that our model estimates distortion within 1.5 dB of actual simulation values in terms of peak-signal-to-noise ratio. Muhammad F. Sabir, Robert W. Heath Jr., Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2009 | Complex Wavelet Structural Similarity: A New Image Similarity IndexabstractWe introduce a new measure of image similarity called the complex wavelet structural similarity (CW-SSIM) index and show its applicability as a general purpose image similarity index. The key idea behind CW-SSIM is that certain image distortions lead to consistent phase changes in the local wavelet coefficients, and that a consistent phase shift of the coefficients does not change the structural content of the image. By conducting four case studies, we have demonstrated the superiority of the CW-SSIM index against other indices (e.g., Dice, Hausdorff distance) commonly used for assessing the similarity of a given pair of images. In addition, we show that the CW-SSIM index has a number of advantages. It is robust to small rotations and translations. It provides useful comparisons even without a preprocessing image registration step, which is essential for other indices. Moreover, it is computationally less expensive. Mehul P. Sampat, Zhou Wang 0001, Shalini Gupta, Alan C. Bovik, Mia K. Markey |
IEEE Trans. Image Process. | 4 |
| 2009 | Indexes for Three-Class Classification Performance Assessment - An Empirical ComparisonabstractAssessment of classifier performance is critical for fair comparison of methods, including considering alternative models or parameters during system design. The assessment must not only provide meaningful data on the classifier efficacy, but it must do so in a concise and clear manner. For two-class classification problems, receiver operating characteristic analysis provides a clear and concise assessment methodology for reporting performance and comparing competing systems. However, many other important biomedical questions cannot be posed as "two-class" classification tasks and more than two classes are often necessary. While several methods have been proposed for assessing the performance of classifiers for such multiclass problems, none has been widely accepted. The purpose of this paper is to critically review methods that have been proposed for assessing multiclass classifiers. A number of these methods provide a classifier performance index called the volume under surface (VUS). Empirical comparisons are carried out using 4 three-class case studies, in which three popular classification techniques are evaluated with these methods. Since the same classifier was assessed using multiple performance indexes, it is possible to gain insight into the relative strengths and weakness of the measures. We conclude that: 1) the method proposed by Scurfield provides the most detailed description of classifier performance and insight about the sources of error in a given classification task and 2) the methods proposed by He and Nakas also have great practical utility as they provide both the VUS and an estimate of the variance of the VUS. These estimates can be used to statistically compare two classification algorithms. Mehul P. Sampat, Amit C. Patel, Yuhling Wang, Shalini Gupta, Chih-Wen Kan, Alan C. Bovik, Mia K. Markey |
IEEE Trans. Inf. Technol. Biomed. | 6 |
| 2009 | Color Compensation of Multicolor FISH ImagesabstractMulticolor fluorescence in situ hybridization (M-FISH) techniques provide color karyotyping that allows simultaneous analysis of numerical and structural abnormalities of whole human chromosomes. Chromosomes are stained combinatorially in M-FISH. By analyzing the intensity combinations of each pixel, all chromosome pixels in an image are classified. Due to the overlap of excitation and emission spectra and the broad sensitivity of image sensors, the obtained images contain crosstalk between the color channels. The crosstalk complicates both visual and automatic image analysis and may eventually affect the classification accuracy in M-FISH. The removal of crosstalk is possible by finding the color compensation matrix, which quantifies the color spillover between channels. However, there exists no simple method of finding the color compensation matrix from multichannel fluorescence images whose specimens are combinatorially hybridized. In this paper, we present a method of calculating the color compensation matrix for multichannel fluorescence images whose specimens are combinatorially stained. Hyohoon Choi, Kenneth R. Castleman, Alan C. Bovik |
IEEE Trans. Medical Imaging | 3 |
| 2009 | Optimal Channel Adaptation of Scalable Video Over a Multicarrier-Based Multicell EnvironmentabstractTo achieve seamless multimedia streaming services over wireless networks, it is important to overcome inter-cell interference (ICI), particularly in cell border regions. In this regard scalable video coding (SVC) has been actively studied due to its advantage of channel adaptation. We explore an optimal solution for maximizing the expected visual entropy over an orthogonal frequency division multiplexing (OFDM)-based broadband network from the perspective of cross-layer optimization. An optimization problem is parameterized by a set of source and channel parameters that are acquired along the user location over a multicell environment. A suboptimal solution is suggested using a greedy algorithm that allocates the radio resources to the scalable bitstreams as a function of their visual importance. The simulation results show that the greedy algorithm effectively resists ICI in the cell border region, while conventional nonscalable coding suffers severely because of ICI. Jincheol Park, Hyungkeuk Lee, Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Multim. | 4 |
| 2008 | Rate Bounds on SSIM Index of Quantized Image DCT CoefficientsabstractIn this paper, we derive bounds on the structural similarity (SSIM) index as a function of quantization rate for fixed-rate uniform quantization of image discrete cosine transform (DCT) coefficients under the high rate assumption. The space domain SSIM index is first expressed in terms of the DCT coefficients of the space domain vectors. The transform domain SSIM Index is then used to derive bounds on the average SSIM index as a function of quantization rate for Gaussian and Laplacian sources. As an illustrative example, uniform quantization of the DCT coefficients of natural images is considered. We show that the SSIM index between the reference and quantized images fall within the bounds for a large set of natural images. Further, we show using a simple example that the proposed bounds could be very useful for rate allocation problems in practical image and video coding applications. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr., Constantine Caramanis |
DCC | 2 |
| 2008 | SSIM-optimal linear image restorationabstractIn this paper, we present an algorithm for designing a linear equalizer that is optimal with respect to the structural similarity (SSIM) index. The optimization problem is shown to be non-convex, thereby making it non-trivial. The non-convex problem is first converted to a quasi-convex problem and then solved using a combination of first order necessary conditions and bisection search. To demonstrate the usefulness of this solution, it is applied to image denoising and image restoration examples. We show using these examples that optimizing equalizers for the SSIM index does indeed result in higher perceptual image quality compared to equalizers optimized for the ubiquitous mean squared error (MSE). Sumohana S. Channappayya, Alan C. Bovik, Constantine Caramanis, Robert W. Heath Jr. |
ICASSP | 2 |
| 2008 | Fast computation of dense stereo correspondence by stochastic sampling of match qualityabstractWe present a method for computing dense stereo correspondences in calibrated monocular video by iteratively and stochastically sampling match quality values in the disparity search space. Most existing methods exhaustively compute local correspondence quality before searching for a globally optimal solution. Instead, we iteratively refine a correspondence estimate by perturbing it with random noise and formulating an influence at each sample based on the perturbation and its effect on correspondence match quality. Local influence is aggregated to recover consistent trends in match quality caused by the piecewise-continuous structure of the scene. Correspondence estimates for a given frame pair are seeded with the estimates from the previous frame pair, allowing convergence to occur across multiple frame pairs. Thayne R. Coffman, Alan C. Bovik |
ICASSP | 2 |
| 2008 | Perceptual soft thresholding using the structural similarity indexabstractIn this paper, we present a novel algorithm for wavelet domain image denoising using the soft thresholding function. The thresholds are designed to be locally optimal with respect to the structural similarity (SSIM) index. The SSIM Index is first expressed in terms of wavelet transform coefficients of orthogonal wavelet transforms. The wavelet domain representation of the SSIM Index, along with the assumption of a Gaussian prior for the wavelet coefficients is used to formulate the soft thresholding optimization problem. A locally optimal solution is found using a quasi-Newton approach. This solution is applied to denoise images in the wavelet domain. The visual quality of the images denoised using the proposed algorithm is shown to be higher compared to the MSE-optimal soft thresholding denoising solution, as measured by the SSIM Index. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr. |
ICIP | 2 |
| 2008 | Automated facial feature detection and face recognition using Gabor features on range and portrait imagesabstractIn this paper, we present a novel identity verification system based on Gabor features extracted from range (3D) representations of faces. Multiple landmarks (fiducials) on a face are automatically detected using these Gabor features. Once the landmarks are identified, the Gabor features on all fiducials of a face are concatenated to form a feature vector for that particular face. Linear discriminant analysis (LDA) is used to reduce the dimensionality of the feature vector while maximizing the discrimination power. These novel features were tested on 1196 range images. The same features were also extracted from portrait images, and the accuracies of both modalities were compared. A superior verification accuracy was obtained using the range data, and a highly competitive accuracy to that of other techniques in the literature was also obtained for the portrait data. Sina Jahanbin, Hyohoon Choi, Rana Jahanbin, Alan C. Bovik |
ICIP | 4 |
| 2008 | Fixation selection by maximization of texure and contrast informationabstractWe present information-theoretic underpinnings of a computation theory of low-level visual fixations in natural images. In continuation of our prior work on optimal contrast-based fixations [1], we develop an optimum texture- based fixation selection algorithm based on a recent theory of non-stationarity measurement in natural images [2]. Thereafter we propose a simple coupling of the optimal texture-based and contrast-based fixation features to produce a new algorithm called CONTEXT, which exhibits robust performance for fixation selection in natural images. The performance of the fixation algorithms are evaluated for natural images by comparison to randomized fixation strategies via actual human fixations performed on the images. The fixation patterns obtained outperform randomized, GAFFE-based [3], and Itti [4] fixation strategies in terms of matching human fixation patterns. These results also demonstrate the important role that contrast and textural information play in low-level visual processes in the Human Visual System (HVS). Raghu G. Raj, Alan C. Bovik, Lawrence K. Cormack |
ICIP | 2 |
| 2008 | Unifying analysis of full reference image quality assessmentabstractThis paper studies two increasingly popular paradigms for image quality assessment - Structural SIMilarity (SSIM) metrics and Information Fidelity metrics. The relation of the SSIM metric to Mean Squared Error and Human Visual System (HVS) based models of quality assessment are studied. The SSIM model is shown to be equivalent to models of contrast gain control of the HVS. We study the information theoretic metrics and show that the Information Fidelity Criterion (IFC) is a monotonic function of the structure term of the SSIM index applied in the sub-band filtered domain. Our analysis of the Visual Information Fidelity (VIF) criterion shows that improvements in VIF include incorporation of a contrast comparison, in addition to the structure comparison in IFC. Our analysis attempts to unify quality metrics derived from different first principles and characterize the relative performance of different QA systems. Kalpana Seshadrinathan, Alan C. Bovik |
ICIP | 2 |
| 2008 | Making Video Quality Assessment Models Robust to Bit DepthabstractWe introduce a novel feature set, which we call HDRMAX features, that when included into Video Quality Assessment (VQA) algorithms designed for Standard Dynamic Range (SDR) videos, sensitizes them to distortions of High Dynamic Range (HDR) videos that are inadequately accounted for by these algorithms. While these features are not specific to HDR, and also augment the equality prediction performances of VQA models on SDR content, they are especially effective on HDR. HDRMAX features modify powerful priors drawn from Natural Video Statistics (NVS) models by enhancing their measurability where they visually impact the brightest and darkest local portions of videos, thereby capturing distortions that are often poorly accounted for by existing VQA models. As a demonstration of the efficacy of our approach, we show that, while current state-of-the-art VQA models perform poorly on 10-bit HDR databases, their performances are greatly improved by the inclusion of HDRMAX features when tested on HDR and 10-bit distorted videos. Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Sriram Sethuraman, Alan C. Bovik |
IEEE Signal Process. Lett. | 6 |
| 2008 | Design of Linear Equalizers Optimized for the Structural Similarity IndexabstractWe propose an algorithm for designing linear equalizers that maximize the structural similarity (SSIM) index between the reference and restored signals. The SSIM index has enjoyed considerable application in the evaluation of image processing algorithms. Algorithms, however, have not been designed yet to explicitly optimize for this measure. The design of such an algorithm is nontrivial due to the nonconvex nature of the distortion measure. In this paper, we reformulate the nonconvex problem as a quasi-convex optimization problem, which admits a tractable solution. We compute the optimal solution in near closed form, with complexity of the resulting algorithm comparable to complexity of the linear minimum mean squared error (MMSE) solution, independent of the number of filter taps. To demonstrate the usefulness of the proposed algorithm, it is applied to restore images that have been blurred and corrupted with additive white gaussian noise. As a special case, we consider blur-free image denoising. In each case, its performance is compared to a locally adaptive linear MSE-optimal filter. We show that the images denoised and restored using the SSIM-optimal filter have higher SSIM index, and superior perceptual quality than those restored using the MSE-optimal adaptive linear filter. Through these results, we demonstrate that a) designing image processing algorithms, and, in particular, denoising and restoration-type algorithms, can yield significant gains over existing (in particular, linear MMSE-based) algorithms by optimizing them for perceptual distortion measures, and b) these gains may be obtained without significant increase in the computational complexity of the algorithm. Sumohana S. Channappayya, Alan C. Bovik, Constantine Caramanis, Robert W. Heath Jr. |
IEEE Trans. Image Process. | 2 |
| 2008 | Rate Bounds on SSIM Index of Quantized ImagesabstractIn this paper, we derive bounds on the structural similarity (SSIM) index as a function of quantization rate for fixed-rate uniform quantization of image discrete cosine transform (DCT) coefficients under the high-rate assumption. The space domain SSIM index is first expressed in terms of the DCT coefficients of the space domain vectors. The transform domain SSIM index is then used to derive bounds on the average SSIM index as a function of quantization rate for uniform, Gaussian, and Laplacian sources. As an illustrative example, uniform quantization of the DCT coefficients of natural images is considered. We show that the SSIM index between the reference and quantized images fall within the bounds for a large set of natural images. Further, we show using a simple example that the proposed bounds could be very useful for rate allocation problems in practical image and video coding applications. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr. |
IEEE Trans. Image Process. | 2 |
| 2008 | Nonlinearities in Stereoscopic Phase-DifferencingabstractExploiting the quasi-linear relationship between local phase and disparity, phase-differencing registration algorithms provide a fast, powerful means for disparity estimation. Unfortunately, these phase-differencing techniques suffer a significant impediment: phase nonlinearities. In regions of phase nonlinearity, the signals under consideration possess properties that invalidate the use of phase for disparity estimation. This paper uses the amenable properties of Gaussian white noise images to analytically quantify these properties. The improved understanding gained from this analysis enables us to better understand current methodologies for detecting regions of phase instability. Most importantly, we introduce a new, more effective means for identifying these regions based on the second derivative of phase. James Monaco, Alan C. Bovik, Lawrence K. Cormack |
IEEE Trans. Image Process. | 2 |
| 2008 | MICA: A Multilinear ICA Decomposition for Natural Scene ModelingabstractWe refine the classical independent component analysis (ICA) decomposition using a multilinear expansion of the probability density function of the source statistics. In particular, we introduce a specific nonlinear system that allows us to elegantly capture the statistical dependences between the responses of the multilinear ICA (MICA) filters. The resulting multilinear probability density is analytically tractable and does not require Monte Carlo simulations to estimate the model parameters. We demonstrate the MICA model on natural image textures and envision that the new model will prove useful for analyzing nonstationarity natural images using natural scene statistics models. Raghu G. Raj, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2008 | GAFFE: A Gaze-Attentive Fixation Finding EngineabstractThe ability to automatically detect visually interesting regions in images has many practical applications, especially in the design of active machine vision and automatic visual surveillance systems. Analysis of the statistics of image features at observers' gaze can provide insights into the mechanisms of fixation selection in humans. Using a foveated analysis framework, we studied the statistics of four low-level local image features: luminance, contrast, and bandpass outputs of both luminance and contrast, and discovered that image patches around human fixations had, on average, higher values of each of these features than image patches selected at random. Contrast-bandpass showed the greatest difference between human and random fixations, followed by luminance-bandpass, RMS contrast, and luminance. Using these measurements, we present a new algorithm that selects image regions as likely candidates for fixation. These regions are shown to correlate well with fixations recorded from human observers. Umesh Rajashekar, Ian van der Linde, Alan C. Bovik, Lawrence K. Cormack |
IEEE Trans. Image Process. | 3 |
| 2008 | Feature Normalization via Expectation Maximization and Unsupervised Nonparametric Classification For M-FISH Chromosome ImagesabstractMulticolor fluorescence in situ hybridization (M-FISH) techniques provide color karyotyping that allows simultaneous analysis of numerical and structural abnormalities of whole human chromosomes. Chromosomes are stained combinatorially in M-FISH. By analyzing the intensity combinations of each pixel, all chromosome pixels in an image are classified. Often, the intensity distributions between different images are found to be considerably different and the difference becomes the source of misclassifications of the pixels. Improved pixel classification accuracy is the most important task to ensure the success of the M-FISH technique. In this paper, we introduce a new feature normalization method for M-FISH images that reduces the difference in the feature distributions among different images using the expectation maximization (EM) algorithm. We also introduce a new unsupervised, nonparametric classification method for M-FISH images. The performance of the classifier is as accurate as the maximum-likelihood classifier, whose accuracy also significantly improved after the EM normalization. We would expect that any classifier will likely produce an improved classification accuracy following the EM normalization. Since the developed classification method does not require training data, it is highly convenient when ground truth does not exist. A significant improvement was achieved on the pixel classification accuracy after the new feature normalization. Indeed, the overall pixel classification accuracy improved by 20% after EM normalization. Hyohoon Choi, Alan C. Bovik, Kenneth R. Castleman |
IEEE Trans. Medical Imaging | 2 |
| 2007 | 3D Face Recognition Founded on the Structural Diversity of Human FacesabstractWe present a systematic procedure for selecting facial fiducial points associated with diverse structural characteristics of a human face. We identify such characteristics from the existing literature on anthropometric facial proportions. We also present three dimensional (3D) face recognition algorithms, which employ Euclidean/geodesic distances between these anthropometric fiducial points as features along with linear discriminant analysis classifiers. Furthermore, we show that in our algorithms, when anthropometric distances are replaced by distances between arbitrary regularly spaced facial points, their performances decrease substantially. This demonstrates that incorporating domain specific knowledge about the structural diversity of human faces significantly improves the performance of 3D human face recognition algorithms. Shalini Gupta, Jake K. Aggarwal, Mia K. Markey, Alan C. Bovik |
CVPR | 4 |
| 2007 | The Multilinear ICA Decompositionwith Applications to NSS ModelingabstractWe refine the classical independent component analysis (ICA) decomposition using a multilinear expansion of the probability density function of the source statistics. In particular, to model the source statistics of natural image textures, we introduce a specific non-linear system that allows us to elegantly capture the statistical dependences between the responses of the multilinear ICA (MICA) filters. The resulting multilinear probability density is analytically tractable and does not require Monte Carlo simulations to estimate the model parameters. We demonstrate the success of the MICA model on natural textures and discuss applications to non-stationarity detection and natural scene statistics (NSS) modeling. Raghu G. Raj, Alan C. Bovik |
ICASSP (2) | 2 |
| 2007 | A Structural Similarity Metric for Video Based on Motion ModelsabstractQuality assessment plays a very important role in almost all aspects of multimedia signal processing such as acquisition, coding, display, processing etc. Several objective quality metrics have been proposed for images, but video quality assessment has received relatively little attention and most video quality metrics have been simple extension of metrics for images. In this paper, we propose a novel quality metric for video sequences that utilizes motion information in video sequences, which is the main difference in moving from images to video. This metric is capable of capturing temporal artifacts in video sequences in addition to spatial distortions. Results are presented that demonstrate the efficacy of our quality metric by comparing model performance against subjective scores on the database developed by the video quality experts group. Kalpana Seshadrinathan, Alan C. Bovik |
ICASSP (1) | 2 |
| 2007 | Three Dimensional Face Recognition using Wavelet Decomposition of Range ImagesabstractInterest in face recognition systems has increased significantly due to the emergence of significant commercial opportunities in surveillance and security applications. In this paper we propose a novel technique to extract features from 3D face representations. In this technique, first the nose tip is automatically located on the range image, then the range data from a hexagonal region of interest around this landmark is decomposed using Barycentric wavelet kernels. The dimensionality of the extracted coefficients at each resolution level is reduced using principal component analysis (PCA). These new features are tested on 206 range images, and a high classification accuracy is achieved using a small number of features. The obtained accuracy is competitive to that of other techniques in literature. Sina Jahanbin, Hyohoon Choi, Alan C. Bovik, Kenneth R. Castleman |
ICIP (1) | 3 |
| 2007 | Epipolar Spaces and Optimal Sampling StrategiesabstractIf precise calibration information is unavailable, as is often the case for active binocular vision systems, the determination of epipolar lines becomes untenable. Yet, even without instantaneous knowledge of the geometry, the search for corresponding points can be restricted to areas called epipolar spaces. For each point in one image, we define the corresponding epipolar space in the other image as the union of all associated epipolar lines over all possible system geometries. Epipolar spaces eliminate the need for calibration at the cost of an increased search region. One approach to mitigate this increase is the application of a space variant sampling or foveation strategy. While the application of such strategies to stereo vision tasks is not new, only rarely has a foveation scheme been specifically tailored for a stereo vision task. In this paper we derive a foundation of theorems that provide a means for obtaining optimal sampling schemes for a given set of epipolar spaces. An optimal sampling scheme is defined as a strategy that minimizes the average area per epipolar space. James Monaco, Alan C. Bovik, Lawrence K. Cormack |
ICIP (6) | 2 |
| 2007 | Epipolar Spaces for Active Binocular Vision SystemsabstractDepth recovery for active binocular vision systems is simplified if the camera geometry is known and corresponding points can be restricted to epipolar lines. Unfortunately, computation of epipolar lines requires calibration which can be complex and inaccurate. While it is possible to register images without geometric information, such unconstrained algorithms are usually time consuming and prone to error. In this paper we propose a compromise. Even without the instantaneous knowledge of the system geometry, we can restrict the region of correspondence by imposing limits on the possible range of configurations, and as a result, confine our search for matching points to epipolar spaces. For each point in one image, we define the corresponding epipolar space in the other image as the union of all associated epipolar lines over all possible system geometries. Epipolar spaces eliminate the need for calibration at the cost of an increased search region. James Monaco, Alan C. Bovik, Lawrence K. Cormack |
ICIP (6) | 2 |
| 2007 | Non-Stationarity Detection in Natural ImagesabstractWe present a novel approach for non-stationarity detection in natural images by exploiting the prior knowledge of the independent component structure of scene statistics. Our proposed non-stationarity index is conceptually simple and is intertwined with the probabilistic structure of the image segment being analyzed. It shows consistently good results when applied to natural scenes and, we expect, will find useful applications in computer vision algorithms in as much as the detection of statistically non-stationary locations in images can be an important preliminary step toward the understanding of scene content and in the guiding of visual fixations. Raghu G. Raj, Alan C. Bovik, Wilson S. Geisler |
ICIP (3) | 2 |
| 2007 | New Directions in Image and Video Quality Assessment Plenary Talk
Alan C. Bovik |
MMSP | 1 |
| 2007 | Facial Range Image Matching Using the ComplexWavelet Structural Similarity MetricabstractWe propose a novel 3D face recognition algorithm based on facial range image matching using the complex wavelet structural similarity metric (CW-SSIM) metric. Compared with many existing 3D surface matching methods, CW-SSIM is computationally efficient and is robust to small geometrical distortions. Using a data set that contains 360 3D face models of 12 subjects, we tested the performance of the proposed method and compared it with existing 3D surface matching based face recognition algorithms. Verification and identification performance of each algorithm was evaluated by means of the receiver operating characteristic curve and the cumulative match characteristic curve. Among the algorithms tested, the proposed algorithm based on the CW-SSIM resulted in the best overall performance with an equal error rate of 9.13% and a rank 1 recognition rate of 98.6%, significantly better than all the other algorithms. Besides the introduction of a novel approach for 3D face recognition, this is also the first attempt to expand the application scope of complex wavelet domain similarity measure to range image matching in general Shalini Gupta, Mehul P. Sampat, Mia K. Markey, Alan C. Bovik, Zhou Wang 0001 |
WACV | 4 |
| 2007 | Analyzing Image Structure by Multidimensional Frequency ModulationabstractWe develop a mathematical framework for quantifying and understanding multidimensional frequency modulations in digital images. We begin with the widely accepted definition of the instantaneous frequency vector (IF) as the gradient of the phase and define the instantaneous frequency gradient tensor (IFGT) as the tensor of component derivatives of the IF vector. Frequency modulation bounds are derived and interpreted in terms of the eigendecomposition of the IFGT. Using the IFGT, we derive the ordinary differential equations (ODEs) that describe image flowlines. We study the diagonalization of the ODEs of multidimensional frequency modulation on the IFGT eigenvector coordinate system and suggest that separable transforms can be computed along these coordinates. We illustrate these new methods of image pattern analysis on textured and fingerprint images. We envision that this work will find value in applications involving the analysis of image textures that are nonstationary yet exhibit local regularity. Examples of such textures abound in nature. Marios S. Pattichis, Alan C. Bovik |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Foveated Visual Search for CornersabstractWe cast the problem of corner detection as a corner search process. We develop principles of foveated visual search and automated fixation selection to accomplish the corner search, supplying a case study of both foveated search and foveated feature detection. The result is a new algorithm for finding corners, which is also a corner-based algorithm for aiming computed foveated visual fixations. In the algorithm, long saccades move the fovea to previously unexplored areas of the image, while short saccades improve the accuracy of putative corner locations. The system is tested on two natural scenes. As an interesting comparison study, we compare fixations generated by the algorithm with those of subjects viewing the same images, whose eye movements are being recorded by an eye tracker. The comparison of fixation patterns is made using an information-theoretic measure. Results show that the algorithm is a good locater of corners, but does not correlate particularly well with human visual fixations. Thomas L. Arnow, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2006 | Joint Source-Channel Distortion Modelling for MPEG-4 VideoabstractJoint source-channel coding is becoming more important for wireless multimedia transmission due to high bandwidth requirements of these multimedia sources. Design of all joint source-channel coding schemes require an estimate of distortion at different source coding rates and under different channel conditions. In this paper, we present one such distortion model for estimating distortion due to quantization and channel errors in a joint manner for MPEG-4 compressed video streams. This model takes into account important aspects of video compression such as transform coding, motion compensation, and variable length coding. Results show that our model estimates distortion within 1 dB of actual simulation values in terms of peak-signal-to-noise-ratio Muhammad F. Sabir, Robert W. Heath Jr., Alan C. Bovik |
ICASSP (4) | 3 |
| 2006 | Toroidal Gaussian Filters for Detection and Extraction of Properties of Spiculated MassesabstractWe have invented a new class of linear filters for the detection of spiculated masses and architectural distortions in mammography. We call these Spiculation Filters. These filters are narrow band filters and form a new class of wavelet-type filter banks. In this paper, we show that unmodulated versions of these filters can be used to detect the central mass region of spiculated masses. We refer to these as toroidal gaussian filters. We also show that the physical properties of spiculated masses can be extracted from the responses of the toroidal gaussian filters without segmentation. Mehul P. Sampat, Alan C. Bovik, Mia K. Markey, Gary J. Whitman, Tanya W. Stephens |
ICASSP (2) | 2 |
| 2006 | A Linear Estimator Optimized for the Structural Similarity Index and its Application to Image DenoisingabstractWe use a perceptual distortion metric-the structural similarity (SSIM) index, to derive a new linear estimator for estimating zero-mean Gaussian sources distorted by additive white Gaussian noise (AWGN). We use this estimator in an image denoising application and compare its performance with the traditional linear least squared error (LLSE) estimator. Although images denoised using the SSIM-optimized estimator have a lower peak signal-to-noise ratio (PSNR) compared to their LLSE counterparts, the SSIM-optimized estimator clearly outperforms the LLSE estimator in terms of the visual quality of the denoised images. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr. |
ICIP | 2 |
| 2006 | Segmentation and Fuzzy-Logic Classification of M-FISH Chromosome ImagesabstractMulticolor fluorescence in-situ hybridization (M-FISH) technique provides color karyotyping that allows simultaneous analysis of numerical and structural abnormalities of whole human chromosomes. Currently available M-FISH systems exhibit misclassifications of multiple pixel regions that are often larger than the actual chromosomal rearrangement. This paper presents a novel unsupervised classification method based on fuzzy logic classification and a prior adjusted reclassification method. Utilizing the chromosome boundaries, the initial classification results improved significantly after the prior adjusted reclassification while keeping the translocations intact. This paper also presents a new segmentation method that combines both spectral and edge information. Ten M-FISH images from a publicly available database were used to test our methods. The segmentation accuracy was more than 98% on average. Hyohoon Choi, Kenneth R. Castleman, Alan C. Bovik |
ICIP | 3 |
| 2006 | Foveated Analysis and Selection of Visual Fixations in Natural ScenesabstractThe ability to automatically detect visually interesting regions in images has practical applications in the design of active machine vision systems. Analysis of the statistics of image features at observers gaze can provide insights into the mechanisms of fixation selection in humans. Using a novel foveated analysis framework, in which features were analyzed at the spatial resolution at which they were perceived, we studied the statistics of four low-level local image features: luminance, contrast, center-surround outputs of luminance and contrast, and discovered that the image patches around human fixations had, on average, higher values of each of these features than the image patches selected at random. Center-surround contrast showed the greatest difference between human and random fixations, followed by contrast, center-surround luminance, and luminance. Using these measurements, we present a new algorithm that selects image regions as likely candidates for fixation. These regions are shown to correlate well with fixations recorded from observers. Umesh Rajashekar, Ian van der Linde, Alan C. Bovik, Lawrence K. Cormack |
ICIP | 3 |
| 2006 | Measuring Intra- and Inter-Observer Agreement in Identifying and Localizing Structures in Medical ImagesabstractInter- and intra-observer variability exists in any measurements made on medical images. There are two sources of variability. The first occurs when the observers identify and localize the object of interest, and the second happens when the observers make appropriate measurement on the object of interest. A number of statistical methods are available to quantify the degree of agreement between measurements made by different observers. However, little has been done to develop metrics for quantifying the variability in identifying and localizing the objects of interest prior to measurement. In this paper, we propose to use the complex wavelet structural similarity index (CW-SSIM) method to measure the variability in identifying and localizing structures on images. Performance comparisons using simulated images as well as real mammography images demonstrate the effectiveness and robustness of the CW-SSIM method. Mehul P. Sampat, Zhou Wang 0001, Mia K. Markey, Gary J. Whitman, Tanya W. Stephens, Alan C. Bovik |
ICIP | 6 |
| 2006 | A joint source-channel distortion model for JPEG compressed imagesabstractThe need for efficient joint source-channel coding (JSCC) is growing as new multimedia services are introduced in commercial wireless communication systems. An important component of practical JSCC schemes is a distortion model that can predict the quality of compressed digital multimedia such as images and videos. The usual approach in the JSCC literature for quantifying the distortion due to quantization and channel errors is to estimate it for each image using the statistics of the image for a given signal-to-noise ratio (SNR). This is not an efficient approach in the design of real-time systems because of the computational complexity. A more useful and practical approach would be to design JSCC techniques that minimize average distortion for a large set of images based on some distortion model rather than carrying out per-image optimizations. However, models for estimating average distortion due to quantization and channel bit errors in a combined fashion for a large set of images are not available for practical image or video coding standards employing entropy coding and differential coding. This paper presents a statistical model for estimating the distortion introduced in progressive JPEG compressed images due to quantization and channel bit errors in a joint manner. Statistical modeling of important compression techniques such as Huffman coding, differential pulse-coding modulation, and run-length coding are included in the model. Examples show that the distortion in terms of peak signal-to-noise ratio (PSNR) can be predicted within a 2-dB maximum error over a variety of compression ratios and bit-error rates. To illustrate the utility of the proposed model, we present an unequal power allocation scheme as a simple application of our model. Results show that it gives a PSNR gain of around 6.5 dB at low SNRs, as compared to equal power allocation. Muhammad F. Sabir, Hamid R. Sheikh, Robert W. Heath Jr., Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2006 | Image information and visual qualityabstractMeasurement of visual quality is of fundamental importance to numerous image and video processing applications. The goal of quality assessment (QA) research is to design algorithms that can automatically assess the quality of images or videos in a perceptually consistent manner. Image QA algorithms generally interpret image quality as fidelity or similarity with a "reference" or "perfect" image in some perceptual space. Such "full-reference" QA methods attempt to achieve consistency in quality prediction by modeling salient physiological and psychovisual features of the human visual system (HVS), or by signal fidelity measures. In this paper, we approach the image QA problem as an information fidelity problem. Specifically, we propose to quantify the loss of image information to the distortion process and explore the relationship between image information and visual quality. QA systems are invariably involved with judging the visual quality of "natural" images and videos that are meant for "human consumption." Researchers have developed sophisticated models to capture the statistics of such natural signals. Using these models, we previously presented an information fidelity criterion for image QA that related image quality with the amount of information shared between a reference and a distorted image. In this paper, we propose an image information measure that quantifies the information that is present in the reference image and how much of this reference information can be extracted from the distorted image. Combining these two quantities, we propose a visual information fidelity measure for image QA. We validate the performance of our algorithm with an extensive subjective study involving 779 images and show that our method outperforms recent state-of-the-art image QA algorithms by a sizeable margin in our simulations. The code and the data from the subjective study are available at the LIVE website. Hamid R. Sheikh, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2006 | A Statistical Evaluation of Recent Full Reference Image Quality Assessment AlgorithmsabstractMeasurement of visual quality is of fundamental importance for numerous image and video processing applications, where the goal of quality assessment (QA) algorithms is to automatically assess the quality of images or videos in agreement with human quality judgments. Over the years, many researchers have taken different approaches to the problem and have contributed significant research in this area and claim to have made progress in their respective domains. It is important to evaluate the performance of these algorithms in a comparative setting and analyze the strengths and weaknesses of these methods. In this paper, we present results of an extensive subjective quality assessment study in which a total of 779 distorted images were evaluated by about two dozen human subjects. The "ground truth" image quality data obtained from about 25,000 individual human quality judgments is used to evaluate the performance of several prominent full-reference image quality assessment algorithms. To the best of our knowledge, apart from video quality studies conducted by the Video Quality Experts Group, the study presented in this paper is the largest subjective image quality study in the literature in terms of number of images, distortion types, and number of human judgments per image. Moreover, we have made the data from the study freely available to the research community. This would allow other researchers to easily report comparative results in the future. Hamid R. Sheikh, Muhammad F. Sabir, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2006 | Quality-aware imagesabstractWe propose the concept of quality-aware image, in which certain extracted features of the original (high-quality) image are embedded into the image data as invisible hidden messages. When a distorted version of such an image is received, users can decode the hidden messages and use them to provide an objective measure of the quality of the distorted image. To demonstrate the idea, we build a practical quality-aware image encoding, decoding and quality analysis system, which employs: 1) a novel reduced-reference image quality assessment algorithm based on a statistical model of natural images and 2) a previously developed quantization watermarking-based data hiding technique in the wavelet transform domain. Zhou Wang 0001, Guixing Wu, Hamid R. Sheikh, Eero P. Simoncelli, En-Hui Yang, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2005 | Multiple Description Image Coding Using Natural Scene StatisticsabstractThe statistics of natural scenes in the wavelet domain are accurately characterized by the Gaussian scale mixture (GSM) model. The model lends itself easily to analysis and many applications that use this model are emerging (e.g., denoising, watermark detection). We present an error-resilient image communications application that uses the GSM model and multiple description coding (MDC) to provide error-resilience. We derive a rate-distortion bound for GSM random variables, derive the redundancy rate-distortion function, and finally implement an MD image communication system. Sumohana S. Channappayya, Robert W. Heath Jr., Alan C. Bovik |
ICASSP (2) | 3 |
| 2005 | Frame based multiple description image coding in the wavelet domainabstractMultiple description codes generated by quantized frame expansions have been shown to perform well on erasure channels when compared to traditional channel codes. In this paper we propose a multiple description image coding scheme in the wavelet domain using quantized frame expansions. We form zerotrees from wavelet coefficients and apply a tight frame operator to the zerotrees. We then group appropriate expansions to form packets and evaluate the performance of the scheme over an erasure channel. We compare the performance of the proposed scheme with a conventional channel coding scheme. Sumohana S. Channappayya, Joonsoo Lee, Robert W. Heath Jr., Alan C. Bovik |
ICIP (3) | 4 |
| 2005 | Natural contrast statistics and the selection of visual fixationsabstractIn this paper we address the problem of visual surveillance, which we define as the problem of optimally extracting information from the visual scene with a fixating, foveated imaging system. We are explicitly concerned with eye/camera movement strategies that result in maximizing information extraction from the visual field. Here we demonstrate how a novel characterization of the contrast statistics of natural images can be used for selecting fixation points that minimize the total contrast uncertainty (entropy) of natural images. We demonstrate the performance of the algorithm and compare its performance to ground truth methods. The results show that our algorithm performs favorably in terms of both efficiency and its ability to find salient features in the image. Raghu G. Raj, Wilson S. Geisler, Robert A. Frazor, Alan C. Bovik |
ICIP (3) | 4 |
| 2005 | Detecting spread spectrum watermarks using natural scene statisticsabstractThis paper presents novel techniques for detecting watermarks in images in a known-cover attack framework using natural scene models. Specifically, we consider a class of watermarking algorithms, popularly known as spread spectrum-based techniques. We attempt to classify images as either watermarked or distorted by common signal processing operations like compression, additive noise etc. The basic idea is that the statistical distortion introduced by spread spectrum watermarking is very different from that introduced by other common distortions. Our results are very promising and indicate that this statistical framework is effective in the steganalysis of spread spectrum watermarks. Kalpana Seshadrinathan, Hamid R. Sheikh, Alan C. Bovik |
ICIP (2) | 3 |
| 2005 | High quality, low delay foveated visual communications over mobile channels
Sanghoon Lee 0001, Alan C. Bovik, Young Yong Kim |
J. Vis. Commun. Image Represent. | 2 |
| 2005 | Foveation embedded DCT domain video transcoding
Shizhong Liu, Alan C. Bovik |
J. Vis. Commun. Image Represent. | 2 |
| 2005 | Supervised parametric and non-parametric classification of chromosome images
Mehul P. Sampat, Alan C. Bovik, Jake K. Aggarwal, Kenneth R. Castleman |
Pattern Recognit. | 2 |
| 2005 | Approximating filtered scale-variant signalsabstractWe develop theorems that place limits on the point-wise approximation of the responses of filters, both linear shift invariant (LSI) and linear shift variant (LSV), to input signals and images that are LSV in the following sense: they can be expressed as the outputs of systems with LSV impulse responses, where the shift variance is with respect to the filter scale of a single-prototype fillter. The approximations take the form of LSI approximations to the responses. We develop tight bounds on the approximation errors expressed in terms of filter durations and derivative (Sobolev) norms. Finally, we find application of the developed theory to defoveation of images, deblurring of shift-variant blurs, and shift-variant edge detection. Alan C. Bovik, Raghu G. Raj |
IEEE Trans. Image Process. | 1 |
| 2005 | An Image Model and Segmentation Algorithm for Reflectance Confocal Images of In Vivo Cervical TissueabstractThe automatic segmentation of nuclei in confocal reflectance images of cervical tissue is an important goal toward developing less expensive cervical precancer detection methods. Since in vivo confocal reflectance microscopy is an emerging technology for cancer detection, no prior work has been reported on the automatic segmentation of in vivo confocal reflectance images. However, prior work has shown that nuclear size and nuclear-to-cytoplasmic ratio can determine the presence or extent of cervical precancer. Thus, segmenting nuclei in confocal images will aid in cervical precancer detection. Successful segmentation of images of any type can be significantly enhanced by the introduction of accurate image models. To enable a deeper understanding of confocal reflectance microscopy images of cervical tissue, and to supply a basis for parameter selection in a classification algorithm, we have developed a model that accounts for the properties of the imaging system and of the tissues. Using our model in conjunction with a powerful image enhancement tool (anisotropic median-diffusion), appropriate statistical image modeling of spatial interactions (Gaussian Markov random fields), and a Bayesian framework for classification-segmentation, we have developed an effective algorithm for automatically segmenting nuclei in confocal images of cervical tissue. We have applied our algorithm to an extensive set of cervical images and have found that it detects 90% of hand-segmented nuclei with an average of 6 false positives per frame. Brette L. Luck, Kristen C. Maitland, Alan C. Bovik, Rebecca R. Richards-Kortum |
IEEE Trans. Image Process. | 3 |
| 2005 | No-reference quality assessment using natural scene statistics: JPEG2000abstractMeasurement of image or video quality is crucial for many image-processing algorithms, such as acquisition, compression, restoration, enhancement, and reproduction. Traditionally, image quality assessment (QA) algorithms interpret image quality as similarity with a "reference" or "perfect" image. The obvious limitation of this approach is that the reference image or video may not be available to the QA algorithm. The field of blind, or no-reference, QA, in which image quality is predicted without the reference image or video, has been largely unexplored, with algorithms focusing mostly on measuring the blocking artifacts. Emerging image and video compression technologies can avoid the dreaded blocking artifact by using various mechanisms, but they introduce other types of distortions, specifically blurring and ringing. In this paper, we propose to use natural scene statistics (NSS) to blindly measure the quality of images compressed by JPEG2000 (or any other wavelet based) image coder. We claim that natural scenes contain nonlinear dependencies that are disturbed by the compression process, and that this disturbance can be quantified and related to human perceptions of quality. We train and test our algorithm with data from human subjects, and show that reasonably comprehensive NSS models can help us in making blind, but accurate, predictions of quality. Our algorithm performs close to the limit imposed on useful prediction by the variability between human subjects. Hamid R. Sheikh, Alan C. Bovik, Lawrence K. Cormack |
IEEE Trans. Image Process. | 2 |
| 2005 | An information fidelity criterion for image quality assessment using natural scene statisticsabstractMeasurement of visual quality is of fundamental importance to numerous image and video processing applications. The goal of quality assessment (QA) research is to design algorithms that can automatically assess the quality of images or videos in a perceptually consistent manner. Traditionally, image QA algorithms interpret image quality as fidelity or similarity with a "reference" or "perfecft" image in some perceptual space. Such "full-referenc" QA methods attempt to achieve consistency in quality prediction by modeling salient physiological and psychovisual features of the human visual system (HVS), or by arbitrary signal fidelity criteria. In this paper, we approach the problem of image QA by proposing a novel information fidelity criterion that is based on natural scene statistics. QA systems are invariably involved with judging the visual quality of images and videos that are meant for "human consumption." Researchers have developed sophisticated models to capture the statistics of natural signals, that is, pictures and videos of the visual environment. Using these statistical models in an information-theoretic setting, we derive a novel QA algorithm that provides clear advantages over the traditional approaches. In particular, it is parameterless and outperforms current methods in our testing. We validate the performance of our algorithm with an extensive subjective study involving 779 images. We also show that, although our approach distinctly departs from traditional HVS-based methods, it is functionally similar to them under certain conditions, yet it outperforms them due to improved modeling. The code and the data from the subjective study are available at. Hamid R. Sheikh, Alan C. Bovik, Gustavo de Veciana |
IEEE Trans. Image Process. | 2 |
| 2005 | Maximum-likelihood techniques for joint segmentation-classification of multispectral chromosome imagesabstractTraditional chromosome imaging has been limited to grayscale images, but recently a 5-fluorophore combinatorial labeling technique (M-FISH) was developed wherein each class of chromosomes binds with a different combination of fluorophores. This results in a multispectral image, where each class of chromosomes has distinct spectral components. In this paper, we develop new methods for automatic chromosome identification by exploiting the multispectral information in M-FISH chromosome images and by jointly performing chromosome segmentation and classification. We (1) develop a maximum-likelihood hypothesis test that uses multispectral information, together with conventional criteria, to select the best segmentation possibility; (2) use this likelihood function to combine chromosome segmentation and classification into a robust chromosome identification system; and (3) show that the proposed likelihood function can also be used as a reliable indicator of errors in segmentation, errors in classification, and chromosome anomalies, which can be indicators of radiation damage, cancer, and a wide variety of inherited diseases. We show that the proposed multispectral joint segmentation-classification method outperforms past grayscale segmentation methods when decomposing touching chromosomes. We also show that it outperforms past M-FISH classification techniques that do not use segmentation information. Wade Schwartzkopf, Alan C. Bovik, Brian L. Evans |
IEEE Trans. Medical Imaging | 2 |
| 2004 | Image information and visual qualityabstractMeasurement of image quality is crucial for many image-processing algorithms. Traditionally, image quality assessment algorithms predict visual quality by comparing a distorted image against a reference image, typically by modeling the human visual system (HVS), or by using arbitrary signal fidelity criteria. We adopt a new paradigm for image quality assessment. We propose an information fidelity criterion that quantifies the Shannon information that is shared between the reference and distorted images relative to the information contained in the reference image itself. We use natural scene statistics (NSS) modeling in concert with an image degradation model and an HVS model. We demonstrate the performance of our algorithm by testing it on a data set of 779 images, and show that our method is competitive with state of the art quality assessment methods, and outperforms them in our simulations. Hamid R. Sheikh, Alan C. Bovik |
ICASSP (3) | 2 |
| 2004 | A joint source-channel distortion model for jpeg compressed imagesabstractThe need for efficient joint source-channel coding is growing as new multimedia services are introduced in commercial wireless communication systems. An important component of practical joint source-channel coding schemes is a distortion model to measure the quality of compressed digital multimedia such as images and videos. Unfortunately, models for estimating the distortion due to quantization and channel bit errors in a combined fashion do not appear to be available for practical image or video coding standards. This paper presents a statistical model for estimating the distortion introduced in progressive JPEG compressed images due to both quantization and channel bit errors. Important compression techniques such as Huffman coding, DPCM coding, and run-length coding are included in the model. Examples show that the distortion in terms of peak signal to noise ratio can be predicted within a 2 dB maximum error. Muhammad F. Sabir, Hamid R. Sheikh, Robert W. Heath Jr., Alan C. Bovik |
ICIP | 4 |
| 2004 | Video quality assessment based on structural distortion measurement
Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2004 | Image quality assessment: from error visibility to structural similarityabstractObjective methods for assessing perceptual image quality traditionally attempted to quantify the visibility of errors (differences) between a distorted image and a reference image using a variety of known properties of the human visual system. Under the assumption that human visual perception is highly adapted for extracting structural information from a scene, we introduce an alternative complementary framework for quality assessment based on the degradation of structural information. As a specific example of this concept, we develop a Structural Similarity Index and demonstrate its promise through a set of intuitive examples, as well as comparison to both subjective ratings and state-of-the-art objective methods on a database of images compressed with JPEG and JPEG2000. Zhou Wang 0001, Alan C. Bovik, Hamid R. Sheikh, Eero P. Simoncelli |
IEEE Trans. Image Process. | 2 |
| 2003 | Segmenting cervical epithelial nuclei from confocal images Gaussian Markov random fieldsabstractCervical cancer is always preceded by epithelial lesions which have larger and more densely spaced nuclei than normal tissue. Detecting and removing these lesions prevents the development of cervical cancer. A proposed method to detect precancerous lesion in vivo is to use the nuclear size and density information from fiber optic confocal images of the cervical epithelial tissue to classify the tissue as normal or precancerous. Automatically segmenting nuclei is challenging because they are hard to decipher from the noise in the confocal images. This paper outlines an algorithm to automatically segment cervical epithelial nuclei from fiber optic confocal videos using Gaussian Markov random fields. Gaussian Markov random fields segment images with additive Gaussian noise by modeling the underlying structure of the image. The algorithm described in this paper detects 90% of the nuclei in each frame with a 14% error rate. Brette L. Luck, Alan C. Bovik, Rebecca R. Richards-Kortum |
ICIP (2) | 2 |
| 2003 | Image features that draw fixationsabstractThe ability to automatically detect 'visually interesting' regions in an image has many practical applications especially in the design of active machine vision systems. This paper describes a data-driven approach that uses eye tracking in tandem with principal component analysis to extract low-level image features that attract human gaze. Data analysis on an ensemble of image patches extracted at the observer's point of gaze revealed features that resemble derivatives of the 2D Gaussian operator. Dissimilarities between human and random fixations are investigated by comparing the features extracted at the point of gaze to the general image structure obtained by random sampling in Monte-Carlo simulations. Finally, a simple application where these features are used to predict fixations is illustrated. Umesh Rajashekar, Lawrence K. Cormack, Alan C. Bovik |
ICIP (3) | 3 |
| 2003 | Fast algorithms for foveated video processingabstractThis paper explores the problem of communicating high-quality, foveated video streams in real time. Foveated video exploits the nonuniform resolution of the human visual system by preferentially allocating bits according to the proximity to assumed visual fixation points, thus delivering perceptually high quality at greatly reduced bandwidths. Foveated video streams possess specific data density properties that can be exploited to enhance the efficiency of subsequent video processing. Here, we exploit these properties to construct several efficient foveated video processing algorithms: foveation filtering (local bandwidth reduction), motion estimation, motion compensation, video rate control, and video postprocessing. Our approach leads to enhanced computational efficiency by interpreting nonuniform-density foveated images on the uniform domain and by using a foveation protocol between the encoder and the decoder. Sanghoon Lee 0001, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Foveation scalable video coding with automatic fixation selectionabstractImage and video coding is an optimization problem. A successful image and video coding algorithm delivers a good tradeoff between visual quality and other coding performance measures, such as compression, complexity, scalability, robustness, and security. In this paper, we follow two recent trends in image and video coding research. One is to incorporate human visual system (HVS) models to improve the current state-of-the-art of image and video coding algorithms by better exploiting the properties of the intended receiver. The other is to design rate scalable image and video codecs, which allow the extraction of coded visual information at continuously varying bit rates from a single compressed bitstream. Specifically, we propose a foveation scalable video coding (FSVC) algorithm which supplies good quality-compression performance as well as effective rate scalability. The key idea is to organize the encoded bitstream to provide the best decoded video at an arbitrary bit rate in terms of foveated visual quality measurement. A foveation-based HVS model plays an important role in the algorithm. The algorithm is adaptable to different applications, such as knowledge-based video coding and video communications over time-varying, multiuser and interactive networks. Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2002 | Visual search: structure from noiseabstractIn this paper, we present two techniques to reveal image features that attract the eye during visual search: the discrimination image paradigm and principal component analysis. In preliminary experiments, we employed these techniques to identify image features used to identify simple targets embedded in 1/ƒ noise. Two main findings emerged. First, the loci of fixations were not random but were driven by local image features, even in very noisy displays. Second, subjects often searched for a component feature of a target rather that the target itself, even if the target was a simple geometric form. Moreover, the particular relevant component varied from individual to individual. Also, principal component analysis of the noise patches at the point of fixation reveals global image features used by the subject in the search task. In addition to providing insight into the human visual system, these techniques have relevance for machine vision as well. The efficacy of a foveated machine vision system largely depends on its ability to actively select 'visually interesting' regions in its environment. The techniques presented in this paper provide valuable low-level criteria for executing human-like scanpaths in such machine vision systems. Umesh Rajashekar, Lawrence K. Cormack, Alan C. Bovik |
ETRA | 3 |
| 2002 | A fast and memory efficient video transcoder for low bit rate wireless communicationsabstractWireless video is one of the important applications supported by upcoming 3G mobile communication systems. In this paper, we propose a fast and memory efficient DCT-domain video transcoder to convert a high quality MPEG2 video bit stream into a low bit rate MPEG4 stream with low spatial resolution for wireless video access. Compared to existing approaches, the proposed video transcoder can save more than 50% of required memory. Furthermore, the computational complexity of the proposed method is less than 30% of that required by existing methods. However, the video quality achieved by the proposed method and by existing methods is hardly distinguishable for target bit rates of 384 kb/s and 256 kb/s, as shown in our experimental results. Shizhong Liu, Alan C. Bovik |
ICASSP | 2 |
| 2002 | Foveated multipoint videoconferencing at low bit ratesabstractMultipoint videoconferencing (MPVC) involves three or more participants engaged in video communication over a network. A video server combines the video streams from each participant and then broadcasts the resulting stream to all participants. In this paper, we propose to use foveation, which is non-uniform resolution representation of an image reflecting the sampling in the retina, to reduce the bandwidth requirements of MPVC. We develop foveated MPVC algorithms for variable and constant bit rate MPVC. We show that foveated MPVC can provide considerable bit rate savings, and for the same bit rate, provide improvement in subjective quality. Hamid R. Sheikh, Shizhong Liu, Zhou Wang 0001, Alan C. Bovik |
ICASSP | 4 |
| 2002 | Why is image quality assessment so difficult?abstractImage quality assessment plays an important role in various image processing applications. A great deal of effort has been made in recent years to develop objective image quality metrics that correlate with perceived quality measurement. Unfortunately, only limited success has been achieved. In this paper, we provide some insights on why image quality assessment is so difficult by pointing out the weaknesses of the error sensitivity based framework, which has been used by most image quality assessment approaches in the literature. Furthermore, we propose a new philosophy in designing image quality metrics: The main function of the human eyes is to extract structural information from the viewing field, and the human visual system is highly adapted for this purpose. Therefore, a measurement of structural distortion should be a good approximation of perceived image distortion. Based on the new philosophy, we implemented a simple but effective image quality indexing algorithm, which is very promising as shown by our current results. Zhou Wang 0001, Alan C. Bovik, Ligang Lu |
ICASSP | 2 |
| 2002 | No-reference perceptual quality assessment of JPEG compressed imagesabstractHuman observers can easily assess the quality of a distorted image without examining the original image as a reference. By contrast, designing objective No-Reference (NR) quality measurement algorithms is a very difficult task. Currently, NR quality assessment is feasible only when prior knowledge about the types of image distortion is available. This research aims to develop NR quality measurement algorithms for JPEG compressed images. First, we established a JPEG image database and subjective experiments were conducted on the database. We show that Peak Signal-to-Noise Ratio (PSNR), which requires the reference images, is a poor indicator of subjective quality. Therefore, tuning an NR measurement model towards PSNR is not an appropriate approach in designing NR quality metrics. Furthermore, we propose a computational and memory efficient NR quality assessment model for JPEG images. Subjective test results are used to train the model, which achieves good quality prediction performance. Hamid R. Sheikh, Zhou Wang 0001, Alan C. Bovik |
ICIP (1) | 3 |
| 2002 | Generalized bitplane-by-bitplane shift method for JPEG2000 ROI codingabstractOne interesting feature of the new JPEG2000 image coding standard is support of region of interest (ROI) coding using the maximum shift (Maxshift) method, which allows for arbitrarily shaped ROI image compression without shape coding or explicitly transmitting any shape information to the decoder. The major disadvantage of the Maxshift method is that it cannot adjust the scaling value which determines the degree of relative importance between the ROI and the background wavelet coefficients. The bitplane-by-bitplane shift (BbBShift) method was introduced to support both arbitrary ROI shape and arbitrary scaling without shape coding. We propose a generalized BbBShift (GBbBShift) method, which delivers much more flexibility than both Maxshift and BbBShift for "degree-of-interest" adjustment of the ROI with insignificant effect on coding efficiency and computational complexity. Experiments show that it can provide significantly better visual quality than Maxshift at low bit rates. GBbBShift is not compliant with the current JPEG2000 definitions. In order to use it, a new ROI coding mode would need to be added to the standard. Zhou Wang 0001, Serene Banerjee, Brian L. Evans, Alan C. Bovik |
ICIP (3) | 4 |
| 2002 | Video quality assessment using structural distortion measurementabstractObjective image/video quality measures play important roles in various image/video processing applications, such as compression, communication, printing, analysis, registration, restoration and enhancement. Most proposed quality assessment approaches in the literature are error sensitivity-based methods. We follow a new philosophy in designing image/video quality metrics, which uses structural distortion as an estimation of perceived visual distortion. We develop a new approach for video quality assessment. Experiments on the video quality experts group (VQEG) test data set shows that the new quality measure has higher correlation with subjective quality measurement than the proposed methods in VQEG's Phase I tests for full-reference video quality assessment. Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
ICIP (3) | 3 |
| 2002 | Full-reference video quality assessment considering structural distortion and no-reference quality evaluation of MPEG videoabstractThere has been an increasing need recently to develop objective quality measurement techniques that can predict perceived video quality automatically. This paper introduces two video quality assessment models. The first one requires the original video as a reference and is a structural distortion measurement based approach, which is different from traditional error sensitivity based methods. Experiments on the video quality experts group (VQEG) test data set show that the new quality measure has higher correlation with subjective quality evaluation than the proposed methods in VQEG's Phase I tests for full-reference video quality assessment. The second model is designed for quality estimation of compressed MPEG video stream without referring to the original video sequence. Preliminary experimental results show that it correlates well with our full-reference quality assessment model. Ligang Lu, Zhou Wang 0001, Alan C. Bovik, Jack Kouloheris |
ICME (1) | 3 |
| 2002 | A universal image quality indexabstractWe propose a new universal objective image quality index, which is easy to calculate and applicable to various image processing applications. Instead of using traditional error summation methods, the proposed index is designed by modeling any image distortion as a combination of three factors: loss of correlation, luminance distortion, and contrast distortion. Although the new index is mathematically defined and no human visual system model is explicitly employed, our experiments on various image distortion types indicate that it performs significantly better than the widely used distortion metric mean squared error. Demonstrative images and an efficient MATLAB implementation of the algorithm are available online at http://anchovy.ece.utexas.edu//spl sim/zwang/research/quality_index/demo.html. Zhou Wang 0001, Alan C. Bovik |
IEEE Signal Process. Lett. | 2 |
| 2002 | Bitplane-by-bitplane shift (BbBShift) - a suggestion for JPEG2000 region of interest image codingabstractThe JPEG2000 image coding standard defines two kinds of region of interest (ROI) coding methods-the general scaling based method and the maximum shift (maxshift) method. The former requires shape coding of the ROIs, which leads to increased complexity of codec implementations and limits the choice of ROI shapes (currently, only rectangle and ellipse shapes are defined). The latter allows for arbitrarily shaped ROI coding without explicitly transmitting any shape information to the decoder, but does not have the flexibility to select an arbitrary scaling value to define the relative importance of the ROI and the background wavelet coefficients. We propose a bitplane-by-bitplane shift (BbBShift) method, which supports both arbitrary ROI shape and arbitrary scaling without shape coding. Zhou Wang 0001, Alan C. Bovik |
IEEE Signal Process. Lett. | 2 |
| 2002 | Local bandwidth constrained fast inverse motion compensation for DCT-domain video transcodingabstractDiscrete cosine transform (DCT) based digital video coding standards, such as MPEG and H.26x, are becoming more widely adopted for multimedia applications. Since the standards differ in their format and syntax, video transcoding, where a compressed video bit-stream is converted from one format to another, is of interest for purposes such as channel bandwidth adaptation and video composition. DCT-domain video transcoding is generally more efficient than spatial-domain transcoding. However, since the data is organized block by block in the DCT domain, inverse motion compensation becomes the bottleneck for DCT-domain methods. We propose a novel local bandwidth constrained fast inverse motion compensation algorithm operating in the DCT domain. Relative to Chang's algorithm, we achieve computational improvement of 25%-55% without visual degradation. A by-product of our algorithm is a reduction of blocking artifacts in very low bit-rate compressed video sequences. The proposed algorithm can be combined with other reported fast methods for more computational savings. We also present a look-up-table (LUT) based implementation method by modeling the statistical distribution of the DCT coefficients in natural images and video sequences. By this method, we obtain a further 31%-48% improvement in computation. The memory requirement of the LUT is about 800 kB, which is reasonable. Moreover, the LUT can be shared by multiple DCT-domain video processing applications running on the same computer or video server. Shizhong Liu, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | . Efficient DCT-domain blind measurement and reduction of blocking artifactsabstractBlocking artifacts continue to be among the most serious defects that occur in images and video streams compressed to low bit rates using block discrete cosine transform (DCT)-based compression standards (e.g., JPEG, MPEG, and H.263). It is of interest to be able to numerically assess the degree of blocking artifact in a visual signal, for example, in order to objectively determine the efficacy of a compression method, or to discover the quality of video content being delivered by a web server. We propose new methods for efficiently assessing, and subsequently reducing, the severity of blocking artifacts in compressed image bitstreams. The method is blind, and operates only in the DCT domain. Hence, it can be applied to unknown visual signals, and it is efficient since the signal need not be compressed or decompressed. In the algorithm, blocking artifacts are modeled as 2-D step functions. A fast DCT-domain algorithm extracts all parameters needed to detect the presence of, and estimate the amplitude of blocking artifacts, by exploiting several properties of the human vision system. Using the estimate of blockiness, a novel DCT-domain method is then developed which adaptively reduces detected blocking artifacts. Our experimental results show that the proposed method of measuring blocking artifacts is effective and stable across a wide variety of images. Moreover, the proposed blocking-artifact reduction method exhibits satisfactory performance as compared to other post-processing techniques. The proposed technique has a low computational cost hence can be used for real-time image/video quality monitoring and control, especially in applications where it is desired that the image/video data be processed directly in the DCT-domain. Shizhong Liu, Alan C. Bovik |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Smoothing low SNR molecular images via anisotropic median-diffusionabstractWe propose a new anisotropic diffusion filter for denoising low-signal-to-noise molecular images. This filter, which incorporates a median filter into the diffusion steps, is called an anisotropic median-diffusion filter. This hybrid filter achieved much better noise suppression with minimum edge blurring compared with the original anisotropic diffusion filter when it was tested on an image created based on a molecular image model. The universal quality index, proposed in this paper to measure the effectiveness of denoising, suggests that the anisotropic median-diffusion filter can retain adherence to the original image intensities and contrasts better than other filters. In addition, the performance of the filter is less sensitive to the selection of the image gradient threshold during diffusion, thus making automatic image denoising easier than with the original anisotropic diffusion filter. The anisotropic median-diffusion filter also achieved good denoising results on a piecewise-smooth natural image and real Raman molecular images. Jian Ling, Alan C. Bovik |
IEEE Trans. Medical Imaging | 2 |
| 2002 | Foveated video quality assessmentabstractMost image and video compression algorithms that have been proposed to improve picture quality relative to compression efficiency have either been designed based on objective criteria such as signal-to-noise-ratio (SNR) or have been evaluated, post-design, against competing methods using an objective sample measure. However, existing quantitative design criteria and numerical measurements of image and video quality both fail to adequately capture those attributes deemed important by the human visual system, except, perhaps, at very low error rates. We present a framework for assessing the quality of and determining the efficiency of foveated and compressed images and video streams. Image foveation is a process of nonuniform sampling that accords with the acquisition of visual information at the human retina. Foveated image/video compression algorithms seek to exploit this reduction of sensed information by nonuniformly reducing the resolution of the visual data. We develop unique algorithms for assessing the quality of foveated image/video data using a model of human visual response. We demonstrate these concepts on foveated, compressed video streams using modified (foveated) versions of H.263 that are standard-compliant. We rind that quality vs. compression is enhanced considerably by the foveation approach. Sanghoon Lee 0001, Marios S. Pattichis, Alan C. Bovik |
IEEE Trans. Multim. | 3 |
| 2001 | DCT-domain blind measurement of blocking artifacts in DCT-coded imagesabstractA method for DCT-domain blind measurement of blocking artifacts is proposed. By constituting a new block across any two adjacent blocks, the blocking artifact is modeled as a 2-D step function. A fast DCT-domain algorithm has been derived to constitute the new block and extract all parameters needed. Then an human visual system (HVS) based measurement of blocking artifacts is conducted. Experimental results have shown the effectiveness and stability of our method. The proposed technique can be used for online image/video quality monitoring and control in applications of DCT-domain image/video processing. Alan C. Bovik, Shizhong Liu |
ICASSP | 1 |
| 2001 | Local bandwidth constrained fast inverse motion compensation for DCT-domain video transcodingabstractDCT-based digital video coding standards such as MPEG and H.26x are becoming more widely adopted for multimedia applications. Since the standards differ in their format and syntax, video transcoding, where a pre-coded video bit-stream is converted from one format to another format, is of interest for purposes such as channel bandwidth adaptation and video composition. DCT-domain video transcoding is generally more efficient than spatial domain transcoding. However, since the data is organized block by block in the DCT-domain, inverse motion compensation becomes the bottleneck for DCT-domain methods. We propose a novel local bandwidth constrained fast inverse motion compensation algorithm operating in the DCT-domain. Relative to Chang's (1995)algorithm, the proposed algorithm achieves computational improvement of 25% to 55% without visual degradation. A by-product of the proposed algorithm is a reduction of blocking artifacts in very low bit-rate compressed video sequences. Shizhong Liu, Alan C. Bovik |
ICASSP | 2 |
| 2001 | Real-time foveation techniques for H.263 video encoding in softwareabstractVideo coding techniques employ characteristics of the human visual system (HVS) to achieve high coding efficiency. Lee (2000) and Bovik have exploited foveation, which is a non-uniform resolution representation of an image reflecting the sampling in the retina, for low bit-rate video coding. We develop a fast approximation of the foveation model and demonstrate real-time foveation techniques in the spatial domain and discrete cosine transform (DCT) domain. We incorporate fast DCT domain foveation into the baseline H.263 video encoding standard. We show that DCT-domain foveation requires much lower computational overhead but generates higher bit rates than spatial domain foveation. Our techniques do not require any modifications of the decoder. Hamid R. Sheikh, Shizhong Liu, Brian L. Evans, Alan C. Bovik |
ICASSP | 4 |
| 2001 | Rate scalable video coding using a foveation-based human visual system modelabstractRecently, there have been two interesting trends in image and video coding research. One is to use human visual system (HVS) models to improve the current state-of-the-art coding algorithms by better exploiting the properties of the intended receiver. The other is to design rate-scalable video codecs, which allow the extraction of coded visual information at continuously varying bit rates from a single compressed bitstream. We follow these two trends and propose a foveation scalable video coding (FSVC) algorithm, which supplies good quality-compression performance as well as effective rate scalability to support simple and precise bit rate control. A foveation-based HVS model plays a key role in the algorithm. The algorithm is amenable to the inclusion of various HVS models and adaptable to different video communication applications. Zhou Wang 0001, Ligang Lu, Alan C. Bovik |
ICASSP | 3 |
| 2001 | Look-up-table based DCT domain inverse motion compensationabstractDCT-based digital video coding standards such as MPEG and H.26x have been widely adopted for multimedia applications. Thus video processing in the DCT domain usually proves to be more efficient than in the spatial domain. To directly convert an inter-coded frame into an intra-coded frame in the DCT domain, the problem of DCT domain inverse motion compensation was studied by Chang and Messerschmitt(1995). Since the data is organized block by block in the DCT domain, the DCT domain inverse motion compensation is computationally intensive. In this paper, a look-up-table (LUT) based method for DCT domain inverse motion compensation is proposed by modeling the statistical distribution of the DCT coefficients in typical images and video sequences. Compared to the method of Chang et al., the LUT based method can save more than 50% of the computing time based on experimental results. The memory requirement of the LUT is about 800 KB which is reasonable. Moreover, the LUT can be shared by multiple DCT domain video processing applications running on the same computer. Shizhong Liu, Alan C. Bovik |
ICIP (2) | 2 |
| 2001 | Minimum entropy segmentation applied to multi-spectral chromosome imagesabstractIn the early 1990s, the state-of-the-art in commercial chromosome image acquisition was grayscale. Automated chromosome classification was based on the grayscale image and boundary information obtained during segmentation. Multi-spectral image acquisition was developed in 1990 and commercialized in the mid-1990s. One acquisition method, multiplex fluorescence in-situ hybridization (M-FISH), uses five color dyes. We propose a segmentation algorithm for M-FISH images that minimizes the entropy of classified pixels within possible chromosomes. This method is shown to correctly decompose even difficult clusters of touching and overlapping chromosomes. Finally, an example image is given to illustrate the algorithm. Wade Schwartzkopf, Brian L. Evans, Alan C. Bovik |
ICIP (2) | 3 |
| 2001 | Wavelet-based foveated image quality measurement for region of interest image codingabstractRegion of interest (ROI) image and video compression techniques have been widely used in visual communication applications in an effort to deliver good quality images and videos at limited bandwidths. Most image quality metrics have been developed for uniform resolution images. These metrics are not appropriate for the assessment of ROI coded images, where space-variant resolution is necessary. The spatial resolution of the human visual system (HVS) is highest around the point of fixation and decreases rapidly with increasing eccentricity. Since the ROIs are usually the regions "fixated" by human eyes, the foveation property of the HVS supplies a natural approach for guiding the design of ROI image quality measurement algorithms. We have developed an objective quality metric for ROI coded images in the wavelet transform domain. This metric can serve to mediate the compression and enhancement of ROI coded images and videos. We show its effectiveness by applying it to an embedded foveated image coding system. Zhou Wang 0001, Alan C. Bovik, Ligang Lu |
ICIP (2) | 2 |
| 2001 | Adaptive Frame Prediction for Foveation Scalable Video CodingabstractEmbedded rate scalable video coding allows for the extraction of coded visual information at continuously varying bit rates from a single compressed bitstream. This is a very attractive feature for many multimedia communication applications. Motion estimation (ME) /motion compensation (MC) techniques are widely employed in various video coding systems to reduce temporal information redundancy. One of the major challenging problems in ME/MC based rate scalable video coding is how to generate the prediction frame from the previous frame to match the current frame. This problem is more difficult in rate scalable coding than in fixed rate coding because the decoding data rate is unavailable to the encoder. We propose an adaptive frame prediction scheme for foveation scalable video coding (FSVC), which is a new video coding algorithm that combines a foveation-based human visual system (HVS) model with a wavelet-based rate scalable coding algorithm. The new frame prediction algorithm provides an adaptive mechanism to control the prediction errors while reduce error propagation. Ligang Lu, Zhou Wang 0001, Alan C. Bovik, Jack Kouloheris |
ICME | 3 |
| 2001 | Oriented texture completion by AM-FM reaction-diffusionabstractWe provide an automated method to repair broken, occluded oriented image textures. Our approach is based on partial differential equations (PDEs) and AM-FM image modeling. Reconstruction of the texture occurs via simultaneous PDE-generated diffusion and reaction. In the diffusion process, the image is adaptively smoothed, preserving important boundaries and features. The reaction process produces the reconstructed textural information in the occluded image regions. Gabor (1946) filters are designed and used in the reaction process using an AM-FM dominant component analysis. An AM-FM model of the texture image is constructed, making it possible to localize the reaction filters spatio-spectrally. In contrast to previous disocclusion techniques that depend on interpolation, on continuity of the connected components within the image level sets, or on texture estimation, the reaction-diffusion process proposed here yields a seamless transition between the recreated region and the unoccluded image regions. Using AM-FM dominant component analysis, we avoid the ad hoc parameter selection typified with other reaction-diffusion approaches. As a useful example, we focus on the repair of broken, occluded fingerprints. We also treat several exemplary natural textures to demonstrate the technique's generality. Scott T. Acton, Dipti Prasad Mukherjee, Joseph P. Havlicek, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2001 | Foveated video compression with optimal rate controlabstractPreviously, fovcated video compression algorithms have been proposed which, in certain applications, deliver high-quality video at reduced bit rates by seeking to match the nonuniform sampling of the human retina. We describe such a framework here where foveated video is created by a nonuniform filtering scheme that increases the compressibility of the video stream. We maximize a new foveal visual quality metric. the foveal signal-to-noise ratio (FSNR) to determine the best compression and rate control parameters for a given target bit rate. Specifically, we establish a new optimal rate control algorithm for maximizing the FSNR using a Lagrange multiplier method defined on a curvilinear coordinate system. For optimal rate control, we also develop a piecewise R-D (rate-distortion)/R-Q (rate-quantization) model. A fast algorithm for searching for an optimal Lagrange multiplier lambda* is subsequently presented. For the new models, we show how the reconstructed video quality is affected, where the FSNR is maximized, and demonstrate the coding performance for H.263,+,++/MPEG-4 video coding. For H.263/MPEG video coding, a suboptimal rate control algorithm is developed for fast, high-performance applications. In the simulations, we compare the reconstructed pictures obtained using optimal rate control methods for foveated and normal video. We show that foveated video coding using the suboptimal rate control algorithm delivers excellent performance under 64 kb/s. Sanghoon Lee 0001, Marios S. Pattichis, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2001 | On eigenstructure-based direct multichannel blind image restorationabstractExisting eigenstructure-based direct multichannel blind image restoration techniques include nullspace-based and direct deconvolver estimation techniques. The nullspace-based approach can be formulated as an optimization problem. We show that this formulation implies a new subspace-based approach that uses matrix operations. This new approach has the same advantages as the nullspace-based one but requires less computational complexity. Under some mild conditions, its complexity is equal to that of the FFT. Furthermore, the relation among the nullspace-based approach, the direct deconvolver estimation and the new subspace-based approach is studied. Hung-Ta Pai, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2001 | Multidimensional orthogonal FM transformsabstractThe present a novel class of multidimensional orthogonal FM transforms. The analysis suggests a novel signal-adaptive FM transform possessing interesting energy compaction properties. We show that the proposed signal-adaptive FM transform produces point spectra for multidimensional signals with uniformly distributed samples. This suggests that the proposed transform is suitable for energy compaction and subsequent coding of broadband signals and images that locally exhibit significant level diversity. We illustrate these concepts with simulation experiments. Marios S. Pattichis, Alan C. Bovik, John W. Havlicek, Nicholas D. Sidiropoulos |
IEEE Trans. Image Process. | 2 |
| 2001 | Fingerprint classification using an AM-FM modelabstractResearch on fingerprint classification has primarily focused on finding improved classifiers, image and feature enhancement, and less on the development of novel fingerprint representations. Using an AM-FM representation for each fingerprint, we obtain significant gains in classification performance as compared to the commonly used National Institute of Standards system, for the same classifier. Marios S. Pattichis, George Panayi, Alan C. Bovik, Shun-Pin Hsu |
IEEE Trans. Image Process. | 3 |
| 2001 | Embedded foveation image codingabstractThe human visual system (HVS) is highly space-variant in sampling, coding, processing, and understanding. The spatial resolution of the HVS is highest around the point of fixation (foveation point) and decreases rapidly with increasing eccentricity. By taking advantage of this fact, it is possible to remove considerable high-frequency information redundancy from the peripheral regions and still reconstruct a perceptually good quality image. Great success has been obtained previously by a class of embedded wavelet image coding algorithms, such as the embedded zerotree wavelet (EZW) and the set partitioning in hierarchical trees (SPIHT) algorithms. Embedded wavelet coding not only provides very good compression performance, but also has the property that the bitstream can be truncated at any point and still be decoded to recreate a reasonably good quality image. In this paper, we propose an embedded foveation image coding (EFIC) algorithm, which orders the encoded bitstream to optimize foveated visual quality at arbitrary bit-rates. A foveation-based image quality metric, namely, foveated wavelet image quality index (FWQI), plays an important role in the EFIC system. We also developed a modified SPIHT algorithm to improve the coding efficiency. Experiments show that EFIC integrates foveation filtering with foveated image coding and demonstrates very good coding performance and scalability in terms of foveated image quality measurement. Zhou Wang 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2000 | Unequal Error Protection for Foveation-Based Error Resilience over Mobile NetworksabstractIn this paper, we introduce an unequal error protection technique for foveation-based error resilience over highly error-prone mobile networks. For point-to-point visual communications, visual quality can be significantly increased by using foveation-based error resilience where each frame is divided into foveated and background layers according to the gaze direction of the human eye, and two bitstreams are generated. In an effort to increase the source throughput of the foveated layer, we employ unequal delay-constrained ARQ and RCPC (rate compatible punctured convolutional) codes in H.223 Annex C. In the simulation, the visual quality is increased in the range of 0.3 dB to 1 dB over channel SNR 5 dB to 15 dB. Sanghoon Lee 0001, Christine Podilchuk, Vidhya Krishnan, Alan C. Bovik |
ICIP | 4 |
| 2000 | Modeling and Restoration of Raman Microscopic ImagesabstractPresents a model for Raman microscopic images. The model describes the degradation of Raman signals by non-uniform illumination, by the microscopic system, and by additive signal-dependent Gaussian noise. Using this model, synthetic images were created to validate the model. Based on these synthetic images, an anisotropic diffusion filter was applied to reduce the signal-dependent Gaussian noise and at the same time not blur the objects' boundary. A Wiener filter was used to restore the blurred Raman images by the microscopic system. And, an image ratioing method was used to correct for the non-uniform illumination. After the restoration, the mean absolute error between the restored image and the true image was minimized. Jian Ling, Alan C. Bovik |
ICIP | 2 |
| 2000 | Blind Measurement of Blocking Artifacts in ImagesabstractThe objective measurement of blocking artifacts plays an important role in the design, optimization, and assessment of image and video coding systems. We propose a new approach that can blindly measure blocking artifacts in images without reference to the originals. The key idea is to model the blocky image as a non-blocky image interfered with a pure blocky signal. The task of the blocking effect measurement algorithm is then to detect and evaluate the power of the blocky signal. The proposed approach has the flexibility to integrate human visual system features such as the luminance and the texture masking effects. Zhou Wang 0001, Alan C. Bovik, Brian L. Evans |
ICIP | 2 |
| 2000 | Generalized predictive binary shape coding using polygon approximation
Jong-Il Kim, Alan C. Bovik, Brian L. Evans |
Signal Process. Image Commun. | 2 |
| 2000 | Image quality assessment based on a degradation modelabstractWe model a degraded image as an original image that has been subject to linear frequency distortion and additive noise injection. Since the psychovisual effects of frequency distortion and noise injection are independent, we decouple these two sources of degradation and measure their effect on the human visual system. We develop a distortion measure (DM) of the effect of frequency distortion, and a noise quality measure (NQM) of the effect of additive noise. The NQM, which is based on Peli's (1990) contrast pyramid, takes into account the following: 1) variation in contrast sensitivity with distance, image dimensions, and spatial frequency; 2) variation in the local luminance mean; 3) contrast interaction between spatial frequencies; 4) contrast masking effects. For additive noise, we demonstrate that the nonlinear NQM is a better measure of visual quality than peak signal-to noise ratio (PSNR) and linear quality measures. We compute the DM in three steps. First, we find the frequency distortion in the degraded image. Second, we compute the deviation of this frequency distortion from an allpass response of unity gain (no distortion). Finally, we weight the deviation by a model of the frequency response of the human visual system and integrate over the visible frequencies. We demonstrate how to decouple distortion and additive noise degradation in a practical image restoration system. Niranjan Damera-Venkata, Thomas D. Kite, Wilson S. Geisler, Brian L. Evans, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2000 | Multidimensional quasi-eigenfunction approximations and multicomponent AM-FM modelsabstractWe develop multicomponent AM-FM models for multidimensional signals. The analysis is cast in a general n-dimensional framework where the component modulating functions are assumed to lie in certain Sobolev spaces. For both continuous and discrete linear shift invariant (LSI) systems with AM-FM inputs, powerful new approximations are introduced that provide closed form expressions for the responses in terms of the input modulations. The approximation errors are bounded by generalized energy variances quantifying the localization of the filter impulse response and by Sobolev norms quantifying the smoothness of the modulations. The approximations are then used to develop novel spatially localized demodulation algorithms that estimate the AM and FM functions for multiple signal components simultaneously from the channel responses of a multiband linear filterbank used to isolate components. Two discrete computational paradigms are presented. Dominant component analysis estimates the locally dominant modulations in a signal, which are useful in a variety of machine vision applications, while channelized components analysis delivers a true multidimensional multicomponent signal representation. We demonstrate the techniques on several images of general interest in practical applications, and obtain reconstructions that establish the validity of characterizing images of this type as sums of locally narrowband modulated components. Joseph P. Havlicek, David S. Harding, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2000 | A fast, high-quality inverse halftoning algorithm for error diffused halftonesabstractHalftones and other binary images are difficult to process with causing several degradation. Degradation is greatly reduced if the halftone is inverse halftoned (converted to grayscale) before scaling, sharpening, rotating, or other processing. For error diffused halftones, we present (1) a fast inverse halftoning algorithm and (2) a new multiscale gradient estimator. The inverse halftoning algorithm is based on anisotropic diffusion. It uses the new multiscale gradient estimator to vary the tradeoff between spatial resolution and grayscale resolution at each pixel to obtain a sharp image with a low perceived noise level. Because the algorithm requires fewer than 300 arithmetic operations per pixel and processes 7x7 neighborhoods of halftone pixels, it is well suited for implementation in VLSI and embedded software. We compare the implementation cost, peak signal to noise ratio, and visual quality with other inverse halftoning algorithms. Thomas D. Kite, Niranjan Damera-Venkata, Brian L. Evans, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2000 | Modeling and quality assessment of halftoning by error diffusionabstractDigital halftoning quantizes a graylevel image to one bit per pixel. Halftoning by error diffusion reduces local quantization error by filtering the quantization error in a feedback loop. In this paper, we linearize error diffusion algorithms by modeling the quantizer as a linear gain plus additive noise. We confirm the accuracy of the linear model in three independent ways. Using the linear model, we quantify the two primary effects of error diffusion: edge sharpening and noise shaping. For each effect, we develop an objective measure of its impact on the subjective quality of the halftone. Edge sharpening is proportional to the linear gain, and we give a formula to estimate the gain from a given error filter. In quantifying the noise, we modify the input image to compensate for the sharpening distortion and apply a perceptually weighted signal-to-noise ratio to the residual of the halftone and modified input image. We compute the correlation between the residual and the original image to show when the residual can be considered signal independent. We also compute a tonality measure similar to total harmonic distortion. We use the proposed measures for edge sharpening, noise shaping, and tonality to evaluate the quality of error diffusion algorithms. Thomas D. Kite, Brian L. Evans, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2000 | AM-FM Texture Segmentation in Electron Microscopy Muscle ImagingabstractThis paper describes the application of an amplitude modulation-frequency modulation (AM-FM) image representation in segmenting electron micrographs of skeletal muscle for the recognition of: 1) normal sarcomere ultrastructural pattern and 2) abnormal regions that occur in sarcomeres in various myopathies. A total of 26 electron micrographs from different myopathies were used for this study. It is shown that the AM-FM image representation can identify normal repetitive structures and sarcomeres, with a good degree of accuracy. This system can also detect abnormalities in sarcomeres which alter the normal regular pattern, as seen in muscle pathology, with a recognition accuracy of 75%-84% as compared to a human expert. Marios S. Pattichis, Constantinos S. Pattichis, Maria Avraam, Alan C. Bovik, Kyriacos C. Kyriacou |
IEEE Trans. Medical Imaging | 4 |
| 1999 | Very low bit rate foveated video coding for H.263abstractRecently, foveated video has been introduced as an important emerging method for very low bit rate multimedia applications. In this paper, we develop several rate control algorithms, and measure the performance of foveated video. We utilize H.263 video, and compare the performance with regular video based on the SNRC (signal-to-noise-ratio in curvilinear coordinates). In order to maximize compression, we use a maximum quantization parameter (QP=31) for the regular video, and code a foveated video sequence at the equivalent bit rate. In simulation, we improve the PSNRC to 3.64 (1.62)dB under 30 (14) kbits/sec for P pictures in CIF "News" ("Akiyo") standard video sequence. Sanghoon Lee 0001, Alan C. Bovik |
ICASSP | 2 |
| 1999 | AM-FM texture segmentation in electron microscopic muscle imagingabstractWe segment the structural units of electron microscope muscle images using a novel AM-FM image representation. This novel AM-FM approach is shown to be effective in describing sarcomeres and mitochondrial regions of the electron microscope muscle images. Marios S. Pattichis, Constantinos S. Pattichis, Maria Avraam, Alan C. Bovik, Kyriakos Kyriakou |
ICASSP | 4 |
| 1999 | Fast Rehalftoning and Interpolated Halftoning Algorithms with Flat LWO-Frequencey ResponseabstractWe present raster image processing algorithms for rehalftoning error diffused halftones and producing interpolated error diffused halftones. Rehalftoning converts a halftone created by one method into one created by another method. In interpolated halftoning, interpolation increases the image size before halftoning, e.g. for printing. Both rehalftoning and interpolated halftoning introduce blur and noise in the output image. To compensate for the blur, we use modified error diffusion, which has a variable gain parameter to control the sharpness. We derive optimal formulas for the sharpness control parameter to make the overall frequency response flat. The high-frequency noise is masked by error diffusion. The proposed algorithms yield halftones of high fidelity at a low computational cost. Thomas D. Kite, Brian L. Evans, Alan C. Bovik |
ICIP (3) | 3 |
| 1999 | Motion Estimation and Compensation for Foveated VideoabstractUtilizing the nonuniform resolution property of the human visual system, foveated video can provide high visual quality relative to non-foveated video by allocating more bits to the central foveation area. In this paper, we present a motion estimation and compensation algorithm for foveated video and measure the performance using a new measure of visual fidelity termed foveal mean absolute distortion. The computation redundancy reduction is achieved by subsampling searching area dependent on the local bandwidth in the sense of Nyquist sampling criterion. In addition, we reduce motion compensated errors by increasing temporal correlation when single or multiple foveation points are added or subtracted. Sanghoon Lee 0001, Alan C. Bovik |
ICIP (2) | 2 |
| 1999 | Low Delay Foveated Visual Communications over Wireless ChannelsabstractThe great potential of "foveated imaging" lies in the entropy reduction relative to the original image while minimizing the loss of visual information. Utilizing human foveation combined with video compression, as well as communication and human-machine interface techniques, more efficient multimedia services are expected to be provided in the near future. In this paper, we introduce a prototype for foveated visual communications as one of future human interactive multimedia applications, and demonstrate the benefit of the foveation over fading statistics in the downtown area of Austin, Texas. In order to compare the performance with regular video, we use spatial/temporal resolution and source transmission delay as the evaluation criteria. Sanghoon Lee 0001, Alan C. Bovik, Young Yong Kim |
ICIP (3) | 2 |
| 1999 | Low-Complexity Velocity Estimation in High-Speed Optical Doppler Tomography SystemsabstractOptical Doppler Tomography (ODT) is a noninvasive 3-D optical interferometric imaging technique that measures static and dynamic structures in a sample. To obtain the dynamic structure, e.g. blood flowing in tissue, a velocity estimation algorithm detects the Doppler shift in the received interference fringe data with respect to the carrier frequency. Previous velocity estimation algorithms use conventional Fourier magnitude techniques that do not provide sufficient frequency resolution in fast ODT systems because of the high data acquisition rates and hence short time series. In this paper, we propose a nonlinear algorithm that uses the phase shift between two successive scans of interference fringe data to give a high-resolution estimate of the Doppler shift. The algorithm detects Doppler shifts of 0.1 to 3 kHz with respect to a 1 MHz carrier. In processing 5 frames/s with 100×100 pixels/frame and 32 samples/pixel, i.e. 1.6 million samples/s, the algorithm requires 26 million multiply-accumulates/s. The algorithm works well at 4 bits/sample. The low complexity and small input data size are well-suited for real-time implementation in software. We provide a mathematical analysis of the Doppler shift resolution by modeling the interference fringe data as an AM-FM signal. Milos Milosevic, Wade Schwartzkopf, Thomas E. Milner, Brian L. Evans, Alan C. Bovik |
ICIP (2) | 5 |
| 1999 | A stereo visual pattern image coding system
Danielle Craievich, Barry S. Barnett, Alan C. Bovik |
Image Vis. Comput. | 3 |
| 1999 | Piecewise and local image models for regularized image restoration using cross-validationabstractWe describe two broad classes of useful and physically meaningful image models that can be used to construct novel smoothing constraints for use in the regularized image restoration problem. The two classes, termed piecewise image models (PIMs) and focal image models (LIMs), respectively, capture unique image properties that can be adapted to the image and that reflect structurally significant surface characteristics. Members of the PIM and LIM classes are easily formed into regularization operators that replace differential-type constraints. We also develop an adaptive strategy for selecting the best PIM or LIM for a given problem (from among the defined class), and we explain the construction of the corresponding regularization operators. Considerable attention is also given to determining the regularization parameter via a cross-validation technique, and also to the selection of an optimization strategy for solving the problem. Several results are provided that illustrate the processes of model selection, parameter selection, and image restoration. The overall approach provides a new viewpoint on the restoration problem through the use of new image models that capture salient image features that are not well represented through traditional approaches. Scott T. Acton, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 1999 | Stereoscopic ranging by matching image modulationsabstractWe apply an AM-FM surface albedo model to analyze the projection of surface patterns viewed through a binocular camera system. This is used to support the use of modulation-based stereo matching where local image phase is used to compute stereo disparities. The local image phase is an advantageous feature for image matching, since the problem of computing disparities reduces to identifying local phase shifts between the stereoscopic image data. Local phase shifts, however, are problematic at high frequencies due to phase wrapping when disparities exceed +/-pi. We meld powerful multichannel Gabor image demodulation techniques for multiscale (coarse-to-fine) computation of local image phase with a disparity channel model for depth computation. The resulting framework unifies phase-based matching approaches with AM-FM surface/image models. We demonstrate the concepts in a stereo algorithm that generates a dense, accurate disparity map without the problems associated with phase wrapping. Tieh-Yuh Chen, Alan C. Bovik, Lawrence K. Cormack |
IEEE Trans. Image Process. | 2 |
| 1999 | Comments on "Subband coding of images using asymmetrical filterbanks"abstractIn the above paper, Egger and Li presented a set of two-channel filterbanks, asymmetrical filterbanks (AFB's), for image coding applications. The basic properties of these filters are linear-phase, perfect reconstruction, asymmetric lengths for dual filters, and maximum regularity. In this correspondence, we point out that the proposed AFB's are not new in the sense that the proposed construction is equivalent to the factorization of Lagrange halfband filters, which has been reported by other researchers. In addition, we correct an error in the formulation of constructing AFBs in their paper. Dong Wei 0003, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 1998 | Generically sufficient conditions for exact multichannel blind image restorationabstractWe have previously developed an algorithm and sufficient conditions for exact multichannel blind image restoration. In this paper, we use the resultant matrix theorem and techniques of algebraic geometry to prove that the sufficient conditions hold generically given three blurred versions of the same image and some restrictions on the size of the original image. Moreover, the extension to multichannel blind n-dimensional signal restoration is described. Hung-Ta Pai, John W. Havlicek, Alan C. Bovik |
ICASSP | 3 |
| 1998 | COPERM: transform-domain energy compaction by optimal permutationabstractCOPERM is a novel paradigm for energy compaction and signal compression, whose foundation is a simple but powerful idea: any signal can be transformed to resemble a more desirable signal from a class of "target" signals, by means of a suitable permutation of its samples. The approach is well-suited for transform domain energy compaction prior to transform-domain compression of persistent broadband signals. The associated optimal permutation precoders are surprisingly simple, and the permutation precoding overhead can be made modest-resulting in improved overall rate-distortion performance. Nicholas D. Sidiropoulos, Marios S. Pattichis, Alan C. Bovik, John W. Havlicek |
ICASSP | 3 |
| 1998 | Skewed 2D Hilbert Transforms and Computed AM-FM ModelsabstractComputed AM-FM models represent images in terms of instantaneous amplitude and frequency modulations. However, the instantaneous amplitude and frequency of a real valued image are ambiguous. We apply the directional 2D Hilbert transform to compute a complex extension for a real image. This extension, called the analytic image, admits most of the attractive properties of the 1D analytic signal. However, the analytic image is not unique: for a given real image, taking the Hilbert transform in the horizontal and vertical directions yields different complex extensions and differing computed AM-FM models. We show that these two differing models are essentially equivalent and develop explicit formulations relating them. Joseph P. Havlicek, John W. Havlicek, Ngao D. Mamuya, Alan C. Bovik |
ICIP (1) | 4 |
| 1998 | A High Quality Fast Inverse Halftoning Algorithm for Error Diffused HalftonesabstractWe present an inverse halftoning algorithm for error diffused halftones. At each pixel, the algorithm applies a separable 7/spl times/7 FIR filter parameterized by the computed local horizontal and vertical gradients. All operations are entirely local; only 7 rows of image storage and fewer than 300 operations per pixel are required. The algorithm can be easily implemented in embedded software or hardware. We compare our algorithm with previously reported approaches, and show that it delivers comparable PSNR and subjective quality at a fraction of the computation and memory requirements. A C implementation of the algorithm is available. Thomas D. Kite, Niranjan Damera-Venkata, Brian L. Evans, Alan C. Bovik |
ICIP (2) | 4 |
| 1998 | Rate Control for Foveated MPEG/H.263 VideoabstractGiven a set of target bits, video rate control algorithms that use Lagrange multipliers have been generally known as an optimal solution for maximizing the picture quality in the uniform spatial domain. Even if the SNR (signal-to-noise) of a picture is maximized by the rate control scheme, the visual quality can be enhanced using a suitable algorithm for the human visual system. We establish a new optimal rate control algorithm for maximizing the SNRC (signal-to-noise ratio in curvilinear coordinates) using the Lagrange multiplier. In addition, a target bit allocation technique for foveated video is introduced for simplified rate control over MEPG/H.263 video standards. Sanghoon Lee 0001, Marios S. Pattichis, Alan C. Bovik |
ICIP (2) | 3 |
| 1998 | Maximally Flat Bandwidth Allocation for Variable Bit Rate VideoabstractWhen a variable bit rate (VBR) video bitstream is transmitted into a network, the channel utilization is dependent on the maximum peak value of the traffic pattern and improved by reducing the peak rate. A lower bound on the transmission rate is derived from the given traffic cells. Based on the transmission rate constraint, the optimal bandwidth that is needed to meet the delay requirement during an allocation interval is also derived. From the delay bound and the constraints, we demonstrate that the traffic has been made as flat as possible. Sanghoon Lee 0001, Alan C. Bovik |
ICIP (2) | 2 |
| 1998 | Antisymmetric Biorthogonal Coiflets for Image CodingabstractWavelet techniques have achieved a tremendous success in image data compression. In designing wavelet coding algorithms, the choice of wavelet systems is of great importance for compression performance. We design the entire class of antisymmetric biorthogonal coiflet systems, whose filterbanks have even lengths and are linear phase. We show that one of the novel filterbanks achieves noticeably better rate-distortion performance than several state-of-the-art filterbanks in image coding. Dong Wei 0003, Hung-Ta Pai, Alan C. Bovik |
ICIP (2) | 3 |
| 1998 | FOVEA: A Foveated Vergent Active Stereo System for Dynamic Three-Dimensional Scene RecoveryabstractWe introduce FOVEA: a Foveated Vergent Active stereo vision system. FOVEA actively directs a pair of vergent stereo cameras to fixate on surfaces in a scene, performing multiresolution surface depth recovery at each fixation point, and accumulating and integrating a multiresolution map of surface depth over multiple successive fixations. Several features of the system are novel: a foveated image sampling and processing strategy is shown to greatly simplify the problem of establishing; a probabilistic fixation strategy is developed that is driven by the scene structure; the system uses the fixation strategy to recover local depth maps at a high resolution at multiple fixation points, eventually mapping the entire scene; and finally, the local maps are integrated as they are acquired into a global depth map. William N. Klarquist, Alan C. Bovik |
ICRA | 2 |
| 1998 | Enhancement of Compressed Images by Optimal Shift-Invariant Wavelet Packet Basis
Dong Wei 0003, Alan C. Bovik |
J. Vis. Commun. Image Represent. | 2 |
| 1998 | On the instantaneous frequencies of multicomponent AM-FM signalsabstractWe study the instantaneous frequencies (IFs) of multicomponent AM-FM signals by extending the work on the two-component case by Loughlin and Tacer (see ibid.,, vol. 4, no.5, p.123-25, 1997) to the more general M-component case. A novel necessary and sufficient condition for the valid interpretation of the IF as a nonnegatively weighted average of the IFs of the components is proposed, which we apply as a method of interpreting the IFs of signals having no more than three dominant components at each time. Our quantitative study shows that in general, the IF-based monocomponent AM-FM decomposition is not appropriate for the modeling and analysis of multicomponent AM-FM signals unless all the components are well separated in the time domain. Dong Wei 0003, Alan C. Bovik |
IEEE Signal Process. Lett. | 2 |
| 1998 | Nonlinear image estimation using piecewise and local image modelsabstractWe introduce a new approach to image estimation based on a flexible constraint framework that encapsulates meaningful structural image assumptions. Piecewise image models (PIMs) and local image models (LIMs) are defined and utilized to estimate noise-corrupted images, PIMs and LIMs are defined by image sets obeying certain piecewise or local image properties, such as piecewise linearity, or local monotonicity. By optimizing local image characteristics imposed by the models, image estimates are produced with respect to the characteristic sets defined by the models. Thus, we propose a new general formulation for nonlinear set-theoretic image estimation. Detailed image estimation algorithms and examples are given using two PIMs: piecewise constant (PICO) and piecewise linear (PILI) models, and two LIMs: locally monotonic (LOMO) and locally convex/concave (LOCO) models. These models define properties that hold over local image neighborhoods, and the corresponding image estimates may be inexpensively computed by iterative optimization algorithms. Forcing the model constraints to hold at every image coordinate of the solution defines a nonlinear regression problem that is generally nonconvex and combinatorial. However, approximate solutions may be computed in reasonable time using the novel generalized deterministic annealing (GDA) optimization technique, which is particularly well suited for locally constrained problems of this type. Results are given for corrupted imagery with signal-to-noise ratio (SNR) as low as 2 dB, demonstrating high quality image estimation as measured by local feature integrity, and improvement in SNR. Scott T. Acton, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 1998 | Multiresolution 3-D range segmentation using focus cuesabstractThis paper describes a novel system for computing a three-dimensional (3-D) range segmentation of an arbitrary visible scene using focus information. The process of range segmentation is divided into three steps: an initial range classification, a surface merging process, and a 3-D multiresolution range segmentation. First, range classification is performed to obtain quantized range estimates. The range classification is performed by analyzing focus cues within a Bayesian estimation framework. A combined energy functional measures the degree of focus and the Gibbs distribution of the class field. The range classification provides an initial range segmentation. Second, a statistical merging process is performed to merge the initial surface segments. This gives a range segmentation at a coarse resolution. Third, 3-D multiresolution range segmentation (3-D MRS) is performed to refine the range segmentation into finer resolutions. The proposed range segmentation method does not require initial depth estimates, it allows the analysis of scenes containing multiple objects, and it provides a rich description of the 3-D structure of a scene. Changhoon Yim, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 1998 | FOVEA: a foveated vergent active stereo vision system for dynamic three-dimensional scene recoveryabstractWe introduce FOVEA: a foveated vergent active stereo vision system. FOVEA actively directs a pair of vergent stereo cameras to fixate on surfaces in a scene, performing multiresolution surface depth recovery at each fixation point, and accumulating and integrating a multiresolution map of surface depth over multiple successive fixations. Several features of the system are novel: an active foveated image sampling and processing strategy is shown to greatly simplify the problem of establishing correspondence; a probabilistic fixation strategy is developed that is driven by the scene structure; the system uses the fixation strategy to recover local depth maps at a high resolution at multiple fixation points, eventually mapping the entire scene; and finally, the local maps are integrated as they are acquired into a global depth map. William N. Klarquist, Alan C. Bovik |
IEEE Trans. Robotics Autom. | 2 |
| 1997 | The Analytic ImageabstractWe introduce a novel directional multidimensional Hilbert transform and use it to define the complex-valued analytic image associated with a real-valued image. The analytic image associates a unique pair of instantaneous amplitude and frequency functions with an image, and also admits many of the other important properties of the one-dimensional analytic signal. Joseph P. Havlicek, John W. Havlicek, Alan C. Bovik |
ICIP (2) | 3 |
| 1997 | Digital Halftoning as 2-D Delta-Sigma ModulationabstractThe error diffusion algorithm for digital halftoning is equivalent in form to a noise-shaping feedback coder, a class of delta-sigma modulator. The white noise assumption of the quantizer error is known to be false; in fact, the quantizer error is seen to be highly correlated with the input image. To account for this correlation, we use a gain model for the quantizer. This model accurately predicts the edge sharpening and noise shaping caused by all error diffusion schemes. It also permits an extension of error diffusion to oversampled imagery. Thomas D. Kite, Brian L. Evans, T. L. Sculley, Alan C. Bovik |
ICIP (1) | 4 |
| 1997 | Biorthogonal Quincunx Coifman WaveletsabstractWe define and construct a new family of compactly supported, nonseparable two-dimensional wavelets, "biorthogonal quincunx Coifman wavelets" (BQCWs), from their one-dimensional counterparts using the McClellan transformation. The resulting filter banks possess many interesting properties such as perfect reconstruction, vanishing moments, symmetry, diamond-shaped passbands, and dyadic fractional filter coefficients. We derive explicit formulas for the frequency responses of these filter banks. Both the analysis and synthesis lowpass filters converge to an ideal diamond-shaped halfband lowpass filter as the order of the corresponding BQCW system tends to infinity. Hence, they are promising in image and multidimensional signal processing applications. In addition, the synthesis scaling function in a BQCW system of any order is interpolating (or cardinal), which has been known as a desired merit in numerical analysis. Dong Wei 0003, Brian L. Evans, Alan C. Bovik |
ICIP (2) | 3 |
| 1997 | Adaptive variable baseline stereo for vergence controlabstractActive control of camera positioning is an important tool in the recovery of scene information. In this paper an important innovation is introduced which uses the local spatial frequency content of the acquired images to adaptively control stereo baseline, improving both the reliability and speed of the depth recovery process. This technique was designed to support the operation of an active vergent stereo imaging system by providing a coarse map of depth for camera fixation, providing an initial map of surface depth for accurate vergence. William N. Klarquist, Alan C. Bovik |
ICRA | 2 |