EDBT 2026 Demo / reviewers in the wild / expert
Sumohana S. Channappayya
dblp:49/904
· DBLP profile ↗
50ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-5687-0887ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training-free Adapter for Multi-Modal Image Matching for All-Day Visual Place RecognitionabstractVisual Place Recognition (VPR) identifies an image’s location by matching a query image of an unknown location against geotagged reference images. This has been a problem of interest for the computer vision community for many years. Consequently, many successful methods with impressive performance have been proposed in the literature. However, these works are primarily for the visible spectrum, restricting recognition to RGB images. This study reveals the shortcomings of the popular VPR methods for handling multi-modal RGB-Thermal (RGB-T) retrieval. Additionally, we introduce a straightforward yet effective training-free aggregator that can be integrated with any backbone model. The proposed method, motivated by the rich structures in the images, attempts to capture the correlation between them. We call it Self-Neighbourhood Support Maps (SNSM). Simple to comprehend and implement, experiments on various RGB-T datasets reveal that SNSM performs better than widely used unsupervised techniques such as VLAD by a considerable margin. The source code at: https://github.com/lfovia/SNSM. Anuradha Uggi, Sumohana S. Channappayya |
ICASSP | 2 |
| 2025 | Fed-SMTDA: A Novel Framework for Federated Source-Free Multi-Target Domain Adaptation Using Feature Clustering and Adaptive AggregationabstractFederated Learning (FL) enables collaborative model training across decentralized devices while maintaining data privacy. However, in many real-world scenarios, accessing both source and labeled target data on these devices is not feasible due to privacy concerns, data heterogeneity, and resource constraints. This paper introduces a novel approach to address this challenge, termed Federated Source-Free Multi-Target Domain Adaptation (Fed-SMTDA). In this framework, client devices possess only unlabeled target data, while the server has a pre-trained model derived from source data. Our method leverages the intrinsic structure of the target domain at each client by clustering similar features, facilitating more effective feature extraction and assignment. To improve model performance, we incorporate an entropy regularization term to minimize class confusion, ensuring cleaner decision boundaries. Additionally, we introduce a dynamic aggregation strategy called Weight Adjustment (WA), where the server adjusts the weights assigned to client models during aggregation based on the observed generalization gap across clients. This adaptive approach improves the overall robustness and generalization of the federated model, enabling it to perform effectively across diverse and unlabeled target domains. Fed-SMTDA is evaluated on the OfficeHome and PACS datasets. Its performance is compared against two centralized baselines – one that has full access to labels across domains and is termed Oracle, and another where the model is trained only with the source domain labels, which we call Source-Only. Fed-SMTDA delivers consistently better performance than the Source-Only model, and its performance is upper-bounded by the Oracle model. Challapalli Phanindra Revanth, Sumohana S. Channappayya, C. Krishna Mohan |
IJCNN | 2 |
| 2024 | SPIDER: A Semi-Supervised Continual Learning-based Network Intrusion Detection SystemabstractNetwork intrusion detection (NID) aims to identify unusual network traffic patterns (distribution shifts) that require NID systems to evolve continuously. While prior art emphasizes fully supervised annotated data-intensive continual learning methods for NID, semi-supervised continual learning (SSCL) methods require only limited annotated data. However, the inherent class imbalance (CI) in network traffic can significantly impact the performance of SSCL approaches. Previous approaches to tackle CI issues require storing a subset of labeled training samples from all past tasks in the memory for an extended duration, potentially raising privacy concerns. The proposed Semisupervised Privacy-preserving Intrusion detection with Drift-aware continual LEaRning (SPIDER) is a novel method that combines gradient projection memory (GPM) with SSCL to handle CI effectively without the requirement to store labeled samples from all of the previous tasks. We assess SPIDER’s performance against baselines on six intrusion detection benchmarks formed over a short period and the Anoshift benchmark spanning ten years, which includes natural distribution shifts. Additionally, we validate our approach on standard continual learning image classification benchmarks known for frequent distribution shifts compared to NID benchmarks. SPIDER achieves comparable performance to fully supervised and semisupervised baseline methods, while requiring a maximum of 20% annotated data and reducing the total training time by 2X. Suresh Kumar Amalapuram, Tamma Bheemarjuna Reddy, Sumohana S. Channappayya |
INFOCOM | 3 |
| 2024 | MS-NetVLAD: Multi-Scale NetVLAD for Visual Place RecognitionabstractMany successful Visual Place Recognition (VPR) techniques operate in a contrastive learning framework using features extracted from a Convolutional Neural Network (CNN) backbone. Among these, the NetVLAD is a popular framework that transforms the classical Vector of Locally Aggregated Descriptors (VLAD) method into a modern data-driven model. Introducing learnability in VLAD has led to several variants of NetVLAD, such as Patch-NetVLAD. However, many of these use only thebottleneck featuresof the backbone model, ignoring the rest of the feature hierarchy. A few state-of-the-art models adoptcomplex architecturesto improve the quality of features. In this letter, we propose a simple extension to the NetVLAD that leverages the feature representations from intermediate layers of the CNN backbone in addition to the bottleneck features. We conduct extensive experiments to demonstrate the significance of these intermediate features for VPR. The proposed method, which we call Multi-Scale-NetVLAD (MS-NetVLAD), surpasses the successful NetVLAD and Patch-NetVLAD models by a significant margin. We demonstrate consistent performance improvements on large-scale VPR benchmarks, including Pittsburgh 30 k, Tokyo 24/7, Nordland, and MSLS. This improvement is attributed to the complementary multi-scale features employed by MS-NetVLAD. Importantly, this work reinforces the inherent strength of the NetVLAD framework for VPR. Further, MS-NetVLAD is shown to be competitive with state-of-the-art VPR models such as MixVPR and R2Former. Anuradha Uggi, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 2 |
| 2023 | An Automotive Radar Dataset For Object ClassificationabstractAutonomous and semi-autonomous navigation systems use multiple sensors for perception. Amongst all the sensors, the most prominently used are camera, lidar and radar. In addition to being more expensive than radar, cameras and lidar fail to operate smoothly in adverse weather conditions. Radar, however, can work in various temperatures and weather conditions. Given radar’s capabilities and recent advancements in deep learning, can a low-cost and robust perception solution be achieved? To address this question, we make the following contributions via this work. We present a novel 77 GHz automotive radar dataset of static and moving objects. We also propose using a novel 7 × 5 object representation frame-work for automotive radar data-based object classification. We design a lightweight CNN architecture to classify objects in the automotive radar scene and demonstrate that the proposed CNN delivers strong performance on our dataset. Additionally, we experiment with the convLSTM architecture to exploit temporal characteristics present in the radar data. Further, we evaluate the performance of standard machine learning algorithms on the proposed dataset. Finally, we show that our CNN can also perform well on an open-source automotive radar dataset. Our dataset and codes are available at this link Akshad Shyam, Kusum Komalavally, Monika Gautam, Vamshikrishna Kancharla, Vennela Gudisa, Virendra Patil, Aanandh Balasubramanian, Sumohana S. Channappayya |
ICASSP | 8 |
| 2023 | Selective Binarization based Architecture Design Methodology for Resource-constrained Computation of Deep Neural NetworksabstractIn this paper, we introduced a novel selective binarization based architecture design methodology for the compute-intensive deep neural networks (DNNs) implementation on memory-constrained platforms with insignificant compromise in accuracy. To demonstrate the advantages of our proposed architecture design methodology, we performed a detailed layer-wise performance analysis of a DNN with the help of metrics like memory savings and accuracies. Subsequently, we validated the proposed design methodology by implementing it on the resource-constrained AMD-Xilinx Kintex-7 field-programmable gate array (FPGA) chip for the considered DNN. A thorough analysis of our architecture design methodology results shows a significant memory savings of about 93% or compression of 15.7× with less than a 1% marginal reduction in accuracy. Ramesh Reddy Chandrapu, Dubacharla Gyaneshwar, Sumohana S. Channappayya, Amit Acharyya |
ISCAS | 3 |
| 2023 | Augmented Memory Replay-based Continual Learning Approaches for Network Intrusion DetectionabstractIntrusion detection is a form of anomalous activity detection in communication network traffic. Continual learning (CL) approaches to the intrusion detection task accumulate old knowledge while adapting to the latest threat knowledge. Previous works have shown the effectiveness of memory replay-based CL approaches for this task. In this work, we present two novel contributions to improve the performance of CL-based network intrusion detection in the context of class imbalance and scalability. First, we extend class balancing reservoir sampling (CBRS), a memory-based CL method, to address the problems of severe class imbalance for large datasets. Second, we propose a novel approach titled perturbation assistance for parameter approximation (PAPA) based on the Gaussian mixture model to reduce the number of \textit{virtual stochastic gradient descent (SGD) parameter} computations needed to discover maximally interfering samples for CL. We demonstrate that the proposed approaches perform remarkably better than the baselines on standard intrusion detection benchmarks created over shorter periods (KDDCUP'99, NSL-KDD, CICIDS-2017/2018, UNSW-NB15, and CTU-13) and a longer period with distribution shift (AnoShift). We also validated proposed approaches on standard continual learning benchmarks (SVHN, CIFAR-10/100, and CLEAR-10/100) and anomaly detection benchmarks (SMAP, SMD, and MSL). Further, the proposed PAPA approach significantly lowers the number of virtual SGD update operations, thus resulting in training time savings in the range of 12 to 40\% compared to the maximally interfered samples retrieval algorithm. Suresh Kumar Amalapuram, Sumohana S. Channappayya, Tamma Bheemarjuna Reddy |
NeurIPS | 2 |
| 2022 | No-Reference Video Quality Assessment Using Voxel-Wise fMRI Models of the Visual CortexabstractThe performance of the human visual system is very efficient in many visual tasks such as identifying visual scenes, anticipating future actions based on the past observations, assessing the quality of visual stimuli, etc. A significant amount of effort has been directed towards finding quality aware representations of natural videos to solve the quality prediction task. In this work we present a novel no reference video quality assessment (NR-VQA) algorithm based on the functional Magnetic Resonance Imaging (fMRI) Blood Oxygen Level Dependent (BOLD) signal prediction with voxel-wise encoding models of the human brain. The voxel encoding models are learnt using deep features extracted from the AlexNet model to predict the fMRI response to natural video stimuli. We show that the curvature in the predicted voxel response time series provides good quality discriminability, and forms an important feature for quality prediction. Further, we show that the proposed curvature features in combination with the spatial index, temporal index and NIQE features deliver acceptable performance on the Video Quality Assessment (VQA) task on both synthetic and authentic distortion data-sets. Naga Sailaja Mahankali, Mohan Raghavan, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 3 |
| 2022 | Completely Blind Quality Assessment of User Generated Video ContentabstractIn this work, we address the challenging problem of completely blind video quality assessment (BVQA) of user generated content (UGC). The challenge is twofold since the quality prediction model is oblivious of human opinion scores, and there are no well-defined distortion models for UGC content. Our solution is inspired by a recent computational neuroscience model which hypothesizes that the human visual system (HVS) transforms a natural video input to follow a straighter temporal trajectory in the perceptual domain. A bandpass filter based computational model of the lateral geniculate nucleus (LGN) and V1 regions of the HVS was used to validate the perceptual straightening hypothesis. We hypothesize that distortions in natural videos lead to loss in straightness (or increased curvature) in their transformed representations in the HVS. We provide extensive empirical evidence to validate our hypothesis. We quantify the loss in straightness as a measure of temporal quality, and show that this measure delivers acceptable quality prediction performance on its own. Further, the temporal quality measure is combined with a state-of-the-art blind spatial (image) quality metric to design a blind video quality predictor that we call STraightness Evaluation Metric (STEM). STEM is shown to deliver state-of-the-art performance over the class of BVQA algorithms on five UGC VQA datasets including KoNViD-1K, LIVE-Qualcomm, LIVE-VQC, CVD and YouTube-UGC. Importantly, our solution is completely blind i.e., training-free, generalizes very well, is explainable, has few tunable parameters, and is simple and easy to implement. Parimala Kancharla, Sumohana S. Channappayya |
IEEE Trans. Image Process. | 2 |
| 2021 | Video Quality Prediction Using Voxel-Wise fMRI Models of the Visual CortexabstractIn this work, we address the problem of full-reference video quality prediction. To address this problem, we rely on deep learning based spatio-temporal representations of natural videos. Specifically, we use feature representations derived from a per-voxel deep learning regression model. This model predicts the functional Magnetic Resonance Imaging (fMRI) responses of the visual cortical regions to natural video stimuli. We construct a rudimentary full-reference spatio-temporal quality feature that is simply the L1-norm of the error between the voxel model’s response to the reference and test video stimuli. This feature is shown to correlate well with subjective quality scores. Additionally, we rely on the Multi-Scale Structural Similarity (MS-SSIM) index as the spatial quality feature. We show that the combination of the proposed spatio-temporal feature and the spatial (MS-SSIM) feature delivers competitive performance for both Quality of Experience (QoE) prediction and Video Quality Assessment (VQA) tasks. This finding not only provides corroborative evidence to previous results based on electroencephalograph (EEG) signals on the role of the visual cortex in quality prediction but also opens up interesting directions for perceptually inspired design of objective video quality metrics. Naga Sailaja Mahankali, Sumohana S. Channappayya |
ICASSP | 2 |
| 2021 | Improving the Visual Quality of Video Frame Prediction Models Using the Perceptual Straightening HypothesisabstractWe present a simple and effective method to improve the visual quality of the predicted frames in a frame prediction model. A recent neuroscience study hypothesizes that the perceptual representations of a sequence of frames extracted from a natural video follow a straight temporal trajectory. The perceptual representations of a sequence of video frames are found using a computational model of the LGN and V1 areas of the human visual system. In this work, we leverage the strength of this perceptual straightening model to formulate a novel objective function for video frame prediction. In general, a frame prediction model takes past frames as input and predicts the future frame. We enforce the perceptual straightness constraint through adversarial training by introducing the proposed novel quality aware discriminator loss. Our quality aware discriminator imposes the linear relationship between the perceptual representation of the predicted frame and the perceptual representations of the past frames. Specifically, we claim that imposing a perceptual straightness constraint through the discriminator helps in predicting (i.e., generating) video frames that look more natural and therefore, having a higher perceptual quality. We demonstrate the effectiveness of our proposed objective function on two popular video datasets using two different frame prediction models. These experiments show that our solution is both consistent and stable, thereby allowing it to be integrated with other frame prediction models as well. Parimala Kancharla, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 2 |
| 2021 | Predicting Spatio-Temporal Entropic Differences for Robust No Reference Video Quality AssessmentabstractWe consider the problem of robust no reference (NR) video quality assessment (VQA) where the algorithms need to have good generalization performance when they are trained and tested on different datasets. We specifically address this question in the context of predicting video quality for compression and transmission applications. Motivated by the success of the spatio-temporal entropic differences video quality predictor in this context, we design a framework using convolutional neural networks to predict spatial and temporal entropic differences without the need for a reference or human opinion score. This approach enables our model to capture both spatial and temporal distortions effectively and allows for robust generalization. We evaluate our algorithms on a variety of datasets and show superior cross database performance when compared to state of the art NR VQA algorithms. Shankhanil Mitra, Rajiv Soundararajan, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 3 |
| 2020 | Lqaid: Localized Quality Aware Image Denoising Using Deep Convolutional Neural NetworksabstractIn this paper we propose the Localized Quality Aware Image Denoising (LQAID) technique for image denoising using deep convolutional neural networks (CNNs). LQAID relies on local quality estimates over global cues like noise standard deviation since the perceptual quality of a noisy image is typically spatially varying. Specifically, we use localized quality maps generated using DistNet, a spatial quality map estimation method. These quality maps are used to augment the noisy image and guide the denoising process. The augmented noisy image is denoised using a deep fully convolutional network (FCN) trained using mean square error (MSE) as the loss function. The proposed approach shows state-of-the-art performance both qualitatively and quantitatively on two vision datasets: TID 2008 and BSD500. We also show that the proposed approach possesses excellent generalization ability. Lastly, the proposed approach is completely blind since it neither requires information about the strength of the additive noise nor does it try to explicitly estimate it. Dendi Sathya Veera Reddy, Chander Dev, Narayan Kothari, Sumohana S. Channappayya |
ICASSP | 4 |
| 2020 | Streaming Video QoE Modeling and Prediction: A Long Short-Term Memory ApproachabstractDue to the rate adaptation in hypertext transfer protocol adaptive streaming, the video quality delivered to the client keeps varying with time depending on the end-to-end network conditions. Moreover, the varying network conditions could also lead to the video client running out of the playback content resulting in rebuffering events. These factors affect the user satisfaction and cause degradation of the user quality of experience (QoE). Hence, it is important to quantify the perceptual QoE of the streaming video users and to monitor the same in a continuous manner so that the QoE degradation can be minimized. However, the continuous evaluation of QoE is challenging as it is determined by complex dynamic interactions among the QoE influencing factors. Toward this end, we present long short-term memory (LSTM)-QoE, a recurrent neural network-based QoE prediction model using an LSTM network. The LSTM-QoE is a network of cascaded LSTM blocks to capture the nonlinearities and the complex temporal dependencies involved in the time-varying QoE. Based on an evaluation over several publicly available continuous QoE datasets, we demonstrate that the LSTM-QoE has the capability to model the QoE dynamics effectively. We compare the proposed model with the state-of-the-art QoE prediction models and show that it provides an excellent performance across these datasets. Furthermore, we discuss the state space perspective for the LSTM-QoE and show the efficacy of the state space modeling approaches for the QoE prediction. Nagabhushan Eswara, S. Ashique, Anand Panchbhai, Soumen Chakraborty, Hemanth P. Sethuram, Kiran Kuchi, Abhinav Kumar 0001, Sumohana S. Channappayya |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2020 | No-Reference Video Quality Assessment Using Natural Spatiotemporal Scene StatisticsabstractRobust spatiotemporal representations of natural videos have several applications including quality assessment, action recognition, object tracking etc. In this paper, we propose a video representation that is based on a parameterized statistical model for the spatiotemporal statistics of mean subtracted and contrast normalized (MSCN) coefficients of natural videos. Specifically, we propose an asymmetric generalized Gaussian distribution (AGGD) to model the statistics of MSCN coefficients of natural videos and their spatiotemporal Gabor bandpass filtered outputs. We then demonstrate that the AGGD model parameters serve as good representative features for distortion discrimination. Based on this observation, we propose a supervised learning approach using support vector regression (SVR) to address the no-reference video quality assessment (NRVQA) problem. The performance of the proposed algorithm is evaluated on publicly available video quality assessment (VQA) datasets with both traditional and in-capture/authentic distortions. We show that the proposed algorithm delivers competitive performance on traditional (synthetic) distortions and acceptable performance on authentic distortions. The code for our algorithm will be released at https://www.iith.ac.in/~lfovia/downloads.html. Dendi Sathya Veera Reddy, Sumohana S. Channappayya |
IEEE Trans. Image Process. | 2 |
| 2019 | Quality Aware Generative Adversarial NetworksabstractGenerative Adversarial Networks (GANs) have become a very popular tool for im- plicitly learning high-dimensional probability distributions. Several improvements have been made to the original GAN formulation to address some of its shortcom- ings like mode collapse, convergence issues, entanglement, poor visual quality etc. While a significant effort has been directed towards improving the visual quality of images generated by GANs, it is rather surprising that objective image quality metrics have neither been employed as cost functions nor as regularizers in GAN objective functions. In this work, we show how a distance metric that is a variant of the Structural SIMilarity (SSIM) index (a popular full-reference image quality assessment algorithm), and a novel quality aware discriminator gradient penalty function that is inspired by the Natural Image Quality Evaluator (NIQE, a popular no-reference image quality assessment algorithm) can each be used as excellent regularizers for GAN objective functions. Specifically, we demonstrate state-of- the-art performance using the Wasserstein GAN gradient penalty (WGAN-GP) framework over CIFAR-10, STL10 and CelebA datasets. Parimala Kancharla, Sumohana S. Channappayya |
NeurIPS | 2 |
| 2019 | Generating Image Distortion Maps Using Convolutional Autoencoders With Application to No Reference Image Quality AssessmentabstractWe present two contributions in this work: 1) a reference-free image distortion map generating algorithm for spatially localizing distortions in a natural scene; and 2) no reference image quality assessment (NRIQA) algorithms derived from the generated distortion map. We use a convolutional autoencoder (CAE) for distortion map generation. We rely on distortion maps generated by the SSIM image quality assessment algorithm as the “ground truth” for training the CAE. We train the CAE on a synthetically generated dataset composed of pristine images and their distorted versions. Specifically, the dataset was created by applying standard distortions such as JPEG compression, JP2K compression, additive white Gaussian noise, and blur to the pristine images. SSIM maps are then generated on a per distorted image basis for each of the distorted images in the dataset and are in turn used for training the CAE. We first qualitatively demonstrate the robustness of the proposed distortion map generation algorithm over several images with both traditional and authentic distortions. We also demonstrate the distortion map's effectiveness quantitatively on both standard distortions and authentic distortions by deriving three different NRIQA algorithms. We show that these NRIQA algorithms deliver competitive performance over traditional databases like LIVE Phase II, CSIQ, TID 2013, LIVE MD, and MDID 2013, and databases with authentic distortions like LIVE Wild and KonIQ-10K. In summary, the proposed method generates high-quality distortion maps that are used to design robust NRIQA algorithms. Furthermore, the CAE-based distortion maps generation method can easily be modified to work with other ground truth distortion maps. Dendi Sathya Veera Reddy, Chander Dev, Narayan Kothari, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 4 |
| 2019 | Study of Subjective Quality and Objective Blind Quality Prediction of Stereoscopic VideosabstractWe present a new subjective and objective study on full high-definition (HD) stereoscopic (3D or S3D) video quality. In subjective study, we constructed an S3D video dataset with 12 pristine and 288 test videos, and the test videos are generated by applying the H.264 and H.265 compression, blur and frame freeze artifacts. We also propose a no reference (NR) objective video quality assessment (QA) algorithm that relies on measurements of the statistical dependencies between the motion and disparity subband coefficients of S3D videos. Inspired by the Generalized Gaussian Distribution (GGD) approach in liu2011statistical, we model the joint statistical dependencies between the motion and disparity components as following a Bivariate Generalized Gaussian Distribution (BGGD). We estimate the BGGD model parameters (α,β) and the coherence measure (Ψ) from the eigenvalues of the sample covariance matrix (M) of the BGGD. In turn, we model the BGGD parameters of pristine S3D videos using a Multivariate Gaussian (MVG) distribution. The likelihood of a test video's MVG model parameters coming from the pristine MVG model is computed and shown to play a key role in the overall quality estimation. We also estimate the global motion content of each video by averaging the SSIM scores between pairs of successive video frames. To estimate the test S3D video's spatial quality, we apply the popular 2D NR unsupervised NIQE image QA model on a frame-by-frame basis on both views. The overall quality of a test S3D video is finally computed by pooling the test S3D video's likelihood estimates, global motion strength and spatial quality scores. The proposed algorithm, which is 'completely blind' (requiring no reference videos or training on subjective scores) is called the Motion and Disparity based 3D video quality evaluator (MoDi3D). We show that MoDi3D delivers competitive performance over a wide variety of datasets including the IRCCYN dataset, the WaterlooIVC Phase I dataset, the LFOVIA dataset and our proposed LFOVIAS3DPh2 S3D video dataset. Balasubramanyam Appina, Dendi Sathya Veera Reddy, K. Manasa, Sumohana S. Channappayya, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2018 | No-Reference Stereoscopic Video Quality Assessment Algorithm Using Joint Motion and Depth StatisticsabstractWe propose a supervised no-reference (NR) quality assessment algorithm for assessing the perceptual quality of natural stereoscopic (S3D) videos. We empirically model the joint statistics of motion and depth subband coefficients of an S3D video frame using a Bivaraite Generalized Gaussian Distribution (BGGD). We compute the BGGD model parameters (α, β) to estimate the statistical dependency strength and show the features are quality discriminative. We compute the popular 2D NR image quality assessment (IQA) model NIQE on a frame-by-frame basis for both views to estimate the spatial quality. The frame-level BGGD features and spatial features are consolidated and used with the corresponding S3D videos difference mean opinion score (DMOS) labels for supervised learning using support vector regression (SVR). The overall quality of an S3D video is computed by averaging the frame-level quality predictions of the constituent video frames. The proposed algorithm, dubbed Video QUality Evaluation using MOtion and DEpth Statistics (VQUEMODES) is shown to outperform the state-of-the-art methods when evaluated over the IRCCYN and LFOVIA S3D subjective quality assessment databases. Balasubramanyam Appina, Jalli Akshith, Shanmukha Srinivas Battula, Sumohana S. Channappayya |
ICIP | 4 |
| 2018 | Improving the Visual Quality of Generative Adversarial Network (GAN)-Generated Images Using the Multi-Scale Structural Similarity IndexabstractThis paper presents a simple yet effective method to improve the visual quality of Generative Adversarial Network (GAN) generated images. In typical GAN architectures, the discriminator block is designed mainly to capture the class-specific content from images without explicitly imposing constraints on the visual quality of the generated images. A key insight from the image quality assessment literature is that natural scenes possess a very unique local structural and (hence) statistical signature, and that distortions affect this signature. We translate this insight into a constraint on the loss function of the discriminator in the GAN architecture with the goal of improving the visual quality of the generated images. Specifically, this constraint is based on the Multi-scale Structural Similarity (MS-SSIM) index to guarantee local structural and statistical integrity. We train GAN s (Boundary Equilibrium GANs, to be precise,) using the proposed approach on popular face and car image databases and demonstrate the improvement relative to standard training approaches both visually and quantitatively. Parimala Kancharla, Sumohana S. Channappayya |
ICIP | 2 |
| 2018 | An Evaluation Metric for Object Detection Algorithms in Autonomous Navigation Systems and its Application to a Real-Time Alerting SystemabstractAn autonomous navigation system relies on a number of sensors including radar, LIDAR and a visible light camera for its operation. We focus our attention on the visible light camera in this work. Object detection is the key first step to processing the video input from the camera. Specifically, we address the problem of assessing the performance of object detection algorithms in hazardous driving conditions that an autonomous navigation system is expected to encounter in a realistic scenario. To this end, we propose a novel metric for quantifying the degradation in performance of an object detection algorithm under different weather conditions. Additionally' we introduce a real-time method to detect extreme variations in performance of the algorithm that can be used to issue an alert. We evaluate the performance of our metric and alerting system and demonstrate its utility using the YOLOv2 object detection algorithm trained on the KITTI and virtual KITTI dataset. Harshitha Machiraju, Sumohana S. Channappayya |
ICIP | 2 |
| 2018 | Optical Character Recognition (OCR) for Telugu: Database, Algorithm and ApplicationabstractTelugu is a Dravidian language spoken by more than 80 million people worldwide. The optical character recognition (OCR) of the Telugu script has wide ranging applications including education, health-care, administration etc. The beautiful Telugu script however is very different from Germanic scripts like English and German. This makes the use of transfer learning of Germanic OCR solutions to Telugu a non-trivial task. To address the challenge of OCR for Telugu, we make three contributions in this work: (i) a database of Telugu characters, (ii) a deep learning based OCR algorithm, and (iii) a client server solution for the online deployment of the algorithm. For the benefit of the Telugu people and the research community, our code has been made freely available at this link. Konkimalla Chandra Prakash, Y. M. Srikar, Gayam Trishal, Souraj Mandal, Sumohana S. Channappayya |
ICIP | 5 |
| 2018 | Modeling Continuous Video QoE Evolution: A State Space ApproachabstractA rapid increase in the video traffic together with an increasing demand for higher quality videos has put a significant load on content delivery networks in the recent years. Due to the relatively limited delivery infrastructure, the video users in HTTP streaming often encounter dynamically varying quality over time due to rate adaptation, while the delays in video packet arrivals result in rebuffering events. The user quality-of-experience (QoE) degrades and varies with time because of these factors. Thus, it is imperative to monitor the QoE continuously in order to minimize these degradations and deliver an optimized QoE to the users. Towards this end, we propose a nonlinear state space model for efficiently and effectively predicting the user QoE on a continuous time basis. The QoE prediction using the proposed approach relies on a state space that is defined by a set of carefully chosen time varying QoE determining features. An evaluation of the proposed approach conducted on two publicly available continuous QoE databases shows a superior QoE prediction performance over the state-of-the-art QoE modeling approaches. The evaluation results also demonstrate the efficacy of the selected features and the model order employed for predicting the QoE. Finally, we show that the proposed model is completely state controllable and observable, so that the potential of state space modeling approaches can be exploited for further improving QoE prediction. Nagabhushan Eswara, Hemanth P. Sethuram, Soumen Chakraborty, Kiran Kuchi, Abhinav Kumar 0001, Sumohana S. Channappayya |
ICME | 6 |
| 2018 | Effect of Primitive Features of Content on Perceived Quality of Light Field VisualizationabstractDue to recent advent of light field visualization, ac-quisition/creation, encoding, transmission, rendering and quality assessment of 3D light field content has gained momentum. In particular, large light field displays need content with large field of view, and with high spatial and angular quality. Accordingly, subjective and objective quality evaluation studies have been conducted to examine spatial, angular and spatio-angular aspects of light field visualization. Recently, the effect of various zooming levels of the displayed content, as well as regions of interest on Quality of Experience (QoE) has also been explored. However, there has been no systematic attempt to see how the features of the content itself affect the visualization quality. In this work, we attempt to examine the effects of some primitive features of the content on subjective QoE. The results are based on a subjective study conducted on a large light field display, offering virtually continuous horizontal parallax. Roopak R. Tamboli, Balasubramanyam Appina, Péter A. Kara, Maria G. Martini, Sumohana S. Channappayya, Soumya Jana |
QoMEX | 5 |
| 2018 | A High-angular-resolution Turntable Data-set for Experiments on Light Field Visualization QualityabstractIn this paper, we present a high-angular-resolution data-set created using a turntable arrangement. Seven distinct objects, positioned on an automated turntable, were captured from three camera positions for every half degree of rotation, generating 720 images for each camera position. For each object, the camera positions were registered to the coordinate system of the middle camera. Intrinsic parameters of the camera were also estimated. A data-set of this kind is instrumental for research in a variety of areas, such as light field visualization, manifold learning, visual quality assessment, evaluation of preferred object orientation etc. Due to the availability of three-view stereo, this data-set could be useful for studying view interpolation techniques. Roopak R. Tamboli, M. Shanmukh Reddy, Péter A. Kara, Maria G. Martini, Sumohana S. Channappayya, Soumya Jana |
QoMEX | 5 |
| 2018 | Novel Light Weight Compressed Data Aggregation using sparse measurements for IoT networks
Madapu Amarlingam, Pachamuthu Rajalakshmi, Sumohana S. Channappayya, C. S. Sastry 0001 |
J. Netw. Comput. Appl. | 4 |
| 2018 | Full-Reference 3-D Video Quality Assessment Using Scene Component Statistical DependenciesabstractIn this letter, we present a full-reference (FR) quality assessment algorithm to assess the perceptual quality of natural stereoscopic three-dimensional (S3-D) videos. Toward the end, we rely on an empirical model for the joint statistics of motion and depth subband coefficients of an S3-D video frame. Specifically, we use the recently proposed bivariate generalized Gaussian distribution (BGGD) model for the joint statistics. In this letter, we show that the coherence of the covariance matrix of the BGGD varies in proportion with the perceptual video quality. We compute the coherence scores from the eigenvalues of the covariance matrix to estimate the amount of directional dependency between the motion and depth components. To estimate the overall spatial quality score, we apply off-the-shelf 2-D FR image quality assessment metrics on a frame-by-frame basis on both the views and average the frame wise scores. Finally, we pool the coherence and spatial quality scores to derive the overall quality for the S3-D video. An evaluation of the proposed algorithm on the IRCCYN, WaterlooIVC Phase I, and LFOVIA S3D video databases demonstrates its robust performance. The proposed algorithm is called depth- and motion-based 3-D video quality evaluator. Balasubramanyam Appina, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 2 |
| 2018 | A Continuous QoE Evaluation Framework for Video Streaming Over HTTPabstractA continuous evaluation of the end user's quality-of-experience (QoE) is essential for efficient video streaming. This is crucial for networks with constrained resources that offer time-varying channel quality to its users. In hypertext transfer protocol-based video streaming, the QoE is measured by quantifying the perceptual impact of distortions caused by rate adaptation or interruptions in playback due to rebuffering events. The resulting impact on the QoE due to these distortions has been studied individually in the literature. However, the QoE is determined by an interplay of these distortions, and therefore necessitates a combined study of them. To the best of our knowledge, there is no publicly available database that studies these distortions jointly on a continuous time basis. In this paper, our contributions are twofold. First, we present a database consisting of videos at full high definition and ultrahigh definition resolutions. We consider various levels of rate adaptation and rebuffering distortions together in these videos as experienced in a typical realistic setting. A subjective evaluation of these videos is conducted on a continuous time scale. Second, we present a QoE evaluation framework comprising a learning-based model during playback and an exponential model during rebuffering. Furthermore, we perform an objective evaluation of popular video quality assessment and continuous time QoE metrics over the constructed database. The objective evaluation study demonstrates that the performance of the proposed QoE model is superior to that of the objective metrics. The database is publicly available for download at http://www.iith.ac.in/~lfovia/downloads.html. Nagabhushan Eswara, K. Manasa, Avinash Kommineni, Soumen Chakraborty, Hemanth P. Sethuram, Kiran Kuchi, Abhinav Kumar 0001, Sumohana S. Channappayya |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2018 | Estimating Depth-Salient Edges and Its Application to Stereoscopic Image Quality AssessmentabstractThe human visual system pays attention to salient regions while perceiving an image. When viewing a stereoscopic 3-D (S3D) image, we hypothesize that while most of the contribution to saliency is provided by the 2-D image, a small but significant contribution is provided by the depth component. Further, we claim that only a subset of image edges contribute to depth perception while viewing an S3D image. In this paper, we propose a systematic approach for depth saliency estimation, called salient edges with respect to depth perception (SED) which localizes the depth-salient edges in an S3D image. We demonstrate the utility of SED in full reference stereoscopic image quality assessment. We consider gradient magnitude and inter-gradient maps for predicting structural similarity. A coarse quality map is estimated first by comparing the 2-D saliency and gradient maps of reference and test stereo pairs. We average this quality map to estimate luminance quality and refine this quality map using SED maps for evaluating depth quality. Finally, we combine this luminance and depth quality to obtain an overall stereo image quality. We perform a comprehensive evaluation of our metric on seven publicly available S3D IQA databases. The proposed metric shows competitive performance on all seven databases with state-of-the-art performance on three of them. Sameeulla Khan Md, Sumohana S. Channappayya |
IEEE Trans. Image Process. | 2 |
| 2017 | A full reference stereoscopic video quality assessment metricabstractWe propose a full reference stereo video quality assessment algorithm for assessing the perceptual quality of natural stereo videos. We exploit the separable representation of motion and binocular disparity in the visual cortex and develop a four stage algorithm to measure the quality of a stereoscopic video called FLOSIM3D. First, we compute the temporal features by utilizing an existing 2D VQA metric which measures the temporal annoyance based on patch level statistics such as mean, variance and minimum eigen value and pools them with a frame categorization based non-linear pooling strategy. Second, a structure based 2D Image Quality Assessment (IQA) metric is used to compute the spatial quality of the frames. Next, the loss in depth cues is measured using a structure based metric. Finally, the features for each of the stereo views are pooled to obtain the final stereo video quality score. We demonstrate the state-of-the-art performance of the proposed metric on IRCCYN dataset involving H.264, JP2K compression artifacts. Balasubramanyam Appina, K. Manasa, Sumohana S. Channappayya |
ICASSP | 3 |
| 2017 | No-reference quality assessment of tone mapped High Dynamic Range (HDR) images using transfer learningabstractWe present a transfer learning framework for no-reference image quality assessment (NRIQA) of tonemapped High Dynamic Range (HDR) images. This work is motivated by the observation that quality assessment databases in general, and HDR image databases in particular are “small” relative to the typical requirements for training deep neural networks. Transfer learning based approaches have been successful in such scenarios where learning from a related but larger database is transferred to the smaller database. Specifically, we propose a framework where the successful AlexNet is used to extract image features. This is followed by the application of Principal Component Analysis (PCA) to reduce the dimensionality of the feature vector (from 4096 to 400), given the small database size. A linear regression model is then fit to Mean Opinion Scores (MOS) using L2 regularization to prevent overfitting. We demonstrate state-of-the-art performance of the proposed approach on the ESPL-LIVE database. Abhinau Kumar Venkataramanan, Shashank Gupta 0001, Sai Sheetal Chandra, Shanmuganathan Raman, Sumohana S. Channappayya |
QoMEX | 5 |
| 2016 | A simple and accurate matrix for model based photoacoustic imagingabstractAccurate model-based methods in Photo-Acoustic Tomography (PAT) can reconstruct the image from insufficient and inaccurate measurements. Most of the models either make the simplified assumption of spherical averaging or use accurate models that have computationally burdensome implementations. We present a simple and accurate measurement matrix that is derived from the pseudo-spectral PAT model. The accuracy of the measurement matrix is first validated against the experimental PAT signal. We also compare the model against the standard k-wave measurement model and the spherical averaging model. We then highlight several reconstruction strategies based on the nature of the region of interest to further demonstrate the accuracy of the proposed measurement matrix. Kalloor Joseph Francis, Pachamuthu Rajalakshmi, Sumohana S. Channappayya, Ashutosh Richhariya |
HealthCom | 4 |
| 2016 | An optical flow-based no-reference video quality assessment algorithmabstractWe present an optical flow-based no-reference video quality assessment (NR-VQA) algorithm for assessing the perceptual quality of natural videos. Our algorithm is based on the hypothesis that distortions affect flow statistics both locally and globally. To capture the effects of distortion on optical flow, we measure irregularities at the patch level and at the frame level. At the patch level, we measure intra- and inter-patch level irregularities in the flow magnitude's variance and mean. We also measure the correlation in the patch level flow randomness between successive frames. At the frame level, we measure the normalized mean flow magnitude difference between successive frames. We rely on the robust NIQE algorithm for no-reference spatial quality assessment of the frames. These temporal and spatial features are averaged over all the frames to arrive at a video level feature vector. The video level features and the corresponding DMOS scores are used to train a support vector machine for regression (SVR). This machine is used to estimate the quality score of a test video. The competence of the proposed method is clearly demonstrated on SD and HD video databases that include common distortion types such as compression artifacts, packet loss artifacts, additive noise, and blur. K. Manasa, Sumohana S. Channappayya |
ICIP | 2 |
| 2016 | eTVSQ based video rate adaptation in cellular networks with α-fair resource allocationabstractDue to proliferation of mobile devices, the demand for video in cellular networks has increased exorbitantly. However, cellular networks have limited resources and the wireless medium is time-varying in nature. This necessitates the video streaming protocols to be re-designed taking into account the overall quality of experience (QoE) of the end users. In this paper, we propose a metric called enhanced-time varying subjective quality (eTVSQ) to measure the QoE of the video users. The eTVSQ accounts for time variation in QoE due to both rate adaption in HTTP streaming and playback interruption caused by rebuffering events. Based on this metric, we propose a rate adaptation strategy for HTTP video streaming in the downlink of cellular networks with α-fair resource allocation. The proposed method results in significant performance gains over the traditional throughput based rate adaptation strategy. Nagabhushan Eswara, Sumohana S. Channappayya, Abhinav Kumar 0001, Kiran Kuchi |
WCNC | 2 |
| 2016 | No-reference Stereoscopic Image Quality Assessment Using Natural Scene Statistics
Balasubramanyam Appina, Sameeulla Khan Md, Sumohana S. Channappayya |
Signal Process. Image Commun. | 3 |
| 2016 | Super-multiview content with high angular resolution: 3D quality assessment on horizontal-parallax lightfield display
Roopak R. Tamboli, Balasubramanyam Appina, Sumohana S. Channappayya, Soumya Jana |
Signal Process. Image Commun. | 3 |
| 2016 | An Optical Flow-Based Full Reference Video Quality Assessment AlgorithmabstractWe present a simple yet effective optical flow-based full-reference video quality assessment (FR-VQA) algorithm for assessing the perceptual quality of natural videos. Our algorithm is based on the premise that local optical flow statistics are affected by distortions and the deviation from pristine flow statistics is proportional to the amount of distortion. We characterize the local flow statistics using the mean, the standard deviation, the coefficient of variation (CV), and the minimum eigenvalue ( λ min ) of the local flow patches. Temporal distortion is estimated as the change in the CV of the distorted flow with respect to the reference flow, and the correlation between λ min of the reference and of the distorted patches. We rely on the robust multi-scale structural similarity index for spatial quality estimation. The computed temporal and spatial distortions, thus, are then pooled using a perceptually motivated heuristic to generate a spatio-temporal quality score. The proposed method is shown to be competitive with the state-of-the-art when evaluated on the LIVE SD database, the EPFL Polimi SD database, and the LIVE Mobile HD database. The distortions considered in these databases include those due to compression, packet-loss, wireless channel errors, and rate-adaptation. Our algorithm is flexible enough to allow for any robust FR spatial distortion metric for spatial distortion estimation. In addition, the proposed method is not only parameter-free but also independent of the choice of the optical flow algorithm. Finally, we show that the replacement of the optical flow vectors in our proposed method with the much coarser block motion vectors also results in an acceptable FR-VQA algorithm. Our algorithm is called the flow similarity index. K. Manasa, Sumohana S. Channappayya |
IEEE Trans. Image Process. | 2 |
| 2015 | Distributed compressed sensing for photo-acoustic imagingabstractPhoto-Acoustic Tomography (PAT) combines ultrasound resolution and penetration with endogenous optical contrast of tissue. Real-time PAT imaging is limited by the number of parallel data acquisition channels and pulse repetition rate of the laser. Typical photoacoustic signals afford sparse representation. Additionally, PAT transducer configurations exhibit significant intra- and inter- signal correlation. In this work, we formulate photoacoustic signal recovery in the Distributed Compressed Sensing (DCS) framework to exploit this correlation. Reconstruction using the proposed method achieves better image quality than compressed sensing with significantly fewer samples. Through our results, we demonstrate that DCS has the potential to achieve real-time PAT imaging. Kalloor Joseph Francis, Pachamuthu Rajalakshmi, Sumohana S. Channappayya |
ICIP | 3 |
| 2015 | Full-Reference Stereo Image Quality Assessment Using Natural Stereo Scene StatisticsabstractEmpirical studies of the joint statistics of luminance and disparity images (or wavelet coefficients) of natural stereoscopic scenes have resulted in two important findings: the marginal statistics are modelled well by the generalized Gaussian distribution (GGD) and there exists significant correlation between them. Inspired by these findings, we propose a full-reference image quality assessment algorithm dubbed STeReoscopic Image Quality Evaluator (STRIQE). We show that the parameters of the GGD fits of luminance wavelet coefficients along with correlation values form excellent features. Importantly, we demonstrate that the use of disparity information (via correlation) results in a consistent improvement in the performance of the algorithm. The performance of our algorithm is evaluated over popular datasets and shown to be competitive with the state-of-the-art full-reference algorithms. The efficacy of the algorithm is further highlighted by its near-linear relation with subjective scores, low root mean squared error (RMSE), and consistently good performance over both symmetric and asymmetric distortions. Sameeulla Khan Md, Balasubramanyam Appina, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 3 |
| 2014 | A statistical evaluation of Sparsity-based Distance Measure (SDM) as an image quality assessment algorithmabstractSparsity-based Distance Measure (SDM), a sparse reconstruction-based image similarity measure was recently proposed and shown to have promising applications in image classification, clustering and retrieval. In this paper, we present a statistical evaluation of SDM's performance as an image quality assessment (IQA) algorithm. This evaluation is carried out on the LIVE image database. We show that the SDM performs fairly in comparison with the state-of-the-art while possessing several attractive properties. Specifically, we demonstrate its robustness to rotation (90°, 180°), scaling, and combinations of distortions - properties that are highly desirable of any IQA algorithm. K. V. S. N. L. Manasa Priya, K. Manasa, Sumohana S. Channappayya |
ICASSP | 3 |
| 2008 | Rate Bounds on SSIM Index of Quantized Image DCT CoefficientsabstractIn this paper, we derive bounds on the structural similarity (SSIM) index as a function of quantization rate for fixed-rate uniform quantization of image discrete cosine transform (DCT) coefficients under the high rate assumption. The space domain SSIM index is first expressed in terms of the DCT coefficients of the space domain vectors. The transform domain SSIM Index is then used to derive bounds on the average SSIM index as a function of quantization rate for Gaussian and Laplacian sources. As an illustrative example, uniform quantization of the DCT coefficients of natural images is considered. We show that the SSIM index between the reference and quantized images fall within the bounds for a large set of natural images. Further, we show using a simple example that the proposed bounds could be very useful for rate allocation problems in practical image and video coding applications. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr., Constantine Caramanis |
DCC | 1 |
| 2008 | SSIM-optimal linear image restorationabstractIn this paper, we present an algorithm for designing a linear equalizer that is optimal with respect to the structural similarity (SSIM) index. The optimization problem is shown to be non-convex, thereby making it non-trivial. The non-convex problem is first converted to a quasi-convex problem and then solved using a combination of first order necessary conditions and bisection search. To demonstrate the usefulness of this solution, it is applied to image denoising and image restoration examples. We show using these examples that optimizing equalizers for the SSIM index does indeed result in higher perceptual image quality compared to equalizers optimized for the ubiquitous mean squared error (MSE). Sumohana S. Channappayya, Alan C. Bovik, Constantine Caramanis, Robert W. Heath Jr. |
ICASSP | 1 |
| 2008 | Perceptual soft thresholding using the structural similarity indexabstractIn this paper, we present a novel algorithm for wavelet domain image denoising using the soft thresholding function. The thresholds are designed to be locally optimal with respect to the structural similarity (SSIM) index. The SSIM Index is first expressed in terms of wavelet transform coefficients of orthogonal wavelet transforms. The wavelet domain representation of the SSIM Index, along with the assumption of a Gaussian prior for the wavelet coefficients is used to formulate the soft thresholding optimization problem. A locally optimal solution is found using a quasi-Newton approach. This solution is applied to denoise images in the wavelet domain. The visual quality of the images denoised using the proposed algorithm is shown to be higher compared to the MSE-optimal soft thresholding denoising solution, as measured by the SSIM Index. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr. |
ICIP | 1 |
| 2008 | Design of Linear Equalizers Optimized for the Structural Similarity IndexabstractWe propose an algorithm for designing linear equalizers that maximize the structural similarity (SSIM) index between the reference and restored signals. The SSIM index has enjoyed considerable application in the evaluation of image processing algorithms. Algorithms, however, have not been designed yet to explicitly optimize for this measure. The design of such an algorithm is nontrivial due to the nonconvex nature of the distortion measure. In this paper, we reformulate the nonconvex problem as a quasi-convex optimization problem, which admits a tractable solution. We compute the optimal solution in near closed form, with complexity of the resulting algorithm comparable to complexity of the linear minimum mean squared error (MMSE) solution, independent of the number of filter taps. To demonstrate the usefulness of the proposed algorithm, it is applied to restore images that have been blurred and corrupted with additive white gaussian noise. As a special case, we consider blur-free image denoising. In each case, its performance is compared to a locally adaptive linear MSE-optimal filter. We show that the images denoised and restored using the SSIM-optimal filter have higher SSIM index, and superior perceptual quality than those restored using the MSE-optimal adaptive linear filter. Through these results, we demonstrate that a) designing image processing algorithms, and, in particular, denoising and restoration-type algorithms, can yield significant gains over existing (in particular, linear MMSE-based) algorithms by optimizing them for perceptual distortion measures, and b) these gains may be obtained without significant increase in the computational complexity of the algorithm. Sumohana S. Channappayya, Alan C. Bovik, Constantine Caramanis, Robert W. Heath Jr. |
IEEE Trans. Image Process. | 1 |
| 2008 | Rate Bounds on SSIM Index of Quantized ImagesabstractIn this paper, we derive bounds on the structural similarity (SSIM) index as a function of quantization rate for fixed-rate uniform quantization of image discrete cosine transform (DCT) coefficients under the high-rate assumption. The space domain SSIM index is first expressed in terms of the DCT coefficients of the space domain vectors. The transform domain SSIM index is then used to derive bounds on the average SSIM index as a function of quantization rate for uniform, Gaussian, and Laplacian sources. As an illustrative example, uniform quantization of the DCT coefficients of natural images is considered. We show that the SSIM index between the reference and quantized images fall within the bounds for a large set of natural images. Further, we show using a simple example that the proposed bounds could be very useful for rate allocation problems in practical image and video coding applications. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr. |
IEEE Trans. Image Process. | 1 |
| 2006 | A Linear Estimator Optimized for the Structural Similarity Index and its Application to Image DenoisingabstractWe use a perceptual distortion metric-the structural similarity (SSIM) index, to derive a new linear estimator for estimating zero-mean Gaussian sources distorted by additive white Gaussian noise (AWGN). We use this estimator in an image denoising application and compare its performance with the traditional linear least squared error (LLSE) estimator. Although images denoised using the SSIM-optimized estimator have a lower peak signal-to-noise ratio (PSNR) compared to their LLSE counterparts, the SSIM-optimized estimator clearly outperforms the LLSE estimator in terms of the visual quality of the denoised images. Sumohana S. Channappayya, Alan C. Bovik, Robert W. Heath Jr. |
ICIP | 1 |
| 2005 | Multiple Description Image Coding Using Natural Scene StatisticsabstractThe statistics of natural scenes in the wavelet domain are accurately characterized by the Gaussian scale mixture (GSM) model. The model lends itself easily to analysis and many applications that use this model are emerging (e.g., denoising, watermark detection). We present an error-resilient image communications application that uses the GSM model and multiple description coding (MDC) to provide error-resilience. We derive a rate-distortion bound for GSM random variables, derive the redundancy rate-distortion function, and finally implement an MD image communication system. Sumohana S. Channappayya, Robert W. Heath Jr., Alan C. Bovik |
ICASSP (2) | 1 |
| 2005 | Frame based multiple description image coding in the wavelet domainabstractMultiple description codes generated by quantized frame expansions have been shown to perform well on erasure channels when compared to traditional channel codes. In this paper we propose a multiple description image coding scheme in the wavelet domain using quantized frame expansions. We form zerotrees from wavelet coefficients and apply a tight frame operator to the zerotrees. We then group appropriate expansions to form packets and evaluate the performance of the scheme over an erasure channel. We compare the performance of the proposed scheme with a conventional channel coding scheme. Sumohana S. Channappayya, Joonsoo Lee, Robert W. Heath Jr., Alan C. Bovik |
ICIP (3) | 1 |
| 2001 | Coding of digital imagery for transmission over multiple noisy channelsabstractThis paper presents a multiple description image coding scheme that facilitates the transmission of digital imagery over multiple noisy channels. The proposed scheme divides the image into smaller parts that are transmitted over the individual channels of an inverse multiplexing system. The division or splitting is done in such a fashion that it facilitates the interpolation of lost coefficients in the case of one or more channel failures. At the receiver, the image is reconstructed by proper assembly of the data received from each channel. In case of channel failure, the missing coefficients are estimated from the available data with the use of a novel post-processing scheme. For operation over four noisy channels with various bit error probabilities, we investigate the quantitative and subjective performance of the proposed system for the case of multiple channel failures. Sumohana S. Channappayya, Glen P. Abousleman, Lina J. Karam |
ICASSP | 1 |
| 2001 | Image coding for transmission over multiple noisy channels using punctured convolutional codes and trellis-coded quantizationabstractThis paper presents a multiple description image coding scheme that facilitates the transmission of digital imagery over multiple noisy channels. The proposed scheme divides the image into smaller parts that are transmitted over the individual channels of an inverse multiplexing system. The division or splitting is done in such a fashion that it facilitates the interpolation of lost coefficients in the case of one or more channel failures. To combat the effects of channel noise, a novel joint source-channel coding scheme is employed, which uses punctured convolutional channel codes and trellis-coded quantization. At the receiver, the image is reconstructed by proper assembly of the data received from each channel. In case of channel failure, the missing coefficients are estimated from the available data with the use of a novel post processing scheme. For operation over four noisy channels with various bit error probabilities, we investigate the quantitative and subjective performance with and without channel failures. Glen P. Abousleman, Sumohana S. Channappayya, Lina J. Karam |
ICIP (1) | 2 |