Rajiv Soundararajan

dblp:25/3993 · DBLP profile ↗
← Back
47ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0001-5767-5373ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Theory of computation · 4 · 2 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 RAM-VQA: Restoration Assisted Multi-Modality Video Quality Assessment
abstract
Video Quality Assessment (VQA) strives to computationally emulate human perceptual judgments and has garnered significant attention given its widespread applicability. However, existing methodologies face two primary impediments: (1) limited proficiency in evaluating samples at quality extremes (e.g., severely degraded or near-perfect videos), and (2) insufficient sensitivity to nuanced quality variations arising from a misalignment with human perceptual mechanisms. Although vision-language models offer promising semantic understanding, their reliance on visual encoders pre-trained for high-level tasks often compromises their sensitivity to low-level distortions. To surmount these challenges, we propose the Restoration-Assisted Multi-modality VQA (RAM-VQA) framework. Uniquely, our approach leverages video restoration as a proxy to explicitly model distortion-sensitive features. The framework operates through two synergistic stages: a prompt learning stage that constructs a quality-aware textual space using triple-level references (degraded, restored, and pristine) derived from the restoration process, and a dual-branch evaluation stage that integrates semantic cues with technical quality indicators via spatio-temporal differential analysis. Extensive experiments demonstrate that RAM-VQA achieves state-of-the-art performance across diverse benchmarks, exhibiting superior capability in handling extreme-quality content while ensuring robust generalization.
Pengfei Chen 0003, Jiebin Yan, Rajiv Soundararajan, Giuseppe Valenzise, Leida Li
IEEE Trans. Image Process.3
2026 Simple-RF: Regularizing Sparse Input Radiance Fields with Simpler Solutions
abstract
Neural Radiance Fields (NeRF) show impressive performances in photo-realistic free-view rendering of scenes. Recent improvements such as TensoRF and ZipNeRF employ explicit models for faster optimization and rendering. However, all these radiance fields require a dense sampling of images in the given scene for effective training. Their performances degrade significantly when only a sparse set of views is available. Existing depth priors used to supervise the radiance fields are either sparse or suffer from generalization issues. We seek to learn scene-specific dense depth priors to regularize the radiance fields. Further, we desire a framework of regularizations that can work across different radiance field models. We observe that certain features of the radiance fields, such as positional encoding, number of decomposed tensor components or size of the hash table, cause overfitting in the sparse-input scenario. We design augmented models by reducing the capacity of these features and train them along with the main radiance field. These augmented models learn simpler solutions, which estimate better depth in certain regions. By supervising the main radiance field with such depths, we significantly improve the performance of the radiance fields on popular forward-facing and 360ˆ datasets by employing the above regularization.
Nagabhushan Somraj, Sai Harsha Mupparaju, Adithyan Karanayil, Rajiv Soundararajan
ACM Trans. Graph.4
2025 Vision-Language Model Guided Semi-supervised Learning for No-Reference Video Quality Assessment
abstract
Perceptual assessment of user-generated content videos is an important problem that impacts viewing experience of millions of users. Current no-reference video quality assessment (NR-VQA) algorithms require a large amount of human annotated videos. In this work, we address this problem by specifically designing a dual-model based Semi-supervised Learning (SSL) method for NR-VQA. The first model is based on a popular vision-language model namely CLIP, where we adapt the visual encoder to capture high-level semantic quality information through quality-relevant text prompts. A second model learns complementary low-level spatio-temporal quality using a 3D vision transformer and video fragments. We enable intelligent knowledge transfer between the high-level vision-language and low-level vision-transformer model to pseudo-label the unlabelled videos. Our unified model outperforms existing state-of-the-art SSL methods for VQA across popular VQA databases including inter-database settings.
Shankhanil Mitra, Rajiv Soundararajan
ICASSP2
2025 Bidirectional Flow Fields for Sparse Input Novel View Synthesis of Dynamic Scenes
Kapil Choudhary, Nagabhushan Somraj, Rajiv Soundararajan
ICIP3
2025 AD-GS: Alternating Densification for Sparse-Input 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has shown impressive results in real-time novel view synthesis. However, it often struggles under sparse-view settings, producing undesirable artifacts such as floaters, inaccurate geometry, and overfitting due to limited observations. We find that a key contributing factor is uncontrolled densification, where adding Gaussian primitives rapidly without guidance can harm geometry and cause artifacts. We propose AD-GS, a novel alternating densification framework that interleaves high and low densification phases. During high densification, the model densifies aggressively, followed by photometric loss based training to capture fine-grained scene details. Low densification then primarily involves aggressive opacity pruning of Gaussians followed by regularizing their geometry through pseudo-view consistency and edge-aware depth smoothness. This alternating approach helps reduce overfitting by carefully controlling model capacity growth while progressively refining the scene representation. Extensive experiments on challenging datasets demonstrate that AD-GS significantly improves rendering quality and geometric consistency compared to existing methods. The source code for our model can be found on our project page: https://gurutvapatle.github.io/publications/2025/ADGS.html.
Gurutva Patle, Nilay Girgaonkar, Nagabhushan Somraj, Rajiv Soundararajan
SIGGRAPH Asia4
2025 UnDIVE: Generalized Underwater Video Enhancement Using Generative Priors
abstract
With the rise of marine exploration, underwater imaging has gained significant attention as a research topic. Under-water video enhancement has become crucial for real-time computer vision tasks in marine exploration. However, most existing methods focus on enhancing individual frames and neglect video temporal dynamics, leading to visually poor enhancements. Furthermore, the lack of ground-truth references limits the use of abundant available underwater video data in many applications. To address these issues, we propose a two-stage framework for enhancing underwater videos. The first stage uses a denoising diffusion probabilistic model to learn a generative prior from unlabeled data, capturing robust and descriptive feature representations. In the second stage, this prior is incorporated into a physics-based image formulation for spatial enhancement, while also enforcing temporal consistency between video frames. Our method enables real-time and computationally-efficient processing of high-resolution underwater videos at lower resolutions, and offers efficient enhancement in the presence of diverse water-types. Extensive experiments on four datasets show that our approach generalizes well and outperforms existing enhancement methods. Our code is available at github. com/suhas-srinath/undive.
Suhas Srinath, Aditya Chandrasekar, Hemang Jamadagni, Rajiv Soundararajan, Prathosh A. P.
WACV4
2025 Quality Assessment of Multispectral Remote Sensing Images: A Subjective Study and No Reference Algorithms
abstract
Multispectral remote sensing images suffer from a wide range of degradations due to noise, blur, haze, compression, striping, and other distortions. While there is rich literature in the restoration of such multipsectral images, the quality assessment (QA) of multispectral images has received much less attention. In addition, often a reference image is not available, which motivates the study of no reference (NR) image QA (IQA). While several algorithms for NR-IQA have been designed for natural images, their relevance for multispectral images has not been validated. The major impediment to achieving this is the lack of a subjective QA database for validation. We make three contributions in our work. We first design a multispectral IQA database, namely Multispectral Remote Sensing Image Database (MRSID), with images suffering from various distortions and conduct a subjective study to obtain pairwise image quality rankings. We then benchmark several NR-IQA methods developed for natural images on MRSID. We show that recent vision-language based models can yield excellent performance when fine-tuned on our dataset, especially when multispectral images are trained with images pooled from all bands, achieving an improvement of 15.8% in pairwise preference accuracy over single-band training. Finally, we investigate the reasons for such superior performance and show that improving the alignment of features across different spectral bands can further improve performance, yielding an additional 3.1% gain in pairwise preference accuracy on the cross-scene evaluation split, indicating improved generalization.
Aakanksha Bharti, Nithin C. Babu, M. Naveen Sai Kiran, Aditi Prasad, Neeraj Badal, Rajiv Soundararajan
IEEE Trans. Geosci. Remote. Sens.6
2024 Knowledge Guided Semi-supervised Learning for Quality Assessment of User Generated Videos
abstract
Perceptual quality assessment of user generated content (UGC) videos is challenging due to the requirement of large scale human annotated videos for training. In this work, we address this challenge by first designing a self-supervised Spatio-Temporal Visual Quality Representation Learning (ST-VQRL) framework to generate robust quality aware features for videos. Then, we propose a dual-model based Semi Supervised Learning (SSL) method specifically designed for the Video Quality Assessment (SSL-VQA) task, through a novel knowledge transfer of quality predictions between the two models. Our SSL-VQA method uses the ST-VQRL backbone to produce robust performances across various VQA datasets including cross-database settings, despite being learned with limited human annotated videos. Our model improves the state-of-the-art performance when trained only with limited data by around 10%, and by around 15% when unlabelled data is also used in SSL. Source codes and checkpoints are available at https://github.com/Shankhanil006/SSL-VQA.
Shankhanil Mitra, Rajiv Soundararajan
AAAI2
2024 Learning Generalizable Perceptual Representations for Data-Efficient No-Reference Image Quality Assessment
abstract
No-reference (NR) image quality assessment (IQA) is an important tool in enhancing the user experience in diverse visual applications. A major drawback of state-of-the-art NR-IQA techniques is their reliance on a large number of human annotations to train models for a target IQA application. To mitigate this requirement, there is a need for unsupervised learning of generalizable quality representations that capture diverse distortions. We enable the learning of low-level quality features agnostic to distortion types by introducing a novel quality-aware contrastive loss. Further, we leverage the generalizability of vision-language models by fine-tuning one such model to extract high-level image quality information through relevant text prompts. The two sets of features are combined to effectively predict quality by training a simple regressor with very few samples on a target dataset. Additionally, we design zero-shot quality predictions from both pathways in a completely blind setting. Our experiments on diverse datasets encompassing various distortions show the generalizability of the features and their superior performance in the data-efficient and zero-shot settings.
Suhas Srinath, Shankhanil Mitra, Shika Rao, Rajiv Soundararajan
WACV4
2024 Semi-Supervised Learning of Perceptual Video Quality by Generating Consistent Pairwise Pseudo-Ranks
abstract
Designing learning-based no-reference (NR) video quality assessment (VQA) algorithms for camera-captured videos is cumbersome due to the large number of human annotations of quality. In this work, we propose a semi-supervised learning (SSL) framework exploiting many unlabelled and very limited numbers of authentically distorted labelled videos. Our main contributions are twofold. Leveraging the benefits of consistency regularization and pseudo-labelling, our SSL model generates pairwise pseudo-ranks for the unlabelled videos using a student-teacher model on strong-weak augmented videos. We design the strong-weak augmentations to be quality invariant to use the unlabelled videos effectively in SSL. The generated pseudo-ranks are used along with the limited labels to train our SSL model. Our primary focus in SSL for NR VQA is to learn mapping from video feature representations to quality scores. We compare various feature extraction methods and show that our SSL framework can lead to improved performance on these features. We present a spatial and temporal feature extraction method based on predicting spatial and temporal entropic differences. We show that these features help achieve robust performance when trained with limited data, providing a better baseline to apply SSL. Extensive experiments on three popular VQA datasets demonstrate that the proposed semi-supervised VQA method improves on the performance of existing methods in terms of correlation with human opinion by approximately$15 \! - \! 20 \%$
Shankhanil Mitra, Saiyam Jogani, Rajiv Soundararajan
IEEE Trans. Multim.3
2023 Test Time Adaptation for Blind Image Quality Assessment
abstract
While the design of blind image quality assessment (IQA) algorithms has improved significantly, the distribution shift between the training and testing scenarios often leads to a poor performance of these methods at inference time. This motivates the study of test time adaptation (TTA) techniques to improve their performance at inference time. Existing auxiliary tasks and loss functions used for TTA may not be relevant for quality-aware adaptation of the pre-trained model. In this work, we introduce two novel quality-relevant auxiliary tasks at the batch and sample levels to enable TTA for blind IQA. In particular, we introduce a group contrastive loss at the batch level and a relative rank loss at the sample level to make the model quality aware and adapt to the target data. Our experiments reveal that even using a small batch of images from the test distribution helps achieve significant improvement in performance by updating the batch normalization statistics of the source model.
Subhadeep Roy, Shankhanil Mitra, Soma Biswas, Rajiv Soundararajan
ICCV4
2023 SimpleNeRF: Regularizing Sparse Input Neural Radiance Fields with Simpler Solutions
abstract
Neural Radiance Fields (NeRF) show impressive performance for the photo-realistic free-view rendering of scenes. However, NeRFs require dense sampling of images in the given scene, and their performance degrades significantly when only a sparse set of views are available. Researchers have found that supervising the depth estimated by the NeRF helps train it effectively with fewer views. The depth supervision is obtained either using classical approaches or neural networks pre-trained on a large dataset. While the former may provide only sparse supervision, the latter may suffer from generalization issues. As opposed to the earlier approaches, we seek to learn the depth supervision by designing augmented models and training them along with the NeRF. We design augmented models that encourage simpler solutions by exploring the role of positional encoding and view-dependent radiance in training the few-shot NeRF. The depth estimated by these simpler models is used to supervise the NeRF depth estimates. Since the augmented models can be inaccurate in certain regions, we design a mechanism to choose only reliable depth estimates for supervision. Finally, we add a consistency loss between the coarse and fine multi-layer perceptrons of the NeRF to ensure better utilization of hierarchical sampling. We achieve state-of-the-art view-synthesis performance on two popular datasets by employing the above regularizations. The source code for our model can be found on our project page: https://nagabhushansn95.github.io/publications/2023/SimpleNeRF.html
Nagabhushan Somraj, Adithyan Karanayil, Rajiv Soundararajan
SIGGRAPH Asia3
2023 No Reference Opinion Unaware Quality Assessment of Authentically Distorted Images
abstract
The quality assessment (QA) of camera captured authentically distorted images is important on account of its ubiquitous applications and challenging due to the lack of a reference. While there exists a plethora of supervised no reference (NR) image QA (IQA) algorithms, there is a need to study unsupervised or opinion unaware algorithms on account of their superior generalization performance. We explore self-supervised learning (SSL) for the feature design on authentically distorted images to predict quality without training on human labels. While SSL on synthetic distortions has recently shown promise, there is a need to enrich the feature learning on authentic distortions. The key challenge in achieving this is in the learning of quality sensitive features with mitigated content dependence. We design a self-supervised contrastive learning approach which only requires positives and introduce a content separation loss by estimating a bound on the mutual information between the features learnt and the content information. We show on multiple authentically distorted datasets that our self-supervised features can predict image quality by comparing with a corpus of pristine images and achieve state-of-the-art performance.§
Nithin C. Babu, Vignesh Kannan, Rajiv Soundararajan
WACV3
2023 Semi-Supervised Learning for Low-light Image Restoration through Quality Assisted Pseudo-Labeling
abstract
Convolutional neural networks have been successful in restoring images captured under poor illumination conditions. Nevertheless, such approaches require a large number of paired low-light and ground truth images for training. Thus, we study the problem of semi-supervised learning for low-light image restoration when limited low-light images have ground truth labels. Our main contributions in this work are twofold. We first deploy an ensemble of low-light restoration networks to restore the unlabeled images and generate a set of potential pseudo-labels. We model the contrast distortions in the labeled set to generate different sets of training data and create the ensemble of networks. We then design a contrastive self-supervised learning based image quality measure to obtain the pseudo-label among the images restored by the ensemble. We show that training the restoration network with the pseudo-labels allows us to achieve excellent restoration performance even with very few labeled pairs. We conduct extensive experiments on three popular low-light image restoration datasets to show the superior performance of our semi-supervised low-light image restoration compared to other approaches. Project page is available at https://github.com/sameerIISc/SSL-LLR.
Sameer Malik, Rajiv Soundararajan
WACV2
2022 Low Light Video Enhancement by Learning on Static Videos with Cross-Frame Attention
Shivam Chhirolya, Sameer Malik, Rajiv Soundararajan
BMVC3
2022 Temporal View Synthesis of Dynamic Scenes through 3D Object Motion Estimation with Multi-Plane Images
abstract
The challenge of graphically rendering high frame-rate videos on low compute devices can be addressed through periodic prediction of future frames to enhance the user experience in virtual reality applications. This is studied through the problem of temporal view synthesis (TVS), where the goal is to predict the next frames of a video given the previous frames and the head poses of the previous and the next frames. In this work, we consider the TVS of dynamic scenes in which both the user and objects are moving. We design a framework that decouples the motion into user and object motion to effectively use the available user motion while predicting the next frames. We predict the motion of objects by isolating and estimating the 3D object motion in the past frames and then extrapolating it. We employ multi-plane images (MPI) as a 3D representation of the scenes and model the object motion as the 3D displacement between the corresponding points in the MPI representation. In order to handle the sparsity in MPIs while estimating the motion, we incorporate partial convolutions and masked correlation layers to estimate corresponding points. The predicted object motion is then integrated with the given user or camera motion to generate the next frame. Using a disocclusion infilling module, we synthesize the regions uncovered due to the camera and object motion. We develop a new synthetic dataset for TVS of dynamic scenes consisting of 800 videos at full HD resolution. We show through experiments on our dataset and the MPI Sintel dataset that our model outperforms all the competing methods in the literature.
Nagabhushan Somraj, Pranali Sancheti, Rajiv Soundararajan
ISMAR3
2022 Multiview Contrastive Learning for Completely Blind Video Quality Assessment of User Generated Content
abstract
Completely blind video quality assessment (VQA) refers to a class of quality assessment methods that do not use any reference videos, human opinion scores or training videos from the target database to learn a quality model. The design of this class of methods is particularly important since it can allow for superior generalization in performance across various datasets. We consider the design of completely blind VQA for user generated content. While several deep feature extraction methods have been considered in supervised and weakly supervised settings, such approaches have not been studied in the context of completely blind VQA. We bridge this gap by presenting a self-supervised multiview contrastive learning framework to learn spatio-temporal quality representations. In particular, we capture the common information between frame differences and frames by treating them as a pair of views and similarly obtain the shared representations between frame differences and optical flow. The resulting features are then compared with a corpus of pristine natural video patches to predict the quality of the distorted video. Detailed experiments on multiple camera captured VQA datasets reveal the superior performance of our method over other features when evaluated without training on human scores. Code will be made available at https://github.com/Shankhanil006/VISION.
Shankhanil Mitra, Rajiv Soundararajan
ACM Multimedia2
2022 Revealing Disocclusions in Temporal View Synthesis through Infilling Vector Prediction
abstract
We consider the problem of temporal view synthesis, where the goal is to predict a future video frame from the past frames using knowledge of the depth and relative camera motion. In contrast to revealing the disoccluded regions through intensity based infilling, we study the idea of an infilling vector to infill by pointing to a non-disoccluded region in the synthesized view. To exploit the structure of disocclusions created by camera motion during their infilling, we rely on two important cues, temporal correlation of infilling directions and depth. We design a learning framework to predict the infilling vector by computing a temporal prior that reflects past infilling directions and a normalized depth map as input to the network. We conduct extensive experiments on a large scale dataset we build for evaluating temporal view synthesis in addition to the SceneNet RGB-D dataset. Our experiments demonstrate that our infilling vector prediction approach achieves superior quantitative and qualitative infilling performance compared to other approaches in literature.
Vijayalakshmi Kanchana, Nagabhushan Somraj, Suraj Yadwad, Rajiv Soundararajan
WACV4
2022 Understanding the perceived quality of video predictions
Nagabhushan Somraj, Manoj Surya Kashi, S. P. Arun, Rajiv Soundararajan
Signal Process. Image Commun.4
2021 A low light natural image statistical model for joint contrast enhancement and denoising
Sameer Malik, Rajiv Soundararajan
Signal Process. Image Commun.2
2021 Predicting Spatio-Temporal Entropic Differences for Robust No Reference Video Quality Assessment
abstract
We consider the problem of robust no reference (NR) video quality assessment (VQA) where the algorithms need to have good generalization performance when they are trained and tested on different datasets. We specifically address this question in the context of predicting video quality for compression and transmission applications. Motivated by the success of the spatio-temporal entropic differences video quality predictor in this context, we design a framework using convolutional neural networks to predict spatial and temporal entropic differences without the need for a reference or human opinion score. This approach enables our model to capture both spatial and temporal distortions effectively and allows for robust generalization. We evaluate our algorithms on a variety of datasets and show superior cross database performance when compared to state of the art NR VQA algorithms.
Shankhanil Mitra, Rajiv Soundararajan, Sumohana S. Channappayya
IEEE Signal Process. Lett.2
2020 A Model Learning Approach For Low Light Image Restoration
abstract
We study the problem of low light image restoration through contrast enhancement and denoising. We approach this problem by learning a model that relates a noisy low light and well lit image pair. The low light image is modeled to suffer from contrast distortion and additive noise. In particular, we model the loss of contrast through a global parametric function, which enables the estimation of the underlying noise. We then use a pair of convolutional neural network (CNN) models to learn the noise and the parameters of a function to achieve contrast enhancement. This contrast enhancement function is modeled as a linear combination of multiple gamma enhancers. We show through extensive evaluations that our Low Light Image Model for Enhancement Network (LLIMENet) achieves superior restoration performance when compared to other methods on several publicly available datasets.
Sameer Malik, Rajiv Soundararajan
ICIP2
2020 Robust image retrieval by cascading a deep quality assessment network
abstract
The performance of computer vision algorithms can severely degrade in the presence of a variety of distortions. While image enhancement algorithms have evolved to optimize image quality as measured according to human visual perception , their relevance in maximizing the success of computer vision algorithms operating on the enhanced image has been much less investigated. We consider the problem of image enhancement to combat Gaussian noise and low resolution with respect to the specific application of image retrieval from a dataset. We define the notion of image quality as determined by the success of image retrieval and design a deep convolutional neural network (CNN) to predict this quality. This network is then cascaded with a deep CNN designed for image denoising or super resolution , allowing for optimization of the enhancement CNN to maximize retrieval performance . This framework allows us to couple enhancement to the retrieval problem. We also consider the problem of adapting image features for robust retrieval performance in the presence of distortions. We show through experiments on distorted images of the Oxford and Paris buildings datasets that our algorithms yield improved mean average precision when compared to using enhancement methods that are oblivious to the task of image retrieval. 1
Biju Venkadath Somasundaran, Rajiv Soundararajan, Soma Biswas
Signal Process. Image Commun.2
2019 Llrnet: A Multiscale Subband Learning Approach for Low Light Image Restoration
abstract
We consider the problem of low light image restoration through joint contrast enhancement and denoising. Deep convolutional neural networks (CNNs) based on residual learning have been successful in achieving state of the art performance in image denoising. However, their application to joint contrast enhancement and denoising poses challenges owing to the nature of the distortion process involving both loss of details and noise. Thus, we propose a multiscale learning approach by learning the subbands obtained in a Laplacian pyramid decomposition through a subband CNN (SCNN). The enhanced subbands at multiple scales are then combined to obtain the final restored image using a recomposition CNN (ReCNN). We refer to the overall network involving SCNN and ReCNN as low light restoration network (LLRNet). We show through extensive experiments based on the `See in the Dark' Dataset that our approach produces better quality restored images when compared to other contrast enhancement techniques and CNN based approaches.
Sameer Malik, Rajiv Soundararajan
ICIP2
2019 Prediction of Discomfort due to Egomotion in Immersive Videos for Virtual Reality
abstract
We consider the problem of automatic assessment of visually induced motion sickness in virtual reality applications. In particular, we study the impact on visual discomfort due to camera motion or egomotion present in the video displayed through a head mounted display. We develop a database of 100 short duration videos with different camera trajectories, speeds and shake levels and conduct a large scale subjective study by collecting more than 4000 human ratings of discomfort levels. The videos are generated synthetically by applying different camera trajectories. We then use the subjective study to learn to predict discomfort by designing features describing the camera motion. The features are based on the ground truth camera trajectory and estimate the camera velocity and shake and depth of the visual scene. We show that these features can be effectively used to predict discomfort by obtaining a high correlation with the subjective discomfort scores provided by humans.
Suprith Balasubramanian, Rajiv Soundararajan
ISMAR2
2019 Machine vision quality assessment for robust face detection
Rajiv Soundararajan, Soma Biswas
Signal Process. Image Commun.1
2019 Subjective and Objective Quality Assessment of Stitched Images for Virtual Reality
abstract
We consider the problem of quality assessment (QA) of image stitching algorithms used to generate panoramic images for virtual reality applications. Our contributions are two-fold. We design the Indian Institute of Science Stitched Image QA (ISIQA) database consisting of 264 stitched images and 6600 human quality ratings. The database consists of a variety of artifacts due to stitching such as blur, ghosting, photometric, and geometric distortions. We then devise an objective QA model called the stitched image quality evaluator (SIQE) using the statistics of steerable pyramid decompositions. In particular, we propose a Gaussian mixture model to capture the bivariate statistics of neighboring coefficients of steerable pyramid decompositions and show this to be effective in modeling the increased spatial correlation due to ghosting artifacts. We show through extensive experiments that our quality model correlates very well with subjective scores in the ISIQA database. The ISIQA database as well as the software release of SIQE has been made available online for public use and evaluation purposes.
Pavan C. Madhusudana, Rajiv Soundararajan
IEEE Trans. Image Process.2
2018 Image Denoising for Image Retrieval by Cascading a Deep Quality Assessment Network
abstract
Image denoising algorithms have evolved to optimize image quality as measured according to human visual perception. However, image denoising to maximize the success of computer vision algorithms operating on the denoised image has been much less investigated. We consider the problem of image denoising for Gaussian noise with respect to the specific application of image retrieval from a dataset. We define the notion of image quality as determined by the success of image retrieval and design a deep convolutional neural network (CNN) to predict this quality. This network is then cascaded with a deep CNN designed for image denoising, allowing for optimization of the denoising CNN to maximize retrieval performance. This framework allows us to couple denoising to the retrieval problem. We show through experiments on noisy images of the Oxford and Paris buildings datasets that such an approach yields improved mean average precision when compared to using denoising methods that are oblivious to the task of image retrieval.
Biju Venkadath Somasundaran, Rajiv Soundararajan, Soma Biswas
ICIP2
2018 Generalized Gaussian scale mixtures: A model for wavelet coefficients of natural images
Praful Gupta, Anush K. Moorthy, Rajiv Soundararajan, Alan C. Bovik
Signal Process. Image Commun.3
2017 SpEED-QA: Spatial Efficient Entropic Differencing for Image and Video Quality
abstract
Many image and video quality assessment (I/VQA) models rely on data transformations of image/video frames, which increases their programming and computational complexity. By comparison, some of the most popular I/VQA models deploy simple spatial bandpass operations at a couple of scales, making them attractive for efficient implementation. Here we design reduced-reference image and video quality models of this type that are derived from the high-performance reduced reference entropic differencing (RRED) I/VQA models. A new family of I/VQA models, which we call the spatial efficient entropic differencing for quality assessment (SpEED-QA) model, relies on local spatial operations on image frames and frame differences to compute perceptually relevant image/video quality features in an efficient way. Software for SpEED-QA is available at: http://live.ece.utexas.edu/research/Quality/SpEED_Demo.zip.
Christos G. Bampis, Praful Gupta, Rajiv Soundararajan, Alan C. Bovik
IEEE Signal Process. Lett.3
2017 Evaluating Multiexposure Fusion Using Image Information
abstract
Multiexposure fusion (MEF) refers to image fusion methods that capture high dynamic range natural scenes from a set of low dynamic range camera images. In this letter, we study the problem of designing quality assessment (QA) algorithms to estimate the perceptual quality of images generated by different MEF algorithms. We develop our quality index by evaluating individual quality maps between a given fused test image and individual over-/underexposed images at multiple scales and orientations and then combining these maps across the over-/underexposed images. Our approach works on the premise that the true undistorted reference is contained across the over-/underexposed source images. We identify this true reference based on the notion of perceived image information using natural scene statistical models. It is shown that our approach outperforms the state of the art QA algorithms in terms of correlation with human perception of quality on a publicly available MEF database.
Hisham Rahman, Rajiv Soundararajan, Venkatesh Babu Radhakrishnan
IEEE Signal Process. Lett.2
2016 State Amplification Subject to Masking Constraints
abstract
This paper considers a state dependent broadcast channel with one transmitter, Alice, and two receivers, Bob and Eve. The problem is to effectively convey (“amplify”) the channel state sequence to Bob while “masking” it from Eve. The extent to which the state sequence cannot be masked from Eve is referred to as leakage. This can be viewed as a secrecy problem, where we desire that the channel state itself be minimally leaked to Eve while being communicated to Bob. This paper is aimed at characterizing the tradeoff region between amplification and leakage rates for such a system. An achievable coding scheme is presented, wherein the transmitter transmits a partial state information over the channel to facilitate the amplification process. For the case when Bob observes a stronger signal than Eve, the achievable coding scheme is enhanced with secure refinement. Outer bounds on the tradeoff region are also derived, and used in characterizing some special case results. In particular, the optimal amplification-leakage rate difference, called as differential amplification capacity, is characterized for the reversely degraded discrete memoryless channel, the degraded binary, and the degraded Gaussian channels. In addition, for the degraded Gaussian model, the extremal corner points of the tradeoff region are characterized, and the gap between the outer bound and achievable rate-regions is shown to be less than half a bit for a wide set of channel parameters.
Onur Ozan Koyluoglu, Rajiv Soundararajan, Sriram Vishwanath
IEEE Trans. Inf. Theory2
2013 Making a "Completely Blind" Image Quality Analyzer
abstract
An important aim of research on the blind image quality assessment (IQA) problem is to devise perceptual models that can predict the quality of distorted images with as little prior knowledge of the images or their distortions as possible. Current state-of-the-art “general purpose” no reference (NR) IQA algorithms require knowledge about anticipated distortions in the form of training examples and corresponding human opinion scores. However we have recently derived a blind IQA model that only makes use of measurable deviations from statistical regularities observed in natural images, without training on human-rated distorted images, and, indeed without any exposure to distorted images. Thus, it is “completely blind.” The new IQA model, which we call the Natural Image Quality Evaluator (NIQE) is based on the construction of a “quality aware” collection of statistical features based on a simple and successful space domain natural scene statistic (NSS) model. These features are derived from a corpus of natural, undistorted images. Experimental results show that the new index delivers performance comparable to top performing NR IQA models that require training on large databases of human opinions of distorted images. A software release is available at http://live.ece.utexas.edu/research/quality/niqe_release.zip.
Anish Mittal, Rajiv Soundararajan, Alan C. Bovik
IEEE Signal Process. Lett.2
2013 Video Quality Assessment by Reduced Reference Spatio-Temporal Entropic Differencing
abstract
We present a family of reduced reference video quality assessment (QA) models that utilize spatial and temporal entropic differences. We adopt a hybrid approach of combining statistical models and perceptual principles to design QA algorithms. A Gaussian scale mixture model for the wavelet coefficients of frames and frame differences is used to measure the amount of spatial and temporal information differences between the reference and distorted videos, respectively. The spatial and temporal information differences are combined to obtain the spatio-temporal-reduced reference entropic differences. The algorithms are flexible in terms of the amount of side information required from the reference that can range between a single scalar per frame and the entire reference information. The spatio-temporal entropic differences are shown to correlate quite well with human judgments of quality, as demonstrated by experiments on the LIVE video quality assessment database.
Rajiv Soundararajan, Alan C. Bovik
IEEE Trans. Circuits Syst. Video Technol.1
2013 Estimation With a Helper Who Knows the Interference
abstract
We consider the problem of estimating a signal corrupted by independent interference with the assistance of a cost-constrained helper who knows the interference causally or noncausally. When the interference is known causally, we characterize the minimum distortion incurred in estimating the desired signal. In the noncausal case, we present a general achievable scheme for discrete memoryless systems and novel lower bounds on the distortion for the binary and Gaussian settings. Our Gaussian setting coincides with that of assisted interference suppression introduced by Grover and Sahai. Our lower bound for this setting is based on the relation recently established by Verdú between divergence and minimum mean squared error. We illustrate with a few examples that this lower bound can improve on those previously developed. Our bounds also allow us to characterize the optimal distortion in several interesting regimes. Moreover, we show that causal and noncausal estimation are not equivalent for this problem. Finally, we consider the case where the desired signal is also available at the helper. We develop new lower bounds for this setting that improve on those previously developed, and characterize the optimal distortion up to a constant multiplicative factor for some regimes of interest.
Yeow-Khiang Chia, Rajiv Soundararajan, Tsachy Weissman
IEEE Trans. Inf. Theory2
2012 Estimation with a helper who knows the interference
abstract
We consider the problem of estimating a signal corrupted by independent interference in which the estimator (decoder) is aided by a helper (encoder) with a limited power budget that has knowledge of the interfering signal noncausally. In the Gaussian case, this problem is equivalent to the problem of Assisted Interference Suppression considered by Grover and Sahai and we obtain an improved lower bound for this problem using Verdú's relationship between mismatched estimation and relative entropy. We also extend our analysis to consider the case when the signal to be estimated, in addition to the interference, is known to the helper. We establish a lower bound that improves on those recently derived by Huang and Narayanan.
Yeow-Khiang Chia, Rajiv Soundararajan, Tsachy Weissman
ISIT2
2012 RRED Indices: Reduced Reference Entropic Differencing for Image Quality Assessment
abstract
We study the problem of automatic "reduced-reference" image quality assessment (QA) algorithms from the point of view of image information change. Such changes are measured between the reference- and natural-image approximations of the distorted image. Algorithms that measure differences between the entropies of wavelet coefficients of reference and distorted images, as perceived by humans, are designed. The algorithms differ in the data on which the entropy difference is calculated and on the amount of information from the reference that is required for quality computation, ranging from almost full information to almost no information from the reference. A special case of these is algorithms that require just a single number from the reference for QA. The algorithms are shown to correlate very well with subjective quality scores, as demonstrated on the Laboratory for Image and Video Engineering Image Quality Assessment Database and the Tampere Image Database. Performance degradation, as the amount of information is reduced, is also studied.
Rajiv Soundararajan, Alan C. Bovik
IEEE Trans. Image Process.1
2012 Communicating Linear Functions of Correlated Gaussian Sources Over a MAC
abstract
This paper considers the problem of transmitting linear functions of two correlated Gaussian sources over a two-user additive Gaussian noise multiple access channel. The goal is to recover this linear function within an average mean squared error distortion criterion. Each transmitter has access to only one of the two Gaussian sources and is limited by an average power constraint. In this paper, a lattice coding scheme and two lower bounds on the achievable distortion are presented. The lattice scheme achieves within a constant of a distortion lower bound if the signal-to-noise ratio is greater than a threshold. Furthermore, for the difference of correlated Gaussian sources, uncoded transmission is shown to be worse in performance to lattice coding methods for correlation coefficients above a threshold.
Rajiv Soundararajan, Sriram Vishwanath
IEEE Trans. Inf. Theory1
2012 Sum Rate of the Vacationing-CEO Problem
abstract
The vacationing chief executive officer (CEO) problem combines the salient features of the so-called CEO problem and the multiple-description (MD) problem. In this setting, noisy versions of a source are observed by two encoders, as in the CEO problem. In addition, we require that each encoder generate MDs of the source, as in the MD problem. The vacationing-CEO problem arises in asynchronous multicast networks, and solving it is an essential step in developing a general theory for multiencoder and multidecoder lossy compression. In this paper, an achievable sum rate and two sum rate lower bounds are presented for the quadratic Gaussian vacationing-CEO problem. These bounds exactly determine the optimal sum rate over a wide range of parameters.
Rajiv Soundararajan, Aaron B. Wagner, Sriram Vishwanath
IEEE Trans. Inf. Theory1
2011 RRED indices: Reduced reference entropic differencing framework for image quality assessment
abstract
We study the problem of automatic "reduced reference" image quality assessment algorithms from the point of view of image information change. Algorithms that measure differences between the entropies of wavelet coefficients of reference and distorted images are designed. A family of algorithms are presented, each differing in the amount of data on which information change is predicted and ranging from al most full reference to almost no reference. A special case of this are algorithms that require just a single number from the reference for quality assessment. The algorithms are shown to correlate very well with subjective quality scores as demonstrated on the LIVE Image Quality Assessment Database.
Rajiv Soundararajan, Alan C. Bovik
ICASSP1
2011 Multi-terminal source coding through a relay
abstract
This paper studies a multi -terminal source coding problem, where two terminals possess two (correlated) Gaussian sources to be compressed and delivered to the destination through an intermediate relay. Unlike the CEO and conventional two terminal source coding problems as well as point-to-point relay source coding problem, lattices are found to play an important role in achieving "good" rates for this problem setting. Two achievable strategies compute-and-forward and compress-and forward are used to develop achievable rates for this problem setting. For the symmetric case, the inner and outer bounds developed are shown to be within 1/2 bits of each other.
Rajiv Soundararajan, Sriram Vishwanath
ISIT1
2010 Sum rate of the vacationing CEO problem
abstract
This paper studies a class of source coding problems that combines elements of the CEO problem with the multiple description problem. In this setting, noisy versions of one remote source are observed by two nodes with encoders (which is similar to the CEO problem). However, it differs from the CEO problem in that each node must generate multiple descriptions of the source. This problem is of interest in multiple scenarios in efficient communication over networks. In this paper, an achievable region and an outer bound are presented for this problem, which is shown to be sum rate optimal for a class of distortion constraints.
Rajiv Soundararajan, Aaron B. Wagner, Sriram Vishwanath
ISIT1
2010 Wireless Video Quality Assessment: A Study of Subjective Scores and Objective Algorithms
abstract
Evaluating the perceptual quality of video is of tremendous importance in the design and optimization of wireless video processing and transmission systems. In an endeavor to emulate human perception of quality, various objective video quality assessment (VQA) algorithms have been developed. However, the only subjective video quality database that exists on which these algorithms can be tested is dated and does not accurately reflect distortions introduced by present generation encoders and/or wireless channels. In order to evaluate the performance of VQA algorithms for the specific task of H.264 advanced video coding compressed video transmission over wireless networks, we conducted a subjective study involving 160 distorted videos. Various leading full reference VQA algorithms were tested for their correlation with human perception. The data from the paper has been made available to the research community, so that further research on new VQA algorithms and on the general area of VQA may be carried out.
Anush K. Moorthy, Kalpana Seshadrinathan, Rajiv Soundararajan, Alan C. Bovik
IEEE Trans. Circuits Syst. Video Technol.3
2010 Study of Subjective and Objective Quality Assessment of Video
abstract
We present the results of a recent large-scale subjective study of video quality on a collection of videos distorted by a variety of application-relevant processes. Methods to assess the visual quality of digital videos as perceived by human observers are becoming increasingly important, due to the large number of applications that target humans as the end users of video. Owing to the many approaches to video quality assessment (VQA) that are being developed, there is a need for a diverse independent public database of distorted videos and subjective scores that is freely available. The resulting Laboratory for Image and Video Engineering (LIVE) Video Quality Database contains 150 distorted videos (obtained from ten uncompressed reference videos of natural scenes) that were created using four different commonly encountered distortion types. Each video was assessed by 38 human subjects, and the difference mean opinion scores (DMOS) were recorded. We also evaluated the performance of several state-of-the-art, publicly available full-reference VQA algorithms on the new database. A statistical evaluation of the relative performance of these algorithms is also presented. The database has a dedicated web presence that will be maintained as long as it remains relevant and the data is available online.
Kalpana Seshadrinathan, Rajiv Soundararajan, Alan C. Bovik, Lawrence K. Cormack
IEEE Trans. Image Process.2
2009 Communicating the Difference of Correlated Gaussian Sources over a MAC
abstract
This paper considers the problem of transmitting the difference of two positively correlated Gaussian sources over a two-user additive Gaussian noise multiple access channel (MAC). The goal is to recover this difference within an average mean squared error distortion criterion. Each transmitter has access to only one of the two Gaussian sources and is limited by an average power constraint. In this work, a lattice coding scheme that achieves a distortion within a constant of a distortion lower bound is presented if the signal to noise ratio (SNR) is greater than a threshold. Further, uncoded transmission is shown to be worse in performance to lattice coding methods for correlation coefficients above a threshold. An alternative lattice coding scheme is also presented that can potentially improve on the performance of uncoded transmission.
Rajiv Soundararajan, Sriram Vishwanath
DCC1
2009 Hybrid coding for Gaussian broadcast channels with Gaussian sources
abstract
This paper considers a degraded Gaussian broadcast channel over which Gaussian sources are to be communicated. When the sources are independent, this paper shows that hybrid coding achieves the optimal distortion region, the same as that of separate source and channel coding. It also shows that uncoded transmission is not optimal for this setting. For correlated sources, the paper shows that a hybrid coding strategy has a better distortion region than separate source-channel coding below a certain signal to noise ratio threshold. Thus, hybrid coding is a good choice for Gaussian broadcast channels with correlated Gaussian sources.
Rajiv Soundararajan, Sriram Vishwanath
ISIT1
2008 Adaptive Sum Power Iterative Waterfilling for MIMO Cognitive Radio Channels
abstract
In this paper, the sum capacity of the Gaussian multiple input multiple output (MIMO) cognitive radio channel (MCC) is expressed as a convex problem with finite number of linear constraints, allowing for polynomial time interior point techniques to find the solution. In addition, a specialized class of sum power iterative waterfilling algorithms is determined that exploits the inherent structure of the sum capacity problem. These algorithms not only determine the maximizing sum capacity value, but also the transmit policies that achieve this optimum. The paper concludes by providing numerical results which demonstrate that the algorithm takes very few iterations to converge to the optimum.
Rajiv Soundararajan, Sriram Vishwanath
ICC1