VLDB 2026 Research / reviewers in the wild / expert
Shankhanil Mitra
dblp:285/8320
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-1570-7442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vision-Language Model Guided Semi-supervised Learning for No-Reference Video Quality AssessmentabstractPerceptual assessment of user-generated content videos is an important problem that impacts viewing experience of millions of users. Current no-reference video quality assessment (NR-VQA) algorithms require a large amount of human annotated videos. In this work, we address this problem by specifically designing a dual-model based Semi-supervised Learning (SSL) method for NR-VQA. The first model is based on a popular vision-language model namely CLIP, where we adapt the visual encoder to capture high-level semantic quality information through quality-relevant text prompts. A second model learns complementary low-level spatio-temporal quality using a 3D vision transformer and video fragments. We enable intelligent knowledge transfer between the high-level vision-language and low-level vision-transformer model to pseudo-label the unlabelled videos. Our unified model outperforms existing state-of-the-art SSL methods for VQA across popular VQA databases including inter-database settings. Shankhanil Mitra, Rajiv Soundararajan |
ICASSP | 1 |
| 2024 | Knowledge Guided Semi-supervised Learning for Quality Assessment of User Generated VideosabstractPerceptual quality assessment of user generated content (UGC) videos is challenging due to the requirement of large scale human annotated videos for training. In this work, we address this challenge by first designing a self-supervised Spatio-Temporal Visual Quality Representation Learning (ST-VQRL) framework to generate robust quality aware features for videos. Then, we propose a dual-model based Semi Supervised Learning (SSL) method specifically designed for the Video Quality Assessment (SSL-VQA) task, through a novel knowledge transfer of quality predictions between the two models. Our SSL-VQA method uses the ST-VQRL backbone to produce robust performances across various VQA datasets including cross-database settings, despite being learned with limited human annotated videos. Our model improves the state-of-the-art performance when trained only with limited data by around 10%, and by around 15% when unlabelled data is also used in SSL. Source codes and checkpoints are available at https://github.com/Shankhanil006/SSL-VQA. Shankhanil Mitra, Rajiv Soundararajan |
AAAI | 1 |
| 2024 | Learning Generalizable Perceptual Representations for Data-Efficient No-Reference Image Quality AssessmentabstractNo-reference (NR) image quality assessment (IQA) is an important tool in enhancing the user experience in diverse visual applications. A major drawback of state-of-the-art NR-IQA techniques is their reliance on a large number of human annotations to train models for a target IQA application. To mitigate this requirement, there is a need for unsupervised learning of generalizable quality representations that capture diverse distortions. We enable the learning of low-level quality features agnostic to distortion types by introducing a novel quality-aware contrastive loss. Further, we leverage the generalizability of vision-language models by fine-tuning one such model to extract high-level image quality information through relevant text prompts. The two sets of features are combined to effectively predict quality by training a simple regressor with very few samples on a target dataset. Additionally, we design zero-shot quality predictions from both pathways in a completely blind setting. Our experiments on diverse datasets encompassing various distortions show the generalizability of the features and their superior performance in the data-efficient and zero-shot settings. Suhas Srinath, Shankhanil Mitra, Shika Rao, Rajiv Soundararajan |
WACV | 2 |
| 2024 | Semi-Supervised Learning of Perceptual Video Quality by Generating Consistent Pairwise Pseudo-RanksabstractDesigning learning-based no-reference (NR) video quality assessment (VQA) algorithms for camera-captured videos is cumbersome due to the large number of human annotations of quality. In this work, we propose a semi-supervised learning (SSL) framework exploiting many unlabelled and very limited numbers of authentically distorted labelled videos. Our main contributions are twofold. Leveraging the benefits of consistency regularization and pseudo-labelling, our SSL model generates pairwise pseudo-ranks for the unlabelled videos using a student-teacher model on strong-weak augmented videos. We design the strong-weak augmentations to be quality invariant to use the unlabelled videos effectively in SSL. The generated pseudo-ranks are used along with the limited labels to train our SSL model. Our primary focus in SSL for NR VQA is to learn mapping from video feature representations to quality scores. We compare various feature extraction methods and show that our SSL framework can lead to improved performance on these features. We present a spatial and temporal feature extraction method based on predicting spatial and temporal entropic differences. We show that these features help achieve robust performance when trained with limited data, providing a better baseline to apply SSL. Extensive experiments on three popular VQA datasets demonstrate that the proposed semi-supervised VQA method improves on the performance of existing methods in terms of correlation with human opinion by approximately$15 \! - \! 20 \%$ Shankhanil Mitra, Saiyam Jogani, Rajiv Soundararajan |
IEEE Trans. Multim. | 1 |
| 2023 | Test Time Adaptation for Blind Image Quality AssessmentabstractWhile the design of blind image quality assessment (IQA) algorithms has improved significantly, the distribution shift between the training and testing scenarios often leads to a poor performance of these methods at inference time. This motivates the study of test time adaptation (TTA) techniques to improve their performance at inference time. Existing auxiliary tasks and loss functions used for TTA may not be relevant for quality-aware adaptation of the pre-trained model. In this work, we introduce two novel quality-relevant auxiliary tasks at the batch and sample levels to enable TTA for blind IQA. In particular, we introduce a group contrastive loss at the batch level and a relative rank loss at the sample level to make the model quality aware and adapt to the target data. Our experiments reveal that even using a small batch of images from the test distribution helps achieve significant improvement in performance by updating the batch normalization statistics of the source model. Subhadeep Roy, Shankhanil Mitra, Soma Biswas, Rajiv Soundararajan |
ICCV | 2 |
| 2022 | Multiview Contrastive Learning for Completely Blind Video Quality Assessment of User Generated ContentabstractCompletely blind video quality assessment (VQA) refers to a class of quality assessment methods that do not use any reference videos, human opinion scores or training videos from the target database to learn a quality model. The design of this class of methods is particularly important since it can allow for superior generalization in performance across various datasets. We consider the design of completely blind VQA for user generated content. While several deep feature extraction methods have been considered in supervised and weakly supervised settings, such approaches have not been studied in the context of completely blind VQA. We bridge this gap by presenting a self-supervised multiview contrastive learning framework to learn spatio-temporal quality representations. In particular, we capture the common information between frame differences and frames by treating them as a pair of views and similarly obtain the shared representations between frame differences and optical flow. The resulting features are then compared with a corpus of pristine natural video patches to predict the quality of the distorted video. Detailed experiments on multiple camera captured VQA datasets reveal the superior performance of our method over other features when evaluated without training on human scores. Code will be made available at https://github.com/Shankhanil006/VISION. Shankhanil Mitra, Rajiv Soundararajan |
ACM Multimedia | 1 |
| 2021 | Predicting Spatio-Temporal Entropic Differences for Robust No Reference Video Quality AssessmentabstractWe consider the problem of robust no reference (NR) video quality assessment (VQA) where the algorithms need to have good generalization performance when they are trained and tested on different datasets. We specifically address this question in the context of predicting video quality for compression and transmission applications. Motivated by the success of the spatio-temporal entropic differences video quality predictor in this context, we design a framework using convolutional neural networks to predict spatial and temporal entropic differences without the need for a reference or human opinion score. This approach enables our model to capture both spatial and temporal distortions effectively and allows for robust generalization. We evaluate our algorithms on a variety of datasets and show superior cross database performance when compared to state of the art NR VQA algorithms. Shankhanil Mitra, Rajiv Soundararajan, Sumohana S. Channappayya |
IEEE Signal Process. Lett. | 1 |