Nabajeet Barman

dblp:192/3500 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0003-2587-7370ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 9 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Non-Aligned Reference Image Quality Assessment for Novel View Synthesis
abstract
Evaluating the perceptual quality of Novel View Synthesis (NVS) images remains a key challenge, particularly in the absence of pixel-aligned ground truth references. Full-Reference Image Quality Assessment (FR-IQA) methods fail under misalignment, while No-Reference (NR-IQA) methods struggle with generalization. In this work, we introduce a Non-Aligned Reference (NAR-IQA) framework tailored for NVS, where it is assumed that the reference view shares partial scene content but lacks pixel-level alignment. We constructed a large-scale image dataset containing synthetic distortions targeting Temporal Regions of Interest (TROI) to train our NAR-IQA model. Our model is built on a contrastive learning framework that incorporates LoRA-enhanced DINOv2 embeddings and is guided by supervision from existing IQA methods. We train exclusively on synthetically generated distortions, deliberately avoiding overfitting to specific real NVS samples and thereby enhancing the model’s generalization capability. Our model outperforms state-of-the-art FR-IQA, NR-IQA, and NAR-IQA methods, achieving robust performance on both aligned and non-aligned references. We also conducted a novel user study to gather data on human preferences when viewing non-aligned references in NVS. We find strong correlation between our proposed quality prediction model and the collected subjective ratings. For dataset, and code, please visit our project page: https://stootaghaj.github.io/nova-project/
Abhijay Ghildyal, Rajesh Sureddi, Nabajeet Barman, Saman Zad Tootaghaj, Alan C. Bovik
WACV3
2025 Foundation Models Boost Low-Level Perceptual Similarity Metrics
abstract
For full-reference image quality assessment (FR-IQA) using deep-learning approaches, the perceptual similarity score between a distorted image and a reference image is typically computed as a distance measure between features extracted from a pretrained CNN or more recently, a Transformer network. Often, these intermediate features require further fine-tuning or processing with additional neural network layers to align the final similarity scores with human judgments. So far, most IQA models based on foundation models have primarily relied on the final layer or the embedding for the quality score estimation. In contrast, this work explores the potential of utilizing the intermediate features of these foundation models, which have largely been unexplored so far in the design of low-level perceptual similarity metrics. We demonstrate that the intermediate features are comparatively more effective. Moreover, without requiring any training, these metrics can outperform both traditional and state-of-the-art learned metrics by utilizing distance measures between the features. Code: https://github.com/abhijay9/ZS-IQA
Abhijay Ghildyal, Nabajeet Barman, Saman Zad Tootaghaj
ICASSP2
2025 Triqa: Image Quality Assessment by Contrastive Pretraining on Ordered Distortion Triplets
abstract
Image Quality Assessment (IQA) models aim to predict perceptual image quality in alignment with human judgments. No-Reference (NR) IQA remains particularly challenging due to the absence of a reference image. While deep learning has significantly advanced this field, a major hurdle in developing NR-IQA models is the limited availability of subjectively labeled data. Most existing deep learning-based NR-IQA approaches rely on pre-training on large-scale datasets before fine-tuning for IQA tasks. To further advance progress in this area, we propose a novel approach that constructs a custom dataset using a limited number of reference content images and introduces a no-reference IQA model that incorporates both content and quality features for perceptual quality prediction. Specifically, we train a quality-aware model using contrastive triplet-based learning, enabling efficient training with fewer samples while achieving strong generalization performance across publicly available datasets. Our repository is available at https://github.com/rajeshsureddi/triqa.1
Rajesh Sureddi, Saman Zad Tootaghaj, Nabajeet Barman, Alan C. Bovik
ICIP3
2025 VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
abstract
With video games leading in entertainment revenues, optimizing game development workflows is critical to the industry’s long-term success. Recent advances in vision-language models (VLMs) hold significant potential to automate and enhance various aspects of game development—particularly video game quality assurance (QA), which remains one of the most labor-intensive processes with limited automation. To effectively measure VLM performance in video game QA tasks and evaluate their ability to handle real-world scenarios, there is a clear need for standardized benchmarks, as current ones fall short in addressing this domain. To bridge this gap, we introduce VideoGameQA-Bench - a comprehensive benchmark designed to encompass a wide range of game QA activities, including visual unit testing, visual regression testing, needle-in-a-haystack, glitch detection, and bug report generation for both images and videos.
Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zad Tootaghaj, Nabajeet Barman, Cor-Paul Bezemer
NeurIPS4
2024 Codec Compression Efficiency Evaluation for Ultra Low-Latency Cloud Gaming Applications
abstract
Video streaming applications, both on-demand and real-time, have seen tremendous growth and acceptance in the past two decades. To meet the increasing demand and user expectations of such any time, any device, any place availability of such services, there is a need for higher compression efficiency codecs and intelligent encoding strategies.This paper presents a comparative codec compression efficiency evaluation of the four most popular and widely used codec compression standards (H.264, HEVC, VP9 and AV1) for cloud gaming applications considering ultra-low latency encoding settings on a 4K gaming video dataset. Our results considering multiple resolution-bitrate pair encoding show that under strict ultra-low latency encoding settings, software codec implementations of VP9 and AV1 codecs perform better than H.264 and HEVC in terms of quality and rate savings albeit at an increased encoding time compared to H.264, especially at higher resolutions.
Nabajeet Barman, Saman Zad Tootaghaj, Steven Schmidt 0001, Yuanhan Chen, Man Cheung Kung
QoMEX1
2023 Datasheet for Subjective and Objective Quality Assessment Datasets
abstract
Over the years, many subjective and objective quality assessment datasets have been created and made available to the research community. However, there is no standard process for documenting the various aspects of the dataset, such as details about the source sequences, number of test subjects, test methodology, encoding settings, etc. Such information is often of great importance to the users of the dataset as it can help them get a quick understanding of the motivation and scope of the dataset. Without such a template, it is left to each reader to collate the information from the relevant publication or website, which is a tedious and time-consuming process. In some cases, the absence of a template to guide the documentation process can result in an unintentional omission of some important information. This paper addresses this simple but significant gap by proposing a datasheet template for documenting various aspects of sub-jective and objective quality assessment datasets for multimedia data. The contributions presented in this work aim to simplify the documentation process for existing and new datasets and improve their reproducibility. The proposed datasheet template is available on GitHub1, along with a few sample datasheets of a few open-source audiovisual subjective and objective datasets.
Nabajeet Barman, Yuriy A. Reznik, Maria G. Martini
QoMEX1
2023 A Subjective Dataset for Multi-Screen Video Streaming Applications
abstract
In modern-era video streaming systems, videos are streamed and displayed on a wide range of devices. Such devices vary from large-screen UHD and HDTVs to medium-screen Desktop PCs and Laptops to smaller-screen devices such as mobile phones and tablets. It is well known that a video is perceived differently when displayed on different devices. The viewing experience for a particular video on smaller screen devices such as smartphones and tablets, which have high pixel density, will be different with respect to the case where the same video is played on a large screen device such as a TV or PC monitor. Being able to model such relative differences in perception effectively can help in the design of better quality metrics and in the design of more efficient and optimized encoding profiles, leading to lower storage, encoding, and transmission costs. However, to the best of our knowledge, open-source datasets providing subjective scores for the same content when viewed on multiple devices with different screen sizes do not exist, thus limiting a proper evaluation of the existing quality metrics for such multi-screen video streaming applications. This paper addresses this research gap by presenting a new, open-source dataset consisting of subjective ratings for various encoded video sequences of different resolutions and bitrates (quality) when viewed on three devices of varying screen sizes: TV, Tablet, and Mobile. Along with the subjective scores, an evaluation of some of the most famous and commonly used open-source objective quality metrics is also presented. It is observed that the performance of the metrics varies a lot across different device types, with the recently standardized ITU-T P.1204.3 Model, on average, outperforming their full-reference counterparts. The dataset consisting of the videos, along with their subjective and objective scores, is available freely on Github1.
Nabajeet Barman, Yuriy A. Reznik, Maria G. Martini
QoMEX1
2022 Generalized Westerink-Roufs Model for Predicting Quality of Scaled Video
abstract
Resolution is a fundamental property of encoded video. Understanding the impact of resolution on quality as an independent parameter can help design better, more efficient systems, such as selecting optimum rendition in adaptive video streaming applications. One known quality model that considers resolution for predicting the perceived picture quality is the Westerink and Roufs (WR) model, which establishes the relationship between subjective quality and two parameters of viewing setup: angular resolution and viewing angle. This paper first validates the WR model on recent datasets and shows that it is reasonably accurate. We then propose a generalization of this model, allowing operation in a broader range of parameters and with more graceful saturation in extended regions. We then validate the performance of the proposed Generalized WR model on the new datasets and show that the proposed model achieves even a better fit to the recent datasets. We also demonstrate that the proposed Generalized model can account for the differences in scaling algorithms, including more advanced ML-based methods such as super-resolution. We conclude with a discussion of several possible applications of this model, including its use to guide the rendition selection decisions in streaming players and adapt that decision logic based on the upsampling algorithms used at the player.
Nabajeet Barman, Rahul Vanam, Yuriy A. Reznik
QoMEX1
2022 User Generated HDR Gaming Video Streaming: Dataset, Codec Comparison, and Challenges
abstract
Gaming video streaming services have grown tremendously in the past few years, with higher resolutions, higher frame rates and HDR gaming videos getting increasingly adopted among the gaming community. Since gaming content as such is different from non-gaming content, it is imperative to evaluate the performance of the existing encoders to help understand the bandwidth requirements of such services, as well as further improve the compression efficiency of such encoders. Towards this end, we present in this paper GamingHDRVideoSET, a dataset consisting of eighteen 10-bit UHD-HDR gaming videos and encoded video sequences using four different codecs, together with their objective evaluation results. Additionally, the paper discusses the codec compression efficiency of most widely used practical encoders, i.e., x264 (H.264/AVC), x265 (H.265/HEVC) and libvpx (VP9), as well the recently proposed encoder libaom (AV1), on 10-bit, UHD-HDR content gaming content. Our results show that the latest compression standard AV1 results in the best compression efficiency, followed by HEVC, H.264, and VP9.
Nabajeet Barman, Maria G. Martini
IEEE Trans. Circuits Syst. Video Technol.1
2021 Estimation of Quality Scores From Subjective Tests-Beyond Subjects' MOS
abstract
Subjective tests for the assessment of the quality of experience (QoE) are typically run with a pool of subjects providing their opinion scores using a 5-point scale. The subjects’ mean opinion score (MOS) is generally assumed as the best estimation of the average score in the target population. Indeed, for a large enough sample, we may assume that the mean of the variations across the subjects approaches zero, but this is not the case for the limited number of subjects typically considered in subjective tests. In this paper, we propose an approach based on generalized linear models (GLMs) for estimation of the population average QoE. The motivating dataset is composed of the individual scores assigned by 25 subjects to a set of gaming videos evaluated under different resolutions and compression ratios. The approach recognizes the multinomial nature of the data and allows for correlation between scores of the same subject. The resulting estimated average QoE is shown to follow more credible patterns than the MOS, particularly for higher bitrates, for which the model estimates present more coherent behavior. Similar convincing results are found on a second dataset, showing the validity of the approach.
Sergio Pezzulli, Maria G. Martini, Nabajeet Barman
IEEE Trans. Multim.3
2020 A Large-scale Evaluation of the bitstream-based video-quality model ITU-T P.1204.3 on Gaming Content
abstract
The streaming of gaming content, both passive and interactive, has increased manifolds in recent years. Gaming contents bring with them some peculiarities which are normally not seen in traditional 2D videos, such as the artificial and synthetic nature of contents or repetition of objects in a game. In addition, the perception of gaming content by the user is different from that of traditional 2D videos due to its pecularities and also the fact that users may not often watch such content. Hence, it becomes imperative to evaluate whether the existing video quality models usually designed for traditional 2D videos are applicable to gaming content. In this paper, we evaluate the applicability of the recently standardized bitstream-based video-quality model ITU-T P.1204.3 on gaming content. To analyze the performance of this model, we used 4 different gaming datasets (3 publicly available + 1 internal) not previously used for model training, and compared it with the existing state-of-the-art models. We found that the ITU P.1204.3 model out of the box performs well on these unseen datasets, with an RMSE ranging between 0.38 - 0.45 on the 5-point absolute category rating and Pearson Correlation between 0.85 - 0.93 across all the 4 databases. We further propose a full-HD variant of the P.1204.3 model, since the original model is trained and validated which targets a resolution of 4K/UHD-1. A 50:50 split across all databases is used to train and validate this variant so as to make sure that the proposed model is applicable to various conditions.
Rakesh Rao Ramachandra Rao, Steve Goering, Robert Steger, Saman Zad Tootaghaj, Nabajeet Barman, Stephan Fremerey, Sebastian Möller 0001, Alexander Raake
MMSP5
2020 DEMI: Deep Video Quality Estimation Model using Perceptual Video Quality Dimensions
abstract
Existing works in the field of quality assessment focus separately on gaming and non-gaming content. Along with the traditional modeling approaches, deep learning based approaches have been used to develop quality models, due to their high prediction accuracy. In this paper, we present a deep learning based quality estimation model considering both gaming and non-gaming videos. The model is developed in three phases. First, a convolutional neural network (CNN) is trained based on an objective metric which allows the CNN to learn video artifacts such as blurriness and blockiness. Next, the model is fine-tuned based on a small image quality dataset using blockiness and blurriness ratings. Finally, a Random Forest is used to pool frame-level predictions and temporal information of videos in order to predict the overall video quality. The light-weight, low complexity nature of the model makes it suitable for real-time applications considering both gaming and non-gaming content while achieving similar performance to existing state-of-the-art model NDNetGaming. The model implementation for testing is available on GitHub1.
Saman Zad Tootaghaj, Nabajeet Barman, Rakesh Rao Ramachandra Rao, Steve Goering, Maria G. Martini, Alexander Raake, Sebastian Möller 0001
MMSP2
2020 Quality Assessment of Gaming Videos Compressed via AV1
abstract
With the increasing demand in video traffic, delivering appropriate Quality of Experience may appear as a challenge for highly popular services such as Twitch.tv and YouTube-Gaming. This paper presents the evaluation of three widely used encoders, namely H.264, H.265 and AV1, for the case of gaming videos, assuming streaming scenario. The performance of the codecs has been assessed in terms of objective video quality metrics (PSNR, SSIM and VMAF) and subjective video quality assessment. The encoding settings were kept similar for all the three aforementioned encoders. In the considered setting, AV1 achieves a better level of video quality both in terms of objective and subjective VQA, for almost all bitrates and content considered. The improvement is remarkable in particular for the lower range of bitrates considered.
Darkhan Ashimov, Maria G. Martini, Nabajeet Barman
QoMEX3
2020 Quality Enhancement of Gaming Content using Generative Adversarial Networks
abstract
Recently, streaming of gameplay scenes has gained much attention, as evident with the rise of platforms such as Twitch.tv and Facebook Gaming. These streaming services have to deal with many challenges due to the low quality of source materials caused by client devices, network limitations such as bandwidth and packet loss, as well as low delay requirements. Spatial video artifact such as blockiness and blurriness as a result of as video compression or up-scaling algorithms can significantly impact the Quality of Experience of end-users of passive gaming video streaming applications. In this paper, we investigate solutions to enhance the video quality of compressed gaming content. Recently, several super-resolution enhancement techniques using Generative Adversarial Network (e.g., SRGAN) have been proposed, which are shown to work with high accuracy on non-gaming content. Towards this end, we improved the SRGAN by adding a modified loss function as well as changing the generator network such as layer levels and skip connections to improve the flow of information in the network, which is shown to improve the perceived quality significantly. In addition, we present a performance evaluation of improved SRGAN for the enhancement of frame quality caused by compression and rescaling artifacts for gaming content encoded in multiple resolution-bitrate pairs.
Nasim Jamshidi Avanaki, Saman Zad Tootaghaj, Nabajeet Barman, Steven Schmidt 0001, Maria G. Martini, Sebastian Möller 0001
QoMEX3
2020 An Evaluation of the Next-Generation Image Coding Standard AVIF
abstract
This paper presents a comparative performance evaluation of the newly proposed AV1 Image File Format (AVIF) vs. other state-of-the art image codecs, for natural, synthetic and gaming images. The codecs are compared in terms of Rate-quality curves and BD-Rate savings considering different quality metrics. AVIF results in the best overall performance considering both 4:2:0 and 4:4:4 chroma sub-sampling encoded images.
Nabajeet Barman, Maria G. Martini
QoMEX1
2018 NR-GVQM: A No Reference Gaming Video Quality Metric
abstract
Gaming as a popular system has recently expanded the associated services, by stepping into live streaming services. Live gaming video streaming is not only limited to cloud gaming services, such as Geforce Now, but also include passive streaming, where the players' gameplay is streamed both live and ondemand over services such as Twitch.tv and YouTubeGaming. So far, in terms of gaming video quality assessment, typical video quality assessment methods have been used. However, their performance remains quite unsatisfactory. In this paper, we present a new No Reference (NR) gaming video quality metric called NR-GVQM with performance comparable to state-of-the-art Full Reference (FR) metrics. NR-GVQM is designed by training a Support Vector Regression (SVR) with the Gaussian kernel using nine frame-level indexes such as naturalness and blockiness as input features and Video Multimethod Assessment Fusion (VMAF) scores as the ground truth. Our results based on a publicly available dataset of gaming videos are shown to have a correlation score of 0.98 with VMAF and 0.89 with MOS scores. We further present two approaches to reduce computational complexity.
Saman Zad Tootaghaj, Nabajeet Barman, Steven Schmidt 0001, Maria G. Martini, Sebastian Möller 0001
ISM2
2018 A Comparative Quality Assessment Study for Gaming and Non-Gaming Videos
abstract
Recent years have seen a tremendous increase in video traffic with the rise of Over The Top (OTT) services. Along with traditional Video on demand (VoD) streaming services (e.g., Netflix, YouTube), live video services (e.g., Twitch. tv, YouTubeGaming, Facebook Live) have also resulted in a tremendous share of Internet traffic. Among the live streaming services, gaming video streaming has a major share, with Twitch.tv alone currently responsible for the fourth highest peak Internet traffic in the US. As a consequence of this, and due to the fact that gaming videos are artificial and synthetic, it is worth investigating the specificity of gaming videos in relation to compression and the consequent end user QoE. In this paper, we present an objective and subjective quality comparison study for regular videos and gaming videos, with 30 video sequences (15 per type), encoded using the state of the art encoder HEVC. We discuss the similarity and dissimilarity between the two video types and also discuss how these observations can be used to improve the end user QoE.
Nabajeet Barman, Maria G. Martini, Saman Zad Tootaghaj, Sebastian Möller 0001, Sanghoon Lee 0001
QoMEX1
2017 H.264/MPEG-AVC, H.265/MPEG-HEVC and VP9 codec comparison for live gaming video streaming
abstract
Gaming videos are increasingly being streamed live over the Internet, as evident by the increasing popularity of services such as Twitch.tv and YouTube-Gaming, with Twitch.tv alone consisting of approx. two million streamers and over nine million daily active users. We present here an objective evaluation of eight most popular games encoded using H.264/MPEG-AVC, H.265/MPEG-HEVC and VP9 encoders for live game video streaming applications as currently used by Twitch.tv and YouTube-Gaming. The results are reported in terms of three objective video quality metrics (PSNR, SSIM, VIFp), Bjontegaard-Delta Bitrate (BD-BR) analysis, and encoding duration. For the encoding settings and the encoders used, in terms of BD-BR analysis, H.265/MPEG-HEVC is found to provide the best compression efficiency but is 2.6 times slower than H.264/MPEG-AVC. The magnitude of bitrate savings for VP9 compared to H.264/MPEG-AVC is found to be highly dependent on the content type, with H.264/MPEG-AVC resulting in higher average bitrate savings with an encoding speed four times faster than VP9.
Nabajeet Barman, Maria G. Martini
QoMEX1
2016 Predicting link quality of wireless channel of vehicular users using street and coverage maps
abstract
Limited radio resources and increasing user demands make wireless resource allocation a challenging task. As a user moves, his channel quality varies significantly depending on his location. In this paper we show that wireless link quality can be predicted on a large time scale in terms of Path Loss values by using street and coverage maps. The predicted wireless link quality information can be used by the network operators for optimized resource allocation. The feasibility of the proposed approach is demonstrated by extensive simulations advocating this data driven approach for channel prediction.
Nabajeet Barman, Stefan Valentin, Maria G. Martini
PIMRC1