VLDB 2026 Research / reviewers in the wild / expert
Vinit Jakhetiya
dblp:27/9850
· DBLP profile ↗
47ranked-venue papers
7as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-authorSecurity and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Cross-Modal Mutual Prompt Learning for Video Quality AssessmentabstractEnhancing video quality assessment (VQA) through semantic information integration is a critical research focus. Recent research has employed the Contrastive Language-Image Pre-training (CLIP) model as a foundation to improve semantic perception. However, the image-text alignment inherent in these pre-trained Vision-Language (VL) models frequently results in suboptimal VQA performance. While prompt engineering has recently targeted the language component to address this alignment issue, the unique insights resided in visual analysis is still overlooked for further advancing VQA tasks. Additionally, seeking a trade-off between quality separability and domain invariance in VQA remains largely unresolved within the VL paradigm. In this paper, we introduce a novel cross-modal prompt-based approach to tackle these challenges. Specifically, we propose learnable prompts within the vision branch to foster synergy between visual and language modalities through a language-to-vision coupling function. The multi-view backbone is then carefully crafted with content enhancement and distortion-aware temporal modulation to ensure quality separability. The language prompts, derived from visual representations, are further supported by adaptive weighting mechanisms to optimize the balance between quality separability and domain invariance. Experimental results demonstrate the effectiveness of our proposed method over leading VQA models, showing significant improvements in generalization across diverse datasets. The source code for this work is publicly available athttps://github.com/cpf0079/CM2PL. Pengfei Chen 0003, Leida Li, Jinjian Wu, Jiebin Yan, Vinit Jakhetiya, Aladine Chetouani |
IEEE Trans. Multim. | 5 |
| 2025 | T-GAP: Temporal Granularity Aware Projection Network for Action LocalizationabstractAction recognition and localization in videos pose significant challenges, as they require identifying temporal boundaries and classifying actions within long video sequences. This paper introduces an anchor-free, single-stage framework that predicts both the start and end times of actions while simultaneously classifying the actions, framing the task as a sequence labeling problem. The proposed model utilizes an encoder-decoder architecture with LSTM projections and 1D convolutions to capture rich temporal dependencies. This is followed by hierarchical feature pyramid generation, which is refined using Pyramidal Pooling Aggregation (PPA) and enhanced through a Progressive Feature Enhancement Unit (PFEU) with dilated convolutions to preserve the contextual relationships present in complex video frames. We also propose a novel Temporal Granularity Convolution (TGC) Layer, which refines temporal features. The TGC Layer captures various temporal details using a multi-branch structure consisting of Fine-Grained and Coarse-Grained Temporal Branches, designed to efficiently handle different temporal granularities. Finally, a decoder utilizes these multi-scale features to predict action instances at multiple levels of granularity. The model simultaneously integrates classification and regression heads to predict action labels and temporal boundaries, achieving improved localization and recognition. We have demonstrated the effectiveness of the proposed scheme, using mean average precision (mAP) across various Intersection over Union (IoU) thresholds on the Thumos14, ActivityNet-1.3, and MultiTHUMOS datasets over multiple state-of-the-art (SOTA) methods. Himanshu Singh 0006, Avijit Dey, Badri N. Subudhi, Vinit Jakhetiya, Veerakumar Thangaraj |
AVSS | 4 |
| 2025 | Feature Affinity based Clustering for Test-Time Adaptation for Image Quality AssessmentabstractRecently, Test-Time Adaptation (TTA) algorithms have gained traction in image/video quality assessment (IQA/VQA). These methods use virtual losses, like group contrastive and rank loss, as auxiliary tasks to adapt batch normalization parameters, helping models generalize better to distribution shifts between training and testing datasets. Group contrastive loss clusters images into low- and high-quality groups based on predicted quality scores, maximizing feature space distance between them. However, its effectiveness relies on the base model’s ability to accurately predict quality scores, which can be compromised when distribution shifts occur, leading to suboptimal adaptation and degraded performance. We propose a novel clustering approach based on the assumption that high-quality images contain richer high-level information, which is extracted using a pre-trained VGG-16 model. Images are clustered by comparing the VGG-16 features of the highest quality image in a batch with the others, enabling effective grouping based on feature affinities. These accurate clusters enhance the computation of contrastive loss, improving the adaptation of batch normalization layers. Additionally, we introduce an adaptive rank loss to reduce the impact of rank loss when the base model can distinguish images with varying distortion levels. Experimental results across multiple image quality assessment datasets, including LIVE, CID-2013, KONIQ-10K, and SPAQ, as well as algorithms like MetaIQA, HyperIQA, TReS, and MUSIQ, show that the proposed method consistently performs better than the existing Test-Time Adaptation (TTA) approach. Meghna Kapoor, Vinit Jakhetiya, Badri N. Subudhi, Ankur Bansal, Weisi Lin |
ICME | 2 |
| 2025 | Two streams ResNet-50 network for infrared and visible image fusion
Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya |
Multim. Tools Appl. | 4 |
| 2025 | Underwater surveillance using spatially curated perceptual loss and graph refactored network
Meghna Kapoor, Bhargava N. Satya, Badri N. Subudhi, Vinit Jakhetiya, Ankur Bansal |
Pattern Recognit. | 4 |
| 2024 | Underwater Change Detection Using Multiple Sampling-Based Probabilistic Learner and Feature Preservance DiscriminatorabstractSurveillance can be defined as the process of monitoring the behavior and activities of different objects to generate meaningful insights into a video scene. In the context of underwater, surveillance can be elucidated as one of the processes of detecting and tracking the moving objects present in underwater videos. Many methods have been put forth to separate moving objects from underwater environments. Nevertheless, such methods cannot maintain the minute details that are crucial for determining an object’s boundary. This is mainly due to the intricate natural properties of water and some of its characteristics, such as excessive turbidity, scattering, low visibility, etc. In this regard, we put forth an adversarial learning-based end-to-end deep learning architecture to detect underwater moving objects. The proposed architecture uses two modules for underwater object detection. The initial module is a generator comprised of a probabilistic learner which is based on multiple down-sampling and up-sampling modules. Further, the discriminator network is composed of a multi-level feature concatenation module which can perpetuate specifics at distinct levels. The effectiveness of the proposed method is confirmed using two underwater benchmark datasets by contrasting its outcomes with those of eight state-of-the-art methods. Mehvish Nissar, Badri N. Subudhi, Vinit Jakhetiya, Amit Mishra 0004 |
ICIP | 3 |
| 2024 | Principal Graph Neighborhood Aggregation for Underwater Moving Object Detection
Meghna Kapoor, Badri N. Subudhi, Vinit Jakhetiya, Ankur Bansal |
ICPR (29) | 3 |
| 2024 | Project and Pool: An Action Localization Network for Localizing Actions in Untrimmed Videos
Himanshu Singh 0006, Avijit Dey, Badri N. Subudhi, Vinit Jakhetiya |
ICPR (29) | 4 |
| 2024 | A Multi-Scale Contrast Preserving Encoder-Decoder Architecture for Local Change Detection From Thermal Video ScenesabstractThis article presents a new deep-learning architecture based on an encoder-decoder framework that retains contrast while performing background subtraction (BS) on thermal videos. The proposed scheme consists of three consecutive blocks: the encoder, the Multi-Scale Contrast Preservation (MSCP) block, and the decoder. The encoder network employs a hybrid of convolution and atrous convolution blocks to preserve both sparse and dense features, with a skip connection. The encoder, combined with the MSCP block, maintains multi-scale contrast features with reduced training loss. Furthermore, the decoder network accurately projects the extracted features at different layers into pixel-level detail. The proposed end-to-end model efficiently provides a binary map for the corresponding thermal video scene. The efficiency of the proposed algorithm is validated on two large-scale datasets, namely CDnet 2014 and the Tripura University Video Dataset at Night Time (TU-VDN). Both qualitative and quantitative results demonstrate that MSCP outperforms thirty-eight existing BS schemes. Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya, Thierry Bouwmans |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Non-Subsampled Contourlet Transform and Ground-Truth Score Generation Based Quality Assessment for DIBR-Synthesized ViewsabstractIn recent years, there have been advancements in developing Depth-Image-Based Rendering (DIBR) views. However, the quality of these synthesized views is often degraded by inefficient in-painting techniques and synthesis procedures, leading to geometric and structural distortions. This paper introduces two novel approaches to evaluate the quality of DIBR synthesized views, using full reference (FR) and no-reference (NR) metrics. The proposed FR quality assessment (QA) metric is based on the observation that the deep features of the Non-Subsampled Contourlet Transform (NSCT) maps capture the perceptually important characteristics of the images. By calculating the difference between these deep feature vectors of the reference and distorted views, we determine the quality of the image. Moreover, a lot of existing NR metrics typically divide an image into blocks and assign the same subjective quality scores to each block for training a deep learning model. However, this approach is not suitable for DIBR synthesized views, as distortions are often localized in specific areas rather than affecting the entire view. Consequently, the performance of existing block-based deep-learning algorithms suffers due to the absence of accurate ground truth scores for each image block. To address this limitation, this work proposes an innovative method for calculating ground truth scores for individual image blocks. This process is similar to the proposed FR metric. Firstly, we obtain the deep features of NSCT map of an image block and the quality score for each block is calculated using its and the reference block's feature vector. These block-wise ground truth scores are used to train a deep learning model which serves as an NR metric for estimating the quality of a given test block. Finally, the predicted block-level quality values are aggregated to determine the overall quality of the entire image. Experimental results demonstrate that both the proposed algorithms perform better than the existing objective metrics for DIBR synthesized views. Deebha Mumtaz, Sadbhawna, Vinit Jakhetiya, Badri N. Subudhi, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2024 | Bayesian's probabilistic strategy for feature fusion from visible and infrared images
Veerakumar Thangaraj, Badri N. Subudhi, Vinit Jakhetiya |
Vis. Comput. | 4 |
| 2023 | Two-Streams: Dark and Light Networks with Graph Convolution for Action Recognition from Dark Videos (Student Abstract)abstractIn this article, we propose a two-stream action recognition technique for recognizing human actions from dark videos. The proposed action recognition network consists of an image enhancement network with Self-Calibrated Illumination (SCI) module, followed by a two-stream action recognition network. We have used R(2+1)D as a feature extractor for both streams with shared weights. Graph Convolutional Network (GCN), a temporal graph encoder is utilized to enhance the obtained features which are then further fed to a classification head to recognize the actions in a video. The experimental results are presented on the recent benchmark ``ARID" dark-video database. Saurabh Suman, Nilay Naharas, Badri N. Subudhi, Vinit Jakhetiya |
AAAI | 4 |
| 2023 | Drone-vs-Bird: Drone Detection Using YOLOv7 with CSRT TrackerabstractDrone detection has become a critical component of surveillance, but detecting small drones remains a challenge. In this work, we introduce a new drone detection scheme that uses a customized YOLOv7 combined with a tracker. Using simple YOLOv7 has some limitations, including false positives and difficulty detecting drones in complex environments. To address these issues, we propose two solutions. First, we use a simple method based on the probability or confidence level of detected objects to reduce the number of false positives. Second, we integrate YOLOv7 with object tracking to detect drones in complex environments. Our proposed method was successful in achieving a top-5 ranking in the 6thedition of the Drone-vs-Bird challenge, which was organized by the ICASSP Signal Processing Grand Challenge in 2023. Sahaj Mistry, Shreyas Chatterjee, Ajeet Kumar Verma, Vinit Jakhetiya, Badri N. Subudhi, Sunil Prasad Jaiswal |
ICASSP | 4 |
| 2023 | Measuring Underwater Image Quality with Depthwise-Separable Convolutions and Global-Sparse AttentionabstractUnderwater Image Quality Assessment (UIQA) poses specific challenges due to blur, dispersion of light, and color garbling caused by water turbulence. Images captured by Autonomous Underwater Vehicles (AVS) correlate poorly along their RGB channels, as blue and green channel components dominate. We propose a novel lightweight no-reference (NR) UIQA architecture based on depthwise-separable convolutions and global sparse attention that surpasses the existing deep learning architectures in terms of performance as well as parameter count and Multiply-Accumulate operations (MACs). Proposed model gives 3.64% increase in SRCC and 6.11% increase in PLCC in comparison to state-of-the-art for UEQAB dataset. Our work also illustrates the effectiveness of depthwise separable convolutions for underwater image quality assessment. Our empirical analysis confirms less redundancy among the channels of a feature map when depthwise separable convolutions are used compared to standard convolutions by increasing representational efficiency. Arpita Nema, Vinit Jakhetiya, Sunil Prasad Jaiswal, Badri N. Subudhi, Sharath Chandra Guntuku |
MMSP | 2 |
| 2023 | Kernel-Induced Possibilistic Fuzzy Associate Background Subtraction for Video SceneabstractThe background subtraction (BGS) technique is popularly used for many surveillance systems, segmenting the foreground by subtracting the modeled background from the image sequences. The effectiveness of any BGS technique depends on the robustness of the constructed background model. It is to be noted that many BGS schemes are affected by the inclusion of either noisy pixels in background construction or parameters of generative models. In this regard, we propound an idea of a kernel-induced possibilistic fuzzy associated BGS scheme for local change detection from a fixed camera captured sequence. The proposed scheme follows two stages: background training and foreground segmentation. In the background construction stage, each pixel is modeled using a possibilistic fuzzy cost function in kernel-induced space. The use of the induced kernel function will project the low-dimensional data into a higher dimensional space and the use of the possibilistic function will construct a robust background model based on the density of the data in the temporal domain avoiding the noisy and outlier points. The performance of the proposed scheme is tested on three benchmark databases. The effectiveness of the proposed scheme is evaluated on different performance evaluation measures: precision, recall, F-measure, and average similarity. We corroborate our findings by comparing them against 19 state-of-the-art existing BGS techniques. Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya, Esakkirajan Sankaralingam |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Context Region Identification Based Quality Assessment of 3D Synthesized ViewsabstractPerceptual quality assessment of 3D synthesized views is an open research problem in computer vision. Researchers across the globe have developed several algorithms to identify distortions. At the same time, the existing algorithms cannot quantify the context in which these distortions affect the overall perceptual quality. According to the recently proposed 3D view synthesis algorithm, the choice of context region for the disocclusion plays a vital role in predicting the quality of 3D views. The context region taken from the background of a view produces a perceptually better quality of 3D synthesized views than when the context region is taken from the foreground. With this view, the proposed algorithm aims to identify the context region and incorporate this information for the perceptual quality assessment of 3D synthesized views. We observed that the depth energy maps of the 3D synthesized views vary significantly with the change in the context region and subsequently can identify the context region. Hence, in this work, we propose a new and efficient quality assessment algorithm based upon the variation in the depth of 3D synthesized and reference views, giving two-fold advantages: 1. It can predict the quality based on whether the context region is foreground or not. 2. It is also able to suggest the possible location of distortions. We have proposed two new algorithms for both situations when the context region is foreground or not. The overall predicted score is the direct multiplication of the quality score estimated when the context region is foreground or not. When applied to the established benchmark dataset, the proposed technique performs satisfactorily with the PLCC of 0.7707 and 0.7572 of SRCC. Also, the proposed algorithm can work as a plug-in to improve the performance of the existing algorithms. Sadbhawna Thakur, Vinit Jakhetiya, Badri N. Subudhi, Sunil Prasad Jaiswal, Leida Li, Weisi Lin |
IEEE Trans. Multim. | 2 |
| 2022 | Do We Need a New Large-Scale Quality Assessment Database for Generative Inpainting Based 3D View Synthesis? (Student Abstract)abstractThe advancement in Image-to-Image translation techniques using generative Deep Learning-based approaches has shown promising results for the challenging task of inpainting-based 3D view synthesis. At the same time, even the current 3D view synthesis methods often create distorted structures or blurry textures inconsistent with surrounding areas. We analyzed the recently proposed algorithms for inpainting-based 3D view synthesis and observed that these algorithms no longer produce stretching and black holes. However, the existing databases such as IETR, IRCCyN, and IVY have 3D-generated views with these artifacts. This observation suggests that the existing 3D view synthesis quality assessment algorithms can not judge the quality of most recent 3D synthesized views. With this view, through this abstract, we analyze the need for a new large-scale database and a new perceptual quality metric oriented for 3D views using a test dataset. Sadbhawna, Vinit Jakhetiya, Badri N. Subudhi, Harshit Shakya, Deebha Mumtaz |
AAAI | 2 |
| 2022 | C3D and Localization Model for Locating and Recognizing the Actions from Untrimmed Videos (Student Abstract)abstractIn this article, we proposed a technique for action localization and recognition from long untrimmed videos. It consists of C3D CNN model followed by the action mining using the localization model, where the KNN classifier is used. We segment the video into expressible sub-action known as action-bytes. The pseudo labels have been used to train the localization model, which makes the trimmed videos untrimmed for action-bytes. We present experimental results on the recent benchmark trimmed video dataset “Thumos14”. Himanshu Singh 0006, Tirupati Pallewad, Badri N. Subudhi, Vinit Jakhetiya |
AAAI | 4 |
| 2022 | Social Media Reveals Urban-Rural Differences in Stress across China
Jesse Cui, Tingdan Zhang, Kokil Jaidka, Dandan Pang, Garrick Sherman, Vinit Jakhetiya, Lyle H. Ungar, Sharath Chandra Guntuku |
ICWSM | 6 |
| 2022 | Transformer-based quality assessment model for generalized user-generated multimedia audio content
Deebha Mumtaz, Ajit Jena, Vinit Jakhetiya, Karan Nathwani, Sharath Chandra Guntuku |
INTERSPEECH | 3 |
| 2022 | Encoder and decoder network with ResNet-50 and global average feature pooling for local change detection
Akhilesh Sharma, Vatsalya Bajpai, Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya |
Comput. Vis. Image Underst. | 6 |
| 2022 | Nonintrusive Perceptual Audio Quality Assessment for User-Generated Content Using Deep LearningabstractWith the boom of social media communication, teleconferencing, and online classes, audiovisual communication over bandwidth strained networks has become an integral part of our lives. Consequently, the growing demand for the quality of experience necessitates developing algorithms to measure and enrich user experience. Prior studies have mainly focused on assessing speech quality and intelligibility with reference to audio quality assessment, while other categories in user-generated multimedia (UGM) are less explored. Moreover, frequency-domain properties of speech and UGM audio are significantly different from each other. Furthermore, there is a lack of a standard dataset for the quality assessment of UGM. Considering these limitations, in this article, we first develop the IIT-JMU-UGM audio dataset consisting of 1150 audio clips, with diverse context, content, and types of degradation commonly observed in real-world scenarios and annotated with the subjective quality scores. Finally, we propose a non-intrusive audio quality assessment metric using a stacked gated-recurrent-unit-based deep learning framework. The proposed model outperforms several baseline methods, including state-of-the-art non-intrusive and intrusive approaches. The resulting Pearson’s correlation coefficient of 0.834 indicates that the proposed method efficiently mirrors human auditory perception. Deebha Mumtaz, Vinit Jakhetiya, Karan Nathwani, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Perceptually Unimportant Information Reduction and Cosine Similarity-Based Quality Assessment of 3D-Synthesized ImagesabstractQuality assessment of 3D-synthesized images has traditionally been based on detecting specific categories of distortions such as stretching, black-holes, blurring, etc. However, such approaches have limitations in accurately detecting distortions entirely in 3D synthesized images affecting their performance. This work proposes an algorithm to efficiently detect the distortions and subsequently evaluate the perceptual quality of 3D synthesized images. The process of generation of 3D synthesized images produces a few pixel shift between reference and 3D synthesized image, and hence they are not properly aligned with each other. To address this, we propose using morphological operation (opening) in the residual image to reduce perceptually unimportant information between the reference and the distorted 3D synthesized image. The residual image suppresses the perceptually unimportant information and highlights the geometric distortions which significantly affect the overall quality of 3D synthesized images. We utilized the information present in the residual image to quantify the perceptual quality measure and named this algorithm as Perceptually Unimportant Information Reduction (PU-IR) algorithm. At the same time, the residual image cannot capture the minor structural and geometric distortions due to the usage of erosion operation. To address this, we extract the perceptually important deep features from the pre-trained VGG-16 architectures on the Laplacian pyramid. The distortions in 3D synthesized images are present in patches, and the human visual system perceives even the small levels of these distortions. With this view, to compare these deep features between reference and distorted image, we propose using cosine similarity and named this algorithm as Deep Features extraction and comparison using Cosine Similarity (DF-CS) algorithm. The cosine similarity is based upon their similarity rather than computing the magnitude of the difference of deep features. Finally, the pooling is done to obtain the objective quality scores using simple multiplication to both PU-IR and DF-CS algorithms. Our source code is available online: https://github.com/sadbhawnathakur/3D-Image-Quality-Assessment. Sadbhawna, Vinit Jakhetiya, Shubham Chaudhary 0005, Badri N. Subudhi, Weisi Lin, Sharath Chandra Guntuku |
IEEE Trans. Image Process. | 2 |
| 2022 | Stretching Artifacts Identification for Quality Assessment of 3D-Synthesized ViewsabstractExisting Quality Assessment (QA) algorithms consider identifying "black-holes" to assess perceptual quality of 3D-synthesized views. However, advancements in rendering and inpainting techniques have made black-hole artifacts near obsolete. Further, 3D-synthesized views frequently suffer from stretching artifacts due to occlusion that in turn affect perceptual quality. Existing QA algorithms are found to be inefficient in identifying these artifacts, as has been seen by their performance on the IETR dataset. We found, empirically, that there is a relationship between the number of blocks with stretching artifacts in view and the overall perceptual quality. Building on this observation, we propose a Convolutional Neural Network (CNN) based algorithm that identifies the blocks with stretching artifacts and incorporates the number of blocks with the stretching artifacts to predict the quality of 3D-synthesized views. To address the challenge with existing 3D-synthesized views dataset, which has few samples, we collect images from other related datasets to increase the sample size and increase generalization while training our proposed CNN-based algorithm. The proposed algorithm identifies blocks with stretching distortions and subsequently fuses them to predict perceptual quality without reference, achieving improvement in performance compared to existing no-reference QA algorithms that are not trained on the IETR dataset. The proposed algorithm can also identify the blocks with stretching artifacts efficiently, which can further be used in downstream applications to improve the quality of 3D views. Our source code is available online: https://github.com/sadbhawnathakur/3D-Image-Quality-Assessment. Sadbhawna, Vinit Jakhetiya, Deebha Mumtaz, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Image Process. | 2 |
| 2021 | Detecting Covid-19 and Community Acquired Pneumonia Using Chest CT Scan Images With Deep LearningabstractWe propose a two-stage Convolutional Neural Network (CNN) based classification framework for detecting COVID-19 and Community Acquired Pneumonia (CAP) using the chest Computed Tomography (CT) scan images. In the first stage, an infection - COVID-19 or CAP, is detected using a pre-trained DenseNet architecture. Then, in the second stage, a fine-grained three-way classification is done using EfficientNet architecture. The proposed COVID+CAP-CNN framework achieved a slice-level classification accuracy of over 94% at identifying COVID-19 and CAP. Further, the proposed framework has the potential to be an initial screening tool for differential diagnosis of COVID-19 and CAP, achieving a validation accuracy of over 89.3% at the finer three-way COVID-19, CAP, and healthy classification. Within the IEEE ICASSP 2021 Signal Processing Grand Challenge (SPGC) on COVID-19 Diagnosis, our proposed two-stage classification framework achieved an overall accuracy of 90% and sensitivity of .857, .9, and .942 at distinguishing COVID-19, CAP, and normal individuals respectively, to rank first in the evaluation. Code and model weights are available at https://github.com/shubhamchaudhary2015/ct_covid19_cap_cnn Shubham Chaudhary 0005, Sadbhawna, Vinit Jakhetiya, Badri N. Subudhi, Ujjwal Baid, Sharath Chandra Guntuku |
ICASSP | 3 |
| 2021 | Perceptual Quality Assessment of DIBR Synthesized Views Using Saliency Based Deep FeaturesabstractIn recent years, Depth-Image-Based-Rendering (DIBR) synthesized views have gained popularity due to their numerous visual media applications. Consequently, the research in their quality assessment (QA) has also gained momentum. In this work, we propose an efficient metric to estimate the perceptual quality of DIBR synthesized views via the extraction of Deep-features. These Deep-features are extracted from a pretrained CNN model. Generally, in DIBR synthesized views, geometric distortions arise near the objects due to occlusion, and the human visual system is quite sensitive towards these objects. On the other end, saliency maps are efficiently able to highlight perceptually important objects. With this intuition, instead of extracting deep features directly from DIBR synthesized views, we obtain the refined feature vector from their corresponding saliency maps. Also, most of the pixels with geometric distortions have a nearly similar impact on the perceptual quality of 3D synthesized views. Considering this, we propose to fuse the feature maps using the cosine similarity measure based upon the deviation of one feature vector from another. It may also be emphasized that no training is performed in the proposed algorithm, and all the features are extracted from the pre-trained vanilla VGG-16 architecture. The proposed metric, when applied to the standard database, results in PLCC of 0.762 and SRCC equal to 0.7513, which is better than the existing state-of-the-art QA metrics. Shubham Chaudhary 0005, Alokendu Mazumder, Deebha Mumtaz, Vinit Jakhetiya, Badri N. Subudhi |
ICIP | 4 |
| 2021 | Perceptual Quality Evaluation of Hazy Natural ImagesabstractHaze is an intrusion element that disrupts color fidelity and contrast of outdoor natural images, affecting their perceptual quality. The differential characteristics of hazy images compared to other natural images restrict the generalization of existing image quality assessment (IQA) algorithms. At the same time, efficient IQA algorithms for predicting the perceptual quality of naturally hazed images have not been proposed in the literature due to lack of a relevant dataset. To address this, we build the IIT-JMU Hazy Image Dataset comprising of 1000 high-definition hazy natural images consisting of diverse categories such as landscape, forests, roads, seascapes, and cityscapes, along with their subjective quality ratings. We present an analysis of existing natural-scene-statistics-based IQA algorithms on hazy natural images. In this article, we propose a convolutional-neural-network-based quality assessment algorithm for hazy natural images along with an IQA metric called deep learning-based haze perceptual quality evaluator (DLHPQE). The proposed DLHPQE efficiently predicts the perceptual quality of hazy natural images without a reference. Our results demonstrate that the DLHPQE outperforms existing state-of-the-art no-reference IQAs in terms of several performance parameters such as Pearson linear correlation coefficient, Spearman rank-order correlation coefficient, Kendall's rank-order correlation coefficient, and root-mean-square error. Palak Mahajan, Vinit Jakhetiya, Pawanesh Abrol, Parveen Lehana, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Distortion Specific Contrast Based No-Reference Quality Assessment of DIBR-Synthesized ViewsabstractIn the literature, many 3D-Synthesized Image Quality Assessment (IQA) algorithms are proposed, which are based on predicting the geometric and structural distortions present in the synthesized datasets. With the exponential growth of accurate inpainting algorithms, certain types of distortions, such as Blackholes, has become obsolete. Unfortunately, the existing IQA algorithms are mainly concentrating on efficiently identifying these black holes and subsequently predicting the perceptual quality of 3D synthesized views. The performance of these algorithms is quite weak in the recently proposed IETR dataset. Towards this end, we propose a new completely blind IQA algorithm, which is based on the following key observations: 1. Distortions such as blurriness, blockiness (compression artifacts), and fast fading (object shifting) primarily affect the perceptual quality of 3D-synthesized views. 2. The perceptual characteristics of natural and synthetic synthesized views are quite different; distortions in natural views are perceptually more sensitive than the former. 3. Human Visual System's (HVS) ability to access the perceptual quality of an image also depends on some other properties of the images, such as contrast. All these observations are integrated into the proposed algorithm named Distortion-Specific Contrast-Based (DSCB) IQA. Various experiments validate that the proposed DSCB IQA efficiently competes with human perception and exhibits substantially better results (at least 17% gain in terms of PLCC) when compared to the existing NR IQAs. Sadbhawna, Vinit Jakhetiya, Deebha Mumtaz, Sunil Prasad Jaiswal |
MMSP | 2 |
| 2019 | Frequency-Domain Analysis Based Exploitation Of Color Channels For Color Image DemosaickingabstractColor-difference interpolation (CDI) has been a widely used technique for various color demosaicking methods. CDI-based methods perform interpolation in the color-difference domain assuming that the color-difference signal is a low-pass signal. Recently, a residual interpolation (RI) algorithm, which conducts interpolation in the residual domain, has been developed, and it assumes that the residual domain is flatter or smoother than the channel-difference domain. In this paper, we comprehensively show a frequency domain analysis of these assumptions and observe that it is image dependent and creates artifacts in the interpolated image. With this view, we propose an algorithm that uses the inter-color correlation as well as the residual smoothness among the different channel much better than the existing algorithms. Experimental results emphasize that the proposed algorithm atribute better performances the existing algorithms in terms of both visual and objective quality. Sunil Prasad Jaiswal, Vinit Jakhetiya, Ke Gu 0001, Sharath Chandra Guntuku, Ashutosh Singla |
VCIP | 2 |
| 2019 | A Highly Efficient Blind Image Quality Assessment Metric of 3-D Synthesized Images Using Outlier DetectionabstractWith multitudes of image processing applications, image quality assessment (IQA) has become a prerequisite for obtaining maximally distinctive statistics from images. Despite the widespread research in this domain over several years, existing IQA algorithms have a number of key limitations concerning different image distortion types and algorithms' computational efficiency. Images that are synthesized using depth image-based rendering have applications in various disciplines, such as free viewpoint videos, which enable synthesis of novel realistic images in the referenceless environment. In the literature, very few no-reference (NR) quality assessment metrics of three-dimensional (3-D) synthesized images are proposed, and most of them are computationally expensive, which makes it difficult for them to be deployed in real-time applications. In this paper, we attribute the geometrically distorted pixels as outliers in 3-D synthesized images. This assumption is validated using the three $sigma$ rule-based robust outlyingness ratio. We propose a novel fast and accurate blind IQA metric of 3-D synthesized images using nonlinear median filtering since the median filtering has the capability of identifying and removing outliers. The advantages of the proposed algorithm are twofold. First, it uses a simple technique, i.e., median filtering, to capture the level of geometric and structural distortions (up to some extend). Second, the proposed algorithm has higher computational efficiency. Experiments show the superiority of the proposed NR IQA algorithm over existing state-of-the-art full-, reduced-, and NR IQA methods, in terms of both predicting accuracy and computational complexity. Vinit Jakhetiya, Ke Gu 0001, Trisha Singhal, Sharath Chandra Guntuku, Zhifang Xia, Weisi Lin |
IEEE Trans. Ind. Informatics | 1 |
| 2018 | Just Noticeable Difference for natural images using RMS contrast and feed-back mechanism
Vinit Jakhetiya, Weisi Lin, Sunil Prasad Jaiswal, Ke Gu 0001, Sharath Chandra Guntuku |
Neurocomputing | 1 |
| 2018 | A Prediction Backed Model for Quality Assessment of Screen Content and 3-D Synthesized ImagesabstractIn this paper, we address problems associated with free-energy-principle-based image quality assessment (IQA) algorithms for objectively assessing the quality of Screen Content (SC) and three-dimensional (3-D) synthesized images and also propose a very fast and efficient IQA algorithm to address these issues. These algorithms separate an image into predicted and disorder residual parts and assume disorder residual part does not contribute much to the overall perceptual quality. These algorithms fail for quality estimation of SC images as information of textual regions in SC images are largely separated into the disorder residual part and less information in the predicted part and subsequently, given a negligible emphasis. However, this is in contrast with the characteristics of human vision. Since our eyes are well trained to detect text in daily life. So, our human vision has prior information about text regions and can sense small distortions in these regions. In this paper, we proposed a new reduced-reference IQA algorithm for SC images based upon a more perceptually relevant prediction model and distortion categorization, which overcomes problems with existing free-energy-principle-based predictors. From experiments, it is validated that the proposed model has a better capability of efficiently estimating the quality of SC images as compared to the recently developed reduced-reference IQA algorithms. We also applied the proposed algorithm to judge the quality of 3-D synthesized images and observed that it even achieves better performance than the full-reference IQA metrics specifically designed for the 3-D synthesized views. Vinit Jakhetiya, Ke Gu 0001, Weisi Lin, Qiaohong Li, Sunil Prasad Jaiswal |
IEEE Trans. Ind. Informatics | 1 |
| 2018 | Model-Based Referenceless Quality Metric of 3D Synthesized Images Using Local Image DescriptionabstractNew challenges have been brought out along with the emerging of 3D-related technologies, such as virtual reality, augmented reality (AR), and mixed reality. Free viewpoint video (FVV), due to its applications in remote surveillance, remote education, and so on, based on the flexible selection of direction and viewpoint, has been perceived as the development direction of next-generation video technologies and has drawn a wide range of researchers' attention. Since FVV images are synthesized via a depth image-based rendering (DIBR) procedure in the "blind" environment (without reference images), a reliable real-time blind quality evaluation and monitoring system is urgently required. But existing assessment metrics do not render human judgments faithfully mainly because geometric distortions are generated by DIBR. To this end, this paper proposes a novel referenceless quality metric of DIBR-synthesized images using the autoregression (AR)-based local image description. It was found that, after the AR prediction, the reconstructed error between a DIBR-synthesized image and its AR-predicted image can accurately capture the geometry distortion. The visual saliency is then leveraged to modify the proposed blind quality metric to a sizable margin. Experiments validate the superiority of our no-reference quality method as compared with prevailing full-, reduced-, and no-reference models. Ke Gu 0001, Vinit Jakhetiya, Junfei Qiao 0001, Xiaoli Li 0011, Weisi Lin, Daniel Thalmann |
IEEE Trans. Image Process. | 2 |
| 2017 | Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral CorrelationabstractColor filter array (CFA) interpolation, or three-band demosaicking, is a process of interpolating the missing color samples in each band to reconstruct a full color image. In this paper, we are concerned with the challenging problem of multispectral demosaicking, where each band is significantly undersampled due to the increment in the number of bands. Specifically, we demonstrate a frequency-domain analysis of the subsampled color-difference signal and observe that the conventional assumption of highly correlated spectral bands for estimating undersampled components is not precise. Instead, such a spectral correlation assumption is image dependent and rests on the aliasing interferences among the various color-difference spectra. To address this problem, we propose an adaptive spectral-correlation-based demosaicking (ASCD) algorithm that uses a novel anti-aliasing filter to suppress these interferences, and we then integrate it with an intra-prediction scheme to generate a more accurate prediction for the reconstructed image. Our ASCD is computationally very simple, and exploits the spectral correlation property much more effectively than the existing algorithms. Experimental results conducted on two data sets for multispectral demosaicking and one data set for CFA demosaicking demonstrate that the proposed ASCD outperforms the state-of-the-art algorithms. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Jiahao Pang, Klaus Mueller 0001, Oscar C. Au |
IEEE Trans. Image Process. | 3 |
| 2017 | Maximum a Posterior and Perceptually Motivated Reconstruction Algorithm: A Generic FrameworkabstractMost of the existing image reconstruction algorithms are application specific, and have generalization issues due to the need for parameter tuning and an unknown level of signal distortion. Addressing these problems, in this paper, we propose an efficient perceptually motivated and maximum a posterior (MAP)-based generic framework for image reconstruction. This can be applied to several image/video processing applications, where there is a necessity to improve reconstruction accuracy and suppress visible artifacts, such as denoising, deinterlacing, interpolation, de-blocking of Jpeg/Jpeg-2000, and demosaicing. The gradient magnitudes are noise insensitive to a moderate levels of noise and we propose to utilize this property for finding pixels with similar edge semantics in the neighborhood when neighboring pixels are noisy. With this view, we incorporate the gradient magnitude similarity based image quality assessment metric with the MAP estimation and, in turn, it can better approximate the variance of the MAP, as compared to nonlinear filters. The proposed generic algorithm (without manually tuning any parameters) is shown to produce a better quality of reconstruction when compared to the state-of-the-art application-specific algorithms, for most of the image processing applications. Vinit Jakhetiya, Weisi Lin, Sunil Prasad Jaiswal, Sharath Chandra Guntuku, Oscar C. Au |
IEEE Trans. Multim. | 1 |
| 2016 | Optimized high-frequency based interpolation for multispectral demosaickingabstractMultispectral demosaicking, which is an extension of color demosaicking, is a challenging problem because each band is significantly undersampled and thus precise reconstruction is needed for the restoration of high-frequency components, such as edges, textures etc. In general, existing algorithms borrow high-frequency information either from different bands via inter-color correlation or from within the bands, and produces artifact in the reconstructed image. To meet this inherent shortcoming, we propose to incorporate two different high-frequency components and integrate them optimally in the linear minimum mean square sense (LMMSE) for the precise reconstruction of undersampled components. Experimental results demonstrate that the proposed algorithm based on the optimized high-frequency achieves superior performance compared to existing algorithms both in terms of objective and subjective quality. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Manohar Kuse, Oscar C. Au |
ICIP | 3 |
| 2016 | Personalizing User Interfaces for improving quality of experience in VoD recommender systemsabstractRecommending content to users involves understanding a) what to present and b) how to present them, so as to increase quality of experience (QoE) and thereby, content consumption. This work attempts to address the question of how to present contents in a way so that the user finds it easy to get to desired content. While the process of User Interface (UI) design is dependent on several human factors, there are basic design components and their combination that have to be common to any recommender system user interface. Personalization of the UI design process involves picking the right components and their combination, and presenting a UI to suit the usage behavior of an individual user, so as to enhance the QoE. This work proposes a system that learns from a user's content consumption patterns and makes some recommendations regarding how to present the content for the user (in the context of Video-On-Demand/Live-TV services on Computer displays), so as to enhance the QoE of the recommender system. Sharath Chandra Guntuku, Sujoy Roy, Weisi Lin, Kelvin Ng, Wee Keong Ng, Vinit Jakhetiya |
QoMEX | 6 |
| 2014 | Fast and efficient intra-frame deinterlacing using observation model based bilateral filterabstractRecently, a few bilateral filter based interpolation and intraframe deinterlacing algorithms have been proposed, but these algorithms only use prior information (bilateral filter). In this paper, we propose an efficient and fast intra-frame deinterlacing algorithm using an observation model based bilateral filter (using both likelihood and prior information). Our proposed algorithm is also able to use approximated horizontal pixels for the deinterlacing, which results into the better prediction of the edges. From extensive experiments, it is observed that the proposed algorithm has the capability of provide satisfactory results in terms of both objective and subjective quality. Vinit Jakhetiya, Oscar C. Au, Sunil Prasad Jaiswal, Luheng Jia, Hong Zhang 0024 |
ICASSP | 1 |
| 2014 | Exploitation of inter-color correlation for color image demosaickingabstractImage demosaicking or color filter array interpolation is a process of interpolating missing color samples to reconstruct a full color image. In general, existing algorithms assume that the high frequency components such as edges, texture etc. of different color channels are similar and thus take an advantage of it to estimate the missing samples. In this paper, we efficiently analyze the relationships of intra and inter-color correlation among the channels and observe that such assumption fails in most cases. In view of this observation, we propose a scheme that exploits the correlation between different color channels much effectively than the existing algorithms. Experimental results demonstrate that the proposed algorithm outperforms the existing methods both in terms of Peak Signal to Noise Ratio (PSNR) and visual perception. Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuan Yuan 0002 |
ICIP | 3 |
| 2014 | Symmetrical predictor structure based integrated lossy, near lossless/lossless coding of imagesabstractPrediction based algorithms reported in the literature are not able to integrate lossy and near-lossless/lossless coding and uses only causal pixels (non-symmetrical predictor structure) for prediction. A non-symmetrical predictor structure, however, is not able to efficiently adapt near the intensity varying areas, which results into poor prediction. Hence, we propose a novel two-stage algorithm for lossy, near lossless/lossless compression using a symmetrical predictor structure is proposed. In the first stage, the proposed algorithm encodes and decodes the given image using the JPEG-2000 standard algorithm (lossy coding). This JPEG-2000 decoded image in the first stage, enables us to use the symmetrical predictor (using both causal and non-causal pixels) for prediction in the second stage. A performance evaluation shows that our algorithm is significantly better in terms of compression performance as compared to some of the computationally complex methods. Vinit Jakhetiya, Oscar C. Au, Sunil Prasad Jaiswal, Luheng Jia, Gaurav Mittal |
ISCAS | 1 |
| 2014 | Adaptive Predictor Structure Based Interpolation for Reversible Data Hiding
Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuanfang Guo, Anil Kumar Tiwari |
IWDW | 3 |
| 2013 | Efficient adaptive prediction based reversible image watermarkingabstractIn this paper, we propose a new reversible watermarking algorithm based on additive prediction-error expansion which can recover original image after extracting the hidden data. Embedding capacity of such algorithms depend on the prediction accuracy of the predictor. We observed that the performance of a predictor based on full context prediction is preciser as compared to that of partial context prediction. In view of this observation, we propose an efficient adaptive prediction (EAP) method based on full context, that exploits local characteristics of neighboring pixels much effectively than other prediction methods reported in literature. Experimental results demonstrate that the proposed algorithm has a better embedding capacity and also gives better Peak Signal to Noise Ratio (PSNR) as compared to state-of-the-art reversible watermarking schemes. Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuanfang Guo, Anil Kumar Tiwari, Yue Kong |
ICIP | 3 |
| 2013 | Reconfigurable hardware-friendly CU-group based merge/skip mode for high efficient video codingabstractMerge/skip mode is one of the most important inter prediction tools adopted in the High Efficiency Video Coding (HEVC) standard which is the state-of-the-art video coding standard. It is very efficient in reducing the side information for the blocks within the same object. However, it is difficult for parallel encoding and decoding due to the data dependency problem between neighboring prediction units (PU). Furthermore, different shapes and positions of PUs would result in different definition of the merge/skip candidate list (MCL), which would lead to potentially extra hardware cost and is not easy to be efficiently implemented by the hardware. To deal with this problem, two reconfigurable hardware-friendly MCL construction schemes are proposed in this paper. The first scheme which is called unified MCL (UMCL) uses one candidate list for all PUs inside the motion estimation region (MER), which is regarded as the basic parallel processing unit for the hardware realization. The second scheme which is named boundary MCL (BMCL) allows different candidate lists for the PUs on the boundary of MER. Both of the two schemes can have flexible parallel degree based on the requirement specification. Experimental results show that UMCL reduces the hardware complexity significantly with little coding performance degradation and BMCL achieves significant coding gain while maintaining the hardware complexity. Wei Dai 0002, Oscar C. Au, Feng Zou 0006, Vinit Jakhetiya |
MMSP | 7 |
| 2012 | Adaptive Predictor Structures for Lossless Compression of VideosabstractIn this paper, we propose a prediction algorithm that uses adaptive predictor structures for lossless video coding. The proposed encoder finds cross-correlation coefficient (Rxy) between current frame and motion compensated previous frame and classifies the coefficient value into a small number of bins. For general videos, we propose four bins and associate different predictor structures with each of the bins. Similarly for medical videos, numbers of bins are three and each of these bins is associated with different predictor structures. Ashutosh Singla, Jaya Shukla, Anil Kumar Tiwari, Sunil Prasad Jaiswal, Vinit Jakhetiya |
DCC | 5 |
| 2012 | Interpolation based symmetrical predictor structure for lossless image codingabstractPredictor based algorithms reported in literature uses only causal pixels and hence a non-symmetrical predictor structure for prediction. We observed that the performance of predictor is highly dependent on the predictor structure used. In view of this, we propose a novel interpolation based prediction scheme that enables us to use symmetrical predictor structure. In this sense, we have also used non causal pixels in our scheme. Also, from various interpolation algorithms available, we selected a simple one to ensure decoder simplicity, without any significant loss in performance. From performance evaluation, we found that our algorithm is significantly better in terms of compression performance as compared to some of the computationally complex methods. Vinit Jakhetiya, Sunil Prasad Jaiswal, Anil Kumar Tiwari, Oscar C. Au |
ISCAS | 1 |
| 2012 | Bit-depth expansion using Minimum Risk Based ClassificationabstractBit-depth expansion is an art of converting low bit-depth image into high bit-depth image. Bit-depth of an image represents the number of bits required to represent an intensity value of the image. Bit-depth expansion is an important field since it directly affects the display quality. In this paper, we propose a novel method for bit-depth expansion which uses Minimum Risk Based Classification to create high bit-depth image. Blurring and other annoying artifacts are lowered in this method. Our method gives better objective (PSNR) and superior visual quality as compared to recently developed bit-depth expansion algorithms. Gaurav Mittal, Vinit Jakhetiya, Sunil Prasad Jaiswal, Oscar C. Au, Anil Kumar Tiwari, Dai Wei |
VCIP | 2 |
| 2011 | An efficient image interpolation algorithm based upon the switching and self learned characteristics for natural imagesabstractThis paper presents a new image interpolation technique for enhancement of spatial resolution of images. The proposed algorithm uses the switching of existing Soft-decision Adaptive Interpolation (SAI) algorithm and Single Pass Interpolation Algorithm (SPIA) methods. We learn the error pattern in the interpolation process of SAI method and SPIA Method after interpolating downsampled version of LR image. Then we deviced a mechanism to correct the error pattern. Emperically we found that SAI methods works better on smooth images (variation among the pixels is less) while SPIA method works better on detailed images (more variation among the pixels), because of the type of pixels used in the interpolation. So, a hybrid scheme of combining SAI method and SPIA method is proposed for best prediction of high resolution (HR) image. The proposed algorithm produces the best results in different varieties of images in terms of both PSNR measurement and subjective visual quality. Sunil Prasad Jaiswal, Vinit Jakhetiya, Anil Kumar Tiwari |
ISCAS | 2 |