EDBT 2026 Demo / reviewers in the wild / expert
Badri N. Subudhi
dblp:14/6448 · also Badri Narayan Subudhi
· DBLP profile ↗
52ranked-venue papers
13as first author
32since 2021 · last 2026
0000-0002-4378-0065ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph refactored domain adversarial learning for underwater image enhancement
Meghna Kapoor, Badri N. Subudhi, Thierry Bouwmans, Ankur Bansal |
Pattern Recognit. Lett. | 2 |
| 2025 | T-GAP: Temporal Granularity Aware Projection Network for Action LocalizationabstractAction recognition and localization in videos pose significant challenges, as they require identifying temporal boundaries and classifying actions within long video sequences. This paper introduces an anchor-free, single-stage framework that predicts both the start and end times of actions while simultaneously classifying the actions, framing the task as a sequence labeling problem. The proposed model utilizes an encoder-decoder architecture with LSTM projections and 1D convolutions to capture rich temporal dependencies. This is followed by hierarchical feature pyramid generation, which is refined using Pyramidal Pooling Aggregation (PPA) and enhanced through a Progressive Feature Enhancement Unit (PFEU) with dilated convolutions to preserve the contextual relationships present in complex video frames. We also propose a novel Temporal Granularity Convolution (TGC) Layer, which refines temporal features. The TGC Layer captures various temporal details using a multi-branch structure consisting of Fine-Grained and Coarse-Grained Temporal Branches, designed to efficiently handle different temporal granularities. Finally, a decoder utilizes these multi-scale features to predict action instances at multiple levels of granularity. The model simultaneously integrates classification and regression heads to predict action labels and temporal boundaries, achieving improved localization and recognition. We have demonstrated the effectiveness of the proposed scheme, using mean average precision (mAP) across various Intersection over Union (IoU) thresholds on the Thumos14, ActivityNet-1.3, and MultiTHUMOS datasets over multiple state-of-the-art (SOTA) methods. Himanshu Singh 0006, Avijit Dey, Badri N. Subudhi, Vinit Jakhetiya, Veerakumar Thangaraj |
AVSS | 3 |
| 2025 | Graph Refinement in Latent Space: A Hypergraph Convolution for Underwater Object DetectionabstractUnderwater object detection presents significant challenges due to the intrinsic properties of light in aquatic environments. State-of-the-art methods often fail to capture the subtle details necessary for accurate detection in these scenarios. Recent advancements have shown promising results by reformulating relationships in graph space; however, most SOTA methods typically employ graph structures that are insufficient to represent the complex latent variables inherent in underwater environments. Hence, these models are unable to preserve the actual boundaries of object detection. To address these limitations, in this paper, we propose a novel end-to-end architecture that uses a graph refactoring aware deep learning based encoder-decoder architecture. The proposed approach uses a convolutional backbone to project the image into latent space, where an unsupervised initial graph is constructed. The hypergraph convolution is then utilized to optimize message passing between graph nodes, enhancing the representation of complex relationships of latent space. This helps in the retention of intricate details by modelling two or more latent variables as hyperedge by sharing the information among themselves. Finally, an image generation module maps the enhanced graph representation back to image space. The effectiveness of the proposed method is demonstrated through a comparative analysis against fourteen state-of-the-art methods on the benchmark underwater databases. The code for the proposed scheme can be found at https://github.com/immkapoor/hyper_graph. Meghna Kapoor, Badri N. Subudhi, Ankur Bansal |
ICASSP | 2 |
| 2025 | Feature Affinity based Clustering for Test-Time Adaptation for Image Quality AssessmentabstractRecently, Test-Time Adaptation (TTA) algorithms have gained traction in image/video quality assessment (IQA/VQA). These methods use virtual losses, like group contrastive and rank loss, as auxiliary tasks to adapt batch normalization parameters, helping models generalize better to distribution shifts between training and testing datasets. Group contrastive loss clusters images into low- and high-quality groups based on predicted quality scores, maximizing feature space distance between them. However, its effectiveness relies on the base model’s ability to accurately predict quality scores, which can be compromised when distribution shifts occur, leading to suboptimal adaptation and degraded performance. We propose a novel clustering approach based on the assumption that high-quality images contain richer high-level information, which is extracted using a pre-trained VGG-16 model. Images are clustered by comparing the VGG-16 features of the highest quality image in a batch with the others, enabling effective grouping based on feature affinities. These accurate clusters enhance the computation of contrastive loss, improving the adaptation of batch normalization layers. Additionally, we introduce an adaptive rank loss to reduce the impact of rank loss when the base model can distinguish images with varying distortion levels. Experimental results across multiple image quality assessment datasets, including LIVE, CID-2013, KONIQ-10K, and SPAQ, as well as algorithms like MetaIQA, HyperIQA, TReS, and MUSIQ, show that the proposed method consistently performs better than the existing Test-Time Adaptation (TTA) approach. Meghna Kapoor, Vinit Jakhetiya, Badri N. Subudhi, Ankur Bansal, Weisi Lin |
ICME | 3 |
| 2025 | Graph-based Moving Object Segmentation for underwater videos using semi-supervised learningabstractInternational audience Meghna Kapoor, Wieke Prummel, Jhony-Heriberto Giraldo-Zuluaga, Badri N. Subudhi, Anastasia Zakharova, Thierry Bouwmans, Ankur Bansal |
Comput. Vis. Image Underst. | 4 |
| 2025 | Two streams ResNet-50 network for infrared and visible image fusion
Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya |
Multim. Tools Appl. | 2 |
| 2025 | Underwater surveillance using spatially curated perceptual loss and graph refactored network
Meghna Kapoor, Bhargava N. Satya, Badri N. Subudhi, Vinit Jakhetiya, Ankur Bansal |
Pattern Recognit. | 3 |
| 2024 | Underwater Change Detection Using Multiple Sampling-Based Probabilistic Learner and Feature Preservance DiscriminatorabstractSurveillance can be defined as the process of monitoring the behavior and activities of different objects to generate meaningful insights into a video scene. In the context of underwater, surveillance can be elucidated as one of the processes of detecting and tracking the moving objects present in underwater videos. Many methods have been put forth to separate moving objects from underwater environments. Nevertheless, such methods cannot maintain the minute details that are crucial for determining an object’s boundary. This is mainly due to the intricate natural properties of water and some of its characteristics, such as excessive turbidity, scattering, low visibility, etc. In this regard, we put forth an adversarial learning-based end-to-end deep learning architecture to detect underwater moving objects. The proposed architecture uses two modules for underwater object detection. The initial module is a generator comprised of a probabilistic learner which is based on multiple down-sampling and up-sampling modules. Further, the discriminator network is composed of a multi-level feature concatenation module which can perpetuate specifics at distinct levels. The effectiveness of the proposed method is confirmed using two underwater benchmark datasets by contrasting its outcomes with those of eight state-of-the-art methods. Mehvish Nissar, Badri N. Subudhi, Vinit Jakhetiya, Amit Mishra 0004 |
ICIP | 2 |
| 2024 | Principal Graph Neighborhood Aggregation for Underwater Moving Object Detection
Meghna Kapoor, Badri N. Subudhi, Vinit Jakhetiya, Ankur Bansal |
ICPR (29) | 2 |
| 2024 | Project and Pool: An Action Localization Network for Localizing Actions in Untrimmed Videos
Himanshu Singh 0006, Avijit Dey, Badri N. Subudhi, Vinit Jakhetiya |
ICPR (29) | 3 |
| 2024 | A Multi-Scale Contrast Preserving Encoder-Decoder Architecture for Local Change Detection From Thermal Video ScenesabstractThis article presents a new deep-learning architecture based on an encoder-decoder framework that retains contrast while performing background subtraction (BS) on thermal videos. The proposed scheme consists of three consecutive blocks: the encoder, the Multi-Scale Contrast Preservation (MSCP) block, and the decoder. The encoder network employs a hybrid of convolution and atrous convolution blocks to preserve both sparse and dense features, with a skip connection. The encoder, combined with the MSCP block, maintains multi-scale contrast features with reduced training loss. Furthermore, the decoder network accurately projects the extracted features at different layers into pixel-level detail. The proposed end-to-end model efficiently provides a binary map for the corresponding thermal video scene. The efficiency of the proposed algorithm is validated on two large-scale datasets, namely CDnet 2014 and the Tripura University Video Dataset at Night Time (TU-VDN). Both qualitative and quantitative results demonstrate that MSCP outperforms thirty-eight existing BS schemes. Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya, Thierry Bouwmans |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Non-Subsampled Contourlet Transform and Ground-Truth Score Generation Based Quality Assessment for DIBR-Synthesized ViewsabstractIn recent years, there have been advancements in developing Depth-Image-Based Rendering (DIBR) views. However, the quality of these synthesized views is often degraded by inefficient in-painting techniques and synthesis procedures, leading to geometric and structural distortions. This paper introduces two novel approaches to evaluate the quality of DIBR synthesized views, using full reference (FR) and no-reference (NR) metrics. The proposed FR quality assessment (QA) metric is based on the observation that the deep features of the Non-Subsampled Contourlet Transform (NSCT) maps capture the perceptually important characteristics of the images. By calculating the difference between these deep feature vectors of the reference and distorted views, we determine the quality of the image. Moreover, a lot of existing NR metrics typically divide an image into blocks and assign the same subjective quality scores to each block for training a deep learning model. However, this approach is not suitable for DIBR synthesized views, as distortions are often localized in specific areas rather than affecting the entire view. Consequently, the performance of existing block-based deep-learning algorithms suffers due to the absence of accurate ground truth scores for each image block. To address this limitation, this work proposes an innovative method for calculating ground truth scores for individual image blocks. This process is similar to the proposed FR metric. Firstly, we obtain the deep features of NSCT map of an image block and the quality score for each block is calculated using its and the reference block's feature vector. These block-wise ground truth scores are used to train a deep learning model which serves as an NR metric for estimating the quality of a given test block. Finally, the predicted block-level quality values are aggregated to determine the overall quality of the entire image. Experimental results demonstrate that both the proposed algorithms perform better than the existing objective metrics for DIBR synthesized views. Deebha Mumtaz, Sadbhawna, Vinit Jakhetiya, Badri N. Subudhi, Weisi Lin |
IEEE Trans. Multim. | 4 |
| 2024 | Bayesian's probabilistic strategy for feature fusion from visible and infrared images
Veerakumar Thangaraj, Badri N. Subudhi, Vinit Jakhetiya |
Vis. Comput. | 3 |
| 2023 | Two-Streams: Dark and Light Networks with Graph Convolution for Action Recognition from Dark Videos (Student Abstract)abstractIn this article, we propose a two-stream action recognition technique for recognizing human actions from dark videos. The proposed action recognition network consists of an image enhancement network with Self-Calibrated Illumination (SCI) module, followed by a two-stream action recognition network. We have used R(2+1)D as a feature extractor for both streams with shared weights. Graph Convolutional Network (GCN), a temporal graph encoder is utilized to enhance the obtained features which are then further fed to a classification head to recognize the actions in a video. The experimental results are presented on the recent benchmark ``ARID" dark-video database. Saurabh Suman, Nilay Naharas, Badri N. Subudhi, Vinit Jakhetiya |
AAAI | 3 |
| 2023 | Drone-vs-Bird: Drone Detection Using YOLOv7 with CSRT TrackerabstractDrone detection has become a critical component of surveillance, but detecting small drones remains a challenge. In this work, we introduce a new drone detection scheme that uses a customized YOLOv7 combined with a tracker. Using simple YOLOv7 has some limitations, including false positives and difficulty detecting drones in complex environments. To address these issues, we propose two solutions. First, we use a simple method based on the probability or confidence level of detected objects to reduce the number of false positives. Second, we integrate YOLOv7 with object tracking to detect drones in complex environments. Our proposed method was successful in achieving a top-5 ranking in the 6thedition of the Drone-vs-Bird challenge, which was organized by the ICASSP Signal Processing Grand Challenge in 2023. Sahaj Mistry, Shreyas Chatterjee, Ajeet Kumar Verma, Vinit Jakhetiya, Badri N. Subudhi, Sunil Prasad Jaiswal |
ICASSP | 5 |
| 2023 | Measuring Underwater Image Quality with Depthwise-Separable Convolutions and Global-Sparse AttentionabstractUnderwater Image Quality Assessment (UIQA) poses specific challenges due to blur, dispersion of light, and color garbling caused by water turbulence. Images captured by Autonomous Underwater Vehicles (AVS) correlate poorly along their RGB channels, as blue and green channel components dominate. We propose a novel lightweight no-reference (NR) UIQA architecture based on depthwise-separable convolutions and global sparse attention that surpasses the existing deep learning architectures in terms of performance as well as parameter count and Multiply-Accumulate operations (MACs). Proposed model gives 3.64% increase in SRCC and 6.11% increase in PLCC in comparison to state-of-the-art for UEQAB dataset. Our work also illustrates the effectiveness of depthwise separable convolutions for underwater image quality assessment. Our empirical analysis confirms less redundancy among the channels of a feature map when depthwise separable convolutions are used compared to standard convolutions by increasing representational efficiency. Arpita Nema, Vinit Jakhetiya, Sunil Prasad Jaiswal, Badri N. Subudhi, Sharath Chandra Guntuku |
MMSP | 4 |
| 2023 | Kernel-Induced Possibilistic Fuzzy Associate Background Subtraction for Video SceneabstractThe background subtraction (BGS) technique is popularly used for many surveillance systems, segmenting the foreground by subtracting the modeled background from the image sequences. The effectiveness of any BGS technique depends on the robustness of the constructed background model. It is to be noted that many BGS schemes are affected by the inclusion of either noisy pixels in background construction or parameters of generative models. In this regard, we propound an idea of a kernel-induced possibilistic fuzzy associated BGS scheme for local change detection from a fixed camera captured sequence. The proposed scheme follows two stages: background training and foreground segmentation. In the background construction stage, each pixel is modeled using a possibilistic fuzzy cost function in kernel-induced space. The use of the induced kernel function will project the low-dimensional data into a higher dimensional space and the use of the possibilistic function will construct a robust background model based on the density of the data in the temporal domain avoiding the noisy and outlier points. The performance of the proposed scheme is tested on three benchmark databases. The effectiveness of the proposed scheme is evaluated on different performance evaluation measures: precision, recall, F-measure, and average similarity. We corroborate our findings by comparing them against 19 state-of-the-art existing BGS techniques. Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya, Esakkirajan Sankaralingam |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Context Region Identification Based Quality Assessment of 3D Synthesized ViewsabstractPerceptual quality assessment of 3D synthesized views is an open research problem in computer vision. Researchers across the globe have developed several algorithms to identify distortions. At the same time, the existing algorithms cannot quantify the context in which these distortions affect the overall perceptual quality. According to the recently proposed 3D view synthesis algorithm, the choice of context region for the disocclusion plays a vital role in predicting the quality of 3D views. The context region taken from the background of a view produces a perceptually better quality of 3D synthesized views than when the context region is taken from the foreground. With this view, the proposed algorithm aims to identify the context region and incorporate this information for the perceptual quality assessment of 3D synthesized views. We observed that the depth energy maps of the 3D synthesized views vary significantly with the change in the context region and subsequently can identify the context region. Hence, in this work, we propose a new and efficient quality assessment algorithm based upon the variation in the depth of 3D synthesized and reference views, giving two-fold advantages: 1. It can predict the quality based on whether the context region is foreground or not. 2. It is also able to suggest the possible location of distortions. We have proposed two new algorithms for both situations when the context region is foreground or not. The overall predicted score is the direct multiplication of the quality score estimated when the context region is foreground or not. When applied to the established benchmark dataset, the proposed technique performs satisfactorily with the PLCC of 0.7707 and 0.7572 of SRCC. Also, the proposed algorithm can work as a plug-in to improve the performance of the existing algorithms. Sadbhawna Thakur, Vinit Jakhetiya, Badri N. Subudhi, Sunil Prasad Jaiswal, Leida Li, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2022 | Do We Need a New Large-Scale Quality Assessment Database for Generative Inpainting Based 3D View Synthesis? (Student Abstract)abstractThe advancement in Image-to-Image translation techniques using generative Deep Learning-based approaches has shown promising results for the challenging task of inpainting-based 3D view synthesis. At the same time, even the current 3D view synthesis methods often create distorted structures or blurry textures inconsistent with surrounding areas. We analyzed the recently proposed algorithms for inpainting-based 3D view synthesis and observed that these algorithms no longer produce stretching and black holes. However, the existing databases such as IETR, IRCCyN, and IVY have 3D-generated views with these artifacts. This observation suggests that the existing 3D view synthesis quality assessment algorithms can not judge the quality of most recent 3D synthesized views. With this view, through this abstract, we analyze the need for a new large-scale database and a new perceptual quality metric oriented for 3D views using a test dataset. Sadbhawna, Vinit Jakhetiya, Badri N. Subudhi, Harshit Shakya, Deebha Mumtaz |
AAAI | 3 |
| 2022 | C3D and Localization Model for Locating and Recognizing the Actions from Untrimmed Videos (Student Abstract)abstractIn this article, we proposed a technique for action localization and recognition from long untrimmed videos. It consists of C3D CNN model followed by the action mining using the localization model, where the KNN classifier is used. We segment the video into expressible sub-action known as action-bytes. The pseudo labels have been used to train the localization model, which makes the trimmed videos untrimmed for action-bytes. We present experimental results on the recent benchmark trimmed video dataset “Thumos14”. Himanshu Singh 0006, Tirupati Pallewad, Badri N. Subudhi, Vinit Jakhetiya |
AAAI | 3 |
| 2022 | An End to End Encoder-Decoder Network with Multi-scale Feature Pulling for Detecting Local Changes From Video SceneabstractLocal change detection for moving object detection is an essential step in any computer vision task. The most well-known technique is background subtraction BGS. However, the performance of BGS is strongly dependent on the background construction. The background construction to be robust in the presence of various challenges: dynamic backgrounds, illumination changes, camera jitter, etc. In this paper, we propose a novel encoder-decoder-based end-to-end deep learning framework for BGS. Thus, we explore a VGG-19 deep network with a transfer learning strategy as an encoder that deeply learned and extracted the features at different levels. We herewith propose a Multi-scale Feature Pulling MFP block which can retain the features at the various scales of the challenging video scenes. We also design a decoder network which is a stack of several transposed convolutional layers which precisely predict that each pixel of the target frame belongs to the background or foreground. The efficiency of the proposed algorithm is validated on the CDNet-2014 dataset by comparing its results against seventeen state-of-the-art techniques. Badri N. Subudhi, Thierry Bouwmans, Vinit Jakheytiya, Veerakumar Thangaraj |
AVSS | 2 |
| 2022 | Encoder and decoder network with ResNet-50 and global average feature pooling for local change detection
Akhilesh Sharma, Vatsalya Bajpai, Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya |
Comput. Vis. Image Underst. | 4 |
| 2022 | Multiresolution visual enhancement of hazy underwater scene
Deepak Kumar Rout, Badri N. Subudhi, Veerakumar Thangaraj, Santanu Chaudhury, John J. Soraghan |
Multim. Tools Appl. | 2 |
| 2022 | Nonintrusive Perceptual Audio Quality Assessment for User-Generated Content Using Deep LearningabstractWith the boom of social media communication, teleconferencing, and online classes, audiovisual communication over bandwidth strained networks has become an integral part of our lives. Consequently, the growing demand for the quality of experience necessitates developing algorithms to measure and enrich user experience. Prior studies have mainly focused on assessing speech quality and intelligibility with reference to audio quality assessment, while other categories in user-generated multimedia (UGM) are less explored. Moreover, frequency-domain properties of speech and UGM audio are significantly different from each other. Furthermore, there is a lack of a standard dataset for the quality assessment of UGM. Considering these limitations, in this article, we first develop the IIT-JMU-UGM audio dataset consisting of 1150 audio clips, with diverse context, content, and types of degradation commonly observed in real-world scenarios and annotated with the subjective quality scores. Finally, we propose a non-intrusive audio quality assessment metric using a stacked gated-recurrent-unit-based deep learning framework. The proposed model outperforms several baseline methods, including state-of-the-art non-intrusive and intrusive approaches. The resulting Pearson’s correlation coefficient of 0.834 indicates that the proposed method efficiently mirrors human auditory perception. Deebha Mumtaz, Vinit Jakhetiya, Karan Nathwani, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Perceptually Unimportant Information Reduction and Cosine Similarity-Based Quality Assessment of 3D-Synthesized ImagesabstractQuality assessment of 3D-synthesized images has traditionally been based on detecting specific categories of distortions such as stretching, black-holes, blurring, etc. However, such approaches have limitations in accurately detecting distortions entirely in 3D synthesized images affecting their performance. This work proposes an algorithm to efficiently detect the distortions and subsequently evaluate the perceptual quality of 3D synthesized images. The process of generation of 3D synthesized images produces a few pixel shift between reference and 3D synthesized image, and hence they are not properly aligned with each other. To address this, we propose using morphological operation (opening) in the residual image to reduce perceptually unimportant information between the reference and the distorted 3D synthesized image. The residual image suppresses the perceptually unimportant information and highlights the geometric distortions which significantly affect the overall quality of 3D synthesized images. We utilized the information present in the residual image to quantify the perceptual quality measure and named this algorithm as Perceptually Unimportant Information Reduction (PU-IR) algorithm. At the same time, the residual image cannot capture the minor structural and geometric distortions due to the usage of erosion operation. To address this, we extract the perceptually important deep features from the pre-trained VGG-16 architectures on the Laplacian pyramid. The distortions in 3D synthesized images are present in patches, and the human visual system perceives even the small levels of these distortions. With this view, to compare these deep features between reference and distorted image, we propose using cosine similarity and named this algorithm as Deep Features extraction and comparison using Cosine Similarity (DF-CS) algorithm. The cosine similarity is based upon their similarity rather than computing the magnitude of the difference of deep features. Finally, the pooling is done to obtain the objective quality scores using simple multiplication to both PU-IR and DF-CS algorithms. Our source code is available online: https://github.com/sadbhawnathakur/3D-Image-Quality-Assessment. Sadbhawna, Vinit Jakhetiya, Shubham Chaudhary 0005, Badri N. Subudhi, Weisi Lin, Sharath Chandra Guntuku |
IEEE Trans. Image Process. | 4 |
| 2022 | Stretching Artifacts Identification for Quality Assessment of 3D-Synthesized ViewsabstractExisting Quality Assessment (QA) algorithms consider identifying "black-holes" to assess perceptual quality of 3D-synthesized views. However, advancements in rendering and inpainting techniques have made black-hole artifacts near obsolete. Further, 3D-synthesized views frequently suffer from stretching artifacts due to occlusion that in turn affect perceptual quality. Existing QA algorithms are found to be inefficient in identifying these artifacts, as has been seen by their performance on the IETR dataset. We found, empirically, that there is a relationship between the number of blocks with stretching artifacts in view and the overall perceptual quality. Building on this observation, we propose a Convolutional Neural Network (CNN) based algorithm that identifies the blocks with stretching artifacts and incorporates the number of blocks with the stretching artifacts to predict the quality of 3D-synthesized views. To address the challenge with existing 3D-synthesized views dataset, which has few samples, we collect images from other related datasets to increase the sample size and increase generalization while training our proposed CNN-based algorithm. The proposed algorithm identifies blocks with stretching distortions and subsequently fuses them to predict perceptual quality without reference, achieving improvement in performance compared to existing no-reference QA algorithms that are not trained on the IETR dataset. The proposed algorithm can also identify the blocks with stretching artifacts efficiently, which can further be used in downstream applications to improve the quality of 3D views. Our source code is available online: https://github.com/sadbhawnathakur/3D-Image-Quality-Assessment. Sadbhawna, Vinit Jakhetiya, Deebha Mumtaz, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Image Process. | 4 |
| 2022 | A Fully Automatic Feature-Based Real-Time Traffic Surveillance System Using Data Association in the Probabilistic FrameworkabstractMulti-object tracking involves maintaining several trajectories of different objects moving in the scene throughout the video. With this objective, in this article, a fully automatic and real-time tracking algorithm to track multiple vehicles in a video is proposed. The proposed method specifically tries to address the challenges of occlusion and fast-motion in traffic surveillance. The algorithm begins with automatic detection of the moving targets by an adaptive GMM-based background subtraction method. The trajectories of these detected targets are then built using a three-level multi-motion modeled particle filter framework which allows to deal with the challenges of occlusion and fast-motions of the target. The likelihood model for targets is based on their color distribution and edge oriented histogram features. It is contended that the color distribution feature, which can represent the target appearance and the edge oriented histogram, which can describe the target structure are sufficient to represent it in an unique feature space. Based on the similarity of target likelihood, the locations of the targets are filtered. These filtered locations are then associated with the most likely detections using the proposed low-cost and fast data association algorithm based on Euclidean distance and prevailing motion vector. The performance evaluation of the proposed scheme is carried out based on the six measures: Multi Object Tracking Accuracy, Multi Object Tracking Precision, Mostly Tracked trajectories, Mostly Lost trajectories, Identity Switches and Frames Per Second. The results evaluated on stationary camera shot sequences from benchmark datasets as well as real-time shot videos indicate that the proposed algorithm ensures robust tracking and can be used effectively for real-time surveillance of highways. Pranab Gajanan Bhat, Badri N. Subudhi, Veerakumar Thangaraj, Esakkirajan Sankaralingam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Detecting Covid-19 and Community Acquired Pneumonia Using Chest CT Scan Images With Deep LearningabstractWe propose a two-stage Convolutional Neural Network (CNN) based classification framework for detecting COVID-19 and Community Acquired Pneumonia (CAP) using the chest Computed Tomography (CT) scan images. In the first stage, an infection - COVID-19 or CAP, is detected using a pre-trained DenseNet architecture. Then, in the second stage, a fine-grained three-way classification is done using EfficientNet architecture. The proposed COVID+CAP-CNN framework achieved a slice-level classification accuracy of over 94% at identifying COVID-19 and CAP. Further, the proposed framework has the potential to be an initial screening tool for differential diagnosis of COVID-19 and CAP, achieving a validation accuracy of over 89.3% at the finer three-way COVID-19, CAP, and healthy classification. Within the IEEE ICASSP 2021 Signal Processing Grand Challenge (SPGC) on COVID-19 Diagnosis, our proposed two-stage classification framework achieved an overall accuracy of 90% and sensitivity of .857, .9, and .942 at distinguishing COVID-19, CAP, and normal individuals respectively, to rank first in the evaluation. Code and model weights are available at https://github.com/shubhamchaudhary2015/ct_covid19_cap_cnn Shubham Chaudhary 0005, Sadbhawna, Vinit Jakhetiya, Badri N. Subudhi, Ujjwal Baid, Sharath Chandra Guntuku |
ICASSP | 4 |
| 2021 | Perceptual Quality Assessment of DIBR Synthesized Views Using Saliency Based Deep FeaturesabstractIn recent years, Depth-Image-Based-Rendering (DIBR) synthesized views have gained popularity due to their numerous visual media applications. Consequently, the research in their quality assessment (QA) has also gained momentum. In this work, we propose an efficient metric to estimate the perceptual quality of DIBR synthesized views via the extraction of Deep-features. These Deep-features are extracted from a pretrained CNN model. Generally, in DIBR synthesized views, geometric distortions arise near the objects due to occlusion, and the human visual system is quite sensitive towards these objects. On the other end, saliency maps are efficiently able to highlight perceptually important objects. With this intuition, instead of extracting deep features directly from DIBR synthesized views, we obtain the refined feature vector from their corresponding saliency maps. Also, most of the pixels with geometric distortions have a nearly similar impact on the perceptual quality of 3D synthesized views. Considering this, we propose to fuse the feature maps using the cosine similarity measure based upon the deviation of one feature vector from another. It may also be emphasized that no training is performed in the proposed algorithm, and all the features are extracted from the pre-trained vanilla VGG-16 architecture. The proposed metric, when applied to the standard database, results in PLCC of 0.762 and SRCC equal to 0.7513, which is better than the existing state-of-the-art QA metrics. Shubham Chaudhary 0005, Alokendu Mazumder, Deebha Mumtaz, Vinit Jakhetiya, Badri N. Subudhi |
ICIP | 5 |
| 2021 | A hybrid BPSO-SVM for feature selection and classification of ocular healthabstractAbstract Glaucoma and diabetic retinopathy are the most common eye diseases and the leading cause of blindness around the world. The prime objective of this study is to devise and develop an experimental computer‐aided diagnosis system to provide an efficient way for assisting the ophthalmologist in early detection of ocular diseases such as glaucoma and diabetic retinopathy. The proposed technique follows three stages: Pre‐processing, feature selection and classification. Initially, the fundus image is pre‐processed to extract the green channel image, and the obtained green channel image is further enhanced using contrast limited adaptive histogram equalisation technique. Three different kinds of features: Clinical features, transform domain features and structural features are utilised to extract the relevant information from the enhanced fundus images. To avoid redundant information, an improved feature selection mechanism is used to select the optimum set of features from the extracted features. Subsequently, the selected features are used to train the support vector machine classifier for the classification of the retinal diseases with 10‐fold cross‐validation. The performance of the proposed method is assessed using eight different quantitative evaluation measures. The experimental results demonstrate the effectiveness of the proposed work over prior works for the early detection of ocular diseases. Balraj Keerthiveena, Esakkirajan Sankaralingam, Badri N. Subudhi, Veerakumar Thangaraj |
IET Image Process. | 3 |
| 2021 | Mixed Poisson Gaussian noise reduction in fluorescence microscopy images using modified structure of wavelet transformabstractAbstract Fluorescence microscopy is an important investigation tool of discoveries in the field of biological sciences where the imaging phenomena are limited by the noise. This paper introduces the integration of biorthogonal wavelet filters along with mixed Poisson‐Gaussian unbiased risk estimate (MPGURE) based subband adaptive thresholding function for the restoration of low photon count microscopy images. The proposed algorithm consists of four steps. In the first step, variance stabilization transform along with a multi‐scale Wiener filtering approach is used to filter out the noise and blurring effect. In the second step, deconvolved images are further decomposed by the biorthogonal wavelet filters. The modified wavelet subband structure is used for the identification of noisy and noise‐free subbands. In the next stage, different noisy coefficients are thresholded using the MPGURE‐based thresholding operation. The different thresholded images are combined along with different optimum coefficients. Finally, inverse variance stabilization transformation is applied to obtain the final restored output. Performance of the proposed algorithm is tested on 14 different benchmark image data sets with performance evaluation measures like signal‐to‐noise ratio, peak signal‐to‐noise ratio, mean structural similarity index measur, and correlation coefficient. Simulation results of the proposed algorithm claim better results than other state‐of‐the‐art techniques. Tushar Rasal, Veerakumar Thangaraj, Badri N. Subudhi, Esakkirajan Sankaralingam |
IET Image Process. | 3 |
| 2021 | Perceptual Quality Evaluation of Hazy Natural ImagesabstractHaze is an intrusion element that disrupts color fidelity and contrast of outdoor natural images, affecting their perceptual quality. The differential characteristics of hazy images compared to other natural images restrict the generalization of existing image quality assessment (IQA) algorithms. At the same time, efficient IQA algorithms for predicting the perceptual quality of naturally hazed images have not been proposed in the literature due to lack of a relevant dataset. To address this, we build the IIT-JMU Hazy Image Dataset comprising of 1000 high-definition hazy natural images consisting of diverse categories such as landscape, forests, roads, seascapes, and cityscapes, along with their subjective quality ratings. We present an analysis of existing natural-scene-statistics-based IQA algorithms on hazy natural images. In this article, we propose a convolutional-neural-network-based quality assessment algorithm for hazy natural images along with an IQA metric called deep learning-based haze perceptual quality evaluator (DLHPQE). The proposed DLHPQE efficiently predicts the perceptual quality of hazy natural images without a reference. Our results demonstrate that the DLHPQE outperforms existing state-of-the-art no-reference IQAs in terms of several performance parameters such as Pearson linear correlation coefficient, Spearman rank-order correlation coefficient, Kendall's rank-order correlation coefficient, and root-mean-square error. Palak Mahajan, Vinit Jakhetiya, Pawanesh Abrol, Parveen Lehana, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Edge Preserving Image Fusion using Intensity Variation ApproachabstractIn this article, a novel edge preserving image fusion method is proposed by merging multiple images captured from different imaging sensors. The objective of this paper is to highlight the informative contents of multiple images into a single fused image. Fusion of data from multiple sensors is a difficult task as the imaging modality are different and sensors capturing the data may be affected by sensors noise. As the images captured from multiple sensors possess uncertainty within a pixel due to the multi-valued level of brightness. It is obvious that a deterministic method of fusion may not give a better results. Hence, it is required to explore the use of fuzzy sets theoretic approaches in this regard. The proposed scheme follow three stages. In the first stage of the algorithm, a resultant image is obtained by setting the maximum value between the pixel intensity of visible and infrared sub-images considered within a small spatial neighborhood. The edges of the visible image are preserved in the second stage of the algorithm using a Fuzzy edge technique. Finally the fused image is obtained by combining the obtained resultant image and the edges of the visible image. In order to evaluate the performance of the proposed method quantitatively, and qualitatively experiments were carried out on publicly available benchmark database, "TNO-database". The proposed method is compared with those of eight state-of-the-arts techniques. The experimental results of the proposed method attained state-of-the-art performance in objective assessment and visual quality assessment. Badri N. Subudhi, Veerakumar Thangaraj, Manoj Singh Gaur |
TENCON | 2 |
| 2020 | Automatic lecture video skimming using shot categorization and contrast based features
Badri N. Subudhi, Veerakumar Thangaraj, Esakkirajan Sankaralingam, Santanu Chaudhury |
Expert Syst. Appl. | 1 |
| 2020 | An efficient framework for increasing image quality using DRN Bi-layer enfolded compressor
S. Tamboli Shabanam, V. R. Udupi, Badri N. Subudhi |
Multim. Tools Appl. | 3 |
| 2020 | Walsh-Hadamard-Kernel-Based Features in Particle Filter Framework for Underwater Object TrackingabstractOne of the well-established research domains among computer vision scientists is object tracking. However, not much work has been done in underwater scenarios. This article addresses the problem of visual tracking in the underwater environment with the stationary and nonstationary camera setups. In order to deal with the underwater optical dynamics, a dominant color component-based scene representation is employed in the YCbCr color space. An adaptive approach is devised to select the Walsh-Hadamard (WH) kernels for the efficient extraction of color, edge, and texture strengths, whereas a new feature called range strength is proposed to extract the variation of intensity from underwater sequences in the local neighborhood using the WH kernel. The likelihood of these feature strengths is integrated in a particle filter framework to track the object of interest in underwater sequences. The reference feature strengths used in assigning weights to the particles are updated based on the S$\phi$rensen distance. The coefficients of feature strengths are calculated in such a way that if one feature fails, then its coefficient become insignificant, whereas the more suitable features get higher feature coefficients. The effectiveness of the proposed scheme is evaluated using the underwater video datasets: reefVid, fish4knowledge (F4K), underwaterchangedetection (UWCD), and National Oceanic and Atmospheric Administration (NOAA). The performance evaluation is performed by comparing the scheme with five recent state-of-the-art tracking schemes. The quantitative analysis of the proposed scheme is carried out using three evaluation measures: overall intersection over union, centroid location error, and average tracking error. The performance of the proposed scheme is quite encouraging in the case of sequences with hazy and degraded, partially occluded, and camouflaged challenges. Deepak Kumar Rout, Badri N. Subudhi, Veerakumar Thangaraj, Santanu Chaudhury |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Kernelized Fuzzy Modal Variation for Local Change Detection From Video ScenesabstractBackground subtraction (BGS) is a popular scheme epitomized in the state-of-the-art literature on video processing. In this context, a novel online kernelized fuzzy modal variation based background subtraction scheme for detecting local changes from the sequences of image frames is proposed. In the proposed scheme, the time varying background at different instances of time are modeled using fuzzy set theory. The proposed background subtraction scheme, utilizes the fuzzy modal variation as the cost function for fitting the pixel values of the image frames. The use of kernel based modal variation helps in projecting the pixel values in a higher dimensional space, linearly separating them into object and background classes. The results of the proposed technique is verified on different challenging sequences including dynamic background, camera jitter, noise, blurred scene, etc. The proposed technique is successfully tested over several test sequences with two major databases (all sequences) and it provides better results compared to the twenty one existing state-of-the-art techniques. Badri N. Subudhi, Veerakumar Thangaraj, Esakkirajan Sankaralingam, Ashish Ghosh |
IEEE Trans. Multim. | 1 |
| 2019 | Empirical mode decomposition and adaptive bilateral filter approach for impulse noise removal
Veerakumar Thangaraj, Badri N. Subudhi, Esakkirajan Sankaralingam |
Expert Syst. Appl. | 2 |
| 2019 | Big data analytics for video surveillance
Badri N. Subudhi, Deepak Kumar Rout, Ashish Ghosh |
Multim. Tools Appl. | 1 |
| 2018 | Spatio-contextual Gaussian mixture model for local change detection in underwater video
Deepak Kumar Rout, Badri N. Subudhi, Veerakumar Thangaraj, Santanu Chaudhury |
Expert Syst. Appl. | 2 |
| 2017 | Context model based edge preservation filter for impulse noise removal
Veerakumar Thangaraj, Badri N. Subudhi, Esakkirajan Sankaralingam, Prasanta Kumar Pradhan |
Expert Syst. Appl. | 2 |
| 2017 | Moving object detection using spatio-temporal multilayer compound Markov Random Field and histogram thresholding based change detection
Badri N. Subudhi, Susmita Ghosh, Pradipta Kumar Nanda, Ashish Ghosh |
Multim. Tools Appl. | 1 |
| 2016 | Integration of fuzzy Markov random field and local information for separation of moving objects and shadows
Badri N. Subudhi, Susmita Ghosh, Sung-Bae Cho, Ashish Ghosh |
Inf. Sci. | 1 |
| 2016 | Statistical feature bag based background subtraction for local change detection
Badri N. Subudhi, Susmita Ghosh, Simon C. K. Shiu, Ashish Ghosh |
Inf. Sci. | 1 |
| 2015 | Application of Gibbs-Markov random field and Hopfield-type neural networks for detecting moving objects from video sequences captured by static camera
Badri N. Subudhi, Susmita Ghosh, Ashish Ghosh |
Soft Comput. | 1 |
| 2013 | Spatial constraint Hopfield-type neural networks for detecting changes in remotely sensed multitemporal imagesabstractIn this article a spatio-contextual unsupervised change detection technique for multitemporal, multi spectral remote sensing images is proposed. The technique uses a Gibbs Markov Random Field (GMRF) to model the spatial regularity between the neighboring pixels of the multitemporal difference image. The difference image is generated by change vector analysis (CVA) applied to images acquired on the same area at different times. The change detection problem is solved using the Maximum a posteriori probability (MAP) estimation principle. The MAP estimator of the GMRF used to model the difference image is exponential in nature, thus a modified Hopfield type neural network is exploited for estimating the MAP. In the considered Hopfield type network, a single neuron is assigned to each pixel of the difference image and is assumed to be connected only to its neighbors. Initial values of the neurons are set by histogram thresholding. An Expectation Maximization (EM) algorithm is used to estimate the GMRF model parameters. The proposed technique is validated by testing on different multispectral and multitemporal remote sensing images and compared with existing state-of-the-art techniques. Badri N. Subudhi, Susmita Ghosh, Ashish Ghosh |
ICIP | 1 |
| 2013 | Change detection for moving object segmentation with robust background construction under Wronskian framework
Badri N. Subudhi, Susmita Ghosh, Ashish Ghosh |
Mach. Vis. Appl. | 1 |
| 2013 | Integration of Gibbs Markov Random Field and Hopfield-Type Neural Networks for Unsupervised Change Detection in Remotely Sensed Multitemporal ImagesabstractIn this paper, a spatiocontextual unsupervised change detection technique for multitemporal, multispectral remote sensing images is proposed. The technique uses a Gibbs Markov random field (GMRF) to model the spatial regularity between the neighboring pixels of the multitemporal difference image. The difference image is generated by change vector analysis applied to images acquired on the same geographical area at different times. The change detection problem is solved using the maximum a posteriori probability (MAP) estimation principle. The MAP estimator of the GMRF used to model the difference image is exponential in nature, thus a modified Hopfield type neural network (HTNN) is exploited for estimating the MAP. In the considered Hopfield type network, a single neuron is assigned to each pixel of the difference image and is assumed to be connected only to its neighbors. Initial values of the neurons are set by histogram thresholding. An expectation-maximization algorithm is used to estimate the GMRF model parameters. Experiments are carried out on three-multispectral and multitemporal remote sensing images. Results of the proposed change detection scheme are compared with those of the manual-trial-and-error technique, automatic change detection scheme based on GMRF model and iterated conditional mode algorithm, a context sensitive change detection scheme based on HTNN, the GMRF model, and a graph-cut algorithm. A comparison points out that the proposed method provides more accurate change detection maps than other methods. Ashish Ghosh, Badri N. Subudhi, Lorenzo Bruzzone |
IEEE Trans. Image Process. | 2 |
| 2012 | Object and shadow separation using fuzzy Markov Random Field and local gray level co-occurence matrix based textural featuresabstractIn this article, we propose a novel object detection technique that can separate the moving object from its shadow. In this regard, we initially built a background model by taking median of pixel values in the temporal direction. To suppress the effects of quick change in illumination, and color frequency variation of the textured background, we have extracted the RGB color and ten local features at each pixel location in the target image and background model. For background separation, a difference image is generated by considering pixel by pixel absolute difference of the thirteen dimensional target image frame and the constructed background model. This is followed by a spatial Markov Random Field (MRF) constrained fuzzy clustering to find the moving regions in the target frame. The maximum a'posteriori probability (MAP) of the MRF constrained fuzzy clustering provides a binary image, where the moving objects with the moving cast shadow are identified as one group and the background is obtained as another group. To segment the moving objects from its shadow we explore a three stage shadow analysis technique. It uses analysis of rg color chrominance property of shadow, local gray level feature based shadow processing followed by boundary refinement to separate out the moving objects from its shadows. The performance of the proposed scheme is evaluated by comparing it with the state-of-the-art techniques. Badri N. Subudhi, Susmita Ghosh, Ashish Ghosh |
ISDA | 1 |
| 2012 | Object Detection From Videos Captured by Moving Camera by Fuzzy Edge Incorporated Markov Random Field and Local Histogram MatchingabstractIn this paper, we put forward a novel region matching-based motion estimation scheme to detect objects with accurate boundaries from videos captured by moving camera. Here, a fuzzy edge incorporated Markov random field (MRF) model is considered for spatial segmentation. The algorithm is able to identify even the blurred boundaries of objects in a scene. Expectation Maximization algorithm is used to estimate the MRF model parameters. To reduce the complexity of searching, a new scheme is proposed to get a rough idea of maximum possible shift of objects from one frame to another by finding the amount of shift in positions of the centroid. We propose a χ2-test-based local histogram matching scheme for detecting moving objects from complex scenes from low illumination environment and objects that change size from one frame to another. The proposed scheme is successfully applied for detecting moving objects from video sequences captured in both real-life and controlled environments. It is also noticed that the proposed scheme provides better results with less object background misclassification as compared to existing techniques. Ashish Ghosh, Badri N. Subudhi, Susmita Ghosh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Entropy based region selection for moving object detection
Badri N. Subudhi, Pradipta Kumar Nanda, Ashish Ghosh |
Pattern Recognit. Lett. | 1 |
| 2011 | A Change Information Based Fast Algorithm for Video Object Detection and TrackingabstractIn this paper, we present a novel algorithm for moving object detection and tracking. The proposed algorithm includes two schemes: one for spatio-temporal spatial segmentation and the other for temporal segmentation. A combination of these schemes is used to identify moving objects and to track them. A compound Markov random field (MRF) model is used as the prior image attribute model, which takes care of the spatial distribution of color, temporal color coherence and edge map in the temporal frames to obtain a spatio-temporal spatial segmentation. In this scheme, segmentation is considered as a pixel labeling problem and is solved using the maximum a posteriori probability (MAP) estimation technique. The MRF-MAP framework is computation intensive due to random initialization. To reduce this burden, we propose a change information based heuristic initialization technique. The scheme requires an initially segmented frame. For initial frame segmentation, compound MRF model is used to model attributes and MAP estimate is obtained by a hybrid algorithm [combination of both simulated annealing (SA) and iterative conditional mode (ICM)] that converges fast. For temporal segmentation, instead of using a gray level difference based change detection mask (CDM), we propose a CDM based on label difference of two frames. The proposed scheme resulted in less effect of silhouette. Further, a combination of both spatial and temporal segmentation process is used to detect the moving objects. Results of the proposed spatial segmentation approach are compared with those of JSEG method, and edgeless and edgebased approaches of segmentation. It is noticed that the proposed approach provides a better spatial segmentation compared to the other three methods. Badri N. Subudhi, Pradipta Kumar Nanda, Ashish Ghosh |
IEEE Trans. Circuits Syst. Video Technol. | 1 |