Prabir Kumar Biswas

dblp:b/PKBiswas · also Prabir K. Biswas · DBLP profile ↗
← Back
66ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0002-8922-1306ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 18 since 2021Artificial intelligence and machine learning · 25 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Fast, Simple, and Flexible Scale Informative Feature Transform Module for Arbitrary Scale Image Super-Resolution
abstract
The single-image super-resolution domain has witnessed a significant performance improvement due to the advancement of deep learning models. However, most of the deep learning models use an integer-scale and scale-specific model for super-resolution. Separate scale-specific networks require huge memory during deployment. Therefore, a single model for any random scale image super-resolution is an old age demand. Unlike existing solutions based on implicit representation functions, we propose a fully convolutional arbitrary scale upscaling module. Our proposed module consists of fewer parameters and consumes less memory and inference time than existing ones. As it is based on a simple convolutional neural network, it has the flexibility to be adapted to any other networks for arbitrary scale transformation. We also show that the proposed upscaling module can be extended to super-resolution under homographic transformation. We perform extensive experiments on widely used benchmark datasets, and experimental findings show the comparative performance of our proposed upscaling module as compared to recently developed approaches, while it provides ad-hoc benefits of being simple and computationally inexpensive. The code is available at https://github.com/aupendu/FTSR.
Aupendu Kar, Prabir Kumar Biswas
WACV2
2025 Towards Test Time Adaptation in Low Dose Computed Tomography Denoising Via Bias Modulation
abstract
Test-time adaptation (TTA) is crucial for robust medical image analysis, particularly in low-dose CT denoising where models trained on one noise level often fail to generalize to unseen noise intensities. This work proposes a novel TTA method that adapts the parameters of any pre-trained denoising model to denoise unknown test images. We demonstrate that surprisingly, fine-tuning only the convolutional layer biases achieves optimal TTA performance for this task. Building upon this, we introduce a novel pseudo-labeling strategy: re-corrupting the initially restored images to generate targets for further adaptation. This pseudo-labeling is combined with a knowledge distillation framework to efficiently update the model parameters during testing. Evaluations on low-dose CT scans demonstrate significant improvements in denoising performance, with peak signal-to-noise ratio (PSNR) gains of up to 3 dB observed.
Sutanu Bera, Krishnendu Ghosh, Prabir Kumar Biswas
ICIP3
2025 3D Shape Completion using Multi-resolution Spectral Encoding
abstract
Reconstruction of intricate local patterns and large missing regions during 3D shape completion has the contradictory requirements of computation over a wider context and operations for finer detail restoration. To this end, we propose a multi-resolution spectral encoding based 3D shape completion approach to work on truncated Signed Distance Field (SDF) based shape representations. Our novelty lies in judiciously integrating multi-resolution 3D convolutional blocks that encode the input shape and a spectral module (SM) that captures the shape-wide context, thus addressing the contradictory requirements. SM acts on the features extracted from both partial input scans and shape priors using the multi-resolution convolutional blocks. Our SM contains a 3D convolutional block placed between fast Fourier transform (FFT) and inverse FFT operations, which results in the expansion of the receptive field for the appropriate context computation. Our approach has an attention-based encoder-decoder architecture, where the encoding of a partial scan is acted upon by shape prior encodings to produce attention maps. These attention maps are lever-aged differently in pretraining, and in the later training and inference stages of our approach to produce the reconstructed 3D shape. A surface gradient-based loss function is used in addition to the L1 loss, both in the pretraining and training stages for emphasizing the differences in minute details. These along with an attention refinement operation often leads to complete reconstruction while restoring finer details. Experiments using standard synthetic and real datasets demonstrate the superiority of our approach over the state-of-the-art.
Pallabjyoti Deka, Saumik Bhattacharya, Debashis Sen, Prabir Kumar Biswas
WACV4
2023 Generative Pipeline for Data Augmentation of Unconstrained Document Images with Structural and Textural Degradation (Student Abstract)
abstract
Computer vision applications for document image understanding (DIU) such as optical character recognition, word spotting, enhancement etc. suffer from structural deformations like strike-outs and unconstrained strokes, to name a few. They also suffer from texture degradation due to blurring, aging, or blotting-spots etc. The DIU applications with deep networks are limited to constrained environment and lack diverse data with text-level and pixel-level annotation simultaneously. In this work, we propose a generative framework to produce realistic synthetic handwritten document images with simultaneous annotation of text and corresponding pixel-level spatial foreground information. The proposed approach generates realistic backgrounds with artificial handwritten texts which supplements data-augmentation in multiple unconstrained DIU systems. The proposed framework is an early work to facilitate DIU system-evaluation in both image quality and recognition performance at a go.
Arnab Poddar, Abhishek Kumar Sah, Soumyadeep Dey, Pratik Jawanpuria, Jayanta Mukhopadhyay, Prabir Kumar Biswas
AAAI6
2023 TBM-GAN: Synthetic Document Generation with Degraded Background
Arnab Poddar, Soumyadeep Dey, Pratik Jawanpuria, Jayanta Mukhopadhyay, Prabir Kumar Biswas
ICDAR (2)5
2023 Noise Conditioned Weight Modulation for Robust and Generalizable Low Dose CT Denoising
Sutanu Bera, Prabir Kumar Biswas
MICCAI (10)2
2023 Memory Replay for Continual Medical Image Segmentation Through Atypical Sample Selection
Sutanu Bera, Vinay Ummadi, Debashis Sen, Subhamoy Mandal, Prabir Kumar Biswas
MICCAI (4)5
2023 Self Supervised Low Dose Computed Tomography Image Denoising Using Invertible Network Exploiting Inter Slice Congruence
abstract
The resurgence of deep neural networks has created an alternative pathway for low-dose computed tomography denoising by learning a nonlinear transformation function between low-dose CT (LDCT) and normal-dose CT (NDCT) image pairs. However, those paired LDCT and NDCT images are rarely available in the clinical environment, making deep neural network deployment infeasible. This study proposes a novel method for self-supervised low-dose CT denoising to alleviate the requirement of paired LDCT and NDCT images. Specifically, we have trained an invertible neural network to minimize the pixel-based mean square distance between a noisy slice and the average of its two immediate adjacent noisy slices. We have shown the aforementioned is similar to training a neural network to minimize the distance between clean NDCT and noisy LDCT image pairs. Again, during the reverse mapping of the invertible network, the output image is mapped to the original input image, similar to cycle consistency loss. Finally, the trained invertible network’s forward mapping is used for denoising LDCT images. Extensive experiments on two publicly available datasets showed that our method performs favourably against other existing unsupervised methods.
Sutanu Bera, Prabir Kumar Biswas
WACV2
2023 Texture aware autoencoder pre-training and pairwise learning refinement for improved iris recognition
Manashi Chakraborty, Aritri Chakraborty, Prabir Kumar Biswas, Pabitra Mitra
Multim. Tools Appl.3
2023 Frame-level global context modeling for detection and localization of abnormality
Vikas Kumar 0001, Debdoot Sheet, Prabir Kumar Biswas
Multim. Tools Appl.4
2022 Gated Convolutional Network for Metal Artifact Reduction in Computed Tomography Images
abstract
This study proposes an image domain restoration network for metal artifact reduction in clinical computed tomography images. Specifically, we have proposed a pool and excite module to identify the streaking artifacts in the hidden latent space via learning a sigmoidal mask and a novel gated convolution layer, which utilises the previously learned gating weights for the reduction of metal artifacts. Our formulation of gated convolution is unique and custom-made to deal with metal artifacts. Extensive experiments on real CT images show that our method accomplishes significant improvement over the current state-of-the-art methods without requiring additional data, e.g., projection data, metal trace, etc.
Sutanu Bera, Prabir Kumar Biswas
ICIP2
2022 Multi-Latent GAN Inversion for Unsupervised 3D Shape Completion
abstract
The objective of 3-dimensional point cloud completion is to estimate a plausible complete shape from a given partial point cloud. Most of the data-driven point cloud completion approaches have been proposed in a supervised manner needing one-to-one correspondence between the partial and complete shapes. A promising way to solve the paired data dependency is to use the mapping capability of a pre-trained point cloud generation network to the best possible matching latent vector. However, recovering the composite structural details and complex geometry of a 3D shape is often difficult using a single latent vector alone. In this paper, we propose to employ multiple latent vectors, each of which generates individual feature maps, which are then combined to reconstruct a faithful complete 3D shape corresponding to an available partial shape. Deploying more than one latent vector enables the pre-trained generative network to increase its fidelity by using multiple combinations of feature representations learned by each single latent. Experimental results show that our algorithm performs well compared to the other existing shape completion methods. We also study the completion performance with a varying number of latent codes and the role of each latent vector in the final complete shape generation.
Krishnendu Ghosh, Aupendu Kar, Saumik Bhattacharya, Debashis Sen, Prabir Kumar Biswas
ICIP5
2022 Sub-Aperture Feature Adaptation in Single Image Super-Resolution Model for Light Field Imaging
abstract
With the availability of commercial Light Field (LF) cameras, LF imaging has emerged as an up-and-coming technology in computational photography. However, the spatial resolution is significantly constrained in commercial micro-lens-based LF cameras because of the inherent multiplexing of spatial and angular information. Therefore, it becomes the main bottleneck for other applications of light field cameras. This paper proposes an adaptation module in a pre-trained Single Image Super-Resolution (SISR) network to leverage the powerful SISR model instead of using highly engineered light field imaging domain-specific Super Resolution models. The adaption module consists of a Sub-aperture Shift block and a fusion block. It is an adaptation in the SISR network to further exploit the spatial and angular information in LF images to improve the super-resolution performance. Experimental validation shows that the proposed method outperforms existing light field super-resolution algorithms. It also achieves PSNR gains of more than 1 dB across all the datasets as compared to the same pre-trained SISR models for scale factor 2, and PSNR gains 0.6 − 1 dB for scale factor 4.
Aupendu Kar, Suresh Nehra, Jayanta Mukhopadhyay, Prabir Kumar Biswas
ICIP4
2022 Analyzing and Improving Low Dose CT Denoising Network via HU Level Slicing
Sutanu Bera, Prabir Kumar Biswas
MICCAI (6)2
2021 Fast Bayesian Uncertainty Estimation and Reduction of Batch Normalized Single Image Super-Resolution Network
abstract
Convolutional neural network (CNN) has achieved unprecedented success in image super-resolution tasks in re-cent years. However, the network’s performance depends on the distribution of the training sets and degrades on out-of-distribution samples. This paper adopts a Bayesian approach for estimating uncertainty associated with output and applies it in a deep image super-resolution model to address the concern mentioned above. We use the uncertainty estimation technique using the batch-normalization layer, where stochasticity of the batch mean and variance generate Monte-Carlo (MC) samples. The MC samples, which are nothing but different super-resolved images using different stochastic parameters, reconstruct the image, and provide a confidence or uncertainty map of the reconstruction. We propose a faster approach for MC sample generation, and it allows the variable image size during testing. Therefore, it will be useful for image reconstruction domain. Our experimental findings show that this uncertainty map strongly relates to the quality of reconstruction generated by the deep CNN model and explains its limitation. Furthermore, this paper proposes an approach to reduce the model’s uncertainty for an input image, and it helps to defend the adversarial attacks on the image super-resolution model. The proposed uncertainty reduction technique also improves the performance of the model for out-of-distribution test images. To the best of our knowledge, we are the first to propose an adversarial defense mechanism in any image reconstruction domain.
Aupendu Kar, Prabir Kumar Biswas
CVPR2
2021 Zero-Shot Single Image Restoration Through Controlled Perturbation of Koschmieder's Model
abstract
Real-world image degradation due to light scattering can be described based on the Koschmieder’s model. Training deep models to restore such degraded images is challenging as real-world paired data is scarcely available and synthetic paired data may suffer from domain-shift issues. In this paper, a zero-shot single real-world image restoration model is proposed leveraging a theoretically deduced property of degradation through the Koschmieder’s model. Our zero-shot network estimates the parameters of the Koschmieder’s model, which describes the degradation in the input image, to perform image restoration. We show that a suitable degradation of the input image amounts to a controlled perturbation of the Koschmieder’s model that describes the image’s formation. The optimization of the zero-shot network is achieved by seeking to maintain the relation between its estimates of Koschmieder’s model parameters before and after the controlled perturbation, along with the use of a few no-reference losses. Image dehazing and underwater image restoration are carried out using the proposed zero-shot framework, which in general outperforms the state-of-the-art quantitatively and subjectively on multiple standard real-world image datasets. Additionally, the application of our zero-shot framework for low-light image enhancement is also demonstrated.
Aupendu Kar, Sobhan Kanti Dhara, Debashis Sen, Prabir Kumar Biswas
CVPR4
2021 Structural similarity-based rate control algorithm for 3D video
Harshalatha Y, Prabir Kumar Biswas
Multim. Tools Appl.2
2021 Local instance and context dictionary-based detection and localization of abnormalities
Debdoot Sheet, Prabir Kumar Biswas
Mach. Vis. Appl.3
2021 Color Cast Dependent Image Dehazing via Adaptive Airlight Refinement and Non-Linear Color Balancing
abstract
Hazy images suffer from low visibility since the light gets scattered as it passes through various atmospheric particles. Moreover, such images are prone to color distortion, particularly in real weather conditions like sandstorms. In this letter, an effective dehazing technique is proposed using weighted least squares filtering on dark channel prior and color correction that involves automatic detection of color cast images. We show that the spread of the hue in a hazy image can differentiate a color cast image from a non-cast one. We propose a measure using the same for categorizing hazy images as cast and non-cast ones. Our novel color correction is performed by color balancing using a non-linear transformation followed by a cast-adaptive airlight refinement. Subjective and quantitative evaluations show that our method outperforms the state-of-the-art. It removes cast satisfactorily and reduces haze substantially while maintaining the naturalness of the image. Moreover, it produces visually pleasing images without halo artifacts.
Sobhan Kanti Dhara, Mayukh Roy, Debashis Sen, Prabir Kumar Biswas
IEEE Trans. Circuits Syst. Video Technol.4
2021 Lightweight Modules for Efficient Deep Learning Based Image Restoration
abstract
Low level image restoration is an integral component of modern artificial intelligence (AI) driven camera pipelines. Most of these frameworks are based on deep neural networks which present a massive computational overhead on resource constrained platform like a mobile phone. In this paper, we propose several lightweight low-level modules which can be used to create a computationally low cost variant of a given baseline model. Recent works for efficient neural networks design have mainly focused on classification. However, low-level image processing falls under the `image-to-image' translation genre which requires some additional computational modules not present in classification. This paper seeks to bridge this gap by designing generic efficient modules which can replace essential components used in contemporary deep learning based image restoration networks. We also present and analyse our results highlighting the drawbacks of applying depthwise separable convolutional kernel (a popular method for efficient classification network) for sub-pixel convolution based upsampling (a popular upsampling strategy for low-level vision applications). This shows that concepts from domain of classification cannot always be seamlessly integrated into `image-to-image' translation tasks. We extensively validate our findings on three popular tasks of image inpainting, denoising and super-resolution. Our results show that proposed networks consistently output visually similar reconstructions compared to full capacity baselines with significant reduction of parameters, memory footprint and execution speeds on contemporary mobile devices.
Avisek Lahiri, Sourav Bairagya, Sutanu Bera, Siddhant Haldar, Prabir Kumar Biswas
IEEE Trans. Circuits Syst. Video Technol.5
2021 Noise Conscious Training of Non Local Neural Network Powered by Self Attentive Spectral Normalized Markovian Patch GAN for Low Dose CT Denoising
abstract
The explosive rise of the use of Computer tomography (CT) imaging in medical practice has heightened public concern over the patient's associated radiation dose. On the other hand, reducing the radiation dose leads to increased noise and artifacts, which adversely degrades the scan's interpretability. In recent times, the deep learning-based technique has emerged as a promising method for low dose CT(LDCT) denoising. However, some common bottleneck still exists, which hinders deep learning-based techniques from furnishing the best performance. In this study, we attempted to mitigate these problems with three novel accretions. First, we propose a novel convolutional module as the first attempt to utilize neighborhood similarity of CT images for denoising tasks. Our proposed module assisted in boosting the denoising by a significant margin. Next, we moved towards the problem of non-stationarity of CT noise and introduced a new noise aware mean square error loss for LDCT denoising. The loss mentioned above also assisted to alleviate the laborious effort required while training CT denoising network using image patches. Lastly, we propose a novel discriminator function for CT denoising tasks. The conventional vanilla discriminator tends to overlook the fine structural details and focus on the global agreement. Our proposed discriminator leverage self-attention and pixel-wise GANs for restoring the diagnostic quality of LDCT images. Our method validated on a publicly available dataset of the 2016 NIH-AAPM-Mayo Clinic Low Dose CT Grand Challenge performed remarkably better than the existing state of the art method. The corresponding source code is available at: https://github.com/reach2sbera/ldct_nonlocal.
Sutanu Bera, Prabir Kumar Biswas
IEEE Trans. Medical Imaging2
2020 Prior Guided GAN Based Semantic Inpainting
abstract
Contemporary deep learning based semantic inpainting can be approached from two directions. First, and the more explored, approach is to train an offline deep regression network over the masked pixels with an additional refinement by adversarial training. This approach requires a single feed-forward pass for inpainting at inference. Another promising, yet unexplored approach is to first train a generative model to map a latent prior distribution to natural image manifold and during inference time search for the best-matching prior to reconstruct the signal. The primary aversion towards the latter genre is due to its inference time iterative optimization and difficulty to scale to higher resolution. In this paper, going against the general trend, we focus on the second paradigm of inpainting and address both of its mentioned problems. Most importantly, we learn a data driven parametric network to directly predict a matching prior for a given masked image. This converts an iterative paradigm to a single feed forward inference pipeline with around 800X speedup. We also regularize our network with structural prior (computed from the masked image itself) which helps in better preservation of pose and size of the object to be inpainted. Moreover, to extend our model for sequence reconstruction, we propose a recurrent net based grouped latent prior learning. Finally, we leverage recent advancements in high resolution GAN training to scale our inpainting network to 256X256. Experiments (spanning across resolutions from 64X64 to 256X256) conducted on SVHN, Standford Cars, CelebA, CelebA-HQ and ImageNet image datasets, and FaceForensics video datasets reveal that we consistently improve upon contemporary benchmarks from both schools of approaches.
Avisek Lahiri, Arnav Kumar Jain, Sanskar Agrawal, Pabitra Mitra, Prabir Kumar Biswas
CVPR5
2020 Unsupervised Pre-Trained, Texture Aware and Lightweight Model for Deep Learning Based Iris Recognition Under Limited Annotated Data
abstract
In this paper, we present a texture aware lightweight deep learning framework for iris recognition. Our contributions are primarily three fold. Firstly, to address the dearth of labelled iris data, we propose a reconstruction loss guided unsupervised pre-training stage followed by supervised refinement. This drives the network weights to focus on discriminative iris texture patterns. Next, we propose several texture aware improvisations inside a Convolution Neural Net to better leverage iris textures. Finally, we show that our systematic training and architectural choices enable us to design an efficient framework with upto 100× fewer parameters than contemporary deep learning baselines yet achieve better recognition performance for within and cross dataset evaluations.
Manashi Chakraborty, Mayukh Roy, Prabir Kumar Biswas, Pabitra Mitra
ICIP3
2020 Retinal Vessel Segmentation Under Extreme Low Annotation: A Gan Based Semi-Supervised Approach
abstract
Contemporary deep learning based medical image segmentation algorithms require hours of annotation labor by domain experts. These data hungry deep models perform sub-optimally in the presence of limited amount of labeled data. In this paper, we present a data efficient learning framework using the recent concept of Generative Adversarial Networks; this allows a deep neural network to perform significantly better than its fully supervised counterpart in low annotation regime. The proposed method is an extension of our previous work with the addition of a new unsupervised adversarial loss and a structured prediction based architecture. Though generic, we demonstrate the efficacy of our approach for retinal blood vessels segmentation from fundus images on DRIVE and STARE datasets. We experiment with extreme low annotation budget and we show, that under this constrained data setting, the proposed method outperforms our previous method and other fully supervised benchmark models. In addition, our systematic ablation studies suggest some key observations for successfully training GAN based semi-supervised algorithms with an encoder-decoder style network architecture.
Avisek Lahiri, Vineet Jain, Arnab Mondal, Prabir Kumar Biswas
ICIP4
2020 Spatiotemporal deep networks for detecting abnormality in videos
Debdoot Sheet, Prabir Kumar Biswas
Multim. Tools Appl.3
2019 Faster Unsupervised Semantic Inpainting: A GAN Based Approach
abstract
In this paper, we propose to improve the inference speed and visual quality of contemporary baseline of Generative Adversarial Networks (GAN) based unsupervised semantic inpainting. This is made possible with better initialization of the core iterative optimization involved in the framework. To our best knowledge, this is also the first attempt of GAN based video inpainting with consideration to temporal cues. On single image inpainting, we achieve about 4.5-5× speedup and 80 × on videos compared to baseline. Simultaneously, our method has better spatial and temporal reconstruction qualities as found on three image and one video dataset.
Avisek Lahiri, Arnav Kumar Jain, Divyasri Nadendla, Prabir Kumar Biswas
ICIP4
2019 Unsupervised Categorization of Forest-Cover Using Multi-Spectral and Hybrid Polarimetric Sar Images
abstract
In this paper, we propose to distinguish forest-cover in an unsupervised fashion by a combination of passive multi-spectral imagery and active hybrid polarized SAR data. At first, multi-spectral imagery (MSI) is used to separate general vegetation region (e.g., forest, mature grassland, and pre-harvest agricultural fields) from the imaged scene using spectral slopes based rules and support vector machine technique. Then, hybrid polarimetric SAR image of the same region (acquired with a common time stamp) is clustered into three scatter classes, namely, surface, volume, and dihedral, using Stokes parameters based m - δ decomposition. Forest cover is extracted by bi-labeled pixels of the study site that correspond to vegetation (in MSI) and volume scatter (in SAR), which forms a community level classification of forest region. Further, using Wishart derived mean-shift clustering technique, we segregate possible categories of forest clusters within the mapped forest region to obtain a sub-community level classification. Discernible spectral and scattering characteristics of remotely sensed images are explored in our work for identifying forest regions and their possible categories. The proposed method is automated by freeing the manual supervision in selecting seed pixels for training any machine learning technique.
Shashaank M. Aswatha, Rajeswari Mahapatra, Jayanta Mukhopadhyay, Prabir Kumar Biswas, Subhas Aikat, Arundhati Misra 0001
IGARSS4
2019 Unsupervised Adversarial Visual Level Domain Adaptation for Learning Video Object Detectors From Images
abstract
Deep learning based object detectors require thousands of diversified bounding box and class annotated examples. Though image object detectors have shown rapid progress in recent years with release of multiple large scale static image datasets, object detection on videos still remains an open problem due to unavailability of annotated video frames. Having a robust video object detector is an essential component for video understanding and curating large scale automated annotations in videos. Domain difference between images and videos makes the transferability of image object detectors to videos sub-optimal. The most common solution is to use weakly supervised annotations where a video frame has to be tagged for presence/absence of object categories. This still takes up manual effort. In this paper we take a step forward to attain zero supervision on video domain by adapting the concept of unsupervised adversarial image-to-image translation to perturb static high quality images to be visually indistinguishable from set of video frmes. We assume the presence of a fully annotated static image dataset and an unannotated video frames. Object detector is trained on adversarially transformed image dataset using the annotations of original dataset. Experiments on Youtube-Objects and Youtube-Objects-Subset datasets with two contemporary baseline object detectors reveal that such unsupervised pixel level domain adaptation boosts the generalization performance on video frames compared to direct application of image object detector. Also we achieve competitive performance compared to recent baselines of weakly supervised methods. This paper can be seen as an application of image translation for cross domain object detection.
Avisek Lahiri, Sri Charan Ragireddy, Prabir Kumar Biswas, Pabitra Mitra
WACV3
2018 Global abnormal events detection in crowded scenes using context location and motion-rich spatio-temporal volumes
abstract
Global abnormal events form unique and distinct motion characteristics and category of anomalies at image level rather than pixel level with less complexity compared to local abnormal events. However, traditional anomaly detection approaches focused more on pixel‐level feature extraction from foreground pixels and combine global and local anomaly detection in a single algorithm with equal degree of computational complexity. In this paper, we propose a novel framework for global anomaly detection via block‐level feature extraction using context location (CL) and motion‐rich STVs (MRSTVs). The histogram of optical flow orientation and motion magnitude features from spatio‐temporal volumes (STVs) are used as global feature descriptor to capture motion characteristics of normal and abnormal events. Simple and cost‐effective one‐class SVM classifier is employed to learn normal behaviour from MRSTVs during training and detect abnormal STVs from test data. Thereafter, a spatio‐temporal post‐processing technique detects frame‐level abnormal behaviour and reduces false alarm rate. We define CL to detect abnormal behaviour in an unexpected region. The proposed approach omits pixel‐level feature extraction and background modelling by considering MRSTVs, thus enhances detection rate and reduces computational complexity. We have conducted experiments on widely used UMN and PETS2009 datasets to compare the performance of proposed approach with existing methods.
N. Patil, Prabir Kumar Biswas
IET Image Process.2
2018 SSIM-based joint-bit allocation for 3D video coding
Harshalatha Y, Prabir Kumar Biswas
Multim. Tools Appl.2
2018 Forward Stagewise Additive Model for Collaborative Multiview Boosting
abstract
Multiview assisted learning has gained significant attention in recent years in supervised learning genre. Availability of high-performance computing devices enables learning algorithms to search simultaneously over multiple views or feature spaces to obtain an optimum classification performance. This paper is a pioneering attempt of formulating a mathematical foundation for realizing a multiview aided collaborative boosting architecture for multiclass classification. Most of the present algorithms apply multiview learning heuristically without exploring the fundamental mathematical changes imposed on traditional boosting. Also, most of the algorithms are restricted to two class or view setting. Our proposed mathematical framework enables collaborative boosting across any finite-dimensional view spaces for multiclass learning. The boosting framework is based on a forward stagewise additive model, which minimizes a novel exponential loss function. We show that the exponential loss function essentially captures the difficulty of a training sample space instead of the traditional "1/0" loss. The new algorithm restricts a weak view from overlearning and thereby preventing overfitting. The model is inspired by our earlier attempt on collaborative boosting, which was devoid of mathematical justification. The proposed algorithm is shown to converge much nearer to global minimum in the exponential loss space and thus supersedes our previous algorithm. This paper also presents analytical and numerical analyses of convergence and margin bounds for multiview boosting algorithms and we show that our proposed ensemble learning manifests lower error bound and higher margin compared with our previous model. Also, the proposed model is compared with traditional boosting and recent multiview boosting algorithms. In the majority of instances, the new algorithm manifests a faster rate of convergence on training set error and also simultaneously offers better generalization performance. The kappa-error diagram analysis reveals the robustness of the proposed boosting framework to labeling noise.
Avisek Lahiri, Biswajit Paria, Prabir Kumar Biswas
IEEE Trans. Neural Networks Learn. Syst.3
2016 Spectral slopes for automated classification of land cover in landsat images
abstract
In the literature, various techniques for supervised/ semi-supervised classification of satellite imageries require manual selection of samples for each class. In this paper, we propose a spectral-slope based classification technique, which automates the process of initial labeling of a set of sample points. These are subsequently used in a supervised classifier as training samples and it performs the task of classification over all the pixels in the image. We demonstrate the effectiveness of our proposed classification technique in summarizing the changes in temporal image sets. For selecting the training samples from the satellite imageries, a set of rules is proposed by using the spectral-slope properties. We classify the land-cover into three classes, namely, water, vegetation, and vegetation-void, and validate the classification results using very high resolution satellite imagery. The approach has also been used in the analysis of images acquired by different sensors operating under similar wavelength ranges.
Shashaank M. Aswatha, Jayanta Mukhopadhyay, Prabir Kumar Biswas
ICIP3
2016 Rate distortion optimization using SSIM for 3D video coding
abstract
Coding efficiency can be enhanced through rate-distortion optimization (RDO) that provides a trade-off between bit-rate and distortion. In this paper, we have proposed Structural SIMilarity (SSIM) based RDO for 3D video coding improvement. SSIM index is a quality metric that gives better approximation to visual quality. Most of the existing literature on 3D video coding employs sum-of-squared error (SSE) as a measure of distortion, which does not always correlate to visual quality. In order to overcome this gap, SSIM-based RDO is implemented in this paper. Lagrange multiplier is modified to obtain optimum rate along with a reduction in distortion which improves the perceptual quality of the video. The entropy of macroblock (MB) is also considered in the scaling of Lagrange multiplier to increase RDO performance. The proposed algorithm is implemented in 3DV- ATM reference software. Experimental results show an improvement in the perceptual quality of the synthesized sequences with bitrate reduction of 6 - 15%.
Harshalatha Y, Prabir Kumar Biswas
ICPR2
2016 A novel moving object segmentation framework utilizing camera motion recognition for H.264 compressed videos
Manish Okade, Prabir Kumar Biswas
J. Vis. Commun. Image Represent.2
2016 Robust Learning-Based Camera Motion Characterization Scheme With Applications to Video Stabilization
abstract
This paper investigates a novel learning-based camera motion characterization scheme along with its application to the video stabilization problem. The proposed characterization scheme represents the compressed domain block motion vectors (MVs) using polar angle and magnitude histograms. Discriminative features from these two histograms are extracted and fed to a supervised learning-based hierarchical classifier for recognizing the six camera motion patterns. A comparative analysis with an existing scheme is carried out to support and validate the proposed characterization scheme. The proposed scheme works at the frame level by classifying the inter-frame camera motion patterns. This scheme is extended to classify the video segments and a novel application to video stabilization is investigated. An experimental analysis of a number of test sequences captured using a handheld video camera shows that by characterizing the smooth and jittery motions, selective video stabilization could be carried out only on those video segments that have been degraded. This approach of selective video stabilization saves considerable amount of computational time compared with running the stabilization algorithm on the entire video sequence, as proposed in the literature. The proposed strategy of using a classification scheme prior to applying the video stabilization routine offers a new paradigm to the conventional video stabilization problem. Experimental validation carried out using exhaustive search motion estimation obtained block MVs, and H.264/Advanced Video Coding-obtained MVs shows that by using the idea of selective video stabilization, up to 62% reduction in processing time can be achieved compared with video stabilization approaches wherein the entire video sequence is processed.
Manish Okade, Gaurav Patel, Prabir Kumar Biswas
IEEE Trans. Circuits Syst. Video Technol.3
2015 Noise-aided dynamic range compression using selective processing in a statistics-dependent stochastic resonance model
abstract
This paper presents a noise-aided dynamic range compression algorithm using a stochastic resonance model in spatial domain. An input statistics-dependent stochastic resonance (ISSR) model, that is designed for contrast enhancement of dark images, is used here to enhance an image with both bright and dark areas. The underilluminated regions of such an image are selected as the De Vries Rose region from a human visual system-based segmentation algorithm, and then processed using the ISSR model. It is observed that by semi-adaptively changing the processing parameters with iteration, the processed dark regions and the unprocessed bright regions of an image smoothly merge producing a quality of dynamic range compression in the image. The performance of the proposed algorithm is characterized using image quality index for tone-mapped images and a no-reference perceptual quality measure. Results and comparative analysis suggest notable performance of the proposed algorithm with fewer iteration.
Rajlaxmi Chouhan, Prabir Kumar Biswas
VCIP2
2014 A New Framework for Multiclass Classification Using Multiview Assisted Adaptive Boosting
Avisek Lahiri, Prabir Kumar Biswas
ACCV (3)2
2014 Image enhancement and dynamic range compression using novel intensity-specific stochastic resonance-based parametric image enhancement model
abstract
This paper presents a noise-aided image enhancement algorithm focussed on addressing images that have a large dynamic range, i.e., images with both dark and bright regions. The application of a new mathematical model, in a shifted double-well system exhibiting stochastic resonance, is investigated for such images. The new mathematical model addresses the shortcomings of earlier SR-based enhancement model by deriving parameters purely from input values (instead of input statistics). This model is specific to spatial domain pixel representation and operates on a revised iterative equation. This iterative processing is here applied selectively to the under-illuminated regions of the image, characterized as the De Vries-Rose (DVR) region of a human psychovisual model. The idea of suitably modifying the existing universal image quality index is also proposed for its participation in iteration termination, and to gauge the property of dynamic range compression. While the iterative algorithm is terminated using the revised image quality index, entropy maximization, and contrast quality of DVR region with constraints on perceptual quality, the performance of the proposed algorithm is also characterized by observing color enhancement and subjective scores on visual quality.
Rajlaxmi Chouhan, Prabir Kumar Biswas
ICIP2
2014 Video stabilization using maximally stable extremal region features
Manish Okade, Prabir Kumar Biswas
Multim. Tools Appl.2
2013 Enhancement of dark and low-contrast images using dynamic stochastic resonance
abstract
In this study, a dynamic stochastic resonance (DSR)‐based technique in spatial domain has been proposed for the enhancement of dark‐ and low‐contrast images. Stochastic resonance (SR) is a phenomenon in which the performance of a system (low‐contrast image) can be improved by addition of noise. However, in the proposed work, the internal noise of an image has been utilised to produce a noise‐induced transition of a dark image from a state of low contrast to that of high contrast. DSR is applied in an iterative fashion by correlating the bistable system parameters of a double‐well potential with the intensity values of a low‐contrast image. Optimum output is ensured by adaptive computation of performance metrics – relative contrast enhancement factor ( F ), perceptual quality measures and colour enhancement factor. When compared with the existing enhancement techniques such as adaptive histogram equalisation, gamma correction, single‐scale retinex, multi‐scale retinex, modified high‐pass filtering, edge‐preserving multi‐scale decomposition and automatic controls of popular imaging tools, the proposed technique gives significant performance in terms of contrast and colour enhancement as well as perceptual quality. Comparison with a spatial domain SR‐based technique has also been illustrated.
Rajlaxmi Chouhan, Rajib Kumar Jha, Prabir Kumar Biswas
IET Image Process.3
2012 Internal noise-induced contrast enhancement of dark images
abstract
A contrast enhancement technique using scaling of internal noise of a dark image in discrete cosine transform (DCT) domain has been proposed in this paper. The mechanism of enhancement is attributed to noise-induced transition of DCT coefficients from a poor state to an enhanced state. This transition is effected by the internal noise present due to lack of sufficient illumination and can be modeled by a general bistable system exhibiting dynamic stochastic resonance. The proposed technique adopts a local adaptive processing and significantly enhances the image contrast and color information while ascertaining good perceptual quality. When compared with the existing enhancement techniques such as adaptive histogram equalization, gamma correction, single-scale retinex, multi-scale retinex, modified high-pass filtering, multi-contrast enhancement, multi-contrast enhancement with dynamic range compression, color enhancement by scaling, edge-preserving multi-scale decomposition and automatic controls of popular imaging tool, the proposed technique gives remarkable performance in terms of relative contrast enhancement, colorfulness and visual quality of enhanced image.
Rajib Kumar Jha, Rajlaxmi Chouhan, Prabir Kumar Biswas, Kiyoharu Aizawa
ICIP3
2012 Fast Video Stabilization in the Compressed Domain
abstract
Video stabilization is an important technique in present day digital cameras as most of the cameras are hand-held, mounted on moving platforms or subjected to atmospheric vibrations. Motion estimation is a bottleneck in the stabilization pipeline as it consumes about 90% of the processing time. In this paper we propose to perform the stabilization task in the compressed domain using the motion information readily available in the compressed bit stream. The motion vector based global motion estimation technique is used to estimate the camera motion parameters from the block motion vectors. Smoothing of the motion parameters to retain the desired motion is performed using a gaussian filter followed by motion compensation to construct the stabilized frame. The focus of our work is on using the information in coded video streams to reduce the computational complexity and processing time. We compare the proposed scheme with existing pixel domain counterparts and show how it achieves speedup in processing time at the same time maintaining satisfactory stabilization performance.
Manish Okade, Prabir Kumar Biswas
ICME2
2012 Image denoising using dynamic stochastic resonance in wavelet domain
abstract
A dynamic stochastic resonance (DSR)-based technique in discrete wavelet transform (DWT) domain for noise suppression in digital images has been proposed in this paper. The initial results on investigation of this concept for denoising of images corrupted by gaussian noise have been presented. Though traditionally noise is considered as undesirable, it has been utilized in the proposed technique to reduce its own effect. In the iterative DSR step, an input noisy image is subjected to independent noise of different standard deviations that iteratively tunes the detail wavelet coefficients, such that the overall effect is the suppression of the degradation due to its own noise. The results are quantified in terms of Noise Mean Value (NMV), Noise Standard Deviation (NSD), and Mean Square Difference (MSD). When compared with the conventional techniques for gaussian denoising, such as gaussian low pass filtering, and soft thresholding of wavelet coefficients, the DSR-based technique is found to give marginally better noise reduction in most of the cases.
Rajlaxmi Chouhan, Rajib Kumar Jha, Prabir Kumar Biswas
ISDA3
2012 Fast camera motion estimation using discrete wavelet transform on block motion vectors
abstract
In this paper we propose to use the discrete wavelet transform on the block motion vector field for estimating the camera (global) motion parameters in the compressed domain. By taking the wavelet transform of the block motion vector field we get the decomposition of the motion vectors into wavelet sub-bands. The LL sub-band gives us the average motion which we assume is mainly due to the background (camera) motion. By using gradient descent regression on LL sub-band wavelet coefficients we extract the camera motion parameters. Our results show that LL sub-band is enough to estimate the camera motion parameters unlike previous methods which utilize the entire block motion vector field to carry out the estimation process. Our experimental results demonstrate significant gain in processing time at the cost of marginal drop of estimation accuracy for the camera motion parameters. Our proposed technique finds its application in video indexing and video shot segmentation where fast camera motion estimation is warranted.
Manish Okade, Prabir Kumar Biswas
PCS2
2011 Image filtering in the block DCT domain using symmetric convolution
Kapinaiah Viswanath, Jayanta Mukhopadhyay, Prabir Kumar Biswas
J. Vis. Commun. Image Represent.3
2010 New packet aggregation schemes for multimedia applications in WLAN
abstract
In IEEE 802.11 standard MAC layer and Physical layer overheads are necessary for the proper synchronization of transmitter and receiver besides sharing the wireless medium efficiently. To increase the network efficiency, the ratio of header size to payload size in a packet has to be reduced. The new IEEE WLAN amendment, IEEE 802.11n, allows aggregation of packets to increase the payload size. Here a node aggregates the packets of different applications to compose a larger packet before it is send to the access point (AP). However, this method fails in the case of delay sensitive multimedia applications. In this paper, we propose three methods that use packet aggregation to improve the system throughput of real time multimedia applications in a WLAN. These schemes use capture effect, power control scheme and directionality of antenna to allow concurrent packet transmissions in a WLAN. Our simulation studies show that these schemes increase the system throughput considerably.
Mangalathu Jibukumar, Raja Datta, Prabir Kumar Biswas
NOMS3
2010 CoopMACA: a cooperative MAC protocol using packet aggregation
Mangalathu Jibukumar, Raja Datta, Prabir Kumar Biswas
Wirel. Networks3
2009 Transcoding in the block DCT space
abstract
Transcoding enables the transformation of multimedia content to adapt to a diverse nature of client/user requirements. In this paper, we propose a technique for transcoding wavelet coefficients to block DCT coefficients in the transform domain. In the first step, the wavelet coefficients are transformed into upsampled block DCT coefficients. Subsequently these transformed coefficients are synthesized by filtering in the block DCT space. To reduce the transcoding complexity, we perform upsampling and filtering operations in a single combined step. The proposed transcoding approach restricts all operations of the synthesis process in the block DCT space. The proposed approach achieves the same quality of reconstruction (Except rounding errors in processing) as that of spatial domain technique with reduced complexity for a given wavelet synthesis filter bank.
Kapinaiah Viswanath, Jayanta Mukhopadhyay, Prabir Kumar Biswas
ICIP3
2009 A Granular Reflex Fuzzy Min-Max Neural Network for Classification
abstract
Granular data classification and clustering is an upcoming and important issue in the field of pattern recognition. Conventionally, computing is thought to be manipulation of numbers or symbols. However, human recognition capabilities are based on ability to process nonnumeric clumps of information (information granules) in addition to individual numeric values. This paper proposes a granular neural network (GNN) called granular reflex fuzzy min-max neural network (GrRFMN) which can learn and classify granular data. GrRFMN uses hyperbox fuzzy set to represent granular data. Its architecture consists of a reflex mechanism inspired from human brain to handle class overlaps. The network can be trained online using granular or point data. The neuron activation functions in GrRFMN are designed to tackle data of different granularity (size). This paper also addresses an issue to granulate the training data and learn from it. It is observed that such a preprocessing of data can improve performance of a classifier. Experimental results on real data sets show that the proposed GrRFMN can classify granules of different granularity more correctly. Results are compared with general fuzzy min-max neural network (GFMN) proposed by Gabrys and Bargiela and with some classical methods.
Abhijeet V. Nandedkar, Prabir Kumar Biswas
IEEE Trans. Neural Networks2
2007 Texture image retrieval using rotated wavelet filters
Manesh Kokare, Prabir Kumar Biswas, Biswanath N. Chatterji
Pattern Recognit. Lett.2
2007 A Fuzzy Min-Max Neural Network Classifier With Compensatory Neuron Architecture
abstract
This paper proposes a fuzzy min-max neural network classifier with compensatory neurons (FMCNs). FMCN uses hyperbox fuzzy sets to represent the pattern classes. It is a supervised classification technique with new compensatory neuron architecture. The concept of compensatory neuron is inspired from the reflex system of human brain which takes over the control in hazardous conditions. Compensatory neurons (CNs) imitate this behavior by getting activated whenever a test sample falls in the overlapped regions amongst different classes. These neurons are capable to handle the hyperbox overlap and containment more efficiently. Simpson used contraction process based on the principle of minimal disturbance, to solve the problem of hyperbox overlaps. FMCN eliminates use of this process since it is found to be erroneous. FMCN is capable to learn the data online in a single pass through with reduced classification and gradation errors. One of the good features of FMCN is that its performance is less dependent on the initialization of expansion coefficient, i.e., maximum hyperbox size. The paper demonstrates the performance of FMCN by comparing it with fuzzy min-max neural network (FMNN) classifier and general fuzzy min-max neural network (GFMN) classifier, using several examples.
Abhijeet V. Nandedkar, Prabir Kumar Biswas
IEEE Trans. Neural Networks2
2006 Rotation-Invariant Texture Image Retrieval Using Rotated Complex Wavelet Filters
abstract
This paper proposes a novel approach for rotation-invariant texture image retrieval by using set of dual-tree rotated complex wavelet filter (DT-RCWF) and DT complex wavelet transform (DT-CWT) jointly, which obtains texture features in 12 different directions. Two-dimensional RCWFs are nonseparable and oriented, which improves characterization of oriented textures. Robust and efficient isotropic rotationally invariant features are extracted from DT-RCWF and DT-CWT decomposed subbands. This paper demonstrates the effectiveness of this new set of features on four different sets of rotated and nonrotated databases. Experimental results indicate that the proposed method improves retrieval accuracy from 83.17% to 93.71% on a small size (208 images) nonrotated database D1, from 82.71% to 90.86% on a small size (208 images) rotated database D2, from 72.18% to 76.09% on a medium-size (640 images) rotated database D3, and from 64.17% to 78.93% on a large size (1856 images) rotated database D4, compared with the discrete wavelet transform-based approach. New method also retains comparable levels of computational complexity.
Manesh Kokare, Prabir Kumar Biswas, Biswanath N. Chatterji
IEEE Trans. Syst. Man Cybern. Part B2
2005 A General Fuzzy Min Max Neural Network with Compensatory Neuron Architecture
Abhijeet V. Nandedkar, Prabir Kumar Biswas
KES (3)2
2005 Texture image retrieval using new rotated complex wavelet filters
abstract
A new set of two-dimensional (2-D) rotated complex wavelet filters (RCWFs) are designed with complex wavelet filter coefficients, which gives texture information strongly oriented in six different directions (45 degrees apart from complex wavelet transform). The 2-D RCWFs are nonseparable and oriented, which improves characterization of oriented textures. Most texture image retrieval systems are still incapable of providing retrieval result with high retrieval accuracy and less computational complexity. To address this problem, we propose a novel approach for texture image retrieval by using a set of dual-tree rotated complex wavelet filter (DT-RCWF) and dual-tree-complex wavelet transform (DT-CWT) jointly, which obtains texture features in 12 different directions. The information provided by DT-RCWF complements the information generated by DT-CWT. Features are obtained by computing the energy and standard deviation on each subband of the decomposed image. To check the retrieval performance, texture database D1 of 1856 textures from Brodatz album and database D2 of 640 texture images from VisTex image database is created. Experimental results indicates that the proposed method improves retrieval rate from 69.61% to 77.75% on database D1, and from 64.83% to 82.81% on database D2, in comparing with traditional discrete wavelet transform based approach. The proposed method also retains comparable levels of computational complexity.
Manesh Kokare, Prabir Kumar Biswas, Biswanath N. Chatterji
IEEE Trans. Syst. Man Cybern. Part B2
2004 Rotation invariant texture features using rotated complex wavelet for content based image retrieval
abstract
A new rotationally invariant texture feature extraction method is introduced that utilizes the dual tree rotated complex wavelet filters (DT-RCWF) and dual tree complex wavelet transform (DT-CWT) jointly. A new two-dimensional rotated complex wavelet filter is designed with a complex wavelet filter coefficient. Decomposing the image with DT-RCWF and DT-CWT jointly gives shift invariant subbands oriented in twelve different directions. Isotropic rotationally invariant features are extracted from these subbands. The performance of image retrieval with the proposed features on rotated and nonrotated image databases is compared with the existing method. Experimental results show that the proposed rotation-invariant texture features are more robust and outperform the other existing methods.
Manesh Kokare, Prabir Kumar Biswas, Biswanath N. Chatterji
ICIP2
2004 Cosine-modulated wavelet based texture features for content-based image retrieval
Manesh Kokare, Biswanath N. Chatterji, Prabir Kumar Biswas
Pattern Recognit. Lett.3
2003 Rotation invariant texture classification using even symmetric Gabor filters
Ramchandra Manthalkar, Prabir Kumar Biswas, Biswanath N. Chatterji
Pattern Recognit. Lett.2
2003 Rotation and scale invariant texture features using discrete wavelet packet transform
Ramchandra Manthalkar, Prabir Kumar Biswas, Biswanath N. Chatterji
Pattern Recognit. Lett.2
2002 Dimensionality reduction of tree structured wavelet transform texture features for content based image retrieval
abstract
Dimensionality reduction methods are of interest in applications such as content-based image and video retrieval. The focus of this paper is on the dimensionality reduction of feature vectors of tree structured wavelet decomposition for improving the retrieval speed. We have investigated a novel idea of reduction of feature dimension by concatenating the inter scale approximate coefficients and intra scale detail coefficient of each individual channel. The results are quite impressive; in an experiment using Brodatz texture database, feature vector length is reduced by a factor of three, which doubles retrieval speed without significantly reducing retrieval performance.
Manesh Kokare, Biswanath N. Chatterji, Prabir Kumar Biswas
ICARCV3
2000 Analysis of fuzzy thresholding schemes
C. V. Jawahar, Prabir Kumar Biswas
Pattern Recognit.2
1998 Fractal dimension estimation for texture images: A parallel approach
Manoj Kumar Biswas, Tirthankar Ghose, Sudipta Guha, Prabir Kumar Biswas
Pattern Recognit. Lett.4
1997 Investigations on fuzzy thresholding based on fuzzy clustering
C. V. Jawahar, Prabir Kumar Biswas
Pattern Recognit.2
1995 An SIMD algorithm for range image segmentation
Prabir Kumar Biswas, S. S. Biswas, Biswanath N. Chatterji
Pattern Recognit.1
1995 Detection of clusters of distinct geometry: A step towards generalised fuzzy clustering
C. V. Jawahar, Prabir Kumar Biswas
Pattern Recognit. Lett.2
1993 Component labeling in pyramid architecture
Prabir Kumar Biswas, Jayanta Mukhopadhyay, Biswanath N. Chatterji
Pattern Recognit.1
1992 Qualitative Description of Three-Dimensional Scenes
abstract
This paper describes a system which obtains a structural scene description of 3-D objects from range images. The system uses a hierarchical approach to obtain higher level primitives from lower level ones. Instead of a detailed mathematical approach, qualitative reasoning by rule based deduction is used to obtain the scene description. The rule bases are also hierarchical and several special control strategies like rule pruning, windowing (or zoning) and fact inhibition are used to considerably improve the speed of the system. Experimental results and performance of the system on actual range images are presented.
Prabir Kumar Biswas, Jayanta Mukhopadhyay, Biswanath N. Chatterji, P. P. Chakrabarti 0001
Int. J. Pattern Recognit. Artif. Intell.1