VLDB 2026 Research / reviewers in the wild / expert
Saumik Bhattacharya
dblp:154/8318
· DBLP profile ↗
33ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0003-1273-7969ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 19 · 18 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | $\ell _{0}$ℓ0-Regularized Sparse Coding-Based Interpretable Network for Multi-Modal Image FusionabstractMulti-modal image fusion (MMIF) enhances the information content of the fused image by combining the unique as well as common features obtained from different modality sensor images, improving visualization, object detection, and many more tasks. In this work, we introduce an interpretable network for the MMIF task, named FNet, based on an $\ell _{0}$ℓ0-regularized multi-modal convolutional sparse coding (MCSC) model. Specifically, for solving the $\ell _{0}$ℓ0-regularized CSC problem, we design a learnable $\ell _{0}$ℓ0-regularized sparse coding (LZSC) block in a principled manner through deep unfolding. Given different modality source images, FNet first separates the unique and common features from them using the LZSC block and then these features are combined to generate the final fused image. Additionally, we propose an $\ell _{0}$ℓ0-regularized MCSC model for the inverse fusion process. Based on this model, we introduce an interpretable inverse fusion network named IFNet, which is utilized during FNet's training. Extensive experiments show that FNet achieves high-quality fusion results across eight different MMIF datasets. Furthermore, we show that FNet enhances downstream object detection and semantic segmentation in visible-thermal image pairs. We have also visualized the intermediate results of FNet, which demonstrates the good interpretability of our network. Gargi Panda, Soumitra Kundu, Saumik Bhattacharya, Aurobinda Routray |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | INN-PAR: Invertible Neural Network for PPG to ABP ReconstructionabstractNon-invasive and continuous blood pressure (BP) monitoring is essential for the early prevention of many cardiovascular diseases. Estimating arterial blood pressure (ABP) from photoplethysmography (PPG) has emerged as a promising solution. However, existing deep learning approaches for PPG-to-ABP reconstruction (PAR) encounter certain information loss, impacting the precision of the reconstructed signal. To overcome this limitation, we introduce an invertible neural network for PPG to ABP reconstruction (INN-PAR), which employs a series of invertible blocks to jointly learn the mapping between PPG and its gradient with the ABP signal and its gradient. INN-PAR efficiently captures both forward and inverse mappings simultaneously, thereby preventing information loss. By integrating signal gradients into the learning process, INN-PAR enhances the network’s ability to capture essential high-frequency details, leading to more accurate signal reconstruction. Moreover, we propose a multi-scale convolution module (MSCM) within the invertible block, enabling the model to learn features across multiple scales effectively. We have experimented on two benchmark datasets, which show that INN-PAR significantly outperforms the state-of-the-art methods in both waveform reconstruction and BP measurement accuracy. Codes can be found at: https://github.com/soumitra1992/INNPAR-PPG2ABP. Soumitra Kundu, Gargi Panda, Saumik Bhattacharya, Aurobinda Routray, Rajlakshmi Guha |
ICASSP | 3 |
| 2025 | SINET: Sparsity-driven Interpretable Neural Network for Underwater Image EnhancementabstractImproving the quality of underwater images is essential for advancing marine research and technology. This work introduces a sparsity-driven interpretable neural network (SINET) for the underwater image enhancement (UIE) task. Unlike pure deep learning methods, our network architecture is based on a novel channel-specific convolutional sparse coding (CCSC) model, ensuring good interpretability of the underlying image enhancement process. The key feature of SINET is that it estimates the salient features from the three color channels using three sparse feature estimation blocks (SFEBs). The architecture of SFEB is designed by unrolling an iterative algorithm for solving the ℓ1regulaized convolutional sparse coding (CSC) problem. Our experiments show that SINET surpasses state-of-the-art PSNR value by 1.05 dB with 3873 times lower computational complexity. Code can be found at: https://github.com/gargi884/SINET-UIE/tree/main. Gargi Panda, Soumitra Kundu, Saumik Bhattacharya, Aurobinda Routray |
ICASSP | 3 |
| 2025 | FASTER: A Font-Agnostic Scene Text Editing and Rendering FrameworkabstractScene Text Editing (STE) is a challenging research prob-lem, that primarily aims towards modifying existing texts in an image while preserving the background and the font style of the original text. Despite its utility in numerous real-world applications, existing style-transfer-based approaches have shown sub-par editing performance due to (1) complex image backgrounds, (2) diverse font attributes, and (3) varying word lengths within the text. To address such limitations, in this paper, we propose a novel font-agnostic scene text editing and rendering framework, named FASTER, for simultaneously generating text in arbitrary styles and locations while preserving a natural and realistic appearance and structure. A combined fusion of target mask generation and style transfer units, with a cascaded self-attention mech-anism has been proposed to focus on multi-level text region edits to handle varying word lengths. Extensive evaluation on a real-world database withfurther subjective human eval-uation study indicates the superiority of FASTER in both scene text editing and rendering tasks, in terms of model per-formance and efficiency. The code and pre-trained models have been released in our Gi thub repo. Alloy Das, Sanket Biswas, Prasun Roy, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein, Josep Lladós 0001, Saumik Bhattacharya |
WACV | 8 |
| 2025 | 3D Shape Completion using Multi-resolution Spectral EncodingabstractReconstruction of intricate local patterns and large missing regions during 3D shape completion has the contradictory requirements of computation over a wider context and operations for finer detail restoration. To this end, we propose a multi-resolution spectral encoding based 3D shape completion approach to work on truncated Signed Distance Field (SDF) based shape representations. Our novelty lies in judiciously integrating multi-resolution 3D convolutional blocks that encode the input shape and a spectral module (SM) that captures the shape-wide context, thus addressing the contradictory requirements. SM acts on the features extracted from both partial input scans and shape priors using the multi-resolution convolutional blocks. Our SM contains a 3D convolutional block placed between fast Fourier transform (FFT) and inverse FFT operations, which results in the expansion of the receptive field for the appropriate context computation. Our approach has an attention-based encoder-decoder architecture, where the encoding of a partial scan is acted upon by shape prior encodings to produce attention maps. These attention maps are lever-aged differently in pretraining, and in the later training and inference stages of our approach to produce the reconstructed 3D shape. A surface gradient-based loss function is used in addition to the L1 loss, both in the pretraining and training stages for emphasizing the differences in minute details. These along with an attention refinement operation often leads to complete reconstruction while restoring finer details. Experiments using standard synthetic and real datasets demonstrate the superiority of our approach over the state-of-the-art. Pallabjyoti Deka, Saumik Bhattacharya, Debashis Sen, Prabir Kumar Biswas |
WACV | 2 |
| 2025 | A Novel Infogain and Multi-Axial Wavelet-Based Transformer for Personality Trait Question AnsweringabstractVisual Question Answering (VQA) is one of the attractive topics in the field of multimedia, affective, and empathic computing to garner user interest. Unlike existing models which aim at addressing challenges of VQA for the scene images, this work aims at developing a new model for Personality Trait Question Answering (PQA). It uses Twitter account information, which includes shared images, profile pictures, banners, text in the images, and descriptions of the images. Motivated by the accomplishments of the transformer, for encoding visual features of the images, a new InfoGain Multi-Axial Wavelet Vision Transformer (IgMaWaViT) is explored here. For encoding textual features in the images and descriptions, a new Information Gain BERT (InfoBert) method is introduced, which can handle the variable length encoding of text by choosing the optimal discriminator. Furthermore, the model fuses encodings of images and text according to the questions on different personality traits for question answering. The model is called InfoGain Multi-Axial Wavelet Vision Transformer for Personality Traits Question Answering (IgMaWaViT-PQA). To validate the efficacy of the proposed model, a dataset has been constructed, and it is used along with standard datasets for experimentation. Comprehensive experiments show that the proposed model is better than the state-of-the-art models. The code is available at the link: https://github.com/biswaskunal29/InfoGain_MultiAxial_PQA . Kunal Biswas, Palaiahnakote Shivakumara, Saumik Bhattacharya, Umapada Pal 0001, Ram Sarkar |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2025 | Decorrelation-Based Self-Supervised Visual Representation Learning for Writer IdentificationabstractSelf-supervised learning has developed rapidly over the last decade and has been applied in many areas of computer vision. Decorrelation-based self-supervised pretraining has shown great promise among non-contrastive algorithms, yielding performance at par with supervised and contrastive self-supervised baselines. In this work, we explore the decorrelation-based paradigm of self-supervised learning and apply the same to learning disentangled stroke features for writer identification. Here, we propose a modified formulation of the decorrelation-based framework named SWIS which was proposed for signature verification by standardizing the features along each dimension on top of the existing framework. We show that the proposed framework outperforms the contemporary self-supervised learning framework on the writer identification benchmark by 0.89%, 0.56%, and 1.23% on word-level images and 1.15%, 0.10%, and 0.39% on page-level images on IAM, CVL ,and Firemaker datasets, respectively. The proposed framework achieves word level accuracy of 87.94%, 84.80%, 93.32%, 74.24% and page level accuracy of 97.09%, 95.58%, 96.87%, 98.40% on AHAWP, IAM, CVL and Firemaker datasets, respectively, outperforming several recent supervised methods as well. Arkadip Maitra, Shree Mitra, Siladittya Manna, Saumik Bhattacharya, Umapada Pal 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
Alloy Das, Sanket Biswas, Umapada Pal 0001, Josep Lladós 0001, Saumik Bhattacharya |
ICPR (20) | 5 |
| 2024 | A New StyleGAN Latent Space Based Model for Image Style Transfer
Rakesh Dey, Palaiahnakote Shivakumara, Saumik Bhattacharya, Sukalpa Chanda, Umapada Pal 0001 |
ICPR (11) | 3 |
| 2024 | λ-Color: Amplifying Long-Range Dependencies for Image Colorization
Subhankar Ghosh, Saumik Bhattacharya, Prasun Roy, Umapada Pal 0001, Michael Blumenstein |
ICPR (22) | 2 |
| 2024 | Correlation Weighted Prototype-Based Self-supervised One-Shot Segmentation of Medical Images
Siladittya Manna, Saumik Bhattacharya, Umapada Pal 0001 |
ICPR (10) | 2 |
| 2024 | d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 2 |
| 2024 | Semantically Consistent Person Image Generation
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 2 |
| 2024 | Harnessing the Power of Multi-Lingual Datasets for Pre-training: Towards Enhancing Text Spotting PerformanceabstractThe adaptation capability to a wide range of domains is crucial for scene text spotting models when deployed to real-world conditions. However, existing SOTA approaches usually incorporate scene text detection and recognition simply by pretraining on natural scene text datasets, which do not directly exploit the intermediate feature representations between multiple domains. Here, we investigate the problem of domain-adaptive scene text spotting, i.e., training a model on multi-domain source data such that it can directly adapt to target domains rather than being specialized for a specific domain or scenario. Further, we investigate a transformer baseline called Swin-TESTR to focus on solving scene-text spotting for both regular and arbitraryshaped text along with an exhaustive evaluation. The results demonstrate the potential of intermediate representations to gain significant performance on text spotting benchmarks across multiple domains (e.g. language, synth-to-real, and documents). both in terms of accuracy and efficiency. Alloy Das, Sanket Biswas, Ayan Banerjee 0002, Josep Lladós 0001, Umapada Pal 0001, Saumik Bhattacharya |
WACV | 6 |
| 2024 | TIC: text-guided image colorization using conditional generative modelabstractAbstract Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make the colorization pipeline automatic, these processes often produce unrealistic results due to a lack of conditioning. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To the best of our knowledge, this is one of the first attempts to incorporate textual conditioning in the colorization pipeline. To do so, a novel deep network has been proposed that takes two inputs (the grayscale image and the respective encoded text description) and tries to predict the relevant color gamut. As the respective textual descriptions contain color information of the objects present in the scene, the text encoding helps to improve the overall quality of the predicted colors. The proposed model has been evaluated using different metrics like SSIM, PSNR, LPISPS and achieved scores of 0.917, 23.27,0.223, respectively. These quantitative metrics have shown that the proposed method outperforms the SOTA techniques in most of the cases. Subhankar Ghosh, Prasun Roy, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
Multim. Tools Appl. | 3 |
| 2024 | An end-to-end model for multi-view scene text recognition
Ayan Banerjee 0002, Palaiahnakote Shivakumara, Saumik Bhattacharya, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2023 | SelfDocSeg: A Self-supervised Vision-Based Approach Towards Document Segmentation
Subhajit Maity, Sanket Biswas, Siladittya Manna, Ayan Banerjee 0002, Josep Lladós 0001, Saumik Bhattacharya, Umapada Pal 0001 |
ICDAR (1) | 6 |
| 2023 | Multi-scale attention guided pose transfer
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001 |
Pattern Recognit. | 2 |
| 2022 | TIPS: Text-Induced Pose Synthesis
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ECCV (38) | 3 |
| 2022 | Multi-Latent GAN Inversion for Unsupervised 3D Shape CompletionabstractThe objective of 3-dimensional point cloud completion is to estimate a plausible complete shape from a given partial point cloud. Most of the data-driven point cloud completion approaches have been proposed in a supervised manner needing one-to-one correspondence between the partial and complete shapes. A promising way to solve the paired data dependency is to use the mapping capability of a pre-trained point cloud generation network to the best possible matching latent vector. However, recovering the composite structural details and complex geometry of a 3D shape is often difficult using a single latent vector alone. In this paper, we propose to employ multiple latent vectors, each of which generates individual feature maps, which are then combined to reconstruct a faithful complete 3D shape corresponding to an available partial shape. Deploying more than one latent vector enables the pre-trained generative network to increase its fidelity by using multiple combinations of feature representations learned by each single latent. Experimental results show that our algorithm performs well compared to the other existing shape completion methods. We also study the completion performance with a varying number of latent codes and the role of each latent vector in the final complete shape generation. Krishnendu Ghosh, Aupendu Kar, Saumik Bhattacharya, Debashis Sen, Prabir Kumar Biswas |
ICIP | 3 |
| 2022 | SWIS: Self-Supervised Representation Learning for Writer Independent Offline Signature VerificationabstractWriter independent offline signature verification is one of the most challenging tasks in pattern recognition as there is often a scarcity of training data. To handle such data scarcity problem, in this paper, we propose a novel self-supervised learning (SSL) framework for writer independent offline signature verification. To our knowledge, this is the first attempt to utilize self-supervised setting for the signature verification task. The objective of self-supervised representation learning from the signature images is achieved by minimizing the cross-covariance between two random variables belonging to different feature directions and ensuring a positive cross-covariance between the random variables denoting the same feature direction. This ensures that the features are decorrelated linearly and the redundant information is discarded. Through experimental results on different data sets, we obtained encouraging results. Siladittya Manna, Soumitri Chattopadhyay, Saumik Bhattacharya, Umapada Pal 0001 |
ICIP | 3 |
| 2022 | SURDS: Self-Supervised Attention-guided Reconstruction and Dual Triplet Loss for Writer Independent Offline Signature VerificationabstractOffline Signature Verification (OSV) is a fundamental biometric task across various forensic, commercial and legal applications. The underlying task at hand is to carefully model fine-grained features of the signatures to distinguish between genuine and forged ones, which differ only in minute deformities. This makes OSV more challenging compared to other verification problems. In this work, we propose a two-stage deep learning framework that leverages self-supervised representation learning as well as metric learning for writer-independent OSV. First, we train an image reconstruction network using an encoder-decoder architecture that is augmented by a 2D spatial attention mechanism using signature image patches. Next, the trained encoder backbone is fine-tuned with a projector head using a supervised metric learning framework, whose objective is to optimize a novel dual triplet loss by sampling negative samples from both within the same writer class as well as from other writers. The intuition behind this is to ensure that a signature sample lies closer to its positive counterpart compared to negative samples from both intra-writer and cross-writer sets. This results in robust discriminative learning of the embedding space. To the best of our knowledge, this is the first work of using self-supervised learning frameworks for OSV. The proposed two-stage framework has been evaluated on two publicly available offline signature datasets and compared with various state-of-the-art methods. It is noted that the proposed method provided promising results outperforming several existing pieces of work. The code is publicly available at: https://github.com/soumitri2001/SURDS-SSL-OSV. Soumitri Chattopadhyay, Siladittya Manna, Saumik Bhattacharya, Umapada Pal 0001 |
ICPR | 3 |
| 2022 | Attention W-Net: Improved Skip Connections for Better RepresentationsabstractSegmentation of macro and microvascular structures in fundoscopic retinal images plays a crucial role in the detection of multiple retinal and systemic diseases, yet it is a difficult problem to solve. Most neural network approaches face several issues such as lack of enough parameters, overfitting and/or incompatibility between internal feature-spaces. We propose Attention W-Net, a new U-Net based architecture for retinal vessel segmentation to address these problems. In this architecture, we have two main contributions: Attention Block and regularisation measures. Our Attention Block uses attention between encoder and decoder features, resulting in higher compatibility upon addition. Our regularisation measures include augmentation and modifications to the ResNet Block used, which greatly prevent overfitting. We observe an F1 and AUC of 0.8407 and 0.9833 on the DRIVE and 0.8174 and 0.9865 respectively on the CHASE-DB1 datasets — a sizeable improvement over its backbone as well as competitive performance among contemporary state-of-the-art methods. Shikhar Mohan, Saumik Bhattacharya, Sayantari Ghosh |
ICPR | 2 |
| 2022 | Scene Aware Person Image Generation through Global Contextual ConditioningabstractPerson image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results. Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ICPR | 3 |
| 2022 | Compressed Sensing MRI Reconstruction with Co-VeGAN: Complex-Valued Generative Adversarial NetworkabstractCompressed sensing (CS) is extensively used to reduce magnetic resonance imaging (MRI) acquisition time. State-of-the-art deep learning-based methods have proven effective in obtaining fast, high-quality reconstruction of CS-MR images. However, they treat the inherently complex-valued MRI data as real-valued entities by extracting the magnitude content or concatenating the complex-valued data as two real-valued channels for processing. In both cases, the phase content is discarded. To address the fundamental problem of real-valued deep networks, i.e. their inability to process complex-valued data, we propose a complex-valued generative adversarial network (Co-VeGAN) framework, which is the first-of-its-kind generative model exploring the use of complex-valued weights and operations. Further, since real-valued activation functions do not generalize well to the complex-valued space, we propose a novel complex-valued activation function that is sensitive to the input phase and has a learnable profile. Extensive evaluation of the proposed approach1on different datasets demonstrates that it significantly outperforms the existing CS-MRI reconstruction techniques. Bhavya Vasudeva, Puneesh Deora, Saumik Bhattacharya, Pyari Mohan Pradhan |
WACV | 3 |
| 2022 | Self-supervised representation learning for detection of ACL tear injury in knee MR videos
Siladittya Manna, Saumik Bhattacharya, Umapada Pal 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | LoOp: Looking for Optimal Hard Negative Embeddings for Deep Metric LearningabstractDeep metric learning has been effectively used to learn distance metrics for different visual tasks like image retrieval, clustering, etc. In order to aid the training process, existing methods either use a hard mining strategy to extract the most informative samples or seek to generate hard synthetics using an additional network. Such approaches face different challenges and can lead to biased embeddings in the former case, and (i) harder optimization (ii) slower training speed (iii) higher model complexity in the latter case. In order to overcome these challenges, we propose a novel approach that looks for optimal hard negatives (LoOp) in the embedding space, taking full advantage of each tuple by calculating the minimum distance between a pair of positives and a pair of negatives. Unlike mining-based methods, our approach considers the entire space between pairs of embeddings to calculate the optimal hard negatives. Extensive experiments combining our approach and representative metric learning losses reveal a significant boost in performance on three benchmark datasets1. Bhavya Vasudeva, Puneesh Deora, Saumik Bhattacharya, Umapada Pal 0001, Sukalpa Chanda |
ICCV | 3 |
| 2020 | STEFANN: Scene Text Editor Using Font Adaptive Neural NetworkabstractTextual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our knowledge, there is no significant prior work that aims to modify the textual information in an image. The ability to edit text directly on images has several advantages including error correction, text restoration and image reusability. In this paper, we propose a method to modify text in an image at character-level. We approach the problem in two stages. At first, the unobserved character (target) is generated from an observed character (source) being modified. We propose two different neural network architectures - (a) FANnet to achieve structural consistency with source font and (b) Colornet to preserve source color. Next, we replace the source character with the generated character maintaining both geometric and visual consistency with neighboring characters. Our method works as a unified platform for modifying text in images. We present the effectiveness of our method on COCO-Text and ICDAR datasets both qualitatively and quantitatively. Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001 |
CVPR | 2 |
| 2018 | Visual Saliency Detection Using Spatiotemporal DecompositionabstractWe propose a novel technique for detection of visual saliency in dynamic video based on video decomposition. The decomposition obtains the sparse features in a particular orientation by exploiting the spatiotemporal discontinuities present in a video cube. A weighted sum of the sparse features along three orthogonal directions determines the salient regions in the video cubes. The weights computed using the frame correlation along three directions are based on the characteristic of human visual system that identifies the sparsest feature as the most salient feature in a video. Unlike the existing methods, which detect the salient region as blob, the proposed approach detects the exact boundaries of salient region with minimum false detection. The experimental results confirm that the detected salient regions of a video closely resemble the salient regions detected by actual tracking of human eyes. The algorithm is tested on different types of video contents and compared with the several state-of-the-art methods to establish the effectiveness of the proposed method. Saumik Bhattacharya, K. S. Venkatesh, Sumana Gupta |
IEEE Trans. Image Process. | 1 |
| 2017 | Improved scene capture in unfavorable lighting conditionsabstractReal world scenes have huge intensity variations which are not in the control of the capture process. While human eye has an excellent dynamic range that enables us to visualize precise contrast variations and dynamically adapts to illumination variations, the dynamic range of conventional imaging devices is limited because of the physical constraints of the sensors. As a result of limited capabilities of the sensors, image saturation is observed often when lighting conditions are unfavorable (very bright, dark or uneven). In such scenarios, the captured image will have some optimally illuminated parts while some parts may undergo saturation (underexposure or overexposure). This makes the captured scene visually unappealing and the capture suffers from significant information loss. In this work, we propose an imaging solution to recover the scene information lost due to saturation, and hence, produce a better quality image ensuring no or minimal saturation. Megha Nawhal, Saumik Bhattacharya, K. S. Venkatesh |
ICIP | 2 |
| 2017 | Spatiotemporal Colorization of Video Using 3D Steerable PyramidsabstractWe propose a new technique for video colorization based on spatiotemporal color propagation in the 3D video volume, utilizing the dominant orientation response obtained from the steerable pyramid decomposition of the video. The volumetric color diffusion from the sources that are marked by scribbles occurs across spatiotemporally smooth regions, and the prevention of leakage is facilitated by the spatiotemporal discontinuities in the output of steerable filters, representing the object boundaries and motion boundaries. Unlike most existing methods, our approach dispenses with the need of motion vectors for interframe color transfer and provides a general framework for image and video colorization. The experimental results establish the effectiveness of the proposed approach in colorizing videos having different types of motion and visual content even in the presence of occlusion, in terms of accuracy and computational requirements. Somdyuti Paul, Saumik Bhattacharya, Sumana Gupta |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Visual saliency detection using video decompositionabstractEstimation of salient regions in an input video is an active area of research due to its wide applications. In this paper, we propose a novel algorithm to estimate the eye gaze movement in a video using motion, color and structural cues with minimum outliers. The algorithm is generalized to capture salient information for the videos taken under different camera motions. The entire algorithm is parallelizable and ensures faster estimation of salient regions. Using different standard datasets, the estimations of proposed algorithm are compared with state-of-the-art approaches. It is observed that the proposed method produces estimations closer to the ground-truth eye tracker data with minimum outliers. Saumik Bhattacharya, Sumana Gupta, K. S. Venkatesh |
ICIP | 1 |
| 2016 | Dehazing of color image using stochastic enhancementabstractImages captured in presence of fog, haze or snow usually suffer from poor contrast and visibility. In this paper we propose a novel dehazing method to increase visibility from a single view without using any prior knowledge about the outdoor scene. The proposed method estimates a visibility map of the scene from the input image and uses stochastic iterative algorithm to remove fog and haze. The method can be applied to color and grayscale images. Experimental results show that the proposed algorithm outperforms most of the state-of-the-art algorithms in terms of contrast, colorfulness and visibility. Saumik Bhattacharya, Sumana Gupta, K. S. Venkatesh |
ICIP | 1 |