VLDB 2026 Research / reviewers in the wild / expert
Subhasis Chaudhuri
dblp:c/SubhasisChaudhuri
· DBLP profile ↗
139ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-1680-0016ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 68 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | X-JEPA: A Novel Joint Learning Cross-Modal Predictive Alignment Framework for Remote Sensing Image RetrievalabstractThe growing scale and heterogeneity of remote sensing (RS) imagery demand robust, scalable frameworks for content-based image retrieval across sensor modalities. We introduce X-JEPA, a novel predictive self-supervised architecture explicitly designed for cross-modal remote sensing image retrieval (RS-CMIR), and the first to extend joint embedding predictive paradigms beyond unimodal domains. Unlike prior contrastive or reconstruction-based methods, X-JEPA formulates representation learning as a latent forecasting task: predicting the semantic embedding of a target modality given context from another. To enforce modality-invariant alignment, we propose a geometry-aware Prediction Space Alignment (PSA) loss, which captures the structure of the latent space without requiring pixel-level reconstruction or modality pairing. We evaluate X-JEPA on two large-scale benchmarks—BEN-14K (Sentinel-1/Sentinel-2) and fMoW (RGB/Sentinel) across both unimodal and cross-modal retrieval tasks. X-JEPA consistently outperforms state-of-the-art self-supervised baselines, including MAE, SatMAE, CrossMAE, CSMAE-SESD, CROMA, SkySense, DeCUR, and REJEPA, achieving up to 11.0% F1-score improvement in cross-modal retrieval and 9.8% in unimodal settings. Despite its high retrieval accuracy, the model remains lightweight, requiring fewer parameters and yielding 8–10% F1-score gains on average, establishing a new state-of-the-art for scalable, sensor-agnostic RS-CMIR.1 Shabnam Choudhury, Yash Salunkhe, Vaibhav Rajan, Subhasis Chaudhuri, Biplab Banerjee |
WACV | 4 |
| 2026 | Vision-informed Semantic Text Alignment for Open-set Recognition in Remote SensingabstractExisting Open-Set Recognition (OSR) methods struggle in remote sensing (RS) as their reliance on unimodal visual features fails to resolve the severe inter-class similarity inherent in overhead imagery. To address this, we propose ViSTA-RS, a novel multimodal framework that leverages semantic context from language to disambiguate visually similar scenes. Our approach first constructs semantically-rich class prototypes by jointly encoding images with generated text captions using a Vision-Language Model. We then introduce a reconstruction-based mechanism where an image’s visual embedding is expressed as a weighted combination of these semantic prototypes. The magnitude of the reconstruction error serves as a robust novelty score, with a statistically principled threshold determined by Extreme Value Theory (EVT). This alignment of multimodal semantics with prototype reconstruction is uniquely suited for the fine-grained nature of RS data. On four challenging benchmarks, ViSTA-RS sets a new state-of-the-art, improving the AUROC for unknown detection by a significant 6.7% over leading baselines while maintaining high accuracy on known classes. Siddhant Gole, Akash Pal, Ankit Jha, Subhasis Chaudhuri, Biplab Banerjee |
WACV | 4 |
| 2026 | SCOPE: Segmenting common objects with prompt-conditioned encoding and SAM distillationabstractCo-segmentation aims to identify and segment common objects across a set of related images, requiring consistent semantic understanding despite contextual variations. While foundation models such as SAM have demonstrated remarkable success across diverse vision tasks, their potential for co-segmentation remains largely unexplored. In this paper, we propose a novel framework that distills feature-level knowledge from the Segment Anything Model (SAM) into a Swin backbone, enhancing semantic consistency and generalization. To better align with the co-segmentation objective, we integrate learnable prompts into the Swin backbone. The resulting hierarchical features are processed using intra-image and inter-image attention mechanisms to capture correlations within and across the images. These features are further refined by a noise suppression module, and the SAM decoder is used to produce high-quality segmentation masks. We train the network using a contrastive loss tailored for co-segmentation, alongside conventional objectives. Extensive experiments on four challenging benchmarks-PASCAL-VOC, Internet, iCoseg, and MSRC-demonstrate that our method achieves state-of-the-art performance. Shruthi Akkala, Tanisha Chawada, Saikat Dutta 0002, Subhasis Chaudhuri, Biplab Banerjee |
Pattern Recognit. Lett. | 4 |
| 2025 | Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud Segmentationabstract3D point cloud segmentation is essential across a range of applications; however, conventional methods often struggle in evolving environments, particularly when tasked with identifying novel categories under limited supervision. Few-Shot Learning (FSL) and Class Incremental Learning (CIL) have been adapted previously to address these challenges in isolation, yet the combined paradigm of Few-Shot Class Incremental Learning (FSCIL) remains largely unexplored for point cloud segmentation. To address this gap, we introduce Hyperbolic Ideal Prototypes Optimization (HIPO), a novel framework that harnesses hyperbolic embeddings for FSCIL in 3D point clouds. HIPO employs the Poincaré Hyperbolic Sphere as its embedding space, integrating Ideal Prototypes enriched by CLIP-derived class semantics, to capture the hierarchical structure of 3D data. By enforcing orthogonality among prototypes and maximizing representational margins, HIPO constructs a resilient embedding space that mitigates forgetting and enables the seamless integration of new classes, thereby effectively countering overfitting. Extensive evaluations on S3DIS, ScanNetv2, and cross-dataset scenarios demonstrate HIPO’s strong performance, significantly surpassing existing approaches in both in-domain and cross-dataset FSCIL tasks for 3D point cloud segmentation. Tanuj Sur, Samrat Mukherjee, Kaizer Rahaman, Subhasis Chaudhuri, Muhammad Haris Khan, Biplab Banerjee |
CVPR | 4 |
| 2025 | UIDAPLE: Unsupervised Incremental Domain Adaptation through Adaptive Prompt LearningabstractContinual learning poses significant challenges for deep neural networks, notably catastrophic forgetting, particularly when faced with shifting data distributions that compromise previously acquired knowledge. This paper tackles these issues within the Unsupervised Incremental Domain Adaptation (UIDA) framework, where the initial source domain is labeled, but subsequent domains are not. Existing methods often struggle with limited cross-domain generalization and adaptation capabilities. As a remedy, we introduce UIDAPLE, a novel approach that utilizes a unified prompt across all domains, leveraging the foundation model CLIP to obviate the need for isolated domain treatments. Specifically, UIDAPLE implements supervised prompt learning in the labeled source domain and extends this learning to unlabeled domains through confidence-based adaptation. We also present an efficient parameter alignment strategy that maintains semantic coherence across domains, effectively balancing stability and plasticity to combat catastrophic forgetting. Extensive evaluations on two benchmark datasets reveal that UIDAPLE markedly surpasses other UIDA techniques in performance. Samrat Mukherjee, Tanuj Sur, Saurish Seksaria, Subhasis Chaudhuri, Gemma Roig, Biplab Banerjee |
ICASSP | 4 |
| 2024 | DARK: Few-Shot Remote-Sensing Colorization Using Label-Conditioned Color InjectionabstractSatellite image colorization is a broad challenging problem in the domain of remote sensing (RS) having huge potential applications. The problem becomes even more complicated under the few-shot setting yet it has barely been studied to date. In this paper, we propose a colorization framework for the RS scene for synthesizing optical images from their panchromatic (PAN) counterparts using color injection and attention fusion mechanism. Our proposed model ensures that the synthesized optical images are coherent with the structural variability of panchromatic images while constraining the realistic appearance in the optical domain from a few training image pairs. To accomplish the same, we introduce a novel Dual Attention fusion of Receptive Kernels (DARK) which considers the spatial nuances along with color injection conditioned on prior label allocation. DARK is a multi spectral-spatial feature generator that selectively accentuates important cross-spatial features based on attention fusion. We also employ a prior distribution constraint on color embedding generation for introducing vibrant yet diverse variance in a color generation. Our approach achieves state-of-the-art results on the publicly available EuroSAT and PatternNet datasets while demonstrating significant speedups. We showcase our results quantitatively by comparing the PSNR, mean squared error(MSE), and cosine similarity of generated images and qualitatively via visual perception. Rupak Bose, Anshul Shrivastava, Biplab Banerjee, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Remeshing-free Graph-based Finite Element Method for Fracture SimulationabstractFracture produces new mesh fragments that introduce additional degrees of freedom in the system dynamics. Existing finite element method (FEM) based solutions suffer from an explosion in computational cost as the system matrix size increases. We solve this problem by presenting a graph-based FEM model for fracture simulation that is remeshing-free and easily scales to high-resolution meshes. Our algorithm models fracture on the graph induced in a volumetric mesh with tetrahedral elements. We relabel the edges of the graph using a computed damage variable to initialize and propagate fracture. We prove that non-linear, hyper-elastic strain energy is expressible entirely in terms of the edge lengths of the induced graph. This allows us to reformulate the system dynamics for the relabeled graph without changing the size of system dynamics matrix and thus prevents the computational cost from blowing up. The fractured surface has to be reconstructed explicitly only for visualization purposes. We simulate standard laboratory experiments from structural mechanics and compare the results with corresponding real-world experiments. We fracture objects made of a variety of brittle and ductile materials, and show that our technique offers stability and speed that is unmatched in current literature. Avirup Mandal, Parag Chaudhuri, Subhasis Chaudhuri |
Comput. Graph. Forum | 3 |
| 2023 | Theoretical Analysis of Null Foley-Sammon Transform and its ImplicationsabstractNull Foley-Sammon Transform (NFST) has received increasing attention in the machine learning and pattern recognition literature. NFST finds a discriminative nullspace where all samples of the same class get mapped into a single point. It has a closed form solution and is free of parameters to tune. NFST has been leveraged in many areas including novelty detection, person or vehicle re-identification and achieved state-of-the-art results. Motivated from its attractive properties and its effectiveness in wide range of applications, in this paper we focus on the theoretical analysis of NFST. In previous literature, NFST was shown to exist in small sample size (SSS) case. We first prove that NFST can exist in non-SSS case also, under certain conditions. Thereby, we extend the domain of applicability of NFST to a more general case. Secondly, we perform analysis of the singular points of NFST, revealing important insights on their identities and existence. Thirdly, we show the theoretical relation between NFST of SSS data and NFST of the non-SSS data obtained by PCA. Fourthly, we show that this theoretical relation can be exploited to obtain an efficient algorithm for computing NFST on high dimensional SSS data. Finally, we perform extensive experiments to validate our theoretical analysis. T. M. Feroz Ali, Subhasis Chaudhuri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | 3D-NVS: A 3D Supervision Approach for Next View SelectionabstractWe present a classification-based approach for the best next view selection and show how we can plausibly obtain a supervisory signal for this task. The proposed approach is end-to-end trainable and aims to get the best possible 3D reconstruction quality with an actively selected second view, given a passively chosen initial view. The proposed model consists of two stages: a classifier and a reconstructor network trained directly from ground truth voxels, as opposed to exhaustively selecting ground truth of best pair views. While testing, the proposed method assumes no prior knowledge of the underlying 3D shape for selecting the next best view. We demonstrate the proposed method’s effectiveness via detailed experiments on synthetic and real images and show how it provides improved reconstruction quality than the existing state of the art 3D reconstruction and the next best view prediction techniques. Kumar Ashutosh, Saurabh Kumar 0005, Subhasis Chaudhuri |
ICPR | 3 |
| 2022 | Simulating Fracture in Anisotropic Materials Containing ImpuritiesabstractFracture simulation of real-world materials is an exceptionally challenging problem due to complex material properties like anisotropic elasticity and the presence of material impurities. We present a graph-based finite element method to simulate dynamic fracture in anisotropic materials. We further enhance this model by developing a novel probabilistic damage mechanics for modelling materials with impurities using a random graph-based formulation. We demonstrate how this formulation can be used by artists for directing and controlling fracture. We simulate and render fractures for a diverse set of materials to demonstrate the potency and robustness of our methods. Avirup Mandal, Parag Chaudhuri, Subhasis Chaudhuri |
MIG | 3 |
| 2022 | Semantics-Driven Generative Replay for Few-Shot Class Incremental LearningabstractWe deal with the problem of few-shot class incremental learning (FSCIL), which requires a model to continuously recognize new categories for which limited training data are available. Existing FSCIL methods depend on prior knowledge to regularize the model parameters for combating catastrophic forgetting. Devising an effective prior in a low-data regime, however, is not trivial. The memory-replay based approaches from the fully-supervised class incremental learning (CIL) literature cannot be used directly for FSCIL as the generative memory-replay modules of CIL are hard to train from few training samples. However, generative replay can tackle both the stability and plasticity of the models simultaneously by generating a large number of class-conditional samples. Convinced by this fact, we propose a generative modeling-based FSCIL framework using the paradigm of memory-replay in which a novel conditional few-shot generative adversarial network (GAN) is incrementally trained to produce visual features while ensuring the stability-plasticity trade-off through novel loss functions and combating the mode-collapse problem effectively. Furthermore, the class-specific synthesized visual features from the few-shot GAN are constrained to match the respective latent semantic prototypes obtained from a well-defined semantic space. We find that the advantages of this semantic restriction is two-fold, in dealing with forgetting, while making the features class-discernible. The model requires a single per-class prototype vector to be maintained in a dynamic memory buffer. Experimental results on the benchmark and large-scale CiFAR-100, CUB-200, and Mini-ImageNet confirm the superiority of our model over the current FSCIL state of the art. Aishwarya Agarwal, Biplab Banerjee, Fabio Cuzzolin, Subhasis Chaudhuri |
ACM Multimedia | 4 |
| 2022 | FRIDA - Generative feature replay for incremental domain adaptation
Sayan Rakshit, Anwesh Mohanty, Ruchika Chavhan, Biplab Banerjee, Gemma Roig, Subhasis Chaudhuri |
Comput. Vis. Image Underst. | 6 |
| 2021 | SemGIF: A Semantics Guided Incremental Few-shot Learning Framework with Generative Replay
S. Divakar Bhat, Biplab Banerjee, Subhasis Chaudhuri |
BMVC | 3 |
| 2021 | A Unified Batch Selection Policy for Active Metric Learning
Priyadarshini Kumari, Siddhartha Chaudhuri, Vivek S. Borkar, Subhasis Chaudhuri |
ECML/PKDD (2) | 4 |
| 2021 | ADA-AT/DT: An Adversarial Approach for Cross-Domain and Cross-Task Knowledge TransferabstractWe deal with the problem of cross-task and cross-domain knowledge transfer in the realm of scene understanding for autonomous vehicles. We consider the scenario where supervision is available for a pair of tasks in a source domain while it is available for only one of the tasks in the target domain. Given that, the goal is to perform inference for the task in the target which is devoid of any training information. We argue that the only reported work in learning across tasks and domains (AT/DT) [26] faces the problem of domain shift between the source and target domains, hindering predictions on the target domain when the transfer of knowledge is learned on a statistically different yet related source domain. As a remedy, we develop a novel framework called ADA-AT/DT based on the adversarial training strategy to ensure that the domain-gaps are minimized for the common cross-domain supervised task. This, in effect, helps in realizing a domain-independent task-transfer function that eventually helps in performing improved inference in the target domain. We demonstrate that our proposed method significantly outperforms [26] by using models with 81% fewer trainable parameters. In addition, we perform experiments on a transformation mapping similar to U-Net to ensure maximum exploitation of features for task transfer. Extensive experiments have been performed on four different domains (Synthia, CityScapes, Carla, and KITTI) for two visual tasks (depth estimation and semantic segmentation) to confirm the superiority of our method. Ruchika Chavhan, Ankit Jha, Biplab Banerjee, Subhasis Chaudhuri |
WACV | 4 |
| 2021 | Improved Landcover Classification using Online Spectral Data Hallucination
Saurabh Kumar 0005, Biplab Banerjee, Subhasis Chaudhuri |
Neurocomputing | 3 |
| 2020 | Multi-source Open-Set Deep Adversarial Domain Adaptation
Sayan Rakshit, Dipesh Tamboli, Pragati Shuddhodhan Meshram, Biplab Banerjee, Gemma Roig, Subhasis Chaudhuri |
ECCV (26) | 6 |
| 2020 | MT-UNET: A Novel U-Net Based Multi-Task Architecture For Visual Scene UnderstandingabstractWe tackle the problem of deep end-to-end multi-task learning (MTL) for jointly performing image segmentation and depth estimation from monocular images. It is proven already that learning several related tasks together helps in attaining improved performance per task than training them autonomously. To this end, we follow the typical U-Net based encoder-decoder architecture (MT-UNet) where the densely connected deep convolutional neural network (CNN) based feature encoder is shared among the tasks while the soft attention based task-specific decoder modules produce the desired outputs. Additionally, we encourage cross-talk (CT) between the tasks by introducing cross-task skip connections at the decoder end with adaptive weight learning for the task-specific loss functions in the final cost measure. We validate the proposed framework on the challenging CityScapes and NYUv2 datasets, where our method sharply outperforms the current state-of-the-art. Ankit Jha, Awanish Kumar, Shivam Pande, Biplab Banerjee, Subhasis Chaudhuri |
ICIP | 5 |
| 2020 | Directed Variational Cross-encoder Network for Few-shot Multi-image Co-segmentationabstractIn this paper, we propose a novel framework for multi-image co-segmentation using class agnostic meta-learning strategy by generalizing to new classes given only a small number of training samples for each new class. We have developed a novel encoder-decoder network termed as DVICE (Directed Variational Inference Cross Encoder), which learns a continuous embedding space to ensure better similarity learning. We employ a combination of the proposed DVICE network and a novel few-shot learning approach to tackle the small sample size problem encountered in co-segmentation with small datasets like iCoseg and MSRC. Furthermore, the proposed framework does not use any semantic class labels and is entirely class agnostic. Through exhaustive experimentation over multiple datasets using only a small volume of training data, we have demonstrated that our approach outperforms all existing state-of-the-art techniques. Sayan Banerjee, S. Divakar Bhat, Subhasis Chaudhuri, Rajbabu Velmurugan |
ICPR | 3 |
| 2020 | GuCNet: A Guided Clustering-based Network for Improved ClassificationabstractWe deal with the problem of semantic classification of challenging and highly-cluttered dataset. We present a novel, and yet a very simple classification technique by leveraging the ease of classifiability of any existing well separable dataset for guidance. Since the guide dataset which may or may not have any semantic relationship with the experimental dataset, forms well separable clusters in the feature set, the proposed network tries to embed class-wise features of the challenging dataset to those distinct clusters of the guide set, making them more separable. Depending on the availability, we propose two types of guide sets: one using texture (image) guides and another using prototype vectors representing cluster centers. Experimental results obtained on the challenging benchmark RSSCN, LSUN, and TU-Berlin datasets establish the efficacy of the proposed method as we outperform the existing state-of-the-art techniques by a considerable margin. Ushasi Chaudhuri, Syomantak Chaudhuri, Subhasis Chaudhuri |
ICPR | 3 |
| 2020 | A Novel Actor Dual-Critic Model for Remote Sensing Image CaptioningabstractWe deal with the problem of generating textual captions from optical remote sensing (RS) images using the notion of deep reinforcement learning. Due to the high inter-class similarity in reference sentences describing remote sensing data, jointly encoding the sentences and images encourages prediction of captions that are semantically more precise than the ground truth in many cases. To this end, we introduce an Actor Dual-Critic training strategy where a second critic model is deployed in the form of an encoder-decoder RNN to encode the latent information corresponding to the original and generated captions. While all actor-critic methods use an actor to predict sentences for an image and a critic to provide rewards, our proposed encoder-decoder RNN guarantees high-level comprehension of images by sentence-to-image translation. We observe that the proposed model generates sentences on the test data highly similar to the ground truth and is successful in generating even better captions in many critical cases. Extensive experiments on the benchmark Remote Sensing Image Captioning Dataset (RSICD) and the UCM-captions dataset confirm the superiority of the proposed approach in comparison to the previous state-of-the-art where we obtain a gain of sharp increments in both the ROUGE-L and CIDEr measures. Ruchika Chavhan, Biplab Banerjee, Xiao Xiang Zhu 0001, Subhasis Chaudhuri |
ICPR | 4 |
| 2020 | Batch Decorrelation for Active Metric LearningabstractWe present an active learning strategy for training parametric models of distance metrics, given triplet-based similarity assessments: object $x_i$ is more similar to object $x_j$ than to $x_k$. In contrast to prior work on class-based learning, where the fundamental goal is classification and any implicit or explicit metric is binary, we focus on perceptual metrics that express the degree of (dis)similarity between objects. We find that standard active learning approaches degrade when annotations are requested for batches of triplets at a time: our studies suggest that correlation among triplets is responsible. In this work, we propose a novel method to decorrelate batches of triplets, that jointly balances informativeness and diversity while decoupling the choice of heuristic for each criterion. Experiments indicate our method is general, adaptable, and outperforms the state-of-the-art. Priyadarshini Kumari, Ritesh Goru, Siddhartha Chaudhuri, Subhasis Chaudhuri |
IJCAI | 4 |
| 2020 | Generalized Zero-Shot Learning using Generated Proxy Unseen Samples and Entropy SeparationabstractThe recent generative model-driven Generalized Zero-shot Learning (GZSL) techniques overcome the prevailing issue of the model bias towards the seen classes by synthesizing the visual samples of the unseen classes through leveraging the corresponding semantic prototypes. Although such approaches significantly improve the GZSL performance due to data augmentation, they violate the principal assumption of GZSL regarding the unavailability of semantic information of unseen classes during training. In this work, we propose to use a generative model (GAN) for synthesizing the visual proxy samples while strictly adhering to the standard assumptions of the GZSL. The aforementioned proxy samples are generated by exploring the early training regime of the GAN. We hypothesize that such proxy samples can effectively be used to characterize the average entropy of the label distribution of the samples from the unseen classes. Further, we train a classifier on the visual samples from the seen classes and proxy samples using entropy separation criterion such that an average entropy of the label distribution is low and high, respectively, for the visual samples from the seen classes and the proxy samples. Such entropy separation criterion generalizes well during testing where the samples from the unseen classes exhibit higher entropy than the entropy of the samples from the seen classes. Subsequently, low and high entropy samples are classified using supervised learning and ZSL rather than GZSL. We show the superiority of the proposed method by experimenting on AWA1, CUB, HMDB51, and UCF101 datasets. Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri, Fabio Cuzzolin |
ACM Multimedia | 3 |
| 2020 | DeFraudNet: End2End Fingerprint Spoof Detection using Patch Level AttentionabstractIn recent years, fingerprint recognition systems have made remarkable advancements in the field of biometric security as it plays an important role in personal, national and global security. In spite of all these notable advancements, the fingerprint recognition technology is still susceptible to spoof attacks which can significantly jeopardize the user security. The cross sensor and cross material spoof detection still pose a challenge with a myriad of spoof materials emerging every day, compromising sensor interoperability and robustness. This paper proposes a novel method for fingerprint spoof detection using both global and local fingerprint feature descriptors. These descriptors are extracted using DenseNet which significantly improves cross-sensor, cross-material and cross-dataset performance. A novel patch attention network is used for finding the most discriminative patches and also for network fusion. We evaluate our method on four publicly available datasets: LivDet 2011, 2013, 2015 and 2017. A set of comprehensive experiments are carried out to evaluate cross-sensor, cross-material and cross-dataset performance over these datasets. The proposed approach achieves an average accuracy of 99.52%, 99.16% and 99.72% on LivDet 2017, 2015 and 2011 respectively outperforming the current state-of-the-art results by 3% and 4% for LivDet 2015 and 2011 respectively. B. V. S. Anusha, Sayan Banerjee, Subhasis Chaudhuri |
WACV | 3 |
| 2020 | Multi-timescale Trajectory Prediction for Abnormal Human Activity DetectionabstractA classical approach to abnormal activity detection is to learn a representation for normal activities from the training data and then use this learned representation to detect abnormal activities while testing. Typically, the methods based on this approach operate at a fixed timescale - either a single time-instant (e.g. frame-based) or a constant time duration (e.g. video-clip based). But human abnormal activities can take place at different timescales. For example, jumping is a short-term anomaly and loitering is a long-term anomaly in a surveillance scenario. A single and pre-defined timescale is not enough to capture the wide range of anomalies occurring with different time duration. In this paper, we propose a multi-timescale model to capture the temporal dynamics at different timescales. In particular, the proposed model makes future and past predictions at different timescales for a given input pose trajectory. The model is multi-layered where intermediate layers are responsible to generate predictions corresponding to different timescales. These predictions are combined to detect abnormal activities. In addition, we also introduce a single-camera abnormal activity dataset for research use that contains 483,566 annotated frames. Our experiments show that the proposed model can capture the anomalies of different time duration and outperforms existing methods. Royston Rodrigues, Neha Bhargava, Rajbabu Velmurugan, Subhasis Chaudhuri |
WACV | 4 |
| 2019 | Generalized Zero-shot Learning using Open Set Recognition
Omkar Gune, Amit More, Biplab Banerjee, Subhasis Chaudhuri |
BMVC | 4 |
| 2019 | CoSegNet: Image Co-segmentation using a Conditional Siamese Convolutional NetworkabstractThe objective in image co-segmentation is to jointly segment unknown common objects from a given set of images. In this paper, we propose a novel deep convolution neural network based end-to-end co-segmentation model. It is composed of a metric learning and decision network leading to a novel conditional siamese encoder-decoder network for estimating a co-segmentation mask. The role of the metric learning network is to find an optimum latent feature space where objects of the same class are closer and that of different classes are separated by a certain margin. Depending on the extracted features, the decision network decides whether input images have common objects or not and the encoder-decoder network produces a cosegmentation mask accordingly. Key aspects of the architecture are as follows. First, it is completely class agnostic and does not require any semantic information. Second, in addition to producing masks, the decoder network also learns similarity across image pairs that improves co-segmentation significantly. Experimental results reflect an excellent performance of our method compared to state of-the-art methods on challenging co-segmentation datasets. Sayan Banerjee, Avik Hati, Subhasis Chaudhuri, Rajbabu Velmurugan |
IJCAI | 3 |
| 2019 | DeepQuantizedCS: Quantized Compressive Video Recovery using Deep Convolutional NetworksabstractThis work proposes a deep learning based approach to sparse signal recovery from compressively sensed (temporally or spectrally collapsed) and single bit quantized measurements. We demonstrate the effectiveness and applicability of this technique with the recovery of video and hyperspectral volumes from such compressed data. The compressively sensed data is represented by single bit quantization using an ordered dithering scheme with modifications to this whole compressive acquisition pipeline for efficiency and ease of practical implementation. All this allows us to have a compressive acquisition setup which doubles as an extremely simple encoder, without a decoder in the loop and which is power, memory, and computationally very efficient, and is suitable for onboard compression applications. When used as a compression engine, the proposed pipeline, unlike existing methods, requires only basic elements namely, adders, multipliers and comparators to offer a significant compression ratio without a need for costly high precision ADCs or transform coding ASICs in the workflow. Saurabh Kumar 0005, Yagnesh Badiyani, Subhasis Chaudhuri |
ACM Multimedia | 3 |
| 2019 | On QoS-compliant telehaptic communication over shared networks
Vineet Gokhale, Jayakrishnan Nair 0001, Subhasis Chaudhuri, Jan Fesl |
Comput. Networks | 3 |
| 2019 | Graph convolutional network for multi-label VHR remote sensing scene recognition
Nagma Khan, Ushasi Chaudhuri, Biplab Banerjee, Subhasis Chaudhuri |
Neurocomputing | 4 |
| 2019 | A Pseudo-likelihood Approach for Geo-localization of Events from Crowd-sourced Sensor-MetadataabstractEvents such as live concerts, protest marches, and exhibitions are often video recorded by many people at the same time, typically using smartphone devices. In this work, we address the problem of geo-localizing such events from crowd-generated data. Traditional approaches for solving such a problem using multiple video sequences of the event would require highly complex computer vision (CV) methods, which are computation intensive and are not robust under the environment where visual data are collected through crowd-sourced medium. In the present work, we approach the problem in a probabilistic framework using only the sensor metadata obtained from smartphones. We model the event location and camera locations and orientations (camera parameters) as the hidden states in a Hidden Markov Model. The sensor metadata from GPS and the digital compass from user smartphones are used as the observations associated with the hidden states of the model. We have used a suitable potential function to capture the complex interaction between the hidden states (i.e., event location and camera parameters). The non-Gaussian densities involved in the model, such as the potential function involving hidden states, make the maximum-likelihood estimation intractable. We propose a pseudo-likelihood-based approach to maximize the approximate-likelihood, which provides a tractable solution to the problem. The experimental results on the simulated as well as real data show correct event geo-localization using the proposed method. When compared with several baselines the proposed method shows a superior performance. The overall computation time required is much smaller, since only the sensor metadata are used instead of visual data. Amit More, Subhasis Chaudhuri |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Structure Aligning Discriminative Latent Embedding for Zero-Shot Learning
Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri |
BMVC | 3 |
| 2018 | Maximum Margin Metric Learning over Discriminative Nullspace for Person Re-identification
T. M. Feroz Ali, Subhasis Chaudhuri |
ECCV (13) | 2 |
| 2018 | Co-Segmentation of Non-Homogeneous Image SetsabstractIn this paper, we formulate image co-segmentation as a classification problem in an unsupervised framework with the classes being the common foreground and the remaining regions in the image set. We first find a set of superpixels across all images with high feature similarity such that the constituent superpixels in individual images are spatially compact and label them as seed for the common foreground. Those superpixels with high background probability are labeled as respective seeds for multiple background classes. Seed computation here is unsupervised and automated unlike some semi-supervised methods. Then, we compute discriminative features that separate the initially labeled classes using linear discriminant analysis. We use these features to perform spatially constrained label propagation and obtain labels for the unlabeled regions, and iterate this process till the seed regions grow to the common object. Experimental results demonstrate excellent robustness properties even while processing non-homogeneous image sets where the common object is present only in majority of the images. Avik Hati, Subhasis Chaudhuri, Rajbabu Velmurugan |
ICIP | 2 |
| 2018 | Class Specific Coders for Hyper-Spectral Image ClassificationabstractIn this paper, we introduce the paradigm of class specific coders (CSC) for classification of hyper-spectral images (HSI). Apparently, CSC are defined as a set of distinct encoder-decoder (henceforth called a coder) networks where a given coder is trained on the samples of a particular class. In contrast to auto-encoders (AE) which learn an identity mapping of data in an unsupervised fashion, the CSC model, on the other hand, learns re-constructive mappings for all possible pairs of training samples for each class in separate coders. Further, for reducing redundancy, it is ensured that the latent space dimensions of the coders are orthogonal to each other. Once the CSC are trained, we introduce a feature encoding by concatenating the latent representations of all the class specific coders. Experimental results obtained on the benchmark Botswana and Indian Pines datasets ensure that the CSC model outperforms a number of AE variants significantly. Sanatan Sharma, Akashdeep Goel, Omkar Gune, Biplab Banerjee, Subhasis Chaudhuri |
ICIP | 5 |
| 2018 | Scene Recognition From Optical Remote Sensing Images Using Mid-Level Deep Feature MiningabstractWe solve the problem of scene recognition from very high-resolution optical satellite remote sensing (RS) images by exploring the notion of mid-level feature mining. The existing mid-level feature extraction techniques are based on applying feature encodings over a set of discriminatively selected localized feature descriptors from the images. Such techniques inherently suffer from two shortcomings: 1) the local descriptors are not enough discriminative, since they are mostly based on scale invariant feature transform (SIFT) like ad hoc features and 2) the definition of a robust ranking function to select discriminative local features is nontrivial. As a remedy, we propose a pattern mining-based approach for an efficient discovery of mid-level visual elements, which considers convolutional neural network features of the category-independent region proposals extracted from the images as the local descriptors. While the region proposals depict better semantic information than the SIFT like features, the proposed pattern mining strategy can efficiently highlight the correlations between such local descriptors and the class labels. Experimental results suggest that the proposed technique outperforms a number of existing mid-level feature descriptors for the standard optical RS data sets. Biplab Banerjee, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Tensorization of Multifrequency PolSAR Data for Classification Using an Autoencoder NetworkabstractA novel tensorization framework is proposed, which utilizes the Kronecker product to combine multifrequency polarimetric synthetic aperture radar data in conjunction with an artificial neural network (ANN) for classification. The ANN comprises of two stages, where an unsupervised stochastic sampling autoencoder learns an efficient representation and a supervised feed forward network performs classification. The proposed framework is demonstrated using multifrequency (C-, L-, and P-bands) data sets collected by the AIRSAR system. The classification performance of single tensor product of dual- and triple-band combinations is evaluated. It is observed that the classification accuracy of the tensor products outperforms single, as well as, the simple augmentation of the frequency bands. Shaunak De, Debanshu Ratha, Dikshya Ratha, Avik Bhattacharya, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Multilabel Remote Sensing Image Retrieval Using a Semisupervised Graph-Theoretic MethodabstractConventional supervised content-based remote sensing (RS) image retrieval systems require a large number of already annotated images to train a classifier for obtaining high retrieval accuracy. Most systems assume that each training image is annotated by a single label associated to the most significant semantic content of the image. However, this assumption does not fit well with the complexity of RS images, where an image might have multiple land-cover classes (i.e., multilabels). Moreover, annotating images with multilabels is costly and time consuming. To address these issues, in this paper, we introduce a semisupervised graph-theoretic method in the framework of multilabel RS image retrieval problems. The proposed method is based on four main steps. The first step segments each image in the archive and extracts the features of each region. The second step constructs an image neighborhood graph and uses a correlated label propagation algorithm to automatically assign a set of labels to each image in the archive by exploiting only a small number of training images annotated with multilabels. The third step associates class labels with image regions by a novel region labeling strategy, whereas the final step retrieves the images similar to a given query image by a subgraph matching strategy. Experiments carried out on an archive of aerial images show the effectiveness of the proposed method when compared with the state-of-the-art RS content-based image retrieval methods. Bindita Chaudhuri, Begüm Demir, Subhasis Chaudhuri, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | An image texture insensitive method for saliency detection
Avik Hati, Subhasis Chaudhuri, Rajbabu Velmurugan |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Congestion Control for Network-Aware Telehaptic CommunicationabstractTelehaptic applications involve delay-sensitive multimedia communication between remote locations with distinct Quality of Service (QoS) requirements for different media components. These QoS constraints pose a variety of challenges, especially when the communication occurs over a shared network, with unknown and time-varying cross-traffic. In this work, we propose a transport layer congestion control protocol for telehaptic applications operating over shared networks, termed as Dynamic Packetization Module (DPM). DPM is a lossless, network-aware protocol that tunes the telehaptic packetization rate based on the level of congestion in the network. To monitor the network congestion, we devise a novel network feedback module , which communicates the end-to-end delays encountered by the telehaptic packets to the respective transmitters with negligible overhead. Via extensive simulations, we show that DPM meets the QoS requirements of telehaptic applications over a wide range of network cross-traffic conditions. We also report qualitative results of a real-time telepottery experiment with several human subjects, which reveal that DPM preserves the quality of telehaptic activity even under heavily congested network scenarios. Finally, we compare the performance of DPM with several previously proposed telehaptic communication protocols and demonstrate that DPM outperforms these protocols. Vineet Gokhale, Jayakrishnan Nair 0001, Subhasis Chaudhuri |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2016 | Image Co-segmentation Using Maximum Common Subgraph Matching and Region Co-growing
Avik Hati, Subhasis Chaudhuri, Rajbabu Velmurugan |
ECCV (6) | 2 |
| 2016 | Region-Based Retrieval of Remote Sensing Images Using an Unsupervised Graph-Theoretic ApproachabstractThis letter introduces a novel unsupervised graph-theoretic approach in the framework of region-based retrieval of remote sensing (RS) images. The proposed approach is characterized by two main steps: (1) modeling each image by a graph, which provides region-based image representation combining both local information and related spatial organization, and (2) retrieving the images in the archive that are most similar to the query image by evaluating graph-based similarities. In the first step, each image is initially segmented into distinct regions and then modeled by an attributed relational graph, where nodes and edges represent region characteristics and their spatial relationships, respectively. In the second step, a novel inexact graph matching strategy, which jointly exploits a subgraph isomorphism algorithm and a spectral graph embedding technique, is applied to match corresponding graphs and to retrieve images in the order of graph similarity. Experiments carried out on an archive of aerial images point out that the proposed approach significantly improves the retrieval performance compared to the state-of-the-art unsupervised RS image retrieval methods.(RS) images. Bindita Chaudhuri, Begüm Demir, Lorenzo Bruzzone, Subhasis Chaudhuri |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Estimation of resolvability of user response in kinesthetic perception of jump discontinuitiesabstractPerceptually adaptive sampling based on Weber's law has been used to reduce the haptic packet rate for efficient haptic data transmission over Internet. In a teleoperation, perceptually significant force samples are transmitted from a robot to a human operator. If the time spacing between two consecutive perceptually sampled kinesthetic force stimuli is less than the minimum time spacing (temporal resolution Tr) required in perceiving the jump discontinuity, then the second force stimulus will not be perceived even if it is well above the just noticeable difference. Hence, there is no need to transmit the second force sample to the operator. Thus, for the transmission in a teleoperation, the temporal resolution Trneeds also to be considered while effecting perceptually adaptive sampling. To the best of our knowledge, it has not been studied in the literature. In this work, we propose a statistical method to estimate the temporal resolution Tr. In order to achieve this, we design an appropriate experimental set up, and record the haptic responses for several users extensively. We also study the effect of perceptual fatigue on the temporal resolution Tr, and validate all results using the classical psychometric approach. Amit Bhardwaj, Subhasis Chaudhuri |
World Haptics | 2 |
| 2015 | Salient object carvingabstractIn this paper, we propose an unsupervised two-stage algorithm to extract salient objects from images. In the first stage, the image is segmented into superpixels that are grouped together through k-means clustering, based on histogram features of superpixels. The saliency of each cluster is calculated using inter-cluster and intra-cluster feature dissimilarities. In the second stage, we use seam carving to obtain an object level segmentation of the image in the form of a bounding box around the salient object. We propose an automated approach for seam carving based on a novel energy function obtained by combining the saliency output with a texture removed input image. We compute the optimal number of seams to be removed to extract the salient object instead of manually providing it. The performance of the proposed method is demonstrated by processing different types of images. Avik Hati, Subhasis Chaudhuri, Rajbabu Velmurugan |
ICIP | 2 |
| 2015 | Low bit-rate compression of video and light-field data using coded snapshots and learned dictionariesabstractThe method of coded snapshots has been proposed recently for compressive acquisition of video data to overcome the space-time trade-of inherent in video acquisition. The method involves modulation of the light entering the video camera at different time instants during the exposure period by means of a different and randomly generated code pattern at each of those time instants, followed by integration across time, leading to a single coded snapshot image. Given this image and knowledge of the random codes, it is possible to reconstruct the underlying video frames - by means of sparse coding on a suitably learned dictionary. In this paper, we apply a modified version of this idea, proposed formerly in the compressive sensing literature, to the task of compression of videos and light-field data. At low bit rates, we demonstrate markedly better reconstruction fidelity for the same storage costs, in comparison to JPEG2000 and MPEG-4 (H.264) on light-field and video data respectively. Our technique can cope with overlapping blocks of image data, thereby leading to suppression of block artifacts. Chandrajit Choudhury, Tarun Yellamraju, Ajit Rajwade 0001, Subhasis Chaudhuri |
MMSP | 4 |
| 2015 | A New Self-Training-Based Unsupervised Satellite Image Classification Technique Using Cluster Ensemble StrategyabstractThis letter addresses the problem of unsupervised land-cover classification of remotely sensed multispectral satellite images from the perspective of cluster ensembles and self-learning. The cluster ensembles combine multiple data partitions generated by different clustering algorithms into a single robust solution. A cluster-ensemble-based method is proposed here for the initialization of the unsupervised iterative expectation-maximization (EM) algorithm which eventually produces a better approximation of the cluster parameters considering a certain statistical model is followed to fit the data. The method assumes that the number of land-cover classes is known. A novel method for generating a consistent labeling scheme for each clustering of the consensus is introduced for cluster ensembles. A maximum likelihood classifier is henceforth trained on the updated parameter set obtained from the EM step and is further used to classify the rest of the image pixels. The self-learning classifier, although trained without any external supervision, reduces the effect of data overlapping from different clusters which otherwise a single clustering algorithm fails to identify. The clustering performance of the proposed method on a medium resolution and a very high spatial resolution image have effectively outperformed the results of the individual clustering of the ensemble. Biplab Banerjee, Francesca Bovolo, Avik Bhattacharya, Lorenzo Bruzzone, Subhasis Chaudhuri, B. Krishna Mohan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Convergence analysis of a quadratic upper bounded TV regularizer based blind deconvolution
M. R. Renu, Subhasis Chaudhuri, Rajbabu Velmurugan |
Signal Process. | 2 |
| 2015 | A Novel Graph-Matching-Based Approach for Domain Adaptation in Classification of Remote Sensing Image PairabstractThis paper addresses the problem of land-cover classification of remotely sensed image pairs in the context of domain adaptation. The primary assumption of the proposed method is that the training data are available only for one of the images (source domain), whereas for the other image (target domain), no labeled data are available. No assumption is made here on the number and the statistical properties of the land-cover classes that, in turn, may vary from one domain to the other. The only constraint is that at least one land-cover class is shared by the two domains. Under these assumptions, a novel graph theoretic cross-domain cluster mapping algorithm is proposed to detect efficiently the set of land-cover classes which are common to both domains as well as the additional or missing classes in the target domain image. An interdomain graph is introduced, which contains all of the class information of both images, and subsequently, an efficient subgraph-matching algorithm is proposed to highlight the changes between them. The proposed cluster mapping algorithm initially clusters the target domain data into an optimal number of groups given the available source domain training samples. To this end, a method based on information theory and a kernel-based clustering algorithm is proposed. Considering the fact that the spectral signature of land-cover classes may overlap significantly, a postprocessing step is applied to refine the classification map produced by the clustering algorithm. Two multispectral data sets with medium and very high geometrical resolution and one hyperspectral data set are considered to evaluate the robustness of the proposed technique. Two of the data sets consist of multitemporal image pairs, while the remaining one contains images of spatially disjoint geographical areas. The experiments confirm the effectiveness of the proposed framework in different complex scenarios. Biplab Banerjee, Francesca Bovolo, Avik Bhattacharya, Lorenzo Bruzzone, Subhasis Chaudhuri, Krishna Mohan Buddhiraju |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | A Technique for Simultaneous Visualization and Segmentation of Hyperspectral DataabstractIn this paper, we propose an optimization-based method for simultaneous fusion and unsupervised segmentation of hyperspectral remote sensing images by exploiting redundancy in the data. The hyperspectral data set is visualized as a single image obtained by weighted addition of all spectral points at each pixel location in the data set. The weights are optimized to improve those statistical characteristics of the fused image, which invoke an enhanced response from a human observer. A piecewise-constant smoothness constraint is imposed on the weights instead of the fused image by minimization of its 3-D total-variation norm, thus preventing the fused image from blurring. The optimal recovery of the weight matrix additionally provides useful information in segmenting the hyperspectral data set spatially. We provide ample experimental results to substantiate the usefulness of the proposed method. Abhimitra Meka, Subhasis Chaudhuri |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Design and Analysis of Predictive Sampling of Haptic SignalsabstractIn this article, we identify adaptive sampling strategies for haptic signals. Our approach relies on experiments wherein we record the response of several users to haptic stimuli. We then learn different classifiers to predict the user response based on a variety of causal signal features. The classifiers that have good prediction accuracy serve as candidates to be used in adaptive sampling. We compare the resultant adaptive samplers based on their rate-distortion tradeoff using synthetic as well as natural data. For our experiments, we use a haptic device with a maximum force level of 3 N and 10 users. Each user is subjected to several piecewise constant haptic signals and is required to click a button whenever he perceives a change in the signal. For classification, we not only use classifiers based on level crossings and Weber’s law but also random forests using a variety of causal signal features. The random forest typically yields the best prediction accuracy and a study of the importance of variables suggests that the level crossings and Weber’s classifier features are most dominant. The classifiers based on level crossings and Weber’s law have good accuracy (more than 90%) and are only marginally inferior to random forests. The level crossings classifier consistently outperforms the one based on Weber’s law even though the gap is small. Given their simple parametric form, the level crossings and Weber’s law--based classifiers are good candidates to be used for adaptive sampling. We study their rate-distortion performance and find that the level crossing sampler is superior. For example, for haptic signals obtained while exploring various rendered objects, for an average sampling rate of 10 samples per second, the level crossings adaptive sampler has a mean square error about 3dB less than the Weber sampler. Amit Bhardwaj, Subhasis Chaudhuri, Onkar Dabeer |
ACM Trans. Appl. Percept. | 2 |
| 2014 | Novel Speed-Up Strategies for Non-Local Means Denoising With Patch and Edge Patch Based DictionariesabstractIn this paper, a novel technique to speed-up a nonlocal means (NLM) filter is proposed. In the original NLM filter, most of its computational time is spent on finding distances for all the patches in the search window. Here, we build a dictionary in which patches with similar photometric structures are clustered together. Dictionary is built only once with high resolution images belonging to different scenes. Since the dictionary is well organized in terms of indexing its entries, it is used to search similar patches very quickly for efficient NLM denoising. We achieve a substantial reduction in computational cost compared with the original NLM method, especially when the search window of NLM is large, without much affecting the PSNR. Second, we show that by building a dictionary for edge patches as opposed to intensity patches, it is possible to reduce the dictionary size; thus, further improving the computational speed and memory requirement. The proposed method preclassifies similar patches with the same distance measure as used by NLM method. The proposed algorithm is shown to outperform other prefiltering based fast NLM algorithms computationally as well as qualitatively. Hemalata Bhujle, Subhasis Chaudhuri |
IEEE Trans. Image Process. | 2 |
| 2013 | Scalable rendering of variable density point cloud dataabstractIn this paper, we present a novel proxy based method of adaptive haptic rendering of a variable density 3D point cloud data at different levels of detail without pre-computing the mesh structure. We also incorporate features like rotation, translation and friction to provide a better realistic experience to the user. Instead of a point proxy, a spherical proxy of variable radius is used which avoids the sinking of proxy during the haptic interaction of sparse data. The radius of the proxy is adaptively varied depending upon the local density of the point data using kernel bandwidth estimation. During the interaction, the proxy moves in small steps tangentially over the point cloud such that the new position always minimizes the distance between the proxy and the haptic interaction point (HIP). The raw point cloud data re-sampled in a regular 3D lattice of voxels are loaded to the haptic space after proper smoothing to avoid aliasing effects. The rendering technique is experimented with several subjects and it is observed that this functionality supplements the user's experience by allowing the user to interact with an object at multiple resolutions. Priyadarshini Kumari, K. G. Sreeni, Subhasis Chaudhuri |
World Haptics | 3 |
| 2012 | A Novel Domain Adaptation Bayesian Classifier for Updating Land-Cover Maps With Class Differences in Source and Target DomainsabstractThis paper addresses the problem of land-cover map updating by classification of multitemporal remote-sensing images in the context of domain adaptation (DA). The basic assumptions behind the proposed approach are twofold. The first one is that training data (ground reference information) are available for one of the considered multitemporal acquisitions (source domain) whereas they are not for the other (target domain). The second one is that multitemporal acquisitions (i.e., target and source domains) may be characterized by different sets of classes. Unlike other approaches available in the literature, the proposed DA Bayesian classifier based on maximum a posteriori decision rule (DA-MAP) automatically identifies whether there exist differences between the set of classes in the target and source domains and properly handles these differences in the updating process. The proposed method was tested in different scenarios of increasing complexity related to multitemporal image classification. Experimental results on medium-resolution and very high resolution multitemporal remote-sensing data sets confirm the effectiveness and the reliability of the proposed DA-MAP classifier. Kanchan Bahirat, Francesca Bovolo, Lorenzo Bruzzone, Subhasis Chaudhuri |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2011 | Object Mining for Large Video dataabstractWe propose a method for achieving a novel concise graph-based representation for retrieval of objects from large video data. The emphasis in this paper is towards achieving a compact representation of video data for faster retrieval. Specifically, we use information available from scripts and subtitles in order to group all occurrences of an object in video data, which provides a separate representation for each scene. Further, based on the premise that the number of objects in a shot are typically much less than the number of video frames in that shot, we propose a graph-based representation in which vertices represent objects rather than video frames. Key advantages of the proposed approach include faster retrieval, efficiency in performing tasks such as spatial re-ranking and graph partitioning and a single representation for both retrieval and summarization applications. We demonstrate efficacy of the proposed approach in retrieval and summarization applications over video data consisting of episodes of a popular TV series Friends. Ronak Shah, Rishabh Iyer 0001, Subhasis Chaudhuri |
BMVC | 3 |
| 2011 | An optimization-based approach to fusion of multi-exposure, low dynamic range images
Ketan Kotwal, Subhasis Chaudhuri |
FUSION | 2 |
| 2011 | Modeling of PSF for refractive index variation in fluorescence microscopyabstractIn fluorescence microscopy, thickness of the specimen is significant compared to the depth of focus of a high numerical aperture (Na) lens of the imaging system. An aberration is introduced when there is a variation in refractive index in the specimen. Aberrations decrease the resolution and contrast of the image. Deconvolution is a method used to restore the image computationally, and for a successful deconvolution an accurate knowledge of the imaging system's point spread function (PSF) is a must. Here, we derive an appropriate model to obtain shift varying PSF for a thick specimen. Deconvolution results obtained for such a specimen are significantly better compared to those obtained using the standard Gibson and Lanni PSF model. Sameer Hiware, Pradyot Porwal, Rajbabu Velmurugan, Subhasis Chaudhuri |
ICIP | 4 |
| 2011 | High dynamic range imaging under noisy observationsabstractWe propose a radiance domain denoising frame work for the high dynamic range (HDR) imaging problem. The proposed method uses a maximum aposteriori probability (MAP) based reconstruction of the HDR image with total variation (TV) as the prior to avoid unnecessary smoothing of the radiance field. To make the computation with TV prior efficient, we extend the majorize-minimize method of upper bounding the total variation by a quadratic function to our case which has a nonlinear term arising from the camera response function. A theoretical justification for doing radiance domain denoising as opposed to image domain denoising is also provided. Our method yields better results, with the edges well preserved and noise reduced considerably. Renu M. Rameshan, Subhasis Chaudhuri, Rajbabu Velmurugan |
ICIP | 2 |
| 2011 | Adaptive smoothness based robust active contours
Viswanathan Srikrishnan, Subhasis Chaudhuri |
Image Vis. Comput. | 2 |
| 2011 | Reconstruction of high contrast images for dynamic scenes
Shanmuganathan Raman, Subhasis Chaudhuri |
Vis. Comput. | 2 |
| 2010 | Crowd Motion Analysis Using Linear Cyclic PursuitabstractCrowd motion analysis, where there is interdependence amongst the constituent elements, is a relatively unexplored application area in computer vision. In this work, we propose a fast method for short-term crowd motion prediction using a sparse set of particles. We study the dynamics of a crowd motion model and linear cyclic pursuit. We show that linear cyclic pursuit naturally captures the repulsive and attractive forces acting on the individual crowd member. The pursuit parameters are estimated from videos in an online manner using a feature tracker. Short term trajectory prediction is done by numerical solution of estimated cyclic pursuit equation. We demonstrate the suitability of the proposed technique through extensive experimentations. Viswanathan Srikrishnan, Subhasis Chaudhuri |
ICPR | 2 |
| 2010 | Visualization of Hyperspectral Images Using Bilateral FilteringabstractThis paper presents a new approach for hyperspectral image visualization. A bilateral filtering-based approach is presented for hyperspectral image fusion to generate an appropriate resultant image. The proposed approach retains even the minor details that exist in individual image bands, by exploiting the edge-preserving characteristics of a bilateral filter. It does not introduce visible artifacts in the fused image. A hierarchical fusion scheme has also been proposed for implementation purposes to accommodate a large number of hyperspectral image bands. The proposed scheme provides computational and storage efficiency without affecting the quality and performance of the fusion. It also facilitates the midband visualization of a subset of the hyperspectral image cube. Quantitative performance results are presented to indicate the effectiveness of the proposed method. Ketan Kotwal, Subhasis Chaudhuri |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2009 | Poisson compositingabstractMost of the real world scenes have a very high dynamic range. However the common capture and display devices can handle only a limited dynamic range. General approach to solve this problem is to use multi-exposure images and composite them in the irradiance domain to get a High Dynamic Range (HDR) image [Reinhard et al. 2005]. The generated image will be able to represent the real world scene faithfully. However, it needs to be tone-mapped to a Low Dynamic Range (LDR) image for visualization in common displays and printers. Generation of the high-quality LDR image of the scene directly from multi-exposure images even in the absence of any knowledge of camera response function and the exposure settings of the camera is of interest to graphics community. We propose a gradient domain compositing technique to solve the above problem and call it Poisson Compositing. We compare the proposed methodology with similar existing techniques and show that the proposed method is very fast and accurate. Shanmuganathan Raman, Subhasis Chaudhuri |
SIGGRAPH ASIA Sketches | 2 |
| 2009 | Application Of Papoulis-Gerchberg Method In Image Super-Resolution and InpaintingabstractIn this paper, we study the Papoulis–Gerchberg (PG) method and its applications to domains of image restoration such as super-resolution (SR) and inpainting. We show that the method performs well under certain conditions. We then suggest improvements to the method to achieve better SR and inpainting results. The modification applied to the SR process also allows us to apply the method to a larger class of images by doing away with some of the restrictions inherent in the classical PG method. We also present results to demonstrate the performance of the proposed techniques. Priyam Chatterjee, Sujata Mukherjee, Subhasis Chaudhuri, Guna Seetharaman |
Comput. J. | 3 |
| 2009 | Stabilization of Parametric Active Contours Using a Tangential Redistribution TermabstractDepending on implementation, active contours have been classified as geometric or parametric active contours. Parametric contours, irrespective of representation, are known to suffer from the problem of irregular bunching and spacing out of curve points during the curve evolution. In a spline-based implementation of active contours, this leads to occasional formation of loops locally, and subsequently the curve blows up due to instabilities. In this paper, we analyze the reason for this problem and propose a solution to alleviate the same. We propose an ordinary differential equation (ODE) for controlling the curve parametrization during evolution by including a tangential force. We show that the solution of the proposed ODE is bounded. We demonstrate the effectiveness of the proposed method for segmentation and tracking tasks on closed as well as open contours. Viswanathan Srikrishnan, Subhasis Chaudhuri |
IEEE Trans. Image Process. | 2 |
| 2008 | Recovery of relative depth from a single observation using an uncalibrated (real-aperture) cameraabstractIn this paper we investigate the challenging problem of recovering the depth layers in a scene from a single defocused observation. The problem is definitely solvable if there are multiple observations. In this paper we show that one can perceive the depth in the scene even from a single observation. We use the inhomogeneous reverse heat equation to obtain an estimate of the blur, thereby preserving the depth information characterized by the defocus. However, the reverse heat equation, due to its parabolic nature, is divergent. We stabilize the reverse heat equation by considering the gradient degeneration as an effective stopping criterion. The amount of (inverse) diffusion is actually a measure of relative depth. Because of ill-posedness we propose a graph-cuts based method for inferring the depth in the scene using the amount of diffusion as a data likelihood and a smoothness condition on the depth in the scene. The method is verified experimentally on a varied set of test cases. Vinay P. Namboodiri, Subhasis Chaudhuri |
CVPR | 2 |
| 2008 | A Theoretical Analysis of Data Reduction Using the Weber QuantizerabstractWe present a theoretical analysis of a perceptual coding approach, the so called Weber quantizer. Extensive studies performed by experimental psychologists and physiologists have unveiled one major conclusion: human perception often follows Weber's law. Ernst Weber was an experimental physiologist who in 1834 first discovered the following implication DeltaI = kl, where DeltaI is the so called difference threshold or the just noticeable difference (JND). It describes the smallest amount of change of an (arbitrary) stimulus I which can be detected just as often as it cannot be detected and defines the Weber bound at [(1 - k)I, (1 + k)I]. Julius Kammerl, Peter Hinterseer, Subhasis Chaudhuri, Eckehard G. Steinbach |
DCC | 3 |
| 2008 | Regularized depth from defocusabstractIn the area of depth estimation from images an interesting approach has been structure recovery from defocus cue. Towards this end, there have been a number of approaches [4,6]. Here we propose a technique to estimate the regularized depth from defocus using diffusion. The coefficient of the diffusion equation is modeled using a pair-wise Markov random field (MRF) ensuring spatial regularization to enhance the robustness of the depth estimated. This framework is solved efficiently using a graph-cuts based techniques. The MRF representation is enhanced by incorporating a smoothness prior that is obtained from a graph based segmentation of the input images. The method is demonstrated on a number of data sets and its performance is compared with state of the art techniques. Vinay P. Namboodiri, Subhasis Chaudhuri, Sunil Hadap |
ICIP | 2 |
| 2008 | New Closed-Form Bounds on the Partition Function
Krishnamurthy Dvijotham, Soumen Chakrabarti, Subhasis Chaudhuri |
ECML/PKDD (1) | 3 |
| 2008 | A Context-Sensitive Clustering Technique Based on Graph-Cut Initialization and Expectation-Maximization AlgorithmabstractThis letter presents a multistage clustering technique for unsupervised classification that is based on the following: 1) a graph-cut procedure to produce initial segments that are made up of pixels with similar spatial and spectral properties; 2) a fuzzy c-means algorithm to group these segments into a fixed number of classes; 3) a proper implementation of the expectation-maximization (EM) algorithm to estimate the statistical parameters of classes on the basis of the initial seeds that are achieved at convergence by the fuzzy c-means algorithm; and 4) the Bayes rule for minimum error to perform the final classification on the basis of the distributions that are estimated with the EM algorithm. Experimental results confirm the effectiveness of the proposed technique. Mayank Tyagi, Francesca Bovolo, Ankit K. Mehra, Subhasis Chaudhuri, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2008 | New closed-form bounds on the partition function
Krishnamurthy Dvijotham, Soumen Chakrabarti, Subhasis Chaudhuri |
Mach. Learn. | 3 |
| 2008 | Multidimensional Probability Density Function Matching for Preprocessing of Multitemporal Remote Sensing ImagesabstractThis paper addresses the problem of matching the statistical properties of the distributions of two (or more) multi-spectral remote sensing images acquired on the same geographical area at different times. An N-D probability density function (pdf) matching technique for the preprocessing of multitemporal images is introduced in the remote sensing domain by defining and analyzing three important application scenarios: 1) supervised classification; 2) partially supervised classification; and 3) change detection. Unlike other methods adopted in remote sensing applications, the procedure considered performs the matching process by properly taking into account the correlation among spectral channels, thus retaining the data correlation structure after the pdf matching. Experimental results obtained on real multitemporal remote sensing data sets confirm the validity of the presented technique in all the considered scenarios. Shilpa Inamdar, Francesca Bovolo, Lorenzo Bruzzone, Subhasis Chaudhuri |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2007 | Automated Billboard Insertion in Video
Hitesh Shah, Subhasis Chaudhuri |
ACCV (1) | 2 |
| 2007 | Shape Recovery Using Stochastic Heat FlowabstractWe consider the problem of depth estimation from multiple images based on the defocus cue. For a Gaussian defocus blur, the observations can be shown to be the solution of a deterministic but inhomogeneous diffusion process. However, the diffusion process does not sufficiently address the case in which the Gaussian kernel is deformed. This deformation happens due to several factors like self-occlusion, possible aberrations and imperfections in the aperture. These issues can be solved by incorporating a stochastic perturbation into the heat diffusion process. The resultant flow is that of an inhomogeneous heat diffusion perturbed by a stochastic curvature driven motion. The depth in the scene is estimated from the coefficient of the stochastic heat equation without actually knowing the departure from the Gaussian assumption. Further, the proposed method also takes into account the non-convex nature of the diffusion process. The method provides a strong theoretical framework for handling the depth from defocus problem. 1 Vinay P. Namboodiri, Subhasis Chaudhuri |
BMVC | 2 |
| 2007 | On Stabilisation of Parametric Active ContoursabstractParametric active contours have been used extensively in computer vision for different tasks like segmentation and tracking. However, all parametric contours are known to suffer from the problem of frequent bunching and spacing out of curve points locally during the curve evolution. In a spline based implementation of active contours, this leads to occasional formation of loops locally, and subsequently the curve blows up due to instabilities. It has been shown earlier that in addition to usual evolution along the normal direction, the curve should also be evolved in the tangential direction for stability purposes. In this paper, we provide a mathematical basis for selecting such a suitable tangential component for stabilisation. We prove the boundedness of the evolved curve in this paper, and provide the physical significance. We demonstrate the usefulness of the proposed method with a number of experiments. Viswanathan Srikrishnan, Subhasis Chaudhuri, Sumantra Dutta Roy, Daniel Sevcovic |
CVPR | 2 |
| 2007 | A Matte-less, Variational Approach to Automatic Scene CompositingabstractIn this paper, we consider the problem of compositing a scene from multiple images. Multiple images, for example, can be obtained by varying the exposure of the camera, by changing the object at focus, or by simply sampling a video sequence at arbitrary time instants. We develop this problem in an optimization framework and then adopt a variational approach to derive a generalized algorithm which will be able to solve diverse applications depending on the nature of the input images. Our approach has distinct advantages over the existing digital compositing techniques, such as alpha matting and alpha blending, which require an explicit preparation of the matte while there is no such requirement in the proposed technique. We demonstrate the usefulness of our approach through results from diverse applications in computer vision. Shanmuganathan Raman, Subhasis Chaudhuri |
ICCV | 2 |
| 2007 | Image Restoration using Geometrically Stabilized Reverse Heat EquationabstractBlind restoration of blurred images is a classical ill-posed problem. There has been considerable interest in the use of partial differential equations to solve this problem. The blurring of an image has traditionally been modeled by Witkin [10] and Koenderink [4] by the heat equation. This has been the basis of the Gaussian scale space. However, a similar theoretical formulation has not been possible for deblurring of images due to the ill-posed nature of the reverse heat equation. Here we consider the stabilization of the reverse heat equation. We do this by damping the distortion along the edges by adding a normal component of the heat equation in the forward direction. We use a stopping criterion based on the divergence of the curvature in the resulting reverse heat flow. The resulting stabilized reverse heat flow makes it possible to solve the challenging problem of blind space varying deconvolution. The method is justified by a varied set of experimental results. Vinay P. Namboodiri, Subhasis Chaudhuri |
ICIP (4) | 2 |
| 2007 | Segmentation and region of interest based image retrieval in low depth of field observations
Rajashekhar, Subhasis Chaudhuri |
Image Vis. Comput. | 2 |
| 2007 | Retrieval of images of man-made structures based on projective invariance
Rajashekhar, Subhasis Chaudhuri, Vinay P. Namboodiri |
Pattern Recognit. | 2 |
| 2007 | On defocus, diffusion and depth estimation
Vinay P. Namboodiri, Subhasis Chaudhuri |
Pattern Recognit. Lett. | 2 |
| 2006 | Novel View Synthesis Using Locally Adaptive Depth Regularization
Hitesh Shah, Subhasis Chaudhuri |
ACCV (1) | 2 |
| 2006 | Alias-Free Interpolation
C. Victor Jiji, Prakash Neethu, Subhasis Chaudhuri |
ECCV (4) | 3 |
| 2006 | Perception-Based Compression of Haptic Data Streams Using Kalman FiltersabstractIn order to realize truly immersive and stable telepresence and teleaction over the Internet it is necessary to keep delay, data rate, and packet rate of haptic data streams as low as possible. In addition, the compression scheme for haptic data, which is necessary to achieve those goals, must be fast enough to work on a sample by sample basis to not add further delay. This paper presents an approach that reduces haptic data traffic in networked telepresence and teleaction systems to a small fraction of the original rate without impairing performance by using fast Kalman filters on the input signals combined with model based prediction of haptic signals. Our approach reduces the number of transmitted packets to 9.8% (velocity) and 6.2% (force) of the original rate without impairing immersiveness Peter Hinterseer, Eckehard G. Steinbach, Subhasis Chaudhuri |
ICASSP (5) | 3 |
| 2006 | Simultaneous estimation of super-resolved depth map and intensity field using photometric cue
Manjunath V. Joshi, Subhasis Chaudhuri |
Comput. Vis. Image Underst. | 2 |
| 2006 | Automatic illumination correction for scene enhancement and object tracking
Arvind Nayak, Subhasis Chaudhuri |
Image Vis. Comput. | 2 |
| 2006 | Region-based CBIR in GIS with local space filling curves to spatial representation
Adel Hafiane, Subhasis Chaudhuri, Guna Seetharaman, Bertrand Y. Zavidovique |
Pattern Recognit. Lett. | 2 |
| 2006 | A Model-Based Approach to Multiresolution Fusion in Remotely Sensed ImagesabstractIn this paper, a model-based approach to multiresolution fusion of remotely sensed images is presented. Given a high spatial resolution panchromatic (Pan) image and a lowspatial resolution multispectral (MS) image acquired on the same geographical area, the presented method aims to enhance the spatial resolution of the MS image to the resolution of the Pan observation. The proposed fusion technique utilizes the spatial correlation of each of the high-resolution MS channels by using an autoregressive (AR) model, whose parameters are learnt from the analysis of the Pan data. Under the assumption that the parameters of the AR model for the Pan image are the same as those that represent the MS images due to spectral correlation, the proposed technique exploits the learnt parameter values in the context of a proper regularization technique to estimate the high spatial resolution fields for the MS bands. This results in a combination of the spectral characteristics of the low-resolution MS data with the high spatial resolution of the Pan image. The main advantages of the proposed technique are: 1) unlike standard methods proposed in the literature, it requires no registration between the Pan and the MS images; 2) it models effectively the texture of the scene during the fusion process; 3) it shows very small spectral distortion (as it is less affected, compared to standard methods, by the specific digital numbers of pixels in the Pan image, since it exploits the learnt parameters from the Pan image rather than the actual Pan digital numbers for fusion); and 4) it can be used in critical situations in which the Pan and the MS images are acquired (also by different sensors) in slightly different areas. Quantitative experimental results obtained using Landsat-7 Enhanced Thematic Mapper Plus (ETM+) and Quickbird images point out the effectiveness of the proposed method Manjunath V. Joshi, Lorenzo Bruzzone, Subhasis Chaudhuri |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2006 | Qualitative Visual Environment RetrievalabstractA system for retrieval of an unstructured environment under static and dynamic scenarios is proposed. The use of cylindrical mosaics or omnidirectional images is exploited for providing a rich description about the surrounding environment spanning 360 degrees. The environment description is based on defining the attributes of the nodes of a graph derived from the angular partitions of the captured images. Content-based image retrieval for each of these partitions is performed on an exemplar image database to annotate the nodes of the graph. The complete environment description is recovered by collating the retrieval results over all the partitions based on a simple voting scheme. This offers a qualitative description of the location in a totally natural and unstructured surrounding. The experiments yield quite promising results. Rajashekhara, Amit B. Prabhudesai, Subhasis Chaudhuri |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2005 | Shock Filters Based on Implicit Cluster SeparationabstractOne of the classic problems in low level vision is image restoration. An important contribution toward this effort has been the development of shock filters by Osher and Rudin (1990). It performs image deblurring using hyperbolic partial differential equations. In this paper we relate the notion of cluster separation from the field of pattern recognition to the shock filter formulation. A kind of shock filter is proposed based on the idea of gradient based separation of clusters. The proposed formulation is general enough as it can allow various models of density functions in the cluster separation process. The efficacy of the method is demonstrated through various examples. Vinay P. Namboodiri, Subhasis Chaudhuri |
CVPR (1) | 2 |
| 2005 | A modified FCM with optimal Peano scans for image segmentationabstractThis paper describes a new method for fuzzy segmentation based on spatial constraints. Taking into account the neighborhood influence two techniques are used. First a new feature is derived from Peano scans to represent a spatial relationship among neighbors. Second we incorporate a regularization term to fuzzy C-means algorithm (FCM). The algorithm is tested on both synthetic and multispectral images. Experimental results are presented and discussed. They show the effectiveness of the method. Adel Hafiane, Bertrand Y. Zavidovique, Subhasis Chaudhuri |
ICIP (3) | 3 |
| 2005 | A context-sensitive Bayesian technique for the partially supervised classification of multitemporal imagesabstractAn advanced context-sensitive classification technique that exploits a temporal series of remote sensing images for a regular updating of land-cover maps is proposed. This technique extends the use of spatio-contextual information to the framework of partially supervised approaches (that are capable of addressing the updating problem under the realistic, though critical, constraint that no ground-truth information is available for some of the images to be classified). The proposed classifier is based on an iterative partially supervised algorithm that jointly estimates the class-conditional densities and the prior model for the class labels on the image to be classified by taking into account spatio-contextual information. Experimental results point out that the proposed technique is effective and that it significantly outperforms the context-insensitive partially supervised approaches presented in the literature. Roberto Cossu, Subhasis Chaudhuri, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2005 | Multispectral panoramic mosaicing
Udhav Bhosle, Sumantra Dutta Roy, Subhasis Chaudhuri |
Pattern Recognit. Lett. | 3 |
| 2005 | A learning-based method for image super-resolution from zoomed observationsabstractWe propose a technique for super-resolution imaging of a scene from observations at different camera zooms. Given a sequence of images with different zoom factors of a static scene, we obtain a picture of the entire scene at a resolution corresponding to the most zoomed observation. The high-resolution image is modeled through appropriate parameterization, and the parameters are learned from the most zoomed observation. Assuming a homogeneity of the high-resolution field, the learned model is used as a prior while super-resolving the scene. We suggest the use of either a Markov random field (MRF) or an simultaneous autoregressive (SAR) model to parameterize the field based on the computation one can afford. We substantiate the suitability of the proposed method through a large number of experimentations on both simulated and real data. Manjunath V. Joshi, Subhasis Chaudhuri, Rajkiran Panuganti |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Robust shape based two hand trackerabstractThis paper presents a robust shape-based on-line tracker for simultaneously tracking the motion of both hands, that is robust to cases of background clutter, other moving objects, occlusions of one hand by the other and a wide range of illumination variations. The tracker is based on an online predictive eigentracking framework. This framework allows efficient tracking of articulate objects, which change in appearance across views. We show results of successful tracking across all possible cases of motion dynamics of both hands during occlusion and a wide range of illumination conditions. Ketan Barhate, Kaustubh Patwardhan, Sumantra Dutta Roy, Subhasis Chaudhuri, Santanu Chaudhury |
ICIP | 4 |
| 2004 | Zoom based super-resolution through sar model fitting
Manjunath V. Joshi, Subhasis Chaudhuri |
ICIP | 2 |
| 2004 | Image retrieval based on projective invarianceabstractWe propose an image retrieval scheme based on projectively invariant features. Since cross-ratio is the fundamental invariant feature under projective transformations for points, we use that as the basic feature parameter. We compute the cross-ratios of point sets in quadruplets and a discrete representation of the distribution of the cross-ratio is obtained from the computed values. The distribution is used as the feature for retrieval purposes. The method is very effective in retrieving images, like buildings, having similar planar 3D structures. S. Rajashekhar, Subhasis Chaudhuri, Vinay P. Namboodiri |
ICIP | 2 |
| 2004 | Guest editorial
Subhasis Chaudhuri, Andrew Zisserman |
Image Vis. Comput. | 1 |
| 2004 | Content based image retrieval using motif cooccurrence matrix
N. Jhanwar, Subhasis Chaudhuri, Guna Seetharaman, Bertrand Y. Zavidovique |
Image Vis. Comput. | 2 |
| 2004 | Super-resolution imaging: use of zoom as a cue
Manjunath V. Joshi, Subhasis Chaudhuri, Rajkiran Panuganti |
Image Vis. Comput. | 2 |
| 2004 | Depth Estimation and Image Restoration Using Defocused Stereo PairsabstractWe propose a method for estimating depth from images captured with a real aperture camera by fusing defocus and stereo cues. The idea is to use stereo-based constraints in conjunction with defocusing to obtain improved estimates of depth over those of stereo or defocus alone. The depth map as well as the original image of the scene are modeled as Markov random fields with a smoothness prior, and their estimates are obtained by minimizing a suitable energy function using simulated annealing. The main advantage of the proposed method, despite being computationally less efficient than the standard stereo or DFD method, is simultaneous recovery of depth as well as space-variant restoration of the original focused image of the scene. A. N. Rajagopalan 0001, Subhasis Chaudhuri, Uma Mudenagudi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Gesture recognition using position and appearance featuresabstractIn this paper a scheme for recognizing hand gestures is presented using the output of a condensation tracker. The tracker is used to obtain a set of features. These features consisting of temporal evolution of the spatial moments form high dimensional feature vectors. The principal components of the feature trajectories are used to recognize the gestures. Tushar Agrawal, Subhasis Chaudhuri |
ICIP (2) | 2 |
| 2003 | Self-induced color correction for skin tracking under varying illuminationabstractAn important challenge of any skin color based tracking system is to accommodate varying illumination conditions. We present a method for automatic transformation of the color planes to match the skin color model learnt for a fixed illumination. The first couple of initial frames are used to automatically extract the palm region which in turn serves as an observed skin color palette which needs to be color-transformed to match a similar palette under a known illumination. A neural network implementing back-propagation learning rule then performs the color correction for the entire sequence. We use a condensation algorithm for tracking after the color correction for improved performance. Arvind Nayak, Subhasis Chaudhuri |
ICIP (3) | 2 |
| 2003 | Simultaneous Estimation of Super-Resolved Scene and Depth Map from Low Resolution Defocused ObservationsabstractThis paper presents a novel technique to simultaneously estimate the depth map and the focused image of a scene, both at a super-resolution, from its defocused observations. Super-resolution refers to the generation of high spatial resolution images from a sequence of low resolution images. Hitherto, the super-resolution technique has been restricted mostly to the intensity domain. In this paper, we extend the scope of super-resolution imaging to acquire depth estimates at high spatial resolution simultaneously. Given a sequence of low resolution, blurred, and noisy observations of a static scene, the problem is to generate a dense depth map at a resolution higher than one that can be generated from the observations as well as to estimate the true high resolution focused image. Both the depth and the image are modeled as separate Markov random fields (MRF) and a maximum a posteriori estimation method is used to recover the high resolution fields. Since there is no relative motion between the scene and the camera, as is the case with most of the super-resolution and structure recovery techniques, we do away with the correspondence problem. Deepu Rajan, Subhasis Chaudhuri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Mixed H2/ H∞ algorithm for exponentially windowed adaptive filteringabstractThe RLS or its equivalent H2algorithms achieve the best average performance. But these suffer from poor worst case performance. Stochastic gradient based algorithms (like LMS) achieve best worst case performance, but a poor average performance. In this paper we propose switching criteria for acquiring advantages of both of these algorithms. The proposed algorithm uses a nonlinear combination of H2optimal and H∞optimal estimation. We present a mixed H2/H∞algorithm employing an exponential window and the achievable bound for this estimation strategy. Sandip D. Kothari, Ravi N. Banavar, Subhasis Chaudhuri |
ICASSP | 3 |
| 2001 | Simultaneous Tracking of Both Hands by Estimation of Erroneous ObservationsabstractThe articulate motion of the hand makes it very difficult to track the hands while performing a gesture. Simultaneous tracking of both hands needs to deal with large interframe variations in shape, clutter and mutual occlusion. In this paper, we present a robust method for localizing the hand region by tracking even in the presence of severe occlusion. We develop a model for tracking two rectangular windows, each bounding one of the hands, using the condensation algorithm. We propose a new method for dealing with occlusions by estimating the occluded observations in terms of the non-occluded observations and their predicted values, yielding very reliable results. 1 James P. Mammen, Subhasis Chaudhuri, Tushar Agrawal |
BMVC | 2 |
| 2001 | Generation of super-resolution images from blurred observations using Markov random fieldsabstractThis paper presents a new technique for generating a high resolution image from a blurred image sequence; this is also referred to as super-resolution restoration of images. The image sequence consists of decimated, blurred and noisy versions of the high resolution image. The high resolution image is modeled as a Markov random field (MRF) and a maximum a posteriori (MAP) estimation technique is used. A simple gradient descent method is used to optimize the functional. Further, line fields are introduced in the cost function and optimization using Graduated Non-Convexity (GNC) is shown to yield improved results. Lastly, we present results of optimization using Simulated Annealing (SA). Deepu Rajan, Subhasis Chaudhuri |
ICASSP | 2 |
| 2001 | Depth From Defocus in Presence of Partial Self Occlusion
Sundeep Singh Bhasin, Subhasis Chaudhuri |
ICCV | 2 |
| 2001 | Simultaneous Estimation of Super-Resolved Intensity and Depth Maps from Low Resolution Defocused Observations of a Scene
Deepu Rajan, Subhasis Chaudhuri |
ICCV | 2 |
| 2001 | Generalized interpolation and its application in super-resolution imaging
Deepu Rajan, Subhasis Chaudhuri |
Image Vis. Comput. | 2 |
| 2000 | A Perceptually Organized Method for Image InterpolationabstractPerception of an image depends on its visual representation. In this paper we present a perceptually organized scheme for image expansion or scattered data interpolation. This is done by decomposing the image (or data) into appropriate perceptual groups, carrying out the interpolation in individual groups as per the perceptual necessity and subsequently transforming the interpolated values back to the image domain. Various perceptual properties of the image, such as the 3D shape of an object, textural homogeneity, local variations in scene reflectivity, visual transparency can be better preserved during the interpolation process. Deepu Rajan, Subhasis Chaudhuri |
ICPR | 2 |
| 2000 | Locating Human Faces in a Cluttered Scene
A. N. Rajagopalan 0001, K. Sunil Kumar, Jayashree Karlekar, R. Manivasakan, M. Milind Patil, Uday B. Desai, P. G. Poonacha, Subhasis Chaudhuri |
Graph. Model. | 8 |
| 2000 | Visual understanding of dynamic hand gestures
Mohammed Yeasin, Subhasis Chaudhuri |
Pattern Recognit. | 2 |
| 2000 | Toward automatic robot programming: learning human skill from visual dataabstractWe propose a novel approach to program a robot by demonstrating the task multiple number of times in front of a binocular vision system. We track artificially-induced features appearing in the image plane due to nonimpedimental color stickers attached at different fingertips and wrist joint, in a simultaneous feature detection and tracking framework. A Kalman filter does the tracking by recursively predicting the tentative feature location and a higher order statistics (HOS)-based data clustering algorithm extracts the feature. A fast and efficient algorithm for the vision system thus developed processes a binocular video sequence to obtain the trajectories and the orientation information of the end effector from the images of a human hand. The concept of trajectory bundle is introduced to avoid singularities and to obtain an optimal path. Mohammed Yeasin, Subhasis Chaudhuri |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1999 | Simultaneous Depth Recovery and Image Restoration from Defocused ImagesabstractWe propose a method for simultaneous recovery of depth and restoration of scene intensity, given two defocused images of a scene. The space-variant blur parameter and the focused image of the scene are modeled as Markov random fields (MRFs). Line fields are included to preserve discontinuities. The joint posterior distribution of the blur parameter and the intensity process is examined for locality property and we derive an important result that the posterior is again Markov. The result enables us to obtain the maximum a posterior (MAP) estimates of the blur parameter and the focused image, within reasonable computational limits. The estimates of depth and the quality of the restored image are found to be quite good, even in the presence of discontinuities. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
CVPR | 2 |
| 1999 | Dynamic hand gesture understanding-a new approachabstractThe analysis of a dynamic hand gesture requires processing a spatio-temporal image sequence. The actual length of the sequence varies with each instantiation of the gesture. We propose a novel, vision based system for automatic interpretation of a limited set of dynamic hand gestures. This involves extracting the temporal signature of the hand motion from the performed gesture and is subsequently analyzed by a finite-state machine to automatically interpret the performed gesture. Mohammed Yeasin, Subhasis Chaudhuri |
ICASSP | 2 |
| 1999 | Depth Estimation using Defocused Stereo Image PairsabstractIn this paper we propose a new method for estimating depth using a fusion of defocus and stereo, that relaxes the assumption of a pinhole model of the camera. It avoids the correspondence problem of stereo. The main advantage of this algorithm is simultaneous recovery of depth and image restoration. The depth (blur or disparity) in the scene and the intensity process in the focused image are individually modeled as Markov random fields (MRF). It avoids the windowing of data and allows incorporation of multiple observations in the estimation procedure. The accuracy of depth estimation and the quality of the restored image are improved compared to the depth from defocus method, and a dense depth map is estimated without correspondence and interpolation as in the case of stereo. Uma Mudenagudi, Subhasis Chaudhuri |
ICCV | 2 |
| 1999 | An MRF Model-Based Approach to Simultaneous Recovery of Depth and Restoration from Defocused ImagesabstractIn this paper, we propose a MAP-Markov random field (MRF) based scheme for recovering the depth and the focused image of a scene from two defocused images. The space-variant blur parameter and the focused image of the scene are both modeled as MRFs and their MAP estimates are obtained using simulated annealing. The scheme is amenable to the incorporation of smoothness constraints on the spatial variations of the blur parameter as well as the scene intensity. It also allows for inclusion of line fields to preserve discontinuities. The performance of the proposed scheme is tested on synthetic as well as real data and the estimates of the depth are found to be better than that of the existing window-based depth from defocus technique. The quality of the space-variant restored image of the scene is quite good even under severe space-varying blurring conditions. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | MRF model-based identification of shift-variant point spread function for a class of imaging systems
A. N. Rajagopalan 0001, Subhasis Chaudhuri |
Signal Process. | 2 |
| 1998 | Optimal Recovery of Depth from Defocused Images Using an MRF ModelabstractA MAP-MRF based scheme is proposed for simultaneous recovery of the depth and the focused image of a scene from two defocused images. The space-variant blur parameter and the focused image of the scene are both modeled as MRFs and their MAP estimates are obtained using simulated annealing. The performance of the proposed scheme is tested on synthetic as well as real data and the estimates of the depth are found to be better than that of existing window-based techniques. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
ICCV | 2 |
| 1998 | Finding Faces in PhotographsabstractTwo new schemes are presented for finding human faces in a photograph. The first scheme approximates the unknown distributions of the face and the face-like manifolds wing higher order statistics (HOS). An HOS-based data clustering algorithm is also proposed. In the second scheme, the face to non-face and non-face to face transitions are learnt using a hidden Markov model (HMM). The HMM parameters are estimated corresponding to a given photograph and the faces are located by examining the optimal state sequence of the HMM. Experimental results are presented on the performance of both the schemes. A. N. Rajagopalan 0001, K. Sunil Kumar, Jayashree Karlekar, R. Manivasakan, M. Milind Patil, Uday B. Desai, P. G. Poonacha, Subhasis Chaudhuri |
ICCV | 8 |
| 1998 | Automatic Generation of Robot Program Code: Learning from Perceptual DataabstractWe propose a novel approach to program a robot by demonstrating the task multiple number of times in front of a vision system. Here we integrate human dexterity with sensory data using computer vision techniques in a single platform. A simultaneous feature detection and tracking framework is used to track various features (finger tips and the wrist joint). A Kalman filter does the tracking by predicting the tentative feature location and a HOS-based data clustering algorithm extracts the feature. Color information of the features are used for establishing correspondences. A fast, efficient and robust algorithm for the vision system thus developed process a binocular video sequence to obtain the trajectories and the orientation information of the end effector. The concept of a trajectory bundle is introduced to avoid singularities and to obtain an optimal path. Mohammed Yeasin, Subhasis Chaudhuri |
ICCV | 2 |
| 1998 | Recursive Estimation of Illuminant Motion from Flow Field and Simultaneous Recovery of Shape
S. G. Deshpande, Subhasis Chaudhuri |
Comput. Vis. Image Underst. | 2 |
| 1998 | Performance Analysis of Maximum Likelihood Estimator for Recovery of Depth from Defocused Images and Optimal Selection of Camera Parameters
A. N. Rajagopalan 0001, Subhasis Chaudhuri |
Int. J. Comput. Vis. | 2 |
| 1998 | Direct parametric object detection in tomographic images
M. Kamath, Subhasis Chaudhuri, Uday B. Desai |
Image Vis. Comput. | 2 |
| 1998 | A recursive algorithm for maximum likelihood-based identification of blur from multiple observationsabstractA maximum likelihood-based method is proposed for blur identification from multiple observations of a scene. When the relations among the blurring functions are known, the estimate of blur obtained using the proposed method is very good. Since direct computation of the likelihood function becomes difficult as the number of images increases, we propose an algorithm to compute the likelihood function recursively. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
IEEE Trans. Image Process. | 2 |
| 1997 | Optimal Selection of Camera Parameters for Recovery of Depth from Defocused ImagesabstractIn the depth from defocus (DFD) method two defocused images of a scene are obtained by capturing the scene with different sets of camera parameters. An arbitrary selection of the camera settings can result in observed images whose relative blurring is insufficient to yield a good estimate of the depth. In this paper, we study the effect of the degree of relative blurring on the accuracy of the estimate of the depth by addressing the DFD problem in a maximum likelihood-based framework. We propose a criterion for optimal selection of camera parameters to obtain an improved estimate of the depth. The optimality criterion is based on the Cramer-Rao bound of the variance of the error in the estimate of blur. Simulations as well as experimental results on real images are presented for validation. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
CVPR | 2 |
| 1997 | Maximum likelihood estimation of blur from multiple observationsabstractA limitation of the existing maximum likelihood (ML) based methods for blur identification is that the estimate of blur is poor when the blurring is severe. In this paper, we propose an ML-based method for blur identification from multiple observations of a scene. When the relations among the blurring functions of these observations are known, we show that the estimate of blur obtained by using the proposed method is very good. The improvement is particularly significant under severe blurring conditions. With an increase in the number of images, direct computation of the likelihood function, however, becomes difficult as it involves calculating the determinant and the inverse of the cross-correlation matrix. To tackle this problem, we propose an algorithm that computes the likelihood function recursively as more observations are added. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
ICASSP | 2 |
| 1997 | Robust estimation of multi-component motion in image sequences using the epipolar constraintabstractGiven two frames of a dynamic scene with several rigid body objects undergoing different motions in the three-dimensional space, we robustly estimate the motion and structure of each object. The least median of squares (LMedS) estimator is integrated into a robust 3D motion parameter estimation and scene structure recovery framework to deal with the multi-motion problem. Experimental results underline the capability of the approach to deal successfully with multi-component motion. We apply the approach presented in this paper to the problem of automatic insertion of artificial objects in real image sequences. Eckehard G. Steinbach, Subhasis Chaudhuri, Bernd Girod |
ICASSP | 2 |
| 1997 | Data-Driven Multi-Frame 3D Motion EstimationabstractWe investigate how the temporal evolution of image point correspondences over multiple frames can be exploits for 3D motion and structure estimation. In comparison to other multi-frame approaches we do not separate the establishment of feature point correspondences and the 3D motion parameter computation. As a result our approach does not suffer from the limitations of the traditional two step approaches where even small errors in point correspondence can lead to large errors in 3D motion parameters. Experimental results on synthetic and real video sequences show that the estimation error is reduced when processing more than two frames simultaneously. Eckehard G. Steinbach, Subhasis Chaudhuri, Bernd Girod |
ICIP (1) | 2 |
| 1997 | Space-Variant Approaches to Recovery of Depth from Defocused Images
A. N. Rajagopalan 0001, Subhasis Chaudhuri |
Comput. Vis. Image Underst. | 2 |
| 1997 | A Variational Approach to Recovering Depth From Defocused ImagesabstractIn this paper, we propose a regularized solution to the depth from defocus (DFD) problem using the space-frequency representation (SFR) framework. A smoothness constraint is imposed on the estimates of the blur parameter, and a variational approach to the DFD problem is developed. Among the numerous SFRs, we study the applicability of the complex spectrogram and the Wigner distribution, in particular, for depth recovery. The performance of the proposed variational method is tested on both synthetic and real images. The method yields good results, and the quality of the estimates is significantly better than that obtained without the smoothness constraint on the blur parameter. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Robust detection of skew in document imagesabstractWe describe a robust yet fast algorithm for skew detection in binary document images. The method is based on interline cross-correlation in the scanned image. Instead of finding the correlation for the entire image, it is calculated over small regions selected randomly. The proposed method does not require prior segmentation of the document into text and graphics regions. The maximum median of cross-correlation is used as the criterion to obtain the skew, and a Monte Carlo sampling technique is chosen to determine the number of regions over which the correlations have to be calculated. Experimental results on detecting skews in various types of documents containing different linguistic scripts are presented here. Avanindra Chaudhuri, Subhasis Chaudhuri |
IEEE Trans. Image Process. | 2 |
| 1996 | Recursive estimation of illuminant motion from flow fieldabstractWe consider the problem of estimating the motion of a light source, given a sequence of images of a stationary, unknown object illuminated by it. The apparent flow field induced on the image plane is used to recover the motion parameters. The motion is assumed to be uniform between the successive image frames. A recursive total least squares estimator is developed in this paper to track the motion of the illuminant. We present typical simulation results to validate the proposed scheme. 1. INTRODUCTION Various schemes exist in the literature to track the motion of an object from a sequence of its images [1],[2]. A lot of work has also been done in the estimation of illuminant direction from a single image [3],[4],[5]. However, the problem addressed in this paper is that of estimation of the motion of the illuminant itself by observing an image sequence of an unknown object illuminated by it. The camera and the object are both fixed and the light source illuminating the object is moving.... S. G. Deshpande, Subhasis Chaudhuri |
ICIP (3) | 2 |
| 1996 | Recursive Estimation of Motion Parameters
Subhasis Chaudhuri, Shankar Chatterjee |
Comput. Vis. Image Underst. | 1 |
| 1995 | A block shift-variant blur model for recovering depth from defocused imagesabstractThe recovery of depth from defocus involves calculating the depth of various points in a scene by modeling the effect that the focal parameters of the camera have on images acquired with a small depth of field. We propose a method that, instead of analyzing an image region in isolation, uses a block shift-variant interactive blur model to account for the interaction among neighboring subimages. Simulation results are presented on the performance of the model. A. N. Rajagopalan 0001, Subhasis Chaudhuri |
ICIP (3) | 2 |
| 1995 | Automated assembling of images: image montage preparation
Pankaj Dani, Subhasis Chaudhuri |
Pattern Recognit. | 2 |
| 1991 | Motion analysis of a homogeneously deformable object using subset correspondences
Subhasis Chaudhuri, Shankar Chatterjee |
Pattern Recognit. | 1 |
| 1991 | Performance analysis of total least squares methods in three-dimensional motion estimationabstractAn algorithm is presented to obtain the total least squares (TLS) estimates of the motion parameters of an object from range/stereo data or perspective views in a closed form. TLS estimates are suitable when data in both time frames are corrupted by noise, which is an appropriate model for motion analysis in practice. The robustness of different linear least squares methods is analyzed for the estimation of motion parameters against the sensor noise and possible mismatches in establishing object feature point correspondence. As the errors in point correspondence increase, the performance of an ordinary least squares (LS) estimator was found to deteriorate much faster than that of the TLS estimator. The Cramer-Rao lower bound (CRLB) of the error covariance matrix was derived for the TLS model under the assumption of uncorrelated additive Gaussian noise. The CRLB for the TLS model is shown to be always higher than that for the LS model.> Subhasis Chaudhuri, Shankar Chatterjee |
IEEE Trans. Robotics Autom. | 1 |
| 1990 | A fast method for global surface interpolationabstractA fast algorithm is presented for a globally smooth interpolation for a visual surface from scattered range data. This method is based on matching lower order spatial moments, where the reconstructed surface is given by a linear combination of Legendre polynomials. This method can handle data that do not lie on a regular grid. The value of the function at points where the grid is not defined, or where the original data is missing, can be approximated from the constituent polynomials. Some simulation experiments evaluate the performance of the proposed scheme. An arbitrarily shaped surface is shown on a 35*35 grid, and the depth information is available only at a few selected grid points. The result of the implementation of Grimson's (1981) algorithm where the surface is constrained to pass through the given data set is given. The proposed method runs much faster than Grimson's algorithm. The coefficients of the Legendre polynomials can be precomputed.> Subhasis Chaudhuri, Shankar Chatterjee |
ICASSP | 1 |
| 1989 | Estimation of motion parameters for a deformable object from range dataabstractIf the correspondence between two sets of points representing the coordinates of different points of an object undergoing rotational motion and deformation is known, the parameters can be estimated using different least-squares estimators. The total-least-squares (TLS) method is very appropriate when the observation and the data matrices are both perturbed by random noise. For Gaussian-distributed noise, the TLS solution is equivalent to maximum-likelihood estimation. The mean-square error in TLS is always smaller than in an ordinary least-squares (LS) estimator. The scope is analyzed of TLS in estimating the generalized motion parameters, as is the feasibility of decomposing the generalized motion parameters in terms of rotation and deformation parameters. The performance of TLS is compared to that of the LS estimator.> Subhasis Chaudhuri, Shankar Chatterjee |
CVPR | 1 |