Soma Biswas

dblp:82/2665 · DBLP profile ↗
← Back
79ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-9068-7023ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 36 · 10 first-author · 10 since 2021Security and privacy · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Fully Unsupervised Self-debiasing of Text-to-Image Diffusion Models
abstract
Text-to-image (T2I) diffusion models have achieved widespread success due to their ability to generate high-resolution, photorealistic images. These models are trained on large-scale datasets, like LAION-5B, often scraped from the internet. However, since this data contains numerous biases, the models inherently learn and reproduce them, resulting in stereotypical outputs. We introduce SelfDebias, a fully unsupervised test-time debiasing method applicable to any diffusion model that uses a UNet as its noise predictor. SelfDebias identifies semantic clusters in an image encoder’s embedding space and uses these clusters to guide the diffusion process during inference, minimizing the KL divergence between the output distribution and the uniform distribution. Unlike supervised approaches, SelfDebias does not require human-annotated datasets or external classifiers trained for each generated concept. Instead, it is designed to automatically identify semantic modes. Extensive experiments show that SelfDebias generalizes across prompts and diffusion model architectures, including both conditional and unconditional models. It not only effectively debiases images along key demographic dimensions while maintaining the visual fidelity of the generated images, but also more abstract concepts for which identifying biases is also challenging.
Korada Sri Vardhana, Shrikrishna Lolla, Soma Biswas
WACV3
2025 Semantic Prototype-Guided Sampling for Long-Tailed Generalized Category Discovery
abstract
In this work, we address the challenging, real-world data imbalance problem in the context of Generalized Category Discovery (GCD). Here, given a large set of labeled and unlabeled data, the goal is to simultaneously classify known classes and also discover novel classes. Towards this goal, we propose an effective framework, ProtoGuide, which consists of two novel modules built on top of a vision transformer with contrastive learning, namely (i) Text-Enhanced Prototype Generation: where we utilize the text embedding of each labeled class names along with the augmented images for computing semantically meaningful class prototypes; and (ii) Prototype-guided sampling: where we dynamically sample examples guided by the current class prototypes for the minority classes based on the predicted imbalance. These modules along with utilizing the nearest neighbor information for each instance in an unsupervised contrastive manner allow us to develop strong feature representations, thereby enhancing clustering performance. Extensive experiments on multiple benchmarks show that ProtoGuide significantly outperforms the state-of-the-art for this challenging task.
Nitu Pathak, Anwesha Banerjee, Soma Biswas
ICIP3
2025 GeNGuide: Generalized Neighbour Guidance Framework for Noisy Class Incremental Learning
abstract
This work addresses the challenging, real-world problem of Noisy-Label Class Incremental Learning, where label noise adversely affects the performance at each incremental task. Towards this goal, we propose a two-stage approach leveraging pre-trained models effectively and information from neighbouring samples. In the first stage, we introduce a neighbour-guidance cross-entropy loss function for feature learning, where the contributions of each data sample is computed based on the labels of neighbouring samples in the feature space, which also evolves as training progresses. This adaptive loss prioritizes reliable samples and mitigates the influence of noisy ones. In the second stage, classifiers are refined using weighted class statistics, which is also guided by the neighbourhood information. This step further enhances the model’s performance by aligning classifiers with the learned class distributions. The proposed GeNGuide (Generalized Neighbour Guidance) framework works seamlessly without any modifications for several scenarios of class-incremental learning, namely (i) when the class-labels have varying amount of noise, including the noise-less case; (ii) when, in addition to the presence of noisy labels, the data is imbalanced, which makes the problem even more challenging. To the best of our knowledge, this is the first work which takes a step towards building generalized models which achieves state-of-the-art performance on several challenging class-incremental learning protocols, thereby justifying its effectiveness. https://github.com/avnCode/GeNGuide.git
Avnish Kumar, Jayateja Kalla, Soma Biswas
IJCNN3
2025 Can Out-of-Domain Data Help to Learn Domain-Specific Prompts for Multimodal Misinformation Detection?
Amartya Bhattacharya, Debarshi Brahma, Suraj Nagaje Mahadev, Anmol Asati, Vikas Verma, Soma Biswas
WACV6
2025 TACLE: Task and Class-Aware Exemplar-Free Semi-Supervised Class Incremental Learning
abstract
We propose a novel TACLE (TAsk and CLass-awarE) framework for the relatively unexplored and challenging problem of exemplar-free semi-supervised class incremen-tal learning. In this scenario, at each new task, the model has to learn new classes from both (few) labeled and unlabeled data without access to exemplars from previous classes. In addition to leveraging the capabilities of pretrained models, TACLE proposes a novel task-adaptive threshold, thereby maximizing the utilization of the available unlabeled data as incremental learning progresses. Additionally, to enhance the performance of the under-represented classes within each task, we propose a class-aware weighted cross-entropy loss. We also exploit the unlabeled data for classifier alignment, which further enhances the model performance. Extensive experiments on benchmark datasets, namely CIFAR10, CIFAR100, and ImageNet-Subset100 demonstrate the effectiveness of the proposed TACLE framework. We further showcase its effectiveness when the unlabeled data is imbalanced and also for the extreme case of one labeled example per class. code: https://github.com/rokmr/TACLE
Jayateja Kalla, Soma Biswas
WACV3
2025 FRAUD-Net: Fraud News Detection Using Sample Uncertainty & Domain Aware Generalized Network
abstract
Due to the widespread impact of social media, efficiently detecting out-of-context misinformation, where a real image is paired with a fake caption has become imperative. Towards this goal, we propose a novel framework FRAUD-Net, which incorporates several unexplored aspects of this task in the model formulation. Keeping in mind that the image and textual evidences collected using the input image-text pair by web search play a crucial role in this task, we build a generalized model, which can handle missing or variable number of evidences, as expected in real-world scenarios. We achieve this by effectively utilizing the common semantic latent space of Visual Language Models and transformer attention blocks. In addition, we observe that explicit domain information not only allows the model to learn better, but also dynamically adjusts to the diverse domains (like politics, healthcare, etc.) during testing. We also propose a mechanism to handle noisy training data by analyzing prediction consistency which helps in model training. Extensive experiments on the large-scale NewsClippings and Verite benchmark datasets showcase the effectiveness of the proposed framework compared to the state-of-the-art techniques for this challenging task.
Devendra Patel, Vikas Verma, Shreyas Kumar Tah, Shwetabh Biswas, Soma Biswas
WACV5
2024 CoVLM: Leveraging Consensus from Vision-Language Models for Semi-supervised Multi-modal Fake News Detection
Devank, Jayateja Kalla, Soma Biswas
ACCV (6)3
2024 AggSS: An Aggregated Self-Supervised Approach for Class Incremental Learning
Jayateja Kalla, Soma Biswas
BMVC2
2024 Adaprompt: Prompt Tuning with Adaptive Neighbours for Generalized Category Discovery
abstract
In this work, we address the challenging task of generalized category discovery, where the goal is to correctly classify unlabelled objects from previously seen classes, while also categorizing instances of completely new, unseen classes. Inspired by the remarkable success of prompt tuning of Vision Transformer models for this task, we propose two novel modifications to further improve their effectiveness. First, we propose to simultaneously perform parametric classification and representation learning with prompt tuning in an end-to-end learning framework, instead of a two-stage process of representation learning followed by clustering. Second, we introduce an adaptive neighborhood strategy for positive mining and contrastive learning, eliminating the need for large memory banks for affinity learning. We demonstrate the effectiveness and efficiency of our approach in comparison with the state-of-the-art methods on four benchmark datasets.
Liyana Sahir, Anwesha Banerjee, Soma Biswas
ICIP3
2024 AMEND: Adaptive Margin and Expanded Neighborhood for Efficient Generalized Category Discovery
abstract
Generalized Category Discovery aims to discover and cluster images from previously unseen classes, in addition to classifying images from seen classes correctly. In this work, we propose a simple, yet effective framework for this task, which not only performs on-par or better with the current approaches but is also significantly more efficient in terms of computational requirements. Our first contribution is to use expanded neighborhood information in contrastive learning to generate robust and generalizable features. To generate more discriminative feature representations, especially for fine-grained datasets and confusing classes, we propose a class-wise adaptive margin regularizer that aims at increasing the angular separation among the prototypes of all classes. Extensive experiments on three generic as well as four fine-grained benchmark datasets show the usefulness of the proposed Adaptive Margin and Expanded Neighborhood (AMEND) framework.
Anwesha Banerjee, Liyana Sahir Kallooriyakath, Soma Biswas
WACV3
2024 PhISH-Net: Physics Inspired System for High Resolution Underwater Image Enhancement
abstract
Underwater imaging presents numerous challenges due to refraction, light absorption, and scattering, resulting in color degradation, low contrast, and blurriness. Enhancing underwater images is crucial for high-level computer vision tasks, but existing methods either neglect the physics-based image formation process or require expensive computations. In this paper, we propose an effective framework that combines a physics-based Underwater Image Formation Model (UIFM) with a deep image enhancement approach based on the retinex model. Firstly, we remove backscatter by estimating attenuation coefficients using depth information. Then, we employ a retinex model-based deep image enhancement module to enhance the images. To ensure adherence to the UIFM, we introduce a novel Wideband Attenuation prior. The proposed PhISH-Net framework achieves real-time processing of high-resolution underwater images using a lightweight neural network and a bilateral-grid-based upsampler. Extensive experiments on two underwater image datasets demonstrate the superior performance of our method compared to state-of-the-art techniques. Additionally, qualitative evaluation on a cross-dataset scenario confirms its generalization capability. Our contributions lie in combining the physics-based UIFM with deep image enhancement methods, introducing the wideband attenuation prior, and achieving superior performance and efficiency.
Aditya Chandrasekar, Manogna Sreenivas, Soma Biswas
WACV3
2024 Robust Feature Learning and Global Variance-Driven Classifier Alignment for Long-Tail Class Incremental Learning
abstract
This paper introduces a two-stage framework designed to enhance long-tail class incremental learning, enabling the model to progressively learn new classes, while mitigating catastrophic forgetting in the context of long-tailed data distributions. Addressing the challenge posed by the under-representation of tail classes in long-tail class incremental learning, our approach achieves classifier alignment by leveraging global variance as an informative measure and class prototypes in the second stage. This process effectively captures class properties and eliminates the need for data balancing or additional layer tuning. Alongside traditional class incremental learning losses in the first stage, the proposed approach incorporates mixup classes to learn robust feature representations, ensuring smoother boundaries. The proposed framework can seamlessly integrate as a module with any class incremental learning method to effectively handle long-tail class incremental learning scenarios. Extensive experimentation on the CIFAR-100 and ImageNet-Subset datasets validates the approach’s efficacy, showcasing its superiority over state-of-the-art techniques across various long-tail CIL settings. Code is available at https://github.com/JAYATEJAK/GVAlign.
Jayateja Kalla, Soma Biswas
WACV2
2024 pSTarC: Pseudo Source Guided Target Clustering for Fully Test-Time Adaptation
abstract
Test Time Adaptation (TTA) is a pivotal concept in machine learning, enabling models to perform well in real-world scenarios, where test data distribution differs from training. In this work, we propose a novel approach called pseudo Source guided Target Clustering (pSTarC) addressing the relatively unexplored area of TTA under real-world domain shifts. This method draws inspiration from target clustering techniques and exploits the source classifier for generating pseudo-source samples. The test samples are strategically aligned with these pseudo-source samples, facilitating their clustering and thereby enhancing TTA performance. pSTarC operates solely within the fully test-time adaptation protocol, removing the need for actual source data. Experimental validation on a variety of domain shift datasets, namely VisDA, Office-Home, DomainNet-126, CIFAR-100C verifies pSTarC’s effectiveness. This method exhibits significant improvements in prediction accuracy along with efficient computational requirements. Furthermore, we also demonstrate the universality of the pSTarC framework by showing its effectiveness for the continuous TTA framework.
Manogna Sreenivas, Goirik Chakrabarty, Soma Biswas
WACV3
2024 Generalized semi-supervised class incremental learning in presence of outliers
Jayateja Kalla, Prishruit Punia, Titir Dutta, Soma Biswas
Multim. Tools Appl.4
2023 Test Time Adaptation for Blind Image Quality Assessment
abstract
While the design of blind image quality assessment (IQA) algorithms has improved significantly, the distribution shift between the training and testing scenarios often leads to a poor performance of these methods at inference time. This motivates the study of test time adaptation (TTA) techniques to improve their performance at inference time. Existing auxiliary tasks and loss functions used for TTA may not be relevant for quality-aware adaptation of the pre-trained model. In this work, we introduce two novel quality-relevant auxiliary tasks at the batch and sample levels to enable TTA for blind IQA. In particular, we introduce a group contrastive loss at the batch level and a relative rank loss at the sample level to make the model quality aware and adapt to the target data. Our experiments reveal that even using a small batch of images from the test distribution helps achieve significant improvement in performance by updating the batch normalization statistics of the source model.
Subhadeep Roy, Shankhanil Mitra, Soma Biswas, Rajiv Soundararajan
ICCV3
2022 SEIC: Semantic Embedding with Intermediate Classes for Zero-Shot Domain Generalization
Soma Biswas
ACCV (5)2
2022 Handling Class-Imbalance for Improved Zero-Shot Domain Generalization
Ahmad Arfeen, Titir Dutta, Soma Biswas
BMVC3
2022 Novel Class Discovery Without Forgetting
K. J. Joseph, Sujoy Paul, Gaurav Aggarwal, Soma Biswas, Piyush Rai, Kai Han 0001, Vineeth N. Balasubramanian
ECCV (24)4
2022 S3C: Self-Supervised Stochastic Classifiers for Few-Shot Class-Incremental Learning
Jayateja Kalla, Soma Biswas
ECCV (25)2
2021 Universal Cross-Domain Retrieval: Generalizing Across Classes and Domains
abstract
In this work, for the first time, we address the problem of universal cross-domain retrieval, where the test data can belong to classes or domains which are unseen during training. Due to dynamically increasing number of categories and practical constraint of training on every possible domain, which requires large amounts of data, generalizing to both unseen classes and domains is important. Towards that goal, we propose SnMpNet (Semantic Neighbourhood and Mixture Prediction Network), which incorporates two novel losses to account for the unseen classes and domains encountered during testing. Specifically, we introduce a novel Semantic Neighborhood loss to bridge the knowledge gap between seen and unseen classes and ensure that the latent space embedding of the unseen classes is semantically meaningful with respect to its neighboring classes. We also introduce a mix-up based supervision at image-level as well as semantic-level of the data for training with the Mixture Prediction loss, which helps in efficient retrieval when the query belongs to an unseen domain. These losses are incorporated on the SE-ResNet50 backbone to obtain SnMpNet. Extensive experiments on two large-scale datasets, Sketchy Extended and DomainNet, and thorough comparisons with state-of-the-art justify the effectiveness of the proposed model.
Soumava Paul, Titir Dutta, Soma Biswas
ICCV3
2021 SML: Semantic meta-learning for few-shot semantic segmentation☆
Ayyappa Kumar Pambala, Titir Dutta, Soma Biswas
Pattern Recognit. Lett.3
2021 StyleGuide: Zero-Shot Sketch-Based Image Retrieval Using Style-Guided Image Generation
abstract
The goal of zero-shot sketch-based image retrieval is to retrieve relevant images from a search set against a hand-drawn sketch query, which belongs to a class, previously unseen by the model. The knowledge gap between such unseen and seen classes along with the domain-gap between the query and search-set makes the problem extremely challenging. In this work, we address this problem by proposing a novel retrieval methodology,StyleGuideusing style-guidedfake-image generation. In addition, we further study the scenario of generalized zero-shot sketch-based image retrieval, where the search set contains images from both seen and unseen categories. Specifically, we propose a detection approach for unseen class samples in the search-set, based on pre-computed seen class-prototypes, to obtain a refined search-set for a particular unseen-class query. Thus, the query sketch needs to be compared only to those image data which are more likely to belong to the unseen classes, resulting in improved retrieval performance. Extensive experiments on two large-scale sketch-image datasets, Sketchy extended and TU-Berlin show that the proposed approach performs better or comparable to the state-of-the-art for ZS-SBIR and gives significant improvements over the state-of-the-art for generalized ZS-SBIR.
Titir Dutta, Soma Biswas
IEEE Trans. Multim.3
2020 Adaptive Margin Diversity Regularizer for Handling Data Imbalance in Zero-Shot SBIR
Titir Dutta, Soma Biswas
ECCV (5)3
2020 Cross-Modal Retrieval With Noisy Labels
abstract
Cross-modal retrieval is an important field of study for design of algorithms to effectively retrieve items from one modality when provided with a query from another modality. Recent progress in this field have shown that supervised algorithms perform significantly better than their unsupervised counterparts by utilizing the label information. In real scenarios, the labels are obtained through manual or automatic annotation, and thus are prone to errors. In this work, we systematically study the effect of label corruption on the performance of standard cross-modal algorithms. We propose a very simple, yet effective pre-processing framework which can help to mitigate the performance degradation suffered due to label corruption. First, the potentially more promising modality is automatically chosen, on which two different versions of a noise-resistant classification algorithm is trained to generate the pseudo-labels of the noisy cross-modal training data. The generated pseudo-labels can then be used by any cross-modal supervised approach to improve its performance. Extensive experiments across four cross-modal datasets with different types of label corruption show that the proposed framework gives impressive improvements for this important problem.
Devraj Mandal, Soma Biswas
ICIP2
2020 Label Prediction Framework For Semi-Supervised Cross-Modal Retrieval
abstract
Cross-modal data matching refers to retrieval of data from one modality, when given a query from another modality. In general, supervised algorithms achieve better retrieval performance compared to their unsupervised counterpart, as they can learn better representative features by leveraging the available label information. However, this comes at the cost of requiring huge amount of labeled examples, which may not always be available. In this work, we propose a novel framework in a semi-supervised cross-modal retrieval setting, which can predict the labels of the unlabeled data using complementary information from different modalities. The proposed framework can be used as an add-on with any baseline cross-modal algorithm to give significant performance improvement, even in case of limited labeled data. Extensive evaluation using several baseline algorithms across three different datasets show the effectiveness of our label prediction framework.
Devraj Mandal, Pramod Rao, Soma Biswas
ICIP3
2020 Multi-class Novelty Detection Using Mix-up Technique
abstract
Multi-class novelty detection is increasingly becoming an important area of research due to the continuous increase in the number of object categories. It tries to answer the pertinent question: given a test sample, should we even try to classify it? We propose a novel solution using the concept of mix-up technique for novelty detection, termed as Segregation Network. During training, a pair of examples are selected from the training data and an interpolated data point using their convex combination is constructed. We develop a suitable loss function to train our model to predict its constituent classes. During testing, each input query is combined with the known class prototypes to generate mixed samples which are then passed through the trained network. Our model which is trained to reveal the constituent classes can then be used to determine whether the sample is novel or not. The intuition is that if a query comes from a known class and is mixed with the set of known class prototypes, then the prediction of the trained model for the correct class should be high. In contrast, for a query from a novel class, the predictions for all the known classes should be low. The proposed model is trained using only the available known class data and does not need access to any auxiliary dataset or attributes. Extensive experiments on two benchmark datasets, namely Caltech 256 and Stanford Dogs and comparisons with the state-of-the-art algorithms justifies the usefulness of our approach.
Supritam Bhattacharjee, Devraj Mandal, Soma Biswas
WACV3
2020 s-SBIR: Style Augmented Sketch based Image Retrieval
abstract
Sketch-based image retrieval (SBIR) is gaining increasing popularity because of its flexibility to search natural images using unrestricted hand-drawn sketch query. Here, we address a related, but relatively unexplored problem, where the users can also specify their preferred styles of the images they want to retrieve, e.g., color, shape, etc., as keywords, whose information is not present in the sketch. The contribution of this work is three-fold. First, we propose a deep network for the problem of style-augmented SBIR (or s-SBIR) having three main components - category module, style module and mixer module, which are trained in an end-to-end manner. Second, we propose a quintuplet loss, which takes into consideration both the category and style, while giving appropriate importance to the two components. Third, we propose a normalized composite evaluation metric or ncMAP which can quantitatively evaluate s-SBIR approaches. Extensive experiments on subsets of two benchmark image-sketch datasets, Sketchy and TU-Berlin show the effectiveness of the proposed approach.
Titir Dutta, Soma Biswas
WACV2
2020 A Novel Self-Supervised Re-labeling Approach for Training with Noisy Labels
abstract
The major driving force behind the immense success of deep learning models is the availability of large datasets along with their clean labels. This is very difficult to obtain and thus has motivated research on training deep neural networks in the presence of label noise. In this work, we build upon the seminal work in this area, Co-teaching and propose a simple, yet efficient approach termed mCT-S2R (modified co-teaching with self-supervision and relabeling) for this task. Firstly, to deal with significant amount of noise in the labels, we propose to use self-supervision to generate robust features without using any labels. Furthermore, using a parallel network architecture, an estimate of the clean labeled portion of the data is obtained. Finally, using this data, a portion of the estimated noisy labeled portion is re-labeled, before resuming the network training with the augmented data. Extensive experiments on three standard datasets show the effectiveness of the proposed framework.
Devraj Mandal, Shrisha Bharadwaj, Soma Biswas
WACV3
2020 Generative Model with Semantic Embedding and Integrated Classifier for Generalized Zero-Shot Learning
abstract
Generative models have achieved impressive performance for the generalized zero-shot learning task by learning the mapping from attributes to feature space. In this work, we propose to derive semantic inferences from images and use them for the generation, which enables us to capture the bidirectional information i.e., visual to semantic and semantic to visual spaces. Specifically, we propose a Semantic Embedding module which not only gives image specific semantic information to the generative model for generation of better features, but also makes sure that the generated features can be mapped to the correct semantic space. We also propose an Integrated Classifier, which is trained along with the generator. This module not only eliminates the requirement of additional classifier for new object categories which is required by the existing generative approaches, but also facilitates the generation of more discriminative and useful features. This approach can be used seamlessly for the task of few-shot learning. Extensive experiments on four benchmark datasets, namely, CUB, SUN, AWA1, AWA2 for both zero-shot learning and few-shot setting show the effectiveness of the proposed approach.
Ayyappa Kumar Pambala, Titir Dutta, Soma Biswas
WACV3
2020 Robust image retrieval by cascading a deep quality assessment network
abstract
The performance of computer vision algorithms can severely degrade in the presence of a variety of distortions. While image enhancement algorithms have evolved to optimize image quality as measured according to human visual perception , their relevance in maximizing the success of computer vision algorithms operating on the enhanced image has been much less investigated. We consider the problem of image enhancement to combat Gaussian noise and low resolution with respect to the specific application of image retrieval from a dataset. We define the notion of image quality as determined by the success of image retrieval and design a deep convolutional neural network (CNN) to predict this quality. This network is then cascaded with a deep CNN designed for image denoising or super resolution , allowing for optimization of the enhancement CNN to maximize retrieval performance . This framework allows us to couple enhancement to the retrieval problem. We also consider the problem of adapting image features for robust retrieval performance in the presence of distortions. We show through experiments on distorted images of the Oxford and Paris buildings datasets that our algorithms yield improved mean average precision when compared to using enhancement methods that are oblivious to the task of image retrieval. 1
Biju Venkadath Somasundaran, Rajiv Soundararajan, Soma Biswas
Signal Process. Image Commun.3
2020 Semi-Supervised Cross-Modal Retrieval With Label Prediction
abstract
Cross-modal retrieval tasks with image-text, audio-image, etc. are gaining increasing importance due to an abundance of data from multiple modalities. In general, supervised approaches give significant improvement over their unsupervised counterparts at the additional cost of labeling or annotation of the training data. Recently, semi-supervised methods are becoming popular as they provide an elegant framework to balance the conflicting requirement of labeling cost and accuracy. In this work, we propose a novel deep semi-supervised framework, which can seamlessly handle both labeled as well as unlabeled data. The network has two important components: (a) first, the labels for the unlabeled portion of the training data are predicted using the label prediction component, and then (b) a common representation for both the modalities is learned for performing cross-modal retrieval. The two parts of the network are trained sequentially one after the other. Extensive experiments on three benchmark datasets, Wiki, Pascal VOC, and NUS-WIDE demonstrate that the proposed framework outperforms the state-of-the-art for both supervised and semi-supervised settings.
Devraj Mandal, Pramod Rao, Soma Biswas
IEEE Trans. Multim.3
2019 Style-Guided Zero-Shot Sketch-based Image Retrieval
Titir Dutta, Soma Biswas
BMVC2
2019 Do I Know You? A Two-Stage Framework for Novelty Detection
abstract
In this work, we address the problem of novelty detection in the context of image classification, where the goal is to classify whether a query belongs to one of the classes seen during training or to a novel class. Given a network trained on a set of training classes, we utilize the activations from the final fully connected layer in a two-stage framework to determine whether a given query is seen or novel. In the first stage, for a given query, analyzing the top retrieved samples gives us a set of probable categories for that query. For the second stage, a comparator network is designed which can compare features from two samples and classify them into similar or dissimilar pairs. Prototype exemplars from each of the probable categories computed in the first stage are compared with the query feature to determine whether they come from the same or different category. The scores from both the stages are fused using a simple yet effective fuzzy aggregation operator to give a final decision on whether the given query is from a seen or novel class. Extensive experiments on three benchmark databases, MNIST, HASYv2, and Fashion-MNIST demonstrate the effectiveness of the proposed framework.
Supritam Bhattacharjee, Sivaram Prasad Mudunuri, Soma Biswas
ICIP3
2019 Autoencoder based novelty detection for generalized zero shot learning
abstract
The problem of generalized zero-shot learning deals with the classification of test examples for which training data may or may not be available. Existing baseline algorithms connect the seen and unseen set of categories by learning functions to project the image data into the attribute space or vice versa. However, since the classification framework is trained only on the seen set of categories, the recognition performance is typically biased and algorithms have great difficulty in recognizing novel classes. In this work, we investigate the usefulness of a novelty detector to recognize a given data as coming from the seen or novel set. The proposed novelty detector is based on an autoencoder network structure with reconstruction and triplet cosine embedding losses which can be effectively trained using only the seen data and its categories. Experiments over a variety of benchmark datasets and zero-shot algorithms show the efficacy of the proposed approach.
Supritam Bhattacharjee, Devraj Mandal, Soma Biswas
ICIP3
2019 Principle-to-program: Neural Fashion Recommendation with Multi-modal Input
abstract
Outfit recommendation automatically pairs user-specified reference clothing with the most suitable complement from online shops. Wearing aesthetically is a criterion for matching such fashion items. Fashion style tells a lot about one's personality and emerges from how people assemble clothing outfit from seemingly disjoint items into a cohesive concept. Experts share fashion tips showcasing their compositions to public where each item has both an image and textual meta-data. Also, retrieving products from online shopping catalogs in response to such real-world image query is essential for outfit recommendation. Our earlier tutorial focused on style and compatibility in fashion recommendation mostly based on metric and deep learning approaches. Herein, we cover several other aspects of fashion recommendation using visual signals (e.g., cross-scenario retrieval, attribute classification) and combine text input (e.g., interpretable embedding) as well. Each section concludes walking through programs executed on Jupyter workstation using real-world data sets.
Muthusamy Chelliah, Soma Biswas, Lucky Dhakad
ACM Multimedia2
2019 Cross-modal retrieval in challenging scenarios using attributes
Titir Dutta, Soma Biswas
Pattern Recognit. Lett.2
2019 Machine vision quality assessment for robust face detection
Rajiv Soundararajan, Soma Biswas
Signal Process. Image Commun.2
2019 Dictionary Alignment With Re-Ranking for Low-Resolution NIR-VIS Face Recognition
abstract
Recently, near-infrared (NIR) images are increasingly being captured for recognizing faces in low-light/night-time conditions. Matching these images against the controlled high-resolution visible facial images usually present in the database is a challenging task. In surveillance scenarios, the NIR images can have very low resolution and also non-frontal pose which makes the problem even more challenging. In this paper, we propose an orthogonal dictionary alignment approach for addressing this problem. We also propose a re-ranking approach to further improve the recognition performance for each probe by combining the rank list given by the proposed algorithm with that given by another complementary feature/algorithm. Finally, we have also collected our own database Heterogeneous face recognition across Pose and Resolution (HPR) that has facial images captured from two surveillance quality NIR cameras and one high-resolution visible camera, with significant variations in head pose and resolution. Extensive experiments on the modified CASIA NIR-VIS 2.0 database, the Surveillance Camera face database, and our HPR database show the effectiveness of the proposed approaches and the collected database.
Sivaram Prasad Mudunuri, Shashanka Venkataramanan, Soma Biswas
IEEE Trans. Inf. Forensics Secur.3
2019 Generalized Zero-Shot Cross-Modal Retrieval
abstract
Cross-modal retrieval is an important research area due to its wide range of applications, and several algorithms have been proposed to address this task. We feel that it is the right time to take a step back and analyze the current status of research in this area. As new object classes are continuously being discovered over time, it is necessary to design algorithms that can generalize to data from previously unseen classes. Towards that goal, our first contribution is to establish protocols for generalized zero-shot cross-modal retrieval and analyze the generalization ability of the standard cross-modal algorithms. Second, we propose a semantic-aware ranking algorithm that can be used as an add-on to any existing cross-modal approach to improve its performance on both seen and unseen classes. Finally, we propose a modification of the standard evaluation metric (MAP for single-label data and NDCG for multi-label data), which we feel is a more intuitive measure of the cross-modal retrieval performance. Extensive experiments on two single-label and three multi-label cross-modal datasets show the effectiveness of the proposed approach.
Titir Dutta, Soma Biswas
IEEE Trans. Image Process.2
2019 Generalized Semantic Preserving Hashing for Cross-Modal Retrieval
abstract
Cross-modal retrieval is gaining importance due to the availability of large amounts of multimedia data. Hashing-based techniques provide an attractive solution to this problem when the data size is large. For cross-modal retrieval, data from the two modalities may be associated with a single label or multiple labels, and in addition, may or may not have a one-to-one correspondence. This work proposes a simple hashing framework which has the capability to work with different scenarios while effectively capturing the semantic relationship between the data items. The work proceeds in two stages in which the first stage learns the optimum hash codes by factorizing an affinity matrix, constructed using the label information. In the second stage, ridge regression and kernel logistic regression is used to learn the hash functions for mapping the input data to the bit domain. We also propose a novel iterative solution for cases where the training data is very large, or when the whole training data is not available at once. Extensive experiments on single label data set like Wiki and multi-label datasets like MirFlickr, NUS-WIDE, Pascal, and LabelMe, and comparisons with the state-of-the-art, shows the usefulness of the proposed approach.
Devraj Mandal, Kunal N. Chaudhury, Soma Biswas
IEEE Trans. Image Process.3
2018 GrowBit: Incremental Hashing for Cross-Modal Retrieval
Devraj Mandal, Yashas Annadani, Soma Biswas
ACCV (4)3
2018 Preserving Semantic Relations for Zero-Shot Learning
abstract
Zero-shot learning has gained popularity due to its potential to scale recognition models without requiring additional training data. This is usually achieved by associating categories with their semantic information like attributes. However, we believe that the potential offered by this paradigm is not yet fully exploited. In this work, we propose to utilize the structure of the space spanned by the attributes using a set of relations. We devise objective functions to preserve these relations in the embedding space, thereby inducing semanticity to the embedding space. Through extensive experimental evaluation on five benchmark datasets, we demonstrate that inducing semanticity to the embedding space is beneficial for zero-shot learning. The proposed approach outperforms the state-of-the-art on the standard zero-shot setting as well as the more realistic generalized zero-shot setting. We also demonstrate how the proposed approach can be useful for making approximate semantic inferences about an image belonging to a category for which attribute information is not available.
Yashas Annadani, Soma Biswas
CVPR2
2018 Coarse to Fine Training for Low-Resolution Heterogeneous Face Recognition
abstract
Recently, near-infrared (NIR) images are being increasingly used for recognizing facial images across illumination variations and in low-light conditions. In surveillance scenarios, the captured NIR may have low-resolution which results in significant loss of discriminative information along with uncontrolled pose. In this work, we address the challenging task of matching these low-resolution (LR) uncontrolled NIR images with high-resolution (HR) controlled visible (VIS) images usually present in the database. Since the probe and gallery images differ significantly in terms of pose, resolution and spectral properties, we employ a two-stage approach. First, the images are transformed into a common space using metric learning such that the images of the same subject are pushed closer and those of different subjects are pushed apart. We then define an objective function which can simultaneously push both LR NIR and HR VIS samples towards the centroids of the HR VIS samples. We show that the approach is general and can be used for other data like RGB-D and also for matching across pose. Extensive experiments conducted on five datasets shows the effectiveness of our approach.
Sivaram Prasad Mudunuri, Soma Biswas
ICIP2
2018 Image Denoising for Image Retrieval by Cascading a Deep Quality Assessment Network
abstract
Image denoising algorithms have evolved to optimize image quality as measured according to human visual perception. However, image denoising to maximize the success of computer vision algorithms operating on the denoised image has been much less investigated. We consider the problem of image denoising for Gaussian noise with respect to the specific application of image retrieval from a dataset. We define the notion of image quality as determined by the success of image retrieval and design a deep convolutional neural network (CNN) to predict this quality. This network is then cascaded with a deep CNN designed for image denoising, allowing for optimization of the denoising CNN to maximize retrieval performance. This framework allows us to couple denoising to the retrieval problem. We show through experiments on noisy images of the Oxford and Paris buildings datasets that such an approach yields improved mean average precision when compared to using denoising methods that are oblivious to the task of image retrieval.
Biju Venkadath Somasundaran, Rajiv Soundararajan, Soma Biswas
ICIP3
2017 Generalized Semantic Preserving Hashing for N-Label Cross-Modal Retrieval
abstract
Due to availability of large amounts of multimedia data, cross-modal matching is gaining increasing importance. Hashing based techniques provide an attractive solution to this problem when the data size is large. Different scenarios of cross-modal matching are possible, for example, data from the different modalities can be associated with a single label or multiple labels, and in addition may or may not have one-to-one correspondence. Most of the existing approaches have been developed for the case where there is one-to-one correspondence between the data of the two modalities. In this paper, we propose a simple, yet effective generalized hashing framework which can work for all the different scenarios, while preserving the semantic distance between the data points. The approach first learns the optimum hash codes for the two modalities simultaneously, so as to preserve the semantic similarity between the data points, and then learns the hash functions to map from the features to the hash codes. Extensive experiments on single label dataset like Wiki and multi-label datasets like NUS-WIDE, Pascal and LabelMe under all the different scenarios and comparisons with the state-of-the-art shows the effectiveness of the proposed approach.
Devraj Mandal, Kunal N. Chaudhury, Soma Biswas
CVPR3
2017 Label consistent matrix factorization based hashing for cross-modal retrieval
abstract
Matrix factorization-based hashing has been very effective in addressing the cross-modal retrieval task. In this work, we propose a novel supervised hashing approach utilizing the concepts of matrix factorization which can seamlessly incorporate the label information. In the proposed approach, the latent factors for each individual modality are generated and then converted to the more discriminative label space using modality specific linear transformations. In the first stage of the approach, the hash codes are learnt using an alternating minimization algorithm and in the next stage, modality specific hash functions are learned to convert the original features of the cross-modal data into the hash code domain. In addition, we also propose an extension of the approach for handling very large amounts of data during the training stage. Extensive experiments performed on the single label Wiki, and the multi-labeled MirFlickr and NUS-WIDE datasets show the effectiveness of the proposed approach.
Devraj Mandal, Soma Biswas
ICIP2
2017 Aligned discriminative pose robust descriptors for face and object recognition
abstract
Face and object recognition in uncontrolled scenarios due to pose and illumination variations, low resolution, etc. is a challenging research area. Here we propose a novel descriptor, Aligned Discriminative Pose Robust (ADPR) descriptor, for matching faces and objects across pose which is also robust to resolution and illumination variations. We generate virtual intermediate pose subspaces from training examples at a few poses and compute the alignment matrices of those subspaces with the frontal subspace. These matrices are then used to align the generated subspaces with the frontal one. An image is represented by a feature set obtained by projecting its low-level feature on these aligned subspaces and applying a discriminative transform. Finally, concatenating all the features we generate the ADPR descriptor. We perform experiments on face and object databases across pose, pose and resolution, and compare with state-of-the-art methods including deep learning approaches to show the effectiveness of our descriptor.
Soubhik Sanyal, Devraj Mandal, Soma Biswas
ICIP3
2017 Dictionary Alignment for Low-Resolution and Heterogeneous Face Recognition
abstract
Cross-domain matching is a challenging problem with several applications like face recognition across pose and resolution, heterogeneous face recognition, etc. Coupled dictionary learning has emerged as a powerful technique for addressing such problems. A novel approach based on aligning two orthogonal dictionaries constructed independently from the two domains is proposed in this work. Once the dictionaries are constructed, the correspondence between the dictionary atoms of the two domains are computed using bipartite graph matching in a common space. A Mahalanobis metric is then derived from sparse coefficient vectors of the aligned dictionaries of the two domains such that the coefficients from data of same class move closer and that of different classes move apart. Unlike other coupled dictionary learning approaches, one-to-one paired training data is not required in the proposed approach. Extensive experiments on MultiPIE, SCFace and MBGC database for face recognition across pose and resolution, CASIA NIRVIS 2.0 database for matching visible to near-infrared face images show the usefulness of the proposed approach for different applications.
Sivaram Prasad Mudunuri, Soma Biswas
WACV2
2017 Abnormality detection in crowd videos by tracking sparse components
Soma Biswas
Mach. Vis. Appl.1
2017 Bayesian Modeling of Temporal Coherence in Videos for Entity Discovery and Summarization
abstract
A video is understood by users in terms of entities present in it. Entity Discovery is the task of building appearance model for each entity (e.g., a person), and finding all its occurrences in the video. We represent a video as a sequence of tracklets, each spanning 10-20 frames, and associated with one entity. We pose Entity Discovery as tracklet clustering, and approach it by leveraging Temporal Coherence (TC): the property that temporally neighboring tracklets are likely to be associated with the same entity. Our major contributions are the first Bayesian nonparametric models for TC at tracklet-level. We extend Chinese Restaurant Process (CRP) to TC-CRP, and further to Temporally Coherent Chinese Restaurant Franchise (TC-CRF) to jointly model entities and temporal segments using mixture components and sparse distributions. For discovering persons in TV serial videos without meta-data like scripts, these methods show considerable improvement over state-of-the-art approaches to tracklet clustering in terms of clustering accuracy, cluster purity and entity coverage. The proposed methods can perform online tracklet clustering on streaming videos unlike existing approaches, and can automatically reject false tracklets. Finally we discuss entity-driven video summarization- where temporal segments of the video are selected based on the discovered entities, to create a semantically meaningful summary.
Adway Mitra, Soma Biswas, Chiranjib Bhattacharyya
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Discriminative pose-free descriptors for face and object matching
Soubhik Sanyal, Sivaram Prasad Mudunuri, Soma Biswas
Pattern Recognit.3
2017 Query specific re-ranking for improved cross-modal retrieval
Devraj Mandal, Soma Biswas
Pattern Recognit. Lett.2
2017 Simultaneous Semi-Coupled Dictionary Learning for Matching in Canonical Space
abstract
Cross-modal recognition and matching with privileged information are important challenging problems in the field of computer vision. The cross-modal scenario deals with matching across different modalities and needs to take care of the large variations present across and within each modality. The privileged information scenario deals with the situation that all the information available during training may not be available during the testing stage, and hence, algorithms need to leverage the extra information from the training stage itself. We show that for multi-modal data, either one of the above situations may arise if one modality is absent during testing. Here, we propose a novel framework, which can handle both these scenarios seamlessly with applications to matching multi-modal data. The proposed approach jointly uses data from the two modalities to build a canonical representation, which encompasses information from both the modalities. We explore four different types of canonical representations for different types of data. The algorithm computes dictionaries and canonical representation for data from both the modalities, such that the transformed sparse coefficients of both the modalities are equal to that of the canonical representation. The sparse coefficients are finally matched using Mahalanobis metric. Extensive experiments on different data sets, involving RGBD, text-image, and audio-image data, show the effectiveness of the proposed framework.
Nilotpal Das, Devraj Mandal, Soma Biswas
IEEE Trans. Image Process.3
2016 Low Resolution Face Recognition Across Variations in Pose and Illumination
abstract
We propose a completely automatic approach for recognizing low resolution face images captured in uncontrolled environment. The approach uses multidimensional scaling to learn a common transformation matrix for the entire face which simultaneously transforms the facial features of the low resolution and the high resolution training images such that the distance between them approximates the distance had both the images been captured under the same controlled imaging conditions. Stereo matching cost is used to obtain the similarity of two images in the transformed space. Though this gives very good recognition performance, the time taken for computing the stereo matching cost is significant. To overcome this limitation, we propose a reference-based approach in which each face image is represented by its stereo matching cost from a few reference images. Experimental evaluation on the real world challenging databases and comparison with the state-of-the-art super-resolution, classifier based and cross modal synthesis techniques show the effectiveness of the proposed algorithm.
Sivaram Prasad Mudunuri, Soma Biswas
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 A coupled discriminative dictionary and transformation learning approach with applications to cross domain matching
Sivaram Prasad Mudunuri, Soma Biswas
Pattern Recognit. Lett.2
2016 Generalized Coupled Dictionary Learning Approach With Applications to Cross-Modal Matching
abstract
Coupled dictionary learning (CDL) has recently emerged as a powerful technique with wide variety of applications ranging from image synthesis to classification tasks. In this paper, we extend the existing CDL approaches in two aspects to make them more suitable for the task of cross-modal matching. Data coming from different modalities may or may not be paired. For example, for image-text retrieval problem, 100 images of a class are available as opposed to only 50 samples of text data for training. Current CDL approaches are not designed to handle such scenarios, where classes of data points in one modality correspond to classes of data points in the other modality. Given the data from the two modalities, first two dictionaries are learnt for the respective modalities, so that the data have a sparse representation with respect to their own dictionaries. Then, the sparse coefficients from the two modalities are transformed in such a manner that data from the same class are maximally correlated, while that from different classes have very less correlation. This way of modeling the coupling between the sparse representations of the two modalities makes this approach work seamlessly for paired as well as unpaired data. The discriminative coupling term also makes the approach better suited for classification tasks. Experiments on different publicly available cross-modal data sets, namely, CUHK photosketch face data set, HFB visible and near-infrared facial images data set, IXMAS multiview action recognition data set, wiki image and text data set and Multiple Features data set, show that this generalized CDL approach performs better than the state-of-the-art for both paired as well as unpaired data.
Devraj Mandal, Soma Biswas
IEEE Trans. Image Process.2
2015 Discriminative Pose-Free Descriptors for Face and Object Matching
abstract
Pose invariant matching is a very important and challenging problem with various applications like recognizing faces in uncontrolled scenarios, matching objects taken from different view points, etc. In this paper, we propose a discriminative pose-free descriptor (DPFD) which can be used to match faces/objects across pose variations. Training examples at very few representative poses are used to generate virtual intermediate pose subspaces. An image or image region is then represented by a feature set obtained by projecting it on all these subspaces and a discriminative transform is applied on this feature set to make it suitable for classification tasks. Finally, this discriminative feature set is represented by a single feature vector, termed as DPFD. The DPFD of images taken from different viewpoints can be directly compared for matching. Extensive experiments on recognizing faces across pose, pose and resolution on the Multi-PIE and Surveillance Cameras Face datasets and comparisons with state-of-the-art approaches show the effectiveness of the proposed approach. Experiments on matching general objects across viewpoints show the generalizability of the proposed approach beyond faces.
Soubhik Sanyal, Sivaram Prasad Mudunuri, Soma Biswas
ICCV3
2015 EntScene: Nonparametric Bayesian Temporal Segmentation of Videos Aimed at Entity-Driven Scene Detection
Adway Mitra, Chiranjib Bhattacharyya, Soma Biswas
IJCAI3
2015 Temporally Coherent CRP: A Bayesian Non-Parametric Approach for Clustering Tracklets with applications to Person Discovery in Videos
abstract
Tracklet Clustering is central to several Computer vision tasks [17][20]. A video can be represented as a sequence of tracklets, each spanning over 10–20 successive video frames, and each tracklet is associated with one entity (eg. person in case of TV-serial videos). Tracklets are instances of data-types exhibiting rich spatio-temporal structure. Existing approaches model tracklets by deploying detailed parametric models with a large number of parameters, making the inference unwieldy. The task of Person Discovery in long TV-series videos (40–45 minutes) with many persons can be naturally posed as tracklet clustering, and existing approaches give unsatisfactory performance on it. In this paper we attempt to leverage Temporal Coherence(TC) of videos to improve tracklet clustering. TC is the fundamental property of videos that each tracklet is likely to be associated with the same entity as its predecessor or successor. We propose the first Bayesian nonparametric approach for modelling TC, which can automatically infer the number of clusters to be formed. The major contribution of this paper is Temporally Coherent Chinese Restaurant Process (TC-CRP), which extends CRP by using TC. On the task of discovering persons in TV serials via tracklet clustering, without meta-data such as scripts, TC-CRP shows up to 25% improvement in cluster purity compared to state-of-the-art parametric models, and upto 36% improvement in number of persons discovered. We use a simple representation of tracklets: a vector of very generic features (like pixel intensity) which can correspond to any type of entity (not necessarily person), and empirically demonstrate the utility of TC-CRP for discovering entities like cars and planes. Moreover, unlike existing approaches TC-CRP can perform online tracklet clustering on streaming videos with very little performance deterioration, and can also automatically reject outliers (tracklets resulting from false detections).
Adway Mitra, Soma Biswas, Chiranjib Bhattacharyya
SDM2
2014 Lesion Detection in Breast Ultrasound Images Using Tissue Transition Analysis
abstract
Breast cancer is one of the leading cause of cancer related deaths in women and early detection is crucial for reducing mortality rates. In this paper, we present a novel and fully automated approach based on tissue transition analysis for lesion detection in breast ultrasound images. Every candidate pixel is classified as belonging to the lesion boundary, lesion interior or normal tissue based on its descriptor value. The tissue transitions are modeled using a Markov chain to estimate the likelihood of a candidate lesion region. Experimental evaluation on a clinical dataset of 135 images show that the proposed approach can achieve high sensitivity (95 %) with modest (3) false positives per image. The approach achieves very similar results (94 % for 3 false positives) on a completely different clinical dataset of 159 images without retraining, highlighting the robustness of the approach.
Soma Biswas, Xiaoxing Li, Rakesh Mullick, Vivek Vaidya 0001
ICPR1
2013 Pose-Robust Recognition of Low-Resolution Face Images
abstract
Face images captured by surveillance cameras usually have poor resolution in addition to uncontrolled poses and illumination conditions, all of which adversely affect the performance of face matching algorithms. In this paper, we develop a completely automatic, novel approach for matching surveillance quality facial images to high-resolution images in frontal pose, which are often available during enrollment. The proposed approach uses multidimensional scaling to simultaneously transform the features from the poor quality probe images and the high-quality gallery images in such a manner that the distances between them approximate the distances had the probe images been captured in the same conditions as the gallery images. Tensor analysis is used for facial landmark localization in the low-resolution uncontrolled probe images for computing the features. Thorough evaluation on the Multi-PIE dataset and comparisons with state-of-the-art super-resolution and classifier-based approaches are performed to illustrate the usefulness of the proposed approach. Experiments on surveillance imagery further signify the applicability of the framework. We also show the usefulness of the proposed approach for the application of tracking and recognition in surveillance videos.
Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn, Kevin W. Bowyer
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Digital Curvatures Applied to 3D Object Analysis and Recognition: A Case Study
Li Chen 0002, Soma Biswas
IWCIA2
2012 A sparse representation approach to face matching across plastic surgery
abstract
Plastic surgery procedures can significantly alter facial appearance, thereby posing a serious challenge even to the state-of-the-art face matching algorithms. In this paper, we propose a novel approach to address the challenges involved in automatic matching of faces across plastic surgery variations. In the proposed formulation, part-wise facial characterization is combined with the recently popular sparse representation approach to address these challenges. The sparse representation approach requires several images per subject in the gallery to function effectively which is often not available in several use-cases, as in the problem we address in this work. The proposed formulation utilizes images from sequestered non-gallery subjects with similar local facial characteristics to fulfill this requirement. Extensive experiments conducted on a recently introduced plastic surgery database [17] consisting of 900 subjects highlight the effectiveness of the proposed approach.
Gaurav Aggarwal, Soma Biswas, Patrick J. Flynn, Kevin W. Bowyer
WACV2
2012 Predicting good, bad and ugly match Pairs
abstract
Several sources of variation in facial appearance that affect face matching performance have long been investigated. The recently introduced GBU challenge problem [1] indicates that there can be significant variation in performance across different partitions of the data, even when the impact of most known factors is eliminated or significantly reduced by the data collection and experimentation protocol. The GBU challenge problem consists of three partitions which are called the Good (easy to match), the Bad (average matching difficulty) and the Ugly (difficult to match). In this paper, we investigate various image and facial characteristics that can account for the observed significant difference in performance across these partitions. Given a match pair, we aim to predict the partition it belongs to. Partial Least Squares (PLS)-based regression is used to perform the prediction task. Our analysis indicates that the match pairs from the three partitions differ from each other in terms of simple but often ignored factors like image sharpness, hue, saturation and extent of facial expressions.
Gaurav Aggarwal, Soma Biswas, Patrick J. Flynn, Kevin W. Bowyer
WACV2
2012 Face Recognition from Video: a Review
abstract
Driven by key law enforcement and commercial applications, research on face recognition from video sources has intensified in recent years. The ensuing results have demonstrated that videos possess unique properties that allow both humans and automated systems to perform recognition accurately in difficult viewing conditions. However, significant research challenges remain as most video-based applications do not allow for controlled recordings. In this survey, we categorize the research in this area and present a broad and deep review of recently proposed methods for overcoming the difficulties encountered in unconstrained settings. We also draw connections between the ways in which humans and current algorithms recognize faces. An overview of the most popular and difficult publicly available face video databases is provided to complement these discussions. Finally, we cover key research challenges and opportunities that lie ahead for the field as a whole.
Jeremiah R. Barr, Kevin W. Bowyer, Patrick J. Flynn, Soma Biswas
Int. J. Pattern Recognit. Artif. Intell.4
2012 Multidimensional Scaling for Matching Low-Resolution Face Images
abstract
Face recognition performance degrades considerably when the input images are of Low Resolution (LR), as is often the case for images taken by surveillance cameras or from a large distance. In this paper, we propose a novel approach for matching low-resolution probe images with higher resolution gallery images, which are often available during enrollment, using Multidimensional Scaling (MDS). The ideal scenario is when both the probe and gallery images are of high enough resolution to discriminate across different subjects. The proposed method simultaneously embeds the low-resolution probe images and the high-resolution gallery images in a common space such that the distance between them in the transformed space approximates the distance had both the images been of high resolution. The two mappings are learned simultaneously from high-resolution training images using an iterative majorization algorithm. Extensive evaluation of the proposed approach on the Multi-PIE data set with probe image resolution as low as 8 6 pixels illustrates the usefulness of the method. We show that the proposed approach improves the matching performance significantly as compared to performing matching in the low-resolution domain or using super-resolution techniques to obtain a higher resolution test image prior to recognition. Experiments on low-resolution surveillance images from the Surveillance Cameras Face Database further highlight the effectiveness of the approach.
Soma Biswas, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Dictionary-Based Face Recognition Under Variable Lighting and Pose
abstract
We present a face recognition algorithm based on simultaneous sparse approximations under varying illumination and pose. A dictionary is learned for each class based on given training examples which minimizes the representation error with a sparseness constraint. A novel test image is projected onto the span of the atoms in each learned dictionary. The resulting residual vectors are then used for classification. To handle variations in lighting conditions and pose, an image relighting technique based on pose-robust albedo estimation is used to generate multiple frontal images of the same person with variable lighting. As a result, the proposed algorithm has the ability to recognize human faces with high accuracy even when only a single or a very few images per person are provided for training. The efficiency of the proposed method is demonstrated using publicly available databases available databases and it is shown that this method is efficient and can perform significantly better than many competitive face recognition algorithms.
Vishal M. Patel, Tao Wu 0009, Soma Biswas, P. Jonathon Phillips, Rama Chellappa
IEEE Trans. Inf. Forensics Secur.3
2011 Pose-robust recognition of low-resolution face images
abstract
Face images captured by surveillance cameras usually have poor resolution in addition to uncontrolled poses and illumination conditions which adversely affect performance of face matching algorithms. In this paper, we develop a novel approach for matching surveillance quality facial images to high resolution images in frontal pose which are often available during enrollment. The proposed approach uses Multidimensional Scaling to simultaneously transform the features from the poor quality probe images and the high quality gallery images in such a manner that the distances between them approximate the distances had the probe images been captured in the same conditions as the gallery images. Thorough evaluation on the Multi-PIE dataset and comparisons with state-of-the-art super-resolution and classifier based approaches are performed to illustrate the usefulness of the proposed approach. Experiments on real surveillance images further signify the applicability of the framework.
Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn
CVPR1
2011 Face recognition in low-resolution videos using learning-based likelihood measurement model
abstract
Low-resolution surveillance videos with uncontrolled pose and illumination present a significant challenge to both face tracking and recognition algorithms. Considerable appearance difference between the probe videos and high-resolution controlled images in the gallery acquired during enrollment makes the problem even harden In this paper, we extend the simultaneous tracking and recognition framework [22] to address the problem of matching high-resolution gallery images with surveillance quality probe videos. We propose using a learning-based likelihood measurement model to handle the large appearance and resolution difference between the gallery images and probe videos. The measurement model consists of a mapping which transforms the gallery and probe features to a space in which their inter-Euclidean distances approximate the distances that would have been obtained had all the descriptors been computed from good quality frontal images. Experimental results on real surveillance quality videos and comparisons with related approaches show the effectiveness of the proposed framework.
Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn
IJCB1
2011 Difficult imaging covariates or difficult subjects? - An empirical investigation
abstract
The performance of face recognition algorithms is affected both by external factors and internal subject characteristics [I]. Reliably identifying these factors and understanding their behavior on performance can potentially serve two important goals to predict the performance of the algorithms at novel deployment sites and to design appropriate acquisition environments at prospective sites to optimize performance. There have been a few recent efforts in this direction that focus on identifying factors that affect face recognition performance but there has been no extensive study regarding the consistency of the effects various factors have on algorithms when other covariates vary. To give an example, a smiling target image has been reported to be better than a neutral expression image, but is this true across all possible illumination conditions, head poses, gender, etc.? In this paper, we perform rigorous experiments to provide answers to such questions. Our investigation indicates that controlled lighting and smiling expression are the most favorable conditions that consistently give superior performance even when other factors are allowed to vary. We also observe that internal subject characterization using biometric menagerie-based classification shows very weak consistency when external conditions are allowed to vary.
Jeffrey R. Paone, Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn
IJCB2
2011 Illumination robust dictionary-based face recognition
abstract
In this paper, we present a face recognition method based on simultaneous sparse approximations under varying illumination. Our method consists of two main stages. In the first stage, a dictionary is learned for each face class based on given training examples which minimizes the representation error with a sparseness constraint. In the second stage, a test image is projected onto the span of the atoms in each learned dictionary. The resulting residual vectors are then used for classification. Furthermore, to handle changes in lighting conditions, we use a relighting approach based on a non-stationary stochastic filter to generate multiple images of the same person with different lighting. As a result, our algorithm has the ability to recognize human faces with good accuracy even when only a single or a very few images are provided for training. The effectiveness of the proposed method is demonstrated on publicly available databases and it is shown that this method is efficient and can perform significantly better than many competitive face recognition algorithms.
Vishal M. Patel, Tao Wu 0009, Soma Biswas, P. Jonathon Phillips, Rama Chellappa
ICIP3
2010 Pose-robust albedo estimation from a single image
abstract
We present a stochastic filtering approach to perform albedo estimation from a single non-frontal face image. Albedo estimation has far reaching applications in various computer vision tasks like illumination-insensitive matching, shape recovery, etc. We extend the formulation proposed in that assumes face in known pose and present an algorithm that can perform albedo estimation from a single image even when pose information is inaccurate. 3D pose of the input face image is obtained as a byproduct of the algorithm. The proposed approach utilizes class-specific statistics of faces to iteratively improve albedo and pose estimates. Illustrations and experimental results are provided to show the effectiveness of the approach. We highlight the usefulness of the method for the task of matching faces across variations in pose and illumination. The facial pose estimates obtained are also compared against ground truth.
Soma Biswas, Rama Chellappa
CVPR1
2010 The role of geometry in age estimation
abstract
Understanding and modeling of aging in human faces is an important problem in many real-world applications such as biometrics, authentication, and synthesis. In this paper, we consider the role of geometric attributes of faces, as described by a set of landmark points on the face, in age perception. Towards this end, we show that the space of landmarks can be interpreted as a Grassmann manifold. Then the problem of age estimation is posed as a problem of function estimation on the manifold. The warping of an average face to a given face is quantified as a velocity vector that transforms the average to a given face along a smooth geodesic in unit-time. This deformation is then shown to contain important information about the age of the face. We show in experiments that exploiting geometric cues in a principled manner provides comparable performance to several systems that utilize both geometric and textural cues. We show results on age estimation using the standard FG-Net dataset and a passport dataset which illustrate the effectiveness of the approach.
Pavan Turaga, Soma Biswas, Rama Chellappa
ICASSP2
2010 An Efficient and Robust Algorithm for Shape Indexing and Retrieval
abstract
Many shape matching methods are either fast but too simplistic to give the desired performance or promising as far as performance is concerned but computationally demanding. In this paper, we present a very simple and efficient approach that not only performs almost as good as many state-of-the-art techniques but also scales up to large databases. In the proposed approach, each shape is indexed based on a variety of simple and easily computable features which are invariant to articulations, rigid transformations, etc. The features characterize pairwise geometric relationships between interest points on the shape. The fact that each shape is represented using a number of distributed features instead of a single global feature that captures the shape in its entirety provides robustness to the approach. Shapes in the database are ordered according to their similarity with the query shape and similar shapes are retrieved using an efficient scheme which does not involve costly operations like shape-wise alignment or establishing correspondences. Depending on the application, the approach can be used directly for matching or as a first step for obtaining a short list of candidate shapes for more rigorous matching. We show that the features proposed to perform shape indexing can be used to perform the rigorous matching as well, to further improve the retrieval performance.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
IEEE Trans. Multim.1
2009 Robust Estimation of Albedo for Illumination-Invariant Matching and Shape Recovery
abstract
We present a nonstationary stochastic filtering framework for the task of albedo estimation from a single image. There are several approaches in the literature for albedo estimation, but few include the errors in estimates of surface normals and light source direction to improve the albedo estimate. The proposed approach effectively utilizes the error statistics of surface normals and illumination direction for robust estimation of albedo, for images illuminated by single and multiple light sources. The albedo estimate obtained is subsequently used to generate albedo-free normalized images for recovering the shape of an object. Traditional Shape-from-Shading (SFS) approaches often assume constant/piecewise constant albedo and known light source direction to recover the underlying shape. Using the estimated albedo, the general problem of estimating the shape of an object with varying albedo map and unknown illumination source is reduced to one that can be handled by traditional SFS approaches. Experimental results are provided to show the effectiveness of the approach and its application to illumination-invariant matching and shape recovery. The estimated albedo maps are compared with the ground truth. The maps are used as illumination-invariant signatures for the task of face recognition across illumination variations. The recognition results obtained compare well with the current state-of-the-art approaches. Impressive shape recovery results are obtained using images downloaded from the Web with little control over imaging conditions. The recovered shapes are also used to synthesize novel views under novel illumination conditions.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Symmetric Objects are Hardly Ambiguous
abstract
Given any two images taken under different illumination conditions, there always exist a physically realizable object which is consistent with both the images even if the lighting in each scene is constrained to be a known point light source at infinity. In this paper, we show that images are much less ambiguous for the class of bilaterally symmetric Lambertian objects. In fact, the set of such objects can be partitioned into equivalence classes such that it is always possible to distinguish between two objects belonging to different equivalence classes using just one image per object. The conditions required for two objects to belong to the same equivalence class are very restrictive, thereby leading to the conclusion that images of symmetric objects are hardly ambiguous. The observation leads to an illumination-invariant matching algorithm to compare images of bilaterally symmetric Lambertian objects. Experiments on real data are performed to show the implications of the theoretical result even when the symmetry and Lambertian assumptions are not strictly satisfied.
Gaurav Aggarwal, Soma Biswas, Rama Chellappa
CVPR2
2007 Efficient Indexing For Articulation Invariant Shape Matching And Retrieval
abstract
Most shape matching methods are either fast but too simplistic to give the desired performance or promising as far as performance is concerned but computationally demanding. In this paper, we present a very simple and efficient approach that not only performs almost as good as many state-of-the-art techniques but also scales up to large databases. In the proposed approach, each shape is indexed based on a variety of simple and easily computable features which are invariant to articulations and rigid transformations. The features characterize pairwise geometric relationships between interest points on the shape, thereby providing robustness to the approach. Shapes are retrieved using an efficient scheme which does not involve costly operations like shape-wise alignment or establishing correspondences. Even for a moderate size database of 1000 shapes, the retrieval process is several times faster than most techniques with similar performance. Extensive experimental results are presented to illustrate the advantages of our approach as compared to the best in the field.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
CVPR1
2007 Robust Estimation of Albedo for Illumination-invariant Matching and Shape Recovery
abstract
In this paper, we propose a non-stationary stochastic filtering framework for the task of albedo estimation from a single image. There are several approaches in literature for albedo estimation, but few include the errors in estimates of surface normals and light source directions to improve the albedo estimate. The proposed approach effectively utilizes the error statistics of surface normals and illumination direction for robust estimation of albedo. The albedo estimate obtained is further used to generate albedo-free normalized images for recovering the shape of an object. Illustrations and experiments are provided to show the efficacy of the approach and its application to illumination-invariant matching and shape recovery.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
ICCV1
2006 Invariant Geometric Representation of 3D Point Clouds for Registration and Matching
abstract
Though implicit representations of surfaces have often been used for various computer graphics tasks like modeling and morphing of objects, it has rarely been used for registration and matching of 3D point clouds. Unlike in graphics, where the goal is precise reconstruction, we use isosurfaces to derive a smooth and approximate representation of the underlying point cloud which helps in generalization. Implicit surfaces are generated using a variational interpolation technique. Implicit function values on a set of concentric spheres around the 3D point cloud of object are used as features for matching. Geometric-invariance is achieved by decomposing implicit values based feature set into various spherical harmonics. The decomposition provides a compact representation of 3D point clouds while achieving rotation invariance.
Soma Biswas, Gaurav Aggarwal, Rama Chellappa
ICIP1