VLDB 2026 Research / reviewers in the wild / expert
Anoop M. Namboodiri
dblp:16/143
· DBLP profile ↗
51ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-4638-0833ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 10 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-authorSecurity and privacy · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fusion2Print: Deep Flash-NoFlash Fusion for Contactless Fingerprint Matching
Roja Sahoo, Anoop M. Namboodiri |
ICPR (6) | 2 |
| 2024 | Image Attribution by Generating ImagesabstractWe introduce GPNN-CAM, a novel method for CNN explanation, that bridges two distinct areas of computer vision: Image Attribution, which aims to explain a predictor by highlighting image regions it finds important, and Single Image Generation (SIG), that focuses on learning how to generate variations of a single sample.GPNN-CAM leverages samples generated by Generative Patch Nearest Neighbors (GPNN) into a Class Activation Map (CAM) flavored attribution scheme. Our findings reveal that the incorporation of these samples yields remarkably effective results, enabling GPNN-CAM to demonstrate superior performance across multiple classifier architectures, and datasets. Aniket Singh, Anoop M. Namboodiri |
ICASSP | 2 |
| 2024 | CLIP4Sketch: Enhancing Sketch to Mugshot Matching through Dataset Augmentation using Diffusion ModelsabstractForensic sketch-to-mugshot matching is a challenging task in face recognition, primarily hindered by the scarcity of annotated forensic sketches and the modality gap between sketches and photographs. To address this, we propose CLIP4Sketch, a novel approach that leverages diffusion models to generate a large and diverse set of sketch images, which helps in enhancing the performance of face recognition systems in sketch-to-mugshot matching. Our method utilizes Denoising Diffusion Probabilistic Models (DDPMs) to generate sketches with explicit control over identity and style. We combine CLIP and Adaface embeddings of a reference mugshot, along with textual descriptions of style, as the conditions to the diffusion model. We demonstrate the efficacy of our approach by generating a comprehensive dataset of sketches corresponding to mugshots and training a face recognition model on our synthetic data. Our results show significant improvements in sketch-to-mugshot matching accuracy over training on an existing, limited amount of real face sketch data, validating the potential of diffusion models in enhancing the performance of face recognition systems across modalities. We also compare our dataset with datasets generated using GAN-based methods to show its superiority. Kushal Kumar Jain, Steven A. Grosz, Anoop M. Namboodiri, Anil K. Jain 0001 |
IJCB | 3 |
| 2024 | Saliency As A Schedule: Intuitive Image AttributionabstractImage Attribution seeks to reveal the importance of image regions in the classifier’s final decision. Of the various ways to tackle this problem, the optimization based perspective is particularly intuitive: It applies the attribution as a mask on the image and reduces the attribution task to a loss, that can be optimized using gradient descent. Previous work has considered the goal as searching for the single best mask. Under this setup, however, there is a tendency towards trivial solutions of large masks with reduced discernment of the relative importances of regions. This has typically required auxiliary loss terms to control the area of the mask, however, their strength relative to the primary loss needs to be tuned. We challenge this necessity, by re-imagining attribution as an ordering of pixels according to importance. This ordering may be interpreted as a schedule which determines which locations get seen earlier and which later, allowing us to create a trajectory of masks from completely OFF to completely ON. We optimize through this sequence of masks of “all” areas and not just a single mask as in previous methods. We explore this setting, which we dub Saliency-as-Schedule (SaS), and demonstrate its effectiveness through experiments in a variety of settings, involving multiple datasets and CNN architectures. Further, we also propose a novel attribution task, feature saliency, where we use SaS to rate the influence of image regions on the intermediate feature maps of a CNN, and not just the class logit. Our findings suggest that SaS is a promising direction for the attribution problem. Our code will be available at https://github.com/tumbleweed/SaliencyasSchedule Aniket Singh, Anoop M. Namboodiri |
ICIP | 2 |
| 2024 | Using Multiscale Information for Improved Optimization-Based Image Attribution
Aniket Singh, Anoop M. Namboodiri |
ICPR (3) | 2 |
| 2023 | AdvGen: Physical Adversarial Attack on Face Presentation Attack Detection SystemsabstractEvaluating the risk level of adversarial images is essential for safely deploying face authentication models in the real world. Popular approaches for physical-world attacks, such as print or replay attacks, suffer from some limitations, like including physical and geometrical artifacts. Recently adversarial attacks have gained attraction, which try to digitally deceive the learning strategy of a recognition system using slight modifications to the captured image. While most previous research assumes that the adversarial image could be digitally fed into the authentication systems, this is not always the case for systems deployed in the real world. This paper demonstrates the vulnerability of face authentication systems to adversarial images in physical world scenarios. We propose AdvGen, an automated Generative Adversarial Network, to simulate print and replay attacks and generate adversarial images that can fool state-of-the-art PADs in a physical domain attack setting. Using this attack strategy, the attack success rate goes up to 82.01%. We test AdvGen extensively on four datasets and ten state-of-the-art PADs. We also demonstrate the effectiveness of our attack by conducting experiments in a realistic, physical environment. Sai Amrit Patnaik, Shivali Chansoriya, Anoop M. Namboodiri, Anil K. Jain 0001 |
IJCB | 3 |
| 2022 | One-Shot Sensor and Material Translator : A Bilinear Decomposer for Fingerprint Presentation Attack GeneralizationabstractAutomatic fingerprint recognition systems are currently under the constant threat of presentation attacks (PAs). Existing fingerprint presentation attack detection (FPAD) solutions improve cross-sensor and cross-material generalization by utilizing style-transfer-based augmentation wrappers over a two-class PAD classifier. These solutions synthesize data by learning the style as a single entity, containing both sensor and material characteristics. However, these strategies necessitate learning the entire style upon adding a new sensor for an already known material or vice versa. We propose a bilinear decomposition-based wrapper called OSMT to improve cross-sensor and cross-material FPAD. OSMT uses one PA fingerprint to learn the corresponding sensor and material representations by disentanglement. Our approach also reduces the computational complexity by generating compact representations and utilizing lesser combinations of sensors and materials to produce several styles. We present the improvement in PAD performance using our technique on the publicly available LivDet datasets (2015, 2017, 2019 and 2021). Gowri Lekshmy, Anoop M. Namboodiri |
IJCB | 2 |
| 2022 | SIAN: Secure Iris Authentication using NoiseabstractBiometric noise is often discarded in many biometric template protection systems. However, the noise ratio between two templates encodes specific correlational properties that template protection schemes can exploit. Biometric authentication usually occurs between mutually distrusting parties, which calls for privacy-preserving techniques. In this paper, we propose a novel biometric authentication protocol, SIAN(Secure Iris Authentication using Noise), adapting secure two-party computation and incorporating uncertainty constraints from biometric noise for security. We evaluate it on three iris datasets: MMU v1, Ubiris v1, and IITD v1, and observe a low EER degradation. The proposed protocol has information-theoretic security and low computational complexity, making it suitable for practical real-time applications. Praguna Manvi, Achintya Desai, K. Srinathan 0001, Anoop M. Namboodiri |
IJCB | 4 |
| 2022 | Transformer based Fingerprint Feature ExtractionabstractFingerprint feature extraction is a task that is solved using either a global or a local representation. State-of-the-art global approaches use heavy deep learning models to process the full fingerprint image at once, which makes the corresponding approach memory intensive. On the other hand, local approaches involve minutiae based patch extraction, multiple feature extraction steps and an expensive matching stage, which make the corresponding approach time intensive. However, both these approaches provide useful and sometimes exclusive insights for solving the problem. Using both approaches together for extracting fingerprint representations is semantically useful but quite inefficient. Our convolutional transformer based approach with an in-built minutiae extractor provides a time and memory efficient solution to extract a global as well as a local representation of the fingerprint. The use of these representations along with a smart matching process gives us state-of-the-art performance across multiple databases. The project page can be found at https://saraansh1999.github.io/global-plus-local-fp-transformer. Saraansh Tandon, Anoop M. Namboodiri |
ICPR | 2 |
| 2021 | A Unified Model for Fingerprint Authentication and Presentation Attack DetectionabstractTypical fingerprint recognition systems are comprised of a spoof detection module and a subsequent recognition module, running one after the other. In this paper, we reformulate the workings of a typical fingerprint recognition system. In particular, we posit that both spoof detection and fingerprint recognition are correlated tasks. Therefore, rather than performing the two tasks separately, we propose a joint model for spoof detection and matching1to simultaneously perform both tasks without compromising the accuracy of either task. We demonstrate the capability of our joint model to obtain an authentication accuracy (1:1 matching) of TAR = 100% @ FAR = 0.1% on the FVC 2006 DB2A dataset while achieving a spoof detection ACE of 1.44% on the LiveDet 2015 dataset, both maintaining the performance of stand-alone methods. In practice, this reduces the time and memory requirements of the fingerprint recognition system by 50% and 40%, respectively; a significant advantage for recognition systems running on resource-constrained devices and communication channels. Additya Popli, Saraansh Tandon, Joshua J. Engelsma, Naoyuki Onoe, Atsushi Okubo, Anoop M. Namboodiri |
IJCB | 6 |
| 2020 | STQ-Nets: Unifying Network Binarization and Structured Pruning
Sri Aurobindo Munagala, Ameya Prabhu, Anoop M. Namboodiri |
BMVC | 3 |
| 2020 | Mutual Information based Method for Unsupervised Disentanglement of Video RepresentationabstractVideo Prediction is a challenging but interesting task of predicting future frames from a given set context frames that belong to a video sequence. Video prediction models have prospective applications in maneuver planning, healthcare, autonomous navigation and simulation. One of the major challenges in future frame generation is the high dimensional nature of visual data. To handle this, we propose a Mutual Information Predictive Auto-Encoder (MIPAE) framework that reduces the task of predicting high dimensional video frames by factorising video representations into content and low dimensional pose latent variables. Our approach leverages the temporal structure in the latent generative factors of video sequences by applying a novel mutual information loss to learn disentangled video representations. A standard LSTM network is used to predict these low dimensional pose representations. Content and the predicted pose representations are decoded to generate future frames. We also propose a metric based on mutual information gap (MIG) to quantitatively access the effectiveness of disentanglement on DSprites and MPI3D-real datasets. MIG scores corroborate the visual superiority of frames predicted by MIPAE. We also compare our method quantitatively on LPIPS, SSIM and PSNR evaluation metrics. P. Aditya Sreekar, Ujjwal Tiwari, Anoop M. Namboodiri |
ICPR | 3 |
| 2020 | Reducing the Variance of Variational Estimates of Mutual Information by Limiting the Critic's Hypothesis Space to RKHSabstractMutual information (MI) is an information-theoretic measure of dependency between two random variables. Several methods to estimate MI from samples of two random variables with unknown underlying probability distributions have been proposed in the literature. Recent methods realize parametric probability distributions or critic as a neural network to approximate unknown density ratios. These approximated density ratios are used to estimate different variational lower bounds of MI. While, these estimation methods are reliable when the true MI is low, they tend to produce high variance estimates when the true MI is high. We argue that the high variance characteristics is due to the uncontrolled complexity of the critic's hypothesis space. In support of this argument, we use the data-driven Rademacher complexity of the hypothesis space associated with the critic's architecture to analyse generalization error bound of variational lower bound estimates of MI. In the proposed work, we show that it is possible to negate the high variance characteristics of these estimators by constraining the critic's hypothesis space to Reproducing Hilbert Kernel Space (RKHS), which corresponds to a kernel learned using Automated Spectral Kernel Learning (ASKL). By analysing the generalization error bounds, we augment the overall optimisation objective with effective regularisation term. We empirically demonstrate the efficacy of this regularization in enforcing proper bias variance tradeoff on four different variational lower bounds of MI, namely NWJ, MINE, JS and SMILE. P. Aditya Sreekar, Ujjwal Tiwari, Anoop M. Namboodiri |
ICPR | 3 |
| 2020 | Understanding Dynamic Scenes using Graph Convolution NetworksabstractWe present a novel Multi-Relational Graph Convolutional Network (MRGCN) based framework to model on-road vehicle behaviors from a sequence of temporally ordered frames as grabbed by a moving monocular camera. The input to MRGCN is a multi-relational graph where the graph's nodes represent the active and passive agents/objects in the scene, and the bidirectional edges that connect every pair of nodes are encodings of their Spatio-temporal relations.We show that this proposed explicit encoding and usage of an intermediate spatio-temporal interaction graph to be well suited for our tasks over learning end-end directly on a set of temporally ordered spatial relations. We also propose an attention mechanism for MRGCNs that conditioned on the scene dynamically scores the importance of information from different interaction types.The proposed framework achieves significant performance gain over prior methods on vehicle-behavior classification tasks on four datasets. We also show a seamless transfer of learning to multiple datasets without resorting to fine-tuning. Such behavior prediction methods find immediate relevance in a variety of navigation tasks such as behavior planning, state estimation, and applications relating to the detection of traffic violations over videos. Sravan Mylavarapu, Mahtab Sandhu, Priyesh Vijayan, K. Madhava Krishna, Balaraman Ravindran, Anoop M. Namboodiri |
IROS | 6 |
| 2020 | Towards Accurate Vehicle Behaviour Classification With Multi-Relational Graph Convolutional NetworksabstractUnderstanding on-road vehicle behaviour from a temporal sequence of sensor data is gaining in popularity. In this paper, we propose a pipeline for understanding vehicle behaviour from a monocular image sequence or video. A monocular sequence along with scene semantics, optical flow and object labels are used to get spatial information about the object (vehicle) of interest and other objects (semantically contiguous set of locations) in the scene. This spatial information is encoded by a Multi-Relational Graph Convolutional Network (MR-GCN), and a temporal sequence of such encodings is fed to a recurrent network to label vehicle behaviours. The proposed framework can classify a variety of vehicle behaviours to high fidelity on datasets that are diverse and include European, Chinese and Indian on-road scenes. The framework also provides for seamless transfer of models across datasets without entailing re-annotation, retraining and even fine-tuning. We show comparative performance gain over baseline Spatio-temporal classifiers and detail a variety of ablations to showcase the efficacy of the framework. Sravan Mylavarapu, Mahtab Sandhu, Priyesh Vijayan, K. Madhava Krishna, Balaraman Ravindran, Anoop M. Namboodiri |
IV | 6 |
| 2020 | Region Pooling with Adaptive Feature Fusion for End-to-End Person RecognitionabstractCurrent approaches for person recognition train an ensemble of region specific convolutional neural networks for representation learning, and then adopt naive fusion strategies to combine their features or predictions during testing. In this paper, we propose an unified end-to-end architecture that generates a complete person representation based on pooling and aggregation of features from multiple body regions. Our network takes a person image and the predetermined locations of body regions as input, and generates common feature maps that are shared across all the regions. Multiple features corresponding to different regions are then pooled and combined with an aggregation block, where the adaptive weights required for aggregation are obtained through an attention mechanism. Evaluations on three person recognition datasets - PIPA, Soccer and Hannah show that a single model trained end-to-end is computationally faster, requires fewer parameters and achieves improved performance over separately trained models. Vijay Kumar 0004, Anoop M. Namboodiri, C. V. Jawahar |
WACV | 2 |
| 2019 | IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained EnvironmentsabstractWhile several datasets for autonomous navigation have become available in recent years, they have tended to focus on structured driving environments. This usually corresponds to well-delineated infrastructure such as lanes, a small number of well-defined categories for traffic participants, low variation in object or background appearance and strong adherence to traffic rules. We propose DS, a novel dataset for road scene understanding in unstructured environments where the above assumptions are largely not satisfied. It consists of 10,004 images, finely annotated with 34 classes collected from 182 drive sequences on Indian roads. The label set is expanded in comparison to popular benchmarks such as Cityscapes, to account for new classes. It also reflects label distributions of road scenes significantly different from existing datasets, with most classes displaying greater within-class diversity. Consistent with real driving behaviors, it also identifies new classes such as drivable areas besides the road. We propose a new four-level label hierarchy, which allows varying degrees of complexity and opens up possibilities for new training methods. Our empirical study provides an in-depth analysis of the label characteristics. State-of-the-art methods for semantic segmentation achieve much lower accuracies on our dataset, demonstrating its distinction compared to Cityscapes. Finally, we propose that our dataset is an ideal opportunity for new problems such as domain adaptation, few-shot learning and behavior prediction in road scenes. Girish Varma, Anbumani Subramanian, Anoop M. Namboodiri, Manmohan Krishna Chandraker, C. V. Jawahar |
WACV | 3 |
| 2018 | Deep Expander Networks: Efficient Deep Networks from Graph Theory
Ameya Prabhu, Girish Varma, Anoop M. Namboodiri |
ECCV (13) | 3 |
| 2018 | Cross-Modal Style TransferabstractWe, humans, have the ability to easily imagine scenes that depict sentences such as “Today is a beautiful sunny day” or “There is a Christmas feel, in the air”. While it is hard to precisely describe what one person may imagine, the essential high-level themes associated with such sentences largely remains the same. The ability to synthesize novel images that depict the feel of a sentence is very useful in a variety of applications such as education, advertisement, and entertainment. While existing papers tackle this problem given a style image, we aim to provide a far more intuitive and easy to use solution that synthesizes novel renditions of an existing image, conditioned on a given sentence. We present a method for cross-modal style transfer between an English sentence and an image, to produce a new image that imbibes the essential theme of the sentence. We do this by modifying the style transfer mechanism used in image style transfer to incorporate a style component derived from the given sentence. We demonstrate promising results using the YFCC100m dataset. Sahil Chelaramani, Abhishek Jha 0001, Anoop M. Namboodiri |
ICIP | 3 |
| 2018 | Hybrid Binary Networks: Optimizing for Accuracy, Efficiency and MemoryabstractBinarization is an extreme network compression approach that provides large computational speedups along with energy and memory savings, albeit at significant accuracy costs. We investigate the question of where to binarize inputs at layer-level granularity and show that selectively binarizing the inputs to specific layers in the network could lead to significant improvements in accuracy while preserving most of the advantages of binarization. We analyze the binarization tradeoff using a metric that jointly models the input binarization-error and computational cost and introduce an efficient algorithm to select layers whose inputs are to be binarized. Practical guidelines based on insights obtained from applying the algorithm to a variety of models are discussed. Experiments on Imagenet dataset using AlexNet and ResNet-18 models show 3-4% improvements in accuracy over fully binarized networks with minimal impact on compression and computational speed. The improvements are even more substantial on sketch datasets like TU-Berlin, where we match state-of-the-art accuracy as well, getting over 8% increase in accuracies. We further show that our approach can be applied in tandem with other forms of compression that deal with individual layers or overall model compression (e.g., SqueezeNets). Unlike previous quantization approaches, we are able to binarize the weights in the last layers of a network, which often have a large number of parameters, resulting in significant improvement in accuracy over fully binarized models. Ameya Prabhu, Vishal Batchu, Rohit Gajawada, Sri Aurobindo Munagala, Anoop M. Namboodiri |
WACV | 5 |
| 2018 | Distribution-Aware Binarization of Neural Networks for Sketch RecognitionabstractDeep neural networks are highly effective at a range of computational tasks. However, they tend to be computationally expensive, especially in vision-related problems, and also have large memory requirements. One of the most effective methods to achieve significant improvements in computational/spatial efficiency is to binarize the weights and activations in a network. However, naive binarization results in accuracy drops when applied to networks for most tasks. In this work, we present a highly generalized, distribution-aware approach to binarizing deep networks that allows us to retain the advantages of a binarized network, while reducing accuracy drops. We also develop efficient implementations for our proposed approach across different architectures. We present a theoretical analysis of the technique to show the effective representational power of the resulting layers, and explore the forms of data they model best. Experiments on popular datasets show that our technique offers better accuracies than naive binarization, while retaining the same benefits that binarization provides - with respect to run-time compression, reduction of computational costs, and power consumption. Ameya Prabhu, Vishal Batchu, Sri Aurobindo Munagala, Rohit Gajawada, Anoop M. Namboodiri |
WACV | 5 |
| 2017 | Pose-Aware Person RecognitionabstractPerson recognition methods that use multiple body regions have shown significant improvements over traditional face-based recognition. One of the primary challenges in full-body person recognition is the extreme variation in pose and view point. In this work, (i) we present an approach that tackles pose variations utilizing multiple models that are trained on specific poses, and combined using pose-aware weights during testing. (ii) For learning a person representation, we propose a network that jointly optimizes a single loss over multiple body regions. (iii) Finally, we introduce new benchmarks to evaluate person recognition in diverse scenarios and show significant improvements over previously proposed approaches on all the benchmarks including the photo album setting of PIPA. Vijay Kumar 0004, Anoop M. Namboodiri, Manohar Paluri, C. V. Jawahar |
CVPR | 2 |
| 2017 | Learning deep and compact models for gesture recognitionabstractWe look at the problem of developing a compact and accurate model for gesture recognition from videos in a deep-learning framework. Towards this we propose a joint 3DCNN-LSTM model that is end-to-end trainable and is shown to be better suited to capture the dynamic information in actions. The solution achieves close to state-of-the-art accuracy on the ChaLearn dataset, with only half the model size. We also explore ways to derive a much more compact representation in a knowledge distillation framework followed by model compression. The final model is less than 1 MB in size, which is less than one hundredth of our initial model, with a drop of 7% in accuracy, and is suitable for real-time gesture recognition on mobile devices. Koustav Mullick, Anoop M. Namboodiri |
ICIP | 2 |
| 2016 | Panoramic Stereo Videos with a Single CameraabstractWe present a practical solution for generating 360° stereo panoramic videos using a single camera. Current approaches either use a moving camera that captures multiple images of a scene, which are then stitched together to form the final panorama, or use multiple cameras that are synchronized. A moving camera limits the solution to static scenes, while multi-camera solutions require dedicated calibrated setups. Our approach improves upon the existing solutions in two significant ways: It solves the problem using a single camera, thus minimizing the calibration problem and providing us the ability to convert any digital camera into a panoramic stereo capture device. It captures all the light rays required for stereo panoramas in a single frame using a compact custom designed mirror, thus making the design practical to manufacture and easier to use. We analyze several properties of the design as well as present panoramic stereo and depth estimation results. Rajat Aggarwal, Amrisha Vohra, Anoop M. Namboodiri |
CVPR | 3 |
| 2016 | Leveraging multiple tasks to regularize fine-grained classificationabstractFine-grained classification is an extremely challenging problem in computer vision, compounded by subtle differences in shape, pose, illumination and appearance. While convolutional neural networks have become the versatile jack-of-all-trades tool in modern computer vision, approaches for fine-grained recognition still rely on localization of keypoints and parts to learn discriminative features for recognition. In order to achieve this, most approaches use a localization module and subsequently learn classifiers for the inferred locations, thus necessitating large amounts of manual annotations for bounding boxes and keypoints. In order to tackle this problem, we aim to leverage the (taxonomic and/or semantic) relationships present among fine-grained classes. The ontology tree is a free source of labels that can be used as auxiliary tasks to train a multi-task loss. Additional tasks can act as regularizers, and increase the generalization capabilities of the network. Multiple tasks try to take the network in diverging directions, and the network has to reach a common minimum by adapting and learning features common to all tasks in its shared layers. We train a multi-task network using auxiliary tasks extracted from taxonomical or semantic hierarchies, using a novel method to update task-wise learning rates to ensure that the related tasks aid and unrelated tasks does not hamper performance on the primary task. Experiments on the popular CUB-200-2011 dataset show that employing super-classes in an end-to-end model improves performance, compared to methods employing additional expensive annotations such as keypoints and bounding boxes and/or using multi-stage pipelines. Riddhiman Dasgupta, Anoop M. Namboodiri |
ICPR | 2 |
| 2015 | Semantic Classification of Boundaries of an RGBD ImageabstractThe problem of labeling the edges present in a single color image as convex, concave, and occluding entities is one of the fundamental problems in computer vision. It has been shown that this information can contribute to segmentation, reconstruction and recognition problems. Recently, it has been shown that this classification is not straightforward even using RGBD data. This makes us wonder whether this apparent simple cue has more information than a depth map? In this paper, we propose a novel algorithm using random forest for classifying edges into convex, concave and occluding entities. We release a data set with more than 500 RGBD images with pixel-wise ground labels. Our method produces promising results and achieves an F-score of 0.84 on the data set. Nishit Soni, Anoop M. Namboodiri, C. V. Jawahar, Srikumar Ramalingam |
BMVC | 2 |
| 2015 | Visual Phrases for Exemplar Face DetectionabstractRecently, exemplar based approaches have been successfully applied for face detection in the wild. Contrary to traditional approaches that model face variations from a large and diverse set of training examples, exemplar-based approaches use a collection of discriminatively trained exemplars for detection. In this paradigm, each exemplar casts a vote using retrieval framework and generalized Hough voting, to locate the faces in the target image. The advantage of this approach is that by having a large database that covers all possible variations, faces in challenging conditions can be detected without having to learn explicit models for different variations. Current schemes, however, make an assumption of independence between the visual words, ignoring their relations in the process. They also ignore the spatial consistency of the visual words. Consequently, every exemplar word contributes equally during voting regardless of its location. In this paper, we propose a novel approach that incorporates higher order information in the voting process. We discover visual phrases that contain semantically related visual words and exploit them for detection along with the visual words. For spatial consistency, we estimate the spatial distribution of visual words and phrases from the entire database and then weigh their occurrence in exemplars. This ensures that a visual word or a phrase in an exemplar makes a major contribution only if it occurs at its semantic location, thereby suppressing the noise significantly. We perform extensive experiments on standard FDDB, AFW and G-album datasets and show significant improvement over previous exemplar approaches. Vijay Kumar 0004, Anoop M. Namboodiri, C. V. Jawahar |
ICCV | 2 |
| 2015 | Online handwriting recognition using depth sensorsabstractIn this work, we propose an online handwriting solution, where the data is captured with the help of depth sensors. Users may write in the air and our method recognizes it in real time using the proposed feature representation. Our method uses an efficient fingertip tracking approach and reduces the necessity of pen-up/pen-down switching. We validate our method on two depth sensors, Kinect and Leap Motion Controller. On a dataset collected from 20 users, we achieve a recognition accuracy of 97.59% for character recognition. We also demonstrate how this system can be extended for lexicon recognition with reliable performance. We have also prepared a dataset containing 1,560 characters and 400 words with the intention of providing common benchmark for handwritten character recognition using depth sensors and related research. Rajat Aggarwal, Sirnam Swetha, Anoop M. Namboodiri, Jayanthi Sivaswamy, C. V. Jawahar |
ICDAR | 3 |
| 2014 | Estimating Floor Regions in Cluttered Indoor Scenes from First Person Camera ViewabstractThe ability to detect floor regions from an image enables a variety of applications such as indoor scene understanding, mobility assessment, robot navigation, path planning and surveillance. In this work, we propose a framework for estimating floor regions in cluttered indoor environments. The problem of floor detection and segmentation is challenging in situations where floor and non-floor regions have similar appearances. It is even harder to segment floor regions when clutter, specular reflections, shadows and textured floors are present within the scene. Our framework utilizes a generic classifier trained from appearance cues as well as floor density estimates, both trained from a variety of indoor images. The results of the classifier is then adapted to a specific test image where we integrate appearance, position and geometric cues in an iterative framework. A Markov Random Field framework is used to integrate the cues to segment floor regions. In contrast to previous settings that relied on optical flow, depth sensors or multiple images in a calibrated setup, our method can work on a single image. It is also more flexible as we avoid assumptions like Manhattan world scene or restricting clutter only to wall-floor boundaries. Experimental results on the public MIT Scene dataset as well as a more challenging dataset that we acquired, demonstrate the robustness and efficiency of our framework on the above mentioned complex situations. Sanchit Aggarwal, Anoop M. Namboodiri, C. V. Jawahar |
ICPR | 2 |
| 2014 | Face Recognition in Videos by Label PropagationabstractWe consider the problem of automatic identification of faces in videos such as movies, given a dictionary of known faces from a public or an alternate database. This has applications in video indexing, content based search, surveillance, and real time recognition on wearable computers. We propose a two stage approach for this problem. First, we recognize the faces in a video using a sparse representation framework using l1-minimization and select a few key-frames based on a robust confidence measure. We then use transductive learning to propagate the labels from the key-frames to the remaining frames by incorporating constraints simultaneously in temporal and feature spaces. This is in contrast to some of the previous approaches where every test frame/track is identified independently, ignoring the correlation between the faces in video tracks. Having a few key frames belonging to few subjects for label propagation rather than a large dictionary of actors reduces the amount of confusion. We evaluate the performance of our algorithm on Movie Trailer face dataset and five movie clips, and achieve a significant improvement in labeling accuracy compared to previous approaches. Vijay Kumar 0004, Anoop M. Namboodiri, C. V. Jawahar |
ICPR | 2 |
| 2013 | Ink-Bleed Reduction Using Layer SeparationabstractWe present a novel method for reducing the effects of ink-bleed in handwritten documents. We go beyond the existing works on ink bleed detection and removal. We consider each pixel in a document as a result of combination of foreground, ink-bleed and background. We carry of a decomposition of the document image into separate foreground ink, ink-bleed, and background Layers. We propose an efficient MRF formulation to achieve this separation. Degradation model for the ink and paper is proposed. The ability to extract the contributions of the three components to each pixel allows us to recover finer details of the writing. Quantitative and qualitative results on a set of historic manuscripts as well as synthetically generated documents demonstrate the effectiveness of our approach. Shrikant Baronia, Anoop M. Namboodiri |
ICDAR | 2 |
| 2013 | Sparse Document Image Coding for RestorationabstractSparse representation based image restoration techniques have shown to be successful in solving various inverse problems such as denoising, in painting, and super-resolution, etc. on natural images and videos. In this paper, we explore the use of sparse representation based methods specifically to restore the degraded document images. While natural images form a very small subset of all possible images admitting the possibility of sparse representation, document images are significantly more restricted and are expected to be ideally suited for such a representation. However, the binary nature of textual document images makes dictionary learning and coding techniques unsuitable to be applied directly. We leverage the fact that different characters possess similar strokes, curves, and edges, and learn a dictionary that gives sparse decomposition for patches. Experimental results show significant improvement in image quality and OCR performance on documents collected from a variety of sources such as magazines and books. This method is therefore, ideally suited for restoring highly degraded images in repositories such as digital libraries. Vijay Kumar 0004, Amit Bansal, Goutam Hari Tulsiyan, Anand Mishra 0001, Anoop M. Namboodiri, C. V. Jawahar |
ICDAR | 5 |
| 2013 | A Ballistic Stroke Representation of Online Handwriting for RecognitionabstractRobust segmentation of ballistic strokes from online handwritten traces is critical in parameter estimation of stroke based models for applications such as recognition, synthesis, and writer identification. In this paper we propose a new method for segmenting ballistic strokes from online handwriting. Traditional methods of ballistic stroke segmentation rely on detection of local minima of pen speed. Unfortunately, this approach is highly sensitive to noise, in sensing and in both spatial and temporal dimensions. We decompose the problem into two steps, where the spatial noise is filtered out in the first step. The ballistic stroke boundaries are then detected at the local curvature maxima, which we show to be invariant to temporal sampling noise. We also propose a bag-of-strokes representation based on ballistic stroke segmentation for online character recognition that improves the state-of-the-art recognition accuracies on multiple datasets. S. Prabhu Teja, Anoop M. Namboodiri |
ICDAR | 2 |
| 2012 | Improving realism of 3D texture using component based modelingabstract3D textures are often described by parametric functions for each pixel, that models the variation in its appearance with respect to varying lighting direction. However, parametric models such as Polynomial Texture Maps (PTMs) tend to smoothen the changes in appearance. We propose a technique to effectively model natural material surfaces and their interactions with changing light conditions. We show that the direct and global components of the image have different nature, and when modeled separately, leads to a more accurate and compact model of the 3D surface texture. For a given lighting position, both components are computed separately and combined to render a new image. This method models sharp shadows and specularities, while preserving the structural relief and surface color. Thus rendered image have enhanced photorealism as compared to images rendered by existing single pixel models such as PTMs. Siddharth Kherada, Prateek Pandey, Anoop M. Namboodiri |
WACV | 3 |
| 2011 | Fingerprint feature extraction from gray scale images by ridge tracingabstractThis paper deals with extraction of fingerprint features directly from gray scale images by the method of ridge tracing. While doing so, we make substantial use of contextual information gathered during the tracing process. Narrow bandpass based filtering methods for fingerprint image enhancement are extremely robust as noisy regions do not affect the result of cleaner ones. However, these method often generate artifacts whenever the underlying image does not fit the filter model, which may be due to the presence of noise and singularities. The proposed method allows us to use the contextual information to better handle such noisy regions. Moreover, the various parameters used in the algorithm have been made adaptive in order to circumvent human supervision. The experimental results from our algorithm have been compared with those from Gabor based filtering and feature extraction, as well as with the original ridge tracing work from Maio and Maltoni [11]. The results clearly indicate that the proposed approach makes ridge tracing more robust to noise and makes the extracted features more reliable. Devansh Arpit, Anoop M. Namboodiri |
IJCB | 2 |
| 2011 | Fingerprint enhancement using Hierarchical Markov Random FieldsabstractWe propose a novel approach to enhance the finger print image and extract features such as directional fields, minutiae and singular points reliably using a Hierarchical Markov Random Field Model. Unlike traditional finger print enhancement techniques, we use previously learned prior patterns from a set of clean fingerprints to restore a noisy one. We are able to recover the ridge and valley structure from degraded and noisy fingerprint images by formulating it as a hierarchical interconnected MRF that processes the information at multiple resolutions. The top layer incorporates the compatibility between an observed degraded fingerprint patch and prior training patterns in addition to ridge continuity across neighboring patches. A second layer accounts for spatial smoothness of the orientation field and its discontinuity at the singularities. Further layers could be used for incorporating higher level priors such as the class of the fingerprint. The strength of the pro posed approach lies in its flexibility to model possible variations in fingerprint images as patches and from its ability to incorporate contextual information at various resolutions. Experimental results (both quantitative and qualitative) clearly demonstrate the effectiveness of this approach. Reddy K. N. V. Rama, Anoop M. Namboodiri |
IJCB | 2 |
| 2011 | A Semi-supervised SVM Framework for Character RecognitionabstractIn order to incorporate various writing styles or fonts in a character recognizer, it is critical that a large amount of labeled data is available, which is difficult to obtain. In this work, we present a semi-supervised SVM based framework that can incorporate the unlabeled data for improvement of recognition performance. Existing semi supervised learning methods for SVMs work well only for two-class problems. We propose a method to extend this to large-class problems by incorporating a participation term into the optimization process. The proposed system uses a Decision Directed Acyclic Graphs (DDAG) of SVM classifiers, which have proven to be very effective for such recognition problems. We present experimental results on three different digits dataset with varying complexity, as well as additional multi-class datasets from the UCI repository for comparison with existing approaches. In addition we show that approximate annotations at the word or sentence level can be used for evaluation as well as active learning to further improve the recognition results. Anoop M. Namboodiri |
ICDAR | 2 |
| 2010 | A Hybrid Model for Recognition of Online Handwriting in Indian ScriptsabstractWe present a complete online handwritten character recognition system for Indian languages that handles the ambiguities in segmentation as well as recognition of the strokes. The recognition is based on a generative model of handwriting formation, coupled with a discriminative model for classification of strokes. Such an approach can seamlessly integrate language and script information in the generative model and deal with similar strokes using the discriminative stroke classification model. The recognition is performed in a purely bottom-up fashion, starting with the strokes, and the ambiguities at each stage are reserved and transferred to the next stage for obtaining the most probable results at each stage. We also present the results of various pre-processing, feature selection and classification studies on a large data set collected from native language writers in two different Indian languages: Malayalam and Telugu. The system achieves a stroke level accuracy of 95.78% and 95.12% on Malayalam and Telugu data, respectively. The akshara level accuracy of the system is around 78% on a corpus of 60, 492 words from 367 writers. Anoop M. Namboodiri |
ICFHR | 2 |
| 2010 | Video Based Palmprint RecognitionabstractThe use of camera as a biometric sensor is desirable due to its ubiquity and low cost, especially for mobile devices. Palm print is an effective modality in such cases due to its discrimination power, ease of presentation and the scale and size of texture for capture by commodity cameras. However, the unconstrained nature of pose and lighting introduces several challenges in the recognition process. Even minor changes in pose of the palm can induce significant changes in the visibility of the lines. We turn this property to our advantage by capturing a short video, where the natural palm motion induces minor pose variations, providing additional texture information. We propose a method to register multiple frames of the video without requiring correspondence, while being efficient. Experimental results on a set of different 100 palms show that the use of multiple frames reduces the error rate from 12.75% to 4.7%. We also propose a method for detection of poor quality samples due to specularities and motion blur, which further reduces the EER to 1.8%. Chhaya Methani, Anoop M. Namboodiri |
ICPR | 2 |
| 2010 | Blind authentication: a secure crypto-biometric verification protocolabstractConcerns on widespread use of biometric authentication systems are primarily centered around template security, revocability, and privacy. The use of cryptographic primitives to bolster the authentication process can alleviate some of these concerns as shown by biometric cryptosystems. In this paper, we propose aprovably secureandblindbiometric authentication protocol, which addresses the concerns of user's privacy, template protection, and trust issues. The protocol is blind in the sense that it reveals only the identity, and no additional information about the user or the biometric to the authenticating server or vice-versa. As the protocol is based on asymmetric encryption of the biometric data, it captures the advantages of biometric authentication as well as the security of public key cryptography. The authentication protocol can run over public networks and provide nonrepudiable identity verification. The encryption also provides template protection, the ability to revoke enrolled templates, and alleviates the concerns on privacy in widespread use of biometrics. The proposed approach makes no restrictive assumptions on the biometric data and is hence applicable to multiple biometrics. Such a protocol has significant advantages over existing biometric cryptosystems, which use a biometric to secure a secret key, which in turn is used for authentication. We analyze the security of the protocol under various attack scenarios. Experimental results on four biometric datasets (face, iris, hand geometry, and fingerprint) show that carrying out the authentication in the encrypted domain does not affect the accuracy, while the encryption key acts as an additional layer of security. Maneesh Upmanyu, Anoop M. Namboodiri, K. Srinathan 0001, C. V. Jawahar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Contextual restoration of severely degraded document imagesabstractWe propose an approach to restore severely degraded document images using a probabilistic context model. Unlike traditional approaches that use previously learned prior models to restore an image, we are able to learn the text model from the degraded document itself, making the approach independent of script, font, style, etc. We model the contextual relationship using an MRF. The ability to work with larger patch sizes allows us to deal with severe degradations including cuts, blobs, merges and vandalized documents. Our approach can also integrate document restoration and super-resolution into a single framework, thus directly generating high quality images from degraded documents. Experimental results show significant improvement in image quality on document images collected from various sources including magazines and books, and comprehensively demonstrate the robustness and adaptability of the approach. It works well with document collections such as books, even with severe degradations, and hence is ideally suited for repositories such as digital libraries. Jyotirmoy Banerjee, Anoop M. Namboodiri, C. V. Jawahar |
CVPR | 2 |
| 2009 | Efficient privacy preserving video surveillanceabstractWidespread use of surveillance cameras in offices and other business establishments, pose a significant threat to the privacy of the employees and visitors. The challenge of introducing privacy and security in such a practical surveillance system has been stifled by the enormous computational and communication overhead required by the solutions. In this paper, we propose an efficient framework to carry out privacy preserving surveillance. We split each frame into a set of random images. Each image by itself does not convey any meaningful information about the original frame, while collectively, they retain all the information. Our solution is derived from a secret sharing scheme based on the Chinese Remainder Theorem, suitably adapted to image data. Our method enables distributed secure processing and storage, while retaining the ability to reconstruct the original data in case of a legal requirement. The system installed in an office like environment can effectively detect and track people, or solve similar surveillance tasks. Our proposed paradigm is highly efficient compared to Secure Multiparty Computation, making privacy preserving surveillance, practical. Maneesh Upmanyu, Anoop M. Namboodiri, K. Srinathan 0001, C. V. Jawahar |
ICCV | 2 |
| 2009 | Learning and Adaptation for Improving Handwritten Character RecognizersabstractWriter independent handwriting recognition systems are limited in their accuracy, primarily due the large variations in writing styles of most characters. Samples from a single character class can be thought of as emanating from multiple sources, corresponding to each writing style. This also makes the inter-class boundaries, complex and disconnected in the feature space. Multiple kernel methods have emerged as a potential framework to model such decision boundaries effectively, which can be coupled with maximal margin learning algorithms. We show that formulating the problem in the above framework improves the recognition accuracy. We also propose a mechanism to adapt the resulting classifier by modifying the weights of the support vectors as well as that of the individual kernels. Experimental results are presented on a data set of 16,000 alphabets collected from 470 writers using a digitizing tablet. Naveen Chandra Tewari, Anoop M. Namboodiri |
ICDAR | 2 |
| 2009 | Retrieval of online handwriting by synthesis and matching
C. V. Jawahar, A. Balasubramanian, Million Meshesha, Anoop M. Namboodiri |
Pattern Recognit. | 4 |
| 2008 | Projected Texture for Object Classification
Avinash Sharma 0001, Anoop M. Namboodiri |
ECCV (3) | 2 |
| 2008 | Robust image registration with illumination, blur and noise variations for super-resolutionabstractSuper-resolution reconstruction algorithms assume the availability of exact registration and blur parameters. Inaccurate estimation of these parameters adversely affects the quality of the reconstructed image. However, traditional approaches for image registration are either sensitive to image degradations such as variations in blur, illumination and noise, or are limited in the class of image transformations that can be estimated. We propose an accurate registration algorithm that uses the local phase information, which is robust to the above degradations. We derive the theoretical error rate of the estimates in presence of non-ideal band-pass behavior of the filter and show that the error converges to zero over iterations. We also show the invariance of local phase to a class of blur kernels. Experimental results on images taken under varying conditions clearly demonstrates the robustness of our approach. Himanshu Arora, Anoop M. Namboodiri, C. V. Jawahar |
ICASSP | 2 |
| 2007 | On Using Classical Poetry Structure for Indian Language Post-ProcessingabstractPost-processors are critical to the performance of language recognizers like OCRs, speech recognizers, etc. Dictionary-based post-processing commonly employ either an algorithmic approach or a statistical approach. Other linguistic features are not exploited for this purpose. The language analysis is also largely limited to the prose form. This paper proposes a framework to use the rich metric and formal structure of classical poetic forms in Indian languages for post-processing a recognizer like an OCR engine. We show that the structure present in the form of the vrtta and prasa can be efficiently used to disambiguate some cases that may be difficult for an OCR. The approach is efficient, and complementary to other post-processing approaches and can be used in conjunction with them. Anoop M. Namboodiri, P. J. Narayanan, C. V. Jawahar |
ICDAR | 1 |
| 2005 | Document Understanding System Using Stochastic Context-Free GrammarsabstractWe present a document understanding system in which the arrangement of lines of text and block separators within a document are modeled by stochastic context free grammars. A grammar corresponds to a document genre; our system may be adapted to a new genre simply by replacing the input grammar. The system incorporates an optical character recognition system that outputs characters, their positions and font sizes. These features are combined to form a document representation of lines of text and separators. Lines of text are labeled as tokens using regular expression matching. The maximum likelihood parse of this stream of tokens and separators yields a functional labeling of the document lines. We describe business card and business letter applications. John C. Handley, Anoop M. Namboodiri, Richard Zanibbi |
ICDAR | 2 |
| 2004 | Online Handwritten Script Recognition
Anoop M. Namboodiri, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Indexing and Retrieval of On-line Handwritten DocumentsabstractRecent advances in on-line data capturing technologiesand its widespread deployment in devices like PDAsand notebook PCs is creating large amounts of handwrittendata that need to be archived and retrieved efficiently.Word-spotting, which is based on a direct comparison ofa handwritten keyword to words in the document, is commonlyused for indexing and retrieval. We propose a stringmatching-based method for word-spotting in on-line documents.The retrieval algorithm achieves a precision of92.3% at a recall rate of 90% on a database of 6,672 wordswritten by 10 different writers. Indexing experiments showan accuracy of 87.5% using a database of 3,872 on-linewords. Anil K. Jain 0001, Anoop M. Namboodiri |
ICDAR | 2 |
| 2001 | Structure in On-line DocumentsabstractWe present a hierarchical approach for extracting homogeneous regions in on-line documents. The problem of identifying and processing ruled and unruled tables, text and drawings is addressed. The on-line document is first segmented into regions with only text strokes and regions with both text and non-text strokes. The text region is further classified as unruled table or plain text. Stroke clustering is used to segment the non-text regions. Each nontext segment is then classified as drawing, ruled table or underlined keyword using stroke properties. The individual regions are processed and the results are assembled to identify the structure of the on-line document. Anil K. Jain 0001, Anoop M. Namboodiri, Jayashree Subrahmonia |
ICDAR | 2 |