Sumantra Dutta Roy

dblp:44/5419 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-2141-5067ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Biometric characteristics of hand gestures through joint decomposition of cross-subject and cross-session biases
Aman Verma, Gaurav Jaswal, Seshan Srirangarajan, Sumantra Dutta Roy
Pattern Recognit. Lett.4
2024 Quantifying Biometric Characteristics of Hand Gestures Through Feature Space Probing and Identity-Level Cross-Gesture Disentanglement
abstract
We present the delta-gesture biometrics quantification assessment (DGBQA) framework which estimates the biometric characteristics of hand gestures. The proposed framework is aimed at learning generic motion-representations of gestures instead of subject-specific details from a large number of identities. It also enables the biometric scores to be estimated for a set of gestures at a time instead of having to estimate these one at a time. In the first step, it formulates a feature space which is identity and gesture aware, and in the second step, it proceeds to compute biometric scores using inter-subject and intra-subject distance measures in the feature space. However, due to the inclusion of identity-aware objective, the identity details tend to be shared across gestures. We refer to this as identity sharing and this can lead to the score for different gestures being dependent on each other. To address this issue, we introduce an identity-level cross-gesture disentanglement loss$(\mathscr{L}_{ICGD})$which encourages the different gestures belonging to the same identity to be orthogonal in the feature space. We demonstrate the efficacy of the proposed biometric quantification framework and the disentanglement loss function through extensive experiments on four datasets and using standard as well as proposed novel evaluation metrics. Our analysis indicates that gestures involving multiple coarse movements are better for biometrics.
Aman Verma, Gaurav Jaswal, Seshan Srirangarajan, Sumantra Dutta Roy
FG4
2024 Unsupervised domain alignment of fingerprint denoising models using pseudo annotations
Indu Joshi, Tushar Prakash, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
Multim. Tools Appl.5
2024 On characterizing the evolution of embedding space of neural networks using algebraic topology
Suryaka Suresh, Bishshoy Das, Vinayak Abrol, Sumantra Dutta Roy
Pattern Recognit. Lett.4
2023 Altering Backward Pass Gradients to Improve Convergence (S)
abstract
In standard neural network training, the gradients in the backward pass are determined by the forward pass.As a result, the two stages are coupled.This is how most neural networks are trained hitherto.Gradient modification in the backward pass has seldom been studied in the literature.In this paper we explore decoupled training, where we alter the gradients in the backward pass.We propose a simple yet powerful method called PowerGrad Transform (PGT), that alters the gradients before the weight update in the backward pass and significantly enhances the predictive performance of a convolutional neural network.PGT trains networks to arrive at a better optima at convergence.It is computationally efficient, and adds no additional cost to either memory or compute, but results in improved final accuracies on both the training and test datasets.Power-Grad Transform is easy to integrate into existing training routines, requiring just a few lines of code.With decoupled training, our method improves baseline accuracies for ResNet-50 by 0.73%, for SE-ResNet-50 by 0.66% and by more than 1.0% for the non-normalized ResNet-18 network on the ImageNet classification task.
Bishshoy Das, Milton Mondal, Brejesh Lall, Shiv Dutt Joshi, Sumantra Dutta Roy
SEKE5
2023 Feature independent Filter Pruning by Successive Layers analysis
Milton Mondal, Bishshoy Das, Brejesh Lall, Pushpendra Singh 0002, Sumantra Dutta Roy, Shiv Dutt Joshi
Comput. Vis. Image Underst.5
2022 Knowledge Diversification in Ensembles of Identical Neural Networks
Bishshoy Das, Sumantra Dutta Roy
BMVC2
2022 Adaptive CNN filter pruning using global importance metric
Milton Mondal, Bishshoy Das, Sumantra Dutta Roy, Pushpendra Singh 0002, Brejesh Lall, Shiv Dutt Joshi
Comput. Vis. Image Underst.3
2022 On restoration of degraded fingerprints
Indu Joshi, Ayush Utkarsh, Pravendra Singh, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
Multim. Tools Appl.5
2021 Data Uncertainty Guided Noise-aware Preprocessing Of Fingerprints
abstract
The effectiveness of fingerprint-based authentication systems on good quality fingerprints is established long back. However, the performance of standard fingerprint matching systems on noisy and poor quality fingerprints is far from satisfactory. Towards this, we propose a data uncertainty-based framework which enables the state-of-the-art fingerprint pre-processing models to quantify noise present in the input image and identify fingerprint regions with background noise and poor ridge clarity. Quantification of noise helps the model two folds: firstly, it makes the objective function adaptive to the noise in a particular input fingerprint and consequently, helps to achieve robust performance on noisy and distorted fingerprint regions. Secondly, it provides a noise variance map which indicates noisy pixels in the input fingerprint image. The predicted noise variance map enables the end-users to understand erroneous predictions due to noise present in the input image. Extensive experimental evaluation on 13 publicly available fingerprint databases, across different architectural choices and two fingerprint processing tasks demonstrate effectiveness of the proposed framework.
Indu Joshi, Ayush Utkarsh, Riya Kothari, Vinod K. Kurmi, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
IJCNN6
2021 Sensor-invariant Fingerprint ROI Segmentation Using Recurrent Adversarial Learning
abstract
A fingerprint region of interest (roi) segmentation algorithm is designed to separate the foreground fingerprint from the background noise. All the learning based state-of-the-art fingerprint roi segmentation algorithms proposed in the literature are benchmarked on scenarios when both training and testing databases consist of fingerprint images acquired from the same sensors. However, when testing is conducted on a different sensor, the segmentation performance obtained is often unsatisfactory. As a result, every time a new fingerprint sensor is used for testing, the fingerprint roi segmentation model needs to be re-trained with the fingerprint image acquired from the new sensor and its corresponding manually marked ROI. Manually marking fingerprint ROI is expensive because firstly, it is time consuming and more importantly, requires domain expertise. In order to save the human effort in generating annotations required by state-of-the-art, we propose a fingerprint roi segmentation model which aligns the features of fingerprint images derived from the unseen sensor such that they are similar to the ones obtained from the fingerprints whose ground truth roi masks are available for training. Specifically, we propose a recurrent adversarial learning based feature alignment network that helps the fingerprint roi segmentation model to learn sensor-invariant features. Consequently, sensor-invariant features learnt by the proposed roi segmentation model help it to achieve improved segmentation performance on fingerprints acquired from the new sensor. Experiments on publicly available FVC databases demonstrate the efficacy of the proposed work.
Indu Joshi, Ayush Utkarsh, Riya Kothari, Vinod K. Kurmi, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
IJCNN6
2019 Latent Fingerprint Enhancement Using Generative Adversarial Networks
abstract
Latent fingerprints recognition is very useful in law enforcement and forensics applications. However, automated matching of latent fingerprints with a gallery of live scan images is very challenging due to several compounding factors such as noisy background, poor ridge structure, and overlapping unstructured noise. In order to efficiently match latent fingerprints, an effective enhancement module is a necessity so that it can facilitate correct minutiae extraction. In this research, we propose a Generative Adversarial Network based latent fingerprint enhancement algorithm to enhance the poor quality ridges and predict the ridge information. Experiments on two publicly available datasets, IIITD-MOLF and IIITD-MSLFD show that the proposed enhancement algorithm improves the fingerprints quality while preserving the ridge structure. It helps the standard feature extraction and matching algorithms to boost latent fingerprints matching performance.
Indu Joshi, Adithya Anand, Mayank Vatsa, Richa Singh 0001, Sumantra Dutta Roy, Prem Kumar Kalra
WACV5
2018 Novel and improved stage estimation in Parkinson's disease using clinical scales and machine learning
R. Prashanth, Sumantra Dutta Roy
Neurocomputing2
2018 Restricted affine motion compensation and estimation in video coding with particle filtering and importance sampling: a multi-resolution approach
Mithilesh Kumar Jha, Ravi Chaudhary, Sumantra Dutta Roy, Mona Mathur, Brejesh Lall
Multim. Syst.3
2017 High-Accuracy Classification of Parkinson's Disease Through Shape Analysis and Surface Fitting in 123I-Ioflupane SPECT Imaging
abstract
Early and accurate identification of Parkinsonian syndromes (PS) involving presynaptic degeneration from nondegenerative variants such as scans without evidence of dopaminergic deficit (SWEDD) and tremor disorders is important for effective patient management as the course, therapy, and prognosis differ substantially between the two groups. In this study, we use single photon emission computed tomography (SPECT) images from healthy normal, early PD, and SWEDD subjects, as obtained from the Parkinson's Progression Markers Initiative (PPMI) database, and process them to compute shape- and surface-fitting-based features. We use these features to develop and compare various classification models that can discriminate between scans showing dopaminergic deficit, as in PD, from scans without the deficit, as in healthy normal or SWEDD. Along with it, we also compare these features with striatal binding ratio (SBR)-based features, which are well established and clinically used, by computing a feature-importance score using random forests technique. We observe that the support vector machine (SVM) classifier gives the best performance with an accuracy of 97.29%. These features also show higher importance than the SBR-based features. We infer from the study that shape analysis and surface fitting are useful and promising methods for extracting discriminatory features that can be used to develop diagnostic models that might have the potential to help clinicians in the diagnostic process.
Prashanth Ravindran, Sumantra Dutta Roy, Pravat K. Mandal, Shantanu Ghosh
IEEE J. Biomed. Health Informatics2
2016 Pose estimation of texture-less cylindrical objects in bin picking using sensor fusion
abstract
We propose an approach for emptying of bin using a combination of Image and Range sensor. Offering a complete solution: calibration, segmentation and pose estimation, along with approachability analysis for the estimated pose. The work is novel in the sense that the objects to be picked are featureless and uniformly black in colour, hence existing approaches are not directly applicable. A key point involves optimal utilization of range data acquired from the laser scanner for 3-D segmentation using localized geometric information. This information guides segmentation of the image for better object pose estimation, used for pick-and-drop. We analytically assure the approachability of the object to avoid collision of the manipulator with the bin. Disturbance of objects caused during pick up has been modelled, which allows pickup of multiple pellets based on information from a single range scan. This eliminates the necessity of repeated scanning and data conditioning. The proposed method offers high object detection rate and pose estimation accuracy. The innovative techniques aimed at reducing the average pickup time makes it suitable for robust industrial operation.
Mayank Roy, Riby Abraham Boby, Shraddha Chaudhary, Santanu Chaudhury, Sumantra Dutta Roy, Subir Kumar Saha
IROS5
2015 DEMD-based video coding for textured videos in an H.264/MPEG framework
Mithilesh Kumar Jha, Sumantra Dutta Roy, Brejesh Lall
Pattern Recognit. Lett.2
2015 Camera-based document image matching using multi-feature probabilistic information fusion
Sumantra Dutta Roy, Kavita Bhardwaj, Rhishabh Garg, Santanu Chaudhury
Pattern Recognit. Lett.1
2014 Newspaper Article Extraction Using Hierarchical Fixed Point Model
abstract
This paper presents a novel learning based framework to extract articles from newspaper images using a Fixed-Point Model. The input to the system comprises blocks of text and graphics, obtained using standard image processing techniques. The fixed point model uses contextual information and features of each block to learn the layout of newspaper images and attains a contraction mapping to assign a unique label to every block. We use a hierarchical model which works in two stages. In the first stage, a semantic label (heading, sub-heading, text-blocks, image and caption) is assigned to each segmented block. The labels are then used as input to the next stage to group the related blocks into news articles. Experimental results show the applicability of our algorithm in newspaper labeling and article extraction.
Anukriti Bansal, Santanu Chaudhury, Sumantra Dutta Roy, J. B. Srivastava
Document Analysis Systems3
2014 Automatic classification and prediction models for early Parkinson's disease diagnosis from SPECT imaging
R. Prashanth, Sumantra Dutta Roy, Pravat K. Mandal, Shantanu Ghosh
Expert Syst. Appl.2
2008 Parametric video compression using appearance space
abstract
The novelty of the approach presented in this paper is the unique object-based video coding framework for videos obtained from a static camera. As opposed to most existing methods, the proposed method does not require explicit 2D or 3D models of objects and hence is general enough to satisfy the need for varying types of objects in the scene. The proposed system detects and tracks an object in the scene by learning the appearance model of each object online using nontraditional uniform norm based subspace. At the same time the object is coded using the projection coefficients to the orthonormal basis of the subspace learnt. The tracker incorporates a predictive framework based upon Kalman filter for predicting the five motion parameters. The proposed method shows substantially better compression than MPEG2 based coding with almost no additional complexity.
Santanu Chaudhury, Subarna Tripathi, Sumantra Dutta Roy
ICPR3
2007 On Stabilisation of Parametric Active Contours
abstract
Parametric active contours have been used extensively in computer vision for different tasks like segmentation and tracking. However, all parametric contours are known to suffer from the problem of frequent bunching and spacing out of curve points locally during the curve evolution. In a spline based implementation of active contours, this leads to occasional formation of loops locally, and subsequently the curve blows up due to instabilities. It has been shown earlier that in addition to usual evolution along the normal direction, the curve should also be evolved in the tangential direction for stability purposes. In this paper, we provide a mathematical basis for selecting such a suitable tangential component for stabilisation. We prove the boundedness of the evolved curve in this paper, and provide the physical significance. We demonstrate the usefulness of the proposed method with a number of experiments.
Viswanathan Srikrishnan, Subhasis Chaudhuri, Sumantra Dutta Roy, Daniel Sevcovic
CVPR3
2007 Hand gesture modelling and recognition involving changing shapes and trajectories, using a Predictive EigenTracker
Kaustubh Patwardhan, Sumantra Dutta Roy
Pattern Recognit. Lett.2
2005 Multispectral panoramic mosaicing
Udhav Bhosle, Sumantra Dutta Roy, Subhasis Chaudhuri
Pattern Recognit. Lett.2
2005 Recognizing large isolated 3-D objects through next view planning using inner camera invariants
abstract
Most model-based three-dimensional (3-D) object recognition systems use information from a single view of an object. However, a single view may not contain sufficient features to recognize it unambiguously. Further, two objects may have all views in common with respect to a given feature set, and may be distinguished only through a sequence of views. A further complication arises when in an image, we do not have a complete view of an object. This paper presents a new online scheme for the recognition and pose estimation of a large isolated 3-D object, which may not entirely fit in a camera's field of view. We consider an uncalibrated projective camera, and consider the case when the internal parameters of the camera may be varied either unintentionally, or on purpose. The scheme uses a probabilistic reasoning framework for recognition and next-view planning. We show results of successful recognition and pose estimation even in cases of a high degree of interpretation ambiguity associated with the initial view.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
IEEE Trans. Syst. Man Cybern. Part B1
2004 Robust shape based two hand tracker
abstract
This paper presents a robust shape-based on-line tracker for simultaneously tracking the motion of both hands, that is robust to cases of background clutter, other moving objects, occlusions of one hand by the other and a wide range of illumination variations. The tracker is based on an online predictive eigentracking framework. This framework allows efficient tracking of articulate objects, which change in appearance across views. We show results of successful tracking across all possible cases of motion dynamics of both hands during occlusion and a wide range of illumination conditions.
Ketan Barhate, Kaustubh Patwardhan, Sumantra Dutta Roy, Subhasis Chaudhuri, Santanu Chaudhury
ICIP3
2004 On line predictive appearance-based tracking
abstract
We present a novel predictive statistical framework to improve the performance of an eigentracker. In addition, we use fast and efficient eigenspace updates to learn new views of the object being tracked on the fly. We also incorporate a new importance sampling mechanism which increases the robustness of the eigentracker and enables it to track nonconvex objects better. Our eigentracker is flexible-it is possible to use it symbolically with other trackers. We show its successful application in hand gesture analysis; and face and person tracking.
Namita Gupta, Pooja Mittal, Kaustubh Patwardhan, Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
ICIP4
2004 Active recognition through next view planning: a survey
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
Pattern Recognit.1
2003 Aspect graph construction with noisy feature detectors
abstract
Many three-dimensional (3D) object recognition strategies use aspect graphs to represent objects in the model base. A crucial factor in the success of these object recognition strategies is the accurate construction of the aspect graph, its ease of creation, and the extent to which it can represent all views of the object for a given setup. Factors such as noise and nonadaptive thresholds may introduce errors in the feature detection process. This paper presents a characterization of errors in aspect graphs, as well as an algorithm for estimating aspect graphs, given noisy sensor data. We present extensive results of our strategies applied on a reasonably complex experimental set, and demonstrate applications to a robust 3D object recognition problem.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
IEEE Trans. Syst. Man Cybern. Part B1
2001 Recognizing Large 3-D Objects through Next View Planning using an Uncalibrated Camera
abstract
We present a new on-line scheme for the recognition and pose estimation of a large isolated 3-D object, which may not entirely fit in a camera's field of view. We do not assume any knowledge of the internal parameters of the camera, or their constancy. We use a probabilistic reasoning framework for recognition and next view planning. We show results of successful recognition and pose estimation even in cases of a high degree of interpretation ambiguity associated with the initial view.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
ICCV1
2000 Isolated 3D object recognition through next view planning
abstract
In many cases, a single view of an object may not contain sufficient features to recognize it unambiguously. This paper presents a new online recognition scheme based on next view planning for the identification of an isolated 3D object using simple features. The scheme uses a probabilistic reasoning framework for recognition and planning. Our knowledge representation scheme encodes feature based information about objects as well as the uncertainty in the recognition process. This is used both in the probability calculations as well as in planning the next view. Results clearly demonstrate the effectiveness of our strategy for a reasonably complex experimental set.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
IEEE Trans. Syst. Man Cybern. Part A1
1999 Robot Localization using Uncalibrated Camera Invariants
abstract
We describe a set of image measurements which are invariant to the camera internals but are location variant. We show that using these measurements it is possible to calculate the self-localization of a robot using known landmarks and uncalibrated cameras. We also show that it is possible to compute, using uncalibrated cameras, the Euclidean structure of 3-D world points using multiple views from known positions. We are free to alter the internal parameters of the camera during these operations. Our initial experiments demonstrate the applicability of the method.
Michael Werman, MaoLin Qiu, Subhashis Banerjee, Sumantra Dutta Roy
CVPR4