VLDB 2026 Research / reviewers in the wild / expert
Sukhendu Das
dblp:96/1248
· DBLP profile ↗
46ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-2823-9211ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Guided 3D Gaussian Splatting for Sparse View Reconstruction and Segmentation
S. Meena Padnekar, Kaushik Mitra, Sukhendu Das |
ICPR (13) | 3 |
| 2026 | Distilling auxiliary RGB-T features for unsupervised semantic segmentation
S. Meena Padnekar, Kaushik Mitra, Sukhendu Das |
Image Vis. Comput. | 3 |
| 2025 | Conditional Diffusion Transformer for Unified Distortion Correction and RectificationabstractImage distortion correction and rectification are essential tasks in computer vision, addressing challenges such as lens aberrations, geometric distortions, and perspective errors that compromise image quality and usability. Despite advancements, existing methods often struggle with handling complex cases involving multiple distortions simultaneously. These tasks hold significant importance in applications like document processing, medical imaging, and autonomous systems. We propose a novel Conditioned-Guided Diffusion Transformer designed to correct diverse and complex distortions. Our approach combines the strengths of conditional diffusion transformer architectures, leveraging the visual inputs to guide the rectification process. The approach integrates conditional diffusion transformers with a coarse-to-fine rectification strategy. Initially, the input image undergoes pre-processing using Thin Plate Spline transformations, estimating control points and generating a mesh grid for coarse rectification. The semi-rectified image is then refined through a novel Flow Prediction Module, which computes the flow field required for further warping. Finally, a pretrained Latent Conditional Diffusion Transformer produces the fully rectified output. Comprehensive evaluations on benchmark datasets reveal that proposed work consistently surpasses state-of-the-art methods across various distortion types, including fisheye, barrel, skew, and perspective distortions. Its adaptability and superior results highlight it’s potential as a robust solution for real-world distortion correction and rectification tasks. Pooja Kumari 0001, Sukhendu Das |
ICIP | 2 |
| 2025 | Conditional Stable Diffusion for Distortion Correction and Image Rectification
Pooja Kumari 0001, Sukhendu Das |
Pattern Recognit. Lett. | 2 |
| 2024 | Multi-Scale Semantic Enrichment and Dual Angular Margin Contrast for Few-Shot Class Incremental Learning
Riya Verma, Sukhendu Das |
BMVC | 2 |
| 2024 | Stochastic Binary Network for Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) is the unsupervised domain adaptation with label shift. UniDA aims to classify unlabeled target samples into one of the "known" categories or into a single "unknown" category. Its main challenge lies in detecting private classes from both domains and performing alignment between the common classes. Current methods employ various techniques and loss functions to address these challenges. However, these methods commonly represent classifiers as point weight vectors, which are prone to overfitting by the source domain samples due to the lack of supervision from the target domain. Consequently, these classifiers struggle to separate target samples into known and unknown categories effectively. To address this, we introduce a novel framework called Stochastic Binary Network for Universal Domain Adaptation (STUN). STUN uses a Stochastic binary classifier for each class, whose weight is modeled as Gaus-sian distribution, enabling to sample an arbitrary number of classifiers while keeping the model size same as of two classifiers. Consistency between these sampled classifiers is used to derive the confidence scores for both source and target samples, which facilitates the alignment of common classes using weighted adversarial learning. Finally, we use deep discriminative clustering to formulate a loss function for solving the problem of fragmented feature distributions in the target domain. Extensive ablation studies and state-of-the-art results across three standard benchmark datasets show the efficacy of our framework. Saurabh Kumar Jain, Sukhendu Das |
WACV | 2 |
| 2024 | Am I readable? Transfer learning based document image rectification
Pooja Kumari 0001, Sukhendu Das |
Int. J. Document Anal. Recognit. | 2 |
| 2021 | V3GAN: Decomposing Background, Foreground and Motion for Video Generation
Arti Keshari, Sonam Gupta 0001, Sukhendu Das |
BMVC | 3 |
| 2021 | Where to Look?: Mining Complementary Image Regions for Weakly Supervised Object LocalizationabstractHumans possess an innate capability of recognizing objects and their corresponding parts and confine their attention to that location in a visual scene where the object is spatially present. Recently, efforts to train machines to mimic this ability of humans in the form of weakly supervised object localization, using training labels only at the image-level, have garnered a lot of attention. Nonetheless, one of the well-known problems that most of the existing methods suffer from is localizing only the most discriminative part of an object. Such methods provide very little or no focus on other pertinent parts of the object. In this paper, we propose a novel way of scrupulously localizing objects using training with labels as for the entire image by mining information from complementary regions in an image. Primarily, we adapt to regional dropout at complementary spatial locations to create two intermediate images. With the help of a novel Channel-wise Assisted Attention Module (CAAM) coupled with a Spatial Self-Attention Module (SSAM), we parallely train our model to leverage the information from complementary image regions for excellent localization. Finally, we fuse the attention maps generated by the two classifiers using our Attention-based Fusion Loss. Several experimental studies manifest the superior performance of our proposed approach. Our method demonstrates a significant increase in localization performance over the existing state-of-the-art methods on CUB-200-2011 and ILSVRC 2016 datasets. Sadbhavana Babar, Sukhendu Das |
WACV | 2 |
| 2020 | See the Sound, Hear the PixelsabstractFor every event occurring in the real world, most often a sound is associated with the corresponding visual scene. Humans possess an inherent ability to automatically map the audio content with visual scenes leading to an effortless and enhanced understanding of the underlying event. This triggers an interesting question: Can this natural correspondence between video and audio, which has been diminutively explored so far, be learned by a machine and modeled jointly to localize the sound source in a visual scene? In this paper, we propose a novel algorithm that addresses the problem of localizing sound source in unconstrained videos, which uses efficient fusion and attention mechanisms. Two novel blocks namely, Audio Visual Fusion Block (AVFB) and Segment-Wise Attention Block (SWAB) have been developed for this purpose. Quantitative and qualitative evaluations show that it is feasible to use the same algorithm with minor modifications to serve the purpose of sound localization using three different types of learning: supervised, weakly supervised and unsupervised. A novel Audio Visual Triplet Gram Matrix Loss (AVTGML) has been proposed as a loss function to learn the localization in an unsupervised way. Our empirical evaluations demonstrate a significant increase in performance over the existing state-of-the-art methods, serving as a testimony to the superiority of our proposed approach. Janani Ramaswamy, Sukhendu Das |
WACV | 2 |
| 2019 | What's There in the DarkabstractScene Parsing is an important cog for modern autonomous driving systems. Most of the works in semantic segmentation pertains to day-time scenes with favourable weather and illumination conditions. In this paper, we propose a novel deep architecture, NiSeNet, that performs semantic segmentation of night scenes using a domain mapping approach of synthetic to real data. It is a dual-channel network, where we designed a Real channel using DeepLabV3+ coupled with an MSE loss to preserve the spatial information. In addition, we used an Adaptive channel reducing the domain gap between synthetic and real night images, which also complements the failures of Real channel output. Apart from the dual channel, we introduced a novel fusion scheme to fuse the outputs of two channels. In addition to that, we compiled a new dataset Urban Night Driving Dataset (UNDD); it consists of 7125 unlabelled day and night images; additionally, it has 75 night images with pixel-level annotations having classes equivalent to Cityscapes dataset. We evaluated our approach on the Berkley Deep Drive dataset, the challenging Mapillary dataset and UNDD dataset to exhibit that the proposed method outperforms the state-of-the-art techniques in terms of accuracy and visual quality. Sauradip Nag, Saptakatha Adak, Sukhendu Das |
ICIP | 3 |
| 2019 | Motion-Based Occlusion-Aware Pixel Graph Network for Video Object Segmentation
Saptakatha Adak, Sukhendu Das |
ICONIP (2) | 2 |
| 2019 | Visual Saliency Detection via Convolutional Gated Recurrent Units
Sayanti Bardhan, Sukhendu Das, Shibu Jacob |
ICONIP (2) | 2 |
| 2019 | SACIC: A Semantics-Aware Convolutional Image Captioner Using Multi-level Pervasive Attention
Sandeep Narayan Parameswaran, Sukhendu Das |
ICONIP (3) | 2 |
| 2019 | Directional Attention based Video Frame Prediction using Graph Convolutional NetworksabstractThis paper proposes a novel network architecture for video frame prediction based on Graph Convolutional Neural Networks (GCNN). Most recent methods often fail in situations where multiple close-by objects at different scales move in random directions with variable speeds. We overcome this by modeling the scene as a space-time graph with intermediate features from the pixels (or a local region) as vertices and the relationships among them as edges. Our main contribution lies within posing the frame generation problem with our proposed space-time graph, which enables the network to learn the spatial as well as temporal inter-pixel relationships independent of each other, thus making the system invariant to velocity differences among the moving objects present in the scene. Moreover, we also propose a novel directional attention mechanism for the graph based model to efficiently learn a significance score based on directional relationship between pixels in the original scene. We also show that the proposed model generalizes better on the much more challenging task of predicting semantic scene segmentation of future scenes, even without access to any raw RGB frames. We perform several proxy tasks such as comparison of the quality of the semantic segmentation produced on the generated frames and comparing the accuracies for the task of recognizing actions in case of the dataset consisting of human actions. We use the popular Cityscapes traffic scene segmentation dataset as well as UCF-101 and Penn Action containing human actions to quantitatively and qualitatively evaluate the proposed framework over the recent state-of-the-art. Prateep Bhattacharjee, Sukhendu Das |
IJCNN | 2 |
| 2018 | Predicting Video Frames Using Feature Based Locally Guided Objectives
Prateep Bhattacharjee, Sukhendu Das |
ACCV (4) | 2 |
| 2018 | Panorama from Representative Frames of Unconstrained Videos Using DiffeoMeshes
Geethu Miriam Jacob, Sukhendu Das |
ACCV (3) | 2 |
| 2018 | Large Parallax Image Stitching Using an Edge-Preserving Diffeomorphic Warping Process
Geethu Miriam Jacob, Sukhendu Das |
ACIVS | 2 |
| 2018 | Mutual variation of information on transfer-CNN for face recognition with degraded probe samples
Samik Banerjee, Sukhendu Das |
Neurocomputing | 2 |
| 2018 | LR-GAN for degraded Face Recognition
Samik Banerjee, Sukhendu Das |
Pattern Recognit. Lett. | 2 |
| 2017 | Moving Object Segmentation in Jittery Videos by Stabilizing Trajectories Modeled in Kendall's Shape Space
Geethu Miriam Jacob, Sukhendu Das |
BMVC | 2 |
| 2017 | Temporal Coherency based Criteria for Predicting Video Frames using Deep Multi-stage Generative Adversarial NetworksabstractPredicting the future from a sequence of video frames has been recently a sought after yet challenging task in the field of computer vision and machine learning. Although there have been efforts for tracking using motion trajectories and flow features, the complex problem of generating unseen frames has not been studied extensively. In this paper, we deal with this problem using convolutional models within a multi-stage Generative Adversarial Networks (GAN) framework. The proposed method uses two stages of GANs to generate a crisp and clear set of future frames. Although GANs have been used in the past for predicting the future, none of the works consider the relation between subsequent frames in the temporal dimension. Our main contribution lies in formulating two objective functions based on the Normalized Cross Correlation (NCC) and the Pairwise Contrastive Divergence (PCD) for solving this problem. This method, coupled with the traditional L1 loss, has been experimented with three real-world video datasets, viz. Sports-1M, UCF-101 and the KITTI. Performance analysis reveals superior results over the recent state-of-the-art methods. Prateep Bhattacharjee, Sukhendu Das |
NIPS | 2 |
| 2017 | Moving object segmentation for jittery videos, by clustering of stabilized latent trajectories
Geethu Miriam Jacob, Sukhendu Das |
Image Vis. Comput. | 2 |
| 2016 | Supervised framework for automatic recognition and retrieval of interaction: a framework for classification and retrieving videos with similar human interactionsabstractThis study presents supervised framework for automatic recognition and retrieval of interactions (SAFARRIs), a supervised learning framework to recognise interactions such as pushing, punching, and hugging, between a pair of human performers in a video shot. The primary contribution of the study is to extend the vectors of locally aggregated descriptors (VLADs) as a compact and discriminative video encoding representation, to solve the complex class partitioning problem of recognising human interaction. An initial codebook is generated from the training set of video shots, by extracting feature descriptors around the spatiotemporal interest points computed across frames. A bag of action words is generated by encoding the first‐order statistics of the visual words using VLAD. Support vector machine classifiers (1 against all) are trained using these codebooks. The authors have verified SAFARRI's accuracy for classification and retrieval (query by example). SAFARRI is free from tracking or recognition of body parts and capable of identifying the region of interaction in video shots. It gives superior retrieval and classification performances over recently proposed methods, on two publicly available human interaction datasets. Chiranjoy Chattopadhyay, Sukhendu Das |
IET Comput. Vis. | 2 |
| 2016 | Minimising disparity in distribution for unsupervised domain adaptation by preserving the local spatial arrangement of dataabstractDomain adaptation is used for machine learning tasks, when the distribution of the training (obtained from source domain) set differs from that of the testing (referred as target domain) set. In the work presented in this study, the problem of unsupervised domain adaptation is solved using a novel optimisation function to minimise the global and local discrepancies between the transformed source and the target domains. The dissimilarity in data distributions is the major contributor to the global discrepancy between the two domains. The authors propose two techniques to preserve the local structural information of source domain: (i) identify closest pair of instances in source domain and minimise the distances between these pairs of instances after transformation; (ii) preserve the naturally occurring clusters present in source domain during transformation. This cost function and constraints yield a non‐linear optimisation problem, used to estimate the weight matrix. An iterative framework solves the optimisation problem, providing a sub‐optimal solution. Next, using orthogonality constraint, an optimisation task is formulated in the Stiefel manifold. Performance analysis using real‐world datasets show that the proposed methods perform better than a few recently published state‐of‐the‐art methods. Suranjana Samanta, Sukhendu Das |
IET Comput. Vis. | 2 |
| 2015 | Score Normalization in Multimodal Systems using Generalized Extreme Value DistributionabstractIn multimodal biometric systems, human identification is performed by fusing information in different ways like sensor-level, feature-level, score-level, rank-level and decision-level. Score-level fusion is preferred over other levels of fusion because of its low complexity and sufficient availability of information for fusion. However, the scores obtained from different unimodal systems are heterogeneous in nature and hence they all require normalization before fusion. In this paper, we propose a clientcentric score normalization technique based on extreme value theory (EVT), exploiting the properties of Generalized Extreme Value (GEV) distribution. The novelty lies in the application of extreme value theory over the tail of the complete score distribution (genuine and impostor scores), assuming that the genuine scores form extreme values (tail) with respect to the entire set of scores. Normalization is then performed by estimating the cumulative density function of GEV distribution, using the parameter set obtained from genuine data. For evaluation, the proposed method is compared with state-of-the-art methods on two publicly available multimodal databases: i) NIST BSSR1 [22] multimodal biometric score database and ii) Database created from Face Recognition Grand Challenge V2.0 [23] and LG4000 iris images [24], to show the efficiency of the proposed method. Renu Sharma, Sukhendu Das, Padmaja Joshi |
BMVC | 2 |
| 2015 | Unsupervised domain adaptation using eigenanalysis in kernel space for categorisation tasksabstractThis study describes a new technique of unsupervised domain adaptation based on eigenanalysis in kernel space, for the purpose of categorisation tasks. The authors propose a transformation of data in source domain, such that the eigenvectors and eigenvalues of the transformed source domain become similar to that of the target domain. They extend this idea to the reproducing kernel Hilbert space, which enables to deal with non‐linear transformation of source domain. They also propose a measure to obtain the appropriate number of eigenvectors needed for transformation. Results on object, video and text categorisations tasks using real‐world datasets show that the proposed method produces better results when compared with a few recent state‐of‐art methods of domain adaptation. Suranjana Samanta, Sukhendu Das |
IET Image Process. | 2 |
| 2014 | Modeling Sequential Domain Shift through Estimation of Optimal Sub-spaces for Categorization
Suranjana Samanta, Tirumarai Selvan, Sukhendu Das |
BMVC | 3 |
| 2014 | Dictionary based framework for face recognition, designed mutually for single training sample (STS) and degraded set (DS)abstractAvailability of a single training sample (STS) or degraded set (DS) of training and testing samples restricts the success of face recognition in real-world applications. We propose a unified framework for handling both these challenges simultaneously by using a data dictionary, which is a combination of training dictionary and intra-class variation dictionary. The training dictionary is assembled by the single representative sample per class. Variations between the training samples and a query image are captured by the intra-class variation dictionary. Misalignment of the query image is handled by aligning it with respect to the representative samples. A few moderately aligned and warped face images obtained from the query image are then sparsely represented using the data dictionary with additional constraints on their variance which reduces the obligation of a perfectly aligned query image. The experiments results on AR and LFW datasets validate our claim of superior performance in STS and DS as compared to the other recent methods. Renu Sharma, Sukhendu Das, Padmaja Joshi |
IJCB | 2 |
| 2014 | Hierarchy of visual features for object recognitionabstractMost approaches for object recognition (OR) use a single feature descriptor to identify the object class from a query image. However, specifically in case of variations in appearance, scale and illumination, the performance of features not only vary depending on the class, but also on the query sample. We propose a biological inspired framework for OR using concepts from feature integration theory (FIT). Our model uses a hierarchy of visual features for OR. The key components in the proposed approach are: (i) SALCUT - unsupervised segmentation for salient object localization; (ii) optimal feature selection - identify appropriate features for each class, at each level of feature hierarchy, for a test instance; (iii) feature combination - which happens at higher levels of feature hierarchy, if features selected at the lower level are unable to classify a test instance. Our method outperforms several state-of-the-art techniques, when validated using two real-world datasets. Nitin Gupta 0005, Sukhendu Das, Sutanu Chakraborti |
ICIP | 2 |
| 2014 | Unsupervised domain adaptation using manifold alignment for object and event categorizationabstractThis paper describes a method of cross-domain object and event categorization, using the concept of domain adaptation. Here, a classifier is trained using samples from the source/ auxiliary domain and performance is observed on a set of test samples taken from a different domain, termed as the target domain. To overcome the difference between the two domains, we aim to find an optimal sub-space such that the instances from both the domains follow similar distributions when projected onto the sub-space. Along with the distributions, the underlying manifolds of the two domains are aligned in the sub-space to reduce the difference in structure of the data from the two domains. The local spatial arrangement of the instances in both the domains are also preserved in the optimal sub-space. Results show that the proposed method of unsupervised domain adaptation provides better classification accuracy than a few state of the art methods. Suranjana Samanta, Sukhendu Das |
ICIP | 2 |
| 2014 | Physics based virtual cutting using j-integral method for gaming applicationsabstractPowerful graphic cards have enabled the game engine developers to add deformable assets. Many games require the players to cut/chop/slash game assets. To render interaction of deformable assets with sharp weapons they use pre-defined fracture patterns. These pre-defined fracture patterns are used to break/cut objects and the use of physics is limited due to computational costs of the virtual cutting process. In this work, we present a low cost solution for performing physics based virtual cutting on deformable assets. Our aim is to provide a highly tunable physics based virtual cutting algorithm on GPU to meet the varying needs of a game engine. Prateek Shrivastava, Sukhendu Das |
MIG | 2 |
| 2013 | Domain Adaptation Based on Eigen-Analysis and Clustering, for Object Categorization
Suranjana Samanta, Sukhendu Das |
CAIP (1) | 2 |
| 2012 | Enhancing the MST-CSS Representation Using Robust Geometric Features, for Efficient Content Based Video Retrieval (CBVR)abstractMulti-Spectro-Temporal Curvature Scale Space (MST-CSS) had been proposed as a video content descriptor in an earlier work, where the peak and saddle points were used for feature points. But these are inadequate to capture the salient features of the MST-CSS surface, producing poor retrieval results. To overcome these, we propose EMST-CSS (Enhanced MST-CSS) as a better feature representation with an improved matching method for CBVR (Content Based Video Retrieval). Comparative study with the existing MST-CSS representation and two state-of-the-art methods for CBVR shows enhanced performance on one synthetic and two real-world datasets. Chiranjoy Chattopadhyay, Sukhendu Das |
ISM | 2 |
| 2012 | A Motion-Sketch Based Video Retrieval Using MST-CSS RepresentationabstractIn this work, we propose a framework for a robust Content Based Video Retrieval (CBVR) system with free hand query sketches, using the Multi-Spectro Temporal-Curvature Scale Space (MST-CSS) representation. Our designed interface allows sketches to be drawn to depict the shape of the object in motion and its trajectory. We obtain the MST-CSS feature representation using these cues and match with a set of MST-CSS features generated offline from the video clips in the database (gallery). Results are displayed in rank ordered similarity. Experimentation with benchmark datasets shows promising results. Chiranjoy Chattopadhyay, Sukhendu Das |
ISM | 2 |
| 2011 | Use of Salient Features for the Design of a Multistage Framework to Extract Roads From High-Resolution Multispectral Satellite ImagesabstractThe process of road extraction from high-resolution satellite images is complex, and most researchers have shown results on a few selected set of images. Based on the satellite data acquisition sensor and geolocation of the region, the type of processing varies and users tune several heuristic parameters to achieve a reasonable degree of accuracy. We exploit two salient features of roads, namely, distinct spectral contrast and locally linear trajectory, to design a multistage framework to extract roads from high-resolution multispectral satellite images. We trained four Probabilistic Support Vector Machines separately using four different categories of training samples extracted from urban/suburban areas. Dominant Singular Measure is used to detect locally linear edge segments as potential trajectories for roads. This complimentary information is integrated using an optimization framework to obtain potential targets for roads. This provides decent results in situations only when the roads have few obstacles (trees, large vehicles, and tall buildings). Linking of disjoint segments uses the local gradient functions at the adjacent pair of road endings. Region part segmentation uses curvature information to remove stray nonroad structures. Medial-Axis-Transform-based hypothesis verification eliminates connected nonroad structures to improve the accuracy in road detection. Results are evaluated with a large set of multispectral remotely sensed images and are compared against a few state-of-the-art methods to validate the superior performance of our proposed method. Sukhendu Das, T. T. Mirnalinee, Koshy Varghese |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2010 | Estimation of orientation of a textured planar surface using projective equations and separable analysis with M-channel wavelet decomposition
Thomas Greiner, Shivani G. Rao, Sukhendu Das |
Pattern Recognit. | 3 |
| 2010 | MST-CSS (Multi-Spectro-Temporal Curvature Scale Space), a Novel Spatio-Temporal Representation for Content-Based Video RetrievalabstractWe present a novel spatio-temporal descriptor to efficiently represent a video object for the purpose of content-based video retrieval. Features from spatial along with temporal information are integrated in a unified framework for the purpose of retrieval of similar video shots. A sequence of orthogonal processing, using a pair of 1-D multiscale and multispectral filters, on the space-time volume (STV) of a video object (VOB) produces a gradually evolving (smoother) surface. Zero-crossing contours (2-D) computed using the mean curvature on this evolving surface are stacked in layers to yield a hilly (3-D) surface, for a joint multispectro-temporal curvature scale space (MST-CSS) representation of the video object. Peaks and valleys (saddle points) are detected on the MST-CSS surface for feature representation and matching. Computation of the cost function for matching a query video shot with a model involves matching a pair of 3-D point sets, with their attributes (local curvature), and 3-D orientations of the finally smoothed STV surfaces. Experiments have been performed with simulated and real-world video shots using precision-recall metric for our performance study. The system is compared with a few state-of-the-art methods, which use shape and motion trajectory for VOB representation. Our unified approach has shown better performance than other approaches that use combined match-costs obtained with separate shape and motion trajectory representations and our previous work on a simple joint spatio-temporal descriptor (3-D-CSS). A. Dyana, Sukhendu Das |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Unsupervised texture segmentation using feature selection and fusionabstractThis paper describes a method of unsupervised color texture segmentation by efficiently combining different features obtained from multi-channel and multi-resolution filters. The DWT and DCT features are extracted separately from 3 color bands of the image and then fused together for optimal performance. The features are then ranked according to a selection criteria. We propose a new correlation measure for the task of feature ranking. To select the best combination of features to be used, we use the property of cluster scatter of a selected set of features. Finally, the optimum number of ranked order features are used for segmentation using a fuzzy C-Means classifier. The performance of the proposed segmentation method is verified using standard benchmark datasets. Suranjana Samanta, Sukhendu Das |
ICIP | 2 |
| 2009 | Trajectory representation using Gabor features for motion-based video retrieval
A. Dyana, Sukhendu Das |
Pattern Recognit. Lett. | 2 |
| 2008 | Symmetry-based face pose estimation from a single uncalibrated viewabstractIn this paper, a geometric method for estimating the face pose (roll and yaw angles) from a single uncalibrated view is presented. The symmetric structure of the human face is exploited by taking the mirror image (horizontal flip) of a test face image as a virtual second view. Facial feature point correspondences are established between the given test and its mirror image using an active appearance model. Thus, the face pose estimation problem is cast as a two-view rotation estimation problem. By using the bilateral symmetry, roll and yaw angles are estimated without the need for camera calibration. The proposed pose estimation method is evaluated on synthetic and natural face datasets, and the results are compared with an eigenspace-based method. It is shown that the proposed symmetry-based method shows performance that is comparable to the eigenspace-based method for both synthetic and real face image datasets. Vinod Pathangay, Sukhendu Das, Thomas Greiner |
FG | 2 |
| 2008 | Integrating region and edge information for texture segmentation using a modified constraint satisfaction neural network
Lalit Gupta, Utthara Gosa Mangai, Sukhendu Das |
Image Vis. Comput. | 3 |
| 2008 | Enhancing decision combination of face and fingerprint by exploitation of individual classifier space: An approach to multi-modal biometry
Arpita Patra, Sukhendu Das |
Pattern Recognit. | 2 |
| 2007 | System-on-programmable-chip implementation for on-line face recognition
A. Pavan Kumar, V. Kamakoti 0001, Sukhendu Das |
Pattern Recognit. Lett. | 3 |
| 2006 | Generic Object Recognition Using a Combination of ICA and Shape CuesabstractThis paper addresses the problem of Generic Object Recognition by modeling the perceptual capability of human beings. In contrast to the traditional approaches, we have approached the recognition problem by proposing a framework which involves two stages of processing. First, an intelligent generic recognizer based on independent component analysis (ICA) is employed to reduce the search space to a few rank-ordered samples. It is shown that ICA captures the appearance characteristics of objects. Shape cues (distance transform based matching) are then used to verify the result of the appearance-based classifier and identify the correct object class and pose. Experiments were conducted using objects with complex appearance and shape characteristics. Sensitivity of recognition to the number of independent components and number of learning samples is analyzed on COIL-100 database. The performance of the generic classifier using ICA with and without shape matching is also analyzed. Manisha Kalra, Sukhendu Das, Amitava Datta |
AVSS | 2 |
| 2004 | Face Recognition Using Weighted Modular Principle Component Analysis
A. Pavan Kumar, Sukhendu Das, V. Kamakoti 0001 |
ICONIP | 2 |