Santanu Chaudhury

dblp:31/2344 · DBLP profile ↗
← Back
152ranked-venue papers
8as first author
26since 2021 · last 2025
0000-0002-5488-7773ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 106 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 36 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Survival Prediction in Lung Cancer through Multi-Modal Representation Learning
abstract
Survival prediction is a crucial task associated with cancer diagnosis and treatment planning. This paper presents a novel approach to survival prediction by harnessing comprehensive informationfrom CT and PET scans, along with associated Genomic data. Current methods rely on either a single modality or the integration of multiple modalities for prediction without adequately addressing associations across patients or modalities. We aim to develop a robust predictive model for survival outcomes by integrating multimodal imaging data with genetic information while accounting for associations across patients and modali-ties. We learn representations for each modality via a self-supervised module and harness the semantic similarities across the patients to ensure the embeddings are aligned closely. However, optimizing solely for global relevance is inadequate, as many pairs sharing similar high-level semantics, such as tumor type, are inadvertently pushed apart in the embedding space. To address this issue, we use a cross-patient module (CPM) designed to harness inter-subject correspondences. The CPM module aims to bring together embeddings from patients with similar disease characteristics. Our experimental evaluation of the dataset of Non-Small Cell Lung Cancer (NSCLC) patients demonstrates the effectiveness of our approach in predicting survival outcomes, outperforming state-of-the-art methods.
Aiman Farooq, Deepak Mishra 0003, Santanu Chaudhury
WACV3
2024 LISR: Learning Linear 3D Implicit Surface Representation Using Compactly Supported Radial Basis Functions
abstract
Implicit 3D surface reconstruction of an object from its partial and noisy 3D point cloud scan is the classical geometry processing and 3D computer vision problem. In the literature, various 3D shape representations have been developed, differing in memory efficiency and shape retrieval effectiveness, such as volumetric, parametric, and implicit surfaces. Radial basis functions provide memory-efficient parameterization of the implicit surface. However, we show that training a neural network using the mean squared error between the ground-truth implicit surface and the linear basis-based implicit surfaces does not converge to the global solution. In this work, we propose locally supported compact radial basis functions for a linear representation of the implicit surface. This representation enables us to generate 3D shapes with arbitrary topologies at any resolution due to their continuous nature. We then propose a neural network architecture for learning the linear implicit shape representation of the 3D surface of an object. We learn linear implicit shapes within a supervised learning framework using ground truth Signed-Distance Field (SDF) data for guidance. The classical strategies face difficulties in finding linear implicit shapes from a given 3D point cloud due to numerical issues (requires solving inverse of a large matrix) in basis and query point selection. The proposed approach achieves better Chamfer distance and comparable F-score than the state-of-the-art approach on the benchmark dataset. We also show the effectiveness of the proposed approach by using it for the 3D shape completion task.
Atharva Pandey, Vishal Yadav, Rajendra Nagar, Santanu Chaudhury
AAAI4
2024 GDP: Generic Document Pretraining to Improve Document Understanding
Akkshita Trivedi, Akarsh Upadhyay, Rudrabha Mukhopadhyay, Santanu Chaudhury
ICDAR (1)4
2024 Pneumonia Classification in Chest X-Ray Images Using Explainable Slot-Attention Mechanism
Shipra Madan, Santanu Chaudhury, Tapan Kumar Gandhi
ICPR (5)2
2024 Neural Encoding of Odors: Translating Odors into Unique Digital Representation with EEG Signals
Archana Yadav, Vishakha Pareek, Akshay Agarwal 0001, Santanu Chaudhury
ICPR (7)4
2024 Multi-armed bandit based online model selection for concept-drift adaptation
abstract
Abstract Ensemble methods are among the most effective concept‐drift adaptation techniques due to their high learning performance and flexibility. However, they are computationally expensive and pose a challenge in applications involving high‐speed data streams. In this paper, we present a computationally efficient heterogeneous classifier ensemble entitled OMS‐MAB which uses online model selection for concept‐drift adaptation by posing it as a non‐stationary multi‐armed bandit (MAB) problem. We use a MAB to select a single adaptive learner within the ensemble for learning and prediction while systematically exploring promising alternatives. Each ensemble member is made drift resistant using explicit drift detection and is represented as an arm of the MAB. An exploration factor controls the trade‐off between predictive performance and computational resource requirements, eliminating the need to continuously train and evaluate all the ensemble members. A rigorous evaluation on 20 benchmark datasets and 9 algorithms indicates that the accuracy of OMS‐MAB is statistically at par with state‐of‐the‐art (SOTA) ensembles. Moreover, it offers a significant reduction in execution time and model size in comparison to several SOTA ensemble methods, making it a promising ensemble for resource constrained stream‐mining problems.
Jobin Wilson, Santanu Chaudhury, Brejesh Lall
Expert Syst. J. Knowl. Eng.2
2024 Multimodal biometric user authentication using improved decentralized fuzzy vault scheme based on Blockchain network
Shreyansh Sharma, Anil K. Saini, Santanu Chaudhury
J. Inf. Secur. Appl.3
2023 On AI-Assisted Pneumoconiosis Detection from Chest X-rays
abstract
According to theWorld Health Organization, Pneumoconiosis affects millions of workers globally, with an estimated 260,000 deaths annually. The burden of Pneumoconiosis is particularly high in low-income countries, where occupational safety standards are often inadequate, and the prevalence of the disease is increasing rapidly. The reduced availability of expert medical care in rural areas, where these diseases are more prevalent, further adds to the delayed screening and unfavourable outcomes of the disease. This paper aims to highlight the urgent need for early screening and detection of Pneumoconiosis, given its significant impact on affected individuals, their families, and societies as a whole. With the help of low-cost machine learning models, early screening, detection, and prevention of Pneumoconiosis can help reduce healthcare costs, particularly in low-income countries. In this direction, this research focuses on designing AI solutions for detecting different kinds of Pneumoconiosis from chest X-ray data. This will contribute to the Sustainable Development Goal 3 of ensuring healthy lives and promoting well-being for all at all ages, and present the framework for data collection and algorithm for detecting Pneumoconiosis for early screening. The baseline results show that the existing algorithms are unable to address this challenge. Therefore, it is our assertion that this research will improve state-of-the-art algorithms of segmentation, semantic segmentation, and classification not only for this disease but in general medical image analysis literature.
Yasmeena Akhter, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa, Santanu Chaudhury
IJCAI5
2023 Dr. MTL: Driver Recommendation using Federated Multi-Task Learning
abstract
Online vehicle-for-hire service firms such as Ola, Uber, Gett, Lyft, Hailo, Didi, and GrabCab rely on driver recommendations. Several factors should be considered when suggesting a driver for the trip, including ride cost, driver-vehicle-passenger safety, reliability, robustness, and comfort. Multi-task learning has become more widespread in recommender systems in recent years. Our study employs federated multi-objective multitask learning to develop a driver recommendation system that considers drivers’ cognitive-physical stress and driving behavior while respecting users’ privacy. We use publicly available datasets from UAH-DriveSet, HCI Lab, and PhysioNet for training, testing, and evaluating our system. Our main contribution is a unique framework for driver suggestion based on data-driven federated multi-task learning that achieves the best F-measure accuracy. Our proposed system outperforms baseline and state-of-the-art deep learning models. It detects driver stress and behavior with 95% and 96% F-measures, respectively.
Jayant Vyas, Bhumika, Debasis Das 0001, Santanu Chaudhury
VTC Fall4
2023 AnoLeaf: Unsupervised Leaf Disease Segmentation via Structurally Robust Generative Inpainting
abstract
Plant diseases severely limits agriculture production, necessitating the high-throughput monitoring of plant leaves. Currently, this is formulated as an automatic disease segmentation task addressed via deep learning frameworks. These deep leaning frameworks trained with leaf image data in a supervised paradigm have few limitations, mainly: (1) training datasets are heavily imbalanced towards healthy leaf images, (2) disease region annotation is labour-intensive and (3) due to the heterogeneity of disease symptoms, these frameworks lacks generalisability. In this paper, we reformulate disease segmentation as an anomaly localisation task. Specifically, we introduce a novel unsupervised framework (AnoLeaf) based on an edge-guided in-painting that optimises the learning of contextual attention on only healthy leaf images. The network utilisation on diseased leaf images results in reconstruction of its healthy counterparts, generating an inpainting error. The contextual attention maps reinforce the inpainting error to effectively localise the disease. Thus, AnoLeaf alleviates the acquisition and annotation of rare disease images. Additional experiments on MVTec anomaly detection dataset further demonstrate its generalisability.
Swati Bhugra, Vinay Kaushik, Brejesh Lall, Santanu Chaudhury
WACV5
2023 A survey on biometric cryptosystems and their applications
Shreyansh Sharma, Anil K. Saini, Santanu Chaudhury
Comput. Secur.3
2023 Towards explainable deep visual saliency models
Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury
Comput. Vis. Image Underst.4
2023 Federated learning based driver recommendation for next generation transportation system
Jayant Vyas, Bhumika, Debasis Das 0001, Santanu Chaudhury
Expert Syst. Appl.4
2023 Homogeneous-Heterogeneous Hybrid Ensemble for concept-drift adaptation
Jobin Wilson, Santanu Chaudhury, Brejesh Lall
Neurocomputing2
2023 Explainable few-shot learning with visual explanations on a low resource pneumonia dataset
Shipra Madan, Santanu Chaudhury, Tapan Kumar Gandhi
Pattern Recognit. Lett.2
2022 Lighter and Faster Two-Pathway CMRNet for Video Saliency Prediction
abstract
Existing dynamic saliency prediction models face challenges like inefficient spatio-temporal feature integration, ineffective multi-scale feature extraction, and lacking domain adaptation because of huge pre-trained backbone networks. In this paper, we propose a two pathway architecture with effective feature integration of spatial and temporal domains at multiple scales for video saliency prediction. Frame and optical flow pathways extract features from video frame and optical flow maps, respectively using a series of cross-concatenated multi-scale residual (CMR) blocks. We name this network as two-pathway CMRNet (TP-CMRNet). Every CMR block follows a feature fusion and attention module for merging features from two pathways and guiding the network to weigh salient regions, respectively. A bi-directional LSTM module is used for learning the task by looking at previous and next video frames. We build a simple decoder for feature reconstruction into the final attention map. TP-CMRNet is comprehensively evaluated using three benchmark datasets: DHF1K, Hollywood-2, and UCF sports. We observe that our model performs at par with other deep dynamic models. In particular, we outperform all the other models with a lesser number of model parameters and lower inference time.
Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury
ICIP4
2022 Multiresolution visual enhancement of hazy underwater scene
Deepak Kumar Rout, Badri N. Subudhi, Veerakumar Thangaraj, Santanu Chaudhury, John J. Soraghan
Multim. Tools Appl.4
2022 Multi-task driven explainable diagnosis of COVID-19 using chest X-ray images
Aakarsh Malhotra, Surbhi Mittal, Puspita Majumdar, Saheb Chhabra, Kartik Thakral, Mayank Vatsa, Richa Singh 0001, Santanu Chaudhury, Ashwin Pudrod, Anjali Agrawal
Pattern Recognit.8
2022 Temporal Causal Modelling on Large Volume Enterprise Data
abstract
Structural Causal Modelling (SCM) with its intervention analysis is one of the promising modelling approach that assists in data driven decision making. SCM not only overcomes the black box modelling associated with most of the classification algorithms but also gives enterprises an opportunity to perform intervention analysis without having to perform randomized controlled experiments. But the large volume of enterprises' data pose challenges in learning the causal structure as existing algorithms are not suitable to learn from data present in Distributed File System (DFS). Hence algorithm presented in this paper, proposes a novel variation to PC-Stable algorithm to efficiently learn the causal structure from data present in DFS - thus enabling temporal causal modelling on large volume time-series data. The proposed learning algorithm is used to determine the causal story associated with churn in telecommunication industry and flight delay in airline industry. Our model identifies and quantifies the respective causal factors for unfavourable events churn and flight delay.
Ram Mohan, Santanu Chaudhury, Brejesh Lall
IEEE Trans. Big Data2
2022 DriveBFR: Driver Behavior and Fuel-Efficiency-Based Recommendation System
abstract
Despite the tremendous growth of the transportation sector, the availability of systems that ensure safe, efficient, sustainable transportation reduces traffic congestion, maintenance costs, the off-road time of the vehicle, enhances driver’s experiences, and ensures a more reliable journey are very limited. The fast evolution of our economy, lack of driver training, and the grown affordability of our society are reasons for this mismatch in developing economies. We think that the inconsistency will increase and unfavorably affect our traffic structure unless intelligent algorithm-based solutions are developed and deployed. This article presents a system for providing safe, accurate, comfortable, reliable, fuel-efficient, and economical driving behavior using Machine Learning techniques like the hidden Markov model (HMM). Our proposed system recommends subsequent trips using a multi-objective optimization (MOO) technique for the driver. It provides suggestions concerning speed limits and alerts based on the driver’s behavior score and fuel efficiency. We used a publicly available UAH-DriveSet dataset captured by the driving monitoring app DriveSafe for all of our experiments. The results reveal that the proposed model predicts behavior with 95% accuracy and calculates fuel efficiency to improve driving quality and experience. This system recommends safer, more comfortable, more reliable, more efficient, and economical rides, beneficial for everyone in our society.
Jayant Vyas, Debasis Das 0001, Santanu Chaudhury
IEEE Trans. Comput. Soc. Syst.3
2022 Block Sparse Variational Bayes Regression Using Matrix Variate Distributions With Application to SSVEP Detection
abstract
Due to the nonsparse representation, the use of compressed sensing (CS) for physiological signals, such as a multichannel electroencephalogram (EEG), has been a challenge. We present a generalized Bayesian CS framework that is capable of handling representations that arise in the spatiotemporal setting. The proposed model utilizes the standard linear Gaussian observation model associated with the hierarchical modeling of data using the matrix-variate Gaussian scale mixture (GSM). It deploys various random and deterministic parameters to incorporate the knowledge of spatial and temporal correlation present in data. By varying distributions over random parameters, a family of generalized hyperbolic matrix variate distributions is derived. For estimation, we rely on variational Bayes (VB) for random parameters and expectation-maximization (EM) for deterministic parameters. Furthermore, the model is compared with recent developments in matrix-variate distribution-based modeling of data, and we briefly discuss its extension to finite mixtures of skewed distributions. Finally, the framework is applied to the steady-state visual evoked potential (SSVEP)-based EEG benchmark data set, and a comparative study is conducted to show its effectiveness for the frequency detection task. One of the crucial features of the proposed model is that it simultaneously processes multichannel signals with low computational cost and time, making it suitable for real-time systems, especially in a resource-constrained environment.
Santanu Chaudhury, Jayadeva
IEEE Trans. Neural Networks Learn. Syst.2
2021 Video Classification using SlowFast Network via Fuzzy rule
abstract
Anomalous events occur rarely and are challenging to model. Therefore, automatic recognition of abnormal activities in surveillance videos is a non-trivial task. Though with the availability of video datasets of abnormal activities, there has been some progress, recognition of abnormal activities in real-time with high confidence remains unsolved. Existing video-based anomaly detection techniques using traditional machine learning and deep-learning are compute-intensive and give low recognition accuracy. This paper presents a robust and computationally efficient deep learning-based framework to recognize different real-world anomalies from the video. The proposed scheme uses a Fuzzy rule to summarize the video to scale the problem into fewer frames and the slow-fast neural network for classification. Intuitively, the designed pipeline aims to solve two significant problems that arise with video classification; one is to reduce the redundant frames and avoid the computation of optical flow for a video that has a substantial computational requirement. The proposed scheme tested on the UCF-crime dataset and has achieved recognition accuracy of 53%.
Rituraj, Aruna Tiwari, Santanu Chaudhury, Sanjay Singh 0001, Sumeet Saurav
FUZZ-IEEE3
2021 Lighter and Faster Cross-Concatenated Multi-Scale Residual Block Based Network for Visual Saliency Prediction
abstract
Existing deep architectures for visual saliency prediction face problems like inefficient feature encoding, larger inference times, and a huge number of model parameters. One possible solution is to make the local and global contextual feature extraction computationally less intensive by a novel lighter architecture. In this work, we propose an end-to-end learnable, inter-scale information sharing residual block based architecture for saliency prediction. A series of these blocks are used for efficient multi-scale feature extraction followed by a dilated inception module (DIM) and a novel decoder. We name this network as cross-concatenated multi-scale residual (CMR) block based network, CMRNet. We comprehensively evaluate our architecture on three datasets: SALICON, MIT1003, and MIT300. Experimental results show that our model works at par with other state-of-the-art models. Especially, our model outperforms all the other models with a smaller inference time and a lesser number of model parameters.
Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury
ICIP4
2021 Automated detection of COVID-19 on a small dataset of chest CT images using metric learning
abstract
Coronavirus disease has caused unprecedented chaos across the globe causing potentially fatal pneumonia, since the beginning of 2020. Researchers from different communities are working in conjunction with front-line doctors and policymakers to better understand the disease. The key to prevent the spread is a rapid diagnosis, prioritized isolation, and fastidious contact tracing. Recent studies have confirmed the presence of underlying patterns on chest CT for patients with COVID-19. We present a completely automated framework to detect COVID-19 using chest CT scans, only needing a small number of training samples. We present a few-shot learning technique based on the Triplet network in comparison to the conventional deep learning techniques which require a substantial amount of training examples. We used 140 chest CT images for training and the rest for testing from a total of 2482 images for both COVID-19 and non-COVID-19 cases from a publicly available dataset. The model trained with chest CT images achieves an AUC of 0.94, separates the two classes into distinct clusters; thereby giving correct prediction accuracy on the evaluation dataset.
Shipra Madan, Santanu Chaudhury, Tapan Kumar Gandhi
IJCNN2
2021 An Ontology Representation Language for Multimedia Event Applications
abstract
This paper presents formalization of a new Multimedia Web Ontology Language (E-MOWL) to handle events with media depictions. The temporal, spatial and entity aspects that are implicitly linked to an event are represented through this language to model the context of events. The already existing Multimedia Web Ontology Language (MOWL) can be leveraged for perceptual modelling of a domain, where the concepts manifest into media patterns in the multimedia document and helps in semantic processing of the contents. The language E-MOWL provides a rich method for representing knowledge corresponding to a specific domain wherein the context specifies the intended meaning of each element of the domain of discourse; an element in different context may correspond to different functional role. The context information associated with an event ties the audiovisual data with event related aspects. All these aspects when considered altogether provide the evidence and contribute towards recognizing an event from multimedia documents. The language also enables reasoning with the uncertainty associated with the events and is organized in the form of Bayesian Network (BN). The media items that are semantically relevant can be assimilated together on the basis of their association with events. We have demonstrated the efficacy of our approach by utilizing an ontology for the entertainment category in news domain to offer an application \textit{news aggregation} and event-based book recommendations.
Nisha Pahal, Brejesh Lall, Santanu Chaudhury
J. Web Eng.3
2021 Deep learning-based gas identification and quantification with auto-tuning of hyper-parameters
Vishakha Pareek, Santanu Chaudhury
Soft Comput.2
2020 A Hierarchical Framework for Leaf Instance Segmentation: Application to Plant Phenotyping
abstract
Image based analysis of plants is a high-throughput and non-invasive approach to study plant traits. The quantitative estimation of many plant traits (leaf area index, biomass etc.) from plant images is primarily based on accurate segmentation of individual leaves. This is a challenging task due to the presence of overlapped leaves and lack of discernible boundaries between them. To overcome these limitations, state-of-the-art supervised deep learning algorithms have been recently employed. However, the annotations of individual leaf instances is time consuming, in addition the variability in leaf shapes and its arrangement among different plant species limits the broad utilisation of these algorithms. To relieve this bottleneck, we propose a novel framework that relies on a graph based formulation to extract leaf shape knowledge for the task of leaf instance segmentation. These shape priors are generated based on leaf shape characteristics independent of plant species. Evaluation of the proposed framework on multiple plant datasets i.e. Arabidopsis, Komatsuna and salad demonstrates its broad utility.
Swati Bhugra, Kanish Garg, Santanu Chaudhury, Brejesh Lall
ICPR3
2020 Collaborative Human Machine Attention Module for Character Recognition
abstract
The deep learning models, which include attention mechanisms, are shown to enhance the performance and efficiency of the various computer vision tasks such as pattern recognition, object detection, face recognition, etc. Although the visual attention mechanism is the source of inspiration for these models, recent attention models consider 'attention' as a pure machine vision optimization problem, and visual attention remains the most neglected aspect. Therefore, this paper presents a collaborative human and machine attention module which considers both visual and network's attention. The proposed module is inspired by the dorsal (`where') pathways of visual processing and can be integrated with any convolutional neural network (CNN) model. First, the module computes the spatial attention map from the input feature maps, which is then combined with the visual attention maps. The visual attention maps are created using eye-fixations obtained by performing an eye-tracking experiment with human participants. The visual attention map covers the highly salient and discriminating image regions as humans tend to focus on such regions, whereas the other relevant image regions are processed by spatial attention map. The combination of these two maps results in the finer refinement in feature maps, resulting in improved performance. The comparative analysis reveals that our model not only shows significant improvement over the baseline model but also outperforms the other models. We hope that our findings using a collaborative human-machine attention module will be helpful in other computer vision tasks as well.
Chetan Ralekar, Tapan Kumar Gandhi, Santanu Chaudhury
ICPR3
2020 Using Scene Graphs for Detecting Visual Relationships
abstract
In this paper we solve the problem of detecting relationships between pairs of objects in an image. We develop spatially aware word embeddings using scene graphs and use joint feature representations containing visual, spatial and semantic embeddings from the input images to train a deep network on the task of relationship detection. Further, we propose to utilize context aligned scene graph embeddings from the train set, without requiring explicit availability of scene graphs at test time. We show that the proposed method outperforms the state-of-the-art methods for predicate detection and provides competing results on relationship detection. We also show the generalization ability of the proposed method by performing predictions under zero shot settings. Further, we also provide an exhaustive empirical evaluation on each component of the proposed network.
Anurag Tripathi, Siddharth Srivastava 0004, Brejesh Lall, Santanu Chaudhury
ICPR4
2020 Eye Movement State Trajectory Estimator based on Ancestor Sampling
abstract
Human gaze dynamics mainly concern about the sequence of the occurrence of three eye movements: fixations, saccades, and microsaccades. In this paper, we correlate them as three different states to velocities of eye movements. We build a state trajectory estimator based on ancestor sampling (ST EAS) model, which captures the features of the human temporal gaze pattern to identify the kind of visual stimuli. We used a gaze dataset of 72 viewers watching 60 video clips which are equally split into four visual categories. Uniformly sampled velocity vectors from the training set, are used to find the best suitable parameters of the proposed statistical model. Then, the optimized model is used for both gaze data classification and video retrieval on the test set. We observed 93.265% of classification accuracy and a mean reciprocal rank of 0.888 for video retrieval on the test set. Hence, this model can be used for viewer independent video indexing for providing viewers an easier way to navigate through the contents.
Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury
MMSP4
2020 Automatic lecture video skimming using shot categorization and contrast based features
Badri N. Subudhi, Veerakumar Thangaraj, Esakkirajan Sankaralingam, Santanu Chaudhury
Expert Syst. Appl.4
2020 Walsh-Hadamard-Kernel-Based Features in Particle Filter Framework for Underwater Object Tracking
abstract
One of the well-established research domains among computer vision scientists is object tracking. However, not much work has been done in underwater scenarios. This article addresses the problem of visual tracking in the underwater environment with the stationary and nonstationary camera setups. In order to deal with the underwater optical dynamics, a dominant color component-based scene representation is employed in the YCbCr color space. An adaptive approach is devised to select the Walsh-Hadamard (WH) kernels for the efficient extraction of color, edge, and texture strengths, whereas a new feature called range strength is proposed to extract the variation of intensity from underwater sequences in the local neighborhood using the WH kernel. The likelihood of these feature strengths is integrated in a particle filter framework to track the object of interest in underwater sequences. The reference feature strengths used in assigning weights to the particles are updated based on the S$\phi$rensen distance. The coefficients of feature strengths are calculated in such a way that if one feature fails, then its coefficient become insignificant, whereas the more suitable features get higher feature coefficients. The effectiveness of the proposed scheme is evaluated using the underwater video datasets: reefVid, fish4knowledge (F4K), underwaterchangedetection (UWCD), and National Oceanic and Atmospheric Administration (NOAA). The performance evaluation is performed by comparing the scheme with five recent state-of-the-art tracking schemes. The quantitative analysis of the proposed scheme is carried out using three evaluation measures: overall intersection over union, centroid location error, and average tracking error. The performance of the proposed scheme is quite encouraging in the case of sequences with hazy and degraded, partially occluded, and camouflaged challenges.
Deepak Kumar Rout, Badri N. Subudhi, Veerakumar Thangaraj, Santanu Chaudhury
IEEE Trans. Ind. Informatics4
2019 An End-to-End Trainable Framework for Joint Optimization of Document Enhancement and Recognition
abstract
Recognizing text from degraded and low-resolution document images is still an open challenge in the vision community. Existing text recognition systems require a certain resolution and fails if the document is of low-resolution or heavily degraded or noisy. This paper presents an end-to-end trainable deep-learning based framework for joint optimization of document enhancement and recognition. We are using a generative adversarial network (GAN) based framework to perform image denoising followed by deep back projection network (DBPN) for super-resolution and use these super-resolved features to train a bidirectional long short term memory (BLSTM) with Connectionist Temporal Classification (CTC) for recognition of textual sequences. The entire network is end-to-end trainable and we obtain improved results than state-of-the-art for both the image enhancement and document recognition tasks. We demonstrate results on both printed and handwritten degraded document datasets to show the generalization capability of our proposed robust framework.
Anupama Ray, Avinash Upadhyay, Megh Makwana, Santanu Chaudhury, Akkshita Trivedi, Ajay Pratap Singh, Anil K. Saini
ICDAR5
2019 Guided Compositional Generative Adversarial Networks
abstract
In this paper, we propose to synthesize natural images from a set of input objects. The proposed technique generates a scene which has high correlation with the provided set of input objects while also maintaining the natural placement of objects within the scene. The technique constitutes of a generative adversarial network trained on a large corpus of objects and natural scenes. This is in contrast with earlier works where the objective was to generate a natural scene from a noise vector or conditioning the network over a variable. However, such methods have limitations in their ability to control the objects within the generated images. On the contrary, we show that by training a Generative Adversarial Network with raw image pixels as input, we can generate scenes which constitute the objects as well as generate the surrounding environment suitable for the combination of the input objects. We provide qualitative and quantitative results on challenging MS-COCO dataset to show the effectiveness of the proposed technique.
Anurag Tripathi, Siddharth Srivastava 0004, Brejesh Lall, Santanu Chaudhury
SMC4
2019 Memorability-based image compression
abstract
This study is concerned with achieving the image compression using the concept of memorability. The authors have used memorability of an image, as a perceptual measure while image coding. In the proposed approach, a region‐of‐interest‐based memorability preserving image compression algorithm which is accomplished via two sub‐processes namely, memorability prediction and image compression is introduced. The memorability of images is predicted using convolutional neural network and restricted Boltzmann machine features. Based on these features, the memorability score of individual patches in an image is calculated and these scores are used to generate the memorability map. These memorability map values are used for optimised image compression. In order to validate the results, an eye tracking experiment with human participants is performed. The comparative analysis shows that the memorability‐based compression outperforms the state‐of‐the‐art compression techniques.
Meera Thapar Khanna, Chetan Ralekar, Anurika Goel, Santanu Chaudhury, Brejesh Lall
IET Image Process.4
2019 Image super resolution using distributed locality sensitive hashing for manifold learning
Anurag Tripathi, Abhinav Gupta 0006, Santanu Chaudhury
Multim. Tools Appl.3
2018 Despeckling CNN with Ensembles of Classical Outputs
abstract
Ultrasound (US) image despeckling is a problem of high clinical importance. Machine learning solutions to the problem are considered impractical due to the unavailability of speckle-free US image dataset. On the other hand, the classical approaches, which are able to provide the desired outputs, have limitations like input dependent parameter tuning. In this work, a convolutional neural network (CNN) is developed which learns to remove speckle from US images using the outputs of these classical approaches. It is observed that the existing approaches can be combined in a complementary manner to generate an output better than their individual outputs. Thus, the CNN is trained using the individual outputs as well as the output ensembles. It eliminates the cumbersome process of parameter tuning required by the existing approaches for every new input. Further, the proposed CNN is able to outperform the state-of-the-art despeckling approaches and produces the outputs even better than the ensembles for certain images.
Deepak Mishra 0003, Sarthak Tyagi, Santanu Chaudhury, Mukul Sarkar, Arvinder Singh Soin
ICPR3
2018 Automatic Quantification of Stomata for High-Throughput Plant Phenotyping
abstract
Stomatal morphology is a key phenotypic trait for plants' response analysis under various environmental stresses (e.g. drought, salinity etc.). Stomata exhibit diverse characteristics with respect to orientation, size, shape and varying degree of papillae occlusion. Thus, the biologists currently rely on manual or semi-automatic approaches to accurately compute its morphological traits based on scanning electron microscopic (SEM) images of leaf surface. In contrast to these subjective and low-throughput methods, we propose a novel automated framework for stomata quantification. It is realized based on a hybrid approach where the candidate stomata region is first detected by a convolutional neural network (CNN) and the occlusion is dealt with an inpainting algorithm. In addition, we propose stomata segmentation based quantification framework to solve the problem of shape, scale and occlusion in an end-to-end manner. The performance of the proposed automated frameworks is evaluated by comparing the derived traits with manually computed morphological traits of stomata. With no prior information about its size and location, the hybrid and end-to-end machine learning frameworks shows a correlation of 0.94 and 0.93, respectively on rice stomata images. Furthermore, they successfully enable wheat stomata quantification showing generalizability in terms of cultivars.
Swati Bhugra, Deepak Mishra 0003, Anupama Anupama, Santanu Chaudhury, Brejesh Lall, Archana Chugh
ICPR4
2018 Variational Bayes Block Sparse Modeling with Correlated Entries
abstract
This paper addresses the problem of Bayesian Block Sparse Modeling when coefficients within the blocks are correlated. In contrast to the current hierarchical methods which do not exploit correlation structure within the blocks, we propose a three level hierarchical estimation framework. It employs heavy-tailed priors for block sparse modeling and variational inference for Bayesian estimation. This paper also describes the relationship between proposed framework and some of the existing Block Sparse Bayesian Learning (SBL) methods and show that these SBL methods can be viewed as its special cases. Extensive experimental results for synthetic signals are provided, demonstrating the superior performance of the proposed framework in terms of failure rate, relative reconstruction error, to name a few. We also demonstrate the applicability of this framework in telemonitoring of Fetal Electrocardiogram.
Santanu Chaudhury, Jayadeva
ICPR2
2018 Segmentation of Vascular Regions in Ultrasound Images: A Deep Learning Approach
abstract
Vascular region segmentation in ultrasound images is necessary for applications like automatic registration, and surgical navigation. In this paper, a pipelined network comprising of a convolutional neural network (CNN) followed by unsupervised clustering is proposed to perform vessel segmentation in liver ultrasound images. The work is motivated by the tremendous success of CNNs in object detection and localization. CNN here is trained to localize vascular regions, which are subsequently segmented by the clustering. The proposed network results in 99.14% pixel accuracy and 69.62% mean region intersection over union on 132 images. These values are better than some existing methods.
Deepak Mishra 0003, Santanu Chaudhury, Mukul Sarkar, Sidharth Manohar, Arvinder Singh Soin
ISCAS2
2018 Joint Image Classification and Annotation Prediction Using Iterative Learning on Local Neighbourhood
abstract
Image annotation (tag) and classification play a critical role in many computer vision applications, such as image retrieval, scene understanding, scene description etc. While, databases such as ImageNet have high quality labels for images, in real world, a large number of images have missing labels or tags that completely describe the contents of an image. To solve this problem, in this paper, we work on the hypothesis that class and tag information are correlated and propose a joint optimization for image classification and annotation. We construct a unified cost function to learn the class scoring vectors as well as tag scoring vectors. The proposed approach achieves state-of-the-art results on benchmark datasets for joint tag prediction and classification.
Anurag Tripathi, Siddharth Srivastava 0004, Santanu Chaudhury, Brejesh Lall
SMC3
2018 Clustering short temporal behaviour sequences for customer segmentation using LDA
abstract
Abstract Customer segmentation based on temporal variation of subscriber preferences is useful for communication service providers (CSPs) in applications such as targeted campaign design, churn prediction, and fraud detection. Traditional clustering algorithms are inadequate in this context, as a multidimensional feature vector represents a subscriber profile at an instant of time, and grouping of subscribers needs to consider variation of subscriber preferences across time. Clustering in this case usually requires complex multivariate time series analysis‐based models. Because conventional time series clustering models have limitations around scalability and ability to accurately represent temporal behaviour sequences (TBS) of users, that may be short, noisy, and non‐stationary, we propose a latent Dirichlet allocation (LDA) based model to represent temporal behaviour of mobile subscribers as compact and interpretable profiles. Our model makes use of the structural regularity within the observable data corresponding to a large number of user profiles and relaxes the strict temporal ordering of user preferences in TBS clustering. We use mean‐shift clustering to segment subscribers based on their discovered profiles. Further, we mine segment‐specific association rules from the discovered TBS clusters, to aid marketers in designing intelligent campaigns that match segment preferences. Our experiments on real world data collected from a popular Asian communication service provider gave encouraging results.
Jobin Wilson, Santanu Chaudhury, Brejesh Lall
Expert Syst. J. Knowl. Eng.2
2018 Spatio-contextual Gaussian mixture model for local change detection in underwater video
Deepak Kumar Rout, Badri N. Subudhi, Veerakumar Thangaraj, Santanu Chaudhury
Expert Syst. Appl.4
2018 Ultrasound Image Enhancement Using Structure Oriented Adversarial Network
abstract
In this letter, we aim to develop a deep adversarial despeckling approach to enhance the quality of ultrasound images. Most of the existing approaches target a complete removal of speckle, which produces oversmooth outputs and results in loss of structural details. In contrast, the proposed approach reduces the speckle extent without altering the structural and qualitative attributes of the ultrasound images. A despeckling residual neural network (DRNN) is trained with an adversarial loss imposed by a discriminator. The discriminator tries to differentiate between the despeckled images generated by the DRNN and the set of high-quality images. Further to prevent the developed network from oversmoothing, a structural loss term is used along with the adversarial loss. Experimental evaluations show that the proposed DRNN outperforms the state-of-the-art despeckling approaches in terms of the structural similarity index measure, peak signal to noise ratio, edge preservation index, and speckle region's signal to noise ratio.
Deepak Mishra 0003, Santanu Chaudhury, Mukul Sarkar, Arvinder Singh Soin
IEEE Signal Process. Lett.2
2018 Edge Probability and Pixel Relativity-Based Speckle Reducing Anisotropic Diffusion
abstract
Anisotropic diffusion filters are one of the best choices for speckle reduction in the ultrasound images. These filters control the diffusion flux flow using local image statistics and provide the desired speckle suppression. However, inefficient use of edge characteristics results in either oversmooth image or an image containing misinterpreted spurious edges. As a result, the diagnostic quality of the images becomes a concern. To alleviate such problems, a novel anisotropic diffusion-based speckle reducing filter is proposed in this paper. A probability density function of the edges along with pixel relativity information is used to control the diffusion flux flow. The probability density function helps in removing the spurious edges and the pixel relativity reduces the oversmoothing effects. Furthermore, the filtering is performed in superpixel domain to reduce the execution time, wherein a minimum of 15% of the total number of image pixels can be used. For performance evaluation, 31 frames of three synthetic images and 40 real ultrasound images are used. In most of the experiments, the proposed filter shows a better performance as compared to the state-of-the-art filters in terms of the speckle region's signal-to-noise ratio and mean square error. It also shows a comparative performance for figure of merit and structural similarity measure index. Furthermore, in the subjective evaluation, performed by the expert radiologists, the proposed filter's outputs are preferred for the improved contrast and sharpness of the object boundaries. Hence, the proposed filtering framework is suitable to reduce the unwanted speckle and improve the quality of the ultrasound images.
Deepak Mishra 0003, Santanu Chaudhury, Mukul Sarkar, Arvinder Singh Soin
IEEE Trans. Image Process.2
2017 A Noise-Resilient Super-Resolution Framework to Boost OCR Performance
abstract
Recognizing text from noisy low-resolution (LR) images is extremely challenging and is an open problem for the computer vision community. Super-resolving a noisy LR text image results in noisy High Resolution (HR) text image, as super-resolution (SR) leads to spatial correlation in the noise, and further cannot be de-noised successfully. Traditional noise-resilient text image super-resolution methods utilize a denoising algorithm prior to text SR but denoising process leads to loss of some high frequency details, and the output HR image has missing information (texture details and edges). This paper proposes a noise-resilient SR framework for text images and recognizes the text using a deep BLSTM network trained on high resolution images. The proposed end-to-end deep learning based framework for noise-resilient text image SR simultaneously perform image denoising and super-resolution as well as preserves missing details. Stacked sparse denoising auto-encoder (SSDA) is learned for LR text image denoising, and our proposed coupled deep convolutional auto-encoder (CDCA) is learned for text image super-resolution. The pretrained weights for both these networks serve as initial weights to the end-to-end framework during finetuning, and the network is jointly optimized for both the tasks. We tested on several Indian Language datasets and the OCR performance of the noise-resilient super-resolved images is at par with the original HR images.
Anupama Ray, Santanu Chaudhury, Brejesh Lall
ICDAR3
2017 Deep learning based frameworks for image super-resolution and noise-resilient super-resolution
abstract
Our paper is motivated from the advancement in deep learning algorithms for various computer vision problems. We are proposing a novel end-to-end deep learning based framework for image super-resolution. This framework simultaneously calculates the convolutional features of low-resolution (LR) and high-resolution (HR) image patches and learns the non-linear function that maps these convolutional features of LR image patches to their corresponding HR image patches convolutional features. Here, proposed deep learning based image super-resolution architecture is termed as coupled deep convolutional auto-encoder (CDCA) which provides state-of-the-art results. Super-resolution of a noisy/distorted LR images results in noisy/distorted HR images, as super-resolution process gives rise to spatial correlation in the noise, and further, it cannot be de-noised successfully. Traditional noise resilient image super-resolution methods utilize a de-noising algorithm prior to super-resolution but de-noising process gives rise to loss of some high-frequency information (edges and texture details) and super-resolution of the resultant image provides HR image with missing edges and texture information. We are also proposing a novel end-to-end deep learning based framework to obtain noise resilient image super-resolution. Proposed end-to-end deep learning based framework for noise resilient super-resolution simultaneously perform image de-noising and super-resolution as well as preserves textural details. First, stacked sparse de-noising auto-encoder (SSDA) was learned for LR image de-noising and proposed CDCA was learned for image superresolution. Then, both image de-noising and super-resolution networks were cascaded. This cascaded deep learning network was employed as one integral network where pre-trained weights were serving as initial weights. The integral network was end-to-end trained or fine-tuned on a database having noisy, LR image as an input and target as an HR image. In fine-tuning, all layers of the combined end-to-end network was jointly optimized to perform image de-noising and super-resolution simultaneously. Experimental results show that proposed noise resilient super-resolution framework outperforms the conventional and state-of-the-art approaches in terms of PSNR and SSIM metrics.
Santanu Chaudhury, Brejesh Lall
IJCNN2
2017 An IoT approach for context-aware smart traffic management using ontology
abstract
This paper exhibits a novel context-aware service framework for IoT based Smart Traffic Management using ontology to regulate smooth traffic flow in smart cities by analyzing real-time traffic environment. The proposed approach makes smarter use of transport networks to achieve objectives related to performance of transport system. This requires efficient traffic planning measures which relate to the actions designed to adjust the demand and capacity of the network in time and space by use of IoT technologies. The adoption of sensors and IoT devices in Smart Traffic System helps to capture the user's preferences and context information which can be in the form of travel time, weather conditions or real-life driving patterns. We have employed multimedia ontology to derive higher level descriptions of traffic conditions and vehicles from perceptual observation of traffic information which provides important grounds for our proposed IoT framework. The multimedia ontology encoded in Multimedia Web Ontology Language(MOWL) helps to define classes, properties, and structure of a possible traffic environment to provide insights across the transportation network. MOWL supports Dynamic Bayesian networks (DBN) to deal with time-series data and uncertainties linked with context observations which fits the definition of an intelligent IoT system. Thus, our proposed smart traffic framework aggregates information corresponding to traffic domain such as traffic videos captured using CCTV cameras and allows automatic prediction of dynamically changing situations which helps to make traffic authorities more responsive. We have illustrated use of our approach by utilizing contextual information, to assess real-time congestion situation on roads thus allowing to visualize planning services. Once the congestion situation is predicted, alternate congestion free routes which are in accordance with the coveted criteria are suggested that can be propagated through text-messages or e-mails to the users.
Deepti Goel, Santanu Chaudhury, Hiranmay Ghosh
WI2
2017 A Novel Hybrid Kinect-Variety-Based High-Quality Multiview Rendering Scheme for Glass-Free 3D Displays
abstract
This paper presents a new hybrid Kinect-variety-based synthesis scheme that renders artifact-free multiple views for autostereoscopic/automultiscopic displays. The proposed approach does not explicitly require dense scene depth information for synthesizing novel views from arbitrary viewpoints. Instead, the integrated framework first constructs a consistent minimal image–space parameterization of the underlying 3D scene. The compact representation of scene structure is formed using only implicit sparse depth information of a few reference scene points extracted from raw RGB depth data. The views from arbitrary positions can be inferred by moving the novel camera in parameterized space by enforcing Euclidean constraints on reference scene images under a full-perspective projection model. Unlike the state-of-the-art depth image-based rendering (DIBR) methods, in which input depth map accuracy is crucial for high-quality output, our proposed algorithm does not depend on precise per-pixel geometry information. Therefore, it simply sidesteps to recover and refine the incomplete or noisy depth estimates with advanced filling or upscaling techniques. Our approach performs fairly well in unconstrained indoor/outdoor environments, where the performance of range sensors or dense depth-based algorithms could be seriously affected due to scene complex geometric conditions. We demonstrate that the proposed hybrid scheme provides guarantees on the completeness, optimality with respect to the inter-view consistency of the algorithm. In the experimental validation, we performed a quantitative evaluation as well as subjective assessment of the scene with complex geometric or surface properties. A comparison with the latest representative DIBR methods is additionally performed to demonstrate the superior performance of the proposed scheme.
Santanu Chaudhury, Brejesh Lall
IEEE Trans. Circuits Syst. Video Technol.2
2016 Link Prediction in Heterogeneous Social Networks
abstract
A heterogeneous social network is characterized by multiple link types which makes the task of link prediction in such networks more involved. In the last few years collective link prediction methods have been proposed for the problem of link prediction in heterogeneous networks. These methods capture the correlation between different types of links and utilize this information in the link prediction task. In this paper we pose the problem of link prediction in heterogeneous networks as a multi-task, metric learning (MTML) problem. For each link-type (relation) we learn a corresponding distance measure, which utilizes both network and node features. These link-type specific distance measures are learnt in a coupled fashion by employing the Multi-Task Structure Preserving Metric Learning (MT-SPML) setup. We further extend the MT-SPML method to account for task correlations, robustness to non-informative features and non-stationary degree distribution across networks. Experiments on the Flickr and DBLP network demonstrates the effectiveness of our proposed approach vis-à-vis competitive baselines.
Sumit Negi, Santanu Chaudhury
CIKM2
2016 Automatic Selection of Parameters for Document Image Enhancement Using Image Quality Assessment
abstract
Performance of most of the recognition engines for document images is effected by quality of the image being processed and the selection of parameter values for the pre-processing algorithm. Usually the choice of such parameters is done empirically. In this paper, we propose a novel framework for automatic selection of optimal parameters for pre-processing algorithm by estimating the quality of the document image. Recognition accuracy can be used as a metric for document quality assessment. We learn filters that capture the script properties and degradation to predict recognition accuracy. An EM based framework has been formulated to iteratively learn optimal parameters for document image pre-processing. In the E-step, we estimate the expected accuracy using the current set of parameters and filters. In the M-step we compute parameters to maximize the expected recognition accuracy found in E-step. The experiments validate the efficacy of the proposed methodology for document image pre-processing applications.
Ritu Garg, Santanu Chaudhury
DAS2
2016 Partial Multi-View Clustering using Graph Regularized NMF
abstract
Real-world datasets consist of data representations (views) from different sources which often provide information complementary to each other. Multi-view learning algorithms aim at exploiting the complementary information present in different views for clustering and classification tasks. Several multi-view clustering methods that aim at partitioning objects into clusters based on multiple representations of the object have been proposed. Almost all of the proposed methods assume that each example appears in all views or at least there is one view containing all examples. In real-world settings this assumption might be too restrictive. Recent work on Partial View Clustering addresses this limitation by proposing a Non-negative Matrix Factorization based approach called PVC. Our work extends the PVC work in two directions. First, the current PVC algorithm is designed specifically for two-view datasets. We extend this algorithm for the k partial-view scenario. Second, we extend our k partial-view algorithm to include view specific graph laplacian regularization. This enables the proposed algorithm to exploit the intrinsic geometry of the data distribution in each view. The proposed method, which is referred to as GPMVC (Graph Regularized Partial Multi-View Clustering), is compared against 7 baseline methods (including PVC) on 5 publicly available text and image datasets. In all settings the proposed GPMVC method outperforms all baselines. For the purpose of reproducibility, we provide access to our code.
Nishant Rai, Sumit Negi, Santanu Chaudhury, Om Deshmukh
ICPR3
2016 Recognition based text localization from natural scene images
abstract
With the rapid increase of multimedia data, textual content in an image has become a very important source of information for several applications like navigation, image search and retrieval, image understanding, captioning, machine translation and several others. Scene text localization is the first step towards such applications and most current methods focus on generating a small set of high precision detectors rather than obtaining large set of detections covering all text patches. In this work we propose a novel hybrid framework for text localization which uses character level recognition recursively in a feedback mechanism to refine text patches and reduce false positives. We use popular MSER algorithm at multiple scales as an initial region proposal algorithm and several filtering stages recursively to improve precision as well as maximize recall. We aim at achieving high recall rather than achieving higher precision since several robust word recognition systems are already available. The word recognition systems are mature enough to produce highly accurate results if provided with maximum amount of regions rather than providing small set of highly precise text patches and losing several other text regions. The main contribution of this paper is the use of character recognizer within a novel feedback mechanism to recursively search for text regions in the neighborhood of previously detected text patches. Using 3 publicly available benchmark datasets (ICDAR2011, MSRA TD-500 and OSTD), we demonstrate the efficacy of the proposed framework for text localization.
Anupama Ray, Archit Shah, Santanu Chaudhury
ICPR3
2016 Pose estimation of texture-less cylindrical objects in bin picking using sensor fusion
abstract
We propose an approach for emptying of bin using a combination of Image and Range sensor. Offering a complete solution: calibration, segmentation and pose estimation, along with approachability analysis for the estimated pose. The work is novel in the sense that the objects to be picked are featureless and uniformly black in colour, hence existing approaches are not directly applicable. A key point involves optimal utilization of range data acquired from the laser scanner for 3-D segmentation using localized geometric information. This information guides segmentation of the image for better object pose estimation, used for pick-and-drop. We analytically assure the approachability of the object to avoid collision of the manipulator with the bin. Disturbance of objects caused during pick up has been modelled, which allows pickup of multiple pellets based on information from a single range scan. This eliminates the necessity of repeated scanning and data conditioning. The proposed method offers high object detection rate and pose estimation accuracy. The innovative techniques aimed at reducing the average pickup time makes it suitable for robust industrial operation.
Mayank Roy, Riby Abraham Boby, Shraddha Chaudhary, Santanu Chaudhury, Sumantra Dutta Roy, Subir Kumar Saha
IROS4
2016 Video analytics revisited
abstract
Video, rich in visual real‐time content, is however, difficult to interpret and analyse. Video collections necessarily have large data volume. Video analytics strives to automatically discover patterns and correlations present in the large volume of video data, which can help the end‐user to take informed and intelligent decisions as well as predict the future based on the patterns discovered across space and time. In this study, the authors discuss various issues and problems in video analytics, proposed solutions and present some of the important current applications of video analytics.
Ayesha Choudhary, Santanu Chaudhury
IET Comput. Vis.2
2016 Iris recognition based on sparse representation and k-nearest subspace with genetic algorithm
Ashok K. Bhateja, Santanu Chaudhury, Nitin Agrawal 0002
Pattern Recognit. Lett.3
2015 Document indexing framework for retrieval of degraded document images
abstract
With the availability of large collection of document images in Indian languages, image based retrieval has gained popularity. The performance of such systems is effected by the presence of degraded and noisy images. Moreover, Optical character recognition systems for Indian scripts are not yet robust, leading to noisy OCR'ed text. Information retrieval system designed using inputs from both modalities (image features and OCR based recognition data) will lead to better retrieval performance in contrast to usage of individual modality. In this paper we present a indexing methodology that uses multiple kernel learning to combine features from different modalities by joint optimization of search time and accuracy. The evaluation of the proposed methodology is demonstrated on document images of Bangla and Devanagari script.
Ritu Garg, Ehtesham Hassan, Santanu Chaudhury
ICDAR3
2015 A hypothesize-and-verify framework for text recognition using deep recurrent neural networks
abstract
Deep LSTM is an ideal candidate for text recognition. However text recognition involves some initial image processing steps like segmentation of lines and words which can induce error to the recognition system. Without segmentation, learning very long range context is difficult and becomes computationally intractable. Therefore, alternative soft decisions are needed at the pre-processing level. This paper proposes a hybrid text recognizer using a deep recurrent neural network with multiple layers of abstraction and long range context along with a language model to verify the performance of the deep neural network. In this paper we construct a multi-hypotheses tree architecture with candidate segments of line sequences from different segmentation algorithms at its different branches. The deep neural network is trained on perfectly segmented data and tests each of the candidate segments, generating unicode sequences. In the verification step, these unicode sequences are validated using a sub-string match with the language model and best first search is used to find the best possible combination of alternative hypothesis from the tree structure. Thus the verification framework using language models eliminates wrong segmentation outputs and filters recognition errors.
Anupama Ray, Sai Rajeswar, Santanu Chaudhury
ICDAR3
2015 OCR for bilingual documents using language modeling
abstract
Script based features are highly discriminative for text segmentation and recognition. Thus they are widely used in Optical Character Recognition(OCR) problems. But usage of script dependent features restricts the adaptation of such architectures directly for another script. With script independent systems, this problem can be solved to a certain extent for monolingual documents. But the problem aggravates in case of multilingual documents as it is very difficult for a single classifier to learn many scripts. Generally a script identification module identifies text segments and accordingly the script-dependent classifier is selected. This paper presents a unified framework of language model and multiple preprocessing hypotheses for word recognition from bilingual document images. Prior to text recognition, preprocessing steps such as binarization and segmentation are required for ease of recognition. But these steps induce huge combinatorial error propagating to final recognition accuracy. In this paper we use multiple preprocessing routines as alternate hypotheses and use a language model to verify each alternative and choose the best recognized sequence. We test this architecture for word recognition of Kannada-English and Telugu-English bilingual documents and achieved better recognition rates than single methods using same classifier.
Anupama Ray, Sai Rajeswar, Santanu Chaudhury
ICDAR3
2015 Camera-based document image matching using multi-feature probabilistic information fusion
Sumantra Dutta Roy, Kavita Bhardwaj, Rhishabh Garg, Santanu Chaudhury
Pattern Recognit. Lett.4
2014 Semantic clustering-based cross-domain recommendation
abstract
Cross-domain recommendation systems exploit tags, textual descriptions or ratings available for items in one domain to recommend items in multiple domains. Handling unstructured/ unannotated item information is, however, a challenge. Topic modeling offer a popular method for deducing structure in such data corpora. In this paper, we introduce the concept of a common latent semantic space, spanning multiple domains, using topic modeling of semantic clustered vocabularies of distinct domains. The intuition here is to use explicitly-determined semantic relationships between non-identical, but possibly semantically equivalent, words in multiple domain vocabularies, in order to capture relationships across information obtained in distinct domains. The popular WordNet based ontology is used to measure semantic relatedness between textual words. The experimental results shows that there is a marked improvement in the precision of predicting user preferences for items in one domain when given the preferences in another domain.
Santanu Chaudhury, Sumeet Agarwal
CIDM4
2014 Newspaper Article Extraction Using Hierarchical Fixed Point Model
abstract
This paper presents a novel learning based framework to extract articles from newspaper images using a Fixed-Point Model. The input to the system comprises blocks of text and graphics, obtained using standard image processing techniques. The fixed point model uses contextual information and features of each block to learn the layout of newspaper images and attains a contraction mapping to assign a unique label to every block. We use a hierarchical model which works in two stages. In the first stage, a semantic label (heading, sub-heading, text-blocks, image and caption) is assigned to each segmented block. The labels are then used as input to the next stage to group the related blocks into news articles. Experimental results show the applicability of our algorithm in newspaper labeling and article extraction.
Anukriti Bansal, Santanu Chaudhury, Sumantra Dutta Roy, J. B. Srivastava
Document Analysis Systems2
2014 A Robust Online Signature Based Cryptosystem
abstract
Cryptography is the backbone for the security systems. The main challenge in use of the Cryptosystems is maintaining the confidentiality of the cryptographic key. A Cryptosystem which encrypts the data using biometric features improves the security of the data and overcomes the problems of key management and key confidentiality. Fuzzy Vault Scheme proposed by Juels and Sudan [1] binds the secret key and the biometric template, so that extraction of the secret without the biometric data is infeasible. Physical signature is a biometric that is widely accepted and is used for proving the authenticity of a person in legal documents, bank transactions etc. Electronic devices such as digital tablets capture azimuth, altitude and pressure along with x any y coordinates at fixed time interval. This paper describes a robust online signature based cryptosystem to hide the secret by binding it with invariant online signature templates. The invariant templates of the signature are derived from artificial neural network based classifier. The entire signature is divided into fixed number of time slices. Important features are extracted based on the consistency of the feature in the slices of the genuine signature. Binary back propagation based neural network for each feature, each subset of slices for a user is trained by a weighted back propagation algorithm. The decisions of these networks are combined using AdaBoost algorithm. The proposed scheme is highly robust as it works well for all kinds of signatures and is independent of the number of zero crossing and high curvature points in the signature trajectory.
Ashok K. Bhateja, Santanu Chaudhury, P. K. Saxena
ICFHR2
2014 Design of Multi-kernel Distance Based Hashing with Multiple Objectives for Image Indexing
abstract
Approximate nearest neighbor (ANN) search provides computationally viable option for retrieval from large document collection. Hashing based techniques are widely regarded as most efficient methods for ANN based retrieval. It has been established that by combination of multiple features in a multiple kernel learning setup can significantly improve the effectiveness of hash codes. The paper presents a novel image indexing method based on multiple kernel learning, which combines multiple features by combinatorial optimization of time and search complexity. The framework is built upon distance based hashing, where the existing kernel distance based hashing formulation adopts linear combination of kernels in tune with optimum search accuracy. In this direction, a novel multiobjective formulation for optimizing the search time as well as accuracy is proposed which is subsequently solved in Genetic algorithm based solution framework for obtaining the pareto-optimal solutions. We have performed extensive experimental evaluation of proposed concepts on different datasets showing improvement in comparison with the existing methods.
Vaibhav Gaur, Ehtesham Hassan, Santanu Chaudhury
ICPR3
2014 Discovering User-Communities and Associated Topics from YouTube
abstract
Most of the popular multimedia sharing web-sites such as YouTube, Flickr etc not only allow users to author and upload content but also facilitate "social" networking amongst users. These social interactions can be in the form of - user-to-user interactions i.e. adding existing users to friend or contact list or user-to-content interactions : commenting on a video or picture, marking a picture/video as "favorite", subscribing to a user created "channel" etc. Analyzing these social interactions jointly with the content metadata (such as the description of the video, keywords associated with the image/video etc) can reveal interesting insights about user activity on these social media platforms. In this paper, we propose an unsupervised method that jointly models "social" interaction and content metadata in YouTube to discover user-communities and the nature of topics beings discussed in these communities. We report the effectiveness of the proposed method on real-world dataset.
Sumit Negi, Ramnath Balasubramanyan, Santanu Chaudhury
ICPR3
2014 Identifying Diverse Set of Images in Flickr
abstract
Unlike traditional multimedia content, content generated on social media platforms such as YouTube, Flickr etc are usually annotated with rich set of social tags such as keywords, textual description, category information, author's profile etc. In this paper we investigate the use of such social tag information for visual diversification of image search results in Flickr. We model search result diversity as an instance of the p-dispersion problem where the objective is to choose p out of n given points, so that the minimum distance between any pair of chosen points is maximized. The distance metric used in the p-dispersion problem is learnt from the data itself by combining candidate similarity measures as defined on the social tags. We demonstrate the effectiveness of our proposed method on a real-world data set.
Sumit Negi, Santanu Chaudhury
ICPR2
2014 Kinect-Variety Fusion: A Novel Hybrid Approach for Artifacts-Free 3DTV Content Generation
abstract
This paper presents a novel low-cost hybrid Kinect-variety based content generation scheme for 3DTV displays. The integrated framework constructs an efficient consistent image-space parameterization of 3D scene structure using only sparse depth information of few reference scene points. Under full-perspective camera model, the enforced Euclidean constraints simplify the synthesis of high quality novel multiview content for distinct camera motions. The algorithm does not rely on complete precise scene geometry information, and are unaffected by scene complex geometric properties, unconstrained environmental variations and illumination conditions. It, therefore, performs fairly well under a wider set of operation condition where the 3D range sensors fail or reliability of depth-based algorithms are suspect. The robust integration of vision algorithm and visual sensing scheme complement each other's shortcomings. It opens new opportunities for envisioning vision-sensing applications in uncontrolled environments. We demonstrate that proposed robust integration provides guarantees on the completeness and consistency of the algorithm. This leads to improved reliability on an extensive set of experimental results.
Santanu Chaudhury, Brejesh Lall
ICPR2
2014 Feature combination for binary pattern classification
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
Int. J. Document Anal. Recognit.2
2014 A flexible architecture for multi-view 3DTV based on uncalibrated cameras
Santanu Chaudhury, Brejesh Lall, M. S. Venkatesh
J. Vis. Commun. Image Represent.2
2014 Off-line hand written input based identity determination using multi kernel feature combination
Ehtesham Hassan, Santanu Chaudhury, Nivedita Yadav, Prem Kumar Kalra, Madan Gopal
Pattern Recognit. Lett.2
2013 Space-Time Parameterized Variety Manifolds: A Novel Approach for Arbitrary Multi-perspective 3D View Generation
abstract
This paper presents a novel image variety-Based approach that elegantly models the space of a broad class of perspective and non-perspective stereo varieties within a single, unified framework. The basic concept of parameterized variety presented earlier by Genc and Ponce [1] is extended to represent the nonlinear space of images. An efficient algebraic framework is constructed to parameterize the variety associated with full perspective cameras. The algorithm seeks the manifolds that constrain this space of six-dimensional variety to generate compelling multi-perspective 3D effects from arbitrary virtual viewpoints. Combining geometric space of multiple uncalibrated perspective views with appearance space in a globally optimized way leads to numerous potential applications, especially in content creation for multi-perspective 3DTV. The proposed approach works for uncalibrated static/dynamic scenes, containing parallax and unstructured object motion. It even seamlessly deals with images or video sequences that do not share a common origin, thus provides an effective tool for montaging, indexing and virtual navigation.
Santanu Chaudhury, Brejesh Lall
3DV2
2013 Greedy Search for Active Learning of OCR
abstract
Active learning and crowd sourcing are becoming increasingly popular in the machine learning community for fast and cost effective generation of labels for large volumes of data. However, such labels may be noisy. So, it becomes important to ignore the noisy labels for building of a good classifier. We propose a framework for finding the best possible augmentation of a classifier for the character recognition problem using minimum number of crowd labeled samples. The approach inherently rejects the noisy data and tries to accept a subset of correctly labeled data to maximize the classifier performance.
Ritu Garg, Santanu Chaudhury
ICDAR3
2013 Multi-modal Information Integration for Document Retrieval
abstract
The paper proposes a novel multi-modal document image retrieval framework by exploiting the information of text and graphics regions. The framework applies multiple kernel learning based hashing formulation for generation of composite document indexes using different modalities. The existing multimedia management methods for imaged text documents have not addressed the requirement of old and degraded documents. In the subsequent contribution, we propose novel multi-modal document indexing framework for retrieval of old and degraded text documents by combining OCR'ed text and image based representation using learning. The evaluation of proposed concepts is demonstrated on sampled magazine cover pages, and documents of Devanagari script.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
ICDAR2
2013 Character Recognition Using Conditional Random Field Based Recognition Engine
abstract
The paper presents a novel script independent CRF based inferencing framework for character recognition. In this framework we consider a word as a sequence of connected components. The connected components are obtained using different binarization schemes and different possible sequences are considered using a tree structure. CRF uses contextual information to learn perfect primitive sequences and finds the most probable labeling of the sequence of primitives using multiple hypothesis tree to form the correct sequence of alphabets. This approach is particularly suitable for degraded printed document images as it considers multiple alternate hypotheses for correct decision.
Anupama Ray, Ankit Chandawala, Santanu Chaudhury
ICDAR3
2013 Most Discriminative Primitive Selection for Identity Determination Using Handwritten Devanagari Script
abstract
Writer recognition based on peculiarity of hand-writing is an important aspect of any forensic analysis. We present an approach for selecting best discriminative primitives for writer recognition. After selecting the primitives we also propose a hybrid system by combining both writer recognition and handwriting recognition for improved accuracy. We have also validated the performance of selected primitives on publically available dataset. We have performed this study on the Devanagri script. Experimental results verified the effectiveness of the proposed franework.
Nivedita Yadav, Santanu Chaudhury, Prem Kumar Kalra
ICDAR2
2013 Inferring actor communities from videos
Sumit Negi, Ramnath Balasubramanyan, Santanu Chaudhury
INTERSPEECH3
2013 Towards Adaptive Multi-agent Planning in Cyber Physical Space
abstract
Cyber physical space is potentially hosting innumerable spatio-temporal data streams due to increasing use of social networking platforms as real-time information dissemination system and world-wide deployment of sensors for continuous monitoring of physical phenomena. In this paper we address the problem of how cyber physical space can be used for sensing and responding to global calamities such as earthquake. The paper proposes a novel approach of generating online adaptive response in assisting search-and-rescue operations using situation awareness built from real-time heterogeneous spatio-temporal data streams. Online adaptive response is achieved by using agent-based cooperative task sharing and modeling agent decision making as self-organized emergent behavior based on concepts of complex adaptive system. An implemented simulation platform use concepts of situation modeling, domain task network, contract net protocol based negotiation and complex adaptive system to generate adaptive plans. Preliminary simulation results are promising as we have been able to demonstrate a repertoire of self-organized emergent behaviors.
Sumant Mukherjee, Santanu Chaudhury
Web Intelligence2
2013 Word shape descriptor-based document image indexing: a new DBH-based approach
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
Int. J. Document Anal. Recognit.2
2013 Visual saliency guided video compression algorithm
Rupesh Gupta, Meera Thapar Khanna, Santanu Chaudhury
Signal Process. Image Commun.3
2013 MOWL: An ontology representation language for web-based multimedia applications
abstract
Several multimedia applications need to reason with concepts and their media properties in specific domain contexts. Media properties of concepts exhibit some unique characteristics that cannot be dealt with conceptual modeling schemes followed in the existing ontology representation and reasoning schemes. We have proposed a new perceptual modeling technique for reasoning with media properties observed in multimedia instances and the latent concepts. Our knowledge representation scheme uses a causal model of the world where concepts manifest in media properties with uncertainties. We introduce a probabilistic reasoning scheme for belief propagation across domain concepts through observation of media properties. In order to support the perceptual modeling and reasoning paradigm, we propose a new ontology language, Multimedia Web Ontology Language (MOWL). Our primary contribution in this article is to establish the need for the new ontology language and to introduce the semantics of its novel language constructs. We establish the generality of our approach with two disperate knowledge-intensive applications involving reasoning with media properties of concepts.
Anupama Mallik, Hiranmay Ghosh, Santanu Chaudhury, Gaurav Harit
ACM Trans. Multim. Comput. Commun. Appl.3
2012 Parameterized Variety Based View Synthesis Scheme for Multi-view 3DTV
Santanu Chaudhury, Brejesh Lall
ACCV (4)2
2012 Predicting User-to-content Links in Flickr Groups
abstract
The last few years have seen an exponential increase in the amount of multimedia content that is available online thanks to collaborative-online communities such as Flickr, You Tube etc. As opposed to "pure" social networking services these collaborative-online communities not only allow users to create new social links (e.g. add other users to their friend or contact list) but also allow users to contribute multimedia content and engage in content-driven interactions (called user-to-content interactions). A good example of this can be seen in Flickr, in general and Flickr Group in particular where users can comment on or "like" an image contributed by another user. This paper looks at the task of predicting the formation of such user-to-content links in Flickr Groups. More specifically, "what is the chance that a user will comment/like an image contributed by another user?". Our proposed method for predicting user-to-content links takes into account both community effect and content effect. Our results on real-world Flickr Group data reveals that the proposed method shows good performance for the user-to-content link prediction task.
Sumit Negi, Santanu Chaudhury
ASONAM2
2012 Discovering Activities and Their Temporal Significance
abstract
In this paper, we address the problem of discovering activities and their temporal significance in an area under surveillance. Discovering activities along with its expectation of occurrence at a particular time plays an important role in many surveillance applications. We propose an unsupervised model, called Time pLSA model, that extends the probabilistic Latent Semantic Analysis (pLSA) model to jointly capture the activities and their behaviour over time. We use adaptive background subtraction to detect spatio-temporal patches, which are used as feature representation for activity patterns. Each of these patches are associated with the time slot in which they occur. Multinomial distributions are used to model both activities as distribution over spatio-temporal patches and time significance as distribution over the time-line. We demonstrate the effectiveness of our approach on a real life surveillance feed of an outdoor scene.
Ayesha Choudhary, Tanveer A. Faruquie, Subhashis Banerjee, Santanu Chaudhury
AVSS4
2012 Finding Subgroups in a Flickr Group
abstract
Information management systems today face a tremendous challenge considering the growing popularity of social media repositories involving images and video. Considering the growing volume of multimedia content in such online media-sharing communities there is an increasing need for novel ways of organizing content. In this paper we consider the problem of organizing images in a given Flickr Group by discovering latent subgroups. A Flickr Group can be visualized as a collection of such subgroups where each subgroup represents a distinct theme. We model the task of discovering subgroups as that of finding highly correlated topics from a dataset containing images and associated tags. The proposed probabilistic model employs a more flexible prior distribution to model topic-topic correlations and utilizes both tag and image information for discovering such subgroups. Our experiments on Flickr Group data demonstrate that the model is able to successfully discover subgroups without any supervision.
Sumit Negi, Santanu Chaudhury
ICME2
2012 Characterizing user-subgroups in Flickr Group: A block LDA based approach
Sumit Negi, Ramnath Balasubramanyan, Santanu Chaudhury
ICPR3
2012 A free viewpoint 3DTV system based on parameterized variety model
Santanu Chaudhury, Brejesh Lall
ICPR2
2012 A Ubiquitous Image Tagging System Using User Context
abstract
Tagging is nowadays the most predominant technique to make resources searchable. These allow users to create and manage tags to annotate and categorize content. In this paper, we propose an approach to tag images in a user's collection based upon user's personal profile, his/her social context and the context defined by his/her prior image collection. We apply LDA for context modeling. In this scheme, tag similarity and tag relevance are jointly estimated so that they can profit from each other. We have used an Adaptive Context Model created from user related sources to tag images. Experimental validation with user's mobile as well as website based image collection has established effectiveness of the approach.
Shatabdi Kundu, Santanu Chaudhury
Web Intelligence2
2012 Feature Combination in Kernel Space for Distance Based Image Hashing
abstract
The paper presents a novel feature based indexing scheme for image collections. The scheme presents the extension of distance based hashing to kernel space for generating the indexing structure based on similarity in kernel space. The objective of the scheme is to incorporate multiple features for defining the image indexing space using the concept of multiple kernel learning. However, the indexing problems are defined with unique learning objective; therefore, a novel application of genetic algorithm is presented for the optimization task. The extensive evaluation of the proposed concept is performed for developing word based document indexing application of Devanagari, Bengali, and English scripts. In addition, the efficacy of the proposed concept is shown by experimental evaluations on handwritten digits and natural image collection.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
IEEE Trans. Multim.2
2011 A CRF Based Scheme for Overlapping Multi-colored Text Graphics Separation
abstract
In this paper, we propose a novel framework for segmentation of documents with complex layouts. The document segmentation is performed by combination of clustering and conditional random fields (CRF) based modeling. The bottom-up approach for segmentation assigns each pixel to a cluster plane based on color intensity. A CRF based discriminative model is learned to extract the local neighborhood information in different cluster/color planes. The final category assignment is done by a top-level CRF based on the semantic correlation learned across clusters. The proposed framework has been extensively tested on multi-colored document images with text overlapping graphics/image.
Ritu Garg, Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
ICDAR3
2011 A Compression Scheme for Handwritten Patterns Based on Curve Fitting
abstract
We present here an idea of compression of user fed data from a touch screen input interface for storage and transmission over relatively lower bandwidth. The input is taken in the form of hand-written text, graphics, symbols or patterns and recorded as strokes in order of their temporal occurrence. The patterns are segmented into primitive forms each of which is then modeled with third order B-Spline Curves. The number of control points driving the Spline Curve is determined beforehand by recognizing the dominant points in the pattern. The significant reduction of redundancy in data can be exploited in wide application base including low-cost handheld device communication. This algorithm hence proposes a language independent tool for recording the handwriting of user in its original essence. Results obtained from preliminary testing on MATLAB and Android platform show significant improvement in compression ratio over the traditional storage and compression schemes.
Manish Bansal, Santanu Chaudhury
ICDAR3
2011 Document Image Indexing Using Edit Distance Based Hashing
abstract
We present a novel word image based document indexing scheme by combination of string matching and hashing. The word image representation is defined by string codes obtained by unsupervised learning over graphical primitives. The indexing framework is defined by distance based hashing function which does the object projection to hash space by preserving their distances. We have used edit distance based string matching for defining the hashing function and for approximate nearest neighbor based retrieval. The application of the proposed indexing framework is presented for two document image collections belonging to Devanagari and Bengali script.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
ICDAR2
2011 Searching OCR'ed Text: An LDA Based Approach
abstract
Indexing and retrieval performance over digitized document collection significantly depends on the performance of available Optical Character Recognition (OCR). The paper presents a novel document indexing framework which attends the document digitization errors in the indexing process to improve the overall retrieval accuracy. The proposed indexing framework is based on topic modeling using Latent Dirichlet Allocation (LDA). The OCR's confidence in correctly recognizing a symbol is propagated in topic learning process such that semantic grouping of word examples carefully distinguishes between commonly confusing words. We present a novel application of Lucene with topic modeling for document indexing application. The experimental evaluation of the proposed framework is presented on document collection belonging to Devanagari script.
Ehtesham Hassan, Vikram Garg, S. K. Mirajul Haque, Santanu Chaudhury, Madan Gopal
ICDAR4
2010 Use of MKL as symbol classifier for Gujarati character recognition
abstract
The present work is part of ongoing effort to improve the performance of Gujarati character recognition. In the recent advancement in kernel methods, the novel concept of multiple kernel learning(MKL) has given improved results for many problems. In this paper, we present novel application of MKL for Gujarati character recognition. We have applied three different feature representations for symbols obtained after zone wise segmentation of Gujarati text. The MKL based classification is proposed, where the MKL is used for learning optimal combination of different features for classification. In addition MKL based classification results for different features is also presented. The multiclass classification is performed in Decision DAG framework. The comparison results in 1-Vs-1 framework and using KNN classifier is also presented. The experiments have shown substantial improvement in earlier results.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal, Jignesh Dholakia
Document Analysis Systems2
2010 Identity Determination with Offline Handwritten Input Using Multi Kernel Feature Combination
abstract
The paper presents three novel features for handwritten data based identity recognition. A novel framework for combining the features for identification is presented. The framework combines the features in kernel space in MKL based framework. The application of features individually and in combination is presented for writer recognition and signature verification. The writer recognition results have been presented for Devanagari script input and signature verification results have been presented for open dataset [1]. The experiments have shown encouraging results.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
ICFHR2
2010 Texture mapping based video compression scheme
abstract
We present a video coding technique based on texture synthesis via stitching small patches of previous images as done in texture transfer. Each macroblock is identified as a texture or non texture block by means of an edge based criterion. The texture blocks are reconstructed using the luminance values as the control block information to guide the texture transfer algorithm. By modifying the texture transfer algorithm we ensure both spatial as well as temporal consistency. The non texture blocks and the luminance values for texture blocks can be coded by any encoder and thus our method can act as a pre-processor to any video codec e.g. H.264 and add to the compression obtained.
Aditya Khandelia, Santanu Chaudhury
ICIP2
2010 Document Image Retrieval Using Feature Combination in Kernel Space
abstract
The paper presents application of multiple features for word based document image indexing and retrieval. A novel framework to perform Multiple Kernel Learning for indexing using the Kernel based Distance Based Hashing is proposed. The Genetic Algorithm based framework is used for optimization. Two different features representing the structural organization of word shape are defined. The optimal combination of both the features for indexing is learned by performing MKL. The retrieval results for document collection belonging to Devanagari script are presented.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
ICPR2
2009 Shape Descriptor Based Document Image Indexing and Symbol Recognition
abstract
In this paper we present a novel shape descriptor based on shape context, which in combination with hierarchical distance based hashing is used for word and graphical pattern based document image indexing and retrieval. The shape descriptor represents the relative arrangement of points sampled on the boundary of the shape of object. We also demonstrate the applicability of the novel shape descriptor for classification of characters and symbols. For indexing, we provide anew formulation for distance based hierarchical locality sensitive hashing. Experiments have yielded promising results.
Ehtesham Hassan, Santanu Chaudhury, Madan Gopal
ICDAR2
2008 Video summarization with supervised learning
abstract
We present a video summarization technique based on supervised learning. Within a class of videos of similar nature, user provides the desired summaries for a subset of videos. Based on this supervised information, the summaries for other videos in the same class are generated. We derive frame-transitional features and subsequently represent each frame transition as a state. We then formulate a loss functional to quantify the discrepency between state transitional probabilities in the original video and that in the intended summary video, and optimize this functional. We experimentally validate the performance of the technique using cross-validation scores on two different class of videos, and demonstrate that the proposed technique is able to produce high quality summarization capturing the user perception.
Jayanta Basak, Varun Luthra, Santanu Chaudhury
ICPR3
2008 Parametric video compression using appearance space
abstract
The novelty of the approach presented in this paper is the unique object-based video coding framework for videos obtained from a static camera. As opposed to most existing methods, the proposed method does not require explicit 2D or 3D models of objects and hence is general enough to satisfy the need for varying types of objects in the scene. The proposed system detects and tracks an object in the scene by learning the appearance model of each object online using nontraditional uniform norm based subspace. At the same time the object is coded using the projection coefficients to the orthonormal basis of the subspace learnt. The tracker incorporates a predictive framework based upon Kalman filter for predicting the five motion parameters. The proposed method shows substantially better compression than MPEG2 based coding with almost no additional complexity.
Santanu Chaudhury, Subarna Tripathi, Sumantra Dutta Roy
ICPR1
2008 Bag-of-features kernel eigen spaces for classification
abstract
We present a classifier unifying local features based representation and subspace based learning. We also propose a novel method to merge kernel eigen spaces (KES) in feature space. Subspace methods have traditionally been used with the full appearance of the image. Recently local features based bag-of-features (BoF) representation has performed impressively on classification tasks. We use KES with BoF vectors to construct class specific subspaces and use the distance of a query vector from the database KESs as the classification criteria. The use of local features makes our approach invariant to illumination, rotation, scale, small affine transformation and partial occlusions. The system allows hierarchy by merging the KES in the feature space. The classifier performs competitively on the challenging Caltech-101 dataset under normal and simulated occlusion conditions. We show hierarchy on a dataset of videos collected over the internet.
Gaurav Sharma 0004, Santanu Chaudhury, J. B. Srivastava
ICPR2
2007 Word image based latent semantic indexing for conceptual querying in document image databases
abstract
In this paper we present an application of latent semantic analysis (LSA) for indexing and retrieval of document images with text. The query is specified as a set of word images and the documents which best match with the query representation in the the latent semantic space are retrieved. We show through extensive experiments on a large database that use of LSA for document images provides improvements in retrieval precision as is the case with electronic text documents.
Sameek Banerjee, Gaurav Harit, Santanu Chaudhury
ICDAR3
2007 Enhancement of Old Manuscript Images
abstract
In this paper we address the issue of enhancement in the quality of scanned images of old manuscripts. Small por- tions of the text in these manuscripts have degraded with time and are not readable. We propose a segmentation based histogram matching scheme for enhancing these de- graded text regions. To automatically identify the degraded text we use a matched wavelet based text extraction algo- rithm followed by MRF(Markov Random Field) post pro- cessing. Additionally we do background clearing to improve the quality of results. This method does not require any a priori information about the font, font size, background texture or geometric transformation. We have tested our method on a variety of manuscript images. The results show proposed method to be a robust, versatile and effective tool for enhancement of manuscript images.
Arbind K. Gupta, Ronak Gupta, Santanu Chaudhury, Shashank Joshi
ICDAR4
2007 Pàtrà: A Novel Document Architecture for Integrating Handwriting with Audio-Visual Information
abstract
In this paper we present P`atr`a - an integrated docu- ment architecture which incorporates handwritten illustra- tions captured and rendered in a temporal fashion synchro- nized with audio, video, text, and image data. The architec- ture of P`atr`a permits non-linear growth in the form of mul- tiple hierarchically organized play streams. Semantic meta- data is also an integral part of P`atr`a which serves a useful purpose of organizing such documents in a collection. We have developed an email application in which the users are provided with an authoring and rendering environment to compose, view, and reply to messages in the form of P`atr`a.
Gaurav Harit, V. Mankar, Santanu Chaudhury
ICDAR3
2007 On the view synthesis of man-made scenes using uncalibrated cameras
Geetika Sharma, Santanu Chaudhury, J. B. Srivastava
Pattern Recognit. Lett.2
2007 Video Shot Characterization Using Principles of Perceptual Prominence and Perceptual Grouping in Spatio-Temporal Domain
abstract
We present a novel approach for applying perceptual grouping principles to the spatio-temporal domain of video. Our perceptual grouping scheme, applied on blobs, makes use of a specified spatio-temporal coherence model. The grouping scheme identifies the blob cliques or perceptual clusters in the scene. We propose a computational model for analyzing a video shot based on a novel principle of perceptual prominence. The principle of perceptual prominence captures the key aspects of mise-en-scene required for characterizing a video scene.
Gaurav Harit, Santanu Chaudhury
IEEE Trans. Circuits Syst. Video Technol.2
2007 Text Extraction and Document Image Segmentation Using Matched Wavelets and MRF Model
abstract
In this paper, we have proposed a novel scheme for the extraction of textual areas of an image using globally matched wavelet filters. A clustering-based technique has been devised for estim ating globally matched wavelet filters using a collection of groundtruth images. We have extended our text extraction scheme for the segmentation of document images into text, background, and picture components (which include graphics and continuous tone images). Multiple, two-class Fisher classifiers have been used for this purpose. We also exploit contextual information by using a Markov random field formulation-based pixel labeling scheme for refinement of the segmentation results. Experimental results have established effectiveness of our approach.
Rajat Gupta, Nitin Khanna, Santanu Chaudhury, Shiv Dutt Joshi
IEEE Trans. Image Process.4
2006 Video Scene Interpretation Using Perceptual Prominence and Mise-en-scène Features
Gaurav Harit, Santanu Chaudhury
ACCV (2)2
2006 View Synthesis of Scenes with Multiple Independently Translating Objects from Uncalibrated Views
Geetika Sharma, Santanu Chaudhury, J. B. Srivastava
ACCV (1)2
2006 A Microscopic Framework For Distributed Object-Recognition & Pose-Estimation
abstract
Effective self-organization schemes lead to the creation of autonomous and reliable robot teams that can outperform a single, sophisticated robot on several tasks. We present here a novel, vision-based microscopic framework for active and distributed object-recognition and pose-estimation using a team of robots of simple construction. The team performs the task of locating a given object(s) in an unknown territory, recognizing it with sufficient confidence and estimating its pose. The larger goal is to experiment with probabilistic frameworks and graph-theoretic methods in the design of robot teams to achieve autonomous self-organization independent of the task at hand. We have chosen 3D object recognition as a first problem area to evaluate the effectiveness of our system design. The system comprises a probabilistic framework for the successful detection of the object in a coordinated manner and adaptive measures in case of machinery failures or presence of obstacles. A pose estimation method for the detected object and graph theoretic solutions for optimal field coverage by the robots are also presented. Each robot is provided with a part-based, spatial model of the object. The object to be recognized is taken to be much bigger than the robots and need not fit completely into the field of view of the robot cameras. We assume no knowledge of the internal parameters of the robot cameras and perform no camera calibration procedures. Initial simulation results corroborate our system design and field coverage methods
Sathyanarayan Anand, Ahmed Kirmani, Siddharth Shrivastava, Santanu Chaudhury, Basabi Bhaumik
ICARCV4
2006 Using Multimedia Ontology for Generating Conceptual Annotations and Hyperlinks in Video Collections
abstract
To enable seamless integration of video information on the semantic Web, we require that the knowledge of a video domain be formally specified in ontology. We present a novel approach for defining video domain concepts in an ontology using properties that can be observed from the media. We use the ontology specified knowledge for recognizing concepts relevant to a video scene by making observations for the media properties of concepts as well as making inferences from other ontological concept definitions and relations. For this purpose we introduce new language constructs to OWL (Web Ontology Language), which are used to specify the inherently uncertain nature of media observations. The new constructs also allow additional semantics concerned with the association of media properties with concepts. We propose the use of Bayesian network as the reasoning mechanism for doing inferencing tasks in the presence of uncertainty. The video is annotated with the relevant concepts defined in the ontology. These conceptual annotations are used to create hyperlinks in the video collection
Gaurav Harit, Santanu Chaudhury, Hiranmay Ghosh
Web Intelligence2
2006 Clustering in video data: Dealing with heterogeneous semantics of features
Gaurav Harit, Santanu Chaudhury
Pattern Recognit.2
2006 Multiple Exemplar-Based Facial Image Retrieval Using Independent Component Analysis
abstract
In this paper, we design a content-based image retrieval system where multiple query examples can be used to indicate the need to retrieve not only images similar to the individual examples, but also those images which actually represent a combination of the content of query images. We propose a scheme for representing content of an image as a combination of features from multiple examples. This scheme is exploited for developing a multiple example-based retrieval engine. We have explored the use of machine learning techniques for generating the most appropriate feature combination scheme for a given class of images. The combination scheme can be used for developing purposive query engines for specialized image databases. Here, we have considered facial image databases. The effectiveness of the image retrieval system is experimentally demonstrated on different databases.
Jayanta Basak, Koustav Bhattacharya, Santanu Chaudhury
IEEE Trans. Image Process.3
2005 Ontology Guided Access to Document Images
abstract
In this paper, we propose a scheme for accessing document images using ontology. We make use of an extension of OWL (ontology language for Web) to allow encoding of ontologies for document images. We experimentally demonstrate that reasoning with the concepts defined in ontology and their observation models provide a mechanism to support conceptual querying and automated hyperlinking of document images.
Gaurav Harit, Santanu Chaudhury, Jagrati Paranjpe
ICDAR2
2005 Improved Geometric Feature Graph: A Script Independent Representation of Word Images for Compression, and Retrieval
abstract
In this paper, we discuss a new representation scheme for word images which exploits the structural features. The word image features are represented in the form of a graph called as the geometric feature graph (GFG). The GFG is encoded in the form of a string which serves as a compressed representation of the word image skeleton. We demonstrate reconstruction, and retrieval of word images for 3 different scripts using the GFG string.
Gaurav Harit, Richa Jain, Santanu Chaudhury
ICDAR3
2005 Locating Text in Images using Matched Wavelets
abstract
In this paper we have proposed a novel scheme for locating text regions in an image. The method is based on multiresolution wavelet analysis. We used matched wavelets to capture textural characteristics of image regions. A clustering based approach has been proposed for estimating globally matched wavelets (GMWs) for a given collection of images. Using these GMWs, we generate feature vectors for segmentation and identification of text regions in an image. Our method, unlike most of the other methods, does not require any a priori information about the font, font size, scripts, geometric transformation, distortion or background texture. We have tested our method on various categories of images like license plates, posters, hand written documents and document images etc. The results show proposed method to be a robust, versatile and effective tool for text extraction from images.
Nitin Khanna, Santanu Chaudhury, Shiv Dutt Joshi
ICDAR3
2005 Novel view synthesis using a translating camera
Geetika Sharma, Ankita Kumar, Shakti Kamal, Santanu Chaudhury, J. B. Srivastava
Pattern Recognit. Lett.4
2005 Recognizing large isolated 3-D objects through next view planning using inner camera invariants
abstract
Most model-based three-dimensional (3-D) object recognition systems use information from a single view of an object. However, a single view may not contain sufficient features to recognize it unambiguously. Further, two objects may have all views in common with respect to a given feature set, and may be distinguished only through a sequence of views. A further complication arises when in an image, we do not have a complete view of an object. This paper presents a new online scheme for the recognition and pose estimation of a large isolated 3-D object, which may not entirely fit in a camera's field of view. We consider an uncalibrated projective camera, and consider the case when the internal parameters of the camera may be varied either unintentionally, or on purpose. The scheme uses a probabilistic reasoning framework for recognition and next-view planning. We show results of successful recognition and pose estimation even in cases of a high degree of interpretation ambiguity associated with the initial view.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
IEEE Trans. Syst. Man Cybern. Part B2
2004 Robust shape based two hand tracker
abstract
This paper presents a robust shape-based on-line tracker for simultaneously tracking the motion of both hands, that is robust to cases of background clutter, other moving objects, occlusions of one hand by the other and a wide range of illumination variations. The tracker is based on an online predictive eigentracking framework. This framework allows efficient tracking of articulate objects, which change in appearance across views. We show results of successful tracking across all possible cases of motion dynamics of both hands during occlusion and a wide range of illumination conditions.
Ketan Barhate, Kaustubh Patwardhan, Sumantra Dutta Roy, Subhasis Chaudhuri, Santanu Chaudhury
ICIP5
2004 On line predictive appearance-based tracking
abstract
We present a novel predictive statistical framework to improve the performance of an eigentracker. In addition, we use fast and efficient eigenspace updates to learn new views of the object being tracked on the fly. We also incorporate a new importance sampling mechanism which increases the robustness of the eigentracker and enables it to track nonconvex objects better. Our eigentracker is flexible-it is possible to use it symbolically with other trackers. We show its successful application in hand gesture analysis; and face and person tracking.
Namita Gupta, Pooja Mittal, Kaustubh Patwardhan, Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
ICIP5
2004 Active recognition through next view planning: a survey
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
Pattern Recognit.2
2004 Distributed and Reactive Query Planning in R-MAGIC: An Agent-Based Multimedia Retrieval System
abstract
We present the planning scheme for a cooperative agent-based multimedia retrieval architecture that integrates a heterogeneous set of repositories into a coherent information system. The agents in the system collaborate in context of a conceptual query to formulate unique retrieval strategies for the different collections. The retrieval plan makes need-based use of independent content analysis tools available on the network. The retrieval strategies for the repositories so formulated satisfy the specified constraints on quality of results and the response time requirements. The retrieval plan is reactively updated based on the retrieval performance at the individual repositories. We present some experimental results to show the effectiveness of the planning scheme for repositories with different characteristics and the scalability of the architecture. We present a prototype implementation of this architecture that integrates a set of dissimilar collections of multimedia data on Indian cultural heritage. A comparison of the retrieval results with some existing Internet search tools proves the effectiveness of the architecture.
Hiranmay Ghosh, Santanu Chaudhury
IEEE Trans. Knowl. Data Eng.2
2003 Devising Interactive Access Techniques for Indian Language Document Images
abstract
exist only in paper form. Web based interactive access techniques for images of these documents can ensure wider dissemination and easy availability. In this paper, we have proposed an access mechanism based on word based indexing and personalized annotation. The word based indexing scheme exploits typical structural characteristics of Indian scripts. We have combined this word indexing technique with personalized annotation based hyperlinking and query scheme for providing an interactive access interface to a collection of Indian language documents.
Santanu Chaudhury, Geetika Sethi, Anand Vyas, Gaurav Harit
ICDAR1
2003 A framework for video representation and transcoding using appearance spaces
abstract
We present a novel scheme for object-based video sequence presentation using appearance spaces. Our scheme enables fully automatic extraction of semantic video objects for a class of sequences, and their supervised organization in an object-class hierarchy. The hierarchy can be used for generic classification of query video objects, and transcoding using semantics of video objects.
Gaurav Harit, Santanu Chaudhury, Pramod Kumar Sharma 0001
ICME2
2003 Recognition of dynamic hand gestures
Aditya Ramamoorthy, Namrata Vaswani, Santanu Chaudhury, Subhashis Banerjee
Pattern Recognit.3
2003 Aspect graph construction with noisy feature detectors
abstract
Many three-dimensional (3D) object recognition strategies use aspect graphs to represent objects in the model base. A crucial factor in the success of these object recognition strategies is the accurate construction of the aspect graph, its ease of creation, and the extent to which it can represent all views of the object for a given setup. Factors such as noise and nonadaptive thresholds may introduce errors in the feature detection process. This paper presents a characterization of errors in aspect graphs, as well as an algorithm for estimating aspect graphs, given noisy sensor data. We present extensive results of our strategies applied on a reasonably complex experimental set, and demonstrate applications to a robust 3D object recognition problem.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
IEEE Trans. Syst. Man Cybern. Part B2
2002 Indexing for local appearance-based recognition of planar objects
Gaurav Jain, Santanu Chaudhury
Pattern Recognit. Lett.3
2001 Recognizing Large 3-D Objects through Next View Planning using an Uncalibrated Camera
abstract
We present a new on-line scheme for the recognition and pose estimation of a large isolated 3-D object, which may not entirely fit in a camera's field of view. We do not assume any knowledge of the internal parameters of the camera, or their constancy. We use a probabilistic reasoning framework for recognition and next view planning. We show results of successful recognition and pose estimation even in cases of a high degree of interpretation ambiguity associated with the initial view.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
ICCV2
2001 A Model Guided Document Image Analysis Scheme
abstract
This paper presents a new model-based document image segmentation scheme that uses XML-DTDs (eXtensible Markup Language Document Type Definitions). Given a document image, the algorithm has the ability to select the appropriate model. A new wavelet-based tool has been designed for distinguishing text from non-text regions and characterization of font sizes. Our model-based analysis scheme makes use of this tool for identifying the logical components of a document image.
Gaurav Harit, Santanu Chaudhury, Neeti Vohra, Shiv Dutt Joshi
ICDAR2
2001 Reconstruction Based Recognition of Scenes with Multiple Repeated Components
Ragini Choudhury, Santanu Chaudhury, J. B. Srivastava
Comput. Vis. Image Underst.2
2001 Reconstruction-Based Recognition of Scenes with Translationally Repeated Quadrics
abstract
This paper addresses the problem of invariant-based recognition of quadric configurations from a single image. These configurations consist of a pair of rigidly connected translationally repeated quadric surfaces. This problem is approached via a reconstruction framework. A new mathematical framework, using relative affine structure, on the lines of Luong and Vieville (1996), has been proposed. Using this mathematical framework, translationally repeated objects have been projectively reconstructed, from a single image, with four image point correspondences of the distinguished points on the object and its translate. This has been used to obtain a reconstruction of a pair of translationally repeated quadrics. We have proposed joint projective invariants of a pair of proper quadrics. For the purpose of recognition of quadric configurations, we compute these invariants for the pair of reconstructed quadrics. Experimental results on synthetic and real images, establish the discriminatory power and stability of the proposed invariant-based recognition strategy. As a specific example, we have applied this technique for discriminating images of monuments which are characterized by translationally repeated domes modeled as quadrics.
Ragini Choudhury, J. B. Srivastava, Santanu Chaudhury
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Image based document authentication using DCT
Pramod Kumar Sharma 0001, Santanu Chaudhury
Pattern Recognit. Lett.3
2001 A fuzzy theoretic approach for video segmentation using syntactic features
Rakesh Singh Jadon, Santanu Chaudhury, Kanad K. Biswas
Pattern Recognit. Lett.2
2000 Isolated 3D object recognition through next view planning
abstract
In many cases, a single view of an object may not contain sufficient features to recognize it unambiguously. This paper presents a new online recognition scheme based on next view planning for the identification of an isolated 3D object using simple features. The scheme uses a probabilistic reasoning framework for recognition and planning. Our knowledge representation scheme encodes feature based information about objects as well as the uncertainty in the recognition process. This is used both in the probability calculations as well as in planning the next view. Results clearly demonstrate the effectiveness of our strategy for a reasonably complex experimental set.
Sumantra Dutta Roy, Santanu Chaudhury, Subhashis Banerjee
IEEE Trans. Syst. Man Cybern. Part A2
1999 Trainable Script Identification Strategies for Indian Languages
abstract
Identification of the script in an image of a document page is of primary importance for a system processing multi-lingual documents. In this paper three trainable classification schemes have been proposed for identification of Indian scripts. The first scheme is based upon a frequency domain representation of the horizontal profile of the textual blocks. The other two schemes use connected components extracted from the textual region. We have proposed a novel Gabor filter-based feature extraction scheme for the connected components. We have also found that frequency distribution of the width-to-height ratio of the connected components can also be used for script recognition. It has been experimentally found that the Gabor filter-based scheme provides the most reliable performance. However, the other two techniques are computationally more efficient.
Santanu Chaudhury, Rabindra Sheth
ICDAR1
1999 Recognition of partially occluded objects using neural network based indexing
Navin Rajpal, Santanu Chaudhury, Subhashis Banerjee
Pattern Recognit.2
1999 An MIMD algorithm for constant curvature feature extraction using curvature based data partitioning
Santanu Chaudhury, Anjana Roy, Lipika Dey
Pattern Recognit. Lett.1
1997 Signature verification using multiple neural classifiers
Reena Bajaj, Santanu Chaudhury
Pattern Recognit.2
1997 Matching structural shape descriptions using genetic algorithms
Montek Singh, Amitabha Chatterjee, Santanu Chaudhury
Pattern Recognit.3
1996 An Abductive Reasoning Based Image Interpretation System
abstract
This paper describes an abductive reasoning based inferencing engine for image interpretation. The inferencing strategy finds an acceptable and consistent explanation of the features detected in the image in terms of the objects known a priori. The inferencing scheme assumes representation of the domain knowledge about the objects in terms of local and/or relational features. The inferencing system can be applied for different types of image interpretation problems like 2-D and 3-D object recognition, aerial image interpretation, etc. In this paper, we illustrate functioning of the system with the help of a 2-D object recognition problem.
Santanu Chaudhury, Arbind K. Gupta, Parthasarathy Guturu
Int. J. Pattern Recognit. Artif. Intell.1
1996 On adaptive trajectory tracking of a robot manipulator using inversion of its neural emulator
abstract
This paper is concerned with the design of a neuro-adaptive trajectory tracking controller. The paper presents a new control scheme based on inversion of a feedforward neural model of a robot arm. The proposed control scheme requires two modules. The first module consists of an appropriate feedforward neural model of forward dynamics of the robot arm that continuously accounts for the changes in the robot dynamics. The second module implements an efficient network inversion algorithm that computes the control action by inverting the neural model. In this paper, a new extended Kalman filter (EKF) based network inversion scheme is proposed. The scheme is evaluated through comparison with two other schemes of network inversion: gradient search in input space and Lyapunov function approach. Using these three inversion schemes the proposed controller was implemented for trajectory tracking control of a two-link manipulator. Simulation results in all cases confirm the efficacy of control input prediction using network inversion. Comparison of the inversion algorithms in terms of tracking accuracy showed the superior performance of the EKF based inversion scheme over others.
Laxmidhar Behera, Madan Gopal, Santanu Chaudhury
IEEE Trans. Neural Networks3
1994 A connectionist approach for clustering with applications in image analysis
abstract
A new neural network strategy for clustering is presented. The network works on the histogram and the process is similar to mode separation. The number of clusters are autonomously detected by the network and it overcomes some major difficulties encountered by mode separation techniques. Clustering is done by first selecting the prototypes and then assigning patterns to one of the prototypes based on its distance from the prototype and the distribution of data. The network does not employ weight learning and is therefore faster than existing unsupervised learning networks. The network was applied to a wide class of problems including gray level image reduction, color segmentation and remotely sensed image segmentation. The experimental results obtained are promising.>
V. V. Vinod, Santanu Chaudhury, J. Mukherjee, Sujoy Ghose
IEEE Trans. Syst. Man Cybern.2
1993 Matching of Structural Shape Descriptions with Hopfield Net
abstract
Structural description of objects comprised descriptions of the parts and spatial relations between the parts. This paper presents a Hopfield net based scheme for matching structural shape descriptions. The current formulation of the matching scheme is general enough to take care of partial mismatch between the individual parts and spatial constraints between these parts. In addition, a transformation of the shape descriptions has been suggested with which shape descriptions containing asymmetrical spatial constraints between the parts can be matched using symmetric interconnection weights for the Hopfield net. The Hopfield net based formulation has been extended to consider the problem of finding the best match of the test shape descriptions with one of the stored prototypes. The matching scheme has been experimentally applied for recognition of hand-tools and symbols. In both cases, the network produced encouraging recognition results.
Jayanta Basak, Santanu Chaudhury, Sankar K. Pal, D. Dutta Majumder
Int. J. Pattern Recognit. Artif. Intell.2
1993 Abductive formalism for two-dimensional object recognition
Santanu Chaudhury, Parthasarathy Guturu
Inf. Sci.1
1993 Bengali alpha-numeric character recognition using curvature features
Abhijit Dutta, Santanu Chaudhury
Pattern Recognit.2
1993 A new approach for aggregating edge points into line segments
Arbind K. Gupta, Santanu Chaudhury, Parthasarathy Guturu
Pattern Recognit.2
1993 A connectionist model for category perception: theory and implementation
abstract
A connectionist model for learning and recognizing objects (or object classes) is presented. The learning and recognition system uses confidence values for the presence of a feature. The network can recognize multiple objects simultaneously when the corresponding overlapped feature train is presented at the input. An error function is defined, and it is minimized for obtaining the optimal set of object classes. The model is capable of learning each individual object in the supervised mode. The theory of learning is developed based on some probabilistic measures. Experimental results are presented. The model can be applied for the detection of multiple objects occluding each other.
Jayanta Basak, Late C. A. Murthy, Santanu Chaudhury, D. Dutta Majumder
IEEE Trans. Neural Networks3
1992 A connectionist network for simultaneous perception of multiple categories
abstract
A connectionist network is presented for simultaneous perception of multiple categories. These categories provide an adequate explanation of the input features originating from multiple classes. The network optimises an appropriately defined error function for making the inference. A supervised learning algorithm is presented for learning the association between the features and each individual category.>
Jayanta Basak, Late C. A. Murthy, Santanu Chaudhury, D. Dutta Majumder
ICPR (2)3
1992 A Hough transform based approach to polyline approximation of object boundaries
abstract
A technique is presented for obtaining a polygonal approximation of object contours directly from the edge images. The technique is based on a new formulation of the Hough transform (HT) for aggregation of edge points into line segments. The space requirement of the HT is brought down by considering a different parameterization of straight lines. In this method, the process of edge linking and boundary approximation are combined into a single algorithm. Consequently, the scheme is computationally more efficient than the classical boundary approximation techniques which require use of a separate edge linking algorithm. Experimental results highlight the effectiveness of this method for approximating object boundaries of polygonal as well as curved shapes present in the images of complex multi-object scenes.>
Arbind K. Gupta, Santanu Chaudhury, Parthasarathy Guturu
ICPR (3)2
1992 A connectionist approach for gray level image segmentation
abstract
A connectionist network is presented for segmenting gray level images. The network detects the local peaks in the inverted histogram which will correspond to the bottoms of the valleys in the actual histogram. The neural network implementation successfully uses circumstantial evidence and detects multiple winners over the entire range of gray values such that these winners correspond to multiple thresholds for segmenting the image. The dynamics of the network has been analyzed and the conditions for convergence have been established. Experimental results obtained by applying the network for segmenting two X-ray images are presented.>
V. V. Vinod, Santanu Chaudhury, J. Mukherjee, Sujoy Ghose
ICPR (3)2
1992 A connectionist approach for peak detection in Hough space
V. V. Vinod, Santanu Chaudhury, Sujoy Ghose, J. Mukherjee
Pattern Recognit.2
1990 Recognition of Partial Planar Shapes in Limited Memory Environments
abstract
Industrial vision systems should be capable of recognising noisy objects, partially occluded objects and randomly located and/or oriented objects. This paper considers the problem of recognition of partially occluded planar shapes using contour segment-based features. None of the techniques suggested in the literature for solving the above problem guarantee reliable results for problem instances which require memory in excess of what is available. In this paper, a heuristic search-based recognition algorithm is presented, which guarantees reliable recognition results even when memory is limited. This algorithm identifies an object, the maximum portion of whose contour is visible in a conglomerate of objects. For increasing efficiency of the method, a two-stage recognition scheme has been designed. In the first phase, a relevant subset of the known model shapes is chosen and in the second stage, matching between the unknown shape and elements of the relevant subset is attempted using the above approach. The technique is general in the sense that it can be used with any kind of contour features. To evaluate the efficiency of the method, experimentation was carried out using polygonal approximations of the object contours. Results are cited for establishing the effectiveness of the approach.
Santanu Chaudhury, Parthasarathy Guturu
Int. J. Pattern Recognit. Artif. Intell.1
1990 Recognition of occluded objects with heuristic search
Santanu Chaudhury, Arup Acharyya, Parthasarathy Guturu
Pattern Recognit.1