Ujjwal Bhattacharya

dblp:b/UBhattacharya · DBLP profile ↗
← Back
64ranked-venue papers
15as first author
14since 2021 · last 2026
0000-0002-8546-6453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 14 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 18 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiFGANet: Lightweight Frequency and Gradient Aware Network for Robust Image Classification
Jotiraditya Banerjee, Utathya Aich, Ujjwal Bhattacharya
ICPR (15)3
2026 CAD-GAN+: Classifier-Filtered Synthetic CMRI Generation Towards Robust Detection of CAD
Shubham Basak, Nushrat Hussain, Ujjwal Bhattacharya
ICPR (12)3
2025 Cross-Modal Attention with Adaptive and Hierarchical Fusion for Robust RGB-T Image Segmentation for Safe Driving
abstract
Autonomous vehicles (AVs) rely on multiple sensors to ensure safe driving, particularly in complex environments with varying illumination and adverse weather conditions. RGB images perform well under ideal lighting conditions but often suffer from degraded quality in low-light scenarios, whereas thermal images maintain consistent quality across varying lighting conditions. To leverage the complementary strengths of these two imaging modalities, we propose a novel RGB-Thermal image segmentation framework based on an encoder-decoder architecture. At the core of the framework lies the Cross-Modal Attention Module (CMAM), which consists of a sequence of cross-modal attention blocks and are then passed through an Adaptive Aggregation Module (AAM). Finally, the output feature maps at different spatial resolutions are aggregated hierarchically to mitigate suboptimal results. Our proposed method has been evaluated on the publicly available MFNet dataset, where it outperforms SOTA segmentation methods by at least 1.1% in mAcc and 2.8% in mIoU, contributing to advancements in autonomous driving safety technologies.
Saheli Hazra, Ujjwal Bhattacharya
ICIP2
2025 Reflective Teacher: Semi-Supervised Multimodal 3D Object Detection in Bird's-Eye-View via Uncertainty Measure
abstract
Applying pseudo labeling techniques has been found to be advantageous in semi-supervised 3D object detection (SSO D) in Bird’ s-Eye-View (BEV) for autonomous driving, particularly where labeled data is limited. In the literature, Exponential Moving Average (EMA) has been used for adjustments of the weights of teacher network by the student network. However, the same induces catastrophic forgetting in the teacher network. In this work, we address this issue by introducing a novel concept of Reflective Teacher where the student is trained by both labeled and pseudo labeled data while its knowledge is progressively passed to the teacher through a regularizer to ensure retention of previous knowledge. Additionally, we propose Geometry Aware BEV Fusion (GA-BEVFusion) for efficient alignment of multi-modal BEV features, thus reducing the disparity between the modalities-camera and LiDAR. This helps to map the precise geometric information embedded among LiDAR points reliably with the spatial priors for extraction of semantic information from camera images. Our experiments on the nuScenes and Waymo datasets demonstrate: 1) improved performance over state-of-the-art methods in both fully supervised and semi-supervised settings; 2) Reflective Teacher achieves equivalent performance with only 25% and 22% of labeled datafor nuScenes and Waymo datasets respectively, in contrast to other fully supervised methods that utilize the full labeled dataset.
Saheli Hazra, Rohit Choudhary, Ganesh Sistu, Ciarán Eising, Ujjwal Bhattacharya
WACV7
2025 Deep determination of cardiac condition from phonocardiograms
Shubham Basak, Ujjwal Bhattacharya
Neural Comput. Appl.2
2024 MMPrune4U: Regularizing Multimodal Feature Distortion in Weight Pruning for Deep Neural Network Compression
Kaixin Xu, Nushrat Hussain, Ziyuan Zhao, Weisi Lin, Ujjwal Bhattacharya
BMVC7
2024 YOLO Assisted A* Algorithm for Robust Line Segmentation of Degraded Document Images
Ahana Kundu, Ujjwal Bhattacharya
ICDAR (4)2
2023 Revisiting Modality Imbalance In Multimodal Pedestrian Detection
abstract
Multimodal learning, particularly for pedestrian detection, has recently received emphasis due to its capability to function equally well in several critical autonomous driving scenarios such as low-light, night-time, and adverse weather conditions. However, in most cases, the training distribution largely emphasizes the contribution of one specific input that makes the network biased towards one modality. Hence, the generalization of such models becomes a significant problem where the non-dominant input modality during training could be contributing more to the course of inference. Here, we introduce a novel training setup with regularizer in the multimodal architecture to resolve the problem of this disparity between the modalities. Specifically, our regularizer term helps to make the feature fusion method more robust by considering both the feature extractors equivalently important during the training to extract the multimodal distribution which is referred to as removing the imbalance problem. Furthermore, our decoupling concept of output stream helps the detection task by sharing the spatial sensitive information mutually. Extensive experiments of the proposed method on KAIST and UTokyo datasets shows improvement of the respective state-of-the-art performance.
Ganesh Sistu, Jonathan Horgan, Ujjwal Bhattacharya, Edward Jones, Martin Glavin, Ciarán Eising
ICIP5
2023 Weakly supervised deep metric learning on discrete metric spaces for privacy-preserved clustering
Chandan Biswas, Debasis Ganguly, Dwaipayan Roy 0001, Ujjwal Bhattacharya
Inf. Process. Manag.4
2022 Privacy-aware supervised classification: An informative subspace based multi-objective approach
Chandan Biswas, Debasis Ganguly, Partha Sarathi Mukherjee, Ujjwal Bhattacharya, Yufang Hou 0001
Pattern Recognit.4
2022 A Deep OCR for Degraded Bangla Documents
abstract
Despite the significant success of document image analysis techniques, efficient Optical Character Recognition (OCR) of degraded document images still remains an open problem. Although a body of work has been reported on degraded document recognition for English language, only little attention has been paid to Indic scripts. In this work, we focus on developing a degraded OCR for Bangla, a major Indian language. In general, an OCR system includes segmentation of the foreground text part from the background followed by recognition of the extracted text. The text segmentation module aims to assign the foreground or background label to each pixel of the document image. In this paper, we present a new OCR system which is particularly suitable for degraded quality Bangla document images. The contribution is two fold. In the first phase, we use a semi-supervised Markov Random Field (MRF)- based Generative Adversarial Network (GAN) model (which we call MRF-GAN ) for foreground segmentation of texts from degraded text. In the proposed MRF-GAN , we extend the concept of GAN to a multitask learning mechanism where discriminator-classifier networks differentiate between real/fake images and also assign a foreground or background label to each pixel. In the second phase, we propose to use a new encoder-decoder based recognizer that incorporates an attention-based character to a word prediction model, which has the capability of minimizing Word Error Rate (WER) . We optimize this network using a Multitask based Transfer Learning scheme (MTTL) . We perform experiments on a publicly available degraded Bangla document image dataset as well as on a new degraded printed Hindi document image dataset, which has been created as a part of the present study. Results of the experimentations demonstrate the efficacy of the proposed OCR.
Ayan Chaudhury, Partha Sarathi Mukherjee, Chandan Biswas, Ujjwal Bhattacharya
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2022 Spatio-Contextual Deep Network-Based Multimodal Pedestrian Detection for Autonomous Driving
abstract
Pedestrian Detection is the most critical module of an Autonomous Driving system. Although a camera is commonly used for this purpose, its quality degrades severely in low-light night time driving scenarios. On the other hand, the quality of a thermal camera image remains unaffected in similar conditions. This paper proposes an end-to-end multimodal fusion model for pedestrian detection using RGB and thermal images. Its novel spatio-contextual deep network architecture is capable of exploiting the multimodal input efficiently. It consists of two distinct deformable ResNeXt-50 encoders for feature extraction from the two modalities. Fusion of these two encoded features takes place inside a multimodal feature embedding module (MuFEm) consisting of several groups of a pair of Graph Attention Network and a feature fusion unit. The output of the last feature fusion unit of MuFEm is subsequently passed to two CRFs for their spatial refinement. Further enhancement of the features is achieved by applying channel-wise attention and extraction of contextual information with the help of four RNNs traversing in four different directions. Finally, these feature maps are used by a single-stage decoder to generate the bounding box of each pedestrian and the score map. We have performed extensive experiments of the proposed framework on three publicly available multimodal pedestrian detection benchmark datasets, namely KAIST, CVC-14, and UTokyo. The results on each of them improved the respective state-of-the-art performance. A short video giving an overview of this work along with its qualitative results can be seen athttps://youtu.be/FDJdSifuuCs. Our source code will be released upon publication of the paper.
Kinjal Dasgupta, Ujjwal Bhattacharya, Senthil Kumar Yogamani
IEEE Trans. Intell. Transp. Syst.4
2021 AudViSum: Self-Supervised Deep Reinforcement Learning for Diverse Audio-Visual Summary Generation
Sanjoy Chowdhury, Aditya Patra, Subhrajyoti Dasgupta, Ujjwal Bhattacharya
BMVC4
2021 Listen To The Pixels
abstract
Performing sound source separation and visual object segmentation jointly in naturally occurring videos is a notoriously difficult task, especially in the absence of annotated data. In this study, we leverage the concurrency between audio and visual modalities in an attempt to solve the joint audio-visual segmentation problem in a self-supervised manner. Human beings interact with the physical world through a few sensory systems such as vision, auditory, movement, etc. The usefulness of the interplay of such systems lies in the concept of degeneracy [1]. It tells us that the cross-modal signals can educate each other without the presence of an external supervisor. In this work, we efficiently exploit this fact that learning from one modality inherently helps to find patterns in others by introducing a novel audio-visual fusion technique. Also, to the best of our knowledge, we are the first to address the partially occluded sound source segmentation task. Our study shows that the proposed model significantly outperforms existing state-of-the-art methods in both visual and audio source separation tasks.
Sanjoy Chowdhury, Subhrajyoti Dasgupta, Ujjwal Bhattacharya
ICIP4
2020 An End-To-End Framework For Pose Estimation Of Occluded Pedestrians
abstract
Pose estimation in the wild is a challenging problem, particularly in situations of(i) occlusions of varying degrees, and (ii) crowded outdoor scenes. Most of the existing studies of pose estimation did not report the performance in similar situations. Moreover, pose annotations for occluded parts of the human figures have not been provided in any of the relevant standard datasets, which in turn creates further difficulties to the required studies for pose estimation of the entire Figure for occluded humans. Well known pedestrian detection datasets such as CityPersons contains samples of outdoor scenes but it does not include pose annotations. Here we propose a novel multi-task framework for end-to-end training towards the entire pose estimation of pedestrians including in situations of any kind of occlusion. To tackle this problem, we make use of a pose estimation dataset, MS-COCO, and employ unsupervised adversarial instance-level domain adaptation for estimating the entire pose of occluded pedestrians. The experimental studies show that the proposed framework outperforms the SOTA results for pose estimation, instance segmentation and pedestrian detection in cases of heavy occlusions (HO) and reasonable + heavy occlusions (R+HO) on the two benchmark datasets.
Perla Sai Raj Kishore, Ujjwal Bhattacharya
ICIP3
2020 Scale-Invariant Multi-Oriented Text Detection in Wild Scene Image
abstract
Automatic detection of scene texts in the wild is a challenging problem, particularly due to the difficulties in handling (i) occlusions of varying percentages, (ii) widely different scales and orientations, (iii) severe degradation in the image quality etc. Here, we propose a deep architecture consisting of a novel Feature Representation Block (FRB) capable of efficient abstraction of the input information. Also, we consider curriculum learning with respect to difficulties in image samples and gradual increase in pixel-wise blurring. Our framework is capable of detecting texts of different scales and orientations. It can tackle blurring from multiple possible sources, non-uniform illumination as well as partial occlusions of varying percentages. Text detection performance of the proposed framework on various benchmark datasets including ICDAR 2015, ICDAR 2017 MLT, MSRA-TD500 and COCO-Text has significantly improved the respective state-of-the-art results.
Kinjal Dasgupta, Ujjwal Bhattacharya
ICIP3
2020 DenseRecognition of Spoken Languages
abstract
In the present study, we have considered a large number (27) of Indian languages for recognition from their speech signals of different sources. A dense convolutional network architecture (DenseNet) has been used for this classification task. Dynamic elimination of low energy frames from the input speech signal has been considered as a preprocessing operation. Mel-spectrogram of pre-processed speech signal is fed as input to the DenseNet architecture. Language recognition performance of this architecture has been compared with that of several state-of-the-art deep architectures which include a convolutional neural network (CNN), ResNet, CNN-BLSTM and DenseNet-BLSTM hybrid architectures. Additionally, we obtained recognition performances of a stacked BLSTM architecture fed with different sets of handcrafted features for comparison purpose. Simulations for both speaker independent and speaker dependent scenarios have been performed on two different standard datasets which include (i) IITKGP-MLILSC dataset of news clips in 27 different Indian languages and (ii) Linguistic Data Consortium (LDC) dataset of telephonic conversations in 5 different Indian languages. In each case, recognition performance of the DenseNet architecture along with Mel-spectrogram features has been found to be significantly better than all other frameworks implemented in this study.
Jaybrata Chakraborty, Bappaditya Chakraborty, Ujjwal Bhattacharya
ICPR3
2020 Stratified Multi-Task Learning for Robust Spotting of Scene Texts
abstract
Gaining control over the dynamics of multi-task learning should help to unlock the potential of the deep network to a great extent. In the existing multi-task learning (MTL) approaches of deep network, all the parameters of its feature encoding part are subjected to adjustments corresponding to each of the underlying sub-tasks. On the other hand, different functional areas of the human brain are responsible for distinct functions such as the Broca's area of the cerebrum is responsible for speech formation whereas its Wernicke's area is related to the language development etc. Inspired by this fact, in the present study, we used here a Feature Representation Block (FRB) of connection weights spanned over a few successive layers of a deep multi-task learning architecture and stratify the same into distinct subsets for their adjustments exclusively corresponding to different sub-tasks. Additionally, we have introduced a novel regularization component for controlled training of this Feature Representation Block. The purpose of the development of this learning framework is efficient end-to-end recognition of scene texts. Simulation results of the proposed strategy on various benchmark scene text datasets such as ICDAR 2015, ICDAR 2017 MLT, COCO-Text and MSRA-TD500 have improved respective SOTA performance.
Kinjal Dasgupta, Ujjwal Bhattacharya
ICPR3
2020 CardioGAN: An Attention-based Generative Adversarial Network for Generation of Electrocardiograms
abstract
Electrocardiogram (ECG) signal is studied to obtain crucial information about the condition of a patient's heart. Machine learning based automated medical diagnostic systems that may help to evaluate the condition of the heart from this signal are required to be trained using large volumes of labelled training samples and the same may increase the chance of compromising with the patients' privacy. To solve this issue, generation of synthetic electrocardiogram signals by learning only from the general distributions of the available real training samples have been attempted in the literature. However, these studies did not pay necessary attention to the specific vital details of these signals, such as the P wave, the QRS complex, and the T wave. This shortcoming often results in the generation of unrealistic synthetic signals, such as a signal which does not contain one or more of the above components. In the present study, a novel deep generative architecture, termed as CardioGAN, based on generative adversarial network and powered by the effective attention mechanism has been designed which is capable of learning the intricate inter-dependencies among the various parts of real samples leading to the generation of more realistic electrocardiogram signals. Also, it helps in reducing the risk of breaching the privacy of patients. Extensive experimentation performed by us establishes that the proposed method achieves a better performance in generating synthetic electrocardiogram signals in comparison to the existing methods.
Subhrajyoti Dasgupta, Ujjwal Bhattacharya
ICPR3
2019 ClueNet : A Deep Framework for Occluded Pedestrian Pose Estimation
Perla Sai Raj Kishore, Partha Sarathi Mukherjee, Ujjwal Bhattacharya
BMVC4
2019 Privacy Preserving Approximate K-means Clustering
abstract
Privacy preserving computation is of utmost importance in a cloud computing environment where a client often requires to send sensitive data to servers offering computing services over untrusted networks. Eavesdropping over the network or malware at the server may lead to leaking sensitive information from the data. To prevent this, we propose to encode the input data in such a way that, firstly, it should be difficult to decode it back to the true data, and secondly, the computational results obtained with the encoded data should not be substantially different from those obtained with the true data. Specifically, the computational activity that we focus on is the K-means clustering, which is widely used for many data mining tasks. Our proposed variant of the K-means algorithm is capable of privacy preservation in the sense that it requires as input only binary encoded data, and is not allowed to access the true data vectors at any stage of the computation. During intermediate stages of K-means computation, our algorithm is able to effectively process the inputs with incomplete information seeking to yield outputs relatively close to the complete information (non-encoded) case. Evaluation on real datasets show that the proposed methods yields comparable clustering effectiveness in comparison to the standard K-means algorithm on image clustering (MNIST-8M dataset), and in fact outperforms the standard K-means on text clustering (ODPtweets dataset).
Chandan Biswas, Debasis Ganguly, Dwaipayan Roy 0001, Ujjwal Bhattacharya
CIKM4
2018 Does Deeper Network Lead to Better Accuracy: A Case Study on Handwritten Devanagari Characters
abstract
Deep neural network architectures have been used successfully in various document analysis studies. Its strength in producing human like performance has already been explored in handwritten English numeral recognition task. In this context, a natural question that often arises in a practitioner's mind: does an increase in the depth of the network eventually lead to an improved recognition performance on unknown samples? A goal of the present work is to search for an answer of the same through a case study of a larger class handwriting recognition problem. Here, we have studied recognition of handwritten Devanagari characters. In this study, we have implemented convolutional neural network (CNN) architectures of five different depths. We have also implemented additional neural architectures by adding two Bidirectional Long Short Term Memory (BLSTM) layers between the convolutional stack and the fully connected part of each of these five CNN networks. Simulations have been performed on two different databases of handwritten Devanagari characters consisting of 30408 and 36172 samples and a combined set consisting of 58451 samples. The recognition accuracy obtained in the best case improves significantly the existing state-of-the-art of this handwriting recognition problem. Also, further analysis of our simulation results provides an answer to the above question. Additionally, we have trained a BLSTM network alone using the Histogram of Oriented Gradient (HOG) features. Performance of this architecture failed to compete with the performance of CNN-BLSTM hybrid architecture.
Bappaditya Chakraborty, Bikash Shaw, Jayanta Aich, Ujjwal Bhattacharya, Swapan K. Parui
DAS4
2018 An Efficient Feature Vector for Segmentation-Free Recognition of Online Cursive Handwriting Based on a Hybrid Deep Neural Network
abstract
It presents our recent study on segmentation-free recognition of unconstrained cursive online handwriting of Devanagari and Bangla, the two most popular Indic scripts. Here, we have devised an efficient algorithm for obtaining the core region of a handwritten word sample and compute a robust feature set which includes certain measures based on the knowledge of this core region. It also includes first and second order discrete derivatives which estimate certain geometric properties of discrete curves along the pen trajectory of the input online handwritten word sample. The proposed feature vector is obtained at each point on the trajectory of the preprocessed word sample based on a window (called vicinity) centered at the point. A hybrid deep neural network architecture consisting of a convolutional neural network (CNN), a bidirectional recurrent neural network (BiRNN) having Long Short Term Memory (LSTM) cells and a connectionist temporal classification (CTC) layer receives this high level feature vector as input for labelling the character sequence in the online word sample. This study shows that the proposed hybrid architecture recognizes online handwriting more efficiently than a BLSTM network alone.
Partha Sarathi Mukherjee, Ujjwal Bhattacharya, Swapan K. Parui
DAS2
2018 A Hybrid Deep Architecture for Robust Recognition of Text Lines of Degraded Printed Documents
abstract
During the last 20 years, significant research studies have been undertaken for automatic recognition of printed documents. The same is true for Bangla, a major Indian script. All these studies were mainly centered on comparatively well-behaved good quality printed documents. However, many of the large archives include significant volumes of older documents which are so degraded in their present form that they cannot be reasonably transcribed using the existing OCR (Optical Character Recognition) approaches. On the other hand, automatic recognition of printed contents of these documents has significant application potentials such as generation of descriptive metadata, full-text searching, information extraction etc. The contributions made in the present study are (i) creation of a moderately large annotated database of degraded Bangla documents towards their recognition studies, (ii) development of a Gaussian mixture model based strategy for extraction of text components from complex noisy background of such documents and (iii) development of a line level recognition scheme for degraded Bangla documents. We have studied two different CNN-BLSTM-CTC hybrid architectures for this recognition problem. The winning architecture uses the first convolution layer of the CNN in a fashion similar to the inception model of deep learning methodologies.
Chandan Biswas, Partha Sarathi Mukherjee, Koyel Ghosh, Ujjwal Bhattacharya, Swapan K. Parui
ICPR4
2018 Document Image Classification with Intra-Domain Transfer Learning and Stacked Generalization of Deep Convolutional Neural Networks
abstract
In this article, a region-based Deep Convolutional Neural Network framework is presented for document structure learning. The contribution of this work involves efficient training of region based classifiers and effective ensembling for document image classification. A primary level of `inter-domain' transfer learning is used by exporting weights from a pre-trained VGG16 architecture on the ImageNet dataset to train a document classifier on whole document images. Exploiting the nature of region based influence modelling, a secondary level of `intra-domain' transfer learning is used for rapid training of deep learning models for image segments. Finally, a stacked generalization based ensembling is utilized for combining the predictions of the base deep neural network models. The proposed method achieves state-of-the-art accuracy of 92.21% on the popular RVL-CDIP document image dataset, exceeding the benchmarks set by the existing algorithms.
Saikat Roy, Ujjwal Bhattacharya, Swapan K. Parui
ICPR3
2018 An HMM framework based on spherical-linear features for online cursive handwriting recognition
Oendrila Samanta, Anandarup Roy 0001, Swapan K. Parui, Ujjwal Bhattacharya
Inf. Sci.4
2017 A Hybrid Model for End to End Online Handwriting Recognition
abstract
Automatic recognition of online handwritten words in a generic mode has significant application potentials. However, this recognition job is challenging for unconstrained handwriting data. The challenge is more serious for Indic scripts like Devanagari or Bangla due to the inherent cursiveness of their characters, large sizes of respective alphabets, existence of several groups of shape similar characters etc. On the other hand, with the recent development of powerful machine learning tools, major research initiatives in this area of pattern recognition studies have been observed. Feature extraction and classification are two major modules of such a recognizer. Deep architectures of convolutional neural network (CNN) models have been found to be efficient in extraction of useful features from raw signal. On the other hand, a recurrent neural network (RNN) along with connectionist temporal classification (CTC) has been shown to be able to label unsegmented sequence data. In the present article, we propose a hybrid layered architecture consisting of three networks CNN, RNN and CTC for recognition of online handwriting without use of any specific lexicon. In this study, we have also observed that feeding hand-crafted features to the CNN at the first level of the proposed model provides better performance than feeding the raw signal to the CNN. We have simulated the proposed model on two large databases of Devanagari and Bangla online unconstrained handwritten words. The recognition accuracies provided by the proposed model are encouraging.
Partha Sarathi Mukherjee, Bappaditya Chakraborty, Ujjwal Bhattacharya, Swapan K. Parui
ICDAR3
2017 A sigma-lognormal model-based approach to generating large synthetic online handwriting sample databases
Ujjwal Bhattacharya, Réjean Plamondon, Souvik Dutta Chowdhury, Pankaj Goyal, Swapan K. Parui
Int. J. Document Anal. Recognit.1
2016 An End-to-End System for Bangla Online Handwriting Recognition
abstract
A few studies of online Bangla handwriting recognition such as isolated character recognition or limited vocabulary cursive word recognition are found in the literature. However, development of an end-to-end recognition system of unconstrained online Bangla handwritten texts has not been duly attempted so far. In the present report, we describe a similar system which takes a piece of continuous online handwritten Bangla texts as the input. It first segments the input texts into individual lines, each line into its constituent words and each word into sub-strokes. In the present study, 152 different symbols which include basic characters, character modifiers, frequently used conjunct characters, a few special characters and numerals have been considered. The entire set of sub-strokes obtained from the training sample set has been exhaustively studied by 3 experts and 76 different shapes of sub-strokes have been identified based on consensus among these experts. Also, it has been observed that a character may produce at most 3 sub-strokes. Since a piece of Bangla texts often contains either Bangla or English numerals, the present character set consists of both the numeral set and 3 numeral shapes are common to both the scripts. The proposed recognition system uses two classifiers, one for characters and the other for sub-strokes. Sub-strokes are fed to the character classifier in their temporal order. A single sub-stroke followed by two consecutive sub-strokes and finally three successive sub-strokes are passed to the character classifier and the first two top responses of the character classifier among the three cases are compared. If the difference is less than a threshold, the response of sub-stroke classifier is used to reach a final decision. The proposed system provided 94.3% character level accuracy on a test set consisting of 33,453 word samples written by 31 writers.
Soumik Bhattacharya, Durjoy Sen Maitra, Ujjwal Bhattacharya, Swapan K. Parui
ICFHR3
2016 Generalized stacking of layerwise-trained Deep Convolutional Neural Networks for document image classification
abstract
This article presents our recent study of a lightweight Deep Convolutional Neural Network (DCNN) architecture for document image classification. Here, we concentrated on training of a committee of generalized, compact and powerful base DCNNs. A support vector machine (SVM) is used to combine the outputs of individual DCNNs. The main novelty of the present study is introduction of supervised layerwise training of DCNN architecture in document classification tasks for better initialization of weights of individual DCNNs. Each DCNN of the committee is trained for a specific part or the whole document. Also, here we used the principle of generalized stacking for combining the normalized outputs of all the members of the DCNN committee. The proposed document classification strategy has been tested on the well-known Tobacco3482 document image dataset. Results of our experimentations show that the proposed strategy involving a considerably smaller network architecture can produce comparable document classification accuracies in competition with the state-of-the-art architectures making it more suitable for use in comparatively low configuration mobile devices.
Saikat Roy, Ujjwal Bhattacharya
ICPR3
2016 Multilingual scene character recognition with co-occurrence of histogram of oriented gradients
Shangxuan Tian, Ujjwal Bhattacharya, Shijian Lu, Bolan Su, Xiaohua Wei, Yue Lu 0001, Chew Lim Tan
Pattern Recognit.2
2015 CNN based common approach to handwritten character recognition of multiple scripts
abstract
There are many scripts in the world, several of which are used by hundreds of millions of people. Handwritten character recognition studies of several of these scripts are found in the literature. Different hand-crafted feature sets have been used in these recognition studies. However, convolutional neural network (CNN) has recently been used as an efficient unsupervised feature vector extractor. Although such a network can be used as a unified framework for both feature extraction and classification, it is more efficient as a feature extractor than as a classifier. In the present study, we performed certain amount of training of a 5-layer CNN for a moderately large class character recognition problem. We used this CNN trained for a larger class recognition problem towards feature extraction of samples of several smaller class recognition problems. In each case, a distinct Support Vector Machine (SVM) was used as the corresponding classifier. In particular, the CNN of the present study is trained using samples of a standard 50-class Bangla basic character database and features have been extracted for 5 different 10-class numeral recognition problems of English, Devanagari, Bangla, Telugu and Oriya each of which is an official Indian script. Recognition accuracies are comparable with the state-of-the-art.
Durjoy Sen Maitra, Ujjwal Bhattacharya, Swapan K. Parui
ICDAR2
2015 Script independent online handwriting recognition
abstract
The most general form of handwriting style is mixed cursive and this is the most difficult type in view of its automatic recognition. Similar handwriting style are prevalent in various scripts such as English, Arabic, Bengali etc. Handwriting recognition for such a script gets further difficult whenever its alphabet consists of a large number of characters like Bengali which has around 350 characters. Hidden Markov models (HMM) are the most popularly used architectures for similar recognition problems. However, the task becomes easy if the underlying lexicon depending upon the specific application is provided. In such situations, holistic or word-based recognition approach is adopted which does not require recognition of the constituent characters. On the other hand, the same task gets complicated as the lexicon size increases and / or it consists of many similar shape words. In a recent study [1] of similar situation, a fully connected non-homogeneous HMM has been used where its observation sequence was generated through explicit segmentation of the input word. In the present study, we have explored that the performance of this HMM-based recognition scheme is independent of both the script and the particular intelligent segmentation strategy. We implemented a novel segmentation scheme based on Discrete Curve Evolution algorithm [2] and two other existing segmentation methods on standard databases of English, Arabic and Bangla to arrive at the above conclusion. Statistical hypothesis testings of the simulation results further confirm the above claim.
Oendrila Samanta, Anandarup Roy 0001, Ujjwal Bhattacharya, Swapan K. Parui
ICDAR3
2014 Holistic Recognition of Online Handwritten Words Based on an Ensemble of SVM Classifiers
abstract
In this paper, we present our recent study of a data driven approach to combining multiple SVM classifiers with RBF kernels each being trained with a distinct feature vector. The SVM classifiers in our ensemble are ranked based on their increasing order of average performance on the validation sample sets. The outputs of the SVM classifiers are combined based on a weighted average strategy which uses the above ranks of the underlying SVMs to determine the respective weights. In the present study, we design four sets of different feature vectors representing online handwritten words. Simple concatenation of these feature vectors does not help much in improving the recognition accuracy compared to the best performing feature vector among the four. Thus, we train distinct SVM classifiers with different feature vectors and combine their outputs at the final stage. The proposed recognition strategy is implemented on a limited vocabulary recognition problem of unconstrained mixed cursive online handwritten Bangla words. It improves existing recognition accuracies on a moderately large database of similar word samples.
A. Srimany, Sumit Dutta Chowdhury, Ujjwal Bhattacharya, Swapan K. Parui
Document Analysis Systems3
2014 Automatic Detection of Handwritten Texts from Video Frames of Lectures
abstract
Automatic recognition of handwritten texts in video lectures has important applications. In video lectures, the presenter usually writes on white / colored board. The video camera often captures the writing board along with certain other objects possibly including the presenter itself. Recognition of handwritten texts from such a video frame requires prior detection of the region of texts in the frame. In this article, we present our recent study of text localization in such video lecture frames. Here, we use Scale Invariant Feature Transform (SIFT) descriptors densely over the entire region of the frame. The descriptors are located on a regular grid of 5 pixels following the usual practice and considered a uniform patch size of 60 × 60 pixels as its support on the basis of an empirical study. This SIFT descriptor at each location (grid point) is fed as a 128-dimensional input feature vector to a Multilayer Perceptron (MLP) network which gives response for each grid point as either text or non-text. Depending on certain aggregate response at each pixel we localize text regions in the input video frame. Next, we employ K-means clustering to detect the text components present in the localized region of the video frame. Finally, two simple rules are applied to decide certain possible detected text components as noise. We obtained encouraging simulation results of this approach on a variety of video lecture frames.
Purnendu Banerjee, Ujjwal Bhattacharya, Bidyut B. Chaudhuri
ICFHR2
2014 A Machine Learning Approach to Detection of Core Region of Online Handwritten Bangla Word Samples
abstract
Core region detection of handwritten cursive words is an important step towards their automatic recognition. Several preprocessing operations such as height normalization, slant estimation etc. Are often based on this core region. This is particularly useful for word recognition of major Indian scripts, which have large character sets. The main parts of majority of these characters belong to the core region that is bounded above by a headline and bounded below by an imaginary base line. Only a few such characters or their parts appear either above or below the core region. A few approaches are available in the literature for detection of such a core region of offline handwritten word samples of Latin script. Also, a similar region is often determined for recognition of images of printed Indian scripts. However, none of these approaches have studied detection of core region of an unconstrained online handwritten word. In this article, we propose a novel method for detection of the core region of online handwritten word samples of Bangla, a major Indian script. For this we first perform smoothing on the samples and then segment a stroke into sub strokes. We compute certain novel positional features from each such sub stroke. Using these features, a multilayer perceptron (MLP) is trained by back propagation (BP) algorithm. On the basis of the output of the MLP, we determine the position of both the headline and the baseline. We have tested this approach on a recently developed large database of online unconstrained handwriting Bangla word samples. The proposed approach would also work on similar samples of Devanagari, another major Indian script. Experimental results are encouraging.
S. Baral, Soumik Bhattacharya, Anirban Chakraborty 0002, Ujjwal Bhattacharya, Swapan K. Parui
ICFHR4
2014 Stroke Level User-Adaptation for Stroke Order Free Online Handwriting Recognition
abstract
In this article, a novel lightweight user-adaptive online handwriting recognition scheme has been presented. The present recognition approach is stroke order free. It is based on prior identification of the set of various strokes of different shapes used in writing the characters of the underlying alphabet utilizing a representative sample database. In this approach a very small number of prototypes of each stroke shape is used along with certain weighted DTW distance based nearest neighbour classifier to recognize the strokes in the input character. Individual characters are identified using a look-up table (LUT) each row of which corresponds to one character composed of a distinct set of stroke shapes. This LUT is formed using the representative sample set of the underlying character set. If a stroke does not find a close match in the training set or if the set of strokes for an input character does not find a corresponding entry in the LUT, user adaptation takes place using a modified Learning Vector Quantization (LVQ) method. The proposed scheme has been implemented in an Android-based Bangla handwriting recognizer [1] for handheld devices and its recognition performance is encouraging.
Debarshi Dutta, Aruni Roy Chowdhury, Ujjwal Bhattacharya, Swapan K. Parui
ICFHR3
2014 Combination of Features for Efficient Recognition of Offline Handwritten Devanagari Words
abstract
In this article, we describe our recent study of a novel combination of two feature vectors for holistic recognition of offline handwritten word images. In the literature, both contour and skeleton based feature representations have been studied for offline handwriting recognition purpose. However, to the best of our knowledge, there is no such study in which combination of the two feature representations have been considered for the purpose. In the proposed recognition scheme, we use multiclass SVM as the classifier. We have implemented the proposed approach for holistic recognition of Devanagari handwritten town names and tested its performance on a large handwritten word sample database of 100 Indian town names written in Devanagari. Experimental results show sharp improvement in recognition accuracy over the use of any of the individual feature representation schemes. The proposed approach is script independent and can be used for development of a holistic handwritten word image recognition of any script.
Bikash Shaw, Ujjwal Bhattacharya, Swapan K. Parui
ICFHR2
2014 A Global-to-Local Approach to Binarization of Degraded Document Images
abstract
This article deals with binarization of degraded document images. In the proposed approach, Canny edge image of the input degraded document image is obtained after blurring it with a Gaussian filter. Next, the gray values of the two pixels of the input image at the left and right of each edge pixel are noted to form a histogram of these gray values which possesses two distinct peaks and the lowest valley between them provides the global threshold value. Each pixel with gray value greater than the above threshold is turned as background pixel. A small square window is considered around each non-background pixel and certain simple statistics are computed on the gray values of the pixels of this small window based on which the said pixel is turned either background or foreground. Such a local thresholding method at the latter stage can efficiently handle various degradations in the document. The binarized image so obtained is finally subjected to certain common post-processing operations. The proposed method has been compared with a few existing binarization techniques.
Barun Biswas, Ujjwal Bhattacharya, Bidyut B. Chaudhuri
ICPR2
2014 Smoothing of HMM parameters for efficient recognition of online handwriting
Oendrila Samanta, Ujjwal Bhattacharya, Swapan K. Parui
Pattern Recognit.2
2013 Online Handwriting Recognition Using Levenshtein Distance Metric
abstract
In this article, we propose a novel scheme for online handwritten character recognition based on Levenshtein distance metric. Both shape and position information are considered in our feature representation scheme. The shape information is encoded by a string of quantized values of angular displacements between successive sample points along the trajectory of the handwritten character. The consecutive occurrences of same value in such a string are removed retaining only one of them. Next, each element in the resulting string is assigned an integral weight value proportional to the length of the segment of the trajectory represented by the corresponding element. Similarly, position information is encoded by another string of quantized positional information along with their respective weight values. We formulated a distance function based on Levenshtein metric to compute the similarity between an unknown character sample and each training sample. Here, we have also studied the effect of pruning the training sample set based on the above distance between individual training samples of the same character class. The proposed approach has been simulated on different publicly available sample databases of online handwritten characters. The recognition accuracies are acceptable.
Sumit Dutta Chowdhury, Ujjwal Bhattacharya, Swapan K. Parui
ICDAR2
2012 On the Enhancement and Binarization of Mobile Captured Vehicle Identification Number for an Embedded Solution
abstract
An embedded solution for automatic detection of Vehicle Identification Numbers (VIN) captured by a mobile camera has a number of real world applications. But the performance of available open source Optical Character Recognition (OCR) systems on VIN images captured by mobile phones is extremely poor because of the image quality affected by various noises. In a recent study of such images, we have observed that the performance of existing open source OCR systems can be improved by applying several image enhancement techniques on these images before sending them to the OCR engine. In this article, we have presented such a method that improves the recognition accuracy from 5.89% up to 82.3%.
Tanushyam Chattopadhyay, Ujjwal Bhattacharya, Bidyut B. Chaudhuri
Document Analysis Systems2
2012 A Semi-automatic Annotation Scheme for Bangla Online Mixed Cursive Handwriting Samples
abstract
Requirement of annotated handwriting samples for the development of relevant recognition algorithms is an established fact. Although such annotated databases of unconstrained handwriting exist for several scripts of a few languages, the same is not true for any of the scripts of India. As far as Indian scripts are concerned a few databases of handwritten isolated characters are publicly available. These include samples of both online and offline handwriting. However, no such publicly available database of unconstrained handwriting in any of the Indian scripts exists. On the other hand, unconstrained handwriting in Bangla, the second most popular among Indian scripts, is mixed cursive in nature unlike the other scripts of India. Thus, annotation of Bangla unconstrained handwriting samples needs special consideration. During the last few years our group at the Indian Statistical Institute, Kolkata has been working towards the development of a large annotated database of online Bangla handwriting samples and has developed a GUI-based semi-automatic scheme for their annotation at character boundary levels and a scheme for XML representation of such annotated data. The present system implemented for annotation of unconstrained handwriting of Bangla may easily be customized for other scripts. Currently this system is in use for annotation of a large database of Bangla unconstrained online handwriting.
Ujjwal Bhattacharya, R. Banerjee, S. Baral, R. De, Swapan K. Parui
ICFHR1
2012 HMM Based Online Handwritten Bangla Character Recognition Using Dirichlet Distributions
abstract
A reasonably large database of online handwritten Bangla characters has been developed. Such a handwritten character sample is composed of one or more strokes. Seventy five such stroke classes have been identified on the basis of the varying handwriting styles present in the character database. Each character sample is a sequence of strokes emanating from these stroke classes. Another database of handwritten Bangla strokes has been developed from the character database. This is the first such database for Bangla script. Certain stroke level features are defined on the basis of certain extremum points which represent the stroke shape reasonably well. The proposed character classification method is a two-stage approach. First, a probability distribution is estimated for each stroke class using the stroke features and then an HMM based character classifier is designed using each stroke class as a state. The parameters of both the stroke class distributions and the character class HMMs are estimated on the basis of the training set having 29,951 character samples. The character level recognition accuracy obtained by the proposed method on the test set having 8,616 samples, is 91.85%.
Chandan Biswas, Ujjwal Bhattacharya, Swapan K. Parui
ICFHR2
2012 Building a Personal Handwriting Recognizer on an Android Device
abstract
The wide usage of touch-screen based mobile devices has led to a large volume of the users preferring touch-based interaction with the machine, as opposed to traditional input via keyboards/mice. To exploit this, we focus on the Android platform to design a personalized handwriting recognition system that is acceptably fast, light-weight, possessing a user-friendly interface with minimally-intrusive correction and auto-personalization mechanisms. Since cursive writing on smaller screens is not usual, here we study non-cursive handwriting only. The recognition is done at character level using nearest-neighbor matching to a small, automatically user-adaptive and dynamically updating library of character-class template gestures.
Debarshi Dutta, Aruni Roy Chowdhury, Ujjwal Bhattacharya, Swapan K. Parui
ICFHR3
2012 Scene text detection using sparse stroke information and MLP
Aruni Roy Chowdhury, Ujjwal Bhattacharya, Swapan K. Parui
ICPR2
2012 Offline recognition of handwritten Bangla characters: an efficient two-stage approach
Ujjwal Bhattacharya, Malayappan Shridhar, Swapan K. Parui, P. K. Sen, Bidyut B. Chaudhuri
Pattern Anal. Appl.1
2010 Online Bangla Word Recognition Using Sub-Stroke Level Features and Hidden Markov Models
abstract
For automatic recognition of Bangla script, only a few studies are reported in the literature, which is in contrast to the role of Bangla as one of the world's major scripts. In this paper we present a new approach to online Bangla handwriting recognition and one of the first to consider cursively written words instead of isolated characters. Our method uses a sub-stroke level feature representation of the script and a writing model based on hidden Markov models. As for the latter an appropriate internal structure is crucial, we investigate different approaches to defining model structures for a highly compositional script like Bangla. In experimental evaluations of a writer independent Bangla word recognition task we show that the use of context-dependent sub-word units achieves quite promising results and significantly outperforms alternatively structured models.
Gernot A. Fink, Szilárd Vajda, Ujjwal Bhattacharya, Swapan K. Parui, Bidyut B. Chaudhuri
ICFHR3
2010 On-line Handwriting Recognition of Indian Scripts - The First Benchmark
abstract
Online handwriting recognition of Indian scripts has been drawing increasing attention in recent years. Related research has gained further momentum due to recent planned funding by the Govt. of India towards technology development of Indian languages and scripts. Standard databases of handwritten characters of a few Indian scripts have already become available. These include online handwritten character databases of Bangla, Devanagari, Tamil and Telugu and these are available free of cost on request. In the present paper, we present benchmark recognition results of the above databases of four most popular scripts of the Indian subcontinent based on two existing feature extraction methods viz. point-float and direction code histogram features and three classifiers viz. Nearest Neighbour (NN), Multilayer Perceptron (MLP) and Hidden Markov Model (HMM) to test the effectiveness of the existing classification methods and provide benchmark results for future online handwriting recognition research of these Indic scripts.
T. Mondal, Ujjwal Bhattacharya, Swapan K. Parui, D. Mandalapu
ICFHR2
2009 Devanagari and Bangla Text Extraction from Natural Scene Images
abstract
With the increasing popularity of digital cameras attached with various handheld devices, many new computational challenges have gained significance. One such problem is extraction of texts from natural scene images captured by such devices. The extracted text can be sent to OCR or to a text-to-speech engine for recognition. In this article, we propose a novel and effective scheme based on analysis of connected components for extraction of Devanagari and Bangla texts from camera captured scene images. A common unique feature of these two scripts is the presence of headline and the proposed scheme uses mathematical morphology operations for their extraction. Additionally, we consider a few criteria for robust filtering of text components from such scene images. Moreover, we studied the problem of binarization of such scene images and observed that there are situations when repeated binarization by a well-known global thresholding approach is effective. We tested our algorithm on a repository of 100 scene images containing texts of Devanagari and / or Bangla.
Ujjwal Bhattacharya, Swapan K. Parui, Srikanta Mondal
ICDAR1
2009 Handwritten Numeral Databases of Indian Scripts and Multistage Recognition of Mixed Numerals
abstract
This article primarily concerns the problem of isolated handwritten numeral recognition of major Indian scripts. The principal contributions presented here are (a) pioneering development of two databases for handwritten numerals of two most popular Indian scripts, (b) a multistage cascaded recognition scheme using wavelet based multiresolution representations and multilayer perceptron classifiers and (c) application of (b) for the recognition of mixed handwritten numerals of three Indian scripts Devanagari, Bangla and English. The present databases include respectively 22,556 and 23,392 handwritten isolated numeral samples of Devanagari and Bangla collected from real-life situations and these can be made available free of cost to researchers of other academic Institutions. In the proposed scheme, a numeral is subjected to three multilayer perceptron classifiers corresponding to three coarse-to-fine resolution levels in a cascaded manner. If rejection occurred even at the highest resolution, another multilayer perceptron is used as the final attempt to recognize the input numeral by combining the outputs of three classifiers of the previous stages. This scheme has been extended to the situation when the script of a document is not known a priori or the numerals written on a document belong to different scripts. Handwritten numerals in mixed scripts are frequently found in Indian postal mails and table-form documents.
Ujjwal Bhattacharya, Bidyut B. Chaudhuri
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Online handwritten Bangla character recognition using HMM
abstract
We describe here a novel scheme for recognition of online handwritten basic characters of Bangla, an Indian script used by more than 200 million people. There are 50 basic characters in Bangla and we have used a database of 24,500 online handwritten isolated character samples written by 70 persons. Samples in this database are composed of one or more strokes and we have collected all the strokes obtained from the training samples of the 50 character classes. These strokes are manually grouped into 54 classes based on the shape similarity of the graphemes that constitute the ideal character shapes. Strokes are recognized by using hidden Markov models (HMM). One HMM is constructed for each stroke class. A second stage of classification is used for recognition of characters using stroke classification results along with 50 look-up-tables (for 50 character classes).
Swapan K. Parui, Koushik Guin, Ujjwal Bhattacharya, Bidyut B. Chaudhuri
ICPR3
2007 Direction Code Based Features for Recognition of Online Handwritten Characters of Bangla
abstract
In the present article, we describe a novel direction code based feature extraction approach for recognition of online Bangla handwritten basic characters. We have implemented the proposed approach on a database of 7043 online handwritten Bangla (a major script of the Indian subcontinent) character samples, which has been developed by us. This is a 50-class recognition problem and we achieved 93.90% and 83.61% recognition accuracies respectively on its training and test sets.
Ujjwal Bhattacharya, B. K. Gupta, Swapan K. Parui
ICDAR1
2007 A Two Stage Recognition Scheme for Handwritten Tamil Characters
abstract
India is a multilingual multiscript country with more than 18 languages and 10 different major scripts. Not enough research work towards recognition of handwritten characters of these Indian scripts has been done. Tamil, an official as well as popular script of the southern part of India, Singapore, Malaysia, and Sri Lanka has a large character set which includes many compound characters. Only a few works towards handwriting recognition of this large character set has been reported in the literature. Recently, HP Labs India developed a database of handwritten Tamil characters. In the present paper, we describe an off-line recognition approach based on this database. The proposed method consists of two stages. In the first stage, we apply an unsupervised clustering method to create a smaller number of groups of handwritten Tamil character classes. In the second stage, we consider a supervised classification technique in each of these smaller groups for final recognition. The features considered in the two stages are different. The proposed two-stage recognition scheme provided acceptable classification accuracies on both the training and test sets of the present database.
Ujjwal Bhattacharya, Swapan K. Parui
ICDAR1
2005 Fusion of Combination Rules of an Ensemble of MLP Classifiers for Improved Recognition Accuracy of Handprinted Bangla Numerals
abstract
In handwritten character recognition problem, the input images are often affected by distortions and noise. Thus such images at different resolutions include different variations in the input data. In the present work, we considered wavelet transform to obtain multi-resolution representation of each input character image. At each resolution level, we considered three MLPs with different numbers of nodes in their hidden layers and combined the outputs produced by all the MLPs of the whole ensemble by using weighted sum rule, product rule and majority voting. The set of misclassified samples produced by one combination rule is neither a subset nor a superset of a similar set produced by another rule. So, majority voting has been used for the second and final round to produce final outputs after combining the results of the three combinations of the first stage. The proposed approach produced 99.10% correct recognition rate on the test set of Bangia (a major Indian script) numeral database.
Ujjwal Bhattacharya, Bidyut B. Chaudhuri
ICDAR1
2005 Databases for Research on Recognition of Handwritten Characters of Indian Scripts
abstract
Three image databases of handwritten isolated numerals of three different Indian scripts namely Devnagari, Bangla and Oriya are described in this paper. Grayscale images of 22556 Devnagari numerals written by 1049 persons, 12938 Bangla numerals written by 556 persons and 5970 Oriya numerals written by 356 persons form the respective databases. These images were scanned from three different kinds of handwritten documents - postal mails, job application form and another set of forms specially designed by the collectors for the purpose. The only restriction imposed on the writers is to write each numeral within a rectangular box. These databases are free from the limitations that they are neither developed in laboratory environments nor they are non-uniformly distributed over different classes. Also, for comparison purposes, each database has been properly divided into respective training and test sets.
Ujjwal Bhattacharya, Bidyut B. Chaudhuri
ICDAR1
2004 Recognition of Bangla Handwritten Characters Using an MLP Classifier Based on Stroke Features
Tapan Kumar Bhowmik, Ujjwal Bhattacharya, Swapan K. Parui
ICONIP2
2003 A Majority Voting Scheme for Multiresolution Recognition of Handprinted Numerals
abstract
This paper proposes a simple voting scheme for off-line recognition of handprinted numerals. One of the main features of the proposed scheme is that this is not script dependent. Another interesting feature is that it is sufficiently fast for real-life applications. In contrast to the usual practices, here we studied the efficiency of a majority voting approach when all the classifiers involved are multilayer perceptrons (MLP) of different sizes and respective features are based on wavelet transforms at different resolution levels. The rationale for this approach is to explore how one can improve the recognition performance without adding much to the requirements for computational time and resources. For simplicity and efficiency, in the present work, we considered only three coarse-to-fine resolution levels of wavelet representation. We primarily simulated the proposed technique on a database of off-line handprinted Bangla (a major Indian script) numerals. We achieved 97.16% correct recognition rate on a test set of 5000 Bangla numerals. In this simulation we used two other disjoint sets (one for training and the other for validation purpose) of sizes 6000 and 1000 respectively. We have also tested our approach on MNIST database for handwritten English digits. The result is comparable with state-of-the-art technologies.
Ujjwal Bhattacharya, Bidyut B. Chaudhuri
ICDAR1
2002 A Hybrid Scheme for Handprinted Numeral Recognition Based on a Self-Organizing Network and MLP Classifiers
abstract
This paper proposes a novel approach to automatic recognition of handprinted Bangla (an Indian script) numerals. A modified Topology Adaptive Self-Organizing Neural Network is proposed to extract a vector skeleton from a binary numeral image. Simple heuristics are considered to prune artifacts, if any, in such a skeletal shape. Certain topological and structural features like loops, junctions, positions of terminal nodes, etc. are used along with a hierarchical tree classifier to classify handwritten numerals into smaller subgroups. Multilayer perceptron (MLP) networks are then employed to uniquely classify the numerals belonging to each subgroup. The system is trained using a sample data set of 1800 numerals and we have obtained 93.26% correct recognition rate and 1.71% rejection on a separate test set of another 7760 samples. In addition, a validation set consisting of 1440 samples has been used to determine the termination of the training algorithm of the MLP networks. The proposed scheme is sufficiently robust with respect to considerable object noise.
Ujjwal Bhattacharya, Tanmoy Kanti Das, Amitava Datta, Swapan K. Parui, Bidyut B. Chaudhuri
Int. J. Pattern Recognit. Artif. Intell.1
2000 Shape Extraction of Volumetric Images of Filamentous Bacteria Using Topology Adaptive Self Organization
abstract
The study of the filamentous objects in waste water has gained momentum due to its significant effect in environmental pollution. The paper describes a neural network based skeleton extraction technique for volumetric images of these biofilm objects. These objects require huge computer storage space. One way to economize the storage space is to represent such images in the form of a vector skeleton (a piecewise linear approximation). Such a skeleton preserves the essential structure of the object. The proposed neural network does not start with a predefined net topology. The topology evolves during the learning process on the basis of the input. The present technique has certain advantages over the conventional 3-D thinning techniques. It achieves data reduction at a higher rate. Also, the proposed technique is highly robust to noise and arbitrary rotations of an image.
Ujjwal Bhattacharya, Amitava Datta, Swapan K. Parui, Bidyut B. Chaudhuri, Volkmar Liebscher, Karsten Rodenacker
ICPR1
2000 Efficient training and improved performance of multilayer perceptron in pattern classification
Bidyut B. Chaudhuri, Ujjwal Bhattacharya
Neurocomputing2
1997 An MLP-based texture segmentation method without selecting a feature set
Ujjwal Bhattacharya, Bidyut B. Chaudhuri, Swapan K. Parui
Image Vis. Comput.1
1996 An MLP-based texture segmentation technique which does not require a feature set
abstract
In this paper we describe a texture segmentation approach without feature computation based on a multilayer perceptron network (MLP). Thus, the users need not bother about the selection and then computation of feature set and hence real-time segmentation may be possible. The basic motivation of the work is the fact that human vision does not consciously compute features for distinguishing different textures in a scene. A single hidden layer MLP network has been found to be most suitable with heuristically chosen input and hidden layer sizes. A method has been used to speedup the learning of the MLP network. The result of segmentation by a trained network usually results in misclassification in the form of speckles. For the removal of such noise an edge-preserving-noise-smoothing technique is proposed. The final segmentation accuracy is well comparable with that of other existing techniques.
Ujjwal Bhattacharya, Bidyut B. Chaudhuri, Swapan K. Parui
ICPR1
1995 On the rate of convergence of perceptron learning
Ujjwal Bhattacharya, Swapan K. Parui
Pattern Recognit. Lett.1