EDBT 2026 Demo / reviewers in the wild / expert
Umapada Pal 0001
dblp:p/UmapadaPal · also U. Pal 0001
· DBLP profile ↗
95ranked-venue papers in the field
14as first author
16since 2021 · last 2026
0000-0002-5426-2618ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 94 (14 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVR: Diffusion-Based Multi-View Reasoning for Scene Text Detection
Debayan Das Gupta, Palaiahnakote Shivakumara, Palash Ghosal, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 4 |
| 2025 | Personality Trait Prediction from Twitter Data Using Text and Image Features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002 |
ICDAR (1) | 3 |
| 2025 | A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
Ritabrata Chakraborty, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 3 |
| 2025 | A New Fourier-Attention Guided Approach for Domain-Agnostic Text Localization
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
ICDAR (3) | 3 |
| 2025 | Doc2GraphFormer: Bridging Structured Graph Learning with Transformer Attention for Efficient Document Understanding
Souparni Mazumder, Sanket Biswas, Aniket Pal, Alloy Das, Umapada Pal 0001, Josep Lladós 0001 |
ICDAR (4) | 5 |
| 2024 | GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (3) | 4 |
| 2024 | A New Unsupervised Approach for Text Localization in Shaky and Non-shaky Scene Video
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Cheng-Lin Liu 0001 |
ICDAR (5) | 3 |
| 2023 | SwinDocSegmenter: An End-to-End Unified Domain Adaptive Transformer for Document Instance Segmentation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (1) | 4 |
| 2023 | Gaussian Kernels Based Network for Multiple License Plate Number Detection in Day-Night Images
Soumi Das, Palaiahnakote Shivakumara, Umapada Pal 0001, Ramachandra Raghavendra |
ICDAR (5) | 3 |
| 2023 | SelfDocSeg: A Self-supervised Vision-Based Approach Towards Document Segmentation
Subhajit Maity, Sanket Biswas, Siladittya Manna, Ayan Banerjee 0002, Josep Lladós 0001, Saumik Bhattacharya, Umapada Pal 0001 |
ICDAR (1) | 7 |
| 2023 | Scene Text Recognition with Image-Text Matching-Guided Dictionary
Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu 0001, Umapada Pal 0001 |
ICDAR (6) | 5 |
| 2023 | ICDAR 2023 Competition on Video Text Reading for Dense and Small Text
Weijia Wu 0001, Yuzhong Zhao, Zhuang Li 0002, Zheng Shou 0001, Umapada Pal 0001, Dimosthenis Karatzas, Xiang Bai |
ICDAR (2) | 6 |
| 2021 | ICDAR 2021 Competition on Script Identification in the Wild
Abhijit Das 0001, Miguel A. Ferrer, Aythami Morales, Moisés Díaz Cabrera, Umapada Pal 0001, Donato Impedovo, Wentao Yang 0003, Kensho Ota, Tadahito Yao, Le Quang Hung, Nguyen Quoc Cuong, Seungjae Kim, Abdeljalil Gattal |
ICDAR (4) | 5 |
| 2021 | DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (3) | 4 |
| 2021 | DCINN: Deformable Convolution and Inception Based Neural Network for Tattoo Text Detection Through Skin Region
Tamal Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ramachandra Raghavendra, Sukalpa Chanda |
ICDAR (2) | 3 |
| 2021 | Automatic Signature-Based Writer Identification in Mixed-Script Scenarios
Sk Md Obaidullah, Mridul Ghosh, Himadri Mukherjee, Kaushik Roy 0004, Umapada Pal 0001 |
ICDAR (2) | 5 |
| 2020 | A New Context-Based Method for Restoring Occluded Text in Natural Scene Images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Daniel P. Lopresti |
DAS | 3 |
| 2020 | A New Common Points Detection Method for Classification of 2D and 3D Texts in Video/Scene Images
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ahlad Kumar, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti |
DAS | 5 |
| 2019 | Age Estimation using Disconnectedness Features in HandwritingabstractReal-time applications of handwriting analysis have increased drastically in the fields of forensic and information security because of accurate cues. One of such applications is human age estimation based on handwriting for the purpose of immigrant checking. In this paper, we have proposed a new method for age estimation using handwriting analysis using Hu invariant moments and disconnectedness features. To make the proposed method robust to both ruled and un-ruled documents, we propose to explore intersection point detection in Canny edge images of each input document, which results in text components. For each text component pair, we propose Hu invariant moments for extracting disconnectedness features, which in fact measure multi-shape components based on distance, shape and mutual position analysis of components. Furthermore, iterative k-means clustering is proposed for the classification of different age groups. Experimental results on our dataset and some standard datasets, namely, IAM and KHATT, show that the proposed method is effective and outperforms the state-of-the-art methods. V. Basavaraja, Palaiahnakote Shivakumara, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
ICDAR | 4 |
| 2019 | Zero Shot Learning Based Script Identification in the WildabstractThe text recognition system for natural images or video frames containing multilingual text needs a method to first identify the written script and then recognize the word in the identified script. However, the occurrence of some scripts is rare as compared to others. Due to the availability of a few samples of the rare script, the supervised learning of the deep neural networks is difficult. To overcome this problem, we have proposed a zero-shot learning based method for script identification. We have also proposed architecture for script identification which fuses the global feature vector and the semantic embedding vector. The semantic embedding of the script is obtained by using the spatial dependency of the stroke's sequence via the recurrent neural network. The proposed architecture shows superior results as compared to the baseline approaches. Prateek Keserwani, Kanjar De, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICDAR | 4 |
| 2019 | CRNN Based Jersey-Bib Number/Text Recognition in Sports and Marathon ImagesabstractThe primary challenge in tracing the participants in sports and marathon video or images is to detect and localize the jersey/Bib number that may present in different regions of their outfit captured in cluttered environment conditions. In this work, we proposed a new framework based on detecting the human body parts such that both Jersey Bib number and text is localized reliably. To achieve this, the proposed method first detects and localize the human in a given image using Single Shot Multibox Detector (SSD). In the next step, different human body parts namely, Torso, Left Thigh, Right Thigh, that generally contain a Bib number or text region is automatically extracted. These detected individual parts are processed individually to detect the Jersey Bib number/text using a deep CNN network based on the 2-channel architecture based on the novel adaptive weighting loss function. Finally, the detected text is cropped out and fed to a CNN-RNN based deep model abbreviated as CRNN for recognizing jersey/Bib/text. Extensive experiments are carried out on the four different datasets including both bench-marking dataset and a new dataset. The performance of the proposed method is compared with the state-of-the-art methods on all four datasets that indicates the improved performance of the proposed method on all four datasets. Sauradip Nag, Ramachandra Raghavendra, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Mohan Kankanhalli |
ICDAR | 4 |
| 2019 | ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition - RRC-MLT-2019abstractWith the growing cosmopolitan culture of modern cities, the need of robust Multi-Lingual scene Text (MLT) detection and recognition systems has never been more immense. With the goal to systematically benchmark and push the state-of-the-art forward, the proposed competition builds on top of the RRC-MLT-2017 with an additional end-to-end task, an additional language in the real images dataset, a large scale multi-lingual synthetic dataset to assist the training, and a baseline End-to-End recognition method. The real dataset consists of 20,000 images containing text from 10 languages. The challenge has 4 tasks covering various aspects of multi-lingual scene text: (a) text detection, (b) cropped word script classification, (c) joint text detection and script classification and (d) end-to-end detection and recognition. In total, the competition received 60 submissions from the research and industrial communities. This paper presents the dataset, the tasks and the findings of the presented RRC-MLT-2019 challenge. Nibal Nayef, Cheng-Lin Liu 0001, Jean-Marc Ogier, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal 0001, Jean-Christophe Burie |
ICDAR | 10 |
| 2018 | Evaluation of Gist Operator for Document Image RetrievalabstractAs digitised documents normally contain a large variety of structures, a page segmentation- and layout-free method for document image retrieval is preferable. In this research work, therefore, wavelet transform as a transform-based approach is initially used to provide different under-sampled images from the original image. Then, Gist operator, as a feature extraction technique, is employed to extract a set of global features from the original image as well as the sub-images obtained from the wavelet transform. Moreover, the column-wise variances of the values in each sub-image are computed and they are then concatenated to obtain another set of features. Considering each feature set, locality-sensitive hashing is employed to compute similarity distances between a query and the document images in the database. Finally, a classifier fusion technique using the mean function is taken into account to provide a document image retrieval result. The combination of these features and a clustering score fusion strategy provides higher document image retrieval accuracy. Two different databases of the document image are considered for experimentation. The results obtained from the experimental study are detailed and the results are encouraging. Fahimeh Alaei, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
DAS | 3 |
| 2017 | Stroke-Order Normalization for Online Bangla Handwriting RecognitionabstractStroke order variation within characters is one of the difficult problems in online Bangla handwriting recognition. Moreover, in Bangla, character parts are written in zone-wise manner. Character parts written in the middle zone are generally cursive, while character parts in the upper and lower zones are written using delayed strokes. As online recognition depends on the order of writing, words written with different stroke-order are treated as different words to the online recognizer. To the best of our knowledge, no work has been reported on stroke-order normalization for any Indic script though it is an important aspect of online recognition. In this paper, we propose a stroke-order normalization method for Bangla online recognition using offline and online information. Here, at first, based on the offline information, sub-strokes in a word are ordered according to their relative positions. This results in similar stroke-order among the different instances of the same word. Next, online information of each ordered sub-stroke is used for feature extraction. This normalization approach has several significant advantages, e.g. (i) characters/words having any stroke order can be recognized, (ii) number of word classes is reduced, etc. We have tested our method on a dataset of 6000 words and obtained 74.65% and 90.53% word recognition accuracies, respectively, before and after stroke-order normalization. Thus, stroke-order normalization has enhanced the recognition result drastically (15.88%). Nilanjana Bhattacharya 0001, Umapada Pal 0001, Partha Pratim Roy 0001 |
ICDAR | 2 |
| 2017 | ICDAR2017 Robust Reading Challenge on Multi-Lingual Scene Text Detection and Script Identification - RRC-MLTabstractText detection and recognition in a natural environment are key components of many applications, ranging from business card digitization to shop indexation in a street. This competition aims at assessing the ability of state-of-the-art methods to detect Multi-Lingual Text (MLT) in scene images, such as in contents gathered from the Internet media and in modern cities where multiple cultures live and communicate together. This competition is an extension of the Robust Reading Competition (RRC) which has been held since 2003 both in ICDAR and in an online context. The proposed competition is presented as a new challenge of the RRC. The dataset built for this challenge largely extends the previous RRC editions in many aspects: the multi-lingual text, the size of the dataset, the multi-oriented text, the wide variety of scenes. The dataset is comprised of 18,000 images which contain text belonging to 9 languages. The challenge is comprised of three tasks related to text detection and script classification. We have received a total of 16 participations from the research and industrial communities. This paper presents the dataset, the tasks and the findings of this RRC-MLT challenge. Nibal Nayef, Imen Bizid, Hyunsoo Choi, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal 0001, Christophe Rigaud, Joseph Chazalon, Wafa Khlif, Muhammad Muzzamil Luqman, Jean-Christophe Burie, Cheng-Lin Liu 0001, Jean-Marc Ogier |
ICDAR | 8 |
| 2017 | New Fuzzy-Mass Based Features for Video Image Type CategorizationabstractDue to the large variety of video type collections, it becomes difficult to achieve good text detection and recognition accuracy. We propose a new fuzzy-mass based method for classifying (categorizing) text frames from different types of video. For each frame of a video type, we formulate Fuzzy logic to identify straight and curved edge components from edge images. We then estimate mass locally and globally by drawing consecutive ellipses over edge images with respect to straight and curved edge components. Further, we extract features based on spatial proximity between centroid of classified straight/curved edge components and that of the whole image. This results local features. Next, the features are extracted for the whole image without ellipse drawing, which results in global features. The combination of both local and global features is then fed to an SVM classifier for video type classification. Experimental results on the proposed and existing classification methods show that the proposed classification outperforms the stat of art methods. Furthermore, experiments on before and after classification with several text detection and binarization methods show that the proposed classification is significant in improving text detection and recognition performance. Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Umapada Pal 0001, Tong Lu 0002 |
ICDAR | 5 |
| 2017 | Temporal Integration for Word-Wise Caption and Scene Text IdentificationabstractGenerally video consists of edited text (i.e., caption text) and natural text (i.e., scene text), and these two texts differ from one another in nature as well as characteristics. Such different behaviors of caption and scene texts lead to poor accuracy for text recognition in video. In this paper, we explore wavelet decomposition and temporal coherency for the classification of caption and scene text. We propose wavelet of high frequency sub-bands to separate text candidates that are represented by high frequency coefficients in an input word. The proposed method studies the distribution of text candidates over word images based on the fact that the standard deviation of text candidates is high at the first zone, low at the middle zone and high at the third zone. This is extracted by mapping standard deviation values to 8 equal sized bins formed based on the range of standard deviation values. The correlation among bins at the first and second levels of wavelets is explored to differentiate caption and scene text and for determining the number of temporal frames to be analyzed. The properties of caption and scene texts are validated with the chosen temporal frames to find the stable property for classification. Experimental results on three standard datasets (ICDAR 2015, YVT and License Plate Video) show that the proposed method outperforms the existing methods in terms of classification rate and improves recognition rate significantly based on classification results. Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ainuddin Wahid Abdul Wahab |
ICDAR | 3 |
| 2017 | Fourier-Residual for Printer IdentificationabstractPrinter identification is challenging due to advanced software technologies in the field of forgery detection. This paper presents a new idea of using the Fourier transform residual for the identification of documents printed by different printers. The proposed approach first convolves a Laplacian mask with a Fourier transform in the frequency domain to smoothen the edges. Next, we apply an inverse Fourier transform to reconstruct images from smoothed information (RFL). Similarly, the proposed approach reconstructs images using gray information of the input image (RFG). Then the residual is calculated by subtracting RFG from RFL. The set of statistical features, texture and spatial features are extracted from residual images for printer identification. Experimental results with the existing method on our dataset and a standard dataset show that the proposed approach outperforms the existing approach on both the datasets in terms of classification rate, recall, precision and F-measure. Palaiahnakote Shivakumara, Tong Lu 0002, M. Basavanna, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 5 |
| 2016 | Performance of an Off-Line Signature Verification Method Based on Texture Features on a Large Indic-Script Signature DatasetabstractIn this paper, a signature verification method based on texture features involving off-line signatures written in two different Indian scripts is proposed. Both Local Binary Patterns (LBP) and Uniform Local Binary Patterns (ULBP), as powerful texture feature extraction techniques, are used for characterizing off-line signatures. The Nearest Neighbour (NN) technique is considered as the similarity metric for signature verification in the proposed method. To evaluate the proposed verification approach, a large Bangla and Hindi off-line signature dataset (BHSig260) comprising 6240 (260×24) genuine signatures and 7800 (260×30) skilled forgeries was introduced and further used for experimentation. We further used the GPDS-100 signature dataset for a comparison. The experiments were conducted, and the verification accuracies were separately computed for the LBP and ULBP texture features. There were no remarkable changes in the results obtained applying the LBP and ULBP features for verification when the BHSig260 and GPDS-100 signature datasets were used for experimentation. Srikanta Pal, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
DAS | 3 |
| 2016 | New Sharpness Features for Image Type Classification Based on Textual InformationabstractAchieving good recognition results from a single method for text lines in video/natural scene images captured by high resolution cameras or low resolution mobile cameras, and images in web pages, is often hard. In this paper, we propose new sharpness based features of textual portion of each input text line image using HSI color space for the classification of an input image into one of the four classes (video, scene, mobile or born digital). This helps in choosing an appropriate method based on the class type of the input text for its improved recognition rate. For a given input text line image, the proposed method obtains H, S and I images. Then Canny edge images are obtained for H, S and I spaces, which results in text candidates. We perform sliding window operation over the text candidate image of each text line of each color space to estimate new sharpness by calculating stroke width and gradient information. The sharpness values of the text lines of the three color spaces are then fed to k-means clustering with maximum, minimum and average guesses, which results in three respective clusters. The mean of each cluster for respective color spaces outputs a feature vector having nine feature values for image classification with the help of an SVM classifier. Experimental results on standard datasets, namely, ICDAR 2013, ICDAR 2015 video, ICDAR 2015 natural scene data, ICDAR 2013 born digital data and the images captured by a mobile camera (our own data) show that the proposed classification method helps in improving recognition results. Raghunandan K. Srinivas, Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002 |
DAS | 4 |
| 2015 | A comparative study of features for handwritten Bangla text recognitionabstractRecognition of Bangla handwritten text is difficult due to its complex nature of having modifiers and headlines features. This paper presents a comparative study of different features namely LGH (Local Gradient of Histogram), PHOG (Pyramid Histogram of Oriented Gradient), GABOR, G-PHOG (Combined GABOR and PHOG) and profile feature by Marti-Bunke when applied in middle zone recognition of Bangla words using Hidden Markov Model (HMM) based framework. For this purpose, a zone segmentation method is applied to extract the busy (middle) zones of handwritten words and features are extracted from the middle zone. The system has been tested on a sufficiently large and variation-rich dataset consisting of 11,253 training and 3,856 testing data. From the experiment, it has been noted that PHOG feature outperforms other features in middle zone recognition. Since PHOG feature outperform others, we use this feature for full word recognition, For this purpose initially upper and lower zone components are recognized by PHOG features and SVM classifier. Finally, the zone-wise results are combined by the context information of the corresponding components in each zone to obtain the word level recognition. Ayan Kumar Bhunia, Ayan Das 0001, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICDAR | 4 |
| 2015 | A new method based on bag of filters for character recognition in scene images by learningabstractAchieving a good recognition rate for scene characters is a big challenge due to non-uniform illumination effects, perspective distortions, multiple colors or contrasts, different fonts and their various sizes, background or orientation variations, etc. Unlike the existing recognition methods that use binary information or the features extracted from different domains, the proposed method explores gray information in the form of a filter bank to extract the discriminative power for all the 62 scene character classes. We propose a sliding window (patch) operation over a character image for learning the global features, which represent the structures of character images of all the classes by reconstructing a filter bank from the original data. We introduce shareable constrains to activate class-specific filters from the filter bank. Further, we propose constraints by studying the nearest neighbor patches and exemplar selection to maximize the gap between inter-classes and minimize the gap between intra-classes. The method is evaluated and compared with several existing recognition methods in terms of character recognition rate. Experimental results show that the proposed method outperforms the existing methods. Qisu Li, Tong Lu 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Chew Lim Tan |
ICDAR | 4 |
| 2015 | ICDAR2015 competition on signature verification and writer identification for on- and off-line skilled forgeries (SigWIcomp2015)abstractThis paper presents the results of the ICDAR 2015 competition on signature verification and writer identification for on- and off-line skilled forgeries jointly organized by PR-researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. Two modalities (signatures and handwritten text) are considered and training and evaluation data are collected and provided by FHEs and PR-researchers. Four tasks are defined for four different languages; Bengali off-line signature verification, Italian off-line signature verification, German on-line signature verification, and English handwritten text based writer identification. In total, 40 systems have participated in this competition. The participants of the signatures modality were motivated to report their results in Likelihood Ratios (LRs). This has made the systems even more interesting for application in forensic casework. For evaluating the performance of the systems, we have used the forensically substantial Cost of Log Likelihood Ratios (Ĉllr) in the case of signatures, and the F-measure in the case of handwritten text. Muhammad Imran Malik, Sheraz Ahmed, Angelo Marcelli, Umapada Pal 0001, Michael Blumenstein, Linda Alewijnse, Marcus Liwicki |
ICDAR | 4 |
| 2015 | Date field extraction from handwritten documents using HMMsabstractAutomatic document interpretation and retrieval is an important task to access handwritten digitized document repositories. In documents, the date is an important field and it has various applications such as date-wise document indexing/retrieval. In this paper a framework has been proposed for automatic date field extraction from handwritten documents. In order to design the system, sliding window-wise Local Gradient Histogram (LGH)-based features and a character-level Hidden Markov Model (HMM)-based approach have been applied for segmentation and recognition. Individual date components such as month-word (month written in word form i.e. January, Jan, etc.), numeral, punctuation and contraction categories are segmented and labelled from a text line. Next, a Histogram of Gradient (HoG)-based features and a Support Vector Machine (SVM)- based classifier have been used to improve the results obtained from the HMM-based recognition system. Subsequently, both numeric and semi-numeric regular expressions of date patterns have been considered for undertaking date pattern extraction in labelled components. The experiments are performed on an English document dataset and the encouraging results obtained from the approach indicate the effectiveness of the proposed system. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 3 |
| 2015 | Performance evaluation of DTW and its variants for word spotting in degraded documentsabstractIn word spotting literature, classical DTW has been widely employed. However there exists several other improved versions of DTW along with other robust sequence matching techniques. Very few of them have been studied in the context of word spotting and this scarcity of research work is the motivation of the paper. This paper presents a comparative study of classical Dynamic Time Warping (DTW) technique and many of its improved modifications, as well as other sequence matching techniques in the context of word spotting. An experimental study on historical documents is performed to evaluate the behavior of DTW's variants and other sequence matching techniques. A detailed comparative analysis along with wide range of experimentation is performed, which shows that classical DTW remains a good choice when there are no segmentation problems and when features are very local. In case of word segmentation errors, Continuous Dynamic Programming (CDP) seems to be a better choice. This research work has introduced several other improved sequence matching algorithms in the context of word spotting, which show interesting and improved results. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICDAR | 4 |
| 2015 | Exemplary Sequence Cardinality: An effective application for word spottingabstractIn this paper, a new sequence matching algorithm called as Exemplary Sequence Cardinality (ESC) is proposed. ESC combines several abilities of other sequence matching algorithms e.g. DTW, SSDTW, CDP, FSM, MVM, OSB1. Depending on the application domain, ESC can be tuned to behave such as these different sequence matching algorithms. Its generality and robustness comes from its ability to find subsequences (as in CDP and SSDTW), to skip outliers inside the target sequences (as in MVM and FSM) and also in the query sequence (as in OSB ) and it has the ability to have many to one and one to many correspondences (as in DTW) between the elements of the query and the target sequences. It's special characteristic of skipping noisy elements from query sequence along with other afore mentioned properties gives it an edge over FSM. In case of word spotting application, the outliers skipping capability of ESC makes it less sensible to local variations in the spelling of words, and also to noise present in the query and/or in the target word images. Due to it's capability of sub-sequence matching, the ESC algorithm has the ability to retrieve a query inside a line or piece of line. Finally, its multiple matching facilities (many to one and one to many matching) is proven to be well advantageous in case of different length of target and query sequences due to the variability in scale, font, type/size factors. By experimenting on printed historical document images, we have demonstrated the interest of proposed ESC algorithm in specific cases when incorrect word segmentation and word level local variations occur regularly. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICDAR | 4 |
| 2015 | ICDAR2015 Competition on Video Script Identification (CVSI 2015)abstractThis paper presents the final results of the ICDAR 2015 Competition on Video Script Identification. A description and performance of the participating systems in the competition are reported. The general objective of the competition is to evaluate and benchmark the available methods on word-wise video script identification. It also provides a platform for researchers around the globe to particularly address the video script identification problem and video text recognition in general. The competition was organised around four different tasks involving various combinations of scripts comprising tri-script and multi-script scenarios. The dataset used in the competition comprised ten different scripts. In total, six systems were received from five participants over the tasks offered. This report details the competition dataset specifications, evaluation criteria, summary of the participating systems and their performance across different tasks. The systems submitted by Google Inc. were the winner of the competition for all the tasks, whereas the systems received from Huazhong University of Science and Technology (HUST) and Computer Vision Center (CVC) were very close competitors. Nabin Sharma, Ranju Mandal, Rabi Sharma, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 4 |
| 2015 | Multi-lingual text recognition from video framesabstractText recognition from video frames is a challenging task due to low resolution, blur, complex and coloured backgrounds, noise, to mention a few. Consequently, the traditional ways of text recognition from scanned documents having simple backgrounds fails when applied to video text. Although there are various techniques available for text recognition from handwritten and printed documents with simple backgrounds, text recognition from video frames has not been comprehensively investigated, especially for multi-lingual videos. In this paper, we present a technique for multi-lingual video text recognition which involves script identification in the first stage, followed by word and character recognition, and finally the results are refined using a post-processing technique. Considering the inherent problems in videos, a Spatial Pyramid Matching (SPM) based technique, using patch-based SIFT descriptors and SVM classifier, is employed for script identification. In the next stage, a Hidden Markov Model (HMM) based approach is used for word and character recognition, which utilizes the context information. Finally, a lexicon-based post-processing technique is applied to verify and refine the word recognition results. The proposed method was tested on a dataset comprising of 4800 words from three different scripts, namely, Roman (English), Hindi and Bengali. The script identification results obtained are encouraging. The word and character recognition results are also encouraging considering the complexity and problems associated with video text processing. Nabin Sharma, Ranju Mandal, Rabi Sharma, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 5 |
| 2015 | A complete automatic short answer assessment system with student identificationabstractThere are only a few studies undertaken in developing automatic assessment systems using handwriting recognition, even though a successful system would undoubtedly benefit the education system as schools and universities in many countries still employ paper-based examinations. To the best of the authors' knowledge, there is no existing work on an automatic off-line short answer assessment system comprising a student identification component. Hence in this paper, the authors propose a system towards this, where a new feature extraction technique called the Enhanced Water Reservoir, Loop and Gaussian Grid Feature, as well as other enhanced feature extraction techniques were utilised. Artificial Neural Networks and Support Vector Machines were employed as the classifiers; they were used for the investigation, and a comparison of the recognition and accuracy rates of the proposed systems, as well as the feature extraction techniques, was undertaken. The proposed assessment system achieved a recognition rate of 87.12% with 91.12% assessment accuracy, and the student identification component obtained a recognition rate of 99.52% with a 100% identification accuracy rate. Hemmaphan Suwanwiwat, Michael Blumenstein, Umapada Pal 0001 |
ICDAR | 3 |
| 2015 | Multi-lingual date field extraction for automatic document retrieval by machine
Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
Inf. Sci. | 3 |
| 2014 | Design of Unsupervised Feature Extraction System for On-line Bangla Handwriting RecognitionabstractDifferent systems for handwriting recognition use different features to represent the input text. Even after decades of research, no favorable decision on a best-practice exists and many features are carefully hand-crafted. To facilitate the design phase for on-line handwriting systems, in this paper, we propose an unsupervised feature generation approach based on dissimilarity space embedding (DSE) of local neighborhoods around the points along the trajectory. DSE has high capability of discriminative representation and hence beneficial for classification. We compare the approach with a state-of-the-art feature extraction method and demonstrate its superiority. Volkmar Frinken, Nilanjana Bhattacharya 0001, Umapada Pal 0001 |
Document Analysis Systems | 3 |
| 2014 | Multi-oriented Text Recognition in Graphical Documents Using HMMabstractThe text lines in graphical documents (e.g., maps, engineering drawings), artistic documents etc., are often annotated in curve lines to illustrate different locations or symbols. For the optical character recognition of such documents, individual text lines from the documents need to be extracted and recognized. Due to presence of multi-oriented characters in such non-structured layout, word recognition is a challenging task. In this paper, we present an approach towards the recognition of scale and orientation invariant text words in graphical documents using Hidden Markov Models (HMM). First, a line extraction method is applied to segment text lines and the method is based on the foreground and background information of the text components. To effectively utilize the background information, a water reservoir concept is used here. For recognition of curved text lines, a path of sliding window is estimated and features extracted from the sliding window are fed to the HMM system for recognition. Local gradient histogram (LGH) based frame-wise feature is used in HMM. The experimental results are evaluated on a dataset of graphical words and we have obtained encouraging results. Partha Pratim Roy 0001, Sangheeta Roy, Umapada Pal 0001 |
Document Analysis Systems | 3 |
| 2013 | Near Convex Region Adjacency Graph and Approximate Neighborhood String Matching for Symbol Spotting in Graphical DocumentsabstractThis paper deals with a sub graph matching problem in Region Adjacency Graph (RAG) applied to symbol spotting in graphical documents. RAG is a very important, efficient and natural way of representing graphical information with a graph but this is limited to cases where the information is well defined with perfectly delineated regions. What if the information we are interested in is not confined within well defined regions? This paper addresses this particular problem and solves it by defining near convex grouping of oriented line segments which results in near convex regions. Pure convexity imposes hard constraints and can not handle all the cases efficiently. Hence to solve this problem we have defined a new type of convexity of regions, which allows convex regions to have concavity to some extend. We call this kind of regions Near Convex Regions (NCRs). These NCRs are then used to create the Near Convex Region Adjacency Graph (NCRAG) and with this representation we have formulated the problem of symbol spotting in graphical documents as a sub graph matching problem. For sub graph matching we have used the Approximate Edit Distance Algorithm (AEDA) on the neighborhood string, which starts working after finding a key node in the input or target graph and iteratively identifies similar nodes of the query graph in the neighborhood of the key node. The experiments are performed on artificial, real and distorted datasets. Anjan Dutta 0001, Josep Lladós 0001, Horst Bunke, Umapada Pal 0001 |
ICDAR | 4 |
| 2013 | A System for Bangla Online Handwritten TextabstractRecognition of Bangla compound characters has rarely got attention from researchers. This paper deals with segmentation and recognition of online handwritten Bangla cursive text containing basic and compound characters and all types of modifiers. Here, at first, we segment cursive words into primitives. Next primitives are recognized. A primitive may represent a character/compound character or a part of a character/compound character having meaningful structural information or a part incurred while joining two characters. We manually analyzed all the input texts written by different groups of people to create a ground truth set of distinct classes of primitives for result verification and we obtained 251 valid primitive classes. For automatic segmentation of text into primitives, we discovered some rules analyzing different joining patterns of Bangla characters. Applying these rules and using combination of online and offline information the segmentation technique was proposed. We achieved correct primitive segmentation rate of 97.89% from the 4984 online words. Directional features were used in SVM for recognition and we achieved average primitive recognition rate of 97.45%. Nilanjana Bhattacharya 0001, Umapada Pal 0001, Fumitaka Kimura |
ICDAR | 2 |
| 2013 | LBP Based Line-Wise Script IdentificationabstractScript identification is an important step in multi-script document analysis. As different textures present in text portion of a script are the main distinct features of the script, in this paper, we proposed a new algorithm for printed script identification based on texture analysis. Since local patterns is a unifying concept for traditional statistical and structural approaches of texture analysis, here the basic idea is to use the histogram of the local patterns as description of the script stroke directions distribution which is the characteristic of every script. As local pattern, the basic version of the Local Binary Patterns (LBP) and a modified version of the Orientation of the Local Binary Patterns (OLBP) are proposed. A Least Square Support Vector Machine (LS-SVM) is used as identifier. The scheme has been verified on two databases. The first or training database is a database with 200 sheets of 10 different scripts. The scripts font is provided by the Google translator. The second or test database has been obtained by scanning different newspapers and books. It contains 5 common scripts among 10 different scripts of the first database. From the experiment we obtained encouraging results. Miguel A. Ferrer, Aythami Morales, Umapada Pal 0001 |
ICDAR | 3 |
| 2013 | Handwritten Musical Document Retrieval Using Music-Score SpottingabstractIn this paper, we present a novel approach for retrieval of handwritten musical documents using a query sequence/word of musical scores. In our algorithm, the musical score-words are described as sequences of symbols generated from a universal codebook vocabulary of musical scores. Staff lines are removed first from musical documents using structural analysis of staff lines and symbol codebook vocabulary is created in offline. Next, using this symbol codebook the music symbol information in each document image is encoded. Given a query sequence of musical symbols in a musical score-line, the symbols in the query are searched in each of these encoded documents. Finally, a sub-string matching algorithm is applied to find query words. For codebook, two different feature extraction methods namely: Zernike Moments and 400 dimensional gradient features are tested and two unsupervised classifiers using SOM and K-Mean are evaluated. The results are compared with a baseline approach of DTW. The performance is measured on a collection of handwritten musical documents and results are promising. Rakesh Malik, Partha Pratim Roy 0001, Umapada Pal 0001, Fumitaka Kimura |
ICDAR | 3 |
| 2013 | A Fast Word Retrieval Technique Based on Kernelized Locality Sensitive HashingabstractIn this paper, we have presented a new and faster word retrieval approach, which is able to deal with heterogeneous document image collections. A certain amount of image features (statistical and Gabor Wavelet) are extracted, which inherently represent word's images. These features are used for generating hash table for fast retrieval of similar image from a very large image dataset. The decomposition and embedding of high-dimensional features and complex distance functions into a low-dimensional Hamming space helps to efficiently search items. However, existing methods do not apply for high-dimensional kernelized data when the underlying features' embedding for the kernel is unknown. The generalization of locality sensitive hashing (LSH) for arbitrary kernel is presented in the paper. The proposed algorithm provides sub-linear time similarity search and works for a wide class of similarity functions. Tanmoy Mondal, Nicolas Ragot, Jean-Yves Ramel, Umapada Pal 0001 |
ICDAR | 4 |
| 2013 | Word-Wise Script Identification from Video FramesabstractScript identification is an essential step for the efficient use of the appropriate OCR in multilingual document images. There are various techniques available for script identification from printed and handwritten document images, but script identification from video frames has not been explored much. This paper presents a study of some pre-processing techniques and features for word-wise script identification from video frames. Traditional features, namely Zernike moments, Gabor and gradient, have performed well for handwritten and printed documents having simple backgrounds and adequate resolution for OCR. Video frames are mostly coloured and suffer from low resolution, blur, background noise, to mention a few. In this paper, an attempt has been made to explore whether the traditional script identification techniques can be useful in video frames. Three feature extraction techniques, namely Zernike moments, Gabor and gradient features, and SVM classifiers were considered for analyzing three popular scripts, namely English, Bengali and Hindi. Some pre-processing techniques such as super resolution and skeletonization of the original word images were used in order to overcome the inherent problems with video. Experiments show that the super resolution technique with gradient features has performed well, and an accuracy of 87.5% was achieved when testing on 896 words from three different scripts. The study also reveals that the use of proper pre-processing approaches can be helpful in applying traditional script identification techniques to video frames. Nabin Sharma, Sukalpa Chanda, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 3 |
| 2013 | A New Method for Character Segmentation from Multi-oriented Video WordsabstractThis paper presents a two-stage method for multi-oriented video character segmentation. Words segmented from video text lines are considered for character segmentation in the present work. Words can contain isolated or non-touching characters, as well as touching characters. Therefore, the character segmentation problem can be viewed as a two stage problem. In the first stage, text cluster is identified and isolated (non-touching) characters are segmented. The orientation of each word is computed and the segmentation paths are found in the direction perpendicular to the orientation. Candidate segmentation points computed using the top distance profile are used to find the segmentation path between the characters considering the background cluster. In the second stage, the segmentation results are verified and a check is performed to ascertain whether the word component contains touching characters or not. The average width of the components is used to find the touching character components. For segmentation of the touching characters, segmentation points are then found using average stroke width information, along with the top and bottom distance profiles. The proposed method was tested on a large dataset and was evaluated in terms of precision, recall and f-measure. A comparative study with existing methods reveals the superiority of the proposed method. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
ICDAR | 3 |
| 2013 | ICDAR 2013 Handwriting Segmentation ContestabstractThis paper presents the results of the Handwriting Segmentation Contest that was organized in the context of the ICDAR2013. The general objective of the contest was to use well established evaluation practices and procedures to record recent advances in off-line handwriting segmentation. Two benchmarking datasets, one for text line and one for word segmentation, were created in order to test and compare all submitted algorithms as well as some state-of-the-art methods for handwritten document image segmentation in realistic circumstances. Handwritten document images were produced by many writers in two Latin based languages (English and Greek) and in one Indian language (Bangla, the second most popular language in India). These images were manually annotated in order to produce the ground truth which corresponds to the correct text line and word segmentation results. The datasets of previously organized contests (ICDAR2007, ICDAR2009 and ICFHR2010 Handwriting Segmentation Contests) along with a dataset of Bangla document images were used as training dataset. Eleven methods are submitted in this competition. A brief description of the submitted algorithms, the evaluation criteria and the segmentation results obtained from the submitted methods are also provided in this manuscript. Nikolaos Stamatopoulos, Basilios Gatos, Georgios Louloudis, Umapada Pal 0001, Alireza Alaei |
ICDAR | 4 |
| 2013 | A Two-Stage Approach for Word Spotting in Graphical DocumentsabstractPresence of multi-oriented characters, connected characters with graphical lines, intersection of text and symbols with graphical lines/curves etc. are very common in graphical documents. As a result word spotting in graphical documents is still a challenging task that we try to solve (partially) in this paper. The proposed approach proceeds in two stages. In the first stage, recognition of isolated components is done using rotation invariant features and an SVM classifier. The characters having good recognition score and match in the query string are first selected for initial spotting. Because of structural complexity of graphical documents as well as of touching components, we may miss some of the query characters during initial spotting in some documents. In that case, based on the position, size and orientation of the recognized characters in the input document image, regions where missing characters may be located (candidate regions) are defined. In the second stage, Scale Invariant Feature Transform (SIFT) is used to find those missing characters in the candidate regions for possible spotting. Finally, using the position, size, orientation as well as intercharacter gap information of the recognized components, spotting is validated. Experimental results demonstrate that the method is efficient to locate a query word in multi-oriented and/or touching graphical documents. Arundhati Tarafdar, Umapada Pal 0001, Partha Pratim Roy 0001, Nicolas Ragot, Jean-Yves Ramel |
ICDAR | 2 |
| 2013 | Tamil Handwritten City Name Database Development and Recognition for Postal AutomationabstractAlthough there are some reports on offline Tamil isolated handwritten character recognition, to our knowledge there is only two reports on Tamil off-line handwritten word recognition. Also no city name dataset is available for Tamil script. In this paper we present a Tamil offline city name dataset, we developed, and propose a scheme for recognition. Because of the different writing style of various individuals, some of the characters in a Tamil city name may touch and accurate segmentation of such touching into individual characters is a difficult task. Avoiding proper segmentation here, we consider a city name string as a word and the recognition problem is treated as lexicon driven word recognition. In the proposed method, binarized city names are pre-segmented into primitives (individual character or its parts). Primitive components of each city name are then merged into possible characters to get the best city name using dynamic programming. For merging, total likelihood of characters is used as the objective function and character likelihood is computed based on Modified Quadratic Discriminant Function (MQDF), where direction features are applied. A dataset of 265 Tamil city names is developed. and the database will be available freely to the researchers. From the experiment of the proposed scheme 96.89% city name accuracy is obtained from this dataset. S. Thadchanamoorthy, Nihal D. Kodikara, H. L. Premaretne, Umapada Pal 0001, Fumitaka Kimura |
ICDAR | 4 |
| 2012 | Recognition of Similar Shaped Handwritten Characters Using Logistic RegressionabstractRecognition of similar shaped characters is a difficult problem and in character recognition systems most of the errors occur in similar shaped characters. In this article we propose a generic method to differentiate between two similar shaped characters, which works well not only when the characters are rotated about its center, but also in the presence of noise. Rotation is taken care of by contour distance based approach and recognition is done based on logistic regression. We consider a training data set to estimate the parameters of the logistic model, and using these parameters we classify the test object. We have considered pairs of similar shape characters of Bengali script for testing our algorithm. Kinjal Basu 0001, Radhika Nangia, Umapada Pal 0001 |
Document Analysis Systems | 3 |
| 2012 | Text Independent Writer Identification for Oriya ScriptabstractAutomatic identification of an individual based on his/her handwriting characteristics is an important forensic tool. In a computational forensic scenario, presence of huge amount of text/information in a questioned document cannot be ensured. Lack of data threatens system reliability in such cases. We here propose a writer identification system for Oriya script which is capable of performing reasonably well even with small amount of text. Experiments with curvature feature are reported here, using Support Vector Machine (SVM) as classifier. We got promising results of 94.00% writer identification accuracy at first top choice and 99% when considering first three top choices. Sukalpa Chanda, Katrin Franke, Umapada Pal 0001 |
Document Analysis Systems | 3 |
| 2012 | Off-Line Bangla Signature VerificationabstractIn the field of information security, biometric systems play an important role. Within biometrics, automatic signature identification and verification has been a strong research area because of the social and legal acceptance and extensive use of the written signature as an individual authentication. Signature verification is a process in which the questioned signature is examined in detail in order to determine whether it belongs to the claimed person or not. Despite substantial research in the field of signature verification involving Western signatures, very few works have been dedicated to non-Western signatures such as Chinese, Japanese, Arabic, or Persian etc. In this paper, the performance of an off-line signature verification system involving Bangla signatures, whose style is distinct from Western scripts, was investigated. The Gaussian Grid feature extraction technique was employed for feature extraction and Support Vector Machines (SVMs) were considered for classification. The Bangla signature database employed in the experiments consisted of 3000 forgeries and 2400 genuine signatures. An encouraging accuracy of 90.4% was obtained from the experiments. Srikanta Pal, Vu Nguyen 0002, Michael Blumenstein, Umapada Pal 0001 |
Document Analysis Systems | 4 |
| 2012 | Recent Advances in Video Based Document Processing: A ReviewabstractExtraction and recognition of text present in video has become a very popular research area in the last decade. Generally, text present in video frames is of different size, orientation, style, etc. with complex backgrounds, noise, low resolution and contrast. These factors make the automatic text extraction and recognition in video frames a challenging task. A large number of techniques have been proposed by various researchers in the recent past to address the problem. This paper presents a review of various state-of-the-art techniques proposed towards different stages (e.g. detection, localization, extraction, etc.) of text information processing in video frames. Looking at the growing popularity and the recent developments in the processing of text in video frames, this review imparts details of current trends and potential directions for further research activities to assist researchers. Nabin Sharma, Umapada Pal 0001, Michael Blumenstein |
Document Analysis Systems | 2 |
| 2012 | A New Method for Arbitrarily-Oriented Text Detection in VideoabstractText detection in video frames plays a vital role in enhancing the performance of information extraction systems because the text in video frames helps in indexing and retrieving video efficiently and accurately. This paper presents a new method for arbitrarily-oriented text detection in video, based on dominant text pixel selection, text representatives and region growing. The method uses gradient pixel direction and magnitude corresponding to Sobel edge pixels of the input frame to obtain dominant text pixels. Edge components in the Sobel edge map corresponding to dominant text pixels are then extracted and we call them text representatives. We eliminate broken segments of each text representatives to get candidate text representatives. Then the perimeter of candidate text representatives grows along the text direction in the Sobel edge map to group the neighboring text components which we call word patches. The word patches are used for finding the direction of text lines and then the word patches are expanded in the same direction in the Sobel edge map to group the neighboring word patches and to restore missing text information. This results in extraction of arbitrarily-oriented text from the video frame. To evaluate the method, we considered arbitrarily-oriented data, non-horizontal data, horizontal data, Hua's data and ICDAR-2003 competition data (Camera images). The experimental results show that the proposed method outperforms the existing method in terms of recall and f-measure. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
Document Analysis Systems | 3 |
| 2012 | An Effective Staff Detection and Removal Technique for Musical DocumentsabstractMusical staff line detection and removal techniques detect the staff positions in musical documents and segment musical score from musical documents by removing those staff lines. It is an important preprocessing step for ensuing the Optical Music Recognition tasks. This paper proposes an effective staff line detection and removal method that makes use of the global information of the musical document and models the staff line shape. It first estimates the staff height and space, and then models the shape of the staff line by examining the orientation of the staff pixels. At last the estimated model is used to find out the location of staff lines and hence to remove those detected staff lines. The proposed technique is simple, robust, and involves few parameters. It has been tested on the dataset of the recent staff removal competition held under the International Conference of Document Analysis and Recognition(ICDAR) 2011. Experimental results show the effectiveness and robustness of our proposed technique on musical documents with various types of deformations. Bolan Su, Shijian Lu, Umapada Pal 0001, Chew Lim Tan |
Document Analysis Systems | 3 |
| 2011 | A Benchmark Kannada Handwritten Document Dataset and Its SegmentationabstractResearch towards Indian handwritten document analysis achieved increasing attention in recent years. In pattern recognition and especially in handwritten document recognition, standard databases play vital roles for evaluating performances of algorithms and comparing results obtained by different groups of researchers. For Indian languages, there is a lack of standard database of handwritten texts to evaluate performance of different document recognition approaches and for comparison purpose. In this paper, an unconstrained Kannada handwritten text database (KHTD) is introduced. The KHTD contains 204 handwritten documents of four different categories written by 51 native speakers of Kannada. Total number of text-lines and words in the dataset are 4298 and 26115, respectively. In most of text-pages of the KHTD contains either an overlapping or a touching text-lines and the average number of text-lines in each document on the database is 21. Two types of ground truths based on pixels information and content information are generated for the database. Providing these two types of ground truths for the KHTD, it can be utilized in many areas of document image processing such as sentence recognition/understanding, text-line segmentation, word segmentation, word recognition, and character segmentation. To provide a framework for other researches, recent text-line segmentation results on this dataset are also reported. The KHTD is available for research purposes. Alireza Alaei, P. Nagabhushan, Umapada Pal 0001 |
ICDAR | 3 |
| 2011 | A New Text-Line Alignment Approach Based on Piece-Wise Painting Algorithm for Handwritten DocumentsabstractBecause of writing styles of different individuals, some of the text-lines may be curved in shape. For recognition of such text-lines, their proper alignment is necessary. In this paper, we propose a text-line alignment technique based on painting algorithm. Here at first, Piece-wise Painting Algorithm (PPA) is used to get a number of black and white rectangular patches all along the text-line for text-line alignment. Identifying the degree of oscillation of the input text-line, some candidate pixels are also obtained based on horizontal projection and center points of the black patches. Using the degree of oscillation of the input text image and the candidate pixels a curve or straight line is fit to trace the baseline. Subsequently, all components of the text-line are deskewed based on analyzing the characteristic of the fit curve or line to align the components with respect to the horizontal imaginary baseline. The proposed algorithm was evaluated with 128 Persian handwritten text-lines containing 4317 sub words. Experimental analysis showed that 92.31% of the sub words were accurately aligned. Further, the proposed algorithm was tested with another Persian handwritten text-lines dataset [6] and remarkable results were achieved. Alireza Alaei, P. Nagabhushan, Umapada Pal 0001 |
ICDAR | 3 |
| 2011 | A Painting Based Technique for Skew Estimation of Scanned DocumentsabstractIn this paper, we propose an efficient skew estimation technique based on Piece-wise Painting Algorithm (PPA) for scanned documents. Here we, at first, employ the PPA on the document image horizontally and vertically. Applying the PPA on both the directions, two painted images (one for horizontally painted and other for vertically painted) are obtained. Next, based on statistical analysis some regions with specific height (width) from horizontally (vertically) painted images are selected and top (left), middle (middle) and bottom (right) points of such selected regions are categorized in 6 separate lists. Utilizing linear regression, a few lines are drawn using the lists of points. A new majority voting approach is also proposed to find the best-fit line amongst all the lines. The skew angle of the document image is estimated from the slope of the best-fit line. The proposed technique was tested extensively on a dataset containing various categories of documents. Experimental results showed that the proposed technique achieved more accurate results than the state-of-the-art methodologies. Alireza Alaei, Umapada Pal 0001, P. Nagabhushan, Fumitaka Kimura |
ICDAR | 2 |
| 2011 | Identification of Indic Scripts on Torn-DocumentsabstractQuestioned Document Examination processes often encompass analysis of torn documents. To aid a forensic expert, automatic classification of content type in torn documents might be useful. This helps a forensic expert to sort out similar document fragments from a pile of torn documents. One parameter of similarity could be the script of the text. In this article we propose a method to identify the script in document fragments. Torn documents are normally characterized by text with arbitrary orientation. We use Zernike moment - based feature that is rotation invariant together with Support Vector Machine (SVM) to classify the script type. Subsequently gradient features are used for comparative analysis of results between rotation dependent and rotation invariant feature type. We achieved an overall script-identification accuracy of 81.39% when dealing with 11 different scripts at character/connected-component level and 94.65% at word level. Sukalpa Chanda, Katrin Franke, Umapada Pal 0001 |
ICDAR | 3 |
| 2011 | Symbol Spotting in Line Drawings through Graph Paths HashingabstractIn this paper we propose a symbol spotting technique through hashing the shape descriptors of graph paths (Hamiltonian paths). Complex graphical structures in line drawings can be efficiently represented by graphs, which ease the accurate localization of the model symbol. Graph paths are the factorized substructures of graphs which enable robust recognition even in the presence of noise and distortion. In our framework, the entire database of the graphical documents is indexed in hash tables by the locality sensitive hashing (LSH) of shape descriptors of the paths. The hashing data structure aims to execute an approximate k-NN search in a sub-linear time. The spotting method is formulated by a spatial voting scheme to the list of locations of the paths that are decided during the hash table lookup process. We perform detailed experiments with various dataset of line drawings and the results demonstrate the effectiveness and efficiency of the technique. Anjan Dutta 0001, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR | 3 |
| 2011 | Database Development and Recognition of Handwritten Devanagari Legal Amount WordsabstractA dataset containing 26,720 handwritten legal amount words written in Hindi and Marathi languages (Devanagari script) is presented in this paper along with a training-free technique to recognize such handwritten legal amounts present on Indian bank cheques. The recognition of handwritten legal amount words in Hindi and Marathi languages is a challenging because of the similar size and shape of many words in the lexicon. Moreover, many words have same suffixes or prefixes. The recognition technique proposed is a combination of two approaches. The first approach is based on gradient, structural and cavity (GSC) features along with a binary vector matching (BVM) technique. The second approach is based on vertical projection profile (VPP) feature and dynamic time warping (DTW). A number of highly matched words in both the approaches are considered for the recognition step in the combined approach based on a ranking scheme. Syntactical knowledge related to the languages is also used to achieve higher reliability. To the best of our knowledge, this is the first work of its kind in recognizing handwritten legal amounts written in Hindi and Marathi. Researchers interested in the dataset can contact the authors to get it through a shared link. R. Jayadevan, Satish R. Kolhe, Pradeep M. Patil, Umapada Pal 0001 |
ICDAR | 4 |
| 2011 | Signature Segmentation from Machine Printed Documents Using Conditional Random FieldabstractAutomatic separation of signatures from a document page involves difficult challenges due to the free-flow nature of handwriting, overlapping/touching of signature parts with printed text, noise, etc. In this paper, we have proposed a novel approach for the segmentation of signatures from machine printed signed documents. The algorithm first locates the signature block in the document using word level feature extraction. Next, the signature strokes that touch or overlap with the printed texts are separated. A stroke level classification is then performed using skeleton analysis to separate the overlapping strokes of printed text from the signature. Gradient based features and Support Vector Machine (SVM) are used in our scheme. Finally, a Conditional Random Field (CRF) model energy minimization concept based on approximated labeling by graph cut is applied to label the strokes as "signature" or "printed text" for accurate segmentation of signatures. Signature segmentation experiment is performed in "tobacco" dataset1 and we have obtained encouraging results. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICDAR | 3 |
| 2011 | Handwritten Street Name Recognition for Indian Postal AutomationabstractAlthough for postal automation there are many pieces of work towards street name recognition on non-Indian languages, to the best of our knowledge there is no work on street name recognition on Indian languages. In this paper we proposed a scheme for recognition of Indian street name written in Bangla script. Because of the writing style of different individuals some of the characters in a street name may touch with its neighboring characters. Accurate segmentation of such touching into individual characters is a difficult task. To avoid such segmentation, here we consider a street name string as word and the street name recognition problem is treated as lexicon driven word recognition. Some of the street names may contain two or more words and we have concatenated these words to have a single word. In the proposed method, at first, street names are binarized and pre-segmented into possible primitive components (individual characters or its parts) analyzing their cavity portions. Pre-segmented components of a street name are then merged into possible characters to get the best street name. Dynamic programming (DP) is applied for the merging using total likelihood of characters as the objective function. To compute the likelihood of a character, modified quadratic discriminant function (MQDF) is used. Our proposed system shows 99.03% reliability with 18.80% rejection, and 0.79% error rates when tested on 4450 handwritten Bangla street name samples. Umapada Pal 0001, Rami Kumar Roy, Fumitaka Kimura |
ICDAR | 1 |
| 2011 | A New Gradient Based Character Segmentation Method for Video Text RecognitionabstractThe current OCR cannot segment words and characters from video images due to complex background as well as low resolution of video images. To have better accuracy, this paper presents a new gradient based method for words and character segmentation from text line of any orientation in video frames for recognition. We propose a Max-Min clustering concept to obtain text cluster from the normalized absolute gradient feature matrix of the video text line image. Union of the text cluster with the output of Canny operation of the input video text line is proposed to restore missing text candidates. Then a run length algorithm is applied on the text candidate image for identifying word gaps. We propose a new idea for segmenting characters from the restored word image based on the fact that the text height difference at the character boundary column is smaller than that of the other columns of the word image. We have conducted experiments on a large dataset at two levels (word and character level) in terms of recall, precision and f-measure. Our experimental setup involves 3527 characters of English and Chinese, and this dataset is selected from TRECVID database of 2005 and 2006. Palaiahnakote Shivakumara, Souvik Bhowmick, Bolan Su, Chew Lim Tan, Umapada Pal 0001 |
ICDAR | 5 |
| 2010 | Query driven word retrieval in graphical documentsabstractIn this paper, we present an approach towards the retrieval of words from graphical document images. In graphical documents, due to presence of multi-oriented characters in non-structured layout, word indexing is a challenging task. The proposed approach uses recognition results of individual components to form character pairs with the neighboring components. An indexing scheme is designed to store the spatial description of components and to access them efficiently. Given a query text word (ascii/unicode format), the character pairs present in it are searched in the document. Next the retrieved character pairs are linked sequentially to form character string. Dynamic programming is applied to find different instances of query words. A string edit distance is used here to match the query word as the objective function. Recognition of multi-scale and multi-oriented character component is done using Support Vector Machine classifier. To consider multi-oriented character strings the features used in the SVM are invariant to character orientation. Experimental results show that the method is efficient to locate a query word from multi-oriented text in graphical documents. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
Document Analysis Systems | 2 |
| 2010 | A new wavelet-median-moment based method for multi-oriented video text detectionabstractIn this paper, we present a new method based on wavelet-median-moments and a novel idea of angle projection for detecting multi-oriented text in video. The proposed method uses wavelet decomposition first to obtain three high frequency sub-bands (LH, HL and HH) and then median moments are computed on the average sub-bands of the three high frequency sub-bands to brighten the text pixels. K-means clustering (K=2) is used for obtaining text pixels from the wavelet-median-moments features (WMMF). Text candidates are obtained by mapping the output of K-means on Sobel edge map of the original input frame. To deal with multi-oriented text, we introduce a new idea of Angle Projection (AP) based on boundary growing and nearest neighbor concepts from the text candidates instead of conventional projection profiles. The proposed method is experimented on horizontal text data, non-horizontal text data, temporal data, non-text data and camera based images (scene text data of ICDAR 2003 competition) to show that the proposed method is superior to existing methods. Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001 |
Document Analysis Systems | 4 |
| 2009 | Fine Classification of Unconstrained Handwritten Persian/Arabic Numerals by Removing Confusion amongst Similar ClassesabstractIn this paper, we propose two types of feature sets based on modified chain-code direction frequencies in the contour pixels of input image and modified transition features (horizontally and vertically). A multi-level support vector machine (SVM) is proposed as classifier to recognize Persian isolated digits. In first level, we combine similar shaped numerals into a single group and as result; we obtain 7 classes instead of 10 classes. We compute 196-dimension chain-code direction frequencies as features to discriminate 7 classes. In the second level, classes containing more than one numeral because of high resemblance in their shapes are considered. We use modified transition features (horizontally and vertically) for discriminating between two overlapping classes (0 and 1). To separate another overlapping group containing three numerals 2, 3 and 4 we first eliminate common parts of these digits (tail) and then compute chain code features. We employ SVM classifier for the classification and evaluate our scheme on 80,000 handwritten samples of Persian numerals [10]. Using 60,000 samples for training, we tested our scheme on other 20,000 samples and obtained 99.02% accuracy. Alireza Alaei, P. Nagabhushan, Umapada Pal 0001 |
ICDAR | 3 |
| 2009 | Two-stage Approach for Word-wise Script IdentificationabstractA two-stage approach for word-wise identification of English (Roman), Devnagari and Bengali (Bangla) scripts is proposed. This approach balances the tradeoff between recognition accuracy and processing speed. The 1st stage allows identifying scripts with high speed, yet less accuracy when dealing with noisy data. The advanced 2nd stage processes only those samples that yield low recognition confidence in the first stage. For both stages a rough character segmentation is performed and features are computed on segmented character components. Features used in the 1st stage are a 64-dimensional chain-code-histogram feature, while 400-dimensional gradient features are used in the 2nd stage. Final classification of a word to a particular script is done via majority voting of each recognized character component of the word. Extensive experiments with various confidence scores were conducted and reported here. The overall recognition accuracy and speed is remarkable. Correct classification of 98.51% on 11,123 test words is achieved, even when the recognition-confidence is as high as 95% at both stages. Sukalpa Chanda, Srikanta Pal, Katrin Franke, Umapada Pal 0001 |
ICDAR | 4 |
| 2009 | Indian Multi-Script Full Pin-code String Recognition for Postal AutomationabstractUnder three-language formula, the destination address block of postal document of an Indian state is generally written in three languages: English, Hindi and the State official language. Because of inter-mixing of these scripts in postal address writings, it is very difficult to identify the script by which a pin-code is written. Also, because of the writing style of different individuals some of the digits in a pin-code string may touch with its neighboring digits. Accurate segmentation of such touching components into individual digits is a difficult task. To avoid such difficulties, in this paper we proposed a tri-lingual (English, Hindi and Bangla) 6-digit full pin-code string recognition. We obtained 99.01% reliability from our proposed system when error and rejection rates are 0.83% and 15.27%, respectively. Umapada Pal 0001, Rami Kumar Roy, Kaushik Roy 0004, Fumitaka Kimura |
ICDAR | 1 |
| 2009 | Comparative Study of Devnagari Handwritten Character Recognition Using Different Feature and ClassifiersabstractIn recent years research towards Indian handwritten character recognition is getting increasing attention. Many approaches have been proposed by the researchers towards handwritten Indian character recognition and many recognition systems for isolated handwritten numerals/characters are available in the literature. To get idea of the recognition results of different classifiers and to provide new benchmark for future research, in this paper a comparative study of Devnagari handwritten character recognition using twelve different classifiers and four sets of feature is presented. Projection distance, subspace method, linear discriminant function, support vector machines, modified quadratic discriminant function, mirror image learning, Euclidean distance, nearest neighbour, k-Nearest neighbour, modified projection distance, compound projection distance, and compound modified quadratic discriminant function are used as different classifiers. Feature sets used in the classifiers are computed based on curvature and gradient information obtained from binary as well as gray-scale images. Umapada Pal 0001, Tetsushi Wakabayashi, Fumitaka Kimura |
ICDAR | 1 |
| 2009 | Seal Detection and Recognition: An Approach for Document IndexingabstractReliable indexing of documents having seal instances can be achieved by recognizing seal information. This paper presents a novel approach for detecting and classifying such multi-oriented seals in these documents. First, Hough Transform based methods are applied to extract the seal regions in documents. Next, isolated text characters within these regions are detected. Rotation and size invariant features and a Support Vector Machine based classifier have been used to recognize these detected text characters. Next, for each pair of character, we encode their relative spatial organization using their distance and angular position with respect to the centre of the seal, and enter this code into a hash table. Given an input seal, we recognize the individual text characters and compute the code for pair-wise character based on the relative spatial organization. The code obtained from the input seal helps to retrieve model hypothesis from the hash table. The seal model to which we get maximum hypothesis is selected for the recognition of the input seal. The methodology is tested to index seal in rotation and size invariant environment and we obtained encouraging results. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
ICDAR | 2 |
| 2009 | Multi-Oriented and Multi-Sized Touching Character Segmentation Using Dynamic ProgrammingabstractIn this paper, we present a scheme towards the segmentation of English multi-oriented touching strings into individual characters. When two or more characters touch, they generate a big cavity region at the background portion. Using Convex Hull information, we use these background information to find some initial points to segment a touching string into possible primitive segments (a primitive segment consists of a single character or a part of a character). Next these primitive segments are merged to get optimum segmentation and dynamic programming is applied using total likelihood of characters as the objective function. SVM classifier is used to find the likelihood of a character. To consider multi-oriented touching strings the features used in the SVM are invariant to character orientation. Circular ring and convex hull ring based approach has been used along with angular information of the contour pixels of the character to make the feature rotation invariant. From the experiment, we obtained encouraging results. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Mathieu Delalandre |
ICDAR | 2 |
| 2009 | F-ratio Based Weighted Feature Extraction for Similar Shape Character RecognitionabstractRecognition of handwritten similar shaped character is a difficult problem and in character recognition system most of the errors occur from similar shaped characters. In this paper we proposed a novel feature extraction technique to improve the recognition results of two similar shaped characters. The technique is based on F-ratio (Fisher Ratio), a statistical measure defined by the ratio to the between-class variance and within-class variance. F-ratio modifies the feature vector of two similar shape characters by weighting the feature elements. This weighting scheme enhances the feature elements that belongs to the distinguishable portions of the similar shaped characters and reduces the feature elements of the common portion of the characters, so that similar shaped characters can be identified easily. We considered pair of handwritten similar shape characters of different scripts like Arabic/Persian, Devnagari English, Bangla, Oriya, Tamil, Kannada, Telugu etc. and we noted that f-ratio based feature weighting shows better recognition results. Tetsushi Wakabayashi, Umapada Pal 0001, Fumitaka Kimura, Yasuji Miyake |
ICDAR | 2 |
| 2008 | Multi-Oriented English Text Line Extraction Using Background and Foreground InformationabstractIn graphical documents (map, engineering drawing), artistic documents etc. there exist many printed materials where text lines are not parallel to each other and they are multi-oriented and curve in nature. For the OCR of such documents we need to extract individual text lines from the documents. Extraction of individual text lines from multi-oriented and/or curved text document is a difficult problem. In this paper, we propose a novel method to extract individual text lines from such document pages and the method is based on the foreground and background information of the characters of the text. To take care of background information, water reservoir concept is used here. In the proposed scheme at first, individual components are detected and grouped into 3-character clusters using their inter-component distance, size and positional information. Applying concept of graph, initial 3-character clusters are merged to have larger cluster group. Using inter-character background information, orientations of the extreme characters of a larger cluster are decided and based on these orientation, two candidate regions are formed from the cluster. Finally, with the help of these candidate regions, individual lines are extracted. From the experiment, we obtained encouraging result. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Fumitaka Kimura |
Document Analysis Systems | 2 |
| 2007 | SVM Based Scheme for Thai and English Script IdentificationabstractIn some Thai documents, a single text line of a document page may contain both Thai and English scripts. For the optical character recognition (OCR) of such a document page it is better to identify, at first, Thai and English script portions and then to use individual OCR system of the respective scripts on these identified portions. In this paper, a SVM based method is proposed for identification of word-wise printed English and Thai scripts from a single line of a document page. Here, at first, the document is segmented into lines and then lines are segmented into character groups (words). In the proposed scheme, we identify the script of the individual character group combining different character features obtained from structural shape, profile, component overlapping information, topological properties, water reservoir concept etc. Based on the experiment on 6110 data we obtained 99.36% script identification accuracy from the proposed scheme. Sukalpa Chanda, Oriol Ramos Terrades, Umapada Pal 0001 |
ICDAR | 3 |
| 2007 | Off-Line Handwritten Character Recognition of Devnagari ScriptabstractIn this paper we present a system towards the recognition of off-line handwritten characters of Devnagari, the most popular script in India. The features used for recognition purpose are mainly based on directional information obtained from the arc tangent of the gradient. To get the feature, at first, a 2times2 mean filtering is applied 4 times on the gray level image and a non-linear size normalization is done on the image. The normalized image is then segmented to 49times49 blocks and a Roberts filter is applied to obtain gradient image. Next, the arc tangent of the gradient (direction of gradient) is initially quantized into 32 directions and the strength of the gradient is accumulated with each of the quantized direction. Finally, the blocks and the directions are down sampled using Gaussian filter to get 392 dimensional feature vector. A modified quadratic classifier is applied on these features for recognition. We used 36172 handwritten data for testing our system and obtained 94.24% accuracy using 5-fold cross-validation scheme. Umapada Pal 0001, Nabin Sharma, Tetsushi Wakabayashi, Fumitaka Kimura |
ICDAR | 1 |
| 2007 | Handwritten Numeral Recognition of Six Popular Indian ScriptsabstractIndia is a multi-lingual multi-script country but there is not much work towards handwritten character recognition of Indian languages. In this paper we propose a modified quadratic classifier based scheme towards the recognition of off-line handwritten numerals of six popular Indian scripts. Here we consider Devnagari, Bangla, Telugu, Oriya, Kannada and Tamil scripts for our experiment. The features used in the classifier are obtained from the directional information of the numerals. For feature computation, the bounding box of a numeral is segmented into blocks and the directional features are computed in each of the blocks. These blocks are then down sampled by a Gaussian filter and the features obtained from the down sampled blocks are fed to a modified quadratic classifier for recognition. Here we have used two sets of feature. We have used 64 dimensional features for high-speed recognition and 400 dimensional features for high-accuracy recognition in our proposed system. A five-fold cross validation technique has been used for result computation and we obtained 99.56%, 98.99%, 99.37%, 98.40%, 98.71% and 98.51% accuracy from Devnagari, Bangla, Telugu, Oriya, Kannada, and Tamil scripts, respectively. Umapada Pal 0001, Nabin Sharma, Tetsushi Wakabayashi, Fumitaka Kimura |
ICDAR | 1 |
| 2005 | Recognition of Indian Multi-oriented and Curved TextabstractIn stylistic (artistic) documents text lines of a single page may have different orientations or the text lines may be curve in shape. As a result, it is difficult to detect the skew of such documents and hence character segmentation as well as recognition of such documents is a complex task. In this paper, we propose a novel scheme towards the recognition of Indian stylistic documents. Here, at first, using water reservoir concept based features the characters are segmented from stylistic documents without any skew correction. Next, individual characters are recognized. For recognition, contour distances of the outer contour points of the characters are calculated from the centroid. These contour distances are then arranged in a particular order to get size and rotation invariant feature. Finally, computing statistical feature on these arranged contour distances the input character is recognized. Umapada Pal 0001, Nilamadhaba Tripathy |
ICDAR | 1 |
| 2005 | Oriya Handwritten Numeral Recognition SysteabstractThis paper deals with recognition of off-line unconstrained Oriya handwritten numerals. To take care of variability involved in the writing style of different individuals, the features are mainly considered from the contour of the numerals. At first, the bounding box of a numeral is segmented into few blocks and chain code histogram is computed in each of the blocks. Features are mainly based on the direction chain code histogram of the contour points of these blocks. Neural network (NN) classifier and quadratic classifier are used separately for recognition and the results obtained from these two classifiers are compared. We tested the result on 3850 data collected from different individuals of various background and we obtained 90.38% (94.81%) recognition accuracy from NN (quadratic) classifier with a rejection rate of about 1.84% (1.31%), respectively. Kaushik Roy 0004, Tandra Pal 0001, Umapada Pal 0001, Fumitaka Kimura |
ICDAR | 3 |
| 2005 | A System for Indian Postal AutomationabstractIn this paper, we present a system towards Indian postal automation based on the recognition of pin-code and city name of the postal document. In the proposed system, at first, non-text blocks (postal stamp, postal seal etc.) are detected and destination address block (DAB) is identified from the document. Next, lines and words of the DAB are segmented. Since India is a multi-lingual and multi-script country, the address part may be written by combination of two scripts. To identify the script by which a word is written, we propose a water reservoir based technique. It is very difficult to identify the script by which the pin-code portion is written. So, we have used two-stage artificial neural network (NN) based general classifiers for the recognition of pin-code digits written in English/Bangla. For recognition of city names, we propose an NSHP-HMM (non-symmetric half plane-hidden Markov model) based technique. Kaushik Roy 0004, Szilárd Vajda, Abdel Belaïd, Umapada Pal 0001, Bidyut B. Chaudhuri |
ICDAR | 4 |
| 2004 | Word-Wise Script Identification from Indian Documents
Suranjit Sinha, Umapada Pal 0001, Bidyut B. Chaudhuri |
Document Analysis Systems | 2 |
| 2003 | Segmentation of Bangla Unconstrained Handwritten TextabstractTo take care of variability involved in the writing style of different individuals in this paper we propose a robust scheme to segment unconstrained handwritten Bangla texts into lines, words and characters. For line segmentation, at first, we divide the text into vertical stripes. Stripe width of a document is computed by statistical analysis of the text height in the document. Next we determine horizontal histogram of these stripes and the relationship of the minimal values of the histograms is used to segment text lines. Based on vertical projection profile lines are segmented into words. Segmentation of characters from handwritten word is very tricky as the characters are seldom vertically separable. We use a concept based on water reservoir principle for the purpose. Here we, at first, identify isolated and connected (touching) characters in a word. Next touching characters of the word are segmented based on the reservoir base area points and structural feature of the component. 1. Umapada Pal 0001, Sagarika Datta |
ICDAR | 1 |
| 2003 | Recognition of Printed Urdu ScriptabstractThis paper deals with an Optical Character Recognition system for printed Urdu, a popular Indian script. The development of OCR for this script is difficult because (i) a large number of characters have to be recognized (ii) there are many similar shaped characters. In the proposed system individual characters are recognized using a combination of topological, contour and water reservoir concept based features. The feature detection methods are simple and robust. A prototype of the system has been tested on printed Urdu characters and currently achieves 97.8 % character level accuracy on average. 1. Umapada Pal 0001 |
ICDAR | 1 |
| 2003 | Multi-Script Line identification from Indian DocumentabstractA document page may contain two or more different scripts. For Optical Character Recognition (OCR) of such a document page, it is necessary to separate different scripts before feeding them to their individual OCR system. In this paper an automatic scheme is presented to identify text lines of different Indian scripts from a document. For the separation task at first the scripts are grouped into a few classes according to script characteristics. Next feature based on water reservoir principle, contour tracing, profile etc. are employed to identify them without any expensive OCR-like algorithms. At present, the system has an overall accuracy of about 97.52%. 1. Umapada Pal 0001, Suranjit Sinha, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 2001 | Water Reservoir Based Approach for Touching Numeral SegmentationabstractDeals with a scheme for automatic segmentation of unconstrained handwritten connected numerals. The scheme is mainly based on features obtained from a new concept based on a water reservoir. A reservoir is a metaphor to illustrate the region where numerals touch. The reservoir is obtained by considering accumulation of water poured from the top or from the bottom of the numerals. At first, considering the reservoir location and size, touching positions (top, middle and bottom) are decided. Next, by analyzing the reservoir boundary, touching position and topological features of the touching pattern, the best cutting point is determined. Finally, combined with morphological structural features the cutting path for segmentation is generated. Abdel Belaïd, Christophe Choisy, Umapada Pal 0001 |
ICDAR | 3 |
| 2001 | Automatic Recognition of Printed Oriya ScriptabstractThe paper deals with an optical character recognition system for printed Oriya, a popular Indian script. The development of OCR for this script is difficult because a large number of characters have to be recognized. In the proposed system, the digitized document image is first passed through preprocessing modules like skew correction, line segmentation, zone detection, word and character segmentation, etc. These modules have been developed by combining some conventional techniques with some newly proposed ones. Next, individual characters are recognized using a combination of stroke and run-number based features, along with features obtained from the concept of a water reservoir. The feature detection methods are simple and robust. A prototype of the system has been tested on a variety of printed Oriya material, and currently achieves 96.3% character level accuracy on average. Bidyut B. Chaudhuri, Umapada Pal 0001, Mandar Mitra |
ICDAR | 2 |
| 2001 | Automatic Identification of English, Chinese, Arabic, Devnagari and Bangla Script LineabstractIn a general situation, a document page may contain several scriptforms. For optical character recognition (OCR) of such a document page, it is necessary to separate the scripts before feeding them to their individual OCR systems. An automatic technique for the identification of printed Roman, Chinese, Arabic, Devnagari and Bangla text lines from a single document is proposed. Shape based features, statistical features and some features obtained from the concept of a water reservoir are used for script identification. The proposed scheme has an accuracy of about 97.33%. Umapada Pal 0001, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 2001 | Multi-Skew Detection of Indian Script DocumentsabstractThere are many documents where text lines are not parallel to each other i.e. these lines have different inclinations with the horizontal lines (multi-skew documents). For the OCR of such a document we have to estimate the skew angle of individual text lines because a single rotation cannot de-skew all text lines of the document. In this paper, we describe a robust technique for multi-skew angle detection from Indian documents containing the most popular Indian scripts Devnagari and Bangla. Most characters in these scripts have horizontal lines at the top, called head-lines. The character head-lines usually connect one another in a word and the word appears as a single component. In the proposed method, the connected components are at first labeled and selected. The upper envelopes of selected components are found by column-wise scanning from the top of the component. Portions of the upper envelope satisfying the properties of a digital straight line are detected. They are then clustered into groups belonging to single text lines. Estimates from these individual clusters give the skew angle of each text line. The proposed multi-skew detection technique has an accuracy about 98.3%. Umapada Pal 0001, Mandar Mitra, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 1999 | Script Line Separation from Indian Multi-Script DocumentsabstractIn a multi-lingual country like India, a document page may contain more than one script form. Under the three-language formula, the document may be printed in English, Devnagari and one of the other official Indian languages. For OCR of such a document page, it is necessary to separate these three script forms before feeding them to the OCRs of individual scripts. In this paper, an automatic technique of separating the text lines using script characteristics and shape based features is presented. At present, the system has an overall accuracy of about 98.5%. Umapada Pal 0001, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 1999 | Automatic Separation of Machine-Printed and Hand-Written Text LinesabstractThere are many types of documents where machine-printed and hand-written texts appear intermixed. Since the optical character recognition (OCR) methodologies for machine-printed and hand-written texts are different, it is necessary to separate these two types of text before feeding them to the respective OCR systems. In this paper, we present such a scheme for both Bangla and Devnagari characters. The scheme is based on the structural and statistical features of the machine-printed and hand-written text lines. The classification scheme has an accuracy of about 98.3%. Umapada Pal 0001, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 1997 | An OCR System to Read Two Indian Language Scripts: Bangla and Devnagari (Hindi)abstractAn OCR system is proposed that can read two Indian language scripts: Bangla and Devnagari (Hindi), the most popular ones in the Indian subcontinent. These scripts, having the same origin in ancient Brahmi script, have many features in common and hence a single system can be modeled to recognize them. In the proposed model, document digitization, skew detection, text line segmentation and zone separation, word and character segmentation, character grouping into basic, modifier and compound character category are done for both scripts by the same set of algorithms. The feature sets and classification tree as well as the knowledge base required for error correction (such as lexicon) differ for Bangla and Devnagari. The system shows a good performance for single font scripts printed on clear documents. Bidyut B. Chaudhuri, Umapada Pal 0001 |
ICDAR | 2 |
| 1997 | Automatic Separation of Words in Multi-lingual Multi-script Indian DocumentsabstractIn a multi-lingual country like India, a document may contain more than one script forms. For such a document it is necessary to separate different script forms before feeding them to OCRs of individual script. In this paper an automatic word segmentation approach is described which can separate Roman, Bangla and Devnagari scripts present in a single document. The approach has a tree structure where at first Roman script words are separated using the 'headline' feature. The headline is common in Bangla and Devnagari but absent in Roman. Next, Bangla and Devnagari words are separated using some finer characteristics of the character set although recognition of individual character is avoided. At present, the system has an overall accuracy of 96.09%. Umapada Pal 0001, Bidyut B. Chaudhuri |
ICDAR | 1 |