C. V. Jawahar

dblp:j/CVJawahar · also Cheerakkuzhi Veluthemana Jawahar · DBLP profile ↗
← Back
72ranked-venue papers in the field
1as first author
18since 2021 · last 2026
0000-0001-6767-7057ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 67 (1 first)Information Retrieval & Web Search · 4Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Can VLMs Understand Handwritten Mathematical Documents?
Shree Mitra, Ajoy Mondal, C. V. Jawahar
ICDAR (3)3
2025 Adapting Vision-Language Models for Hindi OCR
Shaon Bhattacharyya, Prantik Deb, Ajoy Mondal, C. V. Jawahar
ICDAR (3)5
2025 AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
Suyash Maniyar, Vishvesh Trivedi, Ajoy Mondal, Anand Mishra 0001, C. V. Jawahar
ICDAR (1)5
2025 Attend to What I Say: Highlighting Relevant Content on Slides
K. M. Megha Mariam, C. V. Jawahar
ICDAR (4)2
2025 ICDAR 2025 Handwritten Notes Understanding Challenge
Aniket Pal, Sanket Biswas, Alloy Das, Ayush Lodh, Priyanka Banerjee, Soumitri Chattopadhyay, Ajoy Mondal, Dimosthenis Karatzas, Josep Lladós 0001, C. V. Jawahar
ICDAR (5)10
2025 EviFiVQA: A Benchmark for Evidence-Grounded Multi-hop Reasoning in Financial VQA
Sachin Raja, Ajoy Mondal, C. V. Jawahar
ICDAR (4)3
2025 UniLayDet: Simple Multi-dataset Document Layout Analysis
Prasidh Srikumar, Ajoy Mondal, C. V. Jawahar
ICDAR (1)3
2024 ICDAR 2024 Competition on Reading Documents Through Aria Glasses
Soumya Jahagirdar, Ajoy Mondal, Yuheng (Carl) Ren, Omkar M. Parkhi, C. V. Jawahar
ICDAR (6)5
2024 ICDAR 2024 Competition on Recognition and VQA on Handwritten Documents
Ajoy Mondal, Vijay Mahadevan, R. Manmatha, C. V. Jawahar
ICDAR (6)4
2024 Bridging the Gap in Resource for Offline English Handwritten Text Recognition
Ajoy Mondal, Krishna Tulsyan, C. V. Jawahar
ICDAR (2)3
2024 Indic Scene Text on the Roadside
Ajoy Mondal, Krishna Tulsyan, C. V. Jawahar
ICDAR (5)3
2023 ICDAR 2023 Competition on Indic Handwriting Text Recognition
Ajoy Mondal, C. V. Jawahar
ICDAR (2)2
2023 ICDAR 2023 Competition on Visual Question Answering on Business Document Images
Sachin Raja, Ajoy Mondal, C. V. Jawahar
ICDAR (2)3
2023 ICDAR 2023 Competition on RoadText Video Text Detection, Tracking and Recognition
George Tom, Minesh Mathew, Sergi Garcia-Bordils, Dimosthenis Karatzas, C. V. Jawahar
ICDAR (2)5
2023 Reading Between the Lanes: Text VideoQA on the Road
George Tom, Minesh Mathew, Sergi Garcia-Bordils, Dimosthenis Karatzas, C. V. Jawahar
ICDAR (6)5
2022 Read While You Drive - Multilingual Text Tracking on the Road
Sergi Garcia-Bordils, George Tom, Sangeeth Reddy, Minesh Mathew, Marçal Rusiñol, C. V. Jawahar, Dimosthenis Karatzas
DAS6
2021 iiit-indic-hw-words: A Dataset for Indic Handwritten Text Recognition
Santhoshini Gongidi, C. V. Jawahar
ICDAR (4)2
2021 ICDAR 2021 Competition on Document Visual Question Answering
Rubèn Tito, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas
ICDAR (4)3
2020 Graph Representation Ensemble Learning
abstract
Representation learning on graphs has been gaining attention due to its wide applicability in predicting missing links and classifying and recommending nodes. Most embedding methods aim to preserve specific properties of the original graph in the low dimensional space. However, real-world graphs have a combination of several features that are difficult to characterize and capture by a single approach. In this work, we introduce the problem of graph representation ensemble learning and provide a first of its kind framework to aggregate multiple graph embedding methods efficiently. We provide analysis of our framework and analyze - theoretically and empirically - the dependence between state-of-the-art embedding methods. We test our models on the node classification task on four realworld graphs and show that proposed ensemble approaches can outperform the state-of-the-art methods by up to 20% on macro-F1. We further show that the strategy is even more beneficial for underrepresented classes with an improvement of up to 40%.
Palash Goyal, Sachin Raja, Sujit Rokka Chhetri, Arquimedes Canedo, Ajoy Mondal, Jaya Shree, C. V. Jawahar
ASONAM8
2020 Fused Text Recogniser and Deep Embeddings Improve Word Recognition and Retrieval
Siddhant Bansal, Praveen Krishnan, C. V. Jawahar
DAS3
2020 Adapting OCR with Limited Supervision
Deepayan Das, C. V. Jawahar
DAS2
2020 IIIT-AR-13K: A New Dataset for Graphical Object Detection in Documents
Ajoy Mondal, Peter Lipps, C. V. Jawahar
DAS3
2020 A Benchmark System for Indian Language Text Recognition
Krishna Tulsyan, Nimisha Srivastava, Ajoy Mondal, C. V. Jawahar
DAS4
2019 ICDAR 2019 Competition on Scene Text Visual Question Answering
abstract
This paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual Question Answering system up to date, namely the incorporation of scene text to answer questions asked about an image. The competition introduces a new dataset comprising 23,038 images annotated with 31,791 question / answer pairs where the answer is always grounded on text instances present in the image. The images are taken from 7 different public computer vision datasets, covering a wide range of scenarios. The competition was structured in three tasks of increasing difficulty, that require reading the text in a scene and understanding it in the context of the scene, to correctly answer a given question. A novel evaluation metric is presented, which elegantly assesses both key capabilities expected from an optimal model: text recognition and image understanding. A detailed analysis of results from different participants is showcased, which provides insight into the current capabilities of VQA systems that can read. We firmly believe the dataset proposed in this challenge will be an important milestone to consider towards a path of more robust and general models that can exploit scene text to achieve holistic image understanding.
Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas
ICDAR7
2019 A Cost Efficient Approach to Correct OCR Errors in Large Document Collections
abstract
Word error rate of an OCR is often higher than its character error rate. This is especially true when OCRs are designed by recognizing characters. High word accuracies are critical for many practical applications like content creation and text-to-speech systems. In order to detect and correct the misrecognised words, it is common for an OCR to employ a post-processor module to improve the word accuracy. However, conventional approaches to post-processing like looking up a dictionary or using a statistical language model (SLM), are still limited. In many such scenarios, it is often required to remove the outstanding errors manually. We observe that the traditional post-processing schemes look at error words sequentially, since OCRs process documents one at a time. We propose a cost-efficient model to address the error words in batches rather than correcting them individually. We exploit the fact that a collection of documents (eg. a book), unlike a single document, has a structure leading to repetition of words. Such words, if efficiently grouped together and corrected together, can lead to a significant reduction in the effort. Error correction can be fully automatic or with a human in the loop. We compare the performance of our method with various baseline approaches including the case where all the errors are removed by a human. We demonstrate the efficacy of our solution empirically by reporting more than 70% reduction in the human effort with near perfect error correction. We validate our method on books in both English and Hind.
Deepayan Das, Jerin Philip, Minesh Mathew, C. V. Jawahar
ICDAR4
2019 ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
abstract
The ICDAR 2019 Challenge on "Scanned receipts OCR and key information extraction" (SROIE) covers important aspects related to the automated analysis of scanned receipts. The SROIE tasks play a key role in many document analysis systems and hold significant commercial potential. Although a lot of work has been published over the years on administrative document analysis, the community has advanced relatively slowly, as most datasets have been kept private. One of the key contributions of SROIE to the document analysis community is to offer a first, standardized dataset of 1000 whole scanned receipt images and annotations, as well as an evaluation procedure for such tasks. The Challenge is structured around three tasks, namely Scanned Receipt Text Localization (Task 1), Scanned Receipt OCR (Task 2) and Key Information Extraction from Scanned Receipts (Task 3). The competition opened on 10th February, 2019 and closed on 5th May, 2019. We received 29, 24 and 18 valid submissions received for the three competition tasks, respectively. This report presents the competition datasets, define the tasks and the evaluation protocols, offer detailed submission statistics, as well as an analysis of the submitted performance. While the tasks of text localization and recognition seem to be relatively easy to tackle, it is interesting to observe the variety of ideas and approaches proposed for the information extraction task. According to the submissions' performance we believe there is still margin for improving information extraction performance, although the current dataset would have to grow substantially in following editions. Given the success of the SROIE competition evidenced by the wide interest generated and the healthy number of submissions from academic, research institutes and industry over different countries, we consider that the SROIE competition can evolve into a useful resource for the community, drawing further attention and promoting research and development efforts in this field.
Kai Chen 0006, Jianhua He 0001, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar
ICDAR7
2019 Textual Description for Mathematical Equations
abstract
Reading of mathematical expression or equation in the document images is very challenging due to the large variability of mathematical symbols and expressions. In this paper, we pose reading of mathematical equation as a task of the generation of the textual description which interprets the internal meaning of this equation. Inspired by the natural image captioning problem in computer vision, we present a mathematical equation description ( MED ) model, a novel end-to-end trainable deep neural network based approach that learns to generate a textual description for reading mathematical equation images. Our MED model consists of a convolution neural network as an encoder that extracts features of input mathematical equation images and a recurrent neural network with attention mechanism which generates description related to the input mathematical equation images. Due to the unavailability of mathematical equation image data sets with their textual descriptions, we generate two data sets for experimental purpose. To validate the effectiveness of our MED model, we conduct a real-world experiment to see whether the students are able to write equations by only reading or listening their textual descriptions or not. Experiments conclude that the students are able to write most of the equations correctly by reading their textual descriptions only.
Ajoy Mondal, C. V. Jawahar
ICDAR2
2019 Towards Automated Evaluation of Handwritten Assessments
abstract
Automated evaluation of handwritten answers has been a challenging problem for scaling the education system for many years. Speeding up the evaluation remains as the major bottleneck for enhancing the throughput of instructors. This paper describes an effective method for automatically evaluating the short descriptive handwritten answers from the digitized images. Our goal is to evaluate a student's handwritten answer by assigning an evaluation score that is comparable to the human-assigned scores. Existing works in this domain mainly focused on evaluating handwritten essays with handcrafted, non-semantic features. Our contribution is two-fold: 1) we model this problem as a self-supervised, feature-based classification problem, which can fine-tune itself for each question without any explicit supervision. 2) We introduce the usage of semantic analysis for auto-evaluation in handwritten text space using the combination of Information Retrieval and Extraction (IRE) and, Natural Language Processing (NLP) methods to derive a set of useful features. We tested our method on three datasets created from various domains, using the help of students of different age groups. Experiments show that our method performs comparably to that of human evaluators.
Vijay Rowtula, Subba Reddy Oota, C. V. Jawahar
ICDAR3
2019 Graphical Object Detection in Document Images
abstract
Graphical elements: particularly tables and figures contain a visual summary of the most valuable information contained in a document. Therefore, localization of such graphical objects in the document images is the initial step to understand the content of such graphical objects or document images. In this paper, we present a novel end-to-end trainable deep learning based framework to localize graphical objects in the document images called as Graphical Object Detection ( GOD ). Our framework is data-driven and does not require any heuristics or meta-data to locate graphical objects in the document images. The GOD explores the concept of transfer learning and domain adaptation to handle scarcity of labeled training images for graphical object detection task in the document images. Performance analysis carried out on the various public benchmark data sets: ICDAR -2013, ICDAR - POD2017 and UNLV shows that our model yields promising results as compared to state-of-the-art techniques.
Ranajit Saha, Ajoy Mondal, C. V. Jawahar
ICDAR3
2019 ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
abstract
Chinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinese characters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12.
Rui Zhang 0056, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Nan Li 0071, Dong Wang 0004, Minghui Liao
ICDAR7
2019 Self-Supervised Visual Representations for Cross-Modal Retrieval
abstract
Cross-modal retrieval methods have been significantly improved in last years with the use of deep neural networks and large-scale annotated datasets such as ImageNet and Places. However, collecting and annotating such datasets requires a tremendous amount of human effort and, besides, their annotations are limited to discrete sets of popular visual classes that may not be representative of the richer semantics found on large-scale cross-modal retrieval datasets. In this paper, we present a self-supervised cross-modal retrieval framework that leverages as training data the correlations between images and text on the entire set of Wikipedia articles. Our method consists in training a CNN to predict: (1) the semantic context of the article in which an image is more probable to appear as an illustration, and (2) the semantic context of its caption. Our experiments demonstrate that the proposed method is not only capable of learning discriminative visual representations for solving vision tasks like classification, but that the learned representations are better for cross-modal retrieval when compared to supervised pre-training of the network on the ImageNet dataset.
Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar
ICMR5
2018 Offline Handwriting Recognition on Devanagari Using a New Benchmark Dataset
abstract
Handwriting recognition (HWR) in Indic scripts, like Devanagari is very challenging due to the subtleties in the scripts, variations in rendering and the cursive nature of the handwriting. Lack of public handwriting datasets in Indic scripts has long stymied the development of offline handwritten word recognizers and made comparison across different methods a tedious task in the field. In this paper, we release a new handwritten word dataset for Devanagari, IIIT-HW-Dev to alleviate some of these issues. We benchmark the IIIT-HW-Dev dataset using a CNN-RNN hybrid architecture. Furthermore, using this architecture, we empirically show that usage of synthetic data and cross lingual transfer learning helps alleviate the issue of lack of training data. We use this proposed pipeline on a public dataset, RoyDB and achieve state of the art results.
Kartik Dutta, Praveen Krishnan, Minesh Mathew, C. V. Jawahar
DAS4
2018 Word Spotting and Recognition Using Deep Embedding
abstract
Deep convolutional features for word images and textual embedding schemes have shown great success in word spotting. In this work, we follow these motivations to propose an End2End embedding framework which jointly learns both the text and image embeddings using state of the art deep convolutional architectures. The three major contributions of this work are: (i) an End2End embedding scheme to learn a common representation for word images and its labels, (ii) building a state of art word image descriptor and demonstrating its utility as off-the-shelf features for word spotting, and (iii) use of synthetic data as a complementary modality to further enhance word spotting and recognition. On the challenging IAM handwritten dataset, we report a mAP of 0.9509 for query-by-string based retrieval task. Under lexicon based word recognition, our proposed method report a 2.66 and 5.10 CER and WER respectively.
Praveen Krishnan, Kartik Dutta, C. V. Jawahar
DAS3
2016 Multilingual OCR for Indic Scripts
abstract
In Indian scenario, a document analysis system has to support multiple languages at the same time. With emerging multilingualism in urban India, often bilingual, trilingual or even more languages need to be supported. This demands development of a multilingual OCR system which can work seamlessly across Indic scripts. In our approach the script is identified at word level, prior to the recognition of the word. An end-to-end RNN based architecture which can detect the script and recognize the text in a segmentation-free manner is proposed for this purpose. We demonstrate the approach for 12 Indian languages and English. It is observed that, even with the similar architecture, performance on Indian languages are poorer compared to English. We investigate this further. Our approach is evaluated on a large corpus comprising of thousands of pages. The Hindi OCR is compared with other popular OCRs for the language, as a further testimony for the efficacy of our method.
Minesh Mathew, Ajeet Kumar Singh, C. V. Jawahar
DAS3
2016 A Simple and Effective Solution for Script Identification in the Wild
abstract
We present an approach for automatically identifying the script of the text localized in the scene images. Our approach is inspired by the advancements in mid-level features. We represent the text images using mid-level features which are pooled from densely computed local features. Once text images are represented using the proposed mid-level feature representation, we use an off-the-shelf classifier to identify the script of the text image. Our approach is efficient and requires very less labeled data. We evaluate the performance of our method on a recently introduced CVSI dataset, demonstrating that the proposed approach can correctly identify script of 96.70% of the text images. In addition, we also introduce and benchmark a more challenging Indian Language Scene Text (ILST) dataset for evaluating the performance of our method.
Ajeet Kumar Singh, Anand Mishra 0001, Pranav Dabral, C. V. Jawahar
DAS4
2016 Error Detection in Indic OCRs
abstract
A good post processing module is an indispensable part of an OCR pipeline. In this paper, we propose a novel method for error detection in Indian language OCR output. Our solution uses a recurrent neural network (RNN) for classification of a word as an error or not. We propose a generic error detection method and demonstrate its effectiveness on four popular Indian languages. We divide the words into their constituent aksharas and use their bigram and trigram level information to build a feature representation. In order to train the classifier on incorrect words, we use the mis-recognized words in the output of the OCR. In addition to RNN, we also explore the effectiveness of a generative model such as GMM for our task and demonstrate an improved performance by combining both the approaches. We tested our method on four popular Indian languages and report an average error detection performance above 80%.
V. S. Vinitha, C. V. Jawahar
DAS2
2016 Diverse Yet Efficient Retrieval using Locality Sensitive Hashing
abstract
Typical retrieval systems have three requirements: a) Accurate retrieval, i.e., the method should have high precision, b) Diverse retrieval, i.e., the obtained set of samples should be diverse, and c) Retrieval time should be small. However, most of the existing methods address only one or two of the above mentioned requirements. In this work, we present a method based on randomized locality sensitive hashing which tries to address all of the above requirements simultaneously. While earlier hashing-based approaches considered approximate retrieval to be acceptable only for the sake of efficiency, we argue that one can further exploit approximate retrieval to provide impressive trade-offs between accuracy and diversity. We also extend our method to the problem of multi-label prediction, where the goal is to output a diverse and accurate set of labels for a given document in real-time. Finally, we present empirical results on image and text retrieval tasks and show that our method retrieves diverse and accurate images/labels while ensuring 100x-speed-up over the existing diverse retrieval approaches.
Vidyadhar Rao, Prateek Jain 0002, C. V. Jawahar
ICMR3
2015 Online handwriting recognition using depth sensors
abstract
In this work, we propose an online handwriting solution, where the data is captured with the help of depth sensors. Users may write in the air and our method recognizes it in real time using the proposed feature representation. Our method uses an efficient fingertip tracking approach and reduces the necessity of pen-up/pen-down switching. We validate our method on two depth sensors, Kinect and Leap Motion Controller. On a dataset collected from 20 users, we achieve a recognition accuracy of 97.59% for character recognition. We also demonstrate how this system can be extended for lexicon recognition with reliable performance. We have also prepared a dataset containing 1,560 characters and 400 words with the intention of providing common benchmark for handwritten character recognition using depth sensors and related research.
Rajat Aggarwal, Sirnam Swetha, Anoop M. Namboodiri, Jayanthi Sivaswamy, C. V. Jawahar
ICDAR5
2015 Efficient word image retrieval using fast DTW distance
abstract
Dynamic time warping (DTW) is a popular distance measure used for recognition free document image retrieval. However, it has quadratic complexity and hence is computationally expensive for large scale word image retrieval. In this paper, we use a fast approximation to the DTW distance, which makes word retrieval efficient. For a pair of sequences, to compute their DTW distance, we need to find the optimal alignment from all the possible alignments. This is a computationally expensive operation. In this work, we learn a small set of global principal alignments from the training data and avoid the computation of alignments for query images. Thus, our proposed approximation is significantly faster compared to DTW distance, and gives 40 times speed up. We approximate the DTW distance as a sum of multiple weighted Eulidean distances which are known to be amenable to indexing and efficient retrieval. We show the speed up of proposed approximation on George Washington collection and multi-language datasets containing words from English and two Indian languages.
Gattigorla Nagendar, C. V. Jawahar
ICDAR2
2015 Unsupervised feature learning for optical character recognition
abstract
Most of the popular optical character recognition (OCR) architectures use a set of handcrafted features and a powerful classifier for isolated character classification. Success of these methods often depend on the suitability of these features for the language of interest. In recent years, whole word recognition based on Recurrent Neural Networks (RNN) has gained popularity. These methods use simple features such as raw pixel values or profiles. Success of these methods depend on the learning capabilities of these networks to encode the script and language information. In this work, we investigate the possibility of learning an appropriate set of features for designing OCR for a specific language. We learn the language specific features from the data with no supervision. This enables the seamless adaptation of the architecture across languages. In this work, we learn features using a stacked Restricted Boltzman Machines (RBM) and use it with the RNN based recognition solution. We validate our method on five different languages. In addition, these novel features also result in better convergence rate of the RNNs.
Devendra K. Sahu, C. V. Jawahar
ICDAR2
2015 Can RNNs reliably separate script and language at word and line level?
abstract
In this work, we investigate the utility of Recurrent Neural Networks (RNNs) for script and language identification. Both these problems have been attempted in the past with representations computed from the distribution of connected components or characters (e.g. texture, n-gram). Often these features are computed from a larger segment (a paragraph or a page). We argue that one can predict the script or language with minimal evidence (e.g. given only a word or a line) very accurately with the help of a pre-trained RNN. We propose a simple and generic solution for the task of script and language identification which do not require any special tuning. Our method represents the word images as a sequence of feature vectors, and employ the RNNs for the identification. We verify the method on a large corpus of more than 15.03M words from 55K document images comprising 15 scripts and languages. We report an accurate script and language identification at word and line level.
Ajeet Kumar Singh, C. V. Jawahar
ICDAR2
2014 Towards a Robust OCR System for Indic Scripts
abstract
The current Optical Character Recognition OCR systems for Indic scripts are not robust enough for recognizing arbitrary collection of printed documents. Reasons for this limitation includes the lack of resources (e.g. not enough examples with natural variations, lack of documentation available about the possible font/style variations) and the architecture which necessitates hard segmentation of word images followed by an isolated symbol recognition. Variations among scripts, latent symbol to UNICODE conversion rules, non-standard fonts/styles and large degradations are some of the major reasons for the unavailability of robust solutions. In this paper, we propose a web based OCR system which (i) follows a unified architecture for seven Indian languages, (ii) is robust against popular degradations, (iii) follows a segmentation free approach, (iv) addresses the UNICODE re-ordering issues, and (v) can enable continuous learning with user inputs and feedbacks. Our system is designed to aid the continuous learning while being usable i.e., we capture the user inputs (say example images) for further improving the OCRs. We use the popular BLSTM based transcription scheme to achieve our target. This also enables incremental training and refinement in a seamless manner. We report superior accuracy rates in comparison with the available OCRs for the seven Indian languages.
Praveen Krishnan, Naveen Sankaran, Ajeet Kumar Singh, C. V. Jawahar
Document Analysis Systems4
2013 Detection of Cut-and-Paste in Document Images
abstract
Many documents are created by Cut-And-Paste (CAP) of existing documents. In this paper, we proposed a novel technique to detect CAP in document images. This can help in detecting unethical CAP in document image collections. Our solution is recognition free, and scalable to large collection of documents. Our formulation is also independent of the imaging process (camera based or scanner based) and does not use any language specific information for matching across documents. We model the solution as finding a mixture of homographies, and design a linear programming (LP) based solution to compute the same. Our method is presently limited by the fact that we do not support detection of CAP in documents formed by editing of the textual content. Our experiments demonstrate that without loss of generality (i.e. without assuming the number of source documents), we can correctly detect and match the CAP content in a questioned document image by simultaneously comparing with large number of images in the database. We achieve the CAP detection accuracy of as high as 90%, even when the spatial extent of the CAP content in a document image is as small as 15% of the entire image area.
Ankit Gandhi, C. V. Jawahar
ICDAR2
2013 Whole is Greater than Sum of Parts: Recognizing Scene Text Words
abstract
Recognizing text in images taken in the wild is a challenging problem that has received great attention in recent years. Previous methods addressed this problem by first detecting individual characters, and then forming them into words. Such approaches often suffer from weak character detections, due to large intra-class variations, even more so than characters from scanned documents. We take a different view of the problem and present a holistic word recognition framework. In this, we first represent the scene text image and synthetic images generated from lexicon words using gradient-based features. We then recognize the text in the image by matching the scene and synthetic image features with our novel weighted Dynamic Time Warping (wDTW) approach. We perform experimental analysis on challenging public datasets, such as Street View Text and ICDAR 2003. Our proposed method significantly outperforms our earlier work in Mishra et al. (CVPR 2012), as well as many other recent works, such as Novikova et al. (ECCV 2012), Wang et al. al.(ICPR 2012), Wang et al.(ICCV 2011).
Vibhor Goel, Anand Mishra 0001, Karteek Alahari, C. V. Jawahar
ICDAR4
2013 Bringing Semantics in Word Image Retrieval
abstract
Performance of the recognition free approaches for document retrieval, heavily depends on the exact or approximate matching of images (in some feature space) to retrieve documents containing the same word. However, the harder problem in information retrieval is to effectively bring semantics into the retrieval pipeline. This is further challenging when the matching is based on visual features. In this work, we investigate this problem, and suggest a solution by directly transferring the semantics from the textual domain. Our retrieval framework uses (i) the language resources like Word Net and (ii) an annotated corpus of document images, to retrieve semantically relevant words from a large word image database. We demonstrate the method on two languages - English and Hindi, and quantitatively evaluate the performance on annotated word image databases of more than a Million images.
Praveen Krishnan, C. V. Jawahar
ICDAR2
2013 Sparse Document Image Coding for Restoration
abstract
Sparse representation based image restoration techniques have shown to be successful in solving various inverse problems such as denoising, in painting, and super-resolution, etc. on natural images and videos. In this paper, we explore the use of sparse representation based methods specifically to restore the degraded document images. While natural images form a very small subset of all possible images admitting the possibility of sparse representation, document images are significantly more restricted and are expected to be ideally suited for such a representation. However, the binary nature of textual document images makes dictionary learning and coding techniques unsuitable to be applied directly. We leverage the fact that different characters possess similar strokes, curves, and edges, and learn a dictionary that gives sparse decomposition for patches. Experimental results show significant improvement in image quality and OCR performance on documents collected from a variety of sources such as magazines and books. This method is therefore, ideally suited for restoring highly degraded images in repositories such as digital libraries.
Vijay Kumar 0004, Amit Bansal, Goutam Hari Tulsiyan, Anand Mishra 0001, Anoop M. Namboodiri, C. V. Jawahar
ICDAR6
2013 Character N-Gram Spotting on Handwritten Documents Using Weakly-Supervised Segmentation
abstract
In this paper, we present a solution towards building a retrieval system over handwritten document images that i) is recognition-free, ii) allows text-querying, iii) can retrieve at sub-word level, iv) can search for out-of-vocabulary words. Unlike previous approaches that operate at either character or word levels, we use character n-gram images (CNG-img) as the retrieval primitive. CNG-img are sequences of character segments, that are represented and matched in the image-space. The word-images are now treated as a bag-of-CNG-img, that can be indexed and matched in the feature space. This allows for recognition-free search (query-by-example), which can retrieve morphologically similar words that have matching sub-words. Further, to enable query-by-keyword, we build an automated scheme to generate labeled exemplars for characters and character n-grams, from unconstrained handwritten documents. We pose this problem as one of weakly-supervised learning, where character/n-gram labeling is obtained automatically from the word labels. The resulting retrieval system can answer queries from an unlimited. vocabulary. The approach is demonstrated on the George Washington collection, results show major improvement in retrieval performance as compared to word-recognition and word-spotting methods.
Udit Roy, Naveen Sankaran, K. Pramod Sankar, C. V. Jawahar
ICDAR4
2013 Error Detection in Highly Inflectional Languages
abstract
Error detection in OCR output using dictionaries and statistical language models (SLMs) have become common practice for some time now, while designing post-processors. Multiple strategies have been used successfully in English to achieve this. However, this has not yet translated towards improving error detection performance in many inflectional languages, specially Indian languages. Challenges such as large unique word list, lack of linguistic resources, lack of reliable language models, etc. are some of the reasons for this. In this paper, we investigate the major challenges in developing error detection techniques for highly inflectional Indian languages. We compare and contrast several attributes of English with inflectional languages such as Telugu and Malayalam. We make observations by analyzing statistics computed from popular corpora and relate these observations to the error detection schemes. We propose a method which can detect errors for Telugu and Malayalam, with an F-Score comparable to some of the less inflectional languages like Hindi. Our method learns from the error patterns and SLMs.
Naveen Sankaran, C. V. Jawahar
ICDAR2
2013 Devanagari Text Recognition: A Transcription Based Formulation
abstract
Optical Character Recognition (OCR) problems are often formulated as isolated character (symbol) classification task followed by a post-classification stage (which contains modules like Unicode generation, error correction etc.) to generate the textual representation, for most of the Indian scripts. Such approaches are prone to failures due to (i) difficulties in designing reliable word-to-symbol segmentation module that can robustly work in presence of degraded (cut/fused) images and (ii) converting the outputs of the classifiers to a valid sequence of Unicodes. In this paper, we propose a formulation, where the expectations on these two modules is minimized, and the harder recognition task is modelled as learning of an appropriate sequence to sequence translation scheme. We thus formulate the recognition as a direct transcription problem. Given many examples of feature sequences and their corresponding Unicode representations, our objective is to learn a mapping which can convert a word directly into a Unicode sequence. This formulation has multiple practical advantages: (i) This reduces the number of classes significantly for the Indian scripts. (ii) It removes the need for a reliable word-to-symbol segmentation. (ii) It does not require strong annotation of symbols to design the classifiers, and (iii) It directly generates a valid sequence of Unicodes. We test our method on more than 6000 pages of printed Devanagari documents from multiple sources. Our method consistently outperforms other state of the art implementations.
Naveen Sankaran, Aman Neelappa, C. V. Jawahar
ICDAR3
2013 Document Specific Sparse Coding for Word Retrieval
abstract
Bag of words (BoW) based retrieval is an efficient method to compare the visual similarity between two images. Recognition free methods based on BoW have shown to outperform OCR based methods. We further improve the performance by defining a document specific sparse coding scheme for representing visual words (interest points) in document images. Our method is motivated by the successful use of sparsity in signal representation by exploiting the neighbourhood properties. In addition to providing insights into the design of the coding scheme, we also verify the method on two data sets and compare with the recent methods. We have also developed text query based search solution, and we report performance comparable to image based search.
Ravi Shekhar, C. V. Jawahar
ICDAR2
2012 Robust Recognition of Degraded Documents Using Character N-Grams
abstract
In this paper we present a novel recognition approach that results in a 15% decrease in word error rate on heavily degraded Indian language document images. OCRs have considerably good performance on good quality documents, but fail easily in presence of degradations. Also, classical OCR approaches perform poorly over complex scripts such as those for Indian languages. We address these issues by proposing to recognize character n-gram images, which are basically groupings of consecutive character/component segments. Our approach is unique, since we use the character n-grams as a primitive for recognition rather than for post processing. By exploiting the additional context present in the character n-gram images, we enable better disambiguation between confusing characters in the recognition phase. The labels obtained from recognizing the constituent n-grams are then fused to obtain a label for the word that emitted them. Our method is inherently robust to degradations such as cuts and merges which are common in digital libraries of scanned documents. We also present a reliable and scalable scheme for recognizing character n-gram images. Tests on English and Malayalam document images show considerable improvement in recognition in the case of heavily degraded documents.
Shrey Dutta, Naveen Sankaran, K. Pramod Sankar, C. V. Jawahar
Document Analysis Systems4
2012 Word Image Retrieval Using Bag of Visual Words
abstract
This paper presents a Bag of Visual Words (BoVW) based approach to retrieve similar word images from a large database, efficiently and accurately. We show that a text retrieval system can be adapted to build a word image retrieval solution. This helps in achieving scalability. We demonstrate the method on more than 1 Million word images with a sub-second retrieval time. We validate the method on four Indian languages, and report a mean average precision of more than 0.75. We represent the word images as histogram of visual words present in the image. Visual words are quantized representation of local regions, and for this work, SIFT descriptors at interest points are used as feature vectors. To address the lack of spatial structure in the BoVW representation, we re-rank the retrieved list. This significantly improves the performance.
Ravi Shekhar, C. V. Jawahar
Document Analysis Systems2
2012 Video retrieval by mimicking poses
abstract
We describe a method for real time video retrieval where the task is to match the 2D human pose of a query. A user can form a query by (i) interactively controlling a stickman on a web based GUI, (ii) uploading an image of the desired pose, or (iii) using the Kinect and acting out the query himself. The method is scalable and is applied to a dataset of 18 films totaling more than three million frames. The real time performance is achieved by searching for approximate nearest neighbors to the query using a random forest of K-D trees. Apart from the query modalities, we introduce two other areas of novelty. First, we show that pose retrieval can proceed using a low dimensional representation. Second, we show that the precision of the results can be improved substantially by combining the outputs of independent human pose estimation algorithms. The performance of the system is assessed quantitatively over a range of pose queries.
Nataraj Jammalamadaka, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari, C. V. Jawahar
ICMR5
2011 LSH based outlier detection and its application in distributed setting
abstract
In this paper, we give an approximate algorithm for distance based outlier detection using Locality Sensitive Hashing (LSH) technique. We propose an algorithm for the centralized case wherein the entire dataset is locally available for processing. However, in case of very large datasets collected from various input sources, often the data is distributed across the network. Accordingly, we show that our algorithm can be effectively extended to a constant round protocol with low communication costs, in a distributed setting with horizontal partitioning.
Madhuchand Rushi Pillutla, Nisarg Raval, Piyush Bansal, K. Srinathan 0001, C. V. Jawahar
CIKM5
2011 BLSTM Neural Network Based Word Retrieval for Hindi Documents
abstract
Retrieval from Hindi document image collections is a challenging task. This is partly due to the complexity of the script, which has more than 800 unique ligatures. In addition, segmentation and recognition of individual characters often becomes difficult due to the writing style as well as degradations in the print. For these reasons, robust OCRs are non existent for Hindi. Therefore, Hindi document repositories are not amenable to indexing and retrieval. In this paper, we propose a scheme for retrieving relevant Hindi documents in response to a query word. This approach uses BLSTM neural networks. Designed to take contextual information into account, these networks can handle word images that can not be robustly segmented into individual characters. By zoning the Hindi words, we simplify the problem and obtain high retrieval rates. Our simplification suits the retrieval problem, while it does not apply to recognition. Our scalable retrieval scheme avoids explicit recognition of characters. An experimental evaluation on a dataset of word images gathered from two complete books demonstrates good accuracy even in the presence of printing variations and degradations. The performance is compared with baseline methods.
Raman Jain, Volkmar Frinken, C. V. Jawahar, R. Manmatha
ICDAR3
2011 An MRF Model for Binarization of Natural Scene Text
abstract
Inspired by the success of MRF models for solving object segmentation problems, we formulate the binarization problem in this framework. We represent the pixels in a document image as random variables in an MRF, and introduce a new energy (or cost) function on these variables. Each variable takes a foreground or background label, and the quality of the binarization (or labelling) is determined by the value of the energy function. We minimize the energy function, i.e. find the optimal binarization, using an iterative graph cut scheme. Our model is robust to variations in foreground and background colours as we use a Gaussian Mixture Model in the energy function. In addition, our algorithm is efficient to compute, and adapts to a variety of document images. We show results on word images from the challenging ICDAR 2003 dataset, and compare our performance with previously reported methods. Our approach shows significant improvement in pixel level accuracy as well as OCR accuracy.
Anand Mishra 0001, Karteek Alahari, C. V. Jawahar
ICDAR3
2011 Character n-Gram Spotting in Document Images
abstract
In this paper, we present a novel approach to search and retrieve from document image collections, without explicit recognition. Existing recognition-free approaches such as word-spotting cannot scale to arbitrarily large vocabulary and document image collections. In this paper we put forth a framework that overcomes three issues of word-spotting: i) retrieving word images not labeled during indexing, ii) allow for query and retrieval of morphological variations of words and iii) scale the retrieval to large collections. We propose a character n-gram spotting framework, where word-images are considered as a bag of visual n-grams. The character n-grams are represented in a visual-feature space and indexed for quick retrieval. In the retrieval phase, the query word is expanded to its constituent n-grams, which are used to query the previously built index. A ranking mechanism is proposed that combines the retrieval results from the multiple lists corresponding to each n-gram. The approach is demonstrated on a size-able collection of English and Malayalam books. With a mean AP of 0.64, the performance of the retrieval system was found to be very promising.
M. Sudha Praveen, K. Pramod Sankar, C. V. Jawahar
ICDAR3
2010 Towards more effective distance functions for word image matching
abstract
Matching word images has many applications in document recognition and retrieval systems. Dynamic Time Warping (DTW) is popularly used to estimate the similarity between word images. Word images are represented as sequences of feature vectors, and the cost associated with dynamic programming based alignment is considered as the dissimilarity between them. However, such approaches are computationally costly when compared to fixed length matching schemes. In this paper, we explore systematic methods for identifying appropriate distance metrics for a given database or language. This is achieved by learning query specific distance functions which can be computed online efficiently. We show that a weighted Euclidean distance can outperform DTW for matching word images. This class of distance functions are also ideal for scalability and large scale matching. Our results are validated with mean Average Precision (mAP) on a fully annotated data set of 160K word images. We then show that the learnt distance functions can even be extended to a new database to obtain accurate retrieval.
Raman Jain, C. V. Jawahar
Document Analysis Systems2
2010 A post-processing scheme for malayalam using statistical sub-character language models
abstract
Most of the Indian scripts do not have any robust commercial OCRs. Many of the laboratory prototypes report reasonable results at recognition/classification stage. However, word level accuracies are still poor. It is well known that word accuracy decreases as the number of characters in a word increase. For Malayalam, the average number of characters in a word is almost twice that of English. Moreover, the number of words required to cover 80% of the Malayalam language is more than forty times that of other Indian languages such as Hindi. Hence a direct dictionary based post-processing scheme is not suitable for Malayalam.
Karthika Mohan, C. V. Jawahar
Document Analysis Systems2
2010 Nearest neighbor based collection OCR
abstract
Conventional optical character recognition (OCR) systems operate on individual characters and words, and do not normally exploit document or collection context. We describe a Collection OCR which takes advantage of the fact that multiple examples of the same word (often in the same font) may occur in a document or collection. The idea here is that an OCR or a reCAPTCHA like process generates a partial set of recognized words. In the second stage, a nearest neighbor algorithm compares the remaining word-images to those already recognized and propagates labels from the nearest neighbors. It is shown that by using an approximate fast nearest neighbor algorithm based on Hierarchical K-Means (HKM), we can do this accurately and efficiently. It is also shown that profile based features perform much better than SIFT and Pyramid Histogram of Gradient (PHOG) features. We believe that this is because profile features are more robust to word degradations (common in our documents). This approach is applied to a collection of Telugu books - a language for which no commercial OCR exists. We show from a selection of 33 Telugu books that starting with OCR labels for only 30% of the collection we can recognize the remaining 70% of the words in the collection with 70% accuracy using this approach. Since the approach makes no language specific assumptions, it should be applicable to a large number of languages. In particular we are interested in its applicability to Indic languages and scripts.
K. Pramod Sankar, C. V. Jawahar, R. Manmatha
Document Analysis Systems2
2009 Robust Recognition of Documents by Fusing Results of Word Clusters
abstract
The word error rate of any optical character recognition system (OCR) is usually substantially below its component or character error rate. This is especially true of Indic languages in which a word consists of many components. Current OCRs recognize each character or word separately and do not take advantage of document level constraints. We propose a document level OCR which incorporates information from the entire document to reduce word error rates. Word images are first clustered using a locality sensitive hashing technique. Individual words are then recognized using a (regular) OCR. The OCR outputs of word images in a cluster are then corrected probabilistically by comparing with the OCR outputs of other members of the same cluster. The approach may be applied to improve the accuracy of any OCR run on documents in any language. In particular, we demonstrate it for Telugu, where the use of language models for post-processing is not promising. We show a relative improvement of 28% for long words and 12% for all words which appear at least twice in the corpus.
Venkat Rasagna, Anand Kumar 0001, C. V. Jawahar, R. Manmatha
ICDAR3
2008 Super-Resolution of Text Images Using Edge-Directed Tangent Field
abstract
This paper presents an edge-directed super-resolution algorithm for document images without using any training set. This technique creates an image with smooth regions in both the foreground and the background, while allowing sharp discontinuities across and smoothness along the edges. Our method preserves sharp corners in text images by using the local edge direction, which is computed first by evaluating the gradient field and then taking its tangent. Super-resolution of document images is characterized by bimodality, smoothness along the edges as well as subsampling consistency. These characteristics are enforced in a Markov random field (MRF) framework by defining an appropriate energy function. In our method, subsampling of super-resolution image will return the original low-resolution one, proving the correctness of the method. The super-resolution image, is generated by iteratively reducing this energy function. Experimental results on a variety of input images, demonstrate the effectiveness of our method for document image super-resolution.
Jyotirmoy Banerjee, C. V. Jawahar
Document Analysis Systems2
2007 Content-level Annotation of Large Collection of Printed Document Images
abstract
A large annotated corpus is critical to the development of robust optical character recognizers (OCRs). However, creation of annotated corpora is a tedious task. It is laborious, especially when the annotation is at the character level. In this paper, we propose an efficient hierarchical approach for annotation of large collection of printed document images. We align document images with independently keyed-in text. The method is model-driven and is intended to annotate large collection of documents, scanned in three different resolutions, at character level. We employ an XML representation for storage of the annotation information. APIs are provided for access at content level for easy use in training and evaluation of OCRs and other document understanding tasks.
Anand Kumar 0001, C. V. Jawahar
ICDAR2
2007 On Segmentation of Documents in Complex Scripts
abstract
Document image segmentation algorithms primarily aim at separating text and graphics in presence of complex layouts. However, for many non-Latin scripts, segmentation becomes a challenge due to the characteristics of the script. In this paper, we empirically demonstrate that successful algorithms for Latin scripts may not be very effective for Indic and complex scripts. We explain this based on the differences in the spatial distribution of symbols in the scripts. We argue that the visual information used for segmentation needs to be enhanced with other information like script models for accurate results.
K. S. Sesh Kumar, C. V. Jawahar
ICDAR3
2007 On Using Classical Poetry Structure for Indian Language Post-Processing
abstract
Post-processors are critical to the performance of language recognizers like OCRs, speech recognizers, etc. Dictionary-based post-processing commonly employ either an algorithmic approach or a statistical approach. Other linguistic features are not exploited for this purpose. The language analysis is also largely limited to the prose form. This paper proposes a framework to use the rich metric and formal structure of classical poetic forms in Indian languages for post-processing a recognizer like an OCR engine. We show that the structure present in the form of the vrtta and prasa can be efficiently used to disambiguate some cases that may be difficult for an OCR. The approach is efficient, and complementary to other post-processing approaches and can be used in conjunction with them.
Anoop M. Namboodiri, P. J. Narayanan, C. V. Jawahar
ICDAR3
2006 Retrieval from Document Image Collections
A. Balasubramanian, Million Meshesha, C. V. Jawahar
Document Analysis Systems3
2006 A Semi-automatic Adaptive OCR for Digital Libraries
Sachin Rawat, K. S. Sesh Kumar, Million Meshesha, Indraneel Deb Sikdar, A. Balasubramanian, C. V. Jawahar
Document Analysis Systems6
2006 Digitizing a Million Books: Challenges for Document Analysis
K. Pramod Sankar, Vamshi Ambati, Lakshmi Pratha, C. V. Jawahar
Document Analysis Systems4
2005 Discriminant Substrokes for Online Handwriting Recognition
abstract
A discriminant-based framework for automatic recognition of online handwriting data is presented in this paper. We identify the substrokes that are more useful in discriminating between two online strokes. A similarity/dissimilarity score is computed based on the discriminatory potential of various parts of the stroke for the classification task. The discriminatory potential is then converted to the relative importance of the substroke. Experimental verification on online data such as numerals, characters supports our claims. We achieve an average reduction of 41% in the classification error rate on many test sets of similar character pairs.
Karteek Alahari, Satya Lahari Putrevu, C. V. Jawahar
ICDAR3
2005 Configurable Hybrid Architectures for Character Recognition Applications
abstract
Character recognition is a multiclass problem with typically large number of classes. Hierarchical classifiers are found to be suitable for such classification problems due to their low space and time complexities. However, most hierarchical classifiers employ similar classifiers at different levels of the hierarchy. This may not be ideal, as the complexity of classification is not same at all the stages. In this paper, we propose an algorithm to build a hybrid hierarchical classifier by choosing the classifiers of appropriate complexity at each level of the hierarchy. We demonstrate that hierarchical combination of complex classifiers improves the classification performance significantly.
M. N. S. S. K. Pavan Kumar, C. V. Jawahar
ICDAR2
2005 Recognition of Printed Amharic Documents
abstract
In Africa, there are a number of languages with their own indigenous scripts. This paper presents an OCR for Amharic scripts. Amharic is the official and working language of Ethiopia. This is possibly the first attempt towards the development of an OCR system for Amharic. Research in the recognition of Amharic script faces major challenges due to (i) the use of more than 300 characters in writing and (ii) existence of a large set of visually similar characters. In this paper, we propose a two-stage feature extraction scheme using PCA and LDA, followed by a decision DAG classifier with SVMs as the nodes. Recognition results are presented to demonstrate the performance on the various printing variations (fonts, styles and sizes) and real-life degraded documents such as books, magazines and newspapers.
Million Meshesha, C. V. Jawahar
ICDAR2
2003 A Bilingual OCR for Hindi-Telugu Documents and its Applications
abstract
This paper describes the character recognition process from printed documents containing Hindi and Telugu text. Hindi and Telugu are among the most popular languages in India. The bilingual recognizer is based on Principal Component Analysis followed by support vector classification. This attains an overall accuracy of approximately 96.7%. Extensive experimentation is carried out on an independent test set of approximately 200000 characters. Applications based on this OCR are sketched.
C. V. Jawahar, M. N. S. S. K. Pavan Kumar, S. S. Ravi Kiran
ICDAR1