VLDB 2026 Research / reviewers in the wild / expert
Venu Govindaraju
dblp:g/VenuGovindaraju · also Venugopal Govindaraju
· DBLP profile ↗
215ranked-venue papers
12as first author
27since 2021 · last 2026
0000-0002-5318-7409ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 164 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 5 first-author · 19 since 2021Databases, data management, data science and information retrieval · 64 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 2 since 2021Security and privacy · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Forget Less by Learning from Parents Through Hierarchical RelationshipsabstractCustom Diffusion Models (CDMs) offer impressive capabilities for personalization in generative modeling, yet they remain vulnerable to catastrophic forgetting when learning new concepts sequentially. Existing approaches primarily focus on minimizing interference between concepts, often neglecting the potential for positive inter-concept interactions. In this work, we present Forget Less by Learning from Parents (FLLP), a novel framework that introduces a parent-child inter-concept learning mechanism in hyperbolic space to mitigate forgetting. By embedding concept representations within a Lorentzian manifold, naturally suited to modeling tree-like hierarchies, we define parent-child relationships in which previously learned concepts serve as guidance for adapting to new ones. Our method not only preserves prior knowledge but also supports continual integration of new concepts. We validate FLLP on three public datasets and one synthetic benchmark, showing consistent improvements in both robustness and generalization. Arjun Ramesh Kaushik, Naresh Kumar Devulapally, Vishnu Suresh Lokhande, Nalini K. Ratha, Venu Govindaraju |
AAAI | 5 |
| 2026 | Forget Less by Learning Together through Concept ConsolidationabstractCustom Diffusion Models (CDMs) have gained significant attention due to their remarkable ability to personalize generative processes. However, existing CDMs suffer from catastrophic forgetting when continuously learning new concepts. Most prior works attempt to mitigate this issue under the sequential learning setting with a fixed order of concept inflow and neglect inter-concept interactions. In this paper, we propose a novel framework -Forget Less by Learning Together (FL2T) - that enables concurrent and order-agnostic concept learning while addressing catastrophic forgetting. Specifically, we introduce a set-invariant inter-concept learning module where proxies guide feature selection across concepts, facilitating improved knowledge retention and transfer. By leveraging inter-concept guidance, our approach preserves old concepts while efficiently incorporating new ones. Extensive experiments, across three datasets, demonstrates that our method significantly improves concept retention and mitigates catastrophic forgetting, highlighting the effectiveness of inter-concept catalytic behavior in incremental concept learning of ten tasks with at least 2% gain on average CLIP Image Alignment scores. Arjun Ramesh Kaushik, Naresh Kumar Devulapally, Vishnu Suresh Lokhande, Nalini K. Ratha, Venu Govindaraju |
WACV | 5 |
| 2026 | Learning Action Hierarchies via Hybrid Geometric DiffusionabstractTemporal action segmentation is a critical task in video understanding, where the goal is to assign action labels to each frame in a video. While recent advances leverage iterative refinement-based strategies, they fail to explicitly utilize the hierarchical nature of human actions. In this work, we propose HybridTAS - a novel framework that incorporates a hybrid of Euclidean and hyperbolic geometries into the denoising process of diffusion models to exploit the hierarchical structure of actions. Hyperbolic geometry naturally provides tree-like relationships between embeddings, enabling us to guide the action label denoising process in a coarse-to-fine manner: higher diffusion timesteps are influenced by abstract, high-level action categories (root nodes), while lower timesteps are refined using fine-grained action classes (leaf nodes). Extensive experiments on three benchmark datasets, GTEA, 50Salads, and Breakfast, demonstrate that our method achieves state-of-the-art performance, validating the effectiveness of hyperbolic-guided denoising for the temporal action segmentation task. Arjun Ramesh Kaushik, Nalini K. Ratha, Venu Govindaraju |
WACV | 3 |
| 2026 | AF-CANet: Augmentation-free, curriculum-guided attention UNet framework for few-shot document layout analysis
Hadia Showkat Kawoosa, Sahana Rangasrinivasan, Srirangaraj Setlur, Puneet Goyal, Venu Govindaraju |
Pattern Recognit. Lett. | 5 |
| 2025 | From Scribbles to Text: A Novel Transformer-Based Recognition Model for Child Handwriting
Sahana Rangasrinivasan, M. S. Sumi Suresh, Srirangaraj Setlur, Bharat Jayaraman, Venu Govindaraju |
ICDAR (1) | 5 |
| 2025 | AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social RobotabstractThe social robot’s open API allows users to customize open-domain interactions. However, it remains inaccessible to those without programming experience. We introduce AutoMisty, the first LLM-powered multi-agent framework that converts natural-language commands into executable Misty robot code by decomposing high-level instructions, generating sub-task code, and integrating everything into a deployable program. Each agent employs a two-layer optimization mechanism: first, a self-reflective loop that instantly validates and automatically executes the generated code, regenerating whenever errors emerge; second, human review for refinement and final approval, ensuring alignment with user preferences and preventing error propagation. To evaluate AutoMisty’s effectiveness, we designed a benchmark task set spanning four levels of complexity and conducted experiments in a real Misty robot environment. Extensive evaluations demonstrate that AutoMisty not only consistently generates high-quality code but also enables precise code control, significantly outperforming direct reasoning with ChatGPT-4o and ChatGPT-o1. All code, optimized APIs, and experimental videos will be publicly released through the webpage: AutoMisty. Lu Dong 0004, Sahana Rangasrinivasan, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju |
IROS | 6 |
| 2025 | SCOT: Self-Supervised Contrastive Pretraining for Zero-Shot Compositional RetrievalabstractCompositional image retrieval (CIR) is a multimodal learning task where a model combines a query image with a user-provided text modification to retrieve a target image. CIR finds applications in a variety of domains including product retrieval (e-commerce) and web search. Existing methods primarily focus on fully-supervised learning, wherein models are trained on datasets of labeled triplets such as FashionIQ and CIRR. This poses two significant challenges: (i) curating such triplet datasets is labor intensive; and (ii) models lack generalization to un-seen objects and domains. In this work, we propose SCOT (Self-supervised COmpositional Training), a novel zero-shot compositional pretraining strategy that combines existing large image-text pair datasets with the generative capabilities of large language models to contrastively train an embedding composition network. Specifically, we show that the text embedding from a large-scale contrastively-pretrained vision-language model can be utilized as proxy target supervision during compositional pretraining, replacing the target image embedding. In zero-shot settings, this strategy surpasses SOTA zero-shot compositional re-trieval methods as well as many fully-supervised methods on standard benchmarks such as FashionIQ and CIRR. Our code and models are available at https://github.com/yahoo/SCOT. Bhavin Jawade, João V. B. Soares, Kapil Thadani, Deen Dayal Mohan, Amir Erfan Eshratifar, Benjamin Culpepper, Paloma de Juan, Srirangaraj Setlur, Venu Govindaraju |
WACV | 9 |
| 2024 | GestSpoof: Gesture Based Spatio-Temporal Representation Learning for Robust Fingerprint Presentation Attack DetectionabstractFingerprint spoof attacks represent one of the most prevalent forms of biometric presentation attacks. While significant progress has been made in framing fingerprint spoof detection as a general image classification problem, limited attention has been given to treating it as a temporal learning problem. The distinctions in the elastic properties between authentic and synthetically created counterfeit fingerprints can be more accurately captured under motion-induced gestures during acquisition. In this study, we introduce a novel method for detecting fake fingerprints by deliberately introducing distortions through sliding and twisting motions during acquisition. As widely used spoof datasets such as those from LivDet 2009 to 2021 or MSU FPAD lack the temporal information essential for this investigation, we assembled a new dataset focused on distortion-based fake and real fingerprints, encompassing various types of spoof materials and diverse distortions. This gesture-equipped dataset comprises more than 3680 videos gathered from 184 unique fingers. Additionally, we present a novel spatial-temporal multi-modal network for detecting fingerprint spoofs using intentional-distortion. Our proposed approach yields significantly improved results compared to traditional static classification-based methods for spoof detection, across various metrics and for both known and unknown (generalization) scenarios, thereby highlighting the substantial impact that introducing gestures can have on enhancing fingerprint spoof detection. The dataset can be downloaded from here: https://www.buffalo.edu/cubs/research/datasets/gestspoof-dataset.html Bhavin Jawade, Shreeram Subramanya, Atharv Dabhade, Srirangaraj Setlur, Venu Govindaraju |
FG | 5 |
| 2024 | A Comparative Study of Video-Based Human Representations for American Sign Language Alphabet GenerationabstractSign language is a complex visual language, and automatic interpretations of sign language can facilitate communication involving deaf individuals. As one of the essential components of sign language, fingerspelling connects the natural spoken languages to the sign language and expands the scale of sign language vocabulary. In practice, it is challenging to analyze fingerspelling alphabets due to their signing speed and small motion range. The usage of synthetic data has the potential of further improving fingerspelling alphabets analysis at scale. In this paper, we evaluate how different video-based human representations perform in a framework for Alphabet Generation for American Sign Language (ASL). We tested three mainstream video-based human representations: two-stream inflated 3D ConvNet, 3D landmarks of body joints, and rotation matrices of body joints. We also evaluated the effect of different skeleton graphs and selected body joints. The generation process of ASL fingerspelling used a transformer-based Conditional Variational Autoencoder. To train the model, we collected ASL alphabet signing videos from 17 signers with dynamic alphabet signing. The generated alphabets were evaluated using automatic metrics of quality such as FID, and we also considered supervised metrics by recognizing the generated entries using Spatio-Temporal Graph Convolutional Networks. Our experiments show that using the rotation matrices of the upper body joints and the signing hand give the best results for the generation of ASL alphabet signing. Going forward, our goal is to produce articulated fingerspelling words by combining individual alphabets learned in this work. Lipisha Chaudhary, Lu Dong 0004, Srirangaraj Setlur, Venu Govindaraju, Ifeoma Nwogu |
FG | 5 |
| 2024 | Fine-Grained Engine Fault Sound Event Detection Using Multimodal SignalsabstractSound event detection (SED) is an active area of audio research that aims to detect the temporal occurrence of sounds. In this paper, we apply SED to engine fault detection by introducing a multimodal SED framework that detects fine-grained engine faults of automobile engines using audio and accelerometer-recorded vibration. We first introduce the problem of engine fault SED on a dataset collected from a large variety of vehicles with expertly-labeled engine fault sound events. Next, we propose a SED model to temporally detect ten fine-grained engine faults that occur within vehicle engines and further explore a pretraining strategy using a large-scale weakly-labeled engine fault dataset. Through multiple evaluations, we show our proposed framework is able to effectively detect engine fault sound events. Finally, we investigate the interaction and characteristics of each modality and show that fusing features from audio and vibration improves overall engine fault SED capabilities. Dennis Fedorishin, Livio Forte, Philip Schneider, Srirangaraj Setlur, Venu Govindaraju |
ICASSP | 5 |
| 2024 | Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and VideosabstractA "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although match cuts are often visual, certain match cuts involve the fluid transition of audio, where sounds from different sources merge into one indistinguishable transition between two shots. In this paper, we explore the ability to automatically find and create "audio match cuts" within videos and movies. We create a self-supervised audio representation for audio match cutting and develop a coarse-to-fine audio match pipeline that recommends matching shots and creates the blended audio. We further annotate a dataset for the proposed audio match cut task and compare the ability of multiple audio representations to find audio match cut candidates. Finally, we evaluate multiple methods to blend two matching audio candidates with the goal of creating a smooth transition. Project page and examples are available at: https://denfed.github.io/audiomatchcut/ Dennis Fedorishin, Lie Lu, Srirangaraj Setlur, Venu Govindaraju |
ICASSP | 4 |
| 2024 | CHART-Info 2024: A Dataset for Chart Analysis and Recognition
Kenny Davila, Rupak Lazarus, Nicole Rodríguez Alcántara, Srirangaraj Setlur, Venu Govindaraju, Ajoy Mondal, C. V. Jawahar |
ICPR (19) | 6 |
| 2024 | Maximizing Coverage over a Surveillance Region Using a Specific Number of Cameras
M. S. Sumi Suresh, Vivek Menon, Srirangaraj Setlur, Venu Govindaraju |
ICPR (22) | 4 |
| 2024 | ProxyFusion: Face Feature Aggregation Through Sparse ExpertsabstractFace feature fusion is indispensable for robust face recognition, particularly in scenarios involving long-range, low-resolution media (unconstrained environments) where not all frames or features are equally informative. Existing methods often rely on large intermediate feature maps or face metadata information, making them incompatible with legacy biometric template databases that store pre-computed features. Additionally, real-time inference and generalization to large probe sets remains challenging.
To address these limitations, we introduce a linear time O(N) proxy based sparse expert selection and pooling approach for context driven feature-set attention. Our approach is order invariant on the feature-set, generalizes to large sets, is compatible with legacy template stores, and utilizes significantly less parameters making it suitable real-time inference and edge use-cases. Through qualitative experiments, we demonstrate that ProxyFusion learns discriminative information for importance weighting of face features without relying on intermediate features. Quantitative evaluations on challenging low-resolution face verification datasets such as IARPA BTS3.1 and DroneSURF show the superiority of ProxyFusion in unconstrained long-range face recognition setting.
Our code and pretrained models are available at: https://github.com/bhavinjawade/ProxyFusion Bhavin Jawade, Alexander Stone, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju |
NeurIPS | 6 |
| 2023 | CoNAN: Conditional Neural Aggregation Network For Unconstrained Face Feature FusionabstractFace recognition from image sets acquired under unregulated and uncontrolled settings, such as at large distances, low resolutions, varying viewpoints, illumination, pose, and atmospheric conditions, is challenging. Face feature aggregation, which involves aggregating a set of N feature representations present in a template into a single global representation, plays a pivotal role in such recognition systems. Existing works in traditional face feature aggregation either utilize metadata or high-dimensional intermediate feature representations to estimate feature quality for aggregation. However, generating high-quality metadata or style information is not feasible for extremely low-resolution faces captured in long-range and high altitude settings. To overcome these limitations, we propose a feature distribution conditioning approach called CoNAN for template aggregation. Specifically, our method aims to learn a context vector conditioned over the distribution information of the incoming feature set, which is utilized to weigh the features based on their estimated informativeness. The proposed method produces state-of-the-art results on long-range unconstrained face recognition datasets such as BTS, and DroneSURF, validating the advantages of such an aggregation strategy. Bhavin Jawade, Deen Dayal Mohan, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
IJCB | 5 |
| 2023 | RealCQA: Scientific Chart Question Answering as a Test-Bed for First-Order Logic
Saleem Ahmed, Bhavin Jawade, Shubham Pandey, Srirangaraj Setlur, Venu Govindaraju |
ICDAR (3) | 5 |
| 2023 | SpaDen: Sparse and Dense Keypoint Estimation for Real-World Chart Understanding
Saleem Ahmed, Pengyu Yan, David S. Doermann, Srirangaraj Setlur, Venu Govindaraju |
ICDAR (2) | 5 |
| 2023 | Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source LocalizationabstractLearning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to localize the source of the sound. In a video, oftentimes, the objects exhibiting movement are the ones generating the sound. In this work, we capture this characteristic by modeling the optical flow in a video as a prior to better aid in localizing the sound source. We further demonstrate that the addition of flow-based attention substantially improves visual sound source localization. Finally, we benchmark our method on standard sound source localization datasets and achieve state-of-the-art performance on the Soundnet Flickr and VGG Sound Source datasets. Code: https://github.com/denfed/heartheflow. Dennis Fedorishin, Deen Dayal Mohan, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju |
WACV | 5 |
| 2023 | NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal EmbeddingsabstractCross-modal retrieval is a fundamental vision-language task with a broad range of practical applications. Text-to-image matching is the most common form of cross-modal retrieval where, given a large database of images and a textual query, the task is to retrieve the most relevant set of images. Existing methods utilize dual encoders with an attention mechanism and a ranking loss for learning embeddings that can be used for retrieval based on cosine similarity. Despite the fact that these methods attempt to perform semantic alignment across visual regions and textual words using tailored attention mechanisms, there is no explicit supervision from the training objective to enforce such alignment. To address this, we propose NAPReg, a novel regularization formulation that projects high-level semantic entities i.e Nouns into the embedding space as shared learnable proxies. We show that using such a formulation allows the attention mechanism to learn better word-region alignment while also utilizing region information from other samples to build a more generalized latent representation for semantic concepts. Experiments on three benchmark datasets i.e. MS-COCO, Flickr30k and Flickr8k demonstrate that our method achieves state-of-the-art results in cross-modal metric learning for text-image and image-text retrieval tasks. Code: https://github.com/bhavinjawade/NAPReq Bhavin Jawade, Deen Dayal Mohan, Naji Mohamed Ali, Srirangaraj Setlur, Venu Govindaraju |
WACV | 5 |
| 2022 | Attribute De-biased Vision Transformer (AD-ViT) for Long-Term Person Re-identificationabstractPerson re-identification (re-ID) aims to retrieve images of the same identity from a gallery of person images across cameras and viewpoints. However, most works in person re-ID assume a short-term setting characterized by invariance in appearance. In contrast, a high visual variance can be frequently seen in a long-term setting due to changes in apparel and accessories, which makes the task more challenging. Therefore, learning identity-specific features agnostic of temporally variant features is crucial for robust long-term person Re-ID. To this end, we propose an Attribute De-biased Vision Transformer (AD-ViT) to provide direct supervision to learn identity-specific features. Specifically, we produce attribute labels for person instances and utilize them to guide our model to focus on identity features through gradient reversal. Our experiments on two long-term re-ID datasets - LTCC and NKUP show that the proposed work consistently outperforms current state-of-the-art methods. Kyung Won Lee, Bhavin Jawade, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju |
AVSS | 5 |
| 2022 | RidgeBase: A Cross-Sensor Multi-Finger Contactless Fingerprint DatasetabstractContactless fingerprint matching using smartphone cameras can alleviate major challenges of traditional fingerprint systems including hygienic acquisition, portability and presentation attacks. However, development of practical and robust contactless fingerprint matching techniques is constrained by the limited availablity of large scale real-world datasets. To motivate further advances in contactless fingerprint matching across sensors, we introduce the RidgeBase benchmark dataset. RidgeBase consists of more than 15,000 contactless and contact-based fingerprint image pairs acquired from 88 individuals under different background and lighting conditions using two smartphone cameras and one flatbed contact sensor. Unlike existing datasets, RidgeBase is designed to promote research under different matching scenarios that include Single Finger Matching and Multi-Finger Matching for both contactless-to-contactless (CL2CL) and contact-to-contactless (C2CL) verification and identification. Furthermore, due to the high intra-sample variance in contactless fingerprints belonging to the same finger, we propose a set-based matching protocol inspired by the advances in facial recognition datasets. This protocol is specifically designed for pragmatic contactless fingerprint matching that can account for variances in focus, polarity and finger-angles. We report qualitative and quantitative baseline results for different protocols using a COTS fingerprint matcher (Verifinger) and a Deep CNN based approach on the RidgeBase dataset. The dataset can be downloaded here: https://www.buffalo.edu/cubs/research/datasets/ridgebase-benchmark-dataset.html Bhavin Jawade, Deen Dayal Mohan, Srirangaraj Setlur, Nalini K. Ratha, Venu Govindaraju |
IJCB | 5 |
| 2022 | Synthetic Data Generation for Semantic Segmentation of Lecture Videos
Kenny Davila, James Molina, Srirangaraj Setlur, Venu Govindaraju |
ICFHR | 5 |
| 2022 | ICPR 2022: Challenge on Harvesting Raw Tables from Infographics (CHART-Infographics)abstractThe outcomes of the third Challenge on HArvesting Raw Tables from Infographics (ICPR 2022 CHART-Infographics) are presented in this work. Recognizing charts is a difficult process which we divided into the following task: Chart Image Classification (Task 1), Text Detection and Recognition (Task 2), Text Role Classification (Task 3), Axis Analysis (Task 4), Legend Analysis (Task 5), Plot Element Detection and Classification (Task 6.a), Data Extraction (Task 6.b), and End-to-End Data Extraction (Task 7). We have provided a novel dataset for training reusing all available data from previous challenges, and we also provide a brand new testing dataset for the evaluation of submissions. Both datasets were constructed by manually annotating charts extracted from the Open Access section of the PubMed Central. A total of 9 teams registered out of which 5 submitted results for different tasks of the challenge. Many submissions are based on state-of-the-art methods from computer vision, but the final scores imply that more work will be required to solve the chart recognition problem. The data, annotation tools, and evaluation scripts have been publicly released for academic use. Kenny Davila, Saleem Ahmed, David A. Mendoza, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 6 |
| 2022 | Large-Scale Acoustic Automobile Fault Detection: Diagnosing Engines Through SoundabstractIn this paper we present AMPNet, an acoustic abnormality detection model deployed at ACV Auctions to automatically identify engine faults of vehicles listed on the ACV Auctions platform. We investigate the problem of engine fault detection and discuss our approach of deep-learning based audio classification on a large-scale automobile dataset collected at ACV Auctions. Specifically, we discuss our data collection pipeline and its challenges, dataset preprocessing and training procedures, and deployment of our trained models into a production setting. We perform empirical evaluations of AMPNet and demonstrate that our framework is able to successfully capture various engine anomalies agnostic of vehicle type. Finally we demonstrate the effectiveness and impact of AMPNet in the real world, specifically showing a 20.85% reduction in vehicle arbitrations on ACV Auctions' live auction platform. Dennis Fedorishin, Justas Birgiolas, Deen Dayal Mohan, Livio Forte, Philip Schneider, Srirangaraj Setlur, Venu Govindaraju |
KDD | 7 |
| 2021 | Bayesian Personalized-Wardrobe Model (BP-WM) for Long-Term Person Re-IdentificationabstractLong-term surveillance applications often involve having to re-identify individuals over several days. The task is made even more challenging due to changes in appearance features such as clothing over a longitudinal time-span of days or longer. In this paper, we propose a novel approach called Bayesian Personalized-Wardrobe Model (BPWM) for long-term person re-identification (re-ID) by employing a Bayesian Personalized Ranking (BPR) for clothing features extracted from video sequences. In contrast to previous long-term person re-ID works, we exploit the fact that people typically choose their attire based on their personal preferences and that knowing a person’s chosen wardrobe can be used as a soft-biometric to distinguish identities in the long-term. We evaluate the performance of our proposed BP-WM on the extended Indoor Long-term Re-identification Wardrobe (ILRW) dataset. Experimental results show that our method achieves state-of-the-art performance and that BP-WM can be used as a reliable soft-biometric for person re-identification. Kyung Won Lee, Nishant Sankaran, Deen Dayal Mohan, Kenny Davila, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
AVSS | 7 |
| 2021 | TADPool: Target Adaptive Pooling for Set Based Face RecognitionabstractA majority of the modern methods used for template aggregation of set-based face recognition systems rely on learning to quantify the quality of images present in a template. While focusing on weighting the feature embedding based on this quality factor, they have overlooked aggregation strategies that can adapt the template's features to the paired template involved in matching. In this paper, we explore the potential of such adaptive methods for feature aggregation. We propose a template feature aggregation strategy that tailors a template's image set to mirror the properties exhibited by the target template. The proposed method produces state-of-the-art results on standard unconstrained face recognition datasets such as IJB-A, IJB-C and YouTubeFaces, validating the advantages of such an aggregation strategy. Nishant Sankaran, Deen Dayal Mohan, Sergey Tulyakov, Srirangaraj Setlur, Venu Govindaraju |
FG | 5 |
| 2021 | Chart Mining: A Survey of Methods for Automated Chart AnalysisabstractCharts are useful communication tools for the presentation of data in a visually appealing format that facilitates comprehension. There have been many studies dedicated to chart mining, which refers to the process of automatic detection, extraction and analysis of charts to reproduce the tabular data that was originally used to create them. By allowing access to data which might not be available in other formats, chart mining facilitates the creation of many downstream applications. This paper presents a comprehensive survey of approaches across all components of the automated chart mining pipeline, such as (i) automated extraction of charts from documents; (ii) processing of multi-panel charts; (iii) automatic image classifiers to collect chart images at scale; (iv) automated extraction of data from each chart image, for popular chart types as well as selected specialized classes; (v) applications of chart mining; and (vi) datasets for training and evaluation, and the methods that were used to build them. Finally, we summarize the main trends found in the literature and provide pointers to areas for further research in chart mining. Kenny Davila, Srirangaraj Setlur, David S. Doermann, Bhargava Urala Kota, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Moving in the Right Direction: A Regularization for Deep Metric LearningabstractDeep metric learning leverages carefully designed sampling strategies and loss functions that aid in optimizing the generation of a discriminable embedding space. While effective sampling of pairs is critical for shaping the metric space during training, the relative interactions between pairs, and consequently the forces exerted on these pairs that direct their displacement in the embedding space can significantly impact the formation of well separated clusters. In this work, we identify a shortcoming of existing loss formulations which fail to consider more optimal directions of pair displacements as another criterion for optimization. We propose a novel direction regularization to explicitly account for the layout of sampled pairs and attempt to introduce orthogonality in the representations. The proposed regularization is easily integrated into existing loss functions providing considerable performance improvements. We experimentally validate our hypothesis on the Cars-196, CUB-200 and InShop datasets and outperform existing methods to yield state-of-the-art results on these datasets. Deen Dayal Mohan, Nishant Sankaran, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
CVPR | 5 |
| 2020 | Learning Guided Attention Masks for Facial Action Unit RecognitionabstractHumans have the innate ability to rapidly spot and react to a person's emotional response. For computers to be able to understand expressions in a similar way, the gap between the perception of expressions between the humans and computers needs to be minimized. Inspired by human visual fixations, we propose a guided attention mechanism that facilitates the network to `look' at the most important features of a face. Rather than imposing hard attention, we learn the attention maps from the intermediate representation for Action Units (AUs). We propose a joint attention learning and AU classification module with minimal increase in the network parameters. We demonstrate the efficiency of our approach on three standard datasets: BP4D, MMSE and DISFA and obtain state of the art and near state of the art results respectively. Nagashri N. Lakshminarayana, Srirangaraj Setlur, Venu Govindaraju |
FG | 3 |
| 2020 | Equation Attention Relationship Network (EARN) : A Geometric Deep Metric Framework for Learning Similar Math Expression EmbeddingabstractRepresentational Learning in the form of high dimensional embeddings have been used for multiple pattern recognition applications. There has been a significant interest in building embedding based systems for learning representations in the mathematical domain. At the same time, retrieval of structured information such as mathematical expressions is an important need for modern IR systems. In this work, our motivation is to introduce a robust framework for learning representations for similarity based retrieval of mathematical expressions. Given a query by example, the embedding can find the closest matching expression as a function of euclidean distance between them. We leverage recent advancements in image-based and graph-based deep learning algorithms to learn our similarity embeddings. We do this first, by using unimodal encoders in graph space and image space and then, a multi-modal combination of the same. To overcome the lack of training data, we force the networks to learn a deep metric using triplets generated with a heuristic scoring function. We also adopt a custom strategy for mining hard samples to train our neural networks. Our system produces rankings similar to those generated by the original scoring function, but using only a fraction of the time. Our results establish the viability of using such a multi-modal embedding for this task. Saleem Ahmed, Kenny Davila, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 4 |
| 2020 | Automated Whiteboard Lecture Video Summarization by Content Region Detection and RepresentationabstractLecture videos are rapidly becoming an invaluable source of information for students across the globe. Given the large number of online courses currently available, it is important to condense the information within these videos into a compact yet representative summary that can be used for search-based applications. We propose a framework to summarize whiteboard lecture videos by finding feature representations of detected handwritten content regions to determine unique content. We investigate multi-scale histogram of gradients and embeddings from deep metric learning for feature representation. We explicitly handle occluded, growing and disappearing handwritten content. Our method is capable of producing two kinds of lecture video summaries - the unique regions themselves or so-called key content and keyframes (which contain all unique content in a video segment). We use weighted spatio-temporal conflict minimization to segment the lecture and produce keyframes from detected regions and features. We evaluate both types of summaries and find that we obtain state-of-the-art peformance in terms of number of summary keyframes while our unique content recall and precision are comparable to state-of-the-art. Bhargava Urala Kota, Alexander Stone, Kenny Davila, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 5 |
| 2020 | Domain adaptive representation learning for facial action unit recognition
Nishant Sankaran, Deen Dayal Mohan, Nagashri N. Lakshminarayana, Srirangaraj Setlur, Venu Govindaraju |
Pattern Recognit. | 5 |
| 2019 | Tangent-V: Math Formula Image Search Using Line-of-Sight Graphs
Kenny Davila, Ritvik Joshi, Srirangaraj Setlur, Venu Govindaraju, Richard Zanibbi |
ECIR (1) | 4 |
| 2019 | Multimodal Deep Feature Aggregation for Facial Action Unit Recognition using Visible Images and Physiological SignalsabstractIn this paper we present a feature aggregation method to combine the information from the visible light domain and the physiological signals for predicting the 12 facial action units in the MMSE dataset. Although multimodal affect analysis has gained lot of attention, the utility of physiological signals in recognizing facial action units is relatively unexplored. In this paper we investigate if physiological signals such as Electro Dermal Activity (EDA), Respiration Rate and Pulse Rate can be used as metadata for action unit recognition. We exploit the effectiveness of deep learning methods to learn an optimal combined representation that is derived from the individual modalities. We obtained an improved performance on MMSE dataset further validating our claim. To the best of our knowledge this is the first study on facial action unit recognition using physiological signals. Nagashri N. Lakshminarayana, Nishant Sankaran, Srirangaraj Setlur, Venu Govindaraju |
FG | 4 |
| 2019 | Representation Learning Through Cross-Modality SupervisionabstractLearning robust representations for applications with multiple modalities of input can have a significant impact on its performance. Traditional representation learning methods rely on projecting the input modalities to a common subspace to maximize agreement amongst the modalities for a particular task. We propose a novel approach to representation learning that uses a latent representation decoder to reconstruct the target modality and thereby employs the target modality purely as a supervision signal for discovering correlations between the modalities. Through cross modality supervision, we demonstrate that the learnt representation is able to improve the performance of the task of facial action unit (AU) recognition when compared with the modality specific representations and even their fused counterparts. Our experiments on three AU recognition datasets - MMSE, BP4D and DISFA, show strong performance gains producing state-of-the-art results in spite of the absence of a modality. Nishant Sankaran, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju, Dennis Fedorishin |
FG | 4 |
| 2019 | ICDAR 2019 Competition on Harvesting Raw Tables from Infographics (CHART-Infographics)abstractThis work summarizes the results of the first Competition on Harvesting Raw Tables from Infographics (ICDAR 2019 CHART-Infographics). The complex process of automatic chart recognition is divided into multiple tasks for the purpose of this competition, including Chart Image Classification (Task 1), Text Detection and Recognition (Task 2), Text Role Classification (Task 3), Axis Analysis (Task 4), Legend Analysis (Task 5), Plot Element Detection and Classification (Task 6.a), Data Extraction (Task 6.b), and End-to-End Data Extraction (Task 7). We provided a large synthetic training set and evaluated submitted systems using newly proposed metrics on both synthetic charts and manually-annotated real charts taken from scientific literature. A total of 8 groups registered for the competition out of which 5 submitted results for tasks 1-5. The results show that some tasks can be performed highly accurately on synthetic data, but all systems did not perform as well on real world charts. The data, annotation tools, and evaluation scripts have been publicly released for academic use. Kenny Davila, Bhargava Urala Kota, Srirangaraj Setlur, Venu Govindaraju, Chris Tensmeyer, Ritwick Chaudhry |
ICDAR | 4 |
| 2019 | Content Extraction from Lecture Video via Speaker Action Classification Based on Pose InformationabstractOnline lecture videos are increasingly important e-learning materials for students. Automated content extraction from lecture videos facilitates information retrieval applications that improve access to the lecture material. A significant number of lecture videos include the speaker in the image. Speakers perform various semantically meaningful actions during the process of teaching. Among all the movements of the speaker, key actions such as writing or erasing potentially indicate important features directly related to the lecture content. In this paper, we present a methodology for lecture video content extraction using the speaker actions. Each lecture video is divided into small temporal units called action segments. Using a pose estimator, body and hands skeleton data are extracted and used to compute motion-based features describing each action segment. Then, the dominant speaker action of each of these segments is classified using Random forests and the motion-based features. With the temporal and spatial range of these actions, we implement an alternative way to draw key-frames of handwritten content from the video. In addition, for our fixed camera videos, we also use the skeleton data to compute a mask of the speaker writing locations for the subtraction of the background noise from the binarized key-frames. Our method has been tested on a publicly available lecture video dataset, and it shows reasonable recall and precision results, with a very good compression ratio which is better than previous methods based on content analysis. Kenny Davila, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 4 |
| 2019 | Generalized framework for summarization of fixed-camera lecture videos by detecting and binarizing handwritten content
Bhargava Urala Kota, Kenny Davila, Alexander Stone, Srirangaraj Setlur, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 5 |
| 2019 | Learning deep features for online person tracking using non-overlapping cameras: A survey
Neeti Narayan, Nishant Sankaran, Srirangaraj Setlur, Venu Govindaraju |
Image Vis. Comput. | 4 |
| 2018 | Wardrobe Model for Long Term Re-identification and Appearance PredictionabstractLong-term surveillance applications often involve having to re-identify individuals over several days or weeks. The task is made even more challenging with the lack of sufficient visibility of the subjects faces. We address this problem by modeling the wardrobe of individuals using discriminative features and labels extracted from their clothing information from video sequences. In contrast to previous person re-id works, we exploit that people typically own a limited amount of clothing and that knowing a person's wardrobe can be used as a soft-biometric to distinguish identities. We a) present a new dataset consisting of more than 70,000 images recorded over 30 days of 25 identities; b) model clothing features using CNNs that minimize intra-garments variations while maximizing inter-garments differences; and c) build a reference wardrobe model that captures each persons set of clothes that can be used for re-id. We show that these models open new perspectives to long-term person re-id problem using clothing information. Kyung Won Lee, Nishant Sankaran, Srirangaraj Setlur, Nils Napp, Venu Govindaraju |
AVSS | 5 |
| 2018 | Knowledge Transfer Using Neural Network Based Approach for Handwritten Text RecognitionabstractThe goal of a writer adaptive handwriting recognition system is to build a model that improves the recognition of a generic recognition model for a specific author. In this work, we show how structural representation learned from a generic writer-independent handwriting recognition model can be customized to individual authors. Convolutional Neural Network has shown outstanding performance in learning image-based representation that was used for classification. Additionally, they have been used along with Recurrent Neural Network (RNN) or its variations like, LSTM and GRU layers to analyze and understand sequences in handwriting recognition, sentence analysis, voice recognition etc. In most cases, the CNNs serve as a feature extractor instead of low-level hand-designed features that were used previously for the above-mentioned classification tasks. We design a method to reuse weights from layers trained on the IAM offline handwritten dataset to compute mid-level image representation for text in the Washington and Moore dataset. We show that despite differences in the writing style, fonts across these datasets, the transferred representation is able to capture a spatio-temporal representation leading to significantly improved recognition results. We hypothesize that the performance is solely not dependent on the number of samples and the model is evaluated with varying amount of fine-tuning samples showing promising results backing the hypothesis. Rathin Radhakrishnan Nair, Nishant Sankaran, Bhargava Urala Kota, Sergey Tulyakov, Srirangaraj Setlur, Venu Govindaraju |
DAS | 6 |
| 2018 | Automated Detection of Handwritten Whiteboard Content in Lecture Videos for SummarizationabstractOnline lecture videos are a valuable resource for students across the world. The ability to find videos based on their content could make them even more useful. Methods for automatic extraction of this content reduce the amount of manual effort required to make indexing and retrieval of such videos possible. We adapt a deep learning based method for scene text detection, for the purpose of detection of handwritten text, math expressions and sketches in lecture videos. We detect handwritten elements on the whiteboard to generate a summary of all content over time in the lecture, while also dealing with occluded content due to motion of the lecturer. We train, test on the publicly available AccessMath lecture video dataset and evaluate our framework on the basis of number of summary frames, as well as recall and precision of all whiteboard content in the set of test lecture videos. We found that our method increases the precision of the state-of-the-art while there is potential to increase recall as well. We have added to the existing ground truth in the AccessMath dataset by providing timestamp-based, semantically meaningful bounding box annotations for the handwritten whiteboard content, which has been released. Bhargava Urala Kota, Kenny Davila, Alexander Stone, Srirangaraj Setlur, Venu Govindaraju |
ICFHR | 5 |
| 2018 | Special issue on deep learning for document analysis and recognition
Cheng-Lin Liu 0001, Gernot A. Fink, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 3 |
| 2017 | Score normalization in stratified biometric systemsabstractStratified biometric system can be defined as a system in which the subjects, their templates or matching scores can be separated into two or more categories, or strata, and the matching decisions can be made separately for each stratum. In this paper we investigate the properties of the strat-ifiedbiometric system and, in particular, possible strata creation strategies, score normalization and acceptance decisions, expected performance improvements due to stratification. We perform our experiments on face recognition matching scores from IARPA Janus CS2 dataset. Sergey Tulyakov, Nishant Sankaran, Srirangaraj Setlur, Venu Govindaraju |
IJCB | 4 |
| 2017 | Bayesian background models for keyword spotting in handwritten documents
Venu Govindaraju |
Pattern Recognit. | 2 |
| 2017 | Cognitive-Biometric Recognition From Language Usage: A Feasibility StudyabstractWe propose a novel cognitive biometrics modality based on written language-usage of an individual. This is a feasibility study using the Internet-scale blogs, with tens of thousands of authors to create a cognitive fingerprint for an individual. Existing cognitive biometric modalities involve learning from obtrusive sensors placed on human body. Our modality is based on the characteristic pattern of how individuals express their thoughts through written language. The problems of cognitive authentication (1:1 comparison of genuine versus impostor) and identification (1:n search) are formulated. We detail the algorithms to learn a classifier to distinguish between genuine and impostor classes (for authentication) and multiple classes (for identification). We conclude that a cognitive fingerprint can be successfully learnt, using stylistic (writing style), semantic (themes), and syntactic (grammatical) features extracted from blogs. Our methodology shows promising results (with 79% as the area under the ROC (AUC) in case of authentication). For identification, the individual class accuracies are up to 90%. We performed stricter tests to see how our system performs for unseen user, and report the accuracies of 72% (genuine) and 71% (impostor). Such a study lays the groundwork for building alternative cognitive systems. The modality, presented here, is easy to obtain, unobtrusive and needs no additional hardware. Neeti Pokhriyal, Kshitij Tayal, Ifeoma Nwogu, Venu Govindaraju |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2016 | Understanding Line Plots Using Bayesian NetworkabstractInformation graphics, such as bar charts, graphs, plots etc. in scientific documents primarily facilitate better understanding of information. Graphics are a key component in technical documents as they are simplified representations of complex ideas. When the traditional optical character recognition (OCR) systems is used on digitized documents, we lose the ideas conveyed in these information graphics since OCRs typically work only on text. And although in more recent times, tools have been developed to extract information graphics from pdf files, they still do not intelligently interpret the contents of the extracted graphics. We therefore propose a method for identifying the intended messages of line plots using a Bayesian network. We accomplish this by first extracting a dense set of points in from a line plot and then represent the entire line plot as a sequence of trends. We then implement a Bayesian network for reasoning about the messages conveyed by the line plots and their trends. We validate our approach by performing experiments on a dataset obtained from computer science conference publications and evaluate the performance of the network against the messages generated by human end users. The resulting intended message gives holistic information about the line plot(s) as well as lower level information about the trends that make up the plot. Rathin Radhakrishnan Nair, Nishant Sankaran, Ifeoma Nwogu, Venu Govindaraju |
DAS | 4 |
| 2016 | Online Handwritten Cursive Word Recognition by Combining Segmentation-Free and Segmentation-Based MethodsabstractThis paper describes an online handwritten cursive word recognition approach by combining segmentation-free and segmentation-based methods. To search the optimal segmentation and recognition path as the recognition result, we can attempt two methods: segmentation-free and segmentation-based, where we expand the search space using a character-synchronous beam search strategy. The probable search paths are evaluated by integrating character recognition scores with geometric characteristics of the character patterns in a Conditional Random Field (CRF) model. We make a comparison between online handwritten cursive word recognition using segmentation-free method and that using segmentation-based method, and then attempt combining the two methods to improve performance. Our methods restrict the search paths from the trie lexicon of words and preceding paths during path search. We show this comparison on a publicly available dataset (IAM-OnDB). Bilan Zhu, Arti Shivram, Venu Govindaraju, Masaki Nakagawa |
ICFHR | 3 |
| 2016 | Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep NetworksabstractWhile the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks– \textitInternal Covariate Shift– the current solution has certain drawbacks. For instance, BN depends on batch statistics for layerwise input normalization during training which makes the estimates of mean and standard deviation of input (distribution) to hidden layers inaccurate due to shifting parameter values (especially during initial training epochs). Another fundamental problem with BN is that it cannot be used with batch-size 1 during training. We address these drawbacks of BN by proposing a non-adaptive normalization technique for removing covariate shift, that we call \textitNormalization Propagation. Our approach does not depend on batch statistics, but rather uses a data-independent parametric estimate of mean and standard-deviation in every layer thus being computationally faster compared with BN. We exploit the observation that the pre-activation before Rectified Linear Units follow Gaussian distribution in deep networks, and that once the first and second order statistics of any given dataset are normalized, we can forward propagate this normalization without the need for recalculating the approximate statistics for hidden layers. Devansh Arpit, Yingbo Zhou 0002, Bhargava Urala Kota, Venu Govindaraju |
ICML | 4 |
| 2016 | Why Regularized Auto-Encoders learn Sparse Representation?abstractSparse distributed representation is the key to learning useful features in deep learning algorithms, because not only it is an efficient mode of data representation, but also – more importantly – it captures the generation process of most real world data. While a number of regularized auto-encoders (AE) enforce sparsity explicitly in their learned representation and others don’t, there has been little formal analysis on what encourages sparsity in these models in general. Our objective is to formally study this general problem for regularized auto-encoders. We provide sufficient conditions on both regularization and activation functions that encourage sparsity. We show that multiple popular models (de-noising and contractive auto encoders, e.g.) and activations (rectified linear and sigmoid, e.g.) satisfy these conditions; thus, our conditions help explain sparsity in their learned representation. Thus our theoretical and empirical analysis together shed light on the properties of regularization/activation that are conductive to sparsity and unify a number of existing auto-encoder models and activation functions under the same analytical framework. Devansh Arpit, Yingbo Zhou 0002, Hung Q. Ngo 0001, Venu Govindaraju |
ICML | 4 |
| 2016 | Segmentation of highly unstructured handwritten documents using a neural network techniqueabstractIn recent years there has been a growing interest in digitizing the extensive amounts of books and documents that existed preceding the widespread adoption of digital technologies. Many of these digitizing initiatives deal with huge collections of handwritten documents, for which document image analysis techniques (page segmentation, keyword-spotting, optical character recognition (OCR), etc) are not yet as mature as for printed text. Thus, there is an imminent need to develop techniques to understand, archive, index and search the manuscripts. The antiquated approach of manually transcribing handwritten collections and then using standard text retrieval techniques can be very expensive for large collections. But many of the manuscripts in these collections, unlike machine-printed texts, contain unstructured information, cluttered group of texts and graphics that do not necessarily follow a pre-specified format, thus making it quite challenging to automatically process. Thus, in this paper we present a convolutional neural network (CNN) based implementation that is used to segment pages of handwritten documents into their constituent sections. We showcase a multiscale sliding window based network that is trained to predict the sections of the pages in handwritten manuscripts. The results of the network are post-processed with a novel region growing technique to further improve the segmentation results. The implementation is applied on the Marianne Moore archival collection, a body of handwritten notes and memos by the renowned author Marianne Moore (1887-1972), one of the foremost modernist poets of the early twentieth-century. We present our segmentation results both quantitatively and qualitatively. Rathin Radhakrishnan Nair, Bhargava Urala Kota, Ifeoma Nwogu, Venu Govindaraju |
ICPR | 4 |
| 2016 | Engagement Capacity and Engaging Team Formation for Reach Maximization of Online Social Media PlatformsabstractThe challenges of assessing the "health" of online social media platforms and strategically growing them are recognized by many practitioners and researchers. For those platforms that primarily rely on user-generated content, the reach -- the degree of participation referring to the percentage and involvement of users -- is a key indicator of success. This paper lays a theoretical foundation for measuring engagement as a driver of reach that achieves growth via positive externality effects. The paper takes a game theoretic approach to quantifying engagement, viewing a platform's social capital as a cooperatively created value and finding a fair distribution of this value among the contributors. It introduces engagement capacity, a measure of the ability of users and user groups to engage peers, and formulates the Engaging Team Formation Problem (EngTFP) to identify the sets of users that "make a platform go". We show how engagement capacity can be useful in characterizing forum user behavior and in the reach maximization efforts. We also stress how engagement analysis differs from influence measurement. Computational investigations with Twitter and Health Forum data reveal the properties of engagement capacity and the utility of EngTFP. Alexander G. Nikolaev, Shounak Gore, Venu Govindaraju |
KDD | 3 |
| 2015 | Automated analysis of line plots in documentsabstractInformation graphics, such as graphs and plots, are used in technical documents to convey information to humans and to facilitate greater understanding. Usually, graphics are a key component in a technical document, as they enable the author to convey complex ideas in a simplified visual format. However, in an automatic text recognition system, which are typically used to digitize documents, the ideas conveyed in a graphical format are lost. We contend that the message or extracted information can be used to help better understand the ideas conveyed in the document. In scientific papers, line plots are the most commonly used graphic to represent experimental results in the form of correlation present between values represented on the axes. The contribution of our work is in the series of image processing algorithms that are used to automatically extract relevant information, including text and plot from graphics found in technical documents. We validate the approach by performing the experiments on a dataset of line plots obtained from scientific documents from computer science conference papers and evaluate the variation of a reconstructed curve from the original curve. Our algorithm achieves a classification accuracy of 91% across the dataset and successfully extracts the axes from 92% of line plots. Axes label extraction and line curve tracing are performed successfully in about half the line plots as well. Rathin Radhakrishnan Nair, Nishant Sankaran, Ifeoma Nwogu, Venu Govindaraju |
ICDAR | 4 |
| 2015 | A sigma-lognormal model for character level CAPTCHA generationabstractWord level handwritten CAPTCHA generation involves picking a handwritten word from a pre-existing database and cumulatively applying distortions and noise models. In principle, the addition of distortion and noise makes the CAPTCHA robust to automated attacks. However, the primary drawback of the word level CAPTCHA generation is that it limits us to words that already exist in our data set. If the primary building block of this approach was a character, we could move away from a lexicon based CAPTCHA generation and generate CAPTCHAs which are resistant to a dictionary based attack. In this paper, we propose a Sigma-Lognormal based approach to generate character level CAPTCHAs. Next, we increase the robustness of the model by applying ideas from accents in handwriting to our problem. Finally, we demonstrate the efficacy of our approach by simulating an attack by an automated word recognizer. Chetan Ramaiah, Réjean Plamondon, Venu Govindaraju |
ICDAR | 3 |
| 2014 | Multiclass Learning for Writer Identification Using Error-Correcting CodesabstractWriter Identification can be seen as a multi-class learning problem where number of writers are different classes. One of the fundamental approaches to solve a multi-class problemis by breaking it into binary classification tasks. In this work weare proposing a generic approach for multi-class classification using an ensemble of binary classifiers. We assign a distributedoutput representation to each class in the form of codewords andan ensemble of binary classifiers is created where each classifierpredicts one bit of the codeword. Actual label is determined using Belief Propagation algorithm on a graph constructed from the code matrix. We have performed experiments on a new publiclyavailable IBM-UB-1 dataset for the task of writer identification to show the efficacy of our method. Utkarsh Porwal, Chetan Ramaiah, Venu Govindaraju |
Document Analysis Systems | 4 |
| 2014 | A Hierarchical Framework for Accent Based Writer IdentificationabstractWriter identification is the process of determining the author of a handwritten specimen by utilizing characteristics inherent in the sample. In this work, we apply the concept of accents in handwriting to introduce a novel perspective for writer identification. Analogous to speech, accents in handwriting can be defined as distinctive writing quirks that are unique to a group of people sharing a common native script. Specifically, we postulate that a group of people with a common native script will share certain traits in their handwriting style that are exposed when they write in a different script. We propose a hierarchical framework for the writer identification task, wherein, we first identify the accent of the writer. In the next step, we perform writer identification based on the selected accent. This framework reduces the complexity of the classification task by reducing the number of classes at the prediction stage. Experiments are performed on the UNIPEN dataset and the results lend credibility to our model. Chetan Ramaiah, Venu Govindaraju |
Document Analysis Systems | 2 |
| 2014 | Use of language as a cognitive biometric traitabstractThis paper investigates whether the cognitive state of a person can be learnt and used as a novel biometric trait. We explore the idea of using language written by an author, as his/her cognitive fingerprint. The dataset consists of millions of blogs written by thousands of authors on the Internet. Our proposed method learns a classifier that can distinguish between genuine and impostor authors. Our results are encouraging (we report 72% Area under the ROC curve) and show that users do have a distinctive linguistic style, which is evident even when analyzing a corpora as large and diverse as the Internet. When we tested on new authors that the system had never encountered before, our methodology correctly identified genuine authors with 78% accuracy and impostors with 76% accuracy. Neeti Pokhriyal, Ifeoma Nwogu, Venu Govindaraju |
IJCB | 3 |
| 2014 | A Bayesian Approach to Script Independent Multilingual Keyword SpottingabstractWe propose a script independent Bayesian framework for keyword spotting in multilingual handwritten documents. The approach relies on local character level score and global word level hypothesis scores and learns a Bayesian logistic regression classifier to distinguish between keywords and non-keywords. In a Bayesian formulation of logistic regression, the integral over weights becomes intractable. Variational approximation is used for inference. In order to learn a robust classifier with minimal number of samples, we apply Bayesian active learning framework to request labels for those word images which provide maximum information gain in improving the classifier. We evaluate our system on multilingual datasets, publicly available IAM dataset for English, AMA for Arabic and LAW dataset for Devanagiri. The system is also evaluated on a synthetic multilingual dataset prepared by combining samples from IAM, AMA and LAW datasets. The results are comparable with the state of art multilingual keyword spotting framework. Venu Govindaraju |
ICFHR | 2 |
| 2014 | Bayesian Active Learning for Keyword Spotting in Handwritten DocumentsabstractWe propose the Bayesian Active Learning by Disagreement (BALD) model for keyword spotting in handwritten documents. In the context of keyword spotting in handwritten documents, the background text is all regions in the document that do not contain the keywords. The model tries to learn certain characteristics of the keyword and background text in an active learning framework. It takes into account the local character level scores and global word level scores to distinguish keywords from non-keywords. We propose to apply the bayesian active learning strategy to identify the regions of sample space from which more meaningful labeled samples of keywords and non-keywords can be extracted. This work is an extension to our previous work which used a variational dynamic background model to model the large variations of background text. The approach has been tested on IAM dataset for English. The results show that a decent background model can be learned in a more quicker and efficient manner using the BALD framework. The approach outperforms our prior work and other state of the art approaches. Venu Govindaraju |
ICPR | 2 |
| 2014 | A Sigma-Lognormal Model for Handwritten Text CAPTCHA GenerationabstractPopular CAPTCHA systems consist of garbled printed text character images with significant distortions and noise. It is believed that humans have little difficulty in deciphering the text, whereas automated systems are foiled by the added noise and distortion. However, in recent years, several text based CAPTCHAs have been reported as broken, that is, automated systems can identify the text in the displayed image with a reasonable amount of success. An extension to the text based CAPTCHA concept is to utilize unconstrained handwritten text, which is still considered to be a challenging problem for automated systems. In this work, we present a automated handwritten CAPTCHA generation system by adding distortions to the Sigma-Lognormal representation of a handwritten word sample. In addition, several noise models are also considered. We perform experiments on the UNIPEN dataset and demonstrate the efficacy of the approach. Chetan Ramaiah, Réjean Plamondon, Venu Govindaraju |
ICPR | 3 |
| 2014 | Data Sufficiency for Online Writer Identification: A Comparative Study of Writer-Style Space vs. Feature Space ModelsabstractA key factor in building effective writer identification/verification systems is the amount of data required to build the underlying models. In this research we systematically examine data sufficiency bounds for two broad approaches to online writer identification -- feature space models vs. writer-style space models. We report results from 40 experiments conducted on two publicly available datasets and also test identification performance for the target models using two different feature functions. Our findings show that the writer-style space model gives higher identification performance for a given level of data and further, achieves high performance levels with lesser data costs. This model appears to require as less as 20 words per page to achieve identification performance close to 80% and reaches more than 90% accuracy with higher levels of data enrollment. Arti Shivram, Chetan Ramaiah, Venu Govindaraju |
ICPR | 3 |
| 2014 | Statistical Relational Learning for Handwriting Recognition
Arti Shivram, Tushar Khot, Sriraam Natarajan, Venu Govindaraju |
ILP | 4 |
| 2014 | Dimensionality Reduction with Subspace Structure Preservation
Devansh Arpit, Ifeoma Nwogu, Venu Govindaraju |
NIPS | 3 |
| 2014 | Parallel Feature Selection Inspired by Group Testing
Yingbo Zhou 0002, Utkarsh Porwal, Ce Zhang 0001, Hung Q. Ngo 0001, XuanLong Nguyen, Christopher Ré, Venu Govindaraju |
NIPS | 7 |
| 2014 | Statistical script independent word spotting in offline handwritten documents
Safwan Wshah, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2013 | A Bayesian Framework for Modeling Accents in HandwritingabstractAccent in speech is defined as a distinctive mode of pronunciation that is unique to a geographical region. In a similar way, we define accent in handwriting as distinctive writing characteristics that are unique to a group of people sharing a common native script. In other words, we postulate that a group of people with a common native script will share certain traits in their handwriting that can be ascertained when they write in a different script. In this paper, we establish the existence of accents in handwriting using a hierarchical Bayesian framework. We then demonstrate that the unique trait in handwriting that arises out of the writer's native script is indigenous to that script, which is perceivable when writing in a different script. As a consequence, the ability to identify a person's native script based on the person's handwriting style in another script is introduced. We validated the approach by performing experiments on the UNIPEN dataset, and the experiments lend credibility to our model. Chetan Ramaiah, Arti Shivram, Venu Govindaraju |
ICDAR | 3 |
| 2013 | A Model Based Framework for Table Processing in Degraded Document ImagesabstractThis paper describes a model based framework for detection and extraction of the contents of table cells from degraded handwritten document images that contain tables. Given the very poor quality of the target documents, the table cell detection problem is formulated conceptually as a two-step process. The first step is to identify the location of the table and extract the content of table cells given a model of the structure of the table present in the image. The second step is to identify the model of the table present in a document image from a list of given table models. A model-based representation for tables is introduced and is used for matching table candidates with the given model to identify and extract the contents of table cells. The approach for detecting potential table candidates is based on the detection of horizontal and vertical table line candidates. The table representation is a matrix of horizontal and vertical table line crossings, and the matching algorithm is formulated as a minimization problem where the optimal table candidate is obtained using the minimal distance between the candidate and model table matrices which is then used for extraction of the table cell contents. A similar approach is used to solve the model selection problem where the best fitting location in the document page for each of the candidate models is identified using the distance minimization approach along with a confidence score and the model with the highest confidence score is selected as the correct model. The approach was tested on document page images containing tables from the challenge set of the DARPA MADCAT handwritten document image data. Results indicate that the method is effective for both model selection as well as table cell content extraction. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2013 | IBM_UB_1: A Dual Mode Unconstrained English Handwriting DatasetabstractIn this paper we present a new dual mode, twin-folio structured English handwriting dataset IBM_UB_1. IBM_UB_1 is our first major release from a large multilingual handwriting corpus. Containing over 6000 pages of handwritten matter, this dataset can not only be used for unconstrained handwriting recognition, more importantly, the dataset's unique twin-folio structure presents a natural fit for research on writer identification, keyword spotting, indexing and various forms of handwritten document search and retrieval. We first describe two central characteristics of the dataset - the twin-folio structure and dual modality (online/offline) - and their relevance to current research problems. Secondly, we describe the dataset, its collection and construction, and provide key descriptive statistics. Finally, we evaluate the dataset on two different research domains - handwriting recognition and writer identification - and present related experimental results. Arti Shivram, Chetan Ramaiah, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 4 |
| 2013 | Segmentation Based Online Word Recognition: A Conditional Random Field Driven Beam Search StrategyabstractWe propose a segmentation based online word recognition approach which uses a Conditional Random Field (CRF) driven beam search strategy. An efficient trie-lexicon directed, breadth-first beam search algorithm is employed in a combined segmentation-and-recognition framework to accomplish real-time recognition of online handwritten cursive English words. This framework is developed by building a candidate lattice of primitive segments obtained through over segmentation of the word pattern. The search space for the lattice is expanded by synchronously matching the lattice nodes to likely character patterns from a trie-dictionary constructed out of the target lexicon. The probable paths are evaluated by integrating character recognition scores with physical and spatial characteristics of the handwritten segments in a CRF (conditional random field) model and a beam search strategy is used to prune the set of likely paths. This approach has been benchmarked on the new IBM_UB_1 dataset as well as on the UNIPEN dataset for comparison. Arti Shivram, Bilan Zhu, Srirangaraj Setlur, Masaki Nakagawa, Venu Govindaraju |
ICDAR | 5 |
| 2013 | Online Handwritten Cursive Word Recognition Using Segmentation-Free MRF in Combination with P2DBMN-MQDFabstractThis paper describes an online handwritten English cursive word recognition method using a segmentation-free Markov random field (MRF) model in combination with an offline recognition method which uses pseudo 2D bi-moment normalization (P2DBMN) and modified quadratic discriminant function (MQDF). It extracts feature points along the pen-tip trace from pen-down to pen-up and uses the feature point coordinates as unary features and the differences in coordinates between the neighboring feature points as binary features. Each character is modeled as a MRF and word MRFs are constructed by concatenating character MRFs according to a trie lexicon of words during recognition. Our method expands the search space using a character-synchronous beam search strategy to search the segmentation and recognition paths. This method restricts the search paths from the trie lexicon of words and preceding paths, as well as the lengths of feature points during path search. We also combine it with a P2DBMN-MQDF recognizer that is widely used for Chinese and Japanese character recognition. Bilan Zhu, Arti Shivram, Srirangaraj Setlur, Venu Govindaraju, Masaki Nakagawa |
ICDAR | 4 |
| 2013 | Handwritten text separation from annotated machine printed documents using Markov Random Fields
Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Int. J. Document Anal. Recognit. | 3 |
| 2013 | Language-motivated approaches to action recognition
Manavender R. Malgireddy, Ifeoma Nwogu, Venu Govindaraju |
J. Mach. Learn. Res. | 3 |
| 2013 | Enhancing biometric recognition with spatio-temporal reasoning in smart environments
Vivek Menon, Bharat Jayaraman, Venu Govindaraju |
Pers. Ubiquitous Comput. | 3 |
| 2013 | Labeling Spain With StanfordabstractWe present an end-to-end framework for outdoor scene region decomposition, learned on a small set of randomly selected images that generalizes well to multiple data sets containing images from around the world. We discuss the different aspects of the framework especially a generalized variational inference method with better approximations to the true marginals of a graphical model. Experimentally, we explain why the framework is robust and performs competitively on many diverse scene data sets, including several unseen scene types. We have obtained high pixel-level accuracies (≈ 80%) in three of the four data sets, which include a benchmark data set known as the Stanford background data set. Our model obtained over 70% accuracy on the fourth data set, which contained a number of indoor and close-up images that are significantly different from our training examples. Yingbo Zhou 0002, Ifeoma Nwogu, Venu Govindaraju |
IEEE Trans. Image Process. | 3 |
| 2012 | Ensemble of Biased Learners for Offline Arabic Handwriting RecognitionabstractTechniques and performance of text recognition systems and software has shown great improvement in recent years. OCRs now can read any machine printed document with good accuracy. However, the advancements are primarily for Latin scripts and even for such scripts performance is limited in case of handwritten documents. Little work has been done for cursive scripts such as Arabic and still there is a room for improvement both in terms of accuracy and techniques. This paper presents an algorithm to recognize handwritten Arabic text using an ensemble of biased classifiers in a hierarchical setting. We address the fundamental shortcomings of the traditional Machine Learning paradigms when applied to Arabic scripts. Experiments have been conducted on the AMA Arabic dataset to show the efficacy of our method. Utkarsh Porwal, Arti Shivram, Chetan Ramaiah, Venu Govindaraju |
Document Analysis Systems | 4 |
| 2012 | Accent Detection in Handwriting Based on Writing StylesabstractAccent in handwriting can be defined as the influence of a writer's native script on his/her writing style in another script. In this paper, we approach the problem of detecting the existence of accents in handwriting. We approach this problem using two sets of writers, those who can write only in English, and the other set being multilingual writers who can also write in English. We learn the writing styles that are predominant in each set and use it as features in classification. Latent Dirichlet Allocation is used to learn the distribution over writing styles. Experimental results suggest the existence of accents in handwriting. Chetan Ramaiah, Utkarsh Porwal, Venu Govindaraju |
Document Analysis Systems | 3 |
| 2012 | Keyword Spotting Framework Using Dynamic Background ModelabstractAn important task in Keyword Spotting in handwritten documents is to separate Keywords from Non Keywords. Very often this is achieved by learning a filler or background model. A common method of building a background model is to allow all possible sequences or transitions of characters. However, due to large variation in handwriting styles, allowing all possible sequences of characters as background might result in an increased false reject. A weak background model could result in high false accept. We propose a novel way of learning the background model dynamically. The approach first used in word spotting in speech uses a feature vector of top K local scores per character and top N global scores of matching hypotheses. A two class classifier is learned on these features to classify between Keyword and Non Keyword. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
ICFHR | 4 |
| 2012 | Structural Learning for Writer Identification in Offline HandwritingabstractAvailability of sufficient labeled data is key to the performance of any learning algorithm. However, in document analysis obtaining the large amount of labeled data is difficult. Scarcity of labeled samples is often a main bottleneck in the performance of algorithms for document analysis. However, unlabeled data samples are present in abundance. We propose a semi supervised framework for writer identification for offline handwritten documents that leverages the information hidden in the unlabeled samples. The task of writer identification is a complex one and our framework tries to model the nuances of handwriting with the use of structural learning. This framework models the complexity of learning problem by selecting the best hypotheses space by breaking the main task into several sub tasks. All the hypotheses spaces pertaining to the sub tasks will be used for the best model selection by retrieving a common optimal sub structure that has high correspondence with all of the candidate hypotheses spaces. We have used publically available IAM data set to show the efficacy of our method. Utkarsh Porwal, Chetan Ramaiah, Arti Shivram, Venu Govindaraju |
ICFHR | 4 |
| 2012 | Modeling Writing Styles for Online Writer Identification: A Hierarchical Bayesian ApproachabstractWith the explosive growth of the tablet form factor and greater availability of pen-based direct input, writer identification in online environments is increasingly becoming critical for a variety of downstream applications such as intelligent and adaptive user environments, search, retrieval, indexing and digital forensics. Extant research has approached writer identification by using writing styles as a discriminative function between writers. In contrast, we model writing styles as a shared component of an individualâs handwriting. We develop a theoretical framework for this conceptualization and model this using a three level hierarchical Bayesian model (Latent Dirichlet Allocation). In this text-independent, unsupervised model each writerâs handwriting is modeled as a distribution over finite writing styles that are shared amongst writers. We test our model on a novel online/offline handwriting dataset IBM UB 1 which is being made available to the public. Our experiments show comparable results to current benchmarks and demonstrate the efficacy of explicitly modeling shared writing styles. Arti Shivram, Chetan Ramaiah, Utkarsh Porwal, Venu Govindaraju |
ICFHR | 4 |
| 2012 | Script Independent Word Spotting in Offline Handwritten Documents Based on Hidden Markov ModelsabstractKeyword spotting aims to retrieve all instances of a given keyword from a document in any language. In this paper, we propose a novel script independent line based word spotting framework for offline handwritten documents based on Hidden Markov Models. The methodology simulates the keywords in model space as a sequence of character models and uses the filler models for better representation of background or non-keyword text. We propose a two stage spotting framework where the candidate keywords are further pruned using the character based background and lexicon based background model. The system deals with large vocabulary without the need for word or character segmentation. The system has been evaluated on many public dataset from several languages such as IAM for English, AMA for Arabic and LAW for Devanagari. The system outperforms the modern line based approach on the English, Arabic and Devanagari Datasets. Safwan Wshah, Venu Govindaraju |
ICFHR | 3 |
| 2012 | Handwritten Arabic text recognition using Deep Belief Networks
Utkarsh Porwal, Yingbo Zhou 0002, Venu Govindaraju |
ICPR | 3 |
| 2012 | Multilingual word spotting in offline handwritten documents
Safwan Wshah, Venu Govindaraju |
ICPR | 3 |
| 2012 | Making Sense of All Things Handwritten - From Postal Addresses to Tablet Notes
Venu Govindaraju |
SECRYPT | 1 |
| 2012 | Using a boosted tree classifier for text segmentation in hand-annotated documents
Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Pattern Recognit. Lett. | 3 |
| 2011 | Lie to Me: Deceit detection via online behavioral learningabstractInspired by the the behavioral scientific discoveries of Dr. Paul Ekman in relation to deceit detection, along with the television drama series Lie to Me, also based on Dr. Ekman's work, we use machine learning techniques to study the underlying phenomena expressed when a person tells a lie. We build an automated framework which detects deceit by measuring the deviation from normal behavior, at a critical point in the course of an investigative interrogation. Behavioral psychologists have shown that the eyes (via either gaze aversion or gaze extension) can be good “reflectors” of the inner emotions, when a person tells a high-stake lie. Hence we develop our deceit detection framework around eye movement changes. A dynamic bayesian model of eye movements is trained during a normal course of conversation for each subject, to represent normal behavior. The remaining conversation is broken into sequences and each sequence is tested against the parameters of the model of normal behavior. At the critical points in the interrogations, the deviations from normalcy are observed and used to deduce verity/deceit. An analysis on 40 subjects gave an accuracy of 82.5% which strongly suggests that the latent parameters of eye movements successfully capture behavioral changes and could be viable for use in automated deceit detection. Nisha Bhaskaran, Ifeoma Nwogu, Mark G. Frank, Venu Govindaraju |
FG | 4 |
| 2011 | Combination of multiple samples utilizing identification model in biometric systemsabstractIn some cases, the test person might be asked to provide another authentication attempt besides the first one so that combination of the two input templates might give the system more confidence if the person is genuine or impostor. Instead of simply combining the matching scores which are associated with a single person compared to the two input templates, we investigate the use of matching scores corresponding to all enrolled persons. The dependencies between scores generated by the same input templates are accounted for the proposed combination algorithm. Such combination methods can be extended to large number of classes and input templates. Since matching scores are used, the proposed methods can also be applied on arbitrary biometric modalities. The experiments are conducted on NIST BSSR1 face and FVC2002 fingerprint datasets by using both likelihood ratio and multilayer perceptron combination methods. Sergey Tulyakov, Venu Govindaraju |
IJCB | 3 |
| 2011 | Image Enhancement for Degraded Binary Document ImagesabstractThis paper presents a novel set of image enhancement algorithms for binary images of poorly scanned real world page documents. Problems that are targeted by the methods described include large blobs or clutter noise, salt-and-pepper noise and detection and removal of non-text objects such as form lines or rule-lines. The algorithms described are shown to be very effective in removing clutter noise and pepper noise as well as form lines and rule-lines. A region growing algorithm is also described to enhance the quality of the text and to fix the problems arising from the salt noise which leaves holes in the text and creates broken strokes. The methods were tested on 204 images from the challenge set of the DARPA MADCAT Arabic handwritten document image data. The results indicate that the methods described are robust and are capable of significantly improving the image quality for downstream OCR systems. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2011 | Detecting Figure-Panel Labels in Medical Journal Articles Using MRFabstractWe present a method for figure-panel (subfigure) label detection and recognition in multi-panel figures extracted from biomedical articles. Figures in biomedical articles often comprise several subfigures that are identified by superimposed panel labels ('A', 'B', ...) which are referenced in the figure caption and discussion in the article body. Splitting such multi-panel figures into individual subfigures is a necessary step for improved multimodal biomedical information retrieval. Prior to feature extraction for indexing and retrieval of biomedical figures it is necessary to classify image content in each subfigure by its modality (X-ray, MRI, CT, etc.) and other relevant criteria. Subfigure labels are valuable in associating individual panels with relevant text in captions and discussion. We propose a 4-step panel label detection method based on Markov Random Field (MRF). Experiments on 515 multi-panel figures and analysis of the results show promising results. We present the successes and identify critical challenges. Daekeun You, Sameer K. Antani, Dina Demner-Fushman, Venu Govindaraju, George R. Thoma |
ICDAR | 4 |
| 2011 | A Shared Parameter Model for Gesture and Sub-gesture Analysis
Manavender R. Malgireddy, Ifeoma Nwogu, Subarna Ghosh, Venu Govindaraju |
IWCIA | 4 |
| 2011 | Unconstrained handwritten document retrieval
Huaigu Cao, Venu Govindaraju, Anurag Bhardwaj |
Int. J. Document Anal. Recognit. | 2 |
| 2010 | Latent Dirichlet allocation based writer identification in offline handwritingabstractIn this paper, we describe a novel approach to Writer Identification in Offline handwriting using Latent Dirichlet Allocation. State-of-the-art methods for writer identification employ the traditional feature-classification paradigm which does not provide enough information about the handwriting attributes such as writing style which are key components in any forensic analysis of handwriting. This problem is also compounded due to lack of efficient rules for defining a particular writing style that can capture writer specific characteristics over a large dataset. We propose to address this issue by using a generative model in form of Latent Dirichlet Allocation(LDA) that automatically infers writing styles from handwritten document collection without any pre-defined set of rules. This information is then used to represent each writer as a distribution over multiple writing style for classifying any unknown writer sample. We describe our approach on two different feature sets consisting of contour angle features as well as structural and concavity features. Our experimental results show comparable performance with baseline systems and also demonstrate the efficacy of LDA for learning multiple handwriting styles. Anurag Bhardwaj, Manavender R. Malgireddy, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Document Analysis Systems | 4 |
| 2010 | Overlapped text segmentation using Markov random field and aggregationabstractSeparating machine printed text and handwriting from overlapping text is a challenging problem in the document analysis field and no reliable algorithms have been developed thus far. In this paper, we propose a novel approach for separating handwriting from binary image of overlapped text. Instead of using fixed size training patches, we describe an aggregation method which uses shape context features to extract training samples automatically. We use a Markov Random Field (MRF) to model the overlapped text. The neighbor system is inherited from a coarsening procedure and the prior and likelihood of the MRF is learned based on a distance metric. Experimental results show that the proposed method can achieve 87.97% recall for handwriting and 91.44% recall for machine printed text. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Document Analysis Systems | 3 |
| 2010 | Retrieving Handwriting Styles: A Content Based Approach to Handwritten Document RetrievalabstractLarge scale retrieval of handwritten documents has primarily been focused around searching a query text in the OCR'ed transcription of the document images, which provides a limited view of the complete search process. Recent research advances have led to a number of content based retrieval techniques which expand the search scope to document content level (i.e. image features, meta-information). Based on similar motivations, we propose a new approach to content based retrieval of handwritten document images by retrieving similar handwriting styles corresponding to a handwritten query image. At the core, we formulate this problem as the task of unsupervised writer style classification without the need of any style definitions or grammar. We build upon our previous work in writer style modeling and apply it to learn a style distribution for every handwriting sample in the corpus. Given a query image, all documents are ranked in order of their style distribution similarity. Experimental results conducted on publicly available IAM dataset demonstrate the efficacy of our proposed method over baseline feature based systems. Anurag Bhardwaj, Achint Oommen Thomas, Yun Fu 0001, Venu Govindaraju |
ICFHR | 4 |
| 2010 | Generation of Handwriting by Active Shape Modeling and Global Local Approximation (GLA) AdaptationabstractThe generation of handwriting is a complex task. In order to accommodate for the large variations involved in handwritten words deformable templates need to be used. In this paper we propose a handwriting model, based on Active shape modeling (ASM). In a two-step generation process, a template-based ASM generates characters and a Gaussian mixture regression (GMR) model concatenates the generated characters. For real time generation of cursive handwriting an adaptation of Global local approximation (GLA) methodology is used to fit the generated models. Ashirwad Joseph Chowriappa, Ricardo Rodrigues 0004, Thenkurussi Kesavadas, Venu Govindaraju, Ann M. Bisantz |
ICFHR | 4 |
| 2010 | Leveraging the Mixed-Text Segmentation Problem to Design Secure Handwritten CAPTCHAsabstractIn this paper we present a novel CAPTCHA that is based on the current hard AI problem of mixed-text (handwriting and printed-text) segmentation. The proposed CAPTCHA overlays generated handwritten word images on a generated printed-text background. We first propose a modification that allows for character level perturbations on an existing synthetic handwriting generation technique. These perturbations are parameterized allowing for varying levels of handwritten word complexity. We then use the output from the modified synthetic handwriting generator as the foreground for the mixed-text CAPTCHA. Experiments show that the proposed approach is effective at successfully distinguishing between humans and machines. Human recognition accuracy averages at 0.77 while machine accuracy is below 0.0001. Achint Oommen Thomas, Sulabh Choudhury, Venu Govindaraju |
ICFHR | 3 |
| 2010 | A Robust Iris Localization Method Using an Active Contour Model and Hough TransformabstractIris segmentation is one of the crucial steps in building an iris recognition system since it affects the accuracy of the iris matching significantly. This segmentation should accurately extract the iris region despite the presence of noises such as varying pupil sizes, shadows, specular reflections and highlights. Considering these obstacles, several attempts have been made in robust iris localization and segmentation. In this paper, we propose a robust iris localization method that uses an active contour model and a circular Hough transform. Experimental results on 100 images from CASIA iris image database show that our method achieves 99% accuracy and is about 2.5 times faster than the Daugman's in locating the pupillary and the limbic boundaries. Jaehan Koh, Venu Govindaraju, Vipin Chaudhary |
ICPR | 2 |
| 2010 | Combination of Symmetric Hash Functions for Secure Fingerprint MatchingabstractFingerprint based secure biometric authentication systems have received considerable research attention lately, where the major goal is to provide an anonymous, multipliable and easily revocable methodology for fingerprint verification. In our previous work, we have shown that symmetric hash functions are very effective in providing such secure fingerprint representation and matching since they are independent of order of minutiae triplets as well as location of singular points (e.g. core and delta). In this paper, we extend our prior work by generating a combination of symmetric hash functions, which increases the security of fingerprint matching by an exponential factor. Firstly, we extract kplets from each fingerprint image and generate a unique key for combining multiple hash functions up to an order of (k-1). Each of these keys is generated using the features extracted from minutiae k-plets such as bin index of smallest angles in each k-plet. This combination provides us an extra security in the face of brute force attacks, where the compromise of few hash functions as well do not compromise the overall matching. Our experimental results suggest that the EER obtained using the combination of hash functions (4.98%) is comparable with the baseline system (3.0%), with the added advantage of being more secure. Sergey Tulyakov, Venu Govindaraju |
ICPR | 3 |
| 2010 | A Framework for Hand Gesture Recognition and Spotting Using Sub-gesture ModelingabstractHand gesture interpretation is an open research problem in Human Computer Interaction (HCI), which involves locating gesture boundaries (Gesture Spotting) in a continuous video sequence and recognizing the gesture. Existing techniques model each gesture as a temporal sequence of visual features extracted from individual frames which is not efficient due to the large variability of frames at different timestamps. In this paper, we propose a new sub-gesture modeling approach which represents each gesture as a sequence of fixed sub-gestures (a group of consecutive frames with locally coherent context) and provides a robust modeling of the visual features. We further extend this approach to the task of gesture spotting where the gesture boundaries are identified using a filler model and gesture completion model. Experimental results show that the proposed method outperforms state-of-the-art Hidden Conditional Random Fields (HCRF) based methods and baseline gesture spotting techniques. Manavender R. Malgireddy, Jason J. Corso, Srirangaraj Setlur, Venu Govindaraju, Dinesh Mandalapu |
ICPR | 4 |
| 2010 | Text Separation from Mixed Documents Using a Tree-Structured ClassifierabstractIn this paper, we propose a tree-structured multi-class classifier to identify annotations and overlapping text from machine printed documents. Each node of the tree-structured classifier is a binary weak learner. Unlike normal decision tree(DT) which only considers a subset of training data at each node and is susceptible to over-fitting, we boost the tree using all training data at each node with different weights. The evaluation of the proposed method is presented on a set of machine printed documents which have been annotated by multiple writers in an office/collaborative environment. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
ICPR | 3 |
| 2010 | Removing Rule-Lines from Binary Handwritten Arabic Document Images Using Directional Local ProfileabstractIn this paper, we present a novel approach for detecting and removing pre-printed rule-lines from binary handwritten Arabic document images. The proposed technique is based on a directional local profiling approach for the detection of the rule-line locations. Then a refined adaptive vertical run-length search is designed for removing the rule-line pixels without much damaging to the text. They are also tolerate to the variations in the rule-lines such as broken lines, orientation changes and variation in the thickness of the rule-lines. Analysis of experimental results on the DARPA MADCAT Arabic handwritten document data indicates that the method is robust and is capable of correctly removing rule-lines. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 3 |
| 2010 | A Novel Lexicon Reduction Method for Arabic Handwriting RecognitionabstractIn this paper, we present a method for lexicon size reduction which can be used as an important pre-processing for an off-line Arabic word recognition. The method involves extraction of the dot descriptors and PAWs (Piece of Arabic Word ). Then the number and position of dots and the number of the PAWs are used to eliminate unlikely candidates. The extraction of the dot descriptors is based on defined rules followed by a convolutional neural network for verification. The reduction algorithm makes use of the combination of two features with a dynamic matching scheme. On IFN/ENIT database of 26459 Arabic handwritten word images we achieved a reduction rate of 87% with accuracy above 93%. Safwan Wshah, Venu Govindaraju, Yanfen Cheng |
ICPR | 2 |
| 2010 | Generation and use of handwritten CAPTCHAs
Amalia I. Rusu, Achint Oommen Thomas, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 3 |
| 2010 | On the Difference between Optimal Combination Functions for Verification and Identification SystemsabstractWe have investigated different scenarios of combining pattern matchers. The combination problem can be viewed as a construction of a postprocessing classifier operating on the matching scores of the combined matchers. The optimal combination algorithm for verification systems corresponds to the likelihood ratio combination function. It can be implemented by the direct reconstruction of this function with genuine and impostor score density approximations. However, the optimal combination algorithm for identification systems is difficult to express analytically. We will show that this difficulty is caused by the dependencies between matching scores assigned to different classes by the same classifier. The experiments on the large sets of scores from handwritten word recognizers operating on postal images and biometric matchers (NIST biometric score set BSSR1) confirm the existence of such dependencies and that the optimal combination functions for verification and identification systems are different. Sergey Tulyakov, Chaohong Wu, Venu Govindaraju |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2010 | Multimodal identification and tracking in smart environments
Vivek Menon, Bharat Jayaraman, Venu Govindaraju |
Pers. Ubiquitous Comput. | 3 |
| 2009 | Markov Random Field Based Text Identification from Annotated Machine Printed DocumentsabstractIn this paper, we describe an approach to segment handwritten text, machine printed text and noise from annotated machine printed documents. Three categories of word level features are extracted. We use a modified K-Means clustering algorithm for classification followed by a relabeling procedure using Markov Random Field(MRF) based on a concept of neighboring patches and Belief Propagation(BP) rules. Experimental results on an imbalanced data set show that our approach achieves an overall recall of 96.33%. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram, Kiran Bhuvanagiri |
ICDAR | 3 |
| 2009 | A Steerable Directional Local Profile Technique for Extraction of Handwritten Arabic Text LinesabstractIn this paper, we present a new text line extraction method for handwritten Arabic documents. The proposed technique is based on a generalized adaptive local connectivity map (ALCM) using a steerable directional filter. The algorithm is designed to solve the particularly complex problems seen in handwritten documents such as fluctuating, touching or crossing text lines. The proposed algorithm consists of three steps. Firstly, a steerable filter is used to probe and determine foreground intensity along multiple directions at each pixel while generating the ALCM. The ALCM is then binarized using an adaptive thresholding algorithm to get a rough estimate of the location of the text lines. In the second step, connected component analysis is used to classify text and non text patterns in the generated ALCM to refine the location of the text lines. Finally, the text lines are separated by superimposing the text line patterns in the ALCM on the original document image and extracting the connected components covered by the pattern mask. Analysis of experimental results on the DARPA MADCAT Arabic handwritten document data indicate that the method is robust and is capable of correctly isolating handwritten text lines even on challenging document images. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2009 | Segmentation of Arabic Handwriting Based on both Contour and Skeleton SegmentationabstractWe propose a new algorithm for segmentation of off-line handwritten Arabic words. The algorithm segments the connected letters to smaller segments each of which contains no more than three letters. Each letter may be segmented to at most five pieces. In addition to improving the recognition of Arabic words, another potential application of the proposed segmentation method is to build lexicon of small size, consisting of no more than three letter combinations. Generally, it is very hard to generate lexicon for recognition of unconstraint handwritten Arabic documents due to the large number of words of Arabic language.The algorithm has been tested on over 6300 words from 45 different documents written by 18 writers. The system is able to segment more than 93% of the words into segments, each containing at most one letter, 6% of the words into segments that contains two letters and 3% of the words into segments that contains three letters. Safwan Wshah, Zhixin Shi, Venu Govindaraju |
ICDAR | 3 |
| 2009 | A Hierarchical Classification Model for Document CategorizationabstractWe propose a novel hierarchical classification method for documents categorization in this paper. The approach consists of multiple levels of classification for different hierarchies. Regularized Least Square (RLS)binary classifiers are applied in the middle levels of the hierarchy to classify documents into smaller set of categories and K-nearest-neighbor (KNN) multi-class classifiers are used at the bottom to classify documents into final classes. Experiments on large-scale real world tax documents show that the proposed hierarchical approach outperforms traditional flat classification method. Jianwu Xu, Vartika Singh, Venu Govindaraju, Depankar Neogi |
ICDAR | 3 |
| 2009 | Nested state indexing in pairwise Markov networks for fast handwritten document image rule-line removalabstractThe Markov random field (MRF) has been applied to modeling the connectivity constraints of the text in document images for tasks like binarization and rule-line removal. One challenge of applying the MRF is its high computational cost. This paper presents a method using two nested set of states trained to reduce the computational cost of patch-based MRF. The two sets of states are trained at different levels in coarse-to-fine order. We show effective reduction of run time but very little loss of quality using rule-line removal experiments. Huaigu Cao, Rohit Prasad, Premkumar Natarajan, Venu Govindaraju |
ICIP | 4 |
| 2009 | Using topic models for OCR correction
Faisal Farooq, Anurag Bhardwaj, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 3 |
| 2009 | Devanagari OCR using a recognition driven segmentation framework and stochastic language models
Suryaprakash Kompalli, Srirangaraj Setlur, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 3 |
| 2009 | Automatic recognition of handwritten medical forms for search engines
Robert Milewski, Venu Govindaraju, Anurag Bhardwaj |
Int. J. Document Anal. Recognit. | 2 |
| 2009 | Preprocessing of Low-Quality Handwritten Documents Using Markov Random FieldsabstractThis paper presents a statistical approach to the preprocessing of degraded handwritten forms including the steps of binarization and form line removal. The degraded image is modeled by a Markov Random Field (MRF) where the hidden-layer prior probability is learned from a training set of high-quality binarized images and the observation probability density is learned on-the-fly from the gray-level histogram of the input image. We have modified the MRF model to drop the preprinted ruling lines from the image. We use the patch-based topology of the MRF and Belief Propagation (BP) for efficiency in processing. To further improve the processing speed, we prune unlikely solutions from the search space while solving the MRF. Experimental results show higher accuracy on two data sets of degraded handwritten images than previously used methods. Huaigu Cao, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | A probabilistic method for keyword retrieval in handwritten document images
Huaigu Cao, Anurag Bhardwaj, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2009 | Phrase-based correction model for improving handwriting recognition accuracies
Faisal Farooq, Damien Jose, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2009 | Synthetic handwritten CAPTCHAs
Achint Oommen Thomas, Amalia I. Rusu, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2008 | Lexicon Reduction in Handwriting Recognition Using Topic CategorizationabstractDespite several decades of research in handwriting recognition, the goal of having computers access handwritten information from unconstrained document images is still elusive. Current handwriting recognition systems are only capable of recognizing words that are present in a restricted lexicon typically comprised of 10 to 1000 words. As the size of the lexicon grows, the recognition accuracy falls sharply and is reported to be around 30% for a10K word lexicon. The objective of this research is to raise the accuracy levels on unconstrained handwritten documents by reducing the size of lexicons. We present an innovative method of lexicon reduction by topic categorization of handwritten documents. After categorization of a document into a topic e.g. sports, science etc. we use smaller lexicons that include only words with high mutual information with that topic and hence increase performance of recognizers. In this paper we present different techniques and report results on a publicly available dataset. Faisal Farooq, Gaurav Chandalia, Venu Govindaraju |
Document Analysis Systems | 3 |
| 2008 | Integrating minutiae based fingerprint matching with local mutual informationabstractMinutiae based fingerprint matching algorithms are wildly used in fingerprint identification and verification applications. However, they may suffer from spurious matches because they do not use the rich local image information. In this paper, we extend minutiae based methods to incorporate such local image information. Our method uses local mutual information, a proven similarity measure in various applications, to improve the matching rate. The overall minutiae distribution pattern between two fingerprints is represented by the initial minutiae matching result, while the mutual information measures the similarity between neighborhoods of matched minutiae, thus enhancing the final matching decision. FVC2002 DB1 and DB3 databases are used to test the proposed approach. Experimental result shows the improvement when combining minutiae matching scores with mutual information scores. Sergey Tulyakov, Faisal Farooq, Jason J. Corso, Venu Govindaraju |
ICPR | 5 |
| 2008 | Script Independent Word Spotting in Multilingual Documents
Anurag Bhardwaj, Damien Jose, Venu Govindaraju |
IJCNLP | 3 |
| 2008 | Biometrics Driven Smart Environments: Abstract Framework and Evaluation
Vivek Menon, Bharat Jayaraman, Venu Govindaraju |
UIC | 3 |
| 2008 | Binarization and cleanup of handwritten text from carbon copy medical form images
Robert Milewski, Venu Govindaraju |
Pattern Recognit. | 2 |
| 2008 | Use of Identification Trial Statistics for the Combination of Biometric MatchersabstractCombination functions typically used in biometric identification systems consider as input parameters only those matching scores which are related to a single person in order to derive a combined score for that person. We discuss how such methods can be extended to utilize the matching scores corresponding to all people. The proposed combination methods account for dependencies between scores output by any single participating matcher. Our experiments demonstrate the advantage of using such combination methods when dealing with a large number of classes, as is the case with biometric person identification systems. The experiments are performed on the National Institute of Standards and Technology BSSR1 dataset and the combination methods considered include the likelihood ratio, neural network, and weighted sum. Sergey Tulyakov, Venu Govindaraju |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2007 | Handwritten Carbon Form Preprocessing Based on Markov Random FieldabstractThis paper proposes a statistical approach to degraded handwritten form image preprocessing including binarization and form line removal. The degraded image is modeled by a Markov random field (MRF) where the prior is learnt from a training set of high quality binarized images, and the probabilistic density is learnt on-the-fly from the gray-level histogram of input image. We also modified the MRF model to implement form line removal. Test results of our approach show excellent performance on the data set of handwritten carbon form images. Huaigu Cao, Venu Govindaraju |
CVPR | 2 |
| 2007 | Facial Expression Biometrics Using Tracker Displacement FeaturesabstractIn this paper we investigate a possibility of using the face expression information for person biometrics. The idea of this research is that person's emotional face expressions are repeatable, and face expression features can be used for person identification. In order to avoid using person specific geometric or textural features traditionally used in face biometrics, we restrict ourselves to the tracker displacement features only. In contrast to previous research in facial expression biometrics, we extract features only from the pair of face images, neutral and the apex of emotion expression, instead of using the sequence of images from the video. The experiments, performed on two facial expression databases, confirm that proposed features can indeed be used for biometrics purposes. Sergey Tulyakov, Thomas E. Slowe, Venu Govindaraju |
CVPR | 4 |
| 2007 | Real-time Automatic Deceit Detection from Involuntary Facial ExpressionsabstractBeing the most broadly used tool for deceit measurement, the polygraph is a limited method as it suffers from human operator subjectivity and the fact that target subjects are aware of the measurement, which invites the opportunity to alter their behavior or plan counter-measures in advance. The approach presented in this paper attempts to circumvent these problems by unobtrusively and automatically measuring several prior identified deceit indicators (DIs) based upon involuntary, so-called reliable facial expressions through computer vision analysis of image sequences in real time. Reliable expressions are expressions said by the psychology community to be impossible for a significant percentage of the population to convincingly simulate, without feeling a true inner felt emotion. The strategy is to detect the difference between those expressions which arise from internal emotion, implying verity, and those expressions which are simulated, implying deceit. First, a group of facial action units (AUs) related to the reliable expressions are detected based on distance and texture based features. The DIs then can be measured and finally a decision of deceit or verity will be made accordingly. The performance of this proposed approach is evaluated by its real time implementation for deceit detection. Vartika Singh, Thomas E. Slowe, Sergey Tulyakov, Venu Govindaraju |
CVPR | 5 |
| 2007 | Vector Model Based Indexing and Retrieval of Handwritten Medical FormsabstractA vector model based information retrieval of handwritten medical forms is presented in this paper. In order to improve the IR performance on the erroneous output of handwriting recognition (HR) systems, a variation of the vector model is made to estimate the number of occurrences of terms from word segmentation and recognition probabilities. IR Tests show that our approach outperforms the retrieval of ordinary HR text in terms of mean average precision (MAP), R-Precision, and interpolated 11-point precisions. Huaigu Cao, Venu Govindaraju |
ICDAR | 2 |
| 2007 | PDE-Based Enhancement of Low Quality DocumentsabstractPartial Differential Equations are becoming one of the core tools for low-level image processing. They are especially functional in diffusion processes and variational models. In this paper, we exploit the regional smoothing that occurs in a nonlinear diffusion process and use this to enhance text in a degraded document image. The proposed smoothing method is robust when applied to either a highly corrupted text document or one with little degradation. The technique was tested on historical documents, carbon copies with highly varying grayscale backgrounds and on synthetic noisy documents. The PDE-based method far outperformed other industry-standard binarization techniques when compared quantitatively and qualitatively. Ifeoma Nwogu, Zhixin Shi, Venu Govindaraju |
ICDAR | 3 |
| 2007 | Generalized regression model for sequence matching and clustering
Venu Govindaraju |
Knowl. Inf. Syst. | 2 |
| 2007 | Fingerprint enhancement using STFT analysis
Sharat Chikkerur, Alexander N. Cartwright, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2007 | Symmetric hash functions for secure fingerprint biometric systems
Sergey Tulyakov, Faisal Farooq, Praveer Mansukhani, Venu Govindaraju |
Pattern Recognit. Lett. | 4 |
| 2007 | Introduction to the Special Issue on Recent Advances in Biometric Systems [Guest Editorial]abstractThe fourteen papers in this special section are devoted to recent advancements in biometric systems and application devices. K. W. Boyer, Venu Govindaraju, Nalini K. Ratha |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Extraction of Handwritten Text from Carbon Copy Medical Form Images
Robert Milewski, Venu Govindaraju |
Document Analysis Systems | 2 |
| 2006 | Singularity Preserving Fingerprint Image Adaptive FilteringabstractAccurate and reliable detection of minutiae from the fingerprint images is an important factor in the performance of automatic fingerprint identification systems (AFIS). Fingerprint image quality evaluation and appropriate enhancement technique are critical steps for accuracy of minutiae detection algorithm. Most fingerprint enhancement algorithms rely heavily on local orientation of ridge flows. However, significant orientation changes occur around the delta and core points in the fingerprint images, and this poses a challenge to the enhancement of ridge flows in those high-curvature regions. Instead of identifying the singular points, we calculate an orientation coherence map and determine minimum coherence regions as high-curvature areas. Gaussian filter window sizes are adaptively chosen to smooth the local orientation map. Because the smoothing operation is applied to local ridge shape structures, it efficiently joins broken ridges without destroying essential singularities and enforces continuity of directional fields even in creases. To the best the authors' knowledge, the coherence has not previously been used to estimate ridge curvature, and curvature has not been used to select filter scale in this field. These two strategies are the primary contributions of this paper. Experimental results demonstrate the effectiveness of the proposed method. Chaohong Wu, Venu Govindaraju |
ICIP | 2 |
| 2006 | Offline Arabic Handwriting Recognition: A SurveyabstractThe automatic recognition of text on scanned images has enabled many applications such as searching for words in large volumes of documents, automatic sorting of postal mail, and convenient editing of previously printed documents. The domain of handwriting in the Arabic script presents unique technical challenges and has been addressed more recently than other domains. Many different methods have been proposed and applied to various types of images. This paper provides a comprehensive review of these methods. It is the first survey to focus on Arabic handwriting recognition and the first Arabic character recognition survey to provide recognition rates and descriptions of test data for the approaches discussed. It includes background on the field, discussion of the methods, and future research directions. Liana M. Lorigo, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Hidden Markov Models Combining Discrete Symbols and Continuous Attributes in Handwriting RecognitionabstractPrior arts in handwritten word recognition model either discrete features or continuous features, but not both. This paper combines discrete symbols and continuous attributes into structural handwriting features and model, them by transition-emitting and state-emitting hidden Markov models. The models are rigorously defined and experiments have proven their effectiveness. Hanhong Xue, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | A chaincode based scheme for fingerprint feature extraction
Zhixin Shi, Venu Govindaraju |
Pattern Recognit. Lett. | 2 |
| 2005 | Pre-processing Methods for Handwritten Arabic DocumentsabstractIn order to improve the readability and the automatic recognition of handwritten document images, preprocessing steps are imperative. These steps in addition to conventional steps of noise removal and filtering include text normalization such as baseline correction, slant normalization and skew correction. These steps make the feature extraction process more reliable and effective. Recently Arabic handwriting recognition has received some attention from the research community. Due to the unique nature of the script, the conventional methods do not prove to be effective. In our work, we describe an orientation independent technique for baseline detection of Arabic words. In addition to that we describe, in the rest of the paper, our techniques for slant normalization, slope correction, line and word separation in handwritten Arabic documents. We show how the baseline can be exploited for slope and skew correction before proceeding with the steps of line and word separation. Faisal Farooq, Venu Govindaraju, Michael Perrone |
ICDAR | 2 |
| 2005 | A New Feature Ranking Method in a HMM-Based Handwriting Recognition SystemabstractIn this paper, we propose a new feature ranking method in a recognition system, by introducing the concept of the effectiveness of the distinguishing power of features and considering the correlation among features. To find the subset of most important features, first, the best feature can be identified by its effective distinguishing power and put in an empty feature set. Then, each of the remaining features is ranked based on their effective distinguishing capacity contribution and the highest-ranked feature is added to the selected subset. This process is repeated till the performance of the system reaches its peak or the effective distinguishing contribution falls below a certain value. The application of this method to an existing handwriting recognition system showed strong support for our methodology of feature ranking. Sijun Kang, Venu Govindaraju |
ICDAR | 2 |
| 2005 | Challenges in OCR of Dev anagari DocumentsabstractOCR of Devanagari script presents a wide range of challenges that are not seen in Latin based scripts. This paper outlines the implementation of a neural network based Devanagari OCR. Experimental results on a standard data set are reported and analyzed. Suryaprakash Kompalli, Sankalp Nayak, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 4 |
| 2005 | Similarity-driven Sequence Classification Based on Support Vector MachinesabstractA novel sequence classification method is proposed in the context of support vector machines (SVM). This method is driven by an intuitive similarity measure, namely ER/sup 2/, which directly tells the similarity of two sequences (1- or multi-dimensional). If sequence X is very similar to Y (for instance, the similarity by ER/sup 2/ is above 90%), it is safe to assign X to the same class as Y. ER/sup 2/ is plugged into standard SVM to speed up the decision-making of multi-class classification. The immediate application of the method is in the adaptive online handwriting recognition, where handwritten characters are represented by 2D sequences of X-, Y-coordinates. Experiments on the benchmark database UNIPEN show that the classification driven by ER/sup 2/ can be about three times faster than standard SVM while the classification accuracy is enhanced or comparable. Venu Govindaraju |
ICDAR | 2 |
| 2005 | Segmentation and Pre-Recognition of Arabic HandwritingabstractWe propose a novel algorithm for the segmentation and prerecognition of offline handwritten Arabic text. Our character segmentation method over-segments each word, and then removes extra breakpoints using knowledge of letter shapes. On a test set of 200 images, 92.3% of the segmentation points were detected correctly, with 5.1% instances of over-segmentation. The prerecognition component annotates each detected letter with shape information, to be used for recognition in future work. Liana M. Lorigo, Venu Govindaraju |
ICDAR | 2 |
| 2005 | A Lexicon Reduction Strategy in the Context of Handwritten Medical FormsabstractTraditional handwriting recognition algorithms rely heavily on small lexicons and clean word images. Unfortunately, emergency medical documents do not satisfy either of these conditions. This is a significant road-block that is hampering efforts to rapidly convert valuable offline healthcare handwriting data into digital content that can be efficiently mined for information. This paper describes a strategy whereby given an image representing a noisy handwritten word from a medical document, and a large lexicon consisting of English, medical and pharmacological words, symbols, abbreviations and acronyms, significantly reduces the size of the lexicon while keeping the unknown desired entry within the lexicon. The approach combines geometric interpretations of the word image along with contextual inference of concepts to reduce lexicons for word recognition. The data extracted can then be efficiently and securely disseminated for epidemiological and outbreak detection/analysis. Experimental results on NY State PCR forms are reported. Robert Milewski, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2005 | A Human Interactive Proof Algorithm Using Handwriting RecognitionabstractThe recognition of unconstrained handwriting continues to be a difficult task for computers despite active research for several decades. This is because handwritten text offers great challenges such as: character and word segmentation, character recognition, variation between handwriting styles, different character size and orientation, no font constraints, the type of printing surface, as well as the background clarity. In this paper, we explore the gap in the ability in reading handwritten text between humans and computers to propose solutions for security problems in Web services. We present a new HIP algorithm that uses handwriting recognition task to distinguish between humans and computers. We propose methods to deform handwritten text images to make them indecipherable by computers and explore the cognitive factors that assist humans in reading and understanding. Experimental results on both humans and computers are presented and compared. Amalia I. Rusu, Venu Govindaraju |
ICDAR | 2 |
| 2005 | Multi-scale Techniques for Document Page SegmentationabstractPage segmentation algorithms found in published literatures often rely on some predetermined parameters such as general font sizes, distances between text lines and document scan resolutions. Variations of these parameters in real document images greatly affect the performance of the algorithms. In this paper, we present a novel approach for document page segmentation using a multi-scale technique. An efficient implementation of a local connectivity algorithm transforms a document image into a parameter domain in which a parameter value at a pixel location represents a connectivity property for its neighboring foreground pixels in the original document image. Then a top-down approach with a linear search reveals the document regions at each scale levels as text block, text lines and graphics. We consider our algorithm a transform based multi-scale method. Our ongoing research shows that the algorithm is robust for variations of document parameters. Zhixin Shi, Venu Govindaraju |
ICDAR | 2 |
| 2005 | Text Extraction from Gray Scale Historical Document Images Using Adaptive Local Connectivity MapabstractThis paper presents an algorithm using adaptive local connectivity map for retrieving text lines from the complex handwritten documents such as handwritten historical manuscripts. The algorithm is designed for solving the particularly complex problems seen in handwritten documents. These problems include fluctuating text lines, touching or crossing text lines and low quality image that do not lend themselves easily to binarizations. The algorithm is based on connectivity features similar to local projection profiles, which can be directly extracted from gray scale images. The proposed technique is robust and has been tested on a set of complex historical handwritten documents such as Newton's and Galileo's manuscripts. A preliminary testing shows a successful location rate of above 95% for the test set. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2005 | Combining Matching Scores in Identification ModelabstractThe paper discusses a problem of combining recognition scores for different classes produced by one recognizer during one recognition attempt. This problem arises in identification problems which we define as 1:N classification problems with big or variable N. By using artificial example we show that intuitive solution of making identification decision based solely on the best matching score is frequently suboptimal. Paper presents reasons for such behavior, and draws parallels with score normalization technique used in speaker identification. Two examples of real life applications illustrate the possible benefits of properly combining recognition scores. Sergey Tulyakov, Venu Govindaraju |
ICDAR | 2 |
| 2005 | A minutia-based partial fingerprint recognition system
Tsai-Yang Jea, Venu Govindaraju |
Pattern Recognit. | 2 |
| 2005 | GP-based secondary classifiers
Ankur Teredesai, Venu Govindaraju |
Pattern Recognit. | 2 |
| 2005 | A comparative study on the consistency of features in on-line signature verification
Venu Govindaraju |
Pattern Recognit. Lett. | 2 |
| 2005 | Matching and retrieving sequential patterns using regression
Venu Govindaraju |
Web Intell. Agent Syst. | 2 |
| 2004 | Handwriting Analysis of Pre-Hospital Care ReportsabstractEmergency health care facilities lack automated medical form recognition systems necessary for efficient epidemiological and health surveillance analysis. The task is to extract handwritten text from the New York State (NYS) Pre-Hospital Care Report (PCR) and determine its ASCII translation. Our approach hybridizes image processing and semantic lexicon pruning to compensate for the otherwise enormous lexicon size. In this paper we expand on our IEEE CBMS 2001 paper, which provided a conceptual overview, by probing into our recognizer design and performance measurements. Robert Milewski, Venu Govindaraju |
CBMS | 2 |
| 2004 | Issues in evolving GP based classifiers for a pattern recognition taskabstractThis paper discusses issues when evolving genetic programming (GP) classifiers for a pattern recognition task such as handwritten digit recognition. Developing elegant solutions for handwritten digit classification is a challenging task. Similarly, design and training of classifiers using genetic programming is a relatively new approach in pattern recognition as compared to other traditional techniques. Several strategies for GP training are outlined and the empirical observations are reported. The issues we faced such as training time, a variety of fitness landscapes and accuracy of results are discussed. Care has been taken to test GP using a variety of parameters and on several handwritten digits datasets. Ankur Teredesai, Venu Govindaraju |
IEEE Congress on Evolutionary Computation | 2 |
| 2004 | Document Analysis Systems for Digital Libraries: Challenges and Opportunities
Henry S. Baird, Venu Govindaraju, Daniel P. Lopresti |
Document Analysis Systems | 2 |
| 2004 | DL Architecture for Indic Scripts
Suryaprakash Kompalli, Srirangaraj Setlur, Venu Govindaraju |
Document Analysis Systems | 3 |
| 2004 | Data Mining for Intrusion Detection: Techniques, Applications and SystemsabstractAn intrusion is defined as any set of actions that compromise the integrity, confidentiality or availability of a resource. Intrusion detection is an important task for information infrastructure security. One major challenge in intrusion detection is that we have to identify the camouflaged intrusions from a huge amount of normal communication activities. Data mining is to identify valid, novel, potentially useful, and ultimately understandable patterns in massive data. It is demanding to apply data mining techniques to detect various intrusions. In the last several years, some exciting and important advances have been made in intrusion detection using data mining techniques. Research results have been published and some prototype systems have been established. Inspired by the huge demands from applications, the interactions and collaborations between the communities of security and data mining have been boosted substantially. This seminar will present an interdisciplinary survey of data mining techniques for intrusion detection so that the researchers from computer security and data mining communities can share the experiences and learn from each other. Some data mining based intrusion detection systems will also be reviewed briefly. Moreover, research challenges and problems will be discussed so that future collaborations may be stimulated. For data mining/database researchers and practitioners, the seminar will provide background knowledge and opportunities for applying data mining techniques to intrusion detection and computer security. For computer security researchers and practitioners, it provides knowledge on how data mining can benefit and enhance computer security. We will try to understand and appreciate the following technical issues. Jian Pei 0001, Shambhu J. Upadhyaya, Faisal Farooq, Venu Govindaraju |
ICDE | 4 |
| 2004 | Matching and Retrieving Sequential Patterns Under RegressionabstractSequential pattern matching and retrieving is of real value. For example, finding stocks in the NASDAQ market whose closing prices are always about $β₀ higher than or β₁ times as that of a given company. The probelm reduces to linear pattern retrieval: given query X, find all sequence Y from database S so that Y = β₀ + β₁ with confidence C. In this paper, we novelly introduce SLR (Simple Linear Regression) model [5,7] to solve this problem. We extend 1-dimensional R^2 to ER^2 for multi-dimensional sequence matching, such as on-line handwritten signature. In addition, we develop SLR+FFT pruning techniques based on SLR to speed up retrieval without incurring any false dismissal. Experimental results show that the pruning ratio of SLR+FFT is efficient (can be above 99%). Experiments on real stocks discovered many interesting patterns. Preliminary test on on-line signature recognition using ER^2 as similarity measure also shows high accuracy. Venu Govindaraju |
Web Intelligence | 2 |
| 2003 | Postal address block location by contour clusteringabstractWe have developed a well performing algorithm for locating address blocks in postal parcel images. Both machine printed and handwritten addresses are processed by the algorithm. The algorithm is invariant to the image orientation and scale, and it works with high noise images. It could also serve as an additional step after other address block location algorithms. Venu Govindaraju, Sergey Tulyakov |
ICDAR | 1 |
| 2003 | Text - Image Separation in Devanagari DocumentsabstractIn this paper we present a top-down, projection-profile based algorithm to separate text blocks from image blocks in a Devanagari document. We use a distinctive feature of Devanagari text, called Shirorekha (Header Line) to analyze the pattern produced by Devanagari text in the horizontal profile. The horizontal profile corresponding to a text block possesses certain regularity in frequency, orientation and shows spatial cohesion. The algorithm uses these features to identify text blocks in a document image containing both text and graphics. Swapnil Khedekar, Vemulapati Ramanaprasad, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 4 |
| 2003 | Skew Detection for Complex Document Images Using Fuzzy RunlengthabstractA skew angle estimation approach based on the application of a fuzzy directional runlength is proposed for complex address images. The proposed technique was tested on a variety of USPS parcel images including both machine print and handwritten addresses. The testing results showed a successful rate more than 90% of the test set. Zhixin Shi, Venu Govindaraju |
ICDAR | 2 |
| 2002 | A Stochastic Model Combining Discrete Symbols and Continuous Attributes and Its Application to Handwriting Recognition
Hanhong Xue, Venu Govindaraju |
Document Analysis Systems | 2 |
| 2002 | Large scale address recognition systems Truthing, testing, tools, and other evaluation issues
Srirangaraj Setlur, Alfred Lawson, Venu Govindaraju, Sargur N. Srihari |
Int. J. Document Anal. Recognit. | 3 |
| 2002 | Use of Lexicon Density in Evaluating Word RecognizersabstractWe have developed the notion of lexicon density as a metric to measure the expected accuracy of handwritten word recognizers. Thus far, researchers have used the size of the lexicon as a gauge for the difficulty of the handwritten word recognition task. For example, the literature mentions recognizers with accuracies for lexicons of sizes 10, 100, 1000, and so forth, implying that the difficulty of the task increases (and hence recognition accuracy decreases) with increasing lexicon size across recognizers. Lexicon density is an alternate measure which is quite dependent on the recognizer. There are many applications, such as address interpretation, where such a recognizer-dependent measure can be useful. We have conducted experiments with two different types of recognizers. A segmentation-based and a grapheme-based recognizer have been selected to show how the measure of lexicon density can be developed in general for any recognizer. Experimental results show that the lexicon density measure described is more suitable than lexicon size or a simple string edit distance. Venu Govindaraju, Petr Slavík, Hanhong Xue |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | On the Dependence of Handwritten Word Recognizers on LexiconsabstractThe performance of any word recognizer depends on the lexicon presented. Usually, large lexicons or lexicons containing similar entries pose difficulty for recognizers. However, the literature lacks any quantitative methodology of capturing the precise dependence between word recognizers and lexicons. This paper presents a performance model that views word recognition as a function of character recognition and statistically "discovers" the relation between a word recognizer and the lexicon. It uses model parameters that capture a recognizer's ability of distinguishing characters (of the alphabet) and its sensitivity to lexicon size. These parameters are determined by a multiple regression model which is derived from the performance model. Such a model is very useful in comparing word recognizers by predicting their performance based on the lexicon presented. We demonstrate the performance model with extensive experiments on five different word recognizers, thousands of images, and tens of lexicons. The results show that the model is a good fit not only on the training data but also in predicting the recognizers' performance on testing data. Hanhong Xue, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Use of adaptive segmentation in handwritten phrase recognition
Jaehwa Park, Venu Govindaraju |
Pattern Recognit. | 2 |
| 2002 | Improved k-nearest neighbor classification
Yingquan Wu, Krassimir G. Ianakiev, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2001 | Automated Reading and Mining of Pre-Hospital Care ReportsabstractThe lack of use of high technology in the healthcare delivery system is especially apparent in the emergency medical information systems (EMIS) area. For example, in New York State, all patients who enter the emergency medical service (EMS) are tracked through their pre-hospital care to the emergency room using a pre-hospital care report (PCR). Our goal is to automate the collection of data from the PCR and enable efficient maintenance and dissemination of information. The task involves the automatic extraction and transliteration of the handwritten text in the response boxes provided on the form. The objective is to produce, for each form image, an ASCII transcription of the handwritten contents of the response boxes. These responses could be then used to populate a database. The database itself would then emerge as a valuable resource for enabling data mining and knowledge discovery for the entire medical community. Venu Govindaraju, Robert Milewski |
CBMS | 1 |
| 2001 | Active Handwritten Character Recognition Using Genetic Programming
Ankur Teredesai, Jaehwa Park, Venu Govindaraju |
EuroGP | 3 |
| 2001 | Truthing, Testing and Evaluation Issues in Complex SystemsabstractThis paper describes the issues involved in the design of a system for evaluating improvements in the performance of a real-time address recognition system being used by the United States Postal Service for processing mail-piece images. Evaluation of the performance of recognition systems is normally carried out by measuring the performance of the system on a representative sample of images. Designing a comprehensive and valid testing scenario is a complex task that requires careful attention. Sampling live mail-stream to generate a deck of images representative of the general mail-stream for testing, truthing (generating reference data on a significant number of images), grading and evaluation, and designing tools to facilitate these functions are important topics that need to be addressed. This paper describes the efforts of the United States Postal Service and CEDAR towards developing an infrastructure for sampling, truthing and testing of mail-stream images. Srirangaraj Setlur, Venu Govindaraju, Sargur N. Srihari, Alfred Lawson |
ICDAR | 2 |
| 2001 | Active Digit Classifiers: A Separability Optimization Approach to Emulate CognitionabstractGiven sufficient resources, any classification task is possible with a high accuracy, but to achieve a particular task given finite resources, the problem is to utilize these resources intelligently. Cognitive studies in human vision associate multi-resolution features with high recognition accuracy. We show that classifier development using separability optimization is very similar to emulation of human cognition. The identification of key features leads to optimal resource utilization by the classifier. Evolving such classifiers is the focus of the paper. The resources required for classification can be identified in terms of amount of time required to develop a recognizer amount of processing power required and the number and kind of features extracted. Our digit recognition method strives not only to report high accuracy but also targets generation of simple solutions. The simplicity of a solution can be a measure of the resources utilized. Our methodology is termed as active based on the premise that once the complexity of a classification task is known an intelligent recognizer should incrementally increase the resources needed for classification. Ankur Teredesai, Venu Govindaraju |
ICDAR | 2 |
| 2001 | Probabilistic Model for Segmentation Based Word Recognition with LexiconabstractWe describe the construction of a model for off-line word recognizers based on over-segmentation of the input image and recognition of segment combinations as characters in a given lexicon word. One such recognizer, the Word Model Recognizer (WMR), is used extensively. Based on the proposed model it was possible to improve the performance of WMR. Sergey Tulyakov, Venu Govindaraju |
ICDAR | 2 |
| 2001 | Building Skeletal Graphs for Structural Feature Extraction on Handwriting ImagesabstractPresents a method of building skeletal graphs for handwriting images, aiming at extraction of high-level structural features such as loops, turns, ends, and junctions. Block adjacency graphs are used as the base representation and transformed at locations where deformation occurs to obtain satisfactory skeletal graphs. Then the identification and ordering of structural features are considered based on skeletal graphs. Hanhong Xue, Venu Govindaraju |
ICDAR | 2 |
| 2001 | The Role of Holistic Paradigms in Handwritten Word RecognitionabstractThe holistic paradigm in handwritten word recognition treats the word as a single, indivisible entity and attempts to recognize words from their overall shape, as opposed to their character contents. In this survey, we have attempted to take a fresh look at the potential role of the holistic paradigm in handwritten word recognition. The survey begins with an overview of studies of reading which provide evidence for the existence of a parallel holistic reading process,in both developing and skilled readers. In what we believe is a fresh perspective on handwriting recognition, approaches to recognition are characterized as forming a continuous spectrum based on the visual complexity of the unit of recognition employed and an attempt is made to interpret well-known paradigms of word recognition in this framework. An overview of features, methodologies, representations, and matching techniques employed by holistic approaches is presented. Sriganesh Madhvanath, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Equivalence of Different Methods for Slant and Skew Corrections in Word Recognition ApplicationsabstractNormalization of slant and skew is often used in processing a word image before recognition. In this paper, we prove the theoretical equivalence of different methods for slant and skew corrections. In particular, we show that correcting first for skew by rotation and then for slant by a shear transformation in the horizontal direction is equivalent to first correcting for slant by a shear transformation in the horizontal direction and then for skew by a shear transformation in the vertical direction. Our proof can be easily modified to prove equivalence of other methods for correcting the slant and skew. Petr Slavík, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Syntactic methodology of pruning large lexicons in cursive script recognition
Sriganesh Madhvanath, Venu Krpasundar, Venu Govindaraju |
Pattern Recognit. | 3 |
| 2000 | Active Character Recognition Using 'A*-Like' AlgoritabstractThis paper describes an Active Character Recognition methodology, henceforth referred to as ACR. We present in this paper a method that uses an active heuristic function similar to the one used by A* search algorithm that adaptively determines the length of the feature vector as well as the features themselves used to classify an input pattern. ACR adapts to factors such as the quality of the input pattern, its intrinsic similarities and differences from patterns of other classes it is being compared against and the processing time available. Furthermore, the finer resolution is accorded to only certain "zones" of the input pattern which are deemed important given the classes that are being discriminated. Experimental results support the methodology presented. Recognition rate of ACR is about 96% on the NIST data sets and the speed is better than traditional classification methods. Jaehwa Park, Venu Govindaraju |
CVPR | 2 |
| 2000 | Using Lexical Similarity in Handwritten Word RecognitionabstractRecognition using only visual evidence cannot always be successful due to limitations of information and resources available during training. Considering relation among lexicon entries is sometimes useful for decision making. In this paper we present a method to capture lexical similarity of a lexicon and reliability of a character recognizer which serve to capture the dynamism of the environment. A parameter, lexical similarity, is defined by measuring these two factors as edit distance between lexicon entries and separability of each character's recognition results. Our experiments show that a utility function considering lexical similarity in a decision stage can enhance the performance of a conventional word recognizer. Jaehwa Park, Venu Govindaraju |
CVPR | 2 |
| 2000 | OCR in a Hierarchical Feature SpaceabstractThis paper describes hierarchical OCR, a character recognition methodology that achieves high speed and accuracy by using a multiresolution and hierarchical feature space. Features at different resolutions, from coarse to fine-grained, are implemented by means of a recursive classification scheme. Typically, recognizers have to balance the use of features at many resolutions (which yields a high accuracy), with the burden on computational resources in terms of storage space and processing time. We present in this paper, a method that adaptively determines the degree of resolution necessary in order to classify an input pattern. This leads to optimal use of computational resources. The hierarchical OCR dynamically adapts to factors such as the quality of the input pattern, its intrinsic similarities and differences from patterns of other classes it is being compared against, and the processing time available. Furthermore, the finer resolution is accorded to only certain "zones" of the input pattern which are deemed important given the classes that are being discriminated. Experimental results support the methodology presented. When tested on standard NIST data sets, the hierarchical OCR proves to be 300 times faster than a traditional K-nearest-neighbor classification method, and 10 times taster than a neural network method. The comparison uses the same feature set for all methods. Recognition rate of about 96 percent is achieved by the hierarchical OCR. This is at par with the other two traditional methods. Jaehwa Park, Venu Govindaraju, Sargur N. Srihari |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | Holistic recognition of handwritten character pairs
Venu Govindaraju, Sargur N. Srihari |
Pattern Recognit. | 2 |
| 2000 | Potential improvement of classifier accuracy by using fuzzy measuresabstractTypical digit recognizers classify an unknown digit pattern by computing its distance from the cluster centers in a feature space. In this paper, we propose a methodology that has many salient aspects. First, the classification rule is dependent on the "difficulty" of the unknown sample. Samples "far" from the center, which tend to fall on the boundaries of classes are error prone and, hence, "difficult". An "overlapping zone" is defined in the feature space to identify such difficult samples. A table is precomputed to facilitate an efficient lookup of the class corresponding to all the points in the overlapping zone. The lookup function itself is defined by a modification of the KNN rule. A characteristic function defining the new boundaries is computed using the topology of the set of samples in the overlapping zones. Our two-pronged approach uses different classification schemes with the "difficult" and "easy" samples. The method described has improved the performance of the gradient structural concavity digit recognizer described by Favata et al. (1996). Venu Govindaraju, Krassimir G. Ianakiev |
IEEE Trans. Fuzzy Syst. | 1 |
| 1999 | Recognition of Strings Using Nonstationary Markovian Models: An Application in ZIP Code RecognitionabstractThis paper presents nonstationary Markovian models and their application to recognition of strings of tokens, such as ZIP codes in the US mailstream. Unlike traditional approaches where digits are simply recognized in isolation, the novelty of our approach lies in the manner in which recognitions scores along with domain specific knowledge about the frequency distribution of various combination of digits are all integrated into one unified model. The domain knowledge is derived from postal directory files. This data feeds into the models as n-grams statistics that are seamlessly integrated with recognition scores of digit images. We present the recognition accuracy (90%) achieved on a set of 20,000 ZIP codes. Djamel Bouchaffra, Venu Govindaraju, Sargur N. Srihari |
CVPR | 2 |
| 1999 | Efficient Word Segmentation Driven by Unconstrained Handwritten Phrase RecognitionabstractAn efficient system which finds the best match between an input image and a lexicon is presented. To capture writing style of spacing between words and characters prime stroke analysis based on statistical methods is introduced. A method for estimating bound on number of characters without actual recognition is also presented. For system efficiency, before actual recognition, classified groups of word segments and eligible subset of lexicons are generated as hypotheses. The hypotheses are verified and ordered by a lexicon driven word recognizor. We have tested our approach in the street name recognition/interpretation for US mail stream. Experimental results and encouraging. Jaehwa Park, Venu Govindaraju, Sargur N. Srihari |
ICDAR | 2 |
| 1999 | Information Theoretic Analysis of Postal Address Fields for Automatic Address InterpretationabstractThis paper concerns a study of information content in postal address fields for automatic address interpretation. Information provided by a combination of address components and information interaction among components is characterized in terms of Shannon's entropy. The efficiency of assignment strategies for determining a delivery point code can be compared by the propagation of uncertainty in address components. The quantity of redundancy between components can be computed from the information provided by these components. This information is useful in developing a strategy for selecting a useful component for recovering the value of an uncertain component. The uncertainty of a component based on another known component can be measured by conditional entropy. By ranking the uncertainty quantity, the effective processing flow for determining the value of a candidate component can be constructed. Sargur N. Srihari, Wen-jann Yang, Venu Govindaraju |
ICDAR | 3 |
| 1999 | Multi-experts for Touching Digit String Recognitionabstract84.6% of touching digit strings have only two digits touching, 12.3% have three digits touching and 3.1% have more than three digits touching. We present a multi-expert approach to recognize touching digit pairs (TDP) and touching digit triples (TDT). We combine holistic and traditional segmentation methods. 25,686 TDP training samples and 2,778 TDP testing samples collected from USPS mail are used in our experiment. The holistic method outperforms the traditional segmentation-based methods. The multi-expert combination has the best performance: a correct recognition rate of 91.1% on TDP. Venu Govindaraju, Sargur N. Srihari |
ICDAR | 2 |
| 1999 | An architecture for handwritten text recognition systems
Gyeonghwan Kim, Venu Govindaraju, Sargur N. Srihari |
Int. J. Document Anal. Recognit. | 2 |
| 1999 | A Methodology for Mapping Scores to ProbabilitiesabstractDescribes the derivation of probability of correctness from scores assigned by most recognizers. Derivation of probability values puts the output of different recognizers on the same scale; this makes comparison across recognizers trivial. The authors draw on examples from handwritten word recognition to illustrate their point. Djamel Bouchaffra, Venu Govindaraju, Sargur N. Srihari |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | Postprocessing of Recognized Strings Using Nonstationary Markovian ModelsabstractThis paper presents nonstationary Markovian models and their application to recognition of strings of tokens. Domain specific knowledge is brought to bear on the application of recognizing zip codes in the US mailstream by the use of postal directory files. These files provide a wealth of information on the delivery points (mailstops) corresponding to each zip code. This data feeds into the models as n-grams, statistics that are integrated with recognition scores of digit images. An especially interesting facet of the model is its ability to excite and inhibit certain positions in the n-grams leading to the familiar area of Markov random fields. We empirically illustrate the success of Markovian modeling in postprocessing applications of string recognition. We present the recognition accuracy of the different models on a set of 20000 zip codes. The performance is superior to the present system which ignores all contextual information and simply relies on the recognition scores of the digit recognizers. Djamel Bouchaffra, Venu Govindaraju, Sargur N. Srihari |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | Chaincode Contour Processing for Handwritten Word RecognitionabstractContour representations of binary images of handwritten words afford considerable reduction in storage requirements while providing lossless representation. On the other hand, the one-dimensional nature of contours presents interesting challenges for processing images for handwritten word recognition. Our experiments indicate that significant gains are to be realized in both speed and recognition accuracy by using a contour representation in handwriting applications. Sriganesh Madhvanath, Gyeonghwan Kim, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1999 | Holistic Verification of Handwritten PhrasesabstractIn this paper, we describe a system for rapid verification of unconstrained off-line handwritten phrases using perceptual holistic features of the handwritten phrase image. The system is used to verify handwritten street names automatically extracted from live US mail against recognition results of analytical classifiers. Presented with a binary image of a street name and an ASCII street name, holistic features (reference lines, large gaps and local contour extrema) of the street name hypothesis are "predicted" from the expected features of the constituent characters using heuristic rules. A dynamic programming algorithm is used to match the predicted features with the extracted image features. Classes of holistic features are matched sequentially in increasing order of cost, allowing an ACCEPT/REJECT decision to be arrived at in a time-efficient manner. The system rejects errors with 98 percent accuracy at the 30 percent accept level, while consuming approximately 20/msec per image on the average on a 150 MHz SPARC 10. Sriganesh Madhvanath, Evelyn Kleinberg, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1999 | Local reference lines for handwritten phrase recognition
Sriganesh Madhvanath, Venu Govindaraju |
Pattern Recognit. | 2 |
| 1998 | A Methodology for Deriving Probabilistic Correctness Measures from RecognizersabstractThis paper describes the derivation of probability of correctness from scores assigned by most recognizers. Motivation for this research is three-fold: (i) probability values can be used to rerank the output of any recognizer by using a new set of training data; if the training data is sufficiently large and representative of the test data, the recognition rates are seen to improve significantly, (ii) derivation of probability values puts the output of different recognizers on the same scale; this makes comparison across recognizers trivial, and (iii) word recognition can be readily extended to phrase and sentence recognition because the integration of language models becomes straightforward. We have conducted an extensive set of experiments. The results show a reranking of recognition choices based on the derived probability values leading to an enhancement in performance. Djamel Bouchaffra, Venu Govindaraju, Sargur N. Srihari |
CVPR | 2 |
| 1998 | Postal reply card processingabstractA reply card processing (RCP) service was developed at CEDAR to demonstrate how the United States Postal Service could deliver business and courtesy reply cards to clients electronically instead of physically. The electronic delivery of the more than one billion cards that the USPS processes in a year was to expedite the order fulfilment cycle for participating customers, resulting in cost savings and accelerated response times. CEDAR developed the necessary subsystems which included a postal image management system. As well as providing the networking applications for ensuring the seamless transfer of card images to the postal clients, the task involved advancing the state-of-the-art in forms processing to read the card images and to deliver the information in a form ready for database integration. David C. Bartnik, Venu Govindaraju, Sargur N. Srihari, Binh C. Phan |
ICPR | 2 |
| 1998 | OCR in a hierarchical feature spaceabstractThis paper describes a methodology that allows fast and accurate character recognition while keeping the dimensionality of the feature space relatively small. Higher dimensionality can add to the discriminatory power of a recognizer but pays the price in an increase of computational time. We present a method that achieves high accuracy even with a low-dimensional feature space by simulating a multiresolution feature space. Our approach is supported by promising experimental results. Recognition rate of 98% is achieved on a test set of about 16,000 handwritten numerals. Recognition rates on upper and lower case handprinted characters is about 95%. Jaehwa Park, Venu Govindaraju, Sargur N. Srihari |
SMC | 2 |
| 1998 | Handwritten phrase recognition as applied to street name images
Gyeonghwan Kim, Venu Govindaraju |
Pattern Recognit. | 2 |
| 1997 | Contour-based Image Preprocessing for Holistic Handwritten Word RecognitionabstractThe one-dimensional nature of contour representations presents interesting challenges for processing of images for handwritten word recognition. In this paper, we discuss the issues of determination of upper and lower contours of the word, determination of significant focal extrema on the contour, and determination of reference lines from contour representations of handwritten words. Sriganesh Madhvanath, Venu Govindaraju |
ICDAR | 2 |
| 1997 | The HOVER System for Rapid Holistic Verification of Off-lineHandwritten PhrasesabstractThe authors describe ongoing research on a system for rapid verification of unconstrained off-line handwritten phrases using perceptual holistic features of the handwritten phrase image. The system is used to verify handwritten street names automatically extracted from live US mail against recognition results of analytical classifiers. The system rejects errors with 98% accuracy at the 30% accept level, while consuming approximately 20 msec per image on the average on a 150 MHz SPARC 10. Sriganesh Madhvanath, Evelyn Kleinberg, Venu Govindaraju, Sargur N. Srihari |
ICDAR | 3 |
| 1997 | Bankcheck Recognition Using Cross Validation Between Legal and Courtesy AmountsabstractA bankcheck reading system using cross validation of both the legal and the courtesy amounts is presented in this paper. Some of the challenges posed by the task are (i) segmentation of the legal amount into words, (ii) location of boundaries between dollars and cents amounts, and (iii) high accuracy in terms of recognition performance. Word segmentation in the legal amount is a serious issue because of the nature of the data and patrons' writing habits which tend to clump words together. We have developed a word segmentation algorithm based on the character segmentation results to address this issue. The list of possible amounts generated by the word segmentation hypotheses is used as lexicon for the courtesy amount recognition. The order of magnitude of the amount is estimated during legal amount recognition. We treat the courtesy amount as a numeral string and apply the same word recognition scheme as used for the legal amount. Our approach to check recognition differs from traditional methods in two significant aspects: First, our emphasis on both the legal and the courtesy amounts is balanced. We use an accurate word recognizer which performs equally well on alpha words and digit strings. Second, our combination strategy is serial rather than the commonly used parallel method. Experimental results show that 43.8% of check images are correctly read with an error rate of 0%. Gyeonghwan Kim, Venu Govindaraju |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1997 | Empirical Design of A Multi-Classifier Thresholding/Control Strategy for Recognition of Handwritten Street NamesabstractA central task in the interpretation of handwritten US postal addresses is the off-line recognition of the street name. A lexicon of candidate street names may be extracted from a database of postal delivery points (DPF) by first locating and recognizing numeric fields such as the ZIP code and strewet number. The off-line handwritten word recognition (HWR) task is made difficult by the unconstrained, omni-scriptor nature of the input, and incomplete lexicons resulting from errors in processing numeric fields and intrinsic deficiencies in the DPF. In this paper, we describe an empirical approach to the design of a multi-classifier HWR Thresholding/Control module which forms part of a real-time handwritten address interpretation (HWAI) system. The decisions of two word classifiers are combined in a hierarchical manner to improve recognition and error-rejection performance, while meeting real-time requirements. The design employs logistic regression and agreement for evidence combination, and lexicon reduction for improved throughput as well as performance. The paper concludes with experimental results and directions for future research. Sriganesh Madhvanath, Evelyn Kleinberg, Venu Govindaraju |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1997 | A Lexicon Driven Approach to Handwritten Word Recognition for Real-Time ApplicationsabstractA fast method of handwritten word recognition suitable for real time applications is presented in this paper. Preprocessing, segmentation and feature extraction are implemented using a chain code representation of the word contour. Dynamic matching between characters of a lexicon entry and segment(s) of the input word image is used to rank the lexicon entries in order of best match. Variable duration for each character is defined and used during the matching. Experimental results prove that our approach using the variable duration outperforms the method using fixed duration in terms of both accuracy and speed. Speed of the entire recognition process is about 200 msec on a single SPARC-10 platform and the recognition accuracy is 96.8 percent are achieved for lexicon size of 10, on a database of postal words captured at 212 dpi. Gyeonghwan Kim, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Segmentation and recognition of connected handwritten numeral strings
Zhixin Shi, Venu Govindaraju |
Pattern Recognit. | 2 |
| 1996 | Recognition of Handwritten Phrases as Applied to Street Name ImagesabstractA method for recognition of street name phrases collected from mail pieces is presented in this paper. Some of the challenges posed by the problem are: (i) patron errors, (ii) non-standardized way of abbreviating names, and (iii) variable number of words in a street name image. A neural network has been designed to segment words in a phrase, a street name in this case, using distances between components and style of writing. The network learns the type of spacing (including size) that one should expect between different pairs of characters in handwritten test. Experiments show perfect word segmentation performance at about 85% of cases. Unlike conventional methods, where lexicon entries are expanded to take care of all variations of prefixes and sizes, substring matching is attempted only between the main body of a lexicon entry and the word segments of an image. Efforts to reduce computational complexity are successfully made by the sharing of character segmentation results between the segmentation and recognition phases. 83% phrase recognition accuracy is achieved on a test set. Gyeonghwan Kim, Venu Govindaraju |
CVPR | 2 |
| 1996 | A segmentation and recognition strategy for handwritten phrasesabstractA segmentation and recognition method for handwritten phrases, such as street names, is presented in this paper. Some of the challenges posed by the problem are: (1) identifying correct word gaps from character gaps and (2) minimization of computational complexity during the recognition of potential words. A trainable word segmentation scheme using a neural network is introduced. The network learns the type of spacing (including size) that one should expect between different pairs of characters in handwritten text. The concept of variable duration, which is obtained during the training phase of a word recognition engine we have developed, is expanded to reduce the computational complexity which has been a serious concern in this type of application. Gyeonghwan Kim, Venu Govindaraju, Sargur N. Srihari |
ICPR | 2 |
| 1996 | Locating human faces in photographs
Venu Govindaraju |
Int. J. Comput. Vis. | 1 |
| 1996 | Holistic handwritten word recognition using temporal features derived from off-line images
Venu Govindaraju, Ram Krishnamurthy 0001 |
Pattern Recognit. Lett. | 1 |
| 1996 | Character image enhancement by selective region-growing
Zhixin Shi, Venu Govindaraju |
Pattern Recognit. Lett. | 2 |
| 1995 | Handwritten word recognition for real-time applicationsabstractA fast handwritten word recognition system for real time applications is presented. Preprocessing, segmentation and feature extraction are implemented using chain code representation. Dynamic matching between each character of a lexicon entry and segment(s) of input word image is used for ranking words in the lexicon. Speed of the entire recognition process is about 200 msec on a single SPARC-10 platform for lexicon size of 10. A top choice performance of 96% is achieved on a database of postal words captured at 212 dpi. Gyeonghwan Kim, Venu Govindaraju |
ICDAR | 2 |
| 1995 | Serial classifier combination for handwritten word recognitionabstractThe performance of off-line handwritten word recognition algorithms declines with increasing lexicon size, but may be improved by serial combination of classifiers. The authors address some issues relevant to the design of serial classifier combinations. They present experimental results that show that the performance of a serial combination depends on not only the intrinsic recognition power of the classifiers but also the relative orthogonality of their features. A top-choice recognition rate of 83% is obtained for a lexicon of size 1700 by combining two analytical word classifiers that perform individually at 70%. Even higher recognition rates may be expected from a serial combination of two classifiers with less correlated features, such as a high-performance holistic classifier with an analytical classifier. Sriganesh Madhvanath, Venu Govindaraju |
ICDAR | 2 |
| 1995 | Reading handwritten US census formsabstractCommercial forms-reading systems for extraction of data from forms do not meet acceptable accuracy requirements on forms filled out by hand. In December 1993, NIST called industry and research organizations working in the area of handwriting recognition to participate in a test to determine the state of the art in the area. A database of form images containing actual responses received by the US Census Bureau was provided. The handwritten responses are very loosely constrained in terms of writing style, format of response and choice of text. The sizes of the lexicons provided are very large (about 50000 entries) and yet the coverage is incomplete (about 70%). In this paper we discuss the approach taken by CEDAR to automate the task of reading the census forms. The subtasks of field extraction and phrase recognition are described. Sriganesh Madhvanath, Venu Govindaraju, Vemulapati Ramanaprasad, Dar-Shyang Lee, Sargur N. Srihari |
ICDAR | 2 |
| 1995 | Image quality and readabilityabstractDetermining the readability of documents is an important task. Human readability pertains to the scenario when a document image is ultimately presented to a human to read. Machine readability pertains to the scenario when the document is subjected to an OCR process. In either case, poor image quality might render a document unreadable. A document image which is human readable is often not machine readable. It is often advisable to filter out documents of poor image quality before sending them to either machine or human for reading. This paper is about the design of such a filter. We describe various factors which affect document image quality and the accuracy of predicting the extent of human and machine readability possible using metrics based on document image quality. Venu Govindaraju, Sargur N. Srihari |
ICIP (3) | 1 |
| 1995 | Zero crossings of a non-orthogonal wavelet transform for object locationabstractIn this paper we address the task of segmentation of objects from photographs. A method of extraction of features based on the zero-crossings of a wavelet transform is described. The wavelet transform basis functions are derived from the second derivative of a Gaussian function. The extracted features are then used in a multilevel hypothesis generate and test algorithm to locate the objects of interest. The matching algorithm is based on the springs and templates framework of Fischler and Eschlanger (1973). The zero-crossings of the wavelet coefficients at different scales are combined in the model-matching stage to generate possible candidates. We apply this method to segment human faces from newspaper photographs. Mahesh Venkatraman, Venu Govindaraju |
ICIP (3) | 2 |
| 1994 | Generating manifold samples from a handwritten word
Srirangaraj Setlur, Venu Govindaraju |
Pattern Recognit. Lett. | 2 |
| 1992 | A Computational Model for Face Location Based on Cognitive Principles
Venu Govindaraju, Sargur N. Srihari, David B. Sher |
AAAI | 1 |
| 1992 | Caption-aided face location in newspaper photographsabstractThe human face is an object that is easily located in complex scenes by humans and adults alike. Yet the development of an automated system to perform this task is extremely challenging. This paper is about developing computational procedures to locate human faces in newspaper photographs where scenes are often cluttered making object location non-trivial. On the other hand, the task is made feasible by constraints which follow naturally from rules in photo-journalism. Faces identified by the caption are clearly depicted without occlusion and contrast against the background and sizes of faces fall within a fixed range depending on the dimensions of the photograph and the number of people featuring in it.> Venu Govindaraju, Sargur N. Srihari, David B. Sher |
ICPR (1) | 1 |
| 1990 | A computational model for face locationabstractThe authors adopted a model-based approach, where the shape of the object is defined in terms of several mini-templates. The mini-templates are abstract descriptions of simple geometric features like arcs and corners. Relationships between mini-templates are not rigid. Rather, they are represented by springs that allow deformation of a template in terms of its size and orientation. Cost functionals are determined empirically. The authors expect their system to generate candidate regions in a given photograph associated with a rank of its goodness.> Venu Govindaraju, Sargur N. Srihari, David B. Sher |
ICCV | 1 |
| 1989 | Locating human faces in newspaper photographsabstractA computational approach to locating human faces in newspaper photographs is described. While the computational recognition of a well-framed face as one of a known set of faces has received some attention, the computational location of faces in varying contexts is relatively unexplored. Candidates for the locations of faces are hypothesized by extracting features in the edge-image of the photograph and matching with a model of a face profile. A face is defined by a parametric representation of each of the component parts. Knowledge contained in the caption is represented using semantic networks and is used to reason about the locations of faces of individuals in the photograph.> Venu Govindaraju, David B. Sher, Rohini K. Srihari, Sargur N. Srihari |
CVPR | 1 |
| 1989 | Analysis of textual images using the Hough transform
Sargur N. Srihari, Venu Govindaraju |
Mach. Vis. Appl. | 2 |