Sriganesh Madhvanath

dblp:94/5305 · DBLP profile ↗
← Back
40ranked-venue papers
13as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 12 first-authorDatabases, data management, data science and information retrieval · 16 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 32% Speech recognition and synthesis · 32% Image recognition and object detection · 26%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 44% Haptics and multimodal interaction · 44% Wearable and physiological sensing · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Ubiquitous computing and smart environments › context recognition
activity recognition
0.312018
Deep Temporal Multimodal Fusion for Medical Procedure Monitoring Using Wearable Sensors · IEEE Trans. Multim. 2018
Haptics and multimodal interaction
multimodal fusion
0.312018
Deep Temporal Multimodal Fusion for Medical Procedure Monitoring Using Wearable Sensors · IEEE Trans. Multim. 2018
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › multimodal speech recognition
audio-visual speech recognition
0.312017
Deep Multimodal Representation Learning from Temporal Data · CVPR 2017
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.312017
Deep Multimodal Representation Learning from Temporal Data · CVPR 2017
Computer vision › Image recognition and object detection
handwriting recognition
0.242012
HMM-Based Lexicon-Driven and Lexicon-Free Word Recognition for Online Handwritten Indic Scripts · IEEE Trans. Pattern Anal. Mach. Intell. 2012
The Role of Holistic Paradigms in Handwritten Word Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Holistic Verification of Handwritten Phrases · IEEE Trans. Pattern Anal. Mach. Intell. 1999
Wearable and physiological sensing › motion sensing
motion sensor data
0.112018
Deep Temporal Multimodal Fusion for Medical Procedure Monitoring Using Wearable Sensors · IEEE Trans. Multim. 2018
Computer vision › Video understanding and tracking
activity recognition
0.112017
Deep Multimodal Representation Learning from Temporal Data · CVPR 2017
Computer vision › Image recognition and object detection
text recognition
0.012001
The Role of Holistic Paradigms in Handwritten Word Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2001

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 0.7deep learning · 0.7adaptive sampling · 0.7recurrent neural network · 0.3maximum correlation loss · 0.3attention · 0.3lexicon-driven recognition · 0.1hidden markov model · 0.1bag-of-symbols representation · 0.1holistic paradigm · 0.0feature matching · 0.0chaincode contour processing · 0.0binary image representation · 0.0
YearPublicationVenuePosition
2021 Page-level Optimization of e-Commerce Item Recommendations
abstract
The item details page (IDP) is a web page on an e-commerce website that provides information on a specific product or item listing. Just below the details of the item on this page, the buyer can usually find recommendations for other relevant items. These are typically in the form of a series of modules or carousels, with each module containing a set of recommended items. The selection and ordering of these item recommendation modules are intended to increase discover-ability of relevant items and encourage greater user engagement, while simultaneously showcasing diversity of inventory and satisfying other business objectives. Item recommendation modules on the IDP are often curated and statically configured for all customers, ignoring opportunities for personalization. In this paper, we present a scalable end-to-end production system to optimize the personalized selection and ordering of item recommendation modules on the IDP in real-time by utilizing deep neural networks. Through extensive offline experimentation and online A/B testing, we show that our proposed system achieves significantly higher click-through and conversion rates compared to other existing methods. In our online A/B test, our framework improved click-through rate by 2.48% and purchase-through rate by 7.34% over a static configuration.
Chieh Lo, Krutika Shetty, Changchen He, Kathy Hu, Justin M. Platz, Adam Ilardi, Sriganesh Madhvanath
RecSys9
2019 Zero Shot License Plate Re-Identification
abstract
The problem of person, vehicle or license plate reidentification is generally treated as a multi-shot image retrieval problem. The objective of these tasks is to learn a feature representation of query images (called a "signature") and then use these signatures to match against a database of template image signatures with the aid of a distance metric. In this paper, we propose a novel approach for license plate Re-Id inspired by Zero Shot Learning. The core idea is to generate template signatures for retrieval purposes from a multi-hot text encoding of license plates instead of their images. The proposed method maps license plate images and their license plate numbers to a common embedding space using a Symmetric Triplet loss function so that an image can be queried against its text. In effect, our approach makes it possible to identify license plates whose images have never been seen before, using a large text database of license plate numbers. We show that our system is capable of highly accurate and fast re-identification of license plates, and its performance compares favorably to both OCR-based approaches as well as state of the art image-based Re-ID approaches. In addition to the advantages of avoiding manual image labeling and the ease of creating signature databases, the minimal time and storage requirements enable our system to be deployed even on portable devices.
Mayank Gupta 0002, Abhinav Kumar 0004, Sriganesh Madhvanath
WACV3
2018 Deep Temporal Multimodal Fusion for Medical Procedure Monitoring Using Wearable Sensors
abstract
Process monitoring and verification have a wide range of uses in the medical and healthcare fields. Currently, such tasks are often carried out by a trained specialist, which makes them expensive, inefficient, and time-consuming. Recent advances in automated video- and multimodal-data-based action and activity recognition have made it possible to reduce the extent of manual intervention required to effectively carry out process supervision tasks. In this paper, we propose algorithms for automated egocentric human action and activity recognition from multimodal data, with a target application of monitoring and assisting a user perform a multistep medical procedure. We propose a supervised deep multimodal fusion framework that relies on concurrent processing of motion data acquired with wearable sensors and video data acquired with an egocentric or body-mounted camera. We demonstrate the effectiveness of the algorithm on a public multimodal dataset and conclude that automated process monitoring via the use of multiple heterogeneous sensors is a viable alternative to its manual counterpart. Furthermore, we demonstrate that the application of previously proposed adaptive sampling schemes to the video processing branch of the multimodal framework results in significant performance improvements.
Edgar A. Bernal, Xitong Yang, Qun Li 0003, Jayant Kumar, Sriganesh Madhvanath, Palghat Ramesh, Raja Bala
IEEE Trans. Multim.5
2017 Deep Multimodal Representation Learning from Temporal Data
abstract
In recent years, Deep Learning has been successfully applied to multimodal learning problems, with the aim of learning useful joint representations in data fusion applications. When the available modalities consist of time series data such as video, audio and sensor signals, it becomes imperative to consider their temporal structure during the fusion process. In this paper, we propose the Correlational Recurrent Neural Network (CorrRNN), a novel temporal fusion model for fusing multiple input modalities that are inherently temporal in nature. Key features of our proposed model include: (i) simultaneous learning of the joint representation and temporal dependencies between modalities, (ii) use of multiple loss terms in the objective function, including a maximum correlation loss term to enhance learning of cross-modal information, and (iii) the use of an attention model to dynamically adjust the contribution of different input modalities to the joint representation. We validate our model via experimentation on two different tasks: video-and sensor-based activity classification, and audio-visual speech recognition. We empirically analyze the contributions of different components of the proposed CorrRNN model, and demonstrate its robustness, effectiveness and state-of-the-art performance on multiple datasets.
Xitong Yang, Palghat Ramesh, Radha Chitta, Sriganesh Madhvanath, Edgar A. Bernal, Jiebo Luo 0001
CVPR4
2015 Hybrid active learning for non-stationary streaming data with asynchronous labeling
abstract
Active learning enables supervised classifiers to learn using fewer labeled samples, by actively selecting samples for human labeling. Most Active Learning approaches can be categorized as pool-based or stream-based. Pool-based strategies select instances to be labeled from the available pool of unlabeled data, by evaluating each instance, whereas stream-based strategies examine every instance in the incoming stream of unlabeled data and decide sequentially whether they want that instance to be labeled or not. Stream-based strategies enable the ability to adapt the classifier model more quickly as the incoming data changes, while pool-based strategies often exhibit better learning rates. In this paper, we propose a framework and method for Hybrid Active Learning that integrates pool-based and stream-based strategies to harvest the benefits of both, in a streaming data classification scenario where concept drift may be prevalent, and labeling is asynchronous. In addition, we propose (i) prioritized aggregation of selection to combine selected instances for labeling from the pool-based and stream-based strategies, and (ii) batch period adaptation to dynamically change the triggering pattern of the pool-based strategy based upon the detection of concept drift.
Hyunjoo Kim, Sriganesh Madhvanath
IEEE BigData2
2015 Nudging Grocery Shoppers to Make Healthier Choices
abstract
Despite the rampant increase in obesity rates and concomitant increases in rates of mortality from heart disease, cancer and diabetes, getting the general public to adopt a healthy diet has proven to be challenging for a variety of reasons. In this paper, we describe Foodle, a research project aimed at providing automated, personalized and goal-driven dietary guidance to users based on their grocery receipt data, by leveraging the availability of digital receipts for grocery store purchases. We discuss challenges faced, the current state of the project, and directions for future work.
Elizabeth Wayman, Sriganesh Madhvanath
RecSys2
2014 Allograph modeling for online handwritten characters in devanagari using constrained stroke clustering
abstract
Writer-specific character writing variations such as those of stroke order and stroke number are an important source of variability in the input when handwriting is captured “online” via a stylus and a challenge for robust online recognition of handwritten characters and words. It has been shown by several studies that explicit modeling of character allographs is important for achieving high recognition accuracies in a writer-independent recognition system. While previous approaches have relied on unsupervised clustering at the character or stroke level to find the allographs of a character, in this article we propose the use of constrained clustering using automatically derived domain constraints to find a minimal set of stroke clusters. The allographs identified have been applied to Devanagari character recognition using Hidden Markov Models and Nearest Neighbor classifiers, and the results indicate substantial improvement in recognition accuracy and/or reduction in memory and computation time when compared to alternate modeling techniques.
A. Bharath, Sriganesh Madhvanath
ACM Trans. Asian Lang. Inf. Process.2
2013 Comparison of Phone-Based Distal Pointing Techniques for Point-Select Tasks
Andy Cockburn, Sriganesh Madhvanath
INTERACT (2)3
2012 The blue one to the left: enabling expressive user interaction in a multimodal interface for object selection in virtual 3d environments
abstract
Interaction with virtual 3D environments comes with a host of challenges. For instance, because 3D objects tend to occlude one another, performing object selection by pointing gestures is problematic, and more so when there are many objects in the scene. In the real world we tend to use speech to clarify our intent, by referring to distinctive attributes of the object and/or its absolute or relative location in space. Multimodal interactive systems involving speech and gesture have generally relied on speech for commands and deictic gestures for indicating the target object. In this paper, we present a system which allows object references to be made using gestures and speech, and supports a variety of expressions inspired by real-world usage.
Pulkit Budhiraja, Sriganesh Madhvanath
ICMI2
2012 Designing multiuser multimodal gestural interactions for the living room
abstract
Most work in the space of multimodal and gestural interaction has focused on single user productivity tasks. The design of multimodal, freehand gestural interaction for multiuser lean-back scenarios is a relatively nascent area that has come into focus because of the availability of commodity depth cameras. In this paper, we describe our approach to designing multimodal gestural interaction for multiuser photo browsing in the living room, typically a shared experience with friends and family. We believe that our learnings from this process will add value to the efforts of other researchers and designers interested in this design space.
Sriganesh Madhvanath, Ramadevi Vennelakanti, Anbumani Subramanian, Ankit Shekhawat, Amit Ranjan
ICMI1
2012 Pixene: creating memories while sharing photos
abstract
In this paper we describe Pixene, a photo sharing system that focuses on the capture and subsequent visualization and consumption of interactions around shared photos, where the sharing may be with physically co-present friends and family, or online with one's social network. In the former scenario, the interactions may be richly multimodal and involve pointing and spoken comments. Remote interaction is primarily in the form of 'like's and text comments on social networking sites. Pixene thus acts as a common repository for interactions over photos and brings interactions from a co-located and online photo sharing into a single platform. Pixene also provides a rich photo browsing experience that allows users to view not only the photographs but also the interaction history around them, e.g. who saw it, who did they see it with, what they said in association with different regions of interest, comments and 'like's. In this paper, we describe the features and system design of Pixene.
Ramadevi Vennelakanti, Sriganesh Madhvanath, Anbumani Subramanian, Ajith Sowndararajan, Arun David
ICMI2
2012 HMM-Based Lexicon-Driven and Lexicon-Free Word Recognition for Online Handwritten Indic Scripts
abstract
Research for recognizing online handwritten words in Indic scripts is at its early stages when compared to Latin and Oriental scripts. In this paper, we address this problem specifically for two major Indic scripts--Devanagari and Tamil. In contrast to previous approaches, the techniques we propose are largely data driven and script independent. We propose two different techniques for word recognition based on Hidden Markov Models (HMM): lexicon driven and lexicon free. The lexicon-driven technique models each word in the lexicon as a sequence of symbol HMMs according to a standard symbol writing order derived from the phonetic representation. The lexicon-free technique uses a novel Bag-of-Symbols representation of the handwritten word that is independent of symbol order and allows rapid pruning of the lexicon. On handwritten Devanagari word samples featuring both standard and nonstandard symbol writing orders, a combination of lexicon-driven and lexicon-free recognizers significantly outperforms either of them used in isolation. In contrast, most Tamil word samples feature the standard symbol order, and the lexicon-driven recognizer outperforms the lexicon free one as well as their combination. The best recognition accuracies obtained for 20,000 word lexicons are 87.13 percent for Devanagari when the two recognizers are combined, and 91.8 percent for Tamil using the lexicon-driven technique.
A. Bharath, Sriganesh Madhvanath
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 MozArt: a multimodal interface for conceptual 3D modeling
abstract
There is a need for computer aided design tools that support rapid conceptual level design. In this paper we explore and evaluate how intuitive speech and multitouch input can be combined in a multimodal interface for conceptual 3D modeling. Our system, MozArt, is based on a user's innate abilities - speaking and touching, and has a toolbar/button-less interface for creating and interacting with computer graphics models. We briefly cover the hardware and software technology behind MozArt, and present a pilot study comparing our multimodal system with a conventional multitouch modeling interface with first time CAD users. While a larger study is required to obtain statistically significant comparison regarding efficiency and accuracy of the two interfaces, a majority of the participants preferred the multimodal interface over the multitouch. We summarize lessons learned and discuss directions for future research.
Anirudh Sharma, Sriganesh Madhvanath, Ankit Shekhawat, Mark Billinghurst
ICMI2
2010 JollyMate: Assistive Technology for Young Children with Dyslexia
abstract
In this paper, we describe Jolly mate, a product concept that we have envisioned as assistive technology for young children with Dyslexia. Jolly mate, a digital notepad, emulates the Jolly Phonics system of teaching letter sounds and letter formation to children with dyslexia. Jolly mate in turn uses simple handwritten character recognizers created using the Lipi IDE tool from the Lipi Toolkit project, for detecting when a character has been written incorrectly. In this paper we describe the Jolly mate concept in brief, the Lipi IDE tool used to create the recognizers, and their integration.
Jignesh Khakhar, Sriganesh Madhvanath
ICFHR2
2010 On the Significance of Stroke Size and Position for Online Handwritten Devanagari Word Recognition: An Empirical Study
abstract
Stroke size and position are considered as important information for online recognition of handwritten characters and words in oriental and Indie family of scripts especially because of their multi-stroke and two-dimensional nature. In an Indie script such as Devanagari, the vowel diacritics (matras) can occur at any position around the base consonant and there are even pairs of matras which have similar shapes and differ only in their position with respect to the base consonant. In this paper, we study the relevance of stroke size and position information for the recognition of online handwritten Devanagari words by comparing three different preprocessing schemes. Our experimental results indicate that the word recognition accuracy achieved using a preprocessing scheme that completely disregards the original sizes and positions of the strokes (and symbols) is comparable with the scheme that retains them, when the input is in discrete style, and contextual knowledge in the form of a lexicon is available.
A. Bharath, Sriganesh Madhvanath
ICPR2
2009 A Framework Based on Semi-Supervised Clustering for Discovering Unique Writing Styles
abstract
An online multi-stroke character is often written in many ways. While some vary in the number of strokes they contain, others differ in the ordering of strokes. It is important for a writer-independent recognition system to learn these different styles of writing the character during the training phase in order to better model the training data. Typically, the samples of a character are clustered in an unsupervised manner and each cluster is modeled individually. In this paper, we describe an approach based on dasiasemi-supervised clusteringpsila where basic domain knowledge can be incorporated for better clustering of strokes present across all the characters.Experimental results show improved recognition accuracy when compared to the baseline system.
A. Bharath, Sriganesh Madhvanath
ICDAR2
2009 A Framework for Adaptation of the Active-DTW Classifier for Online Handwritten Character Recognition
abstract
Practical applications of online handwritten character recognition demand robust and highly accurate recognition along with low memory requirements. The Active-DTW classifier proposed by Sridhar et al.combines the advantages of generative and discriminative classifiers to address the similarity of between-class samples, while taking into account the variability of writing styles within the same character class. Active-DTW uses Active Shape Models to model the significant writing styles in a memory-efficient manner.However, in order to create accurate models, a large number of training samples is needed up front, which is not desirable or available in many practical applications. In this paper, we propose a supervised adaptation framework for the Active-DTW classifier which allows recognition to begin with a small number of training samples, and adapts the classifier to the new samples presented to the system during recognition. We compare the performance of Active-DTW using the proposed adaptation framework, with a nearest-neighbor classifier using an LVQ-based adaptation scheme, on the online handwritten Tamil character dataset.
Vandana Roy, Sriganesh Madhvanath, Anand S., Ragunath R. Sharma
ICDAR2
2008 Digital Ink to Form Alignment for Electronic Clipboard Devices
abstract
This paper addresses the problem of aligning handwritten input captured as digital ink to form templates, which occurs when paper forms are filled using "electronic clipboard" devices such as the ACECAD DigiMemo. Due to certain non-linearities in the digitizer hardware, the ink that is obtained from these devices is seldom in exact alignment with the "soft" form template. While not a concern for simple note capturing applications, this poses a serious problem for form processing applications. In this paper, we explore image registration approaches to solve the ink alignment problem. Two kinds of point pattern approaches to solve this alignment problem are explored. One method considers ink as "rigid" and a global transformation matrix is used to align the ink with the form template, while the second assumes only minor distortions and hence adopts a more local approach. We compare the performance of the two approaches and present some preliminary results. A summary of the work in progress and next steps are presented in conclusion.
Varadarajan Jagannadan, Sriganesh Madhvanath
Document Analysis Systems2
2008 FreePad: a novel handwriting-based text input for pen and touch interfaces
abstract
The last decade has seen tremendous growth in mobile devices such as Pocket PCs, mobile phones, Tablet PCs and notebooks. Most of these devices enable interaction through a stylus or touch interface, powered by handwriting recognition (HWR) capability. In this paper, we propose a novel input method that addresses some of the issues that arise due to the constraints posed by these devices in accepting handwriting input. For instance, many of the devices have a small writing area making "continuous" input difficult if not impossible, and the process of handwriting input demands significant user attention. The proposed solution is inspired by touch-typing, and appreciably reduces user's effort in the interaction, and it is especially suited for very small writing areas. The approach has been demonstrated using a prototype system that recognizes handwritten English words, and its accuracy has been evaluated using a standard dataset of handwritten words. A preliminary user study has also been carried out to understand user acceptance of the proposed technique.
A. Bharath, Sriganesh Madhvanath
IUI2
2007 Hidden Markov Models for Online Handwritten Tamil Word Recognition
abstract
Hidden Markov Models (HMM) have long been a popu- lar choice for Western cursive handwriting recognition fol- lowing their success in speech recognition. Even for the recognition of Oriental scripts such as Chinese, Japanese and Korean, Hidden Markov Models are increasingly being used to model substrokes of characters. However, when it comes to Indic script recognition, the published work em- ploying HMMs is limited, and generally focussed on iso- lated character recognition. In this effort, a data-driven HMM-based online handwritten word recognition system for Tamil, an Indic script, is proposed. The accuracies obtained ranged from 98% to 92.2% with different lexicon sizes (1K to 20K words). These initial results are promising and warrant further research in this direction. The results are also encouraging to explore possibilities for adopting the approach to other Indic scripts as well.
A. Bharath, Sriganesh Madhvanath
ICDAR2
2007 Password management using doodles
abstract
The average computer user needs to remember a large number of text username and password combinations for different applications, which places a large cognitive load on the user. Consequently users tend to write down passwords, use easy to remember (and guess) passwords, or use the same password for multiple applications, leading to security risks. This paper describes the use of personalized hand-drawn "doodles" for recall and management of password information. Since doodles can be easier to remember than text passwords, the cognitive load on the user is reduced. Our method involves recognizing doodles by matching them against stored prototypes using handwritten shape matching techniques. We have built a system which manages passwords for web applications through a web browser. In this system, the user logs into a web application by drawing a doodle using a touchpad or digitizing tablet attached to the computer. The user is automatically logged into the web application if the doodle matches the doodle drawn during enrollment. We also report accuracy results for our doodle recognition system, and conclude with a summary of next steps.
Naveen Sundar G., Sriganesh Madhvanath
ICMI2
2005 UPX: A New XML Representation for Annotated Datasets of Online Handwriting Data
abstract
This paper introduces our efforts to create UPX, an XML-based successor to the venerable UNIPEN format for the representation of annotated datasets of online handwriting data. In the first part of the paper, shortcomings of the UNIPEN format are discussed and the goals of UPX are outlined. Prior work related to UPX in the form of the recently proposed hwDataset representation is presented. The second part of the paper summarizes the status of the UPX effort, in particular, experiments to map UNIPEN elements to hwDataset and InkML and identify potential issues with migrating existing UNIPEN data to UPX. This is work in progress, and we invite participation from the handwriting recognition research community and industry to make UPX a reality.
Mudit Agrawal, Kalika Bali, Sriganesh Madhvanath, Louis Vuurpijl
ICDAR3
2005 An Approach to Identify Unique Styles in Online Handwriting Recognition
abstract
We describe a method for identifying different writing styles of online handwritten characters based on clustering. The motivation of this experiment is to develop automatic characterization of different writing styles that arise due to variation in stroke number or stroke ordering. An efficient agglomerative hierarchical clustering technique with the nearest neighbor approach was implemented to cluster strokes. The results obtained from our experiment indicate that the resulting prototypes are unique and essentially capture different writing styles.
A. Bharath, V. Deepu, Sriganesh Madhvanath
ICDAR3
2005 Machine Recognition of Online Handwritten Devanagari Characters
abstract
In this paper, we describe a system for the automatic recognition of isolated handwritten Devanagari characters obtained by linearizing consonant conjuncts. Owing to the large number of characters and resulting demands on data acquisition, we use structural recognition techniques to reduce some characters to others. The residual characters are then classified using the subspace method. Finally the results of structural recognition and feature-based matching are mapped to give final output. The proposed system is evaluated for the writer dependent scenario.
Niranjan Joshi, G. Sita, A. G. Ramakrishnan, V. Deepu, Sriganesh Madhvanath
ICDAR5
2004 Tamil Handwriting Recognition Using Subspace and DTW Based Classifiers
Niranjan Joshi, G. Sita, A. G. Ramakrishnan, Sriganesh Madhvanath
ICONIP4
2004 Elastic Matching Algorithms for Online Tamil Character Recognition
Niranjan Joshi, G. Sita, A. G. Ramakrishnan, Sriganesh Madhvanath
ICONIP4
2004 Experiences in Collection of Handwriting Data for Online Handwriting Recognition in Indic Scripts
Ajay S. Bhaskarabhatla, Sriganesh Madhvanath
LREC2
2004 An XML Representation for Annotated Handwriting Datasets for Online Handwriting Recognition
Ajay S. Bhaskarabhatla, Sriganesh Madhvanath
LREC2
2001 The Role of Holistic Paradigms in Handwritten Word Recognition
abstract
The holistic paradigm in handwritten word recognition treats the word as a single, indivisible entity and attempts to recognize words from their overall shape, as opposed to their character contents. In this survey, we have attempted to take a fresh look at the potential role of the holistic paradigm in handwritten word recognition. The survey begins with an overview of studies of reading which provide evidence for the existence of a parallel holistic reading process,in both developing and skilled readers. In what we believe is a fresh perspective on handwriting recognition, approaches to recognition are characterized as forming a continuous spectrum based on the visual complexity of the unit of recognition employed and an attempt is made to interpret well-known paradigms of word recognition in this framework. An overview of features, methodologies, representations, and matching techniques employed by holistic approaches is presented.
Sriganesh Madhvanath, Venu Govindaraju
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Syntactic methodology of pruning large lexicons in cursive script recognition
Sriganesh Madhvanath, Venu Krpasundar, Venu Govindaraju
Pattern Recognit.1
1999 Extracting Patron Data from Check Images
abstract
We describe on-going research on a system for extraction and recognition of patron data such as name and address from scanned and binarized images of US checks. Extraction of the region of interest is accomplished by a graph algorithm that operates on connected components. OCR on the region of interest involves separation of connected components into lines, and generation of multiple segmentation hypotheses and character choices for each connected component. Interpretation of the raw OCR results as addresses and names is accomplished using Checkmate, a language-based OCR interpretation engine. A description of the different system components is presented in this paper.
Sriganesh Madhvanath, S. McCauliff, K. Moidin Mohiuddin
ICDAR1
1999 Chaincode Contour Processing for Handwritten Word Recognition
abstract
Contour representations of binary images of handwritten words afford considerable reduction in storage requirements while providing lossless representation. On the other hand, the one-dimensional nature of contours presents interesting challenges for processing images for handwritten word recognition. Our experiments indicate that significant gains are to be realized in both speed and recognition accuracy by using a contour representation in handwriting applications.
Sriganesh Madhvanath, Gyeonghwan Kim, Venu Govindaraju
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Holistic Verification of Handwritten Phrases
abstract
In this paper, we describe a system for rapid verification of unconstrained off-line handwritten phrases using perceptual holistic features of the handwritten phrase image. The system is used to verify handwritten street names automatically extracted from live US mail against recognition results of analytical classifiers. Presented with a binary image of a street name and an ASCII street name, holistic features (reference lines, large gaps and local contour extrema) of the street name hypothesis are "predicted" from the expected features of the constituent characters using heuristic rules. A dynamic programming algorithm is used to match the predicted features with the extracted image features. Classes of holistic features are matched sequentially in increasing order of cost, allowing an ACCEPT/REJECT decision to be arrived at in a time-efficient manner. The system rejects errors with 98 percent accuracy at the 30 percent accept level, while consuming approximately 20/msec per image on the average on a 150 MHz SPARC 10.
Sriganesh Madhvanath, Evelyn Kleinberg, Venu Govindaraju
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Local reference lines for handwritten phrase recognition
Sriganesh Madhvanath, Venu Govindaraju
Pattern Recognit.1
1997 Contour-based Image Preprocessing for Holistic Handwritten Word Recognition
abstract
The one-dimensional nature of contour representations presents interesting challenges for processing of images for handwritten word recognition. In this paper, we discuss the issues of determination of upper and lower contours of the word, determination of significant focal extrema on the contour, and determination of reference lines from contour representations of handwritten words.
Sriganesh Madhvanath, Venu Govindaraju
ICDAR1
1997 Pruning Large Lexicons Using Generalized Word Shape Descriptors
abstract
We present a technique for pruning of large lexicons for recognition of cursive script words. The technique involves extraction and representation of downward pen-strokes from the cursive word (off-line or online) to obtain a generalized descriptor which provides a coarse characterization of word shape. The descriptor is matched with ideal descriptors of lexicon entries organized as a trie. When used with a static lexicon of 21,000 words, the accuracy of reduction to 1000 words exceeds 95%.
Sriganesh Madhvanath, Venu Krpasundar
ICDAR1
1997 The HOVER System for Rapid Holistic Verification of Off-lineHandwritten Phrases
abstract
The authors describe ongoing research on a system for rapid verification of unconstrained off-line handwritten phrases using perceptual holistic features of the handwritten phrase image. The system is used to verify handwritten street names automatically extracted from live US mail against recognition results of analytical classifiers. The system rejects errors with 98% accuracy at the 30% accept level, while consuming approximately 20 msec per image on the average on a 150 MHz SPARC 10.
Sriganesh Madhvanath, Evelyn Kleinberg, Venu Govindaraju, Sargur N. Srihari
ICDAR1
1997 Empirical Design of A Multi-Classifier Thresholding/Control Strategy for Recognition of Handwritten Street Names
abstract
A central task in the interpretation of handwritten US postal addresses is the off-line recognition of the street name. A lexicon of candidate street names may be extracted from a database of postal delivery points (DPF) by first locating and recognizing numeric fields such as the ZIP code and strewet number. The off-line handwritten word recognition (HWR) task is made difficult by the unconstrained, omni-scriptor nature of the input, and incomplete lexicons resulting from errors in processing numeric fields and intrinsic deficiencies in the DPF. In this paper, we describe an empirical approach to the design of a multi-classifier HWR Thresholding/Control module which forms part of a real-time handwritten address interpretation (HWAI) system. The decisions of two word classifiers are combined in a hierarchical manner to improve recognition and error-rejection performance, while meeting real-time requirements. The design employs logistic regression and agreement for evidence combination, and lexicon reduction for improved throughput as well as performance. The paper concludes with experimental results and directions for future research.
Sriganesh Madhvanath, Evelyn Kleinberg, Venu Govindaraju
Int. J. Pattern Recognit. Artif. Intell.1
1995 Serial classifier combination for handwritten word recognition
abstract
The performance of off-line handwritten word recognition algorithms declines with increasing lexicon size, but may be improved by serial combination of classifiers. The authors address some issues relevant to the design of serial classifier combinations. They present experimental results that show that the performance of a serial combination depends on not only the intrinsic recognition power of the classifiers but also the relative orthogonality of their features. A top-choice recognition rate of 83% is obtained for a lexicon of size 1700 by combining two analytical word classifiers that perform individually at 70%. Even higher recognition rates may be expected from a serial combination of two classifiers with less correlated features, such as a high-performance holistic classifier with an analytical classifier.
Sriganesh Madhvanath, Venu Govindaraju
ICDAR1
1995 Reading handwritten US census forms
abstract
Commercial forms-reading systems for extraction of data from forms do not meet acceptable accuracy requirements on forms filled out by hand. In December 1993, NIST called industry and research organizations working in the area of handwriting recognition to participate in a test to determine the state of the art in the area. A database of form images containing actual responses received by the US Census Bureau was provided. The handwritten responses are very loosely constrained in terms of writing style, format of response and choice of text. The sizes of the lexicons provided are very large (about 50000 entries) and yet the coverage is incomplete (about 70%). In this paper we discuss the approach taken by CEDAR to automate the task of reading the census forms. The subtasks of field extraction and phrase recognition are described.
Sriganesh Madhvanath, Venu Govindaraju, Vemulapati Ramanaprasad, Dar-Shyang Lee, Sargur N. Srihari
ICDAR1