George Nagy

dblp:02/3226 · DBLP profile ↗
← Back
38ranked-venue papers in the field
9as first author
2since 2021 · last 2021
0000-0002-0521-1443ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 34 (9 first)Information Retrieval & Web Search · 2Database Systems & Data Management · 1Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2021 Competition and Collaboration in Document Analysis and Recognition
Daniel P. Lopresti, George Nagy
ICDAR (1)2
2021 Near-Perfect Relation Extraction from Family Books
George Nagy
ICDAR (3)1
2018 Green Interaction for Extracting Family Information from OCR'd Books
abstract
Repetitively formatted historical books are tokenized and tagged according to eight token types (capitalized words, numbers, punctuation ...). To extract family information, templates of short sequences of tags are generated around frequent proper nouns and specified tokens like "born". Each template is associated with a user-assigned class (head of household, father, mother, spouse, geographic location ...) and a pointer to an overlapping or nearby fragment of text to be extracted. Matching the template against the book text yields class-labeled factoids. In an interaction cycle, new extraction templates are proposed for user approval or editing. Each edit-then-extract cycle typically yields thousands of factoids and a dozen new templates. With five approximately half-hour interactive sessions, 44,000 genealogical factoids were extracted from a 17th century Scottish register of marriages and births and from published 19th-20th century Ohio funeral parlor records. The experience indicates that this method quickly yields quality results with higher F-score than reported for hand-constructed rule templates.
David W. Embley, George Nagy
DAS2
2014 End-to-End Conversion of HTML Tables for Populating a Relational Database
abstract
Automating the conversion of human-readable HTML tables into machine-readable relational tables will enable end-user query processing of the millions of data tables found on the web. Theoretically sound and experimentally successful methods for index-based segmentation, extraction of category hierarchies, and construction of a canonical table suitable for direct input to a relational database are demonstrated on 200 heterogeneous web tables. The methods are scalable: the program generates the 198 Access compatible CSV files in ~0.1s per table (two tables could not be indexed).
George Nagy, Sharad C. Seth, David W. Embley
Document Analysis Systems1
2013 Segmenting Tables via Indexing of Value Cells by Table Headers
abstract
Correct segmentation of a web table into its component regions is the essential first step to understanding tabular data. Our algorithmic solution to the segmentation problem relies on the property that strings defining row and column header paths uniquely index each data cell in the table. We segment the table using only "logical layout analysis" without resorting to any appearance features or natural language understanding. We start with a CSV table that preserves the 2-dimensional structure and contents of the original source table (e.g., an HTML table) but not font size, font weight, and color. The indexing property of table headers implies a four-quadrant partitioning of the table about a minimum index point. The algorithm finds the index point through an efficient guided search. Experimental results on a 200-table benchmark demonstrate the generality of the algorithm in handling a variety of table styles and forms.
Sharad C. Seth, George Nagy
ICDAR2
2012 Adapting the Turing Test for Declaring Document Analysis Problems Solved
abstract
We propose to adapt Turing's seminal 1950 test for machine intelligence to evaluating progress in document analysis systems. Our premise is that a problem can be considered solved if automated and human solutions to the underlying task are indistinguishable to a skeptical human judge. For the domain-specific problems of concern here, we reformulate the test to keep the interaction between judges and human/machine participants to graphical user interfaces that do not require natural language processing, a notable difference from Turing's original formulation. Examples of tasks that may lend themselves to such tests include detecting or identifying specific document components such as logos, photographs, tables, as well as writer and language identification. The administration of the test would be facilitated by commercial crowd-sourcing systems such as Amazon Mechanical Turk, as well as research platforms such as the Lehigh Document Analysis Engine (DAE) that accept arbitrary documents for input, record test results, and provide for trusted execution of submitted programs.
Daniel P. Lopresti, George Nagy
Document Analysis Systems2
2011 When is a Problem Solved?
abstract
Open problems are defined differently in document image analysis than in the physical sciences, theoretical computer science, or mathematics. Instead of a formal definition, problems in DIA are stated in terms of automation of an application area (e.g., postal address reading) or a scientific sub field (e.g., image compression). The notion of a successful solution may be based on (1) the relative accuracy of automated vs. expert solutions (given specific data and degree of manual tuning), (2) the distinguish ability of automated output from human output (a Turing Test), (3) the degree of current community interest (via conferences and journals), and/or (4) economic considerations. Because of the lack of formal definition for DIA problems, heuristics predominate over provably correct algorithms, and full disclosure of implementation details as well as populations and samples is essential. Results on available test sets are often only tangentially related to motivating applications. In addition, interest in automating certain tasks has been evolving rapidly as a result of advances in technology. Further community discussion of these issues may accelerate progress and symbiosis with allied disciplines.
Daniel P. Lopresti, George Nagy
ICDAR2
2011 Data Extraction from Web Tables: The Devil is in the Details
abstract
We present a method based on header paths for efficient and complete extraction of labeled data from tables meant for humans. Although many table configurations yield to the proposed syntactic analysis, some require access to semantic knowledge. Clicking on one or two critical cells per table, through a simple interface, is sufficient to resolve most of these problem tables. Header paths, a purely syntactic representation of visual tables, can be transformed ("factored") into existing representations of structured data such as category trees, relational tables, and RDF triples. From a random sample of 200 web tables from ten large statistical web sites, we generated 376 relational tables and 34,110 subject-predicate-object RDF triples.
George Nagy, Sharad C. Seth, Dongpu Jin, David W. Embley, Spencer Machado, Mukkai S. Krishnamoorthy
ICDAR1
2011 CalliGUI: Interactive Labeling of Calligraphic Character Images
abstract
Calligraphic data entry is accelerated by generating, with a feature-based character classifier, an ordered list of reference candidate labels for each character image. The improvement of labeling throughput depends on the top-N accuracy of the classifier, which in turn is a function of the available already-labeled patterns. Experiments on a database of 13,351 ancient calligraphic characters indicate that clicking on reference labels is more than twice as fast as Pinyin keyboard entry.
George Nagy, Xiafen Zhang
ICDAR1
2011 Towards Improved Paper-Based Election Technology
abstract
Resources are presented for fostering paper-based election technology. They comprise a diverse collection of real and simulated ballot and survey images, and software tools for ballot synthesis, registration, segmentation, and ground truthing. The grids underlying the designated location of voter marks are extracted from 13,315 degraded ballot images. The actual skew angles of sample ballots, recorded as part of complete ballot descriptions compiled with the interactive ground-truthing tool, are compared with their automatically extracted parameters. The average error is 0.1 degrees. These results provide a baseline for the application of digital image analysis to the scrutiny of electoral ballots.
Elisa H. Barney Smith, Daniel P. Lopresti, George Nagy, Ziyan Wu 0001
ICDAR3
2010 Document analysis issues in reading optical scan ballots
abstract
Optical scan voting is considered by many to be the most trustworthy option for conducting elections because it provides an independently verifiable record of each voter’s intent. While op-scan technology has been in use for decades, attempts to improve the machine reading of ballots raises a range of interesting issues in document image analysis. Work thus far has been hindered by a lack of real-world data, since ballots associated with actual elections are kept secure from the public and normally destroyed after a period time. Fortunately, as a result of a recent challenged election in the State of Minnesota, a large collection of op-scan ballot images was made available for public inspection on the World Wide Web. In this paper, we present this unique resource to the document analysis community. We also describe our efforts to annotate the collection, including the latest version of a graphical tool we have developed for collecting ground-truth interpretations, along with the protocol now being employed. The collection, consisting of ballot images, file formats, and associated truth data, is being made openly available to facilitate research in this important area.
Daniel P. Lopresti, George Nagy, Elisa H. Barney Smith
Document Analysis Systems2
2010 Analysis and taxonomy of column header categories for web tables
abstract
We describe a component of a document analysis system for constructing ontologies for domain-specific web tables imported into Excel. This component automates extraction of the Wang Notation for the column header of a table. Using column-header specific rules for XY cutting we convert the geometric structure of the column header to a linear string denoting cell attributes and directions of cuts. The string representation is parsed by a context-free grammar and the parse tree is further processed to produce an abstract data-type representation (the Wang notation tree) of each column category. Experiments were carried out to evaluate this scheme on the original and edited column headers of Excel tables drawn from a collection of 200 used in our earlier work. The transformed headers were obtained by editing the original column headers to conform to the format targeted by our grammar. Forty-four original headers and their reformatted versions were submitted as input to our software system. Our grammar was able to parse and the extract Wang notation tree for all the edited headers, but for only four of the original headers. We suggest extensions to our table grammar that would enable processing a larger fraction of headers without manual editing.
Sharad C. Seth, Ramana Chakradhar Jandhyala, Mukkai S. Krishnamoorthy, George Nagy
Document Analysis Systems4
2009 Camera-Based Ballot Counter
abstract
Portable ballot counters using camera technology and manual paper feed are potentially more reliable and less expensive than scanner based systems. We show that the spatial sampling rate, geometric linearity, point spread function, and photometric transfer function of off-the-shelf consumer cameras are acceptable for ballot imaging. However, scanner illumination is much more uniform than can be economically accomplished for variable size ballots. Therefore flat-field compensation must be designed into the image processing software. We illustrate the mechanical design of a prototype camera based ballot reader based on our comparative observations.
George Nagy, Bryan Clifford, Andrew Berg, Glenn Saunders, Daniel P. Lopresti, Elisa H. Barney Smith
ICDAR1
2009 Style-Based Ballot Mark Recognition
abstract
The push toward voting via hand marked paper ballots has focused attention on the limitations of current optical scan systems. Discrepancies between human and machine interpretations of ballot markings can lead to a loss of trust in the election process. In this paper, a style-based approach to ballot recognition is proposed in which marks are recognized collectively rather than in isolation. The consistency of a voter's style is leveraged to improve the overall accuracy of the system. We compare style-based recognition to various kinds of singlet classifiers and show that it outperforms them by a substantial margin.
Pingping Xiu, Daniel P. Lopresti, Henry S. Baird, George Nagy, Elisa H. Barney Smith
ICDAR4
2008 A Document Analysis System for Supporting Electronic Voting Research
abstract
As a result of well-publicized security concerns with direct recording electronic (DRE) voting, there is a growing call for systems that employ some form of paper artifact to provide a verifiable physical record of a voter's choices. In this paper, we present a system we are developing to support a multi-institution, cross-disciplinary research project examining issues that arise when paper ballots are used in elections. We survey the motivating factors behind our work, discuss the special constraints raised in processing ballots as opposed to more general document images, and describe the current status of our system.
Daniel P. Lopresti, George Nagy, Elisa H. Barney Smith
Document Analysis Systems2
2008 A Conceptual-Model-Based Computational Alembic for a Web of Knowledge
David W. Embley, Stephen W. Liddle, Deryle W. Lonsdale, George Nagy, Yuri A. Tijerino, Robert Clawson, Jordan Crabtree, Yihong Ding, Piyushee Jha, Zonghui Lian, Stephen Lynn, Raghav K. Padmanabhan, Jeff Peters, Cui Tao, Robby Watts, Charla Woodbury, Andrew Zitzelberger
ER4
2006 In search of meaning for time series subsequence clustering: matching algorithms based on a new distance measure
abstract
Recent papers have claimed that the result of K-means clustering for time series subsequences (STS clustering) is independent of the time series that created it. Our paper revisits this claim. In particular, we consider the following question: Given several time series sequences and a set of STS cluster centroids from one of them (generated by the K-means algorithm), is it possible to reliably determine which of the sequences produced these cluster centroids? While recent results suggest that the answer should be NO, we answer this question in the affirmative.We present cluster shape distance, an alternate distance measure for time series subsequence clusters, based on cluster shapes. Given a set of clusters, its shape is the sorted list of the pairwise Euclidean distances between their centroids. We then present two algorithms based on this distance measure, which match a set of STS cluster centroids with the time series that produced it. While the first algorithm creates DQG reuse this term more smaller "fingerprints" for the sequences, the second is more accurate. In our experiments with a dataset of 10 sequences, it produced a correct match 100% of the time.Furthermore, we offer an analysis that explains why our cluster shape distance provides a reliable way to match STS clusters to the original sequences, whereas cluster set distance fails to do so. Our work establishes for the first time a strong relation between the result of K-means STS clustering and the time series sequence that created it, despite earlier predictions that this is not possible.
Dina Q. Goldin, Ricardo Mardales, George Nagy
CIKM3
2006 Notes on Contemporary Table Recognition
David W. Embley, Daniel P. Lopresti, George Nagy
Document Analysis Systems3
2003 Handwriting Recognition Using Position Sensitive Letter N-Gram Matching
abstract
We propose further improvement of a handwriting recognition method that avoids segmentation while able to recognize words that were never seen before in handwritten form. This method is based on the fact that few pairs of English words share exactly the same set of letter bigrams and even fewer share longer n-grams. The lexical n-gram matches between every word in a lexicon and a set of reference words can be precomputed. A position-based match function then detects the matches between the handwritten signal of a query word and each reference word. We show that with a reasonable set of reference words, the recognition of lexicon words exceeds 90%.
Adnan El-Nasan, Sriharsha Veeramachaneni, George Nagy
ICDAR3
2003 Towards a Ptolemaic Model for OCR
abstract
In style-constrained classification often there are only a few samples of each style and class, and the correspondences between styles in the training set and the test set are unknown. To avoid gross misestimates of the classifier parameters it is therefore important to model the pattern distributions accurately. We offer empirical evidence for intuitively appealing assumptions, in feature spaces appropriate for symbolic patterns, for (1) tetrahedral configurations of class means that suggests linear style-adaptive classification, (2) improved estimates of classification boundaries by taking into account the asymmetric configuration of the patterns with respect to the directions toward other classes, and (3) pattern-correlated style variability.
Sriharsha Veeramachaneni, George Nagy
ICDAR2
2003 Ontology Generation from Tables
abstract
We often need to access and reorganize information available in multiple tables in diverse Web pages. To understand tables, we rely on acquired expertise, background information, and practice. Current computerized tools seldom consider the structure and content in the context of other tables with related information. This paper addresses the table processing issue by developing a new framework to table understanding that applies an ontology-based conceptual modeling extraction approach to: (i) understand a table's structure and conceptual content to the extent possible; (ii) discover the constraints that hold between concepts extracted from the table; (iii) match the recognized concepts with ones from a more general specification of related concepts; and (iv) merge the resulting structure with other similar knowledge representations for use in future situations. The result is a formalized method of processing the format and content of tables while incrementally building a relevant reusable conceptual ontology.
Yuri A. Tijerino, David W. Embley, Deryle W. Lonsdale, George Nagy
WISE4
2002 Classifier Adaptation with Non-representative Training Data
Sriharsha Veeramachaneni, George Nagy
Document Analysis Systems2
2001 Word Discrimination Based on Bigram Co-Occurrences
abstract
Very few pairs of English words share exactly the same letter bigrams. This linguistic property can be exploited to bring lexical context into the classification stage of a word recognition system. The lexical n-gram matches between every word in a lexicon and a subset of reference words can be precomputed. If a match function can detect matching segments of at least n-gram length from the feature representation of words, then an unknown word can be recognized by determining the subset of reference words having an n-gram match at the feature level with the unknown word. We show that with a reasonable number of reference words, bigrams represent the best compromise between the recall ability of single letters and the precision of trigrams. Our simulations indicate that using a longer reference list can compensate errors in feature extraction. The algorithm is fast enough, even with a slow processor, for human-computer interaction.
Adnan El-Nasan, Sriharsha Veeramachaneni, George Nagy
ICDAR3
2001 Exploration of Contextual Constraints for Character Pre-Classification
abstract
We present strategies and results for identifying the symbol type (lower-case, upper-case, digit, and punctuation or special symbols) of every character in a text document by using various kinds of information from neighboring characters. In the expectation of reasonable word and character segmentation for shape clustering, we designed several type recognition methods that depend on cluster n-grams, shape codes, and within word context. On an ASCII test corpus of 925 articles that simulates perfect image-level processing, these methods achieve a substantial improvement over default assignment of all characters to lower case.
Tin Kam Ho, George Nagy
ICDAR2
2001 Why Table Ground-Truthing is Hard
abstract
The principle that for every document analysis task there exists a mechanism for creating well-defined ground-truth is a widely held tenet. Past experience with standard datasets providing ground-truth for character recognition and page segmentation tasks supports this belief. In the process of attempting to evaluate several table recognition algorithms we have been developing, however, we have uncovered a number of serious hurdles connected with the ground-truthing of tables. This problem may, in fact, be much more difficult than it appears. We present a detailed analysis of why table ground-truthing is so hard, including the notions that there may exist more than one acceptable "truth" and/or incomplete or partial "truths".
Jianying Hu, Ramanujan S. Kashi, Daniel P. Lopresti, Gordon T. Wilfong, George Nagy
ICDAR5
2001 Advanced Character Recognition 6610
George Nagy
ICDAR1
2001 Style-Consistency in Isogenous Patterns
abstract
In many applications of pattern recognition, patterns appear in groups (fields) that have a common origin. For example, a printed word is afield of character patterns printed in the same font. A common origin induces consistency of style among features measured on patterns. In the presence of multiple styles, the features of co-occurring patterns are statistically dependent through the underlying style. Modeling such dependence among constituent patterns of afield increases the classification accuracy. Effects of style consistency on the distributions of field features (concatenation of pattern features) are modeled by hierarchical mixtures. Each field derives from a mixture of styles, while, within a field, a pattern derives from a class-style conditional mixture of Gaussians. An optimal (least-error) style-conscious classifier processes entire fields of patterns rendered in a consistent but unknown style, based on the model. In a laboratory experiment, style-conscious classification reduced errors on fields of printed digits by nearly 25% over singlet classifiers. Longer fields favor our classification method, because they furnish more information about the underlying style.
Prateek Sarkar, George Nagy
ICDAR2
1999 Cooperative Text and Line-Art Extraction from a Topographic Map
abstract
The black layer is digitized from a USGS topographic map digitized at 1000 dpi. The connected components of this layer are analyzed and separated into line art, text, and icons in two passes. The paired street casings are converted to polylines by vectorization and associated with street labels from the character recognition phase. The accuracy of character recognition is shown to improve by taking account of the frequently occurring overlap of line art with street labels. The experiments show that complete vectorization of the black line-layer bitmap is the major remaining problem.
George Nagy, Ashok Samal, Sharad C. Seth
ICDAR2
1999 Heeding More Than the Top Template
abstract
We present a method of classifying a pattern using information furnished by a ranked list of templates, rather than just the best matching template. We propose a parsimonious model to compute the class-conditional likelihood of a list of templates ranked on the basis of their match scores. We discuss the estimation of parameters used in the model. The results of maximum likelihood classification on isolated digit patterns consistently show a 10-20% relative gain in recognition accuracy when we use more than one top-template.
Prateek Sarkar, George Nagy
ICDAR2
1997 Conference Report: ICDAR' 07
George Nagy
ICDAR1
1997 Automatic Prototype Extracion for Adaptive OCR
abstract
A Bayesian method of isolating character bitmaps from paragraph-length samples of heavily degraded text images is demonstrated. The method requires a transcript of the text, but it is sufficiently robust to tolerate errors in transcripts obtained from multifont commercial OCR software. The resulting prototypes (labeled character images) are used to recognize additional text an the same document.
George Nagy
ICDAR1
1996 Priming the recognizer
George Nagy
DAS1
1995 Joint feature and classifier design for OCR
abstract
Shift-invariant, custom designed n-tuple features are combined with a probabilistic decision tree to classify isolated printed characters. The feature probabilities are estimated using a novel compound Bayesian procedure in order to delay the fall-off in classification accuracy with tree size due to a small sample set. On a ten-class confusion set of eight-point characters, the method yields error rates under 1% with only 3 training samples per class.
Dz-Mou Jung, George Nagy
ICDAR2
1995 Spatial sampling effects in optical character recognition
abstract
In this paper we examine the effects of random-phase spatial sampling on the optical character recognition process. We start by presenting a detailed analysis in the case of 1-dimensional patterns. Empirical data demonstrate that our model is accurate. We then give experimental results for more complex, 2-dimensional patterns (i.e. printed, scanned characters). Spatial sampling seems to account for a significant amount of the variability seen in practice.
Daniel P. Lopresti, Jiangying Zhou, George Nagy, Prateek Sarkar
ICDAR3
1993 Performance metrics for document understanding systems
abstract
Requirements for the objective evaluation of automated data-entry systems are presented. Because the cost of correcting errors dominates the document conversion process, the most important characteristic of an OCR device is accuracy. However, different measures of accuracy (error metrics) are appropriate for different applications, and at the character, word, text-line, text-block, and document levels. For wholly objective assessment, OCR devices must be tested under programmed, rather than interactive, control.>
Junichi Kanai, Thomas A. Nartker, Stephen V. Rice, George Nagy
ICDAR4
1993 Classifier combination for hand-printed digit recognition
abstract
Independent decisions by two high performance nearest-neighbor hand-printed digit classifiers are combined in a principled manner. Three combination methods are investigated: Bayesian combination, Dempster-Shafer evidential reasoning, and dynamic classifier selection. On a test set of 60,000 hand-printed digits, dynamic classifier selection performs slightly better than Bayesian or Dempster-Shafer evidential reasoning, but the lowest error rate is obtained by K-nearest-neighbor combination. Single-parameter classifier combination is used to generate error-reject curves. Essential error-free classification is obtained at the cost of 4% rejects. The zero-reject error rate decreases from 1.18% for the best single classifier system to 0.67% for the combined classifier.>
Michael Sabourin, Amar Mitiche, Danny S. Thomas, George Nagy
ICDAR4
1991 Constrained Integer Approximation to Planar Line Intersection
Shashank K. Mehta, Maharaj Mukherjee, George Nagy
Inf. Process. Lett.3
1989 On the Integration of Lexical and Spatial Data in a Unified High-Level Model
David W. Embley, George Nagy
DASFAA2