Kazem Taghva

dblp:35/5977 · DBLP profile ↗
← Back
20ranked-venue papers
14as first author
3since 2021 · last 2025
0000-0001-8320-0080ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 5 first-authorTheory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Infomod: information-theoretic machine learning model diagnostics
Armin Esmaelizadeh, Sunil Cotterill, Liam Hebert, Lukasz Golab, Kazem Taghva
Distributed Parallel Databases5
2025 Tri-AL: An open source platform for visualization and analysis of clinical trials
Pouyan Nahed, Mina Esmail Zadeh Nojoo Kambar, Kazem Taghva, Lukasz Golab
Inf. Syst.3
2023 InfoMoD: Information-theoretic Model Diagnostics
abstract
Validating and debugging machine learning models is done by testing them on unseen data. Analyzing model performance on various subsets of the data is critical for fairness, trust, bias detection and explainablility. In this paper, we describe a new way to do this. Our solution, called InfoMoD, applies recent work in information-theoretic data summarization to the problem of model diagnostics. Using real-life datasets, we show how InfoMod concisely describes how a model performs across different subsets of the data and produces expected performance indicators for individual test instances.
Armin Esmaeilzadeh, Lukasz Golab, Kazem Taghva
SSDBM3
2011 Acronym Expansion Via Hidden Markov Models
abstract
In this paper, we report on design and implementation of a Hidden Markov Model (HMM) to extract acronyms and their expansions. We also report on the training of this HMM with Maximum Likelihood Estimation (MLE) algorithm using a set of examples. Finally, we report on our testing using standard recall and precision. The HMM achieves a recall and precision of 98% and 92% respectively.
Kazem Taghva, Lakshmi Vyas
ICSEng1
2007 Extracting _Carbon Copy_ Names and Organizations from a Heterogeneous Document Collection
abstract
We describe the development of a tool to identify the names of persons and their corresponding institutions as they appear in the "carbon copy" or "cc" lists of correspondence, memorandums, emails, and other types of documents.
Kazem Taghva, Russell Beckley, Jeffrey S. Coombs
ICDAR1
2006 The Effects of OCR Error on the Extraction of Private Information
Kazem Taghva, Russell Beckley, Jeffrey S. Coombs
Document Analysis Systems1
2005 Document analysis by processing JBIG-encoded images
Emma E. Regentova, Shahram Latifi, De Chen, Kazem Taghva, Dongsheng Yao
Int. J. Document Anal. Recognit.4
2004 The role of manually-assigned keywords in query expansion
Kazem Taghva, Julie Borsack, Thomas A. Nartker, Allen Condit
Inf. Process. Manag.1
2003 A comparison of automatic and manual zoning
Kazem Taghva, Julie Borsack, Steven E. Lumos, Allen Condit
Int. J. Document Anal. Recognit.1
2002 Hairetes: A Search Engine for OCR Documents
Kazem Taghva, Jeffrey S. Coombs
Document Analysis Systems1
2001 OCRSpell: an interactive spelling correction system for OCR errors in text
Kazem Taghva, Eric Stofsky
Int. J. Document Anal. Recognit.1
1999 Recognizing acronyms and their definitions
Kazem Taghva, Jeff Gilbreth
Int. J. Document Anal. Recognit.1
1996 Effects of OCR Errors on Ranking and Feedback Using the Vector Space Model
Kazem Taghva, Julie Borsack, Allen Condit
Inf. Process. Manag.1
1996 Evaluation of Model-Based Retrieval Effectiveness with OCR Text
abstract
We give a comprehensive report on our experiments with retrieval from OCR-generated text using systems based on standard models of retrieval. More specifically, we show that average precision and recall is not affected by OCR errors across systems for several collections. The collections used in these experiments include both actual OCR-generated text and standard information retrieval collections corrupted through the simulation of OCR errors. Both the actual and simulation experiments include full-text and abstract-length documents. We also demonstrate that the ranking and feedback methods associated with these models are generally not robust enough to deal with OCR errors. It is further shown that the OCR errors and garbage strings generated from the mistranslation of graphic objects increase the size of the index by a wide margin. We not only point out problems that can arise from applying OCR text within an information retrieval environment, we also suggest solutions to overcome some of these problems.
Kazem Taghva, Julie Borsack, Allen Condit
ACM Trans. Inf. Syst.1
1995 Post-Editing Through Approximation and Global Correction
abstract
This paper describes a new automatic spelling correction program to deal with OCR generated errors. The method used here is based on three principles: 1. Approximate string matching between the misspellings and the terms occuring in the database as opposed to the entire dictionary 2. Local information obtained from the individual documents 3. The use of a confusion matrix, which contains information inherently specific to the nature of errors caused by the particular OCR device This system is then utilized to process approximately 10,000 pages of OCR generated documents. Among the misspellings discovered by this algorithm, about 87% were corrected.
Kazem Taghva, Julie Borsack, Bryan Bullard, Allen Condit
Int. J. Pattern Recognit. Artif. Intell.1
1994 Results of Applying Probabilistic IR to OCR Text
Kazem Taghva, Julie Borsack, Allen Condit
SIGIR1
1994 The Effects of Noisy Data on Text Retrieval
abstract
We report on the results of our experiments on query evaluation in the presence of noisy data. In particular, an OCR-generated database and its corresponding 99.8% correct version are used to process a set of queries to determine the effect the degraded version will have on retrieval. It is shown that, with the set of scientific documents we use in our testing, the effect is insignificant. We further improve the result by applying an automatic postprocessing system designed to correct the kinds of errors generated by recognition devices. © 1994 John Wiley & Sons, Inc.
Kazem Taghva, Julie Borsack, Allen Condit, Srinivas Erva
J. Am. Soc. Inf. Sci.1
1993 Capturing Strong Reduction in Director String Calculus
Vugranam C. Sreedhar, Kazem Taghva
Theor. Comput. Sci.2
1989 A Novel Way 1o Identify IneguaIity Query Subclasses Which possess the Homomorphism Property
Tianzheng Wu, James L. Clark, Nong Zhou, Kazem Taghva
SEKE4
1986 Some Characterizations of Finitely Specifiable Implicational Dependency Families
Kazem Taghva
Inf. Process. Lett.1