EDBT 2026 Demo / reviewers in the wild / expert
Kazem Taghva
dblp:35/5977
· DBLP profile ↗
20ranked-venue papers
14as first author
3since 2021 · last 2025
0000-0001-8320-0080ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 5 first-authorTheory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Infomod: information-theoretic machine learning model diagnostics
Armin Esmaelizadeh, Sunil Cotterill, Liam Hebert, Lukasz Golab, Kazem Taghva |
Distributed Parallel Databases | 5 |
| 2025 | Tri-AL: An open source platform for visualization and analysis of clinical trials
Pouyan Nahed, Mina Esmail Zadeh Nojoo Kambar, Kazem Taghva, Lukasz Golab |
Inf. Syst. | 3 |
| 2023 | InfoMoD: Information-theoretic Model DiagnosticsabstractValidating and debugging machine learning models is done by testing them on unseen data. Analyzing model performance on various subsets of the data is critical for fairness, trust, bias detection and explainablility. In this paper, we describe a new way to do this. Our solution, called InfoMoD, applies recent work in information-theoretic data summarization to the problem of model diagnostics. Using real-life datasets, we show how InfoMod concisely describes how a model performs across different subsets of the data and produces expected performance indicators for individual test instances. Armin Esmaeilzadeh, Lukasz Golab, Kazem Taghva |
SSDBM | 3 |
| 2011 | Acronym Expansion Via Hidden Markov ModelsabstractIn this paper, we report on design and implementation of a Hidden Markov Model (HMM) to extract acronyms and their expansions. We also report on the training of this HMM with Maximum Likelihood Estimation (MLE) algorithm using a set of examples. Finally, we report on our testing using standard recall and precision. The HMM achieves a recall and precision of 98% and 92% respectively. Kazem Taghva, Lakshmi Vyas |
ICSEng | 1 |
| 2007 | Extracting _Carbon Copy_ Names and Organizations from a Heterogeneous Document CollectionabstractWe describe the development of a tool to identify the names of persons and their corresponding institutions as they appear in the "carbon copy" or "cc" lists of correspondence, memorandums, emails, and other types of documents. Kazem Taghva, Russell Beckley, Jeffrey S. Coombs |
ICDAR | 1 |
| 2006 | The Effects of OCR Error on the Extraction of Private Information
Kazem Taghva, Russell Beckley, Jeffrey S. Coombs |
Document Analysis Systems | 1 |
| 2005 | Document analysis by processing JBIG-encoded images
Emma E. Regentova, Shahram Latifi, De Chen, Kazem Taghva, Dongsheng Yao |
Int. J. Document Anal. Recognit. | 4 |
| 2004 | The role of manually-assigned keywords in query expansion
Kazem Taghva, Julie Borsack, Thomas A. Nartker, Allen Condit |
Inf. Process. Manag. | 1 |
| 2003 | A comparison of automatic and manual zoning
Kazem Taghva, Julie Borsack, Steven E. Lumos, Allen Condit |
Int. J. Document Anal. Recognit. | 1 |
| 2002 | Hairetes: A Search Engine for OCR Documents
Kazem Taghva, Jeffrey S. Coombs |
Document Analysis Systems | 1 |
| 2001 | OCRSpell: an interactive spelling correction system for OCR errors in text
Kazem Taghva, Eric Stofsky |
Int. J. Document Anal. Recognit. | 1 |
| 1999 | Recognizing acronyms and their definitions
Kazem Taghva, Jeff Gilbreth |
Int. J. Document Anal. Recognit. | 1 |
| 1996 | Effects of OCR Errors on Ranking and Feedback Using the Vector Space Model
Kazem Taghva, Julie Borsack, Allen Condit |
Inf. Process. Manag. | 1 |
| 1996 | Evaluation of Model-Based Retrieval Effectiveness with OCR TextabstractWe give a comprehensive report on our experiments with retrieval from OCR-generated text using systems based on standard models of retrieval. More specifically, we show that average precision and recall is not affected by OCR errors across systems for several collections. The collections used in these experiments include both actual OCR-generated text and standard information retrieval collections corrupted through the simulation of OCR errors. Both the actual and simulation experiments include full-text and abstract-length documents. We also demonstrate that the ranking and feedback methods associated with these models are generally not robust enough to deal with OCR errors. It is further shown that the OCR errors and garbage strings generated from the mistranslation of graphic objects increase the size of the index by a wide margin. We not only point out problems that can arise from applying OCR text within an information retrieval environment, we also suggest solutions to overcome some of these problems. Kazem Taghva, Julie Borsack, Allen Condit |
ACM Trans. Inf. Syst. | 1 |
| 1995 | Post-Editing Through Approximation and Global CorrectionabstractThis paper describes a new automatic spelling correction program to deal with OCR generated errors. The method used here is based on three principles: 1. Approximate string matching between the misspellings and the terms occuring in the database as opposed to the entire dictionary 2. Local information obtained from the individual documents 3. The use of a confusion matrix, which contains information inherently specific to the nature of errors caused by the particular OCR device This system is then utilized to process approximately 10,000 pages of OCR generated documents. Among the misspellings discovered by this algorithm, about 87% were corrected. Kazem Taghva, Julie Borsack, Bryan Bullard, Allen Condit |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1994 | Results of Applying Probabilistic IR to OCR Text
Kazem Taghva, Julie Borsack, Allen Condit |
SIGIR | 1 |
| 1994 | The Effects of Noisy Data on Text RetrievalabstractWe report on the results of our experiments on query evaluation in the presence of noisy data. In particular, an OCR-generated database and its corresponding 99.8% correct version are used to process a set of queries to determine the effect the degraded version will have on retrieval. It is shown that, with the set of scientific documents we use in our testing, the effect is insignificant. We further improve the result by applying an automatic postprocessing system designed to correct the kinds of errors generated by recognition devices. © 1994 John Wiley & Sons, Inc. Kazem Taghva, Julie Borsack, Allen Condit, Srinivas Erva |
J. Am. Soc. Inf. Sci. | 1 |
| 1993 | Capturing Strong Reduction in Director String Calculus
Vugranam C. Sreedhar, Kazem Taghva |
Theor. Comput. Sci. | 2 |
| 1989 | A Novel Way 1o Identify IneguaIity Query Subclasses Which possess the Homomorphism Property
Tianzheng Wu, James L. Clark, Nong Zhou, Kazem Taghva |
SEKE | 4 |
| 1986 | Some Characterizations of Finitely Specifiable Implicational Dependency Families
Kazem Taghva |
Inf. Process. Lett. | 1 |