VLDB 2026 Research / reviewers in the wild / expert
Robert Krovetz
dblp:75/664
· DBLP profile ↗
17ranked-venue papers
10as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 6 first-authorArtificial intelligence and machine learning · 7 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
11 papers |
Information retrieval · 84% Data mining · 16% | |
| Software engineering, system software, and programming languages
2 papers |
Software maintenance and evolution · 41% Program analysis · 41% Empirical software engineering · 19% | |
| Artificial intelligence
2 papers |
Knowledge representation and reasoning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis › machine learning for program analysis
code classification |
0.1 | 2 | 2003 | Classification of source code archives · SIGIR 2003 What's the code?: automatic classification of source code archives · KDD 2002 |
Information retrieval › document retrieval
cluster-based retrieval |
0.1 | 1 | 2005 | CLUE: cluster-based retrieval of images by unsupervised learning · IEEE Trans. Image Process. 2005 |
Information retrieval › image retrieval
content-based image retrieval |
0.1 | 1 | 2005 | CLUE: cluster-based retrieval of images by unsupervised learning · IEEE Trans. Image Process. 2005 |
Information retrieval
image retrieval |
0.1 | 1 | 2005 | CLUE: cluster-based retrieval of images by unsupervised learning · IEEE Trans. Image Process. 2005 |
Data mining › clustering
unsupervised learning |
0.1 | 1 | 2005 | CLUE: cluster-based retrieval of images by unsupervised learning · IEEE Trans. Image Process. 2005 |
Information retrieval › web search
web information retrieval |
0.0 | 1 | 2004 | Analysis of lexical signatures for improving information persistence on the World Wide Web · ACM Trans. Inf. Syst. 2004 |
Information retrieval › query understanding
word sense disambiguation |
0.0 | 4 | 1997 | Homonymy and Polysemy in Information Retrieval · ACL 1997 Viewing Morphology as an Inference Process · SIGIR 1993 Lexical Ambiguity and Information Retrieval · ACM Trans. Inf. Syst. 1992 |
Software maintenance and evolution
software reuse |
0.0 | 1 | 2003 | Classification of source code archives · SIGIR 2003 |
Information retrieval
document retrieval |
0.0 | 1 | 2002 | Analysis of lexical signatures for finding lost or related documents · SIGIR 2002 |
Empirical software engineering
mining software repositories |
0.0 | 1 | 2002 | What's the code?: automatic classification of source code archives · KDD 2002 |
Computational social science and digital humanities
computational linguistics |
0.0 | 1 | 2000 | Viewing morphology as an inference process · Artif. Intell. 2000 |
Data mining › text mining
text classification |
0.0 | 2 | 2003 | Classification of source code archives · SIGIR 2003 What's the code?: automatic classification of source code archives · KDD 2002 |
Information retrieval
retrieval models |
0.0 | 2 | 1997 | Homonymy and Polysemy in Information Retrieval · ACL 1997 Panel on the Lexicon and Information Retrieval · SIGIR 1989 |
Information retrieval
indexing |
0.0 | 2 | 1993 | Viewing Morphology as an Inference Process · SIGIR 1993 Word Sense Disambiguation Using Machine-Readable Dictionaries · SIGIR 1989 |
Information retrieval
ranking |
0.0 | 1 | 2004 | Analysis of lexical signatures for improving information persistence on the World Wide Web · ACM Trans. Inf. Syst. 2004 |
Information retrieval
search engines |
0.0 | 1 | 2004 | Analysis of lexical signatures for improving information persistence on the World Wide Web · ACM Trans. Inf. Syst. 2004 |
Information retrieval › text analysis
lexical semantics |
0.0 | 1 | 1993 | Viewing Morphology as an Inference Process · SIGIR 1993 |
Information retrieval › text analysis › text preprocessing
morphological analysis |
0.0 | 1 | 1993 | Viewing Morphology as an Inference Process · SIGIR 1993 |
Information retrieval › query understanding › query analysis
lexical ambiguity |
0.0 | 1 | 1992 | Lexical Ambiguity and Information Retrieval · ACM Trans. Inf. Syst. 1992 |
Information retrieval › indexing
stemming |
0.0 | 1 | 1993 | Viewing Morphology as an Inference Process · SIGIR 1993 |
Information retrieval
text analysis |
0.0 | 1 | 1992 | Corpus Linguistics and Information Retrieval · SIGIR 1992 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.2term frequency · 0.1entropy weighting · 0.1relevance feedback · 0.1graph-theoretic clustering · 0.1test and select · 0.0inverse document frequency · 0.0document frequency · 0.0TF-IDF · 0.0morphology · 0.0test collection analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Test Collection for Part-of-Speech Tagging and Word Sense Disambiguation
Robert Krovetz |
LREC | 1 |
| 2005 | CLUE: cluster-based retrieval of images by unsupervised learningabstractIn a typical content-based image retrieval (CBIR) system, target images (images in the database) are sorted by feature similarities with respect to the query. Similarities among target images are usually ignored. This paper introduces a new technique, cluster-based retrieval of images by unsupervised learning (CLUE), for improving user interaction with image retrieval systems by fully exploiting the similarity information. CLUE retrieves image clusters by applying a graph-theoretic clustering algorithm to a collection of images in the vicinity of the query. Clustering in CLUE is dynamic. In particular, clusters formed depend on which images are retrieved in response to the query. CLUE can be combined with any real-valued symmetric similarity measure (metric or nonmetric). Thus, it may be embedded in many current CBIR systems, including relevance feedback systems. The performance of an experimental image retrieval system using CLUE is evaluated on a database of around 60,000 images from COREL. Empirical results demonstrate improved performance compared with a CBIR system using the same image similarity measure. In addition, results on images returned by Google's Image Search reveal the potential of applying CLUE to real-world image data and integrating CLUE as a part of the interface for keyword-based image retrieval systems. Yixin Chen 0002, James Z. Wang 0001, Robert Krovetz |
IEEE Trans. Image Process. | 3 |
| 2004 | Analysis of lexical signatures for improving information persistence on the World Wide WebabstractA lexical signature (LS) consisting of several key words from a Web document is often sufficient information for finding the document later, even if its URL has changed. We conduct a large-scale empirical study of nine methods for generating lexical signatures, including Phelps and Wilensky's original proposal (PW), seven of our own static variations, and one new dynamic method. We examine their performance on the Web over a 10-month period, and on a TREC data set, evaluating their ability to both (1) uniquely identify the original (possibly modified) document, and (2) locate other relevant documents if the original is lost. Lexical signatures chosen to minimize document frequency (DF) are good at unique identification but poor at finding relevant documents. PW works well on the relatively small TREC data set, but acts almost identically to DF on the Web, which contains billions of documents. Term-frequency-based lexical signatures (TF) are very easy to compute and often perform well, but are highly dependent on the ranking system of the search engine used. The term-frequency inverse-document-frequency- (TFIDF-) based method and hybrid methods (which combine DF with TF or TFIDF) seem to be the most promising candidates among static methods for generating effective lexical signatures. We propose a dynamic LS generator called Test & Select (TS) to mitigate LS conflict. TS outperforms all eight static methods in terms of both extracting the desired document and finding relevant information, over three different search engines. All LS methods show significant performance degradation as documents in the corpus are edited. Seung-Taek Park, David M. Pennock, C. Lee Giles, Robert Krovetz |
ACM Trans. Inf. Syst. | 4 |
| 2003 | Classification of source code archivesabstractThe World Wide Web contains a number of source code archives. Programs are usually classified into various categories within the archive by hand. We report on experiments for automatic classification of source code into these categories. We examined a number of factors that affect classification accuracy. Weighting features by expected entropy loss makes a significant improvement in classification accuracy. We show a Support Vector Machine can be trained to classify source code with a high degree of accuracy. We feel these results show promise for software reuse. Robert Krovetz, Secil Ugurel, C. Lee Giles |
SIGIR | 1 |
| 2002 | Inferring hierarchical descriptionsabstractWe create a statistical model for inferring hierarchical term relationships about a topic, given only a small set of example web pages on the topic, without prior knowledge of any hierarchical information. The model can utilize either the full text of the pages in the cluster or the context of links to the pages. To support the model, we use "ground truth" data taken from the category labels in the Open Directory. We show that the model accurately separates terms in the following classes: self terms describing the cluster, parent terms describing more general concepts, and child terms describing specializations of the cluster. For example, for a set of biology pages, sample parent, self, and child terms are science, biology, and genetics respectively. We create an algorithm to predict parent, self, and child terms using the new model, and compare the predictions to the ground truth data. The algorithm accurately ranks a majority of the ground truth terms highly, and identifies additional complementary terms missing in the Open Directory. Eric J. Glover, David M. Pennock, Steve Lawrence, Robert Krovetz |
CIKM | 4 |
| 2002 | What's the code?: automatic classification of source code archivesabstractThere are various source code archives on the World Wide Web. These archives are usually organized by application categories and programming languages. However, manually organizing source code repositories is not a trivial task since they grow rapidly and are very large (on the order of terabytes). We demonstrate machine learning methods for automatic classification of archived source code into eleven application topics and ten programming languages. For topical classification, we concentrate on C and C++ programs from the Ibiblio and the Sourceforge archives. Support vector machine (SVM) classifiers are trained on examples of a given programming language or programs in a specified category. We show that source code can be accurately and automatically classified into topical categories and can be identified to be in a specific programming language class. Secil Ugurel, Robert Krovetz, C. Lee Giles |
KDD | 2 |
| 2002 | Analysis of lexical signatures for finding lost or related documentsabstractA lexical signature of a web page is often sufficient for finding the page, even if its URL has changed. We conduct a largescale empirical study of eight methods for generating lexi- cal signatures, including Phelps and Wilensky's [14] original proposal (PW) and seven of our own variations. We exmnine their performance on the web and on a TREC data set, evaluating their ability both to uniquely identify the origi- nal document and to locate other relevant documents if the original is lost. Lexical signatures chosen to minimize document frequency (DF) are good at unique identification but poor at finding relevant documents. PW works well on the relatively small TREC data set, but acts almost identically to DF on the web, which contains billions of documents. Term-frequency-based lexical signatures (TF) are very easy to compute and often perform well, but are highly dependent on the ranking system of the search engine used. In general, TFIDF-based method and hybrid methods (which combine DF with TF or TFIDF) seem to be the most promising candidates for generating effective lexical signatures. Seung-Taek Park, David M. Pennock, C. Lee Giles, Robert Krovetz |
SIGIR | 4 |
| 2000 | Persistence of information on the web: Analyzing citations contained in research articlesabstractWe analyze the persistence of information on the web, looking at the percentage of invalid URLs contained in academic articles within the CiteSeer (ResearchIndex) database.The number of URLs contained in the papers has increased from an average of 0.06 in 1993 to 1.6 in 1999.We found that a significant percentage of URLs are now invalid, ranging from 23% for 1999 articles, to 53% for 1994.We also found that for almost all of the invalid URLs, it was possible to locate the information (or highly related information) in an alternate location, primarily with the use of search engines.However, the ability to relocate missing information varied according to search experience and effort expended.Citation practices suggest that more information may be lost in the future unless these practices are improved.We discuss persistent URL standards and their usage, and give recommendations for citing URLs in research articles as well as for finding the new location of invalid URLs. Steve Lawrence, Frans Coetzee, Gary William Flake, David M. Pennock, Robert Krovetz, Finn Årup Nielsen, Andries Kruger, C. Lee Giles |
CIKM | 5 |
| 2000 | Viewing morphology as an inference process
Robert Krovetz |
Artif. Intell. | 1 |
| 1997 | Homonymy and Polysemy in Information RetrievalabstractThis paper discusses research on distinguishing word meanings in the context of information retrieval systems.We conducted experiments with three sources of evidence for making these distinctions: morphology, part-of-speech, and phrases.We have focused on the distinction between homonymy and polysemy (unrelated vs. related meanings).Our results support the need to distinguish homonymy and polysemy.We found: 1) grouping morphological variants makes a significant improvement in retrieval performance, 2) that more than half of all words in a dictionary that differ in part-of-speech are related in meaning, and 3) that it is crucial to assign credit to the component words of a phrase.These experiments provide a better understanding of word-based methods, and suggest where natural language processing can provide further improvements in retrieval performance. Robert Krovetz |
ACL | 1 |
| 1993 | Viewing Morphology as an Inference ProcessabstractMorphology is the area of linguistics concerned with the internal structure of words. Information Retrieval has generally not paid much attention to word structure, other than to account for some of the variability in word forms via the use of stemmers. This paper will describe our experiments to determine the importance of morphology, and the effect that it has on performance. We will also describe the role of morphological analysis in word sense disambiguation, and in identifying lexical semantic relationships in a machine-readable dictionary. We will first provide a brief overview of morphological phenomena, and then describe the experiments themselves. 1 Introduction Morphology is the area of linguistics concerned with the internal structure of words. It is usually broken down into two subclasses: inflectional and derivational. Inflectional morphology describes predictable changes a word undergoes as a result of syntax - the plural and possessive form for nouns, and the past tens... Robert Krovetz |
SIGIR | 1 |
| 1992 | Sense-Linking in a Machine Readable DictionaryabstractDictionaries contain a rich set of relationships between their senses, but often these relationships are only implicit. We report on our experiments to automatically identify links between the senses in a machine-readable dictionary. In particular, we automatically identify instances of zero-affix morphology, and use that information to find specific linkages between senses. This work has provided insight into the performance of a stochastic tagger. Robert Krovetz |
ACL | 1 |
| 1992 | Corpus Linguistics and Information Retrievalabstractnumber of papers about corpus analysis. Robert Krovetz |
SIGIR | 1 |
| 1992 | Lexical Ambiguity and Information RetrievalabstractLexical ambiguity is a pervasive problem in natural language processing. However, little quantitative information is available about the extent of the problem or about the impact that it has on information retrieval systems. We report on an analysis of lexical ambiguity in information retrieval test collections and on experiments to determine the utility of word meanings for separating relevant from nonrelevant documents. The experiments show that there is considerable ambiguity even in a specialized database. Word senses provide a significant separation between relevant and nonrelevant documents, but several factors contribute to determining whether disambiguation will make an improvement in performance. For example, resolving lexical ambiguity was found to have little impact on retrieval effectiveness for documents that have many words in common with the query. Other uses of word sense disambiguation in an information retrieval context are discussed. Robert Krovetz, W. Bruce Croft |
ACM Trans. Inf. Syst. | 1 |
| 1990 | Interactive retrieval of complex documents
W. Bruce Croft, Robert Krovetz, Howard R. Turtle |
Inf. Process. Manag. | 2 |
| 1989 | Panel on the Lexicon and Information Retrieval
Robert Krovetz |
SIGIR | 1 |
| 1989 | Word Sense Disambiguation Using Machine-Readable DictionariesabstractMost approachesto full-text information retrieval currently index documents based on the words they contain, and retrieve them based on the word's frequency of occurrence.This can cause many irrelevant documents to be retrieved because words are often ambiguous.We propose an approach in which documents are indexed by word aenaea, and in which these senses are taken from a machine-readable dictionary.We review some of the work on machine-readable dictionaries and the approaches that have been taken to word sense disambiguation.We then discuss our own approach to the problem based on the use of multiple sources of evidence.We conclude with the results of some experiments that indicate the degree to which lexical ambiguity is a factor in current systems. Robert Krovetz, W. Bruce Croft |
SIGIR | 1 |