Gerard Salton

dblp:s/GerardSalton · also Gerard A. Salton · DBLP profile ↗
← Back
77ranked-venue papers
63as first author
0since 2021 · last 1997
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 58 · 49 first-authorApplied, interdisciplinary, general and emerging computing · 12 · 8 first-authorArtificial intelligence and machine learning · 4 · 4 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorTheory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
26 papers
Information retrieval · 100% Indexing and storage engines · 0%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 91% Electronic design automation · 7% Distributed systems · 2%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
text analysis
0.031995
Selective text utilization and text traversal · Int. J. Hum. Comput. Stud. 1995
Automatic Text Structuring and Retrieval: Experiments in Automatic Encyclopedia Searching · SIGIR 1991
On the Application of Syntactic Methodologies in Automatic Text Analysis · SIGIR 1989
Information retrieval
relevance feedback
0.021995
Optimization of Relevance Feedback Weights · SIGIR 1995
The Effect of Adding Relevance Information in a Relevance Feedback Environment · SIGIR 1994
Information retrieval
retrieval models
0.091991
The SMART Information Retrieval System after 30 years - Panel · SIGIR 1991
Syntactic Approaches to Automatic Book Indexing · ACL 1988
The Use of Extended Boolean Logic in Information Retrieval · SIGMOD Conference 1984
Information retrieval › indexing
document indexing
0.071989
On the Application of Syntactic Methodologies in Automatic Text Analysis · SIGIR 1989
Syntactic Approaches to Automatic Book Indexing · ACL 1988
Recent Trends in Automatic Information Retrieval · SIGIR 1986
Information retrieval › similarity measure
document similarity
0.011995
Selective text utilization and text traversal · Int. J. Hum. Comput. Stud. 1995
Information retrieval › retrieval models
term weighting
0.051986
Recent Trends in Automatic Information Retrieval · SIGIR 1986
Term Weighting in Information Retrieval Using the Term Precision Model · J. ACM 1982
A Comparison of Search Term Weighting: Term Relevance vs. Inverse Document Frequency · SIGIR 1981
Information retrieval › search engines
full-text search
0.011993
Approaches to Passage Retrieval in Full Text Information Systems · SIGIR 1993
Information retrieval › document retrieval
passage retrieval
0.011993
Approaches to Passage Retrieval in Full Text Information Systems · SIGIR 1993
Information retrieval › document retrieval
heterogeneous document retrieval
0.011991
Automatic Text Structuring and Retrieval: Experiments in Automatic Encyclopedia Searching · SIGIR 1991
Information retrieval › retrieval models
vector space model
0.011991
The SMART Information Retrieval System after 30 years - Panel · SIGIR 1991
Information retrieval › retrieval models
probabilistic retrieval model
0.031986
Recent Trends in Automatic Information Retrieval · SIGIR 1986
Some Research Problems in Automatic Information Retrieval · SIGIR 1983
A Comparison of Search Term Weighting: Term Relevance vs. Inverse Document Frequency · SIGIR 1981
Information retrieval
indexing
0.031987
Parallel Architecture in IR · SIGIR 1987
Effective Automatic Indexing Using Term Addition and Deletion · J. ACM 1978
Precision Weighting - An Effective Automatic Indexing Method · J. ACM 1976
Information retrieval › retrieval models
boolean retrieval
0.021985
Automatic Assignment of Soft Boolean Operators · SIGIR 1985
The Use of Extended Boolean Logic in Information Retrieval · SIGMOD Conference 1984
Information retrieval › query reformulation
query expansion
0.021988
On the Use of Spreading Activation Methods in Automatic Information Retrieval · SIGIR 1988
Effective Automatic Indexing Using Term Addition and Deletion · J. ACM 1978
Information retrieval
search engines
0.031988
Syntactic Approaches to Automatic Book Indexing · ACL 1988
Recent Studies in Automatic Text Analysis and Document Retrieval · J. ACM 1973
Computer Evaluation of Indexing and Text Processing · J. ACM 1968
Information retrieval › indexing › text indexing
phrase indexing
0.011988
Syntactic Approaches to Automatic Book Indexing · ACL 1988
Information retrieval › retrieval models › graph-based retrieval
spreading activation
0.011988
On the Use of Spreading Activation Methods in Automatic Information Retrieval · SIGIR 1988
Information retrieval › retrieval models › boolean retrieval
extended boolean retrieval
0.011985
Automatic Assignment of Soft Boolean Operators · SIGIR 1985
Information retrieval
query formulation
0.011985
Automatic Assignment of Soft Boolean Operators · SIGIR 1985
Information retrieval
evaluation
0.031991
The SMART Information Retrieval System after 30 years - Panel · SIGIR 1991
Recent Studies in Automatic Text Analysis and Document Retrieval · J. ACM 1973
Computer Evaluation of Indexing and Text Processing · J. ACM 1968
Information retrieval › query processing
boolean query processing
0.011983
Some Research Problems in Automatic Information Retrieval · SIGIR 1983
Information retrieval › evaluation
test collection
0.011991
The SMART Information Retrieval System after 30 years - Panel · SIGIR 1991
Information retrieval › retrieval models › language model
term dependency models
0.011982
An Evaluation of Term Dependence Models in Information Retrieval · SIGIR 1982
Information retrieval › evaluation
retrieval effectiveness
0.011978
Generation and Search of Clustered Files · ACM Trans. Database Syst. 1978
Information retrieval › retrieval models › term weighting
relevance weighting
0.011976
Precision Weighting - An Effective Automatic Indexing Method · J. ACM 1976
Information retrieval › ranking › text ranking
document ranking
0.011973
Recent Studies in Automatic Text Analysis and Document Retrieval · J. ACM 1973
Information retrieval › retrieval models
associative retrieval
0.011963
Associative Document Retrieval Techniques Using Bibliographic Information · J. ACM 1963
Information retrieval › document processing › document analysis
document representation
0.011963
Associative Document Retrieval Techniques Using Bibliographic Information · J. ACM 1963
Electronic design automation
logic synthesis
0.011960
The Use of Parenthesis-Free Notation for the Automatic Design of Switching Circuits · IRE Trans. Electron. Comput. 1960
Distributed systems
distributed coordination
0.011960
A New Method for the Payment of Bills and the Transfer of Credit · J. ACM 1960

Methods — techniques the papers use, named apart from their topics

syntactic analysis · 0.0text matching · 0.0probabilistic retrieval · 0.0machine readable dictionary · 0.0knowledge base · 0.0spreading activation · 0.0nominal construction identification · 0.0importance weighting · 0.0thesaurus aids · 0.0extended boolean retrieval · 0.0cascading transformations · 0.0
YearPublicationVenuePosition
1997 Automatic Text Structuring and Summarization
abstract
In recent years, information retrieval techniques have been used for automatic generation of semantic hypertext links. This study applies the ideas from the automatic link generation research to attack another important problem in text processing—automatic text summarization. An automatic “general purpose” text summarization tool would be of immense utility in this age of information overload. Using the techniques used (by most automatic hypertext link generation algorithms) for inter-document link generation, we generate intra-document links between passages of a document. Based on the intra-document linkage pattern of a text, we characterize the structure of the text. We apply the knowledge of text structure to do automatic text summarization by passage extraction. We evaluate a set of fifty summaries generated using our techniques by comparing them to paragraph extracts constructed by humans. The automatic summarization methods perform well, especially in view of the fact that the summaries generated by two humans for the same article are surprisingly dissimilar.
Gerard Salton, Amit Singhal 0001, Mandar Mitra, Chris Buckley
Inf. Process. Manag.1
1996 Automatic Text Decomposition and Structuring
abstract
Sophisticated text similarity measurements are used to determine relationships between natural-language texts and text excerpts. The resulting linked hypertext maps can be decomposed into text segments and text themes, and these decompositions are usable to identify different text types and text structures, leading to improved text access and utilization. Examples of text decomposition are given for expository and non-expository texts.
Gerard Salton, James Allan 0001, Amit Singhal 0001
Inf. Process. Manag.1
1996 Document Length Normalization
abstract
In the TREC collection—a large full-text experimental text collection with widely varying document lengths—we observe that the likelihood of a document being judged relevant by a user increases with the document length. We show that a retrieval strategy, such as the vector-space cosine match, that retrieves documents of different lengths with roughly equal chances, will not optimally retrieve useful documents from such a collection. We present a modified technique—pivoted cosine normalization—that attempts to match the likelihood of retrieving documents of all lengths to the likelihood of their relevance, and show that this technique yields significant improvements in retrieval effectiveness.
Amit Singhal 0001, Gerard Salton, Mandar Mitra, Chris Buckley
Inf. Process. Manag.2
1995 Optimization of Relevance Feedback Weights
Chris Buckley, Gerard Salton
SIGIR2
1995 Selective text utilization and text traversal
abstract
Many large collections of full-text documents are currently stored in machine-readable form and processed automatically in various ways. These collections may include different types of documents, such as messages, research articles, and books, and the subject matter may vary widely. To process such collections, robust text analysis methods must be used, capable of handling materials in arbitrary subject areas, and flexible access must be provided to texts and text excerpts of varying size. In this study, global text comparison methods are used to identify similarities between text elements, followed by local context-checking operations that resolve ambiguities and distinguish superficially similar texts from texts that actually cover identical topics. A linked text structure, known as a text relationship map, is then created that relates similar texts at various levels of detail. In particular, text links are available for full texts, as well as text sections, paragraphs, and sentence groups. The relationship graphs are usable as conceptualization tools to illustrate various text manipulation operations and may also serve as browsing maps in situations where searches or text traversal operations are conducted under user control. In this study, the relationship maps are used to identify important text passages, to traverse texts selectively both within particular documents and between documents, and to provide flexible text access to large text collections in response to various kinds of user needs. An automated 29-volume encyclopedia is used as an example to illustrate various possible text accessing and traversal operations. Implementation details are not included in this initial study.
Gerard Salton, James Allan 0001
Int. J. Hum. Comput. Stud.1
1995 Automatic Routing and Retrieval Using Smart: TREC-2
abstract
The Smart information retrieval project emphasizes completely automatic approaches to the understanding and retrieval of large quantities of text. We continue our work in the TREC 2 environment, performing both routing and ad-hoc experiments. The ad-hoc work extends our investigations into combining global similarities, giving an overall indication of how a document matches a query, with local similarities identifying a smaller part of the document that matches the query. The performance of the ad-hoc runs is good, but it is clear we are not yet taking full advantage of the available local information. Our routing experiments use conventional relevance feedback approaches to routing, but with a much greater degree of query expansion than was previously done. The length of a query vector is increased by a factor of 5 to 10 by adding terms found in previously seen relevant documents. This approach improves effectiveness by 30–40% over the original query.
Chris Buckley, James Allan 0001, Gerard Salton
Inf. Process. Manag.3
1994 The Effect of Adding Relevance Information in a Relevance Feedback Environment
Chris Buckley, Gerard Salton, James Allan 0001
SIGIR2
1993 Approaches to Passage Retrieval in Full Text Information Systems
abstract
Large collections of full-text documents are now commonly used in automated information retrieval. When the stored document texts are long, the retrieval of complete documents may not be in the users' best interest. In such circumstance, efficient and effective retrieval results may be obtained by using passage retrieval strategies designed to retrieve text excerpts of varying size in response to statements of user interest.
Gerard Salton, James Allan 0001, Chris Buckley
SIGIR1
1992 The State of Retrieval System Evaluation
Gerard Salton
Inf. Process. Manag.1
1991 The SMART Information Retrieval System after 30 years - Panel
abstract
Article The smart document retrieval project Share on Author: Gerard Salton Cornell University Cornell UniversityView Profile Authors Info & Claims SIGIR '91: Proceedings of the 14th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 1991 Pages 356–358https://doi.org/10.1145/122860.122897Online:01 September 1991Publication History 29citation1,108DownloadsMetricsTotal Citations29Total Downloads1,108Last 12 Months29Last 6 weeks8 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Gerard Salton
SIGIR1
1991 Automatic Text Structuring and Retrieval: Experiments in Automatic Encyclopedia Searching
abstract
Many conventional approaches to text analysis and information retrieval prove ineffective when large text collections must be processed in heterogeneous subject areas. An alternative text manipulation system is outlined useful for the retrieval of large heterogeneous texts, and for the recognition of content similarities between text excerpts, based on flexible text matching procedures carried out in several contexts of different scope. The methods are illustrated by search experiments performed with the 29-volume Funk and Wagnalls encyclopedia.
Gerard Salton, Chris Buckley
SIGIR1
1991 Evaluation of an adaptive linear model
abstract
This article reports on the experimental evaluation of an adaptive linear model that constructs improved query vectors from the user preference judgments on a sample set of documents. The performance of this method is compared with that of the standard relevance feedback techniques. The experimental results seem to demonstrate the effectiveness of the adaptive method. © 1991 John Wiley & Sons, Inc.
S. K. Michael Wong, Yiyu Yao, Gerard Salton, Chris Buckley
J. Am. Soc. Inf. Sci.3
1990 On the application of syntactic methodologies in automatic text analysis
Gerard Salton, Chris Buckley, Maria Smith
Inf. Process. Manag.1
1990 Improving retrieval performance by relevance feedback
abstract
Relevance feedback is an automatic process, introduced over 20 years ago, designed to produce improved query formulations following an initial retrieval operation. The principal relevance feedback methods described over the years are examined briefly, and evaluation data are included to demonstrate the effectiveness of the various methods. Prescriptions are given for conducting text retrieval operations iteratively using relevance feedback. © 1990 John Wiley & Sons, Inc.
Gerard Salton, Chris Buckley
J. Am. Soc. Inf. Sci.1
1989 ASIS Panel on New Developments and Future Prospects for Electronic Databases
Gerard Salton
SIGIR1
1989 On the Application of Syntactic Methodologies in Automatic Text Analysis
abstract
This study summarizes various linguistic approaches proposed for document analysis in information retrieval environments. Included are standard syntactic methods to generate complex content identifiers, and the use of semantic know-how obtained from machine-readable dictionaries and from specially constructed knowledge bases. A particular syntactic analysis methodology is also outlined and its usefulness for the automatic construction of book indexes is examined.
Gerard Salton, Maria Smith
SIGIR1
1988 Syntactic Approaches to Automatic Book Indexing
abstract
Automatic book indexing systems are based on the generation of phrase structures capable of reflecting text content. Some approaches are given for the automatic construction of back-of-book indexes using a syntactic analysis of the available texts, followed by the identification of nominal constructions, the assignment of importance weights to the term phrases, and the choice of phrases as indexing units.
Gerard Salton
ACL1
1988 On the Use of Spreading Activation Methods in Automatic Information Retrieval
abstract
Spreading activation methods have been recommended in information retrieval to expand the search vocabulary and to complement the retrieved document sets. The spreading activation strategy is reminiscent of earlier associative indexing and retrieval systems. Some spreading activation procedures are briefly described, and evaluation output is given, reflecting the effectiveness of one of the proposed procedures.
Gerard Salton, Chris Buckley
SIGIR1
1988 A Simple Blueprint for Automatic Boolean Query Processing
Gerard Salton
Inf. Process. Manag.1
1988 Term-Weighting Approaches in Automatic Text Retrieval
Gerard Salton, Chris Buckley
Inf. Process. Manag.1
1987 Parallel Architecture in IR
Esen A. Ozkarahan, Craig Stanfill, Gerard Salton
SIGIR3
1987 Rejoinder to Nahum Goldmann's letter
Gerard Salton
J. Am. Soc. Inf. Sci.1
1987 The past thirty years in information retrieval
abstract
The documentation literature of the 1950s is reviewed briefly, and some early text processing endeavors are discussed. Various predictions made in 1960 by Mooers about the creative role of computers in information retrieval are then considered, and an attempt is made to explain why some of the more exciting predictions have not been fulfilled. Conclusions are drawn concerning the limits of computer power in text retrieval applications. © 1987 John Wiley & Sons, Inc.
Gerard Salton
J. Am. Soc. Inf. Sci.1
1986 On the Use of Term Associations in Automatic Information Retrieval
Gerard Salton
COLING1
1986 Recent Trends in Automatic Information Retrieval
abstract
Substantial successes were achieved in the early years in automatic indexing and retrieval using single term indexing theories with term weight assignments based on frequency considerations. The development of more refined indexing systems using thesaurus aids and automatically constructed term association maps changed the retrieval effectiveness only slightly. The recent introduction of the relevance concept in the form of probabilistic retrieval models provided a firm basis for term weighting and document ranking practices. However, the probabilistic methods were not helpful in substantially enhancing the retrieval effectiveness.
Gerard Salton
SIGIR1
1986 Enhancement of text representations using related document titles
Gerard Salton
Inf. Process. Manag.1
1985 On The Representation Of Query Term Relations By Soft Boolean Operators
Gerard Salton
EACL1
1985 Automatic Assignment of Soft Boolean Operators
abstract
The conventional bibliographic retrieval systems are based on Boolean query formulations and inverted file implementations. Such systems provide rapid responses in answer to search queries but they are not easy to use by uninitiated patrons. An extended Boolean retrieval strategy has been devised in which the Boolean operators are treated more or less strictly, depending on the setting of a special parameter, known as the p-value. The extended system is much more forgiving than the conventional system, and provides better retrieval effectiveness. In this study various problems associated with the determination of appropriate p-values are discussed, and suggestions are made for an automatic assignment of p-values. Evaluation output is included to illustrate the operations of the suggested procedures.
Gerard Salton, Ellen M. Voorhees
SIGIR1
1985 A note about information science research
abstract
Abstract This note deals with the relationship between information science research and practice. The impression that the field is moribund and that the research output is uniformly inferior is not supported by an examination of the information retrieval literature.
Gerard Salton
J. Am. Soc. Inf. Sci.1
1985 Advanced feedback methods in information retrieval
abstract
Abstract Automatic feedback methods may be used in online information retrieval to generate improved query statements based on information contained in previously retrieved documents. In this study automatic relevance feedback techniques are applied to Boolean query statements. The feedback operations are carried out using both the conventional Boolean logic, as well as an extended logic producing improved retrieval effectiveness. Experimental output is included to evaluate the automatic feedback operations.
Gerard Salton, Edward A. Fox, Ellen M. Voorhees
J. Am. Soc. Inf. Sci.1
1984 The Use of Extended Boolean Logic in Information Retrieval
abstract
An extended Boolean retrieval strategy has previously been introduced in which the individual Boolean operators can be treated more or less strictly, depending on the perceived strength of association of the query terms. The extended Boolean system is illustrated by examples and evaluation output is used to demonstrate the effectiveness of the operations.
Gerard Salton
SIGMOD Conference1
1984 A comparison of two methods for boolean query relevancy feedback
Gerard Salton, Ellen M. Voorhees, Edward A. Fox
Inf. Process. Manag.1
1984 Syntactically based indexing
Gerard Salton
J. Am. Soc. Inf. Sci.1
1983 Some Research Problems in Automatic Information Retrieval
abstract
Information retrieval components are currently incorporated in several types of information systems, including bibliographic retrieval systems, data base management systems and question-answering systems. Some of the problems arising in the real-time environment in which these systems operate are briefly discussed. Certain recent advances in information retrieval research are then mentioned, including the formulation of new probabilistic retrieval models, and the development of automatic document analysis and Boolean query processing techniques.
Gerard Salton
SIGIR1
1983 Automatic query formulations in information retrieval
abstract
Modern information retrieval systems are designed to supply relevant information in response to requests received from the user population. In most retrieval environments the search requests consist of keywords, or index terms, interrelated by appropriate Boolean operators. Since it is difficult for untrained users to generate effective Boolean search requests, trained search intermediaries are normally used to translate original statements of user need into useful Boolean search formulations. Methods are introduced in this study which reduce the role of the search intermediaries by making it possible to generate Boolean search formulations completely automatically from natural language statements provided by the system patrons. Frequency considerations are used automatically to generate appropriate term combinations as well as Boolean connectives relating the terms. Methods are covered to produce automatic query formulations both in a standard Boolean logic system, as well as in an extended Boolean system in which the strict interpretation of the connectives is relaxed. Experimental results are supplied to evaluate the effectiveness of the automatic query formulation process, and methods are described for applying the automatic query formulation process in practice.
Gerard Salton, Chris Buckley, Edward A. Fox
J. Am. Soc. Inf. Sci.1
1982 An Evaluation of Term Dependence Models in Information Retrieval
Gerard Salton, Chris Buckley, Clement T. Yu
SIGIR1
1982 Term Weighting in Information Retrieval Using the Term Precision Model
abstract
At3STRACT It iS known that the use of weighted, as opposed to binary, content identifiers attached to the records of an information file improves the effectiveness of the retrieval operations Under well-defined conditions the term precision offers the best possible term weighting system A mathematscal model is used in the present study to relate the term precision weights to the frequency of occurrence of the terms in a given document collecuon and to the number of relevant documents a user wishes to retrieve in response to a query This provides for the assignment of user-dependent weights to the content identifiers and relates the term precision weights to other well-known term weighting systems Categories and Subject Descriptors.H 3 1 [
Clement T. Yu, Gerard Salton
J. ACM3
1981 A Comparison of Search Term Weighting: Term Relevance vs. Inverse Document Frequency
abstract
The term relevance weighting method has been shown to produce optimal information retrieval queries under well-defined conditions. The parameters needed to generate the term relevance factors cannot unfortunately be estimated accurately in practice; futhermore, in realistic test situations, it appears difficult to obtain improved retrieval results using the term relevance weights over much simpler term weighting systems such as, for example, the inverse document frequency weights.It is shown in this study that the inverse document frequency weights and the term relevance weights are closely related over a wide range of the frequency spectrum. Methods are introduced for estimating the term relevance weights, and experimental results are given comparing the inverse document frequency with the estimated term relevance weights.
Harry Wu, Gerard Salton
SIGIR2
1981 The measurement of term importance in automatic indexing
abstract
Abstract The frequency characteristics of terms in the documents of a collection have been used as indicators of term importance for content analysis and indexing purposes. In particular, very rare or very frequent terms are normally believed to be less effective than medium‐frequency terms. Recently automatic indexing theories have been devised that use not only the term frequency characteristics but also the relevance properties of the terms. The major term‐weighting theories are first briefly reviewed. The term precision and term utility weights that are based on the occurrence characteristics of the terms in the relevant, as opposed to the nonrelevant, documents of a collection are then introduced. Methods are suggested for estimating the relevance properties of the terms based on their overall occurrence characteristics in the collection. Finally, experimental evaluation results are shown comparing the weighting systems using the term relevance properties with the more conventional frequency‐based methodologies.
Gerard Salton, Harry Wu, Clement T. Yu
J. Am. Soc. Inf. Sci.1
1980 A Term Weighting Model Based on Utility Theory
Gerard Salton, Harry Wu
SIGIR1
1980 Automatic term class construction using relevance--A summary of work in automatic pseudoclassification
Gerard Salton
Inf. Process. Manag.1
1980 A progress report on information privacy and data security
abstract
Abstract The role and importance of information privacy in the modern society are briefly described. A number of recent law cases are then examined to illustrate how privacy cases are currently being adjudicated in the United States and to identify the limits of currently available privacy protection. Finally, certain issues are raised regarding the available techniques for insuring data confidentiality and security.
Gerard Salton
J. Am. Soc. Inf. Sci.1
1980 Buck's prime number coding scheme
Gerard Salton, Ana D. Cleveland
J. Am. Soc. Inf. Sci.1
1979 Progress Report on Automatic Information Retrieval
abstract
No abstract available.
Gerard Salton
SIGIR1
1979 Science, shcharansky, and the soviets
Gerard Salton
J. Am. Soc. Inf. Sci.1
1979 Automatic text analysis
Gerard Salton
J. Am. Soc. Inf. Sci.1
1978 Best-match querying in general database systems-a language approach
abstract
We reason in this paper that many queries in general database systems are best-match in nature: to a user, some records are more useful than others, and when the best records cannot be found or when there are not a sufficient number of them, the next best records should be retrieved. The heterogeneity of general databases will require different treatment on such queries than that in some special-purpose systems where the concept of best-match has been exploited. Language features are proposed to augment typical query languages in order to accommodate best-match querying. This is intended as another effort toward the design of convenient, expressive user interface to facilitate decision making via database systems.
Chung-Shu Yang, Gerard Salton
COMPSAC2
1978 Term relevance weights in on-line information retrieval
Gerard Salton, R. K. Waldstein
Inf. Process. Manag.1
1978 Effective Automatic Indexing Using Term Addition and Deletion
abstract
In mformaUon retrieval indexing is the task consisting of the assignment to stored records and mcommg mformatton requests of content ~dent~fiers capable of representing record or query content If the mdexmg is performed automatically and the records are wntten documents, an mmal set of index terms might be chosen by taking words extracted from document roles or abstracts, this mmal term assignment might then be improved by addmg related terms chosen from a thesaurus, by deletmg extraneous or marginal terms, and by replacing smgle terms by term combmaaons and phrases.In the present study formal proofs are given of the retrieval effectiveness under well-defined condmons of mdexmg policies based on the use of single terms, term additions and deletions, and term combmaaons or phrases.
Clement T. Yu, Gerard Salton, Man-Keung Siu
J. ACM2
1978 Generation and Search of Clustered Files
abstract
A classified, or clustered file is one where related, or similar records are grouped into classes, or clusters of items in such a way that all items within a cluster are jointly retrievable. Clustered files are easily adapted to broad and narrow search strategies, and simple file updating methods are available. An inexpensive file clustering method applicable to large files is given together with appropriate file search methods. An abstract model is then introduced to predict the retrieval effectiveness of various search methods in a clustered file environment. Experimental evidence is included to test the versatility of the model and to demonstrate the role of various parameters in the cluster search process.
Gerard Salton, Anita Wong
ACM Trans. Database Syst.1
1976 Automatic indexing using term discrimination and term precision measurements
Gerard Salton, Anita Wong, Clement T. Yu
Inf. Process. Manag.1
1976 Precision Weighting - An Effective Automatic Indexing Method
abstract
A great many automatic indexing methods have been implemented and evaluated over the last few years, and automatic procedures comparable in effectiveness to conventional manual ones are now easy to generate. Two drawbacks of the available automatic indexing methods are the absence of reliable linguistic inputs during the indexing process and the lack of formal, analytical proofs concerning the effectiveness of the proposed methods. The precision weighting procedure described in the present study uses relevance criteria to weight the terms occurring in user queries as a function of the balance between relevant and nonrelevant documents in which these terms occur; this approximates a semantic know-how of term importance. Formal mathematical proofs are given under well-defined conditions of the effectiveness of the method.
Clement T. Yu, Gerard Salton
J. ACM2
1975 A theory of term importance in automatic text analysis
abstract
Abstract A good deal of work has been done over the years in an attempt to use statistical or probabilistic techniques as a basis for automatic indexing and content analysis. (1–10) Unfortunately, many of these methods are lacking in effectiveness, and the more refined procedures are computationally unattractive. A new technique, known as discrimination value analysis, ranks the text words in accordance with how well they are able to discriminate the documents of a collection from each other; that is, the value of a term depends on how much the average separation between individual documents changes when the given term is assigned for content identification. The best words are those which achieve the greatest separation. The discrimination value analysis is computationally simple, and it assigns a specific role in content analysis to single words, juxtaposed words and phrases, and word groups or thesaurus categories. Experimental results are given showing the effectiveness of the technique.
Gerard Salton, Chung-Shu Yang, Clement T. Yu
J. Am. Soc. Inf. Sci.1
1974 Where the Action Is and Was in Information Science
Yehoshua Bar-Hillel, R. Carnap, E. C. Cherry, Eugene Garfield, D. W. King, F. W. Lancaster, J. C. R. Licklider, D. M. Mackay, J. W. Perry, D. J. De S. Price, Gerard Salton, Claude E. Shannon, Mortimer Taube, B. C. Vickery, Anthony E. Cawkell
J. Am. Soc. Inf. Sci.11
1973 Introductory programming at Cornell
abstract
The computer science department at Cornell is a graduate department. Approximately sixty degree candidates are formally enrolled in the computer science degree program, nearly all of them as Ph.D. candidates. There is no formal undergraduate major in computer science, although it is possible for really tenacious undergraduates in the College of Arts and Sciences, and in Engineering to obtain an undergraduate degree in computer science by special petition.
Gerard Salton
SIGCSE1
1973 Experiments in Multi-Lingual Information Retrieval
Gerard Salton
Inf. Process. Lett.1
1973 Automatic processing of current affairs queries
Gerard Salton
Inf. Storage Retr.1
1973 Recent Studies in Automatic Text Analysis and Document Retrieval
abstract
Many experts in mechanized text processing now agree that useful automatic language analysis procedures are largely unavailable and that the existing linguistic methodologies generally produce disappointing results. An attempt is made in the present study to identify those automatic procedures which appear most effective as a replacement for the missing language analysis. A series of computer experiments is described, designed to simulate a conventional document retrieval environment. It is found that a simple duplication, by automatic means, of the standard, manual document indexing and retrieval operations will not produce acceptable output results. New mechanized approaches to document handling are proposed, including document ranking methods, automatic dictionary and word list generation, and user feedback searches. It is shown that the fully automatic methodology is superior in effectiveness to the conventional procedures in normal use.
Gerard Salton
J. ACM1
1973 On the development of information science
abstract
Abstract The citations appearing in two recent comprehensive bibliographies in information science and technology are reviewed, and a comparison is made with bibliographies dating back to 1962. Some conclusions are drawn concerning the development and current state of information science.
Gerard Salton
J. Am. Soc. Inf. Sci.1
1972 Information science and the annual review
Gerard Salton
Inf. Storage Retr.1
1972 Comment on "an evaluation of query expansion by the addition of clustered terms for a document retrieval system"
Gerard Salton
Inf. Storage Retr.1
1972 What Is Computer Science?
abstract
No abstract available.
Gerard Salton
J. ACM1
1972 The "generality" effect and the retrieval evaluation for large collections
abstract
Abstract The retrieval effectiveness of large document collections is normally assessed by using small subsections of the file for test purposes, and extrapolating the data upward to represent the results for the full collection. The accuracy of such an extrapolation unhappily depends on the “generality” of the respective collections. In the present study the role of the generality effect in retrieval system evaluation is assessed, and evaluation results are given for the comparison of several document collections of distinct size and generality in the areas of documentation and aerodynamics.
Gerard Salton
J. Am. Soc. Inf. Sci.1
1972 Thoughts on the unisist feasibility study
Gerard Salton
J. Am. Soc. Inf. Sci.1
1972 A new comparison between conventional indexing (MEDLARS) and automatic text processing (SMART)
abstract
Abstract A new testing process is described designed to compare conventional retrieval (MEDLARS) and automatic text analysis methods (SMART). The results obtained with a collection of documents chosen independently of either SMART or MEDLARS indicate that a simple automatic extraction of keywords from document abstracts produces a 30 to 40 percent loss compared with MEDLARS indexing. A replacement of the unranked Boolean searches used in MEDLARS by the standard ranked output normally provided by SMART reduces the loss to between 15 and 20 percent. When an automatically generated word control list or a thesaurus is used as part of the SMART analysis, the results are comparable in effectiveness to those obtained by the intellectual MEDLARS indexing. Finally, the incorporation of user feedback procedures into SMART furnishes an improvement over the normal MEDLARS output of 15 to 30 percent. One concludes again that no technical justification exists for maintaining controlled, manual indexing in operational retrieval environments.
Gerard Salton
J. Am. Soc. Inf. Sci.1
1972 F. W. lancaster, Vocabulary control for information retrieval. Information Resources Press, Washington, 1972, 233 pages
Gerard Salton
J. Am. Soc. Inf. Sci.1
1971 The Performance of Interactive Information Retrieval
Gerard Salton
Inf. Process. Lett.1
1971 Some Thoughts on Scientific Information Dissemination
abstract
Most of us have been aware for some years of a crisis in scientific information processing, caused by the large mass of available information products, and reflected by
Gerard Salton
J. ACM1
1970 Evaluation problems in interactive information retrieval
Gerard Salton
Inf. Storage Retr.1
1970 On the Role of the ACM Journal
abstract
No abstract available.
Gerard Salton
J. ACM1
1969 Automatic Processing of Foreign Language Documents
Gerard Salton
COLING1
1969 A Policy for JACM
abstract
When the Journal of the Association for Computing Machinery was first issued in 1954, the intention was to make it into the definitive journal in the computer area. This in many ways it has become. Even in some of the applied areas—for example, in information retrieval—many of the articles consistently referred to over the years, and thus considered of fundamental importance, have appeared in the ACM Journal .
Gerard Salton
J. ACM1
1968 Relevance assessments and retrieval system evaluation
Michael E. Lesk, Gerard Salton
Inf. Storage Retr.2
1968 Computer Evaluation of Indexing and Text Processing
abstract
Automatic indexing methods are evaluated and design criteria for modern information systems are derived.
Gerard Salton, Michael E. Lesk
J. ACM1
1963 Associative Document Retrieval Techniques Using Bibliographic Information
abstract
Automatic documentation systems which use the words contained in the individual documents as a principal source of document identifications may not perform satisfactorily under all circumstances.Methods have therefore been devised within the last few years for computing association measures between words and between documents, and for using such associated words, or information contained in associated documents, to supplement and refine the original document identifications.It is suggested in this study that bibliographic citations may provide a simple means for obtaining associated documents to be incorporated in an automatic documentation system.The standard associative retrieval techniques are first briefly reviewed.A computer experiment is then described which tends to confirm the hypothesis that documents exhibiting similar citation sets also deal with similar subject matter.Finally, a fully automatic document retrieval system is proposed which uses bibliographic information in addition to other standard criteria for the identification of document content, and for the detection of relevant information.
Gerard Salton
J. ACM1
1960 A New Method for the Payment of Bills and the Transfer of Credit
abstract
The transfer of credit and the processing of receivables play important roles in most business systems. Present procedures require a dual handling of payments— they are processed both by the payee and again by the banks involved in the transfer of credit. The banks must necessarily process payments made by check; it is therefore reasonable to inquire whether the processing now required of the payee is not largely redundant, and whether it would not therefore be profitable to alter the present system. Three questions deserve examination: Can the present system be simplified? Will simplification lead to net savings? How should the savings be shared?
Gerard Salton
J. ACM1
1960 The Use of Parenthesis-Free Notation for the Automatic Design of Switching Circuits
abstract
A parenthesis-free notation is introduced for the representation of series-parallel switching networks. The notation facilitates the calculation of circuit parameters and permits an unambiguous characterization of the circuit topology. Given certain criteria for feasibility of a switching network related to the circuit parameter values, it is shown how an infeasible series-parallel network can be transformed into an equivalent feasible network by ``cascading'' operations applied to the two-terminal sub-networks of the original network. A systematic method is developed, resulting in an optimum choice of cascading operations such that the number of switching elements required to implement the transformed circuit is minimized relative to cascading.
Eugene L. Lawler, Gerard Salton
IRE Trans. Electron. Comput.2