Gordon Linoff

dblp:44/3920 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 1993
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 72% Query processing and optimization · 15% Data mining · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › sorting
external sorting
0.011993
A practical external sort for shared disk MPPs · SC 1993
Information retrieval
indexing
0.011993
Compression of Indexes with Full Positional Information in Very Large Text Databases · SIGIR 1993
Information retrieval › indexing › index compression
inverted index compression
0.011993
Compression of Indexes with Full Positional Information in Very Large Text Databases · SIGIR 1993
Information retrieval › indexing › inverted index
positional index
0.011993
Compression of Indexes with Full Positional Information in Very Large Text Databases · SIGIR 1993
Information retrieval
document retrieval
0.011992
Classifying News Stories using Memory Based Reasoning · SIGIR 1992
Information retrieval
relevance feedback
0.011992
Classifying News Stories using Memory Based Reasoning · SIGIR 1992
Data mining › text mining
text classification
0.011992
Classifying News Stories using Memory Based Reasoning · SIGIR 1992

Methods — techniques the papers use, named apart from their topics

run-length encoding · 0.0prefix omission · 0.0n-s coding · 0.0memory-based reasoning · 0.0k-nearest neighbor · 0.0
YearPublicationVenuePosition
1993 A practical external sort for shared disk MPPs
abstract
No abstract available.
Xiqing Li, Gordon Linoff, Stephen J. Smith, Craig Stanfill, Kurt H. Thearling
SC2
1993 Compression of Indexes with Full Positional Information in Very Large Text Databases
abstract
This paper describes a combination of compression methods which may be used to reduce the size of inverted indexes for very large text databases. These methods are Prefix Omission, Run-Length Encoding, and a novel family of numeric representations called n-s coding. Using these compression methods on two different text sources (the King James Version of the Bible and a sample of Wall Street Journal Stories), the compressed index occupies less than 40% of the size of the original text, even when both stopwords and numbers are included in the index. The decreased time required for I/O can almost fully compensate for the time needed to uncompress the postings. This research is part of an effort to handle very large text databases on the CM-5, a massively parallel MIMD supercomputer.
Gordon Linoff, Craig Stanfill
SIGIR1
1992 Classifying News Stories using Memory Based Reasoning
abstract
We describe a method for classifying news stories using Memory Based Reasoning (MBR) a k-nearest neighbor method), that does not require manual topic definitions. Using an already coded training database of about 50,000 stories from the Dow Jones Press Release News Wire, and SEEKER [Stanfill] (a text retrieval system that supports relevance feedback) as the underlying match engine, codes are assigned to new, unseen stories with a recall of about 80% and precision of about 70%. There are about 350 different codes to be assigned. Using a massively parallel supercomputer, we leverage the information already contained in the thousands of coded stories and are able to code a story in about 2 seconds. Given SEEKER, the text retrieval system, we achieved these results in about two person-months. We believe this approach is effective in reducing the development time to implement classification systems involving large number of topics for the purpose of classification, message routing etc.
Brij M. Masand, Gordon Linoff, David L. Waltz
SIGIR2