Ray R. Larson

dblp:68/4900 · DBLP profile ↗
← Back
24ranked-venue papers
18as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 24 · 18 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Information retrieval · 88% Data mining · 12%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › document retrieval › domain-specific retrieval
geographic information retrieval
0.122008
Geographic IR and visualization in time and space · SIGIR 2008
Geographic information retrieval (GIR): searching where and what · SIGIR 2004
Information retrieval
cross-language information retrieval
0.122004
Geotemporal querying of multilingual documents · SIGIR 2004
Translingual vocabulary mappings for multilingual information access · SIGIR 2002
Visualization and visual analytics
geospatial visualization
0.112008
Geographic IR and visualization in time and space · SIGIR 2008
Information retrieval
distributed information retrieval
0.122002
A logistic regression approach to distributed IR · SIGIR 2002
Distributed Resource Discovery and Structured Data Searching with Cheshire II · SIGIR 2001
Information retrieval › retrieval models
probabilistic retrieval model
0.142008
Geographic IR and visualization in time and space · SIGIR 2008
Cheshire II: Combining Probabilistic and Boolean Retrieval · SIGIR 1998
A logistic regression approach to distributed IR · SIGIR 2002
Data mining › text mining
information extraction
0.012004
Geotemporal querying of multilingual documents · SIGIR 2004
Information retrieval
retrieval models
0.032006
Cheshire II: Combining Probabilistic and Boolean Retrieval · SIGIR 1998
Cheshire3: retrieving from tera-scale grid-based digital libraries · SIGIR 2006
Text and Image Retrieval in Cheshire II (demonstration abstract) · SIGIR 1999
Data mining › predictive modeling › regression
logistic regression
0.012002
A logistic regression approach to distributed IR · SIGIR 2002
Information retrieval › distributed information retrieval
resource discovery
0.012001
Distributed Resource Discovery and Structured Data Searching with Cheshire II · SIGIR 2001
Information retrieval › search engines
structured data search
0.012001
Distributed Resource Discovery and Structured Data Searching with Cheshire II · SIGIR 2001
Information retrieval › ranking › search ranking
relevance ranking
0.012008
Geographic IR and visualization in time and space · SIGIR 2008
Information retrieval
cross-modal retrieval
0.011999
Text and Image Retrieval in Cheshire II (demonstration abstract) · SIGIR 1999
Information retrieval › document retrieval
metadata-based retrieval
0.011999
Advanced Search Technologies for Unfamiliar Metadata (demonstration abstract) · SIGIR 1999
Information retrieval
multimedia analysis and retrieval
0.011999
Text and Image Retrieval in Cheshire II (demonstration abstract) · SIGIR 1999
Information retrieval › retrieval models
boolean retrieval
0.011998
Cheshire II: Combining Probabilistic and Boolean Retrieval · SIGIR 1998

Methods — techniques the papers use, named apart from their topics

geospatial visualization · 0.2grid computing · 0.1probabilistic retrieval · 0.1geographic information retrieval · 0.0gazetteer matching · 0.0logistic regression · 0.0z39.50 protocol · 0.0text retrieval · 0.0image retrieval · 0.0demonstration · 0.0
YearPublicationVenuePosition
2014 Integrating Data Mining and Data Management Technologies for Scholarly Inquiry
abstract
This short paper discusses the “Integrating Data Mining and Data Management Technologies for Scholarly Inquiry” project. In this “Round Two” Digging Into Data Challenge award, we explored uses and approaches for large-scale data analysis and processing for the Humanities and Social Sciences through the integration of several infrastructure frameworks: Cheshire, iRODS, and Amazon Web Services (EC2 computing and S3 storage). Our “big data” consisted of the entire texts collection of the Internet Archive (approximately 3.6 million volumes) and the entire JSTOR database. We performed surface-level natural language processing on this data to identify noun phrases and further refinements to identify personal, corporate, and geographic names. We then used resources including library and archival authority records to identify variants and merge names. The goal is to create an integrated index of persons, places, and organizations referenced in our collections.
Ray R. Larson, Richard Marciano, Chien-Yi Hou, Shreyas, Paul B. Watry, John Harrison 0002, Luis Aguilar, Jérôme Fuselier
IEEE BigData1
2012 Applying Digital Library Technologies to Nuclear Forensics
Electra Sutton, Chloe Reynolds, Fredric C. Gey, Ray R. Larson
TPDL4
2011 Connecting Archival Collections: The Social Networks and Archival Context Project
Ray R. Larson, Krishna Janakiraman
TPDL1
2010 Introduction to Information Retrieval
Ray R. Larson
J. Assoc. Inf. Sci. Technol.1
2010 Information Retrieval: Searching in the 21st Century; Human Information Retrieval
Ray R. Larson
J. Assoc. Inf. Sci. Technol.1
2008 Biography as events in time and space
abstract
In digital humanities projects, particularly for historical research and cultural heritage, GIS has played an increasingly important role. However, most implementations have concentrated on displays which ignore the temporal dimension or express it as multiple snapshots for fixed or periodic points in time. Our project concentrates on historical biography and expresses a biography as a sequence of life events with in time and space. We utilize named entity recognition and extraction to automatically mark up biographies so that they can be displayed as dynamic maps. In so doing, contextual features and related happenings and people can be overlaid to facilitate serendipitous discovery of unanticipated and seemingly unrelated connections.
Fredric C. Gey, Ryan Shaw 0001, Ray R. Larson, Barry Pateman
GIS3
2008 Geographic IR and visualization in time and space
abstract
This demonstration will show how graphical geospatial query specifications can be used to obtain sets of georeferenced data ranked by probability of relevance, and displayed geographically and temporally in a geospatial browser with temporal support.
Ray R. Larson
SIGIR1
2008 A comparison of geometric approaches to assessing spatial similarity for GIR
abstract
This research compares the geographic information retrieval (GIR) performance of a set of logistic regression models with those of five non‐probabilistic methods that compute a spatial similarity score for a query–document pair. All methods are applied to a test collection of queries and documents indexed spatially by two convex conservative geometric approximations: the minimum bounding box (MBB) and the convex hull. In the comparison, the tested logistic regression models outperform, in terms of standard information retrieval recall and precision measures, all of the non‐probabilistic methods. The retrieval performance achieved by the logistic regression models on MBB approximations is similar to that achieved by the use of the non‐probabilistic methods on convex hulls. Although these results are valid only for the test collection used in this study, they suggest that a logistic regression approach to GIR provides an alternative to the use of higher‐quality geometric representations that are more difficult to obtain, implement, and process. Additionally, this research demonstrates the ability of a probabilistic approach to effectively incorporate information about geographic context in the spatial ranking process.
Patricia Frontiera, Ray R. Larson, John Radke
Int. J. Geogr. Inf. Sci.2
2006 Cheshire3: retrieving from tera-scale grid-based digital libraries
abstract
No abstract available.
Ray R. Larson, Robert Sanderson
SIGIR1
2005 A Fusion Approach to XML Structured Document Retrieval
Ray R. Larson
Inf. Retr.1
2004 Geotemporal querying of multilingual documents
abstract
This demonstration utilizes a geographic information system interface to display multilingual news documents in time and space by extracting place names from text and matching them to a multilingual multi-script gazetteer which identifies the latitude and longitude of the location.
Fredric C. Gey, Aitao Chen, Ray R. Larson, Kim Carl
SIGIR3
2004 Geographic information retrieval (GIR): searching where and what
abstract
No abstract available.
Ray R. Larson, Patricia Frontiera
SIGIR1
2002 Translingual vocabulary mappings for multilingual information access
abstract
No abstract available.
Fredric C. Gey, Aitao Chen, Michael K. Buckland, Ray R. Larson
SIGIR4
2002 A logistic regression approach to distributed IR
abstract
This poster session examines a probabilistic approach to distributed information retrieval using a Logistic Regression algorithm for estimation of collection relevance. The algorithm is compared to other methods for distributed search using test collections developed for distributed search evaluation.
Ray R. Larson
SIGIR1
2001 Distributed Resource Discovery and Structured Data Searching with Cheshire II
abstract
This demonstration will show describe the construction and application of Cross-Domain Information Servers using features of the standard Z39.50 information retrieval protocol[Z39.50]. The system is currently being used to build and search distributed indexes for databases with disparate structured data (SGML and XML). We use the Z39.50 Explain Database to determine the databases and indexes of a given server, then use the Z39.50 SCAN facility to extract the contents of the indexes. This information is used to build collection documents that can be retrieved using probabilistic retrieval algorithms.
Ray R. Larson
SIGIR1
2001 TREC interactive with Cheshire II
Ray R. Larson
Inf. Process. Manag.1
1999 Text and Image Retrieval in Cheshire II (demonstration abstract)
abstract
No abstract available.
Ray R. Larson
SIGIR1
1999 Advanced Search Technologies for Unfamiliar Metadata (demonstration abstract)
abstract
No abstract available.
Barbara A. Norgard, Youngin Kim 0004, Michael K. Buckland, Aitao Chen, Ray R. Larson, Fred Gaey
SIGIR5
1998 Cheshire II: Combining Probabilistic and Boolean Retrieval
abstract
No abstract available.
Ray R. Larson
SIGIR1
1996 Cheshire II: Designing a Next-Generation Online Catalog
abstract
The Cheshire II online catalog system was designed to provide a bridge between the realms of purely bibliographical information and the rapidly expanding full-text and multi-media collections available online. It is based on a number of national and international standards for data description, communication, and interface technology. The system uses a client-server architecture with X window client communication with an SGML-based probabilistic search engine using the Z39.50 information retrieval protocol. © 1996 John Wiley & Sons, Inc.
Ray R. Larson, Jerome McDonough, Paul O'Leary, Lucy Kuntz, Ralph Moon
J. Am. Soc. Inf. Sci.1
1992 Evaluation of Advanced Retrieval Techniques in an Experimental Online Catalog
abstract
Research on the use and users of online catalogs conducted in the early 1980s found that subject searches were the most common form of online catalog search. At the same time, many of the problems experienced by online catalog users have been traced to difficulties with the subject access mechanisms of the online catalog. Numerous proposals have been made for methods intended to improve subject access to online catalog records. These commonly involve enhancing the catalog's bibliographic records with additional terms, or incorporating subject authority files or additional thesauri in the database. Another stream of research has concentrated on applying retrieval techniques derived from information retrieval (IR) research to replace the Boolean search methods of conventional online catalog systems. This study describes the results of retrieval tests using a variety of these search methods in the CHESHIRE experimental online catalog system. © 1992 John Wiley & Sons, Inc.
Ray R. Larson
J. Am. Soc. Inf. Sci.1
1992 Experiments in Automatic Library of Congress Classification
abstract
This article presents the results of research into the automatic selection of Library of Congress Classification numbers based on the titles and subject headings in MARC records. The method used in this study was based on partial match retrieval techniques using various elements of new records (i.e., those to be classified) as “queries,” and a test database of classification clusters generated from previously classified MARC records. Sixty individual methods for automatic classification were tested on a set of 283 new records, using all combinations of four different partial match methods, five query types, and three representations of search terms. The results indicate that if the best method for a particular case can be determined, then up to 86% of the new records may be correctly classified. The single method with the best accuracy was able to select the correct classification for about 46% of the new records. © 1992 John Wiley & Sons, Inc.
Ray R. Larson
J. Am. Soc. Inf. Sci.1
1991 The decline of subject searching: Long-term trends and patterns of index use in an online catalog
abstract
Search index usage in a large university online catalog system over a six-year period (representing about 15.3 million searches) was investigated using transaction monitor data. Mathematical models of trends and patterns in the data were developed and tested using regression techniques. The results of the analyses show a consistent decline in the frequency of subject index use by online catalog users, with a corresponding increase in the frequency of title keyword searching. Significant annual patterns in index usage were also identified. Analysis of the transaction data, and related previous studies of online catalog users, suggest a number of factors contributing to the decline in subject search frequency. Chief among these factors are user difficulties in formulating subject queries with Library of Congress Subject Headings, leading to search failure, and the problem of “information overload” as database size increases. This article presents the models and results of the transaction log analysis, discusses the underlying problems with subject searching contributing to the observed decline, and reviews some proposed improvements to online catalog systems to aid in overcoming these problems. © 1991 John Wiley & Sons, Inc.
Ray R. Larson
J. Am. Soc. Inf. Sci.1
1990 Hypertext hands-on!: An introduction to a new way of organizing and accessing information
Ray R. Larson
J. Am. Soc. Inf. Sci.1