VLDB 2026 Research / reviewers in the wild / expert
Ray R. Larson
dblp:68/4900
· DBLP profile ↗
24ranked-venue papers
18as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 24 · 18 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
10 papers |
Information retrieval · 88% Data mining · 12% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › document retrieval › domain-specific retrieval
geographic information retrieval |
0.1 | 2 | 2008 | Geographic IR and visualization in time and space · SIGIR 2008 Geographic information retrieval (GIR): searching where and what · SIGIR 2004 |
Information retrieval
cross-language information retrieval |
0.1 | 2 | 2004 | Geotemporal querying of multilingual documents · SIGIR 2004 Translingual vocabulary mappings for multilingual information access · SIGIR 2002 |
Visualization and visual analytics
geospatial visualization |
0.1 | 1 | 2008 | Geographic IR and visualization in time and space · SIGIR 2008 |
Information retrieval
distributed information retrieval |
0.1 | 2 | 2002 | A logistic regression approach to distributed IR · SIGIR 2002 Distributed Resource Discovery and Structured Data Searching with Cheshire II · SIGIR 2001 |
Information retrieval › retrieval models
probabilistic retrieval model |
0.1 | 4 | 2008 | Geographic IR and visualization in time and space · SIGIR 2008 Cheshire II: Combining Probabilistic and Boolean Retrieval · SIGIR 1998 A logistic regression approach to distributed IR · SIGIR 2002 |
Data mining › text mining
information extraction |
0.0 | 1 | 2004 | Geotemporal querying of multilingual documents · SIGIR 2004 |
Information retrieval
retrieval models |
0.0 | 3 | 2006 | Cheshire II: Combining Probabilistic and Boolean Retrieval · SIGIR 1998 Cheshire3: retrieving from tera-scale grid-based digital libraries · SIGIR 2006 Text and Image Retrieval in Cheshire II (demonstration abstract) · SIGIR 1999 |
Data mining › predictive modeling › regression
logistic regression |
0.0 | 1 | 2002 | A logistic regression approach to distributed IR · SIGIR 2002 |
Information retrieval › distributed information retrieval
resource discovery |
0.0 | 1 | 2001 | Distributed Resource Discovery and Structured Data Searching with Cheshire II · SIGIR 2001 |
Information retrieval › search engines
structured data search |
0.0 | 1 | 2001 | Distributed Resource Discovery and Structured Data Searching with Cheshire II · SIGIR 2001 |
Information retrieval › ranking › search ranking
relevance ranking |
0.0 | 1 | 2008 | Geographic IR and visualization in time and space · SIGIR 2008 |
Information retrieval
cross-modal retrieval |
0.0 | 1 | 1999 | Text and Image Retrieval in Cheshire II (demonstration abstract) · SIGIR 1999 |
Information retrieval › document retrieval
metadata-based retrieval |
0.0 | 1 | 1999 | Advanced Search Technologies for Unfamiliar Metadata (demonstration abstract) · SIGIR 1999 |
Information retrieval
multimedia analysis and retrieval |
0.0 | 1 | 1999 | Text and Image Retrieval in Cheshire II (demonstration abstract) · SIGIR 1999 |
Information retrieval › retrieval models
boolean retrieval |
0.0 | 1 | 1998 | Cheshire II: Combining Probabilistic and Boolean Retrieval · SIGIR 1998 |
Methods — techniques the papers use, named apart from their topics
geospatial visualization · 0.2grid computing · 0.1probabilistic retrieval · 0.1geographic information retrieval · 0.0gazetteer matching · 0.0logistic regression · 0.0z39.50 protocol · 0.0text retrieval · 0.0image retrieval · 0.0demonstration · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Integrating Data Mining and Data Management Technologies for Scholarly InquiryabstractThis short paper discusses the “Integrating Data Mining and Data Management Technologies for Scholarly Inquiry” project. In this “Round Two” Digging Into Data Challenge award, we explored uses and approaches for large-scale data analysis and processing for the Humanities and Social Sciences through the integration of several infrastructure frameworks: Cheshire, iRODS, and Amazon Web Services (EC2 computing and S3 storage). Our “big data” consisted of the entire texts collection of the Internet Archive (approximately 3.6 million volumes) and the entire JSTOR database. We performed surface-level natural language processing on this data to identify noun phrases and further refinements to identify personal, corporate, and geographic names. We then used resources including library and archival authority records to identify variants and merge names. The goal is to create an integrated index of persons, places, and organizations referenced in our collections. Ray R. Larson, Richard Marciano, Chien-Yi Hou, Shreyas, Paul B. Watry, John Harrison 0002, Luis Aguilar, Jérôme Fuselier |
IEEE BigData | 1 |
| 2012 | Applying Digital Library Technologies to Nuclear Forensics
Electra Sutton, Chloe Reynolds, Fredric C. Gey, Ray R. Larson |
TPDL | 4 |
| 2011 | Connecting Archival Collections: The Social Networks and Archival Context Project
Ray R. Larson, Krishna Janakiraman |
TPDL | 1 |
| 2010 | Introduction to Information Retrieval
Ray R. Larson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Information Retrieval: Searching in the 21st Century; Human Information Retrieval
Ray R. Larson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Biography as events in time and spaceabstractIn digital humanities projects, particularly for historical research and cultural heritage, GIS has played an increasingly important role. However, most implementations have concentrated on displays which ignore the temporal dimension or express it as multiple snapshots for fixed or periodic points in time. Our project concentrates on historical biography and expresses a biography as a sequence of life events with in time and space. We utilize named entity recognition and extraction to automatically mark up biographies so that they can be displayed as dynamic maps. In so doing, contextual features and related happenings and people can be overlaid to facilitate serendipitous discovery of unanticipated and seemingly unrelated connections. Fredric C. Gey, Ryan Shaw 0001, Ray R. Larson, Barry Pateman |
GIS | 3 |
| 2008 | Geographic IR and visualization in time and spaceabstractThis demonstration will show how graphical geospatial query specifications can be used to obtain sets of georeferenced data ranked by probability of relevance, and displayed geographically and temporally in a geospatial browser with temporal support. Ray R. Larson |
SIGIR | 1 |
| 2008 | A comparison of geometric approaches to assessing spatial similarity for GIRabstractThis research compares the geographic information retrieval (GIR) performance of a set of logistic regression models with those of five non‐probabilistic methods that compute a spatial similarity score for a query–document pair. All methods are applied to a test collection of queries and documents indexed spatially by two convex conservative geometric approximations: the minimum bounding box (MBB) and the convex hull. In the comparison, the tested logistic regression models outperform, in terms of standard information retrieval recall and precision measures, all of the non‐probabilistic methods. The retrieval performance achieved by the logistic regression models on MBB approximations is similar to that achieved by the use of the non‐probabilistic methods on convex hulls. Although these results are valid only for the test collection used in this study, they suggest that a logistic regression approach to GIR provides an alternative to the use of higher‐quality geometric representations that are more difficult to obtain, implement, and process. Additionally, this research demonstrates the ability of a probabilistic approach to effectively incorporate information about geographic context in the spatial ranking process. Patricia Frontiera, Ray R. Larson, John Radke |
Int. J. Geogr. Inf. Sci. | 2 |
| 2006 | Cheshire3: retrieving from tera-scale grid-based digital librariesabstractNo abstract available. Ray R. Larson, Robert Sanderson |
SIGIR | 1 |
| 2005 | A Fusion Approach to XML Structured Document Retrieval
Ray R. Larson |
Inf. Retr. | 1 |
| 2004 | Geotemporal querying of multilingual documentsabstractThis demonstration utilizes a geographic information system interface to display multilingual news documents in time and space by extracting place names from text and matching them to a multilingual multi-script gazetteer which identifies the latitude and longitude of the location. Fredric C. Gey, Aitao Chen, Ray R. Larson, Kim Carl |
SIGIR | 3 |
| 2004 | Geographic information retrieval (GIR): searching where and whatabstractNo abstract available. Ray R. Larson, Patricia Frontiera |
SIGIR | 1 |
| 2002 | Translingual vocabulary mappings for multilingual information accessabstractNo abstract available. Fredric C. Gey, Aitao Chen, Michael K. Buckland, Ray R. Larson |
SIGIR | 4 |
| 2002 | A logistic regression approach to distributed IRabstractThis poster session examines a probabilistic approach to distributed information retrieval using a Logistic Regression algorithm for estimation of collection relevance. The algorithm is compared to other methods for distributed search using test collections developed for distributed search evaluation. Ray R. Larson |
SIGIR | 1 |
| 2001 | Distributed Resource Discovery and Structured Data Searching with Cheshire IIabstractThis demonstration will show describe the construction and application of Cross-Domain Information Servers using features of the standard Z39.50 information retrieval protocol[Z39.50]. The system is currently being used to build and search distributed indexes for databases with disparate structured data (SGML and XML). We use the Z39.50 Explain Database to determine the databases and indexes of a given server, then use the Z39.50 SCAN facility to extract the contents of the indexes. This information is used to build collection documents that can be retrieved using probabilistic retrieval algorithms. Ray R. Larson |
SIGIR | 1 |
| 2001 | TREC interactive with Cheshire II
Ray R. Larson |
Inf. Process. Manag. | 1 |
| 1999 | Text and Image Retrieval in Cheshire II (demonstration abstract)abstractNo abstract available. Ray R. Larson |
SIGIR | 1 |
| 1999 | Advanced Search Technologies for Unfamiliar Metadata (demonstration abstract)abstractNo abstract available. Barbara A. Norgard, Youngin Kim 0004, Michael K. Buckland, Aitao Chen, Ray R. Larson, Fred Gaey |
SIGIR | 5 |
| 1998 | Cheshire II: Combining Probabilistic and Boolean RetrievalabstractNo abstract available. Ray R. Larson |
SIGIR | 1 |
| 1996 | Cheshire II: Designing a Next-Generation Online CatalogabstractThe Cheshire II online catalog system was designed to provide a bridge between the realms of purely bibliographical information and the rapidly expanding full-text and multi-media collections available online. It is based on a number of national and international standards for data description, communication, and interface technology. The system uses a client-server architecture with X window client communication with an SGML-based probabilistic search engine using the Z39.50 information retrieval protocol. © 1996 John Wiley & Sons, Inc. Ray R. Larson, Jerome McDonough, Paul O'Leary, Lucy Kuntz, Ralph Moon |
J. Am. Soc. Inf. Sci. | 1 |
| 1992 | Evaluation of Advanced Retrieval Techniques in an Experimental Online CatalogabstractResearch on the use and users of online catalogs conducted in the early 1980s found that subject searches were the most common form of online catalog search. At the same time, many of the problems experienced by online catalog users have been traced to difficulties with the subject access mechanisms of the online catalog. Numerous proposals have been made for methods intended to improve subject access to online catalog records. These commonly involve enhancing the catalog's bibliographic records with additional terms, or incorporating subject authority files or additional thesauri in the database. Another stream of research has concentrated on applying retrieval techniques derived from information retrieval (IR) research to replace the Boolean search methods of conventional online catalog systems. This study describes the results of retrieval tests using a variety of these search methods in the CHESHIRE experimental online catalog system. © 1992 John Wiley & Sons, Inc. Ray R. Larson |
J. Am. Soc. Inf. Sci. | 1 |
| 1992 | Experiments in Automatic Library of Congress ClassificationabstractThis article presents the results of research into the automatic selection of Library of Congress Classification numbers based on the titles and subject headings in MARC records. The method used in this study was based on partial match retrieval techniques using various elements of new records (i.e., those to be classified) as “queries,” and a test database of classification clusters generated from previously classified MARC records. Sixty individual methods for automatic classification were tested on a set of 283 new records, using all combinations of four different partial match methods, five query types, and three representations of search terms. The results indicate that if the best method for a particular case can be determined, then up to 86% of the new records may be correctly classified. The single method with the best accuracy was able to select the correct classification for about 46% of the new records. © 1992 John Wiley & Sons, Inc. Ray R. Larson |
J. Am. Soc. Inf. Sci. | 1 |
| 1991 | The decline of subject searching: Long-term trends and patterns of index use in an online catalogabstractSearch index usage in a large university online catalog system over a six-year period (representing about 15.3 million searches) was investigated using transaction monitor data. Mathematical models of trends and patterns in the data were developed and tested using regression techniques. The results of the analyses show a consistent decline in the frequency of subject index use by online catalog users, with a corresponding increase in the frequency of title keyword searching. Significant annual patterns in index usage were also identified. Analysis of the transaction data, and related previous studies of online catalog users, suggest a number of factors contributing to the decline in subject search frequency. Chief among these factors are user difficulties in formulating subject queries with Library of Congress Subject Headings, leading to search failure, and the problem of “information overload” as database size increases. This article presents the models and results of the transaction log analysis, discusses the underlying problems with subject searching contributing to the observed decline, and reviews some proposed improvements to online catalog systems to aid in overcoming these problems. © 1991 John Wiley & Sons, Inc. Ray R. Larson |
J. Am. Soc. Inf. Sci. | 1 |
| 1990 | Hypertext hands-on!: An introduction to a new way of organizing and accessing information
Ray R. Larson |
J. Am. Soc. Inf. Sci. | 1 |