Martin Klein 0001

dblp:67/1254-1 · DBLP profile ↗
← Back
11ranked-venue papers in the field
5as first author
6since 2021 · last 2024
0000-0003-0130-2097ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 11 (5 first)
YearPublicationVenuePosition
2024 Summarizing Web Archive Corpora via Social Media Storytelling by Automatically Selecting and Visualizing Exemplars
abstract
People often create themed collections to make sense of an ever-increasing number of archived web pages. Some of these collections contain hundreds of thousands of documents. Thousands of collections exist, many covering the same topic. Few collections include standardized metadata. This scale makes understanding a collection an expensive proposition. Our Dark and Stormy Archives (DSA) five-process model implements a novel summarization method to help users understand a collection by combining web archives and social media storytelling. The five processes of the DSA model are: select exemplars, generate story metadata, generate document metadata, visualize the story, and distribute the story. Selecting exemplars produces a set of k documents from the N documents in the collection, where k < < N , thus reducing the number of documents visitors need to review to understand a collection. Generating story and document metadata selects images, titles, descriptions, and other content from these exemplars. Visualizing the story ties this metadata together in a format the visitor can consume. Without distributing the story, it is not shared for others to consume. We present a research study demonstrating that our algorithmic primitives can be combined to select relevant exemplars that are otherwise undiscoverable using a conventional search engine and query generation methods. Having demonstrated improved methods for selecting exemplars, we visualize the story. Previous work established that the social card is the best format for visitors to consume surrogates. The social card combines metadata fields, including the document’s title, a brief description, and a striking image. Social cards are commonly found on social media platforms. We discovered that these platforms perform poorly for mementos and rely on web page authors to supply the necessary values for these metadata fields. With web archives, we often encounter archived web pages that predate the existence of this metadata. To generate this missing metadata and ensure that storytelling is available for these documents, we apply machine learning to generate the images needed for social cards with a Precision@1 of 0.8314. We also provide the length values needed for executing automatic summarization algorithms to generate document descriptions. Applying these concepts helps us create the visualizations needed to fulfill the final processes of story generation. We close this work with examples and applications of this technology.
Shawn M. Jones, Martin Klein 0001, Michele C. Weigle, Michael L. Nelson 0001
ACM Trans. Web2
2023 It's Not Just GitHub: Identifying Data and Software Sources Included in Publications
Emily Escamilla, Lamia Salsabil, Martin Klein 0001, Jian Wu 0006, Michele C. Weigle, Michael L. Nelson 0001
TPDL3
2023 Synthesizing Web Archive Collections into Big Data: Lessons from Mining Data from Web Archives
Shawn M. Jones, Himarsha R. Jayanetti, Martin Klein 0001, Michele C. Weigle, Michael L. Nelson 0001
TPDL3
2022 The Rise of GitHub in Scholarly Publications
Emily Escamilla, Martin Klein 0001, Talya Cooper, Vicky Rampin, Michele C. Weigle, Michael L. Nelson 0001
TPDL2
2022 Creating Structure in Web Archives with Collections: Different Concepts from Web Archivists
Himarsha R. Jayanetti, Shawn M. Jones, Martin Klein 0001, Alex Osbourne, Paul Koerbin, Michael L. Nelson 0001, Michele C. Weigle
TPDL3
2022 Got 404s? Crawling and Analyzing an Institution's Web Domain
Martin Klein 0001, Lyudmila Balakireva
TPDL1
2020 On the Persistence of Persistent Identifiers of the Scholarly Web
Martin Klein 0001, Lyudmila Balakireva
TPDL1
2019 The Memento Tracer Framework: Balancing Quality and Scalability for Web Archiving
Martin Klein 0001, Harihar Shankar, Lyudmila Balakireva, Herbert Van de Sompel
TPDL1
2013 ResourceSync: The NISO/OAI Resource Synchronization Framework
Herbert Van de Sompel, Michael L. Nelson 0001, Martin Klein 0001, Robert Sanderson
TPDL3
2011 Find, New, Copy, Web, Page - Tagging for the (Re-)Discovery of Web Pages
Martin Klein 0001, Michael L. Nelson 0001
TPDL1
2009 Correlation of Term Count and Document Frequency for Google N-Grams
Martin Klein 0001, Michael L. Nelson 0001
ECIR1