EDBT 2026 Demo / reviewers in the wild / expert
Michael L. Nelson 0001
dblp:n/MichaelLNelson
· DBLP profile ↗
36ranked-venue papers in the field
0as first author
11since 2021 · last 2025
0000-0003-3749-8116ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 34Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Problems with archiving and replaying current web advertisementsabstractAbstract Advertisements have always been a part of our cultural heritage, and this also applies to online web advertisements. Unlike print ads, there are serious technical challenges involved in archiving and successfully replaying ads displayed in web pages. To explore these challenges, we created a small dataset of 250 web ads. Ultimately, we collected and archived 279 ads, which we classified into 5 different categories: combination ads, image ads, embedded web page ads, video ads, and text‐only ads. In our sample, combination ads were the most prevalent and, not surprisingly, text‐only ads were the easiest to archive and replay. During the course of this study, we encountered five major problems in archiving and replaying the web ads. We detail the issues uncovered and provide suggestions for ameliorating the replay of some web ads, in addition to other dynamically loaded embedded web resources. Travis Reid, Alex H. Poole, Hyung Wook Choi, Christopher B. Rauch, Mat Kelly, Michael L. Nelson 0001, Michele C. Weigle |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2024 | Retrogressive Document Manipulation of US Federal Environmental WebsitesabstractChanges made to webpages can affect their retrievability. Often this is done with the intention of increasing the page's search engine ranking to improve overall access to information on the page. The Environmental Data and Governance Initiative (EDGI) created a dataset that describes changes on US federal environmental webpages between 2016 and 2020. EDGI noted that many environmental terms were deleted from the pages, but without user data, claims that page retrievability and public information access were lowered are only anecdotal. The Open Resource for Click Analysis in Search (ORCAS) dataset was created during the same time frame, from 2017 to 2020, and enables high quality user intent analysis without compromising on user privacy protection. We present an analysis of the intersection of the EDGI dataset and the ORCAS dataset, matching changes on federal environmental webpages with their associated queries. We use web archives and a change-text indexing system to link changes in term frequency on the pages with the queries. We find that the pages contain fewer query terms in 2020 than in 2016, lowering the pages' retrievability. The analysis provides substantive support of EDGI's claim that federal environmental pages were made less accessible between 2016 and 2020. Lesley Frew, Michael L. Nelson 0001, Michele C. Weigle |
CIKM | 2 |
| 2024 | Summarizing Web Archive Corpora via Social Media Storytelling by Automatically Selecting and Visualizing ExemplarsabstractPeople often create themed collections to make sense of an ever-increasing number of archived web pages. Some of these collections contain hundreds of thousands of documents. Thousands of collections exist, many covering the same topic. Few collections include standardized metadata. This scale makes understanding a collection an expensive proposition. Our Dark and Stormy Archives (DSA) five-process model implements a novel summarization method to help users understand a collection by combining web archives and social media storytelling. The five processes of the DSA model are: select exemplars, generate story metadata, generate document metadata, visualize the story, and distribute the story. Selecting exemplars produces a set of k documents from the N documents in the collection, where k < < N , thus reducing the number of documents visitors need to review to understand a collection. Generating story and document metadata selects images, titles, descriptions, and other content from these exemplars. Visualizing the story ties this metadata together in a format the visitor can consume. Without distributing the story, it is not shared for others to consume. We present a research study demonstrating that our algorithmic primitives can be combined to select relevant exemplars that are otherwise undiscoverable using a conventional search engine and query generation methods. Having demonstrated improved methods for selecting exemplars, we visualize the story. Previous work established that the social card is the best format for visitors to consume surrogates. The social card combines metadata fields, including the document’s title, a brief description, and a striking image. Social cards are commonly found on social media platforms. We discovered that these platforms perform poorly for mementos and rely on web page authors to supply the necessary values for these metadata fields. With web archives, we often encounter archived web pages that predate the existence of this metadata. To generate this missing metadata and ensure that storytelling is available for these documents, we apply machine learning to generate the images needed for social cards with a Precision@1 of 0.8314. We also provide the length values needed for executing automatic summarization algorithms to generate document descriptions. Applying these concepts helps us create the visualizations needed to fulfill the final processes of story generation. We close this work with examples and applications of this technology. Shawn M. Jones, Martin Klein 0001, Michele C. Weigle, Michael L. Nelson 0001 |
ACM Trans. Web | 4 |
| 2023 | It's Not Just GitHub: Identifying Data and Software Sources Included in Publications
Emily Escamilla, Lamia Salsabil, Martin Klein 0001, Jian Wu 0006, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 6 |
| 2023 | Synthesizing Web Archive Collections into Big Data: Lessons from Mining Data from Web Archives
Shawn M. Jones, Himarsha R. Jayanetti, Martin Klein 0001, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 5 |
| 2023 | To Re-experience the Web: A Framework for the Transformation and Replay of Archived Web PagesabstractWhen replaying an archived web page, or memento , the fundamental expectation is that the page should be viewable and function exactly as it did at the archival time. However, this expectation requires web archives upon replay to modify the page and its embedded resources so that all resources and links reference the archive rather than the original server. Although these modifications necessarily change the state of the representation, it is understood that without them the replay of mementos from the archive would not be possible. The process of replaying mementos and the modifications made to the representations by web archives varies between archives. Because of this, there is no standard terminology for describing the replay and needed modifications. In this article, we propose terminology for describing the existing styles of replay and the modifications made on the part of web archives to mementos to facilitate replay. Because of issues discovered with server-side only modifications, we propose a general framework for the auto-generation of client-side rewriting libraries. Finally, we evaluate the effectiveness of using a generated client-side rewriting library to augment the existing replay systems of web archives by crawling mementos replayed from the Internet Archive’s Wayback Machine with and without the generated client-side rewriter. By using the generated client-side rewriter, we were able to decrease the cumulative number of requests blocked by the content security policy of the Wayback Machine for 577 mementos by 87.5% and increased the cumulative number of requests made by 32.8%. We were also able to replay mementos that were previously not replayable from the Internet Archive. Many of the client-side rewriting ideas described in this work have been implemented into Wombat, a client-side URL rewriting system that is used by the Webrecorder, Pywb, and Wayback Machine playback systems. John A. Berlin, Mat Kelly, Michael L. Nelson 0001, Michele C. Weigle |
ACM Trans. Web | 3 |
| 2022 | The Rise of GitHub in Scholarly Publications
Emily Escamilla, Martin Klein 0001, Talya Cooper, Vicky Rampin, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 6 |
| 2022 | Robots Still Outnumber Humans in Web Archives, But Less Than Before
Himarsha R. Jayanetti, Kritika Garg, Sawood Alam, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 4 |
| 2022 | Creating Structure in Web Archives with Collections: Different Concepts from Web Archivists
Himarsha R. Jayanetti, Shawn M. Jones, Martin Klein 0001, Alex Osbourne, Paul Koerbin, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 6 |
| 2022 | A Chromium-Based Memento-Aware Web Browser
Abigail Mabe, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 2 |
| 2021 | Where Did the Web Archive Go?
Mohamed Aturban, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 2 |
| 2020 | Modeling Updates of Scholarly Webpages Using Archived DataabstractThe vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose an approach for modeling the dynamics of change in the web using archived copies of webpages. To evaluate its utility, we conduct a preliminary study on the scholarly web using 19,977 seed URLs of authors’ homepages obtained from their Google Scholar profiles. We first obtain archived copies of these webpages from the Internet Archive (IA), and estimate when their actual updates occurred. Next, we apply maximum likelihood to estimate their mean update frequency (λ) values. Our evaluation shows that λ values derived from a short history of archived data provide a good estimate for the true update frequency in the short-term, and that our method provides better estimations of updates at a fraction of resources compared to the baseline models. Based on this, we demonstrate the utility of archived data to optimize the crawling strategy of web crawlers, and uncover important challenges that inspire future research directions. Yasith Jayawardana, Alexander C. Nwala, Gavindya Jayawardena, Jian Wu 0006, Sampath Jayarathna, Michael L. Nelson 0001, C. Lee Giles |
IEEE BigData | 6 |
| 2019 | Social Cards Probably Provide For Better Understanding Of Web Archive CollectionsabstractUsed by a variety of researchers, web archive collections have become invaluable sources of evidence. If a researcher is presented with a web archive collection that they did not create, how do they know what is inside so that they can use it for their own research? Search engine results and social media links are represented as surrogates, small easily digestible summaries of the underlying page. Search engines and social media have a different focus, and hence produce different surrogates than web archives. Search engine surrogates help a user answer the question "Will this link meet my information need?" Social media surrogates help a user decide "Should I click on this?" Our use case is subtly different. We hypothesize that groups of surrogates together are useful for summarizing a collection. We want to help users answer the question of "What does the underlying collection contain?" But which surrogate should we use? With Mechanical Turk participants, we evaluate six different surrogate types against each other. We find that the type of surrogate does not influence the time to complete the task we presented the participants. Of particular interest are social cards, surrogates typically found on social media, and browser thumbnails, screen captures of web pages rendered in a browser. At p=0.0569, and p=0.0770, respectively, we find that social cards and social cards paired side-by-side with browser thumbnails probably provide better collection understanding than the surrogates currently used by the popular Archive-It web archiving platform. We measure user interactions with each surrogate and find that users interact with social cards less than other types. The results of this study have implications for our web archive summarization work, live web curation platforms, social media, and more. Shawn M. Jones, Michele C. Weigle, Michael L. Nelson 0001 |
CIKM | 3 |
| 2017 | Comparing the Archival Rate of Arabic, English, Danish, and Korean Language Web PagesabstractIt has long been suspected that web archives and search engines favor Western and English language webpages. In this article, we quantitatively explore how well indexed and archived Arabic language webpages are as compared to those from other languages. We began by sampling 15,092 unique URIs from three different website directories: DMOZ (multilingual), Raddadi, and Star28 (the last two primarily Arabic language). Using language identification tools, we eliminated pages not in the Arabic language (e.g., English-language versions of Aljazeera pages) and culled the collection to 7,976 Arabic language webpages. We then used these 7,976 pages and crawled the live web and web archives to produce a collection of 300,646 Arabic language pages. We compared the analysis of Arabic language pages with that of English, Danish, and Korean language pages. First, for each language, we sampled unique URIs from DMOZ; then, using language identification tools, we kept only pages in the desired language. Finally, we crawled the archived and live web to collect a larger sample of pages in English, Danish, or Korean. In total for the four languages, we analyzed over 500,000 webpages. We discovered: (1) English has a higher archiving rate than Arabic, with 72.04% archived. However, Arabic has a higher archiving rate than Danish and Korean, with 53.36% of Arabic URIs archived, followed by Danish and Korean with 35.89% and 32.81% archived, respectively. (2) Most Arabic and English language pages are located in the United States; only 14.84% of the Arabic URIs had an Arabic country code top-level domain (e.g., sa) and only 10.53% had a GeoIP in an Arabic country. Most Danish-language pages were located in Denmark, and most Korean-language pages were located in South Korea. (3) The presence of a webpage in a directory positively impacts indexing and presence in the DMOZ directory, specifically, positively impacts archiving in all four languages. In this work, we show that web archives and search engines favor English pages. However, it is not universally true for all Western-language webpages because, in this work, we show that Arabic webpages have a higher archival rate than Danish language webpages. Lulwah M. Alkwai, Michael L. Nelson 0001, Michele C. Weigle |
ACM Trans. Inf. Syst. | 2 |
| 2016 | Web Archive Profiling Through Fulltext Search
Sawood Alam, Michael L. Nelson 0001, Herbert Van de Sompel, David S. H. Rosenthal |
TPDL | 2 |
| 2016 | InterPlanetary Wayback: Peer-To-Peer Permanence of Web Archives
Mat Kelly, Sawood Alam, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 3 |
| 2015 | Detecting Off-Topic Pages in Web Archives
Yasmin AlNoamany, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 3 |
| 2015 | Characteristics of Social Media Stories
Yasmin AlNoamany, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 3 |
| 2015 | Web Archive Profiling Through CDX Summarization
Sawood Alam, Michael L. Nelson 0001, Herbert Van de Sompel, Lyudmila Balakireva, Harihar Shankar, David S. H. Rosenthal |
TPDL | 2 |
| 2015 | Quantifying Orphaned Annotations in Hypothes.is
Mohamed Aturban, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 2 |
| 2014 | Thumbnail Summarization Techniques for Web Archives
Ahmed Alsum, Michael L. Nelson 0001 |
ECIR | 2 |
| 2013 | Who and What Links to the Internet Archive
Yasmin AlNoamany, Ahmed Alsum, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 4 |
| 2013 | Profiling Web Archive Coverage for Top-Level Domain and Content Language
Ahmed Alsum, Michele C. Weigle, Michael L. Nelson 0001, Herbert Van de Sompel |
TPDL | 3 |
| 2013 | Evaluating the SiteStory Transactional Web Archive with the ApacheBench Tool
Justin F. Brunelle, Michael L. Nelson 0001, Lyudmila Balakireva, Robert Sanderson, Herbert Van de Sompel |
TPDL | 2 |
| 2013 | On the Change in Archivability of Websites Over Time
Mat Kelly, Justin F. Brunelle, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 4 |
| 2013 | Resurrecting My Revolution - Using Social Link Neighborhood in Bringing Context to the Disappearing Web
Hany SalahEldeen, Michael L. Nelson 0001 |
TPDL | 2 |
| 2013 | ResourceSync: The NISO/OAI Resource Synchronization Framework
Herbert Van de Sompel, Michael L. Nelson 0001, Martin Klein 0001, Robert Sanderson |
TPDL | 2 |
| 2012 | Losing My Revolution: How Many Resources Shared on Social Media Have Been Lost?
Hany SalahEldeen, Michael L. Nelson 0001 |
TPDL | 2 |
| 2011 | Find, New, Copy, Web, Page - Tagging for the (Re-)Discovery of Web Pages
Martin Klein 0001, Michael L. Nelson 0001 |
TPDL | 2 |
| 2011 | Music Video Redundancy and Half-Life in YouTube
Matthias Prellwitz, Michael L. Nelson 0001 |
TPDL | 2 |
| 2009 | Correlation of Term Count and Document Frequency for Google N-Grams
Martin Klein 0001, Michael L. Nelson 0001 |
ECIR | 2 |
| 2007 | Search engines and their public interfaces: which apis are the most synchronized?abstractResearchers of commercial search engines often collect datausing the application programming interface (API) or by"scraping" results from the web user interface (WUI), butanecdotal evidence suggests the interfaces produce differentresults. We provide the first in-depth quantitative analysisof the results produced by the Google, MSN and Yahoo APIand WUI interfaces. After submitting a variety of queriesto the interfaces for 5 months, we found significant discrepanciesin several categories. Our findings suggest that theAPI indexes are not older, but they are probably smaller for Google and Yahoo. Researchers may use our findings tobetter understand the differences between the interfaces andchoose the best API for their particular types of queries. Frank McCown, Michael L. Nelson 0001 |
WWW | 2 |
| 2005 | Co-authorship networks in the digital library research community
Xiaoming Liu 0005, Johan Bollen, Michael L. Nelson 0001, Herbert Van de Sompel |
Inf. Process. Manag. | 3 |
| 2002 | Archon - A Digital Library that Federates Physics Collections
Kurt Maly, Mohammad Zubair, Michael L. Nelson 0001, Xiaoming Liu 0005, Hesham Anan, Jinsong Gao, Jianfeng Tang |
Dublin Core Conference | 3 |
| 2000 | Determining the publication impact of a digital libraryabstractWe attempt to assess the publication impact of a digital library (DL) of aerospace scientific and technical information (STI). The Langley Technical Report Server (LTRS) is a digital library of over 1,400 electronic publications authored by NASA Langley Research Center personnel or contractors and has been available in its current World Wide Web (WWW) form since 1994. In this article, we examine calendar year 1997 usage statistics of LTRS and the Center for AeroSpace Information (CASI), a facility that archives and distributes hard copies of NASA and aerospace information. We also perform a citation analysis on some of the top publications distributed by LTRS. We find that although LTRS distributes over 71,000 copies of publications (compared with an estimated 24,000 copies from CASI), citation analysis indicates that LTRS has almost no measurable publication impact. We discuss the caveats of our investigation, speculate on possible different models of usage facilitated by DLs, and suggest “retrieval analysis” as a complementary metric to citation analysis. While our investigation failes to establish a relationship between LTRS and increased citations and raises at least as many questions as it answers, we hope it will serve as an invitation to, and guide for, further research in the use of DLs. Nancy R. Kaplan, Michael L. Nelson 0001 |
J. Am. Soc. Inf. Sci. | 2 |
| 1998 | Evolution of Scientific and Technical Information DistributionabstractWorld Wide Web (WWW) and related information technologies are transforming the distribution of scientific and technical information (STI). We examine 11 recent, functioning digital libraries focusing on the distribution of STI publications, including journal articles, conference papers, and technical reports. We introduce 4 main categories of digital library projects: Based on the architecture (distributed vs. centralized) and the contributor (traditional publisher vs. authoring individual/organization). Many digital library prototypes merely automate existing publishing practices or focus solely on the digitization of the publishing cycle output, not sampling and capturing elements of the input. Still others do not consider for distribution the large body of “gray literature.” We address these deficiencies in the current model of STI exchange by suggesting methods for expanding the scope and target of digital libraries by focusing on a greater source of technical publications and using “buckets,” an object-oriented construct for grouping logically related information objects, to include holdings other than technical publications. © 1998 John Wiley & Sons, Inc. Sandra L. Esler, Michael L. Nelson 0001 |
J. Am. Soc. Inf. Sci. | 2 |