EDBT 2026 Demo / reviewers in the wild / expert
Michele C. Weigle
dblp:38/1669
· DBLP profile ↗
21ranked-venue papers in the field
0as first author
12since 2021 · last 2025
0000-0002-2787-7166ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 20Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Problems with archiving and replaying current web advertisementsabstractAbstract Advertisements have always been a part of our cultural heritage, and this also applies to online web advertisements. Unlike print ads, there are serious technical challenges involved in archiving and successfully replaying ads displayed in web pages. To explore these challenges, we created a small dataset of 250 web ads. Ultimately, we collected and archived 279 ads, which we classified into 5 different categories: combination ads, image ads, embedded web page ads, video ads, and text‐only ads. In our sample, combination ads were the most prevalent and, not surprisingly, text‐only ads were the easiest to archive and replay. During the course of this study, we encountered five major problems in archiving and replaying the web ads. We detail the issues uncovered and provide suggestions for ameliorating the replay of some web ads, in addition to other dynamically loaded embedded web resources. Travis Reid, Alex H. Poole, Hyung Wook Choi, Christopher B. Rauch, Mat Kelly, Michael L. Nelson 0001, Michele C. Weigle |
J. Assoc. Inf. Sci. Technol. | 7 |
| 2024 | Exploring Large Language Models for Analyzing Changes in Web Archive Content: A Retrieval-Augmented Generation ApproachabstractWebsites typically display only their most recent content. However, the dynamic nature of the web leads to frequent updates and deletions. Web archives preserve snapshots of earlier versions for those interested in tracking changes over time. Analyzing these changes often requires a manual process that relies on traditional methods focused on terms or phrase-level differences. This study explores the capability of Large Language Models (LLMs), specifically GPT-4o, through a Retrieval-Augmented Generation (RAG) approach for detecting changes in archived web pages. Using WARC-GPT, a RAG pipeline to interact with Web ARChive (WARC) files, we identify and analyze changes across a small set of U.S. federal environmental web pages that changed between 2016 and 2020. Our findings show that GPT-4o can effectively be used to detect inconsistencies in web archive content, including consideration of the change and the semantic context upon which the changes occurred. Our exploration represents an initial step toward using Artificial Intelligence (AI) for deeper and scalable web change analysis. Jhon G. Botello, Lesley Frew, Jose J. Padilla, Michele C. Weigle |
IEEE Big Data | 4 |
| 2024 | Retrogressive Document Manipulation of US Federal Environmental WebsitesabstractChanges made to webpages can affect their retrievability. Often this is done with the intention of increasing the page's search engine ranking to improve overall access to information on the page. The Environmental Data and Governance Initiative (EDGI) created a dataset that describes changes on US federal environmental webpages between 2016 and 2020. EDGI noted that many environmental terms were deleted from the pages, but without user data, claims that page retrievability and public information access were lowered are only anecdotal. The Open Resource for Click Analysis in Search (ORCAS) dataset was created during the same time frame, from 2017 to 2020, and enables high quality user intent analysis without compromising on user privacy protection. We present an analysis of the intersection of the EDGI dataset and the ORCAS dataset, matching changes on federal environmental webpages with their associated queries. We use web archives and a change-text indexing system to link changes in term frequency on the pages with the queries. We find that the pages contain fewer query terms in 2020 than in 2016, lowering the pages' retrievability. The analysis provides substantive support of EDGI's claim that federal environmental pages were made less accessible between 2016 and 2020. Lesley Frew, Michael L. Nelson 0001, Michele C. Weigle |
CIKM | 3 |
| 2024 | Summarizing Web Archive Corpora via Social Media Storytelling by Automatically Selecting and Visualizing ExemplarsabstractPeople often create themed collections to make sense of an ever-increasing number of archived web pages. Some of these collections contain hundreds of thousands of documents. Thousands of collections exist, many covering the same topic. Few collections include standardized metadata. This scale makes understanding a collection an expensive proposition. Our Dark and Stormy Archives (DSA) five-process model implements a novel summarization method to help users understand a collection by combining web archives and social media storytelling. The five processes of the DSA model are: select exemplars, generate story metadata, generate document metadata, visualize the story, and distribute the story. Selecting exemplars produces a set of k documents from the N documents in the collection, where k < < N , thus reducing the number of documents visitors need to review to understand a collection. Generating story and document metadata selects images, titles, descriptions, and other content from these exemplars. Visualizing the story ties this metadata together in a format the visitor can consume. Without distributing the story, it is not shared for others to consume. We present a research study demonstrating that our algorithmic primitives can be combined to select relevant exemplars that are otherwise undiscoverable using a conventional search engine and query generation methods. Having demonstrated improved methods for selecting exemplars, we visualize the story. Previous work established that the social card is the best format for visitors to consume surrogates. The social card combines metadata fields, including the document’s title, a brief description, and a striking image. Social cards are commonly found on social media platforms. We discovered that these platforms perform poorly for mementos and rely on web page authors to supply the necessary values for these metadata fields. With web archives, we often encounter archived web pages that predate the existence of this metadata. To generate this missing metadata and ensure that storytelling is available for these documents, we apply machine learning to generate the images needed for social cards with a Precision@1 of 0.8314. We also provide the length values needed for executing automatic summarization algorithms to generate document descriptions. Applying these concepts helps us create the visualizations needed to fulfill the final processes of story generation. We close this work with examples and applications of this technology. Shawn M. Jones, Martin Klein 0001, Michele C. Weigle, Michael L. Nelson 0001 |
ACM Trans. Web | 3 |
| 2023 | It's Not Just GitHub: Identifying Data and Software Sources Included in Publications
Emily Escamilla, Lamia Salsabil, Martin Klein 0001, Jian Wu 0006, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 5 |
| 2023 | Synthesizing Web Archive Collections into Big Data: Lessons from Mining Data from Web Archives
Shawn M. Jones, Himarsha R. Jayanetti, Martin Klein 0001, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 4 |
| 2023 | To Re-experience the Web: A Framework for the Transformation and Replay of Archived Web PagesabstractWhen replaying an archived web page, or memento , the fundamental expectation is that the page should be viewable and function exactly as it did at the archival time. However, this expectation requires web archives upon replay to modify the page and its embedded resources so that all resources and links reference the archive rather than the original server. Although these modifications necessarily change the state of the representation, it is understood that without them the replay of mementos from the archive would not be possible. The process of replaying mementos and the modifications made to the representations by web archives varies between archives. Because of this, there is no standard terminology for describing the replay and needed modifications. In this article, we propose terminology for describing the existing styles of replay and the modifications made on the part of web archives to mementos to facilitate replay. Because of issues discovered with server-side only modifications, we propose a general framework for the auto-generation of client-side rewriting libraries. Finally, we evaluate the effectiveness of using a generated client-side rewriting library to augment the existing replay systems of web archives by crawling mementos replayed from the Internet Archive’s Wayback Machine with and without the generated client-side rewriter. By using the generated client-side rewriter, we were able to decrease the cumulative number of requests blocked by the content security policy of the Wayback Machine for 577 mementos by 87.5% and increased the cumulative number of requests made by 32.8%. We were also able to replay mementos that were previously not replayable from the Internet Archive. Many of the client-side rewriting ideas described in this work have been implemented into Wombat, a client-side URL rewriting system that is used by the Webrecorder, Pywb, and Wayback Machine playback systems. John A. Berlin, Mat Kelly, Michael L. Nelson 0001, Michele C. Weigle |
ACM Trans. Web | 4 |
| 2022 | The Rise of GitHub in Scholarly Publications
Emily Escamilla, Martin Klein 0001, Talya Cooper, Vicky Rampin, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 5 |
| 2022 | Robots Still Outnumber Humans in Web Archives, But Less Than Before
Himarsha R. Jayanetti, Kritika Garg, Sawood Alam, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 5 |
| 2022 | Creating Structure in Web Archives with Collections: Different Concepts from Web Archivists
Himarsha R. Jayanetti, Shawn M. Jones, Martin Klein 0001, Alex Osbourne, Paul Koerbin, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 7 |
| 2022 | A Chromium-Based Memento-Aware Web Browser
Abigail Mabe, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 3 |
| 2021 | Where Did the Web Archive Go?
Mohamed Aturban, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 3 |
| 2019 | Social Cards Probably Provide For Better Understanding Of Web Archive CollectionsabstractUsed by a variety of researchers, web archive collections have become invaluable sources of evidence. If a researcher is presented with a web archive collection that they did not create, how do they know what is inside so that they can use it for their own research? Search engine results and social media links are represented as surrogates, small easily digestible summaries of the underlying page. Search engines and social media have a different focus, and hence produce different surrogates than web archives. Search engine surrogates help a user answer the question "Will this link meet my information need?" Social media surrogates help a user decide "Should I click on this?" Our use case is subtly different. We hypothesize that groups of surrogates together are useful for summarizing a collection. We want to help users answer the question of "What does the underlying collection contain?" But which surrogate should we use? With Mechanical Turk participants, we evaluate six different surrogate types against each other. We find that the type of surrogate does not influence the time to complete the task we presented the participants. Of particular interest are social cards, surrogates typically found on social media, and browser thumbnails, screen captures of web pages rendered in a browser. At p=0.0569, and p=0.0770, respectively, we find that social cards and social cards paired side-by-side with browser thumbnails probably provide better collection understanding than the surrogates currently used by the popular Archive-It web archiving platform. We measure user interactions with each surrogate and find that users interact with social cards less than other types. The results of this study have implications for our web archive summarization work, live web curation platforms, social media, and more. Shawn M. Jones, Michele C. Weigle, Michael L. Nelson 0001 |
CIKM | 2 |
| 2017 | Comparing the Archival Rate of Arabic, English, Danish, and Korean Language Web PagesabstractIt has long been suspected that web archives and search engines favor Western and English language webpages. In this article, we quantitatively explore how well indexed and archived Arabic language webpages are as compared to those from other languages. We began by sampling 15,092 unique URIs from three different website directories: DMOZ (multilingual), Raddadi, and Star28 (the last two primarily Arabic language). Using language identification tools, we eliminated pages not in the Arabic language (e.g., English-language versions of Aljazeera pages) and culled the collection to 7,976 Arabic language webpages. We then used these 7,976 pages and crawled the live web and web archives to produce a collection of 300,646 Arabic language pages. We compared the analysis of Arabic language pages with that of English, Danish, and Korean language pages. First, for each language, we sampled unique URIs from DMOZ; then, using language identification tools, we kept only pages in the desired language. Finally, we crawled the archived and live web to collect a larger sample of pages in English, Danish, or Korean. In total for the four languages, we analyzed over 500,000 webpages. We discovered: (1) English has a higher archiving rate than Arabic, with 72.04% archived. However, Arabic has a higher archiving rate than Danish and Korean, with 53.36% of Arabic URIs archived, followed by Danish and Korean with 35.89% and 32.81% archived, respectively. (2) Most Arabic and English language pages are located in the United States; only 14.84% of the Arabic URIs had an Arabic country code top-level domain (e.g., sa) and only 10.53% had a GeoIP in an Arabic country. Most Danish-language pages were located in Denmark, and most Korean-language pages were located in South Korea. (3) The presence of a webpage in a directory positively impacts indexing and presence in the DMOZ directory, specifically, positively impacts archiving in all four languages. In this work, we show that web archives and search engines favor English pages. However, it is not universally true for all Western-language webpages because, in this work, we show that Arabic webpages have a higher archival rate than Danish language webpages. Lulwah M. Alkwai, Michael L. Nelson 0001, Michele C. Weigle |
ACM Trans. Inf. Syst. | 3 |
| 2016 | InterPlanetary Wayback: Peer-To-Peer Permanence of Web Archives
Mat Kelly, Sawood Alam, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 4 |
| 2015 | Detecting Off-Topic Pages in Web Archives
Yasmin AlNoamany, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 2 |
| 2015 | Characteristics of Social Media Stories
Yasmin AlNoamany, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 2 |
| 2015 | Quantifying Orphaned Annotations in Hypothes.is
Mohamed Aturban, Michael L. Nelson 0001, Michele C. Weigle |
TPDL | 3 |
| 2013 | Who and What Links to the Internet Archive
Yasmin AlNoamany, Ahmed Alsum, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 3 |
| 2013 | Profiling Web Archive Coverage for Top-Level Domain and Content Language
Ahmed Alsum, Michele C. Weigle, Michael L. Nelson 0001, Herbert Van de Sompel |
TPDL | 2 |
| 2013 | On the Change in Archivability of Websites Over Time
Mat Kelly, Justin F. Brunelle, Michele C. Weigle, Michael L. Nelson 0001 |
TPDL | 3 |