Lori A. Perine

dblp:289/3034 · DBLP profile ↗
← Back
3ranked-venue papers in the field
2as first author
2since 2021 · last 2024
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (2 first)
YearPublicationVenuePosition
2024 Model Selection for HERITAGE-AI: Evaluating LLMs for Contextual Data Analysis of Maryland's Domestic Traffic Ads (1824-1864)
abstract
The HERITAGE-AI (Harnessing Enhanced Research and Instructional Technologies for Archival Generative Exploration using AI), as part of the IMLS grant initiative, GenAI-4-Archive, aims to analyze sensitive historical datasets ethically using advanced AI technologies. One of the key tasks of this project focuses on selecting the most suitable Large Language Model (LLM) for analyzing the Domestic Traffic Ads (DTA) published in Maryland between 1824 and 1864 by slave traders—a dataset rich in historical significance yet fraught with ethical considerations. Analyzing sensitive historical datasets presents unique ethical and technical challenges. This paper presents a comparative evaluation of leading LLMs to identify the optimal model to meet HERITAGE-AI’s objectives. We survey contemporary models, including OpenAI’s GPT-4o, Anthropic’s Claude Sonnet, Meta’s Llama 3.2, and Google’s Gemini, to identify the most suitable model for Generative AI-based analysis of the DTA dataset. The objective is to select an LLM that can handle the sensitive nature of the data responsibly while providing accurate and insightful analysis. Three critical evaluation criteria, among others, are established for this reason: Sensitivity to Historical Context, Privacy and Security, and Customizability. Our analysis follows a three-step approach: evaluating free versions, paid versions, and enterprise-grade cloud-based implementations of these LLMs. Our findings reveal that while free and paid versions offer varying degrees of accessibility, they fall short in providing the necessary privacy, security, multi-user access, and customization required for analyzing sensitive historical data like the DTA dataset. In the third step, by comparing the cloud-based implementations of Azure OpenAI’s GPT-4o, AWS Bedrock’s Claude, and AWS Bedrock’s Llama3.2 LLMs, Azure openAI GPT-4o emerges as the most suitable option for this project. Although GPT-4o and Claude were close contenders, Gpt-4o demonstrated robust mechanisms due to its high accuracy, ethical sensitivity, robust privacy controls, and scalability in a cloud-based environment. It also offers extensive customizability, allowing for effective integration of the DTA dataset and alignment with the project’s ethical standards. Future work will involve domain experts and community members in implementing Azure OpenAI GPT-4o for the DTA dataset analysis.
Rajesh Kumar Gnanasekaran, Lori A. Perine, Mark Conrad, Richard Marciano
IEEE Big Data2
2024 Historic Black Lives Matter: Recovering Hidden Knowledge in Archives through Interactive Data Visualization
abstract
This paper presents the Historical Black Lives Matter (HBLM) case study, an exploratory application of interactive data visualization to a collection of manumissions documents in the Legacy of Slavery (LoS) project at the Maryland State Archives, with the goal of enhancing discovery and recovering hidden knowledge. The case study extends prior interdisciplinary research on applying computational treatments to LoS collections and contributes to research in Computational Archival Science (CAS), computational thinking, and data visualization to enhance access to archival collections. Three design objectives are addressed: representation of people, user experience, and facilitation of knowledge discovery. The paper is organized to demonstrate a customizable workflow for the process of formulating design based on data visualization principles, implementing designs with open-source tools, and incorporating user evaluation in service to successfully fulfilling the design objectives and related functionality in the final implemented design. Examples of hidden knowledge recovered using the visualizations are presented, providing new insights into Maryland’s antebellum Black population. The data visualization design methods and practices permitted investigation at a more granular level, and enabled communication of a richer narrative. Use of open-source software makes these methods accessible to archivists, information professionals, and researchers, and supports creation of artifacts for research, teaching and learning. Future extensions could incorporate advanced computational techniques to enable map features, network and textual analysis, and dynamic query-based composition of visualizations.
Lori A. Perine
IEEE Big Data1
2020 Computational Treatments to Recover Erased Heritage: A Legacy of Slavery Case Study (CT-LoS)
abstract
Graduate students at the University of Maryland's College of Information Studies (UMD iSchool) collaborated in interdisciplinary teams on a case study to explore application of computational methodologies to datafied collections related to slavery in the Maryland State Archives (MSA). Two research questions were examined: (1) What are the opportunities and limitations for using computational methods and open source tools to characterize data encoded within records of enslavement and to discover new patterns and relationships in that data? (2) How does knowledge of social and cultural systems impact those opportunities and limitations? Computational methods and tools were most effectively used when socio-cultural contextualization and technology's role as a mediator of representation were taken into account. Three additional technical research areas are identified to enhance recovery of heritage hidden in records of enslavement: visualization, graph databases, and ontologies and metadata.
Lori A. Perine, Rajesh Kumar Gnanasekaran, Phillip Nicholas, Alexis Hill, Richard Marciano
IEEE BigData1