Mark Conrad

dblp:273/9287 · DBLP profile ↗
← Back
3ranked-venue papers in the field
1as first author
2since 2021 · last 2024
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (1 first)
YearPublicationVenuePosition
2024 Model Selection for HERITAGE-AI: Evaluating LLMs for Contextual Data Analysis of Maryland's Domestic Traffic Ads (1824-1864)
abstract
The HERITAGE-AI (Harnessing Enhanced Research and Instructional Technologies for Archival Generative Exploration using AI), as part of the IMLS grant initiative, GenAI-4-Archive, aims to analyze sensitive historical datasets ethically using advanced AI technologies. One of the key tasks of this project focuses on selecting the most suitable Large Language Model (LLM) for analyzing the Domestic Traffic Ads (DTA) published in Maryland between 1824 and 1864 by slave traders—a dataset rich in historical significance yet fraught with ethical considerations. Analyzing sensitive historical datasets presents unique ethical and technical challenges. This paper presents a comparative evaluation of leading LLMs to identify the optimal model to meet HERITAGE-AI’s objectives. We survey contemporary models, including OpenAI’s GPT-4o, Anthropic’s Claude Sonnet, Meta’s Llama 3.2, and Google’s Gemini, to identify the most suitable model for Generative AI-based analysis of the DTA dataset. The objective is to select an LLM that can handle the sensitive nature of the data responsibly while providing accurate and insightful analysis. Three critical evaluation criteria, among others, are established for this reason: Sensitivity to Historical Context, Privacy and Security, and Customizability. Our analysis follows a three-step approach: evaluating free versions, paid versions, and enterprise-grade cloud-based implementations of these LLMs. Our findings reveal that while free and paid versions offer varying degrees of accessibility, they fall short in providing the necessary privacy, security, multi-user access, and customization required for analyzing sensitive historical data like the DTA dataset. In the third step, by comparing the cloud-based implementations of Azure OpenAI’s GPT-4o, AWS Bedrock’s Claude, and AWS Bedrock’s Llama3.2 LLMs, Azure openAI GPT-4o emerges as the most suitable option for this project. Although GPT-4o and Claude were close contenders, Gpt-4o demonstrated robust mechanisms due to its high accuracy, ethical sensitivity, robust privacy controls, and scalability in a cloud-based environment. It also offers extensive customizability, allowing for effective integration of the DTA dataset and alignment with the project’s ethical standards. Future work will involve domain experts and community members in implementing Azure OpenAI GPT-4o for the DTA dataset analysis.
Rajesh Kumar Gnanasekaran, Lori A. Perine, Mark Conrad, Richard Marciano
IEEE Big Data3
2021 Computational Archival Science is a Two-Way Street
abstract
Since its definition in 2018 much of the literature written about CAS has been about archives adopting computational theories, methods and resources. Very little has been written about computational professionals adopting archival theories, methods, or resources. The authors believe that CAS could be substantially enriched if some archival theories, methods, and resources were adopted by computational professionals in developing the systems that create and store vast troves of data. For the purposes of this paper, we will focus on two archival resources – ISO 14721 (OAIS) and ISO 16363 (Trustworthy Digital Repositories). These two resources offer recommendations for long term preservation of digital assets, maintaining understandability of those assets through time, and building trustworthy digital repositories to maintain the provenance and integrity of the repository’s collections in such a manner as to provide substantial evidence of the authenticity of the data that it provides to its consumers.
Bruce Ambacher, Mark Conrad
IEEE BigData2
2020 Elevating "Everyday" Voices and People in Archives through the Application of Graph Database Technology
abstract
In a simple experiment using a graph database we demonstrate that it is possible to increase the number of access points to individual items in archival collections. We do this by leveraging existing machine readable and searchable data and metadata to identify and display relationships between persons, places, dates, events, etc. across items and collections. We discuss some of the financial, ethical and representational implications of decisions made in applying technology to archival holdings. Many decisions are made without considering the ethical and representational implications. Our experiment has illustrated some of these ethical and representational implications.
Mark Conrad, Lyneise Williams
IEEE BigData1