Joel Pepper

dblp:309/5537 · DBLP profile ↗
← Back
4ranked-venue papers in the field
3as first author
4since 2021 · last 2025
0000-0002-1601-8729ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (3 first)
YearPublicationVenuePosition
2025 From Analog Records to Computational Research Data: Building the AI-Ready Lab Notebook
Joel Pepper, Zach Siapano, Jacob Furst 0002, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg
IEEE Big Data1
2024 AI-Ready Data: Knowledge Extraction from Archival Lab Notebooks
abstract
Collections of analog lab notebooks are an invaluable source of data about research conditions, steps, and outcomes, and in aggregate have the potential to provide new insights into the successes, failures and pedagogy of research laboratories. Unfortunately, these artifacts are increasingly at risk of being lost from the historical scientific record, given limited archiving and an absence of computational and AI readiness. This paper reports on research addressing this challenge by testing mechanisms for transforming digital scans of analog lab notebooks into AI-ready data resources. The research being pursued is framed by the field of computational archival science (CAS) and the aim to utilize analog, research lab notebook data for scientific study. The paper presents background context on archival lab notebooks and CAS, discusses MOF (metal organic frameworks) and COF (covalent organic frameworks) synthesis – the scientific domain of the lab notebooks under study, and details our research methods. We demonstrate a promising approach that automatically segments pages into discrete entry types, extracts the contents of those entries, refines the output and assesses the automated results. These efforts represent a first step towards developing a framework for both improving the usability of archival lab notebooks, and enabling their contents to be used in subsequent scientific inquiry.
Joel Pepper, Elizabeth Jones, Xintong Zhao, Jacob Furst 0002, Kyle Langlois, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg
IEEE Big Data1
2023 Specimen Outlining: A Computational Archival Science Approach
abstract
Computational archival science (CAS) provides new pathways for research. Biologists, for example, can perform scientific studies by applying AI/ML to digital biological specimen collections and explore questions that were not possible in the analog world. One such approach is the application of computational methods for specimen outlining to assist with specimen identification, morphometry, and other scientific questions. The challenge is to determine how to computationally generate and represent a specimen’s outline. The research presented in this paper addresses this challenge, through the deployment of elliptical Fourier descriptors (EFDs). The paper describes the image processing pipeline for extracting fish outlines, a key morphological feature, and representing the outlines using EFDs. In addition, our research presents the application of machine learning classification on the EFDs. The resulting dataset is well suited for a variety of machine learning-based downstream analyses, including classification by genus and species. Overall, the classification tests produced a 96.3% accuracy, demonstrating the distinguishing nature of the EFDs, and by proxy, the fish outlines as a whole. Broadly, these results indicate the effectiveness of archival specimen usage in machine learning applications, and demonstrate specimen outlining via Fourier descriptors as a computational archival science approach.
David E. Breen, Andrew Senin, Ajani Levere, Joel Pepper, Jane Greenberg
IEEE Big Data4
2022 Metadata Verification: A Workflow for Computational Archival Science
abstract
Researchers seeking to apply computational methods are increasingly turning to scientific digital archives containing images of specimens. Unfortunately, metadata errors can inhibit the discovery and use of scientific archival images. One such case is the NSF-sponsored Biology Guided Neural Network (BGNN) project, where an abundance of metadata errors has significantly delayed development of a proposed, new class of neural networks. This paper reports on research addressing this challenge. We present a prototype workflow for specimen scientific name metadata verification that is grounded in Computational Archival Science (CAS), report on a taxonomy of specimen name metadata error types with preliminary solutions. Our 3-phased workflow includes tag extraction, text processing, and interactive assessment. A baseline test with the prototype workflow identified at least 15 scientific name metadata errors out of 857 manually reviewed, potentially erroneous specimen images, corresponding to a ∼0.2% error rate for the full image dataset. The prototype workflow minimizes the amount of time domain experts need to spend reviewing archive metadata for correctness and AI-readiness before these archival images can be utilized in downstream analysis.
Joel Pepper, Andrew Senin, Dom Jebbia, David E. Breen, Jane Greenberg
IEEE Big Data1