EDBT 2026 Demo / reviewers in the wild / expert
Andrew Senin
dblp:99/8850
· DBLP profile ↗
2ranked-venue papers in the field
0as first author
2since 2021 · last 2023
—ORCID · none
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Specimen Outlining: A Computational Archival Science ApproachabstractComputational archival science (CAS) provides new pathways for research. Biologists, for example, can perform scientific studies by applying AI/ML to digital biological specimen collections and explore questions that were not possible in the analog world. One such approach is the application of computational methods for specimen outlining to assist with specimen identification, morphometry, and other scientific questions. The challenge is to determine how to computationally generate and represent a specimen’s outline. The research presented in this paper addresses this challenge, through the deployment of elliptical Fourier descriptors (EFDs). The paper describes the image processing pipeline for extracting fish outlines, a key morphological feature, and representing the outlines using EFDs. In addition, our research presents the application of machine learning classification on the EFDs. The resulting dataset is well suited for a variety of machine learning-based downstream analyses, including classification by genus and species. Overall, the classification tests produced a 96.3% accuracy, demonstrating the distinguishing nature of the EFDs, and by proxy, the fish outlines as a whole. Broadly, these results indicate the effectiveness of archival specimen usage in machine learning applications, and demonstrate specimen outlining via Fourier descriptors as a computational archival science approach. David E. Breen, Andrew Senin, Ajani Levere, Joel Pepper, Jane Greenberg |
IEEE Big Data | 2 |
| 2022 | Metadata Verification: A Workflow for Computational Archival ScienceabstractResearchers seeking to apply computational methods are increasingly turning to scientific digital archives containing images of specimens. Unfortunately, metadata errors can inhibit the discovery and use of scientific archival images. One such case is the NSF-sponsored Biology Guided Neural Network (BGNN) project, where an abundance of metadata errors has significantly delayed development of a proposed, new class of neural networks. This paper reports on research addressing this challenge. We present a prototype workflow for specimen scientific name metadata verification that is grounded in Computational Archival Science (CAS), report on a taxonomy of specimen name metadata error types with preliminary solutions. Our 3-phased workflow includes tag extraction, text processing, and interactive assessment. A baseline test with the prototype workflow identified at least 15 scientific name metadata errors out of 857 manually reviewed, potentially erroneous specimen images, corresponding to a ∼0.2% error rate for the full image dataset. The prototype workflow minimizes the amount of time domain experts need to spend reviewing archive metadata for correctness and AI-readiness before these archival images can be utilized in downstream analysis. Joel Pepper, Andrew Senin, Dom Jebbia, David E. Breen, Jane Greenberg |
IEEE Big Data | 2 |