Jane Greenberg

dblp:76/6493 · DBLP profile ↗
← Back
39ranked-venue papers in the field
11as first author
11since 2021 · last 2025
0000-0001-7819-5360ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 18 (7 first)Big Data, Cloud & Distributed Data Systems · 13 (1 first)Information Retrieval & Web Search · 8 (3 first)
YearPublicationVenuePosition
2025 Rate-Distortion Guided Knowledge Graph Construction from Lecture Notes Using Gromov-Wasserstein Optimal Transport
Ruhma Hashmi, Michelle Rogers, Jane Greenberg, Brian K. Smith
IEEE Big Data4
2025 From Analog Records to Computational Research Data: Building the AI-Ready Lab Notebook
Joel Pepper, Zach Siapano, Jacob Furst 0002, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg
IEEE Big Data6
2024 AI-Ready Data: Knowledge Extraction from Archival Lab Notebooks
abstract
Collections of analog lab notebooks are an invaluable source of data about research conditions, steps, and outcomes, and in aggregate have the potential to provide new insights into the successes, failures and pedagogy of research laboratories. Unfortunately, these artifacts are increasingly at risk of being lost from the historical scientific record, given limited archiving and an absence of computational and AI readiness. This paper reports on research addressing this challenge by testing mechanisms for transforming digital scans of analog lab notebooks into AI-ready data resources. The research being pursued is framed by the field of computational archival science (CAS) and the aim to utilize analog, research lab notebook data for scientific study. The paper presents background context on archival lab notebooks and CAS, discusses MOF (metal organic frameworks) and COF (covalent organic frameworks) synthesis – the scientific domain of the lab notebooks under study, and details our research methods. We demonstrate a promising approach that automatically segments pages into discrete entry types, extracts the contents of those entries, refines the output and assesses the automated results. These efforts represent a first step towards developing a framework for both improving the usability of archival lab notebooks, and enabling their contents to be used in subsequent scientific inquiry.
Joel Pepper, Elizabeth Jones, Xintong Zhao, Jacob Furst 0002, Kyle Langlois, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg
IEEE Big Data8
2023 Specimen Outlining: A Computational Archival Science Approach
abstract
Computational archival science (CAS) provides new pathways for research. Biologists, for example, can perform scientific studies by applying AI/ML to digital biological specimen collections and explore questions that were not possible in the analog world. One such approach is the application of computational methods for specimen outlining to assist with specimen identification, morphometry, and other scientific questions. The challenge is to determine how to computationally generate and represent a specimen’s outline. The research presented in this paper addresses this challenge, through the deployment of elliptical Fourier descriptors (EFDs). The paper describes the image processing pipeline for extracting fish outlines, a key morphological feature, and representing the outlines using EFDs. In addition, our research presents the application of machine learning classification on the EFDs. The resulting dataset is well suited for a variety of machine learning-based downstream analyses, including classification by genus and species. Overall, the classification tests produced a 96.3% accuracy, demonstrating the distinguishing nature of the EFDs, and by proxy, the fish outlines as a whole. Broadly, these results indicate the effectiveness of archival specimen usage in machine learning applications, and demonstrate specimen outlining via Fourier descriptors as a computational archival science approach.
David E. Breen, Andrew Senin, Ajani Levere, Joel Pepper, Jane Greenberg
IEEE Big Data5
2023 Investigating Data Reusability in Density Functional Theory Studies
abstract
Over the last decade, there has been a significant increase in supporting reproducible computational research (RCR) [1]. The global adoption of the FAIR principles [2] stands as a key indicator of this trend. Specifically, federal and global research funding agencies have increasingly mandated scientific data and related products, such as code and algorithms, be made Findable, Accessible, Interoperable, and Reusable (FAIR) [2].
Rob Fleur, Addy Ireland, Xintong Zhao, Scott McClellan, Eric Paltoo, Channyung Lee, Xiaohua Hu 0001, Elif Ertekin, Jane Greenberg
IEEE Big Data11
2023 When LLM Meets Material Science: An Investigation on MOF Synthesis Labeling
abstract
Recent developments in Large Language Models (LLMs) have advanced the natural language processing (NLP) studies to a new era [1], [2], [4]–[6]. In generic domains, LLMs have become a key component in wide variety of state-of-the-art NLP tasks. In addition, prompt learning enables LLMs-based models to reach robust performance with much smaller training data.
Xintong Zhao, Kyle Langlois, Jacob Furst 0002, Scott McClellan, Rob Fleur, Xiaohua Hu 0001, Fernando J. Uribe-Romo, Diego A. Gómez-Gualdrón, Jane Greenberg
IEEE Big Data10
2022 Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)
abstract
Building a knowledge graph is a time-consuming and costly process which often applies complex natural language processing (NLP) methods for extracting knowledge graph triples from text corpora. Pre-trained large Language Models (PLM) have emerged as a crucial type of approach that provides readily available knowledge for a range of AI applications. However, it is unclear whether it is feasible to construct domain-specific knowledge graphs from PLMs. Motivated by the capacity of knowledge graphs to accelerate data-driven materials discovery, we explored a set of state-of-the-art pre-trained general-purpose and domain-specific language models to extract knowledge triples for metal-organic frameworks (MOFs). We created a knowledge graph benchmark with 7 relations for 1248 published MOF synonyms. Our experimental results showed that domain-specific PLMs consistently outperformed the general-purpose PLMs for predicting MOF related triples. The overall benchmarking results, however, show that using the present PLMs to create domain-specific knowledge graphs is still far from being practical, motivating the need to develop more capable and knowledgeable pre-trained language models for particular applications in materials science.
Jane Greenberg, Xiaohua Hu 0001, Alexander Kalinowski, Xintong Zhao, Scott McClellan, Fernando J. Uribe-Romo, Kyle Langlois, Jacob Furst 0002, Diego A. Gómez-Gualdrón, Fernando Fajardo-Rojas, Katherine Ardila, Semion Saikin, Corey A. Harper, Ron Daniel Jr. 0001
IEEE Big Data2
2022 Metadata Verification: A Workflow for Computational Archival Science
abstract
Researchers seeking to apply computational methods are increasingly turning to scientific digital archives containing images of specimens. Unfortunately, metadata errors can inhibit the discovery and use of scientific archival images. One such case is the NSF-sponsored Biology Guided Neural Network (BGNN) project, where an abundance of metadata errors has significantly delayed development of a proposed, new class of neural networks. This paper reports on research addressing this challenge. We present a prototype workflow for specimen scientific name metadata verification that is grounded in Computational Archival Science (CAS), report on a taxonomy of specimen name metadata error types with preliminary solutions. Our 3-phased workflow includes tag extraction, text processing, and interactive assessment. A baseline test with the prototype workflow identified at least 15 scientific name metadata errors out of 857 manually reviewed, potentially erroneous specimen images, corresponding to a ∼0.2% error rate for the full image dataset. The prototype workflow minimizes the amount of time domain experts need to spend reviewing archive metadata for correctness and AI-readiness before these archival images can be utilized in downstream analysis.
Joel Pepper, Andrew Senin, Dom Jebbia, David E. Breen, Jane Greenberg
IEEE Big Data5
2021 Computational Curation and the Application of Large-Scale Vocabularies
abstract
Paper presents an exploratory case study comparing stemming and lemmatization results for the automatic application of large-scale controlled vocabularies processed against archival encyclopedia entries. The results report relative recall and precision evaluations across both results. Research shows that while stemming has a higher relative recall, lemmatization results in a higher relevance score and eliminates the over-stemming challenges. Results provide insight into improving automatic curation workflows for archival resources.
Sam Grabus, Jane Greenberg
IEEE BigData2
2021 Fine-Tuning BERT Model for Materials Named Entity Recognition
abstract
Scientific literature presents a wellspring of cutting-edge knowledge for materials science, including valuable data (e.g., numerical data from experiment results, material properties and structure). These data are critical for accelerating materials discovery by data-driven machine learning (ML) methods. The challenge is, it is impossible for humans to manually extract and retain this knowledge due to the extensive and growing volume of publications.To this end, we explore a fine-tuned BERT model for extracting knowledge. Our preliminary results show that our fine-tuned Bert model reaches an f-score of 85% for the materials named entity recognition task. The paper covers background, related work, methodology including tuning parameters, and our overall performance evaluation. Our discussion offers insights into our results, and points to directions for next steps.
Xintong Zhao, Jane Greenberg, Xiaohua Hu 0001
IEEE BigData2
2021 Knowledge Graph-Empowered Materials Discovery
abstract
In this position paper, we describe research on knowledge graph-empowered materials science prediction and discovery. The research consists of several key components including ontology mapping, materials data annotation, and information extraction from unstructured scholarly articles. We argue that although big data generated by simulations and experiments have motivated and accelerated the data-driven science, the distribution and heterogeneity of materials science-related big data hinders major advancements in the field. Knowledge graphs, as semantic hubs, integrate disparate data and provide a feasible solution to addressing this challenge. We design a knowledge-graph based approach for data discovery, extraction, and integration in materials science.
Xintong Zhao, Jane Greenberg, Scott McClellan, Yong-Jie Hu, Steven Lopez, Semion Saikin, Xiaohua Hu 0001
IEEE BigData2
2020 A Computational Approach to Historical Ontologies
abstract
This paper presents a use case exploring the application of the Archival Resource Key (ARK) persistent identifier for promoting and maintaining ontologies. In particular we look at improving computation with an in-house ontology server in the context of temporally aligned vocabularies. This effort demonstrates the utility of ARKs in preparing historical ontologies for computational archival science.
Mat Kelly, Jane Greenberg, Christopher B. Rauch, Sam Grabus, Joan P. Boone, John A. Kunze, Peter Melville Logan
IEEE BigData2
2020 Data objects and documenting scientific processes: An analysis of data events in biodiversity data papers
abstract
The data paper, an emerging scholarly genre, describes research data sets and is intended to bridge the gap between the publication of research data and scientific articles. Research examining how data papers report data events, such as data transactions and manipulations, is limited. The research reported on in this article addresses this limitation and investigated how data events are inscribed in data papers. A content analysis was conducted examining the full texts of 82 data papers, drawn from the curated list of data papers connected to the Global Biodiversity Information Facility. Data events recorded for each paper were organized into a set of 17 categories. Many of these categories are described together in the same sentence, which indicates the messiness of data events in the laboratory space. The findings challenge the degrees to which data papers are a distinct genre compared to research articles and they describe data‐centric research processes in a through way. This article also discusses how our results could inform a better data publication ecosystem in the future.
Kai Li 0010, Jane Greenberg, Jillian Dunic
J. Assoc. Inf. Sci. Technol.2
2016 Permanence and Temporal Interoperability of Metadata in the Linked Open Data Environment
Shigeo Sugimoto, Chunqiu Li, Mitsuharu Nagamori, Jane Greenberg
Dublin Core Conference4
2015 Evolution of an Application Profile: Advancing Metadata Best Practices through the Dryad Data Repository
Edward M. Krause, Erin Clary, Adrian Ogletree, Jane Greenberg
Dublin Core Conference4
2015 Advancing Materials Science Semantic Metadata via HIVE
Jane Greenberg, Adrian Ogletree, Garritt J. Tucker
Dublin Core Conference2
2014 Metadata capital: Simulating the predictive value of Self-Generated Health Information (SGHI)
abstract
Metadata is crucial for understanding data, and can be viewed as a form of capital in the context of Big data. This paper reports on research simulating the potential of SGHI (Self-Generated Health Information) for predicting asthma episodes. A data set of 2,000 cases was generated using the Monte Carlo simulation method, with secondary modifications on air quality and geo-location. The research is being pursued as part of a National Consortium for Data Science (NCDS) effort. The research conducted demonstrates that metadata has an inherent “predictive value” and confirms that metadata is crucial for data analytics. The work presented also provides insights into the best direction for future work in this area.
Jane Greenberg, Adrian Ogletree, Angela P. Murillo, Thomas P. Caruso, Herbie Huang
IEEE BigData1
2014 "Lo-Fi to Hi-Fi": A New Metadata Approach in the Third World with the eGranary Digital Library
Deborah Maron, Cliff Missen, Jane Greenberg
Dublin Core Conference3
2013 Metadata Capital in a Data Repository
Jane Greenberg, Shea Swauger, Elena Feinstein
Dublin Core Conference1
2012 Functional and Architectural Requirements for Metadata: Supporting Discovery and Management of Scientific Data
Jian Qin 0001, Alexander Ball, Jane Greenberg
Dublin Core Conference3
2012 Analysis and synthesis of metadata goals for scientific data
abstract
The proliferation of discipline‐specific metadata schemes contributes to artificial barriers that can impede interdisciplinary and transdisciplinary research. The authors considered this problem by examining thedomains,objectives, andarchitecturesof nine metadata schemes used to document scientific data in the physical, life, and social sciences. They used a mixed‐methods content analysis andGreenberg's ( ) metadata objectives, principles, domains, and architectural layout (MODAL) framework, and derived 22 metadata‐related goals from textual content describing each metadata scheme. Relationships are identified between the domains (e.g., scientific discipline and type of data) and the categories of scheme objectives. For each strong correlation (>0.6), a Fisher's exact test for nonparametric data was used to determine significance (p < .05). Significant relationships were found between the domains and objectives of the schemes. Schemes describing observational data are more likely to have “scheme harmonization” (compatibility and interoperability with related schemes) as an objective; schemes with the objective “abstraction” (a conceptual model exists separate from the technical implementation) also have the objective “sufficiency” (the scheme defines a minimal amount of information to meet the needs of the community); and schemes with the objective “data publication” do not have the objective “element refinement.” The analysis indicates that many metadata‐driven goals expressed by communities are independent of scientific discipline or the type of data, although they are constrained by historical community practices and workflows as well as the technological environment at the time of scheme creation. The analysis reveals 11 fundamental metadata goals for metadata documenting scientific data in support of sharing research data across disciplines and domains. The authors report these results and highlight the need for more metadata‐related research, particularly in the context of recent funding agency policy changes.
Craig Willis, Jane Greenberg, Hollie White
J. Assoc. Inf. Sci. Technol.2
2010 INEX+DBPEDIA: a corpus for semantic search evaluation
abstract
This paper presents a new collection based on DBpedia and INEX for evaluating semantic search performance. The proposed corpus is used to calculate the impact of considering document's structure on the retrieval performance of the Lucene and BM25 ranking functions. Results show that BM25 outperforms Lucene in all the considered metrics and that there is room for future improvements, which may be obtained using a hybrid approach combining both semantic technology and information retrieval ranking functions.
José R. Pérez-Agüera, Javier Arroyo, Jane Greenberg, Joaquín Pérez-Iglesias, Víctor Fresno-Fernández
WWW3
2008 Web 2.0 Semantic Systems: Collaborative Learning in Science
Michael Shoffner, Jane Greenberg, Jacob Kramer-Duffield, David Woodbury
Dublin Core Conference2
2008 The Dryad Data Repository: A Singapore Framework Metadata Architecture in a DSpace Environment
Hollie White, Sarah Carrier, Abbey Thompson, Jane Greenberg, Ryan Scherle
Dublin Core Conference4
2007 The DRIADE Project: Phased Application Profile Development in Support of Open Science
Jane Greenberg, Sarah Carrier, Jed Dube
Dublin Core Conference1
2007 The DCMI Tools Application Profile
Thomas Severiens, Jane Greenberg
Dublin Core Conference2
2006 Memex Metadata (M2) for reflective learning
Jane Greenberg, Abe Crystal, Eva M. Méndez Rodríguez, John Oberlin, Michael Shoffner
Dublin Core Conference1
2006 Relevance criteria identified by health information users during Web searches
abstract
Abstract This article focuses on the relevance judgments made by health information users who use the Web. Health information users were conceptualized as motivated information users concerned about how an environmental issue affects their health. Users identified their own environmental health interests and conducted a Web search of a particular environmental health Web site. Users were asked to identify (by highlighting with a mouse) the criteria they use to assess relevance in both Web search engine surrogates and full‐text Web documents. Content analysis of document criteria highlighted by users identified the criteria these users relied on most often. Key criteria identified included (in order of frequency of appearance) research, topic, scope, data, influence, affiliation, Web characteristics, and authority/person. A power‐law distribution of criteria was observed (a few criteria represented most of the highlighted regions, with a long tail of occasionally used criteria). Implications of this work are that information retrieval (IR) systems should be tailored in terms of users' tendencies to rely on certain document criteria, and that relevance research should combine methods to gather richer, contextualized data. Metadata for IR systems, such as that used in search engine surrogates, could be improved by taking into account actual usage of relevance criteria. Such metadata should be user‐centered (based on data from users, as in this study) and context‐appropriate (fit to users' situations and tasks)
Abe Crystal, Jane Greenberg
J. Assoc. Inf. Sci. Technol.2
2005 Growing vocabularies for plant identification and scientific learning
Jane Greenberg, P. Bryan Heidorn, Stephen Seiberling, Alan S. Weakley
Dublin Core Conference1
2004 Architecting a Cross-Disciplinary Thesaurus for the Semantic Web
W. Davenport Robertson, Jane Greenberg
Dublin Core Conference2
2003 Iterative Design of Metadata Creation Tools for Resource Authors
Jane Greenberg, Abe Crystal, W. Davenport Robertson, Ellen M. Leadem
Dublin Core Conference1
2003 Open Source Software Development and Lotka's Law: Bibliometric Patterns in Programming
abstract
Abstract This research applies Lotka's Law to metadata on open source software development. Lotka's Law predicts the proportion of authors at different levels of productivity. Open source software development harnesses the creativity of thousands of programmers worldwide, is important to the progress of the Internet and many other computing environments, and yet has not been widely researched. We examine metadata from the Linux Software Map (LSM), which documents many open source projects, and Sourceforge, one of the largest resources for open source developers. Authoring patterns found are comparable to prior studies of Lotka's Law for scientific and scholarly publishing. Lotka's Law was found to be effective in understanding software development productivity patterns, and offer promise in predicting aggregate behavior of open source developers.
Gregory B. Newby, Jane Greenberg, Paul Jones 0001
J. Assoc. Inf. Sci. Technol.2
2002 Semantic Web Construction: An Inquiry of Authors' Views on Collaborative Metadata Generation
Jane Greenberg, W. Davenport Robertson
Dublin Core Conference1
2002 Abstraction versus Implementation: Issues in Formalizing the NIEHS Application Profile
Corey A. Harper, Jane Greenberg, W. Davenport Robertson, Ellen M. Leadem
Dublin Core Conference2
2001 Author-generated Dublin Core Metadata for Web Resources: A Baseline Study in an Organization
Jane Greenberg, Maria Cristina Pattuelli, Bijan Parsia, W. Davenport Robertson
Dublin Core Conference1
2001 Design and Implementation of the National Institute of Environmental Health Sciences Dublin Core Metadata Schema
W. Davenport Robertson, Ellen M. Leadem, Jed Dube, Jane Greenberg
Dublin Core Conference4
2001 Automatic query expansion via lexical-semantic relationships
abstract
Structured thesauri encode equivalent, hierarchical, and associative relationships and have been developed as indexing/retrieval tools. Despite the fact that these tools provide a rich semantic network of vocabulary terms, they are seldom employed for automatic query expansion (QE) activities. This article reports on an experiment that examined whether thesaurus terms, related to query in a specified semantic way (as synonyms and partial-synonyms (SYNs), narrower terms (NTs), related terms (RTs), and broader terms (BTs)), could be identified as having a more positive impact on retrieval effectiveness when added to a query through automatic QE. The research found that automatic QE via SYNs and NTs increased relative recall with a decline in precision that was not statistically significant, and that automatic QE via RTs and BTs increased relative recall with a decline in precision that was statistically significant. Recall-based and a precision-based ranking orders for automatic QE via semantically encoded thesauri terminology were identified. Mapping results found between end-user query terms and the ProQuest® Controlled Vocabulary (1997) (the thesaurus used in this study) are reported, and future research foci related to the investigation are discussed.
Jane Greenberg
J. Assoc. Inf. Sci. Technol.1
2001 Optimal query expansion (QE) processing methods with semantically encoded structured thesauri terminology
abstract
While researchers have explored the value of structured thesauri as controlled vocabularies for general information retrieval (IR) activities, they have not identified the optimal query expansion (QE) processing methods for taking advantage of the semantic encoding underlying the terminology in these tools. The study reported on in this article addresses this question, and examined whether QE via semantically encoded thesauri terminology is more effective in the automatic or interactive processing environment. The research found that, regardless of end-users' retrieval goals, synonyms and partial synonyms (SYNs) and narrower terms (NTs) are generally good candidates for automatic QE and that related (RTs) are better candidates for interactive QE. The study also examined end-users' selection of semantically encoded thesauri terms for interactive QE, and explored how retrieval goals and QE processes may be combined in future thesauri-supported IR systems.
Jane Greenberg
J. Assoc. Inf. Sci. Technol.1
2001 A quantitative categorical analysis of metadata elements in image-applicable metadata schemas
abstract
Abstract This article reports on a quantitative categorical analysis of metadata elements in the Dublin Core, VRA Core, REACH, and EAD metadata schemas, all of which can be used for organizing and describing images. The study found that each of the examined metadata schemas contains elements that support the discovery, use, authentication, and administration of images, and that the number and proportion of elements supporting functions in these classes varies per schema. The study introduces a new schema comparison methodology and explores the development of a class‐oriented functional metadata schema for controlling images across multiple domains.
Jane Greenberg
J. Assoc. Inf. Sci. Technol.1