VLDB 2026 Research / reviewers in the wild / expert
Fernando J. Uribe-Romo
dblp:324/1976
· DBLP profile ↗
4ranked-venue papers in the field
0as first author
4since 2021 · last 2025
0000-0003-0212-0295ORCID · reported
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Analog Records to Computational Research Data: Building the AI-Ready Lab Notebook
Joel Pepper, Zach Siapano, Jacob Furst 0002, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg |
IEEE Big Data | 4 |
| 2024 | AI-Ready Data: Knowledge Extraction from Archival Lab NotebooksabstractCollections of analog lab notebooks are an invaluable source of data about research conditions, steps, and outcomes, and in aggregate have the potential to provide new insights into the successes, failures and pedagogy of research laboratories. Unfortunately, these artifacts are increasingly at risk of being lost from the historical scientific record, given limited archiving and an absence of computational and AI readiness. This paper reports on research addressing this challenge by testing mechanisms for transforming digital scans of analog lab notebooks into AI-ready data resources. The research being pursued is framed by the field of computational archival science (CAS) and the aim to utilize analog, research lab notebook data for scientific study. The paper presents background context on archival lab notebooks and CAS, discusses MOF (metal organic frameworks) and COF (covalent organic frameworks) synthesis – the scientific domain of the lab notebooks under study, and details our research methods. We demonstrate a promising approach that automatically segments pages into discrete entry types, extracts the contents of those entries, refines the output and assesses the automated results. These efforts represent a first step towards developing a framework for both improving the usability of archival lab notebooks, and enabling their contents to be used in subsequent scientific inquiry. Joel Pepper, Elizabeth Jones, Xintong Zhao, Jacob Furst 0002, Kyle Langlois, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg |
IEEE Big Data | 6 |
| 2023 | When LLM Meets Material Science: An Investigation on MOF Synthesis LabelingabstractRecent developments in Large Language Models (LLMs) have advanced the natural language processing (NLP) studies to a new era [1], [2], [4]–[6]. In generic domains, LLMs have become a key component in wide variety of state-of-the-art NLP tasks. In addition, prompt learning enables LLMs-based models to reach robust performance with much smaller training data. Xintong Zhao, Kyle Langlois, Jacob Furst 0002, Scott McClellan, Rob Fleur, Xiaohua Hu 0001, Fernando J. Uribe-Romo, Diego A. Gómez-Gualdrón, Jane Greenberg |
IEEE Big Data | 8 |
| 2022 | Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)abstractBuilding a knowledge graph is a time-consuming and costly process which often applies complex natural language processing (NLP) methods for extracting knowledge graph triples from text corpora. Pre-trained large Language Models (PLM) have emerged as a crucial type of approach that provides readily available knowledge for a range of AI applications. However, it is unclear whether it is feasible to construct domain-specific knowledge graphs from PLMs. Motivated by the capacity of knowledge graphs to accelerate data-driven materials discovery, we explored a set of state-of-the-art pre-trained general-purpose and domain-specific language models to extract knowledge triples for metal-organic frameworks (MOFs). We created a knowledge graph benchmark with 7 relations for 1248 published MOF synonyms. Our experimental results showed that domain-specific PLMs consistently outperformed the general-purpose PLMs for predicting MOF related triples. The overall benchmarking results, however, show that using the present PLMs to create domain-specific knowledge graphs is still far from being practical, motivating the need to develop more capable and knowledgeable pre-trained language models for particular applications in materials science. Jane Greenberg, Xiaohua Hu 0001, Alexander Kalinowski, Xintong Zhao, Scott McClellan, Fernando J. Uribe-Romo, Kyle Langlois, Jacob Furst 0002, Diego A. Gómez-Gualdrón, Fernando Fajardo-Rojas, Katherine Ardila, Semion Saikin, Corey A. Harper, Ron Daniel Jr. 0001 |
IEEE Big Data | 8 |