VLDB 2026 Research / reviewers in the wild / expert
Scott McClellan
dblp:303/1288
· DBLP profile ↗
4ranked-venue papers in the field
0as first author
4since 2021 · last 2023
0000-0002-1524-8346ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Investigating Data Reusability in Density Functional Theory StudiesabstractOver the last decade, there has been a significant increase in supporting reproducible computational research (RCR) [1]. The global adoption of the FAIR principles [2] stands as a key indicator of this trend. Specifically, federal and global research funding agencies have increasingly mandated scientific data and related products, such as code and algorithms, be made Findable, Accessible, Interoperable, and Reusable (FAIR) [2]. Rob Fleur, Addy Ireland, Xintong Zhao, Scott McClellan, Eric Paltoo, Channyung Lee, Xiaohua Hu 0001, Elif Ertekin, Jane Greenberg |
IEEE Big Data | 4 |
| 2023 | When LLM Meets Material Science: An Investigation on MOF Synthesis LabelingabstractRecent developments in Large Language Models (LLMs) have advanced the natural language processing (NLP) studies to a new era [1], [2], [4]–[6]. In generic domains, LLMs have become a key component in wide variety of state-of-the-art NLP tasks. In addition, prompt learning enables LLMs-based models to reach robust performance with much smaller training data. Xintong Zhao, Kyle Langlois, Jacob Furst 0002, Scott McClellan, Rob Fleur, Xiaohua Hu 0001, Fernando J. Uribe-Romo, Diego A. Gómez-Gualdrón, Jane Greenberg |
IEEE Big Data | 4 |
| 2022 | Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)abstractBuilding a knowledge graph is a time-consuming and costly process which often applies complex natural language processing (NLP) methods for extracting knowledge graph triples from text corpora. Pre-trained large Language Models (PLM) have emerged as a crucial type of approach that provides readily available knowledge for a range of AI applications. However, it is unclear whether it is feasible to construct domain-specific knowledge graphs from PLMs. Motivated by the capacity of knowledge graphs to accelerate data-driven materials discovery, we explored a set of state-of-the-art pre-trained general-purpose and domain-specific language models to extract knowledge triples for metal-organic frameworks (MOFs). We created a knowledge graph benchmark with 7 relations for 1248 published MOF synonyms. Our experimental results showed that domain-specific PLMs consistently outperformed the general-purpose PLMs for predicting MOF related triples. The overall benchmarking results, however, show that using the present PLMs to create domain-specific knowledge graphs is still far from being practical, motivating the need to develop more capable and knowledgeable pre-trained language models for particular applications in materials science. Jane Greenberg, Xiaohua Hu 0001, Alexander Kalinowski, Xintong Zhao, Scott McClellan, Fernando J. Uribe-Romo, Kyle Langlois, Jacob Furst 0002, Diego A. Gómez-Gualdrón, Fernando Fajardo-Rojas, Katherine Ardila, Semion Saikin, Corey A. Harper, Ron Daniel Jr. 0001 |
IEEE Big Data | 7 |
| 2021 | Knowledge Graph-Empowered Materials DiscoveryabstractIn this position paper, we describe research on knowledge graph-empowered materials science prediction and discovery. The research consists of several key components including ontology mapping, materials data annotation, and information extraction from unstructured scholarly articles. We argue that although big data generated by simulations and experiments have motivated and accelerated the data-driven science, the distribution and heterogeneity of materials science-related big data hinders major advancements in the field. Knowledge graphs, as semantic hubs, integrate disparate data and provide a feasible solution to addressing this challenge. We design a knowledge-graph based approach for data discovery, extraction, and integration in materials science. Xintong Zhao, Jane Greenberg, Scott McClellan, Yong-Jie Hu, Steven Lopez, Semion Saikin, Xiaohua Hu 0001 |
IEEE BigData | 3 |