Xintong Zhao

dblp:03/10178 · DBLP profile ↗
← Back
7ranked-venue papers in the field
3as first author
7since 2021 · last 2025
ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6 (3 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 Relational Data Cleaning Meets Artificial Intelligence: A Survey
abstract
Abstract Relational data play a crucial role in various fields, but they are often plagued by low-quality issues such as erroneous and missing values, which can terribly impact downstream applications. To tackle these issues, relational data cleaning with traditional signals, e.g., statistics, constraints, and clusters, have been extensively studied, with interpretability and efficiency. Recently, considering the strong capability of modeling complex relationships, artificial intelligence (AI) techniques have been introduced into the data cleaning field. These AI-based methods either consider multiple cleaning signals, integrate various techniques into the cleaning system, or incorporate neural networks. Among them, methods utilizing deep neural networks are classified as deep learning (DL) based, while those that do not are classified as machine learning (ML) based. In this study, we focus on three essential tasks (i.e., error detection, data repairing, and data imputation) for cleaning relational data, to comprehensively review the representative methods using traditional or AI techniques. By comparing and analyzing two types of methods across five dimensions (cost, generalization, interpretability, efficiency, and effectiveness), we provide insights into their strengths, weaknesses, and suitable application scenarios. Finally, we analyze the challenges and open issues currently faced in data cleaning and discuss possible directions for future studies.
Xintong Zhao, Yu Sun 0027, Shaoxu Song, Xiaojie Yuan
Data Sci. Eng.2
2024 AI-Ready Data: Knowledge Extraction from Archival Lab Notebooks
abstract
Collections of analog lab notebooks are an invaluable source of data about research conditions, steps, and outcomes, and in aggregate have the potential to provide new insights into the successes, failures and pedagogy of research laboratories. Unfortunately, these artifacts are increasingly at risk of being lost from the historical scientific record, given limited archiving and an absence of computational and AI readiness. This paper reports on research addressing this challenge by testing mechanisms for transforming digital scans of analog lab notebooks into AI-ready data resources. The research being pursued is framed by the field of computational archival science (CAS) and the aim to utilize analog, research lab notebook data for scientific study. The paper presents background context on archival lab notebooks and CAS, discusses MOF (metal organic frameworks) and COF (covalent organic frameworks) synthesis – the scientific domain of the lab notebooks under study, and details our research methods. We demonstrate a promising approach that automatically segments pages into discrete entry types, extracts the contents of those entries, refines the output and assesses the automated results. These efforts represent a first step towards developing a framework for both improving the usability of archival lab notebooks, and enabling their contents to be used in subsequent scientific inquiry.
Joel Pepper, Elizabeth Jones, Xintong Zhao, Jacob Furst 0002, Kyle Langlois, Fernando J. Uribe-Romo, David E. Breen, Jane Greenberg
IEEE Big Data3
2023 Investigating Data Reusability in Density Functional Theory Studies
abstract
Over the last decade, there has been a significant increase in supporting reproducible computational research (RCR) [1]. The global adoption of the FAIR principles [2] stands as a key indicator of this trend. Specifically, federal and global research funding agencies have increasingly mandated scientific data and related products, such as code and algorithms, be made Findable, Accessible, Interoperable, and Reusable (FAIR) [2].
Rob Fleur, Addy Ireland, Xintong Zhao, Scott McClellan, Eric Paltoo, Channyung Lee, Xiaohua Hu 0001, Elif Ertekin, Jane Greenberg
IEEE Big Data3
2023 When LLM Meets Material Science: An Investigation on MOF Synthesis Labeling
abstract
Recent developments in Large Language Models (LLMs) have advanced the natural language processing (NLP) studies to a new era [1], [2], [4]–[6]. In generic domains, LLMs have become a key component in wide variety of state-of-the-art NLP tasks. In addition, prompt learning enables LLMs-based models to reach robust performance with much smaller training data.
Xintong Zhao, Kyle Langlois, Jacob Furst 0002, Scott McClellan, Rob Fleur, Xiaohua Hu 0001, Fernando J. Uribe-Romo, Diego A. Gómez-Gualdrón, Jane Greenberg
IEEE Big Data1
2022 Exploring Pre-Trained Language Models to Build Knowledge Graph for Metal-Organic Frameworks (MOFs)
abstract
Building a knowledge graph is a time-consuming and costly process which often applies complex natural language processing (NLP) methods for extracting knowledge graph triples from text corpora. Pre-trained large Language Models (PLM) have emerged as a crucial type of approach that provides readily available knowledge for a range of AI applications. However, it is unclear whether it is feasible to construct domain-specific knowledge graphs from PLMs. Motivated by the capacity of knowledge graphs to accelerate data-driven materials discovery, we explored a set of state-of-the-art pre-trained general-purpose and domain-specific language models to extract knowledge triples for metal-organic frameworks (MOFs). We created a knowledge graph benchmark with 7 relations for 1248 published MOF synonyms. Our experimental results showed that domain-specific PLMs consistently outperformed the general-purpose PLMs for predicting MOF related triples. The overall benchmarking results, however, show that using the present PLMs to create domain-specific knowledge graphs is still far from being practical, motivating the need to develop more capable and knowledgeable pre-trained language models for particular applications in materials science.
Jane Greenberg, Xiaohua Hu 0001, Alexander Kalinowski, Xintong Zhao, Scott McClellan, Fernando J. Uribe-Romo, Kyle Langlois, Jacob Furst 0002, Diego A. Gómez-Gualdrón, Fernando Fajardo-Rojas, Katherine Ardila, Semion Saikin, Corey A. Harper, Ron Daniel Jr. 0001
IEEE Big Data6
2021 Fine-Tuning BERT Model for Materials Named Entity Recognition
abstract
Scientific literature presents a wellspring of cutting-edge knowledge for materials science, including valuable data (e.g., numerical data from experiment results, material properties and structure). These data are critical for accelerating materials discovery by data-driven machine learning (ML) methods. The challenge is, it is impossible for humans to manually extract and retain this knowledge due to the extensive and growing volume of publications.To this end, we explore a fine-tuned BERT model for extracting knowledge. Our preliminary results show that our fine-tuned Bert model reaches an f-score of 85% for the materials named entity recognition task. The paper covers background, related work, methodology including tuning parameters, and our overall performance evaluation. Our discussion offers insights into our results, and points to directions for next steps.
Xintong Zhao, Jane Greenberg, Xiaohua Hu 0001
IEEE BigData1
2021 Knowledge Graph-Empowered Materials Discovery
abstract
In this position paper, we describe research on knowledge graph-empowered materials science prediction and discovery. The research consists of several key components including ontology mapping, materials data annotation, and information extraction from unstructured scholarly articles. We argue that although big data generated by simulations and experiments have motivated and accelerated the data-driven science, the distribution and heterogeneity of materials science-related big data hinders major advancements in the field. Knowledge graphs, as semantic hubs, integrate disparate data and provide a feasible solution to addressing this challenge. We design a knowledge-graph based approach for data discovery, extraction, and integration in materials science.
Xintong Zhao, Jane Greenberg, Scott McClellan, Yong-Jie Hu, Steven Lopez, Semion Saikin, Xiaohua Hu 0001
IEEE BigData1