Pei-Yu Hou

dblp:259/6846 · DBLP profile ↗
← Back
6ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0001-5429-476XORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (1 first)Database Systems & Data Management · 1 (1 first)Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2025 LiteKG: A Lightweight LLM-Assisted Framework for Domain-Specific Knowledge Graph Construction
Yi-Mo Ho, Wei-Pin Ku, Pei-Yu Hou
IEEE Big Data3
2024 FPP-Hunter: Expert-Guided Discovery Of Functional Path Patterns
abstract
In the context of using big data to improve healthcare and life sciences and with a specific focus on drug discovery, we consider the problem of formulating biomedical mechanism-of-action (MOA) hypotheses that can explain how specific drugs treat specific diseases. Our aim is to enable scalable mining and interpretation of MOA hypotheses enabling drug discovery and repurposing on large-scale biomedical knowledge graphs (KGs).The approach that we introduce to address this problem centers on expert-guided generation of candidate MOA hypotheses in the form of regular-path KG patterns between the KG nodes for the entities of interest, such as drugs and diseases. We call those patterns that represent promising candidate MOAs functional path patterns (FPPs), and call the proposed approach FPP-Hunter. The results of a drug-disease case study that we have conducted with the biomedical KG ROBOKOP suggest that the proposed approach has the potential to address scalability challenges in forming promising MOA hypotheses using large-scale KGs, in drug repurposing and potentially beyond.
Daniel R. Korn, Jon-Michael Beasley, Kara Schatz, Pei-Yu Hou, Alexander Tropsha, Rada Chirkova
IEEE Big Data4
2023 BUILD-KG: Integrating Heterogeneous Data Into Analytics-Enabling Knowledge Graphs
abstract
Knowledge graphs (KGs), with their flexible encoding of heterogeneous data, have been increasingly used in a variety of applications. At the same time, domain data are routinely stored in formats such as spreadsheets, text, or figures. Storing such data in KGs can open the door to more complex types of analytics, which might not be supported by the data sources taken in isolation. Giving domain experts the option to use a predefined automated workflow for integrating heterogeneous data from multiple sources into a single unified KG could significantly alleviate their data-integration time and resource burden, while potentially resulting in higher-quality KG data capable of enabling meaningful rule mining and machine learning.In this paper we introduce a domain-agnostic workflow called BUILD-KG for integrating heterogeneous scientific and experimental data from multiple sources into a single unified KG potentially enabling richer analytics. BUILD-KG is broadly applicable, accepting input data in popular structured and unstructured formats. BUILD-KG is also designed to be carried out with end users as humans-in-the-loop, which makes it domain aware. We present the workflow, report on our experiences with applying it to scientific and experimental data in the materials science domain, and provide suggestions for involving domain scientists in BUILD-KG as humans-in-the-loop.
Kara Schatz, Pei-Yu Hou, Alexey V. Gulyuk, Yaroslava G. Yingling, Rada Chirkova
IEEE Big Data2
2023 Provenance-Aware Data Integration and Summarization Querying for Knowledge Graphs
Pei-Yu Hou, Jing Ao, Kara Schatz, Alexey V. Gulyuk, Yaroslava G. Yingling, Rada Chirkova
iiWAS1
2022 Compact Walks: Taming Knowledge-Graph Embeddings with Domain- and Task-Specific Pathways
abstract
Knowledge-graph (KG) embeddings have emerged as a promise in addressing challenges faced by modern biomedical research, including the growing gap between therapeutic needs and available treatments. The popularity of KG embeddings in graph analytics is on the rise, due at least partially to the presumed semanticity of the learned embeddings. Unfortunately, the ability of a node neighborhood picked up by an embedding to capture the node's semantics may depend on the characteristics of the data. One of the reasons for this problem is that KG nodes can be promiscuous, that is, associated with a number of different relationships that are not unique or indicative of the properties of the nodes.
Pei-Yu Hou, Daniel R. Korn, Cleber C. Melo-Filho, David R. Wright 0001, Alexander Tropsha, Rada Chirkova
SIGMOD Conference1
2019 Collaborative Workflow for Analyzing Large-Scale Data for Antimicrobial Resistance: An Experience Report
abstract
In real-life analytics-oriented information-integration projects, the processes of information curation and integration cannot be completely automated. Rather, in each large-scale project the key objectives include maximizing scalability and throughput, while at the same time keeping the processes manageable and productive for the human experts in the loop. In this paper, we describe our experience with addressing these major objectives in the process of building a scalable end-to-end data-extraction, integration, and analytics workflow in the domain of antimicrobial resistance (AMR). The workflow is built using open-source tools, with the aims of enhancing the efficiency and accuracy of data collection and integration, while involving an acceptable level of efforts by collaborative multidisciplinary teams of humans-in-the-loop. We present the components of the proposed workflow, outline the challenges encountered in its development and testing, and discuss the experiences and lessons learned in enabling AMR experts and data analysts to interact with the workflow, with some of the lessons potentially applicable to other application domains.
Pei-Yu Hou, Jing Ao, Andrew J. Rindos, Shivaramu Keelara, Paula J. Fedorka-Cray, Rada Chirkova
IEEE BigData1