VLDB 2026 Research / reviewers in the wild / expert
Daniel R. Korn
dblp:160/4320 · also Daniel Robert Korn
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-1780-9872ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | monarchr: an R package for querying biomedical knowledge graphsabstractSUMMARY: Biomedical knowledge graphs (KGs) aggregate and provide a wealth of information, linking genes and their variants, diseases, phenotypes, and much more. While these data are available in raw and API-hosted form, to date, functionality for working with KGs in the R programming language has been limited. We introduce monarchr, a package for querying and manipulating KG data. Support for the expansive Monarch Initiative KG is built in, and monarchr can accommodate any KG in the Knowledge Graph eXchange (KGX) format. This tidy-inspired interface offers researchers an intuitive, iterative approach to querying and visualizing KG data. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and installation instructions are available at https://github.com/monarch-initiative/monarchr. Shawn T. O'Neil, Brian M. Schilder, Kevin Schaper, Corey Cox, Daniel R. Korn, Sarah Gehrke, Chris Mungall, Melissa A. Haendel |
Bioinform. | 5 |
| 2025 | Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation frameworkabstractBACKGROUND: Computational approaches to support rare disease diagnosis are challenging to build, requiring the integration of complex data types such as ontologies, gene-to-phenotype associations, and cross-species data into variant and gene prioritisation algorithms (VGPAs). However, the performance of VGPAs has been difficult to measure and is impacted by many factors, for example, ontology structure, annotation completeness or changes to the underlying algorithm. Assertions of the capabilities of VGPAs are often not reproducible, in part because there is no standardised, empirical framework and openly available patient data to assess the efficacy of VGPAs-ultimately hindering the development of effective prioritisation tools. RESULTS: In this paper, we present our benchmarking tool, PhEval, which aims to provide a standardised and empirical framework to evaluate phenotype-driven VGPAs. The inclusion of standardised test corpora and test corpus generation tools in the PhEval suite of tools allows open benchmarking and comparison of methods on standardised data sets. CONCLUSIONS: PhEval and the standardised test corpora solve the issues of patient data availability and experimental tooling configuration when benchmarking and comparing rare disease VGPAs. By providing standardised data on patient cohorts from real-world case-reports and controlling the configuration of evaluated VGPAs, PhEval enables transparent, portable, comparable and reproducible benchmarking of VGPAs. As these tools are often a key component of many rare disease diagnostic pipelines, a thorough and standardised method of assessment is essential for improving patient diagnosis and care. Yasemin Bridges, Vinicius de Souza, Katherina G. Cortes, Melissa A. Haendel, Nomi L. Harris, Daniel R. Korn, Nikolaos M. Marinakis, Nicolas Matentzoglu, James Alastair McLaughlin, Chris Mungall, Aaron Odell, David Osumi-Sutherland, Peter N. Robinson, Damian Smedley, Julius O. B. Jacobsen |
BMC Bioinform. | 6 |
| 2025 | Towards Improving the Efficiency of Drug Repurposing by Leveraging Node Promiscuity in Biomedical Knowledge GraphsabstractTo accelerate the time- and labor-intensive processes of drug discovery and repurposing, it is increasingly common to mine knowledge sources for connections between diseases and the drugs that can treat them. In this article we address the scalability challenge in the connection mining, by introducing algorithms that can be used to find plausible mechanistic connections between drugs and the potentially associated diseases in biomedical knowledge graphs. These connections are then presented to biomedical experts as candidate hypotheses for further studies of whether the drugs can be repurposed to treat the diseases. One challenge that has to be addressed in this effort is the processing of promiscuous knowledge-graph nodes, that is, nodes associated with numerous relationships that may not be unique or indicative of the node properties. As it turns out, the multiplicity of relationships involving promiscuous graph nodes may prevent the aforementioned path-finding algorithms from aiding in drug repurposing. To address the promiscuous-node challenge, we introduce promiscuity scores for nodes and paths in knowledge graphs, and incorporate the scores in the proposed path-finding algorithms. We report experimental results that indicate that paths with low-promiscuity scores could be meaningful and of interest to biomedical experts in drug repurposing. Daniel R. Korn, Pei-Yu Hou, Kara Schatz, Jon-Michael Beasley, Alexander Tropsha, Rada Chirkova |
ACM Trans. Comput. Heal. | 1 |
| 2024 | FPP-Hunter: Expert-Guided Discovery Of Functional Path PatternsabstractIn the context of using big data to improve healthcare and life sciences and with a specific focus on drug discovery, we consider the problem of formulating biomedical mechanism-of-action (MOA) hypotheses that can explain how specific drugs treat specific diseases. Our aim is to enable scalable mining and interpretation of MOA hypotheses enabling drug discovery and repurposing on large-scale biomedical knowledge graphs (KGs).The approach that we introduce to address this problem centers on expert-guided generation of candidate MOA hypotheses in the form of regular-path KG patterns between the KG nodes for the entities of interest, such as drugs and diseases. We call those patterns that represent promising candidate MOAs functional path patterns (FPPs), and call the proposed approach FPP-Hunter. The results of a drug-disease case study that we have conducted with the biomedical KG ROBOKOP suggest that the proposed approach has the potential to address scalability challenges in forming promising MOA hypotheses using large-scale KGs, in drug repurposing and potentially beyond. Daniel R. Korn, Jon-Michael Beasley, Kara Schatz, Pei-Yu Hou, Alexander Tropsha, Rada Chirkova |
IEEE Big Data | 1 |
| 2024 | ExEmPLAR (Extracting, Exploring, and Embedding Pathways Leading to Actionable Research): a user-friendly interface for knowledge graph miningabstractSUMMARY: Knowledge graphs are being increasingly used in biomedical research to link large amounts of heterogenous data and facilitate reasoning across diverse knowledge sources. Wider adoption and exploration of knowledge graphs in the biomedical research community is limited by requirements to understand the underlying graph structure in terms of entity types and relationships, represented as nodes and edges, respectively, and learn specialized query languages for graph mining and exploration. We have developed a user-friendly interface dubbed ExEmPLAR (Extracting, Exploring, and Embedding Pathways Leading to Actionable Research) to aid reasoning over biomedical knowledge graphs and assist with data-driven research and hypothesis generation. We explain the key functionalities of ExEmPLAR and demonstrate its use with a case study considering the relationship of Trypanosoma cruzi, the etiological agent of Chagas disease, to frequently associated cardiovascular conditions. AVAILABILITY AND IMPLEMENTATION: ExEmPLAR is freely accessible at https://www.exemplar.mml.unc.edu/. For code and instructions for the using the application, see: https://github.com/beasleyjonm/AOP-COP-Path-Extractor. Jon-Michael Beasley, Daniel R. Korn, Nyssa N. Tucker, Erick T. M. Alves, Eugene N. Muratov, Chris Bizon, Alexander Tropsha |
Bioinform. | 2 |
| 2022 | Workflow for Domain- and Task-Sensitive Curation of Knowledge Graphs, with Use Case of DRKGabstractRecently, knowledge graphs have seen a significant increase in popularity in a wide variety of domains, as they provide a basis for many data-analytics and knowledge-discovery approaches. At the same time, many knowledge graphs are not immediately usable due to their format, unreadable or missing data, and inaccessibility. These issues present barriers to the exploration and use of knowledge graphs for big data analytics and knowledge discovery. In this paper we present a workflow for domain- and task-sensitive curation of large-scale knowledge graphs, and detail our experience with implementing this workflow with the biomedical knowledge graph called Drug Repurposing Knowledge Graph (DRKG). The workflow aims to address usability-related issues of real-life knowledge graphs, by performing data setup and curation that align with the needs of specific tasks and domains. Recognizing that domain experts and anticipated users of a knowledge graph provide invaluable expertise regarding the desired graph format, the proposed workflow involves them as humans-in-the-loop. We present the processes required to execute the workflow, detail our experience in the biomedical domain with the use case of DRKG, and discuss the challenges and lessons learned throughout the experience. We anticipate that the proposed workflow and experiences will be applicable to other domains, and that our workflow will enable and encourage exploration and wider use of large-scale knowledge graphs, thereby improving big data analytics. Kara Schatz, Daniel R. Korn, Alexander Tropsha, Rada Chirkova |
IEEE Big Data | 2 |
| 2022 | Compact Walks: Taming Knowledge-Graph Embeddings with Domain- and Task-Specific PathwaysabstractKnowledge-graph (KG) embeddings have emerged as a promise in addressing challenges faced by modern biomedical research, including the growing gap between therapeutic needs and available treatments. The popularity of KG embeddings in graph analytics is on the rise, due at least partially to the presumed semanticity of the learned embeddings. Unfortunately, the ability of a node neighborhood picked up by an embedding to capture the node's semantics may depend on the characteristics of the data. One of the reasons for this problem is that KG nodes can be promiscuous, that is, associated with a number of different relationships that are not unique or indicative of the properties of the nodes. Pei-Yu Hou, Daniel R. Korn, Cleber C. Melo-Filho, David R. Wright 0001, Alexander Tropsha, Rada Chirkova |
SIGMOD Conference | 2 |
| 2021 | COVID-KOP: integrating emerging COVID-19 data with the ROBOKOP databaseabstractSUMMARY: In response to the COVID-19 pandemic, we established COVID-KOP, a new knowledgebase integrating the existing Reasoning Over Biomedical Objects linked in Knowledge Oriented Pathways (ROBOKOP) biomedical knowledge graph with information from recent biomedical literature on COVID-19 annotated in the CORD-19 collection. COVID-KOP can be used effectively to generate new hypotheses concerning repurposing of known drugs and clinical drug candidates against COVID-19 by establishing respective confirmatory pathways of drug action. AVAILABILITY AND IMPLEMENTATION: COVID-KOP is freely accessible at https://covidkop.renci.org/. For code and instructions for the original ROBOKOP, see: https://github.com/NCATS-Gamma/robokop. Daniel R. Korn, Tesia M. Bobrowski, Michael Li, Yaphet Kebede, Patrick Wang 0003, Phillips Owen, Gaurav Vaidya, Eugene N. Muratov, Rada Chirkova, Chris Bizon, Alexander Tropsha |
Bioinform. | 1 |