VLDB 2026 Research / reviewers in the wild / expert
Daniel R. Harris
dblp:150/7640
· DBLP profile ↗
14ranked-venue papers
8as first author
4since 2021 · last 2023
0000-0001-9139-3433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)abstractDespite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts. Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | geoPIPE: Geospatial Pipeline for Enhancing Open Data for Substance Use Disorders Research
Daniel R. Harris, Nick Anthony, Mojde Mir, Chris Delcher |
AMIA | 1 |
| 2021 | Generalizing Dynamic Dashboards for Community-Engaged Interventions
Daniel R. Harris, Daniel Redmond, Jeffery C. Talbert |
AMIA | 1 |
| 2021 | Integrating Medical and Dental Histories for Translational Science
Darren W. Henderson, Malini S. Kirakodu, James R. Aaron, Tamela J. Harper, Daniel R. Harris, Jeffery C. Talbert, Luciana M. Shaddox |
AMIA | 5 |
| 2020 | Improving the Utility of Tobacco-Related Problem List Entries Using Natural Language Processing
Daniel R. Harris, Darren W. Henderson, Alexandria Corbeau |
AMIA | 1 |
| 2020 | Leveraging Differential Privacy in Geospatial Analyses of Standardized Healthcare DataabstractWe present a collection of geodatabase functions which expedite utilizing differential privacy for privacy-aware geospatial analysis of healthcare data. The healthcare domain has a long history of standardization and research communities have developed open-source common data models to support the larger goals of interoperability, reproducibility, and data sharing; these models also standardize geospatial patient data. However, patient privacy laws and institutional regulations complicate geospatial analyses and dissemination of research findings due to protective restrictions in how data and results are shared. This results in infrastructures with great abilities to organize and store healthcare data, yet which lack the innate ability to produce shareable results that preserve privacy and conform to regulatory requirements. Differential privacy is a model for performing privacy-preserving analytics. We detail our process and findings in inserting an open-source library for differential privacy into a workflow for leveraging a geodatabase for geocoding and analyzing geospatial data stored as part of the Observational Medical Outcomes Partnership (OMOP) common data model. We pilot this process using an open big data repository of addresses. Daniel R. Harris |
IEEE BigData | 1 |
| 2020 | Challenges and Barriers in Applying Natural Language Processing to Medical Examiner Notes from Fatal Opioid Poisoning CasesabstractWe detail the challenges and barriers in applying natural language processing techniques to a collection of medical examiner case investigation notes related to fatal opioid poisonings. Major advances in biomedical informatics have made natural language processing (NLP) of medical texts both a realistic and useful task. Biomedical NLP tools are typically designed to process documents originating from biomedical libraries or electronic health records (EHRs). The usefulness of biomedical NLP tools on texts authored outside of EHRs is unclear, despite an abundance of medicolegal documents existing at the intersection of medicine and law. In particular, we detail our experiences processing unstructured text and extracting semantic concepts using case investigation notes; these notes were authored by trained investigative professionals working in a medical examiner's office and describe cases containing deaths related to fatal opioid poisonings. Applying NLP to case notes is a particularly important step in generalizing the advances of biomedical NLP for other related domains and giving guidance to data scientists working with unstructured data generated outside of EHRs. Daniel R. Harris, Christian Eisinger, Yanning Wang, Chris Delcher |
IEEE BigData | 1 |
| 2019 | bench4gis: Benchmarking Privacy-aware Geocoding with Open Big DataabstractGeocoding, the process of translating addresses to geographic coordinates, is a relatively straight-forward and well-studied process, but limitations due to privacy concerns may restrict usage of geographic data. The impact of these limitations are further compounded by the scale of the data, and in turn, also limits viable geocoding strategies. For example, healthcare data is protected by patient privacy laws in addition to possible institutional regulations that restrict external transmission and sharing of data. This results in the implementation of "in-house" geocoding solutions where data is processed behind an organization's firewall; quality assurance for these implementations is problematic because sensitive data cannot be used to externally validate results. In this paper, we present our software framework called bench4gis which benchmarks privacy-aware geocoding solutions by leveraging open big data as surrogate data for quality assurance; the scale of open big data sets for address data can ensure that results are geographically meaningful for the locale of the implementing institution. Daniel R. Harris, Chris Delcher |
IEEE BigData | 1 |
| 2018 | Retrospective analysis of health claims to evaluate pharmacotherapies with potential for repurposing: Association of bupropion and stimulant use disorder remission
Emily R. Hankosky, Heather Bush, Linda P. Dwoskin, Daniel R. Harris, Darren W. Henderson, Guo-Qiang Zhang 0001, Patricia R. Freeman, Jeffery C. Talbert |
AMIA | 4 |
| 2017 | Proceedings of the 16th Annual UT-KBRIN Bioinformatics Summit 2016: bioinformatics: Burns, TN, USA. April 21-23, 2017abstractMemphis, Tennessee Eric C. Rouchka, Julia H. Chariker, David Tieri, Juw Won Park, Shreedharkumar D. Rajurkar, Nishchal K. Verma, Yan Cui 0001, Mark L. Farman, Bradford Condon, Neil Moore, Jerzy W. Jaromczyk, Jolanta Jaromczyk, Daniel R. Harris, Patrick Calie, Eun Kyong Shin, Robert L. Davis, Arash Shaban-Nejad, Joshua M. Mitchell, Robert M. Flight, Qing Jun Wang, Richard M. Higashi, Teresa W.-M. Fan, Andrew N. Lane, Hunter N. B. Moseley, Liangqun Lu, Bernie J. Daigle, Andrey Smelter, Bailey K. Phan, Nathaniel J. Serpico, Ethan G. Toney, Caroline E. Melton, Jennifer R. Mandel, Bernie J. Daigle Jr., Kazi I. Zaman, Ramin Homayouni, Patrick J. Trainor, Samantha M. Carlisle, Andrew P. DeFilippis, Shesh N. Rai |
BMC Bioinform. | 14 |
| 2014 | An adaptive landscape for training in the essentials of next gen sequencing data acquisition and bioinformatic analysisabstractBackground Recent technological advances in Next Generation Sequencing (NGS) have reduced both the cost and time required to produce Large Data Sets (LDS) of nucleotide sequences. These advances have led to an exponential proliferation of nucleotide sequence data coupled with an exacerbation of a persistent conundrum: the level of difficulty in generating LDS is rapidly decreasing, but the exposure, development and training of students and investigators in the bioinformatic approaches requisite to the proper and correct analysis of such data sets is experiencing a parallel increase in difficulty. Mark L. Farman, Patrick Calie, Jerzy W. Jaromczyk, Jolanta Jaromczyk, Neil Moore, Daniel R. Harris, Christopher L. Schardl |
BMC Bioinform. | 6 |
| 2014 | Using Common Table Expressions to Build a Scalable Boolean Query Generator for Clinical Data WarehousesabstractWe present a custom, Boolean query generator utilizing common-table expressions (CTEs) that is capable of scaling with big datasets. The generator maps user-defined Boolean queries, such as those interactively created in clinical-research and general-purpose healthcare tools, into SQL. We demonstrate the effectiveness of this generator by integrating our study into the Informatics for Integrating Biology and the Bedside (i2b2) query tool and show that it is capable of scaling. Our custom generator replaces and outperforms the default query generator found within the Clinical Research Chart cell of i2b2. In our experiments, 16 different types of i2b2 queries were identified by varying four constraints: date, frequency, exclusion criteria, and whether selected concepts occurred in the same encounter. We generated nontrivial, random Boolean queries based on these 16 types; the corresponding SQL queries produced by both generators were compared by execution times. The CTE-based solution significantly outperformed the default query generator and provided a much more consistent response time across all query types (M = 2.03, SD = 6.64 versus M = 75.82, SD = 238.88 s). Without costly hardware upgrades, we provide a scalable solution based on CTEs with very promising empirical results centered on performance gains. The evaluation methodology used for this provides a means of profiling clinical data warehouse performance. Daniel R. Harris, Darren W. Henderson, Ramakanth Kavuluru, Arnold J. Stromberg, Todd R. Johnson |
IEEE J. Biomed. Health Informatics | 1 |
| 2012 | Improving Scalability and Performance of i2b2 Query Processing Using Common Table Expressions
Darren W. Henderson, Daniel R. Harris, Ramakanth Kavuluru, Todd R. Johnson |
AMIA | 2 |
| 2010 | Experimenting with database segmentation size vs time performance for mpiBLAST on an IBM HS21 blade clusterabstractFigure 1 CPU-time and wait-time composite.Figure 1 shows the summation of CPU-time (blue) and queue wait-time (red) in minutes as the number of nodes and database segments increase. Daniel R. Harris, Jerzy W. Jaromczyk, Christopher L. Schardl |
BMC Bioinform. | 1 |