VLDB 2026 Research / reviewers in the wild / expert
William R. Hogan
dblp:33/3529
· DBLP profile ↗
56ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-9881-1017ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 54 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysisabstractOBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases. Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor |
J. Am. Medical Informatics Assoc. | 90 |
| 2026 | Opportunities for informatics to improve patient experiences: observations and reflections of ACMI fellowsabstractOBJECTIVES: We report on findings from a meeting convened by the American College of Medical Informatics (ACMI) to characterize aspects of the patient experience that could be improved using informatics. MATERIALS AND METHODS: The American College of Medical Informatics fellows were invited to share their experiences as patients and suggest informatics approaches that may improve the patient experience. RESULTS: We identified 4 themes: (1) getting the right care, (2) data sharing and data interoperability, (3) guiding low-cost evaluations, and (4) predictive analytics. DISCUSSION: Despite widespread adoption of health IT, patient experiences remain far from optimal. CONCLUSION: The American College of Medical Informatics fellows identified informatics approaches, applications, and research areas that have the potential to improve patient experiences with health care systems. Howard R. Strasberg, Edward P. Hoffer, Ross Koppel, Kevin B. Johnson, William M. Tierney, Geoffrey W. Rutledge, Elmer V. Bernstam, Jos Aarts, Marion J. Ball, Douglas S. Bell, Bernd Blobel, Suzanne Boren, Iain E. Buchan, James J. Cimino, Lawrence M. Fagan, James Geller, María Adela Grando, David A. Hanauer, William R. Hogan, Andrew S. Kanter, Bonnie Kaplan, Casimir A. Kulikowski, Albert Lai, David McCallie, Vimla Patel, Wanda Pratt, Sarah Collins Rossetti, Edward H. Shortliffe, Hardeep Singh 0005, Dean F. Sittig, William W. Stead, Kim M. Unertl, Mark G. Weiner, Kai Zheng 0002 |
J. Am. Medical Informatics Assoc. | 19 |
| 2024 | Towards Machine-FAIR: Representing software and datasets to facilitate reuse and scientific discovery by machinesabstractOBJECTIVE: To use software, datasets, and data formats in the domain of Infectious Disease Epidemiology as a test collection to evaluate a novel M1 use case, which we introduce in this paper. M1 is a machine that upon receipt of a new digital object of research exhaustively finds all valid compositions of it with existing objects. METHOD: on the test collection and used error analysis to identify needed semantic constraints. RESULTS: search was 61.7%. Error analysis identified needed semantic constraints and needed changes in handling of data services. Most semantic constraints were simple, but one data format was sufficiently complex to be practically impossible to represent semantic constraints over, from which we conclude limitatively that software developers will have to meet the machines halfway by engineering software whose inputs are sufficiently simple that their semantic constraints can be represented, akin to the simple APIs of services. We summarize these insights as M1-FAIR guiding principles for composability and suggest a roadmap for progressively capable devices in the service of reuse and accelerated scientific discovery. CONCLUSION: Algorithmic search of digital repositories for valid workflow compositions has potential to accelerate scientific discovery but requires a scalable solution to the problem of knowledge acquisition about semantic constraints on software inputs. Additionally, practical limitations on the logical complexity of semantic constraints must be respected, which has implications for the design of software. Michael M. Wagner 0001, William R. Hogan, John D. Levander, Matthew Diller |
J. Biomed. Informatics | 2 |
| 2024 | Identifying social determinants of health from clinical narratives: A study of performance, documentation ratio, and potential bias
Zehao Yu 0001, Cheng Peng 0009, Xi Yang 0015, Chong Dang, Prakash Adekkanattu, Braja Gopal Patra, Yifan Peng 0002, Jyotishman Pathak, Debbie L. Wilson, Ching-Yuan Chang, Wei-Hsuan Lo-Ciganic, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
J. Biomed. Informatics | 13 |
| 2023 | The role of health system penetration rate in estimating the prevalence of type 1 diabetes in children and adolescents using electronic health recordsabstractOBJECTIVE: Having sufficient population coverage from the electronic health records (EHRs)-connected health system is essential for building a comprehensive EHR-based diabetes surveillance system. This study aimed to establish an EHR-based type 1 diabetes (T1D) surveillance system for children and adolescents across racial and ethnic groups by identifying the minimum population coverage from EHR-connected health systems to accurately estimate T1D prevalence. MATERIALS AND METHODS: We conducted a retrospective, cross-sectional analysis involving children and adolescents <20 years old identified from the OneFlorida+ Clinical Research Network (2018-2020). T1D cases were identified using a previously validated computable phenotyping algorithm. The T1D prevalence for each ZIP Code Tabulation Area (ZCTA, 5 digits), defined as the number of T1D cases divided by the total number of residents in the corresponding ZCTA, was calculated. Population coverage for each ZCTA was measured using observed health system penetration rates (HSPR), which was calculated as the ratio of residents in the corresponding ZTCA and captured by OneFlorida+ to the overall population in the same ZCTA reported by the Census. We used a recursive partitioning algorithm to identify the minimum required observed HSPR to estimate T1D prevalence and compare our estimate with the reported T1D prevalence from the SEARCH study. RESULTS: Observed HSPRs of 55%, 55%, and 60% were identified as the minimum thresholds for the non-Hispanic White, non-Hispanic Black, and Hispanic populations. The estimated T1D prevalence for non-Hispanic White and non-Hispanic Black were 2.87 and 2.29 per 1000 youth, which are comparable to the reference study's estimation. The estimated prevalence of T1D for Hispanics (2.76 per 1000 youth) was higher than the reference study's estimation (1.48-1.64 per 1000 youth). The standardized T1D prevalence in the overall Florida population was 2.81 per 1000 youth in 2019. CONCLUSION: Our study provides a method to estimate T1D prevalence in children and adolescents using EHRs and reports the estimated HSPRs and prevalence of T1D for different race and ethnicity groups to facilitate EHR-based diabetes surveillance. Piaopiao Li, Tianchen Lyu, Khalid Alkhuzam, Eliot Spector, William T. Donahoo, Sarah Bost, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Desmond A. Schatz, Mark A. Atkinson, Michael J. Haller, Elizabeth Shenkman, Yi Guo 0005, Jiang Bian 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | Clinical concept and relation extraction using prompt-based machine reading comprehensionabstractOBJECTIVE: To develop a natural language processing system that solves both clinical concept extraction and relation extraction in a unified prompt-based machine reading comprehension (MRC) architecture with good generalizability for cross-institution applications. METHODS: We formulate both clinical concept extraction and relation extraction using a unified prompt-based MRC architecture and explore state-of-the-art transformer models. We compare our MRC models with existing deep learning models for concept extraction and end-to-end relation extraction using 2 benchmark datasets developed by the 2018 National NLP Clinical Challenges (n2c2) challenge (medications and adverse drug events) and the 2022 n2c2 challenge (relations of social determinants of health [SDoH]). We also evaluate the transfer learning ability of the proposed MRC models in a cross-institution setting. We perform error analyses and examine how different prompting strategies affect the performance of MRC models. RESULTS AND CONCLUSION: The proposed MRC models achieve state-of-the-art performance for clinical concept and relation extraction on the 2 benchmark datasets, outperforming previous non-MRC transformer models. GatorTron-MRC achieves the best strict and lenient F1-scores for concept extraction, outperforming previous deep learning models on the 2 datasets by 1%-3% and 0.7%-1.3%, respectively. For end-to-end relation extraction, GatorTron-MRC and BERT-MIMIC-MRC achieve the best F1-scores, outperforming previous deep learning models by 0.9%-2.4% and 10%-11%, respectively. For cross-institution evaluation, GatorTron-MRC outperforms traditional GatorTron by 6.4% and 16% for the 2 datasets, respectively. The proposed method is better at handling nested/overlapped concepts, extracting relations, and has good portability for cross-institute applications. Our clinical MRC package is publicly available at https://github.com/uf-hobi-informatics-lab/ClinicalTransformerMRC. Cheng Peng 0009, Xi Yang 0015, Zehao Yu 0001, Jiang Bian 0001, William R. Hogan, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Impacts of Eligibility Criteria on Trial Participants' Age in Alzheimer's Disease Clinical Trials
Aokun Chen, Qian Li 0034, Xing He 0003, Michael Jaffee, William R. Hogan, Fei Wang 0001, Yi Guo 0005, Jiang Bian 0001 |
AMIA | 5 |
| 2022 | The OneFlorida Data Trust: a centralized, translational research data infrastructure of statewide scopeabstractThe OneFlorida Data Trust is a centralized research patient data repository created and managed by the OneFlorida Clinical Research Consortium ("OneFlorida"). It comprises structured electronic health record (EHR), administrative claims, tumor registry, death, and other data on 17.2 million individuals who received healthcare in Florida between January 2012 and the present. Ten healthcare systems in Miami, Orlando, Tampa, Jacksonville, Tallahassee, Gainesville, and rural areas of Florida contribute EHR data, covering the major metropolitan regions in Florida. Deduplication of patients is accomplished via privacy-preserving entity resolution (precision 0.97-0.99, recall 0.75), thereby linking patients' EHR, claims, and death data. Another unique feature is the establishment of mother-baby relationships via Florida vital statistics data. Research usage has been significant, including major studies launched in the National Patient-Centered Clinical Research Network ("PCORnet"), where OneFlorida is 1 of 9 clinical research networks. The Data Trust's robust, centralized, statewide data are a valuable and relatively unique research resource. William R. Hogan, Elizabeth Shenkman, Temple Robinson, Olveen Carasquillo, Patricia S. Robinson, Rebecca Z. Essner, Jiang Bian 0001, Gigi Lipori, Christopher A. Harle, Tanja Magoc, Lizabeth Manini, Tona Mendoza, Sonya White, Alex Loiacono, Jackie Hall, Dave Nelson |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | A Study of Social and Behavioral Determinants of Health in Lung Cancer Patients Using Transformers-based Natural Language Processing Models
Zehao Yu 0001, Xi Yang 0015, Chong Dang, Songzi Wu, Prakash Adekkanattu, Jyotishman Pathak, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001 |
AMIA | 8 |
| 2021 | Developing an Ontology for Social and Behavioral Determinants of Health
Hansi Zhang, Xi Yang 0015, Thomas J. George, William R. Hogan, Jiang Bian 0001, Yonghui Wu 0001 |
AMIA | 4 |
| 2021 | Extracting social determinants of health from electronic health records using natural language processing: a systematic reviewabstractOBJECTIVE: Social determinants of health (SDoH) are nonclinical dispositions that impact patient health risks and clinical outcomes. Leveraging SDoH in clinical decision-making can potentially improve diagnosis, treatment planning, and patient outcomes. Despite increased interest in capturing SDoH in electronic health records (EHRs), such information is typically locked in unstructured clinical notes. Natural language processing (NLP) is the key technology to extract SDoH information from clinical text and expand its utility in patient care and research. This article presents a systematic review of the state-of-the-art NLP approaches and tools that focus on identifying and extracting SDoH data from unstructured clinical text in EHRs. MATERIALS AND METHODS: A broad literature search was conducted in February 2021 using 3 scholarly databases (ACL Anthology, PubMed, and Scopus) following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 6402 publications were initially identified, and after applying the study inclusion criteria, 82 publications were selected for the final review. RESULTS: Smoking status (n = 27), substance use (n = 21), homelessness (n = 20), and alcohol use (n = 15) are the most frequently studied SDoH categories. Homelessness (n = 7) and other less-studied SDoH (eg, education, financial problems, social isolation and support, family problems) are mostly identified using rule-based approaches. In contrast, machine learning approaches are popular for identifying smoking status (n = 13), substance use (n = 9), and alcohol use (n = 9). CONCLUSION: NLP offers significant potential to extract SDoH data from narrative clinical notes, which in turn can aid in the development of screening tools, risk prediction models, and clinical decision support systems. Braja Gopal Patra, Mohit Manoj Sharma, Veer Vekaria, Prakash Adekkanattu, Olga V. Patterson, Benjamin S. Glicksberg, Lauren A. Lepow, Euijung Ryu, Joanna M. Biernacka, Al'ona Furmanchuk, Thomas J. George, William R. Hogan, Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Myrna Weissman, Priya Wickramaratne, J. John Mann, Mark Olfson, Thomas R. Campion Jr., Mark G. Weiner, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 12 |
| 2020 | Developing and Validating a Computable Phenotype for the Identification of Transgender and Gender Nonconforming Individuals and Subgroups
Yi Guo 0005, Xing He 0003, Tianchen Lyu, Hansi Zhang, Yonghui Wu 0001, Xi Yang 0015, Zhaoyi Chen, Merry J. Markham, François Modave, Mengjun Xie, William R. Hogan, Christopher A. Harle, Elizabeth Shenkman, Jiang Bian 0001 |
AMIA | 11 |
| 2020 | Assessing the practice of data quality evaluation in a national clinical data research network through a systematic scoping review in the era of real-world dataabstractOBJECTIVE: To synthesize data quality (DQ) dimensions and assessment methods of real-world data, especially electronic health records, through a systematic scoping review and to assess the practice of DQ assessment in the national Patient-centered Clinical Research Network (PCORnet). MATERIALS AND METHODS: We started with 3 widely cited DQ literature-2 reviews from Chan et al (2010) and Weiskopf et al (2013a) and 1 DQ framework from Kahn et al (2016)-and expanded our review systematically to cover relevant articles published up to February 2020. We extracted DQ dimensions and assessment methods from these studies, mapped their relationships, and organized a synthesized summarization of existing DQ dimensions and assessment methods. We reviewed the data checks employed by the PCORnet and mapped them to the synthesized DQ dimensions and methods. RESULTS: We analyzed a total of 3 reviews, 20 DQ frameworks, and 226 DQ studies and extracted 14 DQ dimensions and 10 assessment methods. We found that completeness, concordance, and correctness/accuracy were commonly assessed. Element presence, validity check, and conformance were commonly used DQ assessment methods and were the main focuses of the PCORnet data checks. DISCUSSION: Definitions of DQ dimensions and methods were not consistent in the literature, and the DQ assessment practice was not evenly distributed (eg, usability and ease-of-use were rarely discussed). Challenges in DQ assessments, given the complex and heterogeneous nature of real-world data, exist. CONCLUSION: The practice of DQ assessment is still limited in scope. Future work is warranted to generate understandable, executable, and reusable DQ measures. Jiang Bian 0001, Tianchen Lyu, Alexander T. Loiacono, Tonatiuh Mendoza Viramontes, Gloria P. Lipori, Yi Guo 0005, Yonghui Wu 0001, Mattia Prosperi, Thomas J. George, Christopher A. Harle, Elizabeth Shenkman, William R. Hogan |
J. Am. Medical Informatics Assoc. | 12 |
| 2020 | Identifying relations of medications with adverse drug events using recurrent convolutional neural networks and gradient boostingabstractOBJECTIVE: To develop a natural language processing system that identifies relations of medications with adverse drug events from clinical narratives. This project is part of the 2018 n2c2 challenge. MATERIALS AND METHODS: We developed a novel clinical named entity recognition method based on an recurrent convolutional neural network and compared it to a recurrent neural network implemented using the long-short term memory architecture, explored methods to integrate medical knowledge as embedding layers in neural networks, and investigated 3 machine learning models, including support vector machines, random forests and gradient boosting for relation classification. The performance of our system was evaluated using annotated data and scripts provided by the 2018 n2c2 organizers. RESULTS: Our system was among the top ranked. Our best model submitted during this challenge (based on recurrent neural networks and support vector machines) achieved lenient F1 scores of 0.9287 for concept extraction (ranked third), 0.9459 for relation classification (ranked fourth), and 0.8778 for the end-to-end relation extraction (ranked second). We developed a novel named entity recognition model based on a recurrent convolutional neural network and further investigated gradient boosting for relation classification. The new methods improved the lenient F1 scores of the 3 subtasks to 0.9292, 0.9633, and 0.8880, respectively, which are comparable to the best performance reported in this challenge. CONCLUSION: This study demonstrated the feasibility of using machine learning methods to extract the relations of medications with adverse drug events from clinical narratives. Xi Yang 0015, Jiang Bian 0001, Ruogu Fang, Ragnhildur I. Bjarnadottir, William R. Hogan, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Clinical concept extraction using transformersabstractOBJECTIVE: The goal of this study is to explore transformer-based models (eg, Bidirectional Encoder Representations from Transformers [BERT]) for clinical concept extraction and develop an open-source package with pretrained clinical models to facilitate concept extraction and other downstream natural language processing (NLP) tasks in the medical domain. METHODS: We systematically explored 4 widely used transformer-based architectures, including BERT, RoBERTa, ALBERT, and ELECTRA, for extracting various types of clinical concepts using 3 public datasets from the 2010 and 2012 i2b2 challenges and the 2018 n2c2 challenge. We examined general transformer models pretrained using general English corpora as well as clinical transformer models pretrained using a clinical corpus and compared them with a long short-term memory conditional random fields (LSTM-CRFs) mode as a baseline. Furthermore, we integrated the 4 clinical transformer-based models into an open-source package. RESULTS AND CONCLUSION: The RoBERTa-MIMIC model achieved state-of-the-art performance on 3 public clinical concept extraction datasets with F1-scores of 0.8994, 0.8053, and 0.8907, respectively. Compared to the baseline LSTM-CRFs model, RoBERTa-MIMIC remarkably improved the F1-score by approximately 4% and 6% on the 2010 and 2012 i2b2 datasets. This study demonstrated the efficiency of transformer-based models for clinical concept extraction. Our methods and systems can be applied to other clinical tasks. The clinical transformer package with 4 pretrained clinical models is publicly available at https://github.com/uf-hobi-informatics-lab/ClinicalTransformerNER. We believe this package will improve current practice on clinical concept extraction and other tasks in the medical domain. Xi Yang 0015, Jiang Bian 0001, William R. Hogan, Yonghui Wu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Assessing the Validity of a a priori Patient-Trial Generalizability Score using Real-world Data from a Large Clinical Data Research Network: A Colorectal Cancer Clinical Trial Case Study
Qian Li 0034, Zhe He 0001, Yi Guo 0005, Hansi Zhang, Thomas J. George, William R. Hogan, Neil Charness, Jiang Bian 0001 |
AMIA | 6 |
| 2019 | Identifying Cancer Patients at Risk for Heart Failure Using Machine Learning Methods
Xi Yang 0015, Nida Waheed, Keith March, Jiang Bian 0001, William R. Hogan, Yonghui Wu 0001 |
AMIA | 6 |
| 2019 | Enhancing the drug ontology with semantically-rich representations of National Drug Codes and RxNorm unique concept identifiersabstractBACKGROUND: The Drug Ontology (DrOn) is a modular, extensible ontology of drug products, their ingredients, and their biological activity created to enable comparative effectiveness and health services researchers to query National Drug Codes (NDCs) that represent products by ingredient, by molecular disposition, by therapeutic disposition, and by physiological effect (e.g., diuretic). It is based on the RxNorm drug terminology maintained by the U.S. National Library of Medicine, and on the Chemical Entities of Biological Interest ontology. Both national drug codes (NDCs) and RxNorm unique concept identifiers (RXCUIS) can undergo changes over time that can obfuscate their meaning when these identifiers occur in historic data. We present a new approach to modeling these entities within DrOn that will allow users of DrOn working with historic prescription data to more easily and correctly interpret that data. RESULTS: We have implemented a full accounting of national drug codes and RxNorm unique concept identifiers as information content entities, and of the processes involved in managing their creation and changes. This includes an OWL file that implements and defines the classes necessary to model these entities. A separate file contains an instance-level prototype in OWL that demonstrates the feasibility of this approach to representing NDCs and RXCUIs and the processes of managing them by retrieving and representing several individual NDCs, both active and inactive, and the RXCUIs to which they are connected. We also demonstrate how historic information about these identifiers in DrOn can be easily retrieved using a simple SPARQL query. CONCLUSIONS: An accurate model of how these identifiers operate in reality is a valuable addition to DrOn that enhances its usefulness as a knowledge management resource for working with historic data. Jonathan P. Bona, Mathias Brochhausen, William R. Hogan |
BMC Bioinform. | 3 |
| 2018 | Combine Factual Medical Knowledge and Distributed Word Representation to Improve Clinical Named Entity Recognition
Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Yi Guo 0005, Hua Xu 0001, William R. Hogan |
AMIA | 6 |
| 2018 | Computable Eligibility Criteria through Ontology-driven Data Access: A Case Study of Hepatitis C Virus Trials
Hansi Zhang, Zhe He 0001, Xing He 0003, Yi Guo 0005, David R. Nelson, François Modave, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Jiang Bian 0001 |
AMIA | 8 |
| 2017 | Comparing and Contrasting A Priori and A Posteriori Generalizability Assessment of Clinical Trials on Type 2 Diabetes Mellitus
Zhe He 0001, Arturo Gonzalez-Izquierdo, Spiros C. Denaxas, Andrei Sura, Yi Guo 0005, William R. Hogan, Elizabeth Shenkman, Jiang Bian 0001 |
AMIA | 6 |
| 2017 | Implementing a Hash-based Privacy-Preserving Entity Resolution Tool in the OneFlorida Clinical Data Research Network
Jiang Bian 0001, Andrei Sura, Gloria P. Lipori, Yi Guo 0005, François Modave, Zhe He 0001, Elizabeth Shenkman, William R. Hogan |
AMIA | 8 |
| 2017 | Building an Obesity and Cancer Semantic Web Knowledge Base
Juan Antonio Lossio-Ventura, William R. Hogan, François Modave, Amanda Hicks, Yi Guo 0005, Zhe He 0001, Mirela Vasconcelos, Jiang Bian 0001 |
AMIA | 2 |
| 2017 | OC-2-KB: A software pipeline to build an evidence-based obesity and cancer knowledge baseabstractObesity has been linked to several types of cancer. Access to adequate health information activates people's participation in managing their own health, which ultimately improves their health outcomes. Nevertheless, the existing online information about the relationship between obesity and cancer is heterogeneous and poorly organized. A formal knowledge representation can help better organize and deliver quality health information. Currently, there are several efforts in the biomedical domain to convert unstructured data to structured data and store them in Semantic Web knowledge bases (KB). In this demo paper, we present, OC-2-KB (Obesity and Cancer to Knowledge Base), a system that is tailored to guide the automatic KB construction for managing obesity and cancer knowledge from free-text scientific literature (i.e., PubMed abstracts) in a systematic way. OC-2-KB has two important modules which perform the acquisition of entities and the extraction then classification of relationships among these entities. We tested the OC-2-KB system on a data set with 23 manually annotated obesity and cancer PubMed abstracts and created a preliminary KB with 765 triples. We conducted a preliminary evaluation on this sample of triples and reported our evaluation results. Juan Antonio Lossio-Ventura, William R. Hogan, François Modave, Yi Guo 0005, Zhe He 0001, Amanda Hicks, Jiang Bian 0001 |
BIBM | 2 |
| 2017 | Towards a privacy preserving cohort discovery framework for clinical research networks
Bradley A. Malin, François Modave, Yi Guo 0005, William R. Hogan, Elizabeth Shenkman, Jiang Bian 0001 |
J. Biomed. Informatics | 5 |
| 2016 | Towards an obesity-cancer knowledge base: Biomedical entity identification and relation detectionabstractObesity is associated with increased risks of various types of cancer, as well as a wide range of other chronic diseases. On the other hand, access to health information activates patient participation, and improve their health outcomes. However, existing online information on obesity and its relationship to cancer is heterogeneous ranging from pre-clinical models and case studies to mere hypothesis-based scientific arguments. A formal knowledge representation (i.e., a semantic knowledge base) would help better organizing and delivering quality health information related to obesity and cancer that consumers need. Nevertheless, current ontologies describing obesity, cancer and related entities are not designed to guide automatic knowledge base construction from heterogeneous information sources. Thus, in this paper, we present methods for named-entity recognition (NER) to extract biomedical entities from scholarly articles and for detecting if two biomedical entities are related, with the long term goal of building a obesity-cancer knowledge base. We leverage both linguistic and statistical approaches in the NER task, which supersedes the state-of-the-art results. Further, based on statistical features extracted from the sentences, our method for relation detection obtains an accuracy of 99.3% and a f-measure of 0.993. Juan Antonio Lossio-Ventura, William R. Hogan, François Modave, Amanda Hicks, Josh Hanna, Yi Guo 0005, Zhe He 0001, Jiang Bian 0001 |
BIBM | 2 |
| 2015 | Mining Twitter as a First Step toward Assessing the Adequacy of Gender Identification Terms on Intake Forms
Amanda Hicks, William R. Hogan, Michael W. Rutherford, Bradley A. Malin, Mengjun Xie, Christiane Fellbaum, Zhijun Yin, Daniel Fabbri, Josh Hanna, Jiang Bian 0001 |
AMIA | 2 |
| 2014 | Social network analysis of biomedical research collaboration networks in a CTSA institution
Jiang Bian 0001, Mengjun Xie, Umit Topaloglu, Teresa Hudson, Hari Eswaran, William R. Hogan |
J. Biomed. Informatics | 6 |
| 2013 | Apollo: Giving application developers a single point of access to public health models using structured vocabularies and Web services
Michael M. Wagner 0001, John D. Levander, Shawn T. Brown, William R. Hogan, Nicholas Millett, Josh Hanna |
AMIA | 4 |
| 2013 | Understanding biomedicai research collaborations through social network analysis: A case studyabstractA recent surge of research on social networks and their characteristics has attracted an increasing amount of interests from the community of biomedicine and biomedical informatics. Social network analysis (SNA) methods have been regarded as an effective tool to assess inter- and intra-institution research collaborations in the Clinical Translational Science Award (CTSA) community. In this paper, we present a case study of SNA on the research collaboration networks (RCNs) at the University of Arkansas for Medical Sciences (UAMS) - a CTSA institution. We have applied graph theoretical analyses to the RCNs prior to and after the CTSA award at UAMS. By virtue of quantitative measures, we have obtained valuable insights into the network dynamics and topological characteristics of the research environment. Moreover, through observing the temporal evolution of the RCNs at UAMS, we are able to demonstrate the effectiveness of the CTSA program and its important role in promoting trans-disciplinary collaborative research within an institution. Jiang Bian 0001, Mengjun Xie, Umit Topaloglu, Teresa Hudson, William R. Hogan |
BIBM | 5 |
| 2013 | Evidence of community structure in Biomedical Research Grant Collaborations
Radhakrishnan Nagarajan, Alex T. Kalinka, William R. Hogan |
J. Biomed. Informatics | 3 |
| 2011 | Towards an ontological theory of substance intolerance and hypersensitivity
William R. Hogan |
J. Biomed. Informatics | 1 |
| 2011 | Natural Language Processing methods and systems for biomedical ontology learning
Kaihong Liu, William R. Hogan, Rebecca S. Jacobson |
J. Biomed. Informatics | 2 |
| 2009 | Knowledge-based variable selection for learning rules from proteomic dataabstractBACKGROUND: The incorporation of biological knowledge can enhance the analysis of biomedical data. We present a novel method that uses a proteomic knowledge base to enhance the performance of a rule-learning algorithm in identifying putative biomarkers of disease from high-dimensional proteomic mass spectral data. In particular, we use the Empirical Proteomics Ontology Knowledge Base (EPO-KB) that contains previously identified and validated proteomic biomarkers to select m/zs in a proteomic dataset prior to analysis to increase performance. RESULTS: We show that using EPO-KB as a pre-processing method, specifically selecting all biomarkers found only in the biofluid of the proteomic dataset, reduces the dimensionality by 95% and provides a statistically significantly greater increase in performance over no variable selection and random variable selection. CONCLUSION: Knowledge-based variable selection even with a sparsely-populated resource such as the EPO-KB increases overall performance of rule-learning for disease classification from high-dimensional proteomic mass spectra. Jonathan L. Lustgarten, Shyam Visweswaran, Robert P. Bowser, William R. Hogan, Vanathi Gopalakrishnan |
BMC Bioinform. | 4 |
| 2009 | Mining aggregates of over-the-counter products for syndromic surveillance
Aurel Cami, Garrick L. Wallstrom, Ashley L. Fowlkes, Cathy A. Panozzo, William R. Hogan |
Pattern Recognit. Lett. | 5 |
| 2008 | EPO-KB: a searchable knowledge base of biomarker to protein linksabstractUNLABELLED: The knowledge base EPO-KB (Empirical Proteomic Ontology Knowledge Base) is based on an OWL ontology that represents current knowledge linking mass-to-charge (m/z) ratios to proteins on multiple platforms including Matrix Assisted Laser/Desorption Ionization (MALDI) and Surface Enhanced Laser/Desorption Ionization (SELDI)--Time of Flight (TOF). At present, it contains information on m/z ratio to protein links that were extracted from 120 published research papers. It has a web interface that allows researchers to query and retrieve putative proteins that correspond to a user-specified m/z ratio. EPO-KB also allows automated entry of additional m/z ratio to protein links and is expandable to the addition of gene to protein and protein to disease links. AVAILABILITY: http://www.dbmi.pitt.edu/EPO-KB Jonathan L. Lustgarten, Chad Kimmel, Henrik Ryberg, William R. Hogan |
Bioinform. | 4 |
| 2007 | Unsupervised clustering of over-the-counter healthcare products into product categories
Garrick L. Wallstrom, William R. Hogan |
J. Biomed. Informatics | 2 |
| 2005 | An Evaluation of Three Policies for Updating Product Categories in the National Retail Data Monitor
William R. Hogan, Garrick L. Wallstrom, Michael M. Wagner 0001 |
AMIA | 1 |
| 2005 | A Multivariate Procedure for Identifying Correlations between Diagnoses and Over-the-counter Products from Historical Datasets
Garrick L. Wallstrom, William R. Hogan |
AMIA | 3 |
| 2005 | Algorithms for rapid outbreak detection: a research synthesis
David L. Buckeridge, Howard S. Burkom, Murray Campbell, William R. Hogan, Andrew W. Moore 0001 |
J. Biomed. Informatics | 4 |
| 2004 | Bayesian Biosurveillance of Disease Outbreaks
Gregory F. Cooper, Denver Dash, John D. Levander, Weng-Keen Wong, William R. Hogan, Michael M. Wagner 0001 |
UAI | 5 |
| 2003 | Telephone Triage: A Timely Data Source for Surveillance of Influenza-like Diseases
Jeremy U. Espino, William R. Hogan, Michael M. Wagner 0001 |
AMIA | 2 |
| 2003 | Detection of Pediatric Respiratory and Gastrointestinal Outbreaks from Free-Text Chief Complaints
Oleg Ivanov, Per H. Gesteland, William R. Hogan, Michael B. Mundorff, Michael M. Wagner 0001 |
AMIA | 3 |
| 2003 | A Framework for Infection Control Surveillance Using Association Rules
Fu-Chiang Tsui, William R. Hogan, Michael M. Wagner 0001, Haobo Ma |
AMIA | 3 |
| 2003 | Detection of Outbreaks from Time Series Data Using Wavelet Transform
Fu-Chiang Tsui, Michael M. Wagner 0001, William R. Hogan |
AMIA | 4 |
| 2003 | Research Paper: Detection of Pediatric Respiratory and Diarrheal Outbreaks from Sales of Over-the-counter Electrolyte ProductsabstractOBJECTIVE: To determine whether sales of electrolyte products contain a signal of outbreaks of respiratory and diarrheal disease in children and, if so, how much earlier a signal relative to hospital diagnoses. DESIGN: Retrospective analysis was conducted of sales of electrolyte products and hospital diagnoses for six urban regions in three states for the period 1998 through 2001. MEASUREMENTS: Presence of signal was ascertained by measuring correlation between electrolyte sales and hospital diagnoses and the temporal relationship that maximized correlation. Earliness was the difference between the date that the exponentially weighted moving average (EWMA) method first detected an outbreak from sales and the date it first detected the outbreak from diagnoses. The coefficient of determination (r2) measured how much variance in earliness resulted from differences in sales' and diagnoses' signal strengths. RESULTS: The correlation between electrolyte sales and hospital diagnoses was 0.90 (95% CI, 0.87-0.93) at a time offset of 1.7 weeks (95% CI, 0.50-2.9), meaning that sales preceded diagnoses by 1.7 weeks. EWMA with a nine-sigma threshold detected the 18 outbreaks on average 2.4 weeks (95% CI, 0.1-4.8 weeks) earlier from sales than from diagnoses. Twelve outbreaks were first detected from sales, four were first detected from diagnoses, and two were detected simultaneously. Only 26% of variance in earliness was explained by the relative strength of the sales and diagnoses signals (r2 = 0.26). CONCLUSION: Sales of electrolyte products contain a signal of outbreaks of respiratory and diarrheal diseases in children and usually are an earlier signal than hospital diagnoses. William R. Hogan, Fu-Chiang Tsui, Oleg Ivanov, Per H. Gesteland, Shaun J. Grannis, J. Marc Overhage, J. Michael Robinson, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2003 | Application of Information Technology: Design of a National Retail Data Monitor for Public Health SurveillanceabstractThe National Retail Data Monitor receives data daily from 10,000 stores, including pharmacies, that sell health care products. These stores belong to national chains that process sales data centrally and utilize Universal Product Codes and scanners to collect sales information at the cash register. The high degree of retail sales data automation enables the monitor to collect information from thousands of store locations in near to real time for use in public health surveillance. The monitor provides user interfaces that display summary sales data on timelines and maps. Algorithms monitor the data automatically on a daily basis to detect unusual patterns of sales. The project provides the resulting data and analyses, free of charge, to health departments nationwide. Future plans include continued enrollment and support of health departments, developing methods to make the service financially self-supporting, and further refinement of the data collection system to reduce the time latency of data receipt and analysis. Michael M. Wagner 0001, J. Michael Robinson, Fu-Chiang Tsui, Jeremy U. Espino, William R. Hogan |
J. Am. Medical Informatics Assoc. | 5 |
| 2002 | Experience with Message Format and Code Set Standards for Early Warning Public Health Surveillance Systems
William R. Hogan, Michael M. Wagner 0001, Fu-Chiang Tsui |
AMIA | 1 |
| 1999 | The use of an explanation algorithm in a clinical event monitor
William R. Hogan, Michael M. Wagner 0001 |
AMIA | 1 |
| 1999 | A feasibility study of two methods for end-user configuration of a clinical event monitor
Fu-Chiang Tsui, Michael M. Wagner 0001, Wayne Wilbright, Aaron Tse, William R. Hogan |
AMIA | 5 |
| 1998 | Mobile workers in healthcare and their information needs: are 2-way pagers the answer?
Stuart A. Eisenstadt, Michael M. Wagner 0001, William R. Hogan, Marvin C. Pankaskie, Fu-Chiang Tsui, Wayne Wilbright |
AMIA | 3 |
| 1998 | Optimal use of communication channels in clinical event monitoring
William R. Hogan, Michael M. Wagner 0001 |
AMIA | 1 |
| 1998 | Preferences of interns and residents for E-mail, paging, or traditional methods for the delivery of different types of clinical information
Michael M. Wagner 0001, Stuart A. Eisenstadt, William R. Hogan, Marvin C. Pankaskie |
AMIA | 3 |
| 1997 | Clinical event monitoring at the University of Pittsburgh
Michael M. Wagner 0001, Marvin C. Pankaskie, William R. Hogan, Fu-Chiang Tsui, Stuart A. Eisenstadt, Eric Rodriguez, John K. Vries |
AMIA | 3 |
| 1997 | Review: Accuracy of Data in Computer-based Patient RecordsabstractData in computer-based patient records (CPRs) have many uses beyond their primary role in patient care, including research and health-system management. Although the accuracy of CPR data directly affects these applications, there has been only sporadic interest in, and no previous review of, data accuracy in CPRs. This paper reviews the published studies of data accuracy in CPRs. These studies report highly variable levels of accuracy. This variability stems from differences in study design, in types of data studied, and in the CPRs themselves. These differences confound interpretation of this literature. We conclude that our knowledge of data accuracy in CPRs is not commensurate with its importance and further studies are needed. We propose methodological guidelines for studying accuracy that address shortcomings of the current literature. As CPR data are used increasingly for research, methods used in research databases to continuously monitor and improve accuracy should be applied to CPRs. William R. Hogan, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 1996 | Research Paper: The Accuracy of Medication Data in an Outpatient Electronic Medical RecordabstractOBJECTIVE: To measure the accuracy of medication records stored in the electronic medical record (EMR) of an outpatient geriatric center. The authors analyzed accuracy from the perspective of a clinician using the data and the perspective of a computer-based medical decision-support system (MDSS). DESIGN: Prospective cohort study. METHODS: The EMR at the geriatric center captures medication data both directly from clinicians and indirectly using encounter forms and data-entry clerks. During a scheduled office visit for medical care, the treating clinician determined whether the medication records for the patient were an accurate representation of the medications that the patient was actually taking. Using the available sources of information (the patient, the patient's vials, any caregivers, and the medical chart), the clinician determined whether the recorded data were correct, whether any data were missing, and the type and cause for each discrepancy found. RESULTS: At the geriatric center, 83% of medication records represented correctly the compound. dose, and schedule of a current medication; 91% represented correctly the compound. 0.37 current medications were missing per patient. The principal cause of errors was the patient (36.1% of errors), who misreported a medication at a previous visit or changed (stopped, started, or dose-adjusted) a medication between visits. The second most frequent cause of errors was failure to capture changes to medications made by outside clinicians, accounting for 25.9% of errors. Transcription errors were a relatively ucommon cause (8.2% of errors). When the accuracy of records from the center was analyzed from the perspective of a MDSS, 90% were correct for compound identity and 1.38 medications were missing or uncoded per patient. The cause of the additional errors of omission was a free-text "comments" field-which it is assumed would be unreadable by current MDSS applications-that was used by clinicians in 18% of records to record the identity of the medication. CONCLUSIONS: Medication records in an outpatient EMR may have significant levels of data error. Based on an analysis of correctable causes of error, the authors conclude that the most effective extension to the EMR studied would be to expand its scope to include all clinicians who can potentially change medications. Even with EMR extensions, however, ineradicable error due to patients and data entry will remain. Several implications of ineradicable error for MDSSs are discussed. The provision of a free-text "comments" field increased the accuracy of medication lists for clinician users at the expense of accuracy for a MDSS. Michael M. Wagner 0001, William R. Hogan |
J. Am. Medical Informatics Assoc. | 2 |