EDBT 2026 Demo / reviewers in the wild / expert
Shaun J. Grannis
dblp:39/6644
· DBLP profile ↗
64ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-8093-6639ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 62 · 12 first-author · 9 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysisabstractOBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases. Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor |
J. Am. Medical Informatics Assoc. | 45 |
| 2026 | Derivation and validation of an algorithm for maternal-child linkage in electronic health recordsabstractINTRODUCTION: We created a probabilistic maternal-child electronic health record (EHR) linkage algorithm to promote clinical research in maternal-child health. METHODS: We used EHR data from 1994 to 2024 to create an XGBoost model to predict maternal-child linkages. The model used standard EHR elements as predictor variables, including first name, last name, birthdate, address, phone number, email, and an EHR-embedded maternal-child indicator as the deterministic outcome. RESULTS: From 82 million unique records, 6.2 billion potential pairs met blocking criteria. Of the potential pairs, 33 364 674 contained the deterministic indicator and were used as cases, and an equal number of controls were randomly sampled. The final model obtained an accuracy of 92%, a precision of 98%, a recall of 87%, and an F1-score of 92%. CONCLUSION: We derived and validated a probabilistic maternal-child linkage algorithm using routinely collected EHR data elements that could benefit future observational research in maternal-child health. Colin M. Rogerson, Christopher W. Bartlett, John P. Price, Lang Li 0001, Eneida A. Mendonça, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Linkability measures to assess the data characteristics for record linkageabstractOBJECTIVES: Accurate record linkage (RL) enables consolidation and de-duplication of data from disparate datasets, resulting in more comprehensive and complete patient data. However, conducting RL with low quality or unfit data can waste institutional resources on poor linkage results. We aim to evaluate data linkability to enhance the effectiveness of record linkage. MATERIALS AND METHODS: We describe a systematic approach using data fitness ("linkability") measures, defined as metrics that characterize the availability, discriminatory power, and distribution of potential variables for RL. We used the isolation forest algorithm to detect abnormal linkability values from 188 sites in Indiana and Colorado, and manually reviewed the data to understand the cause of anomalies. RESULT: We calculated 10 linkability metrics for 11 potential linkage variables (LVs) across 188 sites for a total of 20 680 linkability metrics. Potential LVs such as first name, last name, date of birth, and sex have low missing data rates, while Social Security Number vary widely in completeness among all sites. We investigated anomalous linkability values to identify the cause of many records having identical values in certain LVs, issues with placeholder values disguising data missingness, and orphan records. DISCUSSION: The fitness of a variable for RL is determined by its availability and its discriminatory power to uniquely identify individuals. These results highlight the need for awareness of placeholder values, which inform the selection of variables and methods to optimize RL performance. CONCLUSION: Evaluating linkability measures using the isolation forest algorithm to highlight anomalous findings can help identify fitness-for-use issues that must be addressed before initiating the RL process to ensure high-quality linkage outcomes. Toan Ong, Michael G. Kahn, Lauren R. Lembcke, Lisa M. Schilling, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Assessing the impact of privacy-preserving record linkage on record overlap and patient demographic and clinical characteristics in PCORnet®, the National Patient-Centered Clinical Research NetworkabstractOBJECTIVE: This article describes the implementation of a privacy-preserving record linkage (PPRL) solution across PCORnet®, the National Patient-Centered Clinical Research Network. MATERIAL AND METHODS: Using a PPRL solution from Datavant, we quantified the degree of patient overlap across the network and report a de-duplicated analysis of the demographic and clinical characteristics of the PCORnet population. RESULTS: There were ∼170M patient records across the responding Network Partners, with ∼138M (81%) of those corresponding to a unique patient. 82.1% of patients were found in a single partner and 14.7% were in 2. The percentage overlap between Partners ranged between 0% and 80% with a median of 0%. Linking patients' electronic health records with claims increased disease prevalence in every clinical characteristic, ranging between 63% and 173%. DISCUSSION: The overlap between Partners was variable and depended on timeframe. However, patient data linkage changed the prevalence profile of the PCORnet patient population. CONCLUSIONS: This project was one of the largest linkage efforts of its kind and demonstrates the potential value of record linkage. Linkage between Partners may be most useful in cases where there is geographic proximity between Partners, an expectation that potential linkage Partners will be able to fill gaps in data, or a longer study timeframe. Keith Marsolo, Daniel Kiernan, Sengwee Toh, Jasmin Phua, Darcy Louzao, Kevin Haynes, Mark G. Weiner, Francisco Angulo, L. Charles Bailey, Jiang Bian 0001, Daniel Fort, Shaun J. Grannis, Ashok K. Krishnamurthy 0001, Vinit Nair, Pedro Rivera, Jonathan C. Silverstein, Maryan Zirkle, Thomas Carton |
J. Am. Medical Informatics Assoc. | 12 |
| 2022 | Evaluation of real-world referential and probabilistic patient matching to advance patient identification strategyabstractOBJECTIVE: This study sought both to support evidence-based patient identity policy development by illustrating an approach for formally evaluating operational matching methods, and also to characterize the performance of both referential and probabilistic patient matching algorithms using real-world demographic data. MATERIALS AND METHODS: We assessed matching accuracy for referential and probabilistic matching algorithms using a manually reviewed 30 000 record gold standard reference dataset derived from a large health information exchange containing over 47 million patient registrations. We applied referential and probabilistic algorithms to this dataset and compared the outputs to the gold standard. We computed performance metrics including sensitivity (recall), positive predictive value (precision), and F-score for each algorithm. RESULTS: The probabilistic algorithm exhibited sensitivity, positive predictive value (PPV), and F-score of .6366, 0.9995, and 0.7778, respectively. The referential algorithm exhibited corresponding sensitivity, PPV, and F-score values of 0.9351, 0.9996, and 0.9663, respectively. Treating discordant and limited-data records as nonmatches increased referential match sensitivity to 0.9578. Compared to the more traditional probabilistic approach, referential matching exhibits greater accuracy. CONCLUSIONS: Referential patient matching, an increasingly popular method among health IT vendors, demonstrated notably greater accuracy than a more traditional probabilistic approach without the adaptation of the algorithm to the data that the traditional probabilistic approach usually requires. Health IT policymakers, including the Office of the National Coordinator for Health Information Technology (ONC), should explore strategies to expand the evidence base for real-world matching system performance, given the need for an evidence-based patient identity strategy. Shaun J. Grannis, Jennifer L. Williams, Suranga Nath Kasthurirathne, Molly Murray, Huiping Xu |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | A framework for a consistent and reproducible evaluation of manual review for patient matching algorithmsabstractHealthcare systems are hampered by incomplete and fragmented patient health records. Record linkage is widely accepted as a solution to improve the quality and completeness of patient records. However, there does not exist a systematic approach for manually reviewing patient records to create gold standard record linkage data sets. We propose a robust framework for creating and evaluating manually reviewed gold standard data sets for measuring the performance of patient matching algorithms. Our 8-point approach covers data preprocessing, blocking, record adjudication, linkage evaluation, and reviewer characteristics. This framework can help record linkage method developers provide necessary transparency when creating and validating gold standard reference matching data sets. In turn, this transparency will support both the internal and external validity of recording linkage studies and improve the robustness of new record linkage strategies. Agrayan K. Gupta, Suranga Nath Kasthurirathne, Huiping Xu, Xiaochun Li 0003, Matthew Ruppert, Christopher A. Harle, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Evaluation of Token Collections and Matching Models to Support Privacy-Preserving Record Linkage (PPRL)
Shaun J. Grannis, Abel N. Kho, Jasmin Phua, Suranga Nath Kasthurirathne |
AMIA | 1 |
| 2021 | Evolving Challenges in Patient Matching
Abel N. Kho, Shaun J. Grannis, Adam Culbertson, Molly Murray |
AMIA | 2 |
| 2021 | Leveraging data visualization and a statewide health information exchange to support COVID-19 surveillance and response: Application of public health informaticsabstractOBJECTIVE: We sought to support public health surveillance and response to coronavirus disease 2019 (COVID-19) through rapid development and implementation of novel visualization applications for data amalgamated across sectors. MATERIALS AND METHODS: We developed and implemented population-level dashboards that collate information on individuals tested for and infected with COVID-19, in partnership with state and local public health agencies as well as health systems. The dashboards are deployed on top of a statewide health information exchange. One dashboard enables authorized users working in public health agencies to surveil populations in detail, and a public version provides higher-level situational awareness to inform ongoing pandemic response efforts in communities. RESULTS: Both dashboards have proved useful informatics resources. For example, the private dashboard enabled detection of a local community outbreak associated with a meat packing plant. The public dashboard provides recent trend analysis to track disease spread and community-level hospitalizations. Combined, the tools were utilized 133 637 times by 74 317 distinct users between June 21 and August 22, 2020. The tools are frequently cited by journalists and featured on social media. DISCUSSION: Capitalizing on a statewide health information exchange, in partnership with health system and public health leaders, Regenstrief biomedical informatics experts rapidly developed and deployed informatics tools to support surveillance and response to COVID-19. CONCLUSIONS: The application of public health informatics methods and tools in Indiana holds promise for other states and nations. Yet, development of infrastructure and partnerships will require effort and investment after the current pandemic in preparation for the next public health emergency. Brian E. Dixon, Shaun J. Grannis, Connor McAndrews, Andrea A. Broyles, Waldo Mikels-Carrasco, Ashley Wiensch, Jennifer L. Williams, Umberto Tachinardi, Peter J. Embí |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Novel Application of Data Quality Metrics to Tailor Standardization of Patient Matching Fields
Shaun J. Grannis, Huiping Xu, Toan Ong, Michael G. Kahn, Lauren R. Lembcke, Suranga Nath Kasthurirathne |
AMIA | 1 |
| 2020 | Underrepresented racial minorities in biomedical informatics doctoral programs: graduation trends and academic placement (2002-2017)abstractOBJECTIVE: Biomedical informatics attracts few underrepresented racial minorities (URMs) into PhD programs. We examine graduation trends from 2002 to 2017 to determine how URM representation has changed over time. We also examine academic job placements by race and identify individual and institutional characteristics associated with URM graduates being successfully placed in academic jobs. MATERIALS AND METHODS: We analyze a near census of all research doctoral graduates from US-accredited institutions, surveyed at graduation by the National Science Foundation Survey of Earned Doctorates. Graduates of biomedical informatics-related programs were identified using self-reported primary and secondary disciplines. Data are analyzed using bivariate and multivariable logistic regressions. RESULTS: During the study period, 2426 individuals earned doctoral degrees in biomedical informatics-related disciplines. URM students comprised nearly 12% of graduates, and this proportion did not change over time (2002-2017). URMs included Hispanic (5.7%), Black (3.2%), and others, including multi-racial and indigenous American populations (2.8%). Overall, 82.3% of all graduates accepted academic positions at the time of graduation with significantly more Hispanic graduates electing to go into academia (89.2%; P < .001). URM graduates were more likely to be single (OR = 1.38; P < .05), have a dependent (1.95; P < .01), and not receive full tuition remission (OR = 1.37; P = .05) as a student. URM graduates accepting an academic position were less likely to be a graduate of a private institution (OR = 0.70; P < .05). DISCUSSION AND CONCLUSION: The proportion of URM candidates among biomedical informatics doctoral graduates has not increased over time and remains low. In order to improve URM recruitment and retention within academia, leaders in biomedical informatics should replicate strategies used to improve URM graduation rates in other fields. Kevin K. Wiley, Brian E. Dixon, Shaun J. Grannis, Nir Menachemi |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Comparison of Free-Text Synthetic Data Produced by Three Generative Adversarial Networks for Collaborative Health Data Analytics
Gregory Dexter, Shaun J. Grannis, Suranga Nath Kasthurirathne |
AMIA | 2 |
| 2019 | Local and State Public Health Informatics Workforce Skills and Needs: A Descriptive Analysis using the PH WINS
Brian E. Dixon, Timothy D. McFarlane, Shaun J. Grannis, P. Joseph Gibson |
AMIA | 3 |
| 2019 | An information infrastructure for federated, person-level linkage and query capability across private- and public-sector health data: The Indiana state EMS to HIE ED pilot project
Daniel Hood, Shaun J. Grannis, Peter J. Embí, Josh Martin, Darshan Shah, John Roach, Katie Allen |
AMIA | 2 |
| 2019 | Evaluating the effect of data standardization and validation on patient matching accuracyabstractOBJECTIVE: This study evaluated the degree to which recommendations for demographic data standardization improve patient matching accuracy using real-world datasets. MATERIALS AND METHODS: We used 4 manually reviewed datasets, containing a random selection of matches and nonmatches. Matching datasets included health information exchange (HIE) records, public health registry records, Social Security Death Master File records, and newborn screening records. Standardized fields including last name, telephone number, social security number, date of birth, and address. Matching performance was evaluated using 4 metrics: sensitivity, specificity, positive predictive value, and accuracy. RESULTS: Standardizing address was independently associated with improved matching sensitivities for both the public health and HIE datasets of approximately 0.6% and 4.5%. Overall accuracy was unchanged for both datasets due to reduced match specificity. We observed no similar impact for address standardization in the death master file dataset. Standardizing last name yielded improved matching sensitivity of 0.6% for the HIE dataset, while overall accuracy remained the same due to a decrease in match specificity. We noted no similar impact for other datasets. Standardizing other individual fields (telephone, date of birth, or social security number) showed no matching improvements. As standardizing address and last name improved matching sensitivity, we examined the combined effect of address and last name standardization, which showed that standardization improved sensitivity from 81.3% to 91.6% for the HIE dataset. CONCLUSIONS: Data standardization can improve match rates, thus ensuring that patients and clinicians have better data on which to make decisions to enhance care quality and safety. Shaun J. Grannis, Huiping Xu, Joshua R. Vest, Suranga Nath Kasthurirathne, Na Bo, Ben Moscovitch, Rita Torkzadeh, Josh Rising |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Improving Population Health Reporting through Information Exchange Supported Decision Support: A Controlled Before-and-After Trial
Brian E. Dixon, Zuoyi Zhang, P. Joseph Gibson, Debra Revere, Shaun J. Grannis |
AMIA | 5 |
| 2018 | Navigating the Exposome: Real-World Experiences Connecting Environment, Community, and Behavior to Health
Shaun J. Grannis, Joshua R. Vest, Chirag Patel, Elle Holbrook, Brian E. Dixon |
AMIA | 1 |
| 2018 | Have you heard the one about the doctor who went into the exam room took care of the patient and walked out?
J. Marc Overhage, Titus Schleyer, Shaun J. Grannis, Allan Fong, Raj M. Ratwani |
AMIA | 3 |
| 2018 | The Linchpin of Interoperability: Challenges and Solutions to Patient Record Matching
Robert S. Rudin, Shaun J. Grannis, Micky Tripathi, Jeffery Smith, Ben Moscovitch |
AMIA | 2 |
| 2018 | Examining the Heartland Region Pilot: First Look at the Patient-Centered Data HomeTM Framework
Karmen S. Williams, Shaun J. Grannis |
AMIA | 2 |
| 2018 | Assessing the capacity of social determinants of health data to augment predictive models identifying patients in need of wraparound social servicesabstractIntroduction: A growing variety of diverse data sources is emerging to better inform health care delivery and health outcomes. We sought to evaluate the capacity for clinical, socioeconomic, and public health data sources to predict the need for various social service referrals among patients at a safety-net hospital. Materials and Methods: We integrated patient clinical data and community-level data representing patients' social determinants of health (SDH) obtained from multiple sources to build random forest decision models to predict the need for any, mental health, dietitian, social work, or other SDH service referrals. To assess the impact of SDH on improving performance, we built separate decision models using clinical and SDH determinants and clinical data only. Results: Decision models predicting the need for any, mental health, and dietitian referrals yielded sensitivity, specificity, and accuracy measures ranging between 60% and 75%. Specificity and accuracy scores for social work and other SDH services ranged between 67% and 77%, while sensitivity scores were between 50% and 63%. Area under the receiver operating characteristic curve values for the decision models ranged between 70% and 78%. Models for predicting the need for any services reported positive predictive values between 65% and 73%. Positive predictive values for predicting individual outcomes were below 40%. Discussion: The need for various social service referrals can be predicted with considerable accuracy using a wide range of readily available clinical and community data that measure socioeconomic and public health conditions. While the use of SDH did not result in significant performance improvements, our approach represents a novel and important application of risk predictive modeling. Suranga Nath Kasthurirathne, Joshua R. Vest, Nir Menachemi, Paul K. Halverson, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | Response to letter to the Editor on "Assessing the capacity of social determinants of health data to augment predictive models identifying patients in need of wraparound social services"abstractDear Dr Ohno-Machado, We thank Ancker and her colleagues for the insightful comments on our recent article.1 We wholeheartedly agree that the health informatics community should not conclude that social determinants of health (SDH) are not valuable. The literature in this area is growing, and we believe that SDH will continue to play an increasingly significant role in influencing population health. While the specific elements and local context of our work failed to demonstrate clear benefit, our study represented a single community with a novel outcome. As we noted in our article: a more diverse population or geography may have yielded different results; our results may not be generalizable to different outcomes; and that our SDH and public health measures were contextual. This last point is important, as individual and area level measures are different constructs entirely and an analysis, such as ours, is not subject to the ecological fallacy.2 Correlation between predictors is a significant problem; we believe that this was mitigated by our use of Random Forest, which selects random subsets of features to build an ensemble of trees.3 Consistent with the SDH perspective of social, political, and environmental settings, we endeavored to measure, and account for, patient context. We further agree that context is a frequently changing construct and that regularly updated measures are always better. We also believe that more individual level SDH measures would be an improvement. We applaud the Institute of Medicine (IOM) for recommending that SDH be captured in Electronic Health Records,4 as well as the International Classification of Disease (ICD) for enabling SDH collection by introducing additional SDH codes to ICD-10.5 We anticipate that the availability of patient-level SDH will increase as the adoption of these codes increases. Additionally, other patient-level SDH related to an individual’s family/social support may be inferred from his or her family medical history. Conflict of interest statement. None declared. Suranga Nath Kasthurirathne, Joshua R. Vest, Nir Menachemi, Paul K. Halverson, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | An Overview of Emerging Real-World OpenHIE Use Cases: Successes, Challenges, and Future Opportunities
Shaun J. Grannis, Eric-Jan Manders, Frederick C. Leitner, Annah Ngaruro, Jack Bowie |
AMIA | 1 |
| 2017 | Evaluation of Text Mining Methods to Support Reporting Public Health Notifiable Diseases Using Real-World Clinical Data
Matthias Kochmann, Brian E. Dixon, Suranga Nath Kasthurirathne, Shaun J. Grannis |
AMIA | 5 |
| 2017 | Using EHR and HIE data to identify patients' need for services that address the social determinants of health
Joshua R. Vest, Shaun J. Grannis, Jennifer L. Williams, Dawn P. Haut, Nir Menachemi |
AMIA | 2 |
| 2017 | Discriminative boosted Bayes networks for learning multiple cardiovascular proceduresabstractWe consider the problem of predicting three procedures, viz, EKG, Angioplasty and Valve Replacement procedures jointly from Electronic Health Records (EHR) and develop a discriminative boosted Bayesian network algorithm. Differences between our proposed approach and standard Bayes Net structure learners are (1) we do not assume that the number of features (observations) are uniform across training examples and (2) our method explicitly handles the precision-recall tradeoff. Our empirical evaluations on a real EHR data demonstrates the superiority of this proposed approach to learning these procedures individually. Nandini Ramanan, Shuo Yang 0004, Shaun J. Grannis, Sriraam Natarajan |
BIBM | 3 |
| 2017 | Modeling heart procedures from EHRs: An application of exponential familiesabstractIn order to facilitate better estimations on coronary artery disease conditions of a patient, we aim to predict the number of Angioplasty (a coronary artery procedure) by taking into account all the information from his/her Electronic Health Record (EHR) data. For this purpose, two exponential family members—multinomial distribution and Poisson distribution models—are considered, which treat the target variable as categorical-valued and count-valued respectively. From the perspective of exponential family, we derive the functional gradient boosting approach for these two distributions and analyze their assumptions with real EHR data. Our empirical results show that Poisson models appear to be more faithful for modeling the number of this procedure. Shuo Yang 0004, Fabian Hadiji, Kristian Kersting, Shaun J. Grannis, Sriraam Natarajan |
BIBM | 4 |
| 2017 | Toward better public health reporting using existing off the shelf approaches: The value of medical dictionaries in automated cancer detection using plaintext medical data
Suranga Nath Kasthurirathne, Brian E. Dixon, Judy Gichoya, Huiping Xu, Yuni Xia, Burke W. Mamlin, Shaun J. Grannis |
J. Biomed. Informatics | 7 |
| 2016 | Transmission of ELR messages To Improve Public Health Reporting
Samar Binkheder, Brian E. Dixon, Shaun J. Grannis |
AMIA | 3 |
| 2016 | Methods to Measure and Improve the Quality of Large Scale Health Data: An Application in Public Health Surveillance
Brian E. Dixon, Jon D. Duke, Shaun J. Grannis |
AMIA | 3 |
| 2016 | Toward better public health reporting using existing off the shelf approaches: A comparison of alternative cancer detection approaches using plaintext medical data and non-dictionary based feature selection
Suranga Nath Kasthurirathne, Brian E. Dixon, Judy Gichoya, Huiping Xu, Yuni Xia, Burke W. Mamlin, Shaun J. Grannis |
J. Biomed. Informatics | 7 |
| 2015 | Improving Vaccine-Preventable Disease Reporting through Health Information Exchange
Brian E. Dixon, P. Joseph Gibson, Shaun J. Grannis |
AMIA | 3 |
| 2015 | The Informatics Workforce for Population Health: Challenges, Initiatives and the Path Forward
Brian E. Dixon, Marty LaVenture, Bill Brand, Arthur J. Davidson, Shaun J. Grannis |
AMIA | 5 |
| 2015 | Completeness and Timeliness of Notifiable Disease Surveillance Data Submitted by Providers to Public Health Authorities
Brian E. Dixon, Patrick Lai, Uzay Kirbiyik, Zuoyi Zhang, Debra Revere, Rebecca A. Hills, P. Joseph Gibson, Jennifer L. Williams, Shaun J. Grannis |
AMIA | 9 |
| 2015 | Evaluating the Accuracy of Automated Notifiable Condition Detection in Free-Text Electronic Laboratory Report Results Using Contemporary Text Mining and Machine Learning Methods
Uzay Kirbiyik, Patrick Lai, Brian E. Dixon, Shaun J. Grannis, Suranga Nath Kasthurirathne |
AMIA | 4 |
| 2014 | Demonstrating A Public Health Terrain Data Visualization System
Jeremy Keiper, Shiaofen Fang, Mathew J. Palakal, Yuni Xia, Sam Bloomquist, Shaun J. Grannis, Weizhi Li, Thanh M. Nguyen, Anand Krishnan |
AMIA | 6 |
| 2014 | An Evaluation of Two Methods for Generating Synthetic HL7 Segments Reflecting Real-World Health Information Exchange Transactions
Thomas S. Mwogi, Paul G. Biondich, Shaun J. Grannis |
AMIA | 3 |
| 2014 | Using participatory design to optimize capture of information needed for public health reporting processes
Debra Revere, Jennifer L. Williams, Rebecca A. Hills, Shaun J. Grannis, Brian E. Dixon |
AMIA | 4 |
| 2014 | The long road to semantic interoperability in support of public health: Experiences from two statesabstractProliferation of health information technologies creates opportunities to improve clinical and public health, including high quality, safer care and lower costs. To maximize such potential benefits, health information technologies must readily and reliably exchange information with other systems. However, evidence from public health surveillance programs in two states suggests that operational clinical information systems often fail to use available standards, a barrier to semantic interoperability. Furthermore, analysis of existing policies incentivizing semantic interoperability suggests they have limited impact and are fragmented. In this essay, we discuss three approaches for increasing semantic interoperability to support national goals for using health information technologies. A clear, comprehensive strategy requiring collaborative efforts by clinical and public health stakeholders is suggested as a guide for the long road towards better population health data and outcomes. Brian E. Dixon, Daniel J. Vreeman, Shaun J. Grannis |
J. Biomed. Informatics | 3 |
| 2013 | The Regenstrief Notifiable Condition Detector - Automated Public Health Reporting using Routine Electronic Laboratory Data
Brian E. Dixon, Shaun J. Grannis, Mark Tucker, Debbie Hemler |
AMIA | 2 |
| 2013 | Variation in Information Needs and Quality: Implications for Public Health Surveillance and Biomedical Informatics
Brian E. Dixon, Patrick Lai, Shaun J. Grannis |
AMIA | 3 |
| 2013 | Measuring and Improving the Fitness of Electronic Clinical Data for Reuse in Public Health, Research, and Other Use Cases
Brian E. Dixon, Marc B. Rosenman, Shaun J. Grannis |
AMIA | 3 |
| 2013 | How Fit is Electronic Health Data for its Intended Uses? Exploring Data Quality across Clinical, Public Health, and Research Use Cases
Shaun J. Grannis, Brian E. Dixon, Siaw-Teng Liaw, Michael G. Kahn, Hamish S. F. Fraser |
AMIA | 1 |
| 2013 | A Practical Method for Predicting Frequent Use of Emergency Department Care Using Routinely Available Electronic Registration Data
Jianmin Wu, Huiping Xu, John T. Finnell, Shaun J. Grannis |
AMIA | 4 |
| 2013 | Towards public health decision support: a systematic review of bidirectional communication approachesabstractOBJECTIVE: To summarize the literature describing computer-based interventions aimed at improving bidirectional communication between clinical and public health. MATERIALS AND METHODS: A systematic review of English articles using MEDLINE and Google Scholar. Search terms included public health, epidemiology, electronic health records, decision support, expert systems, and decision-making. Only articles that described the communication of information regarding emerging health threats from public health agencies to clinicians or provider organizations were included. Each article was independently reviewed by two authors. RESULTS: Ten peer-reviewed articles highlight a nascent but promising area of research and practice related to alerting clinicians about emerging threats. Current literature suggests that additional research and development in bidirectional communication infrastructure should focus on defining a coherent architecture, improving interoperability, establishing clear governance, and creating usable systems that will effectively deliver targeted, specific information to clinicians in support of patient and population decision-making. CONCLUSIONS: Increasingly available clinical information systems make it possible to deliver timely, relevant knowledge to frontline clinicians in support of population health. Future work should focus on developing a flexible, interoperable infrastructure for bidirectional communications capable of integrating public health knowledge into clinical systems and workflows. Brian E. Dixon, Roland E. Gamache, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Impact of Selective Mapping Strategies on Automated Laboratory Result Notification to Public Health Authorities
Roland E. Gamache, Shaun J. Grannis, Brian E. Dixon, Daniel J. Vreeman |
AMIA | 2 |
| 2012 | An Evaluation of the Rates of Repeat Notifiable Disease Reporting and Patient Crossover Using a Health Information Exchange-based Automated Electronic Laboratory Reporting System
Judy Gichoya, Brian E. Dixon, John T. Finnell, Daniel J. Vreeman, Roland E. Gamache, Shaun J. Grannis |
AMIA | 6 |
| 2012 | Unintended Consequences of Health Information Exchange: Identifying the Issues and Mitigating the Impact
Julie J. McGowan, Gilad J. Kuperman, Shaun J. Grannis, P. Jon White, Kathy Kenyon |
AMIA | 3 |
| 2012 | Patient Crossover between Healthcare Systems in Indiana
Blaine Y. Takesue, J. Marc Overhage, Shaun J. Grannis, Siu L. Hui, Marc B. Rosenman |
AMIA | 3 |
| 2012 | Evaluation of a clinical decision support algorithm for patient-specific childhood immunization
Vivienne J. Zhu, Shaun J. Grannis, Wanzhu Tu, Marc B. Rosenman, Stephen M. Downs |
Artif. Intell. Medicine | 2 |
| 2010 | Data visualization speeds review of potential adverse drug events in patients on multiple medications
Jon D. Duke, Shaun J. Grannis |
J. Biomed. Informatics | 3 |
| 2010 | A Visual Analytics Approach to Understanding Spatiotemporal HotspotsabstractAs data sources become larger and more complex, the ability to effectively explore and analyze patterns among varying sources becomes a critical bottleneck in analytic reasoning. Incoming data contain multiple variables, high signal-to-noise ratio, and a degree of uncertainty, all of which hinder exploration, hypothesis generation/exploration, and decision making. To facilitate the exploration of such data, advanced tool sets are needed that allow the user to interact with their data in a visual environment that provides direct analytic capability for finding data aberrations or hotspots. In this paper, we present a suite of tools designed to facilitate the exploration of spatiotemporal data sets. Our system allows users to search for hotspots in both space and time, combining linked views and interactive filtering to provide users with contextual information about their data and allow the user to develop and explore their hypotheses. Statistical data models and alert detection algorithms are provided to help draw user attention to critical areas. Demographic filtering can then be further applied as hypotheses generated become fine tuned. This paper demonstrates the use of such tools on multiple geospatiotemporal data sets. Ross Maciejewski, Stephen Rudolph, Ryan Hafen, Ahmad M. Abusalah, Mohamed Yakout, Mourad Ouzzani, William S. Cleveland, Shaun J. Grannis, David S. Ebert |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2009 | A Comparison of Automated Methicillin-Resistant Staphylococcus aureus Identification with Current Infection Control Practice
David Shepherd, F. Jeffrey Friedlin, Shaun J. Grannis, Siu L. Hui, Abel N. Kho |
AMIA | 3 |
| 2009 | Implementing Broad Scale Childhood Immunization Decision Support as a Web Service
Vivienne J. Zhu, Shaun J. Grannis, Marc B. Rosenman, Stephen M. Downs |
AMIA | 2 |
| 2009 | Research Paper: An Empiric Modification to the Probabilistic Record Linkage Algorithm Using Frequency-Based Weight ScalingabstractOBJECTIVE: To incorporate value-based weight scaling into the Fellegi-Sunter (F-S) maximum likelihood linkage algorithm and evaluate the performance of the modified algorithm. Background Because healthcare data are fragmented across many healthcare systems, record linkage is a key component of fully functional health information exchanges. Probabilistic linkage methods produce more accurate, dynamic, and robust matching results than rule-based approaches, particularly when matching patient records that lack unique identifiers. Theoretically, the relative frequency of specific data elements can enhance the F-S method, including minimizing the false-positive or false-negative matches. However, to our knowledge, no frequency-based weight scaling modification to the F-S method has been implemented and specifically evaluated using real-world clinical data. METHODS: The authors implemented a value-based weight scaling modification using an information theoretical model, and formally evaluated the effectiveness of this modification by linking 51,361 records from Indiana statewide newborn screening data to 80,089 HL7 registration messages from the Indiana Network for Patient Care, an operational health information exchange. In addition to applying the weight scaling modification to all fields, we examined the effect of selectively scaling common or uncommon field-specific values. RESULTS: The sensitivity, specificity, and positive predictive value for applying weight scaling to all field-specific values were 95.4, 98.8, and 99.9%, respectively. Compared with nonweight scaling, the modified F-S algorithm demonstrated a 10% increase in specificity with a 3% decrease in sensitivity. CONCLUSION: By eliminating false-positive matches, the value-based weight modification can enhance the specificity of the F-S method with minimal decrease in sensitivity. Vivienne J. Zhu, J. Marc Overhage, James Egg, Stephen M. Downs, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 5 |
| 2008 | Using Natural Language Processing to Improve Accuracy of Automated Notifiable Disease Reporting
F. Jeffrey Friedlin, Shaun J. Grannis, J. Marc Overhage |
AMIA | 2 |
| 2006 | The Indiana Public Health Emergency Surveillance System: Ongoing Progress, Early Findings, and Future Directions
Shaun J. Grannis, Michael Wade, P. Joseph Gibson, J. Marc Overhage |
AMIA | 1 |
| 2006 | Application of Information Technology: A Context-sensitive Approach to Anonymizing Spatial Surveillance Data: Impact on Outbreak DetectionabstractOBJECTIVE: The use of spatially based methods and algorithms in epidemiology and surveillance presents privacy challenges for researchers and public health agencies. We describe a novel method for anonymizing individuals in public health data sets by transposing their spatial locations through a process informed by the underlying population density. Further, we measure the impact of the skew on detection of spatial clustering as measured by a spatial scanning statistic. DESIGN: Cases were emergency department (ED) visits for respiratory illness. Baseline ED visit data were injected with artificially created clusters ranging in magnitude, shape, and location. The geocoded locations were then transformed using a de-identification algorithm that accounts for the local underlying population density. MEASUREMENTS: A total of 12,600 separate weeks of case data with artificially created clusters were combined with control data and the impact on detection of spatial clustering identified by a spatial scan statistic was measured. RESULTS: The anonymization algorithm produced an expected skew of cases that resulted in high values of data set k-anonymity. De-identification that moves points an average distance of 0.25 km lowers the spatial cluster detection sensitivity by less than 4% and lowers the detection specificity less than 1%. CONCLUSION: A population-density-based Gaussian spatial blurring markedly decreases the ability to identify individuals in a data set while only slightly decreasing the performance of a standardly used outbreak detection tool. These findings suggest new approaches to anonymizing data for spatial epidemiology and surveillance. Christopher A. Cassa, Shaun J. Grannis, J. Marc Overhage, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 2 |
| 2005 | How Disease Surveillance Systems Can Serve as Practical Building Blocks for a Health Information Infrastructure: the Indiana Experience
Shaun J. Grannis, Paul G. Biondich, Burke W. Mamlin, M. D. Greg Wilson, Linda Jones, J. Marc Overhage |
AMIA | 1 |
| 2005 | Reviewing and Managing Syndromic Surveillance SaTScan™ Datasets using an Open-Source Data Visualization Tool
Shaun J. Grannis, James Egg, J. Marc Overhage |
AMIA | 1 |
| 2004 | Review Paper: Implementing Syndromic Surveillance: A Practical Guide Informed by the Early ExperienceabstractSyndromic surveillance refers to methods relying on detection of individual and population health indicators that are discernible before confirmed diagnoses are made. In particular, prior to the laboratory confirmation of an infectious disease, ill persons may exhibit behavioral patterns, symptoms, signs, or laboratory findings that can be tracked through a variety of data sources. Syndromic surveillance systems are being developed locally, regionally, and nationally. The efforts have been largely directed at facilitating the early detection of a covert bioterrorist attack, but the technology may also be useful for general public health, clinical medicine, quality improvement, patient safety, and research. This report, authored by developers and methodologists involved in the design and deployment of the first wave of syndromic surveillance systems, is intended to serve as a guide for informaticians, public health managers, and practitioners who are currently planning deployment of such systems in their regions. Kenneth D. Mandl, J. Marc Overhage, Michael M. Wagner 0001, William B. Lober, Paola Sebastiani, Farzad Mostashari, Julie A. Pavlin, Per H. Gesteland, Tracee Treadwell, Eileen Koski, Lori Hutwagner, David L. Buckeridge, Raymond D. Aller, Shaun J. Grannis |
J. Am. Medical Informatics Assoc. | 14 |
| 2003 | Analysis of a Probabilistic Record Linkage Technique without Human Review
Shaun J. Grannis, J. Marc Overhage, Siu L. Hui, Clement J. McDonald |
AMIA | 1 |
| 2003 | Research Paper: Detection of Pediatric Respiratory and Diarrheal Outbreaks from Sales of Over-the-counter Electrolyte ProductsabstractOBJECTIVE: To determine whether sales of electrolyte products contain a signal of outbreaks of respiratory and diarrheal disease in children and, if so, how much earlier a signal relative to hospital diagnoses. DESIGN: Retrospective analysis was conducted of sales of electrolyte products and hospital diagnoses for six urban regions in three states for the period 1998 through 2001. MEASUREMENTS: Presence of signal was ascertained by measuring correlation between electrolyte sales and hospital diagnoses and the temporal relationship that maximized correlation. Earliness was the difference between the date that the exponentially weighted moving average (EWMA) method first detected an outbreak from sales and the date it first detected the outbreak from diagnoses. The coefficient of determination (r2) measured how much variance in earliness resulted from differences in sales' and diagnoses' signal strengths. RESULTS: The correlation between electrolyte sales and hospital diagnoses was 0.90 (95% CI, 0.87-0.93) at a time offset of 1.7 weeks (95% CI, 0.50-2.9), meaning that sales preceded diagnoses by 1.7 weeks. EWMA with a nine-sigma threshold detected the 18 outbreaks on average 2.4 weeks (95% CI, 0.1-4.8 weeks) earlier from sales than from diagnoses. Twelve outbreaks were first detected from sales, four were first detected from diagnoses, and two were detected simultaneously. Only 26% of variance in earliness was explained by the relative strength of the sales and diagnoses signals (r2 = 0.26). CONCLUSION: Sales of electrolyte products contain a signal of outbreaks of respiratory and diarrheal diseases in children and usually are an earlier signal than hospital diagnoses. William R. Hogan, Fu-Chiang Tsui, Oleg Ivanov, Per H. Gesteland, Shaun J. Grannis, J. Marc Overhage, J. Michael Robinson, Michael M. Wagner 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2002 | Analysis of identifier performance using a deterministic linkage algorithm
Shaun J. Grannis, J. Marc Overhage, Clement J. McDonald |
AMIA | 1 |