VLDB 2026 Research / reviewers in the wild / expert
Kavishwar B. Wagholikar
dblp:82/7222
· DBLP profile ↗
35ranked-venue papers
12as first author
4since 2021 · last 2025
0000-0002-6219-861XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 12 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Principles and implementation strategies for equitable and representative academic partnerships in global health informatics researchabstractOBJECTIVE: Developing equitable, sustainable informatics solutions is key to scalability and long-term success for projects in the global health informatics (GHI) domain. This paper presents key strategies for incorporating principles of health equity in the GHI project lifecycle. MATERIALS AND METHODS: The American Medical Informatics Association (AMIA) GHI Working Group organized a collaborative workshop at the 2023 AMIA Annual Symposium that included the presentation of five case studies of how principles of health equity have been incorporated into projects situated in low-and-middle-income countries and with Indigenous communities in the U.S. and best practices for operationalizing these principles into other informatics projects. RESULTS: We present five principles: (1) Inclusion and Participation in Ethical, Sustainable Collaborations; (2) Engaging Community-Based Participatory Research Approaches; (3) Stakeholder Engagement; (4) Scalability and Sustainability; (5) Representation in Knowledge Creation, along with strategies that informatics researchers may use to incorporate these principles into their work. DISCUSSION: Presented case studies and subsequent focus groups yielded key concepts and strategies to promote health equity that may be operationalized across GHI projects. CONCLUSION: Equitable, sustainable, and scalable GHI projects require intentional integration of community and stakeholder perspectives in project development, implementation, and knowledge creation processes. Elizabeth A. Campbell, Oliver J. Bear Don't Walk IV, Hamish S. F. Fraser, Judy Gichoya, Kavishwar B. Wagholikar, Andrew S. Kanter, Felix Holl, Sansanee Craig |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | I2b2-etl: Python application for importing electronic health data into the informatics for integrating biology and the bedside platformabstractMOTIVATION: The i2b2 platform is used at major academic health institutions and research consortia for querying for electronic health data. However, a major obstacle for wider utilization of the platform is the complexity of data loading that entails a steep curve of learning the platform's complex data schemas. To address this problem, we have developed the i2b2-etl package that simplifies the data loading process, which will facilitate wider deployment and utilization of the platform. RESULTS: We have implemented i2b2-etl as a Python application that imports ontology and patient data using simplified input file schemas and provides inbuilt record number de-identification and data validation. We describe a real-world deployment of i2b2-etl for a population-management initiative at MassGeneral Brigham. AVAILABILITY AND IMPLEMENTATION: i2b2-etl is a free, open-source application implemented in Python available under the Mozilla 2 license. The application can be downloaded as compiled docker images. A live demo is available at https://i2b2clinical.org/demo-i2b2etl/ (username: demo, password: Etl@2021). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kavishwar B. Wagholikar, Layne Ainsworth, David Zelle, Kira Chaney, Michael Mendis, Jeffrey G. Klann, Alexander J. Blood, Angela Miller, Rupendra Chulyadyo, Michael Oates, William J. Gordon, Samuel J. Aronson, Benjamin M. Scirica, Shawn N. Murphy |
Bioinform. | 1 |
| 2022 | An objective framework for evaluating unrecognized bias in medical AI models predicting COVID-19 outcomesabstractOBJECTIVE: The increasing translation of artificial intelligence (AI)/machine learning (ML) models into clinical practice brings an increased risk of direct harm from modeling bias; however, bias remains incompletely measured in many medical AI applications. This article aims to provide a framework for objective evaluation of medical AI from multiple aspects, focusing on binary classification models. MATERIALS AND METHODS: Using data from over 56 000 Mass General Brigham (MGB) patients with confirmed severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), we evaluate unrecognized bias in 4 AI models developed during the early months of the pandemic in Boston, Massachusetts that predict risks of hospital admission, ICU admission, mechanical ventilation, and death after a SARS-CoV-2 infection purely based on their pre-infection longitudinal medical records. Models were evaluated both retrospectively and prospectively using model-level metrics of discrimination, accuracy, and reliability, and a novel individual-level metric for error. RESULTS: We found inconsistent instances of model-level bias in the prediction models. From an individual-level aspect, however, we found most all models performing with slightly higher error rates for older patients. DISCUSSION: While a model can be biased against certain protected groups (ie, perform worse) in certain tasks, it can be at the same time biased towards another protected group (ie, perform better). As such, current bias evaluation studies may lack a full depiction of the variable effects of a model on its subpopulations. CONCLUSION: Only a holistic evaluation, a diligent search for unrecognized bias, can provide enough information for an unbiased judgment of AI bias that can invigorate follow-up investigations on identifying the underlying roots of bias and ultimately make a change. Hossein Estiri, Zachary H. Strasser, Sina Rashidian, Jeffrey G. Klann, Kavishwar B. Wagholikar, Thomas H. McCoy Jr., Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record dataabstractOBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites. Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 29 |
| 2020 | High-throughput Phenotyping with EHR Sequences
Shawn N. Murphy, Hossein Estiri, Zachary H. Strasser, Kavishwar B. Wagholikar, Victor M. Castro |
AMIA | 4 |
| 2020 | Polar labeling: silver standard algorithm for training disease classifiersabstractMOTIVATION: Expert-labeled data are essential to train phenotyping algorithms for cohort identification. However expert labeling is time and labor intensive, and the costs remain prohibitive for scaling phenotyping to wider use-cases. RESULTS: We present an approach referred to as polar labeling (PL), to create silver standard for training machine learning (ML) for disease classification. We test the hypothesis that ML models trained on the silver standard created by applying PL on unlabeled patient records, are comparable in performance to the ML models trained on gold standard, created by clinical experts through manual review of patient records. We perform experimental validation using health records of 38 023 patients spanning six diseases. Our results demonstrate the superior performance of the proposed approach. AVAILABILITY AND IMPLEMENTATION: We provide a Python implementation of the algorithm and the Python code developed for this study on Github. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kavishwar B. Wagholikar, Hossein Estiri, Marykate Murphy, Shawn N. Murphy |
Bioinform. | 1 |
| 2019 | Ontologies Enabling Computable Tables
Jeffrey G. Klann, Nich Wattanasin, Michael Mendis, Matthew A. Joss, Hossein Estiri, Kavishwar B. Wagholikar, Shawn N. Murphy |
AMIA | 6 |
| 2019 | Patient Stratification Process for enabling Clinical Interventions
Kavishwar B. Wagholikar, Samuel J. Aronson, Benjamin M. Scirica, Akshay S. Desai, Shawn N. Murphy |
AMIA | 1 |
| 2019 | Dynamic Phenotyping to facilitate Accrual for Prospective Clinical studies: A Case Study in Heart Failure
Kavishwar B. Wagholikar, Christina M. Fischer, Alyssa P. Goodson, Christopher Herrick, Taylor Maclean, Katelyn Smith, Liliana Fera, Thomas Gaziano, Jacqueline Dunning, Joshua Bosque-Hamilton, Lina Matta, Eloy Toscano, Brent Richter, Layne Ainsworth, Michael Oates, Samuel J. Aronson, Calum A. MacRae, Benjamin M. Scirica, Akshay S. Desai, Shawn N. Murphy |
AMIA | 1 |
| 2019 | Stratification of Patient Population for enabling data-driven Clinical Interventions using I2b2
Kavishwar B. Wagholikar, Vishal Vernekar, Akshay Zagade, Yuri Ostrovsky, Shek-Wayne Chan, Alyssa P. Goodson, Ameet Pathak, Corey Glynn, Christopher Herrick, Shawn N. Murphy |
AMIA | 1 |
| 2019 | Plugin for importing spreadsheets into Informatics for Integrating Biology and the Bedside platform
Akshay Zagade, Vishal Vernekar, Shek-Wayne Chan, Kavishwar B. Wagholikar, Rupendra Chulyadyo, Yuri Ostrovsky, Alyssa P. Goodson, Ameet Pathak, Christopher Herrick, Shawn N. Murphy |
AMIA | 4 |
| 2018 | On DXplain vocabulary development: Past and Present
Mitchell J. Feldman, Kavishwar B. Wagholikar, Kathleen Famiglietti, Richard J. Kim, Edward P. Hoffer, Henry C. Chueh |
AMIA | 2 |
| 2018 | Natural Language Processing to Detect High Information Findings for Patients at risk of Missed Diagnosis
Tzu-I Yang, Chia-Ching Chou, Kavishwar B. Wagholikar, Mitchell J. Feldman, Henry C. Chueh |
AMIA | 3 |
| 2018 | Web services for data warehouses: OMOP and PCORnet on i2b2abstractObjective: Healthcare organizations use research data models supported by projects and tools that interest them, which often means organizations must support the same data in multiple models. The healthcare research ecosystem would benefit if tools and projects could be adopted independently from the underlying data model. Here, we introduce the concept of a reusable application programming interface (API) for healthcare and show that the i2b2 API can be adapted to support diverse patient-centric data models. Materials and Methods: We develop methodology for extending i2b2's pre-existing API to query additional data models, using i2b2's recent "multi-fact-table querying" feature. Our method involves developing data-model-specific i2b2 ontologies and mapping these to query non-standard table structure. Results: We implement this methodology to query OMOP and PCORnet models, which we validate with the i2b2 query tool. We implement the entire PCORnet data model and a five-domain subset of the OMOP model. We also demonstrate that additional, ancillary data model columns can be modeled and queried as i2b2 "modifiers." Discussion: i2b2's REST API can be used to query multiple healthcare data models, enabling shared tooling to have a choice of backend data stores. This enables separation between data model and software tooling for some of the more popular open analytic data models in healthcare. Conclusion: This methodology immediately allows querying OMOP and PCORnet using the i2b2 API. It is released as an open-source set of Docker images, and also on the i2b2 community wiki. Jeffrey G. Klann, Lori C. Phillips, Christopher Herrick, Matthew A. Joss, Kavishwar B. Wagholikar, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Applying unsupervised learning to characterize rare observations in clinical data: the DQe-p tool
Hossein Estiri, Jeffrey G. Klann, Kavishwar B. Wagholikar, Shawn N. Murphy |
AMIA | 3 |
| 2017 | Web-Service-Enabled Apps for Research: SMART-on-FHIR for OMOP and PCORNet
Jeffrey G. Klann, Kavishwar B. Wagholikar, Lori C. Phillips, Matthew A. Joss, Shawn N. Murphy |
AMIA | 2 |
| 2017 | Computing Performance Analysis on Clinical Document-level Classification
Wei-Hung Weng, Kavishwar B. Wagholikar, Henry C. Chueh |
AMIA | 2 |
| 2017 | SMART-on-FHIR implemented over i2b2abstractWe have developed an interface to serve patient data from Informatics for Integrating Biology and the Bedside (i2b2) repositories in the Fast Healthcare Interoperability Resources (FHIR) format, referred to as a SMART-on-FHIR cell. The cell serves FHIR resources on a per-patient basis, and supports the "substitutable" modular third-party applications (SMART) OAuth2 specification for authorization of client applications. It is implemented as an i2b2 server plug-in, consisting of 6 modules: authentication, REST, i2b2-to-FHIR converter, resource enrichment, query engine, and cache. The source code is freely available as open source. We tested the cell by accessing resources from a test i2b2 installation, demonstrating that a SMART app can be launched from the cell that accesses patient data stored in i2b2. We successfully retrieved demographics, medications, labs, and diagnoses for test patients. The SMART-on-FHIR cell will enable i2b2 sites to provide simplified but secure data access in FHIR format, and will spur innovation and interoperability. Further, it transforms i2b2 into an apps platform. Kavishwar B. Wagholikar, Joshua C. Mandel, Jeffrey G. Klann, Nich Wattanasin, Michael Mendis, Christopher G. Chute, Kenneth D. Mandl, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001 |
AMIA | 5 |
| 2016 | Cajun Codefest 4.0 on SMART-on-FHIR apps for Diabetes
Kavishwar B. Wagholikar, Eliel Oliveira, Henry Chu, Harshal Shah, Joshua C. Mandel, Jeffrey G. Klann, Sohail Rao, Kenneth D. Mandl, Shawn N. Murphy, Thomas Carton |
AMIA | 1 |
| 2016 | Evaluation of SMART-on-FHIR I2b2 cell using PCORNET data model
Kavishwar B. Wagholikar, Eliel Oliveira, Joshua C. Mandel, Jeffrey G. Klann, Prasad Patil, Kenneth D. Mandl, Shawn N. Murphy, Thomas Carton |
AMIA | 1 |
| 2016 | Improving the Workflow of Curbside Consultation by Using Unstructured Clinical Notes - a Natural Language and Machine Learning-based Approach
Wei-Hung Weng, Avni Khatri, Kavishwar B. Wagholikar, Adam Cohen, Henry C. Chueh |
AMIA | 3 |
| 2016 | Erratum to: Text mining facilitates database curation - extraction of mutation-disease associations from Bio-medical literature
K. E. Ravikumar, Kavishwar B. Wagholikar, Dingcheng Li, Jean-Pierre A. Kocher |
BMC Bioinform. | 2 |
| 2015 | Evaluation of the accuracy of CDS for cervical cancer screening and surveillance
Kathy L. MacLaughlin, K. E. Ravikumar, Kavishwar B. Wagholikar, Marianne R. Scheitel, Rajeev Chaudhry |
AMIA | 3 |
| 2015 | Text mining facilitates database curation - extraction of mutation-disease associations from Bio-medical literatureabstractBACKGROUND: Advances in the next generation sequencing technology has accelerated the pace of individualized medicine (IM), which aims to incorporate genetic/genomic information into medicine. One immediate need in interpreting sequencing data is the assembly of information about genetic variants and their corresponding associations with other entities (e.g., diseases or medications). Even with dedicated effort to capture such information in biological databases, much of this information remains 'locked' in the unstructured text of biomedical publications. There is a substantial lag between the publication and the subsequent abstraction of such information into databases. Multiple text mining systems have been developed, but most of them focus on the sentence level association extraction with performance evaluation based on gold standard text annotations specifically prepared for text mining systems. RESULTS: We developed and evaluated a text mining system, MutD, which extracts protein mutation-disease associations from MEDLINE abstracts by incorporating discourse level analysis, using a benchmark data set extracted from curated database records. MutD achieves an F-measure of 64.3% for reconstructing protein mutation disease associations in curated database records. Discourse level analysis component of MutD contributed to a gain of more than 10% in F-measure when compared against the sentence level association extraction. Our error analysis indicates that 23 of the 64 precision errors are true associations that were not captured by database curators and 68 of the 113 recall errors are caused by the absence of associated disease entities in the abstract. After adjusting for the defects in the curated database, the revised F-measure of MutD in association detection reaches 81.5%. CONCLUSIONS: Our quantitative analysis reveals that MutD can effectively extract protein mutation disease associations when benchmarking based on curated database records. The analysis also demonstrates that incorporating discourse level analysis significantly improved the performance of extracting the protein-mutation-disease association. Future work includes the extension of MutD for full text articles. K. E. Ravikumar, Kavishwar B. Wagholikar, Dingcheng Li, Jean-Pierre A. Kocher |
BMC Bioinform. | 2 |
| 2014 | NLP enhances Quality Care Measures in Heart Failure
K. E. Ravikumar, Kavishwar B. Wagholikar |
AMIA | 2 |
| 2013 | Decision Support can Improve Time Efficiency of Healthcare Providers for Deciding Preventive Care Recommendations
Kavishwar B. Wagholikar, Ronald A. Hankey, Rajeev Chaudhry |
AMIA | 1 |
| 2013 | Comprehensive temporal information detection from clinical text: medical events, time, and TLINK identificationabstractBACKGROUND: Temporal information detection systems have been developed by the Mayo Clinic for the 2012 i2b2 Natural Language Processing Challenge. OBJECTIVE: To construct automated systems for EVENT/TIMEX3 extraction and temporal link (TLINK) identification from clinical text. MATERIALS AND METHODS: The i2b2 organizers provided 190 annotated discharge summaries as the training set and 120 discharge summaries as the test set. Our Event system used a conditional random field classifier with a variety of features including lexical information, natural language elements, and medical ontology. The TIMEX3 system employed a rule-based method using regular expression pattern match and systematic reasoning to determine normalized values. The TLINK system employed both rule-based reasoning and machine learning. All three systems were built in an Apache Unstructured Information Management Architecture framework. RESULTS: Our TIMEX3 system performed the best (F-measure of 0.900, value accuracy 0.731) among the challenge teams. The Event system produced an F-measure of 0.870, and the TLINK system an F-measure of 0.537. CONCLUSIONS: Our TIMEX3 system demonstrated good capability of regular expression rules to extract and normalize time information. Event and TLINK machine learning systems required well-defined feature sets to perform well. We could also leverage expert knowledge as part of the machine learning features to further improve TLINK identification performance. Sunghwan Sohn, Kavishwar B. Wagholikar, Dingcheng Li, Siddhartha Jonnalagadda, Cui Tao, K. E. Ravikumar |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Research and applications: Formative evaluation of the accuracy of a clinical decision support system for cervical cancer screeningabstractOBJECTIVES: We previously developed and reported on a prototype clinical decision support system (CDSS) for cervical cancer screening. However, the system is complex as it is based on multiple guidelines and free-text processing. Therefore, the system is susceptible to failures. This report describes a formative evaluation of the system, which is a necessary step to ensure deployment readiness of the system. MATERIALS AND METHODS: Care providers who are potential end-users of the CDSS were invited to provide their recommendations for a random set of patients that represented diverse decision scenarios. The recommendations of the care providers and those generated by the CDSS were compared. Mismatched recommendations were reviewed by two independent experts. RESULTS: A total of 25 users participated in this study and provided recommendations for 175 cases. The CDSS had an accuracy of 87% and 12 types of CDSS errors were identified, which were mainly due to deficiencies in the system's guideline rules. When the deficiencies were rectified, the CDSS generated optimal recommendations for all failure cases, except one with incomplete documentation. DISCUSSION AND CONCLUSIONS: The crowd-sourcing approach for construction of the reference set, coupled with the expert review of mismatched recommendations, facilitated an effective evaluation and enhancement of the system, by identifying decision scenarios that were missed by the system's developers. The described methodology will be useful for other researchers who seek rapidly to evaluate and enhance the deployment readiness of complex decision support systems. Kavishwar B. Wagholikar, Kathy L. MacLaughlin, Thomas M. Kastner, Petra M. Casey, Michael R. Henry, Robert A. Greenes, Rajeev Chaudhry |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Towards a semantic lexicon for clinical natural language processing
Stephen T. Wu, Dingcheng Li, Siddhartha Jonnalagadda, Sunghwan Sohn, Kavishwar B. Wagholikar, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
AMIA | 6 |
| 2012 | Clinical Decision Support with Natural Language Processing for Cervical Cancer Screening
Kavishwar B. Wagholikar, Kathy L. MacLaughlin, Michael R. Henry, Robert A. Greenes, Ronald A. Hankey, Rajeev Chaudhry |
AMIA | 1 |
| 2012 | Asthma Status Identification with Natural Language Processing
Stephen T. Wu, Young J. Juhn, Sunghwan Sohn, K. E. Ravikumar, Kavishwar B. Wagholikar, Siddhartha Jonnalagadda |
AMIA | 5 |
| 2012 | Coreference analysis in clinical notes: a multi-pass sieve with alternate anaphora resolution modulesabstractOBJECTIVE: This paper describes the coreference resolution system submitted by Mayo Clinic for the 2011 i2b2/VA/Cincinnati shared task Track 1C. The goal of the task was to construct a system that links the markables corresponding to the same entity. MATERIALS AND METHODS: The task organizers provided progress notes and discharge summaries that were annotated with the markables of treatment, problem, test, person, and pronoun. We used a multi-pass sieve algorithm that applies deterministic rules in the order of preciseness and simultaneously gathers information about the entities in the documents. Our system, MedCoref, also uses a state-of-the-art machine learning framework as an alternative to the final, rule-based pronoun resolution sieve. RESULTS: The best system that uses a multi-pass sieve has an overall score of 0.836 (average of B(3), MUC, Blanc, and CEAF F score) for the training set and 0.843 for the test set. DISCUSSION: A supervised machine learning system that typically uses a single function to find coreferents cannot accommodate irregularities encountered in data especially given the insufficient number of examples. On the other hand, a completely deterministic system could lead to a decrease in recall (sensitivity) when the rules are not exhaustive. The sieve-based framework allows one to combine reliable machine learning components with rules designed by experts. CONCLUSION: Using relatively simple rules, part-of-speech information, and semantic type properties, an effective coreference resolution system could be designed. The source code of the system described is available at https://sourceforge.net/projects/ohnlp/files/MedCoref. Siddhartha Jonnalagadda, Dingcheng Li, Sunghwan Sohn, Stephen T. Wu, Kavishwar B. Wagholikar, Manabu Torii |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Clinical decision support with automated text processing for cervical cancer screeningabstractOBJECTIVE: To develop a computerized clinical decision support system (CDSS) for cervical cancer screening that can interpret free-text Papanicolaou (Pap) reports. MATERIALS AND METHODS: The CDSS was constituted by two rulebases: the free-text rulebase for interpreting Pap reports and a guideline rulebase. The free-text rulebase was developed by analyzing a corpus of 49 293 Pap reports. The guideline rulebase was constructed using national cervical cancer screening guidelines. The CDSS accesses the electronic medical record (EMR) system to generate patient-specific recommendations. For evaluation, the screening recommendations made by the CDSS for 74 patients were reviewed by a physician. RESULTS AND DISCUSSION: Evaluation revealed that the CDSS outputs the optimal screening recommendations for 73 out of 74 test patients and it identified two cases for gynecology referral that were missed by the physician. The CDSS aided the physician to amend recommendations in six cases. The failure case was because human papillomavirus (HPV) testing was sometimes performed separately from the Pap test and these results were reported by a laboratory system that was not queried by the CDSS. Subsequently, the CDSS was upgraded to look up the HPV results missed earlier and it generated the optimal recommendations for all 74 test cases. LIMITATIONS: Single institution and single expert study. CONCLUSION: An accurate CDSS system could be constructed for cervical cancer screening given the standardized reporting of Pap tests and the availability of explicit guidelines. Overall, the study demonstrates that free text in the EMR can be effectively utilized through natural language processing to develop clinical decision support tools. Kavishwar B. Wagholikar, Kathy L. MacLaughlin, Michael R. Henry, Robert A. Greenes, Ronald A. Hankey, Rajeev Chaudhry |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Using machine learning for concept extraction on clinical documents from multiple data sourcesabstractOBJECTIVE: Concept extraction is a process to identify phrases referring to concepts of interests in unstructured text. It is a critical component in automated text processing. We investigate the performance of machine learning taggers for clinical concept extraction, particularly the portability of taggers across documents from multiple data sources. METHODS: We used BioTagger-GM to train machine learning taggers, which we originally developed for the detection of gene/protein names in the biology domain. Trained taggers were evaluated using the annotated clinical documents made available in the 2010 i2b2/VA Challenge workshop, consisting of documents from four data sources. RESULTS: As expected, performance of a tagger trained on one data source degraded when evaluated on another source, but the degradation of the performance varied depending on data sources. A tagger trained on multiple data sources was robust, and it achieved an F score as high as 0.890 on one data source. The results also suggest that performance of machine learning taggers is likely to improve if more annotated documents are available for training. CONCLUSION: Our study shows how the performance of machine learning taggers is degraded when they are ported across clinical documents from different sources. The portability of taggers can be enhanced by training on datasets from multiple sources. The study also shows that BioTagger-GM can be easily extended to detect clinical concept mentions with good performance. Manabu Torii, Kavishwar B. Wagholikar |
J. Am. Medical Informatics Assoc. | 2 |