EDBT 2026 Demo / reviewers in the wild / expert
Keith Marsolo
dblp:17/6184 · also Keith A. Marsolo
· DBLP profile ↗
33ranked-venue papers
12as first author
8since 2021 · last 2023
0000-0002-4416-1549ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-authorArtificial intelligence and machine learning · 3 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Potential bias and lack of generalizability in electronic health record data: reflections on health equity from the National Institutes of Health Pragmatic Trials CollaboratoryabstractEmbedded pragmatic clinical trials (ePCTs) play a vital role in addressing current population health problems, and their use of electronic health record (EHR) systems promises efficiencies that will increase the speed and volume of relevant and generalizable research. However, as the number of ePCTs using EHR-derived data grows, so does the risk that research will become more vulnerable to biases due to differences in data capture and access to care for different subsets of the population, thereby propagating inequities in health and the healthcare system. We identify 3 challenges-incomplete and variable capture of data on social determinants of health, lack of representation of vulnerable populations that do not access or receive treatment, and data loss due to variable use of technology-that exacerbate bias when working with EHR data and offer recommendations and examples of ways to actively mitigate bias. Andrew D. Boyd, Rosa Gonzalez-Guarda, Katharine Lawrence, Crystal L. Patil, Miriam O. Ezenwa, Emily C. O'Brien, Hyung Paek, Jordan M. Braciszewski, Oluwaseun Adeyemi, Allison M. Cuthel, Juanita E. Darby, Christina K. Zigler, P. Michael Ho, Keturah R. Faurot, Karen L. Staman, Jonathan W. Leigh, Dana L. Dailey, Andrea Cheville, Guilherme Del Fiol, Mitchell R. Knisely, Corita R. Grudzen, Keith Marsolo, Rachel L. Richesson, Judith M. Schlaeger |
J. Am. Medical Informatics Assoc. | 22 |
| 2023 | Assessing the impact of privacy-preserving record linkage on record overlap and patient demographic and clinical characteristics in PCORnet®, the National Patient-Centered Clinical Research NetworkabstractOBJECTIVE: This article describes the implementation of a privacy-preserving record linkage (PPRL) solution across PCORnet®, the National Patient-Centered Clinical Research Network. MATERIAL AND METHODS: Using a PPRL solution from Datavant, we quantified the degree of patient overlap across the network and report a de-duplicated analysis of the demographic and clinical characteristics of the PCORnet population. RESULTS: There were ∼170M patient records across the responding Network Partners, with ∼138M (81%) of those corresponding to a unique patient. 82.1% of patients were found in a single partner and 14.7% were in 2. The percentage overlap between Partners ranged between 0% and 80% with a median of 0%. Linking patients' electronic health records with claims increased disease prevalence in every clinical characteristic, ranging between 63% and 173%. DISCUSSION: The overlap between Partners was variable and depended on timeframe. However, patient data linkage changed the prevalence profile of the PCORnet patient population. CONCLUSIONS: This project was one of the largest linkage efforts of its kind and demonstrates the potential value of record linkage. Linkage between Partners may be most useful in cases where there is geographic proximity between Partners, an expectation that potential linkage Partners will be able to fill gaps in data, or a longer study timeframe. Keith Marsolo, Daniel Kiernan, Sengwee Toh, Jasmin Phua, Darcy Louzao, Kevin Haynes, Mark G. Weiner, Francisco Angulo, L. Charles Bailey, Jiang Bian 0001, Daniel Fort, Shaun J. Grannis, Ashok K. Krishnamurthy 0001, Vinit Nair, Pedro Rivera, Jonathan C. Silverstein, Maryan Zirkle, Thomas Carton |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Mis-mappings between a producer's quantitative test codes and LOINC codes and an algorithm for correcting themabstractOBJECTIVES: To access the accuracy of the Logical Observation Identifiers Names and Codes (LOINC) mapping to local laboratory test codes that is crucial to data integration across time and healthcare systems. MATERIALS AND METHODS: We used software tools and manual reviews to estimate the rate of LOINC mapping errors among 179 million mapped test results from 2 DataMarts in PCORnet. We separately reported unweighted and weighted mapping error rates, overall and by parts of the LOINC term. RESULTS: Of included 179 537 986 mapped results for 3029 quantitative tests, 95.4% were mapped correctly implying an 4.6% mapping error rate. Error rates were less than 5% for the more common tests with at least 100 000 mapped test results. Mapping errors varied across different LOINC classes. Error rates in chemistry and hematology classes, which together accounted for 92.0% of the mapped test results, were 0.4% and 7.5%, respectively. About 50% of mapping errors were due to errors in the property part of the LOINC name. DISCUSSIONS: Mapping errors could be detected automatically through inconsistencies in (1) qualifiers of the analyte, (2) specimen type, (3) property, and (4) method. Among quantitative test results, which are the large majority of reported tests, application of automatic error detection and correction algorithm could reduce the mapping errors further. CONCLUSIONS: Overall, the mapping error rate within the PCORnet data was 4.6%. This is nontrivial but less than other published error rates of 20%-40%. Such error rate decreased substantially to 0.1% after the application of automatic detection and correction algorithm. Clement J. McDonald, Seo H. Baik, Zhaonian Zheng, Liz Amos, Xiaocheng Luan, Keith Marsolo, Laura G. Qualls |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Standardizing, harmonizing, and protecting data collection to broaden the impact of COVID-19 research: the rapid acceleration of diagnostics-underserved populations (RADx-UP) initiativeabstractOBJECTIVE: The Rapid Acceleration of Diagnostics-Underserved Populations (RADx-UP) program is a consortium of community-engaged research projects with the goal of increasing access to Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) tests in underserved populations. To accelerate clinical research, common data elements (CDEs) were selected and refined to standardize data collection and enhance cross-consortium analysis. MATERIALS AND METHODS: The RADx-UP consortium began with more than 700 CDEs from the National Institutes of Health (NIH) CDE Repository, Disaster Research Response (DR2) guidelines, and the PHENotypes and eXposures (PhenX) Toolkit. Following a review of initial CDEs, we made selections and further refinements through an iterative process that included live forums, consultations, and surveys completed by the first 69 RADx-UP projects. RESULTS: Following a multistep CDE development process, we decreased the number of CDEs, modified the question types, and changed the CDE wording. Most research projects were willing to collect and share demographic NIH Tier 1 CDEs, with the top exception reason being a lack of CDE applicability to the project. The NIH RADx-UP Tier 1 CDE with the lowest frequency of collection and sharing was sexual orientation. DISCUSSION: We engaged a wide range of projects and solicited bidirectional input to create CDEs. These RADx-UP CDEs could serve as the foundation for a patient-centered informatics architecture allowing the integration of disease-specific databases to support hypothesis-driven clinical research in underserved populations. CONCLUSION: A community-engaged approach using bidirectional feedback can lead to the better development and implementation of CDEs in underserved populations during public health emergencies. Gabriel A. Carrillo, Michael Cohen-Wolkowiez, Emily M. D'agostino, Keith Marsolo, Lisa M. Wruck, Laura Johnson, James Topping, Al Richmond, Giselle Corbie, Warren A. Kibbe |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Evaluating fitness-for-use of electronic health records in pragmatic clinical trials: reported practices and recommendationsabstractOBJECTIVE: To empirically explore how pragmatic clinical trials (PCTs) that used real-world data (RWD) assessed study-specific fitness-for-use. METHODS: We conducted interviews and surveys with PCT teams who used electronic health record (EHR) data to ascertain endpoints. The survey cataloged key concerns about RWD, activities used to assess data fitness-for-use, and related barriers encountered by study teams. Patterns and commonalities across trials were used to develop recommendations for study-specific fitness-for-use assessments. RESULTS: Of 15 invited trial teams, 7 interviews were conducted. Of 31 invited trials, 15 responded to the survey. Most respondents had prior experience using RWD (93%). Major concerns about EHR data were data reliability, missingness or incompleteness of EHR elements, variation in data quality across study sites, and presence of implausible or incorrect values. Although many PCTs conducted fitness-for-use activities (eg, data quality assessments, 11/14, 79%), less than a quarter did so before choosing a data source. Fitness-for-use activities, findings, and resulting study design changes were not often publically documented. Overall costs and personnel costs were barriers to fitness-for-use assessments. DISCUSSION: These results support three recommendations for PCTs that use EHR data for endpoint ascertainment. Trials should detail the rationale and plan for study-specific fitness-for-use activities, conduct study-specific fitness-for-use assessments early in the prestudy phase to inform study design changes before the trial begins, and share results of fitness-for-use assessments and description of relevant challenges and facilitators. CONCLUSION: These recommendations can help researchers and end-users of real-world evidence improve characterization of RWD reliability and relevance in the PCT-specific context. Sudha R. Raman, Emily C. O'Brien, Bradley G. Hammill, Adam J. Nelson, Laura J. Fish, Lesley H. Curtis, Keith Marsolo |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Reliable insights regarding medication outcomes from real world data- key principles of an evidence generation framework
Rishi J. Desai, Sebastian Schneeweiss, Keith Marsolo, Shirley V. Wang, Robert Ball |
AMIA | 3 |
| 2021 | U.S. COVID-19 Surveillance in PCORnet®
Sheryl A. Kluberg, Thomas Carton, L. Charles Bailey, Julia A. Fearrington, Keith Marsolo, Kshema M. Nagavedu, Jon Puro, Jason P. Block |
AMIA | 5 |
| 2021 | Enhancing the use of EHR systems for pragmatic embedded research: lessons from the NIH Health Care Systems Research CollaboratoryabstractOBJECTIVE: We identified challenges and solutions to using electronic health record (EHR) systems for the design and conduct of pragmatic research. MATERIALS AND METHODS: Since 2012, the Health Care Systems Research Collaboratory has served as the resource coordinating center for 21 pragmatic clinical trial demonstration projects. The EHR Core working group invited these demonstration projects to complete a written semistructured survey and used an inductive approach to review responses and identify EHR-related challenges and suggested EHR enhancements. RESULTS: We received survey responses from 20 projects and identified 21 challenges that fell into 6 broad themes: (1) inadequate collection of patient-reported outcome data, (2) lack of structured data collection, (3) data standardization, (4) resources to support customization of EHRs, (5) difficulties aggregating data across sites, and (6) accessing EHR data. DISCUSSION: Based on these findings, we formulated 6 prerequisites for PCTs that would enable the conduct of pragmatic research: (1) integrate the collection of patient-centered data into EHR systems, (2) facilitate structured research data collection by leveraging standard EHR functions, usable interfaces, and standard workflows, (3) support the creation of high-quality research data by using standards, (4) ensure adequate IT staff to support embedded research, (5) create aggregate, multidata type resources for multisite trials, and (6) create re-usable and automated queries. CONCLUSION: We are hopeful our collection of specific EHR challenges and research needs will drive health system leaders, policymakers, and EHR designers to support these suggestions to improve our national capacity for generating real-world evidence. Rachel L. Richesson, Keith Marsolo, Brian J. Douthit, Karen L. Staman, P. Michael Ho, Dana L. Dailey, Andrew D. Boyd, Kathleen McTigue, Miriam O. Ezenwa, Judith M. Schlaeger, Crystal L. Patil, Keturah R. Faurot, Leah Tuzzio, Eric B. Larson, Emily C. O'Brien, Christina K. Zigler, Joshua R. Lakin, Alice R. Pressman, Jordan M. Braciszewski, Corita R. Grudzen, Guilherme Del Fiol |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Design and analytic considerations for using patient-reported health data in pragmatic clinical trials: report from an NIH Collaboratory roundtableabstractPragmatic clinical trials often entail the use of electronic health record (EHR) and claims data, but bias and quality issues associated with these data can limit their fitness for research purposes particularly for study end points. Patient-reported health (PRH) data can be used to confirm or supplement EHR and claims data in pragmatic trials, but these data can bring their own biases. Moreover, PRH data can complicate analyses if they are discordant with other sources. Using experience in the design and conduct of multi-site pragmatic trials, we itemize the strengths and limitations of PRH data and identify situational criteria for determining when PRH data are appropriate or ideal to fill gaps in the evidence collected from EHRs. To provide guidance for the scientific rationale and appropriate use of patient-reported data in pragmatic clinical trials, we describe approaches for ascertaining and classifying study end points and addressing issues of incomplete data, data alignment, and concordance. We conclude by identifying areas that require more research. Frank W. Rockhold, Jessica D. Tenenbaum, Rachel L. Richesson, Keith Marsolo, Emily C. O'Brien |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Advancing the Collection and Integration of Patient-reported Outcome Data Implementation Architectures Using FHIR® Technical Specifications
Chun-Ju Hsiao, Stephanie Garcia, Daniella Meeker, Joseph Blumenthal, Keith Marsolo |
AMIA | 5 |
| 2018 | National Patient-Centered Clinical Research Network (PCORnet) Distributed Research Network (DRN): Questions Answered from Preparatory-to-Research to Comparative Effectiveness
Jessica L. Sturtevant, Darcy Louzao, Keith Marsolo, Kathleen McTigue, Lesley H. Curtis |
AMIA | 3 |
| 2018 | Empowering genomic medicine by establishing critical sequencing result data flows: the eMERGE exampleabstractThe eMERGE Network is establishing methods for electronic transmittal of patient genetic test results from laboratories to healthcare providers across organizational boundaries. We surveyed the capabilities and needs of different network participants, established a common transfer format, and implemented transfer mechanisms based on this format. The interfaces we created are examples of the connectivity that must be instantiated before electronic genetic and genomic clinical decision support can be effectively built at the point of care. This work serves as a case example for both standards bodies and other organizations working to build the infrastructure required to provide better electronic clinical decision support for clinicians. Samuel J. Aronson, Lawrence J. Babb, Darren C. Ames, Richard A. Gibbs, Eric Venner, John J. Connelly, Keith Marsolo, Chunhua Weng, Marc S. Williams, Andrea L. Hartzler, Wayne H. Liang, James D. Ralston, Emily Beth Devine, Shawn N. Murphy, Christopher G. Chute, Pedro J. Caraballo, Iftikhar J. Kullo, Robert R. Freimuth, Luke V. Rasmussen, Firas H. Wehbe, Josh F. Peterson, Jamie R. Robinson, Ken Wiley, Casey Overby Taylor |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | Accuracy of the medication list in the electronic health record - implications for care, research, and improvementabstractObjective: Electronic medication lists may be useful in clinical decision support and research, but their accuracy is not well described. Our aim was to assess the completeness of the medication list compared to the clinical narrative in the electronic health record. Methods: We reviewed charts of 30 patients with inflammatory bowel disease (IBD) from each of 6 gastroenterology centers. Centers compared IBD medications from the medication list to the clinical narrative. Results: We reviewed 379 IBD medications among 180 patients. There was variation by center, from 90% patients with complete agreement between the medication list and clinical narrative to 50% agreement. Conclusions: There was a range in the accuracy of the medication list compared to the clinical narrative. This information may be helpful for sites seeking to improve data quality and those seeking to use medication list data for research or clinical decision support. Kathleen E. Walsh, Keith Marsolo, Cori Davis, Theresa Todd, Bernadette Martineau, Carlie Arbaugh, Frederique Verly, Charles Samson, Peter A. Margolis |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | The PCORnet Learning Cycle
Keith Marsolo, Laura G. Qualls, Bradley G. Hammill, Jeffrey S. Brown, Lesley H. Curtis |
AMIA | 1 |
| 2017 | A longitudinal analysis of data quality in a large pediatric data research networkabstractOBJECTIVE: PEDSnet is a clinical data research network (CDRN) that aggregates electronic health record data from multiple children's hospitals to enable large-scale research. Assessing data quality to ensure suitability for conducting research is a key requirement in PEDSnet. This study presents a range of data quality issues identified over a period of 18 months and interprets them to evaluate the research capacity of PEDSnet. MATERIALS AND METHODS: Results were generated by a semiautomated data quality assessment workflow. Two investigators reviewed programmatic data quality issues and conducted discussions with the data partners' extract-transform-load analysts to determine the cause for each issue. RESULTS: The results include a longitudinal summary of 2182 data quality issues identified across 9 data submission cycles. The metadata from the most recent cycle includes annotations for 850 issues: most frequent types, including missing data (>300) and outliers (>100); most complex domains, including medications (>160) and lab measurements (>140); and primary causes, including source data characteristics (83%) and extract-transform-load errors (9%). DISCUSSION: The longitudinal findings demonstrate the network's evolution from identifying difficulties with aligning the data to a common data model to learning norms in clinical pediatrics and determining research capability. CONCLUSION: While data quality is recognized as a critical aspect in establishing and utilizing a CDRN, the findings from data quality assessments are largely unpublished. This paper presents a real-world account of studying and interpreting data quality findings in a pediatric CDRN, and the lessons learned could be used by other CDRNs. Ritu Khare, Levon Utidjian, Byron Ruth, Michael G. Kahn, Evanette Burrows, Keith Marsolo, Nandan Patibandla, Hanieh Razzaghi, Ryan Colvin, Daksha Ranade, Melody Kitzmiller, Daniel Eckrich, L. Charles Bailey |
J. Am. Medical Informatics Assoc. | 6 |
| 2017 | Biases introduced by filtering electronic health records for patients with "complete data"abstractOBJECTIVE: One promise of nationwide adoption of electronic health records (EHRs) is the availability of data for large-scale clinical research studies. However, because the same patient could be treated at multiple health care institutions, data from only a single site might not contain the complete medical history for that patient, meaning that critical events could be missing. In this study, we evaluate how simple heuristic checks for data "completeness" affect the number of patients in the resulting cohort and introduce potential biases. MATERIALS AND METHODS: We began with a set of 16 filters that check for the presence of demographics, laboratory tests, and other types of data, and then systematically applied all 216 possible combinations of these filters to the EHR data for 12 million patients at 7 health care systems and a separate payor claims database of 7 million members. RESULTS: EHR data showed considerable variability in data completeness across sites and high correlation between data types. For example, the fraction of patients with diagnoses increased from 35.0% in all patients to 90.9% in those with at least 1 medication. An unrelated claims dataset independently showed that most filters select members who are older and more likely female and can eliminate large portions of the population whose data are actually complete. DISCUSSION AND CONCLUSION: As investigators design studies, they need to balance their confidence in the completeness of the data with the effects of placing requirements on the data on the resulting patient cohort. Griffin M. Weber, William G. Adams, Elmer V. Bernstam, Jonathan P. Bickel, Kathe P. Fox, Keith Marsolo, Vijay A. Raghavan, Alexander Turchin, Shawn N. Murphy, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 6 |
| 2015 | Identifying and Understanding Data Quality Issues in a Pediatric Distributed Research Network
Ritu Khare, Levon Utidjian, Gregory Schulte, Keith Marsolo, L. Charles Bailey |
AMIA | 4 |
| 2014 | Brief communication: PEDSnet: a National Pediatric Learning Health SystemabstractA learning health system (LHS) integrates research done in routine care settings, structured data capture during every encounter, and quality improvement processes to rapidly implement advances in new knowledge, all with active and meaningful patient participation. While disease-specific pediatric LHSs have shown tremendous impact on improved clinical outcomes, a national digital architecture to rapidly implement LHSs across multiple pediatric conditions does not exist. PEDSnet is a clinical data research network that provides the infrastructure to support a national pediatric LHS. A consortium consisting of PEDSnet, which includes eight academic medical centers, two existing disease-specific pediatric networks, and two national data partners form the initial partners in the National Pediatric Learning Health System (NPLHS). PEDSnet is implementing a flexible dual data architecture that incorporates two widely used data models and national terminology standards to support multi-institutional data integration, cohort discovery, and advanced analytics that enable rapid learning. Christopher B. Forrest, Peter A. Margolis, L. Charles Bailey, Keith Marsolo, Mark A. Del Beccaro, Jonathan A. Finkelstein, David E. Milov, Veronica J. Vieland, Bryan A. Wolf, Feliciano B. Yu, Michael G. Kahn |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Brief communication: Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS): ArchitectureabstractWe describe the architecture of the Patient Centered Outcomes Research Institute (PCORI) funded Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS, http://www.SCILHS.org) clinical data research network, which leverages the $48 billion dollar federal investment in health information technology (IT) to enable a queryable semantic data model across 10 health systems covering more than 8 million patients, plugging universally into the point of care, generating evidence and discovery, and thereby enabling clinician and patient participation in research during the patient encounter. Central to the success of SCILHS is development of innovative 'apps' to improve PCOR research methods and capacitate point of care functions such as consent, enrollment, randomization, and outreach for patient-reported outcomes. SCILHS adapts and extends an existing national research network formed on an advanced IT infrastructure built with open source, free, modular components. Kenneth D. Mandl, Isaac S. Kohane, Douglas MacFadden, Griffin M. Weber, Marc D. Natter, Joshua C. Mandel, Sebastian Schneeweiss, Sarah Weiler, Jeffrey G. Klann, Jonathan P. Bickel, William G. Adams, Yaorong Ge, James Perkins, Keith Marsolo, Elmer V. Bernstam, John Showalter, Alexander Quarshie, Elizabeth O. Ofili, George Hripcsak, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 15 |
| 2014 | Preparing an annotated gold standard corpus to share with extramural investigators for de-identification researchabstractOBJECTIVE: The current study aims to fill the gap in available healthcare de-identification resources by creating a new sharable dataset with realistic Protected Health Information (PHI) without reducing the value of the data for de-identification research. By releasing the annotated gold standard corpus with Data Use Agreement we would like to encourage other Computational Linguists to experiment with our data and develop new machine learning models for de-identification. This paper describes: (1) the modifications required by the Institutional Review Board before sharing the de-identification gold standard corpus; (2) our efforts to keep the PHI as realistic as possible; (3) and the tests to show the effectiveness of these efforts in preserving the value of the modified data set for machine learning model development. MATERIALS AND METHODS: In a previous study we built an original de-identification gold standard corpus annotated with true Protected Health Information (PHI) from 3503 randomly selected clinical notes for the 22 most frequent clinical note types of our institution. In the current study we modified the original gold standard corpus to make it suitable for external sharing by replacing HIPAA-specified PHI with newly generated realistic PHI. Finally, we evaluated the research value of this new dataset by comparing the performance of an existing published in-house de-identification system, when trained on the new de-identification gold standard corpus, with the performance of the same system, when trained on the original corpus. We assessed the potential benefits of using the new de-identification gold standard corpus to identify PHI in the i2b2 and PhysioNet datasets that were released by other groups for de-identification research. We also measured the effectiveness of the i2b2 and PhysioNet de-identification gold standard corpora in identifying PHI in our original clinical notes. RESULTS: Performance of the de-identification system using the new gold standard corpus as a training set was very close to training on the original corpus (92.56 vs. 93.48 overall F-measures). Best i2b2/PhysioNet/CCHMC cross-training performances were obtained when training on the new shared CCHMC gold standard corpus, although performances were still lower than corpus-specific trainings. DISCUSSION AND CONCLUSION: We successfully modified a de-identification dataset for external sharing while preserving the de-identification research value of the modified gold standard corpus with limited drop in machine learning de-identification performance. Louise Deléger, Todd Lingren, Yizhao Ni, Megan Kaiser, Laura Stoutenborough, Keith Marsolo, Michal Kouril, Katalin Molnár, Imre Solti |
J. Biomed. Informatics | 6 |
| 2013 | Large-scale evaluation of automated clinical note de-identification and its impact on information extractionabstractOBJECTIVE: (1) To evaluate a state-of-the-art natural language processing (NLP)-based approach to automatically de-identify a large set of diverse clinical notes. (2) To measure the impact of de-identification on the performance of information extraction algorithms on the de-identified documents. MATERIAL AND METHODS: A cross-sectional study that included 3503 stratified, randomly selected clinical notes (over 22 note types) from five million documents produced at one of the largest US pediatric hospitals. Sensitivity, precision, F value of two automated de-identification systems for removing all 18 HIPAA-defined protected health information elements were computed. Performance was assessed against a manually generated 'gold standard'. Statistical significance was tested. The automated de-identification performance was also compared with that of two humans on a 10% subsample of the gold standard. The effect of de-identification on the performance of subsequent medication extraction was measured. RESULTS: The gold standard included 30 815 protected health information elements and more than one million tokens. The most accurate NLP method had 91.92% sensitivity (R) and 95.08% precision (P) overall. The performance of the system was indistinguishable from that of human annotators (annotators' performance was 92.15%(R)/93.95%(P) and 94.55%(R)/88.45%(P) overall while the best system obtained 92.91%(R)/95.73%(P) on same text). The impact of automated de-identification was minimal on the utility of the narrative notes for subsequent information extraction as measured by the sensitivity and precision of medication name extraction. DISCUSSION AND CONCLUSION: NLP-based de-identification shows excellent performance that rivals the performance of human annotators. Furthermore, unlike manual de-identification, the automated approach scales up to millions of documents quickly and inexpensively. Louise Deléger, Katalin Molnár, Guergana K. Savova, Fei Xia 0004, Todd Lingren, Qi Li 0004, Keith Marsolo, Anil G. Jegga, Megan Kaiser, Laura Stoutenborough, Imre Solti |
J. Am. Medical Informatics Assoc. | 7 |
| 2013 | Informatics and operations - let's get integratedabstractThe widespread adoption of commercial electronic health records (EHRs) presents a significant challenge to the field of informatics. In their current form, EHRs function as a walled garden and prevent the integration of outside tools and services. This impedes the widespread adoption and diffusion of research interventions into the clinic. In most institutions, EHRs are supported by clinical operations staff who are largely separate from their informatics counterparts. This relationship needs to change. Research informatics and clinical operations need to work more closely on the implementation and configuration of EHRs to ensure that they are used to collect high-quality data for research and improvement at the point of care. At the same time, the informatics community needs to lobby commercial EHR vendors to open their systems and design new architectures that allow for the integration of external applications and services. Keith Marsolo |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | An i2b2-based, generalizable, open source, self-scaling chronic disease registryabstractOBJECTIVE: Registries are a well-established mechanism for obtaining high quality, disease-specific data, but are often highly project-specific in their design, implementation, and policies for data use. In contrast to the conventional model of centralized data contribution, warehousing, and control, we design a self-scaling registry technology for collaborative data sharing, based upon the widely adopted Integrating Biology & the Bedside (i2b2) data warehousing framework and the Shared Health Research Information Network (SHRINE) peer-to-peer networking software. MATERIALS AND METHODS: Focusing our design around creation of a scalable solution for collaboration within multi-site disease registries, we leverage the i2b2 and SHRINE open source software to create a modular, ontology-based, federated infrastructure that provides research investigators full ownership and access to their contributed data while supporting permissioned yet robust data sharing. We accomplish these objectives via web services supporting peer-group overlays, group-aware data aggregation, and administrative functions. RESULTS: The 56-site Childhood Arthritis & Rheumatology Research Alliance (CARRA) Registry and 3-site Harvard Inflammatory Bowel Diseases Longitudinal Data Repository now utilize i2b2 self-scaling registry technology (i2b2-SSR). This platform, extensible to federation of multiple projects within and between research networks, encompasses >6000 subjects at sites throughout the USA. DISCUSSION: We utilize the i2b2-SSR platform to minimize technical barriers to collaboration while enabling fine-grained control over data sharing. CONCLUSIONS: The implementation of i2b2-SSR for the multi-site, multi-stakeholder CARRA Registry has established a digital infrastructure for community-driven research data sharing in pediatric rheumatology in the USA. We envision i2b2-SSR as a scalable, reusable solution facilitating interdisciplinary research across diseases. Marc D. Natter, Justin Quan, David M. Ortiz, Athos Bousvaros, Norman T. Ilowite, Christi J. Inman, Keith Marsolo, Andrew J. McMurry, Christy Sandborg, Laura E. Schanberg, Carol A. Wallace, Robert W. Warren, Griffin M. Weber, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 7 |
| 2012 | Building Gold Standard Corpora for Medical Natural Language Processing Tasks
Louise Deléger, Qi Li 0004, Todd Lingren, Megan Kaiser, Katalin Molnár, Laura Stoutenborough, Michal Kouril, Keith Marsolo, Imre Solti |
AMIA | 8 |
| 2012 | Challenges in creating an opt-in biobank with a registrar-based consent process and a commercial EHRabstractResidual clinical samples represent a very appealing source of biomaterial for translational and clinical research. We describe the implementation of an opt-in biobank, with consent being obtained at the time of registration and the decision stored in our electronic health record, Epic. Information on that decision, along with laboratory data, is transferred to an application that signals to biobank staff whether a given sample can be kept for research. Investigators can search for samples using our i2b2 data warehouse. Patient participation has been overwhelmingly positive and much higher than anticipated. Over 86% of patients provided consent and almost 83% requested to be notified of any incidental research findings. In 6 months, we obtained decisions from over 18 000 patients and processed 8000 blood samples for storage in our research biobank. However, commercial electronic health records like Epic lack key functionality required by a registrar-based consent process, although workarounds exist. Keith Marsolo, Jeremy Corsmo, Michael G. Barnes, Carrie Pollick, Jamie Chalfin, Jeremy Nix, Rajesh Ganta |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | On the use of structure and sequence-based features for protein classification and retrieval
Keith Marsolo, Srinivasan Parthasarathy 0001 |
Knowl. Inf. Syst. | 1 |
| 2007 | Spatial Modeling and Classification of Corneal ShapeabstractOne of the most promising applications of data mining is in biomedical data used in patient diagnosis. Any method of data analysis intended to support the clinical decision-making process should meet several criteria: it should capture clinically relevant features, be computationally feasible, and provide easily interpretable results. In an initial study, we examined the feasibility of using Zernike polynomials to represent biomedical instrument data in conjunction with a decision tree classifier to distinguish between the diseased and non-diseased eyes. Here, we provide a comprehensive follow-up to that work, examining a second representation, pseudo-Zernike polynomials, to determine whether they provide any increase in classification accuracy. We compare the fidelity of both methods using residual root-mean-square (rms) error and evaluate accuracy using several classifiers: neural networks, C4.5 decision trees, Voting Feature Intervals, and Naïve Bayes. We also examine the effect of several meta-learning strategies: boosting, bagging, and Random Forests (RFs). We present results comparing accuracy as it relates to dataset and transformation resolution over a larger, more challenging, multi-class dataset. They show that classification accuracy is similar for both data transformations, but differs by classifier. We find that the Zernike polynomials provide better feature representation than the pseudo-Zernikes and that the decision trees yield the best balance of classification accuracy and interpretability. Keith Marsolo, Michael D. Twa, Mark Bullimore, Srinivasan Parthasarathy 0001 |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2006 | Structure-based querying of proteins using waveletsabstractThe ability to retrieve molecules based on structural similarity has use in many applications, from disease diagnosis and treatment to drug discovery and design. In this paper, we present a method to represent protein molecules that allows for the fast, flexible and efficient retrieval of similar structures, based on either global or local attributes. We begin by computing the pair-wise distance between amino acids, transforming each 3D structure into a 2D distance matrix. We normalize this matrix to a specific size and apply a 2D wavelet decomposition to generate a set of approximation coefficients, which serves as our global feature vector. This transformation reduces the overall dimensionality of the data while still preserving spatial features and correlations. We test our method by running queries on three different protein data sets that have been used previously in the literature, basing our comparisons on labels taken from the SCOP database. We find that our method significantly outperforms existing approaches, in terms of retrieval accuracy, memory utilization and execution time. Specifically, using a k-d tree and running a 10-nearest-neighbor search on a dataset of 33,000 proteins against itself, we see an average accuracy of 89% at the SCOP SuperFamily level and a total query time that is up to 350 times faster than previously published techniques. In addition to processing queries based on global similarity, we also propose innovative extensions to effectively match proteins based solely on shared local substructures, allowing for a more flexible query interface. Keith Marsolo, Srinivasan Parthasarathy 0001, Kotagiri Ramamohanarao |
CIKM | 1 |
| 2006 | On the Use of Structure and Sequence-Based Features for Protein Classification and RetrievalabstractThe need to retrieve or classify protein molecules using structure or sequence-based similarity measures underlies a wide range of biomedical applications. In drug discovery, researchers search for proteins that share specific chemical properties as possible sources for new treatment. With folding simulations, similar intermediate structures might be indicative of a common folding pathway. To derive any type of similarity, however, one must have an effective model of the protein that allows for easy comparison. In this work, we present two normalized, stand-alone representations of proteins that enable fast and efficient object retrieval based on sequence or structure. To create our sequence-based representation, we take the frequency and scoring matrices returned by the PSTBIAST alignment algorithm and create a normalized summary using a discrete wavelet transform. Our structural descriptor is constructed using an algorithm we developed previously. First, we transform each 3D structure into a 2D distance matrix by calculating the pair-wise distance between the amino acids of a protein. We normalize this matrix and apply a 2D wavelet decomposition to generate a set of approximation coefficients, which serve as our feature vector. We also concatenate the sequence and structural descriptors together to create a hybrid solution. We evaluate the generality of our models by using them as database indices for nearest-neighbor and range-based retrieval experiments as well as feature vectors for classification using support vector machines. We find that our methods provide excellent performance when compared with the current state-of-the-art techniques of each task. Our results show that the sequence-based representation is on par with, or out-performs, the structure-based representation. Moreover, we find that in the classification context, the hybrid strategy affords a significant improvement over sequence or structure. Keith Marsolo, Srinivasan Parthasarathy 0001 |
ICDM | 1 |
| 2005 | A Model-Based Approach to Visualizing Classification Decisions for Patient Diagnosis
Keith Marsolo, Srinivasan Parthasarathy 0001, Michael D. Twa, Mark Bullimore |
AIME | 1 |
| 2005 | A Multi-Level Approach to SCOP Fold RecognitionabstractThe classification of proteins based on their structure can play an important role in the deduction or discovery of protein function. However, the relatively low number of solved protein structures and the unknown relationship between structure and sequence requires an alternative method of representation for classification to be effective. Furthermore, the large number of potential folds causes problems for many classification strategies, increasing the likelihood that the classifier will reach a local optima while trying to distinguish between all of the possible structural categories. Here we present a hierarchical strategy for structural classification that first partitions proteins based on their SCOP class before attempting to assign a protein fold. Using a well-known dataset derived from the 27 most-populated SCOP folds and several sequence-based descriptor properties as input features, we test a number of classification methods, including Naive Bayes and Boosted C4.5. Our strategy achieves an average fold recognition of 74%, which is significantly higher than the 56-60% previously reported in the literature, indicating the effectiveness of a multi-level approach. Keith Marsolo, Srinivasan Parthasarathy 0001, Chris Ding |
BIBE | 1 |
| 2005 | Classification of Biomedical Data through Model-Based Spatial AveragingabstractEnsemble learning is frequently used to reduce classification error. The more popular techniques draw multiple samples from the training data and employ a voting procedure to aggregate the decisions of the classifiers constructed from those samples. In practice, such ensemble methods have been shown to work well and improve accuracy. Here we present a meta-learning strategy that combines the decisions of classifiers constructed from spatial models taken at multiple resolutions. By varying the resolution from coarse to fine-grained, we are able to partition the data on global features that describe a majority of the objects, as well as small, local features that are present in just a few problem cases. We test our technique on a biomedical dataset containing surface elevation values for diseased and nondiseased corneas. We transform these elevations into a series of coefficients using two different spatial transformations. Using these coefficients, we determine how well they distinguish between the two classes. We find our algorithm can increase the classification accuracy of a single decision tree up to 10% and can also be used in conjunction with traditional meta-learning techniques such as bagging to further improve performance. In an attempt to improve the execution time of the transformation algorithms, we have developed a distributed, grid-based implementation as well. Keith Marsolo, Srinivasan Parthasarathy 0001, Michael D. Twa, Mark Bullimore |
BIBE | 1 |
| 2005 | Alternate Representation of Distance Matrices for Characterization of Protein StructureabstractThe most suitable method for the automated classification of protein structures remains an open problem in computational biology. In order to classify a protein structure with any accuracy, an effective representation must be chosen. Here we present two methods of representing protein structure. One involves representing the distances between the C/sub a/ atoms of a protein as a two-dimensional matrix and creating a model of the resulting surface with Zernike polynomials. The second uses a wavelet-based approach. We convert the distances between a protein's C/sub a/ atoms into a one-dimensional signal which is then decomposed using a discrete wavelet transformation. Using the Zernike coefficients and the approximation coefficients of the wavelet decomposition as feature vectors, we test the effectiveness of our representation with two different classifiers on a dataset of more than 600 proteins taken from the 27 most-populated SCOP folds. We find that the wavelet decomposition greatly outperforms the Zernike model. With the wavelet representation, we achieve an accuracy of approximately 56%, roughly 12% higher than results reported on a similar, but less-challenging dataset. In addition, we can couple our structure-based feature vectors with several sequence-based properties to increase accuracy another 5-7%. Finally, we use a multi-stage classification strategy on the combined features to increase performance to 78%, an improvement in accuracy of more than 15-20% and 34% over the highest reported sequence-based and structure-based classification results, respectively. Keith Marsolo, Srinivasan Parthasarathy 0001 |
ICDM | 1 |