VLDB 2026 Research / reviewers in the wild / expert
Elmer V. Bernstam
dblp:33/2666 · also Elmer Victor Bernstam
· DBLP profile ↗
75ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0001-7643-791XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 74 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A conceptual model for provider scheduling: insights from an EHR implementationabstractOBJECTIVE: We present a set of definitions and a conceptual model to support primary and consulting physician schedule integration into the electronic health record (EHR) in an inpatient setting and show how utilization of this functionality supports patient-centered communication. MATERIALS AND METHODS: Our institution transitioned to the Epic EHR and implemented modules to connect primary and consulting provider schedules from external scheduling systems to secure messaging within an inpatient EHR context. We evaluated legacy functionality, met with provider groups to map their shifts to hospital teams, built a crosswalk tool to extract, transform, and load data from the scheduling systems to the EHR, evaluated the utilization, and assessed issues related to the implementation. We used our experience from the project to develop a set of definitions and a conceptual model for provider scheduling. RESULTS: We met with over 100 groups to map over 2000 shifts to nearly 700 teams across 15 facilities in our health system. Utilization was high with an average of 6500 on-call provider searches per day in the 30 days following implementation. The conceptual model for inpatient provider scheduling defines 11 terms. DISCUSSION: Our definitions and conceptual model sufficiently represent the inpatient provider scheduling domain as evidenced by high initial utilization and few reported defects. The standardized terminology for provider scheduling aids integration of scheduling data into the EHR. CONCLUSION: Our successful integration of real-time scheduling data within the EHR guided development of a provider scheduling conceptual model. Standardized provider scheduling terminology promotes interoperability of scheduling systems. Jeremy B. Hill, Elmer V. Bernstam, Licong Cui, Peter V. Killoran |
J. Am. Medical Informatics Assoc. | 2 |
| 2026 | Opportunities for informatics to improve patient experiences: observations and reflections of ACMI fellowsabstractOBJECTIVES: We report on findings from a meeting convened by the American College of Medical Informatics (ACMI) to characterize aspects of the patient experience that could be improved using informatics. MATERIALS AND METHODS: The American College of Medical Informatics fellows were invited to share their experiences as patients and suggest informatics approaches that may improve the patient experience. RESULTS: We identified 4 themes: (1) getting the right care, (2) data sharing and data interoperability, (3) guiding low-cost evaluations, and (4) predictive analytics. DISCUSSION: Despite widespread adoption of health IT, patient experiences remain far from optimal. CONCLUSION: The American College of Medical Informatics fellows identified informatics approaches, applications, and research areas that have the potential to improve patient experiences with health care systems. Howard R. Strasberg, Edward P. Hoffer, Ross Koppel, Kevin B. Johnson, William M. Tierney, Geoffrey W. Rutledge, Elmer V. Bernstam, Jos Aarts, Marion J. Ball, Douglas S. Bell, Bernd Blobel, Suzanne Boren, Iain E. Buchan, James J. Cimino, Lawrence M. Fagan, James Geller, María Adela Grando, David A. Hanauer, William R. Hogan, Andrew S. Kanter, Bonnie Kaplan, Casimir A. Kulikowski, Albert Lai, David McCallie, Vimla Patel, Wanda Pratt, Sarah Collins Rossetti, Edward H. Shortliffe, Hardeep Singh 0005, Dean F. Sittig, William W. Stead, Kim M. Unertl, Mark G. Weiner, Kai Zheng 0002 |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Review of tools to support Target Trial Emulation
Christina Van Hal, Elmer V. Bernstam, Todd R. Johnson |
J. Biomed. Informatics | 2 |
| 2024 | Understanding enterprise data warehouses to support clinical and translational research: impact, sustainability, demand management, and accessibilityabstractOBJECTIVES: Healthcare organizations, including Clinical and Translational Science Awards (CTSA) hubs funded by the National Institutes of Health, seek to enable secondary use of electronic health record (EHR) data through an enterprise data warehouse for research (EDW4R), but optimal approaches are unknown. In this qualitative study, our goal was to understand EDW4R impact, sustainability, demand management, and accessibility. MATERIALS AND METHODS: We engaged a convenience sample of informatics leaders from CTSA hubs (n = 21) for semi-structured interviews and completed a directed content analysis of interview transcripts. RESULTS: EDW4R have created institutional capacity for single- and multi-center studies, democratized access to EHR data for investigators from multiple disciplines, and enabled the learning health system. Bibliometrics have been challenging due to investigator non-compliance, but one hub's requirement to link all study protocols with funding records enabled quantifying an EDW4R's multi-million dollar impact. Sustainability of EDW4R has relied on multiple funding sources with a general shift away from the CTSA grant toward institutional and industry support. To address EDW4R demand, institutions have expanded staff, used different governance approaches, and provided investigator self-service tools. EDW4R accessibility can benefit from improved tools incorporating user-centered design, increased data literacy among scientists, expansion of informaticians in the workforce, and growth of team science. DISCUSSION: As investigator demand for EDW4R has increased, approaches to tracking impact, ensuring sustainability, and improving accessibility of EDW4R resources have varied. CONCLUSION: This study adds to understanding of how informatics leaders seek to support investigators using EDW4R across the CTSA consortium and potentially elsewhere. Thomas R. Campion Jr., Catherine K. Craven, David A. Dorr, Elmer V. Bernstam, Boyd M. Knosp |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Forecasting acute kidney injury and resource utilization in ICU patients using longitudinal, multimodal models
Yukun Tan, Merve Dede, Vakul Mohanty, Jinzhuang Dou, Holly Hill, Elmer V. Bernstam, Ken Chen 0001 |
J. Biomed. Informatics | 6 |
| 2023 | A deep learning approach to identify missing is-a relations in SNOMED CTabstractOBJECTIVE: SNOMED CT is the largest clinical terminology worldwide. Quality assurance of SNOMED CT is of utmost importance to ensure that it provides accurate domain knowledge to various SNOMED CT-based applications. In this work, we introduce a deep learning-based approach to uncover missing is-a relations in SNOMED CT. MATERIALS AND METHODS: Our focus is to identify missing is-a relations between concept-pairs exhibiting a containment pattern (ie, the set of words of one concept being a proper subset of that of the other concept). We use hierarchically related containment concept-pairs as positive instances and hierarchically unrelated containment concept-pairs as negative instances to train a model predicting whether an is-a relation exists between 2 concepts with containment pattern. The model is a binary classifier leveraging concept name features, hierarchical features, enriched lexical attribute features, and logical definition features. We introduce a cross-validation inspired approach to identify missing is-a relations among all hierarchically unrelated containment concept-pairs. RESULTS: We trained and applied our model on the Clinical finding subhierarchy of SNOMED CT (September 2019 US edition). Our model (based on the validation sets) achieved a precision of 0.8164, recall of 0.8397, and F1 score of 0.8279. Applying the model to predict actual missing is-a relations, we obtained a total of 1661 potential candidates. Domain experts performed evaluation on randomly selected 230 samples and verified that 192 (83.48%) are valid. CONCLUSIONS: The results showed that our deep learning approach is effective in uncovering missing is-a relations between containment concept-pairs in SNOMED CT. Rashmie Abeysinghe, Fengbo Zheng, Elmer V. Bernstam, Jay Shi, Olivier Bodenreider, Licong Cui |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Blockchain-enabled immutable, distributed, and highly available clinical research activity logging system for federated COVID-19 data analysis from multiple institutionsabstractOBJECTIVE: We aimed to develop a distributed, immutable, and highly available cross-cloud blockchain system to facilitate federated data analysis activities among multiple institutions. MATERIALS AND METHODS: We preprocessed 9166 COVID-19 Structured Query Language (SQL) code, summary statistics, and user activity logs, from the GitHub repository of the Reliable Response Data Discovery for COVID-19 (R2D2) Consortium. The repository collected local summary statistics from participating institutions and aggregated the global result to a COVID-19-related clinical query, previously posted by clinicians on a website. We developed both on-chain and off-chain components to store/query these activity logs and their associated queries/results on a blockchain for immutability, transparency, and high availability of research communication. We measured run-time efficiency of contract deployment, network transactions, and confirmed the accuracy of recorded logs compared to a centralized baseline solution. RESULTS: The smart contract deployment took 4.5 s on an average. The time to record an activity log on blockchain was slightly over 2 s, versus 5-9 s for baseline. For querying, each query took on an average less than 0.4 s on blockchain, versus around 2.1 s for baseline. DISCUSSION: The low deployment, recording, and querying times confirm the feasibility of our cross-cloud, blockchain-based federated data analysis system. We have yet to evaluate the system on a larger network with multiple nodes per cloud, to consider how to accommodate a surge in activities, and to investigate methods to lower querying time as the blockchain grows. CONCLUSION: Blockchain technology can be used to support federated data analysis among multiple institutions. Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim 0001, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Christian Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo Koola, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nicholas R. Anderson 0001, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu 0001, Yong K. Choi, Kai Zheng 0002, Yujia Zhou 0003, Rachel A Zucker |
J. Am. Medical Informatics Assoc. | 14 |
| 2023 | Why is biomedical informatics hard? A fundamental framework
Todd R. Johnson, Elmer V. Bernstam |
J. Biomed. Informatics | 2 |
| 2023 | Technical/Algorithm, Stakeholder, and Society (TASS) barriers to the application of artificial intelligence in medicine: A systematic review
Linda T. Li, Lauren C. Haley, Alexandra K. Boyd, Elmer V. Bernstam |
J. Biomed. Informatics | 4 |
| 2022 | ClinicalLayoutLM: A Pre-trained Multi-modal Model for Understanding Scanned Document in Electronic Health RecordsabstractScanned documents (e.g., faxes) are still widely used in clinical practice and are prevalent in Electronic Health Records (EHR). Unlocking information in scanned documents in EHRs is critical for clinical operation and research. However, it is challenging as it requires converting images to texts before applying information extraction technologies. Here we propose a multi-modal approach (ClinicalLayoutLM) that jointly models text extracted from Optical Character Recognition (OCR) and layout/image information to classify scanned clinical documents into different categories (e.g., lab reports and CT scans). Using a clinical corpus of 348, 311 scanned documents, we continually pretrained ClinicalLayoutLM based on LayoutLMv3, a multi-modal model from the open domain. For the task to classify the scanned clinical documents into 16 categories, ClinicalLayoutLM achieved an F1 score of 0.9051, which outperformed the baseline model (0.8840) that was based on text from OCR only. ClinicalLayoutLM is the first of its kind of multi-modal models for clinical documents and we believe it could benefit other clinical natural language processing (NLP) tasks such as layout analysis, information extraction and so on. The code is available at https://github.com/UTHealth-CCB/ClinicalLayoutLM and the pre-trained model is available upon request. Qiang Wei 0002, Xu Zuo, Omer Anjum, Ryan Denlinger, Elmer V. Bernstam, Martin J. Citardi, Hua Xu 0001 |
IEEE Big Data | 6 |
| 2022 | Quantitating and assessing interoperability between electronic health recordsabstractOBJECTIVES: Electronic health records (EHRs) contain a large quantity of machine-readable data. However, institutions choose different EHR vendors, and the same product may be implemented differently at different sites. Our goal was to quantify the interoperability of real-world EHR implementations with respect to clinically relevant structured data. MATERIALS AND METHODS: We analyzed de-identified and aggregated data from 68 oncology sites that implemented 1 of 5 EHR vendor products. Using 6 medications and 6 laboratory tests for which well-accepted standards exist, we calculated inter- and intra-EHR vendor interoperability scores. RESULTS: The mean intra-EHR vendor interoperability score was 0.68 as compared to a mean of 0.22 for inter-system interoperability, when weighted by number of systems of each type, and 0.57 and 0.20 when not weighting by number of systems of each type. DISCUSSION: In contrast to data elements required for successful billing, clinically relevant data elements are rarely standardized, even though applicable standards exist. We chose a representative sample of laboratory tests and medications for oncology practices, but our set of data elements should be seen as an example, rather than a definitive list. CONCLUSIONS: We defined and demonstrated a quantitative measure of interoperability between site EHR systems and within/between implemented vendor systems. Two sites that share the same vendor are, on average, more interoperable. However, even for implementation of the same EHR product, interoperability is not guaranteed. Our results can inform institutional EHR selection, analysis, and optimization for interoperability. Elmer V. Bernstam, Jeremy L. Warner, John C. Krauss, Edward P. Ambinder, Wendy S. Rubinstein, George Komatsoulis, Robert S. Miller, James L. Chen |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Understanding enterprise data warehouses to support clinical and translational research: enterprise information technology relationships, data governance, workforce, and cloud computingabstractOBJECTIVE: Among National Institutes of Health Clinical and Translational Science Award (CTSA) hubs, effective approaches for enterprise data warehouses for research (EDW4R) development, maintenance, and sustainability remain unclear. The goal of this qualitative study was to understand CTSA EDW4R operations within the broader contexts of academic medical centers and technology. MATERIALS AND METHODS: We performed a directed content analysis of transcripts generated from semistructured interviews with informatics leaders from 20 CTSA hubs. RESULTS: Respondents referred to services provided by health system, university, and medical school information technology (IT) organizations as "enterprise information technology (IT)." Seventy-five percent of respondents stated that the team providing EDW4R service at their hub was separate from enterprise IT; strong relationships between EDW4R teams and enterprise IT were critical for success. Managing challenges of EDW4R staffing was made easier by executive leadership support. Data governance appeared to be a work in progress, as most hubs reported complex and incomplete processes, especially for commercial data sharing. Although nearly all hubs (n = 16) described use of cloud computing for specific projects, only 2 hubs reported using a cloud-based EDW4R. Respondents described EDW4R cloud migration facilitators, barriers, and opportunities. DISCUSSION: Descriptions of approaches to how EDW4R teams at CTSA hubs work with enterprise IT organizations, manage workforces, make decisions about data, and approach cloud computing provide insights for institutions seeking to leverage patient data for research. CONCLUSION: Identification of EDW4R best practices is challenging, and this study helps identify a breadth of viable options for CTSA hubs to consider when implementing EDW4R services. Boyd M. Knosp, Catherine K. Craven, David A. Dorr, Elmer V. Bernstam, Thomas R. Campion Jr. |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Closing the loop: automatically identifying abnormal imaging results in scanned documentsabstractOBJECTIVES: Scanned documents (SDs), while common in electronic health records and potentially rich in clinically relevant information, rarely fit well with clinician workflow. Here, we identify scanned imaging reports requiring follow-up with high recall and practically useful precision. MATERIALS AND METHODS: We focused on identifying imaging findings for 3 common causes of malpractice claims: (1) potentially malignant breast (mammography) and (2) lung (chest computed tomography [CT]) lesions and (3) long-bone fracture (X-ray) reports. We train our ClinicalBERT-based pipeline on existing typed/dictated reports classified manually or using ICD-10 codes, evaluate using a test set of manually classified SDs, and compare against string-matching (baseline approach). RESULTS: A total of 393 mammograms, 305 chest CT, and 683 bone X-ray reports were manually reviewed. The string-matching approach had an F1 of 0.667. For mammograms, chest CTs, and bone X-rays, respectively: models trained on manually classified training data and optimized for F1 reached an F1 of 0.900, 0.905, and 0.817, while separate models optimized for recall achieved a recall of 1.000 with precisions of 0.727, 0.518, and 0.275. Models trained on ICD-10-labelled data and optimized for F1 achieved F1 scores of 0.647, 0.830, and 0.643, while those optimized for recall achieved a recall of 1.0 with precisions of 0.407, 0.683, and 0.358. DISCUSSION: Our pipeline can identify abnormal reports with potentially useful performance and so decrease the manual effort required to screen for abnormal findings that require follow-up. CONCLUSION: It is possible to automatically identify clinically significant abnormalities in SDs with high recall and practically useful precision in a generalizable and minimally laborious way. Akshat Kumar, Heath Goodrum, Ashley Kim, Carly Stender, Kirk Roberts, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborativeabstractOBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require. Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai |
J. Am. Medical Informatics Assoc. | 28 |
| 2021 | Understanding Enterprise Data Warehouses to Support Clinical and Translational Research: Initial Findings on Enterprise Information Technology Relationships, Data Governance, Workforce, and Cloud Computing
Boyd M. Knosp, Catherine K. Craven, David A. Dorr, Elmer V. Bernstam, Thomas R. Campion Jr. |
AMIA | 4 |
| 2021 | Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance
Scott A. Malec, Elmer V. Bernstam, Richard D. Boyce, Trevor Cohen |
J. Biomed. Informatics | 3 |
| 2021 | Generalized and transferable patient language representation for phenotyping with limited data
Yuqi Si, Elmer V. Bernstam, Kirk Roberts |
J. Biomed. Informatics | 2 |
| 2020 | SCOR: A secure international informatics infrastructure to investigate COVID-19abstractGlobal pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale. Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux |
J. Am. Medical Informatics Assoc. | 7 |
| 2020 | Recommendations for patient similarity classes: results of the AMIA 2019 workshop on defining patient similarityabstractDefining patient-to-patient similarity is essential for the development of precision medicine in clinical care and research. Conceptually, the identification of similar patient cohorts appears straightforward; however, universally accepted definitions remain elusive. Simultaneously, an explosion of vendors and published algorithms have emerged and all provide varied levels of functionality in identifying patient similarity categories. To provide clarity and a common framework for patient similarity, a workshop at the American Medical Informatics Association 2019 Annual Meeting was convened. This workshop included invited discussants from academics, the biotechnology industry, the FDA, and private practice oncology groups. Drawing from a broad range of backgrounds, workshop participants were able to coalesce around 4 major patient similarity classes: (1) feature, (2) outcome, (3) exposure, and (4) mixed-class. This perspective expands into these 4 subtypes more critically and offers the medical informatics community a means of communicating their work on this important topic. Nathan D. Seligson, Jeremy L. Warner, William S. Dalton, Robert S. Miller, Debra Patt, Kenneth L. Kehl, Matvey Palchuk, Gil Alterovitz, Laura K. Wiley, Ming Huang 0006, Feichen Shen, Yanshan Wang, Khoa A. Nguyen, Anthony F. Wong, Funda Meric-Bernstam, Elmer V. Bernstam, James L. Chen |
J. Am. Medical Informatics Assoc. | 17 |
| 2020 | Predict or draw blood: An integrated method to reduce lab tests
Lishan Yu, Qiuchen Zhang, Elmer V. Bernstam, Xiaoqian Jiang |
J. Biomed. Informatics | 3 |
| 2019 | Continuum of Interoperability in Oncology EHR Implementations
Elmer V. Bernstam, Jeremy L. Warner, John C. Krauss, Edward P. Ambinder, Wendy S. Rubinstein, George Komatsoulis, Robert S. Miller, James L. Chen |
AMIA | 1 |
| 2019 | A frame semantic overview of NLP-based information extraction for cancer-related EHR notes: a scoping review
Surabhi Datta, Elmer V. Bernstam, Kirk Roberts |
AMIA | 2 |
| 2019 | A federated EHR network data completeness tracking systemabstractOBJECTIVE: The study sought to design, pilot, and evaluate a federated data completeness tracking system (CTX) for assessing completeness in research data extracted from electronic health record data across the Accessible Research Commons for Health (ARCH) Clinical Data Research Network. MATERIALS AND METHODS: The CTX applies a systems-based approach to design workflow and technology for assessing completeness across distributed electronic health record data repositories participating in a queryable, federated network. The CTX invokes 2 positive feedback loops that utilize open source tools (DQe-c and Vue) to integrate technology and human actors in a system geared for increasing capacity and taking action. A pilot implementation of the system involved 6 ARCH partner sites between January 2017 and May 2018. RESULTS: The ARCH CTX has enabled the network to monitor and, if needed, adjust its data management processes to maintain complete datasets for secondary use. The system allows the network and its partner sites to profile data completeness both at the network and partner site levels. Interactive visualizations presenting the current state of completeness in the context of the entire network as well as changes in completeness across time were valued among the CTX user base. DISCUSSION: Distributed clinical data networks are complex systems. Top-down approaches that solely rely on technology to report data completeness may be necessary but not sufficient for improving completeness (and quality) of data in large-scale clinical data networks. Improving and maintaining complete (high-quality) data in such complex environments entails sociotechnical systems that exploit technology and empower human actors to engage in the process of high-quality data curating. CONCLUSIONS: The CTX has increased the network's capacity to rapidly identify data completeness issues and empowered ARCH partner sites to get involved in improving the completeness of respective data in their repositories. Hossein Estiri, Jeffrey G. Klann, Sarah Weiler, Ernest Alema-Mensah, R. Joseph Applegate, Galina Lozinski, Nandan Patibandla, William G. Adams, Marc D. Natter, Elizabeth O. Ofili, Brian Ostasiewski, Alexander Quarshie, Gary E. Rosenthal, Elmer V. Bernstam, Kenneth D. Mandl, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 15 |
| 2019 | Bioinformatics service center projects go beyond service
Jeffrey T. Chang, David E. Volk, David G. Gorenstein, David Steffen, Elmer V. Bernstam |
J. Biomed. Informatics | 5 |
| 2019 | A frame semantic overview of NLP-based information extraction for cancer-related EHR notes
Surabhi Datta, Elmer V. Bernstam, Kirk Roberts |
J. Biomed. Informatics | 2 |
| 2019 | Rapamycin-mTOR+BRAF=? Using relational similarity to find therapeutically relevant drug-gene relationships in unstructured text
Safa Fathiamini, Amber M. Johnson, Vijaykumar Holla, Nora S. Sanchez, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen |
J. Biomed. Informatics | 7 |
| 2018 | A FrameNet for Cancer Information in Clinical Narratives: Schema and Annotation
Kirk Roberts, Yuqi Si, Anshul Gandhi, Elmer V. Bernstam |
LREC | 4 |
| 2017 | Biases introduced by filtering electronic health records for patients with "complete data"abstractOBJECTIVE: One promise of nationwide adoption of electronic health records (EHRs) is the availability of data for large-scale clinical research studies. However, because the same patient could be treated at multiple health care institutions, data from only a single site might not contain the complete medical history for that patient, meaning that critical events could be missing. In this study, we evaluate how simple heuristic checks for data "completeness" affect the number of patients in the resulting cohort and introduce potential biases. MATERIALS AND METHODS: We began with a set of 16 filters that check for the presence of demographics, laboratory tests, and other types of data, and then systematically applied all 216 possible combinations of these filters to the EHR data for 12 million patients at 7 health care systems and a separate payor claims database of 7 million members. RESULTS: EHR data showed considerable variability in data completeness across sites and high correlation between data types. For example, the fraction of patients with diagnoses increased from 35.0% in all patients to 90.9% in those with at least 1 medication. An unrelated claims dataset independently showed that most filters select members who are older and more likely female and can eliminate large portions of the population whose data are actually complete. DISCUSSION AND CONCLUSION: As investigators design studies, they need to balance their confidence in the completeness of the data with the effects of placing requirements on the data on the resulting patient cohort. Griffin M. Weber, William G. Adams, Elmer V. Bernstam, Jonathan P. Bickel, Kathe P. Fox, Keith Marsolo, Vijay A. Raghavan, Alexander Turchin, Shawn N. Murphy, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | Literature-Based Discovery of Confounding in Observational Clinical Data
Scott A. Malec, Hua Xu 0001, Elmer V. Bernstam, Sahiti Myneni, Trevor Cohen |
AMIA | 4 |
| 2016 | Extracting genetic alteration information for personalized cancer therapy from ClinicalTrials.govabstractOBJECTIVE: Clinical trials investigating drugs that target specific genetic alterations in tumors are important for promoting personalized cancer therapy. The goal of this project is to create a knowledge base of cancer treatment trials with annotations about genetic alterations from ClinicalTrials.gov. METHODS: We developed a semi-automatic framework that combines advanced text-processing techniques with manual review to curate genetic alteration information in cancer trials. The framework consists of a document classification system to identify cancer treatment trials from ClinicalTrials.gov and an information extraction system to extract gene and alteration pairs from the Title and Eligibility Criteria sections of clinical trials. By applying the framework to trials at ClinicalTrials.gov, we created a knowledge base of cancer treatment trials with genetic alteration annotations. We then evaluated each component of the framework against manually reviewed sets of clinical trials and generated descriptive statistics of the knowledge base. RESULTS AND DISCUSSION: The automated cancer treatment trial identification system achieved a high precision of 0.9944. Together with the manual review process, it identified 20 193 cancer treatment trials from ClinicalTrials.gov. The automated gene-alteration extraction system achieved a precision of 0.8300 and a recall of 0.6803. After validation by manual review, we generated a knowledge base of 2024 cancer trials that are labeled with specific genetic alteration information. Analysis of the knowledge base revealed the trend of increased use of targeted therapy for cancer, as well as top frequent gene-alteration pairs of interest. We expect this knowledge base to be a valuable resource for physicians and patients who are seeking information about personalized cancer therapy. Jun Xu 0007, Hee-Jin Lee, Yonghui Wu 0001, Yaoyun Zhang, Liang-Chin Huang, Amber M. Johnson, Vijaykumar Holla, Ann M. Bailey, Trevor Cohen, Funda Meric-Bernstam, Elmer V. Bernstam, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 12 |
| 2016 | Automated identification of molecular effects of drugs (AIMED)abstractINTRODUCTION: Genomic profiling information is frequently available to oncologists, enabling targeted cancer therapy. Because clinically relevant information is rapidly emerging in the literature and elsewhere, there is a need for informatics technologies to support targeted therapies. To this end, we have developed a system for Automated Identification of Molecular Effects of Drugs, to help biomedical scientists curate this literature to facilitate decision support. OBJECTIVES: To create an automated system to identify assertions in the literature concerning drugs targeting genes with therapeutic implications and characterize the challenges inherent in automating this process in rapidly evolving domains. METHODS: We used subject-predicate-object triples (semantic predications) and co-occurrence relations generated by applying the SemRep Natural Language Processing system to MEDLINE abstracts and ClinicalTrials.gov descriptions. We applied customized semantic queries to find drugs targeting genes of interest. The results were manually reviewed by a team of experts. RESULTS: Compared to a manually curated set of relationships, recall, precision, and F2 were 0.39, 0.21, and 0.33, respectively, which represents a 3- to 4-fold improvement over a publically available set of predications (SemMedDB) alone. Upon review of ostensibly false positive results, 26% were considered relevant additions to the reference set, and an additional 61% were considered to be relevant for review. Adding co-occurrence data improved results for drugs in early development, but not their better-established counterparts. CONCLUSIONS: Precision medicine poses unique challenges for biomedical informatics systems that help domain experts find answers to their research questions. Further research is required to improve the performance of such systems, particularly for drugs in development. Safa Fathiamini, Amber M. Johnson, Alejandro Araya, Vijaykumar Holla, Ann M. Bailey, Beate Litzenburger, Nora S. Sanchez, Yekaterina Khotskaya, Hua Xu 0001, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen |
J. Am. Medical Informatics Assoc. | 12 |
| 2016 | Precision medicine informaticsabstractThis special issue on precision medicine informatics flowed from the AMIA 2015 Translational Bioinformatics Summit theme of “Accelerating Precision Medicine”1 and President Obama’s 2015 State of the Union call “to give all of us access to the personalized information we need to keep ourselves and our families healthier.”2 The goal is to focus on the inherent translational informatics challenges, concerns, and opportunities afforded by precision medicine to provide an accurate, personalized characterization of patient populations based on various characteristics including molecular (eg, genomic, proteomic), clinical (eg, comorbidities), environmental exposures, lifestyle, patient preferences, and other information.4–13 Informatics is a necessary component in a comprehensive initiative to tackle precision medicine. Its roles include (1) managing big data,15–17 (2) creating learning systems for knowledge generation,18–21 (3) providing access for individual involvement,22–25 and (4) ultimately supporting optimal delivery of precision treatments derived from translational research.26–29 The papers submitted for this special issue cover these 4 areas and expand on the themes of informatics emerging in the field of precision medicine. Data become “big” when the scale of the data and analyses challenge our traditional systems in size, velocity, or complexity.8 The large scale of the Million Veterans Program is an example, with systems and analysis methods being developed for application in a repository that has established policies and procedures. The Million Veterans Program can serve as a model for involving thousands of individuals in precision medicine research.30 Similarly, the Precision Medicine Initiative Cohort Program seeks to connect electronic health records with participant-provided data, molecular determinants, environment, and lifestyle patterns to deeply impact our knowledge of health and therapies.31 The ambitious goal of enrolling 1 million volunteers whose demographics reflect the diversity of the US population can only be accomplished with robust and scalable informatics. Tenenbaum et al.3,32 present an informatics perspective and describe key informatics innovations required to advance precision medicine. Big data can also be characterized by the velocity of data, such as in real-time data collection. We have come to expect nearly instant access and tracking of information that improves our quality of life.33–35 With precision medicine, the informatics research community is responding to the challenge and is providing applications relevant to the health of individuals. The Substitutable Medical Applications, Reusable Technologies (SMART) application by Warner et al.36 provides the ability to visualize genomic information for oncology treatment using mobile technology. This can facilitate individualized communication about cancer treatment options in the context of genomic information. Personalizing treatment using the vast volume of available molecular data requires tools that reduce cognitive load.14,37 The tools described by Xu et al.38 create a knowledge base of 2024 genomically informed clinical trials and treatments to support the delivery of personalized cancer therapy. Fathiamini et al.,39 Hintzsche et al.,40 and Cheng et al.41 describe resources that identify variants relevant to therapeutic treatments, variant calls, and natural language processing approaches to create structured information resources. Informatics has the potential to enable the collection, analysis, and reuse of data to facilitate learning knowledge representations for a wide range of diseases using phenotypes, genotypes, and predictive analytics.42–51 Halpern et al.52 present an approach to derive computable phenotypes using machine learning techniques that reduce the need to manually create phenotypes, which can be expensive and time-consuming. The use of machine learning techniques to establish disease-mutation relationships is described by Singhal et al.53 Rioth et al.54 describe automatic parsing and categorizing of molecular profiles gathered in routine care into a database that can inform research and clinical practice. Hoffman et al.55 describe guidelines for incorporating pharmacogenetic tests into clinical decision support systems. The integration of epidemiological evidence into knowledge representations that can inform care is explored by Torosyan et al.56 The integration of information from a wide variety of sources combined with learning systems promises to speed discovery. With patients and physicians making decisions based on genetic tests that identify actionable risks, a patient’s individual attitudes can influence how the information is accepted and used. Strategies and frameworks that address engagement and disparities will need to be developed.57–62 Dye et al.63 show that some minority populations are less likely to want genetic testing or participate in genetic research, which has the potential to exacerbate existing disparities. Precision medicine is focused on the individual patient. Given that molecular measurements are increasingly guiding care, studies will be needed to understand how to improve acceptance and participation by minorities. Adams and Petersen64 discuss ethical, legal, and social issues that can arise when large numbers of individuals participate in genetic research. A framework for addressing ethical, legal, and social challenges can facilitate trust while protecting the individual and advancing research. Ni et al.65 describe a learning framework to predict patient involvement in clinical trials with the goal of improving the effectiveness of recruitment. The evaluation of methods to match genomics to therapeutics is an active area of research.66–74 Eubank et al.75 developed an informatics solution to manage genetic/genomic and clinical data to expedite clinical trials of targeted cancer therapies at Memorial Sloan Kettering Cancer Center. As of August 2015, the system contained data on 159 893 patients, with 64 473 being tracked across 134 research cohorts with 51 192 patients in a genotype-matched eligible pool. The system can serve as a model for cohort programs in cancer clinical trials. Patients will want to know the outcomes for other patients in similar situations, and Warner et al.76 provide a system that tracks clinical outcomes over time. Access to such information is vital to informed decision makers who are involved in keeping themselves and their families healthy. The direct relationship between clinical trials and knowledge representations is demonstrated through a knowledge base for proper cancer drug selection and a The Cancer Genome Atlas (TCGA) analysis of triple-negative breast cancer, where 71.7% of 85 cancer patients had Food and Drug Administration (FDA)-approved “drug-able” genomic targets.77 Informatics tools will be required for precision medicine to both catalyze research on huge populations and enable implementation of precision care now done for isolated cases in routine care. With the powerful response of the translational informatics community to the request for papers for this precision medicine special focus issue, we can chart the main themes of the course ahead. The initial response is the beginning of a transition in health care that is supported through informatics to propel participating individuals into the center of research and care. The needs of such an interactive model of care are broad and will need to be further defined as the initial systems are tested and measure how, when, and where they improve the health and outcomes of individuals and families. The work of L.J.F. was supported in part by National Institutes of Health (NIH) grants 1R01GM108346-01 and U54-GM104941, Health Equity and Rural Outreach Innovation Center grant CIN 13-418, and funding from the Hollings Cancer Center’s support grant P30 CA138313 at the Medical University of South Carolina. The work of E.V.B. was supported by National Cancer Institute (NCI) U01 CA180964, the Sheikh Bin Zayed Al Nahyan Foundation, the Cancer Prevention Research Institute of Texas Precision Oncology Decision Support Core RP150535, National Center for Advancing Translational Sciences (NCATS) grant UL1 TR000371 (Center for Clinical and Translational Sciences), the Bosarge Foundation, and an MD Anderson Cancer Center support grant (NCI P30 CA016672). The work of J.C.D. was supported by R01 LM 010685, R01 GM 103859, and R01 GM 105688 from the NIH. Lewis J. Frey, Elmer V. Bernstam, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 2 |
| 2016 | Improving the utility of MeSH® terms using the TopicalMeSH representation
Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
J. Biomed. Informatics | 2 |
| 2015 | BigMouth: Development of a Scalable Infrastructure to Support Multi Institutional Data Sharing for Dentistry
Krishna Kookal Kumar, Duong Tran, Reuben J. Applegate, Heiko Spallek, Elsbeth Kalenderian, Joel M. White, Elmer V. Bernstam, Muhammad F. Walji |
AMIA | 7 |
| 2015 | Improving Retrieval of PubMed Articles Using the TopicalMeSH Representation
Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
AMIA | 2 |
| 2015 | Patient-centered care, collaboration, communication, and coordination: a report from AMIA's 2013 Policy MeetingabstractIn alignment with a major shift toward patient-centered care as the model for improving care in our health system, informatics is transforming patient-provider relationships and overall care delivery. AMIA's 2013 Health Policy Invitational was focused on examining existing challenges surrounding full engagement of the patient and crafting a research agenda and policy framework encouraging the use of informatics solutions to achieve this goal. The group tackled this challenge from educational, technical, and research perspectives. Recommendations include the need for consumer education regarding rights to data access, the need for consumers to access their health information in real time, and further research on effective methods to engage patients. This paper summarizes the meeting as well as the research agenda and policy recommendations prioritized among the invited experts and stakeholders. Patricia Flatley Brennan, Rupa Valdez, Gregory L. Alexander, Shifali Arora, Elmer V. Bernstam, Margo Edmunds, Nikolai Kirienko, Ross D. Martin, Ida Sim, Diane J. Skiba, S. Trent Rosenbloom |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Expert guided natural language processing using one-class classificationabstractINTRODUCTION: Automatically identifying specific phenotypes in free-text clinical notes is critically important for the reuse of clinical data. In this study, the authors combine expert-guided feature (text) selection with one-class classification for text processing. OBJECTIVES: To compare the performance of one-class classification to traditional binary classification; to evaluate the utility of feature selection based on expert-selected salient text (snippets); and to determine the robustness of these models with respects to irrelevant surrounding text. METHODS: The authors trained one-class support vector machines (1C-SVMs) and two-class SVMs (2C-SVMs) to identify notes discussing breast cancer. Manually annotated visit summary notes (88 positive and 88 negative for breast cancer) were used to compare the performance of models trained on whole notes labeled as positive or negative to models trained on expert-selected text sections (snippets) relevant to breast cancer status. Model performance was evaluated using a 70:30 split for 20 iterations and on a realistic dataset of 10 000 records with a breast cancer prevalence of 1.4%. RESULTS: When tested on a balanced experimental dataset, 1C-SVMs trained on snippets had comparable results to 2C-SVMs trained on whole notes (F = 0.92 for both approaches). When evaluated on a realistic imbalanced dataset, 1C-SVMs had a considerably superior performance (F = 0.61 vs. F = 0.17 for the best performing model) attributable mainly to improved precision (p = .88 vs. p = .09 for the best performing model). CONCLUSIONS: 1C-SVMs trained on expert-selected relevant text sections perform better than 2C-SVMs classifiers trained on either snippets or whole notes when applied to realistically imbalanced data with low prevalence of the positive class. Erel Joffe, Emily J. Pettigrew, Jorge R. Herskovic, Charles F. Bearden, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 5 |
| 2014 | Exploring the Use of SemRep Predications to Help Identify Secondary Drug Targets for Personalized Cancer Therapy
Safa Fathiamini, Amber M. Johnson, Vijaykumar Holla, Ann M. Bailey, Lauren Brusco, Funda Meric-Bernstam, Elmer V. Bernstam, Trevor Cohen |
AMIA | 8 |
| 2014 | A benchmark comparison of deterministic and probabilistic methods for defining manual review datasets in duplicate records reconciliationabstractINTRODUCTION: Clinical databases require accurate entity resolution (ER). One approach is to use algorithms that assign questionable cases to manual review. Few studies have compared the performance of common algorithms for such a task. Furthermore, previous work has been limited by a lack of objective methods for setting algorithm parameters. We compared the performance of common ER algorithms: using algorithmic optimization, rather than manual parameter tuning, and on two-threshold classification (match/manual review/non-match) as well as single-threshold (match/non-match). METHODS: We manually reviewed 20,000 randomly selected, potential duplicate record-pairs to identify matches (10,000 training set, 10,000 test set). We evaluated the probabilistic expectation maximization, simple deterministic and fuzzy inference engine (FIE) algorithms. We used particle swarm to optimize algorithm parameters for a single and for two thresholds. We ran 10 iterations of optimization using the training set and report averaged performance against the test set. RESULTS: The overall estimated duplicate rate was 6%. FIE and simple deterministic algorithms allowed a lower manual review set compared to the probabilistic method (FIE 1.9%, simple deterministic 2.5%, probabilistic 3.6%; p<0.001). For a single threshold, the simple deterministic algorithm performed better than the probabilistic method (positive predictive value 0.956 vs 0.887, sensitivity 0.985 vs 0.887, p<0.001). ER with FIE classifies 98.1% of record-pairs correctly (1/10,000 error rate), assigning the remainder to manual review. CONCLUSIONS: Optimized deterministic algorithms outperform the probabilistic method. There is a strong case for considering optimized deterministic methods for ER. Erel Joffe, Michael J. Byrne, Phillip Reeder, Jorge R. Herskovic, Craig W. Johnson, Allison B. McCoy, Dean F. Sittig, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 8 |
| 2014 | Brief communication: Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS): ArchitectureabstractWe describe the architecture of the Patient Centered Outcomes Research Institute (PCORI) funded Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS, http://www.SCILHS.org) clinical data research network, which leverages the $48 billion dollar federal investment in health information technology (IT) to enable a queryable semantic data model across 10 health systems covering more than 8 million patients, plugging universally into the point of care, generating evidence and discovery, and thereby enabling clinician and patient participation in research during the patient encounter. Central to the success of SCILHS is development of innovative 'apps' to improve PCOR research methods and capacitate point of care functions such as consent, enrollment, randomization, and outreach for patient-reported outcomes. SCILHS adapts and extends an existing national research network formed on an advanced IT infrastructure built with open source, free, modular components. Kenneth D. Mandl, Isaac S. Kohane, Douglas MacFadden, Griffin M. Weber, Marc D. Natter, Joshua C. Mandel, Sebastian Schneeweiss, Sarah Weiler, Jeffrey G. Klann, Jonathan P. Bickel, William G. Adams, Yaorong Ge, James Perkins, Keith Marsolo, Elmer V. Bernstam, John Showalter, Alexander Quarshie, Elizabeth O. Ofili, George Hripcsak, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 16 |
| 2014 | BigMouth: a multi-institutional dental data repositoryabstractFew oral health databases are available for research and the advancement of evidence-based dentistry. In this work we developed a centralized data repository derived from electronic health records (EHRs) at four dental schools participating in the Consortium of Oral Health Research and Informatics. A multi-stakeholder committee developed a data governance framework that encouraged data sharing while allowing control of contributed data. We adopted the i2b2 data warehousing platform and mapped data from each institution to a common reference terminology. We realized that dental EHRs urgently need to adopt common terminologies. While all used the same treatment code set, only three of the four sites used a common diagnostic terminology, and there were wide discrepancies in how medical and dental histories were documented. BigMouth was successfully launched in August 2012 with data on 1.1 million patients, and made available to users at the contributing institutions. Muhammad F. Walji, Elsbeth Kalenderian, Paul C. Stark, Joel M. White, Krishna Kookal Kumar, Dat Phan, Duong Tran, Elmer V. Bernstam, Rachel Badovinac Ramoni |
J. Am. Medical Informatics Assoc. | 8 |
| 2013 | Informatics Infrastructure for Routine Personalized Medicine
Elmer V. Bernstam, Ken Chen 0001, Hua Xu 0001, Funda Meric-Bernstam |
AMIA | 1 |
| 2013 | Optimized Dual Threshold Entity Resolution For Electronic Health Record Databases - Training Set Size And Active Learning
Erel Joffe, Michael J. Byrne, Phillip Reeder, Jorge R. Herskovic, Craig W. Johnson, Allison B. McCoy, Elmer V. Bernstam |
AMIA | 7 |
| 2013 | A Systematic Yet Flexible Systems Analysis Framework
Eliz Markowitz, Todd R. Johnson, Elmer V. Bernstam, Jorge R. Herskovic, Harold W. Thimbleby |
AMIA | 3 |
| 2013 | Twinlist: Novel User Interface Designs for Medication Reconciliation
Catherine Plaisant, Tiffany Chao, Johnny Wu, A. Zachary Hettinger, Jorge R. Herskovic, Todd R. Johnson, Elmer V. Bernstam, Eliz Markowitz, Seth Powsner, Ben Shneiderman |
AMIA | 7 |
| 2013 | Using Early Quizzes to Predict Student Outcomes in Online Introductory Biomedical Informatics Courses
Irmgard Willcockson, Jorge R. Herskovic, Melanie A. Sutton, Robert E. Hoyt, Craig W. Johnson, Todd R. Johnson, Elmer V. Bernstam |
AMIA | 7 |
| 2013 | Healthcare information technology and economicsabstractAt the 2011 American College of Medical Informatics (ACMI) Winter Symposium we studied the overlap between health IT and economics and what leading healthcare delivery organizations are achieving today using IT that might offer paths for the nation to follow for using health IT in healthcare reform. We recognized that health IT by itself can improve health value, but its main contribution to health value may be that it can make possible new care delivery models to achieve much larger value. Health IT is a critically important enabler to fundamental healthcare system changes that may be a way out of our current, severe problem of rising costs and national deficit. We review the current state of healthcare costs, federal health IT stimulus programs, and experiences of several leading organizations, and offer a model for how health IT fits into our health economic future. Thomas H. Payne, David W. Bates, Eta S. Berner, Elmer V. Bernstam, H. Dominic Covvey, Mark E. Frisse, Thomas Graf, Robert A. Greenes, Edward P. Hoffer, Gilad J. Kuperman, Harold P. Lehmann, Louise Liang, Blackford Middleton, Gilbert S. Omenn, Judy G. Ozbolt |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | SYFSA: A framework for Systematic Yet Flexible Systems Analysis
Todd R. Johnson, Eliz Markowitz, Elmer V. Bernstam, Jorge R. Herskovic, Harold W. Thimbleby |
J. Biomed. Informatics | 3 |
| 2012 | Duplicate Patient Records - Implication for Missed Laboratory Results
Erel Joffe, Charles F. Bearden, Michael J. Byrne, Elmer V. Bernstam |
AMIA | 4 |
| 2012 | Deterministic Binary Vectors for Efficient Automated Indexing of MEDLINE/PubMed Abstracts
Manuel Wahle, Dominic Widdows, Jorge R. Herskovic, Elmer V. Bernstam, Trevor Cohen |
AMIA | 4 |
| 2012 | Graph-based signal integration for high-throughput phenotypingabstractBACKGROUND: Electronic Health Records aggregated in Clinical Data Warehouses (CDWs) promise to revolutionize Comparative Effectiveness Research and suggest new avenues of research. However, the effectiveness of CDWs is diminished by the lack of properly labeled data. We present a novel approach that integrates knowledge from the CDW, the biomedical literature, and the Unified Medical Language System (UMLS) to perform high-throughput phenotyping. In this paper, we automatically construct a graphical knowledge model and then use it to phenotype breast cancer patients. We compare the performance of this approach to using MetaMap when labeling records. RESULTS: MetaMap's overall accuracy at identifying breast cancer patients was 51.1% (n=428); recall=85.4%, precision=26.2%, and F1=40.1%. Our unsupervised graph-based high-throughput phenotyping had accuracy of 84.1%; recall=46.3%, precision=61.2%, and F1=52.8%. CONCLUSIONS: We conclude that our approach is a promising alternative for unsupervised high-throughput phenotyping. Jorge R. Herskovic, Devika Subramanian, Trevor Cohen, Pamela A. Bozzo-Silva, Charles F. Bearden, Elmer V. Bernstam |
BMC Bioinform. | 6 |
| 2012 | Focus on information retrieval: Predicting biomedical document access as a function of past useabstractOBJECTIVE: To determine whether past access to biomedical documents can predict future document access. MATERIALS AND METHODS: The authors used 394 days of query log (August 1, 2009 to August 29, 2010) from PubMed users in the Texas Medical Center, which is the largest medical center in the world. The authors evaluated two document access models based on the work of Anderson and Schooler. The first is based on how frequently a document was accessed. The second is based on both frequency and recency. RESULTS: The model based only on frequency of past access was highly correlated with the empirical data (R²=0.932), whereas the model based on frequency and recency had a much lower correlation (R²=0.668). DISCUSSION: The frequency-only model accurately predicted whether a document will be accessed based on past use. Modeling accesses as a function of frequency requires storing only the number of accesses and the creation date for the document. This model requires low storage overheads and is computationally efficient, making it scalable to large corpora such as MEDLINE. CONCLUSION: It is feasible to accurately model the probability of a document being accessed in the future based on past accesses. J. Caleb Goodwin, Todd R. Johnson, Trevor Cohen, Jorge R. Herskovic, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Collaborative knowledge acquisition for the design of context-aware alert systemsabstractOBJECTIVE: To present a framework for combining implicit knowledge acquisition from multiple experts with machine learning and to evaluate this framework in the context of anemia alerts. MATERIALS AND METHODS: Five internal medicine residents reviewed 18 anemia alerts, while 'talking aloud'. They identified features that were reviewed by two or more physicians to determine appropriate alert level, etiology and treatment recommendation. Based on these features, data were extracted from 100 randomly-selected anemia cases for a training set and an additional 82 cases for a test set. Two staff internists assigned an alert level, etiology and treatment recommendation before and after reviewing the entire electronic medical record. The training set of 118 cases (100 plus 18) and the test set of 82 cases were explored using RIDOR and JRip algorithms. RESULTS: The feature set was sufficient to assess 93% of anemia cases (intraclass correlation for alert level before and after review of the records by internists 1 and 2 were 0.92 and 0.95, respectively). High-precision classifiers were constructed to identify low-level alerts (precision p=0.87, recall R=0.4), iron deficiency (p=1.0, R=0.73), and anemia associated with kidney disease (p=0.87, R=0.77). DISCUSSION: It was possible to identify low-level alerts and several conditions commonly associated with chronic anemia. This approach may reduce the number of clinically unimportant alerts. The study was limited to anemia alerts. Furthermore, clinicians were aware of the study hypotheses potentially biasing their evaluation. CONCLUSION: Implicit knowledge acquisition, collaborative filtering and machine learning were combined automatically to induce clinically meaningful and precise decision rules. Erel Joffe, Ofer Havakuk, Jorge R. Herskovic, Vimla L. Patel, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | The SMART Platform: early experience enabling substitutable applications for electronic health recordsabstractOBJECTIVE: The Substitutable Medical Applications, Reusable Technologies (SMART) Platforms project seeks to develop a health information technology platform with substitutable applications (apps) constructed around core services. The authors believe this is a promising approach to driving down healthcare costs, supporting standards evolution, accommodating differences in care workflow, fostering competition in the market, and accelerating innovation. MATERIALS AND METHODS: The Office of the National Coordinator for Health Information Technology, through the Strategic Health IT Advanced Research Projects (SHARP) Program, funds the project. The SMART team has focused on enabling the property of substitutability through an app programming interface leveraging web standards, presenting predictable data payloads, and abstracting away many details of enterprise health information technology systems. Containers--health information technology systems, such as electronic health records (EHR), personally controlled health records, and health information exchanges that use the SMART app programming interface or a portion of it--marshal data sources and present data simply, reliably, and consistently to apps. RESULTS: The SMART team has completed the first phase of the project (a) defining an app programming interface, (b) developing containers, and (c) producing a set of charter apps that showcase the system capabilities. A focal point of this phase was the SMART Apps Challenge, publicized by the White House, using http://www.challenge.gov website, and generating 15 app submissions with diverse functionality. CONCLUSION: Key strategic decisions must be made about the most effective market for further disseminating SMART: existing market-leading EHR vendors, new entrants into the EHR market, or other stakeholders such as health information exchanges. Kenneth D. Mandl, Joshua C. Mandel, Shawn N. Murphy, Elmer V. Bernstam, Rachel Badovinac Ramoni, David A. Kreda, J. Michael McCoy, Ben Adida, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | Cross-terminology mapping challenges: A demonstration using medication terminological systems
Himali Saitwal, David Qing, Elmer V. Bernstam, Christopher G. Chute, Todd R. Johnson |
J. Biomed. Informatics | 4 |
| 2012 | Erratum to "Cross-terminology mapping challenges: A demonstration using medication terminological systems" [J. Biomed. Inform. (2012) 613-625]
Himali Saitwal, David Qing, Elmer V. Bernstam, Christopher G. Chute, Todd R. Johnson |
J. Biomed. Informatics | 4 |
| 2010 | Unintended consequences of health information technology: A need for biomedical informatics
Elmer V. Bernstam, William R. Hersh, Ida Sim, David Eichmann, Jonathan C. Silverstein, Jack W. Smith, Michael J. Becich |
J. Biomed. Informatics | 1 |
| 2010 | What is biomedical informatics?
Elmer V. Bernstam, Jack W. Smith, Todd R. Johnson |
J. Biomed. Informatics | 1 |
| 2009 | Ontology driven integration platform for clinical and translational researchabstractSemantic Web technologies offer a promising framework for integration of disparate biomedical data. In this paper we present the semantic information integration platform under development at the Center for Clinical and Translational Sciences (CCTS) at the University of Texas Health Science Center at Houston (UTHSC-H) as part of our Clinical and Translational Science Award (CTSA) program. We utilize the Semantic Web technologies not only for integrating, repurposing and classification of multi-source clinical data, but also to construct a distributed environment for information sharing, and collaboration online. Service Oriented Architecture (SOA) is used to modularize and distribute reusable services in a dynamic and distributed environment. Components of the semantic solution and its overall architecture are described. Parsa Mirhaji, Mattew Vagnoni, Elmer V. Bernstam, Jack W. Smith |
BMC Bioinform. | 4 |
| 2009 | Research Paper: Predictors of Student Success in Graduate Biomedical Informatics Training: Introductory Course and Program SuccessabstractOBJECTIVE: To predict student performance in an introductory graduate-level biomedical informatics course from application data. DESIGN: A predictive model built through retrospective review of student records using hierarchical binary logistic regression with half of the sample held back for cross-validation. The model was also validated against student data from a similar course at a second institution. MEASUREMENTS: Earning an A grade (Mastery) or a C grade (Failure) in an introductory informatics course. RESULTS: The authors analyzed 129 student records at the University of Texas School of Health Information Sciences at Houston (SHIS) and 106 at Oregon Health and Science University Department of Medical Informatics and Clinical Epidemiology (DMICE). In the SHIS cross-validation sample, the Graduate Record Exam verbal score (GRE-V) correctly predicted Mastery in 69.4%. Undergraduate grade point average (UGPA) and underrepresented minority status (URMS) predicted 81.6% of Failures. At DMICE, GRE-V, UGPA, and prior graduate degree significantly correlated with Mastery. Only GRE-V was a significant independent predictor of Mastery at both institutions. There were too few URMS students and Failures at DMICE to analyze. Course Mastery strongly predicted program performance defined as final cumulative GPA at SHIS (n=19, r=0.634, r2=0.40, p=0.0036) and DMICE (n=106, r=0.603, r2=0.36, p<0.001). CONCLUSIONS: The authors identified predictors of performance in an introductory informatics course including GRE-V, UGPA and URMS. Course performance was a very strong predictor of overall program performance. Findings may be useful for selecting students for admission and identifying students at risk for Failure as early as possible. Irmgard Willcockson, Craig W. Johnson, William R. Hersh, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 4 |
| 2009 | Sequential result refinement for searching the biomedical literature
Len Y. Tanaka, Jorge R. Herskovic, M. Sriram Iyengar, Elmer V. Bernstam |
J. Biomed. Informatics | 4 |
| 2007 | Research paper: A Day in the Life of PubMed: Analysis of a Typical Day's Query LogabstractOBJECTIVE: To characterize PubMed usage over a typical day and compare it to previous studies of user behavior on Web search engines. DESIGN: We performed a lexical and semantic analysis of 2,689,166 queries issued on PubMed over 24 consecutive hours on a typical day. MEASUREMENTS: We measured the number of queries, number of distinct users, queries per user, terms per query, common terms, Boolean operator use, common phrases, result set size, MeSH categories, used semantic measurements to group queries into sessions, and studied the addition and removal of terms from consecutive queries to gauge search strategies. RESULTS: The size of the result sets from a sample of queries showed a bimodal distribution, with peaks at approximately 3 and 100 results, suggesting that a large group of queries was tightly focused and another was broad. Like Web search engine sessions, most PubMed sessions consisted of a single query. However, PubMed queries contained more terms. CONCLUSION: PubMed's usage profile should be considered when educating users, building user interfaces, and developing future biomedical information retrieval systems. Jorge R. Herskovic, Len Y. Tanaka, William R. Hersh, Elmer V. Bernstam |
J. Am. Medical Informatics Assoc. | 4 |
| 2007 | Using hit curves to compare search algorithm performance
Jorge R. Herskovic, M. Sriram Iyengar, Elmer V. Bernstam |
J. Biomed. Informatics | 3 |
| 2006 | Research Paper: Using Citation Data to Improve Retrieval from MEDLINEabstractOBJECTIVE: To determine whether algorithms developed for the World Wide Web can be applied to the biomedical literature in order to identify articles that are important as well as relevant. DESIGN AND MEASUREMENTS A direct comparison of eight algorithms: simple PubMed queries, clinical queries (sensitive and specific versions), vector cosine comparison, citation count, journal impact factor, PageRank, and machine learning based on polynomial support vector machines. The objective was to prioritize important articles, defined as being included in a pre-existing bibliography of important literature in surgical oncology. RESULTS Citation-based algorithms were more effective than noncitation-based algorithms at identifying important articles. The most effective strategies were simple citation count and PageRank, which on average identified over six important articles in the first 100 results compared to 0.85 for the best noncitation-based algorithm (p < 0.001). The authors saw similar differences between citation-based and noncitation-based algorithms at 10, 20, 50, 200, 500, and 1,000 results (p < 0.001). Citation lag affects performance of PageRank more than simple citation count. However, in spite of citation lag, citation-based algorithms remain more effective than noncitation-based algorithms. CONCLUSION Algorithms that have proved successful on the World Wide Web can be applied to biomedical information retrieval. Citation-based algorithms can help identify important articles within large sets of relevant results. Further studies are needed to determine whether citation-based algorithms can effectively meet actual user information needs. Elmer V. Bernstam, Jorge R. Herskovic, Yindalon Aphinyanagphongs, Constantin F. Aliferis, Madurai G. Sriram, William R. Hersh |
J. Am. Medical Informatics Assoc. | 1 |
| 2005 | Using Incomplete Citation Data for MEDLINE Results Ranking
Jorge R. Herskovic, Elmer V. Bernstam |
AMIA | 2 |
| 2005 | Persuasive Email Messages for Patient Communication
Muhammad F. Walji, Kathy A. Johnson-Throop, Todd R. Johnson, Elmer V. Bernstam |
AMIA | 4 |
| 2002 | Comparing communication technology on Chinese, English, and Spanish diabetes web sites
Yanko Michea, Karen Pancheri, Yang Gong, Elmer V. Bernstam |
AMIA | 4 |
| 2002 | Evaluating the prevalence, content and readability of complementary and alternative medicine (CAM) web pages on the internet
Smitha Sagaram, Muhammad F. Walji, Elmer V. Bernstam |
AMIA | 3 |
| 2002 | Saving Lives in Real-time: A Pre-hospital Telementoring Case
R. Douglas Tindall, S. Ward Casscells, R. Matthew Sailors, Elmer V. Bernstam, Carolyn S. Galloway, James H. Duke |
AMIA | 4 |
| 2001 | MedlineQBE (Query-by-Example)
Elmer V. Bernstam |
AMIA | 1 |
| 2001 | Preliminary Evaluation of a Guideline Classification System
Elmer V. Bernstam, Nachman Ash, Mor Peleg, Samson W. Tu, Edward H. Shortliffe, Robert A. Greenes |
AMIA | 1 |
| 2001 | Modification of GEM to Accommodate Changes in a Classification Scheme
Richard Phillips, Elmer V. Bernstam, Kris Mork, Bryant Thomas Karras |
AMIA | 2 |
| 2001 | Sharable Representation of Clinical Guidelines in GLIF: Relationship to the Arden Syntax
Mor Peleg, Aziz A. Boxwala, Elmer V. Bernstam, Samson W. Tu, Robert A. Greenes, Edward H. Shortliffe |
J. Biomed. Informatics | 3 |
| 2000 | Guideline classification to assist modeling, authoring, implementation and retrieval
Elmer V. Bernstam, Nachman Ash, Mor Peleg, Samson W. Tu, Aziz A. Boxwala, Kris Mork, Edward H. Shortliffe, Robert A. Greenes |
AMIA | 1 |
| 2000 | GLIF3: the evolution of a guideline representation format
Mor Peleg, Aziz A. Boxwala, Omolola Ogunyemi, Qing T. Zeng, Samson W. Tu, Ronilda C. Lacson, Elmer V. Bernstam, Nachman Ash, Kris Mork, Lucila Ohno-Machado, Edward H. Shortliffe, Robert A. Greenes |
AMIA | 7 |