Maxim Topaz

dblp:139/1726 · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-2358-9837ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Automating infection indicator extraction in home healthcare through instruction-tuned large language models
abstract
OBJECTIVE: Home healthcare (HHC) clinical notes contain critical infection indicators that clinicians need in structured "indicator + context" pairs. Data sparsity and limited computing resources hinder automated extraction in decentralized HHC settings. This study developed and evaluated a resource-efficient pipeline using instruction-tuned, moderate-sized large language models (LLMs) to address these barriers. To address the data sparsity challenge, we also assessed the impact of a targeted LLM-based data augmentation strategy. MATERIALS AND METHODS: An expert-defined schema of 26 infection indicator categories was developed. We expanded the training set using a 3-stage workflow: targeted annotation, context mutation, and synthetic generation. We adapted 2 moderate-sized models (Gemma-12B and Qwen-14B) via Quantized Low-Rank Adaptation (QLoRA). We compared them to a larger-sized, prompted model and a smaller-sized, fully fine-tuned LLM. We evaluated all models on a held-out test set using partial micro-averaged F1 score, output reliability metrics, and qualitative error analysis. RESULTS: Instruction-tuned moderate-sized LLMs outperformed both baselines. The top-performing model, augmented Gemma-12B, achieved a partial micro-averaged F1 score of 0.879. LLM-based data augmentation enhanced overall performance, improving the identification of rare indicators and the interpretation of negations. The best model maintained a partial F1 score above 0.750 across all indicator categories. It also showed high format adherence, confirming its ability to generate reliable structured outputs. DISCUSSION: Instruction-tuning moderate-sized LLMs with QLoRA and targeted data augmentation enables high-accuracy extraction of infection indicators from HHC notes. CONCLUSION: This resource-efficient pipeline provides a scalable foundation for automated infection surveillance in healthcare settings with limited resources.
Zidu Xu, Jiyoun Song, Shuang Zhou 0012, Danielle Scharp, Mollie Hobensack, Jingjing Shang, Maxim Topaz
J. Am. Medical Informatics Assoc.8
2025 Identifying stigmatizing and positive/preferred language in obstetric clinical notes using natural language processing
abstract
OBJECTIVE: To identify stigmatizing language in obstetric clinical notes using natural language processing (NLP). MATERIALS AND METHODS: We analyzed electronic health records from birth admissions in the Northeast United States in 2017. We annotated 1771 clinical notes to generate the initial gold standard dataset. Annotators labeled for exemplars of 5 stigmatizing and 1 positive/preferred language categories. We used a semantic similarity-based search approach to expand the initial dataset by adding additional exemplars, composing an enhanced dataset. We employed traditional classifiers (Support Vector Machine, Decision Trees, and Random Forest) and a transformer-based model, ClinicalBERT (Bidirectional Encoder Representations from Transformers) and BERT base. Models were trained and validated on initial and enhanced datasets and were tested on enhanced testing dataset. RESULTS: In the initial dataset, we annotated 963 exemplars as stigmatizing or positive/preferred. The most frequently identified category was marginalized language/identities (n = 397, 41%), and the least frequent was questioning patient credibility (n = 51, 5%). After employing a semantic similarity-based search approach, 502 additional exemplars were added, increasing the number of low-frequency categories. All NLP models also showed improved performance, with Decision Trees demonstrating the greatest improvement (21%). ClinicalBERT outperformed other models, with the highest average F1-score of 0.78. DISCUSSION: Clinical BERT seems to most effectively capture the nuanced and context-dependent stigmatizing language found in obstetric clinical notes, demonstrating its potential clinical applications for real-time monitoring and alerts to prevent usages of stigmatizing language use and reduce healthcare bias. Future research should explore stigmatizing language in diverse geographic locations and clinical settings to further contribute to high-quality and equitable perinatal care. CONCLUSION: ClinicalBERT effectively captures the nuanced stigmatizing language in obstetric clinical notes. Our semantic similarity-based search approach to rapidly extract additional exemplars enhanced the performances while reducing the need for labor-intensive annotation.
Jihye Kim Scroggins, Ismael I Hulchafo, Sarah Harkins, Danielle Scharp, Hans Moen, Anahita Davoudi, Kenrick Cato, Michele Tadiello, Maxim Topaz, Veronica Barcelona
J. Am. Medical Informatics Assoc.9
2025 Machine learning-based infection diagnostic and prognostic models in post-acute care settings: a systematic review
abstract
OBJECTIVES: This study aims to (1) review machine learning (ML)-based models for early infection diagnostic and prognosis prediction in post-acute care (PAC) settings, (2) identify key risk predictors influencing infection-related outcomes, and (3) examine the quality and limitations of these models. MATERIALS AND METHODS: PubMed, Web of Science, Scopus, IEEE Xplore, CINAHL, and ACM digital library were searched in February 2024. Eligible studies leveraged PAC data to develop and evaluate ML models for infection-related risks. Data extraction followed the CHARMS checklist. Quality appraisal followed the PROBAST tool. Data synthesis was guided by the socio-ecological conceptual framework. RESULTS: Thirteen studies were included, mainly focusing on respiratory infections and nursing homes. Most used regression models with structured electronic health record data. Since 2020, there has been a shift toward advanced ML algorithms and multimodal data, biosensors, and clinical notes being significant sources of unstructured data. Despite these advances, there is insufficient evidence to support performance improvements over traditional models. Individual-level risk predictors, like impaired cognition, declined function, and tachycardia, were commonly used, while contextual-level predictors were barely utilized, consequently limiting model fairness. Major sources of bias included lack of external validation, inadequate model calibration, and insufficient consideration of data complexity. DISCUSSION AND CONCLUSION: Despite the growth of advanced modeling approaches in infection-related models in PAC settings, evidence supporting their superiority remains limited. Future research should leverage a socio-ecological lens for predictor selection and model construction, exploring optimal data modalities and ML model usage in PAC, while ensuring rigorous methodologies and fairness considerations.
Zidu Xu, Danielle Scharp, Mollie Hobensack, Jiancheng Ye, Jungang Zou, Sirui Ding, Jingjing Shang, Maxim Topaz
J. Am. Medical Informatics Assoc.8
2025 Beyond electronic health record data: leveraging natural language processing and machine learning to uncover cognitive insights from patient-nurse verbal communications
abstract
BACKGROUND: Mild cognitive impairment and early-stage dementia significantly impact healthcare utilization and costs, yet more than half of affected patients remain underdiagnosed. This study leverages audio-recorded patient-nurse verbal communication in home healthcare settings to develop an artificial intelligence-based screening tool for early detection of cognitive decline. OBJECTIVE: To develop a speech processing algorithm using routine patient-nurse verbal communication and evaluate its performance when combined with electronic health record (EHR) data in detecting early signs of cognitive decline. METHOD: We analyzed 125 audio-recorded patient-nurse verbal communication for 47 patients from a major home healthcare agency in New York City. Out of 47 patients, 19 experienced symptoms associated with the onset of cognitive decline. A natural language processing algorithm was developed to extract domain-specific linguistic and interaction features from these recordings. The algorithm's performance was compared against EHR-based screening methods. Both standalone and combined data approaches were assessed using F1-score and area under the curve (AUC) metrics. RESULTS: The initial model using only patient-nurse verbal communication achieved an F1-score of 85 and an AUC of 86.47. The model based on EHR data achieved an F1-score of 75.56 and an AUC of 79. Combining patient-nurse verbal communication with EHR data yielded the highest performance, with an F1-score of 88.89 and an AUC of 90.23. Key linguistic indicators of cognitive decline included reduced linguistic diversity, grammatical challenges, repetition, and altered speech patterns. Incorporating audio data significantly enhanced the risk prediction models for hospitalization and emergency department visits. DISCUSSION: Routine verbal communication between patients and nurses contains critical linguistic and interactional indicators for identifying cognitive impairment. Integrating audio-recorded patient-nurse communication with EHR data provides a more comprehensive and accurate method for early detection of cognitive decline, potentially improving patient outcomes through timely interventions. This combined approach could revolutionize cognitive impairment screening in home healthcare settings.
Maryam Zolnoori, Ali Zolnour, Sasha Vergez, Sridevi Sridharan, Ian Spens, Maxim Topaz, James Noble 0003, Suzanne Bakken, Julia Hirschberg, Kathryn H. Bowles, Nicole Onorato, Margaret V. McDonald
J. Am. Medical Informatics Assoc.6
2025 Rapid review: Growing usage of Multimodal Large Language Models in healthcare
Pallavi Gupta, Zhihong Zhang 0005, Meijia Song, Martin Michalowski, Gregor Stiglic, Maxim Topaz
J. Biomed. Informatics7
2024 Exploring home healthcare clinicians' needs for using clinical decision support systems for early risk warning
abstract
OBJECTIVES: To explore home healthcare (HHC) clinicians' needs for Clinical Decision Support Systems (CDSS) information delivery for early risk warning within HHC workflows. METHODS: Guided by the CDS "Five-Rights" framework, we conducted semi-structured interviews with multidisciplinary HHC clinicians from April 2023 to August 2023. We used deductive and inductive content analysis to investigate informants' responses regarding CDSS information delivery. RESULTS: Interviews with thirteen HHC clinicians yielded 16 codes mapping to the CDS "Five-Rights" framework (right information, right person, right format, right channel, right time) and 11 codes for unintended consequences and training needs. Clinicians favored risk levels displayed in color-coded horizontal bars, concrete risk indicators in bullet points, and actionable instructions in the existing EHR system. They preferred non-intrusive risk alerts requiring mandatory confirmation. Clinicians anticipated risk information updates aligned with patient's condition severity and their visit pace. Additionally, they requested training to understand the CDSS's underlying logic, and raised concerns about information accuracy and data privacy. DISCUSSION: While recognizing CDSS's value in enhancing early risk warning, clinicians highlighted concerns about increased workload, alert fatigue, and CDSS misuse. The top risk factors identified by machine learning algorithms, especially text features, can be ambiguous due to a lack of context. Future research should ensure that CDSS outputs align with clinical evidence and are explainable. CONCLUSION: This study identified HHC clinicians' expectations, preferences, adaptations, and unintended uses of CDSS for early risk warning. Our findings endorse operationalizing the CDS "Five-Rights" framework to optimize CDSS information delivery and integration into HHC workflows.
Zidu Xu, Lauren Evans, Jiyoun Song, Sena Chae, Anahita Davoudi, Kathryn H. Bowles, Margaret V. McDonald, Maxim Topaz
J. Am. Medical Informatics Assoc.8
2024 Utilizing patient-nurse verbal communication in building risk identification models: the missing critical data stream in home healthcare
abstract
BACKGROUND: In the United States, over 12 000 home healthcare agencies annually serve 6+ million patients, mostly aged 65+ years with chronic conditions. One in three of these patients end up visiting emergency department (ED) or being hospitalized. Existing risk identification models based on electronic health record (EHR) data have suboptimal performance in detecting these high-risk patients. OBJECTIVES: To measure the added value of integrating audio-recorded home healthcare patient-nurse verbal communication into a risk identification model built on home healthcare EHR data and clinical notes. METHODS: This pilot study was conducted at one of the largest not-for-profit home healthcare agencies in the United States. We audio-recorded 126 patient-nurse encounters for 47 patients, out of which 8 patients experienced ED visits and hospitalization. The risk model was developed and tested iteratively using: (1) structured data from the Outcome and Assessment Information Set, (2) clinical notes, and (3) verbal communication features. We used various natural language processing methods to model the communication between patients and nurses. RESULTS: Using a Support Vector Machine classifier, trained on the most informative features from OASIS, clinical notes, and verbal communication, we achieved an AUC-ROC = 99.68 and an F1-score = 94.12. By integrating verbal communication into the risk models, the F-1 score improved by 26%. The analysis revealed patients at high risk tended to interact more with risk-associated cues, exhibit more "sadness" and "anxiety," and have extended periods of silence during conversation. CONCLUSION: This innovative study underscores the immense value of incorporating patient-nurse verbal communication in enhancing risk prediction models for hospitalizations and ED visits, suggesting the need for an evolved clinical workflow that integrates routine patient-nurse verbal communication recording into the medical record.
Maryam Zolnoori, Sridevi Sridharan, Ali Zolnour, Sasha Vergez, Margaret V. McDonald, Zoran Kostic, Kathryn H. Bowles, Maxim Topaz
J. Am. Medical Informatics Assoc.8
2023 ADscreen: A speech processing-based screening system for automatic identification of patients with Alzheimer's disease and related dementia
Maryam Zolnoori, Ali Zolnour, Maxim Topaz
Artif. Intell. Medicine3
2023 Predicting emergency department visits and hospitalizations for patients with heart failure in home healthcare using a time series risk model
abstract
OBJECTIVES: Little is known about proactive risk assessment concerning emergency department (ED) visits and hospitalizations in patients with heart failure (HF) who receive home healthcare (HHC) services. This study developed a time series risk model for predicting ED visits and hospitalizations in patients with HF using longitudinal electronic health record data. We also explored which data sources yield the best-performing models over various time windows. MATERIALS AND METHODS: We used data collected from 9362 patients from a large HHC agency. We iteratively developed risk models using both structured (eg, standard assessment tools, vital signs, visit characteristics) and unstructured data (eg, clinical notes). Seven specific sets of variables included: (1) the Outcome and Assessment Information Set, (2) vital signs, (3) visit characteristics, (4) rule-based natural language processing-derived variables, (5) term frequency-inverse document frequency variables, (6) Bio-Clinical Bidirectional Encoder Representations from Transformers variables, and (7) topic modeling. Risk models were developed for 18 time windows (1-15, 30, 45, and 60 days) before an ED visit or hospitalization. Risk prediction performances were compared using recall, precision, accuracy, F1, and area under the receiver operating curve (AUC). RESULTS: The best-performing model was built using a combination of all 7 sets of variables and the time window of 4 days before an ED visit or hospitalization (AUC = 0.89 and F1 = 0.69). DISCUSSION AND CONCLUSION: This prediction model suggests that HHC clinicians can identify patients with HF at risk for visiting the ED or hospitalization within 4 days before the event, allowing for earlier targeted interventions.
Sena Chae, Anahita Davoudi, Jiyoun Song, Lauren Evans, Mollie Hobensack, Kathryn H. Bowles, Margaret V. McDonald, Yolanda Barrón, Sarah Collins Rossetti, Kenrick Cato, Sridevi Sridharan, Maxim Topaz
J. Am. Medical Informatics Assoc.12
2023 Uncovering hidden trends: identifying time trajectories in risk factors documented in clinical notes and predicting hospitalizations and emergency department visits during home health care
abstract
OBJECTIVE: This study aimed to identify temporal risk factor patterns documented in home health care (HHC) clinical notes and examine their association with hospitalizations or emergency department (ED) visits. MATERIALS AND METHODS: Data for 73 350 episodes of care from one large HHC organization were analyzed using dynamic time warping and hierarchical clustering analysis to identify the temporal patterns of risk factors documented in clinical notes. The Omaha System nursing terminology represented risk factors. First, clinical characteristics were compared between clusters. Next, multivariate logistic regression was used to examine the association between clusters and risk for hospitalizations or ED visits. Omaha System domains corresponding to risk factors were analyzed and described in each cluster. RESULTS: Six temporal clusters emerged, showing different patterns in how risk factors were documented over time. Patients with a steep increase in documented risk factors over time had a 3 times higher likelihood of hospitalization or ED visit than patients with no documented risk factors. Most risk factors belonged to the physiological domain, and only a few were in the environmental domain. DISCUSSION: An analysis of risk factor trajectories reflects a patient's evolving health status during a HHC episode. Using standardized nursing terminology, this study provided new insights into the complex temporal dynamics of HHC, which may lead to improved patient outcomes through better treatment and management plans. CONCLUSION: Incorporating temporal patterns in documented risk factors and their clusters into early warning systems may activate interventions to prevent hospitalizations or ED visits in HHC.
Jiyoun Song, Se Hee Min, Sena Chae, Kathryn H. Bowles, Margaret V. McDonald, Mollie Hobensack, Yolanda Barrón, Sridevi Sridharan, Anahita Davoudi, Sungho Oh, Lauren Evans, Maxim Topaz
J. Am. Medical Informatics Assoc.12
2023 Is the patient speaking or the nurse? Automatic speaker type identification in patient-nurse audio recordings
abstract
OBJECTIVES: Patient-clinician communication provides valuable explicit and implicit information that may indicate adverse medical conditions and outcomes. However, practical and analytical approaches for audio-recording and analyzing this data stream remain underexplored. This study aimed to 1) analyze patients' and nurses' speech in audio-recorded verbal communication, and 2) develop machine learning (ML) classifiers to effectively differentiate between patient and nurse language. MATERIALS AND METHODS: Pilot studies were conducted at VNS Health, the largest not-for-profit home healthcare agency in the United States, to optimize audio-recording patient-nurse interactions. We recorded and transcribed 46 interactions, resulting in 3494 "utterances" that were annotated to identify the speaker. We employed natural language processing techniques to generate linguistic features and built various ML classifiers to distinguish between patient and nurse language at both individual and encounter levels. RESULTS: A support vector machine classifier trained on selected linguistic features from term frequency-inverse document frequency, Linguistic Inquiry and Word Count, Word2Vec, and Medical Concepts in the Unified Medical Language System achieved the highest performance with an AUC-ROC = 99.01 ± 1.97 and an F1-score = 96.82 ± 4.1. The analysis revealed patients' tendency to use informal language and keywords related to "religion," "home," and "money," while nurses utilized more complex sentences focusing on health-related matters and medical issues and were more likely to ask questions. CONCLUSION: The methods and analytical approach we developed to differentiate patient and nurse language is an important precursor for downstream tasks that aim to analyze patient speech to identify patients at risk of disease and negative health outcomes.
Maryam Zolnoori, Sasha Vergez, Sridevi Sridharan, Ali Zolnour, Kathryn H. Bowles, Zoran Kostic, Maxim Topaz
J. Am. Medical Informatics Assoc.7
2022 Heart Failure Patient Characteristics and Symptoms Documented in Home Health Care Clinical Notes are Associated with Emergency Department Visits and Hospitalizations
Sena Chae, Jiyoun Song, Yolanda Barrón, Kathryn H. Bowles, Margaret V. McDonald, Sarah Collins Rossetti, Kenrick Cato, Mollie Hobensack, Lauren Evans, Maxim Topaz
AMIA10
2022 Capturing Concerns about Patient Deterioration in Narrative Documentation in Home Healthcare
Mollie Hobensack, Jiyoun Song, Sena Chae, Erin E. Kennedy, Maryam Zolnoori, Kathryn H. Bowles, Margaret V. McDonald, Lauren Evans, Maxim Topaz
AMIA9
2022 Is Auto-generated Transcript of Patient-Nurse Communication Ready to Use for Identifying the Risk for Hospitalizations or Emergency Department Visits in Home Health Care? A Natural Language Processing Pilot Study
Jiyoun Song, Maryam Zolnoori, Danielle Scharp, Sasha Vergez, Margaret V. McDonald, Sridevi Sridharan, Zoran Kostic, Maxim Topaz
AMIA8
2022 ADscreen: An Artificial Intelligence-Based Screening Algorithm for Early Identification of Patients with Alzheimer's Disease and Related Dementia
Maryam Zolnoori, Maxim Topaz
AMIA2
2022 Documentation of hospitalization risk factors in electronic health records (EHRs): a qualitative study with home healthcare clinicians
abstract
OBJECTIVE: To identify the risk factors home healthcare (HHC) clinicians associate with patient deterioration and understand how clinicians respond to and document these risk factors. METHODS: We interviewed multidisciplinary HHC clinicians from January to March of 2021. Risk factors were mapped to standardized terminologies (eg, Omaha System). We used directed content analysis to identify risk factors for deterioration. We used inductive thematic analysis to understand HHC clinicians' response to risk factors and documentation of risk factors. RESULTS: Fifteen HHC clinicians identified a total of 79 risk factors that were mapped to standardized terminologies. HHC clinicians most frequently responded to risk factors by communicating with the prescribing provider (86.7% of clinicians) or following up with patients and caregivers (86.7%). HHC clinicians stated that a majority of risk factors can be found in clinical notes (ie, care coordination (53.3%) or visit (46.7%)). DISCUSSION: Clinicians acknowledged that social factors play a role in deterioration risk; but these factors are infrequently studied in HHC. While a majority of risk factors were represented in the Omaha System, additional terminologies are needed to comprehensively capture risk. Since most risk factors are documented in clinical notes, methods such as natural language processing are needed to extract them. CONCLUSION: This study engaged clinicians to understand risk for deterioration during HHC. The results of our study support the development of an early warning system by providing a comprehensive list of risk factors grounded in clinician expertize and mapped to standardized terminologies.
Mollie Hobensack, Marietta Ojo, Yolanda Barrón, Kathryn H. Bowles, Kenrick Cato, Sena Chae, Erin E. Kennedy, Margaret V. McDonald, Sarah Collins Rossetti, Jiyoun Song, Sridevi Sridharan, Maxim Topaz
J. Am. Medical Informatics Assoc.12
2022 Considerations for development of child abuse and neglect phenotype with implications for reduction of racial bias: a qualitative study
abstract
OBJECTIVE: The study provides considerations for generating a phenotype of child abuse and neglect in Emergency Departments (ED) using secondary data from electronic health records (EHR). Implications will be provided for racial bias reduction and the development of further decision support tools to assist in identifying child abuse and neglect. MATERIALS AND METHODS: We conducted a qualitative study using in-depth interviews with 20 pediatric clinicians working in a single pediatric ED to gain insights about generating an EHR-based phenotype to identify children at risk for abuse and neglect. RESULTS: Three central themes emerged from the interviews: (1) Challenges in diagnosing child abuse and neglect, (2) Health Discipline Differences in Documentation Styles in EHR, and (3) Identification of potential racial bias through documentation. DISCUSSION: Our findings highlight important considerations for generating a phenotype for child abuse and neglect using EHR data. First, information-related challenges include lack of proper previous visit history due to limited information exchanges and scattered documentation within EHRs. Second, there are differences in documentation styles by health disciplines, and clinicians tend to document abuse in different document types within EHRs. Finally, documentation can help identify potential racial bias in suspicion of child abuse and neglect by revealing potential discrepancies in quality of care, and in the language used to document abuse and neglect. CONCLUSIONS: Our findings highlight challenges in building an EHR-based risk phenotype for child abuse and neglect. Further research is needed to validate these findings and integrate them into creation of an EHR-based risk phenotype.
Aviv Y. Landau, Ashley Blanchard, Kenrick Cato, Nia Atkins, Stephanie Salazar, Desmond Upton Patton, Maxim Topaz
J. Am. Medical Informatics Assoc.7
2022 Developing machine learning-based models to help identify child abuse and neglect: key ethical challenges and recommended solutions
abstract
Child abuse and neglect are public health issues impacting communities throughout the United States. The broad adoption of electronic health records (EHR) in health care supports the development of machine learning-based models to help identify child abuse and neglect. Employing EHR data for child abuse and neglect detection raises several critical ethical considerations. This article applied a phenomenological approach to discuss and provide recommendations for key ethical issues related to machine learning-based risk models development and evaluation: (1) biases in the data; (2) clinical documentation system design issues; (3) lack of centralized evidence base for child abuse and neglect; (4) lack of "gold standard "in assessment and diagnosis of child abuse and neglect; (5) challenges in evaluation of risk prediction performance; (6) challenges in testing predictive models in practice; and (7) challenges in presentation of machine learning-based prediction to clinicians and patients. We provide recommended solutions to each of the 7 ethical challenges and identify several areas for further policy and research.
Aviv Y. Landau, Susi Ferrarello, Ashley Blanchard, Kenrick Cato, Nia Atkins, Stephanie Salazar, Desmond Upton Patton, Maxim Topaz
J. Am. Medical Informatics Assoc.8
2022 Clinical notes: An untapped opportunity for improving risk prediction for hospitalization and emergency department visit during home health care
Jiyoun Song, Mollie Hobensack, Kathryn H. Bowles, Margaret V. McDonald, Kenrick Cato, Sarah Collins Rossetti, Sena Chae, Erin E. Kennedy, Yolanda Barrón, Sridevi Sridharan, Maxim Topaz
J. Biomed. Informatics11
2021 Identifying Narrative Documentation of Clinician Concern about Patient Deterioration in Home Healthcare: A Text Mining Study
Mollie Hobensack, Jiyoun Song, Maryam Zolnoori, Marietta Ojo, Kathryn H. Bowles, Sena Chae, Erin E. Kennedy, Margaret V. McDonald, Maxim Topaz
AMIA9
2021 Home Healthcare to Primary Care Data Interoperability Challenges and Opportunities
Paulina S. Sockolow, Kathryn H. Bowles, Maxim Topaz, Edgar Y. Chou
AMIA3
2021 Natural Language Processing Algorithm to Detect Terms Representing Risk of Hospitalization or Emergency Department Visits during Home Health Care
Jiyoun Song, Marietta Ojo, Margaret V. McDonald, Kenrick Cato, Sarah Collins Rossetti, Yolanda Barrón, Sridevi Sridharan, Sena Chae, Mollie Hobensack, Kathryn H. Bowles, Maxim Topaz
AMIA11
2021 Feasibility study of audio recording patient-clinician verbal communications in home healthcare settings
Maryam Zolnoori, Sasha Vergez, Zoran Kostic, Siddhartha Jonnalagadda, Maxim Topaz
AMIA5
2019 Mining fall-related information in clinical notes: Comparison of rule-based and novel word embedding-based machine learning approaches
abstract
BACKGROUND: Natural language processing (NLP) of health-related data is still an expertise demanding, and resource expensive process. We created a novel, open source rapid clinical text mining system called NimbleMiner. NimbleMiner combines several machine learning techniques (word embedding models and positive only labels learning) to facilitate the process in which a human rapidly performs text mining of clinical narratives, while being aided by the machine learning components. OBJECTIVE: This manuscript describes the general system architecture and user Interface and presents results of a case study aimed at classifying fall-related information (including fall history, fall prevention interventions, and fall risk) in homecare visit notes. METHODS: We extracted a corpus of homecare visit notes (n = 1,149,586) for 89,459 patients from a large US-based homecare agency. We used a gold standard testing dataset of 750 notes annotated by two human reviewers to compare the NimbleMiner's ability to classify documents regarding whether they contain fall-related information with a previously developed rule-based NLP system. RESULTS: NimbleMiner outperformed the rule-based system in almost all domains. The overall F- score was 85.8% compared to 81% by the rule based-system with the best performance for identifying general fall history (F = 89% vs. F = 85.1% rule-based), followed by fall risk (F = 87% vs. F = 78.7% rule-based), fall prevention interventions (F = 88.1% vs. F = 78.2% rule-based) and fall within 2 days of the note date (F = 83.1% vs. F = 80.6% rule-based). The rule-based system achieved slightly better performance for fall within 2 weeks of the note date (F = 81.9% vs. F = 84% rule-based). DISCUSSION & CONCLUSIONS: NimbleMiner outperformed other systems aimed at fall information classification, including our previously developed rule-based approach. These promising results indicate that clinical text mining can be implemented without the need for large labeled datasets necessary for other types of machine learning. This is critical for domains with little NLP developments, like nursing or allied health professions.
Maxim Topaz, Ludmila Murga, Katherine M. Gaddis, Margaret V. McDonald, Ofrit Bar-Bachar, Yoav Goldberg, Kathryn H. Bowles
J. Biomed. Informatics1
2018 A value set for documenting adverse reactions in electronic health records
abstract
Objective: To develop a comprehensive value set for documenting and encoding adverse reactions in the allergy module of an electronic health record. Materials and Methods: We analyzed 2 471 004 adverse reactions stored in Partners Healthcare's Enterprise-wide Allergy Repository (PEAR) of 2.7 million patients. Using the Medical Text Extraction, Reasoning, and Mapping System, we processed both structured and free-text reaction entries and mapped them to Systematized Nomenclature of Medicine - Clinical Terms. We calculated the frequencies of reaction concepts, including rare, severe, and hypersensitivity reactions. We compared PEAR concepts to a Federal Health Information Modeling and Standards value set and University of Nebraska Medical Center data, and then created an integrated value set. Results: We identified 787 reaction concepts in PEAR. Frequently reported reactions included: rash (14.0%), hives (8.2%), gastrointestinal irritation (5.5%), itching (3.2%), and anaphylaxis (2.5%). We identified an additional 320 concepts from Federal Health Information Modeling and Standards and the University of Nebraska Medical Center to resolve gaps due to missing and partial matches when comparing these external resources to PEAR. This yielded 1106 concepts in our final integrated value set. The presence of rare, severe, and hypersensitivity reactions was limited in both external datasets. Hypersensitivity reactions represented roughly 20% of the reactions within our data. Discussion: We developed a value set for encoding adverse reactions using a large dataset from one health system, enriched by reactions from 2 large external resources. This integrated value set includes clinically important severe and hypersensitivity reactions. Conclusion: This work contributes a value set, harmonized with existing data, to improve the consistency and accuracy of reaction documentation in electronic health records, providing the necessary building blocks for more intelligent clinical decision support for allergies and adverse reactions.
Foster R. Goss, Kenneth H. Lai, Maxim Topaz, Warren W. Acker, Leigh Kowalski, Joseph M. Plasek, Kimberly G. Blumenthal, Diane L. Seger, Sarah P. Slight, Kin Wah Fung, Frank Y. Chang, David W. Bates, Li Zhou 0007
J. Am. Medical Informatics Assoc.3
2017 The application of machine learning to evaluate the adequacy of information in radiology orders
abstract
Background: Adequate clinical information provided with radiology orders is important for an accurate interpretation of imaging studies. Nonetheless, high percentage of radiology orders lack adequate information. Assessment of the adequacy of the information associated with radiology orders could be achieved manually. However, manual assessment is costly and inefficient. Novel approaches using machine learning and text mining to assess the adequacy of radiology order information could reduce the costs and improve efficiency. We aimed to test the application of machine learning algorithms to identify radiology orders with adequate/inadequate information. Methods: We extracted 1,967 electronic chest computed tomography (CT) orders at an academic tertiary hospital during January 2014, and manually classified them into containing adequate or inadequate information based on the American College of Radiology guidelines. We used text mining (text parsing and vectorization) and machine learning (Naïve Bayes, Support Vector Machines and Decision Tree classifiers) to automate order adequacy classification, and evaluated the system performance against the manual review. Results: Surprisingly, only 30.6% of orders had adequate information when evaluated manually. Non-resident physicians provided the least number of adequate order information (26.7%). Classifiers achieved high classification accuracy. Naïve Bayes classifier performed slightly better overall (Accuracy= .9) than Support Vector Machines (Accuracy= .89) and Decision Trees (J48, Accuracy= .85). Conclusions: High percentage of orders lack adequate information in chest CT. Machine learning classifiers could be utilized to assess the adequacy of radiology order information.
Wasim Al Assad, Maxim Topaz, John Tu, Li Zhou 0007
BIBM2
2016 Designing Next Generation of Clinical Decision Support for Nursing from Hospital to Homecare: AMIA Nursing Informatics Group Pre-Symposium Tutorial
Maxim Topaz, Sarah A. Collins
AMIA1
2016 Expert Recommendations on Redesigning Drug Allergy Alerts in Electronic Health Record Systems
Maxim Topaz, Foster R. Goss, Kimberly G. Blumenthal, Kenneth H. Lai, Diane L. Seger, Sarah P. Slight, Paige G. Wickner, George A. Robinson, Kin Wah Fung, Robert C. McClure, Shelly Spiro, Warren W. Acker, David W. Bates
AMIA1
2016 Nurse Informaticians Report Low Satisfaction and Multi-level Concerns with Electronic Health Records: Results from an International Survey
Maxim Topaz, Charlene Ronquillo, Laura-Maria Peltonen, Lisiane Pruinelli, Raymond Francis Sarmiento, Martha K. Badger, Samira Ali, Adrienne Lewis, Mattias Georgsson, Eunjoo Jeon, Jude L. Tayaben, Chiu-Hsiang Kuo, Tasneem Islam, Janine A. Sommer, Hyunggu Jung, Gabrielle Jacklin Eler, Dari Alhuwail, Ying-Li Lee
AMIA1
2016 Food entries in a large allergy data repository
abstract
OBJECTIVE: Accurate food adverse sensitivity documentation in electronic health records (EHRs) is crucial to patient safety. This study examined, encoded, and grouped foods that caused any adverse sensitivity in a large allergy repository using natural language processing and standard terminologies. METHODS: Using the Medical Text Extraction, Reasoning, and Mapping System (MTERMS), we processed both structured and free-text entries stored in an enterprise-wide allergy repository (Partners' Enterprise-wide Allergy Repository), normalized diverse food allergen terms into concepts, and encoded these concepts using the Systematized Nomenclature of Medicine - Clinical Terms (SNOMED-CT) and Unique Ingredient Identifiers (UNII) terminologies. Concept coverage also was assessed for these two terminologies. We further categorized allergen concepts into groups and calculated the frequencies of these concepts by group. Finally, we conducted an external validation of MTERMS's performance when identifying food allergen terms, using a randomized sample from a different institution. RESULTS: We identified 158 552 food allergen records (2140 unique terms) in the Partners repository, corresponding to 672 food allergen concepts. High-frequency groups included shellfish (19.3%), fruits or vegetables (18.4%), dairy (9.0%), peanuts (8.5%), tree nuts (8.5%), eggs (6.0%), grains (5.1%), and additives (4.7%). Ambiguous, generic concepts such as "nuts" and "seafood" accounted for 8.8% of the records. SNOMED-CT covered more concepts than UNII in terms of exact (81.7% vs 68.0%) and partial (14.3% vs 9.7%) matches. DISCUSSION: Adverse sensitivities to food are diverse, and existing standard terminologies have gaps in their coverage of the breadth of allergy concepts. CONCLUSION: New strategies are needed to represent and standardize food adverse sensitivity concepts, to improve documentation in EHRs.
Joseph M. Plasek, Foster R. Goss, Kenneth H. Lai, Jason J. Lau, Diane L. Seger, Kimberly G. Blumenthal, Paige G. Wickner, Sarah P. Slight, Frank Y. Chang, Maxim Topaz, David W. Bates, Li Zhou 0007
J. Am. Medical Informatics Assoc.10
2016 Rising drug allergy alert overrides in electronic health records: an observational retrospective study of a decade of experience
abstract
OBJECTIVE: There have been growing concerns about the impact of drug allergy alerts on patient safety and provider alert fatigue. The authors aimed to explore the common drug allergy alerts over the last 10 years and the reasons why providers tend to override these alerts. DESIGN: Retrospective observational cross-sectional study (2004-2013). MATERIALS AND METHODS: Drug allergy alert data (n = 611,192) were collected from two large academic hospitals in Boston, MA (USA). RESULTS: Overall, the authors found an increase in the rate of drug allergy alert overrides, from 83.3% in 2004 to 87.6% in 2013 (P < .001). Alarmingly, alerts for immune mediated and life threatening reactions with definite allergen and prescribed medication matches were overridden 72.8% and 74.1% of the time, respectively. However, providers were less likely to override these alerts compared to possible (cross-sensitivity) or probable (allergen group) matches (P < .001). The most common drug allergy alerts were triggered by allergies to narcotics (48%) and other analgesics (6%), antibiotics (10%), and statins (2%). Only slightly more than one-third of the reactions (34.2%) were potentially immune mediated. Finally, more than half of the overrides reasons pointed to irrelevant alerts (i.e., patient has tolerated the medication before, 50.9%) and providers were significantly more likely to override repeated alerts (89.7%) rather than first time alerts (77.4%, P < .001). DISCUSSION AND CONCLUSIONS: These findings underline the urgent need for more efforts to provide more accurate and relevant drug allergy alerts to help reduce alert override rates and improve alert fatigue.
Maxim Topaz, Diane L. Seger, Sarah P. Slight, Foster R. Goss, Kenneth H. Lai, Paige G. Wickner, Kimberly G. Blumenthal, Neil Dhopeshwarkar, Frank Y. Chang, David W. Bates, Li Zhou 0007
J. Am. Medical Informatics Assoc.1
2015 Rising Drug Allergy Alert Overrides in a Computerized Provider Order Entry System: a Decade of Experience
Li Zhou 0007, Maxim Topaz, Diane L. Seger, Sarah P. Slight, Foster R. Goss, Kenneth H. Lai, Paige G. Wickner, Kimberly G. Blumenthal, Neil Dhopeshwarkar, Frank Y. Chang, David W. Bates
AMIA2
2015 Automated misspelling detection and correction in clinical free-text records
abstract
Accurate electronic health records are important for clinical care and research as well as ensuring patient safety. It is crucial for misspelled words to be corrected in order to ensure that medical records are interpreted correctly. This paper describes the development of a spelling correction system for medical text. Our spell checker is based on Shannon's noisy channel model, and uses an extensive dictionary compiled from many sources. We also use named entity recognition, so that names are not wrongly corrected as misspellings. We apply our spell checker to three different types of free-text data: clinical notes, allergy entries, and medication orders; and evaluate its performance on both misspelling detection and correction. Our spell checker achieves detection performance of up to 94.4% and correction accuracy of up to 88.2%. We show that high-performance spelling correction is possible on a variety of clinical documents.
Kenneth H. Lai, Maxim Topaz, Foster R. Goss, Li Zhou 0007
J. Biomed. Informatics2
2014 Health Information Technology Adoption in Home Health ... Research in Progress
Dari Alhuwail, Akif Günes Koru, Ahmad Alaiad 0001, Anthony F. Norcio, Maxim Topaz
AMIA5
2014 The Omaha System: a systematic review of the recent literature
abstract
BACKGROUND: The Omaha System (OS) is one of the oldest of the American Nurses Association recognized standardized terminologies describing and measuring the impact of healthcare services. This systematic review presents the state of science on the use of the OS in practice, research, and education. AIMS: (1) To identify, describe and evaluate the publications on the OS between 2004 and 2011, (2) to identify major trends in the use of the OS in research, practice, and education, and (3) to suggest areas for future research. METHODS: Systematic search in the largest online healthcare databases (PUBMED, CINAHL, Scopus, PsycINFO, Ovid) from 2004 to 2011. Methodological quality of the reviewed research studies was evaluated. RESULTS: 56 publications on the OS were identified and analyzed. The methodological quality of the reviewed research studies was relatively high. Over time, publications' focus shifted from describing clients' problems toward outcomes research. There was an increasing application of advanced statistical methods and a significant portion of authors focused on classification and interoperability research. There was an increasing body of international literature on the OS. Little research focused on the theoretical aspects of the OS, the effective use of the OS in education, or cultural adaptations of the OS outside the USA. CONCLUSIONS: The OS has a high potential to provide meaningful and high quality information about complex healthcare services. Further research on the OS should focus on its applicability in healthcare education, theoretical underpinnings and international validity. Researchers analyzing the OS data should address how they attempted to mitigate the effects of missing data in analyzing their results and clearly present the limitations of their studies.
Maxim Topaz, Nadya Golfenshtein, Kathryn H. Bowles
J. Am. Medical Informatics Assoc.1
2013 "We're all in our own little island": A Qualitative Exploration of Patient Information Exchange during Admission to Home Health Agency
Maxim Topaz, D. Molkina, Akif Günes Koru, Ruth M. Masterson Creber, O. Jarrin, Kavita Radhakrishnan, Melissa O'Connor, Kathryn H. Bowles
AMIA1
2013 Developing Nursing Computer Interpretable Guidelines: a Feasibility Study of Heart Failure Guidelines in Homecare
Maxim Topaz, Erez Shalom, Ruth M. Masterson Creber, Kavita Rhadakrishnan, Kathryn H. Bowles
AMIA1
2012 Impact of Discharge Planning Decision Support on 30 and 60 Day Readmissions
Kathryn H. Bowles, Diane Holland, Sheryl Potashnik, Maxim Topaz, Alexandra L. Hanlon
AMIA4
2012 Association of patient characteristics and telehealth alerts with key medical events experienced by patients with heart failure (HF) in homecare
Kavita Radhakrishnan, Kathryn H. Bowles, Alexandra L. Hanlon, Maxim Topaz
AMIA4