Hilda B. Klasky

dblp:249/3357 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-7235-2521ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Anomaly Detection in Electronic Health Records Across Hospital Networks: Integrating Machine Learning With Graph Algorithms
abstract
In a large hospital system, a network of hospitals relies on electronic health records (EHRs) to make informed decisions regarding their patients in various clinical domains. Consequently, the dependability of the health information technology (HIT) systems responsible for collecting EHR data is of utmost importance for patient safety. Recently, novel methods and tools aimed at identifying anomalies in EHR data to bolster the reliability of HIT systems have been introduced. However, these existing methods and tools primarily concentrate on individual hospitals, which limits our understanding of system-wide anomalous events and their potential impact on patient safety across multiple hospitals. In this article, we introduce a new approach to detecting anomalies in EHR data within a network of hospitals. This is achieved by combining advanced machine learning techniques with graph algorithms to create a tool capable of swiftly identifying and responding to deviations. Our proposed approach employs a combination of five machine learning models, harnessing the unique strengths of each model to provide a more robust detection system. The detected anomalies are then represented as graphs, allowing us to recognize patterns across the hospital network. This aids in identifying anomalies that span multiple medical facilities, potentially indicating broader system-level risks. Extensive real-world testing of our approach demonstrated its ability to offer actionable insights compared to existing methods. Additionally, its scalable design ensures seamless integration into existing HIT infrastructures.
Olufemi A. Omitaomu, Michael A. Langston, Stephen K. Grady, Mohammed M. Olama, Özgür Özmen, Hilda B. Klasky, Angela Laurio, Merry Ward, Jonathan R. Nebeker
IEEE J. Biomed. Health Informatics7
2024 EHR-BERT: A BERT-based model for effective anomaly detection in electronic health records
abstract
OBJECTIVE: Physicians and clinicians rely on data contained in electronic health records (EHRs), as recorded by health information technology (HIT), to make informed decisions about their patients. The reliability of HIT systems in this regard is critical to patient safety. Consequently, better tools are needed to monitor the performance of HIT systems for potential hazards that could compromise the collected EHRs, which in turn could affect patient safety. In this paper, we propose a new framework for detecting anomalies in EHRs using sequence of clinical events. This new framework, EHR-Bidirectional Encoder Representations from Transformers (BERT), is motivated by the gaps in the existing deep-learning related methods, including high false negatives, sub-optimal accuracy, higher computational cost, and the risk of information loss. EHR-BERT is an innovative framework rooted in the BERT architecture, meticulously tailored to navigate the hurdles in the contemporary BERT method; thus, enhancing anomaly detection in EHRs for healthcare applications. METHODS: The EHR-BERT framework was designed using the Sequential Masked Token Prediction (SMTP) method. This approach treats EHRs as natural language sentences and iteratively masks input tokens during both training and prediction stages. This method facilitates the learning of EHR sequence patterns in both directions for each event and identifies anomalies based on deviations from the normal execution models trained on EHR sequences. RESULTS: Extensive experiments on large EHR datasets across various medical domains demonstrate that EHR-BERT markedly improves upon existing models. It significantly reduces the number of false positives and enhances the detection rate, thus bolstering the reliability of anomaly detection in electronic health records. This improvement is attributed to the model's ability to minimize information loss and maximize data utilization effectively. CONCLUSION: EHR-BERT showcases immense potential in decreasing medical errors related to anomalous clinical events, positioning itself as an indispensable asset for enhancing patient safety and the overall standard of healthcare services. The framework effectively overcomes the drawbacks of earlier models, making it a promising solution for healthcare professionals to ensure the reliability and quality of health data.
Olufemi A. Omitaomu, Michael A. Langston, Mohammed M. Olama, Özgür Özmen, Hilda B. Klasky, Angela Laurio, Merry Ward, Jonathan R. Nebeker
J. Biomed. Informatics6
2022 Process mining for healthcare: Characteristics and challenges
abstract
Process mining techniques can be used to analyse business processes using the data logged during their execution. These techniques are leveraged in a wide range of domains, including healthcare, where it focuses mainly on the analysis of diagnostic, treatment, and organisational processes. Despite the huge amount of data generated in hospitals by staff and machinery involved in healthcare processes, there is no evidence of a systematic uptake of process mining beyond targeted case studies in a research context. When developing and using process mining in healthcare, distinguishing characteristics of healthcare processes such as their variability and patient-centred focus require targeted attention. Against this background, the Process-Oriented Data Science in Healthcare Alliance has been established to propagate the research and application of techniques targeting the data-driven improvement of healthcare processes. This paper, an initiative of the alliance, presents the distinguishing characteristics of the healthcare domain that need to be considered to successfully use process mining, as well as open challenges that need to be addressed by the community in the future.
Jorge Munoz-Gama, Niels Martin, Carlos Fernández-Llatas, Owen A. Johnson, Marcos Sepúlveda, Emmanuel Helm, Victor Galvez-Yanjari, Eric Rojas Cordoba, Antonio Martinez-Millana, Davide Aloini, Ilaria Angela Amantea, Robert Andrews 0001, Michael Arias, Iris Beerepoot, Elisabetta Benevento, Andrea Burattin, Daniel Capurro, Josep Carmona 0001, Marco Comuzzi, Benjamin Dalmas, Rene de la Fuente, Chiara Di Francescomarino, Claudio Di Ciccio, Roberto Gatta, Chiara Ghidini, Fernanda Gonzalez-Lopez, Gema Ibáñez-Sánchez, Hilda B. Klasky, Angelina Prima Kurniati, Xixi Lu 0001, Felix Mannhardt, R. S. Mans, Mar Marcos, Renata Medeiros de Carvalho, Marco Pegoraro 0001, Simon K. Poon, Luise Pufahl, Hajo A. Reijers, Simon Remy, Stefanie Rinderle-Ma, Lucia Sacchi, Fernando Seoane, Minseok Song 0001, Alessandro Stefanini, Emilio Sulis, Arthur H. M. ter Hofstede, Pieter J. Toussaint, Vicente Traver 0001, Zoe Valero-Ramon, Inge van de Weerd, Wil M. P. van der Aalst, Rob J. B. Vanwersch, Mathias Weske, Moe Thandar Wynn, Francesca Zerbato
J. Biomed. Informatics28
2022 Detecting anomalous sequences in electronic health records using higher-order tensor networks
abstract
Detecting anomalous sequences is an integral part of building and protecting modern large-scale health information technology (HIT) systems. These HIT systems generate a large volume of records of patients' state and significant events, which provide a valuable resource to help improve clinical decisions, patient care processes, and other issues. However, detecting anomalous sequences in electronic health records (EHR) remains a challenge in healthcare applications for several reasons, including imbalances in the data, complexity of relationships between events in the sequence, and the curse of dimensionality. Conventional anomaly detection methods use the finite sequence of events to discriminate sequences. They fail to incorporate salient event details under variable higher-order dependencies (e.g., duration between events) that can provide better discrimination of sequences in their models. To address this problem, we propose event sequence and subsequence anomaly detection algorithms that (1) use network-based representations of interactions in the data, (2) account for variable higher-order dependencies in the data, and (3) incorporate events duration for adequate discrimination of the data. The proposed approach identifies anomalies by monitoring the change in the graph after the test sequence is removed from the network. The change is quantified using graph distance metrics so that dramatic changes in the network can be attributed to the removed sequence. Furthermore, the proposed subsequence algorithm recommends plausible paths and salient information for the detected anomalous subsequences. Our results show that the proposed event sequence anomaly detection algorithm outperforms the baseline methods for both synthetic data and real-world EHR data.
Olufemi A. Omitaomu, Michael A. Langston, Mohammed M. Olama, Özgür Özmen, Hilda B. Klasky, Angela Laurio, Brian C. Sauer, Merry Ward, Jonathan R. Nebeker
J. Biomed. Informatics6
2021 A new methodological framework for hazard detection models in health information technology systems
abstract
The adoption of health information technology (HIT) has facilitated efforts to increase the quality and efficiency of health care services and decrease health care overhead while simultaneously generating massive amounts of digital information stored in electronic health records (EHRs). However, due to patient safety issues resulting from the use of HIT systems, there is an emerging need to develop and implement hazard detection tools to identify and mitigate risks to patients. This paper presents a new methodological framework to develop hazard detection models and to demonstrate its capability by using the US Department of Veterans Affairs' (VA) Corporate Data Warehouse, the data repository for the VA's EHR. The overall purpose of the framework is to provide structure for research and communication about research results. One objective is to decrease the communication barriers between interdisciplinary research stakeholders and to provide structure for detecting hazards and risks to patient safety introduced by HIT systems through errors in the collection, transmission, use, and processing of data in the EHR, as well as potential programming or configuration errors in these HIT systems. A nine-stage framework was created, which comprises programs about feature extraction, detector development, and detector optimization, as well as a support environment for evaluating detector models. The framework forms the foundation for developing hazard detection tools and the foundation for adapting methods to particular HIT systems.
Olufemi A. Omitaomu, Hilda B. Klasky, Mohammed M. Olama, Özgür Özmen, Laura L. Pullum, Addi Malviya-Thakur, P. Teja Kuruganti, Jean M. Scott, Angela Laurio, Frank Drews, Brian C. Sauer, Merry Ward, Jonathan R. Nebeker
J. Biomed. Informatics2
2020 Adaptive Anomaly Detection for Dynamic Clinical Event Sequences
abstract
Over the past decade, health information technology (IT) has enabled the amount of digital information stored in electronic health records (EHRs) to expand greatly. However, according to some studies, hazards in health IT can lead to changes in clinical decisions, care processes, and care outcomes, as well as other issues. Thus, the effects of health IT hazards on patient safety have been at the forefront of recent patient safety research. Nonetheless, hazard detection in health IT remains a challenge. In this paper, the authors assume that safety-related issues in health IT would exhibit anomalous characteristics in EHR data. Although all hazards will exhibit some anomalous characteristics, not all anomalies can be regarded as hazards. The authors hypothesize that errors in health IT could lead to interruptions in the sequence of clinical actions. To this end, the problem of detecting anomalous sequences in big EHR data is considered. This paper focuses on dynamic event sequences, which are a series of clinical actions in motion. The authors propose an adaptive anomaly detection approach that uses higher-order network representation to detect anomalous sequences. Furthermore, the authors propose a contiguous subsequence anomaly detection approach that identifies abnormal subsequences in the detected anomalous sequences. The proposed approaches are tested by using synthetic and real-world EHR data. The proposed methods outperform existing state of the art anomaly detection techniques. To reduce the computational complexity associated with the operational implementation of the proposed approaches, the Apache Spark environment was leveraged, and a much shorter run time together with improved performance were achieved, especially for data with more than 60,000 sequences.
Olufemi A. Omitaomu, Qing Cao 0001, Mohammed M. Olama, Özgür Özmen, Hilda B. Klasky, Laura L. Pullum, Addi Malviya-Thakur, P. Teja Kuruganti, Jean M. Scott, Angela Laurio, Frank Drews, Brian C. Sauer, Merry Ward, Jonathan R. Nebeker
IEEE BigData6
2020 Accelerated training of bootstrap aggregation-based deep information extraction systems from cancer pathology reports
Hong-Jun Yoon, Hilda B. Klasky, John Gounley, Mohammed M. Alawad, Shang Gao 0008, Eric B. Durbin, Xiao-Cheng Wu, Antoinette Stroup, Jennifer A. Doherty, Linda Coyle, Lynne Penberthy, James Blair Christian, Georgia D. Tourassi
J. Biomed. Informatics2