Chanhee Kim

dblp:322/5565 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2023 Comparison of MIMIC-III and MIMIC-IV for big data analytics of health informatics
abstract
The use of a health big data via Medical Information Mart for Intensive Care (MIMIC) data sets has significant advancements in health informatics and clinical research of electronic health records (EHRs) in hospital systems. MIMIC-III and MIMIC-IV data sets are the two latest publicly available iterations of data from electronic records that possess distinct features, variables, and structures in two separate large relational databases. This study aimed to provide big data analytics comparisons of MIMIC-III and MIMIC-IV for experiential learning of health informatics from EHRs.Both data sets were quantitatively and qualitatively evaluated based on dataset properties (data, list of tables, timeline, patient encounters, software, system), data mining and visual categories (i.e., data size, data structure and schema, features/variables, data quality and completeness, clinical focus, access and use, dashboards, and visualization types), data usability heuristics and utilization (i.e., usability, complexity, granularity, applicability, and capacity). Results showed significant difference of 100 patients across 26 tables for MIMIC-III compared to 2,520 patients across 31 tables for MIMIC-IV data sets, respectively. There were 1716 diagnoses (ICD-9) and 503 procedures (ICD-9) for MIMIC-III. There were 262 diseases categories with >1000 disease instances and >2000 treatments for MIMIC-IV with no ICD codes. Moreover, these data sets contained large (big data) charting events in MIMIC-III of 758,356 rows, and MIMIC-IV eICU with charting events of 1,477,163 for nursing, and respiratory of 176,089 rows, respectively. The results suggest that MIMIC-III provided detailed information for retrospective clinical studies and operations in critical care with high data granularity in terms of re-admission (calculated fields from its admission table), length of stay, prescriptions, caregivers, and diagnosis and procedure (ICD-9). However, it lacked clinical capacity because of no diagnosis or charting event offset times or APACHE IV (Acute Physiology and Chronic Health Evaluation) scores that were in the MIMIC-IV eICU dataset. Hence, MIMIC-IV showed higher data granularity and capacity. MIMIC-IV eICU dataset introduces enhanced data attributes, more sophisticated patient trajectory tracking at the ICU unit level, and improved detailed information from electronic records. Moreover, MIMIC-IV has high usability, complexity, and applicability to critical care but lacking hospital operational data of re-admissions, caregivers, and ICD codes that MIMIC-III contained. However, MIMIC-IV dataset contained complex data schemas of treatment strings, nurse care plans, lab results, medications, microbiology, as well as infusion drug and respiratory charting. The big data analytics of MIMIC-III and MIMIC-IV needs to be further investigated for AI applications to demonstrate its usefulness for dynamic decision-making in health care.
Dillon Chrimes, Chanhee Kim
IEEE Big Data2
2022 Review of Publically Available Health Big Data Sets
abstract
There is a growing interest in using public data for open government policy involving health informatics and healthcare systems. This paper investigated the characteristics of publically available data sets in health informatics that were derived from electronic health records (EHRs), healthcare systems, and a variety of open-government libraries, data marts, or data catalogues.Data used in this study consisted of public data sets that did not require any registration to access online. In total, nine web-based platforms on the Internet were used that included: British Columbia (BC) Data Catalogue, Canadian Institute for Health Information (CIHI), Harvard Dataverse, MIMIC-eICU, FigShare, GitHub, Google Dataset, UCI Machine Learning Repository, and Zenodo. Our initial search across these platforms found over 10,000 public use files that had data sets related to health informatics.We found 558 data sets that matched search criterion that ranged from years 2013-2022. The data source types were mostly found using the health informatics search filters followed by the combination of health informatics and healthcare systems, but fewer data sets were found when using EHR as the criterion. Almost 85% of the total data sets were from 2020-2022. The range of data sizes were 11KB to 7.8MB. The eICU (hosted by MIT’s MIMIC data mart) platform had the largest data set followed by Zenodo, and GitHub. Additionally, any bioinformatics in the 558 data sets were excluded and further classification on the content and usability, and dashboard visualization towards experiential learning resulted in 117 data sets.Of these 117 data sets, we further tested their usability to graph and create a dashboard within 2-5 minutes of loading the data to Tableau© that then used a Data Usability Scale (DUS) scoring developed from the industry standard of System Usability Scale (SUS). Data were deemed usable and useful for >60% average DUS scoring. Finally, 25 sets of data could be used effectively in classroom exercises dealing with electronic records and decision support for health care. Best data for dashboard usability were from MIMIC-eICU, and other websites like Zenodo produced low to high usability. The data sets with low to poor usability were from FigShare, Dataverse, CIHI, and BC Data Catalogue, respectively.Overall, 25 data sets with high usability of data related health informatics and healthcare systems showed 60-85% usability. Moreover, all nine platforms showed ease-of-use search patterns to establish the criteria in a short amount of time. However, more investigation is needed to compare data-to-dashboard visualization for single to multiple files for experiential learning in health informatics.
Dillon Chrimes, Chanhee Kim
IEEE Big Data2