EDBT 2026 Demo / reviewers in the wild / expert
Dillon Chrimes
dblp:120/3464
· DBLP profile ↗
6ranked-venue papers in the field
5as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Big data analytic of US Medicare Claims of healthcare service informatics and billing fraudabstractBig data analysis of health services and health informatics unique to public use files of US Medicare Claims data was the primary objective of this investigation. This study focused on the five most populated US states California, Texas, Florida, New York, and Pennsylvania of the 2021 US Medicare Claims datasets. To establish meaningful insights into the big data of US Medicare Claims in 2021, informative graphs, visualizations, regression analysis, and clustering data mining techniques were carried out.The data comprised two datasets. One dataset was called Medicare Physician & Other Practitioners - by Provider, and the other dataset was Medicare Physician & Other Practitioners - by Provider and Service, which comprised services related to a variety of chronic diseases. Fifteen aggregated HCPCS codes were calculated because there were over 400 HCPCS codes and that needed to be reduced for the analysis. The methodology used traditional data mining techniques that included graphing, visualization with trend lines, and linear regression to calculate residuals to identify outliers. That was, data visualizations were created to present the difference between submitted charges and payment amounts, linear regression, residuals, and subsequent residual outliers. Outliers were investigated for possible identified fraudulent providers or submitted claims.Big data exploration and insightful visualizations led to establish information about submitted claims, payment amounts, total services, provider type and specialty, Healthcare Common Procedure Coding System (HCPCS) codes, and total beneficiaries. The vast amount of health and healthcare informatics in the public use file showed the state of California with highest submitted charges of $45b. Across the top 5 US states the top three provider types by summitted charge was ambulatory surgical center ($23,305,959,194.92) followed by diagnostic radiology, and clinical laboratory. HCPCS codes were highest for repair (i.e. unplanned surgeries) and elective surgery. Sum of beneficiary per day service by 15 aggregated HCPCS codes showed that Patient Care Services ($0.35b), Misc. Testing ($0.21b), Emergency Services ($0.10b), and Vaccine Administration and Medication ($0.10b) were the most frequent HCPCS codes per day. Top three provider type by average risk score: nephrology (4,256,048), infectious disease (2,981,082), and hematopoietic cell transplantation and cellular therapy (2,930,108). Beneficiaries and submitted charges by state showed a linear progression of the total services by the submitted charges. California and Pennsylvania had recurrent and similar outliers in this regression analysis. Furthermore, provider types had clusters identified in the data for each of the five US states that were unique. Some provider types had higher frequency of total services that included pharmacy, independent diagnostic testing facility, and clinical laboratory.PowerBI proved to be a useful tool to use data mining techniques as no other software was able to load the US Medicare Claims Provider Services files and perform data analysis such as clustering. The dataset was limited to one-year coverage in 2021, data errors and technical challenges in PowerBI for more advanced regression analysis all introduced certain constraints in applying more advanced data mining algorithms. In conclusion, we were able to successfully data mine large public US Medicare Claim data using PowerBI and its functionalities such as DAX, measures, action filters, and visualizations.There were different outliers (related to National Provider Identifier) for each visualization with regression trend line for payment amount and total services; submitted charge (claim) and total services; total services and beneficiaries; total services and HCPCS codes; and provider with total beneficiaries. Moreover, 19 suspected fraud claims were identified for professional organizations. Thus, our results clearly showed that the variety of big data can be graphed with outliers. However, the different outliers found based on different data elements (variables and parameters) proved that a more sophisticated (perhaps artificial intelligence algorithms like random forest, boosted gradient or even deep learning) big data analytic is needed for further investigation of possible fraudulent healthcare claims. Maya El-Lanky, Dillon Chrimes |
IEEE Big Data | 2 |
| 2023 | Text mining using clinical terms in electronic records of annual falls of patients in home community careabstractThe number of health informatics in text in electronic health records (EHRs) has increased. In parallel, the text mining of health informatics in clinical notes has also increased in the attempt to utilize big data in EHRs for better patient care. However, little is known on how to utilize the word frequencies of text in clinical notes in EHRs to investigate patient care over a per annual basis.Our methods used a tool called NimbleMiner (NM) that utilized RStudio and generate a relatively “naïve” native lexicon of Simclins (i.e., SIMilar CLINical terms) via the applications with special emphasis on patient falls, which are important events that can lead to decreased health outcomes. The identified Simclins’ frequencies in the file were then determined using a basic Python script for 1-gram words (e.g., “home”) and using Voyant Tools for n+1-gram words (e.g., “living room” or “intensive care unit”).The most frequent words (over an annual basis) in the training corpus were determined to be: “mg” (15,726); “patient” (12,945); “tablet” (10,992); “po” or per os or by mouth (8,904); “blood” (8,824). The Simclins were identified for each domain of interest, the initial prompts for the first search iteration of NM, as well as ChatGPT3.5 output for comparison purposes. There were 15,252 Simclin word frequency related to home and 2,421 related to clinical settings over one year, respectively. Simclin frequencies per year by domain were “fall” (514), “setting” (17,673), computerized provider order entry – CPOE (3,408), and medications (62,317). Furthermore, the significantly larger text of 33,769 on medications (i.e., mg, tablet, medications, aspirin, coumadin, dose, doses, prescribed, and unit_ml_solution) of the total of 46,530 or 72.6% of the corpus.The use of NM and other tools proved that text mining of the clinical notes can provide operational information of important clinical events based on word Simclin frequencies in the clinical notes over an annual basis. The results showed importance of lower frequency words like “fall” as well as higher frequency terms such as “medications” and “prescriptions”. More investigation is needed in terms using Simclins across the health big data of clinical notes of EHRs to better understand the use of text in patient care. Dillon Chrimes, Emile Keruzore |
IEEE Big Data | 1 |
| 2023 | Comparison of MIMIC-III and MIMIC-IV for big data analytics of health informaticsabstractThe use of a health big data via Medical Information Mart for Intensive Care (MIMIC) data sets has significant advancements in health informatics and clinical research of electronic health records (EHRs) in hospital systems. MIMIC-III and MIMIC-IV data sets are the two latest publicly available iterations of data from electronic records that possess distinct features, variables, and structures in two separate large relational databases. This study aimed to provide big data analytics comparisons of MIMIC-III and MIMIC-IV for experiential learning of health informatics from EHRs.Both data sets were quantitatively and qualitatively evaluated based on dataset properties (data, list of tables, timeline, patient encounters, software, system), data mining and visual categories (i.e., data size, data structure and schema, features/variables, data quality and completeness, clinical focus, access and use, dashboards, and visualization types), data usability heuristics and utilization (i.e., usability, complexity, granularity, applicability, and capacity). Results showed significant difference of 100 patients across 26 tables for MIMIC-III compared to 2,520 patients across 31 tables for MIMIC-IV data sets, respectively. There were 1716 diagnoses (ICD-9) and 503 procedures (ICD-9) for MIMIC-III. There were 262 diseases categories with >1000 disease instances and >2000 treatments for MIMIC-IV with no ICD codes. Moreover, these data sets contained large (big data) charting events in MIMIC-III of 758,356 rows, and MIMIC-IV eICU with charting events of 1,477,163 for nursing, and respiratory of 176,089 rows, respectively. The results suggest that MIMIC-III provided detailed information for retrospective clinical studies and operations in critical care with high data granularity in terms of re-admission (calculated fields from its admission table), length of stay, prescriptions, caregivers, and diagnosis and procedure (ICD-9). However, it lacked clinical capacity because of no diagnosis or charting event offset times or APACHE IV (Acute Physiology and Chronic Health Evaluation) scores that were in the MIMIC-IV eICU dataset. Hence, MIMIC-IV showed higher data granularity and capacity. MIMIC-IV eICU dataset introduces enhanced data attributes, more sophisticated patient trajectory tracking at the ICU unit level, and improved detailed information from electronic records. Moreover, MIMIC-IV has high usability, complexity, and applicability to critical care but lacking hospital operational data of re-admissions, caregivers, and ICD codes that MIMIC-III contained. However, MIMIC-IV dataset contained complex data schemas of treatment strings, nurse care plans, lab results, medications, microbiology, as well as infusion drug and respiratory charting. The big data analytics of MIMIC-III and MIMIC-IV needs to be further investigated for AI applications to demonstrate its usefulness for dynamic decision-making in health care. Dillon Chrimes, Chanhee Kim |
IEEE Big Data | 1 |
| 2023 | Big data usability text mining of publicly available YouTube electronic health record (EHR) tutorialsabstractThis study seeks to harness the power of big data through text mining the content of publicly available YouTube tutorials related to Electronic Health Record (EHR) systems. Most information about EHR systems is proprietary and hidden to the public. Nevertheless, understanding the capacity of EHRs from YouTube tutorials can offer insights to the public into health informatics and EHR functionality and usability for their patient care.We established a search strategy of the top three EHR vendors in North America (i.e., Oracle-Cerner, Meditech, and Epic Systems) from YouTube containing metadata and transcripts of EHR-related tutorials. Text mining techniques across a three-phased approach included establishing a search strategy to select the YouTube videos, auto-coding (from Voyant Tools) and manual coding (via standard usability heuristics), and analysis. The EHR components or applications in the tutorials were selected based on search strategy of total volume and quality of videos related to patient registration, scheduling, problem list, documentation, medication/order, and discharge/administration. The tutorials were analyzed for keyword frequencies, thematic patterns, and overall usability in the information presented via video transcriptions. Thematic patterns were established by heuristic categories of navigation, visibility, content, and complex interactions of learnability and memory.Preliminary findings suggest big data text mining led to analyzing 1,788,126 words via YouTube transcripts across 2280 videos. Oracle (Cerner) had a total of 799,653 words in the 1,066 videos. Meditech has 608 videos with 454,968 words. Epic has 466 videos with 436,598 words. There was a good range of content quality from medium to high, with some tutorials offering comprehensive step-by-step workflows like for medication/orders. The heuristics of content and navigation categories suggested that the tutorials covered enough material and word frequencies to be educational towards experiential learning of EHRs.The study underscores the significant role of using YouTube in disseminating knowledge about complex systems like EHRs to the public. Text mining of large, big data corpus can reveal insights into the quality and nature of online tutorials of more targeted and effective educational resources. Dillon Chrimes, Ivan Tang |
IEEE Big Data | 1 |
| 2022 | Big data analytics of predicting annual US Medicare billing claims with health servicesabstractThis paper investigated the use of large public use files (PUFs) of US Medicare claims in the form of big data analytics to predict claim amounts in US dollars (USD) and large spending anomalies across hundreds of health services documented in the data set. There were two main research questions to better understand content and use of PUFs of US Medicare. One question was related to understanding the dataset and the parameters that could predict the total submitted billing claims for one year (i.e. 2017 fiscal year in USD). The second question was to establish whether or not anomalies in health service costs could be detected. Null hypothesis was that there are no significant variables in the general linear model (GLM) of the regression analysis. The hypothesis related to factors of type and frequency of health services, total HCPCS (Healthcare Common Procedural Coding System), population (total beneficiaries), age, provider specialty, chronic disease, states and regions could be significant in the classification and regression models.The 2017 Medicare Claims dataset, publically provided by Centers for Medicare & Medicaid Services (CMS), was 291 MB and consisted of >30 columns and ~1,048,576 rows. The methodology followed data mining techniques to general linear regression to derive model fitting that compared the model residuals. From the residuals, multivariate outlier detection was carried out that included k-means clustering and principal component analysis.The results showed a correlation R2of 52% with health services and submitted Medicare amounts (USD) with thousands of outliers. Total services variable was highly significant with the total amount of submitted claims (maximum of 1025413240). HCPCS was not significant. There was also a strong correlation of Medicare costs to larger population in states with larger cities, especially in California, Florida, New York, and Texas. However, regions, States, cities, zip codes, and other divisions of the US states and regions were not significantly different. Adding other variables (like procedures and chronic diseases) improved the R2correlation only slightly.The results showed that there was significant correlation among services and the total submitted amount. Furthermore, there were many extreme outliers in terms of costs and consistently it was diagnostic radiation services in larger population states that had the highest total amount ($) of Medicare claims. New Your City’s 2017 diagnostic radiology was ten times larger than another city in the US. These large ($) amounts in the outliers and their characteristics did indicate that outliers were detectible.Medicare data comprised longitudinal information on a substantial proportion of the population aged ≥65 years in the US; however, poor model performance did show that there are many gaps in the linkage to chronic conditions of patients, wherein there could be possible misclassification of diagnoses, which would require more detailed investigation spanning other datasets. Similarly, Medicare data, like other forms of health care claims, are not collected for the purposes of research, but to support reimbursement for health care services. Unlike data captured in an electronic health record (EHRs) or as part of a prospective study, information collected as claims for reimbursement for services provided may be influenced by financial incentives and may be more susceptible to misclassification. In conclusion, Medicare claims can be predicted from a variety of parameters and outliers can be detected. However, poor model performance in the big data analytics of billing claims requires further investigation of the data mining techniques. Dillon Chrimes |
IEEE Big Data | 1 |
| 2022 | Review of Publically Available Health Big Data SetsabstractThere is a growing interest in using public data for open government policy involving health informatics and healthcare systems. This paper investigated the characteristics of publically available data sets in health informatics that were derived from electronic health records (EHRs), healthcare systems, and a variety of open-government libraries, data marts, or data catalogues.Data used in this study consisted of public data sets that did not require any registration to access online. In total, nine web-based platforms on the Internet were used that included: British Columbia (BC) Data Catalogue, Canadian Institute for Health Information (CIHI), Harvard Dataverse, MIMIC-eICU, FigShare, GitHub, Google Dataset, UCI Machine Learning Repository, and Zenodo. Our initial search across these platforms found over 10,000 public use files that had data sets related to health informatics.We found 558 data sets that matched search criterion that ranged from years 2013-2022. The data source types were mostly found using the health informatics search filters followed by the combination of health informatics and healthcare systems, but fewer data sets were found when using EHR as the criterion. Almost 85% of the total data sets were from 2020-2022. The range of data sizes were 11KB to 7.8MB. The eICU (hosted by MIT’s MIMIC data mart) platform had the largest data set followed by Zenodo, and GitHub. Additionally, any bioinformatics in the 558 data sets were excluded and further classification on the content and usability, and dashboard visualization towards experiential learning resulted in 117 data sets.Of these 117 data sets, we further tested their usability to graph and create a dashboard within 2-5 minutes of loading the data to Tableau© that then used a Data Usability Scale (DUS) scoring developed from the industry standard of System Usability Scale (SUS). Data were deemed usable and useful for >60% average DUS scoring. Finally, 25 sets of data could be used effectively in classroom exercises dealing with electronic records and decision support for health care. Best data for dashboard usability were from MIMIC-eICU, and other websites like Zenodo produced low to high usability. The data sets with low to poor usability were from FigShare, Dataverse, CIHI, and BC Data Catalogue, respectively.Overall, 25 data sets with high usability of data related health informatics and healthcare systems showed 60-85% usability. Moreover, all nine platforms showed ease-of-use search patterns to establish the criteria in a short amount of time. However, more investigation is needed to compare data-to-dashboard visualization for single to multiple files for experiential learning in health informatics. Dillon Chrimes, Chanhee Kim |
IEEE Big Data | 1 |