Everett Neil Rush

dblp:191/2758 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-5632-5723ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSystems, architecture and hardware · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis
Ziming Gan, Doudou Zhou, Everett Neil Rush, Vidul Ayakulangara Panickan, Yuk-Lam Ho, George Ostrouchov, Shuting Shen, Xin Xiong 0006, Kimberly F. Greco, Chuan Hong, Clara-Lea Bonzel, Jun Wen 0001, Lauren Costa, Tianrun A. Cai, Edmon Begoli, Zongqi Xia, John Michael Gaziano, Katherine P. Liao, Kelly Cho, Tianxi Cai
J. Biomed. Informatics3
2025 DOME: Directional medical embedding vectors from Electronic Health Records
abstract
MOTIVATION: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. METHODS: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. RESULTS: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHR-embedding.
Jun Wen 0001, Hao Xue 0005, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai
J. Biomed. Informatics3
2023 Multimodal representation learning for predicting molecule-disease relations
abstract
MOTIVATION: Predicting molecule-disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule-molecule, molecule-disease and disease-disease semantic dependencies can potentially improve prediction performance. METHODS: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule-disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. RESULTS: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jun Wen 0001, Xiang Zhang 0012, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Marinka Zitnik, Tianxi Cai
Bioinform.3
2022 Knowledge-Driven Online Multimodal Automated Phenotyping System
Molei Liu, Sara Morini Sweet, Xin Xiong 0006, Chuan Hong, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Everett Neil Rush, Yuk-Lam Ho, Kelly Cho, John Michael Gaziano, Katherine P. Liao, Tianxi Cai, Tianrun A. Cai
AMIA7
2022 Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization
abstract
OBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes.
Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai
J. Biomed. Informatics5
2020 Characterizing Sub-Cohorts via Data Normalization and Representation Learning
abstract
The process of identifying a cohort of interest is a very challenging task. It requires manually inspecting many patient records of complex structure that might include medical coding errors and missing data. This paper presents a computational pipeline for refining the process of cohort selection based on medical concepts recorded in the electronic health records (EHRs). The pipeline extracts EHR data for a given cohort and normalizes this data using standard vocabularies. Then a stacked denoising autoencoder is used to embed the normalized patient vectors in a low dimensional space, where the patients are subsequently clustered into sub-cohorts. The goal is to represent the cohort in a standard format and abstract variants of sub-populations. As a use-case, we applied the pipeline to 1.8 million Veterans diagnosed with major depressive disorder (MDD), and identified four meaningful sub-cohorts using the features learned by the autoencoder. Then, each sub-cohort was explored using a set of keywords for interpretation.
Everett Neil Rush, Özgür Özmen, Kathryn Knight, Byung H. Park, Clifton Baker, Makoto Jones, Merry Ward, Jonathan R. Nebeker
CBMS1
2018 Towards Adaptive Parallel Storage Systems
abstract
Disk I/O is a major bottleneck limiting the performance and scalability of data intensive applications. A common way to address disk I/O bottlenecks is using parallel storage systems and utilizing concurrent operation of independent storage components; however, achieving a consistently high parallel I/O performance is challenging due to static configurations. Modern parallel storage systems, especially in the cloud, enterprise data centers, and scientific clusters are commonly shared by various applications generating dynamic and coexisting data access patterns. Nonetheless, these systems generally utilize one-layout-fits-all data placement strategy frequently resulting in suboptimal I/O parallelism. Guided by association rule mining, graph coloring, bin packing, and network flow techniques, this paper proposes a general framework for adaptive parallel storage systems, with the goal of continuously providing a high-degree of I/O parallelism. Evaluation results indicate that the proposed framework is highly successful in adjusting to skewed parallel access patterns for both hard disk drive (HDD) based traditional storage arrays and solid-state drive (SSD) based all-flash arrays. In addition to the storage arrays, the proposed framework is sufficiently generic and can be tailored to various other parallel storage scenarios including but not limited to key-value stores, parallel/distributed file systems, and internal parallelism of SSDs.
Erica Tomes, Everett Neil Rush, Nihat Altiparmak
IEEE Trans. Computers2
2016 Next-gen tools for big scientific data: ARM data center example
abstract
The Atmospheric Radiation Measurement (ARM) Climate Research Facility (www.arm.gov) provides atmospheric observations from diverse climatic regimes around the world. Currently, ARM archives over 22 million user assessable data files, primarily stored in NetCDF file format, with total data volumes close to one Petabyte. In this paper, we will discuss how ARM is currently storing, distributing, cataloging and visualizing such large volumes of multi-dimensional climate observations and model data and also describe their future plan.
Ranjeet Devarakonda, Kyle Dumas, Sheman Beus, Everett Neil Rush, Bhargavi Krishna, Robert Records, Giri Prakash
IEEE BigData4
2016 HPC infrastructure to support the next-generation ARM facility data operations
abstract
The Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) Climate Research Facility is establishing an adaptive data services and operations architecture in support of the Next-Generation ARM Facility as explained in its Decadal Vision. In this paper, we describe the capabilities of the ARM Data Center (ADC) and the upcoming high-performance computing infrastructure in support of this Next-Generation ARM Facility.
Giri Prakash, Jitendra Kumar 0001, Everett Neil Rush, Robert Records, Anthony Clodfelter, Jimmy W. Voyles
IEEE BigData3
2016 Dynamic Data Layout Optimization for High Performance Parallel I/O
abstract
Storage performance bottlenecks are one of the major threats limiting the scalability of I/O intensive applications. Parallel storage systems have the potential to alleviate I/O bottlenecks through concurrent operation of independent storage components if a parallelism-aware data layout can be continuously guaranteed. Existing systems use one-layout-fits-all data placement strategy that frequently results in sub-optimal I/O parallelism. Guided by association rule mining, graph coloring, bin packing, and network flow techniques, this paper proposes a general framework for self-optimizing parallel storage systems, with the goal of continuously providing a high-degree of I/O parallelism that is robust to changes in the parallel access patterns of applications and the coexistence of applications with different parallel access characteristics. Evaluation results indicate that the proposed framework is highly successful in adjusting to skewed parallel access patterns for both traditional hard disk drive (HDD) based storage arrays and solid-state drive (SSD) based all-flash arrays. In addition to the storage arrays, the proposed framework is sufficiently generic to be tailored to various other parallel storage scenarios including but not limited to key-value stores, parallel/distributed file systems, and internal parallelism of SSDs.
Everett Neil Rush, Bryan Harris, Nihat Altiparmak, Ali Saman Tosun
HiPC1
2016 Exploiting Replication for Energy Efficiency of Heterogeneous Storage Systems
abstract
As a result of immense growth of digital data in the last decade, energy consumption has become an important issue in data storage systems. In the US alone, data centers were projected to consume $4 billion (40 TWh) yearly electricity in 2005. This cost had reached to $10 billion (100 TWh) in 2011, and expected to be around $20 billion (200 TWh) in 2016 by doubling itself every 5 years. In addition to the economic burden on companies and research institutions, these large scale data storage systems also have a negative impact on the environment. According to the EPA, generating 1 KWh of electricity in the US results in an average of 1.55 pounds of carbon dioxide emissions. Considering a projected 200 TWh energy requirement for 2016, energy-efficient data storage systems can have a huge economic and environmental impacts on society. This project exploits replication and heterogeneity existing in modern multi-disk storage systems and proposes an energy-efficient and performance-aware replica selection technique to reduce the energy consumption of data storage systems without negatively affecting their performance. Our proposed technique exploits the difference between active and idle energy consumption in heterogeneous disks holding the same replica and selects replicas by balancing energy and performance.
Everett Neil Rush, Nihat Altiparmak
MASCOTS1