John Gounley

dblp:188/2462 · also John P. Gounley · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0001-8424-4982ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit
abstract
Turbulence plays a crucial role in multiphysics applications, including aerodynamics, fusion, and combustion. Accurately capturing turbulence's multiscale characteristics is essential for reliable predictions of multiphysics interactions, but remains a grand challenge even for exascale supercomputers and advanced deep learning models. The extreme-resolution data required to represent turbulence, ranging from billions to trillions of grid points, pose prohibitive computational costs for models based on architectures like vision transformers. To address this challenge, we introduce a multiscale hierarchical Turbulence Transformer that reduces sequence length from billions to a few millions and a novel RingX sequence parallelism approach that enables scalable long-context learning. We perform scaling and science runs on the Frontier supercomputer. Our approach demonstrates excellent performance up to 1.1 EFLOPS on 32,768 AMD GPUs, with a scaling efficiency of 94%. To our knowledge, this is the first AI model for turbulence that can capture small-scale eddies down to the dissipative range.
Junqi Yin, Mijanur Palash, M. Paul Laiu, Muralikrishnan Gopalakrishnan Meena, John Gounley, Stephen de Bruyn Kops, Feiyi Wang, Ramanan Sankaran
IPDPS5
2026 Deformable phrase level attention: A flexible approach for improving AI based medical coding
abstract
OBJECTIVE: Improving the AI-driven automated medical encoding of clinical text plays a vital role in gathering information on the occurrence of diseases to improve population-level health. This work presents a novel attention mechanism designed to enhance text classification models and ensure appropriate classification of medical concepts in unstructured electronic health records. MATERIALS AND METHODS: We developed a deformable, phrase-level attention mechanism to identify important lexical word-level and contextual phrase-level information from clinical text documents. We evaluated conventional and transformer-based deep learning models that we extended with our attention mechanism on the extraction of critical cancer information (e.g., site, subsite, laterality, histology, behavior) from 629,908 electronic pathology reports and on the automated medical encoding of 52,722 hospital discharge summaries. RESULTS: Transformer-based models with the deformable, phrase-level attention mechanism achieved the best performance on the extraction of critical cancer information from pathology reports. Conventional- and transformer-based models show similar or better performance than their baseline counterparts on the automated medical encoding of clinical documents. DISCUSSION: The addition of phrase-level information allowed models extended with our proposed method to outperform standard word-level attention. Our method showed favorable properties for the real-world application in terms of model robustness and phenotyping. These results indicate that our method is promising for automated data harmonization for common data models. CONCLUSION: This work proposes a novel deformable, phrase-level attention mechanism that enhances text classification models in the extraction of medical concepts from clinical text documents. We demonstrate strong performances on two clinical text datasets and showcase real-world deployability of our method.
Christoph Metzner, Shang Gao 0008, Drahomira Herrmannova, John Gounley, Heidi A. Hanson
Artif. Intell. Medicine4
2025 Position-Enhanced Gradient Attack (PEGA) on Medical Language Models
abstract
Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.
Nuo Xu 0013, Christopher Stanley, John Gounley, Heidi A. Hanson, Chang Ge 0002, Caiwen Ding
MMAsia3
2023 Enhancing Adaptive Physics Refinement Simulations Through the Addition of Realistic Red Blood Cell Counts
abstract
Simulations of cancer cell transport require accurately modeling mm-scale and longer trajectories through a circulatory system containing trillions of deformable red blood cells, whose intercellular interactions require submicron fidelity. Using a hybrid CPU-GPU approach, we extend the advanced physics refinement (APR) method to couple a finely-resolved region of explicitly-modeled red blood cells to a coarsely-resolved bulk fluid domain. We further develop algorithms that: capture the dynamics at the interface of differing viscosities, maintain hematocrit within the cell-filled volume, and move the finely-resolved region and encapsulated cells while tracking an individual cancer cell. Comparison to a fully-resolved fluid-structure interaction model is presented for verification. Finally, we use the advanced APR method to simulate cancer cell transport over a mm-scale distance while maintaining a local region of RBCs, using a fraction of the computational power required to run a fully-resolved model.
Sayan Roychowdhury, Samreen T. Mahmud, Aristotle X. Martin, Peter Balogh, Daniel F. Puleri, John Gounley, Erik W. Draeger, Amanda Randles
SC6
2023 Evaluation of pre-training large language models on leadership-class supercomputers
Junqi Yin, Sajal Dash, John Gounley, Feiyi Wang, Georgia D. Tourassi
J. Supercomput.3
2022 High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell Trajectory
abstract
The ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport.
Daniel F. Puleri, Sayan Roychowdhury, Peter Balogh, John Gounley, Erik W. Draeger, Jeff Ames, Adebayo Adebiyi, Simbarashe Chidyagwai, Benjamín Hernández, Seyong Lee, Shirley V. Moore, Jeffrey S. Vetter, Amanda Randles
CLUSTER4
2022 Automating Genetic Algorithm Mutations for Molecules Using a Masked Language Model
abstract
Inspired by the evolution of biological systems, genetic algorithms have been applied to generate solutions for optimization problems in a variety of scientific and engineering disciplines. For a given problem, a suitable genome representation must be defined along with a mutation operator to generate subsequent generations. Unlike natural systems, which display a variety of complex rearrangements (e.g., mobile genetic elements), mutation for genetic algorithms commonly utilizes only random pointwise changes. Furthermore, generalizing beyond pointwise mutations poses a key difficulty as useful genome rearrangements depend on the representation and problem domain. To move beyond the limitations of manually defined pointwise changes, here we propose the use of techniques from masked language models to automatically generate mutations. As a first step, common subsequences within a given population are used to generate a vocabulary. The vocabulary is then used to tokenize each genome. A masked language model is trained on the tokenized data in order to generate possible rearrangements (i.e., mutations). In order to illustrate the proposed strategy, we use string representations of molecules and use a genetic algorithm to optimize for drug-likeness and synthesizability. Our results show that moving beyond random pointwise mutations accelerates genetic algorithm optimization.
Andrew E. Blanchard, Mayanka Chandra Shekar, Shang Gao 0008, John Gounley, Isaac Lyngaas, Jens Glaser, Debsindhu Bhowmik
IEEE Trans. Evol. Comput.4
2022 Propagation Pattern for Moment Representation of the Lattice Boltzmann Method
abstract
A propagation pattern for the moment representation of the regularized lattice Boltzmann method (LBM) in three dimensions is presented. Using effectively lossless compression, the simulation state is stored as a set of moments of the lattice Boltzmann distribution function, instead of the distribution function itself. An efficient cache-aware propagation pattern for this moment representation has the effect of substantially reducing both the storage and memory bandwidth required for LBM simulations. This paper extends recent work with the moment representation by expanding the performance analysis on central processing unit (CPU) architectures, considering how boundary conditions are implemented, and demonstrating the effectiveness of the moment representation on a graphics processing unit (GPU) architecture.
John Gounley, Madhurima Vardhan, Erik W. Draeger, Pedro Valero-Lara, Shirley V. Moore, Amanda Randles
IEEE Trans. Parallel Distributed Syst.1
2021 How Distinct Structural Flexibility within SARS-CoV-2 Spike Protein Reveals Potential Therapeutic Targets
abstract
The emergence and rapid worldwide spread of the novel coronavirus disease, COVID-19, has prompted concerted efforts to find successful treatments. The causative virus, severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), uses its spike (S) protein to gain entry into host cells. Therefore, the S protein presents a viable target to develop a directed therapy. Here, we deployed an integrated artificial intelligence with all-atom molecular dynamics simulation approach to provide new details of the S protein structure. Based on a comprehensive structural analysis of S proteins from SARS-CoV-2 and previous human coronaviruses, we found that the protomer state of S proteins is structurally flexible. Without the presence of a stabilizing beta sheet from another protomer chain, two regions in the S2 domain and the hinge connecting the S1 and S2 subunits lose their secondary structures. Interestingly, the region in the S2 domain was previously identified as an immunodominant site in the SARS-CoV-1 S protein. We anticipate that the molecular details elucidated here will assist in effective therapeutic development for COVID-19.
Serena H. Chen, M. Todd Young, John Gounley, Christopher B. Stanley, Debsindhu Bhowmik
IEEE BigData3
2021 Limitations of Transformers on Clinical Text Classification
abstract
Bidirectional Encoder Representations from Transformers (BERT) and BERT-based approaches are the current state-of-the-art in many natural language processing (NLP) tasks; however, their application to document classification on long clinical texts is limited. In this work, we introduce four methods to scale BERT, which by default can only handle input sequences up to approximately 400 words long, to perform document classification on clinical texts several thousand words long. We compare these methods against two much simpler architectures - a word-level convolutional neural network and a hierarchical self-attention network - and show that BERT often cannot beat these simpler baselines when classifying MIMIC-III discharge summaries and SEER cancer pathology reports. In our analysis, we show that two key components of BERT - pretraining and WordPiece tokenization - may actually be inhibiting BERT's performance on clinical text classification tasks where the input document is several thousand words long and where correctly identifying labels may depend more on identifying a few key words or phrases rather than understanding the contextual meaning of sequences of text.
Shang Gao 0008, Mohammed M. Alawad, M. Todd Young, John Gounley, Noah Schaefferkoetter, Hong-Jun Yoon, Xiao-Cheng Wu, Eric B. Durbin, Jennifer A. Doherty, Antoinette Stroup, Linda Coyle, Georgia D. Tourassi
IEEE J. Biomed. Health Informatics4
2020 Accelerated training of bootstrap aggregation-based deep information extraction systems from cancer pathology reports
Hong-Jun Yoon, Hilda B. Klasky, John Gounley, Mohammed M. Alawad, Shang Gao 0008, Eric B. Durbin, Xiao-Cheng Wu, Antoinette Stroup, Jennifer A. Doherty, Linda Coyle, Lynne Penberthy, James Blair Christian, Georgia D. Tourassi
J. Biomed. Informatics3
2019 Information Extraction from Cancer Pathology Reports with Graph Convolution Networks for Natural Language Texts
abstract
Graph-of-words is a flexible and efficient text representation which addresses well-known challenges, such as word ordering and variation of expressions, to natural language processing. In this paper, we consider the latest graph-based convolutional neural network technique, the Text GraphConvolutional Network (Text GCN), in the context of performingclassification tasks on free-form natural language texts. To do this, we designed a study of multi-task information extraction from medical text documents. We implemented multi-task learning in the Text GCN, performed hyperparameter optimization, and measured the clinical task performances. We evaluated micro and macro-F1 scores of four information extraction tasks,including subsite, laterality, behavior, and histological grades from cancer pathology reports. The scores for the Text GCN significantly outperformed our previous studies with convolutional neural networks, suggesting that the Text GCN model is superior to traditional models in task performance.
Hong-Jun Yoon, John Gounley, M. Todd Young, Georgia D. Tourassi
IEEE BigData2
2019 Multi-physics simulations of particle tracking in arterial geometries with a scalable moving window algorithm
abstract
In arterial systems, cancer cell trajectories determine metastatic cancer locations; similarly, particle trajectories determine drug delivery distribution. Predicting trajectories is challenging, as the dynamics are affected by local interactions with red blood cells, complex hemodynamic flow structure, and downstream factors such as stenoses or blockages. Direct simulation is not possible, as a single simulation of a large arterial domain with explicit red blood cells is currently intractable on even the largest supercomputers. To overcome this limitation, we present a multi-physics adaptive window algorithm, in which individual red blood cells are explicitly modeled in a small region of interest moving through a coupled arterial fluid domain. We describe the coupling between the window and fluid domains, including automatic insertion and deletion of explicit cells and dynamic tracking of cells of interest by the window. We show that this algorithm scales efficiently on heterogeneous architectures and enables us to perform large, highly-resolved particle-tracking simulations that would otherwise be intractable.
Gregory Herschlag, John Gounley, Sayan Roychowdhury, Erik W. Draeger, Amanda Randles
CLUSTER2
2019 Investigating the Role of VR in a Simulation-Based Medical Planning System for Coronary Interventions
Madhurima Vardhan, Harvey Shi, John Gounley, S. James Chen, Andrew Kahn, Jane A. Leopold, Amanda Randles
MICCAI (5)3
2019 Moment representation in the lattice Boltzmann method on massively parallel hardware
abstract
The widely-used lattice Boltzmann method (LBM) for computational fluid dynamics is highly scalable, but also significantly memory bandwidth-bound on current architectures. This paper presents a new regularized LBM implementation that reduces the memory footprint by only storing macroscopic, moment-based data. We show that the amount of data that must be stored in memory during a simulation is reduced by up to 47%. We also present a technique for cache-aware data re-utilization and show that optimizing cache utilization to limit data motion results in a similar improvement in time to solution. These new algorithms are implemented in the hemodynamics solver HARVEY and demonstrated using both idealized and realistic biological geometries. We develop a performance model for the moment representation algorithm and evaluate the performance on Summit.
Madhurima Vardhan, John Gounley, Luiz Hegele, Erik W. Draeger, Amanda Randles
SC2
2019 Performance portability study for massively parallel computational fluid dynamics application on scalable heterogeneous architectures
Seyong Lee, John Gounley, Amanda Randles, Jeffrey S. Vetter
J. Parallel Distributed Comput.2